All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH] block: skip redundant flush for O_DSYNC direct writes
@ 2026-08-12 13:03 Zhenxian Ma
  2026-08-12 16:42 ` Christoph Hellwig
  0 siblings, 1 reply; 3+ messages in thread
From: Zhenxian Ma @ 2026-08-12 13:03 UTC (permalink / raw)
  To: Jens Axboe; +Cc: linux-block, Zhenxian Ma, Zhenxian Ma

For an O_DIRECT | O_DSYNC write to a block device, dio_bio_write_op()
adds REQ_FUA to the bio, so the data is durable when the direct-IO
returns.  blkdev_write_iter() then falls through to
generic_write_sync(), which for O_DSYNC turns into an additional
REQ_PREFLUSH.  That flush is redundant: the FUA write already made
the data persistent.

Skip generic_write_sync() when the direct-IO path already provided
durability via FUA.  The buffered and direct_write_fallback() paths
are unchanged.

Measured on a Seagate ST20000NM007D (20 TB, 7200 rpm, fua=1,
write_cache=write back), Linux v7.2.0-rc7, single-threaded pwrite()
loop opening the raw block device with O_WRONLY | O_DIRECT | O_DSYNC,
4 KiB writes for 60 s:

  Sequential 4 KiB writes:
                            baseline    patched
    IOPS                       119.7     7497.0
    avg latency (us)            8357        133
    p50 latency (us)            8346        127
    p99 latency (us)            8368        395
    p99.9 latency (us)          8728        569

  Random 4 KiB writes (100 GiB span):
                            baseline    patched
    IOPS                       156.1      666.4
    avg latency (us)            6405       1500
    p50 latency (us)            6186       1450
    p99 latency (us)           16133       2285
    p99.9 latency (us)         17250       9916

Signed-off-by: Zhenxian Ma <mzx199711@gmail.com>
Signed-off-by: Zhenxian Ma <mazhenxian@xiaohongshu.com>
---
 block/fops.c | 9 +++++++--
 1 file changed, 7 insertions(+), 2 deletions(-)

diff --git a/block/fops.c b/block/fops.c
index 15783a6180de..9c34a372b476 100644
--- a/block/fops.c
+++ b/block/fops.c
@@ -726,6 +726,7 @@ static ssize_t blkdev_write_iter(struct kiocb *iocb, struct iov_iter *from)
 	bool atomic = iocb->ki_flags & IOCB_ATOMIC;
 	loff_t size = bdev_nr_bytes(bdev);
 	size_t shorted = 0;
+	bool dio_fua_done = false;
 	ssize_t ret;
 
 	if (bdev_read_only(bdev))
@@ -763,9 +764,13 @@ static ssize_t blkdev_write_iter(struct kiocb *iocb, struct iov_iter *from)
 
 	if (iocb->ki_flags & IOCB_DIRECT) {
 		ret = blkdev_direct_write(iocb, from);
-		if (ret >= 0 && iov_iter_count(from))
+		if (ret >= 0 && iov_iter_count(from)) {
 			ret = direct_write_fallback(iocb, from, ret,
 					blkdev_buffered_write(iocb, from));
+		} else if (ret > 0 && iocb_is_dsync(iocb)) {
+			/* FUA from dio_bio_write_op() already made it durable */
+			dio_fua_done = true;
+		}
 	} else {
 		/*
 		 * Take i_rwsem and invalidate_lock to avoid racing with
@@ -777,7 +782,7 @@ static ssize_t blkdev_write_iter(struct kiocb *iocb, struct iov_iter *from)
 		inode_unlock_shared(bd_inode);
 	}
 
-	if (ret > 0)
+	if (ret > 0 && !dio_fua_done)
 		ret = generic_write_sync(iocb, ret);
 	iov_iter_reexpand(from, iov_iter_count(from) + shorted);
 	return ret;
-- 
2.43.5


^ permalink raw reply related	[flat|nested] 3+ messages in thread

* Re: [PATCH] block: skip redundant flush for O_DSYNC direct writes
  2026-08-12 13:03 [PATCH] block: skip redundant flush for O_DSYNC direct writes Zhenxian Ma
@ 2026-08-12 16:42 ` Christoph Hellwig
  2026-08-13  7:35   ` 马振先
  0 siblings, 1 reply; 3+ messages in thread
From: Christoph Hellwig @ 2026-08-12 16:42 UTC (permalink / raw)
  To: Zhenxian Ma; +Cc: Jens Axboe, linux-block, Zhenxian Ma

> +	bool dio_fua_done = false;
>  	ssize_t ret;
>  
>  	if (bdev_read_only(bdev))
> @@ -763,9 +764,13 @@ static ssize_t blkdev_write_iter(struct kiocb *iocb, struct iov_iter *from)
>  
>  	if (iocb->ki_flags & IOCB_DIRECT) {
>  		ret = blkdev_direct_write(iocb, from);
> -		if (ret >= 0 && iov_iter_count(from))
> +		if (ret >= 0 && iov_iter_count(from)) {
>  			ret = direct_write_fallback(iocb, from, ret,
>  					blkdev_buffered_write(iocb, from));
> +		} else if (ret > 0 && iocb_is_dsync(iocb)) {
> +			/* FUA from dio_bio_write_op() already made it durable */
> +			dio_fua_done = true;

This is kinda the wrong way aroud, it should just have a need_sync
variable that is only set for the fallback.

But looking at this, the original logic here actually is very fishy.
If we have a driver that does not actually support FUA, emulating it
for every write command is a bad idea.  So this should be fixed to
only do FUA when actually useful, and then return from the non-fallback
implementations if that actually happened.


^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH] block: skip redundant flush for O_DSYNC direct writes
  2026-08-12 16:42 ` Christoph Hellwig
@ 2026-08-13  7:35   ` 马振先
  0 siblings, 0 replies; 3+ messages in thread
From: 马振先 @ 2026-08-13  7:35 UTC (permalink / raw)
  To: Christoph Hellwig; +Cc: Jens Axboe, linux-block, Zhenxian Ma

On Thu, Aug 13, 2026 at 12:42 AM Christoph Hellwig <hch@infradead.org> wrote:
>
> > +     bool dio_fua_done = false;
> >       ssize_t ret;
> >
> >       if (bdev_read_only(bdev))
> > @@ -763,9 +764,13 @@ static ssize_t blkdev_write_iter(struct kiocb *iocb, struct iov_iter *from)
> >
> >       if (iocb->ki_flags & IOCB_DIRECT) {
> >               ret = blkdev_direct_write(iocb, from);
> > -             if (ret >= 0 && iov_iter_count(from))
> > +             if (ret >= 0 && iov_iter_count(from)) {
> >                       ret = direct_write_fallback(iocb, from, ret,
> >                                       blkdev_buffered_write(iocb, from));
> > +             } else if (ret > 0 && iocb_is_dsync(iocb)) {
> > +                     /* FUA from dio_bio_write_op() already made it durable */
> > +                     dio_fua_done = true;
>
> This is kinda the wrong way aroud, it should just have a need_sync
> variable that is only set for the fallback.

Agreed, that reads much cleaner, I'll modify the flag in v2.

> But looking at this, the original logic here actually is very fishy.
> If we have a driver that does not actually support FUA, emulating it
> for every write command is a bad idea.  So this should be fixed to
> only do FUA when actually useful, and then return from the non-fallback
> implementations if that actually happened.

Let me make sure:
     blkdev_write_iter() then only skips generic_write_sync() when
     the direct-IO path actually issued REQ_FUA (i.e. iocb_is_dsync(iocb)
     && bdev_fua(bdev) && the write completed without falling back to
     buffered).  Is that the direction you had in mind?  If yes, I'll send patch
     v2 later.

Thanks for the review,
Zhenxian

^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-08-13  7:35 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-12 13:03 [PATCH] block: skip redundant flush for O_DSYNC direct writes Zhenxian Ma
2026-08-12 16:42 ` Christoph Hellwig
2026-08-13  7:35   ` 马振先

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.