All of lore.kernel.org
 help / color / mirror / Atom feed
From: Tal Zussman <tz2294@columbia.edu>
To: Jens Axboe <axboe@kernel.dk>, Christoph Hellwig <hch@lst.de>,
	Johannes Thumshirn <johannes.thumshirn@wdc.com>,
	Luis Chamberlain <mcgrof@kernel.org>,
	Hannes Reinecke <hare@suse.de>,
	"Matthew Wilcox (Oracle)" <willy@infradead.org>,
	John Garry <john.g.garry@oracle.com>,
	Christian Brauner <brauner@kernel.org>,
	"Darrick J. Wong" <djwong@kernel.org>,
	Keith Busch <kbusch@kernel.org>,
	"Martin K. Petersen" <martin.petersen@oracle.com>
Cc: linux-block@vger.kernel.org, linux-kernel@vger.kernel.org,
	Sashiko <sashiko-bot@kernel.org>,
	Tal Zussman <tz2294@columbia.edu>
Subject: [PATCH v2 5/7] block: fail atomic writes instead of falling back to buffered I/O
Date: Fri, 28 Aug 2026 09:49:54 -0400	[thread overview]
Message-ID: <20260828-blkdev-fixes-v2-5-32f3f40cebed@columbia.edu> (raw)
In-Reply-To: <20260828-blkdev-fixes-v2-0-32f3f40cebed@columbia.edu>

An IOCB_ATOMIC direct write to a block device can silently lose its
torn-write guarantee in two ways:

  1. blkdev_direct_write() turns an -EBUSY from page cache invalidation
     into a 0 return, so the whole write is retried through
     blkdev_buffered_write(), with no atomicity guarantee.

  2. On a partial page pin, __blkdev_direct_IO_simple() and
     __blkdev_direct_IO_async() submit what was pinned with REQ_ATOMIC
     set and leave the rest to the buffered fallback.

The second case can be triggered deterministically. A 16K
pwritev2(RWF_ATOMIC) whose last page is PROT_NONE, on a scsi_debug
device with atomic_wr=1, completes short with only three of the four
pages written, violating RWF_ATOMIC semantics.

Fail the I/O instead. Return -EAGAIN when page cache invalidation fails
for IOCB_ATOMIC rather than retrying through the page cache, matching
__iomap_dio_rw(), which treats the failure as transient and lets the
caller retry. Release a short atomic pin and return -EFAULT before
submission, which is what a direct write already returns when none of
the buffer can be pinned. A sync atomic write can then never return
short with a remainder, so the buffered fallback is never reached.

ext4 has the same fallback and only warns in it. For block devices both
ways in can be detected before any I/O is submitted, so fail early instead.

Fixes: caf336f81b3a ("block: Add fops atomic write support")
Reported-by: Sashiko <sashiko-bot@kernel.org>
Link: https://sashiko.dev/#/patchset/20260802-blkdev-fixes-v1-0-a82fc549fd74%40columbia.edu?part=2
Assisted-by: Claude:claude-fable-5
Signed-off-by: Tal Zussman <tz2294@columbia.edu>
---
 block/fops.c | 21 ++++++++++++++++++++-
 1 file changed, 20 insertions(+), 1 deletion(-)

diff --git a/block/fops.c b/block/fops.c
index a3a709697b40..8769bb13df1c 100644
--- a/block/fops.c
+++ b/block/fops.c
@@ -87,6 +87,12 @@ static ssize_t __blkdev_direct_IO_simple(struct kiocb *iocb,
 	ret = blkdev_iov_iter_get_pages(&bio, iter, bdev);
 	if (unlikely(ret))
 		goto out;
+	if ((iocb->ki_flags & IOCB_ATOMIC) && iov_iter_count(iter)) {
+		/* a short atomic write would be torn by definition */
+		bio_release_pages(&bio, false);
+		ret = -EFAULT;
+		goto out;
+	}
 	ret = bio.bi_iter.bi_size;
 
 	if (iov_iter_rw(iter) == WRITE)
@@ -352,6 +358,12 @@ static ssize_t __blkdev_direct_IO_async(struct kiocb *iocb,
 		ret = blkdev_iov_iter_get_pages(bio, iter, bdev);
 		if (unlikely(ret))
 			goto out_bio_put;
+		if ((iocb->ki_flags & IOCB_ATOMIC) && iov_iter_count(iter)) {
+			/* a short atomic write would be torn by definition */
+			bio_release_pages(bio, false);
+			ret = -EFAULT;
+			goto out_bio_put;
+		}
 	}
 	dio->size = bio->bi_iter.bi_size;
 
@@ -691,8 +703,15 @@ blkdev_direct_write(struct kiocb *iocb, struct iov_iter *from)
 
 	written = kiocb_invalidate_pages(iocb, count);
 	if (written) {
-		if (written == -EBUSY)
+		/*
+		 * The buffered write fallback cannot provide torn-write
+		 * protection, so atomic writes must fail instead.
+		 */
+		if (written == -EBUSY) {
+			if (iocb->ki_flags & IOCB_ATOMIC)
+				return -EAGAIN;
 			return 0;
+		}
 		return written;
 	}
 

-- 
2.39.5


  parent reply	other threads:[~2026-08-28 13:50 UTC|newest]

Thread overview: 10+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-28 13:49 [PATCH v2 0/7] block device fixes for large block sizes, IOCB_NOWAIT, and direct I/O Tal Zussman
2026-08-28 13:49 ` [PATCH v2 1/7] block: use iomap_dirty_folio for block devices Tal Zussman
2026-08-28 13:49 ` [PATCH v2 2/7] block: take i_rwsem for the direct I/O write fallback Tal Zussman
2026-08-28 13:49 ` [PATCH v2 3/7] block: take i_rwsem for the splice read path Tal Zussman
2026-08-28 13:49 ` [PATCH v2 4/7] block: honor IOCB_NOWAIT in the block device buffered " Tal Zussman
2026-08-28 13:49 ` Tal Zussman [this message]
2026-08-28 13:49 ` [PATCH v2 6/7] block: unpin all pages of a bvec in bio_iov_iter_align_down() Tal Zussman
2026-08-28 14:36   ` Tal Zussman
2026-08-28 13:49 ` [PATCH v2 7/7] block: remove dead metadata handling from the async direct I/O path Tal Zussman
2026-08-28 15:32 ` [PATCH v2 0/7] block device fixes for large block sizes, IOCB_NOWAIT, and direct I/O Tal Zussman

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260828-blkdev-fixes-v2-5-32f3f40cebed@columbia.edu \
    --to=tz2294@columbia.edu \
    --cc=axboe@kernel.dk \
    --cc=brauner@kernel.org \
    --cc=djwong@kernel.org \
    --cc=hare@suse.de \
    --cc=hch@lst.de \
    --cc=johannes.thumshirn@wdc.com \
    --cc=john.g.garry@oracle.com \
    --cc=kbusch@kernel.org \
    --cc=linux-block@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=martin.petersen@oracle.com \
    --cc=mcgrof@kernel.org \
    --cc=sashiko-bot@kernel.org \
    --cc=willy@infradead.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.