From: Tal Zussman <tz2294@columbia.edu>
To: Jens Axboe <axboe@kernel.dk>, Christoph Hellwig <hch@lst.de>,
Johannes Thumshirn <johannes.thumshirn@wdc.com>,
Luis Chamberlain <mcgrof@kernel.org>,
Hannes Reinecke <hare@suse.de>,
"Matthew Wilcox (Oracle)" <willy@infradead.org>,
John Garry <john.g.garry@oracle.com>,
Christian Brauner <brauner@kernel.org>,
"Darrick J. Wong" <djwong@kernel.org>,
Keith Busch <kbusch@kernel.org>,
"Martin K. Petersen" <martin.petersen@oracle.com>
Cc: linux-block@vger.kernel.org, linux-kernel@vger.kernel.org,
Sashiko <sashiko-bot@kernel.org>,
Tal Zussman <tz2294@columbia.edu>
Subject: [PATCH v2 2/7] block: take i_rwsem for the direct I/O write fallback
Date: Fri, 28 Aug 2026 09:49:51 -0400 [thread overview]
Message-ID: <20260828-blkdev-fixes-v2-2-32f3f40cebed@columbia.edu> (raw)
In-Reply-To: <20260828-blkdev-fixes-v2-0-32f3f40cebed@columbia.edu>
Commit c0e473a0d226 ("block: fix race between set_blocksize and read
paths") closed a race between set_blocksize() and block device I/O: with
large sector size support, set_blocksize() can change i_blkbits and the
mapping's minimum folio order while a concurrent reader still holds a
folio of the old, smaller order, leading to crashes. In particular, it
made blkdev_write_iter() wrap buffered writes in inode_lock_shared().
However, the direct I/O fallback path was missed in that conversion.
blkdev_write_iter() passes blkdev_buffered_write() as an argument to
direct_write_fallback() with no lock held. A direct write that completes
only partially then finishes as a buffered write with no protection.
This can cause a BUG by racing partial direct writes against
ioctl(BLKBSZSET). Writer threads issue O_DIRECT pwritev() with a
two-segment iovec whose second segment is an unreadable PROT_NONE
mapping. The direct path then writes the first segment, fails to pin the
second, and returns short, entering the fallback. A second thread keeps
toggling the second segment's protection so that some fallbacks get past
fault_in_iov_iter_readable() and reach the page cache, a third thread
populates the page cache with folios of the current block size via
pread() and readahead(), and a fourth thread toggles the block size
between 512 bytes and 64K with BLKBSZSET. The minimum folio order only
moves with block sizes above PAGE_SIZE, i.e. with
CONFIG_TRANSPARENT_HUGEPAGE raising BLK_MAX_BLOCK_SIZE to 64K.
On a CONFIG_DEBUG_VM kernel this yields the following BUG:
page dumped because: VM_BUG_ON_FOLIO(folio_order(folio) < mapping_min_folio_order(mapping))
kernel BUG at mm/filemap.c:858!
Oops: invalid opcode: 0000 [#1] SMP KASAN NOPTI
RIP: 0010:__filemap_add_folio+0x860/0x8d0
Call Trace:
filemap_add_folio+0xc9/0x1f0
__filemap_get_folio_mpol+0x240/0x660
iomap_write_begin+0xa87/0xd70
iomap_file_buffered_write+0x304/0x6a0
blkdev_write_iter+0x255/0x510
do_iter_readv_writev+0x23d/0x3c0
vfs_writev+0x211/0x7d0
do_pwritev+0x121/0x190
do_syscall_64+0x121/0x630
entry_SYSCALL_64_after_hwframe+0x77/0x7f
The same workload also trips WARN_ON_ONCE(pos >= folio_pos(folio) +
fsize) in iomap_trim_folio_range().
Fix this by calling blkdev_buffered_write() in the fallback path under
inode_lock_shared(), matching the plain buffered-write branch. With the
fix the same workload runs clean.
A short IOCB_NOWAIT direct write reaches the same fallback. Taking
i_rwsem there can now block behind set_blocksize(), and the fallback
already blocks on writeback of the data it copied in
direct_write_fallback(). blkdev_write_iter() already rejects a purely
buffered IOCB_NOWAIT write with -EOPNOTSUPP, so do not enter the
fallback for IOCB_NOWAIT at all: return the bytes the direct path
already wrote, or -EAGAIN if none, and let the caller retry.
The reproducer used was written by an LLM, and is available at [1].
[1] https://gist.github.com/tzussman/69d06bc57d42a42989eb038b1b5aeb74
Fixes: 3c20917120ce ("block/bdev: enable large folio support for large logical block sizes")
Reported-by: Sashiko <sashiko-bot@kernel.org>
Link: https://sashiko.dev/#/patchset/20260730-blk-dontcache-v7-0-3e8e6850068d%40columbia.edu?part=5
Assisted-by: Claude:claude-fable-5
Signed-off-by: Tal Zussman <tz2294@columbia.edu>
---
block/fops.c | 23 ++++++++++++++++++++---
1 file changed, 20 insertions(+), 3 deletions(-)
diff --git a/block/fops.c b/block/fops.c
index c57784773fe1..d5f569333f46 100644
--- a/block/fops.c
+++ b/block/fops.c
@@ -765,9 +765,26 @@ static ssize_t blkdev_write_iter(struct kiocb *iocb, struct iov_iter *from)
if (iocb->ki_flags & IOCB_DIRECT) {
ret = blkdev_direct_write(iocb, from);
- if (ret >= 0 && iov_iter_count(from))
- ret = direct_write_fallback(iocb, from, ret,
- blkdev_buffered_write(iocb, from));
+ if (ret >= 0 && iov_iter_count(from)) {
+ if (iocb->ki_flags & IOCB_NOWAIT) {
+ /*
+ * The buffered fallback blocks on i_rwsem and
+ * on writeback of the data it copied: return
+ * the short direct write instead and let the
+ * caller retry.
+ */
+ if (!ret)
+ ret = -EAGAIN;
+ } else {
+ ssize_t ret2;
+
+ inode_lock_shared(bd_inode);
+ ret2 = blkdev_buffered_write(iocb, from);
+ inode_unlock_shared(bd_inode);
+ ret = direct_write_fallback(iocb, from, ret,
+ ret2);
+ }
+ }
} else {
/*
* Take i_rwsem and invalidate_lock to avoid racing with
--
2.39.5
next prev parent reply other threads:[~2026-08-28 13:50 UTC|newest]
Thread overview: 10+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-28 13:49 [PATCH v2 0/7] block device fixes for large block sizes, IOCB_NOWAIT, and direct I/O Tal Zussman
2026-08-28 13:49 ` [PATCH v2 1/7] block: use iomap_dirty_folio for block devices Tal Zussman
2026-08-28 13:49 ` Tal Zussman [this message]
2026-08-28 13:49 ` [PATCH v2 3/7] block: take i_rwsem for the splice read path Tal Zussman
2026-08-28 13:49 ` [PATCH v2 4/7] block: honor IOCB_NOWAIT in the block device buffered " Tal Zussman
2026-08-28 13:49 ` [PATCH v2 5/7] block: fail atomic writes instead of falling back to buffered I/O Tal Zussman
2026-08-28 13:49 ` [PATCH v2 6/7] block: unpin all pages of a bvec in bio_iov_iter_align_down() Tal Zussman
2026-08-28 14:36 ` Tal Zussman
2026-08-28 13:49 ` [PATCH v2 7/7] block: remove dead metadata handling from the async direct I/O path Tal Zussman
2026-08-28 15:32 ` [PATCH v2 0/7] block device fixes for large block sizes, IOCB_NOWAIT, and direct I/O Tal Zussman
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260828-blkdev-fixes-v2-2-32f3f40cebed@columbia.edu \
--to=tz2294@columbia.edu \
--cc=axboe@kernel.dk \
--cc=brauner@kernel.org \
--cc=djwong@kernel.org \
--cc=hare@suse.de \
--cc=hch@lst.de \
--cc=johannes.thumshirn@wdc.com \
--cc=john.g.garry@oracle.com \
--cc=kbusch@kernel.org \
--cc=linux-block@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=martin.petersen@oracle.com \
--cc=mcgrof@kernel.org \
--cc=sashiko-bot@kernel.org \
--cc=willy@infradead.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.