Linux NFS development
 help / color / mirror / Atom feed
From: Mike Snitzer <snitzer@kernel.org>
To: Chuck Lever <cel@kernel.org>, Jeff Layton <jlayton@kernel.org>
Cc: linux-nfs@vger.kernel.org
Subject: [PATCH v2 1/9] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well
Date: Tue, 29 Sep 2026 19:13:21 -0400	[thread overview]
Message-ID: <20260929231329.22018-2-snitzer@kernel.org> (raw)
In-Reply-To: <20260929231329.22018-1-snitzer@kernel.org>

nfsd_write_dio_iters_init() issues the DIO-aligned middle of a
misaligned WRITE with IOCB_DIRECT.  A file system is free to service
that request with buffered I/O instead: xfs_file_write_iter() retries
via xfs_file_buffered_write() with the same kiocb whenever
xfs_file_dio_write() returns -ENOTBLK, which iomap_dio_rw() produces
when the invalidation it runs before issuing the direct I/O cannot drop
page cache overlapping the range.

That happens routinely in a streaming misaligned WRITE workload with
more than one request in flight.  Request N's direct middle ends inside
the same page that request N+1's buffered prefix starts in.  If N+1's
prefix dirties that page between N's filemap_write_and_wait_range() and
N's invalidate_inode_pages2_range(), the invalidation returns -EBUSY,
iomap returns -ENOTBLK, and XFS re-issues N's entire middle, tens of
kilobytes, as a normal cached buffered write.  Nothing ever drops those
pages.  On a device with dma_alignment=3 (XFS then reports
dio_mem_align=4, which XDR-aligned RPC payloads always satisfy) this is
the only remaining path by which a stream of small misaligned WRITEs in
DIRECT mode fills the page cache.

Set IOCB_DONTCACHE on the middle segment alongside IOCB_DIRECT when the
file system supports FOP_DONTCACHE.  The direct path ignores the flag,
and generic_write_sync() only issues a harmless flusher kick; but if
the file system falls back to buffered I/O the write is now DONTCACHE
rather than cached.  The nfsd_write_direct trace event is unchanged
since it keys on IOCB_DIRECT; the fallback itself remains observable
via the iomap_dio_invalidate_fail event.

Fixes: 06c5c97293e3 ("NFSD: Implement NFSD_IO_DIRECT for NFS WRITE")
Assisted-by: Claude:claude-fable-5-1
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
 Documentation/filesystems/nfs/nfsd-io-modes.rst | 10 ++++++++++
 fs/nfsd/vfs.c                                   |  9 +++++++++
 2 files changed, 19 insertions(+)

diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index 0fd6e82478fe6..dc50c930f9762 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -130,6 +130,16 @@ Misaligned WRITE:
     segments because using normal buffered IO offers significant RMW
     performance benefit when handling streaming misaligned WRITEs.
 
+    The O_DIRECT middle segment also carries the DONTCACHE flag. It has
+    no effect while the IO really is O_DIRECT, but a filesystem may
+    decide on its own to service the segment with buffered IO instead
+    (XFS does so when it cannot invalidate page cache that overlaps the
+    segment, which can happen when another WRITE's buffered start or
+    end segment dirties the shared boundary page at the same time).
+    The flag makes that fallback DONTCACHE buffered IO rather than
+    normal buffered IO. Such fallbacks are visible through the
+    iomap_dio_invalidate_fail trace event; see Tracing below.
+
 Tracing:
     The nfsd_read_direct trace event shows how NFSD expands any
     misaligned READ to the next DIO-aligned block (on either end of the
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index f9131827d391e..5b963f0e3b3f2 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1341,6 +1341,15 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
 	if (iov_iter_bvec_offset(&segments[nsegs].iter) & (mem_align - 1))
 		goto no_dio;
 	segments[nsegs].flags |= IOCB_DIRECT;
+	/*
+	 * Also mark the direct middle DONTCACHE: the file system may fall
+	 * back to buffered I/O on its own (e.g. XFS on -ENOTBLK when it
+	 * cannot invalidate page cache that a concurrent buffered prefix or
+	 * suffix of an adjacent WRITE just dirtied), and it reuses this kiocb
+	 * to do so.  On the direct path itself the flag is inert.
+	 */
+	if (nf->nf_file->f_op->fop_flags & FOP_DONTCACHE)
+		segments[nsegs].flags |= IOCB_DONTCACHE;
 	nsegs++;
 
 	if (suffix)
-- 
2.52.0


  reply	other threads:[~2026-09-29 23:13 UTC|newest]

Thread overview: 14+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-29 23:13 [PATCH v2 0/9] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
2026-09-29 23:13 ` Mike Snitzer [this message]
2026-09-30 21:43   ` [PATCH v2 1/9] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well Chuck Lever
2026-09-29 23:13 ` [PATCH v2 2/9] NFSD: only split a direct-mode WRITE for a worthwhile direct middle Mike Snitzer
2026-09-30 21:44   ` Chuck Lever
2026-09-29 23:13 ` [PATCH v2 3/9] NFSD: do not use direct I/O for a READ smaller than its alignment Mike Snitzer
2026-09-29 23:13 ` [PATCH v2 4/9] NFSD: Enable return of an updated stable_how to NFS clients Mike Snitzer
2026-09-30 21:42   ` Chuck Lever
2026-09-29 23:13 ` [PATCH v2 5/9] NFSD: let a direct-mode WRITE raise stable_how and elide the client's COMMIT Mike Snitzer
2026-09-30 21:46   ` Chuck Lever
2026-09-29 23:13 ` [PATCH v2 6/9] NFSD: persist a synchronous direct-mode WRITE once, after all of its segments Mike Snitzer
2026-09-29 23:13 ` [PATCH v2 7/9] NFSD: keep boundary page of a split direct-mode WRITE until both writers complete Mike Snitzer
2026-09-29 23:13 ` [PATCH v2 8/9] NFSD: add direct_misaligned_dontcache debugfs knob Mike Snitzer
2026-09-29 23:13 ` [PATCH v2 9/9] NFSD: add tracing for how direct-mode READ and WRITE are serviced Mike Snitzer

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260929231329.22018-2-snitzer@kernel.org \
    --to=snitzer@kernel.org \
    --cc=cel@kernel.org \
    --cc=jlayton@kernel.org \
    --cc=linux-nfs@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox