Linux NFS development
 help / color / mirror / Atom feed
From: Mike Snitzer <snitzer@kernel.org>
To: Chuck Lever <cel@kernel.org>, Jeff Layton <jlayton@kernel.org>
Cc: hch@lst.de, linux-nfs@vger.kernel.org
Subject: [PATCH v3 1/9] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well
Date: Thu,  1 Oct 2026 00:54:54 -0400	[thread overview]
Message-ID: <20261001045502.48381-2-snitzer@kernel.org> (raw)
In-Reply-To: <20261001045502.48381-1-snitzer@kernel.org>

nfsd_write_dio_iters_init() issues the DIO-aligned middle of a
misaligned WRITE with IOCB_DIRECT.  A file system may service that
request with buffered I/O instead: xfs_file_write_iter() retries
through xfs_file_buffered_write() with the same kiocb when
xfs_file_dio_write() returns -ENOTBLK, which iomap_dio_rw() does when
it cannot invalidate page cache overlapping the range.

In a streaming misaligned WRITE workload with more than one request in
flight this happens routinely.  Request N's direct middle ends in the
page where request N+1's buffered prefix starts.  If N+1 dirties that
page between N's writeback and N's invalidation, the invalidation
fails and XFS re-issues all of N's middle, tens of kilobytes, as a
cached buffered write that nothing ever drops.

Set IOCB_DONTCACHE on the middle segment alongside IOCB_DIRECT when
the file system supports FOP_DONTCACHE.  The direct path ignores the
flag; a buffered fallback now drops its pages once written back.

nfsd-io-modes.rst said that DONTCACHE is not used for the misaligned
segments, which the write loop already contradicted: every segment
without IOCB_DIRECT gets IOCB_DONTCACHE when the file system supports
it.  Replace that sentence with one that covers all three segments.

Fixes: 06c5c97293e3 ("NFSD: Implement NFSD_IO_DIRECT for NFS WRITE")
Assisted-by: Claude:claude-fable-5-1
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
 Documentation/filesystems/nfs/nfsd-io-modes.rst | 8 +++++---
 fs/nfsd/vfs.c                                   | 3 +++
 2 files changed, 8 insertions(+), 3 deletions(-)

diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index 0fd6e82478fe6..60b0af9b7e49f 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -126,9 +126,11 @@ Misaligned WRITE:
     middle and end as needed. The large middle segment is DIO-aligned
     and the start and/or end are misaligned. Buffered IO is used for the
     misaligned segments and O_DIRECT is used for the middle DIO-aligned
-    segment. DONTCACHE buffered IO is _not_ used for the misaligned
-    segments because using normal buffered IO offers significant RMW
-    performance benefit when handling streaming misaligned WRITEs.
+    segment. If the filesystem supports FOP_DONTCACHE, every segment is
+    marked DONTCACHE. The flag has no effect on the O_DIRECT segment
+    unless the filesystem services it with buffered IO instead, as XFS
+    does when it cannot invalidate page cache that overlaps the segment.
+    The iomap_dio_invalidate_fail trace event reports such a fallback.
 
 Tracing:
     The nfsd_read_direct trace event shows how NFSD expands any
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 4584d5b94feed..1ccdac4693745 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1341,6 +1341,9 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
 	if (iov_iter_bvec_offset(&segments[nsegs].iter) & (mem_align - 1))
 		goto no_dio;
 	segments[nsegs].flags |= IOCB_DIRECT;
+	/* In case the file system falls back to buffered I/O (-ENOTBLK). */
+	if (nf->nf_file->f_op->fop_flags & FOP_DONTCACHE)
+		segments[nsegs].flags |= IOCB_DONTCACHE;
 	nsegs++;
 
 	if (suffix)

base-commit: 32eb1a60b456980761cf7a9cee8f907fdc08afb8
-- 
2.52.0


  reply	other threads:[~2026-10-01  4:55 UTC|newest]

Thread overview: 11+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-01  4:54 [PATCH v3 0/9] NFSD: keep direct-mode I/O out of the page cache Mike Snitzer
2026-10-01  4:54 ` Mike Snitzer [this message]
2026-10-01  4:54 ` [PATCH v3 2/9] NFSD: only split a direct-mode WRITE for a worthwhile direct middle Mike Snitzer
2026-10-01  4:54 ` [PATCH v3 3/9] NFSD: add direct_misaligned_dontcache debugfs knob Mike Snitzer
2026-10-01  4:54 ` [PATCH v3 4/9] NFSD: do not use direct I/O for a READ smaller than its alignment Mike Snitzer
2026-10-01  4:54 ` [PATCH v3 5/9] NFSD: persist a synchronous direct-mode WRITE once, after all of its segments Mike Snitzer
2026-10-01  4:54 ` [PATCH v3 6/9] NFSD: keep boundary page of a split direct-mode WRITE until both writers complete Mike Snitzer
2026-10-01  4:55 ` [PATCH v3 7/9] NFSD: add tracing for how direct-mode READ and WRITE are serviced Mike Snitzer
2026-10-01  4:55 ` [PATCH v3 8/9] NFSD: Enable return of an updated stable_how to NFS clients Mike Snitzer
2026-10-01  4:55 ` [PATCH v3 9/9] NFSD: add direct-mode WRITE settings that persist each WRITE Mike Snitzer
2026-10-01 22:58 ` [PATCH v3 0/9] NFSD: keep direct-mode I/O out of the page cache Mike Snitzer

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261001045502.48381-2-snitzer@kernel.org \
    --to=snitzer@kernel.org \
    --cc=cel@kernel.org \
    --cc=hch@lst.de \
    --cc=jlayton@kernel.org \
    --cc=linux-nfs@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox