Linux NFS development
 help / color / mirror / Atom feed
From: Mike Snitzer <snitzer@hammerspace.com>
To: Chuck Lever <cel@kernel.org>, Jeff Layton <jlayton@kernel.org>
Cc: linux-nfs@vger.kernel.org
Subject: [PATCH 03/10] NFSD: only split a direct-mode WRITE for a worthwhile direct middle
Date: Tue, 29 Sep 2026 13:34:16 -0400	[thread overview]
Message-ID: <20260929173423.16149-4-snitzer@kernel.org> (raw)
In-Reply-To: <20260929173423.16149-1-snitzer@kernel.org>

nfsd_write_dio_iters_init() decides how a WRITE issued in a direct mode
is carried out: split into a buffered prefix, a direct middle and a
buffered suffix, or issued whole as a single buffered segment.

Only split when the split buys direct I/O for a worthwhile middle.  Add
direct_misaligned_num_pages (debugfs knob, default 2) and decline to
split when the aligned middle is smaller than that many pages and the
WRITE has a prefix or a suffix.  A WRITE whose payload memory is
misaligned for the block device already declines: no segment of it can
be direct I/O, so a split would only turn one buffered write into
three.

Decide every segment's flags in nfsd_write_dio_iters_init() rather than
in the write loop.  The buffered prefix and suffix of a split WRITE,
and the whole WRITE when it is not split, are issued IOCB_DONTCACHE
when the file system supports FOP_DONTCACHE and as plain buffered I/O
otherwise: the operator chose a direct mode to keep NFSD out of the
page cache, so a WRITE that cannot be direct gets the next best thing.

Document it in the "Misaligned WRITE" section of nfsd-io-modes.rst.

Suggested-by: Christoph Hellwig <hch@infradead.org>
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
 .../filesystems/nfs/nfsd-io-modes.rst         | 17 ++++-
 fs/nfsd/debugfs.c                             | 14 ++++
 fs/nfsd/nfsd.h                                |  1 +
 fs/nfsd/vfs.c                                 | 73 +++++++++++++------
 4 files changed, 79 insertions(+), 26 deletions(-)

diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index c15983e634a3e..f4e7cee5ee159 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -161,9 +161,9 @@ Misaligned WRITE:
     middle and end as needed. The large middle segment is DIO-aligned
     and the start and/or end are misaligned. Buffered IO is used for the
     misaligned segments and O_DIRECT is used for the middle DIO-aligned
-    segment. DONTCACHE buffered IO is _not_ used for the misaligned
-    segments because using normal buffered IO offers significant RMW
-    performance benefit when handling streaming misaligned WRITEs.
+    segment. If the underlying filesystem supports FOP_DONTCACHE, the
+    misaligned segments use DONTCACHE buffered IO so that their pages
+    are dropped from the page cache once written back.
 
     The O_DIRECT middle segment also carries the DONTCACHE flag. It has
     no effect while the IO really is O_DIRECT, but a filesystem may
@@ -175,6 +175,17 @@ Misaligned WRITE:
     normal buffered IO. Such fallbacks are visible through the
     iomap_dio_invalidate_fail trace event; see Tracing below.
 
+    Whenever no part of a WRITE can use O_DIRECT, the whole WRITE is
+    issued as a single DONTCACHE buffered IO (normal buffered IO if the
+    filesystem lacks FOP_DONTCACHE). This covers: a filesystem that
+    advertises no DIO alignment requirements at all; a WRITE smaller
+    than the larger of the offset and memory alignments; a WRITE whose
+    DIO-aligned middle segment is smaller than
+    /sys/kernel/debug/nfsd/direct_misaligned_num_pages pages (default 2)
+    while also having a misaligned start or end; and a WRITE whose
+    payload memory is not aligned to the block device's dma_alignment,
+    which rules out O_DIRECT for the middle segment as well.
+
 Tracing:
     The nfsd_read_direct trace event shows how NFSD expands any
     misaligned READ to the next DIO-aligned block (on either end of the
diff --git a/fs/nfsd/debugfs.c b/fs/nfsd/debugfs.c
index 995872a1a7f87..afa41bf2d8189 100644
--- a/fs/nfsd/debugfs.c
+++ b/fs/nfsd/debugfs.c
@@ -170,6 +170,17 @@ void nfsd_debugfs_exit(void)
 	nfsd_top_dir = NULL;
 }
 
+/*
+ * /sys/kernel/debug/nfsd/direct_misaligned_num_pages
+ *
+ * The smallest DIO-aligned middle segment, in pages, that is worth
+ * splitting a misaligned direct-mode WRITE into three segments for.  A
+ * WRITE whose middle is smaller than this, and which has a misaligned
+ * start or end, is issued as a single buffered segment instead.
+ *
+ * Default 2.  Not yet tuned by benchmarking.
+ */
+
 void nfsd_debugfs_init(void)
 {
 	nfsd_top_dir = debugfs_create_dir("nfsd", NULL);
@@ -182,6 +193,9 @@ void nfsd_debugfs_init(void)
 
 	debugfs_create_file("io_cache_write", 0644, nfsd_top_dir, NULL,
 			    &nfsd_io_cache_write_fops);
+
+	debugfs_create_u32("direct_misaligned_num_pages", 0644, nfsd_top_dir,
+			   &nfsd_direct_misaligned_num_pages);
 #ifdef CONFIG_NFSD_V4
 	debugfs_create_bool("delegated_timestamps", 0644, nfsd_top_dir,
 			    &nfsd_delegts_enabled);
diff --git a/fs/nfsd/nfsd.h b/fs/nfsd/nfsd.h
index a145294c59c87..a2d72434160af 100644
--- a/fs/nfsd/nfsd.h
+++ b/fs/nfsd/nfsd.h
@@ -142,6 +142,7 @@ enum {
 
 extern u64 nfsd_io_cache_read __read_mostly;
 extern u64 nfsd_io_cache_write __read_mostly;
+extern u32 nfsd_direct_misaligned_num_pages __read_mostly;
 
 extern int nfsd_max_blksize;
 
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 5b963f0e3b3f2..e3ce66bce00d4 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -53,6 +53,7 @@
 bool nfsd_disable_splice_read __read_mostly;
 u64 nfsd_io_cache_read __read_mostly = NFSD_IO_BUFFERED;
 u64 nfsd_io_cache_write __read_mostly = NFSD_IO_BUFFERED;
+u32 nfsd_direct_misaligned_num_pages __read_mostly = 2;
 
 /**
  * nfserrno - Map Linux errnos to NFS errnos
@@ -1302,15 +1303,29 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
 	u32 mem_align = nf->nf_dio_mem_align;
 	size_t prefix, middle, suffix;
 	loff_t offset = iocb->ki_pos;
+	unsigned int dontcache_flags = 0;
 	unsigned int nsegs = 0;
 
+	if (nf->nf_file->f_op->fop_flags & FOP_DONTCACHE)
+		dontcache_flags = IOCB_DONTCACHE;
+
 	/*
-	 * Check if direct I/O is feasible for this write request.
-	 * If alignments are not available, the write is too small,
-	 * or no alignment can be found, fall back to buffered I/O.
+	 * Whenever direct I/O cannot be used for the WRITE, fall back to a
+	 * single DONTCACHE buffered I/O when the file system supports it, so
+	 * the WRITE's pages are dropped from the page cache once written
+	 * back, and to a single cached buffered I/O otherwise.
+	 *
+	 * If the file system doesn't advertise any alignment requirements,
+	 * don't try to issue direct I/O at all.
 	 */
-	if (unlikely(!mem_align || !offset_align) ||
-	    unlikely(total < max(offset_align, mem_align)))
+	if (unlikely(!mem_align || !offset_align))
+		goto no_dio;
+
+	/*
+	 * If the I/O is smaller than the larger of the memory and logical
+	 * offset alignment, no part of it can be direct I/O.
+	 */
+	if (unlikely(total < max(offset_align, mem_align)))
 		goto no_dio;
 
 	prefix_end = round_up(offset, offset_align);
@@ -1321,12 +1336,27 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
 	middle = middle_end - prefix_end;
 	suffix = orig_end - middle_end;
 
-	if (!middle)
+	/*
+	 * If there is no aligned middle section, or the aligned part is too
+	 * small to be worth the split (direct_misaligned_num_pages), issue a
+	 * single buffered I/O write instead of splitting up the write.
+	 */
+	if (!middle ||
+	    ((prefix || suffix) &&
+	     middle < PAGE_SIZE * nfsd_direct_misaligned_num_pages)) {
 		goto no_dio;
+	}
 
-	if (prefix)
-		nfsd_write_dio_seg_init(&segments[nsegs++], bvec,
+	/*
+	 * The prefix and suffix are buffered I/O by definition.  Mark them
+	 * uncached when possible so their folios are dropped once written
+	 * back rather than lingering in the page cache.
+	 */
+	if (prefix) {
+		nfsd_write_dio_seg_init(&segments[nsegs], bvec,
 					nvecs, total, 0, prefix, iocb);
+		segments[nsegs++].flags |= dontcache_flags;
+	}
 
 	nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs,
 				total, prefix, middle, iocb);
@@ -1337,10 +1367,13 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
 	 * bvecs generated from RPC receive buffers are contiguous: After
 	 * the first bvec, all subsequent bvecs start at bv_offset zero
 	 * (page-aligned). Therefore, only the first bvec is checked.
+	 *
+	 * If the memory is not aligned, direct I/O is impossible for the
+	 * middle, so issue the entire write as a single buffered segment:
+	 * splitting would only turn one buffered write into three.
 	 */
 	if (iov_iter_bvec_offset(&segments[nsegs].iter) & (mem_align - 1))
 		goto no_dio;
-	segments[nsegs].flags |= IOCB_DIRECT;
 	/*
 	 * Also mark the direct middle DONTCACHE: the file system may fall
 	 * back to buffered I/O on its own (e.g. XFS on -ENOTBLK when it
@@ -1348,20 +1381,21 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
 	 * suffix of an adjacent WRITE just dirtied), and it reuses this kiocb
 	 * to do so.  On the direct path itself the flag is inert.
 	 */
-	if (nf->nf_file->f_op->fop_flags & FOP_DONTCACHE)
-		segments[nsegs].flags |= IOCB_DONTCACHE;
-	nsegs++;
+	segments[nsegs++].flags |= IOCB_DIRECT | dontcache_flags;
 
-	if (suffix)
-		nfsd_write_dio_seg_init(&segments[nsegs++], bvec, nvecs, total,
+	if (suffix) {
+		nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs, total,
 					prefix + middle, suffix, iocb);
+		segments[nsegs++].flags |= dontcache_flags;
+	}
 
 	return nsegs;
 
 no_dio:
-	/* No DIO alignment possible - pack into single non-DIO segment. */
+	/* No DIO possible - pack into a single uncached (if possible) segment. */
 	nfsd_write_dio_seg_init(&segments[0], bvec, nvecs, total, 0,
 				total, iocb);
+	segments[0].flags |= dontcache_flags;
 	return 1;
 }
 
@@ -1385,16 +1419,9 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
 		if (kiocb->ki_flags & IOCB_DIRECT)
 			trace_nfsd_write_direct(rqstp, fhp, kiocb->ki_pos,
 						segments[i].iter.count);
-		else {
+		else
 			trace_nfsd_write_vector(rqstp, fhp, kiocb->ki_pos,
 						segments[i].iter.count);
-			/*
-			 * Mark the I/O buffer as evict-able to reduce
-			 * memory contention.
-			 */
-			if (nf->nf_file->f_op->fop_flags & FOP_DONTCACHE)
-				kiocb->ki_flags |= IOCB_DONTCACHE;
-		}
 
 		expected = iov_iter_count(&segments[i].iter);
 
-- 
2.52.0


  parent reply	other threads:[~2026-09-29 17:34 UTC|newest]

Thread overview: 17+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-29 17:34 [PATCH 00/10] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
2026-09-29 17:34 ` [PATCH 01/10] NFSD: interlock the use of NFSD_IO_DIRECT for NFS READ and WRITE Mike Snitzer
2026-09-29 18:27   ` Chuck Lever
2026-09-29 19:56     ` Mike Snitzer
2026-09-29 23:17       ` Chuck Lever
2026-09-29 23:30         ` Mike Snitzer
2026-09-30  0:20           ` Chuck Lever
2026-09-30 12:46             ` Mike Snitzer
2026-09-29 17:34 ` [PATCH 02/10] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well Mike Snitzer
2026-09-29 17:34 ` Mike Snitzer [this message]
2026-09-29 17:34 ` [PATCH 04/10] NFSD: do not use direct I/O for a READ smaller than its alignment Mike Snitzer
2026-09-29 17:34 ` [PATCH 05/10] NFSD: Enable return of an updated stable_how to NFS clients Mike Snitzer
2026-09-29 17:34 ` [PATCH 06/10] NFSD: let a direct-mode WRITE raise stable_how and elide the client's COMMIT Mike Snitzer
2026-09-29 17:34 ` [PATCH 07/10] NFSD: persist a synchronous direct-mode WRITE once, after all of its segments Mike Snitzer
2026-09-29 17:34 ` [PATCH 08/10] NFSD: keep boundary page of a split direct-mode WRITE until both writers complete Mike Snitzer
2026-09-29 17:34 ` [PATCH 09/10] NFSD: add direct_misaligned_dontcache debugfs knob Mike Snitzer
2026-09-29 17:34 ` [PATCH 10/10] NFSD: add tracing for how direct-mode READ and WRITE are serviced Mike Snitzer

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260929173423.16149-4-snitzer@kernel.org \
    --to=snitzer@hammerspace.com \
    --cc=cel@kernel.org \
    --cc=jlayton@kernel.org \
    --cc=linux-nfs@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox