Linux NFS development
 help / color / mirror / Atom feed
From: Mike Snitzer <snitzer@kernel.org>
To: Chuck Lever <cel@kernel.org>, Jeff Layton <jlayton@kernel.org>
Cc: linux-nfs@vger.kernel.org
Subject: [PATCH v2 8/9] NFSD: add direct_misaligned_dontcache debugfs knob
Date: Tue, 29 Sep 2026 19:13:28 -0400	[thread overview]
Message-ID: <20260929231329.22018-9-snitzer@kernel.org> (raw)
In-Reply-To: <20260929231329.22018-1-snitzer@kernel.org>

From: Jonathan Flynn <jonathan.flynn@hammerspace.com>

The parts of a direct-mode WRITE that cannot be direct I/O, the
misaligned prefix and suffix of a split WRITE and the whole WRITE when
it is not split, are issued IOCB_DONTCACHE so their pages are dropped
once written back.  That is the right default for the workloads a direct
mode is chosen for, but it is a policy rather than a requirement: a
workload that reads back what it just wrote, or that keeps rewriting the
same partial pages, is better served by those pages staying in the page
cache.

Add a bool debugfs knob, /sys/kernel/debug/nfsd/direct_misaligned_dontcache
(default Y), to choose between the two:

  Y: IOCB_DONTCACHE when the file system supports it, with a split's
     boundary page kept in the page cache until both WRITEs sharing it
     have written it (nfsd_write_dio_boundary_claim()).
  N: ordinary cached buffered I/O; nothing is claimed or marked, and the
     pages stay until reclaim.

The direct middle and the once-per-WRITE persist are unaffected.  The
knob is sampled once per WRITE, so a change takes effect immediately and
without a remount.  It sits beside io_cache_read and io_cache_write, the
other controls over how NFSD issues its I/O.

[snitzer: switched from using modparam to debugfs knob]

Signed-off-by: Jonathan Flynn <jonathan.flynn@hammerspace.com>
[snitzer: documented in nfsd-io-modes.rst]
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
 .../filesystems/nfs/nfsd-io-modes.rst         |  9 ++++++
 fs/nfsd/debugfs.c                             | 21 ++++++++++++++
 fs/nfsd/nfsd.h                                |  1 +
 fs/nfsd/vfs.c                                 | 28 ++++++++++++-------
 4 files changed, 49 insertions(+), 10 deletions(-)

diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index 2ca606ddf5d51..f2c380ca000f7 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -178,6 +178,15 @@ Misaligned WRITE:
     with how far concurrent writers drift apart, not with bytes
     written.
 
+    Whether those pages are dropped at all is a policy choice, selected
+    by /sys/kernel/debug/nfsd/direct_misaligned_dontcache (default Y).
+    Write N to issue the start and end segments, and the whole-WRITE
+    fallbacks, as ordinary cached buffered IO: nothing is claimed or
+    marked and the pages stay until reclaim, which suits a workload that
+    reads back or rewrites what it just wrote. The O_DIRECT middle
+    segment is unaffected. The knob is sampled once per WRITE, so a
+    change takes effect immediately.
+
     The O_DIRECT middle segment also carries the DONTCACHE flag. It has
     no effect while the IO really is O_DIRECT, but a filesystem may
     decide on its own to service the segment with buffered IO instead
diff --git a/fs/nfsd/debugfs.c b/fs/nfsd/debugfs.c
index 279e341a81ec1..2b728eac8532c 100644
--- a/fs/nfsd/debugfs.c
+++ b/fs/nfsd/debugfs.c
@@ -144,6 +144,24 @@ void nfsd_debugfs_exit(void)
  * Default 2.  Not yet tuned by benchmarking.
  */
 
+/*
+ * /sys/kernel/debug/nfsd/direct_misaligned_dontcache
+ *
+ * How a direct-mode WRITE issues the I/O that cannot be direct: the
+ * misaligned start and end of a split WRITE, and the whole WRITE when it
+ * is not split.
+ *
+ * Contents:
+ *   Y: DONTCACHE when the filesystem supports it, with the boundary page
+ *      of a split kept in the page cache only until both WRITEs sharing
+ *      it have written it
+ *   N: ordinary cached buffered IO, left in the page cache until
+ *      reclaim, for A/B comparison against the DONTCACHE path
+ *
+ * Sampled once per WRITE, so it takes effect immediately.  The direct
+ * middle segment is unaffected.
+ */
+
 void nfsd_debugfs_init(void)
 {
 	nfsd_top_dir = debugfs_create_dir("nfsd", NULL);
@@ -159,6 +177,9 @@ void nfsd_debugfs_init(void)
 
 	debugfs_create_u32("direct_misaligned_num_pages", 0644, nfsd_top_dir,
 			   &nfsd_direct_misaligned_num_pages);
+
+	debugfs_create_bool("direct_misaligned_dontcache", 0644, nfsd_top_dir,
+			    &nfsd_direct_misaligned_dontcache);
 #ifdef CONFIG_NFSD_V4
 	debugfs_create_bool("delegated_timestamps", 0644, nfsd_top_dir,
 			    &nfsd_delegts_enabled);
diff --git a/fs/nfsd/nfsd.h b/fs/nfsd/nfsd.h
index 135e319e378d4..dff979ac370ba 100644
--- a/fs/nfsd/nfsd.h
+++ b/fs/nfsd/nfsd.h
@@ -145,6 +145,7 @@ enum {
 extern u64 nfsd_io_cache_read __read_mostly;
 extern u64 nfsd_io_cache_write __read_mostly;
 extern u32 nfsd_direct_misaligned_num_pages __read_mostly;
+extern bool nfsd_direct_misaligned_dontcache __read_mostly;
 
 extern int nfsd_max_blksize;
 
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 5fd850a29694f..7896e2e6c5855 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -54,6 +54,7 @@ bool nfsd_disable_splice_read __read_mostly;
 u64 nfsd_io_cache_read __read_mostly = NFSD_IO_BUFFERED;
 u64 nfsd_io_cache_write __read_mostly = NFSD_IO_BUFFERED;
 u32 nfsd_direct_misaligned_num_pages __read_mostly = 2;
+bool nfsd_direct_misaligned_dontcache __read_mostly = true;
 
 /**
  * nfserrno - Map Linux errnos to NFS errnos
@@ -1373,16 +1374,21 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
 	size_t prefix, middle, suffix;
 	loff_t offset = iocb->ki_pos;
 	unsigned int dontcache_flags = 0;
+	unsigned int buffered_flags;
 	unsigned int nsegs = 0;
 
 	if (nf->nf_file->f_op->fop_flags & FOP_DONTCACHE)
 		dontcache_flags = IOCB_DONTCACHE;
+	/* Buffered segments follow the knob; the direct middle does not. */
+	buffered_flags = READ_ONCE(nfsd_direct_misaligned_dontcache) ?
+			 dontcache_flags : 0;
 
 	/*
 	 * Whenever direct I/O cannot be used for the WRITE, fall back to a
-	 * single DONTCACHE buffered I/O when the file system supports it, so
-	 * the WRITE's pages are dropped from the page cache once written
-	 * back, and to a single cached buffered I/O otherwise.
+	 * single DONTCACHE buffered I/O when the file system supports it (and
+	 * nfsd_direct_misaligned_dontcache is set), so the WRITE's pages are
+	 * dropped from the page cache once written back, and to a single
+	 * cached buffered I/O otherwise.
 	 *
 	 * If the file system doesn't advertise any alignment requirements,
 	 * don't try to issue direct I/O at all.
@@ -1421,13 +1427,15 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
 	 * its page with the neighbouring WRITE; see
 	 * nfsd_write_dio_boundary_claim(), which nfsd_direct_write() calls right
 	 * before issuing each of them, for how the page is held for the
-	 * partner and dropped once both have written it.
+	 * partner and dropped once both have written it.  With
+	 * nfsd_direct_misaligned_dontcache=N both are plain cached writes and
+	 * nothing is held or dropped: the pages stay until reclaim.
 	 */
 	if (prefix) {
 		nfsd_write_dio_seg_init(&segments[nsegs], bvec,
 					nvecs, total, 0, prefix, iocb);
-		segments[nsegs].flags |= dontcache_flags;
-		segments[nsegs++].boundary = !!dontcache_flags;
+		segments[nsegs].flags |= buffered_flags;
+		segments[nsegs++].boundary = !!buffered_flags;
 	}
 
 	nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs,
@@ -1458,8 +1466,8 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
 	if (suffix) {
 		nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs, total,
 					prefix + middle, suffix, iocb);
-		segments[nsegs].flags |= dontcache_flags;
-		segments[nsegs++].boundary = !!dontcache_flags;
+		segments[nsegs].flags |= buffered_flags;
+		segments[nsegs++].boundary = !!buffered_flags;
 	}
 
 	return nsegs;
@@ -1473,8 +1481,8 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
 	 */
 	nfsd_write_dio_seg_init(&segments[0], bvec, nvecs, total, 0,
 				total, iocb);
-	segments[0].flags |= dontcache_flags;
-	segments[0].edges = !!dontcache_flags;
+	segments[0].flags |= buffered_flags;
+	segments[0].edges = !!buffered_flags;
 	return 1;
 }
 
-- 
2.52.0


  parent reply	other threads:[~2026-09-29 23:13 UTC|newest]

Thread overview: 14+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-29 23:13 [PATCH v2 0/9] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
2026-09-29 23:13 ` [PATCH v2 1/9] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well Mike Snitzer
2026-09-30 21:43   ` Chuck Lever
2026-09-29 23:13 ` [PATCH v2 2/9] NFSD: only split a direct-mode WRITE for a worthwhile direct middle Mike Snitzer
2026-09-30 21:44   ` Chuck Lever
2026-09-29 23:13 ` [PATCH v2 3/9] NFSD: do not use direct I/O for a READ smaller than its alignment Mike Snitzer
2026-09-29 23:13 ` [PATCH v2 4/9] NFSD: Enable return of an updated stable_how to NFS clients Mike Snitzer
2026-09-30 21:42   ` Chuck Lever
2026-09-29 23:13 ` [PATCH v2 5/9] NFSD: let a direct-mode WRITE raise stable_how and elide the client's COMMIT Mike Snitzer
2026-09-30 21:46   ` Chuck Lever
2026-09-29 23:13 ` [PATCH v2 6/9] NFSD: persist a synchronous direct-mode WRITE once, after all of its segments Mike Snitzer
2026-09-29 23:13 ` [PATCH v2 7/9] NFSD: keep boundary page of a split direct-mode WRITE until both writers complete Mike Snitzer
2026-09-29 23:13 ` Mike Snitzer [this message]
2026-09-29 23:13 ` [PATCH v2 9/9] NFSD: add tracing for how direct-mode READ and WRITE are serviced Mike Snitzer

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260929231329.22018-9-snitzer@kernel.org \
    --to=snitzer@kernel.org \
    --cc=cel@kernel.org \
    --cc=jlayton@kernel.org \
    --cc=linux-nfs@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox