From: Mike Snitzer <snitzer@hammerspace.com>
To: Chuck Lever <cel@kernel.org>, Jeff Layton <jlayton@kernel.org>
Cc: linux-nfs@vger.kernel.org
Subject: [PATCH 09/10] NFSD: add direct_misaligned_dontcache debugfs knob
Date: Tue, 29 Sep 2026 13:34:22 -0400 [thread overview]
Message-ID: <20260929173423.16149-10-snitzer@kernel.org> (raw)
In-Reply-To: <20260929173423.16149-1-snitzer@kernel.org>
From: Jonathan Flynn <jonathan.flynn@hammerspace.com>
The parts of a direct-mode WRITE that cannot be direct I/O, the
misaligned prefix and suffix of a split WRITE and the whole WRITE when
it is not split, are issued IOCB_DONTCACHE so their pages are dropped
once written back. That is the right default for the workloads a direct
mode is chosen for, but it is a policy rather than a requirement: a
workload that reads back what it just wrote, or that keeps rewriting the
same partial pages, is better served by those pages staying in the page
cache.
Add a bool debugfs knob, /sys/kernel/debug/nfsd/direct_misaligned_dontcache
(default Y), to choose between the two:
Y: IOCB_DONTCACHE when the file system supports it, with a split's
boundary page kept in the page cache until both WRITEs sharing it
have written it (nfsd_write_dio_boundary_claim()).
N: ordinary cached buffered I/O; nothing is claimed or marked, and the
pages stay until reclaim.
The direct middle and the once-per-WRITE persist are unaffected. The
knob is sampled once per WRITE, so a change takes effect immediately and
without a remount. It sits beside io_cache_read and io_cache_write, the
other controls over how NFSD issues its I/O.
Signed-off-by: Jonathan Flynn <jonathan.flynn@hammerspace.com>
[snitzer: documented in nfsd-io-modes.rst]
[snitzer: switched from using modparam to debugfs knob]
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
.../filesystems/nfs/nfsd-io-modes.rst | 9 ++++++
fs/nfsd/debugfs.c | 21 ++++++++++++++
fs/nfsd/nfsd.h | 1 +
fs/nfsd/vfs.c | 28 ++++++++++++-------
4 files changed, 49 insertions(+), 10 deletions(-)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index 7263570668260..b2685dfadbffb 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -213,6 +213,15 @@ Misaligned WRITE:
with how far concurrent writers drift apart, not with bytes
written.
+ Whether those pages are dropped at all is a policy choice, selected
+ by /sys/kernel/debug/nfsd/direct_misaligned_dontcache (default Y).
+ Write N to issue the start and end segments, and the whole-WRITE
+ fallbacks, as ordinary cached buffered IO: nothing is claimed or
+ marked and the pages stay until reclaim, which suits a workload that
+ reads back or rewrites what it just wrote. The O_DIRECT middle
+ segment is unaffected. The knob is sampled once per WRITE, so a
+ change takes effect immediately.
+
The O_DIRECT middle segment also carries the DONTCACHE flag. It has
no effect while the IO really is O_DIRECT, but a filesystem may
decide on its own to service the segment with buffered IO instead
diff --git a/fs/nfsd/debugfs.c b/fs/nfsd/debugfs.c
index 603a608b03c54..1ae2597b9c34a 100644
--- a/fs/nfsd/debugfs.c
+++ b/fs/nfsd/debugfs.c
@@ -186,6 +186,24 @@ void nfsd_debugfs_exit(void)
* Default 2. Not yet tuned by benchmarking.
*/
+/*
+ * /sys/kernel/debug/nfsd/direct_misaligned_dontcache
+ *
+ * How a direct-mode WRITE issues the I/O that cannot be direct: the
+ * misaligned start and end of a split WRITE, and the whole WRITE when it
+ * is not split.
+ *
+ * Contents:
+ * Y: DONTCACHE when the filesystem supports it, with the boundary page
+ * of a split kept in the page cache only until both WRITEs sharing
+ * it have written it
+ * N: ordinary cached buffered IO, left in the page cache until
+ * reclaim, for A/B comparison against the DONTCACHE path
+ *
+ * Sampled once per WRITE, so it takes effect immediately. The direct
+ * middle segment is unaffected.
+ */
+
void nfsd_debugfs_init(void)
{
nfsd_top_dir = debugfs_create_dir("nfsd", NULL);
@@ -201,6 +219,9 @@ void nfsd_debugfs_init(void)
debugfs_create_u32("direct_misaligned_num_pages", 0644, nfsd_top_dir,
&nfsd_direct_misaligned_num_pages);
+
+ debugfs_create_bool("direct_misaligned_dontcache", 0644, nfsd_top_dir,
+ &nfsd_direct_misaligned_dontcache);
#ifdef CONFIG_NFSD_V4
debugfs_create_bool("delegated_timestamps", 0644, nfsd_top_dir,
&nfsd_delegts_enabled);
diff --git a/fs/nfsd/nfsd.h b/fs/nfsd/nfsd.h
index 135e319e378d4..dff979ac370ba 100644
--- a/fs/nfsd/nfsd.h
+++ b/fs/nfsd/nfsd.h
@@ -145,6 +145,7 @@ enum {
extern u64 nfsd_io_cache_read __read_mostly;
extern u64 nfsd_io_cache_write __read_mostly;
extern u32 nfsd_direct_misaligned_num_pages __read_mostly;
+extern bool nfsd_direct_misaligned_dontcache __read_mostly;
extern int nfsd_max_blksize;
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 5fd850a29694f..7896e2e6c5855 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -54,6 +54,7 @@ bool nfsd_disable_splice_read __read_mostly;
u64 nfsd_io_cache_read __read_mostly = NFSD_IO_BUFFERED;
u64 nfsd_io_cache_write __read_mostly = NFSD_IO_BUFFERED;
u32 nfsd_direct_misaligned_num_pages __read_mostly = 2;
+bool nfsd_direct_misaligned_dontcache __read_mostly = true;
/**
* nfserrno - Map Linux errnos to NFS errnos
@@ -1373,16 +1374,21 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
size_t prefix, middle, suffix;
loff_t offset = iocb->ki_pos;
unsigned int dontcache_flags = 0;
+ unsigned int buffered_flags;
unsigned int nsegs = 0;
if (nf->nf_file->f_op->fop_flags & FOP_DONTCACHE)
dontcache_flags = IOCB_DONTCACHE;
+ /* Buffered segments follow the knob; the direct middle does not. */
+ buffered_flags = READ_ONCE(nfsd_direct_misaligned_dontcache) ?
+ dontcache_flags : 0;
/*
* Whenever direct I/O cannot be used for the WRITE, fall back to a
- * single DONTCACHE buffered I/O when the file system supports it, so
- * the WRITE's pages are dropped from the page cache once written
- * back, and to a single cached buffered I/O otherwise.
+ * single DONTCACHE buffered I/O when the file system supports it (and
+ * nfsd_direct_misaligned_dontcache is set), so the WRITE's pages are
+ * dropped from the page cache once written back, and to a single
+ * cached buffered I/O otherwise.
*
* If the file system doesn't advertise any alignment requirements,
* don't try to issue direct I/O at all.
@@ -1421,13 +1427,15 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
* its page with the neighbouring WRITE; see
* nfsd_write_dio_boundary_claim(), which nfsd_direct_write() calls right
* before issuing each of them, for how the page is held for the
- * partner and dropped once both have written it.
+ * partner and dropped once both have written it. With
+ * nfsd_direct_misaligned_dontcache=N both are plain cached writes and
+ * nothing is held or dropped: the pages stay until reclaim.
*/
if (prefix) {
nfsd_write_dio_seg_init(&segments[nsegs], bvec,
nvecs, total, 0, prefix, iocb);
- segments[nsegs].flags |= dontcache_flags;
- segments[nsegs++].boundary = !!dontcache_flags;
+ segments[nsegs].flags |= buffered_flags;
+ segments[nsegs++].boundary = !!buffered_flags;
}
nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs,
@@ -1458,8 +1466,8 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
if (suffix) {
nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs, total,
prefix + middle, suffix, iocb);
- segments[nsegs].flags |= dontcache_flags;
- segments[nsegs++].boundary = !!dontcache_flags;
+ segments[nsegs].flags |= buffered_flags;
+ segments[nsegs++].boundary = !!buffered_flags;
}
return nsegs;
@@ -1473,8 +1481,8 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
*/
nfsd_write_dio_seg_init(&segments[0], bvec, nvecs, total, 0,
total, iocb);
- segments[0].flags |= dontcache_flags;
- segments[0].edges = !!dontcache_flags;
+ segments[0].flags |= buffered_flags;
+ segments[0].edges = !!buffered_flags;
return 1;
}
--
2.52.0
next prev parent reply other threads:[~2026-09-29 17:34 UTC|newest]
Thread overview: 17+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-29 17:34 [PATCH 00/10] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
2026-09-29 17:34 ` [PATCH 01/10] NFSD: interlock the use of NFSD_IO_DIRECT for NFS READ and WRITE Mike Snitzer
2026-09-29 18:27 ` Chuck Lever
2026-09-29 19:56 ` Mike Snitzer
2026-09-29 23:17 ` Chuck Lever
2026-09-29 23:30 ` Mike Snitzer
2026-09-30 0:20 ` Chuck Lever
2026-09-30 12:46 ` Mike Snitzer
2026-09-29 17:34 ` [PATCH 02/10] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well Mike Snitzer
2026-09-29 17:34 ` [PATCH 03/10] NFSD: only split a direct-mode WRITE for a worthwhile direct middle Mike Snitzer
2026-09-29 17:34 ` [PATCH 04/10] NFSD: do not use direct I/O for a READ smaller than its alignment Mike Snitzer
2026-09-29 17:34 ` [PATCH 05/10] NFSD: Enable return of an updated stable_how to NFS clients Mike Snitzer
2026-09-29 17:34 ` [PATCH 06/10] NFSD: let a direct-mode WRITE raise stable_how and elide the client's COMMIT Mike Snitzer
2026-09-29 17:34 ` [PATCH 07/10] NFSD: persist a synchronous direct-mode WRITE once, after all of its segments Mike Snitzer
2026-09-29 17:34 ` [PATCH 08/10] NFSD: keep boundary page of a split direct-mode WRITE until both writers complete Mike Snitzer
2026-09-29 17:34 ` Mike Snitzer [this message]
2026-09-29 17:34 ` [PATCH 10/10] NFSD: add tracing for how direct-mode READ and WRITE are serviced Mike Snitzer
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260929173423.16149-10-snitzer@kernel.org \
--to=snitzer@hammerspace.com \
--cc=cel@kernel.org \
--cc=jlayton@kernel.org \
--cc=linux-nfs@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox