From: Mike Snitzer <snitzer@hammerspace.com>
To: Chuck Lever <cel@kernel.org>, Jeff Layton <jlayton@kernel.org>
Cc: linux-nfs@vger.kernel.org
Subject: [PATCH 06/10] NFSD: let a direct-mode WRITE raise stable_how and elide the client's COMMIT
Date: Tue, 29 Sep 2026 13:34:19 -0400 [thread overview]
Message-ID: <20260929173423.16149-7-snitzer@kernel.org> (raw)
In-Reply-To: <20260929173423.16149-1-snitzer@kernel.org>
A WRITE serviced in a direct mode is already persistent when NFSD
replies: the aligned middle is O_DIRECT, and a synchronous WRITE fsyncs
whatever was buffered before the reply is sent. The reply still says
UNSTABLE, so the client dutifully sends a COMMIT for data that is
already on stable storage, and NFSD answers it with an fsync that has
nothing left to write.
Add two io_cache_write modes that issue direct I/O exactly like
NFSD_IO_DIRECT and raise the stable_how of the reply to a floor:
NFSD_IO_DIRECT_WRITE_DATA_SYNC (3): at least NFS_DATA_SYNC
NFSD_IO_DIRECT_WRITE_FILE_SYNC (4): at least NFS_FILE_SYNC
A client that asked for more is left alone, and the WRITE is persisted
to the level the reply reports before that reply is sent. A client told
FILE_SYNC has no reason to COMMIT and sends none.
What this removes is the COMMIT traffic, and it is worth most where a
client's writes are carved into several WRITE RPCs, because each piece
is then committed separately. Measured on a pNFS flexfiles share where
every write straddles two data servers, so every write becomes two WRITE
RPCs and, under NFSD_IO_DIRECT, two COMMITs: the two modes do identical
durability work, one fsync per COMMIT against one fsync per WRITE on
identical WRITE counts, and the COMMIT RPCs alone cost NFSD_IO_DIRECT
31% more server CPU and 42% more client CPU for the same bytes, about
20 us of server CPU per COMMIT plus a client cost that grows with the
range committed. Where only one write in 22 is split the same effect
is a couple of cores on each side and no resolvable throughput
difference, and writes that fit a single RPC send no COMMIT in either
mode.
Choose 4 when the export services WRITEs in a direct mode and clients
split their writes; choose 3 to promise only the data, leaving a client
that needs metadata durability to COMMIT for it. Both are inert for
READ, and for a WRITE the client already marked FILE_SYNC.
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
.../filesystems/nfs/nfsd-io-modes.rst | 17 ++++++++--
fs/nfsd/debugfs.c | 5 +++
fs/nfsd/nfsd.h | 2 ++
fs/nfsd/vfs.c | 33 +++++++++++++++++--
4 files changed, 52 insertions(+), 5 deletions(-)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index bc1c0f1a7b7ca..548eca7cdb518 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -25,12 +25,14 @@ Based on the configured settings, NFSD's IO will either be:
- cached using page cache (NFSD_IO_BUFFERED=0)
- cached but removed from page cache on completion (NFSD_IO_DONTCACHE=1)
- not cached stable_how=NFS_UNSTABLE (NFSD_IO_DIRECT=2)
+- not cached stable_how=NFS_DATA_SYNC (NFSD_IO_DIRECT_WRITE_DATA_SYNC=3)
+- not cached stable_how=NFS_FILE_SYNC (NFSD_IO_DIRECT_WRITE_FILE_SYNC=4)
-To set an NFSD IO mode, write a supported value (0 - 2) to the
+To set an NFSD IO mode, write a supported value (0 - 4) to the
corresponding IO operation's debugfs interface, e.g.::
echo 2 > /sys/kernel/debug/nfsd/io_cache_read
- echo 2 > /sys/kernel/debug/nfsd/io_cache_write
+ echo 4 > /sys/kernel/debug/nfsd/io_cache_write
To check which IO mode NFSD is using for READ or WRITE, simply read the
corresponding IO operation's debugfs interface, e.g.::
@@ -38,6 +40,17 @@ corresponding IO operation's debugfs interface, e.g.::
cat /sys/kernel/debug/nfsd/io_cache_read
cat /sys/kernel/debug/nfsd/io_cache_write
+The two NFSD_IO_DIRECT_WRITE_*_SYNC modes raise the stable_how of every
+WRITE to at least NFS_DATA_SYNC or NFS_FILE_SYNC, persist the WRITE
+accordingly before replying, and return the raised value to the client;
+a client that asked for a higher stable_how is left alone. With
+NFSD_IO_DIRECT_WRITE_FILE_SYNC the client sends no COMMIT. Against
+NFSD_IO_DIRECT the durability work is the same, one fsync per WRITE
+instead of one per COMMIT; what NFSD_IO_DIRECT adds is the COMMIT RPCs
+themselves, tens of microseconds of server CPU each plus a client cost
+that grows with the range committed, which matters in proportion to how
+many of a client's WRITEs need a COMMIT.
+
If you experiment with NFSD's IO modes on a recent kernel and have
interesting results, please report them to linux-nfs@vger.kernel.org
diff --git a/fs/nfsd/debugfs.c b/fs/nfsd/debugfs.c
index afa41bf2d8189..603a608b03c54 100644
--- a/fs/nfsd/debugfs.c
+++ b/fs/nfsd/debugfs.c
@@ -113,6 +113,9 @@ DEFINE_DEBUGFS_ATTRIBUTE(nfsd_io_cache_read_fops, nfsd_io_cache_read_get,
* Contents:
* %0: NFS WRITE will use buffered IO
* %1: NFS WRITE will use dontcache (buffered IO w/ dropbehind)
+ * %2: NFS WRITE will use direct IO with stable_how=NFS_UNSTABLE
+ * %3: NFS WRITE will use direct IO with stable_how=NFS_DATA_SYNC
+ * %4: NFS WRITE will use direct IO with stable_how=NFS_FILE_SYNC
*
* This setting takes immediate effect for all NFS versions,
* all exports, and in all NFSD net namespaces.
@@ -132,6 +135,8 @@ static int nfsd_io_cache_write_set(void *data, u64 val)
case NFSD_IO_BUFFERED:
case NFSD_IO_DONTCACHE:
case NFSD_IO_DIRECT:
+ case NFSD_IO_DIRECT_WRITE_DATA_SYNC:
+ case NFSD_IO_DIRECT_WRITE_FILE_SYNC:
nfsd_io_cache_write = val;
/*
* Adjust nfsd_io_cache_{read,write} to avoid
diff --git a/fs/nfsd/nfsd.h b/fs/nfsd/nfsd.h
index a2d72434160af..135e319e378d4 100644
--- a/fs/nfsd/nfsd.h
+++ b/fs/nfsd/nfsd.h
@@ -138,6 +138,8 @@ enum {
NFSD_IO_BUFFERED,
NFSD_IO_DONTCACHE,
NFSD_IO_DIRECT,
+ NFSD_IO_DIRECT_WRITE_DATA_SYNC,
+ NFSD_IO_DIRECT_WRITE_FILE_SYNC,
};
extern u64 nfsd_io_cache_read __read_mostly;
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 82eba97656c4e..924c5992dc32e 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1399,17 +1399,42 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
return 1;
}
+/*
+ * Raise the stability of this WRITE to at least @floor_iocb_flags, and
+ * record what was achieved in @iocb_flags so the reply can report it.
+ * A client that asked for more is left alone.
+ */
+static void
+nfsd_write_raise_stability(int floor_iocb_flags, struct kiocb *kiocb,
+ int *iocb_flags)
+{
+ if ((*iocb_flags & floor_iocb_flags) == floor_iocb_flags)
+ return; /* already at or above the floor */
+
+ *iocb_flags |= floor_iocb_flags;
+ kiocb->ki_flags |= floor_iocb_flags;
+}
+
static noinline_for_stack int
nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
- struct nfsd_file *nf, unsigned int nvecs,
+ struct nfsd_file *nf, int *iocb_flags, unsigned int nvecs,
unsigned long *cnt, struct kiocb *kiocb)
{
struct nfsd_write_dio_seg segments[3];
+ int floor_iocb_flags = 0;
struct file *file = nf->nf_file;
unsigned int nsegs, i;
ssize_t host_err;
size_t expected;
+ if (nfsd_io_cache_write == NFSD_IO_DIRECT_WRITE_FILE_SYNC)
+ floor_iocb_flags = IOCB_DSYNC | IOCB_SYNC;
+ else if (nfsd_io_cache_write == NFSD_IO_DIRECT_WRITE_DATA_SYNC)
+ floor_iocb_flags = IOCB_DSYNC;
+ if (floor_iocb_flags)
+ nfsd_write_raise_stability(floor_iocb_flags, kiocb,
+ iocb_flags);
+
nsegs = nfsd_write_dio_iters_init(nf, rqstp->rq_bvec, nvecs,
kiocb, *cnt, segments);
@@ -1513,8 +1538,10 @@ nfsd_vfs_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
switch (nfsd_io_cache_write) {
case NFSD_IO_DIRECT:
- host_err = nfsd_direct_write(rqstp, fhp, nf, nvecs,
- cnt, &kiocb);
+ case NFSD_IO_DIRECT_WRITE_DATA_SYNC:
+ case NFSD_IO_DIRECT_WRITE_FILE_SYNC:
+ host_err = nfsd_direct_write(rqstp, fhp, nf, iocb_flags,
+ nvecs, cnt, &kiocb);
break;
case NFSD_IO_DONTCACHE:
if (file->f_op->fop_flags & FOP_DONTCACHE)
--
2.52.0
next prev parent reply other threads:[~2026-09-29 17:34 UTC|newest]
Thread overview: 17+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-29 17:34 [PATCH 00/10] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
2026-09-29 17:34 ` [PATCH 01/10] NFSD: interlock the use of NFSD_IO_DIRECT for NFS READ and WRITE Mike Snitzer
2026-09-29 18:27 ` Chuck Lever
2026-09-29 19:56 ` Mike Snitzer
2026-09-29 23:17 ` Chuck Lever
2026-09-29 23:30 ` Mike Snitzer
2026-09-30 0:20 ` Chuck Lever
2026-09-30 12:46 ` Mike Snitzer
2026-09-29 17:34 ` [PATCH 02/10] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well Mike Snitzer
2026-09-29 17:34 ` [PATCH 03/10] NFSD: only split a direct-mode WRITE for a worthwhile direct middle Mike Snitzer
2026-09-29 17:34 ` [PATCH 04/10] NFSD: do not use direct I/O for a READ smaller than its alignment Mike Snitzer
2026-09-29 17:34 ` [PATCH 05/10] NFSD: Enable return of an updated stable_how to NFS clients Mike Snitzer
2026-09-29 17:34 ` Mike Snitzer [this message]
2026-09-29 17:34 ` [PATCH 07/10] NFSD: persist a synchronous direct-mode WRITE once, after all of its segments Mike Snitzer
2026-09-29 17:34 ` [PATCH 08/10] NFSD: keep boundary page of a split direct-mode WRITE until both writers complete Mike Snitzer
2026-09-29 17:34 ` [PATCH 09/10] NFSD: add direct_misaligned_dontcache debugfs knob Mike Snitzer
2026-09-29 17:34 ` [PATCH 10/10] NFSD: add tracing for how direct-mode READ and WRITE are serviced Mike Snitzer
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260929173423.16149-7-snitzer@kernel.org \
--to=snitzer@hammerspace.com \
--cc=cel@kernel.org \
--cc=jlayton@kernel.org \
--cc=linux-nfs@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox