From: Mike Snitzer <snitzer@kernel.org>
To: Chuck Lever <cel@kernel.org>, Jeff Layton <jlayton@kernel.org>
Cc: hch@lst.de, linux-nfs@vger.kernel.org
Subject: [PATCH v3 9/9] NFSD: add direct-mode WRITE settings that persist each WRITE
Date: Thu, 1 Oct 2026 00:55:02 -0400 [thread overview]
Message-ID: <20261001045502.48381-10-snitzer@kernel.org> (raw)
In-Reply-To: <20261001045502.48381-1-snitzer@kernel.org>
Under NFSD_IO_DIRECT an UNSTABLE WRITE is issued O_DIRECT without
IOCB_DSYNC and answered UNSTABLE. The data is not yet durable: the
device cache is not flushed, and block allocation, unwritten extent
conversion and the size update are not committed. The client's COMMIT
is what makes it so. This patch does not change NFSD_IO_DIRECT.
Add two io_cache_write modes that issue direct I/O like NFSD_IO_DIRECT
but persist every WRITE before replying, to at least a floor, and
report the stability achieved:
NFSD_IO_DIRECT_WRITE_DATA_SYNC (3): at least NFS_DATA_SYNC
NFSD_IO_DIRECT_WRITE_FILE_SYNC (4): at least NFS_FILE_SYNC
A client that asked for more is left alone. A client told FILE_SYNC
sends no COMMIT.
The cost is a flush and log force on every WRITE, where NFSD_IO_DIRECT
has one per COMMIT, so these are opt-in. They cost more than
NFSD_IO_DIRECT when one COMMIT covers many WRITEs. They cost the same
flushes, and save the COMMIT RPCs, when each COMMIT covers about one
WRITE. That is the case on a pNFS flexfiles share where a client's
writes straddle two data servers, so that each data server receives
one WRITE and one COMMIT per write. Measured there, with identical
WRITE counts and flush counts, the COMMIT RPCs cost NFSD_IO_DIRECT 31%
more server CPU and 42% more client CPU for the same bytes. Where only
one write in 22 was split, there was no resolvable difference.
A Linux client change that sends FILE_SYNC for a WRITE that is alone
in its flush to a given server would remove those COMMITs without a
server setting, for clients that have it.
Mode 3 promises only the data, leaving a client that needs metadata
durability to COMMIT for it. Both modes are inert for READ and for a
WRITE the client already marked FILE_SYNC.
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
.../filesystems/nfs/nfsd-io-modes.rst | 15 ++++++--
fs/nfsd/debugfs.c | 5 +++
fs/nfsd/nfsd.h | 2 ++
fs/nfsd/vfs.c | 35 ++++++++++++++++---
4 files changed, 50 insertions(+), 7 deletions(-)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index 11938caee14ca..87266e43e4d1c 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -25,12 +25,14 @@ Based on the configured settings, NFSD's IO will either be:
- cached using page cache (NFSD_IO_BUFFERED=0)
- cached but removed from page cache on completion (NFSD_IO_DONTCACHE=1)
- not cached stable_how=NFS_UNSTABLE (NFSD_IO_DIRECT=2)
+- not cached stable_how=NFS_DATA_SYNC (NFSD_IO_DIRECT_WRITE_DATA_SYNC=3)
+- not cached stable_how=NFS_FILE_SYNC (NFSD_IO_DIRECT_WRITE_FILE_SYNC=4)
-To set an NFSD IO mode, write a supported value (0 - 2) to the
+To set an NFSD IO mode, write a supported value (0 - 4) to the
corresponding IO operation's debugfs interface, e.g.::
echo 2 > /sys/kernel/debug/nfsd/io_cache_read
- echo 2 > /sys/kernel/debug/nfsd/io_cache_write
+ echo 4 > /sys/kernel/debug/nfsd/io_cache_write
To check which IO mode NFSD is using for READ or WRITE, simply read the
corresponding IO operation's debugfs interface, e.g.::
@@ -38,6 +40,15 @@ corresponding IO operation's debugfs interface, e.g.::
cat /sys/kernel/debug/nfsd/io_cache_read
cat /sys/kernel/debug/nfsd/io_cache_write
+NFSD_IO_DIRECT leaves an UNSTABLE WRITE for the client's COMMIT to
+persist. The two NFSD_IO_DIRECT_WRITE_*_SYNC modes instead persist every
+WRITE before replying, to at least NFS_DATA_SYNC or NFS_FILE_SYNC, and
+report that stable_how to the client; a client that asked for more is
+left alone. With NFSD_IO_DIRECT_WRITE_FILE_SYNC the client sends no
+COMMIT. This costs a flush on every WRITE: more than NFSD_IO_DIRECT
+when one COMMIT covers many WRITEs, about the same, without the COMMIT
+RPCs, when each COMMIT covers about one.
+
If you experiment with NFSD's IO modes on a recent kernel and have
interesting results, please report them to linux-nfs@vger.kernel.org
diff --git a/fs/nfsd/debugfs.c b/fs/nfsd/debugfs.c
index d5b714dd0fbc6..bd8ceb5eb2879 100644
--- a/fs/nfsd/debugfs.c
+++ b/fs/nfsd/debugfs.c
@@ -90,6 +90,9 @@ DEFINE_DEBUGFS_ATTRIBUTE(nfsd_io_cache_read_fops, nfsd_io_cache_read_get,
* Contents:
* %0: NFS WRITE will use buffered IO
* %1: NFS WRITE will use dontcache (buffered IO w/ dropbehind)
+ * %2: NFS WRITE will use direct IO with stable_how=NFS_UNSTABLE
+ * %3: NFS WRITE will use direct IO with stable_how=NFS_DATA_SYNC
+ * %4: NFS WRITE will use direct IO with stable_how=NFS_FILE_SYNC
*
* This setting takes immediate effect for all NFS versions,
* all exports, and in all NFSD net namespaces.
@@ -109,6 +112,8 @@ static int nfsd_io_cache_write_set(void *data, u64 val)
case NFSD_IO_BUFFERED:
case NFSD_IO_DONTCACHE:
case NFSD_IO_DIRECT:
+ case NFSD_IO_DIRECT_WRITE_DATA_SYNC:
+ case NFSD_IO_DIRECT_WRITE_FILE_SYNC:
nfsd_io_cache_write = val;
break;
default:
diff --git a/fs/nfsd/nfsd.h b/fs/nfsd/nfsd.h
index 1c17f96012523..913860ab3eb20 100644
--- a/fs/nfsd/nfsd.h
+++ b/fs/nfsd/nfsd.h
@@ -141,6 +141,8 @@ enum {
NFSD_IO_BUFFERED,
NFSD_IO_DONTCACHE,
NFSD_IO_DIRECT,
+ NFSD_IO_DIRECT_WRITE_DATA_SYNC,
+ NFSD_IO_DIRECT_WRITE_FILE_SYNC,
};
extern u64 nfsd_io_cache_read __read_mostly;
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index a8e15193b2160..af05ef942c5d9 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1450,12 +1450,25 @@ nfsd_write_dio_iters_init(struct svc_rqst *rqstp, struct svc_fh *fhp,
return nsegs;
}
+/* Raise this WRITE to at least @floor_iocb_flags, and report that. */
+static void
+nfsd_write_raise_stability(int floor_iocb_flags, struct kiocb *kiocb,
+ int *iocb_flags)
+{
+ if ((*iocb_flags & floor_iocb_flags) == floor_iocb_flags)
+ return;
+
+ *iocb_flags |= floor_iocb_flags;
+ kiocb->ki_flags |= floor_iocb_flags;
+}
+
static noinline_for_stack int
nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
- struct nfsd_file *nf, unsigned int nvecs,
+ struct nfsd_file *nf, int *iocb_flags, unsigned int nvecs,
unsigned long *cnt, struct kiocb *kiocb)
{
struct nfsd_write_dio_seg segments[3];
+ int floor_iocb_flags = 0;
struct file *file = nf->nf_file;
loff_t start = kiocb->ki_pos, seg_pos, seg_last;
bool sync, datasync, complete_first, complete_last;
@@ -1463,6 +1476,14 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
ssize_t host_err;
size_t expected;
+ if (nfsd_io_cache_write == NFSD_IO_DIRECT_WRITE_FILE_SYNC)
+ floor_iocb_flags = IOCB_DSYNC | IOCB_SYNC;
+ else if (nfsd_io_cache_write == NFSD_IO_DIRECT_WRITE_DATA_SYNC)
+ floor_iocb_flags = IOCB_DSYNC;
+ if (floor_iocb_flags)
+ nfsd_write_raise_stability(floor_iocb_flags, kiocb,
+ iocb_flags);
+
/* Persist a synchronous WRITE once, after all of its segments. */
sync = kiocb->ki_flags & IOCB_DSYNC;
datasync = !(kiocb->ki_flags & IOCB_SYNC);
@@ -1538,7 +1559,8 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
* @payload: xdr_buf containing the write payload
* @cnt: IN: number of bytes to write, OUT: number of bytes actually written
* @iocb_flags: IN: VFS IOCB_* flags expressing the requested write
- * stability; OUT: the stability to report to the client
+ * stability; OUT: the stability to report to the client,
+ * which may be raised
* @verf: NFS WRITE verifier
*
* Upon return, caller must invoke fh_put on @fhp.
@@ -1606,8 +1628,10 @@ nfsd_vfs_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
switch (nfsd_io_cache_write) {
case NFSD_IO_DIRECT:
- host_err = nfsd_direct_write(rqstp, fhp, nf, nvecs,
- cnt, &kiocb);
+ case NFSD_IO_DIRECT_WRITE_DATA_SYNC:
+ case NFSD_IO_DIRECT_WRITE_FILE_SYNC:
+ host_err = nfsd_direct_write(rqstp, fhp, nf, iocb_flags,
+ nvecs, cnt, &kiocb);
break;
case NFSD_IO_DONTCACHE:
if (file->f_op->fop_flags & FOP_DONTCACHE)
@@ -1725,7 +1749,8 @@ __be32 nfsd_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
* @payload: xdr_buf containing the write payload
* @cnt: IN: number of bytes to write, OUT: number of bytes actually written
* @iocb_flags: IN: VFS IOCB_* flags expressing the requested write
- * stability; OUT: the stability to report to the client
+ * stability; OUT: the stability to report to the client,
+ * which may be raised
* @verf: NFS WRITE verifier
*
* Upon return, caller must invoke fh_put on @fhp.
--
2.52.0
next prev parent reply other threads:[~2026-10-01 4:55 UTC|newest]
Thread overview: 11+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-01 4:54 [PATCH v3 0/9] NFSD: keep direct-mode I/O out of the page cache Mike Snitzer
2026-10-01 4:54 ` [PATCH v3 1/9] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well Mike Snitzer
2026-10-01 4:54 ` [PATCH v3 2/9] NFSD: only split a direct-mode WRITE for a worthwhile direct middle Mike Snitzer
2026-10-01 4:54 ` [PATCH v3 3/9] NFSD: add direct_misaligned_dontcache debugfs knob Mike Snitzer
2026-10-01 4:54 ` [PATCH v3 4/9] NFSD: do not use direct I/O for a READ smaller than its alignment Mike Snitzer
2026-10-01 4:54 ` [PATCH v3 5/9] NFSD: persist a synchronous direct-mode WRITE once, after all of its segments Mike Snitzer
2026-10-01 4:54 ` [PATCH v3 6/9] NFSD: keep boundary page of a split direct-mode WRITE until both writers complete Mike Snitzer
2026-10-01 4:55 ` [PATCH v3 7/9] NFSD: add tracing for how direct-mode READ and WRITE are serviced Mike Snitzer
2026-10-01 4:55 ` [PATCH v3 8/9] NFSD: Enable return of an updated stable_how to NFS clients Mike Snitzer
2026-10-01 4:55 ` Mike Snitzer [this message]
2026-10-01 22:58 ` [PATCH v3 0/9] NFSD: keep direct-mode I/O out of the page cache Mike Snitzer
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261001045502.48381-10-snitzer@kernel.org \
--to=snitzer@kernel.org \
--cc=cel@kernel.org \
--cc=hch@lst.de \
--cc=jlayton@kernel.org \
--cc=linux-nfs@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox