From: Mike Snitzer <snitzer@kernel.org>
To: Chuck Lever <cel@kernel.org>, Jeff Layton <jlayton@kernel.org>
Cc: hch@lst.de, linux-nfs@vger.kernel.org
Subject: [PATCH v3 5/9] NFSD: persist a synchronous direct-mode WRITE once, after all of its segments
Date: Thu, 1 Oct 2026 00:54:58 -0400 [thread overview]
Message-ID: <20261001045502.48381-6-snitzer@kernel.org> (raw)
In-Reply-To: <20261001045502.48381-1-snitzer@kernel.org>
nfsd_direct_write() may issue a WRITE as up to three segments: a
buffered prefix, a direct middle and a buffered suffix. For a
FILE_SYNC or DATA_SYNC WRITE the kiocb carries IOCB_DSYNC and every
segment inherits it, so generic_write_sync() runs a range
fsync after each segment: up to three cache flushes and log forces per
WRITE, and each one writes back and drops the boundary page it just
touched.
Strip IOCB_DSYNC and IOCB_SYNC from the per-segment flags and persist
the WRITE once with vfs_fsync_range() over the bytes actually written,
after the last segment. Durability is unchanged: the reply is not sent
until the fsync completes, and datasync mirrors the previous per-segment
choice (IOCB_SYNC present means metadata too). An fsync failure is
returned like a write failure.
Besides the fewer flushes, this puts the sync under NFSD's control,
which a later commit uses to keep the boundary pages of a split WRITE
cached until the partner WRITE completes them.
Assisted-by: Claude:claude-fable-5-1
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
Documentation/filesystems/nfs/nfsd-io-modes.rst | 3 +++
fs/nfsd/vfs.c | 15 ++++++++++++++-
2 files changed, 17 insertions(+), 1 deletion(-)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index 12679001c7bec..686b106b96f1b 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -152,6 +152,9 @@ Misaligned WRITE:
- the WRITE payload is not aligned in memory to the block device's
dma_alignment, so the middle cannot be O_DIRECT either.
+ A FILE_SYNC or DATA_SYNC WRITE is persisted once after all of its
+ segments are written, not once per segment.
+
Writing N to /sys/kernel/debug/nfsd/direct_misaligned_dontcache
(default Y) issues the start and end segments, and a WRITE that is
not split, as normal buffered IO instead of DONTCACHE, which suits a
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index e1d294aceb6bd..592415900403f 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1383,16 +1383,22 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
{
struct nfsd_write_dio_seg segments[3];
struct file *file = nf->nf_file;
+ loff_t start = kiocb->ki_pos;
+ bool sync, datasync;
unsigned int nsegs, i;
ssize_t host_err;
size_t expected;
+ /* Persist a synchronous WRITE once, after all of its segments. */
+ sync = kiocb->ki_flags & IOCB_DSYNC;
+ datasync = !(kiocb->ki_flags & IOCB_SYNC);
+
nsegs = nfsd_write_dio_iters_init(nf, rqstp->rq_bvec, nvecs,
kiocb, *cnt, segments);
*cnt = 0;
for (i = 0; i < nsegs; i++) {
- kiocb->ki_flags = segments[i].flags;
+ kiocb->ki_flags = segments[i].flags & ~(IOCB_DSYNC | IOCB_SYNC);
if (kiocb->ki_flags & IOCB_DIRECT)
trace_nfsd_write_direct(rqstp, fhp, kiocb->ki_pos,
segments[i].iter.count);
@@ -1410,6 +1416,13 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
break; /* partial write */
}
+ if (sync && *cnt) {
+ host_err = vfs_fsync_range(file, start, start + *cnt - 1,
+ datasync);
+ if (host_err < 0)
+ return host_err;
+ }
+
return 0;
}
--
2.52.0
next prev parent reply other threads:[~2026-10-01 4:55 UTC|newest]
Thread overview: 11+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-01 4:54 [PATCH v3 0/9] NFSD: keep direct-mode I/O out of the page cache Mike Snitzer
2026-10-01 4:54 ` [PATCH v3 1/9] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well Mike Snitzer
2026-10-01 4:54 ` [PATCH v3 2/9] NFSD: only split a direct-mode WRITE for a worthwhile direct middle Mike Snitzer
2026-10-01 4:54 ` [PATCH v3 3/9] NFSD: add direct_misaligned_dontcache debugfs knob Mike Snitzer
2026-10-01 4:54 ` [PATCH v3 4/9] NFSD: do not use direct I/O for a READ smaller than its alignment Mike Snitzer
2026-10-01 4:54 ` Mike Snitzer [this message]
2026-10-01 4:54 ` [PATCH v3 6/9] NFSD: keep boundary page of a split direct-mode WRITE until both writers complete Mike Snitzer
2026-10-01 4:55 ` [PATCH v3 7/9] NFSD: add tracing for how direct-mode READ and WRITE are serviced Mike Snitzer
2026-10-01 4:55 ` [PATCH v3 8/9] NFSD: Enable return of an updated stable_how to NFS clients Mike Snitzer
2026-10-01 4:55 ` [PATCH v3 9/9] NFSD: add direct-mode WRITE settings that persist each WRITE Mike Snitzer
2026-10-01 22:58 ` [PATCH v3 0/9] NFSD: keep direct-mode I/O out of the page cache Mike Snitzer
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261001045502.48381-6-snitzer@kernel.org \
--to=snitzer@kernel.org \
--cc=cel@kernel.org \
--cc=hch@lst.de \
--cc=jlayton@kernel.org \
--cc=linux-nfs@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox