* [PATCH v2 0/9] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs
@ 2026-09-29 23:13 Mike Snitzer
2026-09-29 23:13 ` [PATCH v2 1/9] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well Mike Snitzer
` (8 more replies)
0 siblings, 9 replies; 14+ messages in thread
From: Mike Snitzer @ 2026-09-29 23:13 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: linux-nfs
Hi,
What this advance is, precisely. It is not "NFSD uses DONTCACHE". It
is NFSD's direct write path - io_cache_write modes 2, 3 and 4, where
an aligned WRITE goes to the filesystem as O_DIRECT - with DONTCACHE
used only for the buffered fragments a misaligned WRITE cannot issue
directly. The page cache is bypassed for the bulk of the data and
bounded for the remainder. That distinction runs through everything
below, and it is exactly what separates this from io_cache_write=1,
which is plain buffered DONTCACHE with no direct I/O at all and does
not reach this code.
nfsd-next already issues what a direct-mode READ or WRITE cannot do
directly as DONTCACHE. Within that path, this series changes when
direct I/O is attempted at all, closes the one case where a direct
middle could still land in the page cache, lets a WRITE report the
stability it actually has so the client skips its COMMIT, keeps the
page two misaligned WRITEs share until both have written it, and adds
tracepoints that show which path each request took.
Patches 1-3 cover what cannot or should not be direct: the direct
middle of a split WRITE is also marked IOCB_DONTCACHE, for when XFS
falls back to buffered I/O (-ENOTBLK); a WRITE is split only when that
buys a worthwhile direct middle (direct_misaligned_num_pages, default
2); and, on the READ side, a READ smaller than its alignment no longer
costs a full aligned device read.
Patch 4 is Chuck's "Enable return of an updated stable_how to NFS
clients", reworked onto the @iocb_flags argument nfsd_write() now takes
in nfsd-next. The Reviewed-by tags from its first posting are dropped
because the argument changed.
Patch 5 adds io_cache_write modes 3 and 4, which issue direct I/O like
NFSD_IO_DIRECT and raise the reply's stable_how to at least DATA_SYNC
or FILE_SYNC. On a pNFS flexfiles share where every write is split
across two data servers, the COMMITs alone cost NFSD_IO_DIRECT 31% more
server CPU and 42% more client CPU for the same bytes.
Patch 6 persists a synchronous direct-mode WRITE once, after all of its
segments, instead of up to three fsyncs per WRITE.
Patch 7 keeps the page that two misaligned WRITEs share in the page
cache until both have written it, so the second one no longer has to
read it back from disk: with 32 interleaved writers, 704 device reads
for 42895 WRITEs where there were 45144 for 45664.
Patch 8 (Jonathan) adds the direct_misaligned_dontcache debugfs knob,
default Y; set to N, the parts of a direct-mode WRITE that cannot be
direct use cached buffered I/O instead of DONTCACHE.
Patch 9 adds tracepoints for how each direct-mode READ and WRITE was
serviced, including why a WRITE was not direct.
Documentation/filesystems/nfs/nfsd-io-modes.rst is updated throughout.
The series applies to cel/nfsd-next (ac04dab23b5f) and each patch
builds cleanly with W=1.
Changes since v1:
- Dropped v1's patch 1, "NFSD: interlock the use of NFSD_IO_DIRECT for
NFS READ and WRITE". io_cache_read and io_cache_write stay
independent: tying them together would rule out defaults that differ
by direction, such as direct WRITE with buffered READ.
- Rewrote the introduction above to say precisely which path this
series changes. v1's said the series makes direct-mode I/O fall back
to DONTCACHE; nfsd-next already does that.
- No change to the remaining nine patches; patch 5 now applies without
the interlock beneath it, which changes only its context.
v1: https://lore.kernel.org/linux-nfs/20260929173423.16149-1-snitzer@kernel.org/
All review appreciated, thanks.
Mike
Chuck Lever (1):
NFSD: Enable return of an updated stable_how to NFS clients
Jonathan Flynn (1):
NFSD: add direct_misaligned_dontcache debugfs knob
Mike Snitzer (7):
NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well
NFSD: only split a direct-mode WRITE for a worthwhile direct middle
NFSD: do not use direct I/O for a READ smaller than its alignment
NFSD: let a direct-mode WRITE raise stable_how and elide the client's COMMIT
NFSD: persist a synchronous direct-mode WRITE once, after all of its segments
NFSD: keep boundary page of a split direct-mode WRITE until both writers complete
NFSD: add tracing for how direct-mode READ and WRITE are serviced
.../filesystems/nfs/nfsd-io-modes.rst | 129 ++++++-
fs/nfsd/debugfs.c | 40 +++
fs/nfsd/nfs3proc.c | 16 +-
fs/nfsd/nfs4proc.c | 15 +-
fs/nfsd/nfsd.h | 4 +
fs/nfsd/nfsproc.c | 3 +-
fs/nfsd/trace.h | 90 +++++
fs/nfsd/vfs.c | 318 +++++++++++++++---
fs/nfsd/vfs.h | 26 +-
fs/nfsd/xdr3.h | 2 +-
10 files changed, 582 insertions(+), 61 deletions(-)
--
2.52.0
^ permalink raw reply [flat|nested] 14+ messages in thread
* [PATCH v2 1/9] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well
2026-09-29 23:13 [PATCH v2 0/9] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
@ 2026-09-29 23:13 ` Mike Snitzer
2026-09-30 21:43 ` Chuck Lever
2026-09-29 23:13 ` [PATCH v2 2/9] NFSD: only split a direct-mode WRITE for a worthwhile direct middle Mike Snitzer
` (7 subsequent siblings)
8 siblings, 1 reply; 14+ messages in thread
From: Mike Snitzer @ 2026-09-29 23:13 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: linux-nfs
nfsd_write_dio_iters_init() issues the DIO-aligned middle of a
misaligned WRITE with IOCB_DIRECT. A file system is free to service
that request with buffered I/O instead: xfs_file_write_iter() retries
via xfs_file_buffered_write() with the same kiocb whenever
xfs_file_dio_write() returns -ENOTBLK, which iomap_dio_rw() produces
when the invalidation it runs before issuing the direct I/O cannot drop
page cache overlapping the range.
That happens routinely in a streaming misaligned WRITE workload with
more than one request in flight. Request N's direct middle ends inside
the same page that request N+1's buffered prefix starts in. If N+1's
prefix dirties that page between N's filemap_write_and_wait_range() and
N's invalidate_inode_pages2_range(), the invalidation returns -EBUSY,
iomap returns -ENOTBLK, and XFS re-issues N's entire middle, tens of
kilobytes, as a normal cached buffered write. Nothing ever drops those
pages. On a device with dma_alignment=3 (XFS then reports
dio_mem_align=4, which XDR-aligned RPC payloads always satisfy) this is
the only remaining path by which a stream of small misaligned WRITEs in
DIRECT mode fills the page cache.
Set IOCB_DONTCACHE on the middle segment alongside IOCB_DIRECT when the
file system supports FOP_DONTCACHE. The direct path ignores the flag,
and generic_write_sync() only issues a harmless flusher kick; but if
the file system falls back to buffered I/O the write is now DONTCACHE
rather than cached. The nfsd_write_direct trace event is unchanged
since it keys on IOCB_DIRECT; the fallback itself remains observable
via the iomap_dio_invalidate_fail event.
Fixes: 06c5c97293e3 ("NFSD: Implement NFSD_IO_DIRECT for NFS WRITE")
Assisted-by: Claude:claude-fable-5-1
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
Documentation/filesystems/nfs/nfsd-io-modes.rst | 10 ++++++++++
fs/nfsd/vfs.c | 9 +++++++++
2 files changed, 19 insertions(+)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index 0fd6e82478fe6..dc50c930f9762 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -130,6 +130,16 @@ Misaligned WRITE:
segments because using normal buffered IO offers significant RMW
performance benefit when handling streaming misaligned WRITEs.
+ The O_DIRECT middle segment also carries the DONTCACHE flag. It has
+ no effect while the IO really is O_DIRECT, but a filesystem may
+ decide on its own to service the segment with buffered IO instead
+ (XFS does so when it cannot invalidate page cache that overlaps the
+ segment, which can happen when another WRITE's buffered start or
+ end segment dirties the shared boundary page at the same time).
+ The flag makes that fallback DONTCACHE buffered IO rather than
+ normal buffered IO. Such fallbacks are visible through the
+ iomap_dio_invalidate_fail trace event; see Tracing below.
+
Tracing:
The nfsd_read_direct trace event shows how NFSD expands any
misaligned READ to the next DIO-aligned block (on either end of the
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index f9131827d391e..5b963f0e3b3f2 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1341,6 +1341,15 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
if (iov_iter_bvec_offset(&segments[nsegs].iter) & (mem_align - 1))
goto no_dio;
segments[nsegs].flags |= IOCB_DIRECT;
+ /*
+ * Also mark the direct middle DONTCACHE: the file system may fall
+ * back to buffered I/O on its own (e.g. XFS on -ENOTBLK when it
+ * cannot invalidate page cache that a concurrent buffered prefix or
+ * suffix of an adjacent WRITE just dirtied), and it reuses this kiocb
+ * to do so. On the direct path itself the flag is inert.
+ */
+ if (nf->nf_file->f_op->fop_flags & FOP_DONTCACHE)
+ segments[nsegs].flags |= IOCB_DONTCACHE;
nsegs++;
if (suffix)
--
2.52.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH v2 2/9] NFSD: only split a direct-mode WRITE for a worthwhile direct middle
2026-09-29 23:13 [PATCH v2 0/9] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
2026-09-29 23:13 ` [PATCH v2 1/9] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well Mike Snitzer
@ 2026-09-29 23:13 ` Mike Snitzer
2026-09-30 21:44 ` Chuck Lever
2026-09-29 23:13 ` [PATCH v2 3/9] NFSD: do not use direct I/O for a READ smaller than its alignment Mike Snitzer
` (6 subsequent siblings)
8 siblings, 1 reply; 14+ messages in thread
From: Mike Snitzer @ 2026-09-29 23:13 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: linux-nfs
nfsd_write_dio_iters_init() decides how a WRITE issued in a direct mode
is carried out: split into a buffered prefix, a direct middle and a
buffered suffix, or issued whole as a single buffered segment.
Only split when the split buys direct I/O for a worthwhile middle. Add
direct_misaligned_num_pages (debugfs knob, default 2) and decline to
split when the aligned middle is smaller than that many pages and the
WRITE has a prefix or a suffix. A WRITE whose payload memory is
misaligned for the block device already declines: no segment of it can
be direct I/O, so a split would only turn one buffered write into
three.
Decide every segment's flags in nfsd_write_dio_iters_init() rather than
in the write loop. The buffered prefix and suffix of a split WRITE,
and the whole WRITE when it is not split, are issued IOCB_DONTCACHE
when the file system supports FOP_DONTCACHE and as plain buffered I/O
otherwise: the operator chose a direct mode to keep NFSD out of the
page cache, so a WRITE that cannot be direct gets the next best thing.
Document it in the "Misaligned WRITE" section of nfsd-io-modes.rst.
Suggested-by: Christoph Hellwig <hch@infradead.org>
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
.../filesystems/nfs/nfsd-io-modes.rst | 17 ++++-
fs/nfsd/debugfs.c | 14 ++++
fs/nfsd/nfsd.h | 1 +
fs/nfsd/vfs.c | 73 +++++++++++++------
4 files changed, 79 insertions(+), 26 deletions(-)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index dc50c930f9762..fef062f24c525 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -126,9 +126,9 @@ Misaligned WRITE:
middle and end as needed. The large middle segment is DIO-aligned
and the start and/or end are misaligned. Buffered IO is used for the
misaligned segments and O_DIRECT is used for the middle DIO-aligned
- segment. DONTCACHE buffered IO is _not_ used for the misaligned
- segments because using normal buffered IO offers significant RMW
- performance benefit when handling streaming misaligned WRITEs.
+ segment. If the underlying filesystem supports FOP_DONTCACHE, the
+ misaligned segments use DONTCACHE buffered IO so that their pages
+ are dropped from the page cache once written back.
The O_DIRECT middle segment also carries the DONTCACHE flag. It has
no effect while the IO really is O_DIRECT, but a filesystem may
@@ -140,6 +140,17 @@ Misaligned WRITE:
normal buffered IO. Such fallbacks are visible through the
iomap_dio_invalidate_fail trace event; see Tracing below.
+ Whenever no part of a WRITE can use O_DIRECT, the whole WRITE is
+ issued as a single DONTCACHE buffered IO (normal buffered IO if the
+ filesystem lacks FOP_DONTCACHE). This covers: a filesystem that
+ advertises no DIO alignment requirements at all; a WRITE smaller
+ than the larger of the offset and memory alignments; a WRITE whose
+ DIO-aligned middle segment is smaller than
+ /sys/kernel/debug/nfsd/direct_misaligned_num_pages pages (default 2)
+ while also having a misaligned start or end; and a WRITE whose
+ payload memory is not aligned to the block device's dma_alignment,
+ which rules out O_DIRECT for the middle segment as well.
+
Tracing:
The nfsd_read_direct trace event shows how NFSD expands any
misaligned READ to the next DIO-aligned block (on either end of the
diff --git a/fs/nfsd/debugfs.c b/fs/nfsd/debugfs.c
index 386fd1c54f527..398e400d64038 100644
--- a/fs/nfsd/debugfs.c
+++ b/fs/nfsd/debugfs.c
@@ -128,6 +128,17 @@ void nfsd_debugfs_exit(void)
nfsd_top_dir = NULL;
}
+/*
+ * /sys/kernel/debug/nfsd/direct_misaligned_num_pages
+ *
+ * The smallest DIO-aligned middle segment, in pages, that is worth
+ * splitting a misaligned direct-mode WRITE into three segments for. A
+ * WRITE whose middle is smaller than this, and which has a misaligned
+ * start or end, is issued as a single buffered segment instead.
+ *
+ * Default 2. Not yet tuned by benchmarking.
+ */
+
void nfsd_debugfs_init(void)
{
nfsd_top_dir = debugfs_create_dir("nfsd", NULL);
@@ -140,6 +151,9 @@ void nfsd_debugfs_init(void)
debugfs_create_file("io_cache_write", 0644, nfsd_top_dir, NULL,
&nfsd_io_cache_write_fops);
+
+ debugfs_create_u32("direct_misaligned_num_pages", 0644, nfsd_top_dir,
+ &nfsd_direct_misaligned_num_pages);
#ifdef CONFIG_NFSD_V4
debugfs_create_bool("delegated_timestamps", 0644, nfsd_top_dir,
&nfsd_delegts_enabled);
diff --git a/fs/nfsd/nfsd.h b/fs/nfsd/nfsd.h
index a145294c59c87..a2d72434160af 100644
--- a/fs/nfsd/nfsd.h
+++ b/fs/nfsd/nfsd.h
@@ -142,6 +142,7 @@ enum {
extern u64 nfsd_io_cache_read __read_mostly;
extern u64 nfsd_io_cache_write __read_mostly;
+extern u32 nfsd_direct_misaligned_num_pages __read_mostly;
extern int nfsd_max_blksize;
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 5b963f0e3b3f2..e3ce66bce00d4 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -53,6 +53,7 @@
bool nfsd_disable_splice_read __read_mostly;
u64 nfsd_io_cache_read __read_mostly = NFSD_IO_BUFFERED;
u64 nfsd_io_cache_write __read_mostly = NFSD_IO_BUFFERED;
+u32 nfsd_direct_misaligned_num_pages __read_mostly = 2;
/**
* nfserrno - Map Linux errnos to NFS errnos
@@ -1302,15 +1303,29 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
u32 mem_align = nf->nf_dio_mem_align;
size_t prefix, middle, suffix;
loff_t offset = iocb->ki_pos;
+ unsigned int dontcache_flags = 0;
unsigned int nsegs = 0;
+ if (nf->nf_file->f_op->fop_flags & FOP_DONTCACHE)
+ dontcache_flags = IOCB_DONTCACHE;
+
/*
- * Check if direct I/O is feasible for this write request.
- * If alignments are not available, the write is too small,
- * or no alignment can be found, fall back to buffered I/O.
+ * Whenever direct I/O cannot be used for the WRITE, fall back to a
+ * single DONTCACHE buffered I/O when the file system supports it, so
+ * the WRITE's pages are dropped from the page cache once written
+ * back, and to a single cached buffered I/O otherwise.
+ *
+ * If the file system doesn't advertise any alignment requirements,
+ * don't try to issue direct I/O at all.
*/
- if (unlikely(!mem_align || !offset_align) ||
- unlikely(total < max(offset_align, mem_align)))
+ if (unlikely(!mem_align || !offset_align))
+ goto no_dio;
+
+ /*
+ * If the I/O is smaller than the larger of the memory and logical
+ * offset alignment, no part of it can be direct I/O.
+ */
+ if (unlikely(total < max(offset_align, mem_align)))
goto no_dio;
prefix_end = round_up(offset, offset_align);
@@ -1321,12 +1336,27 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
middle = middle_end - prefix_end;
suffix = orig_end - middle_end;
- if (!middle)
+ /*
+ * If there is no aligned middle section, or the aligned part is too
+ * small to be worth the split (direct_misaligned_num_pages), issue a
+ * single buffered I/O write instead of splitting up the write.
+ */
+ if (!middle ||
+ ((prefix || suffix) &&
+ middle < PAGE_SIZE * nfsd_direct_misaligned_num_pages)) {
goto no_dio;
+ }
- if (prefix)
- nfsd_write_dio_seg_init(&segments[nsegs++], bvec,
+ /*
+ * The prefix and suffix are buffered I/O by definition. Mark them
+ * uncached when possible so their folios are dropped once written
+ * back rather than lingering in the page cache.
+ */
+ if (prefix) {
+ nfsd_write_dio_seg_init(&segments[nsegs], bvec,
nvecs, total, 0, prefix, iocb);
+ segments[nsegs++].flags |= dontcache_flags;
+ }
nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs,
total, prefix, middle, iocb);
@@ -1337,10 +1367,13 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
* bvecs generated from RPC receive buffers are contiguous: After
* the first bvec, all subsequent bvecs start at bv_offset zero
* (page-aligned). Therefore, only the first bvec is checked.
+ *
+ * If the memory is not aligned, direct I/O is impossible for the
+ * middle, so issue the entire write as a single buffered segment:
+ * splitting would only turn one buffered write into three.
*/
if (iov_iter_bvec_offset(&segments[nsegs].iter) & (mem_align - 1))
goto no_dio;
- segments[nsegs].flags |= IOCB_DIRECT;
/*
* Also mark the direct middle DONTCACHE: the file system may fall
* back to buffered I/O on its own (e.g. XFS on -ENOTBLK when it
@@ -1348,20 +1381,21 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
* suffix of an adjacent WRITE just dirtied), and it reuses this kiocb
* to do so. On the direct path itself the flag is inert.
*/
- if (nf->nf_file->f_op->fop_flags & FOP_DONTCACHE)
- segments[nsegs].flags |= IOCB_DONTCACHE;
- nsegs++;
+ segments[nsegs++].flags |= IOCB_DIRECT | dontcache_flags;
- if (suffix)
- nfsd_write_dio_seg_init(&segments[nsegs++], bvec, nvecs, total,
+ if (suffix) {
+ nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs, total,
prefix + middle, suffix, iocb);
+ segments[nsegs++].flags |= dontcache_flags;
+ }
return nsegs;
no_dio:
- /* No DIO alignment possible - pack into single non-DIO segment. */
+ /* No DIO possible - pack into a single uncached (if possible) segment. */
nfsd_write_dio_seg_init(&segments[0], bvec, nvecs, total, 0,
total, iocb);
+ segments[0].flags |= dontcache_flags;
return 1;
}
@@ -1385,16 +1419,9 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
if (kiocb->ki_flags & IOCB_DIRECT)
trace_nfsd_write_direct(rqstp, fhp, kiocb->ki_pos,
segments[i].iter.count);
- else {
+ else
trace_nfsd_write_vector(rqstp, fhp, kiocb->ki_pos,
segments[i].iter.count);
- /*
- * Mark the I/O buffer as evict-able to reduce
- * memory contention.
- */
- if (nf->nf_file->f_op->fop_flags & FOP_DONTCACHE)
- kiocb->ki_flags |= IOCB_DONTCACHE;
- }
expected = iov_iter_count(&segments[i].iter);
--
2.52.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH v2 3/9] NFSD: do not use direct I/O for a READ smaller than its alignment
2026-09-29 23:13 [PATCH v2 0/9] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
2026-09-29 23:13 ` [PATCH v2 1/9] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well Mike Snitzer
2026-09-29 23:13 ` [PATCH v2 2/9] NFSD: only split a direct-mode WRITE for a worthwhile direct middle Mike Snitzer
@ 2026-09-29 23:13 ` Mike Snitzer
2026-09-29 23:13 ` [PATCH v2 4/9] NFSD: Enable return of an updated stable_how to NFS clients Mike Snitzer
` (5 subsequent siblings)
8 siblings, 0 replies; 14+ messages in thread
From: Mike Snitzer @ 2026-09-29 23:13 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: linux-nfs
nfsd_direct_read() expands a misaligned READ out to DIO-aligned
boundaries: it reads from round_down(offset, dio_read_offset_align) to
round_up(offset + count, dio_read_offset_align) and returns only the
requested bytes from within that window. When the READ is smaller than
the alignment, that window is always at least one full alignment unit,
and two when the READ straddles a boundary, so a few hundred bytes of
payload can cost a 4K or 64K device read.
Decline direct I/O for those. A READ smaller than
dio_read_offset_align now falls through to the DONTCACHE path, which
issues DONTCACHE buffered I/O when the file system supports
FOP_DONTCACHE and normal buffered I/O otherwise. This mirrors the
WRITE side, which already declines direct I/O for a WRITE smaller than
the larger of its offset and memory alignments.
Only dio_read_offset_align is consulted, because the READ path fills
page-aligned pages from rq_bvec and so has no memory alignment to
satisfy. The threshold only bites when the file system advertises a
large alignment; where it reports 512 almost no READ is excluded.
Document it in the "Misaligned READ" section of nfsd-io-modes.rst.
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
Documentation/filesystems/nfs/nfsd-io-modes.rst | 7 +++++++
fs/nfsd/vfs.c | 6 +++---
2 files changed, 10 insertions(+), 3 deletions(-)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index fef062f24c525..d6ebb82f10b48 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -121,6 +121,13 @@ Misaligned READ:
verified to have proper offset/len (logical_block_size) and
dma_alignment checking.
+ A READ smaller than dio_read_offset_align is not issued as O_DIRECT
+ at all. Expanding it would read a whole alignment unit, or two when
+ the READ straddles a boundary, to return those few bytes. Such a
+ READ is issued as DONTCACHE buffered IO instead (normal buffered IO
+ if the filesystem lacks FOP_DONTCACHE), mirroring the WRITE that is
+ smaller than its own alignment.
+
Misaligned WRITE:
If NFSD_IO_DIRECT is used, split any misaligned WRITE into a start,
middle and end as needed. The large middle segment is DIO-aligned
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index e3ce66bce00d4..1d2b03cb42963 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1186,7 +1186,7 @@ __be32 nfsd_iter_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
unsigned int base, u32 *eof)
{
struct file *file = nf->nf_file;
- unsigned long v, total;
+ unsigned long v, total = *count;
struct iov_iter iter;
struct kiocb kiocb;
ssize_t host_err;
@@ -1199,7 +1199,8 @@ __be32 nfsd_iter_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
break;
case NFSD_IO_DIRECT:
/* When dio_read_offset_align is zero, dio is not supported */
- if (nf->nf_dio_read_offset_align && !rqstp->rq_res.page_len)
+ if (nf->nf_dio_read_offset_align && !rqstp->rq_res.page_len &&
+ total >= nf->nf_dio_read_offset_align)
return nfsd_direct_read(rqstp, fhp, nf, offset,
count, eof);
fallthrough;
@@ -1212,7 +1213,6 @@ __be32 nfsd_iter_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
kiocb.ki_pos = offset;
v = 0;
- total = *count;
while (total && v < rqstp->rq_maxpages &&
rqstp->rq_next_page < rqstp->rq_page_end) {
len = min_t(size_t, total, PAGE_SIZE - base);
--
2.52.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH v2 4/9] NFSD: Enable return of an updated stable_how to NFS clients
2026-09-29 23:13 [PATCH v2 0/9] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
` (2 preceding siblings ...)
2026-09-29 23:13 ` [PATCH v2 3/9] NFSD: do not use direct I/O for a READ smaller than its alignment Mike Snitzer
@ 2026-09-29 23:13 ` Mike Snitzer
2026-09-30 21:42 ` Chuck Lever
2026-09-29 23:13 ` [PATCH v2 5/9] NFSD: let a direct-mode WRITE raise stable_how and elide the client's COMMIT Mike Snitzer
` (4 subsequent siblings)
8 siblings, 1 reply; 14+ messages in thread
From: Mike Snitzer @ 2026-09-29 23:13 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: linux-nfs
From: Chuck Lever <chuck.lever@oracle.com>
NFSv3 and newer protocols enable clients to perform a two-phase
WRITE. A client requests an UNSTABLE WRITE, which sends dirty data
to the NFS server, but does not persist it. The server replies that
it performed the UNSTABLE WRITE, and the client is then obligated to
follow up with a COMMIT request before it can remove the dirty data
from its own page cache. The COMMIT reply is the client's guarantee
that the written data has been persisted on the server.
The purpose of this protocol design is to enable clients to send
a large amount of data via multiple WRITE requests to a server, and
then wait for persistence just once. The server is able to start
persisting the data as soon as it gets it, to shorten the length of
time the client has to wait for the final COMMIT to complete.
It's also possible for the server to respond to an UNSTABLE WRITE
request in a way that indicates that the data was persisted anyway.
In that case, the client can skip the COMMIT and remove the dirty
data from its memory immediately. NetApp filers, for example, do
this because they have a battery-backed cache and can guarantee that
written data is persisted quickly and immediately.
NFSD has never implemented this kind of promotion. UNSTABLE WRITE
requests are unconditionally treated as UNSTABLE. However, in a
subsequent patch, nfsd_vfs_write() will be able to promote an
UNSTABLE WRITE to be a FILE_SYNC WRITE. This will be because NFSD
will handle some WRITE requests locally with O_DIRECT, which
persists written data immediately. The FILE_SYNC WRITE response
indicates to the client that no follow-up COMMIT is necessary.
This patch prepares for that change by making the @iocb_flags
argument of nfsd_write() and nfsd_vfs_write() bi-directional. A
caller passes in the IOCB_* flags that express the stability its
client asked for; on return the argument holds the flags that were
actually satisfied. Each protocol version maps that result back to
its own on-the-wire value when it encodes the WRITE reply, using the
new nfsd3_stable_how() and nfsd4_stable_how() helpers, so that NFS
stable_how values stay out of NFSD's generic VFS API. No behavior
change is expected.
[snitzer: reworked onto the @iocb_flags argument introduced by
commit 6dcddbb70b08 ("NFSD: Replace nfsd_write()'s "stable" argument with "iocb_flags""),
which was applied after this patch was first posted.
The original passed a "u32 *stable_how" instead. Carrying the value as
IOCB_* flags keeps the NFSv3 XDR value out of the VFS API, and lets a
later patch raise the achieved stability with a plain bitwise OR rather
than an ordering comparison. The Reviewed-by tags from the original
posting are dropped because the argument's type and direction changed.]
Signed-off-by: Chuck Lever <chuck.lever@oracle.com>
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
fs/nfsd/nfs3proc.c | 16 +++++++++++++---
fs/nfsd/nfs4proc.c | 15 +++++++++++++--
fs/nfsd/nfsproc.c | 3 ++-
fs/nfsd/vfs.c | 18 +++++++++++-------
fs/nfsd/vfs.h | 4 ++--
fs/nfsd/xdr3.h | 2 +-
6 files changed, 42 insertions(+), 16 deletions(-)
diff --git a/fs/nfsd/nfs3proc.c b/fs/nfsd/nfs3proc.c
index 60cd01b6a37d2..f06573759ad6f 100644
--- a/fs/nfsd/nfs3proc.c
+++ b/fs/nfsd/nfs3proc.c
@@ -102,6 +102,15 @@ static int nfsd3_iocb_flags(enum nfs3_stable_how how)
}
}
+static enum nfs3_stable_how nfsd3_stable_how(int iocb_flags)
+{
+ if (iocb_flags & IOCB_SYNC)
+ return NFS_FILE_SYNC;
+ if (iocb_flags & IOCB_DSYNC)
+ return NFS_DATA_SYNC;
+ return NFS_UNSTABLE;
+}
+
static __be32 nfsd3_map_status(__be32 status)
{
switch (status) {
@@ -299,6 +308,7 @@ nfsd3_proc_write(struct svc_rqst *rqstp)
struct nfsd3_writeargs *argp = rqstp->rq_argp;
struct nfsd3_writeres *resp = rqstp->rq_resp;
unsigned long cnt = argp->len;
+ int iocb_flags;
dprintk("nfsd: WRITE(3) %s %d bytes at %Lu%s\n",
SVCFH_fmt(&argp->fh),
@@ -312,11 +322,11 @@ nfsd3_proc_write(struct svc_rqst *rqstp)
return rpc_success;
fh_copy(&resp->fh, &argp->fh);
- resp->committed = argp->stable;
+ iocb_flags = nfsd3_iocb_flags(argp->stable);
resp->status = nfsd_write(rqstp, &resp->fh, argp->offset,
- &argp->payload, &cnt,
- nfsd3_iocb_flags(resp->committed),
+ &argp->payload, &cnt, &iocb_flags,
resp->verf);
+ resp->committed = nfsd3_stable_how(iocb_flags);
resp->count = cnt;
resp->status = nfsd3_map_status(resp->status);
return rpc_success;
diff --git a/fs/nfsd/nfs4proc.c b/fs/nfsd/nfs4proc.c
index bb74eef439388..051d900581609 100644
--- a/fs/nfsd/nfs4proc.c
+++ b/fs/nfsd/nfs4proc.c
@@ -109,6 +109,15 @@ static const struct nfsd_access_maps nfsd4_access_maps = {
.other = nfsd4_otheraccess,
};
+static enum stable_how4 nfsd4_stable_how(int iocb_flags)
+{
+ if (iocb_flags & IOCB_SYNC)
+ return FILE_SYNC4;
+ if (iocb_flags & IOCB_DSYNC)
+ return DATA_SYNC4;
+ return UNSTABLE4;
+}
+
static int nfsd4_iocb_flags(enum stable_how4 how)
{
switch (how) {
@@ -1432,6 +1441,7 @@ nfsd4_write(struct svc_rqst *rqstp, struct nfsd4_compound_state *cstate,
struct nfsd_file *nf = NULL;
__be32 status = nfs_ok;
unsigned long cnt;
+ int iocb_flags;
if (write->wr_offset > (u64)OFFSET_MAX ||
write->wr_offset + write->wr_buflen > (u64)OFFSET_MAX)
@@ -1450,11 +1460,12 @@ nfsd4_write(struct svc_rqst *rqstp, struct nfsd4_compound_state *cstate,
nfs4_put_stid(stid);
}
- write->wr_how_written = write->wr_stable_how;
+ iocb_flags = nfsd4_iocb_flags(write->wr_stable_how);
status = nfsd_vfs_write(rqstp, &cstate->current_fh, nf,
write->wr_offset, &write->wr_payload,
- &cnt, nfsd4_iocb_flags(write->wr_how_written),
+ &cnt, &iocb_flags,
(__be32 *)write->wr_verifier.data);
+ write->wr_how_written = nfsd4_stable_how(iocb_flags);
nfsd_file_put(nf);
write->wr_bytes_written = cnt;
diff --git a/fs/nfsd/nfsproc.c b/fs/nfsd/nfsproc.c
index 09d3608398250..a847b35006781 100644
--- a/fs/nfsd/nfsproc.c
+++ b/fs/nfsd/nfsproc.c
@@ -274,6 +274,7 @@ nfsd_proc_write(struct svc_rqst *rqstp)
struct nfsd_writeargs *argp = rqstp->rq_argp;
struct nfsd_attrstat *resp = rqstp->rq_resp;
unsigned long cnt = argp->len;
+ int iocb_flags = IOCB_DSYNC;
dprintk("nfsd: WRITE %s %u bytes at %d\n",
SVCFH_fmt(&argp->fh),
@@ -281,7 +282,7 @@ nfsd_proc_write(struct svc_rqst *rqstp)
fh_copy(&resp->fh, &argp->fh);
resp->status = nfsd_write(rqstp, &resp->fh, argp->offset,
- &argp->payload, &cnt, IOCB_DSYNC, NULL);
+ &argp->payload, &cnt, &iocb_flags, NULL);
if (resp->status == nfs_ok)
resp->status = fh_getattr(&resp->fh, &resp->stat);
else if (resp->status == nfserr_jukebox)
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 1d2b03cb42963..82eba97656c4e 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1444,7 +1444,9 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
* @offset: Byte offset of start
* @payload: xdr_buf containing the write payload
* @cnt: IN: number of bytes to write, OUT: number of bytes actually written
- * @iocb_flags: VFS IOCB_* flags expressing the requested write stability
+ * @iocb_flags: IN: VFS IOCB_* flags expressing the requested write
+ * stability; OUT: the flags actually satisfied, which may be
+ * higher than requested
* @verf: NFS WRITE verifier
*
* Upon return, caller must invoke fh_put on @fhp.
@@ -1456,7 +1458,7 @@ __be32
nfsd_vfs_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
struct nfsd_file *nf, loff_t offset,
const struct xdr_buf *payload, unsigned long *cnt,
- int iocb_flags, __be32 *verf)
+ int *iocb_flags, __be32 *verf)
{
struct nfsd_net *nn = net_generic(SVC_NET(rqstp), nfsd_net_id);
struct file *file = nf->nf_file;
@@ -1493,11 +1495,11 @@ nfsd_vfs_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
exp = fhp->fh_export;
if (!EX_ISSYNC(exp))
- iocb_flags = 0;
+ *iocb_flags = 0;
init_sync_kiocb(&kiocb, file);
kiocb.ki_pos = offset;
if (likely(!fhp->fh_use_wgather))
- kiocb.ki_flags |= iocb_flags;
+ kiocb.ki_flags |= *iocb_flags;
nvecs = xdr_buf_to_bvec(rqstp->rq_bvec, rqstp->rq_maxpages, payload);
if (nvecs < 0) {
@@ -1538,7 +1540,7 @@ nfsd_vfs_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
goto out_nfserr;
}
- if (iocb_flags && fhp->fh_use_wgather) {
+ if (*iocb_flags && fhp->fh_use_wgather) {
host_err = wait_for_concurrent_writes(file);
if (host_err < 0)
nfsd_maybe_reset_write_verifier(nn, rqstp, host_err);
@@ -1629,7 +1631,9 @@ __be32 nfsd_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
* @offset: Byte offset of start
* @payload: xdr_buf containing the write payload
* @cnt: IN: number of bytes to write, OUT: number of bytes actually written
- * @iocb_flags: VFS IOCB_* flags expressing the requested write stability
+ * @iocb_flags: IN: VFS IOCB_* flags expressing the requested write
+ * stability; OUT: the flags actually satisfied, which may be
+ * higher than requested
* @verf: NFS WRITE verifier
*
* Upon return, caller must invoke fh_put on @fhp.
@@ -1640,7 +1644,7 @@ __be32 nfsd_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
__be32
nfsd_write(struct svc_rqst *rqstp, struct svc_fh *fhp, loff_t offset,
const struct xdr_buf *payload, unsigned long *cnt,
- int iocb_flags, __be32 *verf)
+ int *iocb_flags, __be32 *verf)
{
struct nfsd_file *nf;
__be32 err;
diff --git a/fs/nfsd/vfs.h b/fs/nfsd/vfs.h
index f0cb184643f2f..6b352ca7f02b7 100644
--- a/fs/nfsd/vfs.h
+++ b/fs/nfsd/vfs.h
@@ -154,12 +154,12 @@ __be32 nfsd_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
u32 *eof);
__be32 nfsd_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
loff_t offset, const struct xdr_buf *payload,
- unsigned long *cnt, int iocb_flags,
+ unsigned long *cnt, int *iocb_flags,
__be32 *verf);
__be32 nfsd_vfs_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
struct nfsd_file *nf, loff_t offset,
const struct xdr_buf *payload,
- unsigned long *cnt, int iocb_flags,
+ unsigned long *cnt, int *iocb_flags,
__be32 *verf);
__be32 nfsd_readlink(struct svc_rqst *, struct svc_fh *,
char *, int *);
diff --git a/fs/nfsd/xdr3.h b/fs/nfsd/xdr3.h
index cad875d142313..a2a980559fb72 100644
--- a/fs/nfsd/xdr3.h
+++ b/fs/nfsd/xdr3.h
@@ -153,7 +153,7 @@ struct nfsd3_writeres {
__be32 status;
struct svc_fh fh;
unsigned long count;
- int committed;
+ u32 committed;
__be32 verf[2];
};
--
2.52.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH v2 5/9] NFSD: let a direct-mode WRITE raise stable_how and elide the client's COMMIT
2026-09-29 23:13 [PATCH v2 0/9] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
` (3 preceding siblings ...)
2026-09-29 23:13 ` [PATCH v2 4/9] NFSD: Enable return of an updated stable_how to NFS clients Mike Snitzer
@ 2026-09-29 23:13 ` Mike Snitzer
2026-09-30 21:46 ` Chuck Lever
2026-09-29 23:13 ` [PATCH v2 6/9] NFSD: persist a synchronous direct-mode WRITE once, after all of its segments Mike Snitzer
` (3 subsequent siblings)
8 siblings, 1 reply; 14+ messages in thread
From: Mike Snitzer @ 2026-09-29 23:13 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: linux-nfs
A WRITE serviced in a direct mode is already persistent when NFSD
replies: the aligned middle is O_DIRECT, and a synchronous WRITE fsyncs
whatever was buffered before the reply is sent. The reply still says
UNSTABLE, so the client dutifully sends a COMMIT for data that is
already on stable storage, and NFSD answers it with an fsync that has
nothing left to write.
Add two io_cache_write modes that issue direct I/O exactly like
NFSD_IO_DIRECT and raise the stable_how of the reply to a floor:
NFSD_IO_DIRECT_WRITE_DATA_SYNC (3): at least NFS_DATA_SYNC
NFSD_IO_DIRECT_WRITE_FILE_SYNC (4): at least NFS_FILE_SYNC
A client that asked for more is left alone, and the WRITE is persisted
to the level the reply reports before that reply is sent. A client told
FILE_SYNC has no reason to COMMIT and sends none.
What this removes is the COMMIT traffic, and it is worth most where a
client's writes are carved into several WRITE RPCs, because each piece
is then committed separately. Measured on a pNFS flexfiles share where
every write straddles two data servers, so every write becomes two WRITE
RPCs and, under NFSD_IO_DIRECT, two COMMITs: the two modes do identical
durability work, one fsync per COMMIT against one fsync per WRITE on
identical WRITE counts, and the COMMIT RPCs alone cost NFSD_IO_DIRECT
31% more server CPU and 42% more client CPU for the same bytes, about
20 us of server CPU per COMMIT plus a client cost that grows with the
range committed. Where only one write in 22 is split the same effect
is a couple of cores on each side and no resolvable throughput
difference, and writes that fit a single RPC send no COMMIT in either
mode.
Choose 4 when the export services WRITEs in a direct mode and clients
split their writes; choose 3 to promise only the data, leaving a client
that needs metadata durability to COMMIT for it. Both are inert for
READ, and for a WRITE the client already marked FILE_SYNC.
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
.../filesystems/nfs/nfsd-io-modes.rst | 17 ++++++++--
fs/nfsd/debugfs.c | 5 +++
fs/nfsd/nfsd.h | 2 ++
fs/nfsd/vfs.c | 33 +++++++++++++++++--
4 files changed, 52 insertions(+), 5 deletions(-)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index d6ebb82f10b48..9ee94cfa546d8 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -25,12 +25,14 @@ Based on the configured settings, NFSD's IO will either be:
- cached using page cache (NFSD_IO_BUFFERED=0)
- cached but removed from page cache on completion (NFSD_IO_DONTCACHE=1)
- not cached stable_how=NFS_UNSTABLE (NFSD_IO_DIRECT=2)
+- not cached stable_how=NFS_DATA_SYNC (NFSD_IO_DIRECT_WRITE_DATA_SYNC=3)
+- not cached stable_how=NFS_FILE_SYNC (NFSD_IO_DIRECT_WRITE_FILE_SYNC=4)
-To set an NFSD IO mode, write a supported value (0 - 2) to the
+To set an NFSD IO mode, write a supported value (0 - 4) to the
corresponding IO operation's debugfs interface, e.g.::
echo 2 > /sys/kernel/debug/nfsd/io_cache_read
- echo 2 > /sys/kernel/debug/nfsd/io_cache_write
+ echo 4 > /sys/kernel/debug/nfsd/io_cache_write
To check which IO mode NFSD is using for READ or WRITE, simply read the
corresponding IO operation's debugfs interface, e.g.::
@@ -38,6 +40,17 @@ corresponding IO operation's debugfs interface, e.g.::
cat /sys/kernel/debug/nfsd/io_cache_read
cat /sys/kernel/debug/nfsd/io_cache_write
+The two NFSD_IO_DIRECT_WRITE_*_SYNC modes raise the stable_how of every
+WRITE to at least NFS_DATA_SYNC or NFS_FILE_SYNC, persist the WRITE
+accordingly before replying, and return the raised value to the client;
+a client that asked for a higher stable_how is left alone. With
+NFSD_IO_DIRECT_WRITE_FILE_SYNC the client sends no COMMIT. Against
+NFSD_IO_DIRECT the durability work is the same, one fsync per WRITE
+instead of one per COMMIT; what NFSD_IO_DIRECT adds is the COMMIT RPCs
+themselves, tens of microseconds of server CPU each plus a client cost
+that grows with the range committed, which matters in proportion to how
+many of a client's WRITEs need a COMMIT.
+
If you experiment with NFSD's IO modes on a recent kernel and have
interesting results, please report them to linux-nfs@vger.kernel.org
diff --git a/fs/nfsd/debugfs.c b/fs/nfsd/debugfs.c
index 398e400d64038..279e341a81ec1 100644
--- a/fs/nfsd/debugfs.c
+++ b/fs/nfsd/debugfs.c
@@ -90,6 +90,9 @@ DEFINE_DEBUGFS_ATTRIBUTE(nfsd_io_cache_read_fops, nfsd_io_cache_read_get,
* Contents:
* %0: NFS WRITE will use buffered IO
* %1: NFS WRITE will use dontcache (buffered IO w/ dropbehind)
+ * %2: NFS WRITE will use direct IO with stable_how=NFS_UNSTABLE
+ * %3: NFS WRITE will use direct IO with stable_how=NFS_DATA_SYNC
+ * %4: NFS WRITE will use direct IO with stable_how=NFS_FILE_SYNC
*
* This setting takes immediate effect for all NFS versions,
* all exports, and in all NFSD net namespaces.
@@ -109,6 +112,8 @@ static int nfsd_io_cache_write_set(void *data, u64 val)
case NFSD_IO_BUFFERED:
case NFSD_IO_DONTCACHE:
case NFSD_IO_DIRECT:
+ case NFSD_IO_DIRECT_WRITE_DATA_SYNC:
+ case NFSD_IO_DIRECT_WRITE_FILE_SYNC:
nfsd_io_cache_write = val;
break;
default:
diff --git a/fs/nfsd/nfsd.h b/fs/nfsd/nfsd.h
index a2d72434160af..135e319e378d4 100644
--- a/fs/nfsd/nfsd.h
+++ b/fs/nfsd/nfsd.h
@@ -138,6 +138,8 @@ enum {
NFSD_IO_BUFFERED,
NFSD_IO_DONTCACHE,
NFSD_IO_DIRECT,
+ NFSD_IO_DIRECT_WRITE_DATA_SYNC,
+ NFSD_IO_DIRECT_WRITE_FILE_SYNC,
};
extern u64 nfsd_io_cache_read __read_mostly;
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 82eba97656c4e..924c5992dc32e 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1399,17 +1399,42 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
return 1;
}
+/*
+ * Raise the stability of this WRITE to at least @floor_iocb_flags, and
+ * record what was achieved in @iocb_flags so the reply can report it.
+ * A client that asked for more is left alone.
+ */
+static void
+nfsd_write_raise_stability(int floor_iocb_flags, struct kiocb *kiocb,
+ int *iocb_flags)
+{
+ if ((*iocb_flags & floor_iocb_flags) == floor_iocb_flags)
+ return; /* already at or above the floor */
+
+ *iocb_flags |= floor_iocb_flags;
+ kiocb->ki_flags |= floor_iocb_flags;
+}
+
static noinline_for_stack int
nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
- struct nfsd_file *nf, unsigned int nvecs,
+ struct nfsd_file *nf, int *iocb_flags, unsigned int nvecs,
unsigned long *cnt, struct kiocb *kiocb)
{
struct nfsd_write_dio_seg segments[3];
+ int floor_iocb_flags = 0;
struct file *file = nf->nf_file;
unsigned int nsegs, i;
ssize_t host_err;
size_t expected;
+ if (nfsd_io_cache_write == NFSD_IO_DIRECT_WRITE_FILE_SYNC)
+ floor_iocb_flags = IOCB_DSYNC | IOCB_SYNC;
+ else if (nfsd_io_cache_write == NFSD_IO_DIRECT_WRITE_DATA_SYNC)
+ floor_iocb_flags = IOCB_DSYNC;
+ if (floor_iocb_flags)
+ nfsd_write_raise_stability(floor_iocb_flags, kiocb,
+ iocb_flags);
+
nsegs = nfsd_write_dio_iters_init(nf, rqstp->rq_bvec, nvecs,
kiocb, *cnt, segments);
@@ -1513,8 +1538,10 @@ nfsd_vfs_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
switch (nfsd_io_cache_write) {
case NFSD_IO_DIRECT:
- host_err = nfsd_direct_write(rqstp, fhp, nf, nvecs,
- cnt, &kiocb);
+ case NFSD_IO_DIRECT_WRITE_DATA_SYNC:
+ case NFSD_IO_DIRECT_WRITE_FILE_SYNC:
+ host_err = nfsd_direct_write(rqstp, fhp, nf, iocb_flags,
+ nvecs, cnt, &kiocb);
break;
case NFSD_IO_DONTCACHE:
if (file->f_op->fop_flags & FOP_DONTCACHE)
--
2.52.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH v2 6/9] NFSD: persist a synchronous direct-mode WRITE once, after all of its segments
2026-09-29 23:13 [PATCH v2 0/9] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
` (4 preceding siblings ...)
2026-09-29 23:13 ` [PATCH v2 5/9] NFSD: let a direct-mode WRITE raise stable_how and elide the client's COMMIT Mike Snitzer
@ 2026-09-29 23:13 ` Mike Snitzer
2026-09-29 23:13 ` [PATCH v2 7/9] NFSD: keep boundary page of a split direct-mode WRITE until both writers complete Mike Snitzer
` (2 subsequent siblings)
8 siblings, 0 replies; 14+ messages in thread
From: Mike Snitzer @ 2026-09-29 23:13 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: linux-nfs
nfsd_direct_write() may issue a WRITE as up to three segments: a
buffered prefix, a direct middle and a buffered suffix. For a
FILE_SYNC or DATA_SYNC WRITE (from the client, or as the floor imposed
by NFSD_IO_DIRECT_WRITE_{DATA,FILE}_SYNC) the kiocb carries IOCB_DSYNC
and every segment inherits it, so generic_write_sync() runs a range
fsync after each segment: up to three cache flushes and log forces per
WRITE, and each one writes back and drops the boundary page it just
touched.
Strip IOCB_DSYNC and IOCB_SYNC from the per-segment flags and persist
the WRITE once with vfs_fsync_range() over the bytes actually written,
after the last segment. Durability is unchanged: the reply is not sent
until the fsync completes, and datasync mirrors the previous per-segment
choice (IOCB_SYNC present means metadata too). An fsync failure is
returned like a write failure.
Besides the fewer flushes, this puts the sync under NFSD's control,
which the next change uses to keep the boundary pages of a split WRITE
cached until the partner WRITE completes them.
Assisted-by: Claude:claude-fable-5-1
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
fs/nfsd/vfs.c | 20 +++++++++++++++++++-
1 file changed, 19 insertions(+), 1 deletion(-)
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 924c5992dc32e..827dfe0b5dac7 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1423,6 +1423,8 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
struct nfsd_write_dio_seg segments[3];
int floor_iocb_flags = 0;
struct file *file = nf->nf_file;
+ loff_t start = kiocb->ki_pos;
+ bool sync, datasync;
unsigned int nsegs, i;
ssize_t host_err;
size_t expected;
@@ -1435,12 +1437,21 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
nfsd_write_raise_stability(floor_iocb_flags, kiocb,
iocb_flags);
+ /*
+ * A synchronous WRITE (client FILE_SYNC/DATA_SYNC, or a floor set by
+ * the IO mode) is persisted once, after all of its segments, rather
+ * than by generic_write_sync() after each segment: one cache flush
+ * and log force instead of up to three.
+ */
+ sync = kiocb->ki_flags & IOCB_DSYNC;
+ datasync = !(kiocb->ki_flags & IOCB_SYNC);
+
nsegs = nfsd_write_dio_iters_init(nf, rqstp->rq_bvec, nvecs,
kiocb, *cnt, segments);
*cnt = 0;
for (i = 0; i < nsegs; i++) {
- kiocb->ki_flags = segments[i].flags;
+ kiocb->ki_flags = segments[i].flags & ~(IOCB_DSYNC | IOCB_SYNC);
if (kiocb->ki_flags & IOCB_DIRECT)
trace_nfsd_write_direct(rqstp, fhp, kiocb->ki_pos,
segments[i].iter.count);
@@ -1458,6 +1469,13 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
break; /* partial write */
}
+ if (sync && *cnt) {
+ host_err = vfs_fsync_range(file, start, start + *cnt - 1,
+ datasync);
+ if (host_err < 0)
+ return host_err;
+ }
+
return 0;
}
--
2.52.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH v2 7/9] NFSD: keep boundary page of a split direct-mode WRITE until both writers complete
2026-09-29 23:13 [PATCH v2 0/9] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
` (5 preceding siblings ...)
2026-09-29 23:13 ` [PATCH v2 6/9] NFSD: persist a synchronous direct-mode WRITE once, after all of its segments Mike Snitzer
@ 2026-09-29 23:13 ` Mike Snitzer
2026-09-29 23:13 ` [PATCH v2 8/9] NFSD: add direct_misaligned_dontcache debugfs knob Mike Snitzer
2026-09-29 23:13 ` [PATCH v2 9/9] NFSD: add tracing for how direct-mode READ and WRITE are serviced Mike Snitzer
8 siblings, 0 replies; 14+ messages in thread
From: Mike Snitzer @ 2026-09-29 23:13 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: linux-nfs
A direct-mode WRITE whose ends are not aligned for direct I/O is split
into a buffered prefix, a direct middle and a buffered suffix, and the
prefix and suffix are issued IOCB_DONTCACHE so their pages are dropped
once written back. The page holding such a segment is shared by exactly
two WRITEs, the one ending in it and the one starting in it, which may
arrive in either order, from different clients, at the same time. If it
is dropped as soon as the first writer's data is written back, the
second has to read the block from disk before it can complete it: one
4 KiB read per WRITE. With many clients interleaving small misaligned
O_DIRECT records into one shared file, that was 0.81 device reads per
record (585 GiB read while writing 8.1 TiB).
Keep the page until both writers have had it, with no state in NFSD and
no knowledge of how clients stripe their I/O. Immediately before a
boundary segment is written, nfsd_write_dio_boundary_claim() puts an
empty folio in the page cache if there is not one there already. The
folio is inserted unmarked, which is what makes this work: a DONTCACHE
write only marks a folio it allocated itself, so neither writer's write
marks this one and it never enters WB_DONTCACHE_DIRTY. That counter is
what the DONTCACHE writeback kick targets, and a boundary page counted
there is written back, and then dropped, in the gap between the two
WRITEs that share it.
The add is atomic, so of two concurrent partners exactly one gets the
folio. The other one finds it (-EEXIST) and is therefore the second of
the two: it completes the page, and once its data is in it
nfsd_write_dio_boundary_complete() marks the folio dropbehind. The mark
has to come after the write, because a clean marked folio is dropped by
whatever writeback completes next, and it has to stay unaccounted,
because counting a folio marked while dirty hands the page straight
back to the kick. Whichever writeback cleans the page then drops it:
the WRITE's own sync for FILE_SYNC and DATA_SYNC, the flusher or the
client's COMMIT for UNSTABLE.
A WRITE that cannot use direct I/O at all (memory-misaligned payload,
or too small for a direct middle) is issued as one buffered DONTCACHE
segment. Its first and last pages are shared with the neighbouring
WRITEs in the same way, and are claimed the same way.
Measured with 32 interleaved writers over an emulated 4Kn NVMe,
io_cache_write=4, direct_misaligned_dontcache=Y, each arm starting from
a freshly made file system. 47008-byte records, which split: 704 device
reads for 42895 WRITEs, against 45144 for 45664 WRITEs with the
mechanism compiled out. 6000-byte records, which never get a direct
middle and so are one buffered segment whose end pages are shared: 385
reads against 196589. The DONTCACHE flusher stays idle in both, and the
pages are dropped: 81 of the written file's 524066 pages are still
resident afterwards.
On a four-server pNFS flexfiles rig, three clients writing 47008-byte
records for 240 s against 640 nfsd threads per server, compared with
the same servers before this change: device reads fall from 0.014-0.015
per record to 0.003, page-cache growth over the write phase falls from
about 11 GiB to 2.9 GiB, and write throughput rises by 2.4% to 4.9%
across NFSD_IO_DIRECT and both stable_how floors, with read throughput
unchanged. Records of that size always take the split path, since their
direct middle is at least 38818 bytes, so the figures are the split
path alone; the whole-WRITE fallback needs records below about 12 KB to
come into play. Keeping every boundary page cached instead, which a
later commit makes possible with a direct_misaligned_dontcache=N
debugfs knob, grows the page cache by about 630 GiB over the same
runs.
Reported-by: Jonathan Flynn <jonathan.flynn@hammerspace.com>
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
.../filesystems/nfs/nfsd-io-modes.rst | 28 ++++
fs/nfsd/vfs.c | 131 ++++++++++++++++--
2 files changed, 151 insertions(+), 8 deletions(-)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index 9ee94cfa546d8..2ca606ddf5d51 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -150,6 +150,34 @@ Misaligned WRITE:
misaligned segments use DONTCACHE buffered IO so that their pages
are dropped from the page cache once written back.
+ A FILE_SYNC or DATA_SYNC WRITE (requested by the client, or imposed
+ as the floor by NFSD_IO_DIRECT_WRITE_FILE_SYNC and
+ NFSD_IO_DIRECT_WRITE_DATA_SYNC) is persisted once, after all of its
+ segments have been written, rather than after each segment.
+
+ The page holding a start or end segment is shared by exactly two
+ WRITEs, the one ending in it and the one starting in it, which may
+ arrive in either order, from different clients, and at the same
+ time. Both segments are issued as DONTCACHE buffered IO, so the page
+ would be dropped as soon as the first writer's data is written back,
+ leaving the second to read it back. Immediately before writing a
+ start or end segment NFSD therefore puts an empty page in the page
+ cache if there is not one already. That page is not marked "drop
+ behind" when it is created, so it does not count towards the
+ DONTCACHE writeback backlog and the writeback kick does not write it
+ back, and drop it, between the two WRITEs that share it.
+
+ Whichever WRITE finds the page already there is the second of the
+ two: it completes the page and marks it "drop behind" once its data
+ is in it, so whichever writeback cleans it afterwards (the WRITE's
+ own sync for FILE_SYNC or DATA_SYNC, the flusher or the client's
+ COMMIT for UNSTABLE) drops it. A WRITE that gets no direct middle at
+ all is issued as a single buffered DONTCACHE segment, and its first
+ and last pages are shared and handled the same way. The retained
+ page cache is the set of half-written boundary pages, which grows
+ with how far concurrent writers drift apart, not with bytes
+ written.
+
The O_DIRECT middle segment also carries the DONTCACHE flag. It has
no effect while the IO really is O_DIRECT, but a filesystem may
decide on its own to service the segment with buffered IO instead
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 827dfe0b5dac7..5fd850a29694f 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1271,6 +1271,8 @@ static int wait_for_concurrent_writes(struct file *file)
struct nfsd_write_dio_seg {
struct iov_iter iter;
int flags;
+ bool boundary; /* prefix or suffix of a split */
+ bool edges; /* buffered fallback: partial end pages */
};
static unsigned long
@@ -1290,6 +1292,73 @@ nfsd_write_dio_seg_init(struct nfsd_write_dio_seg *segment,
iov_iter_advance(&segment->iter, start);
iov_iter_truncate(&segment->iter, len);
segment->flags = iocb->ki_flags;
+ segment->boundary = false;
+ segment->edges = false;
+}
+
+/**
+ * nfsd_write_dio_boundary_claim - claim the page a boundary segment shares
+ * @file: the file being written
+ * @pos: any byte offset within the page
+ *
+ * The page holding a misaligned prefix or suffix is shared by exactly two
+ * WRITEs, the one ending in it and the one starting in it, which may arrive
+ * in either order, from different clients, at the same time. Put an empty
+ * folio there if there is not one already: it is inserted unmarked, so
+ * neither writer's DONTCACHE write marks it (a DONTCACHE write only marks a
+ * folio it allocated itself) and it never enters WB_DONTCACHE_DIRTY, which
+ * is what would otherwise arm the DONTCACHE writeback kick and have the
+ * flusher write the page back, and drop it, between the two WRITEs.
+ *
+ * The add is atomic, so of two concurrent partners exactly one gets the
+ * folio. If it cannot be allocated the segment is an ordinary DONTCACHE
+ * write and the page is dropped after writeback as before.
+ *
+ * Return: true if this WRITE is the second of the two and must call
+ * nfsd_write_dio_boundary_complete() once its segment has been written.
+ */
+static bool
+nfsd_write_dio_boundary_claim(struct file *file, loff_t pos)
+{
+ struct address_space *mapping = file->f_mapping;
+ pgoff_t index = pos >> PAGE_SHIFT;
+ gfp_t gfp = mapping_gfp_mask(mapping);
+ struct folio *folio;
+ int err;
+
+ folio = filemap_alloc_folio(gfp, 0, NULL);
+ if (!folio)
+ return false;
+ err = filemap_add_folio(mapping, folio, index, gfp);
+ if (!err) {
+ /* First writer: the page is in place, waiting for the partner. */
+ folio_unlock(folio);
+ folio_put(folio);
+ return false;
+ }
+ folio_put(folio);
+ return err == -EEXIST;
+}
+
+/*
+ * The second writer's data is in the page: mark it so the next writeback
+ * that cleans it drops it. This must follow the write, because a clean
+ * marked folio is dropped by whatever writeback completes next, and it uses
+ * folio_set_dropbehind() rather than an accounted setter, because counting a
+ * folio marked while dirty arms the DONTCACHE writeback kick and the page is
+ * then written back, and dropped, before its partner has written it.
+ */
+static void
+nfsd_write_dio_boundary_complete(struct file *file, loff_t pos)
+{
+ struct address_space *mapping = file->f_mapping;
+ struct folio *folio;
+
+ folio = __filemap_get_folio(mapping, pos >> PAGE_SHIFT, FGP_DONTCACHE, 0);
+ if (IS_ERR(folio))
+ return;
+ folio_set_dropbehind(folio);
+ folio_put(folio);
}
static unsigned int
@@ -1348,14 +1417,17 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
}
/*
- * The prefix and suffix are buffered I/O by definition. Mark them
- * uncached when possible so their folios are dropped once written
- * back rather than lingering in the page cache.
+ * The prefix and suffix are buffered I/O by definition. Each shares
+ * its page with the neighbouring WRITE; see
+ * nfsd_write_dio_boundary_claim(), which nfsd_direct_write() calls right
+ * before issuing each of them, for how the page is held for the
+ * partner and dropped once both have written it.
*/
if (prefix) {
nfsd_write_dio_seg_init(&segments[nsegs], bvec,
nvecs, total, 0, prefix, iocb);
- segments[nsegs++].flags |= dontcache_flags;
+ segments[nsegs].flags |= dontcache_flags;
+ segments[nsegs++].boundary = !!dontcache_flags;
}
nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs,
@@ -1386,16 +1458,23 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
if (suffix) {
nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs, total,
prefix + middle, suffix, iocb);
- segments[nsegs++].flags |= dontcache_flags;
+ segments[nsegs].flags |= dontcache_flags;
+ segments[nsegs++].boundary = !!dontcache_flags;
}
return nsegs;
no_dio:
- /* No DIO possible - pack into a single uncached (if possible) segment. */
+ /*
+ * No DIO possible - pack into a single buffered segment. Where it
+ * does not start or end on a page boundary, its first and last pages
+ * are shared with the neighbouring WRITEs like a prefix or suffix and
+ * are held the same way (nfsd_write_dio_boundary_claim()).
+ */
nfsd_write_dio_seg_init(&segments[0], bvec, nvecs, total, 0,
total, iocb);
segments[0].flags |= dontcache_flags;
+ segments[0].edges = !!dontcache_flags;
return 1;
}
@@ -1423,8 +1502,8 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
struct nfsd_write_dio_seg segments[3];
int floor_iocb_flags = 0;
struct file *file = nf->nf_file;
- loff_t start = kiocb->ki_pos;
- bool sync, datasync;
+ loff_t start = kiocb->ki_pos, seg_pos, seg_last;
+ bool sync, datasync, complete_first, complete_last;
unsigned int nsegs, i;
ssize_t host_err;
size_t expected;
@@ -1461,9 +1540,45 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
expected = iov_iter_count(&segments[i].iter);
+ /*
+ * Claim the boundary page immediately before writing it, not
+ * when the WRITE is split. Nothing is held across the write:
+ * the claim leaves the page in the page cache unmarked, and
+ * claiming it earlier would give the partner WRITE the time a
+ * direct middle takes to complete the page and have it
+ * dropped before this segment writes it.
+ */
+ seg_pos = kiocb->ki_pos;
+ seg_last = seg_pos + expected - 1;
+ complete_first = false;
+ complete_last = false;
+ if (segments[i].boundary) {
+ complete_first = nfsd_write_dio_boundary_claim(file,
+ seg_pos);
+ } else if (segments[i].edges) {
+ /*
+ * A whole-WRITE buffered segment shares its first page
+ * with the previous WRITE if it does not start on a page
+ * boundary, and its last page with the next one if it
+ * does not end on one; a page it covers entirely is its
+ * own.
+ */
+ if (seg_pos & ~PAGE_MASK)
+ complete_first = nfsd_write_dio_boundary_claim(
+ file, seg_pos);
+ if (((seg_last + 1) & ~PAGE_MASK) &&
+ (seg_last >> PAGE_SHIFT) != (seg_pos >> PAGE_SHIFT))
+ complete_last = nfsd_write_dio_boundary_claim(
+ file, seg_last);
+ }
+
host_err = vfs_iocb_iter_write(file, kiocb, &segments[i].iter);
if (host_err < 0)
return host_err;
+ if (complete_first)
+ nfsd_write_dio_boundary_complete(file, seg_pos);
+ if (complete_last)
+ nfsd_write_dio_boundary_complete(file, seg_last);
*cnt += host_err;
if (host_err < (ssize_t)expected)
break; /* partial write */
--
2.52.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH v2 8/9] NFSD: add direct_misaligned_dontcache debugfs knob
2026-09-29 23:13 [PATCH v2 0/9] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
` (6 preceding siblings ...)
2026-09-29 23:13 ` [PATCH v2 7/9] NFSD: keep boundary page of a split direct-mode WRITE until both writers complete Mike Snitzer
@ 2026-09-29 23:13 ` Mike Snitzer
2026-09-29 23:13 ` [PATCH v2 9/9] NFSD: add tracing for how direct-mode READ and WRITE are serviced Mike Snitzer
8 siblings, 0 replies; 14+ messages in thread
From: Mike Snitzer @ 2026-09-29 23:13 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: linux-nfs
From: Jonathan Flynn <jonathan.flynn@hammerspace.com>
The parts of a direct-mode WRITE that cannot be direct I/O, the
misaligned prefix and suffix of a split WRITE and the whole WRITE when
it is not split, are issued IOCB_DONTCACHE so their pages are dropped
once written back. That is the right default for the workloads a direct
mode is chosen for, but it is a policy rather than a requirement: a
workload that reads back what it just wrote, or that keeps rewriting the
same partial pages, is better served by those pages staying in the page
cache.
Add a bool debugfs knob, /sys/kernel/debug/nfsd/direct_misaligned_dontcache
(default Y), to choose between the two:
Y: IOCB_DONTCACHE when the file system supports it, with a split's
boundary page kept in the page cache until both WRITEs sharing it
have written it (nfsd_write_dio_boundary_claim()).
N: ordinary cached buffered I/O; nothing is claimed or marked, and the
pages stay until reclaim.
The direct middle and the once-per-WRITE persist are unaffected. The
knob is sampled once per WRITE, so a change takes effect immediately and
without a remount. It sits beside io_cache_read and io_cache_write, the
other controls over how NFSD issues its I/O.
[snitzer: switched from using modparam to debugfs knob]
Signed-off-by: Jonathan Flynn <jonathan.flynn@hammerspace.com>
[snitzer: documented in nfsd-io-modes.rst]
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
.../filesystems/nfs/nfsd-io-modes.rst | 9 ++++++
fs/nfsd/debugfs.c | 21 ++++++++++++++
fs/nfsd/nfsd.h | 1 +
fs/nfsd/vfs.c | 28 ++++++++++++-------
4 files changed, 49 insertions(+), 10 deletions(-)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index 2ca606ddf5d51..f2c380ca000f7 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -178,6 +178,15 @@ Misaligned WRITE:
with how far concurrent writers drift apart, not with bytes
written.
+ Whether those pages are dropped at all is a policy choice, selected
+ by /sys/kernel/debug/nfsd/direct_misaligned_dontcache (default Y).
+ Write N to issue the start and end segments, and the whole-WRITE
+ fallbacks, as ordinary cached buffered IO: nothing is claimed or
+ marked and the pages stay until reclaim, which suits a workload that
+ reads back or rewrites what it just wrote. The O_DIRECT middle
+ segment is unaffected. The knob is sampled once per WRITE, so a
+ change takes effect immediately.
+
The O_DIRECT middle segment also carries the DONTCACHE flag. It has
no effect while the IO really is O_DIRECT, but a filesystem may
decide on its own to service the segment with buffered IO instead
diff --git a/fs/nfsd/debugfs.c b/fs/nfsd/debugfs.c
index 279e341a81ec1..2b728eac8532c 100644
--- a/fs/nfsd/debugfs.c
+++ b/fs/nfsd/debugfs.c
@@ -144,6 +144,24 @@ void nfsd_debugfs_exit(void)
* Default 2. Not yet tuned by benchmarking.
*/
+/*
+ * /sys/kernel/debug/nfsd/direct_misaligned_dontcache
+ *
+ * How a direct-mode WRITE issues the I/O that cannot be direct: the
+ * misaligned start and end of a split WRITE, and the whole WRITE when it
+ * is not split.
+ *
+ * Contents:
+ * Y: DONTCACHE when the filesystem supports it, with the boundary page
+ * of a split kept in the page cache only until both WRITEs sharing
+ * it have written it
+ * N: ordinary cached buffered IO, left in the page cache until
+ * reclaim, for A/B comparison against the DONTCACHE path
+ *
+ * Sampled once per WRITE, so it takes effect immediately. The direct
+ * middle segment is unaffected.
+ */
+
void nfsd_debugfs_init(void)
{
nfsd_top_dir = debugfs_create_dir("nfsd", NULL);
@@ -159,6 +177,9 @@ void nfsd_debugfs_init(void)
debugfs_create_u32("direct_misaligned_num_pages", 0644, nfsd_top_dir,
&nfsd_direct_misaligned_num_pages);
+
+ debugfs_create_bool("direct_misaligned_dontcache", 0644, nfsd_top_dir,
+ &nfsd_direct_misaligned_dontcache);
#ifdef CONFIG_NFSD_V4
debugfs_create_bool("delegated_timestamps", 0644, nfsd_top_dir,
&nfsd_delegts_enabled);
diff --git a/fs/nfsd/nfsd.h b/fs/nfsd/nfsd.h
index 135e319e378d4..dff979ac370ba 100644
--- a/fs/nfsd/nfsd.h
+++ b/fs/nfsd/nfsd.h
@@ -145,6 +145,7 @@ enum {
extern u64 nfsd_io_cache_read __read_mostly;
extern u64 nfsd_io_cache_write __read_mostly;
extern u32 nfsd_direct_misaligned_num_pages __read_mostly;
+extern bool nfsd_direct_misaligned_dontcache __read_mostly;
extern int nfsd_max_blksize;
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 5fd850a29694f..7896e2e6c5855 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -54,6 +54,7 @@ bool nfsd_disable_splice_read __read_mostly;
u64 nfsd_io_cache_read __read_mostly = NFSD_IO_BUFFERED;
u64 nfsd_io_cache_write __read_mostly = NFSD_IO_BUFFERED;
u32 nfsd_direct_misaligned_num_pages __read_mostly = 2;
+bool nfsd_direct_misaligned_dontcache __read_mostly = true;
/**
* nfserrno - Map Linux errnos to NFS errnos
@@ -1373,16 +1374,21 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
size_t prefix, middle, suffix;
loff_t offset = iocb->ki_pos;
unsigned int dontcache_flags = 0;
+ unsigned int buffered_flags;
unsigned int nsegs = 0;
if (nf->nf_file->f_op->fop_flags & FOP_DONTCACHE)
dontcache_flags = IOCB_DONTCACHE;
+ /* Buffered segments follow the knob; the direct middle does not. */
+ buffered_flags = READ_ONCE(nfsd_direct_misaligned_dontcache) ?
+ dontcache_flags : 0;
/*
* Whenever direct I/O cannot be used for the WRITE, fall back to a
- * single DONTCACHE buffered I/O when the file system supports it, so
- * the WRITE's pages are dropped from the page cache once written
- * back, and to a single cached buffered I/O otherwise.
+ * single DONTCACHE buffered I/O when the file system supports it (and
+ * nfsd_direct_misaligned_dontcache is set), so the WRITE's pages are
+ * dropped from the page cache once written back, and to a single
+ * cached buffered I/O otherwise.
*
* If the file system doesn't advertise any alignment requirements,
* don't try to issue direct I/O at all.
@@ -1421,13 +1427,15 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
* its page with the neighbouring WRITE; see
* nfsd_write_dio_boundary_claim(), which nfsd_direct_write() calls right
* before issuing each of them, for how the page is held for the
- * partner and dropped once both have written it.
+ * partner and dropped once both have written it. With
+ * nfsd_direct_misaligned_dontcache=N both are plain cached writes and
+ * nothing is held or dropped: the pages stay until reclaim.
*/
if (prefix) {
nfsd_write_dio_seg_init(&segments[nsegs], bvec,
nvecs, total, 0, prefix, iocb);
- segments[nsegs].flags |= dontcache_flags;
- segments[nsegs++].boundary = !!dontcache_flags;
+ segments[nsegs].flags |= buffered_flags;
+ segments[nsegs++].boundary = !!buffered_flags;
}
nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs,
@@ -1458,8 +1466,8 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
if (suffix) {
nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs, total,
prefix + middle, suffix, iocb);
- segments[nsegs].flags |= dontcache_flags;
- segments[nsegs++].boundary = !!dontcache_flags;
+ segments[nsegs].flags |= buffered_flags;
+ segments[nsegs++].boundary = !!buffered_flags;
}
return nsegs;
@@ -1473,8 +1481,8 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
*/
nfsd_write_dio_seg_init(&segments[0], bvec, nvecs, total, 0,
total, iocb);
- segments[0].flags |= dontcache_flags;
- segments[0].edges = !!dontcache_flags;
+ segments[0].flags |= buffered_flags;
+ segments[0].edges = !!buffered_flags;
return 1;
}
--
2.52.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH v2 9/9] NFSD: add tracing for how direct-mode READ and WRITE are serviced
2026-09-29 23:13 [PATCH v2 0/9] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
` (7 preceding siblings ...)
2026-09-29 23:13 ` [PATCH v2 8/9] NFSD: add direct_misaligned_dontcache debugfs knob Mike Snitzer
@ 2026-09-29 23:13 ` Mike Snitzer
8 siblings, 0 replies; 14+ messages in thread
From: Mike Snitzer @ 2026-09-29 23:13 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: linux-nfs
When io_cache_read or io_cache_write selects a direct I/O mode, a
request may still be serviced with DONTCACHE buffered or normal
buffered I/O, and a WRITE may be split into up to three segments that
each take a different path depending on the request's offset/length
alignment, the payload's memory alignment, and the filesystem's
FOP_DONTCACHE support. Until now the only visibility was
nfsd_{read,write}_direct for O_DIRECT and nfsd_{read,write}_vector for
everything else, so the DONTCACHE and cached buffered cases were
indistinguishable and the reason a WRITE was not issued as direct I/O
was not recorded at all.
Add:
- nfsd_write_dio_split, emitted once per direct-mode WRITE before any
segment is issued. It records the advertised offset and memory
alignments, the payload's memory offset within its page, the start,
middle and end segment sizes, the number of segments issued, and a
disposition describing which path was taken (direct, no_alignment,
too_small, no_middle, mem_misaligned), with a dontcache flag set
when the buffered segments were issued IOCB_DONTCACHE, so a run with
direct_misaligned_dontcache=N is distinguishable from one with the
file system lacking FOP_DONTCACHE.
- nfsd_write_dontcache and nfsd_read_dontcache, emitted for WRITE
segments and READs serviced with IOCB_DONTCACHE. nfsd_write_vector
and nfsd_read_vector now fire only for normal buffered I/O.
The disposition enum lives in vfs.h so trace.h can reference it, and
TRACE_DEFINE_ENUM entries are provided for user-space decoding.
Document the new events in nfsd-io-modes.rst, including the iomap
iomap_dio_invalidate_fail event that reveals an O_DIRECT segment which
the filesystem silently serviced with buffered I/O.
Assisted-by: Claude:claude-fable-5-1
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
.../filesystems/nfs/nfsd-io-modes.rst | 41 ++++++++-
fs/nfsd/trace.h | 90 +++++++++++++++++++
fs/nfsd/vfs.c | 44 ++++++---
fs/nfsd/vfs.h | 22 +++++
4 files changed, 183 insertions(+), 14 deletions(-)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index f2c380ca000f7..6b1e9e10b47c2 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -211,21 +211,56 @@ Misaligned WRITE:
Tracing:
The nfsd_read_direct trace event shows how NFSD expands any
misaligned READ to the next DIO-aligned block (on either end of the
- original READ, as needed).
+ original READ, as needed). A READ that is serviced with buffered IO
+ instead emits nfsd_read_dontcache (DONTCACHE buffered IO) or
+ nfsd_read_vector (normal buffered IO).
This combination of trace events is useful for READs::
echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_read_vector/enable
+ echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_read_dontcache/enable
echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_read_direct/enable
echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_read_io_done/enable
echo 1 > /sys/kernel/tracing/events/xfs/xfs_file_direct_read/enable
- The nfsd_write_direct trace event shows how NFSD splits a given
- misaligned WRITE into a DIO-aligned middle segment.
+ The nfsd_write_dio_split trace event is emitted once per WRITE
+ serviced in a DIRECT IO mode, before any IO is issued, and records
+ how the WRITE was split: the offset and memory alignments the
+ filesystem advertised, the memory offset of the WRITE payload, the
+ sizes of the start, middle and end segments, the number of segments
+ actually issued, and a disposition naming the reason::
+
+ direct aligned middle segment uses O_DIRECT
+ mem_misaligned payload memory is misaligned; one buffered segment
+ no_alignment filesystem advertises no DIO alignment; one
+ buffered segment
+ too_small WRITE is smaller than the larger of the two
+ alignments; one buffered segment
+ no_middle no (or too small) aligned middle; one buffered
+ segment
+
+ Whether those buffered segments are DONTCACHE or normal buffered IO
+ is reported separately, by dontcache=1 or dontcache=0, because it is
+ the same answer for every disposition: the buffered segments are
+ DONTCACHE when the filesystem supports FOP_DONTCACHE. For the direct
+ disposition it describes the prefix and suffix of the split, the
+ middle being O_DIRECT.
+
+ Each segment then emits one of nfsd_write_direct (O_DIRECT),
+ nfsd_write_dontcache (DONTCACHE buffered IO) or nfsd_write_vector
+ (normal buffered IO) with the segment's offset and length.
This combination of trace events is useful for WRITEs::
echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_write_opened/enable
+ echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_write_dio_split/enable
echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_write_direct/enable
+ echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_write_dontcache/enable
+ echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_write_vector/enable
echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_write_io_done/enable
echo 1 > /sys/kernel/tracing/events/xfs/xfs_file_direct_write/enable
+ echo 1 > /sys/kernel/tracing/events/iomap/iomap_dio_invalidate_fail/enable
+
+ iomap_dio_invalidate_fail indicates an O_DIRECT middle segment that
+ the filesystem silently serviced with normal buffered IO because it
+ could not invalidate overlapping page cache first.
diff --git a/fs/nfsd/trace.h b/fs/nfsd/trace.h
index 2ae7f150a72ce..0b8aa6c5d0772 100644
--- a/fs/nfsd/trace.h
+++ b/fs/nfsd/trace.h
@@ -15,6 +15,7 @@
#include <trace/misc/fsnotify.h>
#include <trace/misc/sunrpc.h>
+#include "vfs.h"
#include "export.h"
#include "nfsfh.h"
#include "xdr4.h"
@@ -501,6 +502,7 @@ DEFINE_EVENT(nfsd_io_class, nfsd_##name, \
DEFINE_NFSD_IO_EVENT(read_start);
DEFINE_NFSD_IO_EVENT(read_splice);
DEFINE_NFSD_IO_EVENT(read_vector);
+DEFINE_NFSD_IO_EVENT(read_dontcache);
DEFINE_NFSD_IO_EVENT(read_direct);
DEFINE_NFSD_IO_EVENT(read_io_done);
DEFINE_NFSD_IO_EVENT(read_done);
@@ -508,11 +510,99 @@ DEFINE_NFSD_IO_EVENT(write_start);
DEFINE_NFSD_IO_EVENT(write_opened);
DEFINE_NFSD_IO_EVENT(write_direct);
DEFINE_NFSD_IO_EVENT(write_vector);
+DEFINE_NFSD_IO_EVENT(write_dontcache);
DEFINE_NFSD_IO_EVENT(write_io_done);
DEFINE_NFSD_IO_EVENT(write_done);
DEFINE_NFSD_IO_EVENT(commit_start);
DEFINE_NFSD_IO_EVENT(commit_done);
+TRACE_DEFINE_ENUM(NFSD_WRITE_DIO_DIRECT);
+TRACE_DEFINE_ENUM(NFSD_WRITE_DIO_MEM_MISALIGNED);
+TRACE_DEFINE_ENUM(NFSD_WRITE_DIO_NO_ALIGN);
+TRACE_DEFINE_ENUM(NFSD_WRITE_DIO_TOO_SMALL);
+TRACE_DEFINE_ENUM(NFSD_WRITE_DIO_NO_MIDDLE);
+
+#define show_nfsd_write_dio_disposition(x) \
+ __print_symbolic(x, \
+ { NFSD_WRITE_DIO_DIRECT, "direct" }, \
+ { NFSD_WRITE_DIO_MEM_MISALIGNED, "mem_misaligned" }, \
+ { NFSD_WRITE_DIO_NO_ALIGN, "no_alignment" }, \
+ { NFSD_WRITE_DIO_TOO_SMALL, "too_small" }, \
+ { NFSD_WRITE_DIO_NO_MIDDLE, "no_middle" })
+
+/**
+ * nfsd_write_dio_split - how an NFSD_IO_DIRECT WRITE was split
+ *
+ * Emitted once per WRITE handled by nfsd_direct_write(), before any
+ * segment is issued. @prefix/@middle/@suffix are the byte counts of the
+ * three candidate segments (zero when not computed); @mem_offset is the
+ * offset within its page of the first byte of the WRITE payload, from
+ * which the middle segment's memory alignment is (@mem_offset + @prefix)
+ * masked by (@mem_align - 1). @dontcache is whether the WRITE's buffered
+ * segments carry IOCB_DONTCACHE: the single segment of every non-direct
+ * disposition, and the prefix and suffix of a "direct" one. It is ORed
+ * into @disposition by the caller, a tracepoint being limited to twelve
+ * arguments, and split back out into its own field here.
+ */
+TRACE_EVENT(nfsd_write_dio_split,
+ TP_PROTO(struct svc_rqst *rqstp,
+ struct svc_fh *fhp,
+ u64 offset,
+ u32 len,
+ u32 offset_align,
+ u32 mem_align,
+ u32 mem_offset,
+ u32 prefix,
+ u32 middle,
+ u32 suffix,
+ u32 nsegs,
+ unsigned int disposition),
+ TP_ARGS(rqstp, fhp, offset, len, offset_align, mem_align, mem_offset,
+ prefix, middle, suffix, nsegs, disposition),
+ TP_STRUCT__entry(
+ __field(u32, xid)
+ __field(u32, fh_hash)
+ __field(u64, offset)
+ __field(u32, len)
+ __field(u32, offset_align)
+ __field(u32, mem_align)
+ __field(u32, mem_offset)
+ __field(u32, prefix)
+ __field(u32, middle)
+ __field(u32, suffix)
+ __field(u32, nsegs)
+ __field(unsigned int, disposition)
+ __field(bool, dontcache)
+ ),
+ TP_fast_assign(
+ __entry->xid = be32_to_cpu(rqstp->rq_xid);
+ __entry->fh_hash = knfsd_fh_hash(&fhp->fh_handle);
+ __entry->offset = offset;
+ __entry->len = len;
+ __entry->offset_align = offset_align;
+ __entry->mem_align = mem_align;
+ __entry->mem_offset = mem_offset;
+ __entry->prefix = prefix;
+ __entry->middle = middle;
+ __entry->suffix = suffix;
+ __entry->nsegs = nsegs;
+ __entry->disposition = disposition & ~NFSD_WRITE_DIO_DONTCACHE;
+ __entry->dontcache = !!(disposition & NFSD_WRITE_DIO_DONTCACHE);
+ ),
+ TP_printk("xid=0x%08x fh_hash=0x%08x offset=%llu len=%u "
+ "offset_align=%u mem_align=%u mem_offset=%u "
+ "prefix=%u middle=%u suffix=%u nsegs=%u disposition=%s "
+ "dontcache=%u",
+ __entry->xid, __entry->fh_hash,
+ __entry->offset, __entry->len,
+ __entry->offset_align, __entry->mem_align,
+ __entry->mem_offset,
+ __entry->prefix, __entry->middle, __entry->suffix,
+ __entry->nsegs,
+ show_nfsd_write_dio_disposition(__entry->disposition),
+ __entry->dontcache)
+);
+
DECLARE_EVENT_CLASS(nfsd_err_class,
TP_PROTO(struct svc_rqst *rqstp,
struct svc_fh *fhp,
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 7896e2e6c5855..3e7265b8de251 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1226,7 +1226,10 @@ __be32 nfsd_iter_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
base = 0;
}
- trace_nfsd_read_vector(rqstp, fhp, offset, *count - total);
+ if (kiocb.ki_flags & IOCB_DONTCACHE)
+ trace_nfsd_read_dontcache(rqstp, fhp, offset, *count - total);
+ else
+ trace_nfsd_read_vector(rqstp, fhp, offset, *count - total);
iov_iter_bvec(&iter, ITER_DEST, rqstp->rq_bvec, v, *count - total);
host_err = vfs_iocb_iter_read(file, &kiocb, &iter);
return nfsd_finish_read(rqstp, fhp, file, offset, count, eof, host_err);
@@ -1363,7 +1366,8 @@ nfsd_write_dio_boundary_complete(struct file *file, loff_t pos)
}
static unsigned int
-nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
+nfsd_write_dio_iters_init(struct svc_rqst *rqstp, struct svc_fh *fhp,
+ struct nfsd_file *nf, struct bio_vec *bvec,
unsigned int nvecs, struct kiocb *iocb,
unsigned long total,
struct nfsd_write_dio_seg segments[3])
@@ -1371,7 +1375,8 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
u32 offset_align = nf->nf_dio_offset_align;
loff_t prefix_end, orig_end, middle_end;
u32 mem_align = nf->nf_dio_mem_align;
- size_t prefix, middle, suffix;
+ size_t prefix = 0, middle = 0, suffix = 0;
+ enum nfsd_write_dio_disposition disposition;
loff_t offset = iocb->ki_pos;
unsigned int dontcache_flags = 0;
unsigned int buffered_flags;
@@ -1393,15 +1398,19 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
* If the file system doesn't advertise any alignment requirements,
* don't try to issue direct I/O at all.
*/
- if (unlikely(!mem_align || !offset_align))
+ if (unlikely(!mem_align || !offset_align)) {
+ disposition = NFSD_WRITE_DIO_NO_ALIGN;
goto no_dio;
+ }
/*
* If the I/O is smaller than the larger of the memory and logical
* offset alignment, no part of it can be direct I/O.
*/
- if (unlikely(total < max(offset_align, mem_align)))
+ if (unlikely(total < max(offset_align, mem_align))) {
+ disposition = NFSD_WRITE_DIO_TOO_SMALL;
goto no_dio;
+ }
prefix_end = round_up(offset, offset_align);
orig_end = offset + total;
@@ -1419,6 +1428,7 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
if (!middle ||
((prefix || suffix) &&
middle < PAGE_SIZE * nfsd_direct_misaligned_num_pages)) {
+ disposition = NFSD_WRITE_DIO_NO_MIDDLE;
goto no_dio;
}
@@ -1452,8 +1462,10 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
* middle, so issue the entire write as a single buffered segment:
* splitting would only turn one buffered write into three.
*/
- if (iov_iter_bvec_offset(&segments[nsegs].iter) & (mem_align - 1))
+ if (iov_iter_bvec_offset(&segments[nsegs].iter) & (mem_align - 1)) {
+ disposition = NFSD_WRITE_DIO_MEM_MISALIGNED;
goto no_dio;
+ }
/*
* Also mark the direct middle DONTCACHE: the file system may fall
* back to buffered I/O on its own (e.g. XFS on -ENOTBLK when it
@@ -1462,6 +1474,7 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
* to do so. On the direct path itself the flag is inert.
*/
segments[nsegs++].flags |= IOCB_DIRECT | dontcache_flags;
+ disposition = NFSD_WRITE_DIO_DIRECT;
if (suffix) {
nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs, total,
@@ -1469,8 +1482,7 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
segments[nsegs].flags |= buffered_flags;
segments[nsegs++].boundary = !!buffered_flags;
}
-
- return nsegs;
+ goto out;
no_dio:
/*
@@ -1483,7 +1495,14 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
total, iocb);
segments[0].flags |= buffered_flags;
segments[0].edges = !!buffered_flags;
- return 1;
+ nsegs = 1;
+out:
+ trace_nfsd_write_dio_split(rqstp, fhp, offset, total,
+ offset_align, mem_align, bvec->bv_offset,
+ prefix, middle, suffix, nsegs,
+ disposition | (buffered_flags ?
+ NFSD_WRITE_DIO_DONTCACHE : 0));
+ return nsegs;
}
/*
@@ -1533,8 +1552,8 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
sync = kiocb->ki_flags & IOCB_DSYNC;
datasync = !(kiocb->ki_flags & IOCB_SYNC);
- nsegs = nfsd_write_dio_iters_init(nf, rqstp->rq_bvec, nvecs,
- kiocb, *cnt, segments);
+ nsegs = nfsd_write_dio_iters_init(rqstp, fhp, nf, rqstp->rq_bvec,
+ nvecs, kiocb, *cnt, segments);
*cnt = 0;
for (i = 0; i < nsegs; i++) {
@@ -1542,6 +1561,9 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
if (kiocb->ki_flags & IOCB_DIRECT)
trace_nfsd_write_direct(rqstp, fhp, kiocb->ki_pos,
segments[i].iter.count);
+ else if (kiocb->ki_flags & IOCB_DONTCACHE)
+ trace_nfsd_write_dontcache(rqstp, fhp, kiocb->ki_pos,
+ segments[i].iter.count);
else
trace_nfsd_write_vector(rqstp, fhp, kiocb->ki_pos,
segments[i].iter.count);
diff --git a/fs/nfsd/vfs.h b/fs/nfsd/vfs.h
index 6b352ca7f02b7..58f9ae26415dc 100644
--- a/fs/nfsd/vfs.h
+++ b/fs/nfsd/vfs.h
@@ -189,4 +189,26 @@ __be32 nfsd_permission(struct svc_cred *cred, struct svc_export *exp,
void nfsd_filp_close(struct file *fp);
+/*
+ * How nfsd_write_dio_iters_init() disposed of an NFSD_IO_DIRECT WRITE.
+ * "DONTCACHE segment" degrades to a cached segment when the file system
+ * lacks FOP_DONTCACHE.
+ */
+/*
+ * Which exit nfsd_write_dio_iters_init() took. Whether the buffered
+ * segments carry IOCB_DONTCACHE is reported separately, by
+ * nfsd_write_dio_split's @dontcache, because it is the same answer for
+ * every exit below.
+ */
+enum nfsd_write_dio_disposition {
+ NFSD_WRITE_DIO_DIRECT, /* aligned middle uses direct I/O */
+ NFSD_WRITE_DIO_MEM_MISALIGNED, /* payload memory misaligned: one segment */
+ NFSD_WRITE_DIO_NO_ALIGN, /* fs advertises no alignment: one segment */
+ NFSD_WRITE_DIO_TOO_SMALL, /* len < max(offset_align, mem_align): one segment */
+ NFSD_WRITE_DIO_NO_MIDDLE, /* no or tiny aligned middle: one segment */
+
+ /* ORed in: the WRITE's buffered segments carry IOCB_DONTCACHE */
+ NFSD_WRITE_DIO_DONTCACHE = 0x80,
+};
+
#endif /* LINUX_NFSD_VFS_H */
--
2.52.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* Re: [PATCH v2 4/9] NFSD: Enable return of an updated stable_how to NFS clients
2026-09-29 23:13 ` [PATCH v2 4/9] NFSD: Enable return of an updated stable_how to NFS clients Mike Snitzer
@ 2026-09-30 21:42 ` Chuck Lever
0 siblings, 0 replies; 14+ messages in thread
From: Chuck Lever @ 2026-09-30 21:42 UTC (permalink / raw)
To: Mike Snitzer; +Cc: Jeff Layton, linux-nfs
On Tue, Sep 29, 2026 at 07:13:24PM -0400, Mike Snitzer wrote:
> From: Chuck Lever <chuck.lever@oracle.com>
>
One note before the review: this patch, as posted, did not apply to
the current cel/nfsd-testing branch without a conflict. The
fs/nfsd/nfsproc.c hunk is against the older nfsd_proc_write(); on
nfsd-testing that function is now the xdrgen-generated variant with
argp->xdrgen.data and a separate count variable. The resolution was
mechanical, but the series needs a rebase before it can be applied.
[ ... ]
> This patch prepares for that change by making the @iocb_flags
> argument of nfsd_write() and nfsd_vfs_write() bi-directional. A
> caller passes in the IOCB_* flags that express the stability its
> client asked for; on return the argument holds the flags that were
> actually satisfied. Each protocol version maps that result back to
> its own on-the-wire value when it encodes the WRITE reply, using the
> new nfsd3_stable_how() and nfsd4_stable_how() helpers, so that NFS
> stable_how values stay out of NFSD's generic VFS API. No behavior
> change is expected.
Is "No behavior change is expected" accurate for exports with the
async option? See the question on the vfs.c hunk below.
Also, "Each protocol version maps that result back to its own
on-the-wire value when it encodes the WRITE reply" does not quite
match the code. NFSv2 passes &iocb_flags and discards the result,
and the NFSv3 and NFSv4 mapping happens in nfsd3_proc_write() and
nfsd4_write() right after the write call, not in the XDR encoders.
Something like:
NFSv3 and NFSv4 convert the result back to their on-the-wire
stable_how with the new nfsd3_stable_how() and nfsd4_stable_how()
helpers. This keeps NFS stable_how values out of NFSD's VFS API.
[ ... ]
> diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
> index 1d2b03cb42963..82eba97656c4e 100644
> --- a/fs/nfsd/vfs.c
> +++ b/fs/nfsd/vfs.c
> @@ -1444,7 +1444,9 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
> * @offset: Byte offset of start
> * @payload: xdr_buf containing the write payload
> * @cnt: IN: number of bytes to write, OUT: number of bytes actually written
> - * @iocb_flags: VFS IOCB_* flags expressing the requested write stability
> + * @iocb_flags: IN: VFS IOCB_* flags expressing the requested write
> + * stability; OUT: the flags actually satisfied, which may be
> + * higher than requested
Nothing in this patch raises the flags, and the only change to the
OUT value below lowers them. It might be clearer for this comment to
describe what the function does now, and let the patch that adds
promotion add "or raised". The same text appears on nfsd_write().
[ ... ]
> @@ -1493,11 +1495,11 @@ nfsd_vfs_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
> exp = fhp->fh_export;
>
> if (!EX_ISSYNC(exp))
> - iocb_flags = 0;
> + *iocb_flags = 0;
Before this patch, the zeroing here hit a local copy, and the reply
echoed the requested stability:
nfsd3_proc_write()
resp->committed = argp->stable;
nfsd4_write()
write->wr_how_written = write->wr_stable_how;
With this change the zero flows back to the caller, so a FILE_SYNC or
DATA_SYNC WRITE to an export with the async option is now answered
UNSTABLE:
nfsd3_proc_write()
resp->committed = nfsd3_stable_how(iocb_flags);
nfsd4_write()
write->wr_how_written = nfsd4_stable_how(iocb_flags);
RFC 1813 section 3.3.7 says that if stable was FILE_SYNC, committed
must also be FILE_SYNC, and anything else is a protocol violation.
Table 20 in RFC 8881 section 18.32.3 allows only FILE_SYNC4 for a
FILE_SYNC4 request. The Linux client takes its "faulty NFS server"
path in nfs_writeback_done() when committed is lower than what it
asked for, and sends a COMMIT it did not send before, which
nfsd_commit() turns into a no-op on an async export.
This is the same problem that stopped the May 2026 posting of this
patch. Jeff withdrew it then because it was not clear that clients
are prepared to send a COMMIT after asking for a stable WRITE:
https://lore.kernel.org/linux-nfs/e13b7b601f0cefda78cf96458e6af6e466cfb5f2.camel@kernel.org/
Nothing in this revision addresses that, so I can't apply the patch
as posted. It changes the reply to every stable WRITE on an async
export, and this series needs it only so that a later patch can
raise the reported stability, never lower it.
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
^ permalink raw reply [flat|nested] 14+ messages in thread
* Re: [PATCH v2 1/9] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well
2026-09-29 23:13 ` [PATCH v2 1/9] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well Mike Snitzer
@ 2026-09-30 21:43 ` Chuck Lever
0 siblings, 0 replies; 14+ messages in thread
From: Chuck Lever @ 2026-09-30 21:43 UTC (permalink / raw)
To: Mike Snitzer; +Cc: Jeff Layton, linux-nfs
Since this one carries a Fixes: tag, the stable folks will pick it
into 7.2.y by itself. The rest of the series won't follow it there.
> diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
> index 0fd6e82478fe..dc50c930f976 100644
> --- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
> +++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
> @@ -130,6 +130,16 @@ Misaligned WRITE:
> segments because using normal buffered IO offers significant RMW
> performance benefit when handling streaming misaligned WRITEs.
>
> + The O_DIRECT middle segment also carries the DONTCACHE flag. It has
In a stable tree this paragraph will sit directly below the existing
"DONTCACHE buffered IO is _not_ used for the misaligned segments"
sentence, which the "also" contradicts. Patch 2/9 removes that
sentence, but 2/9 is not a fix and won't be backported.
Let's replace that sentence itself, so the document reads consistently
in the LTS kernels that carry only this patch.
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
^ permalink raw reply [flat|nested] 14+ messages in thread
* Re: [PATCH v2 2/9] NFSD: only split a direct-mode WRITE for a worthwhile direct middle
2026-09-29 23:13 ` [PATCH v2 2/9] NFSD: only split a direct-mode WRITE for a worthwhile direct middle Mike Snitzer
@ 2026-09-30 21:44 ` Chuck Lever
0 siblings, 0 replies; 14+ messages in thread
From: Chuck Lever @ 2026-09-30 21:44 UTC (permalink / raw)
To: Mike Snitzer; +Cc: Jeff Layton, linux-nfs
commit 26f3b4748d5f5019c9136254f0cef206f60254db
Author: Mike Snitzer <snitzer@kernel.org>
NFSD: only split a direct-mode WRITE for a worthwhile direct middle
This commit adds a direct_misaligned_num_pages debugfs knob and declines
to split a misaligned direct-mode WRITE when its aligned middle is smaller
than that many pages. It also moves the per-segment IOCB flag decision
out of the write loop in nfsd_direct_write() and into
nfsd_write_dio_iters_init().
Link: https://patch.msgid.link/20260929231329.22018-3-snitzer@kernel.org
There are two categories of findings here. First, the commit message and
code comments appear to have some missing information and factual errors.
Second, the patch as written does not fully reflect Christoph's original
suggestion. I'd like the patch to carry some justification of what is
missing or different. Sorry for the length.
> nfsd_write_dio_iters_init() decides how a WRITE issued in a direct mode
> is carried out: split into a buffered prefix, a direct middle and a
> buffered suffix, or issued whole as a single buffered segment.
>
> Only split when the split buys direct I/O for a worthwhile middle. Add
> direct_misaligned_num_pages (debugfs knob, default 2) and decline to
> split when the aligned middle is smaller than that many pages and the
> WRITE has a prefix or a suffix. A WRITE whose payload memory is
> misaligned for the block device already declines: no segment of it can
> be direct I/O, so a split would only turn one buffered write into
> three.
>
> Decide every segment's flags in nfsd_write_dio_iters_init() rather than
> in the write loop. The buffered prefix and suffix of a split WRITE,
> and the whole WRITE when it is not split, are issued IOCB_DONTCACHE
> when the file system supports FOP_DONTCACHE and as plain buffered I/O
> otherwise: the operator chose a direct mode to keep NFSD out of the
> page cache, so a WRITE that cannot be direct gets the next best thing.
>
> Document it in the "Misaligned WRITE" section of nfsd-io-modes.rst.
The first paragraph describes what nfsd_write_dio_iters_init() already
does. What is missing is the cost that makes a small middle not worth a
split. Could the description open with that instead?
The third paragraph reads as if IOCB_DONTCACHE on the prefix, the suffix
and the unsplit WRITE is new. Before this patch, the loop in
nfsd_direct_write() already did this:
else {
trace_nfsd_write_vector(rqstp, fhp, kiocb->ki_pos,
segments[i].iter.count);
/*
* Mark the I/O buffer as evict-able to reduce
* memory contention.
*/
if (nf->nf_file->f_op->fop_flags & FOP_DONTCACHE)
kiocb->ki_flags |= IOCB_DONTCACHE;
}
Every segment without IOCB_DIRECT got IOCB_DONTCACHE when the file system
sets FOP_DONTCACHE. So the flags do not change; only where they are
decided does. Could the paragraph say that, and drop the justification
for behavior that already exists?
Related: the fourth paragraph does not mention that the rst sentence this
patch removes ("DONTCACHE buffered IO is _not_ used for the misaligned
segments ...") was already wrong against that loop. Worth saying so.
> Suggested-by: Christoph Hellwig <hch@infradead.org>
This points at Christoph's whiteboard diff from last November:
https://lore.kernel.org/linux-nfs/aRShjU_Ti7f2Ci7I@infradead.org/
That diff is where the middle-size threshold and the move of the flag
decisions into nfsd_write_dio_iters_init() come from, so the tag is
earned. But that diff also chose cached buffered I/O, on purpose, for
the prefix, the suffix, and a WRITE smaller than its alignment, to leave
their partial pages in place for read-modify-write. Only the case where
the file system advertises no alignment got DONTCACHE. This patch keeps
the opposite policy, the one already in the tree. A reader who follows
the tag finds a diff that disagrees with the code. The description
should say which parts are taken from that diff and that the DONTCACHE
choice is not one of them.
Jonathan's 8/9 is where that choice becomes a setting. Reading 2/9 on
its own, the policy looks fixed; five patches later it turns out to be
a knob that sits beside the one this patch adds. Could 8/9 move up to
sit right after this one? Or, can 2/9 follow 7/9?
As posted it also sets the boundary and edges fields that 7/9 introduces,
so it cannot move as is. But the knob itself needs only the
buffered_flags selection; 7/9 can then add the boundary and edges
assignments when it adds the fields. That puts the two
direct_misaligned_* knobs and their documentation next to each other,
and the description of each can point at the other.
"a split would only turn one buffered write into three": a WRITE with only
a prefix or only a suffix splits into two segments. The same phrase
appears in a vfs.c comment and in the debugfs.c comment below.
> diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
> index dc50c930f976..fef062f24c52 100644
> --- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
> +++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
[ ... ]
> @@ -140,6 +140,17 @@ Misaligned WRITE:
> normal buffered IO. Such fallbacks are visible through the
> iomap_dio_invalidate_fail trace event; see Tracing below.
>
> + Whenever no part of a WRITE can use O_DIRECT, the whole WRITE is
> + issued as a single DONTCACHE buffered IO (normal buffered IO if the
> + filesystem lacks FOP_DONTCACHE). This covers: a filesystem that
> + advertises no DIO alignment requirements at all; a WRITE smaller
> + than the larger of the offset and memory alignments; a WRITE whose
> + DIO-aligned middle segment is smaller than
> + /sys/kernel/debug/nfsd/direct_misaligned_num_pages pages (default 2)
> + while also having a misaligned start or end; and a WRITE whose
> + payload memory is not aligned to the block device's dma_alignment,
> + which rules out O_DIRECT for the middle segment as well.
"Whenever no part of a WRITE can use O_DIRECT" does not describe the third
case in this list. A WRITE whose middle is smaller than
direct_misaligned_num_pages could use O_DIRECT for that middle; NFSD
chooses not to. Could this open with when NFSD does not split, rather
than when it cannot? For example:
If NFSD does not split a misaligned WRITE, it issues the whole
WRITE as a single DONTCACHE buffered IO (normal buffered IO if the
filesystem lacks FOP_DONTCACHE). NFSD does not split a WRITE in
these cases: ...
> diff --git a/fs/nfsd/debugfs.c b/fs/nfsd/debugfs.c
> index 386fd1c54f52..398e400d6403 100644
> --- a/fs/nfsd/debugfs.c
> +++ b/fs/nfsd/debugfs.c
> @@ -128,6 +128,17 @@ void nfsd_debugfs_exit(void)
> nfsd_top_dir = NULL;
> }
>
> +/*
> + * /sys/kernel/debug/nfsd/direct_misaligned_num_pages
> + *
> + * The smallest DIO-aligned middle segment, in pages, that is worth
> + * splitting a misaligned direct-mode WRITE into three segments for. A
> + * WRITE whose middle is smaller than this, and which has a misaligned
> + * start or end, is issued as a single buffered segment instead.
> + *
> + * Default 2. Not yet tuned by benchmarking.
> + */
> +
> void nfsd_debugfs_init(void)
> {
> nfsd_top_dir = debugfs_create_dir("nfsd", NULL);
This comment block sits above nfsd_debugfs_init() with a blank line after
it, attached to nothing. The other knob comments in this file sit
directly above what they describe. Could it move onto the
debugfs_create_u32() call, or onto the variable's definition in vfs.c?
"into three segments" has the same two-or-three problem as the commit
message, and the sentence ends on "for". "Not yet tuned by benchmarking"
is a note about the development process that will go stale in the tree.
Something like:
/*
* /sys/kernel/debug/nfsd/direct_misaligned_num_pages
*
* Minimum size, in pages, of the DIO-aligned middle segment for NFSD
* to split a misaligned direct-mode WRITE. A WRITE with a misaligned
* start or end and a smaller middle is issued as a single buffered
* segment. The default value of this setting is 2.
*/
> diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
> index fa9d9810485d..050a840546fa 100644
> --- a/fs/nfsd/vfs.c
> +++ b/fs/nfsd/vfs.c
[ ... ]
> @@ -1302,15 +1303,29 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
> u32 mem_align = nf->nf_dio_mem_align;
> size_t prefix, middle, suffix;
> loff_t offset = iocb->ki_pos;
> + unsigned int dontcache_flags = 0;
> unsigned int nsegs = 0;
>
> + if (nf->nf_file->f_op->fop_flags & FOP_DONTCACHE)
> + dontcache_flags = IOCB_DONTCACHE;
> +
> /*
> - * Check if direct I/O is feasible for this write request.
> - * If alignments are not available, the write is too small,
> - * or no alignment can be found, fall back to buffered I/O.
> + * Whenever direct I/O cannot be used for the WRITE, fall back to a
> + * single DONTCACHE buffered I/O when the file system supports it, so
> + * the WRITE's pages are dropped from the page cache once written
> + * back, and to a single cached buffered I/O otherwise.
> + *
> + * If the file system doesn't advertise any alignment requirements,
> + * don't try to issue direct I/O at all.
> */
> - if (unlikely(!mem_align || !offset_align) ||
> - unlikely(total < max(offset_align, mem_align)))
> + if (unlikely(!mem_align || !offset_align))
> + goto no_dio;
After this patch, nfsd_write_dio_iters_init() explains what DONTCACHE
does three times: here, above the prefix segment, and at the no_dio
label. Each one says the same thing in different words, and a reader has
to check all three to be sure they agree.
This first one is also in the wrong place. It describes what happens at
no_dio, but it sits above the first of four gotos to that label. And
"Whenever direct I/O cannot be used" is not accurate for the
direct_misaligned_num_pages case below, where direct I/O is possible and
NFSD declines it.
The second paragraph restates the if statement under it.
Could all of this collapse into one comment at the no_dio label?
/*
* Issue the whole WRITE as a single buffered segment, uncached when
* the file system supports FOP_DONTCACHE.
*/
Then nothing needs to be said here, or at most a single line naming the
case.
> +
> + /*
> + * If the I/O is smaller than the larger of the memory and logical
> + * offset alignment, no part of it can be direct I/O.
> + */
> + if (unlikely(total < max(offset_align, mem_align)))
> goto no_dio;
>
> prefix_end = round_up(offset, offset_align);
> @@ -1321,12 +1336,27 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
[ ... ]
> + /*
> + * The prefix and suffix are buffered I/O by definition. Mark them
> + * uncached when possible so their folios are dropped once written
> + * back rather than lingering in the page cache.
> + */
> + if (prefix) {
> + nfsd_write_dio_seg_init(&segments[nsegs], bvec,
> nvecs, total, 0, prefix, iocb);
> + segments[nsegs++].flags |= dontcache_flags;
> + }
"by definition" adds nothing to the first sentence, and the second
sentence is the DONTCACHE explanation again. With the single no_dio
comment above, does this comment add anything over the code under it?
>
> nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs,
> total, prefix, middle, iocb);
> @@ -1337,10 +1367,13 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
> * bvecs generated from RPC receive buffers are contiguous: After
> * the first bvec, all subsequent bvecs start at bv_offset zero
> * (page-aligned). Therefore, only the first bvec is checked.
> + *
> + * If the memory is not aligned, direct I/O is impossible for the
> + * middle, so issue the entire write as a single buffered segment:
> + * splitting would only turn one buffered write into three.
> */
"into three" again; two or three.
Christoph's diff treated this case differently. When the memory is
misaligned but the aligned middle is large enough, it kept the split and
issued the middle as DONTCACHE buffered I/O when the file system supports
it, collapsing to a single write only when it does not. That matters
once 8/9 is applied: with direct_misaligned_dontcache=N, the split is
the only way a large middle stays out of the page cache while the ends
stay in it, and collapsing here makes the whole WRITE cached. Was
dropping that branch deliberate? If it was, the description should say
why.
[ ... ]
> no_dio:
> - /* No DIO alignment possible - pack into single non-DIO segment. */
> + /* No DIO possible - pack into a single uncached (if possible) segment. */
This line runs past 80 columns. The replacement comment suggested above
fits, if the DONTCACHE explanation lands here.
> nfsd_write_dio_seg_init(&segments[0], bvec, nvecs, total, 0,
> total, iocb);
> + segments[0].flags |= dontcache_flags;
> return 1;
> }
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
^ permalink raw reply [flat|nested] 14+ messages in thread
* Re: [PATCH v2 5/9] NFSD: let a direct-mode WRITE raise stable_how and elide the client's COMMIT
2026-09-29 23:13 ` [PATCH v2 5/9] NFSD: let a direct-mode WRITE raise stable_how and elide the client's COMMIT Mike Snitzer
@ 2026-09-30 21:46 ` Chuck Lever
0 siblings, 0 replies; 14+ messages in thread
From: Chuck Lever @ 2026-09-30 21:46 UTC (permalink / raw)
To: Mike Snitzer; +Cc: Jeff Layton, linux-nfs, Trond Myklebust, Anna Schumaker
On 9/29/26 7:13 PM, Mike Snitzer wrote:
> A WRITE serviced in a direct mode is already persistent when NFSD
> replies: the aligned middle is O_DIRECT, and a synchronous WRITE fsyncs
> whatever was buffered before the reply is sent. The reply still says
> UNSTABLE, so the client dutifully sends a COMMIT for data that is
> already on stable storage, and NFSD answers it with an fsync that has
> nothing left to write.
I'm stopping my review here. The premise of this patch is wrong, and
the remaining patches in the series depend on whether we keep this
patch and the previous patch to return a different stable_how to the
client, or drop them.
An O_DIRECT write issued for an UNSTABLE WRITE carries no IOCB_DSYNC.
It bypasses the page cache, but it does not flush the device's write
cache and it does not commit metadata: block allocation, unwritten
extent conversion, and the i_size update all sit in the journal or
in memory until something calls fsync. Under NFSD_IO_DIRECT the data
is not on stable storage when the reply goes out, and the COMMIT's
fsync is the operation that makes it so.
The diff does the right thing to make the raised stable_how truthful:
the floor reaches every segment, and for the direct middle iomap sets
IOMAP_DIO_NEED_SYNC and runs generic_write_sync at completion. But
that means modes 3 and 4 add a flush and journal commit to every
UNSTABLE WRITE, for every client and every export, whether or not a
COMMIT would have followed. A client streaming a large sequential
write to an ordinary export goes from one flush per COMMIT to one
flush per wsize. A per-WRITE O_DSYNC direct write was measured to
cost more than a direct WRITE followed by a COMMIT, and that is why
NFSD's direct UNSTABLE path deliberately leaves IOCB_DSYNC clear.
That is a rule we established when designing the DIRECT WRITE I/O
mode: an UNSTABLE WRITE is never synced by the server on its own
initiative, and a COMMIT is always needed afterwards.
The redundant COMMIT in your flexfiles setup is a client problem, not
a server one. The client starts each flush with FLUSH_COND_STABLE,
meaning "send FILE_SYNC if this flush is a single WRITE", and clears
it whenever pg_moreio is set. In nfs_pageio_add_request, pg_moreio is
set as soon as nfs_pageio_do_add_request declines to coalesce the
next page, and on a pNFS mount the layout's pg_test declines at a
data server boundary. So a write that straddles two DSs sends both
halves UNSTABLE and follows each with a COMMIT, even though each DS
receives exactly one WRITE. The heuristic conflates "more WRITEs in
this flush" with "more WRITEs to this server". If the piece bound for
a given DS fits a single RPC, the client already knows a COMMIT will
follow and should mark that piece FILE_SYNC. Each DS then does the
write and fsync in one round trip under the existing NFSD_IO_DIRECT
mode, and no server knob is needed.
Trond, Anna: does a change in the pg_moreio / FLUSH_COND_STABLE
logic along those lines look reasonable to you?
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
^ permalink raw reply [flat|nested] 14+ messages in thread
end of thread, other threads:[~2026-09-30 21:46 UTC | newest]
Thread overview: 14+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-29 23:13 [PATCH v2 0/9] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
2026-09-29 23:13 ` [PATCH v2 1/9] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well Mike Snitzer
2026-09-30 21:43 ` Chuck Lever
2026-09-29 23:13 ` [PATCH v2 2/9] NFSD: only split a direct-mode WRITE for a worthwhile direct middle Mike Snitzer
2026-09-30 21:44 ` Chuck Lever
2026-09-29 23:13 ` [PATCH v2 3/9] NFSD: do not use direct I/O for a READ smaller than its alignment Mike Snitzer
2026-09-29 23:13 ` [PATCH v2 4/9] NFSD: Enable return of an updated stable_how to NFS clients Mike Snitzer
2026-09-30 21:42 ` Chuck Lever
2026-09-29 23:13 ` [PATCH v2 5/9] NFSD: let a direct-mode WRITE raise stable_how and elide the client's COMMIT Mike Snitzer
2026-09-30 21:46 ` Chuck Lever
2026-09-29 23:13 ` [PATCH v2 6/9] NFSD: persist a synchronous direct-mode WRITE once, after all of its segments Mike Snitzer
2026-09-29 23:13 ` [PATCH v2 7/9] NFSD: keep boundary page of a split direct-mode WRITE until both writers complete Mike Snitzer
2026-09-29 23:13 ` [PATCH v2 8/9] NFSD: add direct_misaligned_dontcache debugfs knob Mike Snitzer
2026-09-29 23:13 ` [PATCH v2 9/9] NFSD: add tracing for how direct-mode READ and WRITE are serviced Mike Snitzer
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox