* [PATCH v3 0/9] NFSD: keep direct-mode I/O out of the page cache
@ 2026-10-01 4:54 Mike Snitzer
2026-10-01 4:54 ` [PATCH v3 1/9] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well Mike Snitzer
` (9 more replies)
0 siblings, 10 replies; 11+ messages in thread
From: Mike Snitzer @ 2026-10-01 4:54 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: hch, linux-nfs
Hi,
This series changes NFSD's direct write path: io_cache_write modes
where an aligned WRITE goes to the filesystem as O_DIRECT, with
DONTCACHE for the parts a misaligned WRITE cannot issue directly.
1/9 marks the direct middle of a split WRITE IOCB_DONTCACHE, for when
XFS re-issues it buffered (-ENOTBLK). It carries Fixes: and rewrites
the stale nfsd-io-modes.rst sentence in place, so a stable tree
carrying only this patch reads consistently.
2/9 splits a misaligned WRITE only for a direct middle of at least
direct_misaligned_num_pages (default 2). 3/9 (Jonathan) adds
direct_misaligned_dontcache, which selects whether the parts that are
not O_DIRECT are DONTCACHE or cached.
4/9 stops expanding a READ smaller than its alignment into a full
aligned device read.
5/9 persists a synchronous direct-mode WRITE once, after all of its
segments, instead of once per segment. 6/9 keeps the page two
misaligned WRITEs share until both have written it: 704 device reads
for 42895 WRITEs where there were 45144 for 45664. 7/9 adds
tracepoints for how each direct-mode READ and WRITE was serviced.
8/9 and 9/9 are separable and come last for that reason. 8/9 is
Chuck's "Enable return of an updated stable_how to NFS clients",
reworked onto @iocb_flags. 9/9 adds two opt-in io_cache_write modes
that persist every direct-mode WRITE before replying and report
NFS_DATA_SYNC or NFS_FILE_SYNC, so the client sends no COMMIT. A
client can get the same result by sending FILE_SYNC itself, so 1-7
stand without them.
The series applies to cel/nfsd-testing (32eb1a60b456) and each patch
builds with W=1 with no warnings.
Changes since v2:
- Rebased onto cel/nfsd-testing; 8/9 now applies to the xdrgen
nfsd_proc_write().
- 1/9: replaces the "DONTCACHE buffered IO is _not_ used" sentence
instead of adding a paragraph that contradicted it.
- 2/9: the description now opens with the cost a small middle adds,
says the IOCB_* flags only move (the write loop already set
IOCB_DONTCACHE), and says which parts of Christoph's diff are taken
and why the DONTCACHE policy and the memory-misaligned collapse are
not. The DONTCACHE explanation is one comment at no_dio; the
debugfs comment sits on the debugfs_create_u32() call; "three
segments" is now "two or three".
- Reordered: direct_misaligned_dontcache moved up to 3/9, beside
direct_misaligned_num_pages, and the boundary-page patch (now 6/9)
adds its own fields. The two stable_how patches moved to the end
(8/9, 9/9), since nothing before them depends on them.
- 8/9: the async export option clears a local copy, so the reply to a
FILE_SYNC or DATA_SYNC WRITE to an async export is never lowered
(RFC 1813 3.3.7). The description says where NFSv3 and NFSv4 map
the result, and no longer says O_DIRECT persists data.
- 9/9: retitled. The description says an UNSTABLE O_DIRECT write is not
durable, that NFSD_IO_DIRECT is unchanged, and that the new modes cost
a flush per WRITE. They are worth it only where each COMMIT covers
about one WRITE.
- Comments and documentation trimmed throughout.
v2: https://lore.kernel.org/linux-nfs/20260929231329.22018-1-snitzer@kernel.org/
v1: https://lore.kernel.org/linux-nfs/20260929173423.16149-1-snitzer@kernel.org/
Thanks,
Mike
Chuck Lever (1):
NFSD: Enable return of an updated stable_how to NFS clients
Jonathan Flynn (1):
NFSD: add direct_misaligned_dontcache debugfs knob
Mike Snitzer (7):
NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well
NFSD: only split a direct-mode WRITE for a worthwhile direct middle
NFSD: do not use direct I/O for a READ smaller than its alignment
NFSD: persist a synchronous direct-mode WRITE once, after all of its segments
NFSD: keep boundary page of a split direct-mode WRITE until both writers complete
NFSD: add tracing for how direct-mode READ and WRITE are serviced
NFSD: add direct-mode WRITE settings that persist each WRITE
.../filesystems/nfs/nfsd-io-modes.rst | 97 ++++++-
fs/nfsd/debugfs.c | 28 ++
fs/nfsd/nfs3proc.c | 16 +-
fs/nfsd/nfs4proc.c | 15 +-
fs/nfsd/nfsd.h | 4 +
fs/nfsd/nfsproc.c | 3 +-
fs/nfsd/trace.h | 85 +++++++
fs/nfsd/vfs.c | 239 +++++++++++++++---
fs/nfsd/vfs.h | 18 +-
fs/nfsd/xdr3.h | 2 +-
10 files changed, 449 insertions(+), 58 deletions(-)
base-commit: 32eb1a60b456980761cf7a9cee8f907fdc08afb8
--
2.52.0
^ permalink raw reply [flat|nested] 11+ messages in thread
* [PATCH v3 1/9] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well
2026-10-01 4:54 [PATCH v3 0/9] NFSD: keep direct-mode I/O out of the page cache Mike Snitzer
@ 2026-10-01 4:54 ` Mike Snitzer
2026-10-01 4:54 ` [PATCH v3 2/9] NFSD: only split a direct-mode WRITE for a worthwhile direct middle Mike Snitzer
` (8 subsequent siblings)
9 siblings, 0 replies; 11+ messages in thread
From: Mike Snitzer @ 2026-10-01 4:54 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: hch, linux-nfs
nfsd_write_dio_iters_init() issues the DIO-aligned middle of a
misaligned WRITE with IOCB_DIRECT. A file system may service that
request with buffered I/O instead: xfs_file_write_iter() retries
through xfs_file_buffered_write() with the same kiocb when
xfs_file_dio_write() returns -ENOTBLK, which iomap_dio_rw() does when
it cannot invalidate page cache overlapping the range.
In a streaming misaligned WRITE workload with more than one request in
flight this happens routinely. Request N's direct middle ends in the
page where request N+1's buffered prefix starts. If N+1 dirties that
page between N's writeback and N's invalidation, the invalidation
fails and XFS re-issues all of N's middle, tens of kilobytes, as a
cached buffered write that nothing ever drops.
Set IOCB_DONTCACHE on the middle segment alongside IOCB_DIRECT when
the file system supports FOP_DONTCACHE. The direct path ignores the
flag; a buffered fallback now drops its pages once written back.
nfsd-io-modes.rst said that DONTCACHE is not used for the misaligned
segments, which the write loop already contradicted: every segment
without IOCB_DIRECT gets IOCB_DONTCACHE when the file system supports
it. Replace that sentence with one that covers all three segments.
Fixes: 06c5c97293e3 ("NFSD: Implement NFSD_IO_DIRECT for NFS WRITE")
Assisted-by: Claude:claude-fable-5-1
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
Documentation/filesystems/nfs/nfsd-io-modes.rst | 8 +++++---
fs/nfsd/vfs.c | 3 +++
2 files changed, 8 insertions(+), 3 deletions(-)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index 0fd6e82478fe6..60b0af9b7e49f 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -126,9 +126,11 @@ Misaligned WRITE:
middle and end as needed. The large middle segment is DIO-aligned
and the start and/or end are misaligned. Buffered IO is used for the
misaligned segments and O_DIRECT is used for the middle DIO-aligned
- segment. DONTCACHE buffered IO is _not_ used for the misaligned
- segments because using normal buffered IO offers significant RMW
- performance benefit when handling streaming misaligned WRITEs.
+ segment. If the filesystem supports FOP_DONTCACHE, every segment is
+ marked DONTCACHE. The flag has no effect on the O_DIRECT segment
+ unless the filesystem services it with buffered IO instead, as XFS
+ does when it cannot invalidate page cache that overlaps the segment.
+ The iomap_dio_invalidate_fail trace event reports such a fallback.
Tracing:
The nfsd_read_direct trace event shows how NFSD expands any
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 4584d5b94feed..1ccdac4693745 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1341,6 +1341,9 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
if (iov_iter_bvec_offset(&segments[nsegs].iter) & (mem_align - 1))
goto no_dio;
segments[nsegs].flags |= IOCB_DIRECT;
+ /* In case the file system falls back to buffered I/O (-ENOTBLK). */
+ if (nf->nf_file->f_op->fop_flags & FOP_DONTCACHE)
+ segments[nsegs].flags |= IOCB_DONTCACHE;
nsegs++;
if (suffix)
base-commit: 32eb1a60b456980761cf7a9cee8f907fdc08afb8
--
2.52.0
^ permalink raw reply related [flat|nested] 11+ messages in thread
* [PATCH v3 2/9] NFSD: only split a direct-mode WRITE for a worthwhile direct middle
2026-10-01 4:54 [PATCH v3 0/9] NFSD: keep direct-mode I/O out of the page cache Mike Snitzer
2026-10-01 4:54 ` [PATCH v3 1/9] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well Mike Snitzer
@ 2026-10-01 4:54 ` Mike Snitzer
2026-10-01 4:54 ` [PATCH v3 3/9] NFSD: add direct_misaligned_dontcache debugfs knob Mike Snitzer
` (7 subsequent siblings)
9 siblings, 0 replies; 11+ messages in thread
From: Mike Snitzer @ 2026-10-01 4:54 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: hch, linux-nfs
Splitting a misaligned direct-mode WRITE issues it as two or three
writes instead of one, and the direct write first writes back and
invalidates the page cache over its range. When the DIO-aligned middle
is only a page or so, the WRITE is mostly its buffered prefix and
suffix anyway, and the direct middle adds that overhead plus the
invalidation that contends with a neighbouring WRITE's buffered segment
in the page they share.
Add direct_misaligned_num_pages (debugfs, default 2) and do not split a
WRITE that has a misaligned start or end and a middle smaller than that
many pages. Such a WRITE is issued as a single buffered segment, like
the other WRITEs NFSD does not split.
Also decide each segment's IOCB_* flags in nfsd_write_dio_iters_init()
instead of in the write loop of nfsd_direct_write(). The flags do not
change: the loop already set IOCB_DONTCACHE on every segment without
IOCB_DIRECT when the file system supports FOP_DONTCACHE.
The threshold and the move of the flag decisions come from Christoph's
suggested diff [1]. Two of its choices are not taken. It issued the
prefix, the suffix and a WRITE smaller than its alignment as cached
buffered I/O, to leave partial pages in place for read-modify-write;
this patch keeps the DONTCACHE policy already in the tree, and the next
patch makes that policy selectable. It also split a WRITE whose
payload is misaligned in memory and issued the middle as DONTCACHE
buffered I/O; no part of such a WRITE can be direct I/O, so it is still
issued whole and treated like every other WRITE that cannot be.
Document in nfsd-io-modes.rst when NFSD does not split a WRITE.
Link: https://lore.kernel.org/linux-nfs/aRShjU_Ti7f2Ci7I@infradead.org/ [1]
Suggested-by: Christoph Hellwig <hch@infradead.org>
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
.../filesystems/nfs/nfsd-io-modes.rst | 13 ++++++
fs/nfsd/debugfs.c | 11 +++++
fs/nfsd/nfsd.h | 1 +
fs/nfsd/vfs.c | 41 +++++++++++--------
4 files changed, 48 insertions(+), 18 deletions(-)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index 60b0af9b7e49f..0a67ce9244d38 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -132,6 +132,19 @@ Misaligned WRITE:
does when it cannot invalidate page cache that overlaps the segment.
The iomap_dio_invalidate_fail trace event reports such a fallback.
+ If NFSD does not split a misaligned WRITE, it issues the whole WRITE
+ as a single DONTCACHE buffered IO (normal buffered IO if the
+ filesystem lacks FOP_DONTCACHE). NFSD does not split a WRITE when:
+
+ - the filesystem advertises no DIO alignment;
+ - the WRITE is smaller than the larger of the offset and memory
+ alignments;
+ - the WRITE has a misaligned start or end and its DIO-aligned middle
+ is smaller than /sys/kernel/debug/nfsd/direct_misaligned_num_pages
+ pages (default 2);
+ - the WRITE payload is not aligned in memory to the block device's
+ dma_alignment, so the middle cannot be O_DIRECT either.
+
Tracing:
The nfsd_read_direct trace event shows how NFSD expands any
misaligned READ to the next DIO-aligned block (on either end of the
diff --git a/fs/nfsd/debugfs.c b/fs/nfsd/debugfs.c
index 386fd1c54f527..0b3ddf28d2849 100644
--- a/fs/nfsd/debugfs.c
+++ b/fs/nfsd/debugfs.c
@@ -140,6 +140,17 @@ void nfsd_debugfs_init(void)
debugfs_create_file("io_cache_write", 0644, nfsd_top_dir, NULL,
&nfsd_io_cache_write_fops);
+
+ /*
+ * /sys/kernel/debug/nfsd/direct_misaligned_num_pages
+ *
+ * Minimum size, in pages, of the DIO-aligned middle segment for NFSD
+ * to split a misaligned direct-mode WRITE. A WRITE with a misaligned
+ * start or end and a smaller middle is issued as a single buffered
+ * segment. The default value of this setting is 2.
+ */
+ debugfs_create_u32("direct_misaligned_num_pages", 0644, nfsd_top_dir,
+ &nfsd_direct_misaligned_num_pages);
#ifdef CONFIG_NFSD_V4
debugfs_create_bool("delegated_timestamps", 0644, nfsd_top_dir,
&nfsd_delegts_enabled);
diff --git a/fs/nfsd/nfsd.h b/fs/nfsd/nfsd.h
index 8e8763ae7035c..70219d26b7404 100644
--- a/fs/nfsd/nfsd.h
+++ b/fs/nfsd/nfsd.h
@@ -145,6 +145,7 @@ enum {
extern u64 nfsd_io_cache_read __read_mostly;
extern u64 nfsd_io_cache_write __read_mostly;
+extern u32 nfsd_direct_misaligned_num_pages __read_mostly;
bool nfsd_v4client(struct svc_rqst *rqstp);
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 1ccdac4693745..0b39f53e041e1 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -53,6 +53,7 @@
bool nfsd_disable_splice_read __read_mostly;
u64 nfsd_io_cache_read __read_mostly = NFSD_IO_BUFFERED;
u64 nfsd_io_cache_write __read_mostly = NFSD_IO_BUFFERED;
+u32 nfsd_direct_misaligned_num_pages __read_mostly = 2;
/**
* nfserrno - Map Linux errnos to NFS errnos
@@ -1302,8 +1303,12 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
u32 mem_align = nf->nf_dio_mem_align;
size_t prefix, middle, suffix;
loff_t offset = iocb->ki_pos;
+ unsigned int dontcache_flags = 0;
unsigned int nsegs = 0;
+ if (nf->nf_file->f_op->fop_flags & FOP_DONTCACHE)
+ dontcache_flags = IOCB_DONTCACHE;
+
/*
* Check if direct I/O is feasible for this write request.
* If alignments are not available, the write is too small,
@@ -1321,12 +1326,16 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
middle = middle_end - prefix_end;
suffix = orig_end - middle_end;
- if (!middle)
+ if (!middle ||
+ ((prefix || suffix) &&
+ middle < PAGE_SIZE * nfsd_direct_misaligned_num_pages))
goto no_dio;
- if (prefix)
- nfsd_write_dio_seg_init(&segments[nsegs++], bvec,
+ if (prefix) {
+ nfsd_write_dio_seg_init(&segments[nsegs], bvec,
nvecs, total, 0, prefix, iocb);
+ segments[nsegs++].flags |= dontcache_flags;
+ }
nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs,
total, prefix, middle, iocb);
@@ -1340,22 +1349,25 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
*/
if (iov_iter_bvec_offset(&segments[nsegs].iter) & (mem_align - 1))
goto no_dio;
- segments[nsegs].flags |= IOCB_DIRECT;
/* In case the file system falls back to buffered I/O (-ENOTBLK). */
- if (nf->nf_file->f_op->fop_flags & FOP_DONTCACHE)
- segments[nsegs].flags |= IOCB_DONTCACHE;
- nsegs++;
+ segments[nsegs++].flags |= IOCB_DIRECT | dontcache_flags;
- if (suffix)
- nfsd_write_dio_seg_init(&segments[nsegs++], bvec, nvecs, total,
+ if (suffix) {
+ nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs, total,
prefix + middle, suffix, iocb);
+ segments[nsegs++].flags |= dontcache_flags;
+ }
return nsegs;
no_dio:
- /* No DIO alignment possible - pack into single non-DIO segment. */
+ /*
+ * Issue the whole WRITE as a single buffered segment, uncached when
+ * the file system supports FOP_DONTCACHE.
+ */
nfsd_write_dio_seg_init(&segments[0], bvec, nvecs, total, 0,
total, iocb);
+ segments[0].flags |= dontcache_flags;
return 1;
}
@@ -1379,16 +1391,9 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
if (kiocb->ki_flags & IOCB_DIRECT)
trace_nfsd_write_direct(rqstp, fhp, kiocb->ki_pos,
segments[i].iter.count);
- else {
+ else
trace_nfsd_write_vector(rqstp, fhp, kiocb->ki_pos,
segments[i].iter.count);
- /*
- * Mark the I/O buffer as evict-able to reduce
- * memory contention.
- */
- if (nf->nf_file->f_op->fop_flags & FOP_DONTCACHE)
- kiocb->ki_flags |= IOCB_DONTCACHE;
- }
expected = iov_iter_count(&segments[i].iter);
--
2.52.0
^ permalink raw reply related [flat|nested] 11+ messages in thread
* [PATCH v3 3/9] NFSD: add direct_misaligned_dontcache debugfs knob
2026-10-01 4:54 [PATCH v3 0/9] NFSD: keep direct-mode I/O out of the page cache Mike Snitzer
2026-10-01 4:54 ` [PATCH v3 1/9] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well Mike Snitzer
2026-10-01 4:54 ` [PATCH v3 2/9] NFSD: only split a direct-mode WRITE for a worthwhile direct middle Mike Snitzer
@ 2026-10-01 4:54 ` Mike Snitzer
2026-10-01 4:54 ` [PATCH v3 4/9] NFSD: do not use direct I/O for a READ smaller than its alignment Mike Snitzer
` (6 subsequent siblings)
9 siblings, 0 replies; 11+ messages in thread
From: Mike Snitzer @ 2026-10-01 4:54 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: hch, linux-nfs
From: Jonathan Flynn <jonathan.flynn@hammerspace.com>
The parts of a direct-mode WRITE that are not O_DIRECT, the misaligned
start and end of a split WRITE and a WRITE that is not split, are
issued IOCB_DONTCACHE when the file system supports it. That suits
the workloads a direct mode is chosen for, but it is a policy, not a
requirement: a workload that reads back or rewrites what it just wrote
is better served by those pages staying cached.
Add /sys/kernel/debug/nfsd/direct_misaligned_dontcache (default Y).
Set to N, those parts are issued as normal buffered I/O. The O_DIRECT
middle keeps IOCB_DONTCACHE either way. It sits beside
direct_misaligned_num_pages, which decides how much of a misaligned
WRITE is O_DIRECT; this knob decides how the rest is cached.
Signed-off-by: Jonathan Flynn <jonathan.flynn@hammerspace.com>
[snitzer: switched from a module parameter to a debugfs knob, moved
next to direct_misaligned_num_pages, documented in nfsd-io-modes.rst]
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
.../filesystems/nfs/nfsd-io-modes.rst | 18 +++++++++++++-----
fs/nfsd/debugfs.c | 12 ++++++++++++
fs/nfsd/nfsd.h | 1 +
fs/nfsd/vfs.c | 13 +++++++++----
4 files changed, 35 insertions(+), 9 deletions(-)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index 0a67ce9244d38..9b1a9e7b09cef 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -126,11 +126,13 @@ Misaligned WRITE:
middle and end as needed. The large middle segment is DIO-aligned
and the start and/or end are misaligned. Buffered IO is used for the
misaligned segments and O_DIRECT is used for the middle DIO-aligned
- segment. If the filesystem supports FOP_DONTCACHE, every segment is
- marked DONTCACHE. The flag has no effect on the O_DIRECT segment
- unless the filesystem services it with buffered IO instead, as XFS
- does when it cannot invalidate page cache that overlaps the segment.
- The iomap_dio_invalidate_fail trace event reports such a fallback.
+ segment. If the filesystem supports FOP_DONTCACHE, the O_DIRECT
+ segment is marked DONTCACHE, and so are the misaligned segments
+ unless direct_misaligned_dontcache is N (below). The flag has no
+ effect on the O_DIRECT segment unless the filesystem services it
+ with buffered IO instead, as XFS does when it cannot invalidate page
+ cache that overlaps the segment. The iomap_dio_invalidate_fail trace
+ event reports such a fallback.
If NFSD does not split a misaligned WRITE, it issues the whole WRITE
as a single DONTCACHE buffered IO (normal buffered IO if the
@@ -145,6 +147,12 @@ Misaligned WRITE:
- the WRITE payload is not aligned in memory to the block device's
dma_alignment, so the middle cannot be O_DIRECT either.
+ Writing N to /sys/kernel/debug/nfsd/direct_misaligned_dontcache
+ (default Y) issues the start and end segments, and a WRITE that is
+ not split, as normal buffered IO instead of DONTCACHE, which suits a
+ workload that reads back or rewrites what it just wrote. The O_DIRECT
+ middle segment keeps its DONTCACHE flag either way.
+
Tracing:
The nfsd_read_direct trace event shows how NFSD expands any
misaligned READ to the next DIO-aligned block (on either end of the
diff --git a/fs/nfsd/debugfs.c b/fs/nfsd/debugfs.c
index 0b3ddf28d2849..d5b714dd0fbc6 100644
--- a/fs/nfsd/debugfs.c
+++ b/fs/nfsd/debugfs.c
@@ -151,6 +151,18 @@ void nfsd_debugfs_init(void)
*/
debugfs_create_u32("direct_misaligned_num_pages", 0644, nfsd_top_dir,
&nfsd_direct_misaligned_num_pages);
+
+ /*
+ * /sys/kernel/debug/nfsd/direct_misaligned_dontcache
+ *
+ * Y: the parts of a direct-mode WRITE that are not O_DIRECT use
+ * DONTCACHE buffered IO when the file system supports it
+ * N: those parts use normal buffered IO
+ *
+ * The default value of this setting is Y.
+ */
+ debugfs_create_bool("direct_misaligned_dontcache", 0644, nfsd_top_dir,
+ &nfsd_direct_misaligned_dontcache);
#ifdef CONFIG_NFSD_V4
debugfs_create_bool("delegated_timestamps", 0644, nfsd_top_dir,
&nfsd_delegts_enabled);
diff --git a/fs/nfsd/nfsd.h b/fs/nfsd/nfsd.h
index 70219d26b7404..1c17f96012523 100644
--- a/fs/nfsd/nfsd.h
+++ b/fs/nfsd/nfsd.h
@@ -146,6 +146,7 @@ enum {
extern u64 nfsd_io_cache_read __read_mostly;
extern u64 nfsd_io_cache_write __read_mostly;
extern u32 nfsd_direct_misaligned_num_pages __read_mostly;
+extern bool nfsd_direct_misaligned_dontcache __read_mostly;
bool nfsd_v4client(struct svc_rqst *rqstp);
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 0b39f53e041e1..1304953c684b7 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -54,6 +54,7 @@ bool nfsd_disable_splice_read __read_mostly;
u64 nfsd_io_cache_read __read_mostly = NFSD_IO_BUFFERED;
u64 nfsd_io_cache_write __read_mostly = NFSD_IO_BUFFERED;
u32 nfsd_direct_misaligned_num_pages __read_mostly = 2;
+bool nfsd_direct_misaligned_dontcache __read_mostly = true;
/**
* nfserrno - Map Linux errnos to NFS errnos
@@ -1304,10 +1305,13 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
size_t prefix, middle, suffix;
loff_t offset = iocb->ki_pos;
unsigned int dontcache_flags = 0;
+ unsigned int buffered_flags;
unsigned int nsegs = 0;
if (nf->nf_file->f_op->fop_flags & FOP_DONTCACHE)
dontcache_flags = IOCB_DONTCACHE;
+ buffered_flags = READ_ONCE(nfsd_direct_misaligned_dontcache) ?
+ dontcache_flags : 0;
/*
* Check if direct I/O is feasible for this write request.
@@ -1334,7 +1338,7 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
if (prefix) {
nfsd_write_dio_seg_init(&segments[nsegs], bvec,
nvecs, total, 0, prefix, iocb);
- segments[nsegs++].flags |= dontcache_flags;
+ segments[nsegs++].flags |= buffered_flags;
}
nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs,
@@ -1355,7 +1359,7 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
if (suffix) {
nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs, total,
prefix + middle, suffix, iocb);
- segments[nsegs++].flags |= dontcache_flags;
+ segments[nsegs++].flags |= buffered_flags;
}
return nsegs;
@@ -1363,11 +1367,12 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
no_dio:
/*
* Issue the whole WRITE as a single buffered segment, uncached when
- * the file system supports FOP_DONTCACHE.
+ * the file system supports FOP_DONTCACHE and
+ * direct_misaligned_dontcache is set.
*/
nfsd_write_dio_seg_init(&segments[0], bvec, nvecs, total, 0,
total, iocb);
- segments[0].flags |= dontcache_flags;
+ segments[0].flags |= buffered_flags;
return 1;
}
--
2.52.0
^ permalink raw reply related [flat|nested] 11+ messages in thread
* [PATCH v3 4/9] NFSD: do not use direct I/O for a READ smaller than its alignment
2026-10-01 4:54 [PATCH v3 0/9] NFSD: keep direct-mode I/O out of the page cache Mike Snitzer
` (2 preceding siblings ...)
2026-10-01 4:54 ` [PATCH v3 3/9] NFSD: add direct_misaligned_dontcache debugfs knob Mike Snitzer
@ 2026-10-01 4:54 ` Mike Snitzer
2026-10-01 4:54 ` [PATCH v3 5/9] NFSD: persist a synchronous direct-mode WRITE once, after all of its segments Mike Snitzer
` (5 subsequent siblings)
9 siblings, 0 replies; 11+ messages in thread
From: Mike Snitzer @ 2026-10-01 4:54 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: hch, linux-nfs
nfsd_direct_read() expands a misaligned READ out to DIO-aligned
boundaries: it reads from round_down(offset, dio_read_offset_align) to
round_up(offset + count, dio_read_offset_align) and returns only the
requested bytes from within that window. When the READ is smaller than
the alignment, that window is always at least one full alignment unit,
and two when the READ straddles a boundary, so a few hundred bytes of
payload can cost a 4K or 64K device read.
Decline direct I/O for those. A READ smaller than
dio_read_offset_align now falls through to the DONTCACHE path, which
issues DONTCACHE buffered I/O when the file system supports
FOP_DONTCACHE and normal buffered I/O otherwise. This mirrors the
WRITE side, which already declines direct I/O for a WRITE smaller than
the larger of its offset and memory alignments.
Only dio_read_offset_align is consulted, because the READ path fills
page-aligned pages from rq_bvec and so has no memory alignment to
satisfy. The threshold only bites when the file system advertises a
large alignment; where it reports 512 almost no READ is excluded.
Document it in the "Misaligned READ" section of nfsd-io-modes.rst.
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
Documentation/filesystems/nfs/nfsd-io-modes.rst | 5 +++++
fs/nfsd/vfs.c | 6 +++---
2 files changed, 8 insertions(+), 3 deletions(-)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index 9b1a9e7b09cef..12679001c7bec 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -121,6 +121,11 @@ Misaligned READ:
verified to have proper offset/len (logical_block_size) and
dma_alignment checking.
+ A READ smaller than dio_read_offset_align is not expanded: it would
+ read one or two whole alignment units to return fewer bytes than
+ one. It is issued as DONTCACHE buffered IO instead (normal buffered
+ IO if the filesystem lacks FOP_DONTCACHE).
+
Misaligned WRITE:
If NFSD_IO_DIRECT is used, split any misaligned WRITE into a start,
middle and end as needed. The large middle segment is DIO-aligned
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 1304953c684b7..e1d294aceb6bd 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1187,7 +1187,7 @@ __be32 nfsd_iter_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
unsigned int base, u32 *eof)
{
struct file *file = nf->nf_file;
- unsigned long v, total;
+ unsigned long v, total = *count;
struct iov_iter iter;
struct kiocb kiocb;
ssize_t host_err;
@@ -1200,7 +1200,8 @@ __be32 nfsd_iter_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
break;
case NFSD_IO_DIRECT:
/* When dio_read_offset_align is zero, dio is not supported */
- if (nf->nf_dio_read_offset_align && !rqstp->rq_res.page_len)
+ if (nf->nf_dio_read_offset_align && !rqstp->rq_res.page_len &&
+ total >= nf->nf_dio_read_offset_align)
return nfsd_direct_read(rqstp, fhp, nf, offset,
count, eof);
fallthrough;
@@ -1213,7 +1214,6 @@ __be32 nfsd_iter_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
kiocb.ki_pos = offset;
v = 0;
- total = *count;
while (total && v < rqstp->rq_maxpages &&
rqstp->rq_next_page < rqstp->rq_page_end) {
len = min_t(size_t, total, PAGE_SIZE - base);
--
2.52.0
^ permalink raw reply related [flat|nested] 11+ messages in thread
* [PATCH v3 5/9] NFSD: persist a synchronous direct-mode WRITE once, after all of its segments
2026-10-01 4:54 [PATCH v3 0/9] NFSD: keep direct-mode I/O out of the page cache Mike Snitzer
` (3 preceding siblings ...)
2026-10-01 4:54 ` [PATCH v3 4/9] NFSD: do not use direct I/O for a READ smaller than its alignment Mike Snitzer
@ 2026-10-01 4:54 ` Mike Snitzer
2026-10-01 4:54 ` [PATCH v3 6/9] NFSD: keep boundary page of a split direct-mode WRITE until both writers complete Mike Snitzer
` (4 subsequent siblings)
9 siblings, 0 replies; 11+ messages in thread
From: Mike Snitzer @ 2026-10-01 4:54 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: hch, linux-nfs
nfsd_direct_write() may issue a WRITE as up to three segments: a
buffered prefix, a direct middle and a buffered suffix. For a
FILE_SYNC or DATA_SYNC WRITE the kiocb carries IOCB_DSYNC and every
segment inherits it, so generic_write_sync() runs a range
fsync after each segment: up to three cache flushes and log forces per
WRITE, and each one writes back and drops the boundary page it just
touched.
Strip IOCB_DSYNC and IOCB_SYNC from the per-segment flags and persist
the WRITE once with vfs_fsync_range() over the bytes actually written,
after the last segment. Durability is unchanged: the reply is not sent
until the fsync completes, and datasync mirrors the previous per-segment
choice (IOCB_SYNC present means metadata too). An fsync failure is
returned like a write failure.
Besides the fewer flushes, this puts the sync under NFSD's control,
which a later commit uses to keep the boundary pages of a split WRITE
cached until the partner WRITE completes them.
Assisted-by: Claude:claude-fable-5-1
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
Documentation/filesystems/nfs/nfsd-io-modes.rst | 3 +++
fs/nfsd/vfs.c | 15 ++++++++++++++-
2 files changed, 17 insertions(+), 1 deletion(-)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index 12679001c7bec..686b106b96f1b 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -152,6 +152,9 @@ Misaligned WRITE:
- the WRITE payload is not aligned in memory to the block device's
dma_alignment, so the middle cannot be O_DIRECT either.
+ A FILE_SYNC or DATA_SYNC WRITE is persisted once after all of its
+ segments are written, not once per segment.
+
Writing N to /sys/kernel/debug/nfsd/direct_misaligned_dontcache
(default Y) issues the start and end segments, and a WRITE that is
not split, as normal buffered IO instead of DONTCACHE, which suits a
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index e1d294aceb6bd..592415900403f 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1383,16 +1383,22 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
{
struct nfsd_write_dio_seg segments[3];
struct file *file = nf->nf_file;
+ loff_t start = kiocb->ki_pos;
+ bool sync, datasync;
unsigned int nsegs, i;
ssize_t host_err;
size_t expected;
+ /* Persist a synchronous WRITE once, after all of its segments. */
+ sync = kiocb->ki_flags & IOCB_DSYNC;
+ datasync = !(kiocb->ki_flags & IOCB_SYNC);
+
nsegs = nfsd_write_dio_iters_init(nf, rqstp->rq_bvec, nvecs,
kiocb, *cnt, segments);
*cnt = 0;
for (i = 0; i < nsegs; i++) {
- kiocb->ki_flags = segments[i].flags;
+ kiocb->ki_flags = segments[i].flags & ~(IOCB_DSYNC | IOCB_SYNC);
if (kiocb->ki_flags & IOCB_DIRECT)
trace_nfsd_write_direct(rqstp, fhp, kiocb->ki_pos,
segments[i].iter.count);
@@ -1410,6 +1416,13 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
break; /* partial write */
}
+ if (sync && *cnt) {
+ host_err = vfs_fsync_range(file, start, start + *cnt - 1,
+ datasync);
+ if (host_err < 0)
+ return host_err;
+ }
+
return 0;
}
--
2.52.0
^ permalink raw reply related [flat|nested] 11+ messages in thread
* [PATCH v3 6/9] NFSD: keep boundary page of a split direct-mode WRITE until both writers complete
2026-10-01 4:54 [PATCH v3 0/9] NFSD: keep direct-mode I/O out of the page cache Mike Snitzer
` (4 preceding siblings ...)
2026-10-01 4:54 ` [PATCH v3 5/9] NFSD: persist a synchronous direct-mode WRITE once, after all of its segments Mike Snitzer
@ 2026-10-01 4:54 ` Mike Snitzer
2026-10-01 4:55 ` [PATCH v3 7/9] NFSD: add tracing for how direct-mode READ and WRITE are serviced Mike Snitzer
` (3 subsequent siblings)
9 siblings, 0 replies; 11+ messages in thread
From: Mike Snitzer @ 2026-10-01 4:54 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: hch, linux-nfs
A direct-mode WRITE whose ends are not aligned for direct I/O is split
into a buffered prefix, a direct middle and a buffered suffix, and the
prefix and suffix are issued IOCB_DONTCACHE so their pages are dropped
once written back. The page holding such a segment is shared by two
WRITEs, the one ending in it and the one starting in it, which may
arrive in either order, from different clients, at the same time. If
it is dropped as soon as the first writer's data is written back, the
second has to read the block from disk before it can complete it: one
4 KiB read per WRITE. With many clients interleaving small misaligned
O_DIRECT records into one shared file, that was 0.81 device reads per
record (585 GiB read while writing 8.1 TiB).
Keep the page until both writers have had it, with no state in NFSD.
Immediately before a boundary segment is written,
nfsd_write_dio_boundary_claim() puts an empty folio in the page cache
if there is not one there already. The folio is inserted unmarked: a
DONTCACHE write only marks a folio it allocated itself, so neither
writer marks this one and it is never counted in WB_DONTCACHE_DIRTY,
which is what the DONTCACHE writeback kick targets.
The add is atomic, so of two concurrent partners exactly one gets the
folio. The other finds it (-EEXIST), is therefore the second of the
two, and once its data is in the page
nfsd_write_dio_boundary_complete() marks the folio dropbehind. The
mark must follow the write, because a clean marked folio is dropped by
whatever writeback completes next, and must not be accounted, because
a counted folio is handed straight back to the kick. Whichever
writeback then cleans the page drops it: the WRITE's own sync for
FILE_SYNC and DATA_SYNC, the flusher or the client's COMMIT for
UNSTABLE.
A WRITE that is not split is issued as one buffered DONTCACHE segment.
Its first and last pages are shared with the neighbouring WRITEs in the
same way, and are claimed the same way. Nothing is claimed when
direct_misaligned_dontcache is N, since those pages are then cached.
Measured with 32 interleaved writers over an emulated 4Kn NVMe,
direct_misaligned_dontcache=Y and io_cache_write=4 (the
NFSD_IO_DIRECT_WRITE_FILE_SYNC mode added in a later commit), each arm
starting from a freshly made file system. 47008-byte records, which
split: 704 device reads for 42895 WRITEs, against 45144 for 45664 WRITEs
without this patch. 6000-byte records, which are not split: 385 reads
against 196589. The DONTCACHE flusher stays idle in both, and 81 of the
written file's 524066 pages are still resident afterwards.
On a four-server pNFS flexfiles rig, three clients writing 47008-byte
records for 240 s against 640 nfsd threads per server: device reads
fall from 0.014-0.015 per record to 0.003, page-cache growth over the
write phase falls from about 11 GiB to 2.9 GiB, and write throughput
rises by 2.4% to 4.9% across NFSD_IO_DIRECT and the two modes added
in a later commit, with read throughput unchanged.
Keeping every boundary page cached instead, with
direct_misaligned_dontcache=N, grows the page cache by about 630 GiB
over the same runs.
Reported-by: Jonathan Flynn <jonathan.flynn@hammerspace.com>
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
.../filesystems/nfs/nfsd-io-modes.rst | 11 +++
fs/nfsd/vfs.c | 87 ++++++++++++++++++-
2 files changed, 94 insertions(+), 4 deletions(-)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index 686b106b96f1b..683e3a81d4de1 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -155,6 +155,17 @@ Misaligned WRITE:
A FILE_SYNC or DATA_SYNC WRITE is persisted once after all of its
segments are written, not once per segment.
+ The page holding a start or end segment is shared with the
+ neighbouring WRITE, which may arrive before, after or at the same
+ time, from any client. So that the second WRITE to such a page does
+ not have to read it back from disk, NFSD puts an empty page there
+ immediately before writing a start or end segment, if there is not
+ one already, without marking it DONTCACHE. The WRITE that finds the
+ page already there marks it DONTCACHE once it has written it, and
+ the next writeback drops it. The first and last pages of a WRITE
+ that is not split are handled the same way. None of this applies
+ when direct_misaligned_dontcache is N.
+
Writing N to /sys/kernel/debug/nfsd/direct_misaligned_dontcache
(default Y) issues the start and end segments, and a WRITE that is
not split, as normal buffered IO instead of DONTCACHE, which suits a
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 592415900403f..5a54ef664da43 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1272,6 +1272,8 @@ static int wait_for_concurrent_writes(struct file *file)
struct nfsd_write_dio_seg {
struct iov_iter iter;
int flags;
+ bool boundary; /* prefix or suffix */
+ bool edges; /* unsplit WRITE */
};
static unsigned long
@@ -1291,6 +1293,54 @@ nfsd_write_dio_seg_init(struct nfsd_write_dio_seg *segment,
iov_iter_advance(&segment->iter, start);
iov_iter_truncate(&segment->iter, len);
segment->flags = iocb->ki_flags;
+ segment->boundary = false;
+ segment->edges = false;
+}
+
+/*
+ * The page holding a misaligned start or end of a WRITE is shared with the
+ * neighbouring WRITE. Put an unmarked folio there if there is none: a
+ * DONTCACHE write marks only a folio it allocates, so neither WRITE marks
+ * this one and no writeback drops it before the other WRITE has written it.
+ *
+ * Return: true if a folio was already there. This WRITE is then the
+ * second of the two and calls nfsd_write_dio_boundary_complete() after
+ * writing the page.
+ */
+static bool
+nfsd_write_dio_boundary_claim(struct file *file, loff_t pos)
+{
+ struct address_space *mapping = file->f_mapping;
+ gfp_t gfp = mapping_gfp_mask(mapping);
+ struct folio *folio;
+ int err;
+
+ folio = filemap_alloc_folio(gfp, 0, NULL);
+ if (!folio)
+ return false;
+ err = filemap_add_folio(mapping, folio, pos >> PAGE_SHIFT, gfp);
+ if (!err)
+ folio_unlock(folio);
+ folio_put(folio);
+ return err == -EEXIST;
+}
+
+/*
+ * Both WRITEs have written the page: mark it so the writeback that cleans
+ * it drops it. folio_set_dropbehind() does not count the folio in
+ * WB_DONTCACHE_DIRTY.
+ */
+static void
+nfsd_write_dio_boundary_complete(struct file *file, loff_t pos)
+{
+ struct folio *folio;
+
+ folio = __filemap_get_folio(file->f_mapping, pos >> PAGE_SHIFT,
+ FGP_DONTCACHE, 0);
+ if (IS_ERR(folio))
+ return;
+ folio_set_dropbehind(folio);
+ folio_put(folio);
}
static unsigned int
@@ -1338,7 +1388,8 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
if (prefix) {
nfsd_write_dio_seg_init(&segments[nsegs], bvec,
nvecs, total, 0, prefix, iocb);
- segments[nsegs++].flags |= buffered_flags;
+ segments[nsegs].flags |= buffered_flags;
+ segments[nsegs++].boundary = !!buffered_flags;
}
nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs,
@@ -1359,7 +1410,8 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
if (suffix) {
nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs, total,
prefix + middle, suffix, iocb);
- segments[nsegs++].flags |= buffered_flags;
+ segments[nsegs].flags |= buffered_flags;
+ segments[nsegs++].boundary = !!buffered_flags;
}
return nsegs;
@@ -1373,6 +1425,7 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
nfsd_write_dio_seg_init(&segments[0], bvec, nvecs, total, 0,
total, iocb);
segments[0].flags |= buffered_flags;
+ segments[0].edges = !!buffered_flags;
return 1;
}
@@ -1383,8 +1436,8 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
{
struct nfsd_write_dio_seg segments[3];
struct file *file = nf->nf_file;
- loff_t start = kiocb->ki_pos;
- bool sync, datasync;
+ loff_t start = kiocb->ki_pos, seg_pos, seg_last;
+ bool sync, datasync, complete_first, complete_last;
unsigned int nsegs, i;
ssize_t host_err;
size_t expected;
@@ -1408,9 +1461,35 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
expected = iov_iter_count(&segments[i].iter);
+ /*
+ * Claim just before the write: an earlier claim gives the
+ * partner WRITE time to complete the page and drop it first.
+ */
+ seg_pos = kiocb->ki_pos;
+ seg_last = seg_pos + expected - 1;
+ complete_first = false;
+ complete_last = false;
+ if (segments[i].boundary) {
+ complete_first = nfsd_write_dio_boundary_claim(file,
+ seg_pos);
+ } else if (segments[i].edges) {
+ /* A page the segment covers entirely is not shared. */
+ if (seg_pos & ~PAGE_MASK)
+ complete_first = nfsd_write_dio_boundary_claim(
+ file, seg_pos);
+ if (((seg_last + 1) & ~PAGE_MASK) &&
+ (seg_last >> PAGE_SHIFT) != (seg_pos >> PAGE_SHIFT))
+ complete_last = nfsd_write_dio_boundary_claim(
+ file, seg_last);
+ }
+
host_err = vfs_iocb_iter_write(file, kiocb, &segments[i].iter);
if (host_err < 0)
return host_err;
+ if (complete_first)
+ nfsd_write_dio_boundary_complete(file, seg_pos);
+ if (complete_last)
+ nfsd_write_dio_boundary_complete(file, seg_last);
*cnt += host_err;
if (host_err < (ssize_t)expected)
break; /* partial write */
--
2.52.0
^ permalink raw reply related [flat|nested] 11+ messages in thread
* [PATCH v3 7/9] NFSD: add tracing for how direct-mode READ and WRITE are serviced
2026-10-01 4:54 [PATCH v3 0/9] NFSD: keep direct-mode I/O out of the page cache Mike Snitzer
` (5 preceding siblings ...)
2026-10-01 4:54 ` [PATCH v3 6/9] NFSD: keep boundary page of a split direct-mode WRITE until both writers complete Mike Snitzer
@ 2026-10-01 4:55 ` Mike Snitzer
2026-10-01 4:55 ` [PATCH v3 8/9] NFSD: Enable return of an updated stable_how to NFS clients Mike Snitzer
` (2 subsequent siblings)
9 siblings, 0 replies; 11+ messages in thread
From: Mike Snitzer @ 2026-10-01 4:55 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: hch, linux-nfs
When io_cache_read or io_cache_write selects a direct I/O mode, a
request may still be serviced with DONTCACHE buffered or normal
buffered I/O, and a WRITE may be split into up to three segments that
each take a different path depending on the request's offset/length
alignment, the payload's memory alignment, the filesystem's
FOP_DONTCACHE support and direct_misaligned_dontcache. Until now the only visibility was
nfsd_{read,write}_direct for O_DIRECT and nfsd_{read,write}_vector for
everything else, so the DONTCACHE and cached buffered cases were
indistinguishable and the reason a WRITE was not issued as direct I/O
was not recorded at all.
Add:
- nfsd_write_dio_split, emitted once per direct-mode WRITE before any
segment is issued. It records the advertised offset and memory
alignments, the payload's offset within its first page, the start,
middle and end segment sizes, the number of segments issued, a
disposition naming the path taken (direct, no_alignment, too_small,
no_middle, mem_misaligned), and whether the buffered segments were
issued IOCB_DONTCACHE.
- nfsd_write_dontcache and nfsd_read_dontcache, emitted for WRITE
segments and READs serviced with IOCB_DONTCACHE. nfsd_write_vector
and nfsd_read_vector now fire only for normal buffered I/O.
The disposition enum lives in vfs.h so trace.h can reference it.
Document the new events in nfsd-io-modes.rst, including the iomap
iomap_dio_invalidate_fail event that reveals an O_DIRECT segment which
the filesystem silently serviced with buffered I/O.
Assisted-by: Claude:claude-fable-5-1
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
.../filesystems/nfs/nfsd-io-modes.rst | 34 +++++++-
fs/nfsd/trace.h | 85 +++++++++++++++++++
fs/nfsd/vfs.c | 48 ++++++++---
fs/nfsd/vfs.h | 14 +++
4 files changed, 166 insertions(+), 15 deletions(-)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index 683e3a81d4de1..11938caee14ca 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -175,21 +175,49 @@ Misaligned WRITE:
Tracing:
The nfsd_read_direct trace event shows how NFSD expands any
misaligned READ to the next DIO-aligned block (on either end of the
- original READ, as needed).
+ original READ, as needed). A READ that is serviced with buffered IO
+ instead emits nfsd_read_dontcache (DONTCACHE buffered IO) or
+ nfsd_read_vector (normal buffered IO).
This combination of trace events is useful for READs::
echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_read_vector/enable
+ echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_read_dontcache/enable
echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_read_direct/enable
echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_read_io_done/enable
echo 1 > /sys/kernel/tracing/events/xfs/xfs_file_direct_read/enable
- The nfsd_write_direct trace event shows how NFSD splits a given
- misaligned WRITE into a DIO-aligned middle segment.
+ The nfsd_write_dio_split trace event is emitted once per WRITE in a
+ DIRECT IO mode, before any IO is issued. It records the alignments
+ the filesystem advertised, the payload's offset in its first page,
+ the sizes of the start, middle and end segments, the number of
+ segments issued, and a disposition::
+
+ direct the aligned middle segment uses O_DIRECT
+ mem_misaligned payload memory is misaligned; not split
+ no_alignment the filesystem advertises no DIO alignment; not split
+ too_small smaller than the larger alignment; not split
+ no_middle no aligned middle, or one smaller than
+ direct_misaligned_num_pages; not split
+
+ dontcache=1 means the buffered segments were DONTCACHE: the start and
+ end segments of a direct WRITE, or the whole of one that was not
+ split.
+
+ Each segment then emits nfsd_write_direct (O_DIRECT),
+ nfsd_write_dontcache (DONTCACHE buffered IO) or nfsd_write_vector
+ (normal buffered IO).
This combination of trace events is useful for WRITEs::
echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_write_opened/enable
+ echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_write_dio_split/enable
echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_write_direct/enable
+ echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_write_dontcache/enable
+ echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_write_vector/enable
echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_write_io_done/enable
echo 1 > /sys/kernel/tracing/events/xfs/xfs_file_direct_write/enable
+ echo 1 > /sys/kernel/tracing/events/iomap/iomap_dio_invalidate_fail/enable
+
+ iomap_dio_invalidate_fail reports an O_DIRECT middle segment that
+ the filesystem serviced with buffered IO instead.
diff --git a/fs/nfsd/trace.h b/fs/nfsd/trace.h
index ad106d627fe72..89b1962dacfb1 100644
--- a/fs/nfsd/trace.h
+++ b/fs/nfsd/trace.h
@@ -15,6 +15,7 @@
#include <trace/misc/fsnotify.h>
#include <trace/misc/sunrpc.h>
+#include "vfs.h"
#include "export.h"
#include "nfsfh.h"
#include "xdr4.h"
@@ -501,6 +502,7 @@ DEFINE_EVENT(nfsd_io_class, nfsd_##name, \
DEFINE_NFSD_IO_EVENT(read_start);
DEFINE_NFSD_IO_EVENT(read_splice);
DEFINE_NFSD_IO_EVENT(read_vector);
+DEFINE_NFSD_IO_EVENT(read_dontcache);
DEFINE_NFSD_IO_EVENT(read_direct);
DEFINE_NFSD_IO_EVENT(read_io_done);
DEFINE_NFSD_IO_EVENT(read_done);
@@ -508,11 +510,94 @@ DEFINE_NFSD_IO_EVENT(write_start);
DEFINE_NFSD_IO_EVENT(write_opened);
DEFINE_NFSD_IO_EVENT(write_direct);
DEFINE_NFSD_IO_EVENT(write_vector);
+DEFINE_NFSD_IO_EVENT(write_dontcache);
DEFINE_NFSD_IO_EVENT(write_io_done);
DEFINE_NFSD_IO_EVENT(write_done);
DEFINE_NFSD_IO_EVENT(commit_start);
DEFINE_NFSD_IO_EVENT(commit_done);
+TRACE_DEFINE_ENUM(NFSD_WRITE_DIO_DIRECT);
+TRACE_DEFINE_ENUM(NFSD_WRITE_DIO_MEM_MISALIGNED);
+TRACE_DEFINE_ENUM(NFSD_WRITE_DIO_NO_ALIGN);
+TRACE_DEFINE_ENUM(NFSD_WRITE_DIO_TOO_SMALL);
+TRACE_DEFINE_ENUM(NFSD_WRITE_DIO_NO_MIDDLE);
+
+#define show_nfsd_write_dio_disposition(x) \
+ __print_symbolic(x, \
+ { NFSD_WRITE_DIO_DIRECT, "direct" }, \
+ { NFSD_WRITE_DIO_MEM_MISALIGNED, "mem_misaligned" }, \
+ { NFSD_WRITE_DIO_NO_ALIGN, "no_alignment" }, \
+ { NFSD_WRITE_DIO_TOO_SMALL, "too_small" }, \
+ { NFSD_WRITE_DIO_NO_MIDDLE, "no_middle" })
+
+/**
+ * nfsd_write_dio_split - how a direct-mode WRITE was split
+ *
+ * Emitted once per direct-mode WRITE, before any segment is issued.
+ * @prefix, @middle and @suffix are zero when not computed. @mem_offset
+ * is the offset of the payload within its first page. The DONTCACHE
+ * bit arrives in @disposition because a tracepoint takes at most twelve
+ * arguments.
+ */
+TRACE_EVENT(nfsd_write_dio_split,
+ TP_PROTO(struct svc_rqst *rqstp,
+ struct svc_fh *fhp,
+ u64 offset,
+ u32 len,
+ u32 offset_align,
+ u32 mem_align,
+ u32 mem_offset,
+ u32 prefix,
+ u32 middle,
+ u32 suffix,
+ u32 nsegs,
+ unsigned int disposition),
+ TP_ARGS(rqstp, fhp, offset, len, offset_align, mem_align, mem_offset,
+ prefix, middle, suffix, nsegs, disposition),
+ TP_STRUCT__entry(
+ __field(u32, xid)
+ __field(u32, fh_hash)
+ __field(u64, offset)
+ __field(u32, len)
+ __field(u32, offset_align)
+ __field(u32, mem_align)
+ __field(u32, mem_offset)
+ __field(u32, prefix)
+ __field(u32, middle)
+ __field(u32, suffix)
+ __field(u32, nsegs)
+ __field(unsigned int, disposition)
+ __field(bool, dontcache)
+ ),
+ TP_fast_assign(
+ __entry->xid = be32_to_cpu(rqstp->rq_xid);
+ __entry->fh_hash = knfsd_fh_hash(&fhp->fh_handle);
+ __entry->offset = offset;
+ __entry->len = len;
+ __entry->offset_align = offset_align;
+ __entry->mem_align = mem_align;
+ __entry->mem_offset = mem_offset;
+ __entry->prefix = prefix;
+ __entry->middle = middle;
+ __entry->suffix = suffix;
+ __entry->nsegs = nsegs;
+ __entry->disposition = disposition & ~NFSD_WRITE_DIO_DONTCACHE;
+ __entry->dontcache = !!(disposition & NFSD_WRITE_DIO_DONTCACHE);
+ ),
+ TP_printk("xid=0x%08x fh_hash=0x%08x offset=%llu len=%u "
+ "offset_align=%u mem_align=%u mem_offset=%u "
+ "prefix=%u middle=%u suffix=%u nsegs=%u disposition=%s "
+ "dontcache=%u",
+ __entry->xid, __entry->fh_hash,
+ __entry->offset, __entry->len,
+ __entry->offset_align, __entry->mem_align,
+ __entry->mem_offset,
+ __entry->prefix, __entry->middle, __entry->suffix,
+ __entry->nsegs,
+ show_nfsd_write_dio_disposition(__entry->disposition),
+ __entry->dontcache)
+);
+
DECLARE_EVENT_CLASS(nfsd_err_class,
TP_PROTO(struct svc_rqst *rqstp,
struct svc_fh *fhp,
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 5a54ef664da43..ec52c67c4380d 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1226,7 +1226,10 @@ __be32 nfsd_iter_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
base = 0;
}
- trace_nfsd_read_vector(rqstp, fhp, offset, *count - total);
+ if (kiocb.ki_flags & IOCB_DONTCACHE)
+ trace_nfsd_read_dontcache(rqstp, fhp, offset, *count - total);
+ else
+ trace_nfsd_read_vector(rqstp, fhp, offset, *count - total);
iov_iter_bvec(&iter, ITER_DEST, rqstp->rq_bvec, v, *count - total);
host_err = vfs_iocb_iter_read(file, &kiocb, &iter);
return nfsd_finish_read(rqstp, fhp, file, offset, count, eof, host_err);
@@ -1344,7 +1347,8 @@ nfsd_write_dio_boundary_complete(struct file *file, loff_t pos)
}
static unsigned int
-nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
+nfsd_write_dio_iters_init(struct svc_rqst *rqstp, struct svc_fh *fhp,
+ struct nfsd_file *nf, struct bio_vec *bvec,
unsigned int nvecs, struct kiocb *iocb,
unsigned long total,
struct nfsd_write_dio_seg segments[3])
@@ -1352,7 +1356,8 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
u32 offset_align = nf->nf_dio_offset_align;
loff_t prefix_end, orig_end, middle_end;
u32 mem_align = nf->nf_dio_mem_align;
- size_t prefix, middle, suffix;
+ size_t prefix = 0, middle = 0, suffix = 0;
+ enum nfsd_write_dio_disposition disposition;
loff_t offset = iocb->ki_pos;
unsigned int dontcache_flags = 0;
unsigned int buffered_flags;
@@ -1368,9 +1373,14 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
* If alignments are not available, the write is too small,
* or no alignment can be found, fall back to buffered I/O.
*/
- if (unlikely(!mem_align || !offset_align) ||
- unlikely(total < max(offset_align, mem_align)))
+ if (unlikely(!mem_align || !offset_align)) {
+ disposition = NFSD_WRITE_DIO_NO_ALIGN;
goto no_dio;
+ }
+ if (unlikely(total < max(offset_align, mem_align))) {
+ disposition = NFSD_WRITE_DIO_TOO_SMALL;
+ goto no_dio;
+ }
prefix_end = round_up(offset, offset_align);
orig_end = offset + total;
@@ -1382,8 +1392,10 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
if (!middle ||
((prefix || suffix) &&
- middle < PAGE_SIZE * nfsd_direct_misaligned_num_pages))
+ middle < PAGE_SIZE * nfsd_direct_misaligned_num_pages)) {
+ disposition = NFSD_WRITE_DIO_NO_MIDDLE;
goto no_dio;
+ }
if (prefix) {
nfsd_write_dio_seg_init(&segments[nsegs], bvec,
@@ -1402,10 +1414,13 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
* the first bvec, all subsequent bvecs start at bv_offset zero
* (page-aligned). Therefore, only the first bvec is checked.
*/
- if (iov_iter_bvec_offset(&segments[nsegs].iter) & (mem_align - 1))
+ if (iov_iter_bvec_offset(&segments[nsegs].iter) & (mem_align - 1)) {
+ disposition = NFSD_WRITE_DIO_MEM_MISALIGNED;
goto no_dio;
+ }
/* In case the file system falls back to buffered I/O (-ENOTBLK). */
segments[nsegs++].flags |= IOCB_DIRECT | dontcache_flags;
+ disposition = NFSD_WRITE_DIO_DIRECT;
if (suffix) {
nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs, total,
@@ -1413,8 +1428,7 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
segments[nsegs].flags |= buffered_flags;
segments[nsegs++].boundary = !!buffered_flags;
}
-
- return nsegs;
+ goto out;
no_dio:
/*
@@ -1426,7 +1440,14 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
total, iocb);
segments[0].flags |= buffered_flags;
segments[0].edges = !!buffered_flags;
- return 1;
+ nsegs = 1;
+out:
+ trace_nfsd_write_dio_split(rqstp, fhp, offset, total,
+ offset_align, mem_align, bvec->bv_offset,
+ prefix, middle, suffix, nsegs,
+ disposition | (buffered_flags ?
+ NFSD_WRITE_DIO_DONTCACHE : 0));
+ return nsegs;
}
static noinline_for_stack int
@@ -1446,8 +1467,8 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
sync = kiocb->ki_flags & IOCB_DSYNC;
datasync = !(kiocb->ki_flags & IOCB_SYNC);
- nsegs = nfsd_write_dio_iters_init(nf, rqstp->rq_bvec, nvecs,
- kiocb, *cnt, segments);
+ nsegs = nfsd_write_dio_iters_init(rqstp, fhp, nf, rqstp->rq_bvec,
+ nvecs, kiocb, *cnt, segments);
*cnt = 0;
for (i = 0; i < nsegs; i++) {
@@ -1455,6 +1476,9 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
if (kiocb->ki_flags & IOCB_DIRECT)
trace_nfsd_write_direct(rqstp, fhp, kiocb->ki_pos,
segments[i].iter.count);
+ else if (kiocb->ki_flags & IOCB_DONTCACHE)
+ trace_nfsd_write_dontcache(rqstp, fhp, kiocb->ki_pos,
+ segments[i].iter.count);
else
trace_nfsd_write_vector(rqstp, fhp, kiocb->ki_pos,
segments[i].iter.count);
diff --git a/fs/nfsd/vfs.h b/fs/nfsd/vfs.h
index 38f7d36bd4da2..4b352a48a8af2 100644
--- a/fs/nfsd/vfs.h
+++ b/fs/nfsd/vfs.h
@@ -229,4 +229,18 @@ __be32 nfsd_permission(struct svc_cred *cred, struct svc_export *exp,
void nfsd_filp_close(struct file *fp);
+/*
+ * The path nfsd_write_dio_iters_init() took for a direct-mode WRITE.
+ * NFSD_WRITE_DIO_DONTCACHE is ORed in when its buffered segments carry
+ * IOCB_DONTCACHE.
+ */
+enum nfsd_write_dio_disposition {
+ NFSD_WRITE_DIO_DIRECT,
+ NFSD_WRITE_DIO_MEM_MISALIGNED,
+ NFSD_WRITE_DIO_NO_ALIGN,
+ NFSD_WRITE_DIO_TOO_SMALL,
+ NFSD_WRITE_DIO_NO_MIDDLE,
+ NFSD_WRITE_DIO_DONTCACHE = 0x80,
+};
+
#endif /* LINUX_NFSD_VFS_H */
--
2.52.0
^ permalink raw reply related [flat|nested] 11+ messages in thread
* [PATCH v3 8/9] NFSD: Enable return of an updated stable_how to NFS clients
2026-10-01 4:54 [PATCH v3 0/9] NFSD: keep direct-mode I/O out of the page cache Mike Snitzer
` (6 preceding siblings ...)
2026-10-01 4:55 ` [PATCH v3 7/9] NFSD: add tracing for how direct-mode READ and WRITE are serviced Mike Snitzer
@ 2026-10-01 4:55 ` Mike Snitzer
2026-10-01 4:55 ` [PATCH v3 9/9] NFSD: add direct-mode WRITE settings that persist each WRITE Mike Snitzer
2026-10-01 22:58 ` [PATCH v3 0/9] NFSD: keep direct-mode I/O out of the page cache Mike Snitzer
9 siblings, 0 replies; 11+ messages in thread
From: Mike Snitzer @ 2026-10-01 4:55 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: hch, linux-nfs
From: Chuck Lever <chuck.lever@oracle.com>
NFSv3 and newer protocols enable clients to perform a two-phase
WRITE. A client requests an UNSTABLE WRITE, which sends dirty data
to the NFS server, but does not persist it. The server replies that
it performed the UNSTABLE WRITE, and the client is then obligated to
follow up with a COMMIT request before it can remove the dirty data
from its own page cache. The COMMIT reply is the client's guarantee
that the written data has been persisted on the server.
The purpose of this protocol design is to enable clients to send
a large amount of data via multiple WRITE requests to a server, and
then wait for persistence just once. The server is able to start
persisting the data as soon as it gets it, to shorten the length of
time the client has to wait for the final COMMIT to complete.
It's also possible for the server to respond to an UNSTABLE WRITE
request in a way that indicates that the data was persisted anyway.
In that case, the client can skip the COMMIT and remove the dirty
data from its memory immediately. NetApp filers, for example, do
this because they have a battery-backed cache and can guarantee that
written data is persisted quickly and immediately.
NFSD has never implemented this kind of promotion. UNSTABLE WRITE
requests are unconditionally treated as UNSTABLE. However, in a
subsequent patch, nfsd_vfs_write() will be able to promote an
UNSTABLE WRITE to be a FILE_SYNC WRITE, when NFSD is configured to
persist each WRITE it services with O_DIRECT before replying. The
FILE_SYNC WRITE response indicates to the client that no follow-up
COMMIT is necessary.
This patch prepares for that change by making the @iocb_flags
argument of nfsd_write() and nfsd_vfs_write() bi-directional. A
caller passes in the IOCB_* flags that express the stability its
client asked for; on return the argument holds the stability to
report. NFSv3 and NFSv4 convert the result back to their on-the-wire
stable_how with the new nfsd3_stable_how() and nfsd4_stable_how()
helpers. This keeps NFS stable_how values out of NFSD's VFS API.
NFSv2 has no stable_how in its reply and ignores the result.
No behavior change is expected. In particular, a WRITE to an export
with the async option is still issued without IOCB_DSYNC and still
reported at the stability the client requested.
[snitzer: reworked onto the @iocb_flags argument introduced by
commit 6dcddbb70b08 ("NFSD: Replace nfsd_write()'s "stable" argument with "iocb_flags""),
which was applied after this patch was first posted. The original
passed a "u32 *stable_how" instead. Carrying the value as IOCB_* flags
keeps the NFSv3 XDR value out of the VFS API, and lets a later commit
raise the reported stability with a plain bitwise OR. The async
export option clears the flags on a local copy only, so the reply to
a FILE_SYNC or DATA_SYNC WRITE is never lowered (RFC 1813 section
3.3.7, RFC 8881 section 18.32.3). O_DIRECT alone does not persist
data, so the description no longer says it does. The Reviewed-by
tags from the original posting are dropped because the argument's
type and direction changed.]
Signed-off-by: Chuck Lever <chuck.lever@oracle.com>
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
fs/nfsd/nfs3proc.c | 16 +++++++++++++---
fs/nfsd/nfs4proc.c | 15 +++++++++++++--
fs/nfsd/nfsproc.c | 3 ++-
fs/nfsd/vfs.c | 17 ++++++++++-------
fs/nfsd/vfs.h | 4 ++--
fs/nfsd/xdr3.h | 2 +-
6 files changed, 41 insertions(+), 16 deletions(-)
diff --git a/fs/nfsd/nfs3proc.c b/fs/nfsd/nfs3proc.c
index 60cd01b6a37d2..f06573759ad6f 100644
--- a/fs/nfsd/nfs3proc.c
+++ b/fs/nfsd/nfs3proc.c
@@ -102,6 +102,15 @@ static int nfsd3_iocb_flags(enum nfs3_stable_how how)
}
}
+static enum nfs3_stable_how nfsd3_stable_how(int iocb_flags)
+{
+ if (iocb_flags & IOCB_SYNC)
+ return NFS_FILE_SYNC;
+ if (iocb_flags & IOCB_DSYNC)
+ return NFS_DATA_SYNC;
+ return NFS_UNSTABLE;
+}
+
static __be32 nfsd3_map_status(__be32 status)
{
switch (status) {
@@ -299,6 +308,7 @@ nfsd3_proc_write(struct svc_rqst *rqstp)
struct nfsd3_writeargs *argp = rqstp->rq_argp;
struct nfsd3_writeres *resp = rqstp->rq_resp;
unsigned long cnt = argp->len;
+ int iocb_flags;
dprintk("nfsd: WRITE(3) %s %d bytes at %Lu%s\n",
SVCFH_fmt(&argp->fh),
@@ -312,11 +322,11 @@ nfsd3_proc_write(struct svc_rqst *rqstp)
return rpc_success;
fh_copy(&resp->fh, &argp->fh);
- resp->committed = argp->stable;
+ iocb_flags = nfsd3_iocb_flags(argp->stable);
resp->status = nfsd_write(rqstp, &resp->fh, argp->offset,
- &argp->payload, &cnt,
- nfsd3_iocb_flags(resp->committed),
+ &argp->payload, &cnt, &iocb_flags,
resp->verf);
+ resp->committed = nfsd3_stable_how(iocb_flags);
resp->count = cnt;
resp->status = nfsd3_map_status(resp->status);
return rpc_success;
diff --git a/fs/nfsd/nfs4proc.c b/fs/nfsd/nfs4proc.c
index 7df60abfbff14..505c7dadf49b8 100644
--- a/fs/nfsd/nfs4proc.c
+++ b/fs/nfsd/nfs4proc.c
@@ -109,6 +109,15 @@ static const struct nfsd_access_maps nfsd4_access_maps = {
.other = nfsd4_otheraccess,
};
+static enum stable_how4 nfsd4_stable_how(int iocb_flags)
+{
+ if (iocb_flags & IOCB_SYNC)
+ return FILE_SYNC4;
+ if (iocb_flags & IOCB_DSYNC)
+ return DATA_SYNC4;
+ return UNSTABLE4;
+}
+
static int nfsd4_iocb_flags(enum stable_how4 how)
{
switch (how) {
@@ -1442,6 +1451,7 @@ nfsd4_write(struct svc_rqst *rqstp, struct nfsd4_compound_state *cstate,
struct nfsd_file *nf = NULL;
__be32 status = nfs_ok;
unsigned long cnt;
+ int iocb_flags;
if (write->wr_offset > (u64)OFFSET_MAX ||
write->wr_offset + write->wr_buflen > (u64)OFFSET_MAX)
@@ -1460,11 +1470,12 @@ nfsd4_write(struct svc_rqst *rqstp, struct nfsd4_compound_state *cstate,
nfs4_put_stid(stid);
}
- write->wr_how_written = write->wr_stable_how;
+ iocb_flags = nfsd4_iocb_flags(write->wr_stable_how);
status = nfsd_vfs_write(rqstp, &cstate->current_fh, nf,
write->wr_offset, &write->wr_payload,
- &cnt, nfsd4_iocb_flags(write->wr_how_written),
+ &cnt, &iocb_flags,
(__be32 *)write->wr_verifier.data);
+ write->wr_how_written = nfsd4_stable_how(iocb_flags);
nfsd_file_put(nf);
write->wr_bytes_written = cnt;
diff --git a/fs/nfsd/nfsproc.c b/fs/nfsd/nfsproc.c
index 1f74008bc06ad..6565d301c7e31 100644
--- a/fs/nfsd/nfsproc.c
+++ b/fs/nfsd/nfsproc.c
@@ -644,12 +644,13 @@ static __be32 nfsd_proc_write(struct svc_rqst *rqstp)
struct kstat *statp = &resp->stat;
unsigned long count = argp->xdrgen.data.len;
struct svc_fh *fhp = &argp->fh;
+ int iocb_flags = IOCB_DSYNC;
nfsd_fhandle_to_svc_fh(fhp, &argp->xdrgen.file);
resp->xdrgen.status = nfsd_write(rqstp, fhp, argp->xdrgen.offset,
&argp->xdrgen.data, &count,
- IOCB_DSYNC, NULL);
+ &iocb_flags, NULL);
if (resp->xdrgen.status == nfs_ok) {
resp->xdrgen.status = fh_getattr(fhp, statp);
if (resp->xdrgen.status == nfs_ok)
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index ec52c67c4380d..a8e15193b2160 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1537,7 +1537,8 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
* @offset: Byte offset of start
* @payload: xdr_buf containing the write payload
* @cnt: IN: number of bytes to write, OUT: number of bytes actually written
- * @iocb_flags: VFS IOCB_* flags expressing the requested write stability
+ * @iocb_flags: IN: VFS IOCB_* flags expressing the requested write
+ * stability; OUT: the stability to report to the client
* @verf: NFS WRITE verifier
*
* Upon return, caller must invoke fh_put on @fhp.
@@ -1549,11 +1550,12 @@ __be32
nfsd_vfs_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
struct nfsd_file *nf, loff_t offset,
const struct xdr_buf *payload, unsigned long *cnt,
- int iocb_flags, __be32 *verf)
+ int *iocb_flags, __be32 *verf)
{
struct nfsd_net *nn = net_generic(SVC_NET(rqstp), nfsd_net_id);
struct file *file = nf->nf_file;
struct super_block *sb = file_inode(file)->i_sb;
+ int stable_flags = *iocb_flags;
struct kiocb kiocb;
struct svc_export *exp;
struct iov_iter iter;
@@ -1586,11 +1588,11 @@ nfsd_vfs_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
exp = fhp->fh_export;
if (!EX_ISSYNC(exp))
- iocb_flags = 0;
+ stable_flags = 0;
init_sync_kiocb(&kiocb, file);
kiocb.ki_pos = offset;
if (likely(!fhp->fh_use_wgather))
- kiocb.ki_flags |= iocb_flags;
+ kiocb.ki_flags |= stable_flags;
nvecs = xdr_buf_to_bvec(rqstp->rq_bvec, rqstp->rq_maxpages, payload);
if (nvecs < 0) {
@@ -1631,7 +1633,7 @@ nfsd_vfs_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
goto out_nfserr;
}
- if (iocb_flags && fhp->fh_use_wgather) {
+ if (stable_flags && fhp->fh_use_wgather) {
host_err = wait_for_concurrent_writes(file);
if (host_err < 0)
nfsd_maybe_reset_write_verifier(nn, rqstp, host_err);
@@ -1722,7 +1724,8 @@ __be32 nfsd_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
* @offset: Byte offset of start
* @payload: xdr_buf containing the write payload
* @cnt: IN: number of bytes to write, OUT: number of bytes actually written
- * @iocb_flags: VFS IOCB_* flags expressing the requested write stability
+ * @iocb_flags: IN: VFS IOCB_* flags expressing the requested write
+ * stability; OUT: the stability to report to the client
* @verf: NFS WRITE verifier
*
* Upon return, caller must invoke fh_put on @fhp.
@@ -1733,7 +1736,7 @@ __be32 nfsd_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
__be32
nfsd_write(struct svc_rqst *rqstp, struct svc_fh *fhp, loff_t offset,
const struct xdr_buf *payload, unsigned long *cnt,
- int iocb_flags, __be32 *verf)
+ int *iocb_flags, __be32 *verf)
{
struct nfsd_file *nf;
__be32 err;
diff --git a/fs/nfsd/vfs.h b/fs/nfsd/vfs.h
index 4b352a48a8af2..c3b48c45e10a7 100644
--- a/fs/nfsd/vfs.h
+++ b/fs/nfsd/vfs.h
@@ -189,12 +189,12 @@ __be32 nfsd_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
u32 *eof);
__be32 nfsd_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
loff_t offset, const struct xdr_buf *payload,
- unsigned long *cnt, int iocb_flags,
+ unsigned long *cnt, int *iocb_flags,
__be32 *verf);
__be32 nfsd_vfs_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
struct nfsd_file *nf, loff_t offset,
const struct xdr_buf *payload,
- unsigned long *cnt, int iocb_flags,
+ unsigned long *cnt, int *iocb_flags,
__be32 *verf);
__be32 nfsd_readlink(struct svc_rqst *, struct svc_fh *,
char *, int *);
diff --git a/fs/nfsd/xdr3.h b/fs/nfsd/xdr3.h
index 22272695e451b..eb9d79e57fc9a 100644
--- a/fs/nfsd/xdr3.h
+++ b/fs/nfsd/xdr3.h
@@ -154,7 +154,7 @@ struct nfsd3_writeres {
__be32 status;
struct svc_fh fh;
unsigned long count;
- int committed;
+ u32 committed;
__be32 verf[2];
};
--
2.52.0
^ permalink raw reply related [flat|nested] 11+ messages in thread
* [PATCH v3 9/9] NFSD: add direct-mode WRITE settings that persist each WRITE
2026-10-01 4:54 [PATCH v3 0/9] NFSD: keep direct-mode I/O out of the page cache Mike Snitzer
` (7 preceding siblings ...)
2026-10-01 4:55 ` [PATCH v3 8/9] NFSD: Enable return of an updated stable_how to NFS clients Mike Snitzer
@ 2026-10-01 4:55 ` Mike Snitzer
2026-10-01 22:58 ` [PATCH v3 0/9] NFSD: keep direct-mode I/O out of the page cache Mike Snitzer
9 siblings, 0 replies; 11+ messages in thread
From: Mike Snitzer @ 2026-10-01 4:55 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: hch, linux-nfs
Under NFSD_IO_DIRECT an UNSTABLE WRITE is issued O_DIRECT without
IOCB_DSYNC and answered UNSTABLE. The data is not yet durable: the
device cache is not flushed, and block allocation, unwritten extent
conversion and the size update are not committed. The client's COMMIT
is what makes it so. This patch does not change NFSD_IO_DIRECT.
Add two io_cache_write modes that issue direct I/O like NFSD_IO_DIRECT
but persist every WRITE before replying, to at least a floor, and
report the stability achieved:
NFSD_IO_DIRECT_WRITE_DATA_SYNC (3): at least NFS_DATA_SYNC
NFSD_IO_DIRECT_WRITE_FILE_SYNC (4): at least NFS_FILE_SYNC
A client that asked for more is left alone. A client told FILE_SYNC
sends no COMMIT.
The cost is a flush and log force on every WRITE, where NFSD_IO_DIRECT
has one per COMMIT, so these are opt-in. They cost more than
NFSD_IO_DIRECT when one COMMIT covers many WRITEs. They cost the same
flushes, and save the COMMIT RPCs, when each COMMIT covers about one
WRITE. That is the case on a pNFS flexfiles share where a client's
writes straddle two data servers, so that each data server receives
one WRITE and one COMMIT per write. Measured there, with identical
WRITE counts and flush counts, the COMMIT RPCs cost NFSD_IO_DIRECT 31%
more server CPU and 42% more client CPU for the same bytes. Where only
one write in 22 was split, there was no resolvable difference.
A Linux client change that sends FILE_SYNC for a WRITE that is alone
in its flush to a given server would remove those COMMITs without a
server setting, for clients that have it.
Mode 3 promises only the data, leaving a client that needs metadata
durability to COMMIT for it. Both modes are inert for READ and for a
WRITE the client already marked FILE_SYNC.
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
.../filesystems/nfs/nfsd-io-modes.rst | 15 ++++++--
fs/nfsd/debugfs.c | 5 +++
fs/nfsd/nfsd.h | 2 ++
fs/nfsd/vfs.c | 35 ++++++++++++++++---
4 files changed, 50 insertions(+), 7 deletions(-)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index 11938caee14ca..87266e43e4d1c 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -25,12 +25,14 @@ Based on the configured settings, NFSD's IO will either be:
- cached using page cache (NFSD_IO_BUFFERED=0)
- cached but removed from page cache on completion (NFSD_IO_DONTCACHE=1)
- not cached stable_how=NFS_UNSTABLE (NFSD_IO_DIRECT=2)
+- not cached stable_how=NFS_DATA_SYNC (NFSD_IO_DIRECT_WRITE_DATA_SYNC=3)
+- not cached stable_how=NFS_FILE_SYNC (NFSD_IO_DIRECT_WRITE_FILE_SYNC=4)
-To set an NFSD IO mode, write a supported value (0 - 2) to the
+To set an NFSD IO mode, write a supported value (0 - 4) to the
corresponding IO operation's debugfs interface, e.g.::
echo 2 > /sys/kernel/debug/nfsd/io_cache_read
- echo 2 > /sys/kernel/debug/nfsd/io_cache_write
+ echo 4 > /sys/kernel/debug/nfsd/io_cache_write
To check which IO mode NFSD is using for READ or WRITE, simply read the
corresponding IO operation's debugfs interface, e.g.::
@@ -38,6 +40,15 @@ corresponding IO operation's debugfs interface, e.g.::
cat /sys/kernel/debug/nfsd/io_cache_read
cat /sys/kernel/debug/nfsd/io_cache_write
+NFSD_IO_DIRECT leaves an UNSTABLE WRITE for the client's COMMIT to
+persist. The two NFSD_IO_DIRECT_WRITE_*_SYNC modes instead persist every
+WRITE before replying, to at least NFS_DATA_SYNC or NFS_FILE_SYNC, and
+report that stable_how to the client; a client that asked for more is
+left alone. With NFSD_IO_DIRECT_WRITE_FILE_SYNC the client sends no
+COMMIT. This costs a flush on every WRITE: more than NFSD_IO_DIRECT
+when one COMMIT covers many WRITEs, about the same, without the COMMIT
+RPCs, when each COMMIT covers about one.
+
If you experiment with NFSD's IO modes on a recent kernel and have
interesting results, please report them to linux-nfs@vger.kernel.org
diff --git a/fs/nfsd/debugfs.c b/fs/nfsd/debugfs.c
index d5b714dd0fbc6..bd8ceb5eb2879 100644
--- a/fs/nfsd/debugfs.c
+++ b/fs/nfsd/debugfs.c
@@ -90,6 +90,9 @@ DEFINE_DEBUGFS_ATTRIBUTE(nfsd_io_cache_read_fops, nfsd_io_cache_read_get,
* Contents:
* %0: NFS WRITE will use buffered IO
* %1: NFS WRITE will use dontcache (buffered IO w/ dropbehind)
+ * %2: NFS WRITE will use direct IO with stable_how=NFS_UNSTABLE
+ * %3: NFS WRITE will use direct IO with stable_how=NFS_DATA_SYNC
+ * %4: NFS WRITE will use direct IO with stable_how=NFS_FILE_SYNC
*
* This setting takes immediate effect for all NFS versions,
* all exports, and in all NFSD net namespaces.
@@ -109,6 +112,8 @@ static int nfsd_io_cache_write_set(void *data, u64 val)
case NFSD_IO_BUFFERED:
case NFSD_IO_DONTCACHE:
case NFSD_IO_DIRECT:
+ case NFSD_IO_DIRECT_WRITE_DATA_SYNC:
+ case NFSD_IO_DIRECT_WRITE_FILE_SYNC:
nfsd_io_cache_write = val;
break;
default:
diff --git a/fs/nfsd/nfsd.h b/fs/nfsd/nfsd.h
index 1c17f96012523..913860ab3eb20 100644
--- a/fs/nfsd/nfsd.h
+++ b/fs/nfsd/nfsd.h
@@ -141,6 +141,8 @@ enum {
NFSD_IO_BUFFERED,
NFSD_IO_DONTCACHE,
NFSD_IO_DIRECT,
+ NFSD_IO_DIRECT_WRITE_DATA_SYNC,
+ NFSD_IO_DIRECT_WRITE_FILE_SYNC,
};
extern u64 nfsd_io_cache_read __read_mostly;
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index a8e15193b2160..af05ef942c5d9 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1450,12 +1450,25 @@ nfsd_write_dio_iters_init(struct svc_rqst *rqstp, struct svc_fh *fhp,
return nsegs;
}
+/* Raise this WRITE to at least @floor_iocb_flags, and report that. */
+static void
+nfsd_write_raise_stability(int floor_iocb_flags, struct kiocb *kiocb,
+ int *iocb_flags)
+{
+ if ((*iocb_flags & floor_iocb_flags) == floor_iocb_flags)
+ return;
+
+ *iocb_flags |= floor_iocb_flags;
+ kiocb->ki_flags |= floor_iocb_flags;
+}
+
static noinline_for_stack int
nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
- struct nfsd_file *nf, unsigned int nvecs,
+ struct nfsd_file *nf, int *iocb_flags, unsigned int nvecs,
unsigned long *cnt, struct kiocb *kiocb)
{
struct nfsd_write_dio_seg segments[3];
+ int floor_iocb_flags = 0;
struct file *file = nf->nf_file;
loff_t start = kiocb->ki_pos, seg_pos, seg_last;
bool sync, datasync, complete_first, complete_last;
@@ -1463,6 +1476,14 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
ssize_t host_err;
size_t expected;
+ if (nfsd_io_cache_write == NFSD_IO_DIRECT_WRITE_FILE_SYNC)
+ floor_iocb_flags = IOCB_DSYNC | IOCB_SYNC;
+ else if (nfsd_io_cache_write == NFSD_IO_DIRECT_WRITE_DATA_SYNC)
+ floor_iocb_flags = IOCB_DSYNC;
+ if (floor_iocb_flags)
+ nfsd_write_raise_stability(floor_iocb_flags, kiocb,
+ iocb_flags);
+
/* Persist a synchronous WRITE once, after all of its segments. */
sync = kiocb->ki_flags & IOCB_DSYNC;
datasync = !(kiocb->ki_flags & IOCB_SYNC);
@@ -1538,7 +1559,8 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
* @payload: xdr_buf containing the write payload
* @cnt: IN: number of bytes to write, OUT: number of bytes actually written
* @iocb_flags: IN: VFS IOCB_* flags expressing the requested write
- * stability; OUT: the stability to report to the client
+ * stability; OUT: the stability to report to the client,
+ * which may be raised
* @verf: NFS WRITE verifier
*
* Upon return, caller must invoke fh_put on @fhp.
@@ -1606,8 +1628,10 @@ nfsd_vfs_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
switch (nfsd_io_cache_write) {
case NFSD_IO_DIRECT:
- host_err = nfsd_direct_write(rqstp, fhp, nf, nvecs,
- cnt, &kiocb);
+ case NFSD_IO_DIRECT_WRITE_DATA_SYNC:
+ case NFSD_IO_DIRECT_WRITE_FILE_SYNC:
+ host_err = nfsd_direct_write(rqstp, fhp, nf, iocb_flags,
+ nvecs, cnt, &kiocb);
break;
case NFSD_IO_DONTCACHE:
if (file->f_op->fop_flags & FOP_DONTCACHE)
@@ -1725,7 +1749,8 @@ __be32 nfsd_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
* @payload: xdr_buf containing the write payload
* @cnt: IN: number of bytes to write, OUT: number of bytes actually written
* @iocb_flags: IN: VFS IOCB_* flags expressing the requested write
- * stability; OUT: the stability to report to the client
+ * stability; OUT: the stability to report to the client,
+ * which may be raised
* @verf: NFS WRITE verifier
*
* Upon return, caller must invoke fh_put on @fhp.
--
2.52.0
^ permalink raw reply related [flat|nested] 11+ messages in thread
* Re: [PATCH v3 0/9] NFSD: keep direct-mode I/O out of the page cache
2026-10-01 4:54 [PATCH v3 0/9] NFSD: keep direct-mode I/O out of the page cache Mike Snitzer
` (8 preceding siblings ...)
2026-10-01 4:55 ` [PATCH v3 9/9] NFSD: add direct-mode WRITE settings that persist each WRITE Mike Snitzer
@ 2026-10-01 22:58 ` Mike Snitzer
9 siblings, 0 replies; 11+ messages in thread
From: Mike Snitzer @ 2026-10-01 22:58 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: hch, linux-nfs
Just FYI, while this v3 is vastly improved (thanks to it addressng
most of Chuck's detailed review of v2) it still has issues that I'm
addressing.
I'll post v4 when its ready, so please don't waste time reviewing v3.
Thanks,
Mike
^ permalink raw reply [flat|nested] 11+ messages in thread
end of thread, other threads:[~2026-10-01 22:58 UTC | newest]
Thread overview: 11+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-01 4:54 [PATCH v3 0/9] NFSD: keep direct-mode I/O out of the page cache Mike Snitzer
2026-10-01 4:54 ` [PATCH v3 1/9] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well Mike Snitzer
2026-10-01 4:54 ` [PATCH v3 2/9] NFSD: only split a direct-mode WRITE for a worthwhile direct middle Mike Snitzer
2026-10-01 4:54 ` [PATCH v3 3/9] NFSD: add direct_misaligned_dontcache debugfs knob Mike Snitzer
2026-10-01 4:54 ` [PATCH v3 4/9] NFSD: do not use direct I/O for a READ smaller than its alignment Mike Snitzer
2026-10-01 4:54 ` [PATCH v3 5/9] NFSD: persist a synchronous direct-mode WRITE once, after all of its segments Mike Snitzer
2026-10-01 4:54 ` [PATCH v3 6/9] NFSD: keep boundary page of a split direct-mode WRITE until both writers complete Mike Snitzer
2026-10-01 4:55 ` [PATCH v3 7/9] NFSD: add tracing for how direct-mode READ and WRITE are serviced Mike Snitzer
2026-10-01 4:55 ` [PATCH v3 8/9] NFSD: Enable return of an updated stable_how to NFS clients Mike Snitzer
2026-10-01 4:55 ` [PATCH v3 9/9] NFSD: add direct-mode WRITE settings that persist each WRITE Mike Snitzer
2026-10-01 22:58 ` [PATCH v3 0/9] NFSD: keep direct-mode I/O out of the page cache Mike Snitzer
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox