* [PATCH 01/10] NFSD: interlock the use of NFSD_IO_DIRECT for NFS READ and WRITE
2026-09-29 17:34 [PATCH 00/10] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
@ 2026-09-29 17:34 ` Mike Snitzer
2026-09-29 18:27 ` Chuck Lever
2026-09-29 17:34 ` [PATCH 02/10] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well Mike Snitzer
` (8 subsequent siblings)
9 siblings, 1 reply; 17+ messages in thread
From: Mike Snitzer @ 2026-09-29 17:34 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: linux-nfs
Now that NFSD supports NFSD_IO_DIRECT for both READ and WRITE it is
much safer to avoid needless buffered vs direct contention if/when
only one of them has been configured to use NFSD_IO_DIRECT.
Mixing direct and buffered I/O to the same file causes needless page
cache invalidation and writeback, so although io_cache_read and
io_cache_write remain separate interfaces, writing either one adjusts
the other so that READ and WRITE are never left on opposite sides of
the buffered/direct divide:
- Setting io_cache_read to NFSD_IO_DIRECT elevates a BUFFERED or
DONTCACHE io_cache_write to NFSD_IO_DIRECT. A WRITE mode that is
already direct is left as it is.
- Setting io_cache_write to any direct mode elevates a BUFFERED or
DONTCACHE io_cache_read to NFSD_IO_DIRECT.
- Setting io_cache_read to NFSD_IO_BUFFERED or NFSD_IO_DONTCACHE
demotes a direct io_cache_write to that same mode.
- Setting io_cache_write to NFSD_IO_BUFFERED or NFSD_IO_DONTCACHE
demotes a direct io_cache_read to that same mode.
- Enabling splice for READ, by writing 0 to disable-splice-read,
forces io_cache_read to NFSD_IO_BUFFERED and demotes a direct
io_cache_write along with it.
The demotion is factored into nfsd_io_cache_write_demote() so that its
three callers stay in sync.
Both DONTCACHE and DIRECT READ must copy into the RPC reply buffer, so
splice for READ is disabled whenever either interface leaves
nfsd_io_cache_read above NFSD_IO_BUFFERED, regardless of which one was
written.
Document the complete interlock in
Documentation/filesystems/nfs/nfsd-io-modes.rst.
Fixes: 06c5c97293e3 ("NFSD: Implement NFSD_IO_DIRECT for NFS WRITE")
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
.../filesystems/nfs/nfsd-io-modes.rst | 35 ++++++++++++
fs/nfsd/debugfs.c | 54 ++++++++++++++++---
2 files changed, 83 insertions(+), 6 deletions(-)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index 0fd6e82478fe6..a8e277bcb08b1 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -41,6 +41,41 @@ corresponding IO operation's debugfs interface, e.g.::
If you experiment with NFSD's IO modes on a recent kernel and have
interesting results, please report them to linux-nfs@vger.kernel.org
+READ and WRITE IO mode interlock
+================================
+
+Although io_cache_read and io_cache_write are separate interfaces, NFSD
+keeps them from being configured such that one of READ or WRITE uses
+DIRECT IO while the other uses the page cache. Mixing DIRECT and
+buffered IO to the same file causes needless page cache invalidation and
+writeback (see the DIRECT IO discussion in the Linux open(2) manpage),
+so writing one interface may adjust the other:
+
+- Setting io_cache_read to NFSD_IO_DIRECT (2) elevates io_cache_write
+ to NFSD_IO_DIRECT (2) if it was BUFFERED or DONTCACHE. A WRITE mode
+ that is already DIRECT (2, 3 or 4) is left unchanged.
+- Setting io_cache_write to any DIRECT mode (2, 3 or 4) elevates
+ io_cache_read to NFSD_IO_DIRECT (2) if it was BUFFERED or DONTCACHE.
+- Setting io_cache_read to NFSD_IO_BUFFERED (0) or NFSD_IO_DONTCACHE (1)
+ while io_cache_write is a DIRECT mode demotes io_cache_write to that
+ same value (0 or 1). A WRITE mode that is already BUFFERED or
+ DONTCACHE is left unchanged.
+- Setting io_cache_write to NFSD_IO_BUFFERED (0) or NFSD_IO_DONTCACHE
+ (1) while io_cache_read is NFSD_IO_DIRECT demotes io_cache_read to
+ that same value (0 or 1).
+
+Setting either interface to a value other than NFSD_IO_BUFFERED also
+disables NFSD's use of splice for READ, because both DONTCACHE and
+DIRECT READ must copy into the RPC reply buffer. This is reflected in
+/sys/kernel/debug/nfsd/disable-splice-read reading as 1. Writing 0 to
+disable-splice-read re-enables splice, which requires buffered READ, so
+it forces io_cache_read back to NFSD_IO_BUFFERED (0) and, if
+io_cache_write was a DIRECT mode, demotes it to NFSD_IO_BUFFERED (0) as
+well.
+
+Always read both interfaces back after writing either of them to
+confirm the resulting configuration.
+
NFSD DONTCACHE
==============
diff --git a/fs/nfsd/debugfs.c b/fs/nfsd/debugfs.c
index 386fd1c54f527..995872a1a7f87 100644
--- a/fs/nfsd/debugfs.c
+++ b/fs/nfsd/debugfs.c
@@ -24,6 +24,17 @@ static int nfsd_dsr_get(void *data, u64 *val)
return 0;
}
+/*
+ * NFS READ is no longer using direct I/O: demote NFS WRITE from direct
+ * I/O to the same buffered mode, to avoid needless buffered vs direct
+ * contention.
+ */
+static void nfsd_io_cache_write_demote(u64 io_mode)
+{
+ if (nfsd_io_cache_write >= NFSD_IO_DIRECT)
+ nfsd_io_cache_write = io_mode;
+}
+
static int nfsd_dsr_set(void *data, u64 val)
{
nfsd_disable_splice_read = (val > 0);
@@ -32,6 +43,7 @@ static int nfsd_dsr_set(void *data, u64 val)
* Must use buffered I/O if splice_read is enabled.
*/
nfsd_io_cache_read = NFSD_IO_BUFFERED;
+ nfsd_io_cache_write_demote(NFSD_IO_BUFFERED);
}
return 0;
}
@@ -62,22 +74,33 @@ static int nfsd_io_cache_read_set(void *data, u64 val)
switch (val) {
case NFSD_IO_BUFFERED:
- nfsd_io_cache_read = NFSD_IO_BUFFERED;
- break;
case NFSD_IO_DONTCACHE:
+ nfsd_io_cache_read = val;
+ nfsd_io_cache_write_demote(val);
+ break;
case NFSD_IO_DIRECT:
+ nfsd_io_cache_read = val;
/*
- * Must disable splice_read when enabling
- * NFSD_IO_DONTCACHE.
+ * Elevate nfsd_io_cache_write if not already
+ * configured to use NFSD_IO_DIRECT.
*/
- nfsd_disable_splice_read = true;
- nfsd_io_cache_read = val;
+ if (nfsd_io_cache_write < NFSD_IO_DIRECT)
+ nfsd_io_cache_write = NFSD_IO_DIRECT;
break;
default:
ret = -EINVAL;
break;
}
+ if (ret == 0) {
+ /*
+ * Must disable splice_read when enabling
+ * NFSD_IO_DONTCACHE and NFSD_IO_DIRECT.
+ */
+ if (nfsd_io_cache_read > NFSD_IO_BUFFERED)
+ nfsd_disable_splice_read = true;
+ }
+
return ret;
}
@@ -110,12 +133,31 @@ static int nfsd_io_cache_write_set(void *data, u64 val)
case NFSD_IO_DONTCACHE:
case NFSD_IO_DIRECT:
nfsd_io_cache_write = val;
+ /*
+ * Adjust nfsd_io_cache_{read,write} to avoid
+ * needless buffered vs direct contention.
+ */
+ if (nfsd_io_cache_write >= NFSD_IO_DIRECT &&
+ nfsd_io_cache_read < NFSD_IO_DIRECT)
+ nfsd_io_cache_read = NFSD_IO_DIRECT;
+ else if (nfsd_io_cache_write < NFSD_IO_DIRECT &&
+ nfsd_io_cache_read == NFSD_IO_DIRECT)
+ nfsd_io_cache_read = nfsd_io_cache_write;
break;
default:
ret = -EINVAL;
break;
}
+ if (ret == 0) {
+ /*
+ * Must disable splice_read when enabling
+ * NFSD_IO_DONTCACHE and NFSD_IO_DIRECT.
+ */
+ if (nfsd_io_cache_read > NFSD_IO_BUFFERED)
+ nfsd_disable_splice_read = true;
+ }
+
return ret;
}
--
2.52.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* Re: [PATCH 01/10] NFSD: interlock the use of NFSD_IO_DIRECT for NFS READ and WRITE
2026-09-29 17:34 ` [PATCH 01/10] NFSD: interlock the use of NFSD_IO_DIRECT for NFS READ and WRITE Mike Snitzer
@ 2026-09-29 18:27 ` Chuck Lever
2026-09-29 19:56 ` Mike Snitzer
0 siblings, 1 reply; 17+ messages in thread
From: Chuck Lever @ 2026-09-29 18:27 UTC (permalink / raw)
To: Mike Snitzer, Jeff Layton; +Cc: linux-nfs
On Tue, Sep 29, 2026, at 10:34 AM, Mike Snitzer wrote:
> Now that NFSD supports NFSD_IO_DIRECT for both READ and WRITE it is
> much safer to avoid needless buffered vs direct contention if/when
> only one of them has been configured to use NFSD_IO_DIRECT.
>
> Mixing direct and buffered I/O to the same file causes needless page
> cache invalidation and writeback, so although io_cache_read and
> io_cache_write remain separate interfaces, writing either one adjusts
> the other so that READ and WRITE are never left on opposite sides of
> the buffered/direct divide:
Jeff and I have been discussing making DIRECT the default for WRITEs and
BUFFERED the default for READs. This patch takes us in the opposite
direction.
Nothing here convinces me that mixing the modes is a bad thing to do.
"Needless page cache invalidation" needs some demonstration, and it
needs to show why the right thing to do is make it impossible to mix
modes rather than explore the issue as one or more bugs that can be
fixed. Or... why not let admins explore this for themselves? Where is
the hazard and why does it need to be forbidden by the admin UI?
Later patches in the series assume DONTCACHE is the fallback for
certain cases. Jeff has measured substantial performance deficits for
that mode, and maybe removing DONTCACHE would be a better direction
to take.
The patch that reports the actual stability of a WRITE was dropped
because it broke something (although I don't remember what). As I
recall, even a direct WRITE needs a subsequent COMMIT. Have you
demonstrated that the client COMMIT is a latency problem or that
the memory that is pinned on the client has a noticeable impact?
We suspect that it might, but every time I've measured, avoiding
COMMIT with NFSD has shown no impact on throughput, which is why
NFSD doesn't already do this optimization.
So even if these patches apply mechanically and benefit your workloads,
it would be helpful if we could step back and understand what you are
trying to achieve and how it should coordinate/align with where Jeff
and I want to see this mechanism go. Can we start with that
conversation?
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH 01/10] NFSD: interlock the use of NFSD_IO_DIRECT for NFS READ and WRITE
2026-09-29 18:27 ` Chuck Lever
@ 2026-09-29 19:56 ` Mike Snitzer
2026-09-29 23:17 ` Chuck Lever
0 siblings, 1 reply; 17+ messages in thread
From: Mike Snitzer @ 2026-09-29 19:56 UTC (permalink / raw)
To: Chuck Lever; +Cc: Jeff Layton, linux-nfs
On Tue, Sep 29, 2026 at 11:27:05AM -0700, Chuck Lever wrote:
>
> On Tue, Sep 29, 2026, at 10:34 AM, Mike Snitzer wrote:
> > Now that NFSD supports NFSD_IO_DIRECT for both READ and WRITE it is
> > much safer to avoid needless buffered vs direct contention if/when
> > only one of them has been configured to use NFSD_IO_DIRECT.
> >
> > Mixing direct and buffered I/O to the same file causes needless page
> > cache invalidation and writeback, so although io_cache_read and
> > io_cache_write remain separate interfaces, writing either one adjusts
> > the other so that READ and WRITE are never left on opposite sides of
> > the buffered/direct divide:
>
> Jeff and I have been discussing making DIRECT the default for WRITEs and
> BUFFERED the default for READs. This patch takes us in the opposite
> direction.
>
> Nothing here convinces me that mixing the modes is a bad thing to do.
> "Needless page cache invalidation" needs some demonstration, and it
> needs to show why the right thing to do is make it impossible to mix
> modes rather than explore the issue as one or more bugs that can be
> fixed. Or... why not let admins explore this for themselves? Where is
> the hazard and why does it need to be forbidden by the admin UI?
The basis for the interlock is that O_DIRECT mode is intended to avoid
bloating memory with page cache and the excess CPU burn of managing
the page cache. Using read caching in conjunction with O_DIRECT
writes knee-caps the wins of O_DIRECT. But sure, you can do it.. and
I suppose the freedom of choice (and enough rope to hang) is perfectly
fine. I can drop the interlock and send v2 after more time for review
of the other patches (but there were some real bugs fixed in the
interlock patch so will require some care to adjust).
direct_misaligned_dontcache=N shows the excessive 630 GiB page cache
bloat (across 4 servers) I've had to confront with a larger scale test
that really sticks its finger in this particular wound (as documented at
end of the "NFSD: keep boundary page of a split direct-mode WRITE
until both writers complete" patch):
On a four-server pNFS flexfiles rig, three clients writing 47008-byte
records for 240 s against 640 nfsd threads per server, compared with
the same servers before this change: device reads fall from 0.014-0.015
per record to 0.003, page-cache growth over the write phase falls from
about 11 GiB to 2.9 GiB, and write throughput rises by 2.4% to 4.9%
across NFSD_IO_DIRECT and both stable_how floors, with read throughput
unchanged. Records of that size always take the split path, since their
direct middle is at least 38818 bytes, so the figures are the split
path alone; the whole-WRITE fallback needs records below about 12 KB to
come into play. Keeping every boundary page cached instead, which a
later commit makes possible with a direct_misaligned_dontcache=N
debugfs knob, grows the page cache by about 630 GiB over the same
runs.
The "11 GiB" is what's left resident in page cache _without_ the
"NFSD: keep boundary page of a split direct-mode WRITE until both
writers complete". _WITH_ that patch (to make more informed use of
DONTCACHE) it drops to the quoted 2.9 GiB while still benefitting from
RMW avoidance buffered IO makes possible -- pretty fantastic result.
> Later patches in the series assume DONTCACHE is the fallback for
> certain cases.
That's the existing default that current upstream code has. But with
this patchset the fallback is buffered, with buffered DONTCACHE the
default unless debugfs direct_misaligned_dontcache=N configured.
It would be quite bad to remove DONTCACHE support because it actually
does offer pretty solid wins (especially for the workload I quoted
above). If pure buffered IO used (direct_misaligned_dontcache=N)
WRITE performance suffers: "only" 31,297 MiB/s for O_DIRECT +
buffered, whereas with O_DIRECT + DONTCACHE, overall result were (more
runs needed to get stddev):
4 - FILE_SYNC floor 35,567 MiB/s
2 - NFSD_IO_DIRECT 35,302 MiB/s
3 - DATA_SYNC floor 36,163 MiB/s
> Jeff has measured substantial performance deficits for that mode,
> and maybe removing DONTCACHE would be a better direction to take.
But to be clear, I'm not referring to DONTCACHE only mode=1, I'm most
interested in hybrid of O_DIRECT+DONTCACHE (more on that at the end
below).
My series has been fully tested with Jeff's more recent DONTCACHE
commits in place:
88d6f128d06d mm: track DONTCACHE dirty pages per bdi_writeback
f3122ce09a51 mm: kick writeback flusher for IOCB_DONTCACHE with targeted dirty tracking
> The patch that reports the actual stability of a WRITE was dropped
> because it broke something (although I don't remember what).
Haven't seen any issues with it. Been carrying it ever since you
posted it.
> As I recall, even a direct WRITE needs a subsequent COMMIT. Have you
> demonstrated that the client COMMIT is a latency problem or that
> the memory that is pinned on the client has a noticeable impact?
>
> We suspect that it might, but every time I've measured, avoiding
> COMMIT with NFSD has shown no impact on throughput, which is why
> NFSD doesn't already do this optimization.
Yes, in the header for "NFSD: let a direct-mode WRITE raise stable_how
and elide the client's COMMIT":
What this removes is the COMMIT traffic, and it is worth most where a
client's writes are carved into several WRITE RPCs, because each piece
is then committed separately. Measured on a pNFS flexfiles share where
every write straddles two data servers, so every write becomes two WRITE
RPCs and, under NFSD_IO_DIRECT, two COMMITs: the two modes do identical
durability work, one fsync per COMMIT against one fsync per WRITE on
identical WRITE counts, and the COMMIT RPCs alone cost NFSD_IO_DIRECT
31% more server CPU and 42% more client CPU for the same bytes, about
20 us of server CPU per COMMIT plus a client cost that grows with the
range committed. Where only one write in 22 is split the same effect
is a couple of cores on each side and no resolvable throughput
difference, and writes that fit a single RPC send no COMMIT in either
mode.
> So even if these patches apply mechanically and benefit your workloads,
> it would be helpful if we could step back and understand what you are
> trying to achieve and how it should coordinate/align with where Jeff
> and I want to see this mechanism go. Can we start with that
> conversation?
OK, please review what I've provided further and we can then have a
more detailed discussion about anything you like.
Claude helped summarize what this series fixes:
What this advance is, precisely. It is not "NFSD uses DONTCACHE". It
is NFSD's direct write path - io_cache_write modes 2, 3 and 4, where
an aligned WRITE goes to the filesystem as O_DIRECT - with DONTCACHE
used only for the buffered fragments a misaligned WRITE cannot issue
directly. The page cache is bypassed for the bulk of the data and
bounded for the remainder. That distinction runs through everything
below, and it is exactly what separates this from io_cache_write=1,
which is plain buffered DONTCACHE with no direct I/O at all and does
not reach this code.
A misaligned WRITE in one of those direct modes is split into a
buffered prefix, an O_DIRECT middle and a buffered suffix. The
boundary pages are shared by exactly two WRITEs that may arrive in
either order, from different clients, at once. Marking both DONTCACHE
bounded the page cache but cost the second writer a read from disk
every time - 0.81 device reads per record on the rig, 585 GiB read
back during an 8.1 TiB write.
Thanks,
Mike
^ permalink raw reply [flat|nested] 17+ messages in thread* Re: [PATCH 01/10] NFSD: interlock the use of NFSD_IO_DIRECT for NFS READ and WRITE
2026-09-29 19:56 ` Mike Snitzer
@ 2026-09-29 23:17 ` Chuck Lever
2026-09-29 23:30 ` Mike Snitzer
0 siblings, 1 reply; 17+ messages in thread
From: Chuck Lever @ 2026-09-29 23:17 UTC (permalink / raw)
To: Mike Snitzer; +Cc: Jeff Layton, linux-nfs
On Tue, Sep 29, 2026, at 12:56 PM, Mike Snitzer wrote:
> On Tue, Sep 29, 2026 at 11:27:05AM -0700, Chuck Lever wrote:
>>
>> On Tue, Sep 29, 2026, at 10:34 AM, Mike Snitzer wrote:
>> > Now that NFSD supports NFSD_IO_DIRECT for both READ and WRITE it is
>> > much safer to avoid needless buffered vs direct contention if/when
>> > only one of them has been configured to use NFSD_IO_DIRECT.
>> >
>> > Mixing direct and buffered I/O to the same file causes needless page
>> > cache invalidation and writeback, so although io_cache_read and
>> > io_cache_write remain separate interfaces, writing either one adjusts
>> > the other so that READ and WRITE are never left on opposite sides of
>> > the buffered/direct divide:
>>
>> Jeff and I have been discussing making DIRECT the default for WRITEs and
>> BUFFERED the default for READs. This patch takes us in the opposite
>> direction.
>>
>> Nothing here convinces me that mixing the modes is a bad thing to do.
>> "Needless page cache invalidation" needs some demonstration, and it
>> needs to show why the right thing to do is make it impossible to mix
>> modes rather than explore the issue as one or more bugs that can be
>> fixed. Or... why not let admins explore this for themselves? Where is
>> the hazard and why does it need to be forbidden by the admin UI?
>
> The basis for the interlock is that O_DIRECT mode is intended to avoid
> bloating memory with page cache and the excess CPU burn of managing
> the page cache. Using read caching in conjunction with O_DIRECT
> writes knee-caps the wins of O_DIRECT.
Jeff and I have never seen that, and it's a counterintuitive result.
Buffered READ with DIRECT WRITE seems to work very well and the server
is easily capable of managing the page cache in this case, since
reclaiming a clean page doesn't mean having to flush dirty data.
Evicting clean pages is not slow.
So I'd like to see a quantification of the penalties in this mixed
mode before adjudicating it as a hazard. That is, you might be right,
but so far I've seen no direct evidence that having a substantial page
cache presence is a general deficit for WRITEs. If there are certain
cases where even caching READs is a problem, then by all means, set the
READ IO mode to DIRECT too in those cases.
I don't see this as an argument for forcing the server's mode setting
in every case.
> But sure, you can do it.. and
> I suppose the freedom of choice (and enough rope to hang) is perfectly
> fine.
"Enough rope to hang" is the usual approach for Linux tunables. I
should also point out that both of these settings are still debugfs,
not the final shape of the administrative UI. It doesn't seem sensible
to me to restrict them at this point, but I'll keep this idea in mind
as we continue to develop our thinking about what the eventual non-
debug admin UI will look like.
> I can drop the interlock and send v2 after more time for review
> of the other patches (but there were some real bugs fixed in the
> interlock patch so will require some care to adjust).
The usual policy for "real bugs" is that those need to be fixed in
separate, backport-able patches before making behavioral changes.
Stable wants the fixes, and does not want the behavior changes, so
these need to be separable.
You can send the bug fixes any time without waiting for a v2.
>> As I recall, even a direct WRITE needs a subsequent COMMIT. Have you
>> demonstrated that the client COMMIT is a latency problem or that
>> the memory that is pinned on the client has a noticeable impact?
>>
>> We suspect that it might, but every time I've measured, avoiding
>> COMMIT with NFSD has shown no impact on throughput, which is why
>> NFSD doesn't already do this optimization.
>
> Yes, in the header for "NFSD: let a direct-mode WRITE raise stable_how
> and elide the client's COMMIT":
>
> What this removes is the COMMIT traffic, and it is worth most where a
> client's writes are carved into several WRITE RPCs, because each piece
> is then committed separately. Measured on a pNFS flexfiles share where
> every write straddles two data servers, so every write becomes two WRITE
> RPCs and, under NFSD_IO_DIRECT, two COMMITs: the two modes do identical
> durability work, one fsync per COMMIT against one fsync per WRITE on
> identical WRITE counts, and the COMMIT RPCs alone cost NFSD_IO_DIRECT
> 31% more server CPU and 42% more client CPU for the same bytes, about
> 20 us of server CPU per COMMIT plus a client cost that grows with the
> range committed. Where only one write in 22 is split the same effect
> is a couple of cores on each side and no resolvable throughput
> difference, and writes that fit a single RPC send no COMMIT in either
> mode.
I asked about latency, above. This paragraph is about CPU utilization.
Based on this description, it seems to me the client has full visibility
of the data to be pushed back and how it's sharded, it has information
about the network RTT (that's where the real throughput impact is for
COMMIT), and it has control over the selection of UNSTABLE vs. FILE_SYNC.
The problem here might be that when a large WRITE payload is sharded
across multiple servers, the client still thinks it is sending them via
UNSTABLE WRITES to one server, and plans for only one COMMIT after the
server completes the WRITEs.
But for the pNFS scenario, the client sends one UNSTABLE WRITE followed
by a COMMIT to each server. If the sharded WRITES are all single RPCs
to distinct servers, then they should each be FILE_SYNC.
The client can be smarter about how it writes data back to multiple
servers, can't it?
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH 01/10] NFSD: interlock the use of NFSD_IO_DIRECT for NFS READ and WRITE
2026-09-29 23:17 ` Chuck Lever
@ 2026-09-29 23:30 ` Mike Snitzer
2026-09-30 0:20 ` Chuck Lever
0 siblings, 1 reply; 17+ messages in thread
From: Mike Snitzer @ 2026-09-29 23:30 UTC (permalink / raw)
To: Chuck Lever; +Cc: Mike Snitzer, Jeff Layton, linux-nfs
On Tue, Sep 29, 2026 at 04:17:50PM -0700, Chuck Lever wrote:
>
>
> On Tue, Sep 29, 2026, at 12:56 PM, Mike Snitzer wrote:
> > On Tue, Sep 29, 2026 at 11:27:05AM -0700, Chuck Lever wrote:
> >>
> >> On Tue, Sep 29, 2026, at 10:34 AM, Mike Snitzer wrote:
> >> > Now that NFSD supports NFSD_IO_DIRECT for both READ and WRITE it is
> >> > much safer to avoid needless buffered vs direct contention if/when
> >> > only one of them has been configured to use NFSD_IO_DIRECT.
> >> >
> >> > Mixing direct and buffered I/O to the same file causes needless page
> >> > cache invalidation and writeback, so although io_cache_read and
> >> > io_cache_write remain separate interfaces, writing either one adjusts
> >> > the other so that READ and WRITE are never left on opposite sides of
> >> > the buffered/direct divide:
> >>
> >> Jeff and I have been discussing making DIRECT the default for WRITEs and
> >> BUFFERED the default for READs. This patch takes us in the opposite
> >> direction.
> >>
> >> Nothing here convinces me that mixing the modes is a bad thing to do.
> >> "Needless page cache invalidation" needs some demonstration, and it
> >> needs to show why the right thing to do is make it impossible to mix
> >> modes rather than explore the issue as one or more bugs that can be
> >> fixed. Or... why not let admins explore this for themselves? Where is
> >> the hazard and why does it need to be forbidden by the admin UI?
> >
> > The basis for the interlock is that O_DIRECT mode is intended to avoid
> > bloating memory with page cache and the excess CPU burn of managing
> > the page cache. Using read caching in conjunction with O_DIRECT
> > writes knee-caps the wins of O_DIRECT.
>
> Jeff and I have never seen that, and it's a counterintuitive result.
> Buffered READ with DIRECT WRITE seems to work very well and the server
> is easily capable of managing the page cache in this case, since
> reclaiming a clean page doesn't mean having to flush dirty data.
> Evicting clean pages is not slow.
>
> So I'd like to see a quantification of the penalties in this mixed
> mode before adjudicating it as a hazard. That is, you might be right,
> but so far I've seen no direct evidence that having a substantial page
> cache presence is a general deficit for WRITEs. If there are certain
> cases where even caching READs is a problem, then by all means, set the
> READ IO mode to DIRECT too in those cases.
>
> I don't see this as an argument for forcing the server's mode setting
> in every case.
Yes, I understand and I've dropped the interlock patch in v2 (just posted).
> > But sure, you can do it.. and
> > I suppose the freedom of choice (and enough rope to hang) is perfectly
> > fine.
>
> "Enough rope to hang" is the usual approach for Linux tunables. I
> should also point out that both of these settings are still debugfs,
> not the final shape of the administrative UI. It doesn't seem sensible
> to me to restrict them at this point, but I'll keep this idea in mind
> as we continue to develop our thinking about what the eventual non-
> debug admin UI will look like.
Ack.
> > I can drop the interlock and send v2 after more time for review
> > of the other patches (but there were some real bugs fixed in the
> > interlock patch so will require some care to adjust).
>
> The usual policy for "real bugs" is that those need to be fixed in
> separate, backport-able patches before making behavioral changes.
> Stable wants the fixes, and does not want the behavior changes, so
> these need to be separable.
>
> You can send the bug fixes any time without waiting for a v2.
The bug was something I found and fixed in code I had been carrying
privately. So not applicable to upstream, sorry for the noise.
> >> As I recall, even a direct WRITE needs a subsequent COMMIT. Have you
> >> demonstrated that the client COMMIT is a latency problem or that
> >> the memory that is pinned on the client has a noticeable impact?
> >>
> >> We suspect that it might, but every time I've measured, avoiding
> >> COMMIT with NFSD has shown no impact on throughput, which is why
> >> NFSD doesn't already do this optimization.
> >
> > Yes, in the header for "NFSD: let a direct-mode WRITE raise stable_how
> > and elide the client's COMMIT":
> >
> > What this removes is the COMMIT traffic, and it is worth most where a
> > client's writes are carved into several WRITE RPCs, because each piece
> > is then committed separately. Measured on a pNFS flexfiles share where
> > every write straddles two data servers, so every write becomes two WRITE
> > RPCs and, under NFSD_IO_DIRECT, two COMMITs: the two modes do identical
> > durability work, one fsync per COMMIT against one fsync per WRITE on
> > identical WRITE counts, and the COMMIT RPCs alone cost NFSD_IO_DIRECT
> > 31% more server CPU and 42% more client CPU for the same bytes, about
> > 20 us of server CPU per COMMIT plus a client cost that grows with the
> > range committed. Where only one write in 22 is split the same effect
> > is a couple of cores on each side and no resolvable throughput
> > difference, and writes that fit a single RPC send no COMMIT in either
> > mode.
>
> I asked about latency, above. This paragraph is about CPU utilization.
Fair point, latency wasn't measured.
> Based on this description, it seems to me the client has full visibility
> of the data to be pushed back and how it's sharded, it has information
> about the network RTT (that's where the real throughput impact is for
> COMMIT), and it has control over the selection of UNSTABLE vs. FILE_SYNC.
>
> The problem here might be that when a large WRITE payload is sharded
> across multiple servers, the client still thinks it is sending them via
> UNSTABLE WRITES to one server, and plans for only one COMMIT after the
> server completes the WRITEs.
>
> But for the pNFS scenario, the client sends one UNSTABLE WRITE followed
> by a COMMIT to each server. If the sharded WRITES are all single RPCs
> to distinct servers, then they should each be FILE_SYNC.
>
> The client can be smarter about how it writes data back to multiple
> servers, can't it?
When the server is operating in direct mode it isn't something exposed
to the client. That the client could be smarter and/or already has
adequate controls to achieve the same FILE_SYNC result is besides the
point. The point is, if the server has already done the work then it
should, within reason, convey as much back to the client to elide
COMMIT work that isn't needed.
Mike
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH 01/10] NFSD: interlock the use of NFSD_IO_DIRECT for NFS READ and WRITE
2026-09-29 23:30 ` Mike Snitzer
@ 2026-09-30 0:20 ` Chuck Lever
2026-09-30 12:46 ` Mike Snitzer
0 siblings, 1 reply; 17+ messages in thread
From: Chuck Lever @ 2026-09-30 0:20 UTC (permalink / raw)
To: Mike Snitzer; +Cc: Mike Snitzer, Jeff Layton, linux-nfs
On Tue, Sep 29, 2026, at 4:30 PM, Mike Snitzer wrote:
> On Tue, Sep 29, 2026 at 04:17:50PM -0700, Chuck Lever wrote:
>>
>>
>> On Tue, Sep 29, 2026, at 12:56 PM, Mike Snitzer wrote:
>> > On Tue, Sep 29, 2026 at 11:27:05AM -0700, Chuck Lever wrote:
>> >>
>> >> On Tue, Sep 29, 2026, at 10:34 AM, Mike Snitzer wrote:
>> >> > Now that NFSD supports NFSD_IO_DIRECT for both READ and WRITE it is
>> >> > much safer to avoid needless buffered vs direct contention if/when
>> >> > only one of them has been configured to use NFSD_IO_DIRECT.
>> >> >
>> >> > Mixing direct and buffered I/O to the same file causes needless page
>> >> > cache invalidation and writeback, so although io_cache_read and
>> >> > io_cache_write remain separate interfaces, writing either one adjusts
>> >> > the other so that READ and WRITE are never left on opposite sides of
>> >> > the buffered/direct divide:
>> >>
>> >> Jeff and I have been discussing making DIRECT the default for WRITEs and
>> >> BUFFERED the default for READs. This patch takes us in the opposite
>> >> direction.
>> >>
>> >> Nothing here convinces me that mixing the modes is a bad thing to do.
>> >> "Needless page cache invalidation" needs some demonstration, and it
>> >> needs to show why the right thing to do is make it impossible to mix
>> >> modes rather than explore the issue as one or more bugs that can be
>> >> fixed. Or... why not let admins explore this for themselves? Where is
>> >> the hazard and why does it need to be forbidden by the admin UI?
>> >
>> > The basis for the interlock is that O_DIRECT mode is intended to avoid
>> > bloating memory with page cache and the excess CPU burn of managing
>> > the page cache. Using read caching in conjunction with O_DIRECT
>> > writes knee-caps the wins of O_DIRECT.
>>
>> Jeff and I have never seen that, and it's a counterintuitive result.
>> Buffered READ with DIRECT WRITE seems to work very well and the server
>> is easily capable of managing the page cache in this case, since
>> reclaiming a clean page doesn't mean having to flush dirty data.
>> Evicting clean pages is not slow.
>>
>> So I'd like to see a quantification of the penalties in this mixed
>> mode before adjudicating it as a hazard. That is, you might be right,
>> but so far I've seen no direct evidence that having a substantial page
>> cache presence is a general deficit for WRITEs. If there are certain
>> cases where even caching READs is a problem, then by all means, set the
>> READ IO mode to DIRECT too in those cases.
>>
>> I don't see this as an argument for forcing the server's mode setting
>> in every case.
>
> Yes, I understand and I've dropped the interlock patch in v2 (just posted).
Please give reviewers a chance to digest and review, as requested in
Documentation/filesystems/nfs/nfsd-maintainer-entry-profile.rst :
"As always, please avoid reposting series revisions more than once
every 24 hours."
Trust me, it saves a lot of confusion.
>> Based on this description, it seems to me the client has full visibility
>> of the data to be pushed back and how it's sharded, it has information
>> about the network RTT (that's where the real throughput impact is for
>> COMMIT), and it has control over the selection of UNSTABLE vs. FILE_SYNC.
>>
>> The problem here might be that when a large WRITE payload is sharded
>> across multiple servers, the client still thinks it is sending them via
>> UNSTABLE WRITES to one server, and plans for only one COMMIT after the
>> server completes the WRITEs.
>>
>> But for the pNFS scenario, the client sends one UNSTABLE WRITE followed
>> by a COMMIT to each server. If the sharded WRITES are all single RPCs
>> to distinct servers, then they should each be FILE_SYNC.
>>
>> The client can be smarter about how it writes data back to multiple
>> servers, can't it?
>
> When the server is operating in direct mode it isn't something exposed
> to the client. That the client could be smarter and/or already has
> adequate controls to achieve the same FILE_SYNC result is besides the
> point. The point is, if the server has already done the work then it
> should, within reason, convey as much back to the client to elide
> COMMIT work that isn't needed.
Your point assumes the other patches in the series are applied to make
DIRECT UNSTABLE WRITEs completely persistent.
The current IOCB flags do not include IOCB_DSYNC on a DIRECT UNSTABLE
WRITE for a very good reason: that makes them slower and more expensive.
The current server logic is working exactly as we designed it last year,
and I'm not enthusiastic about changing that. I expect at least one
other reviewer will have a similar reaction, once he sobers up from
ALPSS.
If the client wants to avoid the COMMIT, the standing rule is to send a
FILE_SYNC WRITE. That benefits all WRITE I/O modes on the server.
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH 01/10] NFSD: interlock the use of NFSD_IO_DIRECT for NFS READ and WRITE
2026-09-30 0:20 ` Chuck Lever
@ 2026-09-30 12:46 ` Mike Snitzer
0 siblings, 0 replies; 17+ messages in thread
From: Mike Snitzer @ 2026-09-30 12:46 UTC (permalink / raw)
To: Chuck Lever; +Cc: hch, Jeff Layton, linux-nfs
On Tue, Sep 29, 2026 at 05:20:03PM -0700, Chuck Lever wrote:
>
>
> On Tue, Sep 29, 2026, at 4:30 PM, Mike Snitzer wrote:
> > On Tue, Sep 29, 2026 at 04:17:50PM -0700, Chuck Lever wrote:
> >>
> >>
> >> On Tue, Sep 29, 2026, at 12:56 PM, Mike Snitzer wrote:
> >> > On Tue, Sep 29, 2026 at 11:27:05AM -0700, Chuck Lever wrote:
> >> >>
> >> >> On Tue, Sep 29, 2026, at 10:34 AM, Mike Snitzer wrote:
> >> >> > Now that NFSD supports NFSD_IO_DIRECT for both READ and WRITE it is
> >> >> > much safer to avoid needless buffered vs direct contention if/when
> >> >> > only one of them has been configured to use NFSD_IO_DIRECT.
> >> >> >
> >> >> > Mixing direct and buffered I/O to the same file causes needless page
> >> >> > cache invalidation and writeback, so although io_cache_read and
> >> >> > io_cache_write remain separate interfaces, writing either one adjusts
> >> >> > the other so that READ and WRITE are never left on opposite sides of
> >> >> > the buffered/direct divide:
> >> >>
> >> >> Jeff and I have been discussing making DIRECT the default for WRITEs and
> >> >> BUFFERED the default for READs. This patch takes us in the opposite
> >> >> direction.
> >> >>
> >> >> Nothing here convinces me that mixing the modes is a bad thing to do.
> >> >> "Needless page cache invalidation" needs some demonstration, and it
> >> >> needs to show why the right thing to do is make it impossible to mix
> >> >> modes rather than explore the issue as one or more bugs that can be
> >> >> fixed. Or... why not let admins explore this for themselves? Where is
> >> >> the hazard and why does it need to be forbidden by the admin UI?
> >> >
> >> > The basis for the interlock is that O_DIRECT mode is intended to avoid
> >> > bloating memory with page cache and the excess CPU burn of managing
> >> > the page cache. Using read caching in conjunction with O_DIRECT
> >> > writes knee-caps the wins of O_DIRECT.
> >>
> >> Jeff and I have never seen that, and it's a counterintuitive result.
> >> Buffered READ with DIRECT WRITE seems to work very well and the server
> >> is easily capable of managing the page cache in this case, since
> >> reclaiming a clean page doesn't mean having to flush dirty data.
> >> Evicting clean pages is not slow.
> >>
> >> So I'd like to see a quantification of the penalties in this mixed
> >> mode before adjudicating it as a hazard. That is, you might be right,
> >> but so far I've seen no direct evidence that having a substantial page
> >> cache presence is a general deficit for WRITEs. If there are certain
> >> cases where even caching READs is a problem, then by all means, set the
> >> READ IO mode to DIRECT too in those cases.
> >>
> >> I don't see this as an argument for forcing the server's mode setting
> >> in every case.
> >
> > Yes, I understand and I've dropped the interlock patch in v2 (just posted).
>
> Please give reviewers a chance to digest and review, as requested in
> Documentation/filesystems/nfs/nfsd-maintainer-entry-profile.rst :
>
> "As always, please avoid reposting series revisions more than once
> every 24 hours."
>
> Trust me, it saves a lot of confusion.
>
>
> >> Based on this description, it seems to me the client has full visibility
> >> of the data to be pushed back and how it's sharded, it has information
> >> about the network RTT (that's where the real throughput impact is for
> >> COMMIT), and it has control over the selection of UNSTABLE vs. FILE_SYNC.
> >>
> >> The problem here might be that when a large WRITE payload is sharded
> >> across multiple servers, the client still thinks it is sending them via
> >> UNSTABLE WRITES to one server, and plans for only one COMMIT after the
> >> server completes the WRITEs.
> >>
> >> But for the pNFS scenario, the client sends one UNSTABLE WRITE followed
> >> by a COMMIT to each server. If the sharded WRITES are all single RPCs
> >> to distinct servers, then they should each be FILE_SYNC.
> >>
> >> The client can be smarter about how it writes data back to multiple
> >> servers, can't it?
> >
> > When the server is operating in direct mode it isn't something exposed
> > to the client. That the client could be smarter and/or already has
> > adequate controls to achieve the same FILE_SYNC result is besides the
> > point. The point is, if the server has already done the work then it
> > should, within reason, convey as much back to the client to elide
> > COMMIT work that isn't needed.
>
> Your point assumes the other patches in the series are applied to make
> DIRECT UNSTABLE WRITEs completely persistent.
>
> The current IOCB flags do not include IOCB_DSYNC on a DIRECT UNSTABLE
> WRITE for a very good reason: that makes them slower and more expensive.
> The current server logic is working exactly as we designed it last year,
> and I'm not enthusiastic about changing that. I expect at least one
> other reviewer will have a similar reaction, once he sobers up from
> ALPSS.
You haven't reviewed the changes, yet you are passing judgement.
Please review. I think you'll find I have only made things better.
And as for Mr ALPSS, the "NFSD: do not use direct I/O for a READ
smaller than its alignment" and "NFSD: only split a direct-mode WRITE
for a worthwhile direct middle" patches were motivated by his feedback
last year! (I have been carrying them since then, albeit in a less
polished form than I have now made available in this series).
> If the client wants to avoid the COMMIT, the standing rule is to send a
> FILE_SYNC WRITE. That benefits all WRITE I/O modes on the server.
I haven't precluded DIRECT UNSTABLE WRITE from operating as it does
now _at all_.
I've merely exposed additional optional controls that require opt-in.
And provided a natural evolution to what we have that has proven
beneficial.
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH 02/10] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well
2026-09-29 17:34 [PATCH 00/10] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
2026-09-29 17:34 ` [PATCH 01/10] NFSD: interlock the use of NFSD_IO_DIRECT for NFS READ and WRITE Mike Snitzer
@ 2026-09-29 17:34 ` Mike Snitzer
2026-09-29 17:34 ` [PATCH 03/10] NFSD: only split a direct-mode WRITE for a worthwhile direct middle Mike Snitzer
` (7 subsequent siblings)
9 siblings, 0 replies; 17+ messages in thread
From: Mike Snitzer @ 2026-09-29 17:34 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: linux-nfs
nfsd_write_dio_iters_init() issues the DIO-aligned middle of a
misaligned WRITE with IOCB_DIRECT. A file system is free to service
that request with buffered I/O instead: xfs_file_write_iter() retries
via xfs_file_buffered_write() with the same kiocb whenever
xfs_file_dio_write() returns -ENOTBLK, which iomap_dio_rw() produces
when the invalidation it runs before issuing the direct I/O cannot drop
page cache overlapping the range.
That happens routinely in a streaming misaligned WRITE workload with
more than one request in flight. Request N's direct middle ends inside
the same page that request N+1's buffered prefix starts in. If N+1's
prefix dirties that page between N's filemap_write_and_wait_range() and
N's invalidate_inode_pages2_range(), the invalidation returns -EBUSY,
iomap returns -ENOTBLK, and XFS re-issues N's entire middle, tens of
kilobytes, as a normal cached buffered write. Nothing ever drops those
pages. On a device with dma_alignment=3 (XFS then reports
dio_mem_align=4, which XDR-aligned RPC payloads always satisfy) this is
the only remaining path by which a stream of small misaligned WRITEs in
DIRECT mode fills the page cache.
Set IOCB_DONTCACHE on the middle segment alongside IOCB_DIRECT when the
file system supports FOP_DONTCACHE. The direct path ignores the flag,
and generic_write_sync() only issues a harmless flusher kick; but if
the file system falls back to buffered I/O the write is now DONTCACHE
rather than cached. The nfsd_write_direct trace event is unchanged
since it keys on IOCB_DIRECT; the fallback itself remains observable
via the iomap_dio_invalidate_fail event.
Fixes: 06c5c97293e3 ("NFSD: Implement NFSD_IO_DIRECT for NFS WRITE")
Assisted-by: Claude:claude-fable-5-1
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
Documentation/filesystems/nfs/nfsd-io-modes.rst | 10 ++++++++++
fs/nfsd/vfs.c | 9 +++++++++
2 files changed, 19 insertions(+)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index a8e277bcb08b1..c15983e634a3e 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -165,6 +165,16 @@ Misaligned WRITE:
segments because using normal buffered IO offers significant RMW
performance benefit when handling streaming misaligned WRITEs.
+ The O_DIRECT middle segment also carries the DONTCACHE flag. It has
+ no effect while the IO really is O_DIRECT, but a filesystem may
+ decide on its own to service the segment with buffered IO instead
+ (XFS does so when it cannot invalidate page cache that overlaps the
+ segment, which can happen when another WRITE's buffered start or
+ end segment dirties the shared boundary page at the same time).
+ The flag makes that fallback DONTCACHE buffered IO rather than
+ normal buffered IO. Such fallbacks are visible through the
+ iomap_dio_invalidate_fail trace event; see Tracing below.
+
Tracing:
The nfsd_read_direct trace event shows how NFSD expands any
misaligned READ to the next DIO-aligned block (on either end of the
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index f9131827d391e..5b963f0e3b3f2 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1341,6 +1341,15 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
if (iov_iter_bvec_offset(&segments[nsegs].iter) & (mem_align - 1))
goto no_dio;
segments[nsegs].flags |= IOCB_DIRECT;
+ /*
+ * Also mark the direct middle DONTCACHE: the file system may fall
+ * back to buffered I/O on its own (e.g. XFS on -ENOTBLK when it
+ * cannot invalidate page cache that a concurrent buffered prefix or
+ * suffix of an adjacent WRITE just dirtied), and it reuses this kiocb
+ * to do so. On the direct path itself the flag is inert.
+ */
+ if (nf->nf_file->f_op->fop_flags & FOP_DONTCACHE)
+ segments[nsegs].flags |= IOCB_DONTCACHE;
nsegs++;
if (suffix)
--
2.52.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [PATCH 03/10] NFSD: only split a direct-mode WRITE for a worthwhile direct middle
2026-09-29 17:34 [PATCH 00/10] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
2026-09-29 17:34 ` [PATCH 01/10] NFSD: interlock the use of NFSD_IO_DIRECT for NFS READ and WRITE Mike Snitzer
2026-09-29 17:34 ` [PATCH 02/10] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well Mike Snitzer
@ 2026-09-29 17:34 ` Mike Snitzer
2026-09-29 17:34 ` [PATCH 04/10] NFSD: do not use direct I/O for a READ smaller than its alignment Mike Snitzer
` (6 subsequent siblings)
9 siblings, 0 replies; 17+ messages in thread
From: Mike Snitzer @ 2026-09-29 17:34 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: linux-nfs
nfsd_write_dio_iters_init() decides how a WRITE issued in a direct mode
is carried out: split into a buffered prefix, a direct middle and a
buffered suffix, or issued whole as a single buffered segment.
Only split when the split buys direct I/O for a worthwhile middle. Add
direct_misaligned_num_pages (debugfs knob, default 2) and decline to
split when the aligned middle is smaller than that many pages and the
WRITE has a prefix or a suffix. A WRITE whose payload memory is
misaligned for the block device already declines: no segment of it can
be direct I/O, so a split would only turn one buffered write into
three.
Decide every segment's flags in nfsd_write_dio_iters_init() rather than
in the write loop. The buffered prefix and suffix of a split WRITE,
and the whole WRITE when it is not split, are issued IOCB_DONTCACHE
when the file system supports FOP_DONTCACHE and as plain buffered I/O
otherwise: the operator chose a direct mode to keep NFSD out of the
page cache, so a WRITE that cannot be direct gets the next best thing.
Document it in the "Misaligned WRITE" section of nfsd-io-modes.rst.
Suggested-by: Christoph Hellwig <hch@infradead.org>
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
.../filesystems/nfs/nfsd-io-modes.rst | 17 ++++-
fs/nfsd/debugfs.c | 14 ++++
fs/nfsd/nfsd.h | 1 +
fs/nfsd/vfs.c | 73 +++++++++++++------
4 files changed, 79 insertions(+), 26 deletions(-)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index c15983e634a3e..f4e7cee5ee159 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -161,9 +161,9 @@ Misaligned WRITE:
middle and end as needed. The large middle segment is DIO-aligned
and the start and/or end are misaligned. Buffered IO is used for the
misaligned segments and O_DIRECT is used for the middle DIO-aligned
- segment. DONTCACHE buffered IO is _not_ used for the misaligned
- segments because using normal buffered IO offers significant RMW
- performance benefit when handling streaming misaligned WRITEs.
+ segment. If the underlying filesystem supports FOP_DONTCACHE, the
+ misaligned segments use DONTCACHE buffered IO so that their pages
+ are dropped from the page cache once written back.
The O_DIRECT middle segment also carries the DONTCACHE flag. It has
no effect while the IO really is O_DIRECT, but a filesystem may
@@ -175,6 +175,17 @@ Misaligned WRITE:
normal buffered IO. Such fallbacks are visible through the
iomap_dio_invalidate_fail trace event; see Tracing below.
+ Whenever no part of a WRITE can use O_DIRECT, the whole WRITE is
+ issued as a single DONTCACHE buffered IO (normal buffered IO if the
+ filesystem lacks FOP_DONTCACHE). This covers: a filesystem that
+ advertises no DIO alignment requirements at all; a WRITE smaller
+ than the larger of the offset and memory alignments; a WRITE whose
+ DIO-aligned middle segment is smaller than
+ /sys/kernel/debug/nfsd/direct_misaligned_num_pages pages (default 2)
+ while also having a misaligned start or end; and a WRITE whose
+ payload memory is not aligned to the block device's dma_alignment,
+ which rules out O_DIRECT for the middle segment as well.
+
Tracing:
The nfsd_read_direct trace event shows how NFSD expands any
misaligned READ to the next DIO-aligned block (on either end of the
diff --git a/fs/nfsd/debugfs.c b/fs/nfsd/debugfs.c
index 995872a1a7f87..afa41bf2d8189 100644
--- a/fs/nfsd/debugfs.c
+++ b/fs/nfsd/debugfs.c
@@ -170,6 +170,17 @@ void nfsd_debugfs_exit(void)
nfsd_top_dir = NULL;
}
+/*
+ * /sys/kernel/debug/nfsd/direct_misaligned_num_pages
+ *
+ * The smallest DIO-aligned middle segment, in pages, that is worth
+ * splitting a misaligned direct-mode WRITE into three segments for. A
+ * WRITE whose middle is smaller than this, and which has a misaligned
+ * start or end, is issued as a single buffered segment instead.
+ *
+ * Default 2. Not yet tuned by benchmarking.
+ */
+
void nfsd_debugfs_init(void)
{
nfsd_top_dir = debugfs_create_dir("nfsd", NULL);
@@ -182,6 +193,9 @@ void nfsd_debugfs_init(void)
debugfs_create_file("io_cache_write", 0644, nfsd_top_dir, NULL,
&nfsd_io_cache_write_fops);
+
+ debugfs_create_u32("direct_misaligned_num_pages", 0644, nfsd_top_dir,
+ &nfsd_direct_misaligned_num_pages);
#ifdef CONFIG_NFSD_V4
debugfs_create_bool("delegated_timestamps", 0644, nfsd_top_dir,
&nfsd_delegts_enabled);
diff --git a/fs/nfsd/nfsd.h b/fs/nfsd/nfsd.h
index a145294c59c87..a2d72434160af 100644
--- a/fs/nfsd/nfsd.h
+++ b/fs/nfsd/nfsd.h
@@ -142,6 +142,7 @@ enum {
extern u64 nfsd_io_cache_read __read_mostly;
extern u64 nfsd_io_cache_write __read_mostly;
+extern u32 nfsd_direct_misaligned_num_pages __read_mostly;
extern int nfsd_max_blksize;
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 5b963f0e3b3f2..e3ce66bce00d4 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -53,6 +53,7 @@
bool nfsd_disable_splice_read __read_mostly;
u64 nfsd_io_cache_read __read_mostly = NFSD_IO_BUFFERED;
u64 nfsd_io_cache_write __read_mostly = NFSD_IO_BUFFERED;
+u32 nfsd_direct_misaligned_num_pages __read_mostly = 2;
/**
* nfserrno - Map Linux errnos to NFS errnos
@@ -1302,15 +1303,29 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
u32 mem_align = nf->nf_dio_mem_align;
size_t prefix, middle, suffix;
loff_t offset = iocb->ki_pos;
+ unsigned int dontcache_flags = 0;
unsigned int nsegs = 0;
+ if (nf->nf_file->f_op->fop_flags & FOP_DONTCACHE)
+ dontcache_flags = IOCB_DONTCACHE;
+
/*
- * Check if direct I/O is feasible for this write request.
- * If alignments are not available, the write is too small,
- * or no alignment can be found, fall back to buffered I/O.
+ * Whenever direct I/O cannot be used for the WRITE, fall back to a
+ * single DONTCACHE buffered I/O when the file system supports it, so
+ * the WRITE's pages are dropped from the page cache once written
+ * back, and to a single cached buffered I/O otherwise.
+ *
+ * If the file system doesn't advertise any alignment requirements,
+ * don't try to issue direct I/O at all.
*/
- if (unlikely(!mem_align || !offset_align) ||
- unlikely(total < max(offset_align, mem_align)))
+ if (unlikely(!mem_align || !offset_align))
+ goto no_dio;
+
+ /*
+ * If the I/O is smaller than the larger of the memory and logical
+ * offset alignment, no part of it can be direct I/O.
+ */
+ if (unlikely(total < max(offset_align, mem_align)))
goto no_dio;
prefix_end = round_up(offset, offset_align);
@@ -1321,12 +1336,27 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
middle = middle_end - prefix_end;
suffix = orig_end - middle_end;
- if (!middle)
+ /*
+ * If there is no aligned middle section, or the aligned part is too
+ * small to be worth the split (direct_misaligned_num_pages), issue a
+ * single buffered I/O write instead of splitting up the write.
+ */
+ if (!middle ||
+ ((prefix || suffix) &&
+ middle < PAGE_SIZE * nfsd_direct_misaligned_num_pages)) {
goto no_dio;
+ }
- if (prefix)
- nfsd_write_dio_seg_init(&segments[nsegs++], bvec,
+ /*
+ * The prefix and suffix are buffered I/O by definition. Mark them
+ * uncached when possible so their folios are dropped once written
+ * back rather than lingering in the page cache.
+ */
+ if (prefix) {
+ nfsd_write_dio_seg_init(&segments[nsegs], bvec,
nvecs, total, 0, prefix, iocb);
+ segments[nsegs++].flags |= dontcache_flags;
+ }
nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs,
total, prefix, middle, iocb);
@@ -1337,10 +1367,13 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
* bvecs generated from RPC receive buffers are contiguous: After
* the first bvec, all subsequent bvecs start at bv_offset zero
* (page-aligned). Therefore, only the first bvec is checked.
+ *
+ * If the memory is not aligned, direct I/O is impossible for the
+ * middle, so issue the entire write as a single buffered segment:
+ * splitting would only turn one buffered write into three.
*/
if (iov_iter_bvec_offset(&segments[nsegs].iter) & (mem_align - 1))
goto no_dio;
- segments[nsegs].flags |= IOCB_DIRECT;
/*
* Also mark the direct middle DONTCACHE: the file system may fall
* back to buffered I/O on its own (e.g. XFS on -ENOTBLK when it
@@ -1348,20 +1381,21 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
* suffix of an adjacent WRITE just dirtied), and it reuses this kiocb
* to do so. On the direct path itself the flag is inert.
*/
- if (nf->nf_file->f_op->fop_flags & FOP_DONTCACHE)
- segments[nsegs].flags |= IOCB_DONTCACHE;
- nsegs++;
+ segments[nsegs++].flags |= IOCB_DIRECT | dontcache_flags;
- if (suffix)
- nfsd_write_dio_seg_init(&segments[nsegs++], bvec, nvecs, total,
+ if (suffix) {
+ nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs, total,
prefix + middle, suffix, iocb);
+ segments[nsegs++].flags |= dontcache_flags;
+ }
return nsegs;
no_dio:
- /* No DIO alignment possible - pack into single non-DIO segment. */
+ /* No DIO possible - pack into a single uncached (if possible) segment. */
nfsd_write_dio_seg_init(&segments[0], bvec, nvecs, total, 0,
total, iocb);
+ segments[0].flags |= dontcache_flags;
return 1;
}
@@ -1385,16 +1419,9 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
if (kiocb->ki_flags & IOCB_DIRECT)
trace_nfsd_write_direct(rqstp, fhp, kiocb->ki_pos,
segments[i].iter.count);
- else {
+ else
trace_nfsd_write_vector(rqstp, fhp, kiocb->ki_pos,
segments[i].iter.count);
- /*
- * Mark the I/O buffer as evict-able to reduce
- * memory contention.
- */
- if (nf->nf_file->f_op->fop_flags & FOP_DONTCACHE)
- kiocb->ki_flags |= IOCB_DONTCACHE;
- }
expected = iov_iter_count(&segments[i].iter);
--
2.52.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [PATCH 04/10] NFSD: do not use direct I/O for a READ smaller than its alignment
2026-09-29 17:34 [PATCH 00/10] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
` (2 preceding siblings ...)
2026-09-29 17:34 ` [PATCH 03/10] NFSD: only split a direct-mode WRITE for a worthwhile direct middle Mike Snitzer
@ 2026-09-29 17:34 ` Mike Snitzer
2026-09-29 17:34 ` [PATCH 05/10] NFSD: Enable return of an updated stable_how to NFS clients Mike Snitzer
` (5 subsequent siblings)
9 siblings, 0 replies; 17+ messages in thread
From: Mike Snitzer @ 2026-09-29 17:34 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: linux-nfs
nfsd_direct_read() expands a misaligned READ out to DIO-aligned
boundaries: it reads from round_down(offset, dio_read_offset_align) to
round_up(offset + count, dio_read_offset_align) and returns only the
requested bytes from within that window. When the READ is smaller than
the alignment, that window is always at least one full alignment unit,
and two when the READ straddles a boundary, so a few hundred bytes of
payload can cost a 4K or 64K device read.
Decline direct I/O for those. A READ smaller than
dio_read_offset_align now falls through to the DONTCACHE path, which
issues DONTCACHE buffered I/O when the file system supports
FOP_DONTCACHE and normal buffered I/O otherwise. This mirrors the
WRITE side, which already declines direct I/O for a WRITE smaller than
the larger of its offset and memory alignments.
Only dio_read_offset_align is consulted, because the READ path fills
page-aligned pages from rq_bvec and so has no memory alignment to
satisfy. The threshold only bites when the file system advertises a
large alignment; where it reports 512 almost no READ is excluded.
Document it in the "Misaligned READ" section of nfsd-io-modes.rst.
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
Documentation/filesystems/nfs/nfsd-io-modes.rst | 7 +++++++
fs/nfsd/vfs.c | 6 +++---
2 files changed, 10 insertions(+), 3 deletions(-)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index f4e7cee5ee159..bc1c0f1a7b7ca 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -156,6 +156,13 @@ Misaligned READ:
verified to have proper offset/len (logical_block_size) and
dma_alignment checking.
+ A READ smaller than dio_read_offset_align is not issued as O_DIRECT
+ at all. Expanding it would read a whole alignment unit, or two when
+ the READ straddles a boundary, to return those few bytes. Such a
+ READ is issued as DONTCACHE buffered IO instead (normal buffered IO
+ if the filesystem lacks FOP_DONTCACHE), mirroring the WRITE that is
+ smaller than its own alignment.
+
Misaligned WRITE:
If NFSD_IO_DIRECT is used, split any misaligned WRITE into a start,
middle and end as needed. The large middle segment is DIO-aligned
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index e3ce66bce00d4..1d2b03cb42963 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1186,7 +1186,7 @@ __be32 nfsd_iter_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
unsigned int base, u32 *eof)
{
struct file *file = nf->nf_file;
- unsigned long v, total;
+ unsigned long v, total = *count;
struct iov_iter iter;
struct kiocb kiocb;
ssize_t host_err;
@@ -1199,7 +1199,8 @@ __be32 nfsd_iter_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
break;
case NFSD_IO_DIRECT:
/* When dio_read_offset_align is zero, dio is not supported */
- if (nf->nf_dio_read_offset_align && !rqstp->rq_res.page_len)
+ if (nf->nf_dio_read_offset_align && !rqstp->rq_res.page_len &&
+ total >= nf->nf_dio_read_offset_align)
return nfsd_direct_read(rqstp, fhp, nf, offset,
count, eof);
fallthrough;
@@ -1212,7 +1213,6 @@ __be32 nfsd_iter_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
kiocb.ki_pos = offset;
v = 0;
- total = *count;
while (total && v < rqstp->rq_maxpages &&
rqstp->rq_next_page < rqstp->rq_page_end) {
len = min_t(size_t, total, PAGE_SIZE - base);
--
2.52.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [PATCH 05/10] NFSD: Enable return of an updated stable_how to NFS clients
2026-09-29 17:34 [PATCH 00/10] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
` (3 preceding siblings ...)
2026-09-29 17:34 ` [PATCH 04/10] NFSD: do not use direct I/O for a READ smaller than its alignment Mike Snitzer
@ 2026-09-29 17:34 ` Mike Snitzer
2026-09-29 17:34 ` [PATCH 06/10] NFSD: let a direct-mode WRITE raise stable_how and elide the client's COMMIT Mike Snitzer
` (4 subsequent siblings)
9 siblings, 0 replies; 17+ messages in thread
From: Mike Snitzer @ 2026-09-29 17:34 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: linux-nfs
From: Chuck Lever <chuck.lever@oracle.com>
NFSv3 and newer protocols enable clients to perform a two-phase
WRITE. A client requests an UNSTABLE WRITE, which sends dirty data
to the NFS server, but does not persist it. The server replies that
it performed the UNSTABLE WRITE, and the client is then obligated to
follow up with a COMMIT request before it can remove the dirty data
from its own page cache. The COMMIT reply is the client's guarantee
that the written data has been persisted on the server.
The purpose of this protocol design is to enable clients to send
a large amount of data via multiple WRITE requests to a server, and
then wait for persistence just once. The server is able to start
persisting the data as soon as it gets it, to shorten the length of
time the client has to wait for the final COMMIT to complete.
It's also possible for the server to respond to an UNSTABLE WRITE
request in a way that indicates that the data was persisted anyway.
In that case, the client can skip the COMMIT and remove the dirty
data from its memory immediately. NetApp filers, for example, do
this because they have a battery-backed cache and can guarantee that
written data is persisted quickly and immediately.
NFSD has never implemented this kind of promotion. UNSTABLE WRITE
requests are unconditionally treated as UNSTABLE. However, in a
subsequent patch, nfsd_vfs_write() will be able to promote an
UNSTABLE WRITE to be a FILE_SYNC WRITE. This will be because NFSD
will handle some WRITE requests locally with O_DIRECT, which
persists written data immediately. The FILE_SYNC WRITE response
indicates to the client that no follow-up COMMIT is necessary.
This patch prepares for that change by making the @iocb_flags
argument of nfsd_write() and nfsd_vfs_write() bi-directional. A
caller passes in the IOCB_* flags that express the stability its
client asked for; on return the argument holds the flags that were
actually satisfied. Each protocol version maps that result back to
its own on-the-wire value when it encodes the WRITE reply, using the
new nfsd3_stable_how() and nfsd4_stable_how() helpers, so that NFS
stable_how values stay out of NFSD's generic VFS API. No behavior
change is expected.
[snitzer: reworked onto the @iocb_flags argument introduced by
commit 6dcddbb70b08 ("NFSD: Replace nfsd_write()'s "stable" argument with "iocb_flags""),
which was applied after this patch was first posted.
The original passed a "u32 *stable_how" instead. Carrying the value as
IOCB_* flags keeps the NFSv3 XDR value out of the VFS API, and lets a
later patch raise the achieved stability with a plain bitwise OR rather
than an ordering comparison. The Reviewed-by tags from the original
posting are dropped because the argument's type and direction changed.]
Signed-off-by: Chuck Lever <chuck.lever@oracle.com>
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
fs/nfsd/nfs3proc.c | 16 +++++++++++++---
fs/nfsd/nfs4proc.c | 15 +++++++++++++--
fs/nfsd/nfsproc.c | 3 ++-
fs/nfsd/vfs.c | 18 +++++++++++-------
fs/nfsd/vfs.h | 4 ++--
fs/nfsd/xdr3.h | 2 +-
6 files changed, 42 insertions(+), 16 deletions(-)
diff --git a/fs/nfsd/nfs3proc.c b/fs/nfsd/nfs3proc.c
index 60cd01b6a37d2..f06573759ad6f 100644
--- a/fs/nfsd/nfs3proc.c
+++ b/fs/nfsd/nfs3proc.c
@@ -102,6 +102,15 @@ static int nfsd3_iocb_flags(enum nfs3_stable_how how)
}
}
+static enum nfs3_stable_how nfsd3_stable_how(int iocb_flags)
+{
+ if (iocb_flags & IOCB_SYNC)
+ return NFS_FILE_SYNC;
+ if (iocb_flags & IOCB_DSYNC)
+ return NFS_DATA_SYNC;
+ return NFS_UNSTABLE;
+}
+
static __be32 nfsd3_map_status(__be32 status)
{
switch (status) {
@@ -299,6 +308,7 @@ nfsd3_proc_write(struct svc_rqst *rqstp)
struct nfsd3_writeargs *argp = rqstp->rq_argp;
struct nfsd3_writeres *resp = rqstp->rq_resp;
unsigned long cnt = argp->len;
+ int iocb_flags;
dprintk("nfsd: WRITE(3) %s %d bytes at %Lu%s\n",
SVCFH_fmt(&argp->fh),
@@ -312,11 +322,11 @@ nfsd3_proc_write(struct svc_rqst *rqstp)
return rpc_success;
fh_copy(&resp->fh, &argp->fh);
- resp->committed = argp->stable;
+ iocb_flags = nfsd3_iocb_flags(argp->stable);
resp->status = nfsd_write(rqstp, &resp->fh, argp->offset,
- &argp->payload, &cnt,
- nfsd3_iocb_flags(resp->committed),
+ &argp->payload, &cnt, &iocb_flags,
resp->verf);
+ resp->committed = nfsd3_stable_how(iocb_flags);
resp->count = cnt;
resp->status = nfsd3_map_status(resp->status);
return rpc_success;
diff --git a/fs/nfsd/nfs4proc.c b/fs/nfsd/nfs4proc.c
index bb74eef439388..051d900581609 100644
--- a/fs/nfsd/nfs4proc.c
+++ b/fs/nfsd/nfs4proc.c
@@ -109,6 +109,15 @@ static const struct nfsd_access_maps nfsd4_access_maps = {
.other = nfsd4_otheraccess,
};
+static enum stable_how4 nfsd4_stable_how(int iocb_flags)
+{
+ if (iocb_flags & IOCB_SYNC)
+ return FILE_SYNC4;
+ if (iocb_flags & IOCB_DSYNC)
+ return DATA_SYNC4;
+ return UNSTABLE4;
+}
+
static int nfsd4_iocb_flags(enum stable_how4 how)
{
switch (how) {
@@ -1432,6 +1441,7 @@ nfsd4_write(struct svc_rqst *rqstp, struct nfsd4_compound_state *cstate,
struct nfsd_file *nf = NULL;
__be32 status = nfs_ok;
unsigned long cnt;
+ int iocb_flags;
if (write->wr_offset > (u64)OFFSET_MAX ||
write->wr_offset + write->wr_buflen > (u64)OFFSET_MAX)
@@ -1450,11 +1460,12 @@ nfsd4_write(struct svc_rqst *rqstp, struct nfsd4_compound_state *cstate,
nfs4_put_stid(stid);
}
- write->wr_how_written = write->wr_stable_how;
+ iocb_flags = nfsd4_iocb_flags(write->wr_stable_how);
status = nfsd_vfs_write(rqstp, &cstate->current_fh, nf,
write->wr_offset, &write->wr_payload,
- &cnt, nfsd4_iocb_flags(write->wr_how_written),
+ &cnt, &iocb_flags,
(__be32 *)write->wr_verifier.data);
+ write->wr_how_written = nfsd4_stable_how(iocb_flags);
nfsd_file_put(nf);
write->wr_bytes_written = cnt;
diff --git a/fs/nfsd/nfsproc.c b/fs/nfsd/nfsproc.c
index 09d3608398250..a847b35006781 100644
--- a/fs/nfsd/nfsproc.c
+++ b/fs/nfsd/nfsproc.c
@@ -274,6 +274,7 @@ nfsd_proc_write(struct svc_rqst *rqstp)
struct nfsd_writeargs *argp = rqstp->rq_argp;
struct nfsd_attrstat *resp = rqstp->rq_resp;
unsigned long cnt = argp->len;
+ int iocb_flags = IOCB_DSYNC;
dprintk("nfsd: WRITE %s %u bytes at %d\n",
SVCFH_fmt(&argp->fh),
@@ -281,7 +282,7 @@ nfsd_proc_write(struct svc_rqst *rqstp)
fh_copy(&resp->fh, &argp->fh);
resp->status = nfsd_write(rqstp, &resp->fh, argp->offset,
- &argp->payload, &cnt, IOCB_DSYNC, NULL);
+ &argp->payload, &cnt, &iocb_flags, NULL);
if (resp->status == nfs_ok)
resp->status = fh_getattr(&resp->fh, &resp->stat);
else if (resp->status == nfserr_jukebox)
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 1d2b03cb42963..82eba97656c4e 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1444,7 +1444,9 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
* @offset: Byte offset of start
* @payload: xdr_buf containing the write payload
* @cnt: IN: number of bytes to write, OUT: number of bytes actually written
- * @iocb_flags: VFS IOCB_* flags expressing the requested write stability
+ * @iocb_flags: IN: VFS IOCB_* flags expressing the requested write
+ * stability; OUT: the flags actually satisfied, which may be
+ * higher than requested
* @verf: NFS WRITE verifier
*
* Upon return, caller must invoke fh_put on @fhp.
@@ -1456,7 +1458,7 @@ __be32
nfsd_vfs_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
struct nfsd_file *nf, loff_t offset,
const struct xdr_buf *payload, unsigned long *cnt,
- int iocb_flags, __be32 *verf)
+ int *iocb_flags, __be32 *verf)
{
struct nfsd_net *nn = net_generic(SVC_NET(rqstp), nfsd_net_id);
struct file *file = nf->nf_file;
@@ -1493,11 +1495,11 @@ nfsd_vfs_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
exp = fhp->fh_export;
if (!EX_ISSYNC(exp))
- iocb_flags = 0;
+ *iocb_flags = 0;
init_sync_kiocb(&kiocb, file);
kiocb.ki_pos = offset;
if (likely(!fhp->fh_use_wgather))
- kiocb.ki_flags |= iocb_flags;
+ kiocb.ki_flags |= *iocb_flags;
nvecs = xdr_buf_to_bvec(rqstp->rq_bvec, rqstp->rq_maxpages, payload);
if (nvecs < 0) {
@@ -1538,7 +1540,7 @@ nfsd_vfs_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
goto out_nfserr;
}
- if (iocb_flags && fhp->fh_use_wgather) {
+ if (*iocb_flags && fhp->fh_use_wgather) {
host_err = wait_for_concurrent_writes(file);
if (host_err < 0)
nfsd_maybe_reset_write_verifier(nn, rqstp, host_err);
@@ -1629,7 +1631,9 @@ __be32 nfsd_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
* @offset: Byte offset of start
* @payload: xdr_buf containing the write payload
* @cnt: IN: number of bytes to write, OUT: number of bytes actually written
- * @iocb_flags: VFS IOCB_* flags expressing the requested write stability
+ * @iocb_flags: IN: VFS IOCB_* flags expressing the requested write
+ * stability; OUT: the flags actually satisfied, which may be
+ * higher than requested
* @verf: NFS WRITE verifier
*
* Upon return, caller must invoke fh_put on @fhp.
@@ -1640,7 +1644,7 @@ __be32 nfsd_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
__be32
nfsd_write(struct svc_rqst *rqstp, struct svc_fh *fhp, loff_t offset,
const struct xdr_buf *payload, unsigned long *cnt,
- int iocb_flags, __be32 *verf)
+ int *iocb_flags, __be32 *verf)
{
struct nfsd_file *nf;
__be32 err;
diff --git a/fs/nfsd/vfs.h b/fs/nfsd/vfs.h
index f0cb184643f2f..6b352ca7f02b7 100644
--- a/fs/nfsd/vfs.h
+++ b/fs/nfsd/vfs.h
@@ -154,12 +154,12 @@ __be32 nfsd_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
u32 *eof);
__be32 nfsd_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
loff_t offset, const struct xdr_buf *payload,
- unsigned long *cnt, int iocb_flags,
+ unsigned long *cnt, int *iocb_flags,
__be32 *verf);
__be32 nfsd_vfs_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
struct nfsd_file *nf, loff_t offset,
const struct xdr_buf *payload,
- unsigned long *cnt, int iocb_flags,
+ unsigned long *cnt, int *iocb_flags,
__be32 *verf);
__be32 nfsd_readlink(struct svc_rqst *, struct svc_fh *,
char *, int *);
diff --git a/fs/nfsd/xdr3.h b/fs/nfsd/xdr3.h
index cad875d142313..a2a980559fb72 100644
--- a/fs/nfsd/xdr3.h
+++ b/fs/nfsd/xdr3.h
@@ -153,7 +153,7 @@ struct nfsd3_writeres {
__be32 status;
struct svc_fh fh;
unsigned long count;
- int committed;
+ u32 committed;
__be32 verf[2];
};
--
2.52.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [PATCH 06/10] NFSD: let a direct-mode WRITE raise stable_how and elide the client's COMMIT
2026-09-29 17:34 [PATCH 00/10] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
` (4 preceding siblings ...)
2026-09-29 17:34 ` [PATCH 05/10] NFSD: Enable return of an updated stable_how to NFS clients Mike Snitzer
@ 2026-09-29 17:34 ` Mike Snitzer
2026-09-29 17:34 ` [PATCH 07/10] NFSD: persist a synchronous direct-mode WRITE once, after all of its segments Mike Snitzer
` (3 subsequent siblings)
9 siblings, 0 replies; 17+ messages in thread
From: Mike Snitzer @ 2026-09-29 17:34 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: linux-nfs
A WRITE serviced in a direct mode is already persistent when NFSD
replies: the aligned middle is O_DIRECT, and a synchronous WRITE fsyncs
whatever was buffered before the reply is sent. The reply still says
UNSTABLE, so the client dutifully sends a COMMIT for data that is
already on stable storage, and NFSD answers it with an fsync that has
nothing left to write.
Add two io_cache_write modes that issue direct I/O exactly like
NFSD_IO_DIRECT and raise the stable_how of the reply to a floor:
NFSD_IO_DIRECT_WRITE_DATA_SYNC (3): at least NFS_DATA_SYNC
NFSD_IO_DIRECT_WRITE_FILE_SYNC (4): at least NFS_FILE_SYNC
A client that asked for more is left alone, and the WRITE is persisted
to the level the reply reports before that reply is sent. A client told
FILE_SYNC has no reason to COMMIT and sends none.
What this removes is the COMMIT traffic, and it is worth most where a
client's writes are carved into several WRITE RPCs, because each piece
is then committed separately. Measured on a pNFS flexfiles share where
every write straddles two data servers, so every write becomes two WRITE
RPCs and, under NFSD_IO_DIRECT, two COMMITs: the two modes do identical
durability work, one fsync per COMMIT against one fsync per WRITE on
identical WRITE counts, and the COMMIT RPCs alone cost NFSD_IO_DIRECT
31% more server CPU and 42% more client CPU for the same bytes, about
20 us of server CPU per COMMIT plus a client cost that grows with the
range committed. Where only one write in 22 is split the same effect
is a couple of cores on each side and no resolvable throughput
difference, and writes that fit a single RPC send no COMMIT in either
mode.
Choose 4 when the export services WRITEs in a direct mode and clients
split their writes; choose 3 to promise only the data, leaving a client
that needs metadata durability to COMMIT for it. Both are inert for
READ, and for a WRITE the client already marked FILE_SYNC.
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
.../filesystems/nfs/nfsd-io-modes.rst | 17 ++++++++--
fs/nfsd/debugfs.c | 5 +++
fs/nfsd/nfsd.h | 2 ++
fs/nfsd/vfs.c | 33 +++++++++++++++++--
4 files changed, 52 insertions(+), 5 deletions(-)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index bc1c0f1a7b7ca..548eca7cdb518 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -25,12 +25,14 @@ Based on the configured settings, NFSD's IO will either be:
- cached using page cache (NFSD_IO_BUFFERED=0)
- cached but removed from page cache on completion (NFSD_IO_DONTCACHE=1)
- not cached stable_how=NFS_UNSTABLE (NFSD_IO_DIRECT=2)
+- not cached stable_how=NFS_DATA_SYNC (NFSD_IO_DIRECT_WRITE_DATA_SYNC=3)
+- not cached stable_how=NFS_FILE_SYNC (NFSD_IO_DIRECT_WRITE_FILE_SYNC=4)
-To set an NFSD IO mode, write a supported value (0 - 2) to the
+To set an NFSD IO mode, write a supported value (0 - 4) to the
corresponding IO operation's debugfs interface, e.g.::
echo 2 > /sys/kernel/debug/nfsd/io_cache_read
- echo 2 > /sys/kernel/debug/nfsd/io_cache_write
+ echo 4 > /sys/kernel/debug/nfsd/io_cache_write
To check which IO mode NFSD is using for READ or WRITE, simply read the
corresponding IO operation's debugfs interface, e.g.::
@@ -38,6 +40,17 @@ corresponding IO operation's debugfs interface, e.g.::
cat /sys/kernel/debug/nfsd/io_cache_read
cat /sys/kernel/debug/nfsd/io_cache_write
+The two NFSD_IO_DIRECT_WRITE_*_SYNC modes raise the stable_how of every
+WRITE to at least NFS_DATA_SYNC or NFS_FILE_SYNC, persist the WRITE
+accordingly before replying, and return the raised value to the client;
+a client that asked for a higher stable_how is left alone. With
+NFSD_IO_DIRECT_WRITE_FILE_SYNC the client sends no COMMIT. Against
+NFSD_IO_DIRECT the durability work is the same, one fsync per WRITE
+instead of one per COMMIT; what NFSD_IO_DIRECT adds is the COMMIT RPCs
+themselves, tens of microseconds of server CPU each plus a client cost
+that grows with the range committed, which matters in proportion to how
+many of a client's WRITEs need a COMMIT.
+
If you experiment with NFSD's IO modes on a recent kernel and have
interesting results, please report them to linux-nfs@vger.kernel.org
diff --git a/fs/nfsd/debugfs.c b/fs/nfsd/debugfs.c
index afa41bf2d8189..603a608b03c54 100644
--- a/fs/nfsd/debugfs.c
+++ b/fs/nfsd/debugfs.c
@@ -113,6 +113,9 @@ DEFINE_DEBUGFS_ATTRIBUTE(nfsd_io_cache_read_fops, nfsd_io_cache_read_get,
* Contents:
* %0: NFS WRITE will use buffered IO
* %1: NFS WRITE will use dontcache (buffered IO w/ dropbehind)
+ * %2: NFS WRITE will use direct IO with stable_how=NFS_UNSTABLE
+ * %3: NFS WRITE will use direct IO with stable_how=NFS_DATA_SYNC
+ * %4: NFS WRITE will use direct IO with stable_how=NFS_FILE_SYNC
*
* This setting takes immediate effect for all NFS versions,
* all exports, and in all NFSD net namespaces.
@@ -132,6 +135,8 @@ static int nfsd_io_cache_write_set(void *data, u64 val)
case NFSD_IO_BUFFERED:
case NFSD_IO_DONTCACHE:
case NFSD_IO_DIRECT:
+ case NFSD_IO_DIRECT_WRITE_DATA_SYNC:
+ case NFSD_IO_DIRECT_WRITE_FILE_SYNC:
nfsd_io_cache_write = val;
/*
* Adjust nfsd_io_cache_{read,write} to avoid
diff --git a/fs/nfsd/nfsd.h b/fs/nfsd/nfsd.h
index a2d72434160af..135e319e378d4 100644
--- a/fs/nfsd/nfsd.h
+++ b/fs/nfsd/nfsd.h
@@ -138,6 +138,8 @@ enum {
NFSD_IO_BUFFERED,
NFSD_IO_DONTCACHE,
NFSD_IO_DIRECT,
+ NFSD_IO_DIRECT_WRITE_DATA_SYNC,
+ NFSD_IO_DIRECT_WRITE_FILE_SYNC,
};
extern u64 nfsd_io_cache_read __read_mostly;
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 82eba97656c4e..924c5992dc32e 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1399,17 +1399,42 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
return 1;
}
+/*
+ * Raise the stability of this WRITE to at least @floor_iocb_flags, and
+ * record what was achieved in @iocb_flags so the reply can report it.
+ * A client that asked for more is left alone.
+ */
+static void
+nfsd_write_raise_stability(int floor_iocb_flags, struct kiocb *kiocb,
+ int *iocb_flags)
+{
+ if ((*iocb_flags & floor_iocb_flags) == floor_iocb_flags)
+ return; /* already at or above the floor */
+
+ *iocb_flags |= floor_iocb_flags;
+ kiocb->ki_flags |= floor_iocb_flags;
+}
+
static noinline_for_stack int
nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
- struct nfsd_file *nf, unsigned int nvecs,
+ struct nfsd_file *nf, int *iocb_flags, unsigned int nvecs,
unsigned long *cnt, struct kiocb *kiocb)
{
struct nfsd_write_dio_seg segments[3];
+ int floor_iocb_flags = 0;
struct file *file = nf->nf_file;
unsigned int nsegs, i;
ssize_t host_err;
size_t expected;
+ if (nfsd_io_cache_write == NFSD_IO_DIRECT_WRITE_FILE_SYNC)
+ floor_iocb_flags = IOCB_DSYNC | IOCB_SYNC;
+ else if (nfsd_io_cache_write == NFSD_IO_DIRECT_WRITE_DATA_SYNC)
+ floor_iocb_flags = IOCB_DSYNC;
+ if (floor_iocb_flags)
+ nfsd_write_raise_stability(floor_iocb_flags, kiocb,
+ iocb_flags);
+
nsegs = nfsd_write_dio_iters_init(nf, rqstp->rq_bvec, nvecs,
kiocb, *cnt, segments);
@@ -1513,8 +1538,10 @@ nfsd_vfs_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
switch (nfsd_io_cache_write) {
case NFSD_IO_DIRECT:
- host_err = nfsd_direct_write(rqstp, fhp, nf, nvecs,
- cnt, &kiocb);
+ case NFSD_IO_DIRECT_WRITE_DATA_SYNC:
+ case NFSD_IO_DIRECT_WRITE_FILE_SYNC:
+ host_err = nfsd_direct_write(rqstp, fhp, nf, iocb_flags,
+ nvecs, cnt, &kiocb);
break;
case NFSD_IO_DONTCACHE:
if (file->f_op->fop_flags & FOP_DONTCACHE)
--
2.52.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [PATCH 07/10] NFSD: persist a synchronous direct-mode WRITE once, after all of its segments
2026-09-29 17:34 [PATCH 00/10] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
` (5 preceding siblings ...)
2026-09-29 17:34 ` [PATCH 06/10] NFSD: let a direct-mode WRITE raise stable_how and elide the client's COMMIT Mike Snitzer
@ 2026-09-29 17:34 ` Mike Snitzer
2026-09-29 17:34 ` [PATCH 08/10] NFSD: keep boundary page of a split direct-mode WRITE until both writers complete Mike Snitzer
` (2 subsequent siblings)
9 siblings, 0 replies; 17+ messages in thread
From: Mike Snitzer @ 2026-09-29 17:34 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: linux-nfs
nfsd_direct_write() may issue a WRITE as up to three segments: a
buffered prefix, a direct middle and a buffered suffix. For a
FILE_SYNC or DATA_SYNC WRITE (from the client, or as the floor imposed
by NFSD_IO_DIRECT_WRITE_{DATA,FILE}_SYNC) the kiocb carries IOCB_DSYNC
and every segment inherits it, so generic_write_sync() runs a range
fsync after each segment: up to three cache flushes and log forces per
WRITE, and each one writes back and drops the boundary page it just
touched.
Strip IOCB_DSYNC and IOCB_SYNC from the per-segment flags and persist
the WRITE once with vfs_fsync_range() over the bytes actually written,
after the last segment. Durability is unchanged: the reply is not sent
until the fsync completes, and datasync mirrors the previous per-segment
choice (IOCB_SYNC present means metadata too). An fsync failure is
returned like a write failure.
Besides the fewer flushes, this puts the sync under NFSD's control,
which the next change uses to keep the boundary pages of a split WRITE
cached until the partner WRITE completes them.
Assisted-by: Claude:claude-fable-5-1
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
fs/nfsd/vfs.c | 20 +++++++++++++++++++-
1 file changed, 19 insertions(+), 1 deletion(-)
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 924c5992dc32e..827dfe0b5dac7 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1423,6 +1423,8 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
struct nfsd_write_dio_seg segments[3];
int floor_iocb_flags = 0;
struct file *file = nf->nf_file;
+ loff_t start = kiocb->ki_pos;
+ bool sync, datasync;
unsigned int nsegs, i;
ssize_t host_err;
size_t expected;
@@ -1435,12 +1437,21 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
nfsd_write_raise_stability(floor_iocb_flags, kiocb,
iocb_flags);
+ /*
+ * A synchronous WRITE (client FILE_SYNC/DATA_SYNC, or a floor set by
+ * the IO mode) is persisted once, after all of its segments, rather
+ * than by generic_write_sync() after each segment: one cache flush
+ * and log force instead of up to three.
+ */
+ sync = kiocb->ki_flags & IOCB_DSYNC;
+ datasync = !(kiocb->ki_flags & IOCB_SYNC);
+
nsegs = nfsd_write_dio_iters_init(nf, rqstp->rq_bvec, nvecs,
kiocb, *cnt, segments);
*cnt = 0;
for (i = 0; i < nsegs; i++) {
- kiocb->ki_flags = segments[i].flags;
+ kiocb->ki_flags = segments[i].flags & ~(IOCB_DSYNC | IOCB_SYNC);
if (kiocb->ki_flags & IOCB_DIRECT)
trace_nfsd_write_direct(rqstp, fhp, kiocb->ki_pos,
segments[i].iter.count);
@@ -1458,6 +1469,13 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
break; /* partial write */
}
+ if (sync && *cnt) {
+ host_err = vfs_fsync_range(file, start, start + *cnt - 1,
+ datasync);
+ if (host_err < 0)
+ return host_err;
+ }
+
return 0;
}
--
2.52.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [PATCH 08/10] NFSD: keep boundary page of a split direct-mode WRITE until both writers complete
2026-09-29 17:34 [PATCH 00/10] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
` (6 preceding siblings ...)
2026-09-29 17:34 ` [PATCH 07/10] NFSD: persist a synchronous direct-mode WRITE once, after all of its segments Mike Snitzer
@ 2026-09-29 17:34 ` Mike Snitzer
2026-09-29 17:34 ` [PATCH 09/10] NFSD: add direct_misaligned_dontcache debugfs knob Mike Snitzer
2026-09-29 17:34 ` [PATCH 10/10] NFSD: add tracing for how direct-mode READ and WRITE are serviced Mike Snitzer
9 siblings, 0 replies; 17+ messages in thread
From: Mike Snitzer @ 2026-09-29 17:34 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: linux-nfs
A direct-mode WRITE whose ends are not aligned for direct I/O is split
into a buffered prefix, a direct middle and a buffered suffix, and the
prefix and suffix are issued IOCB_DONTCACHE so their pages are dropped
once written back. The page holding such a segment is shared by exactly
two WRITEs, the one ending in it and the one starting in it, which may
arrive in either order, from different clients, at the same time. If it
is dropped as soon as the first writer's data is written back, the
second has to read the block from disk before it can complete it: one
4 KiB read per WRITE. With many clients interleaving small misaligned
O_DIRECT records into one shared file, that was 0.81 device reads per
record (585 GiB read while writing 8.1 TiB).
Keep the page until both writers have had it, with no state in NFSD and
no knowledge of how clients stripe their I/O. Immediately before a
boundary segment is written, nfsd_write_dio_boundary_claim() puts an
empty folio in the page cache if there is not one there already. The
folio is inserted unmarked, which is what makes this work: a DONTCACHE
write only marks a folio it allocated itself, so neither writer's write
marks this one and it never enters WB_DONTCACHE_DIRTY. That counter is
what the DONTCACHE writeback kick targets, and a boundary page counted
there is written back, and then dropped, in the gap between the two
WRITEs that share it.
The add is atomic, so of two concurrent partners exactly one gets the
folio. The other one finds it (-EEXIST) and is therefore the second of
the two: it completes the page, and once its data is in it
nfsd_write_dio_boundary_complete() marks the folio dropbehind. The mark
has to come after the write, because a clean marked folio is dropped by
whatever writeback completes next, and it has to stay unaccounted,
because counting a folio marked while dirty hands the page straight
back to the kick. Whichever writeback cleans the page then drops it:
the WRITE's own sync for FILE_SYNC and DATA_SYNC, the flusher or the
client's COMMIT for UNSTABLE.
A WRITE that cannot use direct I/O at all (memory-misaligned payload,
or too small for a direct middle) is issued as one buffered DONTCACHE
segment. Its first and last pages are shared with the neighbouring
WRITEs in the same way, and are claimed the same way.
Measured with 32 interleaved writers over an emulated 4Kn NVMe,
io_cache_write=4, direct_misaligned_dontcache=Y, each arm starting from
a freshly made file system. 47008-byte records, which split: 704 device
reads for 42895 WRITEs, against 45144 for 45664 WRITEs with the
mechanism compiled out. 6000-byte records, which never get a direct
middle and so are one buffered segment whose end pages are shared: 385
reads against 196589. The DONTCACHE flusher stays idle in both, and the
pages are dropped: 81 of the written file's 524066 pages are still
resident afterwards.
On a four-server pNFS flexfiles rig, three clients writing 47008-byte
records for 240 s against 640 nfsd threads per server, compared with
the same servers before this change: device reads fall from 0.014-0.015
per record to 0.003, page-cache growth over the write phase falls from
about 11 GiB to 2.9 GiB, and write throughput rises by 2.4% to 4.9%
across NFSD_IO_DIRECT and both stable_how floors, with read throughput
unchanged. Records of that size always take the split path, since their
direct middle is at least 38818 bytes, so the figures are the split
path alone; the whole-WRITE fallback needs records below about 12 KB to
come into play. Keeping every boundary page cached instead, which a
later commit makes possible with a direct_misaligned_dontcache=N
debugfs knob, grows the page cache by about 630 GiB over the same
runs.
Reported-by: Jonathan Flynn <jonathan.flynn@hammerspace.com>
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
.../filesystems/nfs/nfsd-io-modes.rst | 28 ++++
fs/nfsd/vfs.c | 131 ++++++++++++++++--
2 files changed, 151 insertions(+), 8 deletions(-)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index 548eca7cdb518..7263570668260 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -185,6 +185,34 @@ Misaligned WRITE:
misaligned segments use DONTCACHE buffered IO so that their pages
are dropped from the page cache once written back.
+ A FILE_SYNC or DATA_SYNC WRITE (requested by the client, or imposed
+ as the floor by NFSD_IO_DIRECT_WRITE_FILE_SYNC and
+ NFSD_IO_DIRECT_WRITE_DATA_SYNC) is persisted once, after all of its
+ segments have been written, rather than after each segment.
+
+ The page holding a start or end segment is shared by exactly two
+ WRITEs, the one ending in it and the one starting in it, which may
+ arrive in either order, from different clients, and at the same
+ time. Both segments are issued as DONTCACHE buffered IO, so the page
+ would be dropped as soon as the first writer's data is written back,
+ leaving the second to read it back. Immediately before writing a
+ start or end segment NFSD therefore puts an empty page in the page
+ cache if there is not one already. That page is not marked "drop
+ behind" when it is created, so it does not count towards the
+ DONTCACHE writeback backlog and the writeback kick does not write it
+ back, and drop it, between the two WRITEs that share it.
+
+ Whichever WRITE finds the page already there is the second of the
+ two: it completes the page and marks it "drop behind" once its data
+ is in it, so whichever writeback cleans it afterwards (the WRITE's
+ own sync for FILE_SYNC or DATA_SYNC, the flusher or the client's
+ COMMIT for UNSTABLE) drops it. A WRITE that gets no direct middle at
+ all is issued as a single buffered DONTCACHE segment, and its first
+ and last pages are shared and handled the same way. The retained
+ page cache is the set of half-written boundary pages, which grows
+ with how far concurrent writers drift apart, not with bytes
+ written.
+
The O_DIRECT middle segment also carries the DONTCACHE flag. It has
no effect while the IO really is O_DIRECT, but a filesystem may
decide on its own to service the segment with buffered IO instead
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 827dfe0b5dac7..5fd850a29694f 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1271,6 +1271,8 @@ static int wait_for_concurrent_writes(struct file *file)
struct nfsd_write_dio_seg {
struct iov_iter iter;
int flags;
+ bool boundary; /* prefix or suffix of a split */
+ bool edges; /* buffered fallback: partial end pages */
};
static unsigned long
@@ -1290,6 +1292,73 @@ nfsd_write_dio_seg_init(struct nfsd_write_dio_seg *segment,
iov_iter_advance(&segment->iter, start);
iov_iter_truncate(&segment->iter, len);
segment->flags = iocb->ki_flags;
+ segment->boundary = false;
+ segment->edges = false;
+}
+
+/**
+ * nfsd_write_dio_boundary_claim - claim the page a boundary segment shares
+ * @file: the file being written
+ * @pos: any byte offset within the page
+ *
+ * The page holding a misaligned prefix or suffix is shared by exactly two
+ * WRITEs, the one ending in it and the one starting in it, which may arrive
+ * in either order, from different clients, at the same time. Put an empty
+ * folio there if there is not one already: it is inserted unmarked, so
+ * neither writer's DONTCACHE write marks it (a DONTCACHE write only marks a
+ * folio it allocated itself) and it never enters WB_DONTCACHE_DIRTY, which
+ * is what would otherwise arm the DONTCACHE writeback kick and have the
+ * flusher write the page back, and drop it, between the two WRITEs.
+ *
+ * The add is atomic, so of two concurrent partners exactly one gets the
+ * folio. If it cannot be allocated the segment is an ordinary DONTCACHE
+ * write and the page is dropped after writeback as before.
+ *
+ * Return: true if this WRITE is the second of the two and must call
+ * nfsd_write_dio_boundary_complete() once its segment has been written.
+ */
+static bool
+nfsd_write_dio_boundary_claim(struct file *file, loff_t pos)
+{
+ struct address_space *mapping = file->f_mapping;
+ pgoff_t index = pos >> PAGE_SHIFT;
+ gfp_t gfp = mapping_gfp_mask(mapping);
+ struct folio *folio;
+ int err;
+
+ folio = filemap_alloc_folio(gfp, 0, NULL);
+ if (!folio)
+ return false;
+ err = filemap_add_folio(mapping, folio, index, gfp);
+ if (!err) {
+ /* First writer: the page is in place, waiting for the partner. */
+ folio_unlock(folio);
+ folio_put(folio);
+ return false;
+ }
+ folio_put(folio);
+ return err == -EEXIST;
+}
+
+/*
+ * The second writer's data is in the page: mark it so the next writeback
+ * that cleans it drops it. This must follow the write, because a clean
+ * marked folio is dropped by whatever writeback completes next, and it uses
+ * folio_set_dropbehind() rather than an accounted setter, because counting a
+ * folio marked while dirty arms the DONTCACHE writeback kick and the page is
+ * then written back, and dropped, before its partner has written it.
+ */
+static void
+nfsd_write_dio_boundary_complete(struct file *file, loff_t pos)
+{
+ struct address_space *mapping = file->f_mapping;
+ struct folio *folio;
+
+ folio = __filemap_get_folio(mapping, pos >> PAGE_SHIFT, FGP_DONTCACHE, 0);
+ if (IS_ERR(folio))
+ return;
+ folio_set_dropbehind(folio);
+ folio_put(folio);
}
static unsigned int
@@ -1348,14 +1417,17 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
}
/*
- * The prefix and suffix are buffered I/O by definition. Mark them
- * uncached when possible so their folios are dropped once written
- * back rather than lingering in the page cache.
+ * The prefix and suffix are buffered I/O by definition. Each shares
+ * its page with the neighbouring WRITE; see
+ * nfsd_write_dio_boundary_claim(), which nfsd_direct_write() calls right
+ * before issuing each of them, for how the page is held for the
+ * partner and dropped once both have written it.
*/
if (prefix) {
nfsd_write_dio_seg_init(&segments[nsegs], bvec,
nvecs, total, 0, prefix, iocb);
- segments[nsegs++].flags |= dontcache_flags;
+ segments[nsegs].flags |= dontcache_flags;
+ segments[nsegs++].boundary = !!dontcache_flags;
}
nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs,
@@ -1386,16 +1458,23 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
if (suffix) {
nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs, total,
prefix + middle, suffix, iocb);
- segments[nsegs++].flags |= dontcache_flags;
+ segments[nsegs].flags |= dontcache_flags;
+ segments[nsegs++].boundary = !!dontcache_flags;
}
return nsegs;
no_dio:
- /* No DIO possible - pack into a single uncached (if possible) segment. */
+ /*
+ * No DIO possible - pack into a single buffered segment. Where it
+ * does not start or end on a page boundary, its first and last pages
+ * are shared with the neighbouring WRITEs like a prefix or suffix and
+ * are held the same way (nfsd_write_dio_boundary_claim()).
+ */
nfsd_write_dio_seg_init(&segments[0], bvec, nvecs, total, 0,
total, iocb);
segments[0].flags |= dontcache_flags;
+ segments[0].edges = !!dontcache_flags;
return 1;
}
@@ -1423,8 +1502,8 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
struct nfsd_write_dio_seg segments[3];
int floor_iocb_flags = 0;
struct file *file = nf->nf_file;
- loff_t start = kiocb->ki_pos;
- bool sync, datasync;
+ loff_t start = kiocb->ki_pos, seg_pos, seg_last;
+ bool sync, datasync, complete_first, complete_last;
unsigned int nsegs, i;
ssize_t host_err;
size_t expected;
@@ -1461,9 +1540,45 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
expected = iov_iter_count(&segments[i].iter);
+ /*
+ * Claim the boundary page immediately before writing it, not
+ * when the WRITE is split. Nothing is held across the write:
+ * the claim leaves the page in the page cache unmarked, and
+ * claiming it earlier would give the partner WRITE the time a
+ * direct middle takes to complete the page and have it
+ * dropped before this segment writes it.
+ */
+ seg_pos = kiocb->ki_pos;
+ seg_last = seg_pos + expected - 1;
+ complete_first = false;
+ complete_last = false;
+ if (segments[i].boundary) {
+ complete_first = nfsd_write_dio_boundary_claim(file,
+ seg_pos);
+ } else if (segments[i].edges) {
+ /*
+ * A whole-WRITE buffered segment shares its first page
+ * with the previous WRITE if it does not start on a page
+ * boundary, and its last page with the next one if it
+ * does not end on one; a page it covers entirely is its
+ * own.
+ */
+ if (seg_pos & ~PAGE_MASK)
+ complete_first = nfsd_write_dio_boundary_claim(
+ file, seg_pos);
+ if (((seg_last + 1) & ~PAGE_MASK) &&
+ (seg_last >> PAGE_SHIFT) != (seg_pos >> PAGE_SHIFT))
+ complete_last = nfsd_write_dio_boundary_claim(
+ file, seg_last);
+ }
+
host_err = vfs_iocb_iter_write(file, kiocb, &segments[i].iter);
if (host_err < 0)
return host_err;
+ if (complete_first)
+ nfsd_write_dio_boundary_complete(file, seg_pos);
+ if (complete_last)
+ nfsd_write_dio_boundary_complete(file, seg_last);
*cnt += host_err;
if (host_err < (ssize_t)expected)
break; /* partial write */
--
2.52.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [PATCH 09/10] NFSD: add direct_misaligned_dontcache debugfs knob
2026-09-29 17:34 [PATCH 00/10] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
` (7 preceding siblings ...)
2026-09-29 17:34 ` [PATCH 08/10] NFSD: keep boundary page of a split direct-mode WRITE until both writers complete Mike Snitzer
@ 2026-09-29 17:34 ` Mike Snitzer
2026-09-29 17:34 ` [PATCH 10/10] NFSD: add tracing for how direct-mode READ and WRITE are serviced Mike Snitzer
9 siblings, 0 replies; 17+ messages in thread
From: Mike Snitzer @ 2026-09-29 17:34 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: linux-nfs
From: Jonathan Flynn <jonathan.flynn@hammerspace.com>
The parts of a direct-mode WRITE that cannot be direct I/O, the
misaligned prefix and suffix of a split WRITE and the whole WRITE when
it is not split, are issued IOCB_DONTCACHE so their pages are dropped
once written back. That is the right default for the workloads a direct
mode is chosen for, but it is a policy rather than a requirement: a
workload that reads back what it just wrote, or that keeps rewriting the
same partial pages, is better served by those pages staying in the page
cache.
Add a bool debugfs knob, /sys/kernel/debug/nfsd/direct_misaligned_dontcache
(default Y), to choose between the two:
Y: IOCB_DONTCACHE when the file system supports it, with a split's
boundary page kept in the page cache until both WRITEs sharing it
have written it (nfsd_write_dio_boundary_claim()).
N: ordinary cached buffered I/O; nothing is claimed or marked, and the
pages stay until reclaim.
The direct middle and the once-per-WRITE persist are unaffected. The
knob is sampled once per WRITE, so a change takes effect immediately and
without a remount. It sits beside io_cache_read and io_cache_write, the
other controls over how NFSD issues its I/O.
Signed-off-by: Jonathan Flynn <jonathan.flynn@hammerspace.com>
[snitzer: documented in nfsd-io-modes.rst]
[snitzer: switched from using modparam to debugfs knob]
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
.../filesystems/nfs/nfsd-io-modes.rst | 9 ++++++
fs/nfsd/debugfs.c | 21 ++++++++++++++
fs/nfsd/nfsd.h | 1 +
fs/nfsd/vfs.c | 28 ++++++++++++-------
4 files changed, 49 insertions(+), 10 deletions(-)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index 7263570668260..b2685dfadbffb 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -213,6 +213,15 @@ Misaligned WRITE:
with how far concurrent writers drift apart, not with bytes
written.
+ Whether those pages are dropped at all is a policy choice, selected
+ by /sys/kernel/debug/nfsd/direct_misaligned_dontcache (default Y).
+ Write N to issue the start and end segments, and the whole-WRITE
+ fallbacks, as ordinary cached buffered IO: nothing is claimed or
+ marked and the pages stay until reclaim, which suits a workload that
+ reads back or rewrites what it just wrote. The O_DIRECT middle
+ segment is unaffected. The knob is sampled once per WRITE, so a
+ change takes effect immediately.
+
The O_DIRECT middle segment also carries the DONTCACHE flag. It has
no effect while the IO really is O_DIRECT, but a filesystem may
decide on its own to service the segment with buffered IO instead
diff --git a/fs/nfsd/debugfs.c b/fs/nfsd/debugfs.c
index 603a608b03c54..1ae2597b9c34a 100644
--- a/fs/nfsd/debugfs.c
+++ b/fs/nfsd/debugfs.c
@@ -186,6 +186,24 @@ void nfsd_debugfs_exit(void)
* Default 2. Not yet tuned by benchmarking.
*/
+/*
+ * /sys/kernel/debug/nfsd/direct_misaligned_dontcache
+ *
+ * How a direct-mode WRITE issues the I/O that cannot be direct: the
+ * misaligned start and end of a split WRITE, and the whole WRITE when it
+ * is not split.
+ *
+ * Contents:
+ * Y: DONTCACHE when the filesystem supports it, with the boundary page
+ * of a split kept in the page cache only until both WRITEs sharing
+ * it have written it
+ * N: ordinary cached buffered IO, left in the page cache until
+ * reclaim, for A/B comparison against the DONTCACHE path
+ *
+ * Sampled once per WRITE, so it takes effect immediately. The direct
+ * middle segment is unaffected.
+ */
+
void nfsd_debugfs_init(void)
{
nfsd_top_dir = debugfs_create_dir("nfsd", NULL);
@@ -201,6 +219,9 @@ void nfsd_debugfs_init(void)
debugfs_create_u32("direct_misaligned_num_pages", 0644, nfsd_top_dir,
&nfsd_direct_misaligned_num_pages);
+
+ debugfs_create_bool("direct_misaligned_dontcache", 0644, nfsd_top_dir,
+ &nfsd_direct_misaligned_dontcache);
#ifdef CONFIG_NFSD_V4
debugfs_create_bool("delegated_timestamps", 0644, nfsd_top_dir,
&nfsd_delegts_enabled);
diff --git a/fs/nfsd/nfsd.h b/fs/nfsd/nfsd.h
index 135e319e378d4..dff979ac370ba 100644
--- a/fs/nfsd/nfsd.h
+++ b/fs/nfsd/nfsd.h
@@ -145,6 +145,7 @@ enum {
extern u64 nfsd_io_cache_read __read_mostly;
extern u64 nfsd_io_cache_write __read_mostly;
extern u32 nfsd_direct_misaligned_num_pages __read_mostly;
+extern bool nfsd_direct_misaligned_dontcache __read_mostly;
extern int nfsd_max_blksize;
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 5fd850a29694f..7896e2e6c5855 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -54,6 +54,7 @@ bool nfsd_disable_splice_read __read_mostly;
u64 nfsd_io_cache_read __read_mostly = NFSD_IO_BUFFERED;
u64 nfsd_io_cache_write __read_mostly = NFSD_IO_BUFFERED;
u32 nfsd_direct_misaligned_num_pages __read_mostly = 2;
+bool nfsd_direct_misaligned_dontcache __read_mostly = true;
/**
* nfserrno - Map Linux errnos to NFS errnos
@@ -1373,16 +1374,21 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
size_t prefix, middle, suffix;
loff_t offset = iocb->ki_pos;
unsigned int dontcache_flags = 0;
+ unsigned int buffered_flags;
unsigned int nsegs = 0;
if (nf->nf_file->f_op->fop_flags & FOP_DONTCACHE)
dontcache_flags = IOCB_DONTCACHE;
+ /* Buffered segments follow the knob; the direct middle does not. */
+ buffered_flags = READ_ONCE(nfsd_direct_misaligned_dontcache) ?
+ dontcache_flags : 0;
/*
* Whenever direct I/O cannot be used for the WRITE, fall back to a
- * single DONTCACHE buffered I/O when the file system supports it, so
- * the WRITE's pages are dropped from the page cache once written
- * back, and to a single cached buffered I/O otherwise.
+ * single DONTCACHE buffered I/O when the file system supports it (and
+ * nfsd_direct_misaligned_dontcache is set), so the WRITE's pages are
+ * dropped from the page cache once written back, and to a single
+ * cached buffered I/O otherwise.
*
* If the file system doesn't advertise any alignment requirements,
* don't try to issue direct I/O at all.
@@ -1421,13 +1427,15 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
* its page with the neighbouring WRITE; see
* nfsd_write_dio_boundary_claim(), which nfsd_direct_write() calls right
* before issuing each of them, for how the page is held for the
- * partner and dropped once both have written it.
+ * partner and dropped once both have written it. With
+ * nfsd_direct_misaligned_dontcache=N both are plain cached writes and
+ * nothing is held or dropped: the pages stay until reclaim.
*/
if (prefix) {
nfsd_write_dio_seg_init(&segments[nsegs], bvec,
nvecs, total, 0, prefix, iocb);
- segments[nsegs].flags |= dontcache_flags;
- segments[nsegs++].boundary = !!dontcache_flags;
+ segments[nsegs].flags |= buffered_flags;
+ segments[nsegs++].boundary = !!buffered_flags;
}
nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs,
@@ -1458,8 +1466,8 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
if (suffix) {
nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs, total,
prefix + middle, suffix, iocb);
- segments[nsegs].flags |= dontcache_flags;
- segments[nsegs++].boundary = !!dontcache_flags;
+ segments[nsegs].flags |= buffered_flags;
+ segments[nsegs++].boundary = !!buffered_flags;
}
return nsegs;
@@ -1473,8 +1481,8 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
*/
nfsd_write_dio_seg_init(&segments[0], bvec, nvecs, total, 0,
total, iocb);
- segments[0].flags |= dontcache_flags;
- segments[0].edges = !!dontcache_flags;
+ segments[0].flags |= buffered_flags;
+ segments[0].edges = !!buffered_flags;
return 1;
}
--
2.52.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [PATCH 10/10] NFSD: add tracing for how direct-mode READ and WRITE are serviced
2026-09-29 17:34 [PATCH 00/10] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
` (8 preceding siblings ...)
2026-09-29 17:34 ` [PATCH 09/10] NFSD: add direct_misaligned_dontcache debugfs knob Mike Snitzer
@ 2026-09-29 17:34 ` Mike Snitzer
9 siblings, 0 replies; 17+ messages in thread
From: Mike Snitzer @ 2026-09-29 17:34 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton; +Cc: linux-nfs
When io_cache_read or io_cache_write selects a direct I/O mode, a
request may still be serviced with DONTCACHE buffered or normal
buffered I/O, and a WRITE may be split into up to three segments that
each take a different path depending on the request's offset/length
alignment, the payload's memory alignment, and the filesystem's
FOP_DONTCACHE support. Until now the only visibility was
nfsd_{read,write}_direct for O_DIRECT and nfsd_{read,write}_vector for
everything else, so the DONTCACHE and cached buffered cases were
indistinguishable and the reason a WRITE was not issued as direct I/O
was not recorded at all.
Add:
- nfsd_write_dio_split, emitted once per direct-mode WRITE before any
segment is issued. It records the advertised offset and memory
alignments, the payload's memory offset within its page, the start,
middle and end segment sizes, the number of segments issued, and a
disposition describing which path was taken (direct, no_alignment,
too_small, no_middle, mem_misaligned), with a dontcache flag set
when the buffered segments were issued IOCB_DONTCACHE, so a run with
direct_misaligned_dontcache=N is distinguishable from one with the
file system lacking FOP_DONTCACHE.
- nfsd_write_dontcache and nfsd_read_dontcache, emitted for WRITE
segments and READs serviced with IOCB_DONTCACHE. nfsd_write_vector
and nfsd_read_vector now fire only for normal buffered I/O.
The disposition enum lives in vfs.h so trace.h can reference it, and
TRACE_DEFINE_ENUM entries are provided for user-space decoding.
Document the new events in nfsd-io-modes.rst, including the iomap
iomap_dio_invalidate_fail event that reveals an O_DIRECT segment which
the filesystem silently serviced with buffered I/O.
Assisted-by: Claude:claude-fable-5-1
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
.../filesystems/nfs/nfsd-io-modes.rst | 41 ++++++++-
fs/nfsd/trace.h | 90 +++++++++++++++++++
fs/nfsd/vfs.c | 44 ++++++---
fs/nfsd/vfs.h | 22 +++++
4 files changed, 183 insertions(+), 14 deletions(-)
diff --git a/Documentation/filesystems/nfs/nfsd-io-modes.rst b/Documentation/filesystems/nfs/nfsd-io-modes.rst
index b2685dfadbffb..c69d0c880ea77 100644
--- a/Documentation/filesystems/nfs/nfsd-io-modes.rst
+++ b/Documentation/filesystems/nfs/nfsd-io-modes.rst
@@ -246,21 +246,56 @@ Misaligned WRITE:
Tracing:
The nfsd_read_direct trace event shows how NFSD expands any
misaligned READ to the next DIO-aligned block (on either end of the
- original READ, as needed).
+ original READ, as needed). A READ that is serviced with buffered IO
+ instead emits nfsd_read_dontcache (DONTCACHE buffered IO) or
+ nfsd_read_vector (normal buffered IO).
This combination of trace events is useful for READs::
echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_read_vector/enable
+ echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_read_dontcache/enable
echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_read_direct/enable
echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_read_io_done/enable
echo 1 > /sys/kernel/tracing/events/xfs/xfs_file_direct_read/enable
- The nfsd_write_direct trace event shows how NFSD splits a given
- misaligned WRITE into a DIO-aligned middle segment.
+ The nfsd_write_dio_split trace event is emitted once per WRITE
+ serviced in a DIRECT IO mode, before any IO is issued, and records
+ how the WRITE was split: the offset and memory alignments the
+ filesystem advertised, the memory offset of the WRITE payload, the
+ sizes of the start, middle and end segments, the number of segments
+ actually issued, and a disposition naming the reason::
+
+ direct aligned middle segment uses O_DIRECT
+ mem_misaligned payload memory is misaligned; one buffered segment
+ no_alignment filesystem advertises no DIO alignment; one
+ buffered segment
+ too_small WRITE is smaller than the larger of the two
+ alignments; one buffered segment
+ no_middle no (or too small) aligned middle; one buffered
+ segment
+
+ Whether those buffered segments are DONTCACHE or normal buffered IO
+ is reported separately, by dontcache=1 or dontcache=0, because it is
+ the same answer for every disposition: the buffered segments are
+ DONTCACHE when the filesystem supports FOP_DONTCACHE. For the direct
+ disposition it describes the prefix and suffix of the split, the
+ middle being O_DIRECT.
+
+ Each segment then emits one of nfsd_write_direct (O_DIRECT),
+ nfsd_write_dontcache (DONTCACHE buffered IO) or nfsd_write_vector
+ (normal buffered IO) with the segment's offset and length.
This combination of trace events is useful for WRITEs::
echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_write_opened/enable
+ echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_write_dio_split/enable
echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_write_direct/enable
+ echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_write_dontcache/enable
+ echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_write_vector/enable
echo 1 > /sys/kernel/tracing/events/nfsd/nfsd_write_io_done/enable
echo 1 > /sys/kernel/tracing/events/xfs/xfs_file_direct_write/enable
+ echo 1 > /sys/kernel/tracing/events/iomap/iomap_dio_invalidate_fail/enable
+
+ iomap_dio_invalidate_fail indicates an O_DIRECT middle segment that
+ the filesystem silently serviced with normal buffered IO because it
+ could not invalidate overlapping page cache first.
diff --git a/fs/nfsd/trace.h b/fs/nfsd/trace.h
index 2ae7f150a72ce..0b8aa6c5d0772 100644
--- a/fs/nfsd/trace.h
+++ b/fs/nfsd/trace.h
@@ -15,6 +15,7 @@
#include <trace/misc/fsnotify.h>
#include <trace/misc/sunrpc.h>
+#include "vfs.h"
#include "export.h"
#include "nfsfh.h"
#include "xdr4.h"
@@ -501,6 +502,7 @@ DEFINE_EVENT(nfsd_io_class, nfsd_##name, \
DEFINE_NFSD_IO_EVENT(read_start);
DEFINE_NFSD_IO_EVENT(read_splice);
DEFINE_NFSD_IO_EVENT(read_vector);
+DEFINE_NFSD_IO_EVENT(read_dontcache);
DEFINE_NFSD_IO_EVENT(read_direct);
DEFINE_NFSD_IO_EVENT(read_io_done);
DEFINE_NFSD_IO_EVENT(read_done);
@@ -508,11 +510,99 @@ DEFINE_NFSD_IO_EVENT(write_start);
DEFINE_NFSD_IO_EVENT(write_opened);
DEFINE_NFSD_IO_EVENT(write_direct);
DEFINE_NFSD_IO_EVENT(write_vector);
+DEFINE_NFSD_IO_EVENT(write_dontcache);
DEFINE_NFSD_IO_EVENT(write_io_done);
DEFINE_NFSD_IO_EVENT(write_done);
DEFINE_NFSD_IO_EVENT(commit_start);
DEFINE_NFSD_IO_EVENT(commit_done);
+TRACE_DEFINE_ENUM(NFSD_WRITE_DIO_DIRECT);
+TRACE_DEFINE_ENUM(NFSD_WRITE_DIO_MEM_MISALIGNED);
+TRACE_DEFINE_ENUM(NFSD_WRITE_DIO_NO_ALIGN);
+TRACE_DEFINE_ENUM(NFSD_WRITE_DIO_TOO_SMALL);
+TRACE_DEFINE_ENUM(NFSD_WRITE_DIO_NO_MIDDLE);
+
+#define show_nfsd_write_dio_disposition(x) \
+ __print_symbolic(x, \
+ { NFSD_WRITE_DIO_DIRECT, "direct" }, \
+ { NFSD_WRITE_DIO_MEM_MISALIGNED, "mem_misaligned" }, \
+ { NFSD_WRITE_DIO_NO_ALIGN, "no_alignment" }, \
+ { NFSD_WRITE_DIO_TOO_SMALL, "too_small" }, \
+ { NFSD_WRITE_DIO_NO_MIDDLE, "no_middle" })
+
+/**
+ * nfsd_write_dio_split - how an NFSD_IO_DIRECT WRITE was split
+ *
+ * Emitted once per WRITE handled by nfsd_direct_write(), before any
+ * segment is issued. @prefix/@middle/@suffix are the byte counts of the
+ * three candidate segments (zero when not computed); @mem_offset is the
+ * offset within its page of the first byte of the WRITE payload, from
+ * which the middle segment's memory alignment is (@mem_offset + @prefix)
+ * masked by (@mem_align - 1). @dontcache is whether the WRITE's buffered
+ * segments carry IOCB_DONTCACHE: the single segment of every non-direct
+ * disposition, and the prefix and suffix of a "direct" one. It is ORed
+ * into @disposition by the caller, a tracepoint being limited to twelve
+ * arguments, and split back out into its own field here.
+ */
+TRACE_EVENT(nfsd_write_dio_split,
+ TP_PROTO(struct svc_rqst *rqstp,
+ struct svc_fh *fhp,
+ u64 offset,
+ u32 len,
+ u32 offset_align,
+ u32 mem_align,
+ u32 mem_offset,
+ u32 prefix,
+ u32 middle,
+ u32 suffix,
+ u32 nsegs,
+ unsigned int disposition),
+ TP_ARGS(rqstp, fhp, offset, len, offset_align, mem_align, mem_offset,
+ prefix, middle, suffix, nsegs, disposition),
+ TP_STRUCT__entry(
+ __field(u32, xid)
+ __field(u32, fh_hash)
+ __field(u64, offset)
+ __field(u32, len)
+ __field(u32, offset_align)
+ __field(u32, mem_align)
+ __field(u32, mem_offset)
+ __field(u32, prefix)
+ __field(u32, middle)
+ __field(u32, suffix)
+ __field(u32, nsegs)
+ __field(unsigned int, disposition)
+ __field(bool, dontcache)
+ ),
+ TP_fast_assign(
+ __entry->xid = be32_to_cpu(rqstp->rq_xid);
+ __entry->fh_hash = knfsd_fh_hash(&fhp->fh_handle);
+ __entry->offset = offset;
+ __entry->len = len;
+ __entry->offset_align = offset_align;
+ __entry->mem_align = mem_align;
+ __entry->mem_offset = mem_offset;
+ __entry->prefix = prefix;
+ __entry->middle = middle;
+ __entry->suffix = suffix;
+ __entry->nsegs = nsegs;
+ __entry->disposition = disposition & ~NFSD_WRITE_DIO_DONTCACHE;
+ __entry->dontcache = !!(disposition & NFSD_WRITE_DIO_DONTCACHE);
+ ),
+ TP_printk("xid=0x%08x fh_hash=0x%08x offset=%llu len=%u "
+ "offset_align=%u mem_align=%u mem_offset=%u "
+ "prefix=%u middle=%u suffix=%u nsegs=%u disposition=%s "
+ "dontcache=%u",
+ __entry->xid, __entry->fh_hash,
+ __entry->offset, __entry->len,
+ __entry->offset_align, __entry->mem_align,
+ __entry->mem_offset,
+ __entry->prefix, __entry->middle, __entry->suffix,
+ __entry->nsegs,
+ show_nfsd_write_dio_disposition(__entry->disposition),
+ __entry->dontcache)
+);
+
DECLARE_EVENT_CLASS(nfsd_err_class,
TP_PROTO(struct svc_rqst *rqstp,
struct svc_fh *fhp,
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 7896e2e6c5855..3e7265b8de251 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1226,7 +1226,10 @@ __be32 nfsd_iter_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
base = 0;
}
- trace_nfsd_read_vector(rqstp, fhp, offset, *count - total);
+ if (kiocb.ki_flags & IOCB_DONTCACHE)
+ trace_nfsd_read_dontcache(rqstp, fhp, offset, *count - total);
+ else
+ trace_nfsd_read_vector(rqstp, fhp, offset, *count - total);
iov_iter_bvec(&iter, ITER_DEST, rqstp->rq_bvec, v, *count - total);
host_err = vfs_iocb_iter_read(file, &kiocb, &iter);
return nfsd_finish_read(rqstp, fhp, file, offset, count, eof, host_err);
@@ -1363,7 +1366,8 @@ nfsd_write_dio_boundary_complete(struct file *file, loff_t pos)
}
static unsigned int
-nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
+nfsd_write_dio_iters_init(struct svc_rqst *rqstp, struct svc_fh *fhp,
+ struct nfsd_file *nf, struct bio_vec *bvec,
unsigned int nvecs, struct kiocb *iocb,
unsigned long total,
struct nfsd_write_dio_seg segments[3])
@@ -1371,7 +1375,8 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
u32 offset_align = nf->nf_dio_offset_align;
loff_t prefix_end, orig_end, middle_end;
u32 mem_align = nf->nf_dio_mem_align;
- size_t prefix, middle, suffix;
+ size_t prefix = 0, middle = 0, suffix = 0;
+ enum nfsd_write_dio_disposition disposition;
loff_t offset = iocb->ki_pos;
unsigned int dontcache_flags = 0;
unsigned int buffered_flags;
@@ -1393,15 +1398,19 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
* If the file system doesn't advertise any alignment requirements,
* don't try to issue direct I/O at all.
*/
- if (unlikely(!mem_align || !offset_align))
+ if (unlikely(!mem_align || !offset_align)) {
+ disposition = NFSD_WRITE_DIO_NO_ALIGN;
goto no_dio;
+ }
/*
* If the I/O is smaller than the larger of the memory and logical
* offset alignment, no part of it can be direct I/O.
*/
- if (unlikely(total < max(offset_align, mem_align)))
+ if (unlikely(total < max(offset_align, mem_align))) {
+ disposition = NFSD_WRITE_DIO_TOO_SMALL;
goto no_dio;
+ }
prefix_end = round_up(offset, offset_align);
orig_end = offset + total;
@@ -1419,6 +1428,7 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
if (!middle ||
((prefix || suffix) &&
middle < PAGE_SIZE * nfsd_direct_misaligned_num_pages)) {
+ disposition = NFSD_WRITE_DIO_NO_MIDDLE;
goto no_dio;
}
@@ -1452,8 +1462,10 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
* middle, so issue the entire write as a single buffered segment:
* splitting would only turn one buffered write into three.
*/
- if (iov_iter_bvec_offset(&segments[nsegs].iter) & (mem_align - 1))
+ if (iov_iter_bvec_offset(&segments[nsegs].iter) & (mem_align - 1)) {
+ disposition = NFSD_WRITE_DIO_MEM_MISALIGNED;
goto no_dio;
+ }
/*
* Also mark the direct middle DONTCACHE: the file system may fall
* back to buffered I/O on its own (e.g. XFS on -ENOTBLK when it
@@ -1462,6 +1474,7 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
* to do so. On the direct path itself the flag is inert.
*/
segments[nsegs++].flags |= IOCB_DIRECT | dontcache_flags;
+ disposition = NFSD_WRITE_DIO_DIRECT;
if (suffix) {
nfsd_write_dio_seg_init(&segments[nsegs], bvec, nvecs, total,
@@ -1469,8 +1482,7 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
segments[nsegs].flags |= buffered_flags;
segments[nsegs++].boundary = !!buffered_flags;
}
-
- return nsegs;
+ goto out;
no_dio:
/*
@@ -1483,7 +1495,14 @@ nfsd_write_dio_iters_init(struct nfsd_file *nf, struct bio_vec *bvec,
total, iocb);
segments[0].flags |= buffered_flags;
segments[0].edges = !!buffered_flags;
- return 1;
+ nsegs = 1;
+out:
+ trace_nfsd_write_dio_split(rqstp, fhp, offset, total,
+ offset_align, mem_align, bvec->bv_offset,
+ prefix, middle, suffix, nsegs,
+ disposition | (buffered_flags ?
+ NFSD_WRITE_DIO_DONTCACHE : 0));
+ return nsegs;
}
/*
@@ -1533,8 +1552,8 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
sync = kiocb->ki_flags & IOCB_DSYNC;
datasync = !(kiocb->ki_flags & IOCB_SYNC);
- nsegs = nfsd_write_dio_iters_init(nf, rqstp->rq_bvec, nvecs,
- kiocb, *cnt, segments);
+ nsegs = nfsd_write_dio_iters_init(rqstp, fhp, nf, rqstp->rq_bvec,
+ nvecs, kiocb, *cnt, segments);
*cnt = 0;
for (i = 0; i < nsegs; i++) {
@@ -1542,6 +1561,9 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp,
if (kiocb->ki_flags & IOCB_DIRECT)
trace_nfsd_write_direct(rqstp, fhp, kiocb->ki_pos,
segments[i].iter.count);
+ else if (kiocb->ki_flags & IOCB_DONTCACHE)
+ trace_nfsd_write_dontcache(rqstp, fhp, kiocb->ki_pos,
+ segments[i].iter.count);
else
trace_nfsd_write_vector(rqstp, fhp, kiocb->ki_pos,
segments[i].iter.count);
diff --git a/fs/nfsd/vfs.h b/fs/nfsd/vfs.h
index 6b352ca7f02b7..58f9ae26415dc 100644
--- a/fs/nfsd/vfs.h
+++ b/fs/nfsd/vfs.h
@@ -189,4 +189,26 @@ __be32 nfsd_permission(struct svc_cred *cred, struct svc_export *exp,
void nfsd_filp_close(struct file *fp);
+/*
+ * How nfsd_write_dio_iters_init() disposed of an NFSD_IO_DIRECT WRITE.
+ * "DONTCACHE segment" degrades to a cached segment when the file system
+ * lacks FOP_DONTCACHE.
+ */
+/*
+ * Which exit nfsd_write_dio_iters_init() took. Whether the buffered
+ * segments carry IOCB_DONTCACHE is reported separately, by
+ * nfsd_write_dio_split's @dontcache, because it is the same answer for
+ * every exit below.
+ */
+enum nfsd_write_dio_disposition {
+ NFSD_WRITE_DIO_DIRECT, /* aligned middle uses direct I/O */
+ NFSD_WRITE_DIO_MEM_MISALIGNED, /* payload memory misaligned: one segment */
+ NFSD_WRITE_DIO_NO_ALIGN, /* fs advertises no alignment: one segment */
+ NFSD_WRITE_DIO_TOO_SMALL, /* len < max(offset_align, mem_align): one segment */
+ NFSD_WRITE_DIO_NO_MIDDLE, /* no or tiny aligned middle: one segment */
+
+ /* ORed in: the WRITE's buffered segments carry IOCB_DONTCACHE */
+ NFSD_WRITE_DIO_DONTCACHE = 0x80,
+};
+
#endif /* LINUX_NFSD_VFS_H */
--
2.52.0
^ permalink raw reply related [flat|nested] 17+ messages in thread