From: Mike Snitzer <snitzer@hammerspace.com>
To: Chuck Lever <cel@kernel.org>
Cc: Jeff Layton <jlayton@kernel.org>, linux-nfs@vger.kernel.org
Subject: Re: [PATCH 01/10] NFSD: interlock the use of NFSD_IO_DIRECT for NFS READ and WRITE
Date: Tue, 29 Sep 2026 15:56:53 -0400 [thread overview]
Message-ID: <arwYBaemyvLoG2sT@hammerspace.com> (raw)
In-Reply-To: <37625d4e-9d86-4651-bbc8-73c1c589942d@app.fastmail.com>
On Tue, Sep 29, 2026 at 11:27:05AM -0700, Chuck Lever wrote:
>
> On Tue, Sep 29, 2026, at 10:34 AM, Mike Snitzer wrote:
> > Now that NFSD supports NFSD_IO_DIRECT for both READ and WRITE it is
> > much safer to avoid needless buffered vs direct contention if/when
> > only one of them has been configured to use NFSD_IO_DIRECT.
> >
> > Mixing direct and buffered I/O to the same file causes needless page
> > cache invalidation and writeback, so although io_cache_read and
> > io_cache_write remain separate interfaces, writing either one adjusts
> > the other so that READ and WRITE are never left on opposite sides of
> > the buffered/direct divide:
>
> Jeff and I have been discussing making DIRECT the default for WRITEs and
> BUFFERED the default for READs. This patch takes us in the opposite
> direction.
>
> Nothing here convinces me that mixing the modes is a bad thing to do.
> "Needless page cache invalidation" needs some demonstration, and it
> needs to show why the right thing to do is make it impossible to mix
> modes rather than explore the issue as one or more bugs that can be
> fixed. Or... why not let admins explore this for themselves? Where is
> the hazard and why does it need to be forbidden by the admin UI?
The basis for the interlock is that O_DIRECT mode is intended to avoid
bloating memory with page cache and the excess CPU burn of managing
the page cache. Using read caching in conjunction with O_DIRECT
writes knee-caps the wins of O_DIRECT. But sure, you can do it.. and
I suppose the freedom of choice (and enough rope to hang) is perfectly
fine. I can drop the interlock and send v2 after more time for review
of the other patches (but there were some real bugs fixed in the
interlock patch so will require some care to adjust).
direct_misaligned_dontcache=N shows the excessive 630 GiB page cache
bloat (across 4 servers) I've had to confront with a larger scale test
that really sticks its finger in this particular wound (as documented at
end of the "NFSD: keep boundary page of a split direct-mode WRITE
until both writers complete" patch):
On a four-server pNFS flexfiles rig, three clients writing 47008-byte
records for 240 s against 640 nfsd threads per server, compared with
the same servers before this change: device reads fall from 0.014-0.015
per record to 0.003, page-cache growth over the write phase falls from
about 11 GiB to 2.9 GiB, and write throughput rises by 2.4% to 4.9%
across NFSD_IO_DIRECT and both stable_how floors, with read throughput
unchanged. Records of that size always take the split path, since their
direct middle is at least 38818 bytes, so the figures are the split
path alone; the whole-WRITE fallback needs records below about 12 KB to
come into play. Keeping every boundary page cached instead, which a
later commit makes possible with a direct_misaligned_dontcache=N
debugfs knob, grows the page cache by about 630 GiB over the same
runs.
The "11 GiB" is what's left resident in page cache _without_ the
"NFSD: keep boundary page of a split direct-mode WRITE until both
writers complete". _WITH_ that patch (to make more informed use of
DONTCACHE) it drops to the quoted 2.9 GiB while still benefitting from
RMW avoidance buffered IO makes possible -- pretty fantastic result.
> Later patches in the series assume DONTCACHE is the fallback for
> certain cases.
That's the existing default that current upstream code has. But with
this patchset the fallback is buffered, with buffered DONTCACHE the
default unless debugfs direct_misaligned_dontcache=N configured.
It would be quite bad to remove DONTCACHE support because it actually
does offer pretty solid wins (especially for the workload I quoted
above). If pure buffered IO used (direct_misaligned_dontcache=N)
WRITE performance suffers: "only" 31,297 MiB/s for O_DIRECT +
buffered, whereas with O_DIRECT + DONTCACHE, overall result were (more
runs needed to get stddev):
4 - FILE_SYNC floor 35,567 MiB/s
2 - NFSD_IO_DIRECT 35,302 MiB/s
3 - DATA_SYNC floor 36,163 MiB/s
> Jeff has measured substantial performance deficits for that mode,
> and maybe removing DONTCACHE would be a better direction to take.
But to be clear, I'm not referring to DONTCACHE only mode=1, I'm most
interested in hybrid of O_DIRECT+DONTCACHE (more on that at the end
below).
My series has been fully tested with Jeff's more recent DONTCACHE
commits in place:
88d6f128d06d mm: track DONTCACHE dirty pages per bdi_writeback
f3122ce09a51 mm: kick writeback flusher for IOCB_DONTCACHE with targeted dirty tracking
> The patch that reports the actual stability of a WRITE was dropped
> because it broke something (although I don't remember what).
Haven't seen any issues with it. Been carrying it ever since you
posted it.
> As I recall, even a direct WRITE needs a subsequent COMMIT. Have you
> demonstrated that the client COMMIT is a latency problem or that
> the memory that is pinned on the client has a noticeable impact?
>
> We suspect that it might, but every time I've measured, avoiding
> COMMIT with NFSD has shown no impact on throughput, which is why
> NFSD doesn't already do this optimization.
Yes, in the header for "NFSD: let a direct-mode WRITE raise stable_how
and elide the client's COMMIT":
What this removes is the COMMIT traffic, and it is worth most where a
client's writes are carved into several WRITE RPCs, because each piece
is then committed separately. Measured on a pNFS flexfiles share where
every write straddles two data servers, so every write becomes two WRITE
RPCs and, under NFSD_IO_DIRECT, two COMMITs: the two modes do identical
durability work, one fsync per COMMIT against one fsync per WRITE on
identical WRITE counts, and the COMMIT RPCs alone cost NFSD_IO_DIRECT
31% more server CPU and 42% more client CPU for the same bytes, about
20 us of server CPU per COMMIT plus a client cost that grows with the
range committed. Where only one write in 22 is split the same effect
is a couple of cores on each side and no resolvable throughput
difference, and writes that fit a single RPC send no COMMIT in either
mode.
> So even if these patches apply mechanically and benefit your workloads,
> it would be helpful if we could step back and understand what you are
> trying to achieve and how it should coordinate/align with where Jeff
> and I want to see this mechanism go. Can we start with that
> conversation?
OK, please review what I've provided further and we can then have a
more detailed discussion about anything you like.
Claude helped summarize what this series fixes:
What this advance is, precisely. It is not "NFSD uses DONTCACHE". It
is NFSD's direct write path - io_cache_write modes 2, 3 and 4, where
an aligned WRITE goes to the filesystem as O_DIRECT - with DONTCACHE
used only for the buffered fragments a misaligned WRITE cannot issue
directly. The page cache is bypassed for the bulk of the data and
bounded for the remainder. That distinction runs through everything
below, and it is exactly what separates this from io_cache_write=1,
which is plain buffered DONTCACHE with no direct I/O at all and does
not reach this code.
A misaligned WRITE in one of those direct modes is split into a
buffered prefix, an O_DIRECT middle and a buffered suffix. The
boundary pages are shared by exactly two WRITEs that may arrive in
either order, from different clients, at once. Marking both DONTCACHE
bounded the page cache but cost the second writer a read from disk
every time - 0.81 device reads per record on the rig, 585 GiB read
back during an 8.1 TiB write.
Thanks,
Mike
next prev parent reply other threads:[~2026-09-29 19:56 UTC|newest]
Thread overview: 17+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-29 17:34 [PATCH 00/10] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
2026-09-29 17:34 ` [PATCH 01/10] NFSD: interlock the use of NFSD_IO_DIRECT for NFS READ and WRITE Mike Snitzer
2026-09-29 18:27 ` Chuck Lever
2026-09-29 19:56 ` Mike Snitzer [this message]
2026-09-29 23:17 ` Chuck Lever
2026-09-29 23:30 ` Mike Snitzer
2026-09-30 0:20 ` Chuck Lever
2026-09-30 12:46 ` Mike Snitzer
2026-09-29 17:34 ` [PATCH 02/10] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well Mike Snitzer
2026-09-29 17:34 ` [PATCH 03/10] NFSD: only split a direct-mode WRITE for a worthwhile direct middle Mike Snitzer
2026-09-29 17:34 ` [PATCH 04/10] NFSD: do not use direct I/O for a READ smaller than its alignment Mike Snitzer
2026-09-29 17:34 ` [PATCH 05/10] NFSD: Enable return of an updated stable_how to NFS clients Mike Snitzer
2026-09-29 17:34 ` [PATCH 06/10] NFSD: let a direct-mode WRITE raise stable_how and elide the client's COMMIT Mike Snitzer
2026-09-29 17:34 ` [PATCH 07/10] NFSD: persist a synchronous direct-mode WRITE once, after all of its segments Mike Snitzer
2026-09-29 17:34 ` [PATCH 08/10] NFSD: keep boundary page of a split direct-mode WRITE until both writers complete Mike Snitzer
2026-09-29 17:34 ` [PATCH 09/10] NFSD: add direct_misaligned_dontcache debugfs knob Mike Snitzer
2026-09-29 17:34 ` [PATCH 10/10] NFSD: add tracing for how direct-mode READ and WRITE are serviced Mike Snitzer
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=arwYBaemyvLoG2sT@hammerspace.com \
--to=snitzer@hammerspace.com \
--cc=cel@kernel.org \
--cc=jlayton@kernel.org \
--cc=linux-nfs@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox