From: "Chuck Lever" <cel@kernel.org>
To: "Mike Snitzer" <snitzer@hammerspace.com>
Cc: "Jeff Layton" <jlayton@kernel.org>, linux-nfs@vger.kernel.org
Subject: Re: [PATCH 01/10] NFSD: interlock the use of NFSD_IO_DIRECT for NFS READ and WRITE
Date: Tue, 29 Sep 2026 16:17:50 -0700 [thread overview]
Message-ID: <d5d7635a-fe7a-4664-8797-b581c23cdf9c@app.fastmail.com> (raw)
In-Reply-To: <arwYBaemyvLoG2sT@hammerspace.com>
On Tue, Sep 29, 2026, at 12:56 PM, Mike Snitzer wrote:
> On Tue, Sep 29, 2026 at 11:27:05AM -0700, Chuck Lever wrote:
>>
>> On Tue, Sep 29, 2026, at 10:34 AM, Mike Snitzer wrote:
>> > Now that NFSD supports NFSD_IO_DIRECT for both READ and WRITE it is
>> > much safer to avoid needless buffered vs direct contention if/when
>> > only one of them has been configured to use NFSD_IO_DIRECT.
>> >
>> > Mixing direct and buffered I/O to the same file causes needless page
>> > cache invalidation and writeback, so although io_cache_read and
>> > io_cache_write remain separate interfaces, writing either one adjusts
>> > the other so that READ and WRITE are never left on opposite sides of
>> > the buffered/direct divide:
>>
>> Jeff and I have been discussing making DIRECT the default for WRITEs and
>> BUFFERED the default for READs. This patch takes us in the opposite
>> direction.
>>
>> Nothing here convinces me that mixing the modes is a bad thing to do.
>> "Needless page cache invalidation" needs some demonstration, and it
>> needs to show why the right thing to do is make it impossible to mix
>> modes rather than explore the issue as one or more bugs that can be
>> fixed. Or... why not let admins explore this for themselves? Where is
>> the hazard and why does it need to be forbidden by the admin UI?
>
> The basis for the interlock is that O_DIRECT mode is intended to avoid
> bloating memory with page cache and the excess CPU burn of managing
> the page cache. Using read caching in conjunction with O_DIRECT
> writes knee-caps the wins of O_DIRECT.
Jeff and I have never seen that, and it's a counterintuitive result.
Buffered READ with DIRECT WRITE seems to work very well and the server
is easily capable of managing the page cache in this case, since
reclaiming a clean page doesn't mean having to flush dirty data.
Evicting clean pages is not slow.
So I'd like to see a quantification of the penalties in this mixed
mode before adjudicating it as a hazard. That is, you might be right,
but so far I've seen no direct evidence that having a substantial page
cache presence is a general deficit for WRITEs. If there are certain
cases where even caching READs is a problem, then by all means, set the
READ IO mode to DIRECT too in those cases.
I don't see this as an argument for forcing the server's mode setting
in every case.
> But sure, you can do it.. and
> I suppose the freedom of choice (and enough rope to hang) is perfectly
> fine.
"Enough rope to hang" is the usual approach for Linux tunables. I
should also point out that both of these settings are still debugfs,
not the final shape of the administrative UI. It doesn't seem sensible
to me to restrict them at this point, but I'll keep this idea in mind
as we continue to develop our thinking about what the eventual non-
debug admin UI will look like.
> I can drop the interlock and send v2 after more time for review
> of the other patches (but there were some real bugs fixed in the
> interlock patch so will require some care to adjust).
The usual policy for "real bugs" is that those need to be fixed in
separate, backport-able patches before making behavioral changes.
Stable wants the fixes, and does not want the behavior changes, so
these need to be separable.
You can send the bug fixes any time without waiting for a v2.
>> As I recall, even a direct WRITE needs a subsequent COMMIT. Have you
>> demonstrated that the client COMMIT is a latency problem or that
>> the memory that is pinned on the client has a noticeable impact?
>>
>> We suspect that it might, but every time I've measured, avoiding
>> COMMIT with NFSD has shown no impact on throughput, which is why
>> NFSD doesn't already do this optimization.
>
> Yes, in the header for "NFSD: let a direct-mode WRITE raise stable_how
> and elide the client's COMMIT":
>
> What this removes is the COMMIT traffic, and it is worth most where a
> client's writes are carved into several WRITE RPCs, because each piece
> is then committed separately. Measured on a pNFS flexfiles share where
> every write straddles two data servers, so every write becomes two WRITE
> RPCs and, under NFSD_IO_DIRECT, two COMMITs: the two modes do identical
> durability work, one fsync per COMMIT against one fsync per WRITE on
> identical WRITE counts, and the COMMIT RPCs alone cost NFSD_IO_DIRECT
> 31% more server CPU and 42% more client CPU for the same bytes, about
> 20 us of server CPU per COMMIT plus a client cost that grows with the
> range committed. Where only one write in 22 is split the same effect
> is a couple of cores on each side and no resolvable throughput
> difference, and writes that fit a single RPC send no COMMIT in either
> mode.
I asked about latency, above. This paragraph is about CPU utilization.
Based on this description, it seems to me the client has full visibility
of the data to be pushed back and how it's sharded, it has information
about the network RTT (that's where the real throughput impact is for
COMMIT), and it has control over the selection of UNSTABLE vs. FILE_SYNC.
The problem here might be that when a large WRITE payload is sharded
across multiple servers, the client still thinks it is sending them via
UNSTABLE WRITES to one server, and plans for only one COMMIT after the
server completes the WRITEs.
But for the pNFS scenario, the client sends one UNSTABLE WRITE followed
by a COMMIT to each server. If the sharded WRITES are all single RPCs
to distinct servers, then they should each be FILE_SYNC.
The client can be smarter about how it writes data back to multiple
servers, can't it?
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
next prev parent reply other threads:[~2026-09-29 23:18 UTC|newest]
Thread overview: 17+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-29 17:34 [PATCH 00/10] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
2026-09-29 17:34 ` [PATCH 01/10] NFSD: interlock the use of NFSD_IO_DIRECT for NFS READ and WRITE Mike Snitzer
2026-09-29 18:27 ` Chuck Lever
2026-09-29 19:56 ` Mike Snitzer
2026-09-29 23:17 ` Chuck Lever [this message]
2026-09-29 23:30 ` Mike Snitzer
2026-09-30 0:20 ` Chuck Lever
2026-09-30 12:46 ` Mike Snitzer
2026-09-29 17:34 ` [PATCH 02/10] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well Mike Snitzer
2026-09-29 17:34 ` [PATCH 03/10] NFSD: only split a direct-mode WRITE for a worthwhile direct middle Mike Snitzer
2026-09-29 17:34 ` [PATCH 04/10] NFSD: do not use direct I/O for a READ smaller than its alignment Mike Snitzer
2026-09-29 17:34 ` [PATCH 05/10] NFSD: Enable return of an updated stable_how to NFS clients Mike Snitzer
2026-09-29 17:34 ` [PATCH 06/10] NFSD: let a direct-mode WRITE raise stable_how and elide the client's COMMIT Mike Snitzer
2026-09-29 17:34 ` [PATCH 07/10] NFSD: persist a synchronous direct-mode WRITE once, after all of its segments Mike Snitzer
2026-09-29 17:34 ` [PATCH 08/10] NFSD: keep boundary page of a split direct-mode WRITE until both writers complete Mike Snitzer
2026-09-29 17:34 ` [PATCH 09/10] NFSD: add direct_misaligned_dontcache debugfs knob Mike Snitzer
2026-09-29 17:34 ` [PATCH 10/10] NFSD: add tracing for how direct-mode READ and WRITE are serviced Mike Snitzer
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=d5d7635a-fe7a-4664-8797-b581c23cdf9c@app.fastmail.com \
--to=cel@kernel.org \
--cc=jlayton@kernel.org \
--cc=linux-nfs@vger.kernel.org \
--cc=snitzer@hammerspace.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox