Linux NFS development
 help / color / mirror / Atom feed
From: Mike Snitzer <snitzer@kernel.org>
To: Chuck Lever <cel@kernel.org>
Cc: hch@lst.de, Jeff Layton <jlayton@kernel.org>, linux-nfs@vger.kernel.org
Subject: Re: [PATCH 01/10] NFSD: interlock the use of NFSD_IO_DIRECT for NFS READ and WRITE
Date: Wed, 30 Sep 2026 08:46:48 -0400	[thread overview]
Message-ID: <ar0EuAKO0VvD8kT4@kernel.org> (raw)
In-Reply-To: <0cbfd399-45cd-4df2-888a-bdf8a2935ab0@app.fastmail.com>

On Tue, Sep 29, 2026 at 05:20:03PM -0700, Chuck Lever wrote:
> 
> 
> On Tue, Sep 29, 2026, at 4:30 PM, Mike Snitzer wrote:
> > On Tue, Sep 29, 2026 at 04:17:50PM -0700, Chuck Lever wrote:
> >> 
> >> 
> >> On Tue, Sep 29, 2026, at 12:56 PM, Mike Snitzer wrote:
> >> > On Tue, Sep 29, 2026 at 11:27:05AM -0700, Chuck Lever wrote:
> >> >> 
> >> >> On Tue, Sep 29, 2026, at 10:34 AM, Mike Snitzer wrote:
> >> >> > Now that NFSD supports NFSD_IO_DIRECT for both READ and WRITE it is
> >> >> > much safer to avoid needless buffered vs direct contention if/when
> >> >> > only one of them has been configured to use NFSD_IO_DIRECT.
> >> >> >
> >> >> > Mixing direct and buffered I/O to the same file causes needless page
> >> >> > cache invalidation and writeback, so although io_cache_read and
> >> >> > io_cache_write remain separate interfaces, writing either one adjusts
> >> >> > the other so that READ and WRITE are never left on opposite sides of
> >> >> > the buffered/direct divide:
> >> >> 
> >> >> Jeff and I have been discussing making DIRECT the default for WRITEs and
> >> >> BUFFERED the default for READs. This patch takes us in the opposite
> >> >> direction.
> >> >> 
> >> >> Nothing here convinces me that mixing the modes is a bad thing to do.
> >> >> "Needless page cache invalidation" needs some demonstration, and it
> >> >> needs to show why the right thing to do is make it impossible to mix
> >> >> modes rather than explore the issue as one or more bugs that can be
> >> >> fixed. Or... why not let admins explore this for themselves? Where is
> >> >> the hazard and why does it need to be forbidden by the admin UI?
> >> >
> >> > The basis for the interlock is that O_DIRECT mode is intended to avoid
> >> > bloating memory with page cache and the excess CPU burn of managing
> >> > the page cache.  Using read caching in conjunction with O_DIRECT
> >> > writes knee-caps the wins of O_DIRECT.
> >> 
> >> Jeff and I have never seen that, and it's a counterintuitive result.
> >> Buffered READ with DIRECT WRITE seems to work very well and the server
> >> is easily capable of managing the page cache in this case, since
> >> reclaiming a clean page doesn't mean having to flush dirty data.
> >> Evicting clean pages is not slow.
> >> 
> >> So I'd like to see a quantification of the penalties in this mixed
> >> mode before adjudicating it as a hazard. That is, you might be right,
> >> but so far I've seen no direct evidence that having a substantial page
> >> cache presence is a general deficit for WRITEs. If there are certain
> >> cases where even caching READs is a problem, then by all means, set the
> >> READ IO mode to DIRECT too in those cases.
> >> 
> >> I don't see this as an argument for forcing the server's mode setting
> >> in every case.
> >
> > Yes, I understand and I've dropped the interlock patch in v2 (just posted).
> 
> Please give reviewers a chance to digest and review, as requested in
> Documentation/filesystems/nfs/nfsd-maintainer-entry-profile.rst :
> 
> "As always, please avoid reposting series revisions more than once
>  every 24 hours."
> 
> Trust me, it saves a lot of confusion.
> 
> 
> >> Based on this description, it seems to me the client has full visibility
> >> of the data to be pushed back and how it's sharded, it has information
> >> about the network RTT (that's where the real throughput impact is for
> >> COMMIT), and it has control over the selection of UNSTABLE vs. FILE_SYNC.
> >> 
> >> The problem here might be that when a large WRITE payload is sharded
> >> across multiple servers, the client still thinks it is sending them via
> >> UNSTABLE WRITES to one server, and plans for only one COMMIT after the
> >> server completes the WRITEs.
> >> 
> >> But for the pNFS scenario, the client sends one UNSTABLE WRITE followed
> >> by a COMMIT to each server. If the sharded WRITES are all single RPCs
> >> to distinct servers, then they should each be FILE_SYNC.
> >> 
> >> The client can be smarter about how it writes data back to multiple
> >> servers, can't it?
> >
> > When the server is operating in direct mode it isn't something exposed
> > to the client.  That the client could be smarter and/or already has
> > adequate controls to achieve the same FILE_SYNC result is besides the
> > point.  The point is, if the server has already done the work then it
> > should, within reason, convey as much back to the client to elide
> > COMMIT work that isn't needed.
> 
> Your point assumes the other patches in the series are applied to make
> DIRECT UNSTABLE WRITEs completely persistent.
> 
> The current IOCB flags do not include IOCB_DSYNC on a DIRECT UNSTABLE
> WRITE for a very good reason: that makes them slower and more expensive.
> The current server logic is working exactly as we designed it last year,
> and I'm not enthusiastic about changing that. I expect at least one
> other reviewer will have a similar reaction, once he sobers up from
> ALPSS.

You haven't reviewed the changes, yet you are passing judgement.

Please review.  I think you'll find I have only made things better.

And as for Mr ALPSS, the "NFSD: do not use direct I/O for a READ
smaller than its alignment" and "NFSD: only split a direct-mode WRITE
for a worthwhile direct middle" patches were motivated by his feedback
last year! (I have been carrying them since then, albeit in a less
polished form than I have now made available in this series).

> If the client wants to avoid the COMMIT, the standing rule is to send a
> FILE_SYNC WRITE. That benefits all WRITE I/O modes on the server.

I haven't precluded DIRECT UNSTABLE WRITE from operating as it does
now _at all_.

I've merely exposed additional optional controls that require opt-in.
And provided a natural evolution to what we have that has proven
beneficial.

  reply	other threads:[~2026-09-30 12:46 UTC|newest]

Thread overview: 17+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-29 17:34 [PATCH 00/10] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs Mike Snitzer
2026-09-29 17:34 ` [PATCH 01/10] NFSD: interlock the use of NFSD_IO_DIRECT for NFS READ and WRITE Mike Snitzer
2026-09-29 18:27   ` Chuck Lever
2026-09-29 19:56     ` Mike Snitzer
2026-09-29 23:17       ` Chuck Lever
2026-09-29 23:30         ` Mike Snitzer
2026-09-30  0:20           ` Chuck Lever
2026-09-30 12:46             ` Mike Snitzer [this message]
2026-09-29 17:34 ` [PATCH 02/10] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well Mike Snitzer
2026-09-29 17:34 ` [PATCH 03/10] NFSD: only split a direct-mode WRITE for a worthwhile direct middle Mike Snitzer
2026-09-29 17:34 ` [PATCH 04/10] NFSD: do not use direct I/O for a READ smaller than its alignment Mike Snitzer
2026-09-29 17:34 ` [PATCH 05/10] NFSD: Enable return of an updated stable_how to NFS clients Mike Snitzer
2026-09-29 17:34 ` [PATCH 06/10] NFSD: let a direct-mode WRITE raise stable_how and elide the client's COMMIT Mike Snitzer
2026-09-29 17:34 ` [PATCH 07/10] NFSD: persist a synchronous direct-mode WRITE once, after all of its segments Mike Snitzer
2026-09-29 17:34 ` [PATCH 08/10] NFSD: keep boundary page of a split direct-mode WRITE until both writers complete Mike Snitzer
2026-09-29 17:34 ` [PATCH 09/10] NFSD: add direct_misaligned_dontcache debugfs knob Mike Snitzer
2026-09-29 17:34 ` [PATCH 10/10] NFSD: add tracing for how direct-mode READ and WRITE are serviced Mike Snitzer

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=ar0EuAKO0VvD8kT4@kernel.org \
    --to=snitzer@kernel.org \
    --cc=cel@kernel.org \
    --cc=hch@lst.de \
    --cc=jlayton@kernel.org \
    --cc=linux-nfs@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox