From: Mike Snitzer <snitzer@hammerspace.com>
To: Chuck Lever <cel@kernel.org>, Jeff Layton <jlayton@kernel.org>
Cc: linux-nfs@vger.kernel.org
Subject: [PATCH 00/10] NFSD: keep direct-mode I/O out of the page cache and elide COMMITs
Date: Tue, 29 Sep 2026 13:34:13 -0400 [thread overview]
Message-ID: <20260929173423.16149-1-snitzer@kernel.org> (raw)
Hi,
This series builds on NFSD_IO_DIRECT. When a READ or WRITE in a direct
mode cannot be direct I/O it now falls back to DONTCACHE rather than to
cached buffered I/O, a direct-mode WRITE can report the stability it
actually has so the client skips its COMMIT, and new tracepoints show
which path each request took.
Patch 1 interlocks io_cache_read and io_cache_write so READ and WRITE
are never left on opposite sides of the buffered/direct divide.
Patches 2-4 cover what cannot or should not be direct: the direct
middle of a split WRITE is also marked IOCB_DONTCACHE, for when XFS
falls back to buffered I/O (-ENOTBLK); a WRITE is split only when that
buys a worthwhile direct middle (direct_misaligned_num_pages, default
2); and a READ smaller than its alignment no longer costs a full
aligned device read.
Patch 5 is Chuck's "Enable return of an updated stable_how to NFS
clients", reworked onto the @iocb_flags argument nfsd_write() now takes
in nfsd-next. The Reviewed-by tags from its first posting are dropped
because the argument changed.
Patch 6 adds io_cache_write modes 3 and 4, which issue direct I/O like
NFSD_IO_DIRECT and raise the reply's stable_how to at least DATA_SYNC
or FILE_SYNC. On a pNFS flexfiles share where every write is split
across two data servers, the COMMITs alone cost NFSD_IO_DIRECT 31% more
server CPU and 42% more client CPU for the same bytes.
Patch 7 persists a synchronous direct-mode WRITE once, after all of its
segments, instead of up to three fsyncs per WRITE.
Patch 8 keeps the page that two misaligned WRITEs share in the page
cache until both have written it, so the second one no longer has to
read it back from disk: with 32 interleaved writers, 704 device reads
for 42895 WRITEs where there were 45144 for 45664.
Patch 9 (Jonathan) adds the direct_misaligned_dontcache debugfs knob,
default Y; set to N, the parts of a direct-mode WRITE that cannot be
direct use cached buffered I/O instead of DONTCACHE.
Patch 10 adds tracepoints for how each direct-mode READ and WRITE was
serviced, including why a WRITE was not direct.
Documentation/filesystems/nfs/nfsd-io-modes.rst is updated throughout.
The series applies to cel/nfsd-next (ac04dab23b5f) and each patch
builds cleanly with W=1.
All review appreciated, thanks.
Mike
Chuck Lever (1):
NFSD: Enable return of an updated stable_how to NFS clients
Jonathan Flynn (1):
NFSD: add direct_misaligned_dontcache debugfs knob
Mike Snitzer (8):
NFSD: interlock the use of NFSD_IO_DIRECT for NFS READ and WRITE
NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well
NFSD: only split a direct-mode WRITE for a worthwhile direct middle
NFSD: do not use direct I/O for a READ smaller than its alignment
NFSD: let a direct-mode WRITE raise stable_how and elide the client's COMMIT
NFSD: persist a synchronous direct-mode WRITE once, after all of its segments
NFSD: keep boundary page of a split direct-mode WRITE until both writers complete
NFSD: add tracing for how direct-mode READ and WRITE are serviced
.../filesystems/nfs/nfsd-io-modes.rst | 164 ++++++++-
fs/nfsd/debugfs.c | 94 +++++-
fs/nfsd/nfs3proc.c | 16 +-
fs/nfsd/nfs4proc.c | 15 +-
fs/nfsd/nfsd.h | 4 +
fs/nfsd/nfsproc.c | 3 +-
fs/nfsd/trace.h | 90 +++++
fs/nfsd/vfs.c | 318 +++++++++++++++---
fs/nfsd/vfs.h | 26 +-
fs/nfsd/xdr3.h | 2 +-
10 files changed, 665 insertions(+), 67 deletions(-)
--
2.52.0
next reply other threads:[~2026-09-29 17:34 UTC|newest]
Thread overview: 17+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-29 17:34 Mike Snitzer [this message]
2026-09-29 17:34 ` [PATCH 01/10] NFSD: interlock the use of NFSD_IO_DIRECT for NFS READ and WRITE Mike Snitzer
2026-09-29 18:27 ` Chuck Lever
2026-09-29 19:56 ` Mike Snitzer
2026-09-29 23:17 ` Chuck Lever
2026-09-29 23:30 ` Mike Snitzer
2026-09-30 0:20 ` Chuck Lever
2026-09-30 12:46 ` Mike Snitzer
2026-09-29 17:34 ` [PATCH 02/10] NFSD: mark the direct middle of a split WRITE IOCB_DONTCACHE as well Mike Snitzer
2026-09-29 17:34 ` [PATCH 03/10] NFSD: only split a direct-mode WRITE for a worthwhile direct middle Mike Snitzer
2026-09-29 17:34 ` [PATCH 04/10] NFSD: do not use direct I/O for a READ smaller than its alignment Mike Snitzer
2026-09-29 17:34 ` [PATCH 05/10] NFSD: Enable return of an updated stable_how to NFS clients Mike Snitzer
2026-09-29 17:34 ` [PATCH 06/10] NFSD: let a direct-mode WRITE raise stable_how and elide the client's COMMIT Mike Snitzer
2026-09-29 17:34 ` [PATCH 07/10] NFSD: persist a synchronous direct-mode WRITE once, after all of its segments Mike Snitzer
2026-09-29 17:34 ` [PATCH 08/10] NFSD: keep boundary page of a split direct-mode WRITE until both writers complete Mike Snitzer
2026-09-29 17:34 ` [PATCH 09/10] NFSD: add direct_misaligned_dontcache debugfs knob Mike Snitzer
2026-09-29 17:34 ` [PATCH 10/10] NFSD: add tracing for how direct-mode READ and WRITE are serviced Mike Snitzer
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260929173423.16149-1-snitzer@kernel.org \
--to=snitzer@hammerspace.com \
--cc=cel@kernel.org \
--cc=jlayton@kernel.org \
--cc=linux-nfs@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox