From: Mike Snitzer <snitzer@kernel.org>
To: Trond Myklebust <trondmy@kernel.org>, Anna Schumaker <anna@kernel.org>
Cc: linux-nfs@vger.kernel.org
Subject: [PATCH 5/6] pNFS/flexfiles: don't read through MDS when previous layout forbid it
Date: Mon, 28 Sep 2026 11:54:29 -0400 [thread overview]
Message-ID: <20260928155430.95985-6-snitzer@kernel.org> (raw)
In-Reply-To: <20260928155430.95985-1-snitzer@kernel.org>
ff_layout_pg_init_read() falls back to the MDS whenever
pnfs_update_layout() returns no segment and no error: a bulk
CB_LAYOUTRECALL (FSID, DEVICEID or ALL), LAYOUTGETs blocked behind a
layout being returned, or a LAYOUTGET refused with
NFS4ERR_LAYOUTUNAVAILABLE or NFS4ERR_TOOSMALL, which set the per-iomode
fail bit for PNFS_LAYOUTGET_RETRY_TIMEOUT. That is the read-side twin
of the write path fixed by commit 1d62e659c0bf ("NFSv4/flexfiles:
honor FF_FLAGS_NO_IO_THRU_MDS in pg_get_mirror_count_write"), and it
is reached in exactly the situation that matters most: the file had a
valid layout, the layout had to go back, and the next one has not
arrived yet.
The write side could answer -EAGAIN because writeback owns the retry and
redirties the page. A read has no such owner, so -EAGAIN here would
surface as an I/O error to the application. Wait for a layout instead,
on the same terms the function already uses when the layout lookup
itself returns -EAGAIN: sleep a second and retry, bounded by
pg_maxretrans, which ff_layout_pg_init_read() sets only for soft and
softerr mounts. A hard mount therefore waits, which is what a hard
mount is for, and matches what the FF_FLAGS_NO_IO_THRU_MDS check on the
data server selection a few lines above already does. That includes
waiting out NFS4ERR_LAYOUTUNAVAILABLE: a file that was given a layout
with this flag must not be read off the MDS, whatever the server says
afterwards.
The test asks the inode (NFS_INO_NO_IO_THRU_MDS) rather than the layout
hdr, because in this path there is no segment, and the hdr that carried
the flag may already have been destroyed. Where a segment is in hand -
the data server selection above, ff_layout_read_pagelist() - the
existing per-segment ff_layout_no_fallback_to_mds() test stays, since it
is the more precise question.
Fixes: 260074cd8413 ("pNFS/flexfiles: Add support for FF_FLAGS_NO_IO_THRU_MDS")
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
fs/nfs/flexfilelayout/flexfilelayout.c | 38 ++++++++++++++++++--------
1 file changed, 26 insertions(+), 12 deletions(-)
diff --git a/fs/nfs/flexfilelayout/flexfilelayout.c b/fs/nfs/flexfilelayout/flexfilelayout.c
index 7d45401ee5b11..84d22f715653d 100644
--- a/fs/nfs/flexfilelayout/flexfilelayout.c
+++ b/fs/nfs/flexfilelayout/flexfilelayout.c
@@ -1073,19 +1073,33 @@ ff_layout_pg_init_read(struct nfs_pageio_descriptor *pgio,
if (pgio->pg_error < 0) {
if (pgio->pg_error != -EAGAIN)
return;
- /* Retry getting layout segment if lower layer returned -EAGAIN */
- if (pgio->pg_maxretrans && req->wb_nio++ > pgio->pg_maxretrans) {
- if (NFS_SERVER(pgio->pg_inode)->flags & NFS_MOUNT_SOFTERR)
- pgio->pg_error = -ETIMEDOUT;
- else
- pgio->pg_error = -EIO;
- return;
- }
- pgio->pg_error = 0;
- /* Sleep for 1 second before retrying */
- ssleep(1);
- goto retry;
+ goto retry_nolseg;
}
+ /*
+ * No segment, and no error to report either: pnfs_update_layout()
+ * simply has nothing to give (NFS_LAYOUT_BULK_RECALL, a failed
+ * pnfs_layout_io_test, blocked LAYOUTGETs, a layout being
+ * returned). If the server forbids reading this file through the
+ * MDS there is no fallback to take, so wait for a layout on the
+ * same terms as the -EAGAIN above. The layout hdr that carried
+ * FF_FLAGS_NO_IO_THRU_MDS may itself be gone by now, which is why
+ * this asks the inode and not the hdr.
+ */
+ if (!nfs_no_io_thru_mds(pgio->pg_inode))
+ goto out_mds;
+retry_nolseg:
+ /* Retry getting layout segment if lower layer returned -EAGAIN */
+ if (pgio->pg_maxretrans && req->wb_nio++ > pgio->pg_maxretrans) {
+ if (NFS_SERVER(pgio->pg_inode)->flags & NFS_MOUNT_SOFTERR)
+ pgio->pg_error = -ETIMEDOUT;
+ else
+ pgio->pg_error = -EIO;
+ return;
+ }
+ pgio->pg_error = 0;
+ /* Sleep for 1 second before retrying */
+ ssleep(1);
+ goto retry;
out_mds:
trace_pnfs_mds_fallback_pg_init_read(pgio->pg_inode,
0, NFS4_MAX_UINT64, IOMODE_READ,
--
2.44.0
next prev parent reply other threads:[~2026-09-28 15:54 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-28 15:54 [PATCH 0/6] NFS: LOCALIO fixes and FF_FLAGS_NO_IO_THRU_MDS fixes Mike Snitzer
2026-09-28 15:54 ` [PATCH 1/6] NFS/localio: fix nfs_local_dio_misaligned tracepoint Mike Snitzer
2026-09-28 15:54 ` [PATCH 2/6] NFS/localio: detect a short read or write before the iterator has moved Mike Snitzer
2026-09-28 15:54 ` [PATCH 3/6] NFS/localio: report the stability a DIO WRITE actually has Mike Snitzer
2026-09-28 15:54 ` [PATCH 4/6] pNFS/flexfiles: honor FF_FLAGS_NO_IO_THRU_MDS over the mdsthreshold hint Mike Snitzer
2026-09-28 15:54 ` Mike Snitzer [this message]
2026-09-28 15:54 ` [PATCH 6/6] pNFS/flexfiles: don't reset to MDS for v4 error when previous layout forbid it Mike Snitzer
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260928155430.95985-6-snitzer@kernel.org \
--to=snitzer@kernel.org \
--cc=anna@kernel.org \
--cc=linux-nfs@vger.kernel.org \
--cc=trondmy@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox