From: Mike Snitzer <snitzer@kernel.org>
To: Trond Myklebust <trondmy@kernel.org>, Anna Schumaker <anna@kernel.org>
Cc: linux-nfs@vger.kernel.org
Subject: [PATCH 6/6] pNFS/flexfiles: don't reset to MDS for v4 error when previous layout forbid it
Date: Mon, 28 Sep 2026 11:54:30 -0400 [thread overview]
Message-ID: <20260928155430.95985-7-snitzer@kernel.org> (raw)
In-Reply-To: <20260928155430.95985-1-snitzer@kernel.org>
ff_layout_async_handle_error_v4() ends in -NFS4ERR_RESET_TO_MDS, which
makes ff_layout_read_release() / ff_layout_write_release() call
ff_layout_reset_read() or ff_layout_reset_write(hdr, false), and those
call pnfs_read_done_resend_to_mds() / pnfs_write_done_resend_to_mds().
Those build a page I/O descriptor with force_mds set, so the flexfiles
->pg_init() never runs and no layout is consulted anywhere on the way:
the READ or WRITE goes to the metadata server unconditionally. For a
layout carrying FF_FLAGS_NO_IO_THRU_MDS that is exactly what the server
forbade.
The fall-through into reset: is already safe, because
ff_layout_avoid_mds_available_ds() returns true for a segment with the
flag and so answers -NFS4ERR_RESET_TO_PNFS first. What is not safe is
the goto: the invalid-layout errors (NFS4ERR_PNFS_NO_LAYOUT, STALE,
BADHANDLE, ISDIR, FHEXPIRED, WRONG_TYPE) call pnfs_destroy_layout() and
jump straight to reset:, over that check. That is the case where the
file had a perfectly good layout, something forced it to be given back,
and the client answers by writing to the MDS behind the server's back.
Test at reset: itself, so it covers the goto and any future one, and ask
the inode (NFS_INO_NO_IO_THRU_MDS): the callers that jump here have just
destroyed the layout hdr, so the hdr-level flag is no longer reachable,
while the inode-level one is set for the life of the in-core inode.
-NFS4ERR_RESET_TO_PNFS is the right answer rather than an error. It is
what the fall-through already returns for the same flag, and the
resends it drives - ff_layout_resend_pnfs_read() and
ff_layout_reset_write(hdr, true) - re-enter the pgio path with force_mds
clear, so a fresh LAYOUTGET is taken, which is precisely what a
destroyed layout needs. A LAYOUTGET that fails fatally makes
pnfs_update_layout() return an error, which ->pg_init() reports; a
non-fatal refusal (NFS4ERR_LAYOUTUNAVAILABLE, NFS4ERR_TOOSMALL) returns
no segment instead, and the client keeps retrying through pNFS rather
than falling back.
ff_layout_async_handle_error_v3() needs no such change: it returns only
-NFS4ERR_RESET_TO_PNFS or -EAGAIN.
Fixes: 260074cd8413 ("pNFS/flexfiles: Add support for FF_FLAGS_NO_IO_THRU_MDS")
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
fs/nfs/flexfilelayout/flexfilelayout.c | 17 +++++++++++++++++
1 file changed, 17 insertions(+)
diff --git a/fs/nfs/flexfilelayout/flexfilelayout.c b/fs/nfs/flexfilelayout/flexfilelayout.c
index 84d22f715653d..d4f2b57e9d5b2 100644
--- a/fs/nfs/flexfilelayout/flexfilelayout.c
+++ b/fs/nfs/flexfilelayout/flexfilelayout.c
@@ -1423,6 +1423,23 @@ static int ff_layout_async_handle_error_v4(struct rpc_task *task,
if (ff_layout_avoid_mds_available_ds(lseg))
return -NFS4ERR_RESET_TO_PNFS;
reset:
+ /*
+ * FF_FLAGS_NO_IO_THRU_MDS: never resend through the MDS. The
+ * caller would do so with force_mds set (ff_layout_reset_read(),
+ * ff_layout_reset_write(hdr, false) -> pnfs_*_done_resend_to_mds()),
+ * which builds a descriptor out of the plain MDS page ops and so
+ * consults no layout at all -- this is the last point at which the
+ * policy can still be applied. Retry through pNFS instead, which
+ * takes a fresh LAYOUTGET; that is also the right answer for the
+ * invalid-layout cases that jump here, since they have just called
+ * pnfs_destroy_layout(). Having done so they can no longer ask the
+ * layout hdr about the flag, hence the inode.
+ */
+ if (nfs_no_io_thru_mds(inode)) {
+ dprintk("%s Retry through pNFS, no MDS fallback. Error %d\n",
+ __func__, task->tk_status);
+ return -NFS4ERR_RESET_TO_PNFS;
+ }
dprintk("%s Retry through MDS. Error %d\n", __func__,
task->tk_status);
return -NFS4ERR_RESET_TO_MDS;
--
2.44.0
prev parent reply other threads:[~2026-09-28 15:54 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-28 15:54 [PATCH 0/6] NFS: LOCALIO fixes and FF_FLAGS_NO_IO_THRU_MDS fixes Mike Snitzer
2026-09-28 15:54 ` [PATCH 1/6] NFS/localio: fix nfs_local_dio_misaligned tracepoint Mike Snitzer
2026-09-28 15:54 ` [PATCH 2/6] NFS/localio: detect a short read or write before the iterator has moved Mike Snitzer
2026-09-28 15:54 ` [PATCH 3/6] NFS/localio: report the stability a DIO WRITE actually has Mike Snitzer
2026-09-28 15:54 ` [PATCH 4/6] pNFS/flexfiles: honor FF_FLAGS_NO_IO_THRU_MDS over the mdsthreshold hint Mike Snitzer
2026-09-28 15:54 ` [PATCH 5/6] pNFS/flexfiles: don't read through MDS when previous layout forbid it Mike Snitzer
2026-09-28 15:54 ` Mike Snitzer [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260928155430.95985-7-snitzer@kernel.org \
--to=snitzer@kernel.org \
--cc=anna@kernel.org \
--cc=linux-nfs@vger.kernel.org \
--cc=trondmy@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox