From: Antonio Alvarez Feijoo <antonio.feijoo@suse.com>
To: linux-nfs@vger.kernel.org
Cc: Mike Snitzer <snitzer@kernel.org>,
Anna Schumaker <anna.schumaker@hammerspace.com>,
Trond Myklebust <trond.myklebust@hammerspace.com>,
Antonio Alvarez Feijoo <antonio.feijoo@suse.com>
Subject: [PATCH] NFSv4/flexfiles: add a soft dependency on nfsv3
Date: Wed, 2 Sep 2026 15:35:54 +0200 [thread overview]
Message-ID: <20260902133554.28863-1-antonio.feijoo@suse.com> (raw)
An NFS root mounted over NFSv4 against a Linux NFS server hangs during
boot, either wedging permanently in the switch-root exec or, on a loaded
machine, spinning until the softlockup watchdog fires.
The flexfiles layout driver reaches nfs4_pnfs_ds_connect() on the first
I/O through a layout segment. For an NFSv3 data server that calls
symbol_request(nfs3_set_ds_client), which becomes
request_module("symbol:nfs3_set_ds_client") when nfsv3 is not already
resident. request_module() execs /sbin/modprobe via the usermode
helper, and on an NFS root that binary -- and libkmod, and the libraries
behind it -- must be read from the very filesystem whose I/O is blocked
waiting on this DS connection to be established. The helper cannot make
progress, so neither can the mount.
On a system whose init is on NFS the first such I/O is typically the
execve() of init itself, immediately after pivot_root(), at which point
the initramfs -- and the only reachable copy of modprobe -- is gone:
WARNING: fs/nfs/pnfs_nfs.c:792 at nfs4_pnfs_ds_connect+0x47e/0x490 [nfsv4], CPU#0: systemd/1
Modules linked in: nfs_layout_flexfiles rpcsec_gss_krb5 krb5 auth_rpcgss
nfsv4 dns_resolver nfs lockd grace netfs af_packet [...]
Call Trace:
nfs4_pnfs_ds_connect+0x47e/0x490 [nfsv4]
nfs4_ff_layout_prepare_ds+0x... [nfs_layout_flexfiles]
ff_layout_choose_ds_for_read+0x...
ff_layout_pg_init_read+0x...
nfs_pageio_add_request+0x...
nfs_readahead+0x...
filemap_read+0x...
__kernel_read+0x...
bprm_execve+0x...
do_execveat_common.isra.0+0x...
__x64_sys_execve+0x...
Note the absence of nfsv3 from the module list.
This has been latent since the flexfiles driver was added, because a
fatal DS connect error used to be absorbed: ff_layout_read_pagelist()
returned PNFS_NOT_ATTEMPTED and pnfs_do_read() reissued the I/O through
pnfs_read_through_mds(). The read completed over plain NFSv4 against
the metadata server and nobody noticed that the layout had been
abandoned.
Commit 7a375cafc14e ("NFSv4/flexfiles: honor FF_FLAGS_NO_IO_THRU_MDS on
fatal DS connect errors") removes that fallback for layouts carrying
FF_FLAGS_NO_IO_THRU_MDS, correctly, since silently contradicting the
flag was a bug. But nfsd sets that flag on every layout it hands out
(fs/nfsd/flexfilelayout.c: FF_FLAGS_NO_LAYOUTCOMMIT |
FF_FLAGS_NO_IO_THRU_MDS | FF_FLAGS_NO_READ_IO, unconditionally), so
against a Linux server the no-fallback path is not the appliance corner
case the commit describes -- it is every layout. With no fallback left,
a DS connect that cannot complete is now fatal to the boot.
FF_FLAGS_NO_READ_IO does not save the read: the client only diverts
reads off IOMODE_RW segments, and a read-only NFS root is served
IOMODE_READ segments, where the flag does not apply.
Declare the dependency via MODULE_SOFTDEP("pre: nfsv3") instead of
discovering it at I/O time. The layout driver is itself loaded by
request_module("nfs-layouttype4-4") from find_pnfs_driver() during
mount(), which on an initramfs boot happens while the initramfs and a
working /sbin/modprobe are still in place, so a softdep is resolved at a
point where the usermode helper can actually run.
This pulls in nfsv3 whenever the flexfiles driver is loaded, including
for layouts whose data servers all speak NFSv4.1 and which would never
have needed it. That seemed the better trade against an unbootable
system.
Reported on openSUSE Tumbleweed with kernel 7.2.2 under QEMU,
reproducible with dracut's TEST-60-NFS suite against any knfsd built with
CONFIG_NFSD_FLEXFILELAYOUT=y. Booting with rd.driver.pre=nfsv3, which
loads the module from the initramfs before the pivot, is a complete
workaround and confirms the diagnosis.
Fixes: d67ae825a59d ("pnfs/flexfiles: Add the FlexFile Layout Driver")
Cc: stable@vger.kernel.org
Assisted-by: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Antonio Alvarez Feijoo <antonio.feijoo@suse.com>
---
fs/nfs/flexfilelayout/flexfilelayout.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/fs/nfs/flexfilelayout/flexfilelayout.c b/fs/nfs/flexfilelayout/flexfilelayout.c
index 7fe8b91fa47c..12a4a7908fe3 100644
--- a/fs/nfs/flexfilelayout/flexfilelayout.c
+++ b/fs/nfs/flexfilelayout/flexfilelayout.c
@@ -3079,7 +3079,7 @@ static void __exit nfs4flexfilelayout_exit(void)
}
MODULE_ALIAS("nfs-layouttype4-4");
-
+MODULE_SOFTDEP("pre: nfsv3");
MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("The NFSv4 flexfile layout driver");
--
2.51.0
reply other threads:[~2026-09-02 13:36 UTC|newest]
Thread overview: [no followups] expand[flat|nested] mbox.gz Atom feed
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260902133554.28863-1-antonio.feijoo@suse.com \
--to=antonio.feijoo@suse.com \
--cc=anna.schumaker@hammerspace.com \
--cc=linux-nfs@vger.kernel.org \
--cc=snitzer@kernel.org \
--cc=trond.myklebust@hammerspace.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox