netdev.vger.kernel.org archive mirror
 help / color / mirror / Atom feed
From: Jeff Layton <jlayton@kernel.org>
To: Chuck Lever <cel@kernel.org>, NeilBrown <neil@brown.name>,
	 Olga Kornievskaia <okorniev@redhat.com>,
	Dai Ngo <Dai.Ngo@oracle.com>,  Tom Talpey <tom@talpey.com>,
	Trond Myklebust <trondmy@kernel.org>,
	 Anna Schumaker <anna@kernel.org>,
	"David S. Miller" <davem@davemloft.net>,
	 Eric Dumazet <edumazet@google.com>,
	Jakub Kicinski <kuba@kernel.org>,
	 Paolo Abeni <pabeni@redhat.com>, Simon Horman <horms@kernel.org>,
	 "J. Bruce Fields" <bfields@fieldses.org>,
	Shuah Khan <shuah@kernel.org>
Cc: linux-nfs@vger.kernel.org, linux-kernel@vger.kernel.org,
	 netdev@vger.kernel.org, Trond Myklebust <trondmy@gmail.com>,
	 linux-kselftest@vger.kernel.org,
	Jeff Layton <jlayton@kernel.org>
Subject: [PATCH 4/7] SUNRPC: bound the local rpcbind client timeout to 1s
Date: Mon, 10 Aug 2026 13:38:51 -0400	[thread overview]
Message-ID: <20260810-nfsd-nl-hang-v1-4-2519fdd5bc1a@kernel.org> (raw)
In-Reply-To: <20260810-nfsd-nl-hang-v1-0-2519fdd5bc1a@kernel.org>

The kernel's local rpcbind client runs on the transport defaults: a 10s
major timeout for AF_LOCAL, 60s for the loopback TCP fallback
(xprt_calc_majortimeo() returns to_initval when to_increment is 0).

Those calls are synchronous and run under nfsd_mutex, several per
operation: rpcb_create_local() attempts up to three client creations, and
svc_register() issues one call per program and version. A local rpcbind
that accepts the connection but never replies stalls each of them, and
the accumulated hold is enough to trip the hung-task watchdog on other
NFSD netlink ops (the holder waits killably and evades it):

  INFO: task hung in nfsd_nl_cache_flush_doit

The local rpcbind lives on loopback or an AF_LOCAL socket and answers in
microseconds, so bound its client to one attempt, 1s.

This shortens the stall rather than removing it, and it is not free.
Registration stays synchronous and stays fatal: rpcb_create_local()
failure aborts nfsd_create_serv() via svc_bind(), and svc_register()
failure makes svc_setup_socket() fail, so a rpcbind that is merely slow
to be scheduled can now fail server startup where it previously
succeeded. Making the registration asynchronous is the real fix.

Link: https://syzkaller.appspot.com/bug?extid=c7eae0eb80858a2dba0f
Signed-off-by: Jeff Layton <jlayton@kernel.org>
Assisted-by: LLM
---
 net/sunrpc/rpcb_clnt.c | 12 ++++++++++++
 1 file changed, 12 insertions(+)

diff --git a/net/sunrpc/rpcb_clnt.c b/net/sunrpc/rpcb_clnt.c
index 6aa372188c86..0aa376b82a52 100644
--- a/net/sunrpc/rpcb_clnt.c
+++ b/net/sunrpc/rpcb_clnt.c
@@ -221,6 +221,16 @@ static void rpcb_set_local(struct net *net, struct rpc_clnt *clnt,
 # define SUN_LEN(ptr) (offsetof(struct sockaddr_un, sun_path)		\
 		      + 1 + strlen((ptr)->sun_path + 1))
 
+/*
+ * The kernel's rpcbind client talks only to the local rpcbind, over loopback
+ * or a local AF_LOCAL socket, where a healthy rpcbind answers in microseconds.
+ */
+static const struct rpc_timeout rpcb_local_timeout = {
+	.to_initval	= 1 * HZ,
+	.to_maxval	= 1 * HZ,
+	.to_retries	= 0,
+};
+
 /*
  * Returns zero on success, otherwise a negative errno value
  * is returned.
@@ -238,6 +248,7 @@ static int rpcb_create_af_local(struct net *net,
 		.version	= RPCBVERS_2,
 		.authflavor	= RPC_AUTH_NULL,
 		.cred		= current_cred(),
+		.timeout	= &rpcb_local_timeout,
 		/*
 		 * We turn off the idle timeout to prevent the kernel
 		 * from automatically disconnecting the socket.
@@ -312,6 +323,7 @@ static int rpcb_create_local_net(struct net *net)
 		.version	= RPCBVERS_2,
 		.authflavor	= RPC_AUTH_UNIX,
 		.cred		= current_cred(),
+		.timeout	= &rpcb_local_timeout,
 		.flags		= RPC_CLNT_CREATE_NOPING,
 	};
 	struct rpc_clnt *clnt, *clnt4;

-- 
2.55.0


  parent reply	other threads:[~2026-08-10 17:39 UTC|newest]

Thread overview: 8+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-10 17:38 [PATCH 0/7] nfsd/sunrpc: harden the netlink listener interfaces Jeff Layton
2026-08-10 17:38 ` [PATCH 1/7] NFSD: validate transport name in listener_set before serv creation Jeff Layton
2026-08-10 17:38 ` [PATCH 2/7] NFSD: cap the number of listeners accepted in listener_set Jeff Layton
2026-08-10 17:38 ` [PATCH 3/7] SUNRPC: keep the first error in svc_register() Jeff Layton
2026-08-10 17:38 ` Jeff Layton [this message]
2026-08-10 17:38 ` [PATCH 5/7] NFSD: report listener creation failures through extack Jeff Layton
2026-08-10 17:38 ` [PATCH 6/7] selftests/nfsd: exercise listener_set request validation Jeff Layton
2026-08-10 17:38 ` [PATCH 7/7] selftests/nfsd: add a per-netns rpcbind stub and the listener round-trips Jeff Layton

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260810-nfsd-nl-hang-v1-4-2519fdd5bc1a@kernel.org \
    --to=jlayton@kernel.org \
    --cc=Dai.Ngo@oracle.com \
    --cc=anna@kernel.org \
    --cc=bfields@fieldses.org \
    --cc=cel@kernel.org \
    --cc=davem@davemloft.net \
    --cc=edumazet@google.com \
    --cc=horms@kernel.org \
    --cc=kuba@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-nfs@vger.kernel.org \
    --cc=neil@brown.name \
    --cc=netdev@vger.kernel.org \
    --cc=okorniev@redhat.com \
    --cc=pabeni@redhat.com \
    --cc=shuah@kernel.org \
    --cc=tom@talpey.com \
    --cc=trondmy@gmail.com \
    --cc=trondmy@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).