Linux NFS development
 help / color / mirror / Atom feed
From: Jeff Layton <jlayton@kernel.org>
To: "Steve Dickson" <steved@redhat.com>,
	"Mantas Mikulėnas" <grawity@gmail.com>
Cc: Chuck Lever <cel@kernel.org>,
	linux-nfs@vger.kernel.org,  Jeff Layton <jlayton@kernel.org>
Subject: [PATCH nfs-utils v3 07/11] mountd: give each worker its own netlink command socket
Date: Mon, 14 Sep 2026 09:14:16 -0400	[thread overview]
Message-ID: <20260914-nl-crossmnt-v3-7-a984a6c94829@kernel.org> (raw)
In-Reply-To: <20260914-nl-crossmnt-v3-0-a984a6c94829@kernel.org>

cache_open() opens the nfsd and sunrpc command sockets before
cache_fork_workers() forks, so every worker shares one fd, and a
copy-on-write struct nl_sock carrying the same s_seq_next/s_seq_expect.
Only the notify sockets get nl_socket_disable_seq_check(); the command
sockets keep libnl's default sequence check, which each worker will
happily pass on the other's reply.

Two workers in a command/reply round trip can therefore take each other's
ack or NLMSG_ERROR.  cache_nl_set_reqs() then reports the wrong outcome
and the caller answers the wrong client and path - marking it exported,
retryable, or negatively cached.

Reopen both command sockets in the worker child.  genl_connect() binds a
fresh port id, so each worker gets its own reply stream.  The notify
sockets stay shared: they are receive-only, have the sequence check
disabled, and are nonblocking, so whichever worker gets there first takes
the notification.

Signed-off-by: Jeff Layton <jlayton@kernel.org>
Assisted-by: LLM
---
 support/export/cache.c | 37 +++++++++++++++++++++++++++++++++++--
 1 file changed, 35 insertions(+), 2 deletions(-)

diff --git a/support/export/cache.c b/support/export/cache.c
index 6a5cdd9d9670..cbe3df83b0eb 100644
--- a/support/export/cache.c
+++ b/support/export/cache.c
@@ -3994,6 +3994,29 @@ cache_wait_for_workers(char *prog)
 	}
 }
 
+/*
+ * Replace a command socket inherited across fork().  Returns @old if a new one
+ * cannot be had; sharing it is worse than having one, but not by as much as
+ * having none.
+ */
+static struct nl_sock *nl_cmd_sock_reopen(struct nl_sock *old)
+{
+	struct nl_sock *sock;
+
+	if (!old)
+		return NULL;
+
+	sock = nl_sock_setup();
+	if (!sock) {
+		xlog(L_WARNING, "%s: cannot reopen netlink command socket,"
+		     " sharing the inherited one", __func__);
+		return old;
+	}
+
+	nl_socket_free(old);
+	return sock;
+}
+
 /* Fork num_threads worker children and wait for them */
 int
 cache_fork_workers(char *prog, int num_threads)
@@ -4015,12 +4038,22 @@ cache_fork_workers(char *prog, int num_threads)
 		if (pid == 0) {
 			/* worker child */
 
+			/*
+			 * A command socket carries a reply back to the process
+			 * that sent the request, so it cannot be shared.  Every
+			 * worker inherits the same fd and the same copied
+			 * sequence counters, and two of them mid-round-trip
+			 * will take each other's ack.
+			 */
+			nfsd_nl_cmd_sock = nl_cmd_sock_reopen(nfsd_nl_cmd_sock);
+			sunrpc_nl_cmd_sock =
+				nl_cmd_sock_reopen(sunrpc_nl_cmd_sock);
+
 			/*
 			 * cache_open() drains the netlink downcalls before we
 			 * get here, so anything it deferred is now on the retry
 			 * queues of every worker.  Let the first worker own
-			 * those, or they get answered once per worker over the
-			 * shared command socket.
+			 * those, or they get answered once per worker.
 			 */
 			if (i > 0) {
 				delayed_export_flush();

-- 
2.55.0


  parent reply	other threads:[~2026-09-14 13:14 UTC|newest]

Thread overview: 14+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-14 13:14 [PATCH nfs-utils v3 00/11] mountd: bugfixes for up/downcall interfaces Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 01/11] mountd: factor out the per-path export attribute computation Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 02/11] mountd: handle unmountable paths and junctions in the netlink downcall Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 03/11] mountd: answer requests the kernel rejects on " Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 04/11] mountd: don't leak the parent export's fsid onto crossmnt submounts Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 05/11] mountd: retry unresolvable fsid lookups on the netlink downcall Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 06/11] mountd: bound the junction path before copying it into e_path Jeff Layton
2026-09-14 13:14 ` Jeff Layton [this message]
2026-09-14 13:14 ` [PATCH nfs-utils v3 08/11] mountd: retry export attributes that fail to resolve for a passing reason Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 09/11] mountd: drop a deferred fsid lookup once it has been answered Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 10/11] mountd: bound the retry queues Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 11/11] mountd: don't warn about pipefs submounts nobody asked to export Jeff Layton
2026-09-15  9:24 ` [PATCH nfs-utils v3 00/11] mountd: bugfixes for up/downcall interfaces Mantas Mikulėnas
2026-09-17  7:21 ` Steve Dickson

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260914-nl-crossmnt-v3-7-a984a6c94829@kernel.org \
    --to=jlayton@kernel.org \
    --cc=cel@kernel.org \
    --cc=grawity@gmail.com \
    --cc=linux-nfs@vger.kernel.org \
    --cc=steved@redhat.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox