* [PATCH nfs-utils v3 00/11] mountd: bugfixes for up/downcall interfaces
@ 2026-09-14 13:14 Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 01/11] mountd: factor out the per-path export attribute computation Jeff Layton
` (12 more replies)
0 siblings, 13 replies; 14+ messages in thread
From: Jeff Layton @ 2026-09-14 13:14 UTC (permalink / raw)
To: Steve Dickson, Mantas Mikulėnas; +Cc: Chuck Lever, linux-nfs, Jeff Layton
The main difference in this version is some cleanup to the warning
messages, so that we don't warn about internal events. Mantas, if you're
able to test this version too, that would be great.
Mantas reported a problem with his configuration that was using crossmnt
and autofs to export a number of btrfs filesystems. Some analysis with
Claude also uncovered a few other bugs. This version also fixes a
number of problems that Sashiko flagged. Some of them were preexisting
but seemed like good things to fix.
Signed-off-by: Jeff Layton <jlayton@kernel.org>
---
Changes in v3:
- Stop warning twice per unexportable path
- Keep a warning for the ip_map and unix_gid answers, which have no fallback
- Don't warn about a submount nobody asked to export
- Quieten the same message on the pipefs downcall (preexisting)
- Link to v2: https://lore.kernel.org/r/20260911-nl-crossmnt-v2-0-f4d869951c5e@kernel.org
Changes in v2:
- Retry instead of caching a negative when the failure may be transient:
kernel errno, statfs, fsidd, and an undelivered negative fallback
- Drop a deferred request once it has been answered
- Bound the junction path copied into exportent.e_path (preexisting)
- Give each worker its own netlink command socket (preexisting)
- Bound the retry queues (preexisting)
- Link to v1: https://lore.kernel.org/r/20260909-nl-crossmnt-v1-0-4064b5a3bd85@kernel.org
---
Jeff Layton (11):
mountd: factor out the per-path export attribute computation
mountd: handle unmountable paths and junctions in the netlink downcall
mountd: answer requests the kernel rejects on the netlink downcall
mountd: don't leak the parent export's fsid onto crossmnt submounts
mountd: retry unresolvable fsid lookups on the netlink downcall
mountd: bound the junction path before copying it into e_path
mountd: give each worker its own netlink command socket
mountd: retry export attributes that fail to resolve for a passing reason
mountd: drop a deferred fsid lookup once it has been answered
mountd: bound the retry queues
mountd: don't warn about pipefs submounts nobody asked to export
support/export/cache.c | 1391 +++++++++++++++++++++++++++++++++++++-----------
1 file changed, 1084 insertions(+), 307 deletions(-)
---
base-commit: 630393e9c73667b3fda6a7507cb5171ec0ec38dc
change-id: 20260908-nl-crossmnt-85584b71c721
Best regards,
--
Jeff Layton <jlayton@kernel.org>
^ permalink raw reply [flat|nested] 14+ messages in thread
* [PATCH nfs-utils v3 01/11] mountd: factor out the per-path export attribute computation
2026-09-14 13:14 [PATCH nfs-utils v3 00/11] mountd: bugfixes for up/downcall interfaces Jeff Layton
@ 2026-09-14 13:14 ` Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 02/11] mountd: handle unmountable paths and junctions in the netlink downcall Jeff Layton
` (11 subsequent siblings)
12 siblings, 0 replies; 14+ messages in thread
From: Jeff Layton @ 2026-09-14 13:14 UTC (permalink / raw)
To: Steve Dickson, Mantas Mikulėnas; +Cc: Chuck Lever, linux-nfs, Jeff Layton
dump_to_cache() adjusts the export flags, fsid and uuid before handing
them to the kernel, because the upcall path may be a crossmnt submount
rather than the exported path itself. Pull that into
export_attrs_build() so the netlink downcall can use it too.
No functional change.
Assisted-by: LLM
Signed-off-by: Jeff Layton <jlayton@kernel.org>
---
support/export/cache.c | 143 ++++++++++++++++++++++++++++++-------------------
1 file changed, 88 insertions(+), 55 deletions(-)
diff --git a/support/export/cache.c b/support/export/cache.c
index 059f48a7069f..e8b13409ad7a 100644
--- a/support/export/cache.c
+++ b/support/export/cache.c
@@ -1072,6 +1072,86 @@ static void write_xprtsec(char **bp, int *blen, struct exportent *ep)
qword_addint(bp, blen, p->info->number);
}
+static int can_reexport_via_fsidnum(struct exportent *exp, struct statfs *st)
+{
+ if (st->f_type != 0x6969 /* NFS_SUPER_MAGIC */)
+ return 0;
+
+ return exp->e_reexport == REEXP_PREDEFINED_FSIDNUM ||
+ exp->e_reexport == REEXP_AUTO_FSIDNUM;
+}
+
+/* What to hand the kernel for one (path, export) pair */
+struct export_attrs {
+ int flags;
+ uint32_t fsidnum;
+ int sec_mask; /* mask for the per-flavor flags */
+ int sec_extra; /* extra per-flavor flags */
+ char uuid[16];
+ bool have_uuid;
+};
+
+/*
+ * An upcall path may be a submount below the exported one, when the export
+ * is marked crossmnt. Such a submount is a filesystem in its own right, so
+ * it must not inherit the parent's fsid= or uuid= - if it does, both end up
+ * claiming the same filehandles and the client sees ESTALE.
+ *
+ * Returns 0, or -1 with errno set if @path cannot be exported at all.
+ */
+static int export_attrs_build(struct export_attrs *ea, char *path,
+ struct exportent *exp)
+{
+ int different_fs = strcmp(path, exp->e_path) != 0;
+ int flag_mask = different_fs ? ~NFSEXP_FSID : ~0;
+ int do_fsidnum = 0;
+
+ memset(ea, 0, sizeof(*ea));
+ ea->fsidnum = exp->e_fsid;
+
+ if (different_fs) {
+ struct statfs st;
+
+ if (nfsd_path_statfs(path, &st)) {
+ xlog(L_WARNING, "unable to statfs %s", path);
+ errno = EINVAL;
+ return -1;
+ }
+
+ /* A re-exported submount gets an fsid= of its own instead */
+ if (can_reexport_via_fsidnum(exp, &st)) {
+ do_fsidnum = 1;
+ flag_mask = ~0;
+ }
+ }
+
+ if (do_fsidnum) {
+ uint32_t search_fsidnum = 0;
+
+ if (exp->e_reexport != REEXP_NONE &&
+ reexpdb_fsidnum_by_path(path, &search_fsidnum,
+ exp->e_reexport == REEXP_AUTO_FSIDNUM) == 0) {
+ errno = EINVAL;
+ return -1;
+ }
+ ea->fsidnum = search_fsidnum;
+ ea->flags = exp->e_flags | NFSEXP_FSID;
+ ea->sec_extra = NFSEXP_FSID;
+ } else {
+ ea->flags = exp->e_flags & flag_mask;
+ }
+ ea->sec_mask = flag_mask;
+
+ if (exp->e_uuid && !different_fs) {
+ get_uuid(exp->e_uuid, 16, ea->uuid);
+ ea->have_uuid = true;
+ } else if ((exp->e_flags & flag_mask & NFSEXP_FSID) == 0) {
+ ea->have_uuid = uuid_by_path(path, 0, 16, ea->uuid);
+ }
+
+ return 0;
+}
+
/*
* Netlink-based svc_export cache support.
*
@@ -2501,15 +2581,6 @@ static void cache_sunrpc_nl_process(void)
cache_nl_process_unix_gid();
}
-static int can_reexport_via_fsidnum(struct exportent *exp, struct statfs *st)
-{
- if (st->f_type != 0x6969 /* NFS_SUPER_MAGIC */)
- return 0;
-
- return exp->e_reexport == REEXP_PREDEFINED_FSIDNUM ||
- exp->e_reexport == REEXP_AUTO_FSIDNUM;
-}
-
static int dump_to_cache(int f, char *buf, int blen, char *domain,
char *path, struct exportent *exp, int ttl)
{
@@ -2524,60 +2595,22 @@ static int dump_to_cache(int f, char *buf, int blen, char *domain,
qword_add(&bp, &blen, domain);
qword_add(&bp, &blen, path);
if (exp) {
- int different_fs = strcmp(path, exp->e_path) != 0;
- int flag_mask = different_fs ? ~NFSEXP_FSID : ~0;
- int rc, do_fsidnum = 0;
- uint32_t fsidnum = exp->e_fsid;
-
- if (different_fs) {
- struct statfs st;
-
- rc = nfsd_path_statfs(path, &st);
- if (rc) {
- xlog(L_WARNING, "unable to statfs %s", path);
- errno = EINVAL;
- return -1;
- }
+ struct export_attrs ea;
- if (can_reexport_via_fsidnum(exp, &st)) {
- do_fsidnum = 1;
- flag_mask = ~0;
- }
- }
+ if (export_attrs_build(&ea, path, exp) < 0)
+ return -1;
qword_adduint(&bp, &blen, now + exp->e_ttl);
-
- if (do_fsidnum) {
- uint32_t search_fsidnum = 0;
- if (exp->e_reexport != REEXP_NONE && reexpdb_fsidnum_by_path(path, &search_fsidnum,
- exp->e_reexport == REEXP_AUTO_FSIDNUM) == 0) {
- errno = EINVAL;
- return -1;
- }
- fsidnum = search_fsidnum;
- qword_addint(&bp, &blen, exp->e_flags | NFSEXP_FSID);
- } else {
- qword_addint(&bp, &blen, exp->e_flags & flag_mask);
- }
-
+ qword_addint(&bp, &blen, ea.flags);
qword_addint(&bp, &blen, exp->e_anonuid);
qword_addint(&bp, &blen, exp->e_anongid);
- qword_addint(&bp, &blen, fsidnum);
+ qword_addint(&bp, &blen, ea.fsidnum);
write_fsloc(&bp, &blen, exp);
- write_secinfo(&bp, &blen, exp, flag_mask, do_fsidnum ? NFSEXP_FSID : 0);
- if (exp->e_uuid == NULL || different_fs) {
- char u[16];
- if ((exp->e_flags & flag_mask & NFSEXP_FSID) == 0 &&
- uuid_by_path(path, 0, 16, u)) {
- qword_add(&bp, &blen, "uuid");
- qword_addhex(&bp, &blen, u, 16);
- }
- } else {
- char u[16];
- get_uuid(exp->e_uuid, 16, u);
+ write_secinfo(&bp, &blen, exp, ea.sec_mask, ea.sec_extra);
+ if (ea.have_uuid) {
qword_add(&bp, &blen, "uuid");
- qword_addhex(&bp, &blen, u, 16);
+ qword_addhex(&bp, &blen, ea.uuid, 16);
}
write_xprtsec(&bp, &blen, exp);
xlog(D_AUTH, "granted access to %s for %s",
--
2.55.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH nfs-utils v3 02/11] mountd: handle unmountable paths and junctions in the netlink downcall
2026-09-14 13:14 [PATCH nfs-utils v3 00/11] mountd: bugfixes for up/downcall interfaces Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 01/11] mountd: factor out the per-path export attribute computation Jeff Layton
@ 2026-09-14 13:14 ` Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 03/11] mountd: answer requests the kernel rejects on " Jeff Layton
` (10 subsequent siblings)
12 siblings, 0 replies; 14+ messages in thread
From: Jeff Layton @ 2026-09-14 13:14 UTC (permalink / raw)
To: Steve Dickson, Mantas Mikulėnas; +Cc: Chuck Lever, linux-nfs, Jeff Layton
Two things nfsd_export() does that cache_nl_process_export() did not:
- If is_mountpoint() fails with an error that isn't a plain lookup
failure, we can't tell whether the export is available. Defer the
request instead of answering "not exported".
- When no export matches, check for a junction before denying, so
referrals still work. Only when the client resolved, as nfsd_export()
does: client_check() dereferences the addrinfo for wildcard and
netgroup clients.
Split lookup_nonexport() so both downcalls share the junction lookup, and
factor the per-request work into nl_add_export_req() so the deferred path
can reuse it. Deferred requests go on their own list, deduped on
(client, path) and retried from cache_process() every RETRY_SEC. The
kernel keeps the upcall pending in the meantime.
Assisted-by: LLM
Signed-off-by: Jeff Layton <jlayton@kernel.org>
---
support/export/cache.c | 324 +++++++++++++++++++++++++++++++++++++++----------
1 file changed, 258 insertions(+), 66 deletions(-)
diff --git a/support/export/cache.c b/support/export/cache.c
index e8b13409ad7a..0ad98c539407 100644
--- a/support/export/cache.c
+++ b/support/export/cache.c
@@ -784,6 +784,18 @@ struct delayed {
struct delayed *next;
} *delayed;
+/* Fold one retry queue head's deadline into the running minimum */
+static time_t retry_delay(time_t *last_attempt, time_t now, time_t delay)
+{
+ time_t d;
+
+ if (*last_attempt > now)
+ /* Clock updated - retry immediately */
+ *last_attempt = now - RETRY_SEC;
+ d = *last_attempt + RETRY_SEC - now;
+ return d < delay ? d : delay;
+}
+
static int nfsd_handle_fh(int f, char *bp, int blen)
{
/* request are:
@@ -1162,8 +1174,18 @@ static int export_attrs_build(struct export_attrs *ea, char *path,
* NFSD_CMD_SVC_EXPORT_SET_REQS.
*/
static nfs_export *lookup_export(char *dom, char *path, struct addrinfo *ai);
+static struct exportent *lookup_nonexport_ent(char *dom, char *path,
+ struct addrinfo *ai);
static struct nl_msg *cache_nl_new_msg(int family, int cmd, int flags);
+static void free_junction(struct exportent *eep)
+{
+ if (!eep)
+ return;
+ exportent_release(eep);
+ free(eep);
+}
+
static struct nl_sock *nfsd_nl_notify_sock; /* multicast notifications */
static struct nl_sock *nfsd_nl_cmd_sock; /* GET_REQS / SET_REQS commands */
static int nfsd_nl_family;
@@ -1632,6 +1654,197 @@ static int cache_nl_set_reqs(struct nl_sock *sock, struct nl_msg *msg)
return ret;
}
+static bool nl_msg_has_reqs(struct nl_msg *msg)
+{
+ return genlmsg_attrlen(nlmsg_data(nlmsg_hdr(msg)), 0) > 0;
+}
+
+enum export_result {
+ EXPORT_ANSWERED,
+ EXPORT_RETRY, /* not resolvable yet, ask again later */
+ EXPORT_NOMEM, /* *msgp is gone, caller must give up */
+};
+
+/*
+ * Resolve one svc_export request and append the answer to *msgp, sending
+ * and replacing the message if it fills up.
+ */
+static enum export_result nl_add_export_req(struct nl_msg **msgp, char *dom,
+ char *path)
+{
+ struct addrinfo *ai = NULL;
+ nfs_export *found = NULL;
+ struct exportent *epp = NULL;
+ struct exportent *junction = NULL;
+ enum export_result res = EXPORT_ANSWERED;
+ int ttl = 0;
+
+ if (is_ipaddr_client(dom)) {
+ ai = lookup_client_addr(dom);
+ if (!ai)
+ xlog(D_AUTH, "%s: failed to resolve client %s",
+ __func__, dom);
+ }
+
+ if (ai || !is_ipaddr_client(dom)) {
+ found = lookup_export(dom, path, ai);
+ if (!found) {
+ junction = lookup_nonexport_ent(dom, path, ai);
+ epp = junction;
+ }
+ }
+
+ if (found) {
+ char *mp = found->m_export.e_mountpoint;
+
+ if (mp && !*mp)
+ mp = found->m_export.e_path;
+ errno = 0;
+ if (mp && !is_mountpoint(mp)) {
+ /*
+ * A strange error means we can't tell whether it is a
+ * mountpoint. Retry later rather than answer wrongly.
+ */
+ if (errno != 0 && !path_lookup_error(errno)) {
+ res = EXPORT_RETRY;
+ goto out;
+ }
+ /* Exportpoint is not mounted, so tell kernel it
+ * is not available. This will cause it not to
+ * appear in the V4 Pseudo-root, so a "mount" of
+ * this path will fail, just like with V3.
+ */
+ xlog(L_WARNING,
+ "Cannot export path '%s': not a mountpoint", path);
+ ttl = 60;
+ } else {
+ epp = &found->m_export;
+ }
+ }
+
+ if (nfsd_nl_add_export(*msgp, dom, path, epp, ttl) < 0) {
+ cache_nl_set_reqs(nfsd_nl_cmd_sock, *msgp);
+ nlmsg_free(*msgp);
+ *msgp = cache_nl_new_msg(nfsd_nl_family,
+ NFSD_CMD_SVC_EXPORT_SET_REQS, 0);
+ if (!*msgp) {
+ res = EXPORT_NOMEM;
+ goto out;
+ }
+ if (nfsd_nl_add_export(*msgp, dom, path, epp, ttl) < 0)
+ xlog(L_WARNING, "%s: skipping oversized entry for %s",
+ __func__, path);
+ }
+out:
+ free_junction(junction);
+ nfs_freeaddrinfo(ai);
+ return res;
+}
+
+/*
+ * is_mountpoint() can fail with a strange error - the ETIMEDOUT a re-exported
+ * "softerr" NFS mount can give, say - leaving us unable to say whether the
+ * path is exportable. Set the request aside and try again later.
+ */
+struct delayed_export {
+ char *client;
+ char *path;
+ time_t last_attempt;
+ struct delayed_export *next;
+};
+
+static struct delayed_export *delayed_export;
+
+static void delayed_export_enqueue(struct delayed_export *d)
+{
+ struct delayed_export **dp = &delayed_export;
+
+ d->last_attempt = time(NULL);
+ d->next = NULL;
+ while (*dp)
+ dp = &(*dp)->next;
+ *dp = d;
+}
+
+static void delayed_export_free(struct delayed_export *d)
+{
+ free(d->client);
+ free(d->path);
+ free(d);
+}
+
+static void delayed_export_flush(void)
+{
+ while (delayed_export) {
+ struct delayed_export *d = delayed_export;
+
+ delayed_export = d->next;
+ delayed_export_free(d);
+ }
+}
+
+static void nl_delay_export(char *dom, char *path)
+{
+ struct delayed_export *d;
+
+ for (d = delayed_export; d; d = d->next)
+ if (!strcmp(d->client, dom) && !strcmp(d->path, path))
+ return;
+
+ d = calloc(1, sizeof(*d));
+ if (!d)
+ return;
+
+ d->client = strdup(dom);
+ d->path = strdup(path);
+ if (!d->client || !d->path) {
+ delayed_export_free(d);
+ return;
+ }
+
+ delayed_export_enqueue(d);
+}
+
+/*
+ * Retry the oldest deferred request if it is due. Entries are queued in
+ * time order, so only the head can be ready.
+ */
+static void nl_retry_export(void)
+{
+ struct delayed_export *d = delayed_export;
+ struct nl_msg *msg;
+
+ if (!d || d->last_attempt + RETRY_SEC > time(NULL))
+ return;
+
+ delayed_export = d->next;
+ d->next = NULL;
+
+ auth_reload();
+
+ msg = cache_nl_new_msg(nfsd_nl_family,
+ NFSD_CMD_SVC_EXPORT_SET_REQS, 0);
+ if (!msg) {
+ delayed_export_enqueue(d);
+ return;
+ }
+
+ switch (nl_add_export_req(&msg, d->client, d->path)) {
+ case EXPORT_ANSWERED:
+ if (nl_msg_has_reqs(msg))
+ cache_nl_set_reqs(nfsd_nl_cmd_sock, msg);
+ delayed_export_free(d);
+ break;
+ case EXPORT_RETRY:
+ delayed_export_enqueue(d);
+ break;
+ case EXPORT_NOMEM:
+ delayed_export_enqueue(d);
+ return; /* msg is already gone */
+ }
+ nlmsg_free(msg);
+}
+
static void cache_nl_process_export(void)
{
struct export_req *reqs = NULL;
@@ -1657,57 +1870,19 @@ static void cache_nl_process_export(void)
goto out_free;
for (i = 0; i < nreqs; i++) {
- char *dom = reqs[i].client;
- char *path = reqs[i].path;
- struct addrinfo *ai = NULL;
- nfs_export *found = NULL;
- struct exportent *epp = NULL;
- int ttl = 0;
-
- if (is_ipaddr_client(dom)) {
- ai = lookup_client_addr(dom);
- if (!ai)
- xlog(D_AUTH, "cache_nl_process_export: "
- "failed to resolve client %s", dom);
- }
-
- if (ai || !is_ipaddr_client(dom))
- found = lookup_export(dom, path, ai);
-
- if (found) {
- char *mp = found->m_export.e_mountpoint;
-
- if (mp && !*mp)
- mp = found->m_export.e_path;
- if (mp && !is_mountpoint(mp)) {
- xlog(L_WARNING,
- "Cannot export path '%s': not a mountpoint",
- path);
- ttl = 60;
- } else {
- epp = &found->m_export;
- }
- }
-
- if (nfsd_nl_add_export(msg, dom, path, epp, ttl) < 0) {
- cache_nl_set_reqs(nfsd_nl_cmd_sock, msg);
- nlmsg_free(msg);
- msg = cache_nl_new_msg(nfsd_nl_family,
- NFSD_CMD_SVC_EXPORT_SET_REQS, 0);
- if (!msg) {
- nfs_freeaddrinfo(ai);
- goto out_free;
- }
- if (nfsd_nl_add_export(msg, dom, path,
- epp, ttl) < 0)
- xlog(L_WARNING, "%s: skipping oversized "
- "entry for %s", __func__, path);
+ switch (nl_add_export_req(&msg, reqs[i].client, reqs[i].path)) {
+ case EXPORT_ANSWERED:
+ break;
+ case EXPORT_RETRY:
+ nl_delay_export(reqs[i].client, reqs[i].path);
+ break;
+ case EXPORT_NOMEM:
+ goto out_free;
}
-
- nfs_freeaddrinfo(ai);
}
- cache_nl_set_reqs(nfsd_nl_cmd_sock, msg);
+ if (nl_msg_has_reqs(msg))
+ cache_nl_set_reqs(nfsd_nl_cmd_sock, msg);
nlmsg_free(msg);
out_free:
@@ -2971,29 +3146,32 @@ out:
return exp;
}
-static void lookup_nonexport(int f, char *buf, int buflen, char *dom, char *path,
+static struct exportent *lookup_nonexport_ent(char *dom, char *path,
struct addrinfo *ai)
{
- struct exportent *eep;
-
- eep = lookup_junction(dom, path, ai);
- dump_to_cache(f, buf, buflen, dom, path, eep, 0);
- if (eep == NULL)
- return;
- exportent_release(eep);
- free(eep);
+ return lookup_junction(dom, path, ai);
}
#else /* !HAVE_JUNCTION_SUPPORT */
-static void lookup_nonexport(int f, char *buf, int buflen, char *dom, char *path,
- struct addrinfo *UNUSED(ai))
+static struct exportent *lookup_nonexport_ent(char *UNUSED(dom),
+ char *UNUSED(path), struct addrinfo *UNUSED(ai))
{
- dump_to_cache(f, buf, buflen, dom, path, NULL, 0);
+ return NULL;
}
#endif /* !HAVE_JUNCTION_SUPPORT */
+static void lookup_nonexport(int f, char *buf, int buflen, char *dom, char *path,
+ struct addrinfo *ai)
+{
+ struct exportent *eep;
+
+ eep = lookup_nonexport_ent(dom, path, ai);
+ dump_to_cache(f, buf, buflen, dom, path, eep, 0);
+ free_junction(eep);
+}
+
static void nfsd_export(int f)
{
/* requests are:
@@ -3201,13 +3379,15 @@ int cache_process(fd_set *readfds)
cache_set_fds(readfds);
v4clients_set_fds(readfds);
- if (delayed) {
+ if (delayed || delayed_export) {
time_t now = time(NULL);
- time_t delay;
- if (delayed->last_attempt > now)
- /* Clock updated - retry immediately */
- delayed->last_attempt = now - RETRY_SEC;
- delay = delayed->last_attempt + RETRY_SEC - now;
+ time_t delay = RETRY_SEC;
+
+ if (delayed)
+ delay = retry_delay(&delayed->last_attempt, now, delay);
+ if (delayed_export)
+ delay = retry_delay(&delayed_export->last_attempt, now,
+ delay);
if (delay < 0)
delay = 0;
tv.tv_sec = delay;
@@ -3225,6 +3405,8 @@ int cache_process(fd_set *readfds)
}
}
+ nl_retry_export();
+
switch (selret) {
case -1:
if (errno == EINTR || errno == ECONNREFUSED
@@ -3430,6 +3612,16 @@ cache_fork_workers(char *prog, int num_threads)
if (pid == 0) {
/* worker child */
+ /*
+ * cache_open() drains the netlink downcalls before we
+ * get here, so anything it deferred is now on the retry
+ * queue of every worker. Let the first worker own
+ * those, or they get answered once per worker over the
+ * shared command socket.
+ */
+ if (i > 0)
+ delayed_export_flush();
+
/* Re-enable the default action on SIGTERM et al
* so that workers die naturally when sent them.
* Only the parent unregisters with pmap and
--
2.55.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH nfs-utils v3 03/11] mountd: answer requests the kernel rejects on the netlink downcall
2026-09-14 13:14 [PATCH nfs-utils v3 00/11] mountd: bugfixes for up/downcall interfaces Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 01/11] mountd: factor out the per-path export attribute computation Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 02/11] mountd: handle unmountable paths and junctions in the netlink downcall Jeff Layton
@ 2026-09-14 13:14 ` Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 04/11] mountd: don't leak the parent export's fsid onto crossmnt submounts Jeff Layton
` (9 subsequent siblings)
12 siblings, 0 replies; 14+ messages in thread
From: Jeff Layton @ 2026-09-14 13:14 UTC (permalink / raw)
To: Steve Dickson, Mantas Mikulėnas; +Cc: Chuck Lever, linux-nfs, Jeff Layton
The kernel refuses an svc_export it cannot build a filehandle for - a 9p
submount picked up by crossmnt, say - and it fails the whole SET_REQS
message when it does. cache_nl_process_export() batches every pending
request into one message and never looks at the result, so:
- every entry queued behind the bad one is dropped
- nothing downgrades the bad path to a negative entry
- the request stays pending and the client hangs on the lookup
Pending requests then pile up on each notification:
cache_nl_process_export: 5 pending export requests
cache_nl_set_reqs: SET_REQS failed: -7
nfsd_export() has no such hole: a rejected dump_to_cache() write returns
-1 and it answers negative instead.
Keep the batch for the fast path, but resubmit it one entry at a time
when the kernel rejects it, and answer negative for whichever entries it
still refuses. nfsd_nl_svc_export_set_reqs_doit() commits each entry as
it parses it and stops at the first failure, so the resubmit re-sends
some entries the kernel already took; that is harmless, as the update is
idempotent, and it is the only way to find the one that failed.
nl_add_export_req() no longer flushes the message itself; the caller
owns the batch so it knows what to resubmit.
Only fall back to a negative entry when the kernel actually answered.
libnl folds its own errors into the same NLE_* space as the kernel's, so
cache_nl_set_reqs() now reports the kernel's errno separately: a broken
socket must not cache "not exported" for default_ttl. A negative entry
the kernel already refused - the path or the client's auth_domain went
away - is not resent as another negative either.
An entry too large for a message of its own is a third case. mountd
never sent it, so the kernel never refused it. Log the size failure and
answer negative, rather than blame the filesystem for it.
The resubmit makes a refusal routine, and the kernel tends to notify one
request at a time, so warning about it printed the same failure twice per
bad path. Trace it at D_NETLINK instead, with nl_geterror() and the
kernel's errno rather than a raw number: the -7 above is NLE_INVAL, libnl's
own code for the kernel's EINVAL, not an errno. Callers that have no
fallback keep their warning through the new cache_nl_flush_reqs(), which
ip_map and unix_gid now use, and the -ENOMEM in cache_nl_set_reqs() becomes
-NLE_NOMEM so the value is in the space nl_geterror() reads.
A path nobody asked to export - /proc or /sys below a crossmnt "/" - is
expected to be refused and the admin cannot act on it, so the negative
fallback only warns for a path exported in its own right.
Reported-by: Mantas Mikulėnas <grawity@gmail.com>
Assisted-by: LLM
Signed-off-by: Jeff Layton <jlayton@kernel.org>
---
support/export/cache.c | 367 +++++++++++++++++++++++++++++++++++++++----------
1 file changed, 294 insertions(+), 73 deletions(-)
diff --git a/support/export/cache.c b/support/export/cache.c
index 0ad98c539407..5bdc00c5847f 100644
--- a/support/export/cache.c
+++ b/support/export/cache.c
@@ -1623,18 +1623,59 @@ static struct nl_msg *cache_nl_new_msg(int family, int cmd, int flags)
return msg;
}
-static int cache_nl_set_reqs(struct nl_sock *sock, struct nl_msg *msg)
+/* State for one SET_REQS round trip */
+struct set_reqs_status {
+ int done;
+ int kern_err; /* errno the kernel answered with */
+};
+
+static int nl_set_reqs_finish_cb(struct nl_msg *UNUSED(msg), void *arg)
+{
+ struct set_reqs_status *st = arg;
+
+ st->done = 1;
+ return NL_STOP;
+}
+
+static int nl_set_reqs_error_cb(struct sockaddr_nl *UNUSED(nla),
+ struct nlmsgerr *nlerr, void *arg)
+{
+ struct set_reqs_status *st = arg;
+
+ st->done = 1;
+ st->kern_err = nlerr->error;
+ return NL_STOP;
+}
+
+/*
+ * Send @msg and wait for the kernel to ack it. Returns 0 on success.
+ *
+ * libnl folds the kernel's errno into its own NLE_* space, and a local
+ * failure lands in the same space, so the return value cannot say whether
+ * the kernel looked at the message at all. When @kern_errp is given it is
+ * set to the negative errno the kernel replied with, or left at 0 if it
+ * never got that far.
+ *
+ * A refusal is routine - the batch path uses one to find the entry the
+ * kernel would not take - so only trace it here and leave it to the caller
+ * to decide what is worth a warning.
+ */
+static int cache_nl_set_reqs(struct nl_sock *sock, struct nl_msg *msg,
+ int *kern_errp)
{
+ struct set_reqs_status st = {};
struct nl_cb *cb;
- int done = 0;
int ret;
+ if (kern_errp)
+ *kern_errp = 0;
+
cb = nl_cb_alloc(NL_CB_DEFAULT);
if (!cb)
- return -ENOMEM;
+ return -NLE_NOMEM;
- nl_cb_set(cb, NL_CB_ACK, NL_CB_CUSTOM, nl_finish_cb, &done);
- nl_cb_err(cb, NL_CB_CUSTOM, nl_error_cb, &done);
+ nl_cb_set(cb, NL_CB_ACK, NL_CB_CUSTOM, nl_set_reqs_finish_cb, &st);
+ nl_cb_err(cb, NL_CB_CUSTOM, nl_set_reqs_error_cb, &st);
ret = nl_send_auto(sock, msg);
if (ret < 0) {
@@ -1642,15 +1683,19 @@ static int cache_nl_set_reqs(struct nl_sock *sock, struct nl_msg *msg)
return ret;
}
- while (!done) {
+ while (!st.done) {
ret = nl_recvmsgs(sock, cb);
if (ret < 0)
break;
}
nl_cb_put(cb);
+ if (kern_errp)
+ *kern_errp = st.kern_err;
if (ret < 0)
- xlog(L_WARNING, "%s: SET_REQS failed: %d", __func__, ret);
+ xlog(D_NETLINK, "%s: SET_REQS failed: %s (kernel: %s)",
+ __func__, nl_geterror(ret),
+ st.kern_err ? strerror(-st.kern_err) : "no answer");
return ret;
}
@@ -1659,24 +1704,54 @@ static bool nl_msg_has_reqs(struct nl_msg *msg)
return genlmsg_attrlen(nlmsg_data(nlmsg_hdr(msg)), 0) > 0;
}
+/*
+ * Send a message the caller has no fallback for. There is nothing to try
+ * instead when the kernel refuses an ip_map or unix_gid answer, so warn.
+ */
+static void cache_nl_flush_reqs(struct nl_sock *sock, struct nl_msg *msg,
+ const char *what)
+{
+ int ret;
+
+ if (!nl_msg_has_reqs(msg))
+ return;
+
+ ret = cache_nl_set_reqs(sock, msg, NULL);
+ if (ret < 0)
+ xlog(L_WARNING, "failed to answer %s requests: %s",
+ what, nl_geterror(ret));
+}
+
enum export_result {
- EXPORT_ANSWERED,
+ EXPORT_ANSWERED, /* positive entry added */
+ EXPORT_DENIED, /* negative entry added */
EXPORT_RETRY, /* not resolvable yet, ask again later */
- EXPORT_NOMEM, /* *msgp is gone, caller must give up */
+ EXPORT_FULL, /* did not fit, flush the message and re-add */
};
/*
- * Resolve one svc_export request and append the answer to *msgp, sending
- * and replacing the message if it fills up.
+ * Did the admin ask for @path itself, or did we get here from a crossmnt
+ * parent or the v4 pseudoroot? A filesystem nobody asked to export - /proc
+ * or /sys below a crossmnt "/", say - is expected to be unexportable, and
+ * saying so once per TTL is just noise.
+ */
+static bool export_is_explicit(nfs_export *found, char *path)
+{
+ return found && !strcmp(found->m_export.e_path, path);
+}
+
+/*
+ * Resolve one svc_export request and append the answer to @msg. @explicitp,
+ * when given, says whether @path is an export in its own right.
*/
-static enum export_result nl_add_export_req(struct nl_msg **msgp, char *dom,
- char *path)
+static enum export_result nl_add_export_req(struct nl_msg *msg, char *dom,
+ char *path, bool *explicitp)
{
struct addrinfo *ai = NULL;
nfs_export *found = NULL;
struct exportent *epp = NULL;
struct exportent *junction = NULL;
- enum export_result res = EXPORT_ANSWERED;
+ enum export_result res;
int ttl = 0;
if (is_ipaddr_client(dom)) {
@@ -1694,6 +1769,9 @@ static enum export_result nl_add_export_req(struct nl_msg **msgp, char *dom,
}
}
+ if (explicitp)
+ *explicitp = export_is_explicit(found, path);
+
if (found) {
char *mp = found->m_export.e_mountpoint;
@@ -1722,25 +1800,137 @@ static enum export_result nl_add_export_req(struct nl_msg **msgp, char *dom,
}
}
- if (nfsd_nl_add_export(*msgp, dom, path, epp, ttl) < 0) {
- cache_nl_set_reqs(nfsd_nl_cmd_sock, *msgp);
- nlmsg_free(*msgp);
- *msgp = cache_nl_new_msg(nfsd_nl_family,
- NFSD_CMD_SVC_EXPORT_SET_REQS, 0);
- if (!*msgp) {
- res = EXPORT_NOMEM;
- goto out;
- }
- if (nfsd_nl_add_export(*msgp, dom, path, epp, ttl) < 0)
- xlog(L_WARNING, "%s: skipping oversized entry for %s",
- __func__, path);
- }
+ if (nfsd_nl_add_export(msg, dom, path, epp, ttl) < 0)
+ res = EXPORT_FULL;
+ else
+ res = epp ? EXPORT_ANSWERED : EXPORT_DENIED;
out:
free_junction(junction);
nfs_freeaddrinfo(ai);
return res;
}
+/*
+ * Answer @path negative in a message of its own. Returns EXPORT_ANSWERED once
+ * the kernel has seen the answer - including when it refused it, as a second
+ * attempt would fare no better - and EXPORT_RETRY when the answer never got
+ * that far and is worth sending again.
+ */
+static enum export_result nl_export_negative(char *dom, char *path)
+{
+ struct nl_msg *msg;
+ int kern_err = 0;
+ int ret;
+
+ msg = cache_nl_new_msg(nfsd_nl_family,
+ NFSD_CMD_SVC_EXPORT_SET_REQS, 0);
+ if (!msg)
+ return EXPORT_RETRY;
+
+ if (nfsd_nl_add_export(msg, dom, path, NULL, 0) < 0) {
+ nlmsg_free(msg);
+ return EXPORT_RETRY;
+ }
+
+ ret = cache_nl_set_reqs(nfsd_nl_cmd_sock, msg, &kern_err);
+ nlmsg_free(msg);
+
+ if (ret < 0 && !kern_err)
+ return EXPORT_RETRY;
+ return EXPORT_ANSWERED;
+}
+
+/*
+ * Errors the kernel gives for an export it can never accept. check_export()
+ * answers EINVAL for a filesystem with no export ops, one that needs an fsid=
+ * and has none, and an idmapped mount; ENOTDIR is an inode that is neither a
+ * directory, a symlink, nor a regular file. Anything else it can return -
+ * ENOENT for an auth_domain or a path that is not there yet, ENOMEM, ENODEV
+ * while nfsd shuts down - may well work on the next try.
+ */
+static bool kernel_refused_for_good(int kern_err)
+{
+ switch (kern_err) {
+ case -EINVAL:
+ case -ENOTDIR:
+ return true;
+ }
+ return false;
+}
+
+/*
+ * Answer one request in a message of its own. The kernel rejects an export
+ * it cannot build a filehandle for - a 9p or other filesystem with no export
+ * ops, say - so fall back to a negative entry rather than leaving the request
+ * pending, which is what dump_to_cache() does when the channel write fails.
+ */
+static enum export_result nl_export_one(char *dom, char *path)
+{
+ bool explicit_export = false;
+ enum export_result res;
+ struct nl_msg *msg;
+ int kern_err = 0;
+ bool sent;
+
+ msg = cache_nl_new_msg(nfsd_nl_family,
+ NFSD_CMD_SVC_EXPORT_SET_REQS, 0);
+ if (!msg)
+ return EXPORT_RETRY;
+
+ res = nl_add_export_req(msg, dom, path, &explicit_export);
+ if (res == EXPORT_RETRY) {
+ nlmsg_free(msg);
+ return res;
+ }
+
+ /*
+ * An entry that does not fit a message of its own can never be sent,
+ * and this is not the kernel refusing it. Answer negative rather
+ * than leave the client hung on a request we cannot satisfy.
+ */
+ if (res == EXPORT_FULL) {
+ nlmsg_free(msg);
+ xlog(L_WARNING, "%s: entry for %s is too large to send",
+ __func__, path);
+ return nl_export_negative(dom, path);
+ }
+
+ sent = cache_nl_set_reqs(nfsd_nl_cmd_sock, msg, &kern_err) == 0;
+ nlmsg_free(msg);
+
+ if (sent)
+ return EXPORT_ANSWERED;
+
+ /*
+ * The kernel never answered - a broken socket, or a message we could
+ * not send - so we cannot tell whether the export is usable. Retry
+ * rather than cache a negative entry for default_ttl.
+ */
+ if (!kern_err)
+ return EXPORT_RETRY;
+
+ /*
+ * It may recover from this one. Denying the path would hide a working
+ * export for default_ttl, and a negative entry needs the very
+ * auth_domain the kernel may have just failed to find, so it would
+ * likely be refused as well. Ask again later instead.
+ */
+ if (!kernel_refused_for_good(kern_err)) {
+ xlog(D_GENERAL, "%s: kernel refused %s: %s, will retry",
+ __func__, path, strerror(-kern_err));
+ return EXPORT_RETRY;
+ }
+
+ /* It refused a negative entry; a second one will fare no better */
+ if (res == EXPORT_DENIED)
+ return EXPORT_ANSWERED;
+
+ xlog(explicit_export ? L_WARNING : D_GENERAL,
+ "Cannot export %s, possibly unsupported filesystem"
+ " or fsid= required", path);
+ return nl_export_negative(dom, path);
+}
+
/*
* is_mountpoint() can fail with a strange error - the ETIMEDOUT a re-exported
* "softerr" NFS mount can give, say - leaving us unable to say whether the
@@ -1783,6 +1973,23 @@ static void delayed_export_flush(void)
}
}
+/* Forget any deferred request for @dom and @path; it has been answered */
+static void delayed_export_remove(char *dom, char *path)
+{
+ struct delayed_export **dp = &delayed_export;
+
+ while (*dp) {
+ struct delayed_export *d = *dp;
+
+ if (!strcmp(d->client, dom) && !strcmp(d->path, path)) {
+ *dp = d->next;
+ delayed_export_free(d);
+ return;
+ }
+ dp = &d->next;
+ }
+}
+
static void nl_delay_export(char *dom, char *path)
{
struct delayed_export *d;
@@ -1812,7 +2019,6 @@ static void nl_delay_export(char *dom, char *path)
static void nl_retry_export(void)
{
struct delayed_export *d = delayed_export;
- struct nl_msg *msg;
if (!d || d->last_attempt + RETRY_SEC > time(NULL))
return;
@@ -1822,35 +2028,33 @@ static void nl_retry_export(void)
auth_reload();
- msg = cache_nl_new_msg(nfsd_nl_family,
- NFSD_CMD_SVC_EXPORT_SET_REQS, 0);
- if (!msg) {
+ if (nl_export_one(d->client, d->path) == EXPORT_RETRY)
delayed_export_enqueue(d);
- return;
- }
-
- switch (nl_add_export_req(&msg, d->client, d->path)) {
- case EXPORT_ANSWERED:
- if (nl_msg_has_reqs(msg))
- cache_nl_set_reqs(nfsd_nl_cmd_sock, msg);
+ else
delayed_export_free(d);
- break;
- case EXPORT_RETRY:
- delayed_export_enqueue(d);
- break;
- case EXPORT_NOMEM:
- delayed_export_enqueue(d);
- return; /* msg is already gone */
- }
- nlmsg_free(msg);
+}
+
+/*
+ * Answer [@start, @end) one at a time. Building the batch may already have
+ * deferred some of these, so an entry that resolves this time has to come back
+ * off the retry queue, or it gets answered a second time.
+ */
+static void nl_export_singly(struct export_req *reqs, int start, int end)
+{
+ int i;
+
+ for (i = start; i < end; i++)
+ if (nl_export_one(reqs[i].client, reqs[i].path) == EXPORT_RETRY)
+ nl_delay_export(reqs[i].client, reqs[i].path);
+ else
+ delayed_export_remove(reqs[i].client, reqs[i].path);
}
static void cache_nl_process_export(void)
{
struct export_req *reqs = NULL;
int nreqs = 0;
- struct nl_msg *msg;
- int i;
+ int i = 0;
/* Fetch all pending requests from the kernel */
if (cache_nl_get_export_reqs(&reqs, &nreqs)) {
@@ -1863,29 +2067,46 @@ static void cache_nl_process_export(void)
xlog(D_CALL, "cache_nl_process_export: %d pending export requests", nreqs);
- /* Build the SET_REQS response */
- msg = cache_nl_new_msg(nfsd_nl_family,
- NFSD_CMD_SVC_EXPORT_SET_REQS, 0);
- if (!msg)
- goto out_free;
+ while (i < nreqs) {
+ int start = i;
+ struct nl_msg *msg;
- for (i = 0; i < nreqs; i++) {
- switch (nl_add_export_req(&msg, reqs[i].client, reqs[i].path)) {
- case EXPORT_ANSWERED:
- break;
- case EXPORT_RETRY:
- nl_delay_export(reqs[i].client, reqs[i].path);
+ msg = cache_nl_new_msg(nfsd_nl_family,
+ NFSD_CMD_SVC_EXPORT_SET_REQS, 0);
+ if (!msg)
break;
- case EXPORT_NOMEM:
- goto out_free;
+
+ for (; i < nreqs; i++) {
+ enum export_result res;
+
+ res = nl_add_export_req(msg, reqs[i].client,
+ reqs[i].path, NULL);
+ if (res == EXPORT_FULL)
+ break;
+ if (res == EXPORT_RETRY)
+ nl_delay_export(reqs[i].client, reqs[i].path);
}
- }
- if (nl_msg_has_reqs(msg))
- cache_nl_set_reqs(nfsd_nl_cmd_sock, msg);
- nlmsg_free(msg);
+ /*
+ * One bad entry fails the whole message, so resubmit the
+ * batch singly to find out which and answer the rest.
+ */
+ if (nl_msg_has_reqs(msg) &&
+ cache_nl_set_reqs(nfsd_nl_cmd_sock, msg, NULL) < 0) {
+ xlog(D_CALL, "%s: batch refused, answering %d request%s"
+ " singly", __func__, i - start,
+ i - start == 1 ? "" : "s");
+ nl_export_singly(reqs, start, i);
+ }
+ nlmsg_free(msg);
+
+ /* First entry did not fit an empty message: answer it alone */
+ if (i == start) {
+ nl_export_singly(reqs, i, i + 1);
+ i++;
+ }
+ }
-out_free:
for (i = 0; i < nreqs; i++) {
free(reqs[i].client);
free(reqs[i].path);
@@ -2155,7 +2376,7 @@ static void cache_nl_process_expkey(void)
do_add_expkey:
if (nfsd_nl_add_expkey(msg, dom, fsidtype, fsid,
fsidlen, found_path) < 0) {
- cache_nl_set_reqs(nfsd_nl_cmd_sock, msg);
+ cache_nl_set_reqs(nfsd_nl_cmd_sock, msg, NULL);
nlmsg_free(msg);
msg = cache_nl_new_msg(nfsd_nl_family,
NFSD_CMD_EXPKEY_SET_REQS, 0);
@@ -2176,7 +2397,7 @@ do_add_expkey:
nfs_freeaddrinfo(ai);
}
- cache_nl_set_reqs(nfsd_nl_cmd_sock, msg);
+ cache_nl_set_reqs(nfsd_nl_cmd_sock, msg, NULL);
nlmsg_free(msg);
out_free:
@@ -2466,7 +2687,7 @@ static void cache_nl_process_ip_map(void)
}
if (nl_add_ip_map(msg, class, ipaddr, domain) < 0) {
- cache_nl_set_reqs(sunrpc_nl_cmd_sock, msg);
+ cache_nl_flush_reqs(sunrpc_nl_cmd_sock, msg, "ip_map");
nlmsg_free(msg);
msg = cache_nl_new_msg(sunrpc_nl_family,
SUNRPC_CMD_IP_MAP_SET_REQS, 0);
@@ -2496,7 +2717,7 @@ static void cache_nl_process_ip_map(void)
nfs_freeaddrinfo(tmp);
}
- cache_nl_set_reqs(sunrpc_nl_cmd_sock, msg);
+ cache_nl_flush_reqs(sunrpc_nl_cmd_sock, msg, "ip_map");
nlmsg_free(msg);
out_free:
@@ -2714,7 +2935,7 @@ static void cache_nl_process_unix_gid(void)
if (ret < 0) {
/* Flush current message and retry with a fresh one */
- cache_nl_set_reqs(sunrpc_nl_cmd_sock, msg);
+ cache_nl_flush_reqs(sunrpc_nl_cmd_sock, msg, "unix_gid");
nlmsg_free(msg);
msg = cache_nl_new_msg(sunrpc_nl_family,
SUNRPC_CMD_UNIX_GID_SET_REQS, 0);
@@ -2731,7 +2952,7 @@ static void cache_nl_process_unix_gid(void)
}
}
- cache_nl_set_reqs(sunrpc_nl_cmd_sock, msg);
+ cache_nl_flush_reqs(sunrpc_nl_cmd_sock, msg, "unix_gid");
nlmsg_free(msg);
out_free:
--
2.55.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH nfs-utils v3 04/11] mountd: don't leak the parent export's fsid onto crossmnt submounts
2026-09-14 13:14 [PATCH nfs-utils v3 00/11] mountd: bugfixes for up/downcall interfaces Jeff Layton
` (2 preceding siblings ...)
2026-09-14 13:14 ` [PATCH nfs-utils v3 03/11] mountd: answer requests the kernel rejects on " Jeff Layton
@ 2026-09-14 13:14 ` Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 05/11] mountd: retry unresolvable fsid lookups on the netlink downcall Jeff Layton
` (8 subsequent siblings)
12 siblings, 0 replies; 14+ messages in thread
From: Jeff Layton @ 2026-09-14 13:14 UTC (permalink / raw)
To: Steve Dickson, Mantas Mikulėnas; +Cc: Chuck Lever, linux-nfs, Jeff Layton
nfsd_nl_add_export() sent e_flags and e_fsid straight from the parent
exportent, unlike dump_to_cache() which strips NFSEXP_FSID (and picks a
path-derived uuid) when the upcall path is a submount below the exported
one.
v4root forces "/" to fsid=0, so exporting "/" with crossmnt handed every
submount fsid=0 as well. Both exports then encode the same fsid, the
submount's filehandles decode back to "/", and the client sees
NFS: server X error: fileid changed
fsid 0:150: expected fileid 0x2, got 0x100
followed by ESTALE. Re-export via fsidnum was missing for the same
reason, so a re-exported submount got the parent's fsid too.
Use export_attrs_build() for the flags, fsid and uuid, mask the per-
flavor secinfo flags to match (the kernel rejects the entry otherwise),
and fall back to a negative entry when the export cannot be resolved.
While here, honour the export's own e_ttl on positive entries.
Two consequences worth noting. A crossmnt submount under an fsid=
parent now gets its own filehandles, so clients holding the old
(aliased) ones see ESTALE once - that aliasing was the bug. And a
crossmnt submount on a filesystem with no blkid uuid and no statfs fsid
(tmpfs, say) now has neither fsid= nor uuid, so check_export() refuses
it; the preceding patch turns that into a negative entry instead of a
failed batch, so this one depends on it. Nobody asked to export that
submount, so it is only worth D_GENERAL - the warning stays for a path
exported in its own right.
Reported-by: Mantas Mikulėnas <grawity@gmail.com>
Assisted-by: LLM
Signed-off-by: Jeff Layton <jlayton@kernel.org>
---
support/export/cache.c | 58 ++++++++++++++++++++++++++++++++------------------
1 file changed, 37 insertions(+), 21 deletions(-)
diff --git a/support/export/cache.c b/support/export/cache.c
index 5bdc00c5847f..b11ef5ec5da5 100644
--- a/support/export/cache.c
+++ b/support/export/cache.c
@@ -1500,7 +1500,8 @@ static int nfsd_nl_add_fsloc(struct nl_msg *msg, struct exportent *ep)
return 0;
}
-static int nfsd_nl_add_secinfo(struct nl_msg *msg, struct exportent *ep)
+static int nfsd_nl_add_secinfo(struct nl_msg *msg, struct exportent *ep,
+ struct export_attrs *ea)
{
struct sec_entry *p;
@@ -1520,7 +1521,7 @@ static int nfsd_nl_add_secinfo(struct nl_msg *msg, struct exportent *ep)
if (nla_put_u32(msg, NFSD_A_AUTH_FLAVOR_PSEUDOFLAVOR,
p->flav->fnum) < 0 ||
nla_put_u32(msg, NFSD_A_AUTH_FLAVOR_FLAGS,
- p->flags) < 0)
+ (p->flags | ea->sec_extra) & ea->sec_mask) < 0)
return -1;
nla_nest_end(msg, sec);
}
@@ -1545,15 +1546,26 @@ static int nfsd_nl_add_xprtsec(struct nl_msg *msg, struct exportent *ep)
return 0;
}
+/*
+ * Add one svc_export response. @ea must be the attributes computed by
+ * export_attrs_build() for (@path, @exp), and is ignored when @exp is NULL.
+ * The only failure mode is a full message, so the caller can retry with a
+ * fresh one.
+ */
static int nfsd_nl_add_export(struct nl_msg *msg, char *domain, char *path,
- struct exportent *exp, int ttl)
+ struct exportent *exp, struct export_attrs *ea,
+ int ttl)
{
struct nlattr *nest;
time_t now = time(0);
- char u[16];
+ uint64_t expiry;
+ /* A positive entry carries the export's own ttl */
+ if (exp)
+ ttl = (int)exp->e_ttl;
if (ttl <= 1)
ttl = default_ttl;
+ expiry = now + ttl;
nest = nla_nest_start(msg, NFSD_A_SVC_EXPORT_REQS_REQUESTS);
if (!nest)
@@ -1561,7 +1573,7 @@ static int nfsd_nl_add_export(struct nl_msg *msg, char *domain, char *path,
if (nla_put_string(msg, NFSD_A_SVC_EXPORT_CLIENT, domain) < 0 ||
nla_put_string(msg, NFSD_A_SVC_EXPORT_PATH, path) < 0 ||
- nla_put_u64(msg, NFSD_A_SVC_EXPORT_EXPIRY, now + ttl) < 0)
+ nla_put_u64(msg, NFSD_A_SVC_EXPORT_EXPIRY, expiry) < 0)
goto nla_failure;
if (!exp) {
@@ -1573,26 +1585,19 @@ static int nfsd_nl_add_export(struct nl_msg *msg, char *domain, char *path,
nla_put_u32(msg, NFSD_A_SVC_EXPORT_ANON_GID,
exp->e_anongid) < 0 ||
nla_put_u32(msg, NFSD_A_SVC_EXPORT_FLAGS,
- exp->e_flags) < 0 ||
+ ea->flags) < 0 ||
nla_put_s32(msg, NFSD_A_SVC_EXPORT_FSID,
- exp->e_fsid) < 0)
+ ea->fsidnum) < 0)
goto nla_failure;
if (nfsd_nl_add_fsloc(msg, exp))
goto nla_failure;
- if (exp->e_uuid) {
- get_uuid(exp->e_uuid, 16, u);
- if (nla_put(msg, NFSD_A_SVC_EXPORT_UUID,
- 16, u) < 0)
- goto nla_failure;
- } else if (uuid_by_path(path, 0, 16, u)) {
- if (nla_put(msg, NFSD_A_SVC_EXPORT_UUID,
- 16, u) < 0)
- goto nla_failure;
- }
+ if (ea->have_uuid &&
+ nla_put(msg, NFSD_A_SVC_EXPORT_UUID, 16, ea->uuid) < 0)
+ goto nla_failure;
- if (nfsd_nl_add_secinfo(msg, exp))
+ if (nfsd_nl_add_secinfo(msg, exp, ea))
goto nla_failure;
if (nfsd_nl_add_xprtsec(msg, exp))
@@ -1751,7 +1756,9 @@ static enum export_result nl_add_export_req(struct nl_msg *msg, char *dom,
nfs_export *found = NULL;
struct exportent *epp = NULL;
struct exportent *junction = NULL;
+ struct export_attrs ea = {};
enum export_result res;
+ bool explicit_export;
int ttl = 0;
if (is_ipaddr_client(dom)) {
@@ -1769,8 +1776,9 @@ static enum export_result nl_add_export_req(struct nl_msg *msg, char *dom,
}
}
+ explicit_export = export_is_explicit(found, path);
if (explicitp)
- *explicitp = export_is_explicit(found, path);
+ *explicitp = explicit_export;
if (found) {
char *mp = found->m_export.e_mountpoint;
@@ -1800,7 +1808,15 @@ static enum export_result nl_add_export_req(struct nl_msg *msg, char *dom,
}
}
- if (nfsd_nl_add_export(msg, dom, path, epp, ttl) < 0)
+ if (epp && export_attrs_build(&ea, path, epp) < 0) {
+ xlog(explicit_export ? L_WARNING : D_GENERAL,
+ "Cannot export %s, possibly unsupported"
+ " filesystem or fsid= required", path);
+ epp = NULL;
+ ttl = 0;
+ }
+
+ if (nfsd_nl_add_export(msg, dom, path, epp, &ea, ttl) < 0)
res = EXPORT_FULL;
else
res = epp ? EXPORT_ANSWERED : EXPORT_DENIED;
@@ -1827,7 +1843,7 @@ static enum export_result nl_export_negative(char *dom, char *path)
if (!msg)
return EXPORT_RETRY;
- if (nfsd_nl_add_export(msg, dom, path, NULL, 0) < 0) {
+ if (nfsd_nl_add_export(msg, dom, path, NULL, NULL, 0) < 0) {
nlmsg_free(msg);
return EXPORT_RETRY;
}
--
2.55.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH nfs-utils v3 05/11] mountd: retry unresolvable fsid lookups on the netlink downcall
2026-09-14 13:14 [PATCH nfs-utils v3 00/11] mountd: bugfixes for up/downcall interfaces Jeff Layton
` (3 preceding siblings ...)
2026-09-14 13:14 ` [PATCH nfs-utils v3 04/11] mountd: don't leak the parent export's fsid onto crossmnt submounts Jeff Layton
@ 2026-09-14 13:14 ` Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 06/11] mountd: bound the junction path before copying it into e_path Jeff Layton
` (7 subsequent siblings)
12 siblings, 0 replies; 14+ messages in thread
From: Jeff Layton @ 2026-09-14 13:14 UTC (permalink / raw)
To: Steve Dickson, Mantas Mikulėnas; +Cc: Chuck Lever, linux-nfs, Jeff Layton
nfsd_fh() defers a request instead of an answer in two cases:
- the fsid names a device that is not present
- the export has a "mountpoint" that is not mounted
The filesystem can appear soon, and an answer now gives a spurious
ESTALE. cache_nl_process_expkey() answered negative in both cases, so
a not-yet-mounted autofs submount failed immediately.
Move the fsid lookup out of nfsd_handle_fh() into lookup_fsid(). Call
it from both downcalls. The netlink path then also gets:
- dev_missing accounting for unmatchable and unmounted exports
- reexpdb_uncover_subvolume() for re-exported fsidnums
- the V4ROOT tie-break and duplicate-filehandle warning
- the fsidtype range check
Deferred netlink requests use their own list. cache_process() retries
them on the same RETRY_SEC cadence as the pipefs ones. The kernel keeps
the upcall pending until then.
expkey has the same batching hole that svc_export had.
nfsd_nl_expkey_set_reqs_doit() stops at the first entry that it refuses.
It refuses an entry when kern_path() on the path fails. It also refuses
an entry when auth_domain_find() on the client fails. The kernel never
processed the entries behind the refused one. Those clients hung.
Resubmit a rejected batch one entry at a time, as the svc_export path
now does, tracing the fallback at D_CALL.
A refused entry gets no negative fallback. The kernel builds a negative
expkey entry from auth_domain_find() alone, so a negative answer would
replace a path that kern_path() refused. But nfsd_nl_add_expkey() sets
an expiry of 0x7fffffff. A path that disappeared for a moment would
become a denial until the next "exportfs -f". Log the entry and leave
it pending for the next upcall, which is what pipefs does when the
channel write fails.
A libnl or socket failure is a different case. The kernel never saw the
message, so the request is still unanswered. cache_nl_set_reqs()
reports the kernel's errno separately, and nl_expkey_one() returns such
a request to the retry list.
Reported-by: Mantas Mikulėnas <grawity@gmail.com>
Assisted-by: LLM
Signed-off-by: Jeff Layton <jlayton@kernel.org>
---
support/export/cache.c | 434 +++++++++++++++++++++++++++++++++----------------
1 file changed, 298 insertions(+), 136 deletions(-)
diff --git a/support/export/cache.c b/support/export/cache.c
index b11ef5ec5da5..79dd0abe9a1b 100644
--- a/support/export/cache.c
+++ b/support/export/cache.c
@@ -796,17 +796,19 @@ static time_t retry_delay(time_t *last_attempt, time_t now, time_t delay)
return d < delay ? d : delay;
}
-static int nfsd_handle_fh(int f, char *bp, int blen)
+enum fsid_lookup {
+ FSID_LOOKUP_ANSWER, /* definitive; *pathp NULL means deny access */
+ FSID_LOOKUP_RETRY, /* not resolvable yet, ask again later */
+ FSID_LOOKUP_IGNORE, /* unusable request, no reply possible */
+};
+
+/*
+ * Find the export path that @fsid refers to for client @dom. On
+ * FSID_LOOKUP_ANSWER the caller owns *pathp.
+ */
+static enum fsid_lookup lookup_fsid(char *dom, int fsidtype, int fsidlen,
+ char *fsid, char **pathp)
{
- /* request are:
- * domain fsidtype fsid
- * interpret fsid, find export point and options, and write:
- * domain fsidtype fsid expiry path
- */
- char *dom;
- int fsidtype;
- int fsidlen;
- char fsid[32];
struct parsed_fsid parsed;
struct exportent *found = NULL;
struct addrinfo *ai = NULL;
@@ -814,21 +816,13 @@ static int nfsd_handle_fh(int f, char *bp, int blen)
nfs_export *exp;
int i;
int dev_missing = 0;
- char buf[RPC_CHAN_BUF_SIZE];
int did_uncover = 0;
- int ret = 0;
+ enum fsid_lookup ret = FSID_LOOKUP_IGNORE;
+
+ *pathp = NULL;
- dom = malloc(blen);
- if (dom == NULL)
- return ret;
- if (qword_get(&bp, dom, blen) <= 0)
- goto out;
- if (qword_get_int(&bp, &fsidtype) != 0)
- goto out;
if (fsidtype < 0 || fsidtype > 7)
goto out; /* unknown type */
- if ((fsidlen = qword_get(&bp, fsid, 32)) <= 0)
- goto out;
if (parse_fsid(fsidtype, fsidlen, fsid, &parsed))
goto out;
@@ -923,7 +917,7 @@ static int nfsd_handle_fh(int f, char *bp, int blen)
* quiet rather than returning stale yet
*/
if (dev_missing) {
- ret = 1;
+ ret = FSID_LOOKUP_RETRY;
goto out;
}
} else if (found->e_mountpoint &&
@@ -935,8 +929,55 @@ static int nfsd_handle_fh(int f, char *bp, int blen)
xlog(L_WARNING, "%s not exported as %d not a mountpoint",
found->e_path, found->e_mountpoint);
*/
+ ret = FSID_LOOKUP_RETRY;
+ goto out;
+ }
+
+ ret = FSID_LOOKUP_ANSWER;
+ *pathp = found_path;
+ found_path = NULL;
+out:
+ if (ret != FSID_LOOKUP_RETRY)
+ xlog(D_CALL, "%s: found %p path %s", __func__,
+ found, found ? found->e_path : NULL);
+ free(found_path);
+ nfs_freeaddrinfo(ai);
+ return ret;
+}
+
+static int nfsd_handle_fh(int f, char *bp, int blen)
+{
+ /* request are:
+ * domain fsidtype fsid
+ * interpret fsid, find export point and options, and write:
+ * domain fsidtype fsid expiry path
+ */
+ char *dom;
+ int fsidtype;
+ int fsidlen;
+ char fsid[32];
+ char *found_path = NULL;
+ char buf[RPC_CHAN_BUF_SIZE];
+ int ret = 0;
+
+ dom = malloc(blen);
+ if (dom == NULL)
+ return ret;
+ if (qword_get(&bp, dom, blen) <= 0)
+ goto out;
+ if (qword_get_int(&bp, &fsidtype) != 0)
+ goto out;
+ if ((fsidlen = qword_get(&bp, fsid, 32)) <= 0)
+ goto out;
+
+ switch (lookup_fsid(dom, fsidtype, fsidlen, fsid, &found_path)) {
+ case FSID_LOOKUP_RETRY:
ret = 1;
goto out;
+ case FSID_LOOKUP_IGNORE:
+ goto out;
+ case FSID_LOOKUP_ANSWER:
+ break;
}
bp = buf; blen = sizeof(buf);
@@ -952,21 +993,16 @@ static int nfsd_handle_fh(int f, char *bp, int blen)
* line.
*/
qword_addint(&bp, &blen, 0x7fffffff);
- if (found)
+ if (found_path)
qword_add(&bp, &blen, found_path);
qword_addeol(&bp, &blen);
if (blen <= 0 || cache_write(f, buf, bp - buf) != bp - buf)
xlog(L_ERROR, "nfsd_fh: error writing reply");
- if (!found)
+ if (!found_path)
xlog(D_AUTH, "denied access to %s", *dom == '$' ? dom+1 : dom);
out:
- if (found_path)
- free(found_path);
- nfs_freeaddrinfo(ai);
+ free(found_path);
free(dom);
- if (!ret)
- xlog(D_CALL, "nfsd_fh: found %p path %s",
- found, found ? found->e_path : NULL);
return ret;
}
@@ -2290,12 +2326,203 @@ nla_failure:
return -1;
}
+/*
+ * An fsid can name a filesystem that isn't mounted yet - an autofs
+ * mountpoint, or a re-exported NFS server that is slow to answer. Set the
+ * request aside and try again later rather than declaring it stale, which
+ * is what nfsd_fh() does with the "delayed" list on pipefs.
+ */
+struct delayed_expkey {
+ char *client;
+ char *fsid;
+ int fsidlen;
+ int fsidtype;
+ time_t last_attempt;
+ struct delayed_expkey *next;
+};
+
+static struct delayed_expkey *delayed_expkey;
+
+static void delayed_expkey_enqueue(struct delayed_expkey *d)
+{
+ struct delayed_expkey **dp = &delayed_expkey;
+
+ d->last_attempt = time(NULL);
+ d->next = NULL;
+ while (*dp)
+ dp = &(*dp)->next;
+ *dp = d;
+}
+
+static void delayed_expkey_free(struct delayed_expkey *d)
+{
+ free(d->client);
+ free(d->fsid);
+ free(d);
+}
+
+static void delayed_expkey_flush(void)
+{
+ while (delayed_expkey) {
+ struct delayed_expkey *d = delayed_expkey;
+
+ delayed_expkey = d->next;
+ delayed_expkey_free(d);
+ }
+}
+
+static void nl_delay_expkey(struct expkey_req *req)
+{
+ struct delayed_expkey *d;
+
+ for (d = delayed_expkey; d; d = d->next)
+ if (d->fsidtype == req->fsidtype &&
+ d->fsidlen == req->fsidlen &&
+ !strcmp(d->client, req->client) &&
+ !memcmp(d->fsid, req->fsid, req->fsidlen))
+ return;
+
+ d = calloc(1, sizeof(*d));
+ if (!d)
+ return;
+
+ d->client = strdup(req->client);
+ d->fsid = malloc(req->fsidlen);
+ if (!d->client || !d->fsid) {
+ delayed_expkey_free(d);
+ return;
+ }
+ memcpy(d->fsid, req->fsid, req->fsidlen);
+ d->fsidlen = req->fsidlen;
+ d->fsidtype = req->fsidtype;
+
+ delayed_expkey_enqueue(d);
+}
+
+enum expkey_result {
+ EXPKEY_ANSWERED,
+ EXPKEY_RETRY, /* not resolvable yet, ask again later */
+ EXPKEY_FULL, /* did not fit, flush the message and re-add */
+};
+
+/* Resolve one expkey request and append the answer to @msg */
+static enum expkey_result nl_add_expkey_req(struct nl_msg *msg,
+ struct expkey_req *req)
+{
+ enum expkey_result res = EXPKEY_ANSWERED;
+ char *found_path = NULL;
+ char *dom = req->client;
+
+ switch (lookup_fsid(dom, req->fsidtype, req->fsidlen, req->fsid,
+ &found_path)) {
+ case FSID_LOOKUP_RETRY:
+ return EXPKEY_RETRY;
+ case FSID_LOOKUP_IGNORE: /* answer negative rather than hang */
+ case FSID_LOOKUP_ANSWER:
+ break;
+ }
+
+ if (nfsd_nl_add_expkey(msg, dom, req->fsidtype, req->fsid,
+ req->fsidlen, found_path) < 0)
+ res = EXPKEY_FULL;
+ else if (!found_path)
+ xlog(D_AUTH, "denied access to %s", *dom == '$' ? dom + 1 : dom);
+
+ free(found_path);
+ return res;
+}
+
+/*
+ * Answer one request in a message of its own. The kernel refuses an entry
+ * whose path or whose client's auth_domain has gone away, and fails the
+ * whole message when it does, so a rejected entry must not take the rest of
+ * a batch down with it.
+ */
+static enum expkey_result nl_expkey_one(struct expkey_req *req)
+{
+ enum expkey_result res;
+ struct nl_msg *msg;
+ int kern_err = 0;
+
+ msg = cache_nl_new_msg(nfsd_nl_family, NFSD_CMD_EXPKEY_SET_REQS, 0);
+ if (!msg)
+ return EXPKEY_RETRY;
+
+ res = nl_add_expkey_req(msg, req);
+ switch (res) {
+ case EXPKEY_ANSWERED:
+ if (cache_nl_set_reqs(nfsd_nl_cmd_sock, msg, &kern_err) < 0) {
+ /*
+ * The kernel never answered - a broken socket, or a
+ * message we could not send - so the request is still
+ * unanswered. Ask again later rather than drop it.
+ */
+ if (!kern_err) {
+ res = EXPKEY_RETRY;
+ break;
+ }
+ /*
+ * Nothing to fall back on: a negative entry needs the
+ * same auth_domain the kernel may have just failed to
+ * find, and we cannot tell that apart from a path it
+ * refused. Leave the request pending for the next
+ * upcall, as pipefs does when the channel write fails.
+ */
+ xlog(L_WARNING, "%s: kernel rejected the fsid answer"
+ " for %s: %s", __func__, req->client,
+ strerror(-kern_err));
+ }
+ break;
+ case EXPKEY_FULL:
+ xlog(L_WARNING, "%s: skipping oversized entry", __func__);
+ break;
+ case EXPKEY_RETRY:
+ break;
+ }
+ nlmsg_free(msg);
+ return res;
+}
+
+static void nl_expkey_singly(struct expkey_req *reqs, int start, int end)
+{
+ int i;
+
+ for (i = start; i < end; i++)
+ if (nl_expkey_one(&reqs[i]) == EXPKEY_RETRY)
+ nl_delay_expkey(&reqs[i]);
+}
+
+/*
+ * Retry the oldest deferred lookup if it is due. Entries are queued in
+ * time order, so only the head can be ready.
+ */
+static void nl_retry_expkey(void)
+{
+ struct delayed_expkey *d = delayed_expkey;
+ struct expkey_req req;
+
+ if (!d || d->last_attempt + RETRY_SEC > time(NULL))
+ return;
+
+ delayed_expkey = d->next;
+ d->next = NULL;
+
+ req.client = d->client;
+ req.fsidtype = d->fsidtype;
+ req.fsid = d->fsid;
+ req.fsidlen = d->fsidlen;
+
+ if (nl_expkey_one(&req) == EXPKEY_RETRY)
+ delayed_expkey_enqueue(d);
+ else
+ delayed_expkey_free(d);
+}
+
static void cache_nl_process_expkey(void)
{
struct expkey_req *reqs = NULL;
int nreqs = 0;
- struct nl_msg *msg;
- int i;
+ int i = 0;
if (cache_nl_get_expkey_reqs(&reqs, &nreqs)) {
xlog(L_WARNING, "cache_nl_process_expkey: failed to get expkey requests");
@@ -2307,116 +2534,45 @@ static void cache_nl_process_expkey(void)
xlog(D_CALL, "cache_nl_process_expkey: %d pending expkey requests", nreqs);
- msg = cache_nl_new_msg(nfsd_nl_family, NFSD_CMD_EXPKEY_SET_REQS, 0);
- if (!msg)
- goto out_free;
-
- for (i = 0; i < nreqs; i++) {
- char *dom = reqs[i].client;
- int fsidtype = reqs[i].fsidtype;
- char *fsid = reqs[i].fsid;
- int fsidlen = reqs[i].fsidlen;
- struct parsed_fsid parsed;
- struct addrinfo *ai = NULL;
- struct exportent *found = NULL;
- char *found_path = NULL;
- nfs_export *exp;
- int j;
-
- if (parse_fsid(fsidtype, fsidlen, fsid, &parsed))
- goto do_add_expkey;
-
- if (is_ipaddr_client(dom)) {
- ai = lookup_client_addr(dom);
- if (!ai)
- goto do_add_expkey;
- }
-
- for (j = 0; j < MCL_MAXTYPES; j++) {
- nfs_export *prev = NULL;
- nfs_export *next_exp;
- void *mnt = NULL;
-
- for (exp = exportlist[j].p_head; exp;
- exp = next_exp) {
- char *path;
-
- if (exp->m_export.e_flags &
- NFSEXP_CROSSMOUNT) {
- if (prev == exp) {
- path = next_mnt(&mnt,
- exp->m_export.e_path);
- if (!path) {
- next_exp = exp->m_next;
- prev = NULL;
- continue;
- }
- next_exp = exp;
- } else {
- prev = exp;
- mnt = NULL;
- path = exp->m_export.e_path;
- next_exp = exp;
- }
- } else {
- path = exp->m_export.e_path;
- next_exp = exp->m_next;
- }
+ while (i < nreqs) {
+ int start = i;
+ struct nl_msg *msg;
- if (!is_ipaddr_client(dom) &&
- !namelist_client_matches(exp, dom))
- continue;
+ msg = cache_nl_new_msg(nfsd_nl_family,
+ NFSD_CMD_EXPKEY_SET_REQS, 0);
+ if (!msg)
+ break;
- switch (match_fsid(&parsed, exp, path)) {
- case 0:
- continue;
- case -1:
- continue;
- }
+ for (; i < nreqs; i++) {
+ enum expkey_result res;
- if (is_ipaddr_client(dom) &&
- !ipaddr_client_matches(exp, ai))
- continue;
+ res = nl_add_expkey_req(msg, &reqs[i]);
+ if (res == EXPKEY_FULL)
+ break;
+ if (res == EXPKEY_RETRY)
+ nl_delay_expkey(&reqs[i]);
+ }
- if (!found ||
- subexport(&exp->m_export, found)) {
- found = &exp->m_export;
- free(found_path);
- found_path = strdup(path);
- if (!found_path)
- goto do_add_expkey;
- }
- }
+ /*
+ * One bad entry fails the whole message, so resubmit the
+ * batch singly to find out which and answer the rest.
+ */
+ if (nl_msg_has_reqs(msg) &&
+ cache_nl_set_reqs(nfsd_nl_cmd_sock, msg, NULL) < 0) {
+ xlog(D_CALL, "%s: batch refused, answering %d request%s"
+ " singly", __func__, i - start,
+ i - start == 1 ? "" : "s");
+ nl_expkey_singly(reqs, start, i);
}
+ nlmsg_free(msg);
-do_add_expkey:
- if (nfsd_nl_add_expkey(msg, dom, fsidtype, fsid,
- fsidlen, found_path) < 0) {
- cache_nl_set_reqs(nfsd_nl_cmd_sock, msg, NULL);
- nlmsg_free(msg);
- msg = cache_nl_new_msg(nfsd_nl_family,
- NFSD_CMD_EXPKEY_SET_REQS, 0);
- if (!msg) {
- free(found_path);
- nfs_freeaddrinfo(ai);
- goto out_free;
- }
- if (nfsd_nl_add_expkey(msg, dom, fsidtype, fsid,
- fsidlen, found_path) < 0)
- xlog(L_WARNING, "%s: skipping oversized "
- "entry", __func__);
+ /* First entry did not fit an empty message: answer it alone */
+ if (i == start) {
+ nl_expkey_singly(reqs, i, i + 1);
+ i++;
}
- if (!found)
- xlog(D_AUTH, "denied access to %s",
- *dom == '$' ? dom + 1 : dom);
- free(found_path);
- nfs_freeaddrinfo(ai);
}
- cache_nl_set_reqs(nfsd_nl_cmd_sock, msg, NULL);
- nlmsg_free(msg);
-
-out_free:
for (i = 0; i < nreqs; i++) {
free(reqs[i].client);
free(reqs[i].fsid);
@@ -3616,12 +3772,15 @@ int cache_process(fd_set *readfds)
cache_set_fds(readfds);
v4clients_set_fds(readfds);
- if (delayed || delayed_export) {
+ if (delayed || delayed_expkey || delayed_export) {
time_t now = time(NULL);
time_t delay = RETRY_SEC;
if (delayed)
delay = retry_delay(&delayed->last_attempt, now, delay);
+ if (delayed_expkey)
+ delay = retry_delay(&delayed_expkey->last_attempt, now,
+ delay);
if (delayed_export)
delay = retry_delay(&delayed_export->last_attempt, now,
delay);
@@ -3642,6 +3801,7 @@ int cache_process(fd_set *readfds)
}
}
+ nl_retry_expkey();
nl_retry_export();
switch (selret) {
@@ -3852,12 +4012,14 @@ cache_fork_workers(char *prog, int num_threads)
/*
* cache_open() drains the netlink downcalls before we
* get here, so anything it deferred is now on the retry
- * queue of every worker. Let the first worker own
+ * queues of every worker. Let the first worker own
* those, or they get answered once per worker over the
* shared command socket.
*/
- if (i > 0)
+ if (i > 0) {
delayed_export_flush();
+ delayed_expkey_flush();
+ }
/* Re-enable the default action on SIGTERM et al
* so that workers die naturally when sent them.
--
2.55.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH nfs-utils v3 06/11] mountd: bound the junction path before copying it into e_path
2026-09-14 13:14 [PATCH nfs-utils v3 00/11] mountd: bugfixes for up/downcall interfaces Jeff Layton
` (4 preceding siblings ...)
2026-09-14 13:14 ` [PATCH nfs-utils v3 05/11] mountd: retry unresolvable fsid lookups on the netlink downcall Jeff Layton
@ 2026-09-14 13:14 ` Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 07/11] mountd: give each worker its own netlink command socket Jeff Layton
` (6 subsequent siblings)
12 siblings, 0 replies; 14+ messages in thread
From: Jeff Layton @ 2026-09-14 13:14 UTC (permalink / raw)
To: Steve Dickson, Mantas Mikulėnas; +Cc: Chuck Lever, linux-nfs, Jeff Layton
create_junction_exportent() strcpy()s the junction pathname into
exportent.e_path, a fixed char[NFS_MAXPATHLEN+1]. Neither downcall
bounds the path first:
- nfsd_export() sizes it from the 32KiB pipefs channel buffer
- the netlink path strdup()s it out of NFSD_A_SVC_EXPORT_PATH
A junction more than NFS_MAXPATHLEN bytes deep therefore corrupts the
heap. Reaching it needs a trusted.junction.nfs xattr, so only server
root can set one up, but the copy should not depend on that.
Reject the path instead, as mkexportent() already does. The callers
handle a NULL exportent: both downcalls answer negative.
Signed-off-by: Jeff Layton <jlayton@kernel.org>
Assisted-by: LLM
---
support/export/cache.c | 6 ++++++
1 file changed, 6 insertions(+)
diff --git a/support/export/cache.c b/support/export/cache.c
index 79dd0abe9a1b..6a5cdd9d9670 100644
--- a/support/export/cache.c
+++ b/support/export/cache.c
@@ -3434,6 +3434,12 @@ static struct exportent *create_junction_exportent(struct exportent *parent,
{
static struct exportent *eep;
+ if (strlen(junction) >= sizeof(eep->e_path)) {
+ xlog(L_ERROR, "%s: junction path %s too long", __func__,
+ junction);
+ return NULL;
+ }
+
eep = (struct exportent *)malloc(sizeof(*eep));
if (eep == NULL)
goto out_nomem;
--
2.55.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH nfs-utils v3 07/11] mountd: give each worker its own netlink command socket
2026-09-14 13:14 [PATCH nfs-utils v3 00/11] mountd: bugfixes for up/downcall interfaces Jeff Layton
` (5 preceding siblings ...)
2026-09-14 13:14 ` [PATCH nfs-utils v3 06/11] mountd: bound the junction path before copying it into e_path Jeff Layton
@ 2026-09-14 13:14 ` Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 08/11] mountd: retry export attributes that fail to resolve for a passing reason Jeff Layton
` (5 subsequent siblings)
12 siblings, 0 replies; 14+ messages in thread
From: Jeff Layton @ 2026-09-14 13:14 UTC (permalink / raw)
To: Steve Dickson, Mantas Mikulėnas; +Cc: Chuck Lever, linux-nfs, Jeff Layton
cache_open() opens the nfsd and sunrpc command sockets before
cache_fork_workers() forks, so every worker shares one fd, and a
copy-on-write struct nl_sock carrying the same s_seq_next/s_seq_expect.
Only the notify sockets get nl_socket_disable_seq_check(); the command
sockets keep libnl's default sequence check, which each worker will
happily pass on the other's reply.
Two workers in a command/reply round trip can therefore take each other's
ack or NLMSG_ERROR. cache_nl_set_reqs() then reports the wrong outcome
and the caller answers the wrong client and path - marking it exported,
retryable, or negatively cached.
Reopen both command sockets in the worker child. genl_connect() binds a
fresh port id, so each worker gets its own reply stream. The notify
sockets stay shared: they are receive-only, have the sequence check
disabled, and are nonblocking, so whichever worker gets there first takes
the notification.
Signed-off-by: Jeff Layton <jlayton@kernel.org>
Assisted-by: LLM
---
support/export/cache.c | 37 +++++++++++++++++++++++++++++++++++--
1 file changed, 35 insertions(+), 2 deletions(-)
diff --git a/support/export/cache.c b/support/export/cache.c
index 6a5cdd9d9670..cbe3df83b0eb 100644
--- a/support/export/cache.c
+++ b/support/export/cache.c
@@ -3994,6 +3994,29 @@ cache_wait_for_workers(char *prog)
}
}
+/*
+ * Replace a command socket inherited across fork(). Returns @old if a new one
+ * cannot be had; sharing it is worse than having one, but not by as much as
+ * having none.
+ */
+static struct nl_sock *nl_cmd_sock_reopen(struct nl_sock *old)
+{
+ struct nl_sock *sock;
+
+ if (!old)
+ return NULL;
+
+ sock = nl_sock_setup();
+ if (!sock) {
+ xlog(L_WARNING, "%s: cannot reopen netlink command socket,"
+ " sharing the inherited one", __func__);
+ return old;
+ }
+
+ nl_socket_free(old);
+ return sock;
+}
+
/* Fork num_threads worker children and wait for them */
int
cache_fork_workers(char *prog, int num_threads)
@@ -4015,12 +4038,22 @@ cache_fork_workers(char *prog, int num_threads)
if (pid == 0) {
/* worker child */
+ /*
+ * A command socket carries a reply back to the process
+ * that sent the request, so it cannot be shared. Every
+ * worker inherits the same fd and the same copied
+ * sequence counters, and two of them mid-round-trip
+ * will take each other's ack.
+ */
+ nfsd_nl_cmd_sock = nl_cmd_sock_reopen(nfsd_nl_cmd_sock);
+ sunrpc_nl_cmd_sock =
+ nl_cmd_sock_reopen(sunrpc_nl_cmd_sock);
+
/*
* cache_open() drains the netlink downcalls before we
* get here, so anything it deferred is now on the retry
* queues of every worker. Let the first worker own
- * those, or they get answered once per worker over the
- * shared command socket.
+ * those, or they get answered once per worker.
*/
if (i > 0) {
delayed_export_flush();
--
2.55.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH nfs-utils v3 08/11] mountd: retry export attributes that fail to resolve for a passing reason
2026-09-14 13:14 [PATCH nfs-utils v3 00/11] mountd: bugfixes for up/downcall interfaces Jeff Layton
` (6 preceding siblings ...)
2026-09-14 13:14 ` [PATCH nfs-utils v3 07/11] mountd: give each worker its own netlink command socket Jeff Layton
@ 2026-09-14 13:14 ` Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 09/11] mountd: drop a deferred fsid lookup once it has been answered Jeff Layton
` (4 subsequent siblings)
12 siblings, 0 replies; 14+ messages in thread
From: Jeff Layton @ 2026-09-14 13:14 UTC (permalink / raw)
To: Steve Dickson, Mantas Mikulėnas; +Cc: Chuck Lever, linux-nfs, Jeff Layton
export_attrs_build() answers -1 both for a filesystem that can never be
exported and for one it could not measure this time:
- nfsd_path_statfs() on a re-exported NFS mount can give ETIMEDOUT with
"softerr", or EIO with plain "soft"
- fsidnum_get_by_path() returns the same false whether fsidd is
restarting, answered with nonsense, or has no fsid for the path
Both then set errno to EINVAL, so both downcalls deny the path for
default_ttl. An fsidd restart takes a working re-export offline until the
negative entry expires.
Report EAGAIN for the cases we cannot decide and keep EINVAL for the rest.
The netlink path turns EAGAIN into EXPORT_RETRY, and nfsd_export() leaves
the request for the next upcall, which is what both already do when
is_mountpoint() fails oddly.
statfs errors are sorted with path_lookup_error(), the same test
is_mountpoint() callers use, plus two that are decidable even though it
does not name them:
- ENOSYS: a filesystem with no statfs will never answer
- 0: a chrooted worker thread ran the stat and its errno never reached
us, which is how the is_mountpoint() callers already read it
Signed-off-by: Jeff Layton <jlayton@kernel.org>
Assisted-by: LLM
---
support/export/cache.c | 41 ++++++++++++++++++++++++++++++++++++++---
1 file changed, 38 insertions(+), 3 deletions(-)
diff --git a/support/export/cache.c b/support/export/cache.c
index cbe3df83b0eb..290ce09bb66e 100644
--- a/support/export/cache.c
+++ b/support/export/cache.c
@@ -1145,7 +1145,9 @@ struct export_attrs {
* it must not inherit the parent's fsid= or uuid= - if it does, both end up
* claiming the same filehandles and the client sees ESTALE.
*
- * Returns 0, or -1 with errno set if @path cannot be exported at all.
+ * Returns 0, or -1 with errno set: EAGAIN if we could not work the attributes
+ * out this time and the caller should ask again, anything else if @path is
+ * genuinely not exportable.
*/
static int export_attrs_build(struct export_attrs *ea, char *path,
struct exportent *exp)
@@ -1161,8 +1163,24 @@ static int export_attrs_build(struct export_attrs *ea, char *path,
struct statfs st;
if (nfsd_path_statfs(path, &st)) {
+ int err = errno;
+
xlog(L_WARNING, "unable to statfs %s", path);
- errno = EINVAL;
+ /*
+ * A strange error - the ETIMEDOUT a "softerr" NFS
+ * re-export gives, or the EIO a plain "soft" one gives
+ * - leaves us unable to say whether the path is
+ * exportable, so ask again later. Two errors are not
+ * of that kind: ENOSYS, because a filesystem with no
+ * statfs will never answer, and 0, which means a
+ * chrooted worker thread ran the statfs and we never
+ * saw its errno - is_mountpoint() callers read that as
+ * definitive too.
+ */
+ if (err != 0 && err != ENOSYS && !path_lookup_error(err))
+ errno = EAGAIN;
+ else
+ errno = EINVAL;
return -1;
}
@@ -1179,7 +1197,13 @@ static int export_attrs_build(struct export_attrs *ea, char *path,
if (exp->e_reexport != REEXP_NONE &&
reexpdb_fsidnum_by_path(path, &search_fsidnum,
exp->e_reexport == REEXP_AUTO_FSIDNUM) == 0) {
- errno = EINVAL;
+ /*
+ * fsidnum_get_by_path() answers the same way whether
+ * fsidd is unreachable, gave us nonsense, or has no
+ * fsid for this path. A restarting fsidd must not
+ * deny a working re-export, so ask again later.
+ */
+ errno = EAGAIN;
return -1;
}
ea->fsidnum = search_fsidnum;
@@ -1845,6 +1869,10 @@ static enum export_result nl_add_export_req(struct nl_msg *msg, char *dom,
}
if (epp && export_attrs_build(&ea, path, epp) < 0) {
+ if (errno == EAGAIN) {
+ res = EXPORT_RETRY;
+ goto out;
+ }
xlog(explicit_export ? L_WARNING : D_GENERAL,
"Cannot export %s, possibly unsupported"
" filesystem or fsid= required", path);
@@ -3637,6 +3665,13 @@ static void nfsd_export(int f)
NULL, 60);
} else if (dump_to_cache(f, buf, sizeof(buf), dom, path,
&found->m_export, 0) < 0) {
+ /*
+ * We could not work out the attributes this time.
+ * Leave the request for the next upcall rather than
+ * denying an export that is probably fine.
+ */
+ if (errno == EAGAIN)
+ goto out;
xlog(L_WARNING,
"Cannot export %s, possibly unsupported filesystem"
" or fsid= required", path);
--
2.55.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH nfs-utils v3 09/11] mountd: drop a deferred fsid lookup once it has been answered
2026-09-14 13:14 [PATCH nfs-utils v3 00/11] mountd: bugfixes for up/downcall interfaces Jeff Layton
` (7 preceding siblings ...)
2026-09-14 13:14 ` [PATCH nfs-utils v3 08/11] mountd: retry export attributes that fail to resolve for a passing reason Jeff Layton
@ 2026-09-14 13:14 ` Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 10/11] mountd: bound the retry queues Jeff Layton
` (3 subsequent siblings)
12 siblings, 0 replies; 14+ messages in thread
From: Jeff Layton @ 2026-09-14 13:14 UTC (permalink / raw)
To: Steve Dickson, Mantas Mikulėnas; +Cc: Chuck Lever, linux-nfs, Jeff Layton
Building an expkey batch can defer an entry, and a later failure of that
batch sends the same entries through nl_expkey_singly(). An entry that
resolves the second time is answered, but its record stays on
delayed_expkey, so nl_retry_expkey() answers it again within RETRY_SEC.
Take it off the queue instead, as nl_export_singly() already does for
delayed_export. The tuple test nl_delay_expkey() dedups on moves into
delayed_expkey_matches() so both users agree on it.
Signed-off-by: Jeff Layton <jlayton@kernel.org>
Assisted-by: LLM
---
support/export/cache.c | 39 +++++++++++++++++++++++++++++++++++----
1 file changed, 35 insertions(+), 4 deletions(-)
diff --git a/support/export/cache.c b/support/export/cache.c
index 290ce09bb66e..6c887cd50327 100644
--- a/support/export/cache.c
+++ b/support/export/cache.c
@@ -2399,15 +2399,39 @@ static void delayed_expkey_flush(void)
}
}
+/* Does @d hold the deferred form of @req? */
+static bool delayed_expkey_matches(struct delayed_expkey *d,
+ struct expkey_req *req)
+{
+ return d->fsidtype == req->fsidtype &&
+ d->fsidlen == req->fsidlen &&
+ !strcmp(d->client, req->client) &&
+ !memcmp(d->fsid, req->fsid, req->fsidlen);
+}
+
+/* Forget any deferred form of @req; it has been answered */
+static void delayed_expkey_remove(struct expkey_req *req)
+{
+ struct delayed_expkey **dp = &delayed_expkey;
+
+ while (*dp) {
+ struct delayed_expkey *d = *dp;
+
+ if (delayed_expkey_matches(d, req)) {
+ *dp = d->next;
+ delayed_expkey_free(d);
+ return;
+ }
+ dp = &d->next;
+ }
+}
+
static void nl_delay_expkey(struct expkey_req *req)
{
struct delayed_expkey *d;
for (d = delayed_expkey; d; d = d->next)
- if (d->fsidtype == req->fsidtype &&
- d->fsidlen == req->fsidlen &&
- !strcmp(d->client, req->client) &&
- !memcmp(d->fsid, req->fsid, req->fsidlen))
+ if (delayed_expkey_matches(d, req))
return;
d = calloc(1, sizeof(*d));
@@ -2511,6 +2535,11 @@ static enum expkey_result nl_expkey_one(struct expkey_req *req)
return res;
}
+/*
+ * Answer [@start, @end) one at a time. Building the batch may already have
+ * deferred some of these, so an entry that resolves this time has to come back
+ * off the retry queue, or it gets answered a second time.
+ */
static void nl_expkey_singly(struct expkey_req *reqs, int start, int end)
{
int i;
@@ -2518,6 +2547,8 @@ static void nl_expkey_singly(struct expkey_req *reqs, int start, int end)
for (i = start; i < end; i++)
if (nl_expkey_one(&reqs[i]) == EXPKEY_RETRY)
nl_delay_expkey(&reqs[i]);
+ else
+ delayed_expkey_remove(&reqs[i]);
}
/*
--
2.55.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH nfs-utils v3 10/11] mountd: bound the retry queues
2026-09-14 13:14 [PATCH nfs-utils v3 00/11] mountd: bugfixes for up/downcall interfaces Jeff Layton
` (8 preceding siblings ...)
2026-09-14 13:14 ` [PATCH nfs-utils v3 09/11] mountd: drop a deferred fsid lookup once it has been answered Jeff Layton
@ 2026-09-14 13:14 ` Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 11/11] mountd: don't warn about pipefs submounts nobody asked to export Jeff Layton
` (2 subsequent siblings)
12 siblings, 0 replies; 14+ messages in thread
From: Jeff Layton @ 2026-09-14 13:14 UTC (permalink / raw)
To: Steve Dickson, Mantas Mikulėnas; +Cc: Chuck Lever, linux-nfs, Jeff Layton
A request is deferred when we cannot yet say whether its path or fsid is
exportable. lookup_fsid() defers on "!found && dev_missing", and
dev_missing counts any export whose "mountpoint" is not mounted, whatever
fsid was asked for. The fsid comes out of the filehandle the client sent,
so on a server with one unmounted mountpoint= export - the case this retry
logic exists for - every fsid a client invents gets a queue entry.
None of the three queues has a limit. The pre-existing "delayed" queue
nfsd_fh() feeds is the worst: it strndup()s the whole upcall message and
does not deduplicate at all, so even a repeated fsid allocates again.
delayed_export and delayed_expkey dedup, but still grow without end, and
the dedup walk makes each insertion O(n).
Cap all three at MAX_DELAYED and drop the new request once full. Queuing
is only an optimisation: the kernel repeats the upcall when the client
retries, so a dropped deferral costs latency, not correctness - the same
trade the existing allocation-failure paths already make.
The cap applies only where a record is created. nfsd_retry_fh(),
nl_retry_export() and nl_retry_expkey() re-queue a record they already
dequeued and own, and must not be refused.
nfsd_fh() had no dedup walk, so it gets a counting one; it ends at the tail
the append needs anyway.
Signed-off-by: Jeff Layton <jlayton@kernel.org>
Assisted-by: LLM
---
support/export/cache.c | 57 +++++++++++++++++++++++++++++++++++++++++++++-----
1 file changed, 52 insertions(+), 5 deletions(-)
diff --git a/support/export/cache.c b/support/export/cache.c
index 6c887cd50327..d1fef23a6c0f 100644
--- a/support/export/cache.c
+++ b/support/export/cache.c
@@ -777,6 +777,38 @@ static struct addrinfo *lookup_client_addr(char *dom)
}
#define RETRY_SEC 120
+
+/*
+ * Cap on each retry queue. A request is deferred when we cannot yet say
+ * whether its path or fsid is exportable, and the client picks the fsid out of
+ * the filehandle it sends, so the queues are reachable from the network: any
+ * export with an unmounted "mountpoint" makes lookup_fsid() defer every fsid it
+ * cannot match, invented ones included. Queuing is only an optimisation - the
+ * kernel repeats the upcall when the client retries - so refusing to grow past
+ * this costs latency, not correctness.
+ */
+#define MAX_DELAYED 1024
+
+/*
+ * Has the queue holding @count entries hit the cap? Warns once when it fills,
+ * and arms the warning again only once it has properly drained, so a queue
+ * sitting at the limit does not turn into a stream of log messages.
+ */
+static bool delayed_is_full(unsigned int count, bool *warned, const char *what)
+{
+ if (count < MAX_DELAYED / 2)
+ *warned = false;
+ if (count < MAX_DELAYED)
+ return false;
+
+ if (!*warned) {
+ *warned = true;
+ xlog(L_WARNING, "%s retry queue is full (%u), dropping requests;"
+ " the kernel will ask again", what, MAX_DELAYED);
+ }
+ return true;
+}
+
struct delayed {
char *message;
time_t last_attempt;
@@ -1008,8 +1040,10 @@ out:
static void nfsd_fh(int f)
{
+ static bool queue_full_warned;
struct delayed *d, **dp;
char inbuf[RPC_CHAN_BUF_SIZE];
+ unsigned int count = 0;
int blen;
blen = cache_read(f, inbuf, sizeof(inbuf));
@@ -1029,6 +1063,11 @@ static void nfsd_fh(int f)
* We cannot tell the kernel to retry, so we have to
* retry ourselves.
*/
+ for (dp = &delayed; *dp; dp = &(*dp)->next)
+ count++;
+ if (delayed_is_full(count, &queue_full_warned, "filehandle"))
+ return;
+
d = malloc(sizeof(*d));
if (!d)
@@ -1041,9 +1080,7 @@ static void nfsd_fh(int f)
d->f = f;
d->last_attempt = time(NULL);
d->next = NULL;
- dp = &delayed;
- while (*dp)
- dp = &(*dp)->next;
+ /* the count above left dp at the tail */
*dp = d;
}
@@ -2072,12 +2109,17 @@ static void delayed_export_remove(char *dom, char *path)
static void nl_delay_export(char *dom, char *path)
{
+ static bool queue_full_warned;
+ unsigned int count = 0;
struct delayed_export *d;
- for (d = delayed_export; d; d = d->next)
+ for (d = delayed_export; d; d = d->next, count++)
if (!strcmp(d->client, dom) && !strcmp(d->path, path))
return;
+ if (delayed_is_full(count, &queue_full_warned, "export"))
+ return;
+
d = calloc(1, sizeof(*d));
if (!d)
return;
@@ -2428,12 +2470,17 @@ static void delayed_expkey_remove(struct expkey_req *req)
static void nl_delay_expkey(struct expkey_req *req)
{
+ static bool queue_full_warned;
+ unsigned int count = 0;
struct delayed_expkey *d;
- for (d = delayed_expkey; d; d = d->next)
+ for (d = delayed_expkey; d; d = d->next, count++)
if (delayed_expkey_matches(d, req))
return;
+ if (delayed_is_full(count, &queue_full_warned, "fsid"))
+ return;
+
d = calloc(1, sizeof(*d));
if (!d)
return;
--
2.55.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH nfs-utils v3 11/11] mountd: don't warn about pipefs submounts nobody asked to export
2026-09-14 13:14 [PATCH nfs-utils v3 00/11] mountd: bugfixes for up/downcall interfaces Jeff Layton
` (9 preceding siblings ...)
2026-09-14 13:14 ` [PATCH nfs-utils v3 10/11] mountd: bound the retry queues Jeff Layton
@ 2026-09-14 13:14 ` Jeff Layton
2026-09-15 9:24 ` [PATCH nfs-utils v3 00/11] mountd: bugfixes for up/downcall interfaces Mantas Mikulėnas
2026-09-17 7:21 ` Steve Dickson
12 siblings, 0 replies; 14+ messages in thread
From: Jeff Layton @ 2026-09-14 13:14 UTC (permalink / raw)
To: Steve Dickson, Mantas Mikulėnas; +Cc: Chuck Lever, linux-nfs, Jeff Layton
nfsd_export() warns whenever dump_to_cache() cannot work out the
attributes for a path, including a crossmnt submount the admin never
asked to export. With "/" exported crossmnt that is one warning per
proc, sys or autofs mount per TTL, and nothing the admin can act on.
Use the same test the netlink downcall now uses: warn only when the
request is for a path exported in its own right.
Assisted-by: LLM
Signed-off-by: Jeff Layton <jlayton@kernel.org>
---
support/export/cache.c | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
diff --git a/support/export/cache.c b/support/export/cache.c
index d1fef23a6c0f..c1201f24bfd4 100644
--- a/support/export/cache.c
+++ b/support/export/cache.c
@@ -3750,7 +3750,8 @@ static void nfsd_export(int f)
*/
if (errno == EAGAIN)
goto out;
- xlog(L_WARNING,
+ xlog(export_is_explicit(found, path) ? L_WARNING
+ : D_GENERAL,
"Cannot export %s, possibly unsupported filesystem"
" or fsid= required", path);
dump_to_cache(f, buf, sizeof(buf), dom, path, NULL, 0);
--
2.55.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* Re: [PATCH nfs-utils v3 00/11] mountd: bugfixes for up/downcall interfaces
2026-09-14 13:14 [PATCH nfs-utils v3 00/11] mountd: bugfixes for up/downcall interfaces Jeff Layton
` (10 preceding siblings ...)
2026-09-14 13:14 ` [PATCH nfs-utils v3 11/11] mountd: don't warn about pipefs submounts nobody asked to export Jeff Layton
@ 2026-09-15 9:24 ` Mantas Mikulėnas
2026-09-17 7:21 ` Steve Dickson
12 siblings, 0 replies; 14+ messages in thread
From: Mantas Mikulėnas @ 2026-09-15 9:24 UTC (permalink / raw)
To: Jeff Layton; +Cc: Chuck Lever, linux-nfs, Steve Dickson
On 14/09/2026 16.14, Jeff Layton wrote:
> The main difference in this version is some cleanup to the warning
> messages, so that we don't warn about internal events. Mantas, if you're
> able to test this version too, that would be great.
>
> Mantas reported a problem with his configuration that was using crossmnt
> and autofs to export a number of btrfs filesystems. Some analysis with
> Claude also uncovered a few other bugs. This version also fixes a
> number of problems that Sashiko flagged. Some of them were preexisting
> but seemed like good things to fix.
>
> Signed-off-by: Jeff Layton <jlayton@kernel.org>
> ---
> Changes in v3:
> - Stop warning twice per unexportable path
> - Keep a warning for the ip_map and unix_gid answers, which have no fallback
> - Don't warn about a submount nobody asked to export
> - Quieten the same message on the pipefs downcall (preexisting)
> - Link to v2: https://lore.kernel.org/r/20260911-nl-crossmnt-v2-0-f4d869951c5e@kernel.org
>
> Changes in v2:
> - Retry instead of caching a negative when the failure may be transient:
> kernel errno, statfs, fsidd, and an undelivered negative fallback
> - Drop a deferred request once it has been answered
> - Bound the junction path copied into exportent.e_path (preexisting)
> - Give each worker its own netlink command socket (preexisting)
> - Bound the retry queues (preexisting)
> - Link to v1: https://lore.kernel.org/r/20260909-nl-crossmnt-v1-0-4064b5a3bd85@kernel.org
>
> ---
> Jeff Layton (11):
> mountd: factor out the per-path export attribute computation
> mountd: handle unmountable paths and junctions in the netlink downcall
> mountd: answer requests the kernel rejects on the netlink downcall
> mountd: don't leak the parent export's fsid onto crossmnt submounts
> mountd: retry unresolvable fsid lookups on the netlink downcall
> mountd: bound the junction path before copying it into e_path
> mountd: give each worker its own netlink command socket
> mountd: retry export attributes that fail to resolve for a passing reason
> mountd: drop a deferred fsid lookup once it has been answered
> mountd: bound the retry queues
> mountd: don't warn about pipefs submounts nobody asked to export
>
> support/export/cache.c | 1391 +++++++++++++++++++++++++++++++++++++-----------
> 1 file changed, 1084 insertions(+), 307 deletions(-)
> ---
> base-commit: 630393e9c73667b3fda6a7507cb5171ec0ec38dc
> change-id: 20260908-nl-crossmnt-85584b71c721
>
> Best regards,
I tested the v3 now, and it seems to be working well.
^ permalink raw reply [flat|nested] 14+ messages in thread
* Re: [PATCH nfs-utils v3 00/11] mountd: bugfixes for up/downcall interfaces
2026-09-14 13:14 [PATCH nfs-utils v3 00/11] mountd: bugfixes for up/downcall interfaces Jeff Layton
` (11 preceding siblings ...)
2026-09-15 9:24 ` [PATCH nfs-utils v3 00/11] mountd: bugfixes for up/downcall interfaces Mantas Mikulėnas
@ 2026-09-17 7:21 ` Steve Dickson
12 siblings, 0 replies; 14+ messages in thread
From: Steve Dickson @ 2026-09-17 7:21 UTC (permalink / raw)
To: Jeff Layton, Mantas Mikulėnas; +Cc: Chuck Lever, linux-nfs
On 9/14/26 9:14 AM, Jeff Layton wrote:
> The main difference in this version is some cleanup to the warning
> messages, so that we don't warn about internal events. Mantas, if you're
> able to test this version too, that would be great.
>
> Mantas reported a problem with his configuration that was using crossmnt
> and autofs to export a number of btrfs filesystems. Some analysis with
> Claude also uncovered a few other bugs. This version also fixes a
> number of problems that Sashiko flagged. Some of them were preexisting
> but seemed like good things to fix.
>
> Signed-off-by: Jeff Layton <jlayton@kernel.org>
> ---
> Changes in v3:
> - Stop warning twice per unexportable path
> - Keep a warning for the ip_map and unix_gid answers, which have no fallback
> - Don't warn about a submount nobody asked to export
> - Quieten the same message on the pipefs downcall (preexisting)
> - Link to v2: https://lore.kernel.org/r/20260911-nl-crossmnt-v2-0-f4d869951c5e@kernel.org
>
> Changes in v2:
> - Retry instead of caching a negative when the failure may be transient:
> kernel errno, statfs, fsidd, and an undelivered negative fallback
> - Drop a deferred request once it has been answered
> - Bound the junction path copied into exportent.e_path (preexisting)
> - Give each worker its own netlink command socket (preexisting)
> - Bound the retry queues (preexisting)
> - Link to v1: https://lore.kernel.org/r/20260909-nl-crossmnt-v1-0-4064b5a3bd85@kernel.org
>
> ---
> Jeff Layton (11):
> mountd: factor out the per-path export attribute computation
> mountd: handle unmountable paths and junctions in the netlink downcall
> mountd: answer requests the kernel rejects on the netlink downcall
> mountd: don't leak the parent export's fsid onto crossmnt submounts
> mountd: retry unresolvable fsid lookups on the netlink downcall
> mountd: bound the junction path before copying it into e_path
> mountd: give each worker its own netlink command socket
> mountd: retry export attributes that fail to resolve for a passing reason
> mountd: drop a deferred fsid lookup once it has been answered
> mountd: bound the retry queues
> mountd: don't warn about pipefs submounts nobody asked to export
>
> support/export/cache.c | 1391 +++++++++++++++++++++++++++++++++++++-----------
> 1 file changed, 1084 insertions(+), 307 deletions(-)
> ---
> base-commit: 630393e9c73667b3fda6a7507cb5171ec0ec38dc
> change-id: 20260908-nl-crossmnt-85584b71c721
>
> Best regards,
Committed... (tag: nfs-utils-2-9-3-rc6)
steved
^ permalink raw reply [flat|nested] 14+ messages in thread
end of thread, other threads:[~2026-09-17 7:27 UTC | newest]
Thread overview: 14+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-14 13:14 [PATCH nfs-utils v3 00/11] mountd: bugfixes for up/downcall interfaces Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 01/11] mountd: factor out the per-path export attribute computation Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 02/11] mountd: handle unmountable paths and junctions in the netlink downcall Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 03/11] mountd: answer requests the kernel rejects on " Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 04/11] mountd: don't leak the parent export's fsid onto crossmnt submounts Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 05/11] mountd: retry unresolvable fsid lookups on the netlink downcall Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 06/11] mountd: bound the junction path before copying it into e_path Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 07/11] mountd: give each worker its own netlink command socket Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 08/11] mountd: retry export attributes that fail to resolve for a passing reason Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 09/11] mountd: drop a deferred fsid lookup once it has been answered Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 10/11] mountd: bound the retry queues Jeff Layton
2026-09-14 13:14 ` [PATCH nfs-utils v3 11/11] mountd: don't warn about pipefs submounts nobody asked to export Jeff Layton
2026-09-15 9:24 ` [PATCH nfs-utils v3 00/11] mountd: bugfixes for up/downcall interfaces Mantas Mikulėnas
2026-09-17 7:21 ` Steve Dickson
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox