* [PATCH RFC v2 0/5] NFS: isolate mTLS client credentials by network namespace
@ 2026-09-25 19:16 Chuck Lever
2026-09-25 19:16 ` [PATCH RFC v2 1/5] NFS: name the init_nfs_fs() error labels Chuck Lever
` (4 more replies)
0 siblings, 5 replies; 6+ messages in thread
From: Chuck Lever @ 2026-09-25 19:16 UTC (permalink / raw)
To: Trond Myklebust, Anna Schumaker, David S. Miller, Eric Dumazet,
Jakub Kicinski, Paolo Abeni, Simon Horman, Jonathan Corbet,
Shuah Khan, Randy Dunlap, Christian Brauner, David Howells,
Sagi Grimberg
Cc: linux-nfs, keyrings, kernel-tls-handshake, netdev, linux-doc,
Chuck Lever
An xprtsec=mtls mount names its client certificate and private key
by keyring serial number, and tlshd reads those keys with its own
credentials. Each key therefore has to grant user read permission,
and any tlshd on the host that learns a serial number can read it.
Nothing separates one network namespace's credentials from
another's.
This series makes the network namespace the isolation domain. The
RFC thread asked whether the user or mount namespace is the better
binding. tlshd services the handshake socket of one network
namespace, so that is the namespace it already lives in.
https://lore.kernel.org/linux-nfs/20260602154740.49861-1-cel@kernel.org/
The keyring is held in struct nfs_net rather than found by name,
because /proc/keys is not namespace scoped. Its serial number
travels with each handshake instead (patches 3-4), and tlshd's
possession of that keyring is what lets a provisioned key grant no
user read permission at all. Both ends of the link exist today, in
tls_handshake_accept() and in tlshd. The handshake genetlink ABI and
the cert_serial= and privkey_serial= mount options do not change.
The keyring is owned by the owner of the network namespace's user
namespace, and grants no permission to anyone else. In a container
with its own user namespace, the container's root can provision the
keyring and its tlshd can link it, while global root on the host
cannot write to it without entering that user namespace. For a
namespace in the initial user namespace the owner is global root,
as before.
Still open is how userspace names the kernel-held keyring. Patch 5
prototypes a request_key type handled in the kernel and tagged
KEY_TYPE_NET_DOMAIN. A keyctl modeled on KEYCTL_GET_PERSISTENT, or a
read-only attribute on the netns-tagged nfs_client sysfs kobject,
would do the same job if the keyrings maintainers prefer one. Keys
that userspace adds are quota-charged by user namespace while the
keyring lives in nfs_net. Confirmation that this is sane when the
two boundaries differ would be welcome.
nfstlskey, the provisioning tool that consumes the key type, is
merged in https://github.com/linux-nfs/ktls-utils/ .
Tested on a v7.3-rc4 kernel carrying this series, on one Fedora VM
acting as both NFS client and NFSD, with tlshd from ktls-utils
1.4.0.
---
Changes in v2:
- Write "serial number" rather than "serial" throughout (Randy)
- Give the .nfs keyring to the netns's user_ns owner (sashiko)
- Link to v1: https://patch.msgid.link/20260918-nfs-mtls-identity-v1-0-197e568d78a7@kernel.org
---
Chuck Lever (5):
NFS: name the init_nfs_fs() error labels
NFS: allocate the .nfs keyring per network namespace
SUNRPC: pass a keyring serial number to the TLS handshake
NFS: name the namespace .nfs keyring in the x509 handshake
NFS: add a key type that reveals the namespace .nfs keyring serial number
Documentation/filesystems/nfs/index.rst | 1 +
Documentation/filesystems/nfs/keyring.rst | 69 ++++++++++
fs/nfs/client.c | 9 +-
fs/nfs/fs_context.c | 1 +
fs/nfs/inode.c | 211 ++++++++++++++++++++++--------
fs/nfs/netns.h | 9 ++
fs/nfs/nfs3client.c | 1 +
fs/nfs/nfs4client.c | 1 +
include/linux/sunrpc/xprt.h | 1 +
net/sunrpc/xprtsock.c | 1 +
10 files changed, 248 insertions(+), 56 deletions(-)
---
base-commit: aa98230e410f0ed212b6788c46b1e4d49e0ff7ca
change-id: 20260917-nfs-mtls-identity-04c14f6f348d
Best regards,
--
Chuck Lever <cel@kernel.org>
^ permalink raw reply [flat|nested] 6+ messages in thread
* [PATCH RFC v2 1/5] NFS: name the init_nfs_fs() error labels
2026-09-25 19:16 [PATCH RFC v2 0/5] NFS: isolate mTLS client credentials by network namespace Chuck Lever
@ 2026-09-25 19:16 ` Chuck Lever
2026-09-25 19:16 ` [PATCH RFC v2 2/5] NFS: allocate the .nfs keyring per network namespace Chuck Lever
` (3 subsequent siblings)
4 siblings, 0 replies; 6+ messages in thread
From: Chuck Lever @ 2026-09-25 19:16 UTC (permalink / raw)
To: Trond Myklebust, Anna Schumaker, David S. Miller, Eric Dumazet,
Jakub Kicinski, Paolo Abeni, Simon Horman, Jonathan Corbet,
Shuah Khan, Randy Dunlap, Christian Brauner, David Howells,
Sagi Grimberg
Cc: linux-nfs, keyrings, kernel-tls-handshake, netdev, linux-doc,
Chuck Lever
The unwind labels in init_nfs_fs() are numbered, and the numbering
already skips out8, so a reader has to count the label block to
learn what each one undoes. Inserting an init step means either
renumbering every label below it or leaving the sequence out of
order, and a goto that picks the wrong number unwinds the wrong
step.
Name each label for the step it undoes, as coding-style.rst asks.
Signed-off-by: Chuck Lever <cel@kernel.org>
---
fs/nfs/inode.c | 40 ++++++++++++++++++++--------------------
1 file changed, 20 insertions(+), 20 deletions(-)
diff --git a/fs/nfs/inode.c b/fs/nfs/inode.c
index 3022454f7698..832923be43a9 100644
--- a/fs/nfs/inode.c
+++ b/fs/nfs/inode.c
@@ -2722,64 +2722,64 @@ static int __init init_nfs_fs(void)
err = nfs_sysfs_init();
if (err < 0)
- goto out10;
+ goto err_keyring;
err = register_pernet_subsys(&nfs_net_ops);
if (err < 0)
- goto out9;
+ goto err_sysfs;
err = nfsiod_start();
if (err)
- goto out7;
+ goto err_pernet;
err = nfs_fs_proc_init();
if (err)
- goto out6;
+ goto err_nfsiod;
err = nfs_init_nfspagecache();
if (err)
- goto out5;
+ goto err_proc;
err = nfs_init_inodecache();
if (err)
- goto out4;
+ goto err_nfspagecache;
err = nfs_init_readpagecache();
if (err)
- goto out3;
+ goto err_inodecache;
err = nfs_init_writepagecache();
if (err)
- goto out2;
+ goto err_readpagecache;
err = nfs_init_directcache();
if (err)
- goto out1;
+ goto err_writepagecache;
err = register_nfs_fs();
if (err)
- goto out0;
+ goto err_directcache;
return 0;
-out0:
+err_directcache:
nfs_destroy_directcache();
-out1:
+err_writepagecache:
nfs_destroy_writepagecache();
-out2:
+err_readpagecache:
nfs_destroy_readpagecache();
-out3:
+err_inodecache:
nfs_destroy_inodecache();
-out4:
+err_nfspagecache:
nfs_destroy_nfspagecache();
-out5:
+err_proc:
nfs_fs_proc_exit();
-out6:
+err_nfsiod:
nfsiod_stop();
-out7:
+err_pernet:
unregister_pernet_subsys(&nfs_net_ops);
-out9:
+err_sysfs:
nfs_sysfs_exit();
-out10:
+err_keyring:
nfs_exit_keyring();
return err;
}
--
2.55.0
^ permalink raw reply related [flat|nested] 6+ messages in thread
* [PATCH RFC v2 2/5] NFS: allocate the .nfs keyring per network namespace
2026-09-25 19:16 [PATCH RFC v2 0/5] NFS: isolate mTLS client credentials by network namespace Chuck Lever
2026-09-25 19:16 ` [PATCH RFC v2 1/5] NFS: name the init_nfs_fs() error labels Chuck Lever
@ 2026-09-25 19:16 ` Chuck Lever
2026-09-25 19:16 ` [PATCH RFC v2 3/5] SUNRPC: pass a keyring serial number to the TLS handshake Chuck Lever
` (2 subsequent siblings)
4 siblings, 0 replies; 6+ messages in thread
From: Chuck Lever @ 2026-09-25 19:16 UTC (permalink / raw)
To: Trond Myklebust, Anna Schumaker, David S. Miller, Eric Dumazet,
Jakub Kicinski, Paolo Abeni, Simon Horman, Jonathan Corbet,
Shuah Khan, Randy Dunlap, Christian Brauner, David Howells,
Sagi Grimberg
Cc: linux-nfs, keyrings, kernel-tls-handshake, netdev, linux-doc,
Chuck Lever
Commit 87268f7a4f1f ("nfs: create a kernel keyring") allocates one
.nfs keyring at module load, and nothing in the NFS client reads it.
One module-wide keyring also cannot isolate x.509 credentials
between network namespaces. Each tlshd instance services the
handshake socket of one network namespace, so a credential
provisioned for that namespace's mounts has to be reachable by that
tlshd and by no other.
Allocate one .nfs keyring per network namespace in nfs_net_init()
and release it in nfs_net_exit(). tlshd finds a keyring by name
through /proc/keys, which is not namespace scoped, so each handshake
request has to carry the keyring serial number instead. Allocate
the keyring under a kernel credential rather than that of the task
creating the namespace, so an LSM labels every namespace's keyring
the same way.
The keyring grants its owner every permission and everyone else
none. Make the owner of the network namespace's user namespace the
keyring's owner, so that in a container with its own user namespace
the container's root can provision the keyring and its tlshd can
link it.
Signed-off-by: Chuck Lever <cel@kernel.org>
---
fs/nfs/inode.c | 83 ++++++++++++++++++++++++++++++++--------------------------
fs/nfs/netns.h | 2 ++
2 files changed, 48 insertions(+), 37 deletions(-)
diff --git a/fs/nfs/inode.c b/fs/nfs/inode.c
index 832923be43a9..2b494fa5ecc0 100644
--- a/fs/nfs/inode.c
+++ b/fs/nfs/inode.c
@@ -2641,11 +2641,52 @@ static int nfsiod_start(void)
unsigned int nfs_net_id;
EXPORT_SYMBOL_GPL(nfs_net_id);
+#ifdef CONFIG_KEYS
+static int nfs_init_keyring(struct net *net)
+{
+ struct nfs_net *nn = net_generic(net, nfs_net_id);
+ struct cred *cred;
+ struct key *keyring;
+
+ cred = prepare_kernel_cred(&init_task);
+ if (!cred)
+ return -ENOMEM;
+ keyring = keyring_alloc(".nfs", net->user_ns->owner,
+ net->user_ns->group, cred,
+ (KEY_POS_ALL & ~KEY_POS_SETATTR) |
+ (KEY_USR_ALL & ~KEY_USR_SETATTR),
+ KEY_ALLOC_NOT_IN_QUOTA, NULL, NULL);
+ put_cred(cred);
+ if (IS_ERR(keyring))
+ return PTR_ERR(keyring);
+ nn->nfs_keyring = keyring;
+ return 0;
+}
+
+static void nfs_exit_keyring(struct nfs_net *nn)
+{
+ key_put(nn->nfs_keyring);
+}
+#else
+static inline int nfs_init_keyring(struct net *net)
+{
+ return 0;
+}
+
+static inline void nfs_exit_keyring(struct nfs_net *nn)
+{
+}
+#endif /* CONFIG_KEYS */
+
static int nfs_net_init(struct net *net)
{
struct nfs_net *nn = net_generic(net, nfs_net_id);
int err;
+ err = nfs_init_keyring(net);
+ if (err)
+ return err;
+
nfs_clients_init(net);
if (!rpc_proc_register(net, &nn->rpcstats)) {
@@ -2663,14 +2704,18 @@ static int nfs_net_init(struct net *net)
rpc_proc_unregister(net, "nfs");
err_proc_rpc:
nfs_clients_exit(net);
+ nfs_exit_keyring(nn);
return err;
}
static void nfs_net_exit(struct net *net)
{
+ struct nfs_net *nn = net_generic(net, nfs_net_id);
+
rpc_proc_unregister(net, "nfs");
nfs_fs_proc_net_exit(net);
nfs_clients_exit(net);
+ nfs_exit_keyring(nn);
}
static struct pernet_operations nfs_net_ops = {
@@ -2680,35 +2725,6 @@ static struct pernet_operations nfs_net_ops = {
.size = sizeof(struct nfs_net),
};
-#ifdef CONFIG_KEYS
-static struct key *nfs_keyring;
-
-static int __init nfs_init_keyring(void)
-{
- nfs_keyring = keyring_alloc(".nfs",
- GLOBAL_ROOT_UID, GLOBAL_ROOT_GID,
- current_cred(),
- (KEY_POS_ALL & ~KEY_POS_SETATTR) |
- (KEY_USR_ALL & ~KEY_USR_SETATTR),
- KEY_ALLOC_NOT_IN_QUOTA, NULL, NULL);
- return PTR_ERR_OR_ZERO(nfs_keyring);
-}
-
-static void nfs_exit_keyring(void)
-{
- key_put(nfs_keyring);
-}
-#else
-static inline int nfs_init_keyring(void)
-{
- return 0;
-}
-
-static inline void nfs_exit_keyring(void)
-{
-}
-#endif /* CONFIG_KEYS */
-
/*
* Initialize NFS
*/
@@ -2716,13 +2732,9 @@ static int __init init_nfs_fs(void)
{
int err;
- err = nfs_init_keyring();
- if (err)
- return err;
-
err = nfs_sysfs_init();
if (err < 0)
- goto err_keyring;
+ return err;
err = register_pernet_subsys(&nfs_net_ops);
if (err < 0)
@@ -2779,8 +2791,6 @@ static int __init init_nfs_fs(void)
unregister_pernet_subsys(&nfs_net_ops);
err_sysfs:
nfs_sysfs_exit();
-err_keyring:
- nfs_exit_keyring();
return err;
}
@@ -2796,7 +2806,6 @@ static void __exit exit_nfs_fs(void)
nfs_fs_proc_exit();
nfsiod_stop();
nfs_sysfs_exit();
- nfs_exit_keyring();
}
/* Not quite true; I just maintain it */
diff --git a/fs/nfs/netns.h b/fs/nfs/netns.h
index 36658579100d..da0854510404 100644
--- a/fs/nfs/netns.h
+++ b/fs/nfs/netns.h
@@ -16,6 +16,7 @@ struct bl_dev_msg {
uint32_t major, minor;
};
+struct key;
struct nfs_netns_client;
struct nfs_net {
@@ -36,6 +37,7 @@ struct nfs_net {
#endif /* CONFIG_NFS_V4 */
struct nfs_netns_client *nfs_client;
spinlock_t nfs_client_lock;
+ struct key *nfs_keyring;
ktime_t boot_time;
struct rpc_stat rpcstats;
#ifdef CONFIG_PROC_FS
--
2.55.0
^ permalink raw reply related [flat|nested] 6+ messages in thread
* [PATCH RFC v2 3/5] SUNRPC: pass a keyring serial number to the TLS handshake
2026-09-25 19:16 [PATCH RFC v2 0/5] NFS: isolate mTLS client credentials by network namespace Chuck Lever
2026-09-25 19:16 ` [PATCH RFC v2 1/5] NFS: name the init_nfs_fs() error labels Chuck Lever
2026-09-25 19:16 ` [PATCH RFC v2 2/5] NFS: allocate the .nfs keyring per network namespace Chuck Lever
@ 2026-09-25 19:16 ` Chuck Lever
2026-09-25 19:16 ` [PATCH RFC v2 4/5] NFS: name the namespace .nfs keyring in the x509 handshake Chuck Lever
2026-09-25 19:16 ` [PATCH RFC v2 5/5] NFS: add a key type that reveals the namespace .nfs keyring serial number Chuck Lever
4 siblings, 0 replies; 6+ messages in thread
From: Chuck Lever @ 2026-09-25 19:16 UTC (permalink / raw)
To: Trond Myklebust, Anna Schumaker, David S. Miller, Eric Dumazet,
Jakub Kicinski, Paolo Abeni, Simon Horman, Jonathan Corbet,
Shuah Khan, Randy Dunlap, Christian Brauner, David Howells,
Sagi Grimberg
Cc: linux-nfs, keyrings, kernel-tls-handshake, netdev, linux-doc,
Chuck Lever
The handshake upcall links the keyring named in ta_keyring into
tlshd's process keyring before tlshd reads the client certificate and
private key. That link is how tlshd gains possession of keys that
grant no user read permission. xprtsock never fills the field, so an
x509 handshake can present only keys that tlshd's own credentials can
read.
Add a keyring serial number to struct xprtsec_parms and pass it to
the x509 handshake, so an RPC client can name the keyring that holds
its credentials.
Signed-off-by: Chuck Lever <cel@kernel.org>
---
fs/nfs/fs_context.c | 1 +
fs/nfs/nfs3client.c | 1 +
fs/nfs/nfs4client.c | 1 +
include/linux/sunrpc/xprt.h | 1 +
net/sunrpc/xprtsock.c | 1 +
5 files changed, 5 insertions(+)
diff --git a/fs/nfs/fs_context.c b/fs/nfs/fs_context.c
index 1967de7d1dff..f8f5f8b2954e 100644
--- a/fs/nfs/fs_context.c
+++ b/fs/nfs/fs_context.c
@@ -1750,6 +1750,7 @@ static int nfs_init_fs_context(struct fs_context *fc)
ctx->minorversion = 0;
ctx->need_mount = true;
ctx->xprtsec.policy = RPC_XPRTSEC_NONE;
+ ctx->xprtsec.keyring_serial = TLS_NO_KEYRING;
ctx->xprtsec.cert_serial = TLS_NO_CERT;
ctx->xprtsec.privkey_serial = TLS_NO_PRIVKEY;
diff --git a/fs/nfs/nfs3client.c b/fs/nfs/nfs3client.c
index 5d97c1d38bb6..c5d8278838b2 100644
--- a/fs/nfs/nfs3client.c
+++ b/fs/nfs/nfs3client.c
@@ -101,6 +101,7 @@ struct nfs_client *nfs3_set_ds_client(struct nfs_server *mds_srv,
.cred = mds_srv->cred,
.xprtsec = {
.policy = RPC_XPRTSEC_NONE,
+ .keyring_serial = TLS_NO_KEYRING,
.cert_serial = TLS_NO_CERT,
.privkey_serial = TLS_NO_PRIVKEY,
},
diff --git a/fs/nfs/nfs4client.c b/fs/nfs/nfs4client.c
index b661f446ea49..05df0fcabfb5 100644
--- a/fs/nfs/nfs4client.c
+++ b/fs/nfs/nfs4client.c
@@ -809,6 +809,7 @@ struct nfs_client *nfs4_set_ds_client(struct nfs_server *mds_srv,
.cred = mds_srv->cred,
.xprtsec = {
.policy = RPC_XPRTSEC_NONE,
+ .keyring_serial = TLS_NO_KEYRING,
.cert_serial = TLS_NO_CERT,
.privkey_serial = TLS_NO_PRIVKEY,
},
diff --git a/include/linux/sunrpc/xprt.h b/include/linux/sunrpc/xprt.h
index a82045804d34..005730623f3b 100644
--- a/include/linux/sunrpc/xprt.h
+++ b/include/linux/sunrpc/xprt.h
@@ -145,6 +145,7 @@ struct xprtsec_parms {
enum xprtsec_policies policy;
/* authentication material */
+ key_serial_t keyring_serial;
key_serial_t cert_serial;
key_serial_t privkey_serial;
};
diff --git a/net/sunrpc/xprtsock.c b/net/sunrpc/xprtsock.c
index 7f60723fa64d..2825dec82d6e 100644
--- a/net/sunrpc/xprtsock.c
+++ b/net/sunrpc/xprtsock.c
@@ -2636,6 +2636,7 @@ static int xs_tls_handshake_sync(struct rpc_xprt *lower_xprt, struct xprtsec_par
goto out_put_xprt;
break;
case RPC_XPRTSEC_TLS_X509:
+ args.ta_keyring = xprtsec->keyring_serial;
args.ta_my_cert = xprtsec->cert_serial;
args.ta_my_privkey = xprtsec->privkey_serial;
rc = tls_client_hello_x509(&args, GFP_KERNEL);
--
2.55.0
^ permalink raw reply related [flat|nested] 6+ messages in thread
* [PATCH RFC v2 4/5] NFS: name the namespace .nfs keyring in the x509 handshake
2026-09-25 19:16 [PATCH RFC v2 0/5] NFS: isolate mTLS client credentials by network namespace Chuck Lever
` (2 preceding siblings ...)
2026-09-25 19:16 ` [PATCH RFC v2 3/5] SUNRPC: pass a keyring serial number to the TLS handshake Chuck Lever
@ 2026-09-25 19:16 ` Chuck Lever
2026-09-25 19:16 ` [PATCH RFC v2 5/5] NFS: add a key type that reveals the namespace .nfs keyring serial number Chuck Lever
4 siblings, 0 replies; 6+ messages in thread
From: Chuck Lever @ 2026-09-25 19:16 UTC (permalink / raw)
To: Trond Myklebust, Anna Schumaker, David S. Miller, Eric Dumazet,
Jakub Kicinski, Paolo Abeni, Simon Horman, Jonathan Corbet,
Shuah Khan, Randy Dunlap, Christian Brauner, David Howells,
Sagi Grimberg
Cc: linux-nfs, keyrings, kernel-tls-handshake, netdev, linux-doc,
Chuck Lever
Keys on the per-namespace .nfs keyring need not grant user read
permission once tlshd possesses the keyring, but nothing tells tlshd
which keyring to possess. An xprtsec=mtls mount can present only keys
that tlshd's own credentials can read, and one namespace's keys are
not isolated from another's.
Record the serial number of the namespace's .nfs keyring in the
nfs_client's xprtsec parameters under the x509 policy, so tlshd
possesses that keyring when it reads the certificate and private
key.
Signed-off-by: Chuck Lever <cel@kernel.org>
---
fs/nfs/client.c | 9 ++++++++-
fs/nfs/netns.h | 7 +++++++
2 files changed, 15 insertions(+), 1 deletion(-)
diff --git a/fs/nfs/client.c b/fs/nfs/client.c
index 60386330aeec..bda1d19db527 100644
--- a/fs/nfs/client.c
+++ b/fs/nfs/client.c
@@ -191,6 +191,13 @@ struct nfs_client *nfs_alloc_client(const struct nfs_client_initdata *cl_init)
clp->cl_principal = "*";
clp->cl_xprtsec = cl_init->xprtsec;
+ /*
+ * Every client in a namespace names the same keyring, so
+ * nfs_match_client() does not compare keyring_serial.
+ */
+ if (clp->cl_xprtsec.policy == RPC_XPRTSEC_TLS_X509)
+ clp->cl_xprtsec.keyring_serial =
+ key_serial(nfs_net_keyring(clp->cl_net));
return clp;
error_cleanup:
@@ -549,7 +556,7 @@ int nfs_create_rpc_client(struct nfs_client *clp,
.version = clp->rpc_ops->version,
.authflavor = flavor,
.cred = cl_init->cred,
- .xprtsec = cl_init->xprtsec,
+ .xprtsec = clp->cl_xprtsec,
.connect_timeout = cl_init->connect_timeout,
.reconnect_timeout = cl_init->reconnect_timeout,
};
diff --git a/fs/nfs/netns.h b/fs/nfs/netns.h
index da0854510404..c8ca7b989f4b 100644
--- a/fs/nfs/netns.h
+++ b/fs/nfs/netns.h
@@ -47,4 +47,11 @@ struct nfs_net {
extern unsigned int nfs_net_id;
+static inline struct key *nfs_net_keyring(struct net *net)
+{
+ struct nfs_net *nn = net_generic(net, nfs_net_id);
+
+ return nn->nfs_keyring;
+}
+
#endif
--
2.55.0
^ permalink raw reply related [flat|nested] 6+ messages in thread
* [PATCH RFC v2 5/5] NFS: add a key type that reveals the namespace .nfs keyring serial number
2026-09-25 19:16 [PATCH RFC v2 0/5] NFS: isolate mTLS client credentials by network namespace Chuck Lever
` (3 preceding siblings ...)
2026-09-25 19:16 ` [PATCH RFC v2 4/5] NFS: name the namespace .nfs keyring in the x509 handshake Chuck Lever
@ 2026-09-25 19:16 ` Chuck Lever
4 siblings, 0 replies; 6+ messages in thread
From: Chuck Lever @ 2026-09-25 19:16 UTC (permalink / raw)
To: Trond Myklebust, Anna Schumaker, David S. Miller, Eric Dumazet,
Jakub Kicinski, Paolo Abeni, Simon Horman, Jonathan Corbet,
Shuah Khan, Randy Dunlap, Christian Brauner, David Howells,
Sagi Grimberg
Cc: linux-nfs, keyrings, kernel-tls-handshake, netdev, linux-doc,
Chuck Lever
The per-namespace .nfs keyring is linked into no keyring that
userspace possesses, so a tool that provisions x509 credentials, or a
mount helper that searches for them, cannot reach it. Every keyring
operation from userspace takes a serial number, and nothing hands
this one out.
Register an "nfs_keyring" key type whose request_key handler runs in
the caller's context, without an upcall, and instantiates the key
with the serial number of the caller's namespace .nfs keyring.
Userspace reads the serial number with one request_key() and one
keyctl_read().
A keyring search runs before the handler does, and a task keeps its
session keyring across setns(). The type carries KEY_TYPE_NET_DOMAIN
so a key instantiated in one namespace does not answer a request from
another. Its preparse accepts no payload but the caller's own serial
number, so a key planted by add_key() cannot misdirect a later
request.
Register the type after register_pernet_subsys() has installed
nfs_net_id, and unregister it before the per-namespace state goes
away. unregister_key_type() waits for in-flight request_key() calls
to drain, so the handler never runs against a freed struct nfs_net.
Signed-off-by: Chuck Lever <cel@kernel.org>
---
Documentation/filesystems/nfs/index.rst | 1 +
Documentation/filesystems/nfs/keyring.rst | 69 +++++++++++++++++++++++
fs/nfs/inode.c | 94 ++++++++++++++++++++++++++++++-
3 files changed, 163 insertions(+), 1 deletion(-)
diff --git a/Documentation/filesystems/nfs/index.rst b/Documentation/filesystems/nfs/index.rst
index a29a212b5b4d..dab5f16aaad1 100644
--- a/Documentation/filesystems/nfs/index.rst
+++ b/Documentation/filesystems/nfs/index.rst
@@ -7,6 +7,7 @@ NFS
:maxdepth: 1
client-identifier
+ keyring
exporting
localio
pnfs
diff --git a/Documentation/filesystems/nfs/keyring.rst b/Documentation/filesystems/nfs/keyring.rst
new file mode 100644
index 000000000000..29367e71a801
--- /dev/null
+++ b/Documentation/filesystems/nfs/keyring.rst
@@ -0,0 +1,69 @@
+.. SPDX-License-Identifier: GPL-2.0
+
+====================================
+The per-namespace NFS client keyring
+====================================
+
+The NFS client holds one keyring, named ".nfs", per network
+namespace. An xprtsec=mtls mount presents the client certificate and
+private key named by its cert_serial= and privkey_serial= mount
+options, and the handshake links the mount's namespace .nfs keyring
+into tlshd before tlshd reads those keys. Keys placed on the .nfs
+keyring therefore need not grant user read permission: tlshd reaches
+them as a possessor, and a tlshd in another network namespace never
+possesses them.
+
+The keyring is owned by the owner of the network namespace's user
+namespace, with KEY_POS_ALL and KEY_USR_ALL less SETATTR. It is not
+charged to any quota. Keys added to it by userspace are charged to
+the user that adds them.
+
+Finding the keyring serial number
+=================================
+
+The .nfs keyring is not linked into any keyring that userspace
+possesses, so a serial number is the only way to name it. The NFS
+client registers a key type, "nfs_keyring", whose sole purpose is to
+hand that serial number out. Its request_key handler runs in the
+caller's context and does not upcall to /sbin/request-key.
+
+To obtain the serial number, request a key of type "nfs_keyring" with
+the description ".nfs" and read its payload::
+
+ id=$(keyctl request2 nfs_keyring .nfs "" @s)
+ serial=$(keyctl print $id)
+
+The contract of the key type is:
+
+ * The description is the string ".nfs". Any other description is
+ rejected with EINVAL before a key is allocated.
+
+ * The callout info must be present. request_key() with a NULL
+ callout_info never invokes a handler and returns ENOKEY when no
+ matching key exists, so use request_key() with an empty string, or
+ ``keyctl request2`` rather than ``keyctl request``. The content of
+ the callout info is ignored.
+
+ * The payload is the keyring serial number as decimal ASCII digits
+ with no terminating NUL. The return value of keyctl_read() gives
+ the length; a buffer of 12 bytes is sufficient.
+
+ * add_key() with this type accepts only that payload. A key carrying
+ any other value is rejected with EINVAL, so a key found in the
+ caller's keyrings always holds the serial number of the caller's
+ namespace.
+
+ * The key type carries KEY_TYPE_NET_DOMAIN. A key instantiated in one
+ network namespace does not answer a request made in another, so a
+ process that keeps its session keyring across setns() receives the
+ serial number of the namespace it is in at the time of the request.
+
+The serial number is not a secret. The keyring's own permissions
+decide what a caller can do with it.
+
+Provisioning credentials
+========================
+
+Add the client certificate and private key to the .nfs keyring as
+keys that grant no user read permission, then pass their serial
+numbers as cert_serial= and privkey_serial= mount options.
diff --git a/fs/nfs/inode.c b/fs/nfs/inode.c
index 2b494fa5ecc0..fe7a05dec02f 100644
--- a/fs/nfs/inode.c
+++ b/fs/nfs/inode.c
@@ -42,6 +42,12 @@
#include <linux/uaccess.h>
#include <linux/iversion.h>
#include <linux/fileattr.h>
+#include <linux/nsproxy.h>
+#include <linux/key-type.h>
+#include <keys/user-type.h>
+#ifdef CONFIG_KEYS
+#include <keys/request_key_auth-type.h>
+#endif
#include "nfs4_fs.h"
#include "callback.h"
@@ -2667,6 +2673,76 @@ static void nfs_exit_keyring(struct nfs_net *nn)
{
key_put(nn->nfs_keyring);
}
+
+static int nfs_keyring_vet_description(const char *desc)
+{
+ return strcmp(desc, ".nfs") ? -EINVAL : 0;
+}
+
+static int nfs_keyring_format_serial(char *buf, size_t len)
+{
+ struct key *keyring = nfs_net_keyring(current->nsproxy->net_ns);
+
+ return snprintf(buf, len, "%d", key_serial(keyring));
+}
+
+/*
+ * A key planted by add_key() answers request_key() before the handler
+ * runs. Accept only the serial number the handler would produce.
+ */
+static int nfs_keyring_preparse(struct key_preparsed_payload *prep)
+{
+ char serial[12];
+ int len;
+
+ len = nfs_keyring_format_serial(serial, sizeof(serial));
+ if (prep->datalen != len || !prep->data ||
+ memcmp(prep->data, serial, len))
+ return -EINVAL;
+ return user_preparse(prep);
+}
+
+static int nfs_keyring_request_key(struct key *authkey, void *aux)
+{
+ struct request_key_auth *rka = get_request_key_auth(authkey);
+ char serial[12];
+ int len, ret;
+
+ len = nfs_keyring_format_serial(serial, sizeof(serial));
+ ret = key_instantiate_and_link(rka->target_key, serial, len,
+ rka->dest_keyring, authkey);
+ if (ret < 0)
+ complete_request_key(authkey, ret);
+ return ret;
+}
+
+/*
+ * A session keyring survives setns(). Without the net domain tag, a
+ * key instantiated in one namespace answers a request from another.
+ */
+static struct key_type key_type_nfs_keyring = {
+ .name = "nfs_keyring",
+ .flags = KEY_TYPE_NET_DOMAIN,
+ .vet_description = nfs_keyring_vet_description,
+ .preparse = nfs_keyring_preparse,
+ .free_preparse = user_free_preparse,
+ .instantiate = generic_key_instantiate,
+ .revoke = user_revoke,
+ .destroy = user_destroy,
+ .describe = user_describe,
+ .read = user_read,
+ .request_key = nfs_keyring_request_key,
+};
+
+static int __init nfs_register_key_type(void)
+{
+ return register_key_type(&key_type_nfs_keyring);
+}
+
+static void nfs_unregister_key_type(void)
+{
+ unregister_key_type(&key_type_nfs_keyring);
+}
#else
static inline int nfs_init_keyring(struct net *net)
{
@@ -2676,6 +2752,15 @@ static inline int nfs_init_keyring(struct net *net)
static inline void nfs_exit_keyring(struct nfs_net *nn)
{
}
+
+static inline int nfs_register_key_type(void)
+{
+ return 0;
+}
+
+static inline void nfs_unregister_key_type(void)
+{
+}
#endif /* CONFIG_KEYS */
static int nfs_net_init(struct net *net)
@@ -2740,10 +2825,14 @@ static int __init init_nfs_fs(void)
if (err < 0)
goto err_sysfs;
- err = nfsiod_start();
+ err = nfs_register_key_type();
if (err)
goto err_pernet;
+ err = nfsiod_start();
+ if (err)
+ goto err_keytype;
+
err = nfs_fs_proc_init();
if (err)
goto err_nfsiod;
@@ -2787,6 +2876,8 @@ static int __init init_nfs_fs(void)
nfs_fs_proc_exit();
err_nfsiod:
nfsiod_stop();
+err_keytype:
+ nfs_unregister_key_type();
err_pernet:
unregister_pernet_subsys(&nfs_net_ops);
err_sysfs:
@@ -2801,6 +2892,7 @@ static void __exit exit_nfs_fs(void)
nfs_destroy_readpagecache();
nfs_destroy_inodecache();
nfs_destroy_nfspagecache();
+ nfs_unregister_key_type();
unregister_pernet_subsys(&nfs_net_ops);
unregister_nfs_fs();
nfs_fs_proc_exit();
--
2.55.0
^ permalink raw reply related [flat|nested] 6+ messages in thread
end of thread, other threads:[~2026-09-25 19:16 UTC | newest]
Thread overview: 6+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-25 19:16 [PATCH RFC v2 0/5] NFS: isolate mTLS client credentials by network namespace Chuck Lever
2026-09-25 19:16 ` [PATCH RFC v2 1/5] NFS: name the init_nfs_fs() error labels Chuck Lever
2026-09-25 19:16 ` [PATCH RFC v2 2/5] NFS: allocate the .nfs keyring per network namespace Chuck Lever
2026-09-25 19:16 ` [PATCH RFC v2 3/5] SUNRPC: pass a keyring serial number to the TLS handshake Chuck Lever
2026-09-25 19:16 ` [PATCH RFC v2 4/5] NFS: name the namespace .nfs keyring in the x509 handshake Chuck Lever
2026-09-25 19:16 ` [PATCH RFC v2 5/5] NFS: add a key type that reveals the namespace .nfs keyring serial number Chuck Lever
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox