From: Christian Brauner <brauner@kernel.org>
To: linux-fsdevel@vger.kernel.org
Cc: Linus Torvalds <torvalds@linux-foundation.org>,
Chris Mason <mason@kernel.org>,
Alexander Viro <viro@zeniv.linux.org.uk>,
Jan Kara <jack@suse.cz>, Jeff Layton <jlayton@kernel.org>,
Aleksa Sarai <cyphar@cyphar.com>,
Amir Goldstein <amir73il@gmail.com>,
bpf@vger.kernel.org,
"Christian Brauner (Amutable)" <brauner@kernel.org>,
stable@vger.kernel.org
Subject: [PATCH 17/17] namespace: don't let a pseudo dentry become the root of a mount
Date: Wed, 30 Sep 2026 15:32:09 +0200 [thread overview]
Message-ID: <20260930-work-mount-fixes-3-v1-17-be34c83956ae@kernel.org> (raw)
In-Reply-To: <20260930-work-mount-fixes-3-v1-0-be34c83956ae@kernel.org>
d_alloc_pseudo() hands out dentries for pipes, sockets and other files
that are never anyone's child or parent. They carry DCACHE_NORCU and
dentry_free() frees them right away because no lockless path walk can
ever reach them. That holds as long as such a dentry isn't the root of
a mount. But bind mounting /proc/self/fd/<fd> of such a file does
exactly that.
Most callers of alloc_file_pseudo() put their files on kernel-internal
mounts and may_copy_tree() refuses those. bpf_token_create() doesn't. It
places the token file on the bpffs mount the caller handed it and that
mount is in the caller's mount namespace so the clone goes through:
mount --bind /proc/self/fd/<token> <file>
open_tree(tokfd, "", AT_EMPTY_PATH | OPEN_TREE_CLONE) + move_mount()
__follow_mount_rcu() then loads ->mnt_root of that mount, reads d_seq
and d_flags of the dentry and only then checks mount_lock. The dentry
can be gone by then. The final mntput() of an unmounted parent unhooks
a child that stayed attached to it under mount_lock alone. When the
child's holder does the final put right after that the root is dput()
and freed immediately while a walker that found the mount hashed still
looks at it.
Refuse to clone a mount with a DCACHE_NORCU dentry as its root. Nothing
sensible can be done with a bind mount of a bpf token anyway.
Fixes: 35f96de04127 ("bpf: Introduce BPF token object")
Cc: stable@vger.kernel.org # v6.9+
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
fs/namespace.c | 4 ++++
1 file changed, 4 insertions(+)
diff --git a/fs/namespace.c b/fs/namespace.c
index 3c90d853e091..fcf42f192aae 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -3026,6 +3026,10 @@ static struct mount *__do_loopback(const struct path *old_path,
if (!may_copy_tree(old_path))
return ERR_PTR(-EINVAL);
+ /* a pseudo dentry is freed without an RCU delay, no walk may find it */
+ if (old_path->dentry->d_flags & DCACHE_NORCU)
+ return ERR_PTR(-EINVAL);
+
if (recurse && !old->mnt_ns)
return ERR_PTR(-EINVAL);
--
2.53.0
prev parent reply other threads:[~2026-09-30 13:33 UTC|newest]
Thread overview: 22+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-30 13:31 [PATCH 00/17] mount: more bugfixes, the Oprah edition Christian Brauner
2026-09-30 13:31 ` [PATCH 01/17] namespace: queue a mount only once for mount notifications Christian Brauner
2026-09-30 13:31 ` [PATCH 02/17] namespace: check a submount for references right before unmounting it Christian Brauner
2026-09-30 13:31 ` [PATCH 03/17] selftests/filesystems: check that a busy submount survives a synchronous umount Christian Brauner
2026-09-30 13:31 ` [PATCH 04/17] namespace: check a recursive bind mount for mount namespace loops Christian Brauner
2026-09-30 13:31 ` [PATCH 05/17] selftests/filesystems: check that a recursive bind mount can't pin the caller's mount namespace Christian Brauner
2026-09-30 13:31 ` [PATCH 06/17] namespace: keep covered mounts covered in OPEN_TREE_NAMESPACE Christian Brauner
2026-09-30 13:31 ` [PATCH 07/17] selftests/filesystems: check that OPEN_TREE_NAMESPACE keeps mounts covered Christian Brauner
2026-09-30 13:32 ` [PATCH 08/17] namespace: look at the topmost mount for a mount namespace file Christian Brauner
2026-09-30 13:32 ` [PATCH 09/17] selftests/filesystems: check that a mount namespace file on top doesn't bury a mount Christian Brauner
2026-09-30 13:32 ` [PATCH 10/17] namespace: check the mounts before reading their parents in pivot_root() Christian Brauner
2026-09-30 13:32 ` [PATCH 11/17] namespace: don't reconfigure internal superblocks via remount and umount Christian Brauner
2026-09-30 13:32 ` [PATCH 12/17] selftests/filesystems: check that the nullfs root can't be reconfigured Christian Brauner
2026-09-30 13:32 ` [PATCH 13/17] namespace: remove the fsnotify marks of a mount namespace in process context Christian Brauner
2026-09-30 15:07 ` Amir Goldstein
2026-09-30 13:32 ` [PATCH 14/17] fsnotify: detach the connector before destroying its marks Christian Brauner
2026-10-01 9:31 ` Christian Brauner
2026-10-01 10:58 ` Amir Goldstein
2026-10-01 12:06 ` Christian Brauner
2026-09-30 13:32 ` [PATCH 15/17] dcache: don't put a mountpoint on a dentry that's being removed Christian Brauner
2026-09-30 13:32 ` [PATCH 16/17] unshare: don't drop active namespace references that were never taken Christian Brauner
2026-09-30 13:32 ` Christian Brauner [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260930-work-mount-fixes-3-v1-17-be34c83956ae@kernel.org \
--to=brauner@kernel.org \
--cc=amir73il@gmail.com \
--cc=bpf@vger.kernel.org \
--cc=cyphar@cyphar.com \
--cc=jack@suse.cz \
--cc=jlayton@kernel.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=mason@kernel.org \
--cc=stable@vger.kernel.org \
--cc=torvalds@linux-foundation.org \
--cc=viro@zeniv.linux.org.uk \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox