* [PATCH 00/17] mount: more bugfixes, the Oprah edition
@ 2026-09-30 13:31 Christian Brauner
2026-09-30 13:31 ` [PATCH 01/17] namespace: queue a mount only once for mount notifications Christian Brauner
` (16 more replies)
0 siblings, 17 replies; 22+ messages in thread
From: Christian Brauner @ 2026-09-30 13:31 UTC (permalink / raw)
To: linux-fsdevel
Cc: Linus Torvalds, Chris Mason, Alexander Viro, Jan Kara,
Jeff Layton, Aleksa Sarai, Amir Goldstein, bpf,
Christian Brauner (Amutable), stable
Yet more bugfixes falling out of my recent work in this area:
- don't let a pseudo dentry become the root of a mount
- don't drop active namespace references that were never taken
- don't put a mountpoint on a dentry that's being removed
- detach the fsnotify connector before destroying its marks
- remove the fsnotify marks of a mount namespace in process context
- don't reconfigure internal superblocks via remount and umount
- check the mounts before reading their parents in pivot_root()
- look at the topmost mount for a mount namespace file
- keep covered mounts covered in OPEN_TREE_NAMESPACE
- check a recursive bind mount for mount namespace loops
- check a submount for references right before unmounting it
- queue a mount only once for mount notifications
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
Christian Brauner (17):
namespace: queue a mount only once for mount notifications
namespace: check a submount for references right before unmounting it
selftests/filesystems: check that a busy submount survives a synchronous umount
namespace: check a recursive bind mount for mount namespace loops
selftests/filesystems: check that a recursive bind mount can't pin the caller's mount namespace
namespace: keep covered mounts covered in OPEN_TREE_NAMESPACE
selftests/filesystems: check that OPEN_TREE_NAMESPACE keeps mounts covered
namespace: look at the topmost mount for a mount namespace file
selftests/filesystems: check that a mount namespace file on top doesn't bury a mount
namespace: check the mounts before reading their parents in pivot_root()
namespace: don't reconfigure internal superblocks via remount and umount
selftests/filesystems: check that the nullfs root can't be reconfigured
namespace: remove the fsnotify marks of a mount namespace in process context
fsnotify: detach the connector before destroying its marks
dcache: don't put a mountpoint on a dentry that's being removed
unshare: don't drop active namespace references that were never taken
namespace: don't let a pseudo dentry become the root of a mount
fs/dcache.c | 4 +-
fs/fs_context.c | 4 +
fs/fsopen.c | 3 -
fs/mount.h | 3 +
fs/namespace.c | 133 +++++++++----
fs/notify/mark.c | 16 +-
include/linux/nsproxy.h | 1 +
kernel/fork.c | 3 +-
kernel/nsproxy.c | 2 +-
.../selftests/filesystems/empty_mntns/.gitignore | 1 +
.../selftests/filesystems/empty_mntns/Makefile | 2 +
.../empty_mntns/internal_sb_reconfigure_test.c | 108 +++++++++++
.../selftests/filesystems/mount_cycle/.gitignore | 2 +
.../selftests/filesystems/mount_cycle/Makefile | 1 +
.../filesystems/mount_cycle/nsfs_rbind_loop_test.c | 193 +++++++++++++++++++
.../mount_cycle/overmount_ns_file_test.c | 187 +++++++++++++++++++
.../selftests/filesystems/open_tree_ns/.gitignore | 1 +
.../selftests/filesystems/open_tree_ns/Makefile | 2 +-
.../open_tree_ns/open_tree_ns_covered_test.c | 183 ++++++++++++++++++
.../filesystems/umount_propagation/Makefile | 2 +-
.../umount_propagation/shrink_submounts_test.c | 205 +++++++++++++++++++++
21 files changed, 1008 insertions(+), 48 deletions(-)
---
base-commit: b4698e50d4601f43d21759203be5b0b3c123e383
change-id: 20260930-work-mount-fixes-3-47562ce9b80d
^ permalink raw reply [flat|nested] 22+ messages in thread
* [PATCH 01/17] namespace: queue a mount only once for mount notifications
2026-09-30 13:31 [PATCH 00/17] mount: more bugfixes, the Oprah edition Christian Brauner
@ 2026-09-30 13:31 ` Christian Brauner
2026-09-30 13:31 ` [PATCH 02/17] namespace: check a submount for references right before unmounting it Christian Brauner
` (15 subsequent siblings)
16 siblings, 0 replies; 22+ messages in thread
From: Christian Brauner @ 2026-09-30 13:31 UTC (permalink / raw)
To: linux-fsdevel
Cc: Linus Torvalds, Chris Mason, Alexander Viro, Jan Kara,
Jeff Layton, Aleksa Sarai, Amir Goldstein, bpf,
Christian Brauner (Amutable), stable
mnt_notify_add() puts a mount on notify_list. It doesn't check whether
the mount is on the list already. That's fine as long as every mount is
queued at most once per namespace_sem hold. It isn't.
propagate_umount() moves a surviving overmount off a stack of mounts
that are going away and queues it as moved. shrink_submounts() and
mark_mounts_for_expiry() call umount_tree() for several mounts under a
single namespace_sem hold. So the second umount_tree() can take down
exactly the mount the first one reparented, reparent it once more and so
end up queueing it a second time.
The second list_add_tail() cuts the mounts queued in between out of the
list while the head still points at the last of them. notify_mnt_list()
then only visits that one mount and the others are freed after the grace
period with notify_list still pointing at them. Every later mount
operation on the host walks freed memory:
BUG: KASAN: slab-use-after-free in __list_add_valid_or_report
Read of size 8 at addr ffff8881003d57e0 by task notify_dq/162
__list_add_valid_or_report+0x15c/0x1a0
mnt_add_to_ns+0x322/0x890
attach_recursive_mnt.isra.0+0xf2a/0x1a00
path_mount+0x139c/0x1d40
This needs a mount with MNT_SHRINKABLE to bind from. That's what
finish_automount() creates (submounts of NFS, AFS and CIFS and the
tracefs mount below debugfs). Bind mounts inherit the flag. A mount
namespace with a FAN_MARK_MNTNS mark does the rest and root in a user
namespace can have both.
Initialize to_notify in alloc_vfsmnt() and leave a mount alone that is
queued already. mnt_notify() looks at the state the mount has when the
list is drained so a single entry per mount is enough. The move event
for a mount that is unmounted under the same hold is lost. That's fine.
The detach is what matters.
Fixes: bf630c401641 ("vfs: add notifications for mount attach and detach")
Cc: stable@vger.kernel.org # v6.15+
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
fs/mount.h | 3 +++
fs/namespace.c | 3 +++
2 files changed, 6 insertions(+)
diff --git a/fs/mount.h b/fs/mount.h
index 85f136786bbc..2223fb141499 100644
--- a/fs/mount.h
+++ b/fs/mount.h
@@ -234,6 +234,9 @@ static inline struct mnt_namespace *to_mnt_ns(struct ns_common *ns)
#ifdef CONFIG_FSNOTIFY
static inline void mnt_notify_add(struct mount *m)
{
+ /* queued already under this namespace_sem hold */
+ if (!list_empty(&m->to_notify))
+ return;
/* Optimize the case where there are no watches */
if ((m->mnt_ns && m->mnt_ns->n_fsnotify_marks) ||
(m->prev_ns && m->prev_ns->n_fsnotify_marks))
diff --git a/fs/namespace.c b/fs/namespace.c
index 5b44eee28f39..2a1e77c0cb6e 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -333,6 +333,9 @@ static struct mount *alloc_vfsmnt(const char *name)
INIT_HLIST_NODE(&mnt->mnt_mp_list);
INIT_HLIST_HEAD(&mnt->mnt_stuck_children);
INIT_HLIST_NODE(&mnt->mnt_ns_visible);
+#ifdef CONFIG_FSNOTIFY
+ INIT_LIST_HEAD(&mnt->to_notify);
+#endif
RB_CLEAR_NODE(&mnt->mnt_node);
mnt->mnt.mnt_idmap = &nop_mnt_idmap;
}
--
2.53.0
^ permalink raw reply related [flat|nested] 22+ messages in thread
* [PATCH 02/17] namespace: check a submount for references right before unmounting it
2026-09-30 13:31 [PATCH 00/17] mount: more bugfixes, the Oprah edition Christian Brauner
2026-09-30 13:31 ` [PATCH 01/17] namespace: queue a mount only once for mount notifications Christian Brauner
@ 2026-09-30 13:31 ` Christian Brauner
2026-09-30 13:31 ` [PATCH 03/17] selftests/filesystems: check that a busy submount survives a synchronous umount Christian Brauner
` (14 subsequent siblings)
16 siblings, 0 replies; 22+ messages in thread
From: Christian Brauner @ 2026-09-30 13:31 UTC (permalink / raw)
To: linux-fsdevel
Cc: Linus Torvalds, Chris Mason, Alexander Viro, Jan Kara,
Jeff Layton, Aleksa Sarai, Amir Goldstein, bpf,
Christian Brauner (Amutable), stable
shrink_submounts() and mark_mounts_for_expiry() first collect all the
mounts they are allowed to unmount and then unmount them.
Whether a mount is busy is decided by propagate_mount_busy() on the
tree. But unmounting the first mount changes the tree that the second
one propagates into. propagate_umount()
moves a surviving overmount off a copy it unmounts and mounts it where
the copy was mounted. If that's where the propagated copy of the second
victim is looked up the overmount becomes a candidate of the second
umount_tree(). It's childless and so trim_one() commits it without
looking at its reference count.
So a synchronous umount pulls out a busy mount that is in use somewhere
else and marks it MNT_SYNC_UMOUNT while it has users:
umount2(/tmp/plshrink/p1, 0) = 0 errno 0
cwd is (unreachable)
The same two umounts requested one after the other fail with EBUSY.
This needs a mount with MNT_SHRINKABLE, i.e., automounted submounts of
NFS, AFS and CIFS or the tracefs mount below debugfs. Bind mounts
inherit the flag.
Check each mount right before it is unmounted under the same hold of
namespace_sem and mount_lock as the umount itself. Make sure that the
algorithm stays linear.
Fixes: 1064f874abc0 ("mnt: Tuck mounts under others instead of creating shadow/side mounts.")
Cc: stable@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
fs/namespace.c | 78 +++++++++++++++++++++++++++++++++++++++-------------------
1 file changed, 53 insertions(+), 25 deletions(-)
diff --git a/fs/namespace.c b/fs/namespace.c
index 2a1e77c0cb6e..23d3bfa9c14d 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -3987,6 +3987,11 @@ void mark_mounts_for_expiry(struct list_head *mounts)
}
while (!list_empty(&graveyard)) {
mnt = list_first_entry(&graveyard, struct mount, mnt_expire);
+ /* an earlier umount_tree() may have moved a busy mount here */
+ if (propagate_mount_busy(mnt, 1)) {
+ list_move(&mnt->mnt_expire, mounts);
+ continue;
+ }
touch_mnt_namespace(mnt->mnt_ns);
umount_tree(mnt, UMOUNT_PROPAGATE|UMOUNT_SYNC);
}
@@ -3994,17 +3999,37 @@ void mark_mounts_for_expiry(struct list_head *mounts)
EXPORT_SYMBOL_GPL(mark_mounts_for_expiry);
+/*
+ * Unmount @mnt if it's a shrinkable mount without children that nobody uses.
+ *
+ * mount_lock must be held for write
+ */
+static bool shrink_submount(struct mount *mnt)
+{
+ if (propagate_mount_busy(mnt, 1))
+ return false;
+ touch_mnt_namespace(mnt->mnt_ns);
+ umount_tree(mnt, UMOUNT_PROPAGATE|UMOUNT_SYNC);
+ return true;
+}
+
/*
* Ripoff of 'select_parent()'
*
- * search the list of submounts for a given mountpoint, and move any
- * shrinkable submounts to the 'graveyard' list.
+ * unmount the shrinkable submounts of @parent that aren't busy, children
+ * before their parent, and say whether anything went
+ *
+ * The cursor into the children of @this_parent survives the umount of a
+ * child mount without child mounts. The mounts that get umounted together with
+ * it are located under receiving mounts of @this_parent and never under
+ * @this_parent itself. The one exception is @this_parent getting unmounted
+ * then the walk starts over.
*/
-static int select_submounts(struct mount *parent, struct list_head *graveyard)
+static bool __shrink_submounts(struct mount *parent)
{
struct mount *this_parent = parent;
struct list_head *next;
- int found = 0;
+ bool shrunk = false;
repeat:
next = this_parent->mnt_mounts.next;
@@ -4023,42 +4048,45 @@ static int select_submounts(struct mount *parent, struct list_head *graveyard)
this_parent = mnt;
goto repeat;
}
-
- if (!propagate_mount_busy(mnt, 1)) {
- list_move_tail(&mnt->mnt_expire, graveyard);
- found++;
- }
+ if (!shrink_submount(mnt))
+ continue;
+ shrunk = true;
+ if (unlikely(this_parent->mnt.mnt_flags & MNT_UMOUNT))
+ return true;
}
/*
* All done at this level ... ascend and resume the search
*/
if (this_parent != parent) {
- next = this_parent->mnt_child.next;
- this_parent = this_parent->mnt_parent;
+ struct mount *mnt = this_parent;
+
+ next = mnt->mnt_child.next;
+ this_parent = mnt->mnt_parent;
+ /* its children are gone, maybe it can go as well */
+ if (shrink_submount(mnt)) {
+ shrunk = true;
+ if (unlikely(this_parent->mnt.mnt_flags & MNT_UMOUNT))
+ return true;
+ }
goto resume;
}
- return found;
+ return shrunk;
}
/*
- * process a list of expirable mountpoints with the intent of discarding any
- * submounts of a specific parent mountpoint
+ * unmount the shrinkable submounts of @mnt that aren't busy
+ *
+ * The busy check and the umount of a mount are adjacent. An umount can
+ * still empty or move a mount in a part of the tree that was walked
+ * already, so walk again until nothing goes.
*
* mount_lock must be held for write
*/
static void shrink_submounts(struct mount *mnt)
{
- LIST_HEAD(graveyard);
- struct mount *m;
-
- /* extract submounts of 'mountpoint' from the expiration list */
- while (select_submounts(mnt, &graveyard)) {
- while (!list_empty(&graveyard)) {
- m = list_first_entry(&graveyard, struct mount,
- mnt_expire);
- touch_mnt_namespace(m->mnt_ns);
- umount_tree(m, UMOUNT_PROPAGATE|UMOUNT_SYNC);
- }
+ for (;;) {
+ if (!__shrink_submounts(mnt))
+ break;
}
}
--
2.53.0
^ permalink raw reply related [flat|nested] 22+ messages in thread
* [PATCH 03/17] selftests/filesystems: check that a busy submount survives a synchronous umount
2026-09-30 13:31 [PATCH 00/17] mount: more bugfixes, the Oprah edition Christian Brauner
2026-09-30 13:31 ` [PATCH 01/17] namespace: queue a mount only once for mount notifications Christian Brauner
2026-09-30 13:31 ` [PATCH 02/17] namespace: check a submount for references right before unmounting it Christian Brauner
@ 2026-09-30 13:31 ` Christian Brauner
2026-09-30 13:31 ` [PATCH 04/17] namespace: check a recursive bind mount for mount namespace loops Christian Brauner
` (13 subsequent siblings)
16 siblings, 0 replies; 22+ messages in thread
From: Christian Brauner @ 2026-09-30 13:31 UTC (permalink / raw)
To: linux-fsdevel
Cc: Linus Torvalds, Chris Mason, Alexander Viro, Jan Kara,
Jeff Layton, Aleksa Sarai, Amir Goldstein, bpf,
Christian Brauner (Amutable)
Add a test for the shrinkable submounts of a synchronous umount:
- P shared, P1 a slave that is shared in turn, P2 its peer
- B on P/options with copies on P1 and P2, R on top of the copy in P2
- P with B moved below P1, R is the working directory
- umount(P1) slides R to where the copy of B is looked up
The umount fails with EBUSY and R stays mounted. Needs the tracefs
automount below debugfs for the shrinkable mounts.
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
.../filesystems/umount_propagation/Makefile | 2 +-
.../umount_propagation/shrink_submounts_test.c | 205 +++++++++++++++++++++
2 files changed, 206 insertions(+), 1 deletion(-)
diff --git a/tools/testing/selftests/filesystems/umount_propagation/Makefile b/tools/testing/selftests/filesystems/umount_propagation/Makefile
index fc0a0783018b..eb85612abf8d 100644
--- a/tools/testing/selftests/filesystems/umount_propagation/Makefile
+++ b/tools/testing/selftests/filesystems/umount_propagation/Makefile
@@ -1,5 +1,5 @@
# SPDX-License-Identifier: GPL-2.0
-TEST_GEN_PROGS := umount_propagation_test
+TEST_GEN_PROGS := umount_propagation_test shrink_submounts_test
CFLAGS += -Wall -O2 -g $(KHDR_INCLUDES)
diff --git a/tools/testing/selftests/filesystems/umount_propagation/shrink_submounts_test.c b/tools/testing/selftests/filesystems/umount_propagation/shrink_submounts_test.c
new file mode 100644
index 000000000000..43efb7faf95c
--- /dev/null
+++ b/tools/testing/selftests/filesystems/umount_propagation/shrink_submounts_test.c
@@ -0,0 +1,205 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * A synchronous umount first unmounts the shrinkable submounts of the
+ * victim that aren't busy. Every one of them has to be checked right
+ * before it is unmounted: unmounting one can slide a busy mount to where
+ * the propagated copy of the next one is looked up.
+ */
+#define _GNU_SOURCE
+#include <errno.h>
+#include <fcntl.h>
+#include <sched.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <unistd.h>
+#include <sys/fanotify.h>
+#include <sys/mount.h>
+#include <sys/stat.h>
+#include <sys/statfs.h>
+#include <sys/vfs.h>
+#include <linux/magic.h>
+
+#include "../../kselftest_harness.h"
+
+#ifndef FAN_REPORT_MNT
+#define FAN_REPORT_MNT 0x00004000
+#endif
+#ifndef FAN_MARK_MNTNS
+#define FAN_MARK_MNTNS 0x00000110
+#endif
+#ifndef FAN_MNT_ATTACH
+#define FAN_MNT_ATTACH 0x01000000
+#endif
+#ifndef FAN_MNT_DETACH
+#define FAN_MNT_DETACH 0x02000000
+#endif
+
+#define DIR_LEN 64
+#define PATH_LEN 128
+
+FIXTURE(shrink_submounts) {
+ char base[DIR_LEN];
+ char automount[PATH_LEN];
+ bool mounted;
+ int fan;
+};
+
+/*
+ * Shrinkable mounts come from an automount. The tracefs mount below debugfs
+ * is one and bind mounts inherit the flag.
+ */
+FIXTURE_SETUP(shrink_submounts)
+{
+ struct stat st;
+ char p[PATH_LEN];
+
+ self->mounted = false;
+ self->fan = -1;
+
+ if (geteuid() != 0)
+ SKIP(return, "test requires CAP_SYS_ADMIN");
+
+ ASSERT_EQ(unshare(CLONE_NEWNS), 0);
+ ASSERT_EQ(mount("", "/", NULL, MS_REC | MS_PRIVATE, NULL), 0);
+
+ snprintf(self->base, sizeof(self->base), "/tmp/shrink_submounts.XXXXXX");
+ ASSERT_NE(mkdtemp(self->base), NULL);
+ ASSERT_EQ(mount("tmpfs", self->base, "tmpfs", 0, NULL), 0);
+ self->mounted = true;
+ ASSERT_EQ(mount(NULL, self->base, NULL, MS_PRIVATE, NULL), 0);
+
+ snprintf(p, sizeof(p), "%s/dbg", self->base);
+ ASSERT_EQ(mkdir(p, 0755), 0);
+ if (mount("debugfs", p, "debugfs", 0, NULL))
+ SKIP(return, "test requires debugfs");
+ snprintf(self->automount, sizeof(self->automount), "%s/dbg/tracing",
+ self->base);
+ snprintf(p, sizeof(p), "%s/dbg/tracing/.", self->base);
+ if (stat(p, &st))
+ SKIP(return, "test requires the tracefs automount");
+}
+
+FIXTURE_TEARDOWN(shrink_submounts)
+{
+ if (self->fan >= 0)
+ close(self->fan);
+ chdir("/");
+ if (self->mounted)
+ umount2(self->base, MNT_DETACH);
+ rmdir(self->base);
+}
+
+static bool mounted_tmpfs(const char *path)
+{
+ struct statfs st;
+
+ return !statfs(path, &st) && st.f_type == TMPFS_MAGIC;
+}
+
+/*
+ * P is a shared bind mount of the automount, P1 a slave of P that is shared
+ * in turn and P2 its peer. B, another bind mount of the automount, goes on
+ * P/options and propagates copies Bc1 and Bc2 onto P1/options and
+ * P2/options. R, a tmpfs and our working directory, sits on top of Bc2 which
+ * is made private first. P, with B on it, is moved to P1/instances.
+ *
+ * A synchronous umount of P1 unmounts the shrinkable submounts Bc1 and B
+ * first. Unmounting Bc1 takes Bc2 along and slides R to P2/options where
+ * the propagated copy of B is looked up next. R is busy, so B has to stay
+ * and the umount fails with EBUSY.
+ */
+TEST_F(shrink_submounts, busy_mount_moved_into_reach)
+{
+ char p[PATH_LEN], p1[PATH_LEN], p2[PATH_LEN], r[PATH_LEN], cwd[PATH_LEN];
+ int nsfd;
+
+ snprintf(p, sizeof(p), "%s/p", self->base);
+ snprintf(p1, sizeof(p1), "%s/p1", self->base);
+ snprintf(p2, sizeof(p2), "%s/p2", self->base);
+ ASSERT_EQ(mkdir(p, 0755), 0);
+ ASSERT_EQ(mkdir(p1, 0755), 0);
+ ASSERT_EQ(mkdir(p2, 0755), 0);
+
+ /* watch the mount namespace so that the detached mounts get queued */
+ self->fan = fanotify_init(FAN_REPORT_MNT, O_RDONLY);
+ if (self->fan >= 0) {
+ nsfd = open("/proc/self/ns/mnt", O_RDONLY | O_CLOEXEC);
+ ASSERT_GE(nsfd, 0);
+ EXPECT_EQ(fanotify_mark(self->fan, FAN_MARK_ADD | FAN_MARK_MNTNS,
+ FAN_MNT_ATTACH | FAN_MNT_DETACH, nsfd, NULL), 0);
+ close(nsfd);
+ }
+
+ ASSERT_EQ(mount(self->automount, p, NULL, MS_BIND, NULL), 0);
+ ASSERT_EQ(mount(NULL, p, NULL, MS_SHARED, NULL), 0);
+ ASSERT_EQ(mount(p, p1, NULL, MS_BIND, NULL), 0);
+ ASSERT_EQ(mount(NULL, p1, NULL, MS_SLAVE, NULL), 0);
+ ASSERT_EQ(mount(NULL, p1, NULL, MS_SHARED, NULL), 0);
+ ASSERT_EQ(mount(p1, p2, NULL, MS_BIND, NULL), 0);
+
+ /* B on P/options, copies on P1/options and P2/options */
+ snprintf(r, sizeof(r), "%s/p/options", self->base);
+ ASSERT_EQ(mount(self->automount, r, NULL, MS_BIND, NULL), 0);
+
+ /* R on top of Bc2 */
+ snprintf(r, sizeof(r), "%s/p2/options", self->base);
+ ASSERT_EQ(mount(NULL, r, NULL, MS_PRIVATE, NULL), 0);
+ ASSERT_EQ(mount("R", r, "tmpfs", 0, NULL), 0);
+ ASSERT_EQ(chdir(r), 0);
+
+ /* P, with B on it, below the victim */
+ snprintf(cwd, sizeof(cwd), "%s/p1/instances", self->base);
+ ASSERT_EQ(mount(p, cwd, NULL, MS_MOVE, NULL), 0);
+
+ ASSERT_TRUE(mounted_tmpfs(r));
+ ASSERT_EQ(umount2(p1, 0), -1);
+ EXPECT_EQ(errno, EBUSY);
+
+ /* R is still mounted and still our working directory */
+ EXPECT_TRUE(mounted_tmpfs(r));
+ ASSERT_NE(getcwd(cwd, sizeof(cwd)), NULL);
+ EXPECT_STREQ(cwd, r);
+}
+
+/*
+ * T is shared and Q, a slave of T, has T moved into it, so Q receives
+ * propagation from its own child. M, another bind mount of the automount, is
+ * on T/options and that is the dentry T sits on in Q. Unmounting M makes T
+ * the propagated victim at that dentry in Q and takes T along. The shrink
+ * walk of V has just unmounted M and continues in the children of T.
+ */
+TEST_F(shrink_submounts, parent_goes_with_child)
+{
+ char v[PATH_LEN], t[PATH_LEN], q[PATH_LEN], p[PATH_LEN];
+ struct stat before, after;
+
+ snprintf(v, sizeof(v), "%s/v", self->base);
+ snprintf(t, sizeof(t), "%s/v/t", self->base);
+ snprintf(q, sizeof(q), "%s/v/q", self->base);
+ ASSERT_EQ(mkdir(v, 0755), 0);
+ ASSERT_EQ(mount("V", v, "tmpfs", 0, NULL), 0);
+ ASSERT_EQ(mount(NULL, v, NULL, MS_PRIVATE, NULL), 0);
+ ASSERT_EQ(stat(v, &before), 0);
+ ASSERT_EQ(mkdir(t, 0755), 0);
+ ASSERT_EQ(mkdir(q, 0755), 0);
+
+ /* T shared, M on T/options before anything receives from T */
+ ASSERT_EQ(mount(self->automount, t, NULL, MS_BIND, NULL), 0);
+ ASSERT_EQ(mount(NULL, t, NULL, MS_SHARED, NULL), 0);
+ snprintf(p, sizeof(p), "%s/v/t/options", self->base);
+ ASSERT_EQ(mount(self->automount, p, NULL, MS_BIND, NULL), 0);
+
+ /* Q, a slave of T, and T moved into Q at the dentry M sits on */
+ ASSERT_EQ(mount(t, q, NULL, MS_BIND, NULL), 0);
+ ASSERT_EQ(mount(NULL, q, NULL, MS_SLAVE, NULL), 0);
+ snprintf(p, sizeof(p), "%s/v/q/options", self->base);
+ ASSERT_EQ(mount(t, p, NULL, MS_MOVE, NULL), 0);
+
+ /* M, T and then Q go, V is empty and can be unmounted */
+ ASSERT_EQ(umount2(v, 0), 0);
+ ASSERT_EQ(stat(v, &after), 0);
+ EXPECT_NE(before.st_dev, after.st_dev);
+}
+
+TEST_HARNESS_MAIN
--
2.53.0
^ permalink raw reply related [flat|nested] 22+ messages in thread
* [PATCH 04/17] namespace: check a recursive bind mount for mount namespace loops
2026-09-30 13:31 [PATCH 00/17] mount: more bugfixes, the Oprah edition Christian Brauner
` (2 preceding siblings ...)
2026-09-30 13:31 ` [PATCH 03/17] selftests/filesystems: check that a busy submount survives a synchronous umount Christian Brauner
@ 2026-09-30 13:31 ` Christian Brauner
2026-09-30 13:31 ` [PATCH 05/17] selftests/filesystems: check that a recursive bind mount can't pin the caller's mount namespace Christian Brauner
` (12 subsequent siblings)
16 siblings, 0 replies; 22+ messages in thread
From: Christian Brauner @ 2026-09-30 13:31 UTC (permalink / raw)
To: linux-fsdevel
Cc: Linus Torvalds, Chris Mason, Alexander Viro, Jan Kara,
Jeff Layton, Aleksa Sarai, Amir Goldstein, bpf,
Christian Brauner (Amutable), stable
do_loopback() refuses to bind mount the file of a mount namespace that
is as old as the caller's or older because a mount namespace that holds
a mount of its own file, or of an ancestor's, would cause a cycle. But
it only checks the dentry the bind mount starts from. With MS_REC
everything below it is copied and do_loopback() passes
CL_COPY_MNT_NS_FILE so mount namespace files below the source are
copied.
That's fine for a source in the caller's own mount namespace. Every
mount namespace file in there passed the same check when it was
bind-mounted. It isn't fine for a source in another mount namespace.
may_copy_tree() accepts a bind mount of any nsfs or pidfs file no matter
what mount namespace it lives in so that /proc/<pid>/ns/<ns> and pidfds
can be bind mounted. If we create a bind-mount stack of mount namespaces
file descriptors the kernel will copy them irrespective of their
ancestoral relationship to the mount namespace in question:
151 146 0:7 net:[4026531833] /tmp/nrc/y rw - nsfs nsfs rw
152 151 0:7 mnt:[4026532293] /tmp/nrc/y rw - nsfs nsfs rw
Mount 152 is a mount of the mount namespace file of mount namespace
4026532293 inside mount namespace 4026532293. That namespace and every
mount in it are leaked... All it takes is a process in an older mount
namespace that stacks the file and repeating it leaks without limit:
after control: Shmem: 380 kB
child: mount(/proc/self/fd/6, MS_BIND|MS_REC): ok
after cycle: Shmem: 65916 kB
do_move_mount() runs check_for_nsfs_mounts() over a detached tree for
exactly that reason. Do the same for the copy before it is grafted.
Fixes: e149ed2b805f ("take the targets of /proc/*/ns/* symlinks to separate fs")
Fixes: ef4144ac2dec ("pidfs: allow bind-mounts")
Cc: stable@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
fs/namespace.c | 6 +++++-
1 file changed, 5 insertions(+), 1 deletion(-)
diff --git a/fs/namespace.c b/fs/namespace.c
index 23d3bfa9c14d..c2f54636ec9d 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -3055,7 +3055,11 @@ static int do_loopback(const struct path *path, const char *old_name,
if (IS_ERR(mnt))
return PTR_ERR(mnt);
- err = graft_tree(mnt, &mp);
+ /* the copy may carry mount namespace files from below the source */
+ if (recurse && !check_for_nsfs_mounts(mnt))
+ err = -EINVAL;
+ else
+ err = graft_tree(mnt, &mp);
if (err) {
lock_mount_hash();
umount_tree(mnt, UMOUNT_SYNC);
--
2.53.0
^ permalink raw reply related [flat|nested] 22+ messages in thread
* [PATCH 05/17] selftests/filesystems: check that a recursive bind mount can't pin the caller's mount namespace
2026-09-30 13:31 [PATCH 00/17] mount: more bugfixes, the Oprah edition Christian Brauner
` (3 preceding siblings ...)
2026-09-30 13:31 ` [PATCH 04/17] namespace: check a recursive bind mount for mount namespace loops Christian Brauner
@ 2026-09-30 13:31 ` Christian Brauner
2026-09-30 13:31 ` [PATCH 06/17] namespace: keep covered mounts covered in OPEN_TREE_NAMESPACE Christian Brauner
` (11 subsequent siblings)
16 siblings, 0 replies; 22+ messages in thread
From: Christian Brauner @ 2026-09-30 13:31 UTC (permalink / raw)
To: linux-fsdevel
Cc: Linus Torvalds, Chris Mason, Alexander Viro, Jan Kara,
Jeff Layton, Aleksa Sarai, Amir Goldstein, bpf,
Christian Brauner (Amutable)
Add a test for a recursive bind mount of a namespace file in another mount
namespace:
- a bind mount of the network namespace file, held through a descriptor
- the caller's own mount namespace file stacked on top of it there
- the recursive bind mount through the descriptor fails with EINVAL
- a plain bind mount of the file still works
The copy would pin the namespace it is put in otherwise.
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
.../selftests/filesystems/mount_cycle/.gitignore | 1 +
.../selftests/filesystems/mount_cycle/Makefile | 1 +
.../filesystems/mount_cycle/nsfs_rbind_loop_test.c | 193 +++++++++++++++++++++
3 files changed, 195 insertions(+)
diff --git a/tools/testing/selftests/filesystems/mount_cycle/.gitignore b/tools/testing/selftests/filesystems/mount_cycle/.gitignore
index 03ca95de7765..8cd722977a32 100644
--- a/tools/testing/selftests/filesystems/mount_cycle/.gitignore
+++ b/tools/testing/selftests/filesystems/mount_cycle/.gitignore
@@ -1,3 +1,4 @@
# SPDX-License-Identifier: GPL-2.0-only
unmounted_tree_test
overmount_reparent_test
+nsfs_rbind_loop_test
diff --git a/tools/testing/selftests/filesystems/mount_cycle/Makefile b/tools/testing/selftests/filesystems/mount_cycle/Makefile
index 88271c82d1c1..32e26336b132 100644
--- a/tools/testing/selftests/filesystems/mount_cycle/Makefile
+++ b/tools/testing/selftests/filesystems/mount_cycle/Makefile
@@ -1,5 +1,6 @@
# SPDX-License-Identifier: GPL-2.0
TEST_GEN_PROGS := unmounted_tree_test overmount_reparent_test
+TEST_GEN_PROGS += nsfs_rbind_loop_test
CFLAGS += -Wall -O2 -g $(KHDR_INCLUDES)
diff --git a/tools/testing/selftests/filesystems/mount_cycle/nsfs_rbind_loop_test.c b/tools/testing/selftests/filesystems/mount_cycle/nsfs_rbind_loop_test.c
new file mode 100644
index 000000000000..0928a584eddc
--- /dev/null
+++ b/tools/testing/selftests/filesystems/mount_cycle/nsfs_rbind_loop_test.c
@@ -0,0 +1,193 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * A recursive bind mount of a namespace file that lives in another mount
+ * namespace copies whatever is stacked on top of it there. If that includes
+ * the file of the caller's own mount namespace, or of an older one, the copy
+ * would pin the namespace it is put in forever. The bind mount has to be
+ * refused, a plain bind mount of the file itself still works.
+ */
+#define _GNU_SOURCE
+#include <errno.h>
+#include <fcntl.h>
+#include <sched.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <unistd.h>
+#include <sys/mount.h>
+#include <sys/stat.h>
+#include <sys/wait.h>
+
+#include "../../kselftest_harness.h"
+
+#define DIR_LEN 64
+#define PATH_LEN 128
+
+/* Child exit codes. */
+enum {
+ CHILD_OK,
+ CHILD_UNSHARE, /* could not create the newer mount namespace */
+ CHILD_PIPE, /* the parent went away */
+ CHILD_TMPFS, /* could not mount the tmpfs in the new namespace */
+ CHILD_REC_ALLOWED, /* the recursive bind mount was not refused */
+ CHILD_REC_ERRNO, /* it was refused with the wrong error */
+ CHILD_PLAIN_REFUSED, /* the plain bind mount of the file was refused */
+};
+
+static int write_file(const char *path, const char *s)
+{
+ ssize_t n = -1;
+ int fd;
+
+ fd = open(path, O_WRONLY | O_CLOEXEC);
+ if (fd >= 0) {
+ n = write(fd, s, strlen(s));
+ close(fd);
+ }
+ return n == (ssize_t)strlen(s) ? 0 : -1;
+}
+
+static int create_file(const char *path)
+{
+ int fd = open(path, O_WRONLY | O_CREAT | O_CLOEXEC, 0644);
+
+ if (fd < 0)
+ return -1;
+ close(fd);
+ return 0;
+}
+
+/* Become root in a new user namespace with a private mount namespace. */
+static int enter_userns(void)
+{
+ uid_t uid = getuid();
+ gid_t gid = getgid();
+ char map[32];
+
+ if (unshare(CLONE_NEWUSER | CLONE_NEWNS))
+ return -1;
+ if (write_file("/proc/self/setgroups", "deny") && errno != ENOENT)
+ return -1;
+ snprintf(map, sizeof(map), "0 %d 1", uid);
+ if (write_file("/proc/self/uid_map", map))
+ return -1;
+ snprintf(map, sizeof(map), "0 %d 1", gid);
+ if (write_file("/proc/self/gid_map", map))
+ return -1;
+ if (setgid(0) || setuid(0))
+ return -1;
+ return mount("", "/", NULL, MS_REC | MS_PRIVATE, NULL);
+}
+
+static int send_msg(int fd, char msg)
+{
+ return write(fd, &msg, 1) == 1 ? 0 : -1;
+}
+
+static char recv_msg(int fd)
+{
+ char msg;
+
+ if (read(fd, &msg, 1) != 1)
+ return 0;
+ return msg;
+}
+
+FIXTURE(nsfs_rbind_loop) {
+ char dir[DIR_LEN];
+ char x[PATH_LEN];
+};
+
+FIXTURE_SETUP(nsfs_rbind_loop)
+{
+ snprintf(self->dir, sizeof(self->dir), "/tmp/nsfs_rbind_loop.XXXXXX");
+ ASSERT_NE(mkdtemp(self->dir), NULL);
+ if (enter_userns()) {
+ rmdir(self->dir);
+ SKIP(return, "test requires user namespaces");
+ }
+ ASSERT_EQ(mount("tmpfs", self->dir, "tmpfs", 0, NULL), 0);
+ snprintf(self->x, sizeof(self->x), "%s/x", self->dir);
+ ASSERT_EQ(create_file(self->x), 0);
+}
+
+FIXTURE_TEARDOWN(nsfs_rbind_loop)
+{
+ umount2(self->dir, MNT_DETACH);
+ rmdir(self->dir);
+}
+
+/*
+ * The child in the newer mount namespace binds the network namespace file
+ * mount of the parent through @fd. Recursively that would copy the mount of
+ * its own mount namespace file that the parent stacked on top.
+ */
+static int newer_ns_child(const char *dir, int fd, int to_parent, int from_parent)
+{
+ char src[32], y[PATH_LEN];
+
+ if (unshare(CLONE_NEWNS))
+ return CHILD_UNSHARE;
+ if (send_msg(to_parent, 'r') || recv_msg(from_parent) != 'g')
+ return CHILD_PIPE;
+
+ snprintf(y, sizeof(y), "%s/y", dir);
+ if (mount("tmpfs", y, "tmpfs", 0, NULL))
+ return CHILD_TMPFS;
+ snprintf(src, sizeof(src), "/proc/self/fd/%d", fd);
+ snprintf(y, sizeof(y), "%s/y/f", dir);
+ if (create_file(y))
+ return CHILD_TMPFS;
+
+ if (!mount(src, y, NULL, MS_BIND | MS_REC, NULL))
+ return CHILD_REC_ALLOWED;
+ if (errno != EINVAL)
+ return CHILD_REC_ERRNO;
+ if (mount(src, y, NULL, MS_BIND, NULL))
+ return CHILD_PLAIN_REFUSED;
+ umount2(y, MNT_DETACH);
+ return CHILD_OK;
+}
+
+TEST_F(nsfs_rbind_loop, own_ns_file_below_foreign_source)
+{
+ int to_child[2], to_parent[2], fd, status;
+ char p[PATH_LEN];
+ pid_t pid;
+
+ snprintf(p, sizeof(p), "%s/y", self->dir);
+ ASSERT_EQ(mkdir(p, 0755), 0);
+
+ /* M, a mount of our network namespace file, held by a descriptor */
+ ASSERT_EQ(mount("/proc/self/ns/net", self->x, NULL, MS_BIND, NULL), 0);
+ fd = open(self->x, O_PATH | O_CLOEXEC);
+ ASSERT_GE(fd, 0);
+
+ ASSERT_EQ(pipe(to_child), 0);
+ ASSERT_EQ(pipe(to_parent), 0);
+ pid = fork();
+ ASSERT_GE(pid, 0);
+ if (pid == 0) {
+ close(to_child[1]);
+ close(to_parent[0]);
+ _exit(newer_ns_child(self->dir, fd, to_parent[1], to_child[0]));
+ }
+ close(to_child[0]);
+ close(to_parent[1]);
+ ASSERT_EQ(recv_msg(to_parent[0]), 'r');
+
+ /* the child's mount namespace file on top of M */
+ snprintf(p, sizeof(p), "/proc/%d/ns/mnt", pid);
+ ASSERT_EQ(mount(p, self->x, NULL, MS_BIND, NULL), 0);
+
+ ASSERT_EQ(send_msg(to_child[1], 'g'), 0);
+ ASSERT_EQ(waitpid(pid, &status, 0), pid);
+ ASSERT_TRUE(WIFEXITED(status));
+ ASSERT_EQ(WEXITSTATUS(status), CHILD_OK);
+
+ close(fd);
+ ASSERT_EQ(umount2(self->x, MNT_DETACH), 0);
+ ASSERT_EQ(umount2(self->x, MNT_DETACH), 0);
+}
+
+TEST_HARNESS_MAIN
--
2.53.0
^ permalink raw reply related [flat|nested] 22+ messages in thread
* [PATCH 06/17] namespace: keep covered mounts covered in OPEN_TREE_NAMESPACE
2026-09-30 13:31 [PATCH 00/17] mount: more bugfixes, the Oprah edition Christian Brauner
` (4 preceding siblings ...)
2026-09-30 13:31 ` [PATCH 05/17] selftests/filesystems: check that a recursive bind mount can't pin the caller's mount namespace Christian Brauner
@ 2026-09-30 13:31 ` Christian Brauner
2026-09-30 13:31 ` [PATCH 07/17] selftests/filesystems: check that OPEN_TREE_NAMESPACE keeps mounts covered Christian Brauner
` (10 subsequent siblings)
16 siblings, 0 replies; 22+ messages in thread
From: Christian Brauner @ 2026-09-30 13:31 UTC (permalink / raw)
To: linux-fsdevel
Cc: Linus Torvalds, Chris Mason, Alexander Viro, Jan Kara,
Jeff Layton, Aleksa Sarai, Amir Goldstein, bpf,
Christian Brauner (Amutable), stable
open_tree(OPEN_TREE_NAMESPACE) creates a new mount namespace from a
copy of the tree at the given path. The caller needs to be privileged
over its current user namespace because the new mount namespace will be
owned by it. It doesn't need to be privileged over the mount namespace
the tree is copied from. An unprivileged user just needs to create a
user namespace first.
The copy follows the rules of a detached bind mount. Without
AT_RECURSIVE only the mount itself is cloned and its children are left
out. With AT_RECURSIVE unbindable mounts are skipped. Both is fine when
the caller has privileges over the source mount namespace because it
could unmount those mounts anyway. Not so for an unprivileged user in
that mount namespace. For them a mount namespace copy is what unshare()
does. copy_mnt_ns() copies everything including unbindable mounts and
lock_mnt_tree() makes sure nothing can be unmounted in the copy. Nothing
that was covered gets revealed.
create_new_namespace() only does the locking and so the covered content
is right there in the new mount namespace:
# mount -t tmpfs none /tmp/otn; mkdir /tmp/otn/covered
# echo hidden > /tmp/otn/covered/under.txt
# mount -t tmpfs none /tmp/otn/covered
uid 1000, after unshare(CLONE_NEWUSER):
fd = open_tree(AT_FDCWD, "/tmp/otn", OPEN_TREE_NAMESPACE);
setns(fd, CLONE_NEWNS);
open("/covered/under.txt", O_RDONLY) -> "hidden"
Covering paths with mounts is how container runtimes mask parts of
/proc and /sys and how admins hide things.
When the caller's user namespace doesn't own the source mount namespace
copy the way unshare() does and refuse a non-recursive copy of a mount
that has anything mounted below the requested directory and copy
unbindable mounts in a recursive copy. A caller that is privileged over
the source mount namespace sees no change.
Fixes: 9b8a0ba68246 ("mount: add OPEN_TREE_NAMESPACE")
Cc: stable@vger.kernel.org # v7.0+
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
fs/namespace.c | 30 ++++++++++++++++++++++++++----
1 file changed, 26 insertions(+), 4 deletions(-)
diff --git a/fs/namespace.c b/fs/namespace.c
index c2f54636ec9d..de3900dd02f6 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -2369,6 +2369,18 @@ bool has_locked_children(struct mount *mnt, struct dentry *dentry)
return __has_locked_children(mnt, dentry);
}
+/* locks: namespace_shared && pinned(mnt) || mount_locked_reader */
+static bool __has_children(struct mount *mnt, struct dentry *dentry)
+{
+ struct mount *child;
+
+ list_for_each_entry(child, &mnt->mnt_mounts, mnt_child) {
+ if (is_subdir(child->mnt_mountpoint, dentry))
+ return true;
+ }
+ return false;
+}
+
/*
* Check that there aren't references to earlier/same mount namespaces in the
* specified subtree. Such references can act as pins for mount namespaces
@@ -3141,12 +3153,18 @@ static struct mnt_namespace *create_new_namespace(struct path *path,
struct mount *mnt;
unsigned int copy_flags = 0;
bool locked = false, recurse = flags & MOUNT_COPY_RECURSIVE;
+ bool foreign = user_ns != ns->user_ns;
if (unlikely(!d_can_lookup(path->dentry)))
return ERR_PTR(-ENOTDIR);
- if (user_ns != ns->user_ns)
- copy_flags |= CL_SLAVE;
+ /*
+ * Without privileges over the mount namespace the copy is made from
+ * nothing mounted below @path may be left out. It would reveal what
+ * it covers. That's what unshare() gives such a caller as well.
+ */
+ if (foreign)
+ copy_flags |= CL_SLAVE | CL_COPY_UNBINDABLE;
new_ns = alloc_mnt_ns(user_ns, false);
if (IS_ERR(new_ns))
@@ -3180,10 +3198,14 @@ static struct mnt_namespace *create_new_namespace(struct path *path,
/*
* We don't emulate unshare()ing a mount namespace. We stick to
* the restrictions of creating detached bind-mounts. It has a
- * lot saner and simpler semantics.
+ * lot saner and simpler semantics. A caller without privileges
+ * over the mount namespace can't leave out any child though.
*/
if (flags & MOUNT_COPY_NEW)
mnt = clone_mnt(real_mount(path->mnt), path->dentry, copy_flags);
+ else if (foreign && !recurse &&
+ __has_children(real_mount(path->mnt), path->dentry))
+ mnt = ERR_PTR(-EINVAL);
else
mnt = __do_loopback(path, recurse, copy_flags);
scoped_guard(mount_writer) {
@@ -3200,7 +3222,7 @@ static struct mnt_namespace *create_new_namespace(struct path *path,
* of the real rootfs we created.
*/
attach_mnt(mnt, new_ns_root, mp.mp);
- if (user_ns != ns->user_ns)
+ if (foreign)
lock_mnt_tree(new_ns_root);
}
--
2.53.0
^ permalink raw reply related [flat|nested] 22+ messages in thread
* [PATCH 07/17] selftests/filesystems: check that OPEN_TREE_NAMESPACE keeps mounts covered
2026-09-30 13:31 [PATCH 00/17] mount: more bugfixes, the Oprah edition Christian Brauner
` (5 preceding siblings ...)
2026-09-30 13:31 ` [PATCH 06/17] namespace: keep covered mounts covered in OPEN_TREE_NAMESPACE Christian Brauner
@ 2026-09-30 13:31 ` Christian Brauner
2026-09-30 13:32 ` [PATCH 08/17] namespace: look at the topmost mount for a mount namespace file Christian Brauner
` (9 subsequent siblings)
16 siblings, 0 replies; 22+ messages in thread
From: Christian Brauner @ 2026-09-30 13:31 UTC (permalink / raw)
To: linux-fsdevel
Cc: Linus Torvalds, Chris Mason, Alexander Viro, Jan Kara,
Jeff Layton, Aleksa Sarai, Amir Goldstein, bpf,
Christian Brauner (Amutable)
Add a test for open_tree(OPEN_TREE_NAMESPACE) from a user namespace that
doesn't own the mount namespace it copies from:
- a file covered by a private or an unbindable tmpfs mount
- the non-recursive copy of the parent mount fails with EINVAL
- the recursive copy keeps the file covered in the new mount namespace
- the owner of the mount namespace keeps the bind mount semantics
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
.../selftests/filesystems/open_tree_ns/.gitignore | 1 +
.../selftests/filesystems/open_tree_ns/Makefile | 2 +-
.../open_tree_ns/open_tree_ns_covered_test.c | 183 +++++++++++++++++++++
3 files changed, 185 insertions(+), 1 deletion(-)
diff --git a/tools/testing/selftests/filesystems/open_tree_ns/.gitignore b/tools/testing/selftests/filesystems/open_tree_ns/.gitignore
index fb12b93fbcaa..76f95c0ae5ef 100644
--- a/tools/testing/selftests/filesystems/open_tree_ns/.gitignore
+++ b/tools/testing/selftests/filesystems/open_tree_ns/.gitignore
@@ -1 +1,2 @@
open_tree_ns_test
+open_tree_ns_covered_test
diff --git a/tools/testing/selftests/filesystems/open_tree_ns/Makefile b/tools/testing/selftests/filesystems/open_tree_ns/Makefile
index 4976ed1d7d4a..fb2aa77b6edb 100644
--- a/tools/testing/selftests/filesystems/open_tree_ns/Makefile
+++ b/tools/testing/selftests/filesystems/open_tree_ns/Makefile
@@ -1,5 +1,5 @@
# SPDX-License-Identifier: GPL-2.0
-TEST_GEN_PROGS := open_tree_ns_test
+TEST_GEN_PROGS := open_tree_ns_test open_tree_ns_covered_test
CFLAGS += -Wall -O0 -g $(KHDR_INCLUDES) $(TOOLS_INCLUDES)
LDLIBS := -lcap
diff --git a/tools/testing/selftests/filesystems/open_tree_ns/open_tree_ns_covered_test.c b/tools/testing/selftests/filesystems/open_tree_ns/open_tree_ns_covered_test.c
new file mode 100644
index 000000000000..1b2f1385326a
--- /dev/null
+++ b/tools/testing/selftests/filesystems/open_tree_ns/open_tree_ns_covered_test.c
@@ -0,0 +1,183 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * open_tree(OPEN_TREE_NAMESPACE) by a caller that isn't privileged over the
+ * mount namespace it copies from must not reveal what the mounts below the
+ * copied mount cover. Without AT_RECURSIVE the copy is refused when there's
+ * anything mounted below the requested directory. With AT_RECURSIVE
+ * unbindable mounts are copied as well.
+ */
+#define _GNU_SOURCE
+#include <errno.h>
+#include <fcntl.h>
+#include <sched.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <unistd.h>
+#include <sys/mount.h>
+#include <sys/stat.h>
+#include <sys/wait.h>
+
+#include "../wrappers.h"
+#include "../../kselftest_harness.h"
+
+#ifndef OPEN_TREE_NAMESPACE
+#define OPEN_TREE_NAMESPACE (1 << 1)
+#endif
+
+#define DIR_LEN 64
+#define PATH_LEN 128
+
+/* Child exit codes. */
+enum {
+ CHILD_OK,
+ CHILD_USERNS, /* could not create the child user namespace */
+ CHILD_NONREC_ALLOWED, /* the non-recursive copy was not refused */
+ CHILD_NONREC_ERRNO, /* it was refused with the wrong error */
+ CHILD_REC_REFUSED, /* the recursive copy failed */
+ CHILD_SETNS, /* could not enter the new mount namespace */
+ CHILD_NO_COVER, /* the covering mount is missing in the copy */
+ CHILD_REVEALED, /* the covered file is visible in the copy */
+};
+
+static int write_file(const char *path, const char *s)
+{
+ ssize_t n = -1;
+ int fd;
+
+ fd = open(path, O_WRONLY | O_CLOEXEC);
+ if (fd >= 0) {
+ n = write(fd, s, strlen(s));
+ close(fd);
+ }
+ return n == (ssize_t)strlen(s) ? 0 : -1;
+}
+
+static int create_file(const char *path, const char *s)
+{
+ ssize_t n = -1;
+ int fd;
+
+ fd = open(path, O_WRONLY | O_CREAT | O_TRUNC | O_CLOEXEC, 0644);
+ if (fd >= 0) {
+ n = write(fd, s, strlen(s));
+ close(fd);
+ }
+ return n == (ssize_t)strlen(s) ? 0 : -1;
+}
+
+/* Become root in a new user namespace, the uid @uid is mapped to 0. */
+static int enter_userns(uid_t uid, gid_t gid)
+{
+ char map[32];
+
+ if (unshare(CLONE_NEWUSER))
+ return -1;
+ if (write_file("/proc/self/setgroups", "deny") && errno != ENOENT)
+ return -1;
+ snprintf(map, sizeof(map), "0 %d 1", uid);
+ if (write_file("/proc/self/uid_map", map))
+ return -1;
+ snprintf(map, sizeof(map), "0 %d 1", gid);
+ if (write_file("/proc/self/gid_map", map))
+ return -1;
+ return setgid(0) || setuid(0) ? -1 : 0;
+}
+
+FIXTURE(open_tree_ns_covered) {
+ char dir[DIR_LEN];
+ char cover[PATH_LEN];
+};
+
+FIXTURE_VARIANT(open_tree_ns_covered) {
+ int propagation;
+};
+
+FIXTURE_VARIANT_ADD(open_tree_ns_covered, private_cover) {
+ .propagation = MS_PRIVATE,
+};
+
+FIXTURE_VARIANT_ADD(open_tree_ns_covered, unbindable_cover) {
+ .propagation = MS_UNBINDABLE,
+};
+
+/*
+ * Root in a user namespace owns a private mount namespace with a tmpfs
+ * on @dir and a second tmpfs covering @dir/covered/under.txt.
+ */
+FIXTURE_SETUP(open_tree_ns_covered)
+{
+ char p[PATH_LEN];
+
+ snprintf(self->dir, sizeof(self->dir), "/tmp/open_tree_ns_covered.XXXXXX");
+ ASSERT_NE(mkdtemp(self->dir), NULL);
+ if (enter_userns(getuid(), getgid()) || unshare(CLONE_NEWNS)) {
+ rmdir(self->dir);
+ SKIP(return, "test requires user namespaces");
+ }
+ ASSERT_EQ(mount("", "/", NULL, MS_REC | MS_PRIVATE, NULL), 0);
+ ASSERT_EQ(mount("tmpfs", self->dir, "tmpfs", 0, NULL), 0);
+
+ snprintf(self->cover, sizeof(self->cover), "%s/covered", self->dir);
+ ASSERT_EQ(mkdir(self->cover, 0755), 0);
+ snprintf(p, sizeof(p), "%s/covered/under.txt", self->dir);
+ ASSERT_EQ(create_file(p, "hidden"), 0);
+ ASSERT_EQ(mount("tmpfs", self->cover, "tmpfs", 0, NULL), 0);
+ ASSERT_EQ(mount(NULL, self->cover, NULL, variant->propagation, NULL), 0);
+}
+
+FIXTURE_TEARDOWN(open_tree_ns_covered)
+{
+ umount2(self->dir, MNT_DETACH);
+ rmdir(self->dir);
+}
+
+/* A caller in a new user namespace that doesn't own the mount namespace. */
+static int foreign_child(const char *dir)
+{
+ struct stat st;
+ int fd;
+
+ if (enter_userns(0, 0))
+ return CHILD_USERNS;
+
+ fd = sys_open_tree(AT_FDCWD, dir, OPEN_TREE_NAMESPACE | OPEN_TREE_CLOEXEC);
+ if (fd >= 0)
+ return CHILD_NONREC_ALLOWED;
+ if (errno != EINVAL)
+ return CHILD_NONREC_ERRNO;
+
+ fd = sys_open_tree(AT_FDCWD, dir,
+ OPEN_TREE_NAMESPACE | OPEN_TREE_CLOEXEC | AT_RECURSIVE);
+ if (fd < 0)
+ return CHILD_REC_REFUSED;
+ if (setns(fd, CLONE_NEWNS))
+ return CHILD_SETNS;
+ if (stat("/covered", &st))
+ return CHILD_NO_COVER;
+ if (!access("/covered/under.txt", F_OK) || errno != ENOENT)
+ return CHILD_REVEALED;
+ return CHILD_OK;
+}
+
+TEST_F(open_tree_ns_covered, foreign_user_namespace)
+{
+ int status, fd;
+ pid_t pid;
+
+ pid = fork();
+ ASSERT_GE(pid, 0);
+ if (pid == 0)
+ _exit(foreign_child(self->dir));
+ ASSERT_EQ(waitpid(pid, &status, 0), pid);
+ ASSERT_TRUE(WIFEXITED(status));
+ ASSERT_EQ(WEXITSTATUS(status), CHILD_OK);
+
+ /* the owner of the mount namespace keeps bind mount semantics */
+ fd = sys_open_tree(AT_FDCWD, self->dir,
+ OPEN_TREE_NAMESPACE | OPEN_TREE_CLOEXEC);
+ ASSERT_GE(fd, 0);
+ close(fd);
+}
+
+TEST_HARNESS_MAIN
--
2.53.0
^ permalink raw reply related [flat|nested] 22+ messages in thread
* [PATCH 08/17] namespace: look at the topmost mount for a mount namespace file
2026-09-30 13:31 [PATCH 00/17] mount: more bugfixes, the Oprah edition Christian Brauner
` (6 preceding siblings ...)
2026-09-30 13:31 ` [PATCH 07/17] selftests/filesystems: check that OPEN_TREE_NAMESPACE keeps mounts covered Christian Brauner
@ 2026-09-30 13:32 ` Christian Brauner
2026-09-30 13:32 ` [PATCH 09/17] selftests/filesystems: check that a mount namespace file on top doesn't bury a mount Christian Brauner
` (8 subsequent siblings)
16 siblings, 0 replies; 22+ messages in thread
From: Christian Brauner @ 2026-09-30 13:32 UTC (permalink / raw)
To: linux-fsdevel
Cc: Linus Torvalds, Chris Mason, Alexander Viro, Jan Kara,
Jeff Layton, Aleksa Sarai, Amir Goldstein, bpf,
Christian Brauner (Amutable), stable
attach_recursive_mnt() preallocates a mountpoint for the root of the
topmost mount of the source so that a mount that already sits at the
destination can be put on top of the source or on top of one of its
propagated copies.
The copies are made without CL_COPY_MNT_NS_FILE so their chain of
overmounts is shorter than the source's. It ends below the first bind
mount of a mount namespace file. So while the loop walks up the chain of
the source it remembers the mountpoint of that mount in "shorter" and
the existing mount is put there for the copies.
But the loop ends at the topmost mount without ever looking at it. If
the topmost mount is the only mount namespace file in the chain
"shorter" stays NULL and mnt_change_mountpoint() attaches the existing
mount of the copy below the root of the mount namespace file. That's a
dentry of nsfs. No path walk in the copy's mount namespace ever gets
there. The mount is gone until its parent goes and the copy that took
its place can't be unmounted synchronously because it has a child:
S bind mount of a file
N bind mount of a mount namespace file on top of S
A shared mount, B a slave of A, Q mounted on B/file
move_mount(S, A/file)
cat B/file -> the content of S instead of Q
umount B/file -> EBUSY
Look at every mount of the chain including the topmost one.
Fixes: 96f5d2e05165 ("attach_recursive_mnt(): unify the mnt_change_mountpoint() logics")
Cc: stable@vger.kernel.org # v6.17+
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
fs/namespace.c | 4 +++-
1 file changed, 3 insertions(+), 1 deletion(-)
diff --git a/fs/namespace.c b/fs/namespace.c
index de3900dd02f6..f66f4609ef9d 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -2617,9 +2617,11 @@ static int attach_recursive_mnt(struct mount *source_mnt,
* Preallocate a mountpoint in case the new mounts need to be
* mounted beneath mounts on the same mountpoint.
*/
- for (top = source_mnt; unlikely(top->overmount); top = top->overmount) {
+ for (top = source_mnt; ; top = top->overmount) {
if (!shorter && is_mnt_ns_file(top->mnt.mnt_root))
shorter = top->mnt_mp;
+ if (likely(!top->overmount))
+ break;
}
err = get_mountpoint(top->mnt.mnt_root, &root);
if (err)
--
2.53.0
^ permalink raw reply related [flat|nested] 22+ messages in thread
* [PATCH 09/17] selftests/filesystems: check that a mount namespace file on top doesn't bury a mount
2026-09-30 13:31 [PATCH 00/17] mount: more bugfixes, the Oprah edition Christian Brauner
` (7 preceding siblings ...)
2026-09-30 13:32 ` [PATCH 08/17] namespace: look at the topmost mount for a mount namespace file Christian Brauner
@ 2026-09-30 13:32 ` Christian Brauner
2026-09-30 13:32 ` [PATCH 10/17] namespace: check the mounts before reading their parents in pivot_root() Christian Brauner
` (7 subsequent siblings)
16 siblings, 0 replies; 22+ messages in thread
From: Christian Brauner @ 2026-09-30 13:32 UTC (permalink / raw)
To: linux-fsdevel
Cc: Linus Torvalds, Chris Mason, Alexander Viro, Jan Kara,
Jeff Layton, Aleksa Sarai, Amir Goldstein, bpf,
Christian Brauner (Amutable)
Add a test for moving a mount with a mount namespace file on top of it:
- S, a bind mount of a file, with a newer mount namespace's file on top
- A shared with a slave B that has Q on B/file
- S moved onto A/file, its copy lands on B/file below Q
B/file keeps reading Q and once Q is unmounted it reads the copy.
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
.../selftests/filesystems/mount_cycle/.gitignore | 1 +
.../selftests/filesystems/mount_cycle/Makefile | 2 +-
.../mount_cycle/overmount_ns_file_test.c | 187 +++++++++++++++++++++
3 files changed, 189 insertions(+), 1 deletion(-)
diff --git a/tools/testing/selftests/filesystems/mount_cycle/.gitignore b/tools/testing/selftests/filesystems/mount_cycle/.gitignore
index 8cd722977a32..d11f5b720d5b 100644
--- a/tools/testing/selftests/filesystems/mount_cycle/.gitignore
+++ b/tools/testing/selftests/filesystems/mount_cycle/.gitignore
@@ -2,3 +2,4 @@
unmounted_tree_test
overmount_reparent_test
nsfs_rbind_loop_test
+overmount_ns_file_test
diff --git a/tools/testing/selftests/filesystems/mount_cycle/Makefile b/tools/testing/selftests/filesystems/mount_cycle/Makefile
index 32e26336b132..49a8402ca858 100644
--- a/tools/testing/selftests/filesystems/mount_cycle/Makefile
+++ b/tools/testing/selftests/filesystems/mount_cycle/Makefile
@@ -1,6 +1,6 @@
# SPDX-License-Identifier: GPL-2.0
TEST_GEN_PROGS := unmounted_tree_test overmount_reparent_test
-TEST_GEN_PROGS += nsfs_rbind_loop_test
+TEST_GEN_PROGS += nsfs_rbind_loop_test overmount_ns_file_test
CFLAGS += -Wall -O2 -g $(KHDR_INCLUDES)
diff --git a/tools/testing/selftests/filesystems/mount_cycle/overmount_ns_file_test.c b/tools/testing/selftests/filesystems/mount_cycle/overmount_ns_file_test.c
new file mode 100644
index 000000000000..c24436a17c0f
--- /dev/null
+++ b/tools/testing/selftests/filesystems/mount_cycle/overmount_ns_file_test.c
@@ -0,0 +1,187 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * A mount namespace file bind-mounted on top of the mount that is moved
+ * onto a shared mount isn't copied to the peers and slaves. The mount that
+ * already sits at the destination in a slave has to end up on top of the
+ * propagated copy, not below the root of the mount namespace file where no
+ * path walk ever finds it.
+ */
+#define _GNU_SOURCE
+#include <errno.h>
+#include <fcntl.h>
+#include <sched.h>
+#include <signal.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <unistd.h>
+#include <sys/mount.h>
+#include <sys/stat.h>
+#include <sys/syscall.h>
+#include <sys/wait.h>
+
+#include "../wrappers.h"
+#include "../../kselftest_harness.h"
+
+#define DIR_LEN 64
+#define PATH_LEN 128
+
+static int write_file(const char *path, const char *s)
+{
+ ssize_t n = -1;
+ int fd;
+
+ fd = open(path, O_WRONLY | O_CLOEXEC);
+ if (fd >= 0) {
+ n = write(fd, s, strlen(s));
+ close(fd);
+ }
+ return n == (ssize_t)strlen(s) ? 0 : -1;
+}
+
+static int create_file(const char *path, const char *s)
+{
+ ssize_t n = -1;
+ int fd;
+
+ fd = open(path, O_WRONLY | O_CREAT | O_TRUNC | O_CLOEXEC, 0644);
+ if (fd >= 0) {
+ n = write(fd, s, strlen(s));
+ close(fd);
+ }
+ return n == (ssize_t)strlen(s) ? 0 : -1;
+}
+
+/* the first bytes of the file at @path, "" if it can't be read */
+static const char *read_file(const char *path, char *buf, size_t len)
+{
+ ssize_t n = -1;
+ int fd;
+
+ fd = open(path, O_RDONLY | O_CLOEXEC);
+ if (fd >= 0) {
+ n = read(fd, buf, len - 1);
+ close(fd);
+ }
+ buf[n > 0 ? n : 0] = '\0';
+ return buf;
+}
+
+/* Become root in a new user namespace with a private mount namespace. */
+static int enter_userns(void)
+{
+ uid_t uid = getuid();
+ gid_t gid = getgid();
+ char map[32];
+
+ if (unshare(CLONE_NEWUSER | CLONE_NEWNS))
+ return -1;
+ if (write_file("/proc/self/setgroups", "deny") && errno != ENOENT)
+ return -1;
+ snprintf(map, sizeof(map), "0 %d 1", uid);
+ if (write_file("/proc/self/uid_map", map))
+ return -1;
+ snprintf(map, sizeof(map), "0 %d 1", gid);
+ if (write_file("/proc/self/gid_map", map))
+ return -1;
+ if (setgid(0) || setuid(0))
+ return -1;
+ return mount("", "/", NULL, MS_REC | MS_PRIVATE, NULL);
+}
+
+FIXTURE(overmount_ns_file) {
+ char dir[DIR_LEN];
+ pid_t child;
+};
+
+FIXTURE_SETUP(overmount_ns_file)
+{
+ self->child = -1;
+ snprintf(self->dir, sizeof(self->dir), "/tmp/overmount_ns_file.XXXXXX");
+ ASSERT_NE(mkdtemp(self->dir), NULL);
+ if (enter_userns()) {
+ rmdir(self->dir);
+ SKIP(return, "test requires user namespaces");
+ }
+ ASSERT_EQ(mount("tmpfs", self->dir, "tmpfs", 0, NULL), 0);
+}
+
+FIXTURE_TEARDOWN(overmount_ns_file)
+{
+ if (self->child > 0) {
+ kill(self->child, SIGKILL);
+ waitpid(self->child, NULL, 0);
+ }
+ umount2(self->dir, MNT_DETACH);
+ rmdir(self->dir);
+}
+
+/*
+ * A is a shared tmpfs and B its slave with Q, a bind mount of a file, on
+ * B/file. S is a bind mount of a file with N, a bind mount of a newer mount
+ * namespace's file, on top of it. S is moved onto A/file. Its copy S' lands
+ * on B/file below Q, without N. B/file keeps reading Q and once Q is
+ * unmounted it reads S'.
+ */
+TEST_F(overmount_ns_file, existing_mount_stays_on_top)
+{
+ char a[PATH_LEN], b[PATH_LEN], s[PATH_LEN], p[PATH_LEN], buf[16];
+ int fd, pfd[2];
+ char c;
+
+ snprintf(a, sizeof(a), "%s/A", self->dir);
+ snprintf(b, sizeof(b), "%s/B", self->dir);
+ snprintf(s, sizeof(s), "%s/s", self->dir);
+ ASSERT_EQ(mkdir(a, 0755), 0);
+ ASSERT_EQ(mkdir(b, 0755), 0);
+ ASSERT_EQ(mkdir(s, 0755), 0);
+
+ /* A shared, B its slave, Q on B/file */
+ ASSERT_EQ(mount("tmpfs", a, "tmpfs", 0, NULL), 0);
+ ASSERT_EQ(mount(NULL, a, NULL, MS_SHARED, NULL), 0);
+ snprintf(p, sizeof(p), "%s/A/file", self->dir);
+ ASSERT_EQ(create_file(p, "A"), 0);
+ ASSERT_EQ(mount(a, b, NULL, MS_BIND, NULL), 0);
+ ASSERT_EQ(mount(NULL, b, NULL, MS_SLAVE, NULL), 0);
+ snprintf(p, sizeof(p), "%s/Q", self->dir);
+ ASSERT_EQ(create_file(p, "Q"), 0);
+ snprintf(b, sizeof(b), "%s/B/file", self->dir);
+ ASSERT_EQ(mount(p, b, NULL, MS_BIND, NULL), 0);
+
+ /* S on s/f, pinned by a file descriptor before N goes on top */
+ ASSERT_EQ(mount("tmpfs", s, "tmpfs", 0, NULL), 0);
+ snprintf(p, sizeof(p), "%s/s/f", self->dir);
+ ASSERT_EQ(create_file(p, "f"), 0);
+ snprintf(s, sizeof(s), "%s/s/S", self->dir);
+ ASSERT_EQ(create_file(s, "S"), 0);
+ ASSERT_EQ(mount(s, p, NULL, MS_BIND, NULL), 0);
+ fd = open(p, O_PATH | O_CLOEXEC);
+ ASSERT_GE(fd, 0);
+
+ /* a newer mount namespace whose file can be bound */
+ ASSERT_EQ(pipe(pfd), 0);
+ self->child = fork();
+ ASSERT_GE(self->child, 0);
+ if (self->child == 0) {
+ if (unshare(CLONE_NEWNS) || write(pfd[1], "r", 1) != 1)
+ _exit(1);
+ pause();
+ _exit(0);
+ }
+ ASSERT_EQ(read(pfd[0], &c, 1), 1);
+ snprintf(s, sizeof(s), "/proc/%d/ns/mnt", self->child);
+ ASSERT_EQ(mount(s, p, NULL, MS_BIND, NULL), 0);
+
+ ASSERT_STREQ(read_file(b, buf, sizeof(buf)), "Q");
+
+ snprintf(a, sizeof(a), "%s/A/file", self->dir);
+ ASSERT_EQ(sys_move_mount(fd, "", AT_FDCWD, a, MOVE_MOUNT_F_EMPTY_PATH), 0);
+ close(fd);
+
+ /* Q is still on top of the copy in B and can be unmounted */
+ EXPECT_STREQ(read_file(b, buf, sizeof(buf)), "Q");
+ EXPECT_EQ(umount2(b, 0), 0);
+ EXPECT_STREQ(read_file(b, buf, sizeof(buf)), "S");
+}
+
+TEST_HARNESS_MAIN
--
2.53.0
^ permalink raw reply related [flat|nested] 22+ messages in thread
* [PATCH 10/17] namespace: check the mounts before reading their parents in pivot_root()
2026-09-30 13:31 [PATCH 00/17] mount: more bugfixes, the Oprah edition Christian Brauner
` (8 preceding siblings ...)
2026-09-30 13:32 ` [PATCH 09/17] selftests/filesystems: check that a mount namespace file on top doesn't bury a mount Christian Brauner
@ 2026-09-30 13:32 ` Christian Brauner
2026-09-30 13:32 ` [PATCH 11/17] namespace: don't reconfigure internal superblocks via remount and umount Christian Brauner
` (6 subsequent siblings)
16 siblings, 0 replies; 22+ messages in thread
From: Christian Brauner @ 2026-09-30 13:32 UTC (permalink / raw)
To: linux-fsdevel
Cc: Linus Torvalds, Chris Mason, Alexander Viro, Jan Kara,
Jeff Layton, Aleksa Sarai, Amir Goldstein, bpf,
Christian Brauner (Amutable), stable
path_pivot_root() reads the parents of new_root and of the caller's root
and checks whether they are shared before it checks that either mount is
in the caller's mount namespace. Only namespace_sem is held. That's fine
for a mount that is in the caller's mount namespace.
But new_root can be a file descriptor to a mount that has been unmounted
and that only the file descriptor keeps alive. If that mount stayed
attached to its parent when it was unmounted nothing pins the parent for
it. The final mntput() of the parent unhooks the children under
mount_lock alone and frees the parent afterwards:
pivot_root() close(fd), last ref on the parent
------------ ---------------------------------
ex_parent = new_mnt->mnt_parent
mntput_no_expire_slowpath()
__umount_mnt(new_mnt)
cleanup_mnt()
call_rcu()
IS_MNT_SHARED(ex_parent)
BUG: KASAN: slab-use-after-free in path_pivot_root+0xf1a/0x1840
Read of size 4 at addr ffff8881047382f8 by task pivot_widen/157
path_pivot_root+0xf1a/0x1840
__x64_sys_pivot_root+0x165/0x190
The same goes for the caller's root via chroot(). Both outcomes of the
check end in EINVAL so nothing but the read itself goes wrong. It's the
same thing commit bb4609405752 ("statmount: read the parent of an
unmounted mount under mount_lock") fixed for statmount().
Check that both mounts are in the caller's namespace before their
parents are read. Every path returns EINVAL either way.
Fixes: e0c9c0afd2fc ("mnt: Update detach_mounts to leave mounts connected")
Cc: stable@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
fs/namespace.c | 5 +++--
1 file changed, 3 insertions(+), 2 deletions(-)
diff --git a/fs/namespace.c b/fs/namespace.c
index f66f4609ef9d..0a7d50db228d 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -4780,14 +4780,15 @@ int path_pivot_root(struct path *new, struct path *old)
new_mnt = real_mount(new->mnt);
root_mnt = real_mount(root.mnt);
+ /* only a mounted mount has a parent that namespace_sem pins */
+ if (!check_mnt(root_mnt) || !check_mnt(new_mnt))
+ return -EINVAL;
ex_parent = new_mnt->mnt_parent;
root_parent = root_mnt->mnt_parent;
if (IS_MNT_SHARED(old_mnt) ||
IS_MNT_SHARED(ex_parent) ||
IS_MNT_SHARED(root_parent))
return -EINVAL;
- if (!check_mnt(root_mnt) || !check_mnt(new_mnt))
- return -EINVAL;
if (new_mnt->mnt.mnt_flags & MNT_LOCKED)
return -EINVAL;
if (d_unlinked(new->dentry))
--
2.53.0
^ permalink raw reply related [flat|nested] 22+ messages in thread
* [PATCH 11/17] namespace: don't reconfigure internal superblocks via remount and umount
2026-09-30 13:31 [PATCH 00/17] mount: more bugfixes, the Oprah edition Christian Brauner
` (9 preceding siblings ...)
2026-09-30 13:32 ` [PATCH 10/17] namespace: check the mounts before reading their parents in pivot_root() Christian Brauner
@ 2026-09-30 13:32 ` Christian Brauner
2026-09-30 13:32 ` [PATCH 12/17] selftests/filesystems: check that the nullfs root can't be reconfigured Christian Brauner
` (5 subsequent siblings)
16 siblings, 0 replies; 22+ messages in thread
From: Christian Brauner @ 2026-09-30 13:32 UTC (permalink / raw)
To: linux-fsdevel
Cc: Linus Torvalds, Chris Mason, Alexander Viro, Jan Kara,
Jeff Layton, Aleksa Sarai, Amir Goldstein, bpf,
Christian Brauner (Amutable), stable
Commit 02587a4af82a ("fs: refuse fspick() on internal superblocks")
blocked fspick() on SB_NOUSER superblocks. But that's not the only way
to reconfigure_super(). mount(MS_REMOUNT) and umount() of the caller's
root without MNT_DETACH. The root of an empty mount namespace is a
nullfs mount and all mount namespaces share that superblock:
nullfs: fspick: FAIL errno=22 (Invalid argument)
nullfs: mount(MS_REMOUNT|MS_RDONLY): ok(0)
nullfs after remount: statfs(/): magic=0x4e554c4c flags=0x21 RDONLY
nullfs: umount2("/", 0): ok(0)
Refuse both like fspick() does. Only root in the initial user namespace
can do this and nullfs is empty and immutable so the flags don't buy
anything. But they show up in statfs() for the root of every mount
namespace on the host.
Fixes: 9d4e752a24f7 ("namespace: allow creating empty mount namespaces")
Cc: stable@vger.kernel.org # v7.1+
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
fs/fs_context.c | 4 ++++
fs/fsopen.c | 3 ---
2 files changed, 4 insertions(+), 3 deletions(-)
diff --git a/fs/fs_context.c b/fs/fs_context.c
index 23ad66cd94e1..a50584df97be 100644
--- a/fs/fs_context.c
+++ b/fs/fs_context.c
@@ -315,6 +315,10 @@ struct fs_context *fs_context_for_reconfigure(struct dentry *dentry,
unsigned int sb_flags,
unsigned int sb_flags_mask)
{
+ /* kernel-internal superblocks are nobody's to reconfigure */
+ if (dentry->d_sb->s_flags & SB_NOUSER)
+ return ERR_PTR(-EINVAL);
+
return alloc_fs_context(dentry->d_sb->s_type, dentry, sb_flags,
sb_flags_mask, FS_CONTEXT_FOR_RECONFIGURE);
}
diff --git a/fs/fsopen.c b/fs/fsopen.c
index 9d5a7a22b529..ae19e5136598 100644
--- a/fs/fsopen.c
+++ b/fs/fsopen.c
@@ -190,9 +190,6 @@ SYSCALL_DEFINE3(fspick, int, dfd, const char __user *, path, unsigned int, flags
ret = -EINVAL;
if (target.mnt->mnt_root != target.dentry)
goto err_path;
- /* kernel-internal superblocks are nobody's to reconfigure */
- if (target.dentry->d_sb->s_flags & SB_NOUSER)
- goto err_path;
fc = fs_context_for_reconfigure(target.dentry, 0, 0);
if (IS_ERR(fc)) {
--
2.53.0
^ permalink raw reply related [flat|nested] 22+ messages in thread
* [PATCH 12/17] selftests/filesystems: check that the nullfs root can't be reconfigured
2026-09-30 13:31 [PATCH 00/17] mount: more bugfixes, the Oprah edition Christian Brauner
` (10 preceding siblings ...)
2026-09-30 13:32 ` [PATCH 11/17] namespace: don't reconfigure internal superblocks via remount and umount Christian Brauner
@ 2026-09-30 13:32 ` Christian Brauner
2026-09-30 13:32 ` [PATCH 13/17] namespace: remove the fsnotify marks of a mount namespace in process context Christian Brauner
` (4 subsequent siblings)
16 siblings, 0 replies; 22+ messages in thread
From: Christian Brauner @ 2026-09-30 13:32 UTC (permalink / raw)
To: linux-fsdevel
Cc: Linus Torvalds, Chris Mason, Alexander Viro, Jan Kara,
Jeff Layton, Aleksa Sarai, Amir Goldstein, bpf,
Christian Brauner (Amutable)
Add a test for the root of an empty mount namespace:
- fspick() of the root fails with EINVAL
- mount(MS_REMOUNT) of the root fails with EINVAL
- umount() of the root fails and doesn't remount it read-only
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
.../selftests/filesystems/empty_mntns/.gitignore | 1 +
.../selftests/filesystems/empty_mntns/Makefile | 2 +
.../empty_mntns/internal_sb_reconfigure_test.c | 108 +++++++++++++++++++++
3 files changed, 111 insertions(+)
diff --git a/tools/testing/selftests/filesystems/empty_mntns/.gitignore b/tools/testing/selftests/filesystems/empty_mntns/.gitignore
index 99f89d329db2..32125b3eaa80 100644
--- a/tools/testing/selftests/filesystems/empty_mntns/.gitignore
+++ b/tools/testing/selftests/filesystems/empty_mntns/.gitignore
@@ -2,3 +2,4 @@
clone3_empty_mntns_test
empty_mntns_test
overmount_chroot_test
+internal_sb_reconfigure_test
diff --git a/tools/testing/selftests/filesystems/empty_mntns/Makefile b/tools/testing/selftests/filesystems/empty_mntns/Makefile
index 22e3fb915e81..b64818b962ca 100644
--- a/tools/testing/selftests/filesystems/empty_mntns/Makefile
+++ b/tools/testing/selftests/filesystems/empty_mntns/Makefile
@@ -4,9 +4,11 @@ CFLAGS += -Wall -O2 -g $(KHDR_INCLUDES) $(TOOLS_INCLUDES)
LDLIBS += -lcap
TEST_GEN_PROGS := empty_mntns_test overmount_chroot_test clone3_empty_mntns_test
+TEST_GEN_PROGS += internal_sb_reconfigure_test
include ../../lib.mk
$(OUTPUT)/empty_mntns_test: ../utils.c
$(OUTPUT)/overmount_chroot_test: ../utils.c
$(OUTPUT)/clone3_empty_mntns_test: ../utils.c
+$(OUTPUT)/internal_sb_reconfigure_test: ../utils.c
diff --git a/tools/testing/selftests/filesystems/empty_mntns/internal_sb_reconfigure_test.c b/tools/testing/selftests/filesystems/empty_mntns/internal_sb_reconfigure_test.c
new file mode 100644
index 000000000000..cb645d1e5a9a
--- /dev/null
+++ b/tools/testing/selftests/filesystems/empty_mntns/internal_sb_reconfigure_test.c
@@ -0,0 +1,108 @@
+// SPDX-License-Identifier: GPL-2.0-or-later
+/*
+ * The root of an empty mount namespace is a nullfs mount. Its superblock is
+ * kernel-internal and shared by every mount namespace. It can't be
+ * reconfigured, neither through fspick() nor through mount(MS_REMOUNT) nor
+ * through umount() of the root which remounts it read-only.
+ */
+#define _GNU_SOURCE
+#include <errno.h>
+#include <fcntl.h>
+#include <sched.h>
+#include <stdio.h>
+#include <string.h>
+#include <sys/mount.h>
+#include <sys/statfs.h>
+#include <sys/statvfs.h>
+#include <sys/syscall.h>
+#include <sys/vfs.h>
+#include <sys/wait.h>
+#include <unistd.h>
+
+#include "../utils.h"
+#include "../wrappers.h"
+#include "empty_mntns.h"
+#include "kselftest_harness.h"
+
+#ifndef __NR_fspick
+#define __NR_fspick 433
+#endif
+
+static int sys_fspick(int dfd, const char *path, unsigned int flags)
+{
+ return syscall(__NR_fspick, dfd, path, flags);
+}
+
+/* Child exit codes. */
+enum {
+ CHILD_OK,
+ CHILD_USERNS, /* could not create the user namespace */
+ CHILD_UNSHARE, /* could not create the empty mount namespace */
+ CHILD_FSPICK, /* fspick() of the root was not refused with EINVAL */
+ CHILD_REMOUNT, /* mount(MS_REMOUNT) was not refused with EINVAL */
+ CHILD_UMOUNT, /* umount() of the root succeeded */
+ CHILD_STATFS, /* statfs() of the root failed */
+ CHILD_RDONLY, /* the root ended up read-only */
+};
+
+static int empty_mntns_child(void)
+{
+ struct statfs st;
+
+ if (enter_userns())
+ return CHILD_USERNS;
+ if (unshare(UNSHARE_EMPTY_MNTNS))
+ return CHILD_UNSHARE;
+
+ if (sys_fspick(AT_FDCWD, "/", 0) >= 0 || errno != EINVAL)
+ return CHILD_FSPICK;
+ if (!mount(NULL, "/", NULL, MS_REMOUNT | MS_RDONLY, NULL) ||
+ errno != EINVAL)
+ return CHILD_REMOUNT;
+ if (!umount2("/", 0))
+ return CHILD_UMOUNT;
+ if (statfs("/", &st))
+ return CHILD_STATFS;
+ if (st.f_flags & ST_RDONLY)
+ return CHILD_RDONLY;
+ return CHILD_OK;
+}
+
+FIXTURE(internal_sb_reconfigure) {};
+
+FIXTURE_SETUP(internal_sb_reconfigure)
+{
+ pid_t pid;
+ int status;
+
+ pid = fork();
+ ASSERT_GE(pid, 0);
+ if (pid == 0) {
+ if (enter_userns())
+ _exit(1);
+ if (unshare(UNSHARE_EMPTY_MNTNS))
+ _exit(1);
+ _exit(0);
+ }
+ ASSERT_EQ(waitpid(pid, &status, 0), pid);
+ if (!WIFEXITED(status) || WEXITSTATUS(status))
+ SKIP(return, "UNSHARE_EMPTY_MNTNS not supported");
+}
+
+FIXTURE_TEARDOWN(internal_sb_reconfigure) {}
+
+TEST_F(internal_sb_reconfigure, nullfs_root)
+{
+ pid_t pid;
+ int status;
+
+ pid = fork();
+ ASSERT_GE(pid, 0);
+ if (pid == 0)
+ _exit(empty_mntns_child());
+ ASSERT_EQ(waitpid(pid, &status, 0), pid);
+ ASSERT_TRUE(WIFEXITED(status));
+ ASSERT_EQ(WEXITSTATUS(status), CHILD_OK);
+}
+
+TEST_HARNESS_MAIN
--
2.53.0
^ permalink raw reply related [flat|nested] 22+ messages in thread
* [PATCH 13/17] namespace: remove the fsnotify marks of a mount namespace in process context
2026-09-30 13:31 [PATCH 00/17] mount: more bugfixes, the Oprah edition Christian Brauner
` (11 preceding siblings ...)
2026-09-30 13:32 ` [PATCH 12/17] selftests/filesystems: check that the nullfs root can't be reconfigured Christian Brauner
@ 2026-09-30 13:32 ` Christian Brauner
2026-09-30 15:07 ` Amir Goldstein
2026-09-30 13:32 ` [PATCH 14/17] fsnotify: detach the connector before destroying its marks Christian Brauner
` (3 subsequent siblings)
16 siblings, 1 reply; 22+ messages in thread
From: Christian Brauner @ 2026-09-30 13:32 UTC (permalink / raw)
To: linux-fsdevel
Cc: Linus Torvalds, Chris Mason, Alexander Viro, Jan Kara,
Jeff Layton, Aleksa Sarai, Amir Goldstein, bpf,
Christian Brauner (Amutable), stable
mnt_ns_release() drops the last passive reference of a mount namespace
and removes its fanotify marks via fsnotify_mntns_delete(). That takes
the mutex of every group with a mark on the namespace and the spinlock
of the connector. Fine from process context. But mnt_ns_tree_remove()
hands the reference the namespace was allocated with to call_rcu() and
so the marks are removed from the RCU softirq:
BUG: sleeping function called from invalid context at kernel/locking/mutex.c:623
in_atomic(): 1, irqs_disabled(): 0, non_block: 0, pid: 0, name: swapper/3
__mutex_lock+0x113/0x24b0
fsnotify_destroy_marks+0x11b/0x3d0
mnt_ns_release_rcu+0x57/0xa0
rcu_core+0x6b6/0x1e40
and lockdep complains about the connector lock being taken from softirq
context. A fanotify group with a FAN_MARK_MNTNS mark on the mount
namespace of another task is all that's needed. Root in a user
namespace can do that for a mount namespace it owns. The task exits and
the namespace is freed with the mark still on it.
Remove the marks in free_mnt_ns() before the namespace is handed to RCU.
That runs in process context once the last active reference is gone. A
mark is added through a file descriptor to the namespace which holds an
active reference so no mark can show up after that.
Fixes: bf630c401641 ("vfs: add notifications for mount attach and detach")
Cc: stable@vger.kernel.org # v6.15+
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
fs/namespace.c | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
diff --git a/fs/namespace.c b/fs/namespace.c
index 0a7d50db228d..3c90d853e091 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -130,7 +130,6 @@ static void mnt_ns_release(struct mnt_namespace *ns)
{
/* keep alive for {list,stat}mount() */
if (ns && refcount_dec_and_test(&ns->passive)) {
- fsnotify_mntns_delete(ns);
put_user_ns(ns->user_ns);
kfree(ns);
}
@@ -4279,6 +4278,8 @@ static void free_mnt_ns(struct mnt_namespace *ns)
if (!is_anon_ns(ns))
ns_common_free(ns);
dec_mnt_namespaces(ns->ucounts);
+ /* the last active reference is gone, no mark can show up anymore */
+ fsnotify_mntns_delete(ns);
mnt_ns_tree_remove(ns);
}
--
2.53.0
^ permalink raw reply related [flat|nested] 22+ messages in thread
* [PATCH 14/17] fsnotify: detach the connector before destroying its marks
2026-09-30 13:31 [PATCH 00/17] mount: more bugfixes, the Oprah edition Christian Brauner
` (12 preceding siblings ...)
2026-09-30 13:32 ` [PATCH 13/17] namespace: remove the fsnotify marks of a mount namespace in process context Christian Brauner
@ 2026-09-30 13:32 ` Christian Brauner
2026-10-01 9:31 ` Christian Brauner
2026-09-30 13:32 ` [PATCH 15/17] dcache: don't put a mountpoint on a dentry that's being removed Christian Brauner
` (2 subsequent siblings)
16 siblings, 1 reply; 22+ messages in thread
From: Christian Brauner @ 2026-09-30 13:32 UTC (permalink / raw)
To: linux-fsdevel
Cc: Linus Torvalds, Chris Mason, Alexander Viro, Jan Kara,
Jeff Layton, Aleksa Sarai, Amir Goldstein, bpf,
Christian Brauner (Amutable), stable
fsnotify_destroy_marks() removes every mark of an object that goes away.
It walks the mark list of the connector and has to drop the connector
lock around fsnotify_destroy_mark() since that sleeps. Afterwards it
continues from the mark it just destroyed. That mark stays on the list
because the function holds a reference to it. But while the lock is
dropped fsnotify_add_mark_list() can add a mark to the very same
connector and put it in front of the current one when group priority
dictates. fsnotify_grab_connector() hands out the connector until it is
detached and it only gets detached after the walk.
So the walk never sees that mark. The connector is detached from the
object with the new mark still attached to it. It never receives an
event nor IN_IGNORED. For inotify that's a watch descriptor that never
reports anything and doesn't show up in fdinfo either:
wd1 = inotify_add_watch(/proc/self/fd/6) = 2
events after the unlink:
wd 1 mask 0x400 IN_DELETE_SELF
wd 1 mask 0x8000 IN_IGNORED
events after write() to the file:
(none)
Detach the connector from the object before the marks are destroyed.
fsnotify_grab_connector() refuses a detached connector so a mark added
from then on gets a connector of its own. Nothing during the destruction
needs the connector. The inode reference of the connector is dropped
after the marks are gone as before.
Fixes: 6b3f05d24d35 ("fsnotify: Detach mark from object list when last reference is dropped")
Cc: stable@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
fs/notify/mark.c | 16 ++++++++++------
1 file changed, 10 insertions(+), 6 deletions(-)
diff --git a/fs/notify/mark.c b/fs/notify/mark.c
index b2640d836a71..f6891d39e42f 100644
--- a/fs/notify/mark.c
+++ b/fs/notify/mark.c
@@ -1112,6 +1112,16 @@ void fsnotify_destroy_marks(fsnotify_connp_t *connp)
conn = fsnotify_grab_connector(connp);
if (!conn)
return;
+ /*
+ * Detach the connector from the object first. Once conn->lock is
+ * dropped a mark could be added in front of the one we're at and the
+ * walk would miss it. fsnotify_grab_connector() refuses a detached
+ * connector so any mark added from now on gets a connector of its
+ * own. This also stops pinning the inode until all mark references
+ * get dropped. It would lead to strange results such as delaying
+ * inode deletion or blocking unmount.
+ */
+ objp = fsnotify_detach_connector_from_object(conn, &type);
/*
* We have to be careful since we can race with e.g.
* fsnotify_clear_marks_by_group() and once we drop the conn->lock, the
@@ -1128,12 +1138,6 @@ void fsnotify_destroy_marks(fsnotify_connp_t *connp)
fsnotify_destroy_mark(mark, mark->group);
spin_lock(&conn->lock);
}
- /*
- * Detach list from object now so that we don't pin inode until all
- * mark references get dropped. It would lead to strange results such
- * as delaying inode deletion or blocking unmount.
- */
- objp = fsnotify_detach_connector_from_object(conn, &type);
spin_unlock(&conn->lock);
if (old_mark)
fsnotify_put_mark(old_mark);
--
2.53.0
^ permalink raw reply related [flat|nested] 22+ messages in thread
* [PATCH 15/17] dcache: don't put a mountpoint on a dentry that's being removed
2026-09-30 13:31 [PATCH 00/17] mount: more bugfixes, the Oprah edition Christian Brauner
` (13 preceding siblings ...)
2026-09-30 13:32 ` [PATCH 14/17] fsnotify: detach the connector before destroying its marks Christian Brauner
@ 2026-09-30 13:32 ` Christian Brauner
2026-09-30 13:32 ` [PATCH 16/17] unshare: don't drop active namespace references that were never taken Christian Brauner
2026-09-30 13:32 ` [PATCH 17/17] namespace: don't let a pseudo dentry become the root of a mount Christian Brauner
16 siblings, 0 replies; 22+ messages in thread
From: Christian Brauner @ 2026-09-30 13:32 UTC (permalink / raw)
To: linux-fsdevel
Cc: Linus Torvalds, Chris Mason, Alexander Viro, Jan Kara,
Jeff Layton, Aleksa Sarai, Amir Goldstein, bpf,
Christian Brauner (Amutable), stable
rmdir(), unlink() and rename() call dont_mount() on the victim and then
detach_mounts() with the victim's inode locked. do_lock_mount() takes
the inode lock of the mountpoint and checks cant_mount() so a mount
can't show up after detach_mounts().
But attach_recursive_mnt() makes a second mountpoint for the root of the
source mount so that the mounts already located at the destination can
be put on top of it. No inode is locked for that one and d_set_mounted()
only refuses a dentry that is unlinked. Between detach_mounts() and
d_delete() the victim is still hashed:
rmrace: b passed dont_mount() and detach_mounts(), sleeping
T2: move_mount(S, /tmp/plcant/x, BENEATH) = 0 errno 0 ()
T1: rmdir(/tmp/plcant/d/b) = 0 errno 0
109 107 0:61 /d/b//deleted /tmp/plcant/x rw,relatime - tmpfs tmpfs rw
108 109 0:63 / /tmp/plcant/x rw,relatime - tmpfs T rw
A bind mount of the directory that's being removed is moved beneath an
existing mount while the rmdir() is located between the two calls. The
mount ends up on the removed directory and nothing will ever detach it.
Check cant_mount() in d_set_mounted() as well. dont_mount() raises the
flag under d_lock before detach_mounts() runs so either the mountpoint
is set first and detach_mounts() finds the mount or the flag is seen and
the mount is refused.
Fixes: 1064f874abc0 ("mnt: Tuck mounts under others instead of creating shadow/side mounts.")
Cc: stable@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
fs/dcache.c | 4 +++-
1 file changed, 3 insertions(+), 1 deletion(-)
diff --git a/fs/dcache.c b/fs/dcache.c
index a66be85f9d01..f44c7b39c3b3 100644
--- a/fs/dcache.c
+++ b/fs/dcache.c
@@ -1583,6 +1583,8 @@ EXPORT_SYMBOL(path_has_submounts);
*
* Only one of d_invalidate() and d_set_mounted() must succeed. For
* this reason take rename_lock and d_lock on dentry and ancestors.
+ * Likewise for dont_mount() which marks a dentry that is being removed
+ * under d_lock.
*/
int d_set_mounted(struct dentry *dentry)
{
@@ -1599,7 +1601,7 @@ int d_set_mounted(struct dentry *dentry)
spin_unlock(&p->d_lock);
}
spin_lock(&dentry->d_lock);
- if (!d_unlinked(dentry)) {
+ if (!d_unlinked(dentry) && !cant_mount(dentry)) {
ret = -EBUSY;
if (!d_mountpoint(dentry)) {
dentry->d_flags |= DCACHE_MOUNTED;
--
2.53.0
^ permalink raw reply related [flat|nested] 22+ messages in thread
* [PATCH 16/17] unshare: don't drop active namespace references that were never taken
2026-09-30 13:31 [PATCH 00/17] mount: more bugfixes, the Oprah edition Christian Brauner
` (14 preceding siblings ...)
2026-09-30 13:32 ` [PATCH 15/17] dcache: don't put a mountpoint on a dentry that's being removed Christian Brauner
@ 2026-09-30 13:32 ` Christian Brauner
2026-09-30 13:32 ` [PATCH 17/17] namespace: don't let a pseudo dentry become the root of a mount Christian Brauner
16 siblings, 0 replies; 22+ messages in thread
From: Christian Brauner @ 2026-09-30 13:32 UTC (permalink / raw)
To: linux-fsdevel
Cc: Linus Torvalds, Chris Mason, Alexander Viro, Jan Kara,
Jeff Layton, Aleksa Sarai, Amir Goldstein, bpf,
Christian Brauner (Amutable), stable
Active references on the namespaces of an nsproxy are taken when the
nsproxy is installed into a task in switch_task_namespaces() and
copy_namespaces() and dropped again by put_nsproxy() through
deactivate_nsproxy().
But ksys_unshare() calls put_nsproxy() on an nsproxy that was never
installed when set_cred_ucounts() fails. The new namespaces of that
nsproxy go from zero to minus one and the namespaces shared with the
caller lose a reference that belongs to the nsproxy the caller keeps
using:
WARNING: kernel/nscommon.c:171 at __ns_ref_active_put+0x1cd/0x230
nsproxy_ns_active_put
deactivate_nsproxy
ksys_unshare
Afterwards the namespaces the caller lives in aren't listed by listns()
anymore. set_cred_ucounts() only fails when alloc_ucounts() can't
allocate and it only allocates when the real uid of the caller differs
from its effective uid.
Commit cefd55bd2159 ("nsproxy: fix free_nsproxy() and simplify
create_new_namespaces()") separated the two cases on purpose.
nsproxy_free() frees an nsproxy that was prepared but never installed
and that's what a failed setns() uses in put_nsset(). Export it and use
it for a failed unshare() as well.
Fixes: a98621a0f187 ("unshare: fix nsproxy leak in ksys_unshare() on set_cred_ucounts() failure")
Cc: stable@vger.kernel.org # v7.1+
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
include/linux/nsproxy.h | 1 +
kernel/fork.c | 3 ++-
kernel/nsproxy.c | 2 +-
3 files changed, 4 insertions(+), 2 deletions(-)
diff --git a/include/linux/nsproxy.h b/include/linux/nsproxy.h
index 5a67648721c7..dc2447f5f092 100644
--- a/include/linux/nsproxy.h
+++ b/include/linux/nsproxy.h
@@ -100,6 +100,7 @@ void exit_cred_namespaces(struct task_struct *tsk);
void switch_task_namespaces(struct task_struct *tsk, struct nsproxy *new);
int exec_task_namespaces(void);
void deactivate_nsproxy(struct nsproxy *ns);
+void nsproxy_free(struct nsproxy *ns);
int unshare_nsproxy_namespaces(unsigned long, struct nsproxy **,
struct cred *, struct fs_struct *);
int __init nsproxy_cache_init(void);
diff --git a/kernel/fork.c b/kernel/fork.c
index da48168c504f..d442cd68a37a 100644
--- a/kernel/fork.c
+++ b/kernel/fork.c
@@ -3338,8 +3338,9 @@ int ksys_unshare(unsigned long unshare_flags)
perf_event_namespaces(current);
bad_unshare_cleanup_nsproxy:
+ /* never installed, so no active references to drop */
if (new_nsproxy)
- put_nsproxy(new_nsproxy);
+ nsproxy_free(new_nsproxy);
bad_unshare_cleanup_cred:
if (new_cred)
put_cred(new_cred);
diff --git a/kernel/nsproxy.c b/kernel/nsproxy.c
index d9d3d5973bf5..3fb1595a4a59 100644
--- a/kernel/nsproxy.c
+++ b/kernel/nsproxy.c
@@ -61,7 +61,7 @@ static inline struct nsproxy *create_nsproxy(void)
return nsproxy;
}
-static inline void nsproxy_free(struct nsproxy *ns)
+void nsproxy_free(struct nsproxy *ns)
{
put_mnt_ns(ns->mnt_ns);
put_uts_ns(ns->uts_ns);
--
2.53.0
^ permalink raw reply related [flat|nested] 22+ messages in thread
* [PATCH 17/17] namespace: don't let a pseudo dentry become the root of a mount
2026-09-30 13:31 [PATCH 00/17] mount: more bugfixes, the Oprah edition Christian Brauner
` (15 preceding siblings ...)
2026-09-30 13:32 ` [PATCH 16/17] unshare: don't drop active namespace references that were never taken Christian Brauner
@ 2026-09-30 13:32 ` Christian Brauner
16 siblings, 0 replies; 22+ messages in thread
From: Christian Brauner @ 2026-09-30 13:32 UTC (permalink / raw)
To: linux-fsdevel
Cc: Linus Torvalds, Chris Mason, Alexander Viro, Jan Kara,
Jeff Layton, Aleksa Sarai, Amir Goldstein, bpf,
Christian Brauner (Amutable), stable
d_alloc_pseudo() hands out dentries for pipes, sockets and other files
that are never anyone's child or parent. They carry DCACHE_NORCU and
dentry_free() frees them right away because no lockless path walk can
ever reach them. That holds as long as such a dentry isn't the root of
a mount. But bind mounting /proc/self/fd/<fd> of such a file does
exactly that.
Most callers of alloc_file_pseudo() put their files on kernel-internal
mounts and may_copy_tree() refuses those. bpf_token_create() doesn't. It
places the token file on the bpffs mount the caller handed it and that
mount is in the caller's mount namespace so the clone goes through:
mount --bind /proc/self/fd/<token> <file>
open_tree(tokfd, "", AT_EMPTY_PATH | OPEN_TREE_CLONE) + move_mount()
__follow_mount_rcu() then loads ->mnt_root of that mount, reads d_seq
and d_flags of the dentry and only then checks mount_lock. The dentry
can be gone by then. The final mntput() of an unmounted parent unhooks
a child that stayed attached to it under mount_lock alone. When the
child's holder does the final put right after that the root is dput()
and freed immediately while a walker that found the mount hashed still
looks at it.
Refuse to clone a mount with a DCACHE_NORCU dentry as its root. Nothing
sensible can be done with a bind mount of a bpf token anyway.
Fixes: 35f96de04127 ("bpf: Introduce BPF token object")
Cc: stable@vger.kernel.org # v6.9+
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
fs/namespace.c | 4 ++++
1 file changed, 4 insertions(+)
diff --git a/fs/namespace.c b/fs/namespace.c
index 3c90d853e091..fcf42f192aae 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -3026,6 +3026,10 @@ static struct mount *__do_loopback(const struct path *old_path,
if (!may_copy_tree(old_path))
return ERR_PTR(-EINVAL);
+ /* a pseudo dentry is freed without an RCU delay, no walk may find it */
+ if (old_path->dentry->d_flags & DCACHE_NORCU)
+ return ERR_PTR(-EINVAL);
+
if (recurse && !old->mnt_ns)
return ERR_PTR(-EINVAL);
--
2.53.0
^ permalink raw reply related [flat|nested] 22+ messages in thread
* Re: [PATCH 13/17] namespace: remove the fsnotify marks of a mount namespace in process context
2026-09-30 13:32 ` [PATCH 13/17] namespace: remove the fsnotify marks of a mount namespace in process context Christian Brauner
@ 2026-09-30 15:07 ` Amir Goldstein
0 siblings, 0 replies; 22+ messages in thread
From: Amir Goldstein @ 2026-09-30 15:07 UTC (permalink / raw)
To: Christian Brauner
Cc: linux-fsdevel, Linus Torvalds, Chris Mason, Alexander Viro,
Jan Kara, Jeff Layton, Aleksa Sarai, bpf, stable
On Wed, Sep 30, 2026 at 3:32 PM Christian Brauner <brauner@kernel.org> wrote:
>
> mnt_ns_release() drops the last passive reference of a mount namespace
> and removes its fanotify marks via fsnotify_mntns_delete(). That takes
> the mutex of every group with a mark on the namespace and the spinlock
> of the connector. Fine from process context. But mnt_ns_tree_remove()
> hands the reference the namespace was allocated with to call_rcu() and
> so the marks are removed from the RCU softirq:
>
> BUG: sleeping function called from invalid context at kernel/locking/mutex.c:623
> in_atomic(): 1, irqs_disabled(): 0, non_block: 0, pid: 0, name: swapper/3
> __mutex_lock+0x113/0x24b0
> fsnotify_destroy_marks+0x11b/0x3d0
> mnt_ns_release_rcu+0x57/0xa0
> rcu_core+0x6b6/0x1e40
>
> and lockdep complains about the connector lock being taken from softirq
> context. A fanotify group with a FAN_MARK_MNTNS mark on the mount
> namespace of another task is all that's needed. Root in a user
> namespace can do that for a mount namespace it owns. The task exits and
> the namespace is freed with the mark still on it.
>
> Remove the marks in free_mnt_ns() before the namespace is handed to RCU.
> That runs in process context once the last active reference is gone. A
> mark is added through a file descriptor to the namespace which holds an
> active reference so no mark can show up after that.
>
> Fixes: bf630c401641 ("vfs: add notifications for mount attach and detach")
> Cc: stable@vger.kernel.org # v6.15+
> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
> ---
> fs/namespace.c | 3 ++-
> 1 file changed, 2 insertions(+), 1 deletion(-)
>
> diff --git a/fs/namespace.c b/fs/namespace.c
> index 0a7d50db228d..3c90d853e091 100644
> --- a/fs/namespace.c
> +++ b/fs/namespace.c
> @@ -130,7 +130,6 @@ static void mnt_ns_release(struct mnt_namespace *ns)
> {
> /* keep alive for {list,stat}mount() */
> if (ns && refcount_dec_and_test(&ns->passive)) {
> - fsnotify_mntns_delete(ns);
> put_user_ns(ns->user_ns);
> kfree(ns);
> }
> @@ -4279,6 +4278,8 @@ static void free_mnt_ns(struct mnt_namespace *ns)
> if (!is_anon_ns(ns))
> ns_common_free(ns);
> dec_mnt_namespaces(ns->ucounts);
> + /* the last active reference is gone, no mark can show up anymore */
> + fsnotify_mntns_delete(ns);
> mnt_ns_tree_remove(ns);
> }
>
The question should be not if new marks can show up, but whether new
fsnotify_mnt_ events can show up.
I think the patch is still correct, but maybe change this comment or remove it.
Otherwise, feel free to add
Reviewed-by: Amir Goldstein <amir73il@gmail.com>
Thanks,
Amir.
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH 14/17] fsnotify: detach the connector before destroying its marks
2026-09-30 13:32 ` [PATCH 14/17] fsnotify: detach the connector before destroying its marks Christian Brauner
@ 2026-10-01 9:31 ` Christian Brauner
2026-10-01 10:58 ` Amir Goldstein
0 siblings, 1 reply; 22+ messages in thread
From: Christian Brauner @ 2026-10-01 9:31 UTC (permalink / raw)
To: linux-fsdevel
Cc: Linus Torvalds, Chris Mason, Alexander Viro, Jan Kara,
Jeff Layton, Aleksa Sarai, Amir Goldstein, bpf, stable
On Wed, Sep 30, 2026 at 03:32:06PM +0200, Christian Brauner wrote:
> fsnotify_destroy_marks() removes every mark of an object that goes away.
> It walks the mark list of the connector and has to drop the connector
> lock around fsnotify_destroy_mark() since that sleeps. Afterwards it
> continues from the mark it just destroyed. That mark stays on the list
> because the function holds a reference to it. But while the lock is
> dropped fsnotify_add_mark_list() can add a mark to the very same
> connector and put it in front of the current one when group priority
> dictates. fsnotify_grab_connector() hands out the connector until it is
> detached and it only gets detached after the walk.
>
> So the walk never sees that mark. The connector is detached from the
> object with the new mark still attached to it. It never receives an
> event nor IN_IGNORED. For inotify that's a watch descriptor that never
> reports anything and doesn't show up in fdinfo either:
>
> wd1 = inotify_add_watch(/proc/self/fd/6) = 2
> events after the unlink:
> wd 1 mask 0x400 IN_DELETE_SELF
> wd 1 mask 0x8000 IN_IGNORED
> events after write() to the file:
> (none)
>
> Detach the connector from the object before the marks are destroyed.
> fsnotify_grab_connector() refuses a detached connector so a mark added
> from then on gets a connector of its own. Nothing during the destruction
> needs the connector. The inode reference of the connector is dropped
> after the marks are gone as before.
>
> Fixes: 6b3f05d24d35 ("fsnotify: Detach mark from object list when last reference is dropped")
> Cc: stable@vger.kernel.org
> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
> ---
> fs/notify/mark.c | 16 ++++++++++------
> 1 file changed, 10 insertions(+), 6 deletions(-)
>
> diff --git a/fs/notify/mark.c b/fs/notify/mark.c
> index b2640d836a71..f6891d39e42f 100644
> --- a/fs/notify/mark.c
> +++ b/fs/notify/mark.c
> @@ -1112,6 +1112,16 @@ void fsnotify_destroy_marks(fsnotify_connp_t *connp)
> conn = fsnotify_grab_connector(connp);
> if (!conn)
> return;
> + /*
> + * Detach the connector from the object first. Once conn->lock is
> + * dropped a mark could be added in front of the one we're at and the
> + * walk would miss it. fsnotify_grab_connector() refuses a detached
> + * connector so any mark added from now on gets a connector of its
> + * own. This also stops pinning the inode until all mark references
> + * get dropped. It would lead to strange results such as delaying
> + * inode deletion or blocking unmount.
> + */
> + objp = fsnotify_detach_connector_from_object(conn, &type);
Folded as a fixup. We just restart the walk after the drop.
diff --git a/fs/notify/mark.c b/fs/notify/mark.c
index f6891d39e42f..e41774b19e00 100644
--- a/fs/notify/mark.c
+++ b/fs/notify/mark.c
@@ -1113,23 +1113,15 @@ void fsnotify_destroy_marks(fsnotify_connp_t *connp)
if (!conn)
return;
/*
- * Detach the connector from the object first. Once conn->lock is
- * dropped a mark could be added in front of the one we're at and the
- * walk would miss it. fsnotify_grab_connector() refuses a detached
- * connector so any mark added from now on gets a connector of its
- * own. This also stops pinning the inode until all mark references
- * get dropped. It would lead to strange results such as delaying
- * inode deletion or blocking unmount.
- */
- objp = fsnotify_detach_connector_from_object(conn, &type);
- /*
- * We have to be careful since we can race with e.g.
- * fsnotify_clear_marks_by_group() and once we drop the conn->lock, the
- * list can get modified. However we are holding mark reference and
- * thus our mark cannot be removed from obj_list so we can continue
- * iteration after regaining conn->lock.
+ * A mark leaves the list with its last reference so the one we hold
+ * keeps the connector alive. Another mark can be added in front of it
+ * while conn->lock is dropped, so start over after every mark and skip
+ * the ones that are already detached.
*/
+restart:
hlist_for_each_entry(mark, &conn->list, obj_list) {
+ if (!(mark->flags & FSNOTIFY_MARK_FLAG_ATTACHED))
+ continue;
fsnotify_get_mark(mark);
spin_unlock(&conn->lock);
if (old_mark)
@@ -1137,7 +1129,14 @@ void fsnotify_destroy_marks(fsnotify_connp_t *connp)
old_mark = mark;
fsnotify_destroy_mark(mark, mark->group);
spin_lock(&conn->lock);
+ goto restart;
}
+ /*
+ * Detach list from object now so that we don't pin inode until all
+ * mark references get dropped. It would lead to strange results such
+ * as delaying inode deletion or blocking unmount.
+ */
+ objp = fsnotify_detach_connector_from_object(conn, &type);
spin_unlock(&conn->lock);
if (old_mark)
fsnotify_put_mark(old_mark);
^ permalink raw reply related [flat|nested] 22+ messages in thread
* Re: [PATCH 14/17] fsnotify: detach the connector before destroying its marks
2026-10-01 9:31 ` Christian Brauner
@ 2026-10-01 10:58 ` Amir Goldstein
2026-10-01 12:06 ` Christian Brauner
0 siblings, 1 reply; 22+ messages in thread
From: Amir Goldstein @ 2026-10-01 10:58 UTC (permalink / raw)
To: Christian Brauner
Cc: linux-fsdevel, Linus Torvalds, Chris Mason, Alexander Viro,
Jan Kara, Jeff Layton, Aleksa Sarai, bpf, stable, Daehyeon Ko
On Thu, Oct 1, 2026 at 11:31 AM Christian Brauner <brauner@kernel.org> wrote:
>
> On Wed, Sep 30, 2026 at 03:32:06PM +0200, Christian Brauner wrote:
> > fsnotify_destroy_marks() removes every mark of an object that goes away.
> > It walks the mark list of the connector and has to drop the connector
> > lock around fsnotify_destroy_mark() since that sleeps. Afterwards it
> > continues from the mark it just destroyed. That mark stays on the list
> > because the function holds a reference to it. But while the lock is
> > dropped fsnotify_add_mark_list() can add a mark to the very same
> > connector and put it in front of the current one when group priority
> > dictates. fsnotify_grab_connector() hands out the connector until it is
> > detached and it only gets detached after the walk.
> >
> > So the walk never sees that mark. The connector is detached from the
> > object with the new mark still attached to it. It never receives an
> > event nor IN_IGNORED. For inotify that's a watch descriptor that never
> > reports anything and doesn't show up in fdinfo either:
> >
> > wd1 = inotify_add_watch(/proc/self/fd/6) = 2
> > events after the unlink:
> > wd 1 mask 0x400 IN_DELETE_SELF
> > wd 1 mask 0x8000 IN_IGNORED
> > events after write() to the file:
> > (none)
> >
> > Detach the connector from the object before the marks are destroyed.
> > fsnotify_grab_connector() refuses a detached connector so a mark added
> > from then on gets a connector of its own. Nothing during the destruction
> > needs the connector. The inode reference of the connector is dropped
> > after the marks are gone as before.
> >
> > Fixes: 6b3f05d24d35 ("fsnotify: Detach mark from object list when last reference is dropped")
> > Cc: stable@vger.kernel.org
> > Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
> > ---
> > fs/notify/mark.c | 16 ++++++++++------
> > 1 file changed, 10 insertions(+), 6 deletions(-)
> >
> > diff --git a/fs/notify/mark.c b/fs/notify/mark.c
> > index b2640d836a71..f6891d39e42f 100644
> > --- a/fs/notify/mark.c
> > +++ b/fs/notify/mark.c
> > @@ -1112,6 +1112,16 @@ void fsnotify_destroy_marks(fsnotify_connp_t *connp)
> > conn = fsnotify_grab_connector(connp);
> > if (!conn)
> > return;
> > + /*
> > + * Detach the connector from the object first. Once conn->lock is
> > + * dropped a mark could be added in front of the one we're at and the
> > + * walk would miss it. fsnotify_grab_connector() refuses a detached
> > + * connector so any mark added from now on gets a connector of its
> > + * own. This also stops pinning the inode until all mark references
> > + * get dropped. It would lead to strange results such as delaying
> > + * inode deletion or blocking unmount.
> > + */
> > + objp = fsnotify_detach_connector_from_object(conn, &type);
>
> Folded as a fixup. We just restart the walk after the drop.
>
> diff --git a/fs/notify/mark.c b/fs/notify/mark.c
> index f6891d39e42f..e41774b19e00 100644
> --- a/fs/notify/mark.c
> +++ b/fs/notify/mark.c
> @@ -1113,23 +1113,15 @@ void fsnotify_destroy_marks(fsnotify_connp_t *connp)
> if (!conn)
> return;
> /*
> - * Detach the connector from the object first. Once conn->lock is
> - * dropped a mark could be added in front of the one we're at and the
> - * walk would miss it. fsnotify_grab_connector() refuses a detached
> - * connector so any mark added from now on gets a connector of its
> - * own. This also stops pinning the inode until all mark references
> - * get dropped. It would lead to strange results such as delaying
> - * inode deletion or blocking unmount.
> - */
> - objp = fsnotify_detach_connector_from_object(conn, &type);
> - /*
> - * We have to be careful since we can race with e.g.
> - * fsnotify_clear_marks_by_group() and once we drop the conn->lock, the
> - * list can get modified. However we are holding mark reference and
> - * thus our mark cannot be removed from obj_list so we can continue
> - * iteration after regaining conn->lock.
> + * A mark leaves the list with its last reference so the one we hold
> + * keeps the connector alive. Another mark can be added in front of it
> + * while conn->lock is dropped, so start over after every mark and skip
> + * the ones that are already detached.
> */
> +restart:
> hlist_for_each_entry(mark, &conn->list, obj_list) {
> + if (!(mark->flags & FSNOTIFY_MARK_FLAG_ATTACHED))
> + continue;
> fsnotify_get_mark(mark);
> spin_unlock(&conn->lock);
> if (old_mark)
> @@ -1137,7 +1129,14 @@ void fsnotify_destroy_marks(fsnotify_connp_t *connp)
> old_mark = mark;
> fsnotify_destroy_mark(mark, mark->group);
> spin_lock(&conn->lock);
> + goto restart;
> }
Isn't the inotify problem described in the commit message already solved by
669bb5af15d41 ("fs: make sure to call fsnotify_inoderemove() only once when
inode is dead”) on vfs-7.4.misc?
Maybe your fix is needed for fanotify mntns watch and related to the change in
patch 13?
I was thinking about setting a flag FSNOTIFY_CONN_FLAG_OBJ_DEAD at the
start of fsnotify_destroy_marks(), which serves as a barrier against adding
new marks.
That is conceptually the same concept as inode_notify_dead() state but
enforced from within fsnotify code.
Thanks,
Amir.
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH 14/17] fsnotify: detach the connector before destroying its marks
2026-10-01 10:58 ` Amir Goldstein
@ 2026-10-01 12:06 ` Christian Brauner
0 siblings, 0 replies; 22+ messages in thread
From: Christian Brauner @ 2026-10-01 12:06 UTC (permalink / raw)
To: Amir Goldstein
Cc: linux-fsdevel, Linus Torvalds, Chris Mason, Alexander Viro,
Jan Kara, Jeff Layton, Aleksa Sarai, bpf, stable, Daehyeon Ko
> Isn't the inotify problem described in the commit message already solved by
> 669bb5af15d41 ("fs: make sure to call fsnotify_inoderemove() only once when
> inode is dead”) on vfs-7.4.misc?
Indeed, I forgot this patch existed. Good, then we can drop this.
^ permalink raw reply [flat|nested] 22+ messages in thread
end of thread, other threads:[~2026-10-01 12:06 UTC | newest]
Thread overview: 22+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-30 13:31 [PATCH 00/17] mount: more bugfixes, the Oprah edition Christian Brauner
2026-09-30 13:31 ` [PATCH 01/17] namespace: queue a mount only once for mount notifications Christian Brauner
2026-09-30 13:31 ` [PATCH 02/17] namespace: check a submount for references right before unmounting it Christian Brauner
2026-09-30 13:31 ` [PATCH 03/17] selftests/filesystems: check that a busy submount survives a synchronous umount Christian Brauner
2026-09-30 13:31 ` [PATCH 04/17] namespace: check a recursive bind mount for mount namespace loops Christian Brauner
2026-09-30 13:31 ` [PATCH 05/17] selftests/filesystems: check that a recursive bind mount can't pin the caller's mount namespace Christian Brauner
2026-09-30 13:31 ` [PATCH 06/17] namespace: keep covered mounts covered in OPEN_TREE_NAMESPACE Christian Brauner
2026-09-30 13:31 ` [PATCH 07/17] selftests/filesystems: check that OPEN_TREE_NAMESPACE keeps mounts covered Christian Brauner
2026-09-30 13:32 ` [PATCH 08/17] namespace: look at the topmost mount for a mount namespace file Christian Brauner
2026-09-30 13:32 ` [PATCH 09/17] selftests/filesystems: check that a mount namespace file on top doesn't bury a mount Christian Brauner
2026-09-30 13:32 ` [PATCH 10/17] namespace: check the mounts before reading their parents in pivot_root() Christian Brauner
2026-09-30 13:32 ` [PATCH 11/17] namespace: don't reconfigure internal superblocks via remount and umount Christian Brauner
2026-09-30 13:32 ` [PATCH 12/17] selftests/filesystems: check that the nullfs root can't be reconfigured Christian Brauner
2026-09-30 13:32 ` [PATCH 13/17] namespace: remove the fsnotify marks of a mount namespace in process context Christian Brauner
2026-09-30 15:07 ` Amir Goldstein
2026-09-30 13:32 ` [PATCH 14/17] fsnotify: detach the connector before destroying its marks Christian Brauner
2026-10-01 9:31 ` Christian Brauner
2026-10-01 10:58 ` Amir Goldstein
2026-10-01 12:06 ` Christian Brauner
2026-09-30 13:32 ` [PATCH 15/17] dcache: don't put a mountpoint on a dentry that's being removed Christian Brauner
2026-09-30 13:32 ` [PATCH 16/17] unshare: don't drop active namespace references that were never taken Christian Brauner
2026-09-30 13:32 ` [PATCH 17/17] namespace: don't let a pseudo dentry become the root of a mount Christian Brauner
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox