Linux filesystem development
 help / color / mirror / Atom feed
* [PATCH RFC v2 0/8] namespace: prevent UMOUNT_CONNECTED reference count cycles
@ 2026-09-24 22:35 Christian Brauner
  2026-09-24 22:35 ` [PATCH RFC v2 1/8] fs: refuse fspick() on internal superblocks Christian Brauner
                   ` (7 more replies)
  0 siblings, 8 replies; 9+ messages in thread
From: Christian Brauner @ 2026-09-24 22:35 UTC (permalink / raw)
  To: Linus Torvalds
  Cc: Jann Horn, Jan Kara, Amir Goldstein, linux-fsdevel,
	Alexander Viro, Christian Brauner (Amutable)

Have barf bags ready, please. Afaict, UMOUNT_CONNECTED as implemented
allows for the creation of reference count cycles. Here's a simple
example:

  mkdir /x; mkfifo /ready /go
  unshare -m sh -c 'mount -t tmpfs tmpfs /x
                    truncate -s 8M /x/img; mkfs.ext4 -q /x/img
                    dev=$(losetup -f --show /x/img)
                    mkdir /x/mp; mount $dev /x/mp
                    echo $dev > /ready; read r < /go' &
  read dev < /ready
  rmdir /x
  echo > /go
  wait
  losetup -d $dev
  losetup -a

Take a directory /x on the host, create a new mount namespace, mount a
tmpfs on /x, use a file on that tmpfs as the backing file for a loop
device, mount that loop device on that tmpfs. Now rmdir /x on the host.
This will lazily unmount the mount on top of /x in the container with
UMOUNT_CONNECTED. Once the namespace exists nothing references the mount
anymore. Now the tmpfs is pinned by the backing file of the loop device
and the loop mount is owned by the tmpfs superblock.

Fun fact, such cycles can be formed by at least the following
subsystems and I have added reproducers for all of them:

(1) a loop mount P from an image on a tmpfs next to it, so that P's
    death shows as the loop device giving up its backing file

(2) autofs with a FIFO on P as its pipe, zram with a device node on P as
    its writeback device, both on a minix image since vfat has neither

(3) ecryptfs with its lower directory on P, under a passphrase token
    added to the session keyring

(4) binfmt_misc in a new user namespace with an 'F' interpreter on P

(5) a fuse server that answers FUSE_INIT with passthrough on and
    registers a file on P as a backing file

(6) zloop with its zone files in a directory on P

(7) a mass storage gadget on the dummy UDC with its LUN file on P,
    mounted from the SCSI disk the gadget shows up as

(8) md with a RAID1 of one loop device and its bitmap file on P, which
    skips while SET_BITMAP_FILE has no way to succeed

(9) rmdir of P's mountpoint from the parent, then the child exits, then
    the device must be free and LOOP_CLR_FD must release the file

I have explored various solution and have branches for most of them. All
suck ass. Highlights include to port everything to use private mounts
similar to what overlayfs does. It's ugly as fuck and it needs a
side-channel to communicate to umount that the underlying thing like the
loop device is still in use so we don't cause spurious EBUSY errors.
It's really not nice.

The underlying mechanism is UMOUNT_CONNECTED (MNT_LOCKED falls into the
same class). With UMOUNT_CONNECTED an unmounted mount stays attached to
its parent. This is used to protect revealing the underlying mount and
is a non-negotiable security mechanism. So now the parent owns that
mount and is put on the parent's final mntput(). That moves it to
mnt_stuck_children and ultimately it's cleaned up by cleanup_mnt().

The fact that ownership of the child mount gets transferred to the
parent turns every reference from a child's superblock back to one of
its ancestors into a cycle.

Don't let the parent own the children. A mount that stays attached
keeps its own reference. namespace_unlock() drops it parents first.

A subtree that nobody refers to now also collapses exactly like a
disconnected one.

So, the problematic case was always a mount that loses its last
reference while it is still attached. That can reveal the covered
directory and that's caused fun exploits. UMOUNT_CONNECTED was always a
sucky mechanism imho.

Instead of that, when a mount loses its last reference but is still
attached to its parent we know that it is a UMOUNT_CONNECTED or
MNT_LOCKED case. Don't take it out of the hash. Instead make it a nullfs
mount. It's an immutable directory that is part of no namespace. Let
that nullfs mount be owned by the parent and put by the parent's final
mntput() the way locked mounts always were.

Ownership can't form a cycle anymore. A vacant mount is pointing at
knullfs. That's nullfs instance that can't go away and doesn't have any
child mounts whatsoever. Nothing leads from a vacant mount back to any
other mount.

Any lookup that still reaches the parent now finds an empty read-only
directory at the mountpoint. Creating anything in it fails with ENOENT,
like every lookup in nullfs does.

The only visible change is for a locked mount under a lazily unmounted
one. It used to stay alive and traversable for as long as something held
the parent. Now it is released with the umount when it isn't referenced
anymore and an empty directory takes its place. What it covered stays
covered either way.

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
Changes in v2:
- Don't take unnecessary references and simplify freeing of vacated mounts.
- Link to v1: https://patch.msgid.link/20260924-work-mount-knullfs-v1-0-ae89b29f7cb3@kernel.org

---
Christian Brauner (8):
      fs: refuse fspick() on internal superblocks
      fsnotify: record the superblock a connector is accounted on
      fs: put the old fs_struct before the old namespaces in unshare()
      namespace: drop the file's reference first in dissolve_on_fput()
      namespace: prevent UMOUNT_CONNECTED reference count cycles
      selftests/filesystems: check that a loop mount below a dead mount is released
      selftests/filesystems: check the two-step cycle over crossed loop images
      selftests/filesystems: check that the holders let go of a dead mount

 fs/file_table.c                                    |    3 +-
 fs/fsopen.c                                        |    3 +
 fs/mount.h                                         |   13 +-
 fs/namei.c                                         |    4 +
 fs/namespace.c                                     |  186 ++-
 fs/notify/fsnotify.h                               |    3 +-
 fs/notify/mark.c                                   |    7 +-
 include/linux/fsnotify_backend.h                   |    2 +
 kernel/fork.c                                      |    9 +-
 tools/testing/selftests/Makefile                   |    1 +
 .../selftests/filesystems/mount_cycle/.gitignore   |    2 +
 .../selftests/filesystems/mount_cycle/Makefile     |    6 +
 .../filesystems/mount_cycle/loop_cycle_test.c      | 1225 ++++++++++++++++++++
 13 files changed, 1422 insertions(+), 42 deletions(-)
---
base-commit: 2d2a2d7aa98741b58f54cacc99b52024e4d865f9
change-id: 20260923-work-mount-knullfs-4daab4eb9f06


^ permalink raw reply	[flat|nested] 9+ messages in thread

* [PATCH RFC v2 1/8] fs: refuse fspick() on internal superblocks
  2026-09-24 22:35 [PATCH RFC v2 0/8] namespace: prevent UMOUNT_CONNECTED reference count cycles Christian Brauner
@ 2026-09-24 22:35 ` Christian Brauner
  2026-09-24 22:35 ` [PATCH RFC v2 2/8] fsnotify: record the superblock a connector is accounted on Christian Brauner
                   ` (6 subsequent siblings)
  7 siblings, 0 replies; 9+ messages in thread
From: Christian Brauner @ 2026-09-24 22:35 UTC (permalink / raw)
  To: Linus Torvalds
  Cc: Jann Horn, Jan Kara, Amir Goldstein, linux-fsdevel,
	Alexander Viro, Christian Brauner (Amutable)

There's a few instances where a internal superblock is reachable by
fspick(). Block that. graft_tree() and fsmount() already refuse them.
fanotify refuses mount and superblock marks on them for the same reason.

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/fsopen.c | 3 +++
 1 file changed, 3 insertions(+)

diff --git a/fs/fsopen.c b/fs/fsopen.c
index ae19e5136598..9d5a7a22b529 100644
--- a/fs/fsopen.c
+++ b/fs/fsopen.c
@@ -190,6 +190,9 @@ SYSCALL_DEFINE3(fspick, int, dfd, const char __user *, path, unsigned int, flags
 	ret = -EINVAL;
 	if (target.mnt->mnt_root != target.dentry)
 		goto err_path;
+	/* kernel-internal superblocks are nobody's to reconfigure */
+	if (target.dentry->d_sb->s_flags & SB_NOUSER)
+		goto err_path;
 
 	fc = fs_context_for_reconfigure(target.dentry, 0, 0);
 	if (IS_ERR(fc)) {

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 9+ messages in thread

* [PATCH RFC v2 2/8] fsnotify: record the superblock a connector is accounted on
  2026-09-24 22:35 [PATCH RFC v2 0/8] namespace: prevent UMOUNT_CONNECTED reference count cycles Christian Brauner
  2026-09-24 22:35 ` [PATCH RFC v2 1/8] fs: refuse fspick() on internal superblocks Christian Brauner
@ 2026-09-24 22:35 ` Christian Brauner
  2026-09-24 22:35 ` [PATCH RFC v2 3/8] fs: put the old fs_struct before the old namespaces in unshare() Christian Brauner
                   ` (5 subsequent siblings)
  7 siblings, 0 replies; 9+ messages in thread
From: Christian Brauner @ 2026-09-24 22:35 UTC (permalink / raw)
  To: Linus Torvalds
  Cc: Jann Horn, Jan Kara, Amir Goldstein, linux-fsdevel,
	Alexander Viro, Christian Brauner (Amutable)

fsnotify keeps a count of watched objects per superblock so that the
event hooks can be skipped on a superblock nobody watches. It computes
the superblock from the object object it watches via
fsnotify_object_sb().

This currently works because an object's superblock doesn't change while
a connector is attached to it. This changes form some of them. So record
the superblock in the connector when it is created and use
that for the accounting.

No functional changes.

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/notify/fsnotify.h             | 3 ++-
 fs/notify/mark.c                 | 7 +++++--
 include/linux/fsnotify_backend.h | 2 ++
 3 files changed, 9 insertions(+), 3 deletions(-)

diff --git a/fs/notify/fsnotify.h b/fs/notify/fsnotify.h
index 58c7bb25e571..0be351275dad 100644
--- a/fs/notify/fsnotify.h
+++ b/fs/notify/fsnotify.h
@@ -54,10 +54,11 @@ static inline struct super_block *fsnotify_object_sb(void *obj,
 	}
 }
 
+/* The sb the connector is accounted on; NULL once it has been detached */
 static inline struct super_block *fsnotify_connector_sb(
 				struct fsnotify_mark_connector *conn)
 {
-	return fsnotify_object_sb(conn->obj, conn->type);
+	return conn->sb;
 }
 
 static inline fsnotify_connp_t *fsnotify_sb_marks(struct super_block *sb)
diff --git a/fs/notify/mark.c b/fs/notify/mark.c
index b2640d836a71..1b6095fcf039 100644
--- a/fs/notify/mark.c
+++ b/fs/notify/mark.c
@@ -425,6 +425,7 @@ static void *fsnotify_detach_connector_from_object(
 	conn->type = FSNOTIFY_OBJ_TYPE_DETACHED;
 	if (sb)
 		fsnotify_update_sb_watchers(sb, conn);
+	conn->sb = NULL;
 
 	return inode;
 }
@@ -791,6 +792,8 @@ static void fsnotify_init_connector(struct fsnotify_mark_connector *conn,
 	conn->prio = 0;
 	conn->type = obj_type;
 	conn->obj = obj;
+	/* the object may move to another sb, the accounting doesn't */
+	conn->sb = fsnotify_object_sb(obj, obj_type);
 }
 
 static struct fsnotify_mark_connector *
@@ -953,8 +956,8 @@ static int fsnotify_add_mark_list(struct fsnotify_mark *mark, void *obj,
 	/* mark should be the last entry.  last is the current last entry */
 	hlist_add_behind_rcu(&mark->obj_list, &last->obj_list);
 added:
-	if (sb)
-		fsnotify_update_sb_watchers(sb, conn);
+	if (conn->sb)
+		fsnotify_update_sb_watchers(conn->sb, conn);
 	/*
 	 * Since connector is attached to object using cmpxchg() we are
 	 * guaranteed that connector initialization is fully visible by anyone
diff --git a/include/linux/fsnotify_backend.h b/include/linux/fsnotify_backend.h
index 618eed4d6d72..3a9ab9e0104b 100644
--- a/include/linux/fsnotify_backend.h
+++ b/include/linux/fsnotify_backend.h
@@ -573,6 +573,8 @@ struct fsnotify_mark_connector {
 		/* Used listing heads to free after srcu period expires */
 		struct fsnotify_mark_connector *destroy_next;
 	};
+	/* sb whose watched_objects account for this connector [lock] */
+	struct super_block *sb;
 	struct hlist_head list;	/* List of marks */
 };
 

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 9+ messages in thread

* [PATCH RFC v2 3/8] fs: put the old fs_struct before the old namespaces in unshare()
  2026-09-24 22:35 [PATCH RFC v2 0/8] namespace: prevent UMOUNT_CONNECTED reference count cycles Christian Brauner
  2026-09-24 22:35 ` [PATCH RFC v2 1/8] fs: refuse fspick() on internal superblocks Christian Brauner
  2026-09-24 22:35 ` [PATCH RFC v2 2/8] fsnotify: record the superblock a connector is accounted on Christian Brauner
@ 2026-09-24 22:35 ` Christian Brauner
  2026-09-24 22:35 ` [PATCH RFC v2 4/8] namespace: drop the file's reference first in dissolve_on_fput() Christian Brauner
                   ` (4 subsequent siblings)
  7 siblings, 0 replies; 9+ messages in thread
From: Christian Brauner @ 2026-09-24 22:35 UTC (permalink / raw)
  To: Linus Torvalds
  Cc: Jann Horn, Jan Kara, Amir Goldstein, linux-fsdevel,
	Alexander Viro, Christian Brauner (Amutable)

unshare(CLONE_NEWNS) always copies the fs_struct and copy_mnt_ns()
points root and pwd of that copy into the new mount namespace. The old
fs_struct pins the old root and pwd until ksys_unshare() frees it at the
end and switch_task_namespaces() will already put the old mount
namespace before.

This causes pointless work in user namespaces because the copied mounts
are locked and stay connected. Switch and free the fs_struct before
switching the namespaces. The old root and pwd are wasted while their
mount namespace is still alive which also keeps both puts on the
mntput() fast path.

setns() already does this right as it updates root and pwd before the
namespaces are switched.

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 kernel/fork.c | 9 ++++++---
 1 file changed, 6 insertions(+), 3 deletions(-)

diff --git a/kernel/fork.c b/kernel/fork.c
index 416758c8a3d4..da48168c504f 100644
--- a/kernel/fork.c
+++ b/kernel/fork.c
@@ -3312,14 +3312,17 @@ int ksys_unshare(unsigned long unshare_flags)
 			shm_init_task(current);
 		}
 
+		if (new_fs) {
+			new_fs = switch_fs_struct(new_fs);
+			if (new_fs)
+				free_fs_struct(no_free_ptr(new_fs));
+		}
+
 		if (new_nsproxy) {
 			switch_task_namespaces(current, new_nsproxy);
 			new_nsproxy = NULL;
 		}
 
-		if (new_fs)
-			new_fs = switch_fs_struct(new_fs);
-
 		if (new_fd) {
 			guard(task_lock)(current);
 			swap(current->files, new_fd);

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 9+ messages in thread

* [PATCH RFC v2 4/8] namespace: drop the file's reference first in dissolve_on_fput()
  2026-09-24 22:35 [PATCH RFC v2 0/8] namespace: prevent UMOUNT_CONNECTED reference count cycles Christian Brauner
                   ` (2 preceding siblings ...)
  2026-09-24 22:35 ` [PATCH RFC v2 3/8] fs: put the old fs_struct before the old namespaces in unshare() Christian Brauner
@ 2026-09-24 22:35 ` Christian Brauner
  2026-09-24 22:35 ` [PATCH RFC v2 5/8] namespace: prevent UMOUNT_CONNECTED reference count cycles Christian Brauner
                   ` (3 subsequent siblings)
  7 siblings, 0 replies; 9+ messages in thread
From: Christian Brauner @ 2026-09-24 22:35 UTC (permalink / raw)
  To: Linus Torvalds
  Cc: Jann Horn, Jan Kara, Amir Goldstein, linux-fsdevel,
	Alexander Viro, Christian Brauner (Amutable)

A detached mount tree is held alive by a file. dissolve_on_fput()
unmounts the detached tree attached to that file.

Currently dissolve_on_fput() leaves the file's reference to the root of
the tree to the caller. So either __fput(), open_detached_copy(), or
fsmount() drop the mount.

That's backwards. The reference keeps the tree around and is dropped
when the tree is already gone. Instead, let dissolve_on_fput() the mount
and drop it after umount_tree() with namespace_sem still held. After
umount_tree() the mntput() is never the last reference and so the root
gets unmounted before its children.

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/file_table.c | 3 ++-
 fs/namespace.c  | 9 ++++++---
 2 files changed, 8 insertions(+), 4 deletions(-)

diff --git a/fs/file_table.c b/fs/file_table.c
index c68b8c0a4097..e2665eaa6f8c 100644
--- a/fs/file_table.c
+++ b/fs/file_table.c
@@ -520,7 +520,8 @@ static void __fput(struct file *file)
 	dput(dentry);
 	if (unlikely(mode & FMODE_NEED_UNMOUNT))
 		dissolve_on_fput(mnt);
-	mntput(mnt);
+	else
+		mntput(mnt);
 out:
 	file_free(file);
 }
diff --git a/fs/namespace.c b/fs/namespace.c
index 580877e46b1a..4ff3cebe6c31 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -2294,9 +2294,11 @@ void drop_collected_paths(const struct path *paths, const struct path *prealloc)
 
 static struct mnt_namespace *alloc_mnt_ns(struct user_namespace *, bool);
 
+/* Consumes the caller's reference to @mnt. */
 void dissolve_on_fput(struct vfsmount *mnt)
 {
-	struct mount *m = real_mount(mnt);
+	struct vfsmount *p __free(mntput) = mnt;
+	struct mount *m = real_mount(p);
 
 	/*
 	 * m used to be the root of anon namespace; if it still is one,
@@ -2321,6 +2323,7 @@ void dissolve_on_fput(struct vfsmount *mnt)
 		lock_mount_hash();
 		umount_tree(m, UMOUNT_CONNECTED);
 		unlock_mount_hash();
+		mntput(no_free_ptr(p));
 	}
 }
 
@@ -3088,7 +3091,7 @@ static struct file *open_detached_copy(struct path *path, unsigned int flags)
 	path->mnt = mntget(&ns->root->mnt);
 	file = dentry_open(path, O_PATH, current_cred());
 	if (IS_ERR(file))
-		dissolve_on_fput(path->mnt);
+		dissolve_on_fput(no_free_ptr(path->mnt));
 	else
 		file->f_mode |= FMODE_NEED_UNMOUNT;
 	return file;
@@ -4545,7 +4548,7 @@ SYSCALL_DEFINE3(fsmount, int, fs_fd, unsigned int, flags,
 	FD_PREPARE(fdf, (flags & FSMOUNT_CLOEXEC) ? O_CLOEXEC : 0,
 		   dentry_open(&new_path, O_PATH, fc->cred));
 	if (fdf.err) {
-		dissolve_on_fput(new_path.mnt);
+		dissolve_on_fput(no_free_ptr(new_path.mnt));
 		return fdf.err;
 	}
 

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 9+ messages in thread

* [PATCH RFC v2 5/8] namespace: prevent UMOUNT_CONNECTED reference count cycles
  2026-09-24 22:35 [PATCH RFC v2 0/8] namespace: prevent UMOUNT_CONNECTED reference count cycles Christian Brauner
                   ` (3 preceding siblings ...)
  2026-09-24 22:35 ` [PATCH RFC v2 4/8] namespace: drop the file's reference first in dissolve_on_fput() Christian Brauner
@ 2026-09-24 22:35 ` Christian Brauner
  2026-09-24 22:35 ` [PATCH RFC v2 6/8] selftests/filesystems: check that a loop mount below a dead mount is released Christian Brauner
                   ` (2 subsequent siblings)
  7 siblings, 0 replies; 9+ messages in thread
From: Christian Brauner @ 2026-09-24 22:35 UTC (permalink / raw)
  To: Linus Torvalds
  Cc: Jann Horn, Jan Kara, Amir Goldstein, linux-fsdevel,
	Alexander Viro, Christian Brauner (Amutable)

Have barf bags ready, please. UMOUNT_CONNECTED as implemented allows for
the creation of reference count cycles. Here's a simple example

  mkdir /x; mkfifo /ready /go
  unshare -m sh -c 'mount -t tmpfs tmpfs /x
                    truncate -s 8M /x/img; mkfs.ext4 -q /x/img
                    dev=$(losetup -f --show /x/img)
                    mkdir /x/mp; mount $dev /x/mp
                    echo $dev > /ready; read r < /go' &
  read dev < /ready
  rmdir /x
  echo > /go
  wait
  losetup -d $dev
  losetup -a

Take a directory /x on the host, create a new mount namespace, mount a
tmpfs on /x, use a file on that tmpfs as the backing file for a loop
device, mount that loop device on that tmpfs. Now rmdir /x on the host.
This will lazily unmount the mount on top of /x in the container with
UMOUNT_CONNECTED. Once the namespace exists nothing references the mount
anymore. Now the tmpfs is pinned by the backing file of the loop device
and the loop mount is owned by the tmpfs superblock.

Fun fact, such cycles can be formed by at least the following
subsystems and I have added reproducers for all of them:

(1) a loop mount P from an image on a tmpfs next to it, so that P's
    death shows as the loop device giving up its backing file

(2) autofs with a FIFO on P as its pipe, zram with a device node on P as
    its writeback device, both on a minix image since vfat has neither

(3) ecryptfs with its lower directory on P, under a passphrase token
    added to the session keyring

(4) binfmt_misc in a new user namespace with an 'F' interpreter on P

(5) a fuse server that answers FUSE_INIT with passthrough on and
    registers a file on P as a backing file

(6) zloop with its zone files in a directory on P

(7) a mass storage gadget on the dummy UDC with its LUN file on P,
    mounted from the SCSI disk the gadget shows up as

(8) md with a RAID1 of one loop device and its bitmap file on P, which
    skips while SET_BITMAP_FILE has no way to succeed

(9) rmdir of P's mountpoint from the parent, then the child exits, then
    the device must be free and LOOP_CLR_FD must release the file

The underlying mechanism is UMOUNT_CONNECTED (MNT_LOCKED falls into the
same class). With UMOUNT_CONNECTED an unmounted mount stays attached to
its parent. This is used to protect revealing the underlying mount and
is a non-negotiable security mechanism. So now the parent owns that
mount and is put on the parent's final mntput(). That moves it to
mnt_stuck_children and ultimately it's cleaned up by cleanup_mnt().

The fact that ownership of the child mount gets transferred to the
parent turns every reference from a child's superblock back to one of
its ancestors into a cycle.

Don't let the parent own the children. A mount that stays attached
keeps its own reference. namespace_unlock() drops it parents first.

A subtree that nobody refers to now also collapses exactly like a
disconnected one.

So, the problematic case was always a mount that loses its last
reference while it is still attached. That can reveal the covered
directory and that's caused fun exploits. UMOUNT_CONNECTED was always a
sucky mechanism imho.

Instead of that, when a mount loses its last reference but is still
attached to its parent we know that it is a UMOUNT_CONNECTED or
MNT_LOCKED case. Don't take it out of the hash. Instead make it a nullfs
mount. It's an immutable directory that is part of no namespace. Let
that nullfs mount be owned by the parent and put by the parent's final
mntput() the way locked mounts always were.

Ownership can't form a cycle anymore. A vacant mount is pointing at
knullfs. That's nullfs instance that can't go away and doesn't have any
child mounts whatsoever. Nothing leads from a vacant mount back to any
other mount.

Any lookup that still reaches the parent now finds an empty read-only
directory at the mountpoint. Creating anything in it fails with ENOENT,
like every lookup in nullfs does.

The only visible change is for a locked mount under a lazily unmounted
one. It used to stay alive and traversable for as long as something held
the parent. Now it is released with the umount when it isn't referenced
anymore and an empty directory takes its place. What it covered stays
covered either way.

Link: https://gist.github.com/mvo5/63ef46482349f3b1c3957d463a0c9c6f
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/mount.h     |  13 ++++-
 fs/namei.c     |   4 ++
 fs/namespace.c | 177 +++++++++++++++++++++++++++++++++++++++++++++++----------
 3 files changed, 162 insertions(+), 32 deletions(-)

diff --git a/fs/mount.h b/fs/mount.h
index 94fcc306d21e..90c32d097673 100644
--- a/fs/mount.h
+++ b/fs/mount.h
@@ -6,6 +6,7 @@
 #include <linux/fs_pin.h>
 
 extern struct file_system_type nullfs_fs_type;
+extern struct vfsmount *knullfs;
 extern struct list_head notify_list;
 
 struct mnt_namespace {
@@ -50,8 +51,14 @@ struct mount {
 	struct vfsmount mnt;
 	union {
 		struct rb_node mnt_node; /* node in the ns->mounts rbtree */
-		struct rcu_head mnt_rcu;
-		struct llist_node mnt_llist;
+		struct {		 /* once it has left its namespace */
+			union {
+				struct rcu_head mnt_rcu;
+				struct llist_node mnt_llist;
+			};
+			/* what a vacant mount stands in for, NULL once released */
+			struct dentry *mnt_old_root;
+		};
 	};
 #ifdef CONFIG_SMP
 	struct mnt_pcp __percpu *mnt_pcp;
@@ -84,7 +91,7 @@ struct mount {
 	struct mountpoint *mnt_mp;	/* where is it mounted */
 	union {
 		struct hlist_node mnt_mp_list;	/* list mounts with the same mountpoint */
-		struct hlist_node mnt_umount;
+		struct hlist_node mnt_umount;	/* in the parent's mnt_stuck_children */
 	};
 #ifdef CONFIG_FSNOTIFY
 	struct fsnotify_mark_connector __rcu *mnt_fsnotify_marks;
diff --git a/fs/namei.c b/fs/namei.c
index 20a6534ea3ef..aca50e2bc42d 100644
--- a/fs/namei.c
+++ b/fs/namei.c
@@ -816,6 +816,10 @@ static bool path_connected(struct vfsmount *mnt, struct dentry *dentry)
 {
 	struct super_block *sb = mnt->mnt_sb;
 
+	/* @mnt was vacated after an RCU walk found @dentry on it */
+	if (unlikely(dentry->d_sb != sb))
+		return false;
+
 	/* Bind mounts can have disconnected paths */
 	if (mnt->mnt_root == sb->s_root)
 		return true;
diff --git a/fs/namespace.c b/fs/namespace.c
index 4ff3cebe6c31..aefa4b00d676 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -80,8 +80,9 @@ static u64 mnt_id_ctr = MNT_UNIQUE_ID_OFFSET;
 static struct hlist_head *mount_hashtable __ro_after_init;
 static struct hlist_head *mountpoint_hashtable __ro_after_init;
 static struct kmem_cache *mnt_cache __ro_after_init;
+struct vfsmount *knullfs __ro_after_init;	/* private nullfs instance */
 static DECLARE_RWSEM(namespace_sem);
-static HLIST_HEAD(unmounted);	/* protected by namespace_sem */
+static LIST_HEAD(unmounted);	/* protected by namespace_sem */
 static LIST_HEAD(ex_mountpoints); /* protected by namespace_sem */
 static struct mnt_namespace *emptied_ns; /* protected by namespace_sem */
 
@@ -227,6 +228,17 @@ static void mnt_free_id(struct mount *mnt)
 	xa_erase(&mnt_id_xa, mnt->mnt_id);
 }
 
+/* mnt_id_ctr is protected by the lock of mnt_id_xa, see mnt_alloc_id(). */
+static u64 mnt_alloc_unique_id(void)
+{
+	u64 id;
+
+	xa_lock(&mnt_id_xa);
+	id = ++mnt_id_ctr;
+	xa_unlock(&mnt_id_xa);
+	return id;
+}
+
 /*
  * Allocate a new peer group ID
  */
@@ -1294,8 +1306,57 @@ static struct mount *clone_mnt(struct mount *old, struct dentry *root,
 	return ERR_PTR(err);
 }
 
+/* What it stood in for is gone already and it holds nothing on knullfs. */
+static void free_vacant_mount(struct mount *mnt)
+{
+	if (!mnt)
+		return;
+
+	VFS_WARN_ON_ONCE(mnt->mnt_pins.first);
+	VFS_WARN_ON_ONCE(!hlist_empty(&mnt->mnt_stuck_children));
+#ifdef CONFIG_FSNOTIFY
+	VFS_WARN_ON_ONCE(rcu_access_pointer(mnt->mnt_fsnotify_marks));
+#endif
+	mnt_free_id(mnt);
+	call_rcu(&mnt->mnt_rcu, delayed_free_vfsmnt);
+}
+
+/* The release is done. Whoever comes second, this or the last put, frees it. */
+static inline void vacant_mount_released(struct mount *mnt)
+{
+	scoped_guard(mount_locked_reader) {
+		mnt->mnt_old_root = NULL;
+		if (!(mnt->mnt.mnt_flags & MNT_DOOMED))
+			mnt = NULL;
+	}
+	free_vacant_mount(mnt);
+}
+
+/* The last put is done. Whoever comes second, this or the release, frees it. */
+static inline struct mount *vacant_mount_put(struct mount *mnt)
+{
+	VFS_WARN_ON_ONCE(mnt_has_parent(mnt));
+	mnt->mnt.mnt_flags |= MNT_DOOMED;
+	mnt_del_instance(mnt);
+	if (mnt->mnt_old_root)
+		return NULL;
+	return mnt;
+}
+
+/*
+ * The only mounts of knullfs' superblock that are ever attached are
+ * vacant mounts.
+ */
+static inline bool is_vacant(const struct mount *mnt)
+{
+	return mnt->mnt.mnt_sb == knullfs->mnt_sb;
+}
+
 static void cleanup_mnt(struct mount *mnt)
 {
+	bool release = is_vacant(mnt);
+	struct dentry *root = release ? mnt->mnt_old_root : mnt->mnt.mnt_root;
+	struct super_block *sb = root->d_sb;
 	struct hlist_node *p;
 	struct mount *m;
 	/*
@@ -1303,9 +1364,10 @@ static void cleanup_mnt(struct mount *mnt)
 	 * up a mnt_want/drop_write() pair.  If this happens, the
 	 * filesystem was probably unable to make r/w->r/o transitions.
 	 * The locking used to deal with mnt_count decrement provides barriers,
-	 * so mnt_get_writers() below is safe.
+	 * so mnt_get_writers() below is safe. A vacant mount is still reachable
+	 * and a failing mnt_want_write() on it bumps the count for a moment.
 	 */
-	WARN_ON(mnt_get_writers(mnt));
+	WARN_ON(!release && mnt_get_writers(mnt));
 	if (unlikely(mnt->mnt_pins.first))
 		mnt_pin_kill(mnt);
 	hlist_for_each_entry_safe(m, p, &mnt->mnt_stuck_children, mnt_umount) {
@@ -1313,10 +1375,35 @@ static void cleanup_mnt(struct mount *mnt)
 		mntput(&m->mnt);
 	}
 	fsnotify_vfsmount_delete(&mnt->mnt);
-	dput(mnt->mnt.mnt_root);
-	deactivate_super(mnt->mnt.mnt_sb);
-	mnt_free_id(mnt);
-	call_rcu(&mnt->mnt_rcu, delayed_free_vfsmnt);
+	dput(root);
+	deactivate_super(sb);
+
+	if (unlikely(release)) {
+		vacant_mount_released(mnt);
+	} else {
+		mnt_free_id(mnt);
+		call_rcu(&mnt->mnt_rcu, delayed_free_vfsmnt);
+	}
+}
+
+/*
+ * @mnt lost its last reference while still attached. Don't reveal what's
+ * beneath so mount knullfs over it.
+ */
+static void vacate_mount(struct mount *mnt)
+{
+	/* a vacant mount is put only after it has been unhashed */
+	VFS_WARN_ON_ONCE(is_vacant(mnt));
+	mnt->mnt_old_root = mnt->mnt.mnt_root;
+	/* knullfs never goes away, a vacant mount holds nothing on it */
+	mnt->mnt.mnt_sb = knullfs->mnt_sb;
+	mnt->mnt.mnt_root = knullfs->mnt_root;
+	mnt_add_instance(mnt, knullfs->mnt_sb);
+	mnt->mnt.mnt_flags |= MNT_READONLY;
+	/* It's a new mount so give it its own mount id. */
+	mnt->mnt_id_unique = mnt_alloc_unique_id();
+	/* One reference for the parent, the release is pending until it's done. */
+	mnt_add_count(mnt, 1);
 }
 
 static void __cleanup_mnt(struct rcu_head *head)
@@ -1338,6 +1425,7 @@ static DECLARE_DELAYED_WORK(delayed_mntput_work, delayed_mntput);
 static void noinline mntput_no_expire_slowpath(struct mount *mnt)
 {
 	LIST_HEAD(list);
+	bool connected = false;
 	int count;
 
 	VFS_BUG_ON(mnt->mnt_ns);
@@ -1360,7 +1448,19 @@ static void noinline mntput_no_expire_slowpath(struct mount *mnt)
 		unlock_mount_hash();
 		return;
 	}
-	mnt->mnt.mnt_flags |= MNT_DOOMED;
+	/* The last put of a vacant mount. */
+	if (unlikely(is_vacant(mnt))) {
+		mnt = vacant_mount_put(mnt);
+		rcu_read_unlock();
+		unlock_mount_hash();
+		free_vacant_mount(mnt);
+		return;
+	}
+	/* Still attached, so it held its own reference: keep the slot filled. */
+	if (unlikely(mnt_has_parent(mnt)))
+		connected = true;
+	else
+		mnt->mnt.mnt_flags |= MNT_DOOMED;
 	rcu_read_unlock();
 
 	mnt_del_instance(mnt);
@@ -1371,9 +1471,13 @@ static void noinline mntput_no_expire_slowpath(struct mount *mnt)
 		struct mount *p, *tmp;
 		list_for_each_entry_safe(p, tmp, &mnt->mnt_mounts,  mnt_child) {
 			__umount_mnt(p, &list);
-			hlist_add_head(&p->mnt_umount, &mnt->mnt_stuck_children);
+			/* only a vacant mount is owned by its parent */
+			if (is_vacant(p))
+				hlist_add_head(&p->mnt_umount, &mnt->mnt_stuck_children);
 		}
 	}
+	if (unlikely(connected))
+		vacate_mount(mnt);
 	unlock_mount_hash();
 	shrink_dentry_list(&list);
 
@@ -1688,13 +1792,12 @@ static bool need_notify_mnt_list(void)
 static void free_mnt_ns(struct mnt_namespace *);
 static void namespace_unlock(void)
 {
-	struct hlist_head head;
-	struct hlist_node *p;
-	struct mount *m;
+	struct mount *m, *n;
 	struct mnt_namespace *ns = emptied_ns;
+	LIST_HEAD(head);
 	LIST_HEAD(list);
 
-	hlist_move_list(&unmounted, &head);
+	list_splice_init(&unmounted, &head);
 	list_splice_init(&ex_mountpoints, &list);
 	emptied_ns = NULL;
 
@@ -1718,13 +1821,14 @@ static void namespace_unlock(void)
 
 	shrink_dentry_list(&list);
 
-	if (likely(hlist_empty(&head)))
+	if (likely(list_empty(&head)))
 		return;
 
 	synchronize_rcu_expedited();
 
-	hlist_for_each_entry_safe(m, p, &head, mnt_umount) {
-		hlist_del(&m->mnt_umount);
+	/* In tree order, so a subtree nobody holds goes without vacant mounts. */
+	list_for_each_entry_safe(m, n, &head, mnt_list) {
+		list_del_init(&m->mnt_list);
 		mntput(&m->mnt);
 	}
 }
@@ -1740,6 +1844,19 @@ enum umount_tree_flags {
 	UMOUNT_CONNECTED = 4,
 };
 
+/*
+ * An unmounted mount that stays attached to its unmounted parent:
+ *
+ *  - is hashed and on the parent's list of children, so a walk on the
+ *    parent finds it and never ends up in what it covered
+ *  - is unreachable from any namespace root
+ *  - holds its own reference, dropped by namespace_unlock(); the parent's
+ *    final mntput() unhashes it without putting it, so a child whose
+ *    superblock pins an ancestor can still shut down
+ *  - is vacated if it loses its last reference while attached
+ *    it releases the filesystem it carried and a knullfs mount is placed on
+ *    the parent which is owned by it
+ */
 static bool disconnect_mount(struct mount *mnt, enum umount_tree_flags how)
 {
 	/* Leaving mounts connected is only valid for lazy umounts */
@@ -1750,10 +1867,7 @@ static bool disconnect_mount(struct mount *mnt, enum umount_tree_flags how)
 	if (!mnt_has_parent(mnt))
 		return true;
 
-	/* Because the reference counting rules change when mounts are
-	 * unmounted and connected, umounted mounts may not be
-	 * connected to mounted mounts.
-	 */
+	/* An unmounted mount may only stay attached to an unmounted parent */
 	if (!(mnt->mnt_parent->mnt.mnt_flags & MNT_UMOUNT))
 		return true;
 
@@ -1824,8 +1938,8 @@ static void umount_tree(struct mount *mnt, enum umount_tree_flags how)
 				umount_mnt(p);
 			}
 		}
-		if (disconnect)
-			hlist_add_head(&p->mnt_umount, &unmounted);
+		/* attached or not, it holds its own reference */
+		list_add_tail(&p->mnt_list, &unmounted);
 
 		/*
 		 * At this point p->mnt_ns is NULL, notification will be queued
@@ -1994,9 +2108,12 @@ void __detach_mounts(struct dentry *dentry)
 		mnt = hlist_entry(mp.node.next, struct mount, mnt_mp_list);
 		if (mnt->mnt.mnt_flags & MNT_UMOUNT) {
 			umount_mnt(mnt);
-			hlist_add_head(&mnt->mnt_umount, &unmounted);
+			/* only a vacant mount is owned by its parent */
+			if (is_vacant(mnt))
+				list_add_tail(&mnt->mnt_list, &unmounted);
+		} else {
+			umount_tree(mnt, UMOUNT_CONNECTED);
 		}
-		else umount_tree(mnt, UMOUNT_CONNECTED);
 	}
 	unpin_mountpoint(&mp);
 }
@@ -6213,10 +6330,12 @@ static void __init init_mount_tree(void)
 	 *
 	 * (1) nullfs with mount id 1
 	 * (2) mutable rootfs with mount id 2
-	 * (3) private nullfs for kthreads (SB_KERNMOUNT)
+	 * (3) private nullfs for kthreads (SB_KERNMOUNT), kept in knullfs
 	 *
 	 * with (2) mounted on top of (1). The init_task's root and pwd
 	 * are pointed at (3) so all kthreads start isolated in nullfs.
+	 * A mount that dies while still attached to its parent is pointed
+	 * at (3) as well, see vacate_mount().
 	 */
 	nullfs_mnt = vfs_kern_mount(&nullfs_fs_type, 0, "nullfs", NULL);
 	if (IS_ERR(nullfs_mnt))
@@ -6248,11 +6367,11 @@ static void __init init_mount_tree(void)
 		init_mnt_ns.nr_mounts++;
 	}
 
-	nullfs_mnt = kern_mount(&nullfs_fs_type);
-	if (IS_ERR(nullfs_mnt))
+	knullfs = kern_mount(&nullfs_fs_type);
+	if (IS_ERR(knullfs))
 		panic("VFS: Failed to create private nullfs instance");
-	root.mnt	= nullfs_mnt;
-	root.dentry	= nullfs_mnt->mnt_root;
+	root.mnt	= knullfs;
+	root.dentry	= knullfs->mnt_root;
 
 	init_task.nsproxy->mnt_ns = &init_mnt_ns;
 	get_mnt_ns(&init_mnt_ns);

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 9+ messages in thread

* [PATCH RFC v2 6/8] selftests/filesystems: check that a loop mount below a dead mount is released
  2026-09-24 22:35 [PATCH RFC v2 0/8] namespace: prevent UMOUNT_CONNECTED reference count cycles Christian Brauner
                   ` (4 preceding siblings ...)
  2026-09-24 22:35 ` [PATCH RFC v2 5/8] namespace: prevent UMOUNT_CONNECTED reference count cycles Christian Brauner
@ 2026-09-24 22:35 ` Christian Brauner
  2026-09-24 22:35 ` [PATCH RFC v2 7/8] selftests/filesystems: check the two-step cycle over crossed loop images Christian Brauner
  2026-09-24 22:35 ` [PATCH RFC v2 8/8] selftests/filesystems: check that the holders let go of a dead mount Christian Brauner
  7 siblings, 0 replies; 9+ messages in thread
From: Christian Brauner @ 2026-09-24 22:35 UTC (permalink / raw)
  To: Linus Torvalds
  Cc: Jann Horn, Jan Kara, Amir Goldstein, linux-fsdevel,
	Alexander Viro, Christian Brauner (Amutable)

A filesystem image on a mount and the loop device mounted below it: the
image pins the mount and the loop mount pins the image. Cover the two
ways the mount above dies with the loop mount left connected:

- rmdir of the mountpoint from another mount namespace

- the last close of a detached tree that holds both, with the image
  opened through the tree and the loop mount moved below it

Neither relies on the mount namespace itself going away, so both stay
valid once put_mnt_ns() disconnects the mounts of a dying namespace
again.

Check in both that no mount namespace shows the device afterwards, that
an exclusive open of the device succeeds, and that LOOP_CLR_FD gives the
backing file up rather than only arming autoclear.

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 tools/testing/selftests/Makefile                   |   1 +
 .../selftests/filesystems/mount_cycle/.gitignore   |   2 +
 .../selftests/filesystems/mount_cycle/Makefile     |   6 +
 .../filesystems/mount_cycle/loop_cycle_test.c      | 377 +++++++++++++++++++++
 4 files changed, 386 insertions(+)

diff --git a/tools/testing/selftests/Makefile b/tools/testing/selftests/Makefile
index 273853937c25..fb3a85d6396d 100644
--- a/tools/testing/selftests/Makefile
+++ b/tools/testing/selftests/Makefile
@@ -43,6 +43,7 @@ TARGETS += filesystems/open_tree_ns
 TARGETS += filesystems/overlayfs
 TARGETS += filesystems/statmount
 TARGETS += filesystems/mount-notify
+TARGETS += filesystems/mount_cycle
 TARGETS += filesystems/nsfs
 TARGETS += filesystems/fuse
 TARGETS += filesystems/move_mount
diff --git a/tools/testing/selftests/filesystems/mount_cycle/.gitignore b/tools/testing/selftests/filesystems/mount_cycle/.gitignore
new file mode 100644
index 000000000000..28c623ef4d9d
--- /dev/null
+++ b/tools/testing/selftests/filesystems/mount_cycle/.gitignore
@@ -0,0 +1,2 @@
+# SPDX-License-Identifier: GPL-2.0-only
+loop_cycle_test
diff --git a/tools/testing/selftests/filesystems/mount_cycle/Makefile b/tools/testing/selftests/filesystems/mount_cycle/Makefile
new file mode 100644
index 000000000000..3becec29d69f
--- /dev/null
+++ b/tools/testing/selftests/filesystems/mount_cycle/Makefile
@@ -0,0 +1,6 @@
+# SPDX-License-Identifier: GPL-2.0
+TEST_GEN_PROGS := loop_cycle_test
+
+CFLAGS += -Wall -O2 -g $(KHDR_INCLUDES)
+
+include ../../lib.mk
diff --git a/tools/testing/selftests/filesystems/mount_cycle/loop_cycle_test.c b/tools/testing/selftests/filesystems/mount_cycle/loop_cycle_test.c
new file mode 100644
index 000000000000..67fc3e420c92
--- /dev/null
+++ b/tools/testing/selftests/filesystems/mount_cycle/loop_cycle_test.c
@@ -0,0 +1,377 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * A mount that another namespace's rmdir detached, or that went down with
+ * the detached tree it was in when the tree's last fd was closed, keeps
+ * its submounts connected, and a connected submount is put by its
+ * parent's final mntput(). A submount whose filesystem keeps a file open
+ * on the parent then holds the parent's count above zero for good:
+ * nothing in userspace refers to either mount any more and nothing can
+ * release them. A loop device is the simplest such filesystem, its
+ * backing file sits on the parent.
+ */
+#define _GNU_SOURCE
+#include <dirent.h>
+#include <errno.h>
+#include <fcntl.h>
+#include <sched.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <sys/ioctl.h>
+#include <sys/mount.h>
+#include <sys/stat.h>
+#include <stdbool.h>
+#include <sys/wait.h>
+#include <unistd.h>
+#include <linux/loop.h>
+
+#include "../wrappers.h"
+#include "../../kselftest_harness.h"
+
+#define IMAGE_SIZE	(1440 * 1024)
+#define SECTOR		512
+
+/* A blank FAT12 floppy image: boot sector, two FATs, an empty root directory. */
+static int write_fat12(int fd)
+{
+	unsigned char sector[SECTOR] = {
+		0xeb, 0x3c, 0x90, 'M', 'S', 'W', 'I', 'N', '4', '.', '1',
+		[11] = 0x00, 0x02,	/* bytes per sector: 512 */
+		[13] = 1,		/* sectors per cluster */
+		[14] = 1, 0,		/* reserved sectors */
+		[16] = 2,		/* FATs */
+		[17] = 0xe0, 0x00,	/* root directory entries: 224 */
+		[19] = 0x40, 0x0b,	/* total sectors: 2880 */
+		[21] = 0xf0,		/* media descriptor */
+		[22] = 9, 0,		/* sectors per FAT */
+		[24] = 18, 0,		/* sectors per track */
+		[26] = 2, 0,		/* heads */
+		[38] = 0x29,		/* extended boot signature */
+		[39] = 0x12, 0x34, 0x56, 0x78,
+		[43] = 'N', 'O', ' ', 'N', 'A', 'M', 'E', ' ', ' ', ' ', ' ',
+		[54] = 'F', 'A', 'T', '1', '2', ' ', ' ', ' ',
+		[510] = 0x55, 0xaa,
+	};
+	unsigned char fat[SECTOR] = { 0xf0, 0xff, 0xff };
+
+	if (pwrite(fd, sector, SECTOR, 0) != SECTOR)
+		return -1;
+	/* the first FAT and the second one, one sector each is enough */
+	if (pwrite(fd, fat, SECTOR, 1 * SECTOR) != SECTOR ||
+	    pwrite(fd, fat, SECTOR, 10 * SECTOR) != SECTOR)
+		return -1;
+	return ftruncate(fd, IMAGE_SIZE);
+}
+
+static int read_sysfs(const char *path, char *buf, size_t size)
+{
+	ssize_t n;
+	int fd;
+
+	fd = open(path, O_RDONLY);
+	if (fd < 0)
+		return -1;
+	n = read(fd, buf, size - 1);
+	close(fd);
+	if (n < 0)
+		return -1;
+	buf[n] = '\0';
+	return 0;
+}
+
+/* Does any mount namespace in the system show a mount of @dev? */
+static bool mounted_anywhere(const char *dev)
+{
+	char path[PATH_MAX], line[4096];
+	struct dirent *de;
+	bool found = false;
+	DIR *proc;
+	FILE *f;
+
+	proc = opendir("/proc");
+	if (!proc)
+		return false;
+	while (!found && (de = readdir(proc))) {
+		if (de->d_name[0] < '0' || de->d_name[0] > '9')
+			continue;
+		snprintf(path, sizeof(path), "/proc/%s/mountinfo", de->d_name);
+		f = fopen(path, "re");
+		if (!f)
+			continue;
+		while (fgets(line, sizeof(line), f)) {
+			if (strstr(line, dev)) {
+				found = true;
+				break;
+			}
+		}
+		fclose(f);
+	}
+	closedir(proc);
+	return found;
+}
+
+/* Wait up to @ms milliseconds for the loop device to give up its backing file. */
+static bool loop_released(const char *sysfs, int ms)
+{
+	char buf[PATH_MAX];
+
+	for (; ms > 0; ms -= 100) {
+		if (read_sysfs(sysfs, buf, sizeof(buf)) < 0)
+			return true;
+		usleep(100000);
+	}
+	return read_sysfs(sysfs, buf, sizeof(buf)) < 0;
+}
+
+FIXTURE(loop_cycle) {
+	char dev[32];		/* the loop device the child set up */
+	char sysfs[64];		/* its backing_file attribute */
+};
+
+FIXTURE_SETUP(loop_cycle)
+{
+	if (geteuid() != 0)
+		SKIP(return, "test requires CAP_SYS_ADMIN");
+	if (access("/dev/loop-control", R_OK | W_OK))
+		SKIP(return, "test requires loop devices");
+
+	ASSERT_EQ(unshare(CLONE_NEWNS), 0);
+	ASSERT_EQ(mount("", "/", NULL, MS_REC | MS_PRIVATE, NULL), 0);
+
+	rmdir("/mnt_dir");
+	ASSERT_EQ(mkdir("/mnt_dir", 0755), 0);
+	ASSERT_EQ(mount("tmpfs", "/mnt_dir", "tmpfs", 0, NULL), 0);
+	self->dev[0] = '\0';
+}
+
+FIXTURE_TEARDOWN(loop_cycle)
+{
+	umount2("/mnt_dir", MNT_DETACH);
+	rmdir("/mnt_dir");
+}
+
+/* Child exit codes. */
+enum {
+	CHILD_OK,
+	CHILD_NS,		/* could not set up the namespace or the tmpfs */
+	CHILD_IMAGE,		/* could not write the image */
+	CHILD_LOOP,		/* could not set up the loop device */
+	CHILD_MOUNT,		/* could not mount it (vfat and msdos both refused) */
+	CHILD_PIPE,		/* the parent went away */
+};
+
+/* Bind a free loop device to the open image @ifd; the device number. */
+static int loop_bind(int ifd)
+{
+	int cfd, lfd, n;
+	char dev[32];
+
+	cfd = open("/dev/loop-control", O_RDWR);
+	if (cfd < 0)
+		return -1;
+	n = ioctl(cfd, LOOP_CTL_GET_FREE);
+	close(cfd);
+	if (n < 0)
+		return -1;
+	snprintf(dev, sizeof(dev), "/dev/loop%d", n);
+	lfd = open(dev, O_RDWR);
+	if (lfd < 0)
+		return -1;
+	if (ioctl(lfd, LOOP_SET_FD, ifd))
+		n = -1;
+	close(lfd);
+	return n;
+}
+
+/* Write an image to @img, bind a loop device to it and mount that at @mp; the device number. */
+static int loop_mount(const char *img, const char *mp)
+{
+	char dev[32];
+	int ifd, n;
+
+	ifd = open(img, O_RDWR | O_CREAT | O_EXCL, 0600);
+	if (ifd < 0 || write_fat12(ifd))
+		return -CHILD_IMAGE;
+	n = loop_bind(ifd);
+	close(ifd);	/* the loop device holds the file from now on */
+	if (n < 0)
+		return -CHILD_LOOP;
+	snprintf(dev, sizeof(dev), "/dev/loop%d", n);
+	if (mkdir(mp, 0755))
+		return -CHILD_MOUNT;
+	if (mount(dev, mp, "vfat", 0, NULL) && mount(dev, mp, "msdos", 0, NULL))
+		return -CHILD_MOUNT;
+	return n;
+}
+
+/*
+ * In its own mount namespace the child mounts a tmpfs on /mnt_dir/vol,
+ * puts a filesystem image on it, binds a loop device to the image and
+ * mounts that loop device below. The loop device's backing file is a
+ * reference on the mount the image is on, held by the loop device, held
+ * by the mounted filesystem, held by the mount below.
+ */
+static int loop_child(int to_parent, int from_parent)
+{
+	char c;
+	int n;
+
+	if (unshare(CLONE_NEWNS) || mount("", "/", NULL, MS_REC | MS_PRIVATE, NULL))
+		return CHILD_NS;
+	if (mkdir("/mnt_dir/vol", 0755) || mount("tmpfs", "/mnt_dir/vol", "tmpfs", 0, NULL))
+		return CHILD_NS;
+	n = loop_mount("/mnt_dir/vol/img", "/mnt_dir/vol/mnt");
+	if (n < 0)
+		return -n;
+
+	if (write(to_parent, &n, sizeof(n)) != sizeof(n))
+		return CHILD_PIPE;
+	/* keep the namespace alive while the parent removes the directory */
+	if (read(from_parent, &c, 1) != 1)
+		return CHILD_PIPE;
+	return CHILD_OK;
+}
+
+/*
+ * An exclusive open of the device fails while a filesystem holds it, and
+ * the only filesystem that ever did is the one mounted below the dead
+ * mount. Give a release in flight a moment. Then the device must clear
+ * right away rather than only be marked for autoclear.
+ */
+static void assert_loop_released(struct __test_metadata *_metadata,
+				 FIXTURE_DATA(loop_cycle) *self)
+{
+	int lfd;
+
+	for (int i = 0; i < 20; i++) {
+		lfd = open(self->dev, O_RDONLY | O_EXCL);
+		if (lfd >= 0)
+			break;
+		usleep(100000);
+	}
+	EXPECT_GE(lfd, 0)
+		TH_LOG("%s is still held by the loop mount below the dead mount: nothing refers to either mount and nothing can release them",
+		       self->dev);
+	if (lfd >= 0)
+		close(lfd);
+
+	lfd = open(self->dev, O_RDWR);
+	ASSERT_GE(lfd, 0);
+	ASSERT_EQ(ioctl(lfd, LOOP_CLR_FD), 0);
+	close(lfd);
+	ASSERT_TRUE(loop_released(self->sysfs, 5000))
+		TH_LOG("%s kept its backing file after LOOP_CLR_FD: the filesystem on it is still mounted somewhere nobody can reach",
+		       self->dev);
+}
+
+/*
+ * rmdir of /mnt_dir/vol from here, where it is not a mountpoint, detaches
+ * the child's tmpfs with the loop mount connected below it. Once the child
+ * is gone nothing refers to either mount. The loop device must then be
+ * free to give up its backing file, which only happens when the mounted
+ * filesystem below the detached tmpfs has been released.
+ */
+TEST_F(loop_cycle, detached_loop_mount_released)
+{
+	int to_parent[2], to_child[2];
+	char buf[PATH_MAX];
+	int status, n = -1;
+	pid_t pid;
+
+	ASSERT_EQ(pipe(to_parent), 0);
+	ASSERT_EQ(pipe(to_child), 0);
+
+	pid = fork();
+	ASSERT_GE(pid, 0);
+	if (pid == 0) {
+		close(to_parent[0]);
+		close(to_child[1]);
+		_exit(loop_child(to_parent[1], to_child[0]));
+	}
+	close(to_parent[1]);
+	close(to_child[0]);
+
+	if (read(to_parent[0], &n, sizeof(n)) != sizeof(n)) {
+		waitpid(pid, &status, 0);
+		if (WIFEXITED(status) && WEXITSTATUS(status) == CHILD_MOUNT)
+			SKIP(return, "test requires a FAT filesystem");
+		ASSERT_EQ(status, 0);
+	}
+	snprintf(self->dev, sizeof(self->dev), "/dev/loop%d", n);
+	snprintf(self->sysfs, sizeof(self->sysfs), "/sys/block/loop%d/loop/backing_file", n);
+	ASSERT_EQ(read_sysfs(self->sysfs, buf, sizeof(buf)), 0);
+	ASSERT_NE(strstr(buf, "/mnt_dir/vol/img"), NULL);
+
+	/* not a mountpoint in this namespace, so the directory can go */
+	ASSERT_EQ(rmdir("/mnt_dir/vol"), 0);
+
+	/* the child leaves: its namespace and every reference it held are gone */
+	ASSERT_EQ(write(to_child[1], "", 1), 1);
+	ASSERT_EQ(waitpid(pid, &status, 0), pid);
+	ASSERT_EQ(status, 0);
+	close(to_parent[0]);
+	close(to_child[1]);
+
+	/* nothing can reach the two mounts any more */
+	ASSERT_EQ(access("/mnt_dir/vol", F_OK), -1);
+	ASSERT_FALSE(mounted_anywhere(self->dev));
+
+	assert_loop_released(_metadata, self);
+}
+
+/*
+ * The same two mounts in a detached tree: a clone of /mnt_dir/vol from
+ * open_tree(), the image opened through the clone so that the loop device
+ * holds the clone, and the loop mount moved below the clone. The last
+ * close of the tree's fd dissolves the tree with the loop mount left
+ * connected below the dead clone, and nothing refers to either afterwards.
+ */
+TEST_F(loop_cycle, dissolved_tree_loop_mount_released)
+{
+	int tfd, ifd, fsfd, mfd, n;
+	char buf[PATH_MAX];
+
+	fsfd = sys_fsopen("vfat", 0);
+	if (fsfd < 0)
+		fsfd = sys_fsopen("msdos", 0);
+	if (fsfd < 0)
+		SKIP(return, "test requires a FAT filesystem");
+
+	ASSERT_EQ(mkdir("/mnt_dir/vol", 0755), 0);
+	ASSERT_EQ(mount("tmpfs", "/mnt_dir/vol", "tmpfs", 0, NULL), 0);
+	tfd = sys_open_tree(AT_FDCWD, "/mnt_dir/vol", OPEN_TREE_CLONE | OPEN_TREE_CLOEXEC);
+	ASSERT_GE(tfd, 0);
+
+	/* the image, opened through the clone: the loop device holds the clone */
+	ifd = openat(tfd, "img", O_RDWR | O_CREAT | O_EXCL, 0600);
+	ASSERT_GE(ifd, 0);
+	ASSERT_EQ(write_fat12(ifd), 0);
+	n = loop_bind(ifd);
+	close(ifd);
+	ASSERT_GE(n, 0);
+	snprintf(self->dev, sizeof(self->dev), "/dev/loop%d", n);
+	snprintf(self->sysfs, sizeof(self->sysfs), "/sys/block/loop%d/loop/backing_file", n);
+
+	/* the loop mount, moved below the clone */
+	ASSERT_EQ(mkdirat(tfd, "mnt", 0755), 0);
+	ASSERT_EQ(sys_fsconfig(fsfd, FSCONFIG_SET_STRING, "source", self->dev, 0), 0);
+	ASSERT_EQ(sys_fsconfig(fsfd, FSCONFIG_CMD_CREATE, NULL, NULL, 0), 0);
+	mfd = sys_fsmount(fsfd, 0, 0);
+	ASSERT_GE(mfd, 0);
+	close(fsfd);
+	ASSERT_EQ(sys_move_mount(mfd, "", tfd, "mnt", MOVE_MOUNT_F_EMPTY_PATH), 0);
+	close(mfd);
+
+	ASSERT_EQ(read_sysfs(self->sysfs, buf, sizeof(buf)), 0);
+	ASSERT_NE(strstr(buf, "img"), NULL);
+
+	/* the last fd of the tree: both mounts die, the loop mount connected */
+	close(tfd);
+
+	/* nothing can reach the two mounts any more */
+	ASSERT_FALSE(mounted_anywhere(self->dev));
+
+	assert_loop_released(_metadata, self);
+}
+
+TEST_HARNESS_MAIN

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 9+ messages in thread

* [PATCH RFC v2 7/8] selftests/filesystems: check the two-step cycle over crossed loop images
  2026-09-24 22:35 [PATCH RFC v2 0/8] namespace: prevent UMOUNT_CONNECTED reference count cycles Christian Brauner
                   ` (5 preceding siblings ...)
  2026-09-24 22:35 ` [PATCH RFC v2 6/8] selftests/filesystems: check that a loop mount below a dead mount is released Christian Brauner
@ 2026-09-24 22:35 ` Christian Brauner
  2026-09-24 22:35 ` [PATCH RFC v2 8/8] selftests/filesystems: check that the holders let go of a dead mount Christian Brauner
  7 siblings, 0 replies; 9+ messages in thread
From: Christian Brauner @ 2026-09-24 22:35 UTC (permalink / raw)
  To: Linus Torvalds
  Cc: Jann Horn, Jan Kara, Amir Goldstein, linux-fsdevel,
	Alexander Viro, Christian Brauner (Amutable)

Two tmpfs mounts in a child namespace, each carrying the image of the
loop mount below the other. Check that both loop devices are released
after:

- rmdir of the first tmpfs' mountpoint from the parent namespace, which
  leaves the loop mount below it connected while its image's mount lives

- rmdir of the second, which takes it down with the loop mount whose
  image is on the first, by then dead, tmpfs

Each dead tmpfs then owns a loop mount whose filesystem pins the other.

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 .../filesystems/mount_cycle/loop_cycle_test.c      | 88 ++++++++++++++++++++++
 1 file changed, 88 insertions(+)

diff --git a/tools/testing/selftests/filesystems/mount_cycle/loop_cycle_test.c b/tools/testing/selftests/filesystems/mount_cycle/loop_cycle_test.c
index 67fc3e420c92..4f4c88397861 100644
--- a/tools/testing/selftests/filesystems/mount_cycle/loop_cycle_test.c
+++ b/tools/testing/selftests/filesystems/mount_cycle/loop_cycle_test.c
@@ -204,6 +204,34 @@ static int loop_mount(const char *img, const char *mp)
 	return n;
 }
 
+/*
+ * Two tmpfs mounts, each carrying the image of the loop mount below the
+ * other: the loop mount below /mnt_dir/vol has its image on /mnt_dir/vol2
+ * and the other way round.
+ */
+static int crossed_child(int to_parent, int from_parent)
+{
+	int n[2];
+	char c;
+
+	if (unshare(CLONE_NEWNS) || mount("", "/", NULL, MS_REC | MS_PRIVATE, NULL))
+		return CHILD_NS;
+	if (mkdir("/mnt_dir/vol", 0755) || mount("tmpfs", "/mnt_dir/vol", "tmpfs", 0, NULL) ||
+	    mkdir("/mnt_dir/vol2", 0755) || mount("tmpfs", "/mnt_dir/vol2", "tmpfs", 0, NULL))
+		return CHILD_NS;
+	n[0] = loop_mount("/mnt_dir/vol2/img", "/mnt_dir/vol/mnt");
+	if (n[0] < 0)
+		return -n[0];
+	n[1] = loop_mount("/mnt_dir/vol/img", "/mnt_dir/vol2/mnt");
+	if (n[1] < 0)
+		return -n[1];
+	if (write(to_parent, n, sizeof(n)) != sizeof(n))
+		return CHILD_PIPE;
+	if (read(from_parent, &c, 1) != 1)
+		return CHILD_PIPE;
+	return CHILD_OK;
+}
+
 /*
  * In its own mount namespace the child mounts a tmpfs on /mnt_dir/vol,
  * puts a filesystem image on it, binds a loop device to the image and
@@ -374,4 +402,64 @@ TEST_F(loop_cycle, dissolved_tree_loop_mount_released)
 	assert_loop_released(_metadata, self);
 }
 
+/*
+ * The cycle in two steps: rmdir of /mnt_dir/vol leaves the loop mount
+ * below it connected while its image's mount, /mnt_dir/vol2, is alive;
+ * then rmdir of /mnt_dir/vol2 takes that one with the loop mount whose
+ * image is on the dead /mnt_dir/vol. Each dead mount now owns a loop
+ * mount whose filesystem pins the other.
+ */
+TEST_F(loop_cycle, crossed_images_released)
+{
+	int to_parent[2], to_child[2];
+	char sysfs[2][64], dev[2][32];
+	int status, n[2] = { -1, -1 };
+	pid_t pid;
+
+	ASSERT_EQ(pipe(to_parent), 0);
+	ASSERT_EQ(pipe(to_child), 0);
+
+	pid = fork();
+	ASSERT_GE(pid, 0);
+	if (pid == 0) {
+		close(to_parent[0]);
+		close(to_child[1]);
+		_exit(crossed_child(to_parent[1], to_child[0]));
+	}
+	close(to_parent[1]);
+	close(to_child[0]);
+
+	if (read(to_parent[0], n, sizeof(n)) != sizeof(n)) {
+		waitpid(pid, &status, 0);
+		if (WIFEXITED(status) && WEXITSTATUS(status) == CHILD_MOUNT)
+			SKIP(return, "test requires a FAT filesystem");
+		ASSERT_EQ(status, 0);
+	}
+	for (int i = 0; i < 2; i++) {
+		snprintf(dev[i], sizeof(dev[i]), "/dev/loop%d", n[i]);
+		snprintf(sysfs[i], sizeof(sysfs[i]), "/sys/block/loop%d/loop/backing_file", n[i]);
+	}
+
+	/* step one: the mount with the first loop mount below it goes */
+	ASSERT_EQ(rmdir("/mnt_dir/vol"), 0);
+
+	/* step two: the other one, with the loop mount whose image is on the first */
+	ASSERT_EQ(rmdir("/mnt_dir/vol2"), 0);
+
+	/* the child leaves: its namespace and every reference it held are gone */
+	ASSERT_EQ(write(to_child[1], "", 1), 1);
+	ASSERT_EQ(waitpid(pid, &status, 0), pid);
+	ASSERT_EQ(status, 0);
+	close(to_parent[0]);
+	close(to_child[1]);
+
+	/* nothing can reach the four mounts any more */
+	for (int i = 0; i < 2; i++) {
+		ASSERT_FALSE(mounted_anywhere(dev[i]));
+		strcpy(self->dev, dev[i]);
+		strcpy(self->sysfs, sysfs[i]);
+		assert_loop_released(_metadata, self);
+	}
+}
+
 TEST_HARNESS_MAIN

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 9+ messages in thread

* [PATCH RFC v2 8/8] selftests/filesystems: check that the holders let go of a dead mount
  2026-09-24 22:35 [PATCH RFC v2 0/8] namespace: prevent UMOUNT_CONNECTED reference count cycles Christian Brauner
                   ` (6 preceding siblings ...)
  2026-09-24 22:35 ` [PATCH RFC v2 7/8] selftests/filesystems: check the two-step cycle over crossed loop images Christian Brauner
@ 2026-09-24 22:35 ` Christian Brauner
  7 siblings, 0 replies; 9+ messages in thread
From: Christian Brauner @ 2026-09-24 22:35 UTC (permalink / raw)
  To: Linus Torvalds
  Cc: Jann Horn, Jan Kara, Amir Goldstein, linux-fsdevel,
	Alexander Viro, Christian Brauner (Amutable)

Extend loop_cycle_test with one case per holder that keeps a file or a
path on a mount and can be mounted below it:

- a loop mount P from an image on a tmpfs next to it, so that P's death
  shows as the loop device giving up its backing file
- autofs with a FIFO on P as its pipe, zram with a device node on P as
  its writeback device, both on a minix image since vfat has neither
- ecryptfs with its lower directory on P, under a passphrase token
  added to the session keyring
- binfmt_misc in a new user namespace with an 'F' interpreter on P
- a fuse server that answers FUSE_INIT with passthrough on and registers
  a file on P as a backing file
- zloop with its zone files in a directory on P
- a mass storage gadget on the dummy UDC with its LUN file on P, mounted
  from the SCSI disk the gadget shows up as
- md with a RAID1 of one loop device and its bitmap file on P, which
  skips while SET_BITMAP_FILE has no way to succeed
- rmdir of P's mountpoint from the parent, then the child exits, then
  the device must be free and LOOP_CLR_FD must release the file

Each case leaks the device on a kernel without the holder's conversion
and releases it with it.

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 .../filesystems/mount_cycle/loop_cycle_test.c      | 760 +++++++++++++++++++++
 1 file changed, 760 insertions(+)

diff --git a/tools/testing/selftests/filesystems/mount_cycle/loop_cycle_test.c b/tools/testing/selftests/filesystems/mount_cycle/loop_cycle_test.c
index 4f4c88397861..09cc824c769e 100644
--- a/tools/testing/selftests/filesystems/mount_cycle/loop_cycle_test.c
+++ b/tools/testing/selftests/filesystems/mount_cycle/loop_cycle_test.c
@@ -23,7 +23,16 @@
 #include <stdbool.h>
 #include <sys/wait.h>
 #include <unistd.h>
+#include <linux/fuse.h>
+#include <linux/keyctl.h>
+#include <linux/major.h>
+#include <linux/raid/md_u.h>
+#include <linux/raid/md_p.h>
 #include <linux/loop.h>
+#include <linux/magic.h>
+#include <sys/syscall.h>
+#include <sys/sysmacros.h>
+#include <sys/uio.h>
 
 #include "../wrappers.h"
 #include "../../kselftest_harness.h"
@@ -34,6 +43,7 @@
 /* A blank FAT12 floppy image: boot sector, two FATs, an empty root directory. */
 static int write_fat12(int fd)
 {
+	struct stat st;
 	unsigned char sector[SECTOR] = {
 		0xeb, 0x3c, 0x90, 'M', 'S', 'W', 'I', 'N', '4', '.', '1',
 		[11] = 0x00, 0x02,	/* bytes per sector: 512 */
@@ -60,9 +70,114 @@ static int write_fat12(int fd)
 	if (pwrite(fd, fat, SECTOR, 1 * SECTOR) != SECTOR ||
 	    pwrite(fd, fat, SECTOR, 10 * SECTOR) != SECTOR)
 		return -1;
+	if (fstat(fd, &st) || S_ISBLK(st.st_mode))
+		return 0;
 	return ftruncate(fd, IMAGE_SIZE);
 }
 
+/*
+ * A blank FAT12 image with 4 KiB sectors and @sectors of them, for a device
+ * with 4 KiB logical blocks (zram) or for a bigger image than a floppy.
+ */
+static int write_fat12_4k(int fd, unsigned int sectors)
+{
+	unsigned int fat_sectors = (sectors * 3 / 2 + 4095) / 4096;
+	unsigned char sector[4096] = {
+		0xeb, 0x3c, 0x90, 'M', 'S', 'W', 'I', 'N', '4', '.', '1',
+		[11] = 0x00, 0x10,	/* bytes per sector: 4096 */
+		[13] = 1,		/* sectors per cluster */
+		[14] = 1, 0,		/* reserved sectors */
+		[16] = 2,		/* FATs */
+		[17] = 128, 0,		/* root directory entries: one sector */
+		[19] = sectors & 0xff, sectors >> 8,
+		[21] = 0xf8,		/* media descriptor */
+		[22] = fat_sectors, 0,
+		[24] = 63, 0,		/* sectors per track */
+		[26] = 255, 0,		/* heads */
+		[38] = 0x29,		/* extended boot signature */
+		[39] = 0x12, 0x34, 0x56, 0x78,
+		[43] = 'N', 'O', ' ', 'N', 'A', 'M', 'E', ' ', ' ', ' ', ' ',
+		[54] = 'F', 'A', 'T', '1', '2', ' ', ' ', ' ',
+		[510] = 0x55, 0xaa,
+	};
+	unsigned char fat[4096] = { 0xf8, 0xff, 0xff };
+	struct stat st;
+
+	if (pwrite(fd, sector, sizeof(sector), 0) != sizeof(sector))
+		return -1;
+	if (pwrite(fd, fat, sizeof(fat), 1 * 4096) != sizeof(fat) ||
+	    pwrite(fd, fat, sizeof(fat), (1 + fat_sectors) * 4096) != sizeof(fat))
+		return -1;
+	if (fstat(fd, &st) || S_ISBLK(st.st_mode))
+		return 0;
+	return ftruncate(fd, (off_t)sectors * 4096);
+}
+
+#define MINIX_BLOCK	1024
+#define MINIX_BLOCKS	4096			/* a 4 MiB image */
+#define MINIX_INODES	512
+#define MINIX_ITABLE	(MINIX_INODES * 32 / MINIX_BLOCK)
+#define MINIX_FIRSTDATA	(2 + 1 + 1 + MINIX_ITABLE)	/* boot, super, imap, zmap, inodes */
+
+/*
+ * A blank minix v1 image, for the holders that need a FIFO or a device
+ * node on the dying mount, which vfat can't hold. Superblock in block 1,
+ * one block each for the inode and zone bitmaps, the inode table, and
+ * the root directory in the first data zone.
+ */
+static int write_minix(int fd)
+{
+	struct {
+		__u16 s_ninodes, s_nzones, s_imap_blocks, s_zmap_blocks;
+		__u16 s_firstdatazone, s_log_zone_size;
+		__u32 s_max_size;
+		__u16 s_magic, s_state;
+	} sb = {
+		.s_ninodes = MINIX_INODES,
+		.s_nzones = MINIX_BLOCKS,
+		.s_imap_blocks = 1,
+		.s_zmap_blocks = 1,
+		.s_firstdatazone = MINIX_FIRSTDATA,
+		.s_max_size = (7 + 512 + 512 * 512) * MINIX_BLOCK,
+		.s_magic = MINIX_SUPER_MAGIC,
+		.s_state = 1,		/* MINIX_VALID_FS */
+	};
+	struct {
+		__u16 i_mode, i_uid;
+		__u32 i_size, i_time;
+		__u8 i_gid, i_nlinks;
+		__u16 i_zone[9];
+	} root = {
+		.i_mode = S_IFDIR | 0755,
+		.i_size = 2 * 16,
+		.i_nlinks = 2,
+		.i_zone = { MINIX_FIRSTDATA },
+	};
+	unsigned char imap[MINIX_BLOCK], zmap[MINIX_BLOCK], dir[MINIX_BLOCK] = {};
+	int i;
+
+	/* bit 0 is reserved in both maps, the root inode and its zone are in use */
+	memset(imap, 0xff, sizeof(imap));
+	for (i = 2; i <= MINIX_INODES; i++)
+		imap[i / 8] &= ~(1 << (i % 8));
+	memset(zmap, 0xff, sizeof(zmap));
+	for (i = 2; i <= MINIX_BLOCKS - MINIX_FIRSTDATA; i++)
+		zmap[i / 8] &= ~(1 << (i % 8));
+	dir[0] = 1;
+	dir[2] = '.';
+	dir[16] = 1;
+	dir[18] = '.';
+	dir[19] = '.';
+
+	if (pwrite(fd, &sb, sizeof(sb), 1 * MINIX_BLOCK) != sizeof(sb) ||
+	    pwrite(fd, imap, sizeof(imap), 2 * MINIX_BLOCK) != sizeof(imap) ||
+	    pwrite(fd, zmap, sizeof(zmap), 3 * MINIX_BLOCK) != sizeof(zmap) ||
+	    pwrite(fd, &root, sizeof(root), 4 * MINIX_BLOCK) != sizeof(root) ||
+	    pwrite(fd, dir, sizeof(dir), MINIX_FIRSTDATA * MINIX_BLOCK) != sizeof(dir))
+		return -1;
+	return ftruncate(fd, (off_t)MINIX_BLOCKS * MINIX_BLOCK);
+}
+
 static int read_sysfs(const char *path, char *buf, size_t size)
 {
 	ssize_t n;
@@ -158,6 +273,8 @@ enum {
 	CHILD_LOOP,		/* could not set up the loop device */
 	CHILD_MOUNT,		/* could not mount it (vfat and msdos both refused) */
 	CHILD_PIPE,		/* the parent went away */
+	CHILD_HOLDER,		/* could not set the holder up below the mount */
+	CHILD_SKIP,		/* the kernel lacks what the holder needs */
 };
 
 /* Bind a free loop device to the open image @ifd; the device number. */
@@ -462,4 +579,647 @@ TEST_F(loop_cycle, crossed_images_released)
 	}
 }
 
+/*
+ * The holders. Each keeps a file or a path on a mount P for as long as
+ * its own filesystem or device lives, and each has that filesystem or
+ * device mounted at C below P. Once another namespace's rmdir has
+ * detached P with C connected below it and the child is gone, P is owned
+ * by nobody, C by P, and P is kept by whatever the holder still holds.
+ *
+ * P is a loop mount so that its death can be observed: the loop device
+ * gives its backing file up when P's superblock goes. The image sits on
+ * a tmpfs next to P, not above it, so the loop device's own reference is
+ * not part of the picture.
+ */
+enum holder {
+	HOLDER_AUTOFS,		/* a FIFO on P as the daemon's pipe */
+	HOLDER_ZRAM,		/* a device node on P as the writeback device */
+	HOLDER_ECRYPTFS,	/* a directory on P as the lower directory */
+	HOLDER_BINFMT_MISC,	/* an executable on P as an 'F' interpreter */
+	HOLDER_FUSE,		/* a file on P as a passthrough backing file */
+	HOLDER_ZLOOP,		/* a directory on P for the zone files */
+	HOLDER_GADGET,		/* a file on P as a mass storage LUN, over dummy_hcd */
+	HOLDER_MD,		/* a file on P as an array's bitmap file */
+};
+
+#define HOLDER_IMG	"/mnt_dir/img/p.img"
+#define HOLDER_MNT	"/mnt_dir/p"
+#define HOLDER_BELOW	"/mnt_dir/p/c"
+
+static int write_file(const char *path, const char *s)
+{
+	int fd = open(path, O_WRONLY);
+	ssize_t n;
+
+	if (fd < 0)
+		return -1;
+	n = write(fd, s, strlen(s));
+	close(fd);
+	return n == (ssize_t)strlen(s) ? 0 : -1;
+}
+
+/* Bind a free loop device to @img; the device number. */
+static int loop_attach(const char *img)
+{
+	int ifd, n;
+
+	ifd = open(img, O_RDWR);
+	if (ifd < 0)
+		return -1;
+	n = loop_bind(ifd);
+	close(ifd);
+	return n;
+}
+
+static int holder_autofs(void)
+{
+	char opts[64];
+	int pfd;
+
+	if (mkfifo(HOLDER_MNT "/pipe", 0600))
+		return CHILD_HOLDER;
+	pfd = open(HOLDER_MNT "/pipe", O_RDWR);
+	if (pfd < 0)
+		return CHILD_HOLDER;
+	snprintf(opts, sizeof(opts), "fd=%d,minproto=5,maxproto=5", pfd);
+	if (mount("autofs", HOLDER_BELOW, "autofs", 0, opts))
+		return errno == ENODEV ? CHILD_SKIP : CHILD_HOLDER;
+	close(pfd);	/* the mount keeps its own */
+	return CHILD_OK;
+}
+
+/*
+ * zram's writeback device has to be a block device node, and that node has
+ * to be on P. Point it at a second loop device.
+ */
+static int holder_zram(void)
+{
+	char buf[64];
+	int fd, n;
+
+	if (access("/sys/block/zram0/backing_dev", W_OK))
+		return CHILD_SKIP;
+	/* somebody else's device, leave it alone */
+	if (read_sysfs("/sys/block/zram0/initstate", buf, sizeof(buf)) || buf[0] != '0')
+		return CHILD_SKIP;
+	fd = open("/mnt_dir/img/wb.img", O_RDWR | O_CREAT | O_EXCL, 0600);
+	if (fd < 0 || ftruncate(fd, IMAGE_SIZE))
+		return CHILD_HOLDER;
+	close(fd);
+	n = loop_attach("/mnt_dir/img/wb.img");
+	if (n < 0)
+		return CHILD_HOLDER;
+	if (mknod(HOLDER_MNT "/wbdev", S_IFBLK | 0600, makedev(7, n)))
+		return CHILD_HOLDER;
+	if (write_file("/sys/block/zram0/backing_dev", HOLDER_MNT "/wbdev"))
+		return CHILD_HOLDER;
+	snprintf(buf, sizeof(buf), "%d", 4 * 1024 * 1024);
+	if (write_file("/sys/block/zram0/disksize", buf))
+		return CHILD_HOLDER;
+	fd = open("/dev/zram0", O_RDWR);
+	if (fd < 0 || write_fat12_4k(fd, 1024))	/* zram has 4 KiB blocks */
+		return CHILD_HOLDER;
+	close(fd);
+	if (mount("/dev/zram0", HOLDER_BELOW, "vfat", 0, NULL) &&
+	    mount("/dev/zram0", HOLDER_BELOW, "msdos", 0, NULL))
+		return CHILD_HOLDER;
+	return CHILD_OK;
+}
+
+/* The kernel's auth token layout, which userspace has to match byte for byte. */
+struct ecryptfs_auth_tok {
+	__u16 version;
+	__u16 token_type;
+	__u32 flags;
+	struct {
+		__u32 flags, encrypted_key_size, decrypted_key_size;
+		__u8 encrypted_key[512], decrypted_key[64];
+	} session_key;
+	__u8 reserved[32];
+	struct {
+		__u32 password_bytes;
+		__s32 hash_algo;
+		__u32 hash_iterations, session_key_encryption_key_bytes, flags;
+		__u8 session_key_encryption_key[64];
+		__u8 signature[17];
+		__u8 salt[8];
+	} password;
+} __attribute__((packed));
+
+#define ECRYPTFS_SIG	"0123456789abcdef"
+
+static int holder_ecryptfs(void)
+{
+	struct ecryptfs_auth_tok tok = {
+		.version = 0x0004,
+		.token_type = 0,		/* ECRYPTFS_PASSWORD */
+		.password.session_key_encryption_key_bytes = 16,
+		.password.flags = 0x02,		/* ECRYPTFS_SESSION_KEY_ENCRYPTION_KEY_SET */
+		.password.signature = ECRYPTFS_SIG,
+	};
+
+	if (syscall(__NR_add_key, "user", ECRYPTFS_SIG, &tok, sizeof(tok),
+		    KEY_SPEC_SESSION_KEYRING) < 0)
+		return CHILD_HOLDER;
+	if (mkdir(HOLDER_MNT "/lower", 0755))
+		return CHILD_HOLDER;
+	if (mount(HOLDER_MNT "/lower", HOLDER_BELOW, "ecryptfs", 0,
+		  "ecryptfs_sig=" ECRYPTFS_SIG ",ecryptfs_cipher=aes,ecryptfs_key_bytes=16"))
+		return errno == ENODEV ? CHILD_SKIP : CHILD_HOLDER;
+	return CHILD_OK;
+}
+
+/*
+ * binfmt_misc instances are per user namespace, so the mount below P is
+ * made from a new one, which gets a copy of P.
+ */
+static int holder_binfmt_misc(void)
+{
+	char buf[4096];
+	int in, out;
+	ssize_t n;
+
+	in = open("/proc/self/exe", O_RDONLY);
+	out = open(HOLDER_MNT "/interp", O_WRONLY | O_CREAT | O_EXCL, 0755);
+	if (in < 0 || out < 0)
+		return CHILD_HOLDER;
+	while ((n = read(in, buf, sizeof(buf))) > 0)
+		if (write(out, buf, n) != n)
+			return CHILD_HOLDER;
+	close(in);
+	close(out);
+
+	if (unshare(CLONE_NEWUSER | CLONE_NEWNS))
+		return errno == EINVAL ? CHILD_SKIP : CHILD_HOLDER;
+	if (write_file("/proc/self/setgroups", "deny") ||
+	    write_file("/proc/self/uid_map", "0 0 1") ||
+	    write_file("/proc/self/gid_map", "0 0 1"))
+		return CHILD_HOLDER;
+	if (mount("binfmt_misc", HOLDER_BELOW, "binfmt_misc", 0, NULL))
+		return errno == ENODEV ? CHILD_SKIP : CHILD_HOLDER;
+	if (write_file(HOLDER_BELOW "/register", ":cycle:E::cyc::" HOLDER_MNT "/interp:F"))
+		return CHILD_HOLDER;
+	return CHILD_OK;
+}
+
+/*
+ * A fuse server that only ever answers FUSE_INIT, with passthrough on, and
+ * then registers a file on P as a backing file. The registration alone
+ * makes the fuse superblock hold the file.
+ */
+static int holder_fuse(void)
+{
+	struct fuse_backing_map map = {};
+	struct fuse_in_header *ih;
+	struct fuse_init_out init = {
+		.major = FUSE_KERNEL_VERSION,
+		.minor = FUSE_KERNEL_MINOR_VERSION,
+		.flags = FUSE_INIT_EXT,
+		.flags2 = FUSE_PASSTHROUGH >> 32,
+		.max_write = 4096,
+		.max_stack_depth = 1,
+	};
+	struct fuse_out_header oh = { .len = sizeof(oh) + sizeof(init) };
+	struct iovec iov[2] = { { &oh, sizeof(oh) }, { &init, sizeof(init) } };
+	char opts[64], buf[8192];
+	int ffd, bfd;
+	ssize_t n;
+
+	bfd = open(HOLDER_MNT "/backing", O_RDWR | O_CREAT | O_EXCL, 0600);
+	if (bfd < 0)
+		return CHILD_HOLDER;
+	ffd = open("/dev/fuse", O_RDWR);
+	if (ffd < 0)
+		return CHILD_SKIP;
+	snprintf(opts, sizeof(opts), "fd=%d,rootmode=40000,user_id=0,group_id=0", ffd);
+	if (mount("fuse", HOLDER_BELOW, "fuse", 0, opts))
+		return errno == ENODEV ? CHILD_SKIP : CHILD_HOLDER;
+
+	n = read(ffd, buf, sizeof(buf));
+	ih = (void *)buf;
+	if (n < (ssize_t)sizeof(*ih) || ih->opcode != FUSE_INIT)
+		return CHILD_HOLDER;
+	oh.unique = ih->unique;
+	if (writev(ffd, iov, 2) != (ssize_t)oh.len)
+		return CHILD_HOLDER;
+
+	map.fd = bfd;
+	if (ioctl(ffd, FUSE_DEV_IOC_BACKING_OPEN, &map) < 0)
+		return errno == EPERM ? CHILD_SKIP : CHILD_HOLDER;
+	close(bfd);	/* the connection keeps its own */
+	return CHILD_OK;	/* ffd stays open until the child exits */
+}
+
+/* Wait for a device node the kernel is about to create. */
+static int open_when_there(const char *dev, int flags, int ms)
+{
+	int fd;
+
+	for (; ms > 0; ms -= 100) {
+		fd = open(dev, flags);
+		if (fd >= 0)
+			return fd;
+		usleep(100000);
+	}
+	return -1;
+}
+
+/*
+ * zloop keeps every zone file open. One conventional zone is enough for a
+ * FAT image, the sequential one stays empty.
+ */
+static int holder_zloop(void)
+{
+	int fd;
+
+	if (access("/dev/zloop-control", W_OK))
+		return CHILD_SKIP;
+	if (mkdir(HOLDER_MNT "/zl", 0755) || mkdir(HOLDER_MNT "/zl/0", 0755))
+		return CHILD_HOLDER;
+	if (write_file("/dev/zloop-control",
+		       "add id=0,capacity_mb=8,zone_size_mb=4,conv_zones=1,base_dir=" HOLDER_MNT "/zl"))
+		return CHILD_HOLDER;
+	fd = open_when_there("/dev/zloop0", O_RDWR, 5000);
+	if (fd < 0 || write_fat12_4k(fd, 1024))	/* 4 KiB blocks, like P */
+		return CHILD_HOLDER;
+	close(fd);
+	if (mount("/dev/zloop0", HOLDER_BELOW, "vfat", 0, NULL) &&
+	    mount("/dev/zloop0", HOLDER_BELOW, "msdos", 0, NULL))
+		return CHILD_HOLDER;
+	return CHILD_OK;
+}
+
+#define GADGET	"/sys/kernel/config/usb_gadget/g1"
+
+/* The disk usb-storage created for the gadget, by the LUN's inquiry string. */
+static int find_gadget_disk(char *dev, size_t len, int ms)
+{
+	char path[PATH_MAX], model[64];
+	struct dirent *de;
+	DIR *d;
+
+	for (; ms > 0; ms -= 100, usleep(100000)) {
+		d = opendir("/sys/block");
+		if (!d)
+			return -1;
+		while ((de = readdir(d))) {
+			if (strncmp(de->d_name, "sd", 2))
+				continue;
+			snprintf(path, sizeof(path), "/sys/block/%s/device/model", de->d_name);
+			if (read_sysfs(path, model, sizeof(model)) ||
+			    strncmp(model, "File-Stor Gadget", 16))
+				continue;
+			snprintf(dev, len, "/dev/%s", de->d_name);
+			closedir(d);
+			return 0;
+		}
+		closedir(d);
+	}
+	return -1;
+}
+
+/*
+ * A mass storage gadget bound to the dummy UDC, so that this kernel is
+ * also the USB host that sees the LUN as a SCSI disk. sd locks the
+ * medium on open, which is what keeps the LUN's file from being ejected.
+ */
+#define GADGET_STEP(x) do {						\
+	if (x) {							\
+		fprintf(stderr, "gadget: %s failed: %s\n", #x, strerror(errno)); \
+		return CHILD_HOLDER;					\
+	}								\
+} while (0)
+
+static int holder_gadget(void)
+{
+	char dev[PATH_MAX];
+	int fd;
+
+	if (access("/sys/kernel/config", F_OK) ||
+	    (mount("configfs", "/sys/kernel/config", "configfs", 0, NULL) && errno != EBUSY))
+		return CHILD_SKIP;
+	if (access("/sys/kernel/config/usb_gadget", F_OK) ||
+	    access("/sys/class/udc/dummy_udc.0", F_OK))
+		return CHILD_SKIP;
+
+	fd = open(HOLDER_MNT "/lun.img", O_RDWR | O_CREAT | O_EXCL, 0600);
+	if (fd < 0 || write_fat12(fd))
+		return CHILD_HOLDER;
+	close(fd);
+
+	GADGET_STEP(mkdir(GADGET, 0755));
+	GADGET_STEP(write_file(GADGET "/idVendor", "0x1d6b"));
+	GADGET_STEP(write_file(GADGET "/idProduct", "0x0104"));
+	GADGET_STEP(mkdir(GADGET "/strings/0x409", 0755));
+	GADGET_STEP(write_file(GADGET "/strings/0x409/serialnumber", "1"));
+	GADGET_STEP(write_file(GADGET "/strings/0x409/manufacturer", "kselftest"));
+	GADGET_STEP(write_file(GADGET "/strings/0x409/product", "cycle"));
+	GADGET_STEP(mkdir(GADGET "/configs/c.1", 0755));
+	GADGET_STEP(mkdir(GADGET "/configs/c.1/strings/0x409", 0755));
+	GADGET_STEP(write_file(GADGET "/configs/c.1/strings/0x409/configuration", "c"));
+	GADGET_STEP(mkdir(GADGET "/functions/mass_storage.0", 0755));
+	GADGET_STEP(write_file(GADGET "/functions/mass_storage.0/lun.0/removable", "1"));
+	GADGET_STEP(write_file(GADGET "/functions/mass_storage.0/lun.0/file", HOLDER_MNT "/lun.img"));
+	GADGET_STEP(symlink(GADGET "/functions/mass_storage.0", GADGET "/configs/c.1/mass_storage.0"));
+	GADGET_STEP(write_file(GADGET "/UDC", "dummy_udc.0"));
+
+	/* usb-storage waits a second before it scans the device */
+	GADGET_STEP(find_gadget_disk(dev, sizeof(dev), 15000));
+	fd = open_when_there(dev, O_RDONLY, 5000);
+	GADGET_STEP(fd < 0);
+	close(fd);
+	if (mount(dev, HOLDER_BELOW, "vfat", 0, NULL) &&
+	    mount(dev, HOLDER_BELOW, "msdos", 0, NULL))
+		return CHILD_HOLDER;
+	return CHILD_OK;
+}
+
+#define BITMAP_MAGIC	0x6d746962
+
+/*
+ * A RAID1 of one loop device, not persistent, with its bitmap in a file
+ * on P. The bitmap file needs a superblock the kernel accepts; sync_size,
+ * uuid and events are not looked at for a non-persistent array.
+ */
+static int holder_md(void)
+{
+	struct {
+		__u32 magic, version;
+		__u8 uuid[16];
+		__u64 events, events_cleared, sync_size;
+		__u32 state, chunksize, daemon_sleep, write_behind;
+	} bsb = {
+		.magic = BITMAP_MAGIC,
+		.version = 4,
+		.chunksize = 64 * 1024,
+		.daemon_sleep = 5,
+	};
+	mdu_array_info_t info = {
+		.level = 1,
+		.raid_disks = 1,
+		.size = 8 * 1024,		/* KiB */
+		.not_persistent = 1,
+	};
+	mdu_disk_info_t disk = {
+		.major = 7,
+		.state = (1 << MD_DISK_ACTIVE) | (1 << MD_DISK_SYNC),
+	};
+	int fd, mdfd, bfd, n;
+	char buf[4096];
+
+	fd = open("/mnt_dir/img/md.img", O_RDWR | O_CREAT | O_EXCL, 0600);
+	if (fd < 0 || ftruncate(fd, 8 * 1024 * 1024))
+		return CHILD_HOLDER;
+	close(fd);
+	n = loop_attach("/mnt_dir/img/md.img");
+	if (n < 0)
+		return CHILD_HOLDER;
+	disk.minor = n;
+
+	bfd = open(HOLDER_MNT "/bitmap", O_RDWR | O_CREAT | O_EXCL, 0600);
+	if (bfd < 0 || write(bfd, &bsb, sizeof(bsb)) != sizeof(bsb) || ftruncate(bfd, 4096))
+		return CHILD_HOLDER;
+
+	if (read_sysfs("/proc/mdstat", buf, sizeof(buf)) || !strstr(buf, "[raid1]")) {
+		fprintf(stderr, "md: no raid1 personality: %s\n", buf);
+		return CHILD_SKIP;
+	}
+	if (access("/dev/md0", F_OK) && mknod("/dev/md0", S_IFBLK | 0600, makedev(9, 0)))
+		return CHILD_HOLDER;
+	mdfd = open("/dev/md0", O_RDWR);
+	if (mdfd < 0)
+		return CHILD_SKIP;
+	/* the bitmap ops are only installed once a bitmap type is chosen */
+	write_file("/sys/block/md0/md/bitmap_type", "bitmap");
+	/* EBUSY: somebody else's array, leave it alone */
+	if (ioctl(mdfd, SET_ARRAY_INFO, &info))
+		return errno == EBUSY ? CHILD_SKIP : CHILD_HOLDER;
+	if (ioctl(mdfd, ADD_NEW_DISK, &disk))
+		return CHILD_HOLDER;
+	/* attach the bitmap file before the array runs, the way mdadm does */
+	if (ioctl(mdfd, SET_BITMAP_FILE, bfd)) {
+		fprintf(stderr, "md: SET_BITMAP_FILE: %s\n", strerror(errno));
+		return errno == EINVAL ? CHILD_SKIP : CHILD_HOLDER;
+	}
+	close(bfd);	/* the array keeps its own */
+	if (ioctl(mdfd, RUN_ARRAY, NULL)) {
+		fprintf(stderr, "md: RUN_ARRAY: %s\n", strerror(errno));
+		return CHILD_HOLDER;
+	}
+	if (write_fat12(mdfd))
+		return CHILD_HOLDER;
+	close(mdfd);
+	if (mount("/dev/md0", HOLDER_BELOW, "vfat", 0, NULL) &&
+	    mount("/dev/md0", HOLDER_BELOW, "msdos", 0, NULL))
+		return CHILD_HOLDER;
+	return CHILD_OK;
+}
+
+/*
+ * In its own mount namespace the child puts the image on a tmpfs next to
+ * P, mounts P from a loop device, and sets the holder up with its own
+ * filesystem or device mounted at C below P.
+ */
+static int holder_child(int to_parent, int from_parent, enum holder holder)
+{
+	bool fifo = holder == HOLDER_AUTOFS || holder == HOLDER_ZRAM;
+	/* room for a 4 MiB zone file or a floppy image on P */
+	bool big = holder == HOLDER_ZLOOP || holder == HOLDER_GADGET;
+	const char *type = fifo ? "minix" : "vfat";
+	char dev[32], c;
+	int ifd, n, ret;
+
+	if (unshare(CLONE_NEWNS) || mount("", "/", NULL, MS_REC | MS_PRIVATE, NULL))
+		return CHILD_NS;
+	if (mkdir("/mnt_dir/img", 0755) || mount("tmpfs", "/mnt_dir/img", "tmpfs", 0, NULL))
+		return CHILD_NS;
+
+	ifd = open(HOLDER_IMG, O_RDWR | O_CREAT | O_EXCL, 0600);
+	if (ifd < 0 || (fifo ? write_minix(ifd) :
+			big ? write_fat12_4k(ifd, 3072) : write_fat12(ifd)))
+		return CHILD_IMAGE;
+	close(ifd);
+	n = loop_attach(HOLDER_IMG);
+	if (n < 0)
+		return CHILD_LOOP;
+	snprintf(dev, sizeof(dev), "/dev/loop%d", n);
+	if (mkdir(HOLDER_MNT, 0755))
+		return CHILD_MOUNT;
+	if (mount(dev, HOLDER_MNT, type, 0, NULL) &&
+	    (fifo || mount(dev, HOLDER_MNT, "msdos", 0, NULL)))
+		return fifo && errno == ENODEV ? CHILD_SKIP : CHILD_MOUNT;
+	if (mkdir(HOLDER_BELOW, 0755))
+		return CHILD_MOUNT;
+
+	switch (holder) {
+	case HOLDER_AUTOFS:
+		ret = holder_autofs();
+		break;
+	case HOLDER_ZRAM:
+		ret = holder_zram();
+		break;
+	case HOLDER_ECRYPTFS:
+		ret = holder_ecryptfs();
+		break;
+	case HOLDER_BINFMT_MISC:
+		ret = holder_binfmt_misc();
+		break;
+	case HOLDER_FUSE:
+		ret = holder_fuse();
+		break;
+	case HOLDER_ZLOOP:
+		ret = holder_zloop();
+		break;
+	case HOLDER_GADGET:
+		ret = holder_gadget();
+		break;
+	case HOLDER_MD:
+		ret = holder_md();
+		break;
+	default:
+		ret = CHILD_HOLDER;
+	}
+	if (ret != CHILD_OK)
+		return ret;
+
+	if (write(to_parent, &n, sizeof(n)) != sizeof(n))
+		return CHILD_PIPE;
+	if (read(from_parent, &c, 1) != 1)
+		return CHILD_PIPE;
+	return CHILD_OK;
+}
+
+/*
+ * A device keeps its file for as long as it is configured, so once nothing
+ * below the dead mount is left the device has to be told to let go. With
+ * the cycle unbroken C's filesystem still holds the device and every one
+ * of these refuses.
+ */
+static int holder_let_go_once(enum holder holder)
+{
+	int fd, ret;
+
+	switch (holder) {
+	case HOLDER_ZRAM:
+		return write_file("/sys/block/zram0/reset", "1");
+	case HOLDER_ZLOOP:
+		return write_file("/dev/zloop-control", "remove id=0");
+	case HOLDER_GADGET:
+		if (mount("configfs", "/sys/kernel/config", "configfs", 0, NULL) && errno != EBUSY)
+			return -1;
+		/* a zero-length write is a no-op for configfs; a newline ejects */
+		return write_file(GADGET "/functions/mass_storage.0/lun.0/file", "\n");
+	case HOLDER_MD:
+		fd = open("/dev/md0", O_RDONLY);
+		if (fd < 0)
+			return -1;
+		ret = ioctl(fd, STOP_ARRAY);
+		close(fd);
+		return ret;
+	default:
+		return 0;
+	}
+}
+
+/*
+ * The release of C's filesystem may still be in flight when the child is
+ * gone, and a device that is still held refuses to let go, so try for a
+ * while. With the cycle unbroken it refuses for good.
+ */
+static void holder_let_go(enum holder holder)
+{
+	for (int i = 0; i < 50; i++) {
+		if (!holder_let_go_once(holder))
+			return;
+		usleep(100000);
+	}
+}
+
+/*
+ * rmdir of /mnt_dir/p from here, where it is a plain directory, detaches
+ * P in the child's namespace with C connected below it. Once the child is
+ * gone the holder's file on P is the only thing left that refers to P,
+ * and it is dropped only when C's filesystem dies, which waits for P.
+ * The loop device backing P tells whether that resolved.
+ */
+static void holder_cycle(struct __test_metadata *_metadata,
+			 FIXTURE_DATA(loop_cycle) *self, enum holder holder)
+{
+	int to_parent[2], from_parent[2], status, n;
+	pid_t pid;
+
+	ASSERT_EQ(pipe(to_parent), 0);
+	ASSERT_EQ(pipe(from_parent), 0);
+	pid = fork();
+	ASSERT_GE(pid, 0);
+	if (pid == 0) {
+		close(to_parent[0]);
+		close(from_parent[1]);
+		_exit(holder_child(to_parent[1], from_parent[0], holder));
+	}
+	close(to_parent[1]);
+	close(from_parent[0]);
+
+	if (read(to_parent[0], &n, sizeof(n)) != sizeof(n)) {
+		waitpid(pid, &status, 0);
+		if (WIFEXITED(status) && WEXITSTATUS(status) == CHILD_SKIP)
+			SKIP(return, "the kernel lacks what this holder needs");
+		ASSERT_TRUE(false)
+			TH_LOG("child failed to set the holder up: exit status %d",
+			       WIFEXITED(status) ? WEXITSTATUS(status) : -1);
+	}
+	snprintf(self->dev, sizeof(self->dev), "/dev/loop%d", n);
+	snprintf(self->sysfs, sizeof(self->sysfs), "/sys/block/loop%d/loop/backing_file", n);
+
+	ASSERT_EQ(rmdir(HOLDER_MNT), 0);
+	ASSERT_EQ(write(from_parent[1], "x", 1), 1);
+	ASSERT_EQ(waitpid(pid, &status, 0), pid);
+	ASSERT_EQ(WEXITSTATUS(status), CHILD_OK);
+
+	holder_let_go(holder);
+	assert_loop_released(_metadata, self);
+	if (holder == HOLDER_GADGET)
+		write_file(GADGET "/UDC", "\n");
+}
+
+TEST_F(loop_cycle, autofs_pipe_on_dead_mount_released)
+{
+	holder_cycle(_metadata, self, HOLDER_AUTOFS);
+}
+
+TEST_F(loop_cycle, zram_writeback_node_on_dead_mount_released)
+{
+	holder_cycle(_metadata, self, HOLDER_ZRAM);
+}
+
+TEST_F(loop_cycle, ecryptfs_lower_on_dead_mount_released)
+{
+	holder_cycle(_metadata, self, HOLDER_ECRYPTFS);
+}
+
+TEST_F(loop_cycle, binfmt_misc_interpreter_on_dead_mount_released)
+{
+	holder_cycle(_metadata, self, HOLDER_BINFMT_MISC);
+}
+
+TEST_F(loop_cycle, fuse_backing_file_on_dead_mount_released)
+{
+	holder_cycle(_metadata, self, HOLDER_FUSE);
+}
+
+TEST_F(loop_cycle, zloop_zone_files_on_dead_mount_released)
+{
+	holder_cycle(_metadata, self, HOLDER_ZLOOP);
+}
+
+TEST_F(loop_cycle, mass_storage_lun_on_dead_mount_released)
+{
+	holder_cycle(_metadata, self, HOLDER_GADGET);
+}
+
+TEST_F(loop_cycle, md_bitmap_file_on_dead_mount_released)
+{
+	holder_cycle(_metadata, self, HOLDER_MD);
+}
+
 TEST_HARNESS_MAIN

-- 
2.53.0


^ permalink raw reply related	[flat|nested] 9+ messages in thread

end of thread, other threads:[~2026-09-24 22:36 UTC | newest]

Thread overview: 9+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-24 22:35 [PATCH RFC v2 0/8] namespace: prevent UMOUNT_CONNECTED reference count cycles Christian Brauner
2026-09-24 22:35 ` [PATCH RFC v2 1/8] fs: refuse fspick() on internal superblocks Christian Brauner
2026-09-24 22:35 ` [PATCH RFC v2 2/8] fsnotify: record the superblock a connector is accounted on Christian Brauner
2026-09-24 22:35 ` [PATCH RFC v2 3/8] fs: put the old fs_struct before the old namespaces in unshare() Christian Brauner
2026-09-24 22:35 ` [PATCH RFC v2 4/8] namespace: drop the file's reference first in dissolve_on_fput() Christian Brauner
2026-09-24 22:35 ` [PATCH RFC v2 5/8] namespace: prevent UMOUNT_CONNECTED reference count cycles Christian Brauner
2026-09-24 22:35 ` [PATCH RFC v2 6/8] selftests/filesystems: check that a loop mount below a dead mount is released Christian Brauner
2026-09-24 22:35 ` [PATCH RFC v2 7/8] selftests/filesystems: check the two-step cycle over crossed loop images Christian Brauner
2026-09-24 22:35 ` [PATCH RFC v2 8/8] selftests/filesystems: check that the holders let go of a dead mount Christian Brauner

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox