From: Christian Brauner <brauner@kernel.org>
To: Linus Torvalds <torvalds@linux-foundation.org>
Cc: Jann Horn <jannh@google.com>, Jan Kara <jack@suse.cz>,
Amir Goldstein <amir73il@gmail.com>,
linux-fsdevel@vger.kernel.org,
Alexander Viro <viro@zeniv.linux.org.uk>,
"Christian Brauner (Amutable)" <brauner@kernel.org>
Subject: [PATCH RFC 0/6] namespace: prevent UMOUNT_CONNECTED reference count cycles
Date: Thu, 24 Sep 2026 00:18:49 +0200 [thread overview]
Message-ID: <20260924-work-mount-knullfs-v1-0-ae89b29f7cb3@kernel.org> (raw)
Have barf bags ready, please. Afaict, UMOUNT_CONNECTED as implemented
allows for the creation of reference count cycles. Here's a simple
example:
mkdir /x; mkfifo /ready /go
unshare -m sh -c 'mount -t tmpfs tmpfs /x
truncate -s 8M /x/img; mkfs.ext4 -q /x/img
dev=$(losetup -f --show /x/img)
mkdir /x/mp; mount $dev /x/mp
echo $dev > /ready; read r < /go' &
read dev < /ready
rmdir /x
echo > /go
wait
losetup -d $dev
losetup -a
Take a directory /x on the host, create a new mount namespace, mount a
tmpfs on /x. Now rmdir /x on the host. This will lazily unmount the
mount on top of /x in the container with UMOUNT_CONNECTED. Once the
namespace exists nothing references the mount anymore. Now the tmpfs is
pinned by the backing file of the loop device and the loop mount is
owned by the tmpfs superblock.
Fun fact, such cycles can be formed by at least the following
subsystems and I have added reproducers for all of them:
(1) a loop mount P from an image on a tmpfs next to it, so that P's
death shows as the loop device giving up its backing file
(2) autofs with a FIFO on P as its pipe, zram with a device node on P as
its writeback device, both on a minix image since vfat has neither
(3) ecryptfs with its lower directory on P, under a passphrase token
added to the session keyring
(4) binfmt_misc in a new user namespace with an 'F' interpreter on P
(5) a fuse server that answers FUSE_INIT with passthrough on and
registers a file on P as a backing file
(6) zloop with its zone files in a directory on P
(7) a mass storage gadget on the dummy UDC with its LUN file on P,
mounted from the SCSI disk the gadget shows up as
(8) md with a RAID1 of one loop device and its bitmap file on P, which
skips while SET_BITMAP_FILE has no way to succeed
(9) rmdir of P's mountpoint from the parent, then the child exits, then
the device must be free and LOOP_CLR_FD must release the file
I have explored various solution and have branches for most of them. All
suck ass. Highlights include to port everything to use private mounts
similar to what overlayfs does. It's ugly as fuck and it needs a
side-channel to communicate to umount that the underlying thing like the
loop device is still in use so we don't cause spurious EBUSY errors.
It's really not nice.
The underlying mechanism is UMOUNT_CONNECTED (MNT_LOCKED falls into the
same class). With UMOUNT_CONNECTED an unmounted mount stays attached to
its parent. This is used to protect revealing the underlying mount and
is a non-negotiable security mechanism. So now the parent owns that
mount and is put on the parent's final mntput(). That moves it to
mnt_stuck_children and ultimately it's cleaned up by cleanup_mnt().
The fact that ownership of the child mount gets transferred to the
parent turns every reference from a child's superblock back to one of
its ancestors into a cycle.
Don't let the parent own the children. A mount that stays attached
keeps its own reference. namespace_unlock() drops it parents first.
A subtree that nobody refers to now also collapses exactly like a
disconnected one.
So, the problematic case was always a mount that loses its last
reference while it is still attached. That can reveal the covered
directory and that's caused fun exploits. UMOUNT_CONNECTED was always a
sucky mechanism imho.
Instead of that, when a mount loses its last reference but is still
attached to its parent we know that it is a UMOUNT_CONNECTED or
MNT_LOCKED case. Don't take it out of the hash. Instead make it a nullfs
mount. It's an immutable directory that is part of no namespace. Let
that nullfs mount be owned by the parent and put by the parent's final
mntput() the way locked mounts always were.
Ownership can't form a cycle anymore. A vacant mount is pointing at
knullfs. That's nullfs instance that can't go away and doesn't have any
child mounts whatsoever. Nothing leads from a vacant mount back to any
other mount.
Any lookup that still reaches the parent now finds an empty read-only
directory at the mountpoint. Creating anything in it fails with ENOENT,
like every lookup in nullfs does.
The only visible change is for a locked mount under a lazily unmounted
one. It used to stay alive and traversable for as long as something held
the parent. Now it is released with the umount when it isn't referenced
anymore and an empty directory takes its place. What it covered stays
covered either way.
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
Christian Brauner (6):
fs: refuse fspick() on internal superblocks
fsnotify: record the superblock a connector is accounted on
namespace: prevent UMOUNT_CONNECTED reference count cycles
selftests/filesystems: check that a loop mount below a dead mount is released
selftests/filesystems: check the two-step cycle over crossed loop images
selftests/filesystems: check that the holders let go of a dead mount
fs/fsopen.c | 3 +
fs/mount.h | 12 +-
fs/namespace.c | 146 ++-
fs/notify/fsnotify.h | 3 +-
fs/notify/mark.c | 7 +-
include/linux/fsnotify_backend.h | 2 +
tools/testing/selftests/Makefile | 1 +
.../selftests/filesystems/mount_cycle/.gitignore | 2 +
.../selftests/filesystems/mount_cycle/Makefile | 6 +
.../filesystems/mount_cycle/loop_cycle_test.c | 1225 ++++++++++++++++++++
10 files changed, 1367 insertions(+), 40 deletions(-)
---
base-commit: 2d2a2d7aa98741b58f54cacc99b52024e4d865f9
change-id: 20260923-work-mount-knullfs-4daab4eb9f06
next reply other threads:[~2026-09-23 22:19 UTC|newest]
Thread overview: 8+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-23 22:18 Christian Brauner [this message]
2026-09-23 22:18 ` [PATCH RFC 1/6] fs: refuse fspick() on internal superblocks Christian Brauner
2026-09-23 22:18 ` [PATCH RFC 2/6] fsnotify: record the superblock a connector is accounted on Christian Brauner
2026-09-24 8:57 ` Amir Goldstein
2026-09-23 22:18 ` [PATCH RFC 3/6] namespace: prevent UMOUNT_CONNECTED reference count cycles Christian Brauner
2026-09-23 22:18 ` [PATCH RFC 4/6] selftests/filesystems: check that a loop mount below a dead mount is released Christian Brauner
2026-09-23 22:18 ` [PATCH RFC 5/6] selftests/filesystems: check the two-step cycle over crossed loop images Christian Brauner
2026-09-23 22:18 ` [PATCH RFC 6/6] selftests/filesystems: check that the holders let go of a dead mount Christian Brauner
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260924-work-mount-knullfs-v1-0-ae89b29f7cb3@kernel.org \
--to=brauner@kernel.org \
--cc=amir73il@gmail.com \
--cc=jack@suse.cz \
--cc=jannh@google.com \
--cc=linux-fsdevel@vger.kernel.org \
--cc=torvalds@linux-foundation.org \
--cc=viro@zeniv.linux.org.uk \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox