From: Christian Brauner <brauner@kernel.org>
To: linux-fsdevel@vger.kernel.org
Cc: Linus Torvalds <torvalds@linux-foundation.org>,
Alexander Viro <viro@zeniv.linux.org.uk>,
Jan Kara <jack@suse.cz>,
"Christian Brauner (Amutable)" <brauner@kernel.org>,
stable@vger.kernel.org
Subject: [PATCH 5/8] fs: don't silently unmount busy mounts
Date: Wed, 23 Sep 2026 14:27:57 +0200 [thread overview]
Message-ID: <20260923-work-mount-fixes-v1-5-f424cf8d3242@kernel.org> (raw)
In-Reply-To: <20260923-work-mount-fixes-v1-0-f424cf8d3242@kernel.org>
The propagate_mount_busy() helper exists to decide whether a synchronous
can suceed. It does that by looking at the copies of the victims in the
mounts its parent propagates to. One of those functions from propagation
hell.
It skips a copy that has child mounts except for the case of a single
child mount covering the copy's root. That algorithm used to be exact.
It isn't anymore.
Originally only childless copies were unmounted during propagation.
And 1064f874abc0 ("mnt: Tuck mounts under others instead of creating
shadow/side mounts.") ended up adding covered mounts to both sides in
one go.
Along came 99b19d16471e ("mnt: In propgate_umount handle visiting mounts
in any order") and that started to unmount more than before:
[1]: A copy of a mount is unmounted when each of its children is either
its overmount or a copy of the victim itself. In the former case
the overmount gets reparented. In the latter case the copy of the
victim will end up being unmounted.
So propagate_umount() ended up being rewritten without adjusting
propagate_mount_busy(). And currently propagate_mount_busy() doesn't
look at copies of the victim that have two or more children nor a single
child that is not the overmount.
So any such copy in [1] is not checked for
references. The result is that a synchronous unmount succeeds, pulling
out a mount that is still in use.
A container can run into this with its own mounts. The host shares a
tree with the container and the host has a mount in that tree:
mount -t tmpfs none /mnt
mount --make-shared /mnt
mount -t tmpfs none /mnt/a
The container's /mnt is a slave copy of the host mount. The container
now takes a detached copy of the tree and mounts that copy beneath its
propagated copy of /mnt/a and keeps the file descriptor to the tree
open:
unshare -m --propagation unchanged
mount --make-rslave /
fd = open_tree(AT_FDCWD, "/mnt", OPEN_TREE_CLONE | AT_RECURSIVE)
move_mount(fd, "", AT_FDCWD, "/mnt/a", MOVE_MOUNT_BENEATH)
grep /mnt /proc/self/mountinfo
775 355 ... /mnt ... master:646
777 775 ... /mnt/a ... master:646
776 777 ... /mnt/a ... master:647
778 777 ... /mnt/a/a ... master:647
The now attached tree (777) is now mounted beneath the container's copy
of /mnt/a. That's also where the host mount sits (as seen from the host
mount namespace). So the propagated copy (776) of the host's mount is on
top of its root and the copy of that (778) is inside it. Now the host
runs:
umount /mnt/a
and it succeeds. The moved tree tree and the copy inside of it are gone
from the container. Now only the propagated copy is left and it got
reparented to where it was before:
grep /mnt /proc/self/mountinfo
775 355 ... /mnt ... master:646
776 775 ... /mnt/a ...
But the file descriptor still refers to the moved tree. So a synchronous
umount detached a mount that is still in use.
Teach propagate_mount_busy() the same checks that propagate_umount()
does. It now checks for a single victim mount without children
for which it can assume that MNT_LOCKED is cleared on every copy like
propagate_umount() does. The copies form chains. IOW, when a copy itself
receives propagation from the victim's parent then the mount at the same
mountpoint inside that copy is a copy as well. Whether a mount receives
propagation from the victim's can simply be checked by walking its
masters. Walk each chain once from its top so we are sure that nested
copies stay linear. Check every reference once of every copy that gets
unmounted.
Fixes: 99b19d16471e ("mnt: In propgate_umount handle visiting mounts in any order")
Cc: stable@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
fs/pnode.c | 108 ++++++++++++++++++++++++++++++++++++++++++++++++++-----------
1 file changed, 90 insertions(+), 18 deletions(-)
diff --git a/fs/pnode.c b/fs/pnode.c
index 5d91c3e58d2a..2cd667958efe 100644
--- a/fs/pnode.c
+++ b/fs/pnode.c
@@ -410,19 +410,99 @@ bool propagation_would_overmount(const struct mount *from,
return false;
}
+/* Does @m receive propagation from @parent? */
+static bool receives_from(struct mount *m, struct mount *parent)
+{
+ if (m == parent)
+ return false;
+ for (; m; m = m->mnt_master)
+ if (m == parent || peers(m, parent))
+ return true;
+ return false;
+}
+
+/*
+ * Does @m receive propagation from the victim's parent as well? If so, then
+ * the mount at the victim's mountpoint inside of @m is a umount candidate as
+ * well. So it's the next candidate in the chain. Otherwise the chain ends at
+ * @m.
+ */
+static struct mount *next_candidate(struct mount *m, struct mount *victim)
+{
+ if (!receives_from(m, victim->mnt_parent))
+ return NULL;
+ return __lookup_mnt(&m->mnt, victim->mnt_mountpoint);
+}
+
+/*
+ * Would propagate_umount() pull out a mount of the chain of candidates that
+ * starts at @c, and does that mount have references beyond its own?
+ *
+ * This mirrors how trim_one(), trim_ancestors() and handle_locked() handle a
+ * synchronous umount:
+ *
+ * - single victim
+ * - without children
+ * - with MNT_LOCKED already cleared on every candidate by propagate_mount_unlock()
+ *
+ * A copy of the victim gets unmounted when each of its children is
+ * the next candidate in the chain or its overmount, unless the next
+ * unmount candidate is not its overmount and some unmount candidate further
+ * down has a child outside the chain. Keep this in sync with
+ * Documentation/filesystems/propagate_umount.txt.
+ */
+static bool chain_busy(struct mount *c, struct mount *victim)
+{
+ struct mount *m, *n, *next, *deepest = NULL;
+ bool above;
+
+ /* the deepest candidate with a child outside the chain */
+ for (m = c; m; m = next) {
+ next = next_candidate(m, victim);
+ list_for_each_entry(n, &m->mnt_mounts, mnt_child) {
+ if (n != next && n != victim) {
+ deepest = m;
+ break;
+ }
+ }
+ }
+
+ above = deepest != NULL; /* @deepest is at or below @m */
+ for (m = c; m; m = next) {
+ bool goes = true;
+
+ next = next_candidate(m, victim);
+ list_for_each_entry(n, &m->mnt_mounts, mnt_child) {
+ if (n != next && n != m->overmount && n != victim) {
+ goes = false;
+ break;
+ }
+ }
+ if (goes && next && next != m->overmount && above && m != deepest)
+ goes = false;
+ if (m == deepest)
+ above = false;
+ if (goes && do_refcount_check(m, 1))
+ return true;
+ }
+ return false;
+}
+
/*
* check if the mount 'mnt' can be unmounted successfully.
* @mnt: the mount to be checked for unmount
* NOTE: unmounting 'mnt' would naturally propagate to all
* other mounts its parent propagates to.
- * Check if any of these mounts that **do not have submounts**
- * have more references than 'refcnt'. If so return busy.
+ * Check if any of the mounts that propagate_umount() would pull out
+ * along with it have more references than their own. If so return busy.
*
* vfsmount lock must be held for write
*/
int propagate_mount_busy(struct mount *mnt, int refcnt)
{
struct mount *parent = mnt->mnt_parent;
+ struct dentry *mp = mnt->mnt_mountpoint;
+ struct mount *m;
/*
* quickly check if the current mount can be unmounted.
@@ -435,24 +515,16 @@ int propagate_mount_busy(struct mount *mnt, int refcnt)
if (mnt == parent)
return 0;
- for (struct mount *m = propagation_next(parent, parent); m;
- m = propagation_next(m, parent)) {
- struct list_head *head;
- struct mount *child = __lookup_mnt(&m->mnt, mnt->mnt_mountpoint);
+ /* the candidates are the mounts at @mp below the receivers */
+ for (m = propagation_next(parent, parent); m;
+ m = propagation_next(m, parent)) {
+ struct mount *c = __lookup_mnt(&m->mnt, mp);
- if (!child)
+ /* each chain once, from its top: skip receivers that are candidates */
+ if (!c || (mnt_has_parent(m) && m->mnt_mountpoint == mp &&
+ receives_from(m->mnt_parent, parent)))
continue;
-
- head = &child->mnt_mounts;
- if (!list_empty(head)) {
- /*
- * a mount that covers child completely wouldn't prevent
- * it being pulled out; any other would.
- */
- if (!list_is_singular(head) || !child->overmount)
- continue;
- }
- if (do_refcount_check(child, 1))
+ if (chain_busy(c, mnt))
return 1;
}
return 0;
--
2.53.0
next prev parent reply other threads:[~2026-09-23 12:28 UTC|newest]
Thread overview: 9+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-23 12:27 [PATCH 0/8] mount: a few gnarly fixes Christian Brauner
2026-09-23 12:27 ` [PATCH 1/8] mount: keep a copied mount unbindable Christian Brauner
2026-09-23 12:27 ` [PATCH 2/8] selftests/filesystems: check that a copied mount namespace keeps unbindable Christian Brauner
2026-09-23 12:27 ` [PATCH 3/8] mount: refuse MOVE_MOUNT_SET_GROUP on an unbindable mount Christian Brauner
2026-09-23 12:27 ` [PATCH 4/8] selftests/move_mount_set_group: check that an unbindable target is refused Christian Brauner
2026-09-23 12:27 ` Christian Brauner [this message]
2026-09-23 12:27 ` [PATCH 6/8] selftests/filesystems: check that a busy propagated copy blocks a synchronous umount Christian Brauner
2026-09-23 12:27 ` [PATCH 7/8] fs: don't let a migrating task hide its reference from do_umount() Christian Brauner
2026-09-23 12:28 ` [PATCH 8/8] docs: update the unmount propagation rule Christian Brauner
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260923-work-mount-fixes-v1-5-f424cf8d3242@kernel.org \
--to=brauner@kernel.org \
--cc=jack@suse.cz \
--cc=linux-fsdevel@vger.kernel.org \
--cc=stable@vger.kernel.org \
--cc=torvalds@linux-foundation.org \
--cc=viro@zeniv.linux.org.uk \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox