All of lore.kernel.org
 help / color / mirror / Atom feed
From: Christian Brauner <brauner@kernel.org>
To: linux-fsdevel@vger.kernel.org
Cc: Linus Torvalds <torvalds@linux-foundation.org>,
	 Chris Mason <mason@kernel.org>,
	Alexander Viro <viro@zeniv.linux.org.uk>,
	 Jan Kara <jack@suse.cz>, Jeff Layton <jlayton@kernel.org>,
	 Aleksa Sarai <cyphar@cyphar.com>,
	Amir Goldstein <amir73il@gmail.com>,
	 bpf@vger.kernel.org,
	"Christian Brauner (Amutable)" <brauner@kernel.org>,
	 stable@vger.kernel.org
Subject: [PATCH 01/17] namespace: queue a mount only once for mount notifications
Date: Wed, 30 Sep 2026 15:31:53 +0200	[thread overview]
Message-ID: <20260930-work-mount-fixes-3-v1-1-be34c83956ae@kernel.org> (raw)
In-Reply-To: <20260930-work-mount-fixes-3-v1-0-be34c83956ae@kernel.org>

mnt_notify_add() puts a mount on notify_list. It doesn't check whether
the mount is on the list already. That's fine as long as every mount is
queued at most once per namespace_sem hold. It isn't.

propagate_umount() moves a surviving overmount off a stack of mounts
that are going away and queues it as moved. shrink_submounts() and
mark_mounts_for_expiry() call umount_tree() for several mounts under a
single namespace_sem hold. So the second umount_tree() can take down
exactly the mount the first one reparented, reparent it once more and so
end up queueing it a second time.

The second list_add_tail() cuts the mounts queued in between out of the
list while the head still points at the last of them. notify_mnt_list()
then only visits that one mount and the others are freed after the grace
period with notify_list still pointing at them. Every later mount
operation on the host walks freed memory:

  BUG: KASAN: slab-use-after-free in __list_add_valid_or_report
  Read of size 8 at addr ffff8881003d57e0 by task notify_dq/162
   __list_add_valid_or_report+0x15c/0x1a0
   mnt_add_to_ns+0x322/0x890
   attach_recursive_mnt.isra.0+0xf2a/0x1a00
   path_mount+0x139c/0x1d40

This needs a mount with MNT_SHRINKABLE to bind from. That's what
finish_automount() creates (submounts of NFS, AFS and CIFS and the
tracefs mount below debugfs). Bind mounts inherit the flag. A mount
namespace with a FAN_MARK_MNTNS mark does the rest and root in a user
namespace can have both.

Initialize to_notify in alloc_vfsmnt() and leave a mount alone that is
queued already. mnt_notify() looks at the state the mount has when the
list is drained so a single entry per mount is enough. The move event
for a mount that is unmounted under the same hold is lost. That's fine.
The detach is what matters.

Fixes: bf630c401641 ("vfs: add notifications for mount attach and detach")
Cc: stable@vger.kernel.org # v6.15+
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/mount.h     | 3 +++
 fs/namespace.c | 3 +++
 2 files changed, 6 insertions(+)

diff --git a/fs/mount.h b/fs/mount.h
index 85f136786bbc..2223fb141499 100644
--- a/fs/mount.h
+++ b/fs/mount.h
@@ -234,6 +234,9 @@ static inline struct mnt_namespace *to_mnt_ns(struct ns_common *ns)
 #ifdef CONFIG_FSNOTIFY
 static inline void mnt_notify_add(struct mount *m)
 {
+	/* queued already under this namespace_sem hold */
+	if (!list_empty(&m->to_notify))
+		return;
 	/* Optimize the case where there are no watches */
 	if ((m->mnt_ns && m->mnt_ns->n_fsnotify_marks) ||
 	    (m->prev_ns && m->prev_ns->n_fsnotify_marks))
diff --git a/fs/namespace.c b/fs/namespace.c
index 5b44eee28f39..2a1e77c0cb6e 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -333,6 +333,9 @@ static struct mount *alloc_vfsmnt(const char *name)
 		INIT_HLIST_NODE(&mnt->mnt_mp_list);
 		INIT_HLIST_HEAD(&mnt->mnt_stuck_children);
 		INIT_HLIST_NODE(&mnt->mnt_ns_visible);
+#ifdef CONFIG_FSNOTIFY
+		INIT_LIST_HEAD(&mnt->to_notify);
+#endif
 		RB_CLEAR_NODE(&mnt->mnt_node);
 		mnt->mnt.mnt_idmap = &nop_mnt_idmap;
 	}

-- 
2.53.0


  reply	other threads:[~2026-09-30 13:32 UTC|newest]

Thread overview: 26+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-30 13:31 [PATCH 00/17] mount: more bugfixes, the Oprah edition Christian Brauner
2026-09-30 13:31 ` Christian Brauner [this message]
2026-09-30 13:31 ` [PATCH 02/17] namespace: check a submount for references right before unmounting it Christian Brauner
2026-09-30 13:31 ` [PATCH 03/17] selftests/filesystems: check that a busy submount survives a synchronous umount Christian Brauner
2026-09-30 13:44   ` sashiko-bot
2026-09-30 13:31 ` [PATCH 04/17] namespace: check a recursive bind mount for mount namespace loops Christian Brauner
2026-09-30 13:31 ` [PATCH 05/17] selftests/filesystems: check that a recursive bind mount can't pin the caller's mount namespace Christian Brauner
2026-09-30 13:31 ` [PATCH 06/17] namespace: keep covered mounts covered in OPEN_TREE_NAMESPACE Christian Brauner
2026-09-30 13:31 ` [PATCH 07/17] selftests/filesystems: check that OPEN_TREE_NAMESPACE keeps mounts covered Christian Brauner
2026-09-30 13:43   ` sashiko-bot
2026-09-30 13:32 ` [PATCH 08/17] namespace: look at the topmost mount for a mount namespace file Christian Brauner
2026-09-30 13:32 ` [PATCH 09/17] selftests/filesystems: check that a mount namespace file on top doesn't bury a mount Christian Brauner
2026-09-30 13:40   ` sashiko-bot
2026-09-30 13:32 ` [PATCH 10/17] namespace: check the mounts before reading their parents in pivot_root() Christian Brauner
2026-09-30 13:32 ` [PATCH 11/17] namespace: don't reconfigure internal superblocks via remount and umount Christian Brauner
2026-09-30 13:32 ` [PATCH 12/17] selftests/filesystems: check that the nullfs root can't be reconfigured Christian Brauner
2026-09-30 13:32 ` [PATCH 13/17] namespace: remove the fsnotify marks of a mount namespace in process context Christian Brauner
2026-09-30 15:07   ` Amir Goldstein
2026-09-30 13:32 ` [PATCH 14/17] fsnotify: detach the connector before destroying its marks Christian Brauner
2026-09-30 13:57   ` sashiko-bot
2026-10-01  9:31   ` Christian Brauner
2026-10-01 10:58     ` Amir Goldstein
2026-10-01 12:06       ` Christian Brauner
2026-09-30 13:32 ` [PATCH 15/17] dcache: don't put a mountpoint on a dentry that's being removed Christian Brauner
2026-09-30 13:32 ` [PATCH 16/17] unshare: don't drop active namespace references that were never taken Christian Brauner
2026-09-30 13:32 ` [PATCH 17/17] namespace: don't let a pseudo dentry become the root of a mount Christian Brauner

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260930-work-mount-fixes-3-v1-1-be34c83956ae@kernel.org \
    --to=brauner@kernel.org \
    --cc=amir73il@gmail.com \
    --cc=bpf@vger.kernel.org \
    --cc=cyphar@cyphar.com \
    --cc=jack@suse.cz \
    --cc=jlayton@kernel.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=mason@kernel.org \
    --cc=stable@vger.kernel.org \
    --cc=torvalds@linux-foundation.org \
    --cc=viro@zeniv.linux.org.uk \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.