Linux filesystem development
 help / color / mirror / Atom feed
From: Christian Brauner <brauner@kernel.org>
To: linux-fsdevel@vger.kernel.org
Cc: Linus Torvalds <torvalds@linux-foundation.org>,
	 Chris Mason <mason@kernel.org>,
	Alexander Viro <viro@zeniv.linux.org.uk>,
	 Jan Kara <jack@suse.cz>, Jeff Layton <jlayton@kernel.org>,
	 Aleksa Sarai <cyphar@cyphar.com>,
	Amir Goldstein <amir73il@gmail.com>,
	 bpf@vger.kernel.org,
	"Christian Brauner (Amutable)" <brauner@kernel.org>,
	 stable@vger.kernel.org
Subject: [PATCH 06/17] namespace: keep covered mounts covered in OPEN_TREE_NAMESPACE
Date: Wed, 30 Sep 2026 15:31:58 +0200	[thread overview]
Message-ID: <20260930-work-mount-fixes-3-v1-6-be34c83956ae@kernel.org> (raw)
In-Reply-To: <20260930-work-mount-fixes-3-v1-0-be34c83956ae@kernel.org>

open_tree(OPEN_TREE_NAMESPACE) creates a new mount namespace from a
copy of the tree at the given path. The caller needs to be privileged
over its current user namespace because the new mount namespace will be
owned by it. It doesn't need to be privileged over the mount namespace
the tree is copied from. An unprivileged user just needs to create a
user namespace first.

The copy follows the rules of a detached bind mount. Without
AT_RECURSIVE only the mount itself is cloned and its children are left
out. With AT_RECURSIVE unbindable mounts are skipped. Both is fine when
the caller has privileges over the source mount namespace because it
could unmount those mounts anyway. Not so for an unprivileged user in
that mount namespace. For them a mount namespace copy is what unshare()
does. copy_mnt_ns() copies everything including unbindable mounts and
lock_mnt_tree() makes sure nothing can be unmounted in the copy. Nothing
that was covered gets revealed.

create_new_namespace() only does the locking and so the covered content
is right there in the new mount namespace:

  # mount -t tmpfs none /tmp/otn; mkdir /tmp/otn/covered
  # echo hidden > /tmp/otn/covered/under.txt
  # mount -t tmpfs none /tmp/otn/covered

  uid 1000, after unshare(CLONE_NEWUSER):
  fd = open_tree(AT_FDCWD, "/tmp/otn", OPEN_TREE_NAMESPACE);
  setns(fd, CLONE_NEWNS);
  open("/covered/under.txt", O_RDONLY)  -> "hidden"

Covering paths with mounts is how container runtimes mask parts of
/proc and /sys and how admins hide things.

When the caller's user namespace doesn't own the source mount namespace
copy the way unshare() does and refuse a non-recursive copy of a mount
that has anything mounted below the requested directory and copy
unbindable mounts in a recursive copy. A caller that is privileged over
the source mount namespace sees no change.

Fixes: 9b8a0ba68246 ("mount: add OPEN_TREE_NAMESPACE")
Cc: stable@vger.kernel.org # v7.0+
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
 fs/namespace.c | 30 ++++++++++++++++++++++++++----
 1 file changed, 26 insertions(+), 4 deletions(-)

diff --git a/fs/namespace.c b/fs/namespace.c
index c2f54636ec9d..de3900dd02f6 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -2369,6 +2369,18 @@ bool has_locked_children(struct mount *mnt, struct dentry *dentry)
 	return __has_locked_children(mnt, dentry);
 }
 
+/* locks: namespace_shared && pinned(mnt) || mount_locked_reader */
+static bool __has_children(struct mount *mnt, struct dentry *dentry)
+{
+	struct mount *child;
+
+	list_for_each_entry(child, &mnt->mnt_mounts, mnt_child) {
+		if (is_subdir(child->mnt_mountpoint, dentry))
+			return true;
+	}
+	return false;
+}
+
 /*
  * Check that there aren't references to earlier/same mount namespaces in the
  * specified subtree.  Such references can act as pins for mount namespaces
@@ -3141,12 +3153,18 @@ static struct mnt_namespace *create_new_namespace(struct path *path,
 	struct mount *mnt;
 	unsigned int copy_flags = 0;
 	bool locked = false, recurse = flags & MOUNT_COPY_RECURSIVE;
+	bool foreign = user_ns != ns->user_ns;
 
 	if (unlikely(!d_can_lookup(path->dentry)))
 		return ERR_PTR(-ENOTDIR);
 
-	if (user_ns != ns->user_ns)
-		copy_flags |= CL_SLAVE;
+	/*
+	 * Without privileges over the mount namespace the copy is made from
+	 * nothing mounted below @path may be left out. It would reveal what
+	 * it covers. That's what unshare() gives such a caller as well.
+	 */
+	if (foreign)
+		copy_flags |= CL_SLAVE | CL_COPY_UNBINDABLE;
 
 	new_ns = alloc_mnt_ns(user_ns, false);
 	if (IS_ERR(new_ns))
@@ -3180,10 +3198,14 @@ static struct mnt_namespace *create_new_namespace(struct path *path,
 	/*
 	 * We don't emulate unshare()ing a mount namespace. We stick to
 	 * the restrictions of creating detached bind-mounts. It has a
-	 * lot saner and simpler semantics.
+	 * lot saner and simpler semantics. A caller without privileges
+	 * over the mount namespace can't leave out any child though.
 	 */
 	if (flags & MOUNT_COPY_NEW)
 		mnt = clone_mnt(real_mount(path->mnt), path->dentry, copy_flags);
+	else if (foreign && !recurse &&
+		 __has_children(real_mount(path->mnt), path->dentry))
+		mnt = ERR_PTR(-EINVAL);
 	else
 		mnt = __do_loopback(path, recurse, copy_flags);
 	scoped_guard(mount_writer) {
@@ -3200,7 +3222,7 @@ static struct mnt_namespace *create_new_namespace(struct path *path,
 		 * of the real rootfs we created.
 		 */
 		attach_mnt(mnt, new_ns_root, mp.mp);
-		if (user_ns != ns->user_ns)
+		if (foreign)
 			lock_mnt_tree(new_ns_root);
 	}
 

-- 
2.53.0


  parent reply	other threads:[~2026-09-30 13:32 UTC|newest]

Thread overview: 22+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-30 13:31 [PATCH 00/17] mount: more bugfixes, the Oprah edition Christian Brauner
2026-09-30 13:31 ` [PATCH 01/17] namespace: queue a mount only once for mount notifications Christian Brauner
2026-09-30 13:31 ` [PATCH 02/17] namespace: check a submount for references right before unmounting it Christian Brauner
2026-09-30 13:31 ` [PATCH 03/17] selftests/filesystems: check that a busy submount survives a synchronous umount Christian Brauner
2026-09-30 13:31 ` [PATCH 04/17] namespace: check a recursive bind mount for mount namespace loops Christian Brauner
2026-09-30 13:31 ` [PATCH 05/17] selftests/filesystems: check that a recursive bind mount can't pin the caller's mount namespace Christian Brauner
2026-09-30 13:31 ` Christian Brauner [this message]
2026-09-30 13:31 ` [PATCH 07/17] selftests/filesystems: check that OPEN_TREE_NAMESPACE keeps mounts covered Christian Brauner
2026-09-30 13:32 ` [PATCH 08/17] namespace: look at the topmost mount for a mount namespace file Christian Brauner
2026-09-30 13:32 ` [PATCH 09/17] selftests/filesystems: check that a mount namespace file on top doesn't bury a mount Christian Brauner
2026-09-30 13:32 ` [PATCH 10/17] namespace: check the mounts before reading their parents in pivot_root() Christian Brauner
2026-09-30 13:32 ` [PATCH 11/17] namespace: don't reconfigure internal superblocks via remount and umount Christian Brauner
2026-09-30 13:32 ` [PATCH 12/17] selftests/filesystems: check that the nullfs root can't be reconfigured Christian Brauner
2026-09-30 13:32 ` [PATCH 13/17] namespace: remove the fsnotify marks of a mount namespace in process context Christian Brauner
2026-09-30 15:07   ` Amir Goldstein
2026-09-30 13:32 ` [PATCH 14/17] fsnotify: detach the connector before destroying its marks Christian Brauner
2026-10-01  9:31   ` Christian Brauner
2026-10-01 10:58     ` Amir Goldstein
2026-10-01 12:06       ` Christian Brauner
2026-09-30 13:32 ` [PATCH 15/17] dcache: don't put a mountpoint on a dentry that's being removed Christian Brauner
2026-09-30 13:32 ` [PATCH 16/17] unshare: don't drop active namespace references that were never taken Christian Brauner
2026-09-30 13:32 ` [PATCH 17/17] namespace: don't let a pseudo dentry become the root of a mount Christian Brauner

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260930-work-mount-fixes-3-v1-6-be34c83956ae@kernel.org \
    --to=brauner@kernel.org \
    --cc=amir73il@gmail.com \
    --cc=bpf@vger.kernel.org \
    --cc=cyphar@cyphar.com \
    --cc=jack@suse.cz \
    --cc=jlayton@kernel.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=mason@kernel.org \
    --cc=stable@vger.kernel.org \
    --cc=torvalds@linux-foundation.org \
    --cc=viro@zeniv.linux.org.uk \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox