All of lore.kernel.org
 help / color / mirror / Atom feed
From: Shakeel Butt <shakeel.butt@linux.dev>
To: Peter Zijlstra <peterz@infradead.org>,
	Ingo Molnar <mingo@redhat.com>, Will Deacon <will@kernel.org>,
	Boqun Feng <boqun@kernel.org>, Waiman Long <longman@redhat.com>
Cc: Paul McKenney <paulmck@kernel.org>,
	Greg Kroah-Hartman <gregkh@linuxfoundation.org>,
	Tejun Heo <tj@kernel.org>,
	Christian Brauner <christian@brauner.io>,
	Sebastian Andrzej Siewior <bigeasy@linutronix.de>,
	Johannes Weiner <hannes@cmpxchg.org>,
	Jonathan Corbet <corbet@lwn.net>,
	Meta kernel team <kernel-team@meta.com>,
	cgroups@vger.kernel.org, driver-core@lists.linux.dev,
	linux-kernel@vger.kernel.org
Subject: [PATCH RESEND 7/7] kernfs, cgroup: track holders of the cgroupfs locks
Date: Thu, 24 Sep 2026 08:50:25 -0700	[thread overview]
Message-ID: <20260924155025.949998-8-shakeel.butt@linux.dev> (raw)
In-Reply-To: <20260924155025.949998-1-shakeel.butt@linux.dev>

Opt the cgroupfs kernfs rwsems, cgroup_mutex and cgroup_threadgroup_rwsem
in to lock holder tracking.

Add KERNFS_ROOT_TRACK_LOCK_HOLDERS, which opts in a root's kernfs_rwsem,
kernfs_iattr_rwsem and kernfs_supers_rwsem, and pass it from
cgroup_setup_root(). Other kernfs users have their own roots, so only
cgroupfs is affected. All three are held across work that can allocate,
so a holder stuck in reclaim can stall many other cgroupfs operations;
kernfs_iattr_rwsem is write-held across kernfs_iattrs(), which allocates
with GFP_KERNEL.

cgroup_mutex is taken by every cgroup create, destroy and migrate, and is
held across the same kind of work, so a stalled holder stalls the whole
cgroupfs control plane.

cgroup_threadgroup_rwsem is taken for read on every fork and exit, and
usually for write on migration, so a stalled holder can stall process
creation on the whole machine.

The last two are statically defined, so opt them in from cgroup_init().
That runs after jump_label_init() and before anything can take either
lock.

signal->cgroup_threadgroup_rwsem, used with favordynmods, is left alone.
Its readers also hold cgroup_threadgroup_rwsem, so they are counted
already, and its writer only blocks the process being migrated.

These opt-ins are for a policy that reads the count. On their own they
only turn on the static branch, so they should go in with that policy.

Signed-off-by: Shakeel Butt <shakeel.butt@linux.dev>
---
 fs/kernfs/dir.c        |  5 +++++
 include/linux/kernfs.h |  7 +++++++
 kernel/cgroup/cgroup.c | 27 ++++++++++++++++++++++++++-
 3 files changed, 38 insertions(+), 1 deletion(-)

diff --git a/fs/kernfs/dir.c b/fs/kernfs/dir.c
index cc6288d5b4cc..1116bb38c1ff 100644
--- a/fs/kernfs/dir.c
+++ b/fs/kernfs/dir.c
@@ -1035,6 +1035,11 @@ struct kernfs_root *kernfs_create_root(struct kernfs_syscall_ops *scops,
 	init_rwsem(&root->kernfs_rwsem);
 	init_rwsem(&root->kernfs_iattr_rwsem);
 	init_rwsem(&root->kernfs_supers_rwsem);
+	if (flags & KERNFS_ROOT_TRACK_LOCK_HOLDERS) {
+		rwsem_track_holder(&root->kernfs_rwsem);
+		rwsem_track_holder(&root->kernfs_iattr_rwsem);
+		rwsem_track_holder(&root->kernfs_supers_rwsem);
+	}
 	INIT_LIST_HEAD(&root->supers);
 	rwlock_init(&root->kernfs_rename_lock);
 
diff --git a/include/linux/kernfs.h b/include/linux/kernfs.h
index 73786b567ccd..ea2efdc5277e 100644
--- a/include/linux/kernfs.h
+++ b/include/linux/kernfs.h
@@ -156,6 +156,13 @@ enum kernfs_root_flag {
 	 * Renames must not change the parent node.
 	 */
 	KERNFS_ROOT_INVARIANT_PARENT		= 0x0010,
+
+	/*
+	 * Track the holders of this root's rwsems (see
+	 * CONFIG_TRACK_LOCK_HOLDERS).  Use it for roots whose locks are
+	 * shared so widely that a stalled holder stalls unrelated work.
+	 */
+	KERNFS_ROOT_TRACK_LOCK_HOLDERS		= 0x0020,
 };
 
 /* type-specific structures for kernfs_node union members */
diff --git a/kernel/cgroup/cgroup.c b/kernel/cgroup/cgroup.c
index 804318ae160e..87d66aae2558 100644
--- a/kernel/cgroup/cgroup.c
+++ b/kernel/cgroup/cgroup.c
@@ -2191,11 +2191,18 @@ int cgroup_setup_root(struct cgroup_root *root, u32 ss_mask)
 	kf_sops = root == &cgrp_dfl_root ?
 		&cgroup_kf_syscall_ops : &cgroup1_kf_syscall_ops;
 
+	/*
+	 * Track the holders of the cgroupfs kernfs locks.  They are held
+	 * across work that can charge memory, so a holder stuck in reclaim
+	 * can stall many other cgroupfs operations.  Other kernfs users have
+	 * their own roots and are not affected.
+	 */
 	root->kf_root = kernfs_create_root(kf_sops,
 					   KERNFS_ROOT_CREATE_DEACTIVATED |
 					   KERNFS_ROOT_SUPPORT_EXPORTOP |
 					   KERNFS_ROOT_SUPPORT_USER_XATTR |
-					   KERNFS_ROOT_INVARIANT_PARENT,
+					   KERNFS_ROOT_INVARIANT_PARENT |
+					   KERNFS_ROOT_TRACK_LOCK_HOLDERS,
 					   root_cgrp);
 	if (IS_ERR(root->kf_root)) {
 		ret = PTR_ERR(root->kf_root);
@@ -6552,6 +6559,24 @@ int __init cgroup_init(void)
 	get_user_ns(init_cgroup_ns.user_ns);
 	cgroup_rt_init();
 
+	/*
+	 * cgroup_mutex is taken by every cgroup create, destroy and migrate
+	 * and is held across work that can allocate, so a holder stuck in
+	 * reclaim stalls the whole cgroupfs control plane.
+	 *
+	 * cgroup_threadgroup_rwsem is taken for read by fork and exit, and
+	 * usually for write by migration, so a stalled holder can stall
+	 * process creation on the whole machine.
+	 *
+	 * Neither can be held yet, as rest_init() has not run.
+	 *
+	 * signal->cgroup_threadgroup_rwsem is not tracked.  Its readers hold
+	 * cgroup_threadgroup_rwsem too, so they are counted already, and its
+	 * writer only blocks the process being migrated.
+	 */
+	mutex_track_holder(&cgroup_mutex);
+	percpu_rwsem_track_holder(&cgroup_threadgroup_rwsem);
+
 	cgroup_lock();
 
 	/*
-- 
2.53.0-Meta


  parent reply	other threads:[~2026-09-24 15:51 UTC|newest]

Thread overview: 11+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-24 15:50 [PATCH RESEND 0/7] locking: opt-in tracking of sleeping lock holders Shakeel Butt
2026-09-24 15:50 ` [PATCH RESEND 1/7] locking: add " Shakeel Butt
2026-09-24 15:50 ` [PATCH RESEND 2/7] locking/rwsem: track holders of opted-in rw_semaphores Shakeel Butt
2026-09-24 15:50 ` [PATCH RESEND 3/7] locking/mutex: track holders of opted-in mutexes Shakeel Butt
2026-09-24 15:50 ` [PATCH RESEND 4/7] locking/percpu-rwsem: track holders of opted-in percpu_rw_semaphores Shakeel Butt
2026-09-24 15:50 ` [PATCH RESEND 5/7] locking/selftests: add KUnit tests for lock holder tracking Shakeel Butt
2026-09-24 15:50 ` [PATCH RESEND 6/7] Documentation/locking: document sleeping " Shakeel Butt
2026-09-24 15:50 ` Shakeel Butt [this message]
2026-09-28 23:34 ` [PATCH RESEND 0/7] locking: opt-in tracking of sleeping lock holders Tejun Heo
2026-10-01 10:41   ` Peter Zijlstra
2026-10-01 11:44     ` Shakeel Butt

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260924155025.949998-8-shakeel.butt@linux.dev \
    --to=shakeel.butt@linux.dev \
    --cc=bigeasy@linutronix.de \
    --cc=boqun@kernel.org \
    --cc=cgroups@vger.kernel.org \
    --cc=christian@brauner.io \
    --cc=corbet@lwn.net \
    --cc=driver-core@lists.linux.dev \
    --cc=gregkh@linuxfoundation.org \
    --cc=hannes@cmpxchg.org \
    --cc=kernel-team@meta.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=longman@redhat.com \
    --cc=mingo@redhat.com \
    --cc=paulmck@kernel.org \
    --cc=peterz@infradead.org \
    --cc=tj@kernel.org \
    --cc=will@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.