Linux Documentation
 help / color / mirror / Atom feed
* [PATCH] rcu: Enable runtime reset of the RCU stall panic count
@ 2026-10-04  1:14 Cunlong Li
  2026-10-04 11:04 ` Bradley Morgan
  2026-10-05 12:56 ` Joel Fernandes
  0 siblings, 2 replies; 10+ messages in thread
From: Cunlong Li @ 2026-10-04  1:14 UTC (permalink / raw)
  To: Jonathan Corbet, Shuah Khan, Randy Dunlap, Paul E. McKenney,
	Frederic Weisbecker, Neeraj Upadhyay, Joel Fernandes,
	Josh Triplett, Boqun Feng, Uladzislau Rezki, Steven Rostedt,
	Mathieu Desnoyers, Lai Jiangshan, Zqiang
  Cc: linux-doc, linux-kernel, rcu, Cunlong Li

The kernel.panic_on_rcu_stall and kernel.max_rcu_stall_to_panic sysctls
count RCU CPU stalls since system boot, and invoke panic() once
max_rcu_stall_to_panic stalls have elapsed.  This count is never reset,
so stalls caused by transient incidents keep consuming the budget of a
long-running system, and a later unrelated stall can immediately
trigger panic() instead of allowing the intended fresh window of
stalls.

This commit therefore introduces the kernel.rcu_stall_panic_count
sysctl.  Reading this file reports the number of stalls counted so far,
and writing 0 to it resets the count.  This allows the cumulative stall
history to be cleared after an incident has been resolved, and also
allows userspace to clear the count periodically, so that panic() is
only triggered by a burst of stalls occurring within a short period.

Signed-off-by: Cunlong Li <shenxiaogll@gmail.com>
---
 Documentation/admin-guide/sysctl/kernel.rst | 15 ++++++++++++++-
 kernel/rcu/tree_stall.h                     | 28 ++++++++++++++++++++++++++--
 2 files changed, 40 insertions(+), 3 deletions(-)

diff --git a/Documentation/admin-guide/sysctl/kernel.rst b/Documentation/admin-guide/sysctl/kernel.rst
index ffea61d448eb..64fe2985e358 100644
--- a/Documentation/admin-guide/sysctl/kernel.rst
+++ b/Documentation/admin-guide/sysctl/kernel.rst
@@ -959,7 +959,20 @@ max_rcu_stall_to_panic
 When ``panic_on_rcu_stall`` is set to 1, this value determines the
 number of times that RCU can stall before panic() is called.
 
-When ``panic_on_rcu_stall`` is set to 0, this value is has no effect.
+When ``panic_on_rcu_stall`` is set to 0, this value has no effect.
+
+rcu_stall_panic_count
+=====================
+
+Indicates the number of RCU CPU stalls that have been counted since
+system boot or since the counter was reset. When ``panic_on_rcu_stall``
+is set to 1, this count is compared against ``max_rcu_stall_to_panic``
+to decide whether panic() should be called.
+
+Writing 0 to this file resets the counter to zero, which restarts the
+``max_rcu_stall_to_panic`` window of stalls. This allows system
+administrators to clear the cumulative stall count after an incident
+has been resolved, without requiring a system restart.
 
 perf_cpu_time_max_percent
 =========================
diff --git a/kernel/rcu/tree_stall.h b/kernel/rcu/tree_stall.h
index 091e7850ab6e..a80f03e1c7ea 100644
--- a/kernel/rcu/tree_stall.h
+++ b/kernel/rcu/tree_stall.h
@@ -19,6 +19,20 @@
 /* panic() on RCU Stall sysctl. */
 static int sysctl_panic_on_rcu_stall __read_mostly;
 static int sysctl_max_rcu_stall_to_panic __read_mostly;
+static unsigned long sysctl_rcu_stall_panic_count;
+
+/* Reset the RCU stall panic count when written to. */
+static int proc_do_rcu_stall_panic_count(const struct ctl_table *table, int write,
+					 void *buffer, size_t *lenp, loff_t *ppos)
+{
+	if (!write)
+		return proc_doulongvec_minmax(table, write, buffer, lenp, ppos);
+
+	WRITE_ONCE(sysctl_rcu_stall_panic_count, 0);
+	*ppos += *lenp;
+
+	return 0;
+}
 
 static const struct ctl_table rcu_stall_sysctl_table[] = {
 	{
@@ -39,6 +53,13 @@ static const struct ctl_table rcu_stall_sysctl_table[] = {
 		.extra1		= SYSCTL_ONE,
 		.extra2		= SYSCTL_INT_MAX,
 	},
+	{
+		.procname	= "rcu_stall_panic_count",
+		.data		= &sysctl_rcu_stall_panic_count,
+		.maxlen		= sizeof(sysctl_rcu_stall_panic_count),
+		.mode		= 0644,
+		.proc_handler	= proc_do_rcu_stall_panic_count,
+	},
 };
 
 static int __init init_rcu_stall_sysctl(void)
@@ -161,7 +182,7 @@ early_initcall(check_cpu_stall_init);
 /* If so specified via sysctl, panic, yielding cleaner stall-warning output. */
 static void panic_on_rcu_stall(const struct cpumask *stalled_mask)
 {
-	static int cpu_stall;
+	unsigned long count;
 
 	/*
 	 * Attempt to kick out the BPF scheduler if it's installed and defer
@@ -170,7 +191,10 @@ static void panic_on_rcu_stall(const struct cpumask *stalled_mask)
 	if (scx_rcu_cpu_stall(stalled_mask))
 		return;
 
-	if (++cpu_stall < sysctl_max_rcu_stall_to_panic)
+	/* A lost RMW update only delays the panic by one stall. */
+	count = READ_ONCE(sysctl_rcu_stall_panic_count) + 1;
+	WRITE_ONCE(sysctl_rcu_stall_panic_count, count);
+	if (count < (unsigned long)READ_ONCE(sysctl_max_rcu_stall_to_panic))
 		return;
 
 	if (sysctl_panic_on_rcu_stall)

---
base-commit: ce1e0223d8ad4211275c82a17ed6d43ab81e13d9
change-id: 20261003-rcu-375496d7704e

Best regards,
-- 
Cunlong Li <shenxiaogll@gmail.com>


^ permalink raw reply related	[flat|nested] 10+ messages in thread
* Re: [PATCH] rcu: Enable runtime reset of the RCU stall panic count
@ 2026-10-07 10:46 Joel Fernandes
  0 siblings, 0 replies; 10+ messages in thread
From: Joel Fernandes @ 2026-10-07 10:46 UTC (permalink / raw)
  To: Cunlong Li
  Cc: Jonathan Corbet, Shuah Khan, Randy Dunlap, Paul E. McKenney,
	Frederic Weisbecker, Neeraj Upadhyay, Josh Triplett, Boqun Feng,
	Uladzislau Rezki, Steven Rostedt, Mathieu Desnoyers,
	Lai Jiangshan, Zqiang, linux-doc@vger.kernel.org,
	linux-kernel@vger.kernel.org, rcu@vger.kernel.org



> On Oct 6, 2026, at 10:56 PM, Cunlong Li <shenxiaogll@gmail.com> wrote:
> 
> On Tue, Oct 06, 2026 at 05:27:51PM -0400, Joel Fernandes wrote:
>> 
>> 
>> On 10/6/2026 9:17 AM, Cunlong Li wrote:
>>>> It sounds like you do know that it is likely transient - but based on this description and the below, it sounds like you have not yet found a use of this patch.
>>>> 
>>>> 
>>> To be precise, we worry that some of the stalls were transient, so a
>>> panic reboot would affect the users.  But we have no hard evidence
>>> either way: every incident ended in a forced reset, and the dmesg was
>>> lost.  Going forward, we plan to enable the NMI and kdump to keep
>>> investigating the cause of the RCU stalls. We will also monitor dmesg
>>> to confirm whether transient RCU stalls actually occur.
>> 
>> 
>> Right, I just worry that if you're not seeing recoverable/transient stalls and
>> you start using this patch, then you might be causing side-effects that don't
>> really help your problem and might be masking the real issue. It would be good
>> to debug this further, and fix the right way IMO.
>> 
>> If you show that this patch indeed helps your problem, we can consider it for
>> 7.5, but there's little bit of time. :-)
> 
> Hi Joel,
> 
> Understood, thanks.  We will keep monitoring dmesg for the root cause.
> Once we observe self-recovering RCU stalls, we will resubmit the patch
> with the evidence.

Thank you!!!

Joel

> 
> Thanks,
> Cunlong
> 
>> 
>> Thanks,
>> 
>> Joel Fernandes
>> 
>> 

^ permalink raw reply	[flat|nested] 10+ messages in thread

end of thread, other threads:[~2026-10-07 10:46 UTC | newest]

Thread overview: 10+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-04  1:14 [PATCH] rcu: Enable runtime reset of the RCU stall panic count Cunlong Li
2026-10-04 11:04 ` Bradley Morgan
2026-10-05 12:56 ` Joel Fernandes
2026-10-06 11:54   ` Cunlong Li
2026-10-06 12:21     ` Joel Fernandes
2026-10-06 12:39     ` Joel Fernandes
2026-10-06 13:17       ` Cunlong Li
2026-10-06 21:27         ` Joel Fernandes
2026-10-07  2:56           ` Cunlong Li
  -- strict thread matches above, loose matches on Subject: below --
2026-10-07 10:46 Joel Fernandes

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox