The Linux Kernel Mailing List
 help / color / mirror / Atom feed
* Re: [PATCH] rcu: Reduce stack usage in show_rcu_gp_kthreads()
       [not found] <20260723100430.17813-1-qiang.zhang@linux.dev>
@ 2026-07-23 18:30 ` Paul E. McKenney
  2026-07-24 13:20 ` David Laight
  1 sibling, 0 replies; 3+ messages in thread
From: Paul E. McKenney @ 2026-07-23 18:30 UTC (permalink / raw)
  To: Zqiang
  Cc: frederic, neeraj.upadhyay, joelagnelf, urezki, boqun, rcu,
	linux-kernel

On Thu, Jul 23, 2026 at 06:04:30PM +0800, Zqiang wrote:
> When CONFIG_KASAN=y and CONFIG_KASAN_STACK=y builds, the
> show_rcu_gp_kthreads() exceeds the 1024-byte frame-size limit:
> 
> make kernel/rcu/tree.o KCFLAGS="-fstack-usage"
>   DESCEND objtool
>   DESCEND bpf/resolve_btfids
>   INSTALL libsubcmd_headers
>   CC      kernel/rcu/tree.o
> In file included from kernel/rcu/tree.c:4998:
> kernel/rcu/tree_stall.h: In function 'show_rcu_gp_kthreads':
> kernel/rcu/tree_stall.h:994:1: warning: the frame size of 1656 bytes is larger than 1024 bytes [-Wframe-larger-than=]
> 
> grep show_rcu kernel/rcu/tree.su
> tree_nocb.h:1622:13:show_rcu_nocb_state 896     dynamic,bounded
> tree_stall.h:933:6:show_rcu_gp_kthreads 1784    dynamic,bounded
> tree_stall.h:1102:13:sysrq_show_rcu     16      static
> 
> Wrap the pr_info() into two noinline_for_stack helpers function:
> show_rcu_state() print rcu_state status, and show_rcu_node()
> print single rcu_node status.
> 
> After apply this change:
> 
> grep show_rcu kernel/rcu/tree.su
> tree_stall.h:955:22:show_rcu_node	696	dynamic,bounded
> tree_stall.h:930:22:show_rcu_state	872	dynamic,bounded
> tree_nocb.h:1622:13:show_rcu_nocb_state	896	dynamic,bounded
> tree_stall.h:972:6:show_rcu_gp_kthreads	544	static
> tree_stall.h:1113:13:sysrq_show_rcu	16	static
> 
> Signed-off-by: Zqiang <qiang.zhang@linux.dev>

Queued and pushed for further review and testing, thank you!

							Thanx, Paul

> ---
>  kernel/rcu/tree_stall.h | 47 +++++++++++++++++++++++++----------------
>  1 file changed, 29 insertions(+), 18 deletions(-)
> 
> diff --git a/kernel/rcu/tree_stall.h b/kernel/rcu/tree_stall.h
> index 45b9856ccd2b..a6dfd036a7e9 100644
> --- a/kernel/rcu/tree_stall.h
> +++ b/kernel/rcu/tree_stall.h
> @@ -927,20 +927,13 @@ bool rcu_check_boost_fail(unsigned long gp_state, int *cpup)
>  }
>  EXPORT_SYMBOL_GPL(rcu_check_boost_fail);
>  
> -/*
> - * Show the state of the grace-period kthreads.
> - */
> -void show_rcu_gp_kthreads(void)
> +static noinline_for_stack void show_rcu_state(void)
>  {
> -	unsigned long cbs = 0;
> -	int cpu;
>  	unsigned long j;
>  	unsigned long ja;
>  	unsigned long jr;
>  	unsigned long js;
>  	unsigned long jw;
> -	struct rcu_data *rdp;
> -	struct rcu_node *rnp;
>  	struct task_struct *t = READ_ONCE(rcu_state.gp_kthread);
>  
>  	j = jiffies;
> @@ -957,21 +950,39 @@ void show_rcu_gp_kthreads(void)
>  		(long)data_race(READ_ONCE(rcu_get_root()->gp_seq_needed)),
>  		data_race(READ_ONCE(rcu_state.gp_max)),
>  		data_race(READ_ONCE(rcu_state.gp_flags)));
> +}
> +
> +static noinline_for_stack void show_rcu_node(struct rcu_node *rnp)
> +{
> +	pr_info("\trcu_node %d:%d ->gp_seq %ld ->gp_seq_needed %ld ->qsmask %#lx %c%c%c%c ->n_boosts %ld\n",
> +		rnp->grplo, rnp->grphi,
> +		(long)data_race(READ_ONCE(rnp->gp_seq)),
> +		(long)data_race(READ_ONCE(rnp->gp_seq_needed)),
> +		data_race(READ_ONCE(rnp->qsmask)),
> +		".b"[!!data_race(READ_ONCE(rnp->boost_kthread_task))],
> +		".B"[!!data_race(READ_ONCE(rnp->boost_tasks))],
> +		".E"[!!data_race(READ_ONCE(rnp->exp_tasks))],
> +		".G"[!!data_race(READ_ONCE(rnp->gp_tasks))],
> +		data_race(READ_ONCE(rnp->n_boosts)));
> +}
> +
> +/*
> + * Show the state of the grace-period kthreads.
> + */
> +void show_rcu_gp_kthreads(void)
> +{
> +	unsigned long cbs = 0;
> +	int cpu;
> +	struct rcu_data *rdp;
> +	struct rcu_node *rnp;
> +
> +	show_rcu_state();
>  	rcu_for_each_node_breadth_first(rnp) {
>  		if (ULONG_CMP_GE(READ_ONCE(rcu_state.gp_seq), READ_ONCE(rnp->gp_seq_needed)) &&
>  		    !data_race(READ_ONCE(rnp->qsmask)) && !data_race(READ_ONCE(rnp->boost_tasks)) &&
>  		    !data_race(READ_ONCE(rnp->exp_tasks)) && !data_race(READ_ONCE(rnp->gp_tasks)))
>  			continue;
> -		pr_info("\trcu_node %d:%d ->gp_seq %ld ->gp_seq_needed %ld ->qsmask %#lx %c%c%c%c ->n_boosts %ld\n",
> -			rnp->grplo, rnp->grphi,
> -			(long)data_race(READ_ONCE(rnp->gp_seq)),
> -			(long)data_race(READ_ONCE(rnp->gp_seq_needed)),
> -			data_race(READ_ONCE(rnp->qsmask)),
> -			".b"[!!data_race(READ_ONCE(rnp->boost_kthread_task))],
> -			".B"[!!data_race(READ_ONCE(rnp->boost_tasks))],
> -			".E"[!!data_race(READ_ONCE(rnp->exp_tasks))],
> -			".G"[!!data_race(READ_ONCE(rnp->gp_tasks))],
> -			data_race(READ_ONCE(rnp->n_boosts)));
> +		show_rcu_node(rnp);
>  		if (!rcu_is_leaf_node(rnp))
>  			continue;
>  		for_each_leaf_node_possible_cpu(rnp, cpu) {
> -- 
> 2.17.1
> 

^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH] rcu: Reduce stack usage in show_rcu_gp_kthreads()
       [not found] <20260723100430.17813-1-qiang.zhang@linux.dev>
  2026-07-23 18:30 ` [PATCH] rcu: Reduce stack usage in show_rcu_gp_kthreads() Paul E. McKenney
@ 2026-07-24 13:20 ` David Laight
  2026-07-24 18:11   ` Paul E. McKenney
  1 sibling, 1 reply; 3+ messages in thread
From: David Laight @ 2026-07-24 13:20 UTC (permalink / raw)
  To: Zqiang
  Cc: paulmck, frederic, neeraj.upadhyay, joelagnelf, urezki, boqun,
	rcu, linux-kernel

On Thu, 23 Jul 2026 18:04:30 +0800
Zqiang <qiang.zhang@linux.dev> wrote:

> When CONFIG_KASAN=y and CONFIG_KASAN_STACK=y builds, the
> show_rcu_gp_kthreads() exceeds the 1024-byte frame-size limit:
> 
...
>  	rcu_for_each_node_breadth_first(rnp) {
>  		if (ULONG_CMP_GE(READ_ONCE(rcu_state.gp_seq), READ_ONCE(rnp->gp_seq_needed)) &&
>  		    !data_race(READ_ONCE(rnp->qsmask)) && !data_race(READ_ONCE(rnp->boost_tasks)) &&
>  		    !data_race(READ_ONCE(rnp->exp_tasks)) && !data_race(READ_ONCE(rnp->gp_tasks)))
>  			continue;
> -		pr_info("\trcu_node %d:%d ->gp_seq %ld ->gp_seq_needed %ld ->qsmask %#lx %c%c%c%c ->n_boosts %ld\n",
> -			rnp->grplo, rnp->grphi,
> -			(long)data_race(READ_ONCE(rnp->gp_seq)),
> -			(long)data_race(READ_ONCE(rnp->gp_seq_needed)),
> -			data_race(READ_ONCE(rnp->qsmask)),
> -			".b"[!!data_race(READ_ONCE(rnp->boost_kthread_task))],
> -			".B"[!!data_race(READ_ONCE(rnp->boost_tasks))],
> -			".E"[!!data_race(READ_ONCE(rnp->exp_tasks))],
> -			".G"[!!data_race(READ_ONCE(rnp->gp_tasks))],
> -			data_race(READ_ONCE(rnp->n_boosts)));

Isn't that code carefully reading all the values twice?

Also the "ab"[!!val] generates far worse code than the more obvious (val ? 'a' : 'b').
For the former gcc indexes a constant string, the latter is done using arithmetic.

I'm sure there is a good reason for data_race(READ_ONCE(xxx)) ...

	David

^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH] rcu: Reduce stack usage in show_rcu_gp_kthreads()
  2026-07-24 13:20 ` David Laight
@ 2026-07-24 18:11   ` Paul E. McKenney
  0 siblings, 0 replies; 3+ messages in thread
From: Paul E. McKenney @ 2026-07-24 18:11 UTC (permalink / raw)
  To: David Laight
  Cc: Zqiang, frederic, neeraj.upadhyay, joelagnelf, urezki, boqun, rcu,
	linux-kernel

On Fri, Jul 24, 2026 at 02:20:48PM +0100, David Laight wrote:
> On Thu, 23 Jul 2026 18:04:30 +0800
> Zqiang <qiang.zhang@linux.dev> wrote:
> 
> > When CONFIG_KASAN=y and CONFIG_KASAN_STACK=y builds, the
> > show_rcu_gp_kthreads() exceeds the 1024-byte frame-size limit:
> > 
> ...
> >  	rcu_for_each_node_breadth_first(rnp) {
> >  		if (ULONG_CMP_GE(READ_ONCE(rcu_state.gp_seq), READ_ONCE(rnp->gp_seq_needed)) &&
> >  		    !data_race(READ_ONCE(rnp->qsmask)) && !data_race(READ_ONCE(rnp->boost_tasks)) &&
> >  		    !data_race(READ_ONCE(rnp->exp_tasks)) && !data_race(READ_ONCE(rnp->gp_tasks)))
> >  			continue;
> > -		pr_info("\trcu_node %d:%d ->gp_seq %ld ->gp_seq_needed %ld ->qsmask %#lx %c%c%c%c ->n_boosts %ld\n",
> > -			rnp->grplo, rnp->grphi,
> > -			(long)data_race(READ_ONCE(rnp->gp_seq)),
> > -			(long)data_race(READ_ONCE(rnp->gp_seq_needed)),
> > -			data_race(READ_ONCE(rnp->qsmask)),
> > -			".b"[!!data_race(READ_ONCE(rnp->boost_kthread_task))],
> > -			".B"[!!data_race(READ_ONCE(rnp->boost_tasks))],
> > -			".E"[!!data_race(READ_ONCE(rnp->exp_tasks))],
> > -			".G"[!!data_race(READ_ONCE(rnp->gp_tasks))],
> > -			data_race(READ_ONCE(rnp->n_boosts)));
> 
> Isn't that code carefully reading all the values twice?

You are quite right, it does, and it did before Zqiang's patch.  I would
welcome a patch that carefully used locals to avoid the double reads,
as that could avoid some debug confusion that might otherwise result
due to low-probability races with updates to these same variables.
However, this function is invoked quite rarely in error conditions,
so I would also avoid holding anything up waiting for such a patch.

> Also the "ab"[!!val] generates far worse code than the more obvious (val ? 'a' : 'b').
> For the former gcc indexes a constant string, the latter is done using arithmetic.

This code is invoked quite rarely, so I am not concerned about the
overhead and I like the compactness.  In the longer term, perhaps
compilers will recognize the pattern and optimize it.  Or not, who
knows?  ;-)

> I'm sure there is a good reason for data_race(READ_ONCE(xxx)) ...

And there is!!!

The data_race() says "ignore the contents from a data-race viewpoint"
and the READ_ONCE() says "carefully avoid destructive optimizations that
might otherwise be applied to this load."

For an example, suppose that we had a global variable x that was accessed
only under a lock.  Except that we have lockless access for debugging
output like that in Zqiang's patch.  We want the normal lock-protected
accesses to be plain C-language accesses so that any accidental (and thus
likely buggy) lockless accesses are spotted by KCSAN, but we very much
to *not* want to acquire the lock when doing the debug output, because
if the lock was held, we might get a debug-output-free hang just when
we are in the greatest need of that debugging output.

So we use normal C-language accesses for the non-debug uses, and we use
data_race(READ_ONCE(x)) for the debugging output.

Does that make sense?

							Thanx, Paul

^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-07-24 18:11 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
     [not found] <20260723100430.17813-1-qiang.zhang@linux.dev>
2026-07-23 18:30 ` [PATCH] rcu: Reduce stack usage in show_rcu_gp_kthreads() Paul E. McKenney
2026-07-24 13:20 ` David Laight
2026-07-24 18:11   ` Paul E. McKenney

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox