Linux Trace Kernel
 help / color / mirror / Atom feed
From: "Paul E. McKenney" <paulmck@kernel.org>
To: "Masami Hiramatsu (Google)" <mhiramat@kernel.org>
Cc: Steven Rostedt <rostedt@goodmis.org>,
	Frederic Weisbecker <frederic@kernel.org>,
	Neeraj Upadhyay <neeraj.upadhyay@kernel.org>,
	Mathieu Desnoyers <mathieu.desnoyers@efficios.com>,
	Josef Bacik <josef@toxicpanda.com>,
	linux-kernel@vger.kernel.org, linux-trace-kernel@vger.kernel.org,
	rcu@vger.kernel.org
Subject: Re: [PATCH] fprobe: Use guard(rcu_sched_notrace) and check rcu_is_watching()
Date: Wed, 30 Sep 2026 11:47:47 -0700	[thread overview]
Message-ID: <1095daf7-fd78-4bbd-87f4-1d487d80fd53@paulmck-laptop> (raw)
In-Reply-To: <179064115227.394389.16910234241400391996.stgit@devnote2>

On Tue, Sep 29, 2026 at 09:19:12AM +0900, Masami Hiramatsu (Google) wrote:
> From: Masami Hiramatsu (Google) <mhiramat@kernel.org>
> 
> unregister_fprobe() and unregister_fprobe_async() (used by BPF
> kprobe-multi) rely on standard RCU grace periods (synchronize_rcu()
> and call_rcu()) to wait until in-flight fprobe handlers complete before
> freeing the fprobe.
> 
> However, if an fprobe handler executes while RCU is not watching (such
> as in the idle loop or nohz_full extended quiescent states), standard
> RCU does not track preemption-disabled sections. Consequently,
> synchronize_rcu() does not wait for those executions, which can lead
> to a use-after-free if the fprobe is freed immediately after
> unregistration. Ensure handlers exit early when !rcu_is_watching().
> 
> Furthermore, fprobe_fgraph_entry() and fprobe_ftrace_entry() previously
> used guard(rcu)() and rcu_read_lock(), which invoke lockdep on every
> hit under CONFIG_PROVE_LOCKING. This adds overhead and can cause lockdep
> recursion if probed functions interact with lockdep.
> 
> Since rhltable_lookup() and rhl_for_each_entry_rcu() use
> rcu_dereference_all_check() (which checks rcu_read_lock_any_held()),
> holding preemption disabled via rcu_read_lock_sched_notrace() is fully
> valid and sufficient so long as rcu_is_watching() is true.
> 
> Define and use guard(rcu_sched_notrace)() across fprobe_ftrace_entry(),
> fprobe_fgraph_entry(), and fprobe_return(). This eliminates fast-path
> rcu_read_lock() and lockdep overhead while guaranteeing safe grace
> period synchronization.
> 
> Reported-by: Sashiko <sashiko-bot@kernel.org>
> Closes: https://sashiko.dev/#/bug/linux-e46bcd68-4a56-4f19-a255-e3772980e5e3
> Fixes: 657b594b2084 ("fprobe: Fix unregister_fprobe() to wait for RCU grace period")
> Cc: stable@vger.kernel.org
> Assisted-by: LLM
> Signed-off-by: Masami Hiramatsu (Google) <mhiramat@kernel.org>

From an RCU perspective:

Reviewed-by: Paul E. McKenney <paulmck@kernel.org>

> ---
>  kernel/trace/fprobe.c |   26 ++++++++++++++++----------
>  1 file changed, 16 insertions(+), 10 deletions(-)
> 
> diff --git a/kernel/trace/fprobe.c b/kernel/trace/fprobe.c
> index 9f2d98181779..da286619c5d8 100644
> --- a/kernel/trace/fprobe.c
> +++ b/kernel/trace/fprobe.c
> @@ -47,6 +47,10 @@ static struct rhltable fprobe_ip_table;
>  static DEFINE_MUTEX(fprobe_mutex);
>  static struct fgraph_ops fprobe_graph_ops;
>  
> +DEFINE_LOCK_GUARD_0(rcu_sched_notrace,
> +		    rcu_read_lock_sched_notrace(),
> +		    rcu_read_unlock_sched_notrace())
> +
>  static u32 fprobe_node_hashfn(const void *data, u32 len, u32 seed)
>  {
>  	return hash_ptr(*(unsigned long **)data, 32);
> @@ -329,16 +333,14 @@ static void fprobe_ftrace_entry(unsigned long ip, unsigned long parent_ip,
>  	struct fprobe *fp;
>  	int bit;
>  
> +	if (!rcu_is_watching())
> +		return;
> +
>  	bit = ftrace_test_recursion_trylock(ip, parent_ip);
>  	if (bit < 0)
>  		return;
>  
> -	/*
> -	 * ftrace_test_recursion_trylock() disables preemption, but
> -	 * rhltable_lookup() checks whether rcu_read_lcok is held.
> -	 * So we take rcu_read_lock() here.
> -	 */
> -	rcu_read_lock();
> +	guard(rcu_sched_notrace)();
>  	head = rhltable_lookup(&fprobe_ip_table, &ip, fprobe_rht_params);
>  
>  	rhl_for_each_entry_rcu(node, pos, head, hlist) {
> @@ -353,7 +355,6 @@ static void fprobe_ftrace_entry(unsigned long ip, unsigned long parent_ip,
>  		else
>  			__fprobe_handler(ip, parent_ip, fp, fregs, NULL);
>  	}
> -	rcu_read_unlock();
>  	ftrace_test_recursion_unlock(bit);
>  }
>  NOKPROBE_SYMBOL(fprobe_ftrace_entry);
> @@ -567,10 +568,13 @@ static int fprobe_fgraph_entry(struct ftrace_graph_ent *trace, struct fgraph_ops
>  	struct fprobe *fp;
>  	int used, ret;
>  
> +	if (!rcu_is_watching())
> +		return 0;
> +
>  	if (WARN_ON_ONCE(!fregs))
>  		return 0;
>  
> -	guard(rcu)();
> +	guard(rcu_sched_notrace)();
>  	head = rhltable_lookup(&fprobe_ip_table, &func, fprobe_rht_params);
>  	reserved_words = 0;
>  	rhl_for_each_entry_rcu(node, pos, head, hlist) {
> @@ -665,13 +669,16 @@ static void fprobe_return(struct ftrace_graph_ret *trace,
>  	int size, curr;
>  	int size_words;
>  
> +	if (!rcu_is_watching())
> +		return;
> +
>  	fgraph_data = (unsigned long *)fgraph_retrieve_data(gops->idx, &size);
>  	if (WARN_ON_ONCE(!fgraph_data))
>  		return;
>  	size_words = SIZE_IN_LONG(size);
>  	ret_ip = ftrace_regs_get_instruction_pointer(fregs);
>  
> -	preempt_disable_notrace();
> +	guard(rcu_sched_notrace)();
>  
>  	curr = 0;
>  	while (size_words > curr) {
> @@ -687,7 +694,6 @@ static void fprobe_return(struct ftrace_graph_ret *trace,
>  		}
>  		curr += size;
>  	}
> -	preempt_enable_notrace();
>  }
>  NOKPROBE_SYMBOL(fprobe_return);
>  
> 

  reply	other threads:[~2026-09-30 18:47 UTC|newest]

Thread overview: 4+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-29  0:19 [PATCH] fprobe: Use guard(rcu_sched_notrace) and check rcu_is_watching() Masami Hiramatsu (Google)
2026-09-30 18:47 ` Paul E. McKenney [this message]
2026-09-30 20:03 ` Steven Rostedt
2026-10-01 13:23   ` Masami Hiramatsu (Google)

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=1095daf7-fd78-4bbd-87f4-1d487d80fd53@paulmck-laptop \
    --to=paulmck@kernel.org \
    --cc=frederic@kernel.org \
    --cc=josef@toxicpanda.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-trace-kernel@vger.kernel.org \
    --cc=mathieu.desnoyers@efficios.com \
    --cc=mhiramat@kernel.org \
    --cc=neeraj.upadhyay@kernel.org \
    --cc=rcu@vger.kernel.org \
    --cc=rostedt@goodmis.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox