* Re: [PATCH RFC v3 03/13] rcu-tasks: Add a Tasks RCU implementation for reader-marked trampolines [not found] ` <20260915-b4-rcu-tasks-preempt-qs-v3-3-0ad30c4c5ee7@toxicpanda.com> @ 2026-09-15 15:14 ` Frederic Weisbecker 2026-09-15 23:56 ` Paul E. McKenney 2026-09-17 20:20 ` Frederic Weisbecker 1 sibling, 1 reply; 20+ messages in thread From: Frederic Weisbecker @ 2026-09-15 15:14 UTC (permalink / raw) To: Josef Bacik Cc: Paul E. McKenney, Neeraj Upadhyay, Joel Fernandes, Boqun Feng, Thomas Gleixner, Peter Zijlstra, Steven Rostedt, Masami Hiramatsu, Mark Rutland, Jiri Olsa, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, x86, Catalin Marinas, Will Deacon, Puranjay Mohan, Xu Kuohai, Andy Lutomirski, Josh Triplett, Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang, Juergen Gross, Luis Chamberlain, Ihor Solodrai, linux-kernel, rcu, linux-trace-kernel, bpf, linux-arm-kernel, xen-devel Le Tue, Sep 15, 2026 at 01:17:30PM +0000, Josef Bacik a écrit : > Tasks RCU waits for every task to pass through a voluntary context > switch, usermode or idle, because a preempted task might be sitting in a > trampoline that is about to be freed and nothing marks it as such. With > PREEMPT_LAZY that is a poor fit for servers: cond_resched() is a no-op, > so a CPU-bound kthread only ever leaves the CPU by preemption, and one > such kthread holds every synchronize_rcu_tasks() caller -- ftrace and > BPF trampoline teardown under their mutexes, the kprobe jump optimizer > under text_mutex and cpus_read_lock() -- hostage for as long as it runs. > > Following the discussion on v2, take the other road: let the > architecture make its trampolines Tasks Trace RCU readers. When an > architecture selects HAVE_RCU_TRAMPOLINE_READERS it promises that every > trampoline whose lifetime Tasks RCU guards enters rcu_read_lock_trace() > (or its assembly equivalent) before calling out and leaves it before > returning, so a task anywhere inside such a call-out, preempted or not, > is an ordinary Tasks Trace reader. > > That leaves the few instructions of trampoline text before the reader > is entered and after it is left (plus, in a later patch, the bytes a > kprobe jump optimization is about to overwrite). A task can only linger > there by being interrupted there, and such text never calls anything > that schedules, so instead of tracking tasks we track CPUs: every pass > through __schedule() is a per-CPU quiescent event, except that the one > context switch that can catch a task at an arbitrary instruction -- a > preemption from irq exit -- first records the interrupted IP in the task > and parks it on a per-CPU list for the duration (reusing the fields and > lists the classic flavor keeps for its exit-path bookkeeping), and, if > the IP is inside such "unmarked" text, puts the task on a short holdout > list; the task takes itself off at its next context switch outside such > a preemption or irq-exit check that finds it elsewhere. Usermode (the > existing tick hook, or a nohz_full CPU in an RCU extended quiescent > state) and idle count as well. rcu_tasks_trampoline_text() does the > classification: anything outside core and module text, a new > .text..rcu_tramp section for C glue that trampolines call before it has > entered the reader (__rcu_trampoline), and an arch hook for things like > static ftrace stubs and return thunks. > > The grace period, run by the existing rcu_tasks kthread so that > call_rcu_tasks(), synchronize_rcu_tasks() and rcu_barrier_tasks() keep > their names and callers, is: wait for every online non-idle CPU to > context switch (nudging stragglers with resched_cpu() after a jiffy), > drain the holdout list as it stood, synchronize_rcu_tasks_trace() for > everything inside the readers, then one more CPU pass and drain for > tasks that have since left the reader into the trailing instructions. > That is bounded by a few jiffies, preempt-off latency and an SRCU grace > period rather than by the longest stretch any task runs without > sleeping, needs no per-task scan, and makes cond_resched_tasks_rcu_qs() > unnecessary on such architectures. As before, idle tasks are not > waited for. rcu_tasks_wait_irq_preempted() walks the parked lists for > the one caller (the kprobe jump optimizer, later in the series) that > makes ordinary text unsafe to be parked in and so has to wait out tasks > that were preempted there before it said so. > > The classic implementation is untouched and remains the default; the > new one is built only as CONFIG_TASKS_RCU_TRAMPOLINE_READERS when the > architecture opts in and uses the generic irq entry code, whose > reschedule check gains the rcu_tasks_irq_resched() call. Nothing > selects it yet. > > Suggested-by: Paul E. McKenney <paulmck@kernel.org> > Suggested-by: Alexei Starovoitov <ast@kernel.org> > Assisted-by: LLM > Signed-off-by: Josef Bacik <josef@toxicpanda.com> > --- > include/asm-generic/vmlinux.lds.h | 11 + > include/linux/rcupdate.h | 32 ++- > include/linux/sched.h | 1 + > kernel/entry/common.c | 8 +- > kernel/fork.c | 1 + > kernel/rcu/Kconfig | 22 ++ > kernel/rcu/tasks.h | 460 +++++++++++++++++++++++++++++++++++++- > kernel/rcu/update.c | 2 + > 8 files changed, 528 insertions(+), 9 deletions(-) > > diff --git a/include/asm-generic/vmlinux.lds.h b/include/asm-generic/vmlinux.lds.h > index b2988aa12f66..86e58c4fe370 100644 > --- a/include/asm-generic/vmlinux.lds.h > +++ b/include/asm-generic/vmlinux.lds.h > @@ -571,6 +571,16 @@ > __cpuidle_text_end = .; \ > __noinstr_text_end = .; > > +/* > + * C glue called directly from Tasks-RCU-protected trampolines, bounded so > + * that rcu_tasks_trampoline_text() can recognise it; see __rcu_trampoline. > + */ > +#define RCU_TRAMP_TEXT \ > + ALIGN_FUNCTION(); \ > + __rcu_tramp_text_start = .; \ > + *(.text..rcu_tramp) \ > + __rcu_tramp_text_end = .; > + > #define TEXT_SPLIT \ > __split_text_start = .; \ > *(.text.split .text.split.[0-9a-zA-Z_]*) \ > @@ -607,6 +617,7 @@ > TEXT_HOT \ > *(TEXT_MAIN .text.fixup) \ > NOINSTR_TEXT \ > + RCU_TRAMP_TEXT \ > *(.ref.text) > > /* sched.text is aling to function alignment to secure we have same > diff --git a/include/linux/rcupdate.h b/include/linux/rcupdate.h > index 44c07a66edff..fb2a3889a696 100644 > --- a/include/linux/rcupdate.h > +++ b/include/linux/rcupdate.h > @@ -50,6 +50,31 @@ token_context_lock_instance(RCU, RCU_BH); > /* Exported common interfaces */ > void call_rcu(struct rcu_head *head, rcu_callback_t func); > void rcu_barrier_tasks(void); > + > +/* > + * Trampoline-reader Tasks RCU (CONFIG_TASKS_RCU_TRAMPOLINE_READERS), see > + * kernel/rcu/tasks.h. rcu_tasks_irq_resched_enter()/_exit() bracket the > + * irq-exit preemption; rcu_tasks_trampoline_text() and the arch_ override > + * classify an interrupted IP; rcu_tasks_wait_irq_preempted() lets a caller > + * wait out tasks already preempted somewhere it is about to make unsafe. > + * __rcu_trampoline places C code that such trampolines call directly, before > + * it has entered its Tasks Trace reader, where that classification can see it. > + */ > +void rcu_tasks_irq_resched_enter(unsigned long ip); > +void rcu_tasks_irq_resched_exit(void); > +bool rcu_tasks_trampoline_text(unsigned long ip); > +bool arch_rcu_tasks_trampoline_text(unsigned long ip); > +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS > +void rcu_tasks_wait_irq_preempted(bool (*inside)(unsigned long ip)); > +#else > +static inline void rcu_tasks_wait_irq_preempted(bool (*inside)(unsigned long ip)) { } > +#endif > +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS > +/* Also keeps instrumentation calls out of the prologue, ahead of the reader. */ > +#define __rcu_trampoline __noinstr_section(".text..rcu_tramp") > +#else > +#define __rcu_trampoline > +#endif > void synchronize_rcu(void); > > /* > @@ -180,11 +205,16 @@ static inline void rcu_nocb_flush_deferred_wakeup(void) { } > #ifdef CONFIG_TASKS_RCU_GENERIC > > # ifdef CONFIG_TASKS_RCU > -# define rcu_tasks_classic_qs(t, preempt) \ > +# ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS > +void rcu_tasks_note_qs(struct task_struct *t, bool preempt); > +# define rcu_tasks_classic_qs(t, preempt) rcu_tasks_note_qs((t), (preempt)) > +# else > +# define rcu_tasks_classic_qs(t, preempt) \ > do { \ > if (!(preempt) && READ_ONCE((t)->rcu_tasks_holdout)) \ > WRITE_ONCE((t)->rcu_tasks_holdout, false); \ > } while (0) > +# endif > void call_rcu_tasks(struct rcu_head *head, rcu_callback_t func); > void synchronize_rcu_tasks(void); > void rcu_tasks_torture_stats_print(char *tt, char *tf); > diff --git a/include/linux/sched.h b/include/linux/sched.h > index 8b3d47a325cc..15beb44caa2c 100644 > --- a/include/linux/sched.h > +++ b/include/linux/sched.h > @@ -957,6 +957,7 @@ struct task_struct { > u8 rcu_tasks_holdout; > u8 rcu_tasks_idx; > int rcu_tasks_idle_cpu; > + unsigned long rcu_tasks_irq_ip; > struct list_head rcu_tasks_holdout_list; > int rcu_tasks_exit_cpu; > struct list_head rcu_tasks_exit_list; > diff --git a/kernel/entry/common.c b/kernel/entry/common.c > index e4acd50bd81a..94318519998c 100644 > --- a/kernel/entry/common.c > +++ b/kernel/entry/common.c > @@ -6,6 +6,7 @@ > #include <linux/jump_label.h> > #include <linux/kmsan.h> > #include <linux/livepatch.h> > +#include <linux/rcupdate.h> > #include <linux/resume_user_mode.h> > #include <linux/tick.h> > > @@ -141,8 +142,13 @@ void raw_irqentry_exit_cond_resched(struct pt_regs *regs) > rcu_irq_exit_check_preempt(); > if (IS_ENABLED(CONFIG_DEBUG_ENTRY)) > WARN_ON_ONCE(!on_thread_stack()); > - if (need_resched() && arch_irqentry_exit_need_resched()) > + if (need_resched() && arch_irqentry_exit_need_resched()) { > + if (IS_ENABLED(CONFIG_TASKS_RCU_TRAMPOLINE_READERS)) > + rcu_tasks_irq_resched_enter(instruction_pointer(regs)); > preempt_schedule_irq(); > + if (IS_ENABLED(CONFIG_TASKS_RCU_TRAMPOLINE_READERS)) > + rcu_tasks_irq_resched_exit(); > + } > } > } > #ifdef CONFIG_PREEMPT_DYNAMIC > diff --git a/kernel/fork.c b/kernel/fork.c > index 416758c8a3d4..8077336bb136 100644 > --- a/kernel/fork.c > +++ b/kernel/fork.c > @@ -1871,6 +1871,7 @@ static inline void rcu_copy_process(struct task_struct *p) > p->rcu_tasks_holdout = false; > INIT_LIST_HEAD(&p->rcu_tasks_holdout_list); > p->rcu_tasks_idle_cpu = -1; > + p->rcu_tasks_irq_ip = 0; > INIT_LIST_HEAD(&p->rcu_tasks_exit_list); > #endif /* #ifdef CONFIG_TASKS_RCU */ > #ifdef CONFIG_TASKS_TRACE_RCU > diff --git a/kernel/rcu/Kconfig b/kernel/rcu/Kconfig > index 332df7a7a634..bbab14bc14c3 100644 > --- a/kernel/rcu/Kconfig > +++ b/kernel/rcu/Kconfig > @@ -107,6 +107,28 @@ config TASKS_RCU > default NEED_TASKS_RCU && PREEMPTION > select IRQ_WORK > > +config HAVE_RCU_TRAMPOLINE_READERS > + bool > + help > + Select this if the architecture uses the generic irq entry code and > + every trampoline whose lifetime Tasks RCU guards on it (ftrace > + trampolines, kprobe out-of-line and optimized-probe slots, BPF > + trampolines, out-of-line ftrace direct-call trampolines) enters a > + Tasks Trace RCU read-side critical section before calling out of > + the trampoline and leaves it before returning, and any core text > + that runs on behalf of such a trampoline outside that reader is > + reported by arch_rcu_tasks_trampoline_text(). The assembly readers > + use the this_cpu_inc() form of SRCU-fast, hence !NEED_SRCU_NMI_SAFE. > + > +config TASKS_RCU_TRAMPOLINE_READERS > + def_bool TASKS_RCU && HAVE_RCU_TRAMPOLINE_READERS && GENERIC_IRQ_ENTRY && !NEED_SRCU_NMI_SAFE > + select TASKS_TRACE_RCU > + help > + Implement the Tasks RCU grace period as a per-CPU pass over > + context switches and irq-exit reschedules outside trampoline text > + plus a Tasks Trace RCU grace period, instead of waiting for every > + task to voluntarily context switch. See kernel/rcu/tasks.h. > + > config FORCE_TASKS_RUDE_RCU > bool "Force selection of Tasks Rude RCU" > depends on RCU_EXPERT > diff --git a/kernel/rcu/tasks.h b/kernel/rcu/tasks.h > index 627295396cd9..3a7c092361a6 100644 > --- a/kernel/rcu/tasks.h > +++ b/kernel/rcu/tasks.h > @@ -152,7 +152,7 @@ static struct rcu_tasks rt_name = \ > .kname = #rt_name, \ > } > > -#ifdef CONFIG_TASKS_RCU > +#if defined(CONFIG_TASKS_RCU) && !defined(CONFIG_TASKS_RCU_TRAMPOLINE_READERS) > > /* Report delay of scan exiting tasklist in rcu_tasks_postscan(). */ > static void tasks_rcu_exit_stall(struct timer_list *unused); > @@ -802,7 +802,7 @@ static void rcu_tasks_torture_stats_print_generic(struct rcu_tasks *rtp, char *t > > #endif // #ifndef CONFIG_TINY_RCU > > -#if defined(CONFIG_TASKS_RCU) > +#if defined(CONFIG_TASKS_RCU) && !defined(CONFIG_TASKS_RCU_TRAMPOLINE_READERS) > > //////////////////////////////////////////////////////////////////////// > // > @@ -897,10 +897,445 @@ static void rcu_tasks_wait_gp(struct rcu_tasks *rtp) > rtp->postgp_func(rtp); > } > > -#endif /* #if defined(CONFIG_TASKS_RCU) */ > +#endif /* #if defined(CONFIG_TASKS_RCU) && !defined(CONFIG_TASKS_RCU_TRAMPOLINE_READERS) */ > > #ifdef CONFIG_TASKS_RCU > > +static int rcu_tasks_lazy_ms = -1; > +module_param(rcu_tasks_lazy_ms, int, 0444); > + > +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS > + > +//////////////////////////////////////////////////////////////////////// > +// > +// Tasks RCU for architectures whose trampolines are Tasks Trace RCU > +// readers (CONFIG_HAVE_RCU_TRAMPOLINE_READERS). > +// > +// On these architectures every piece of text whose lifetime Tasks RCU > +// guards -- ftrace trampolines, kprobe optinsn slots, BPF trampoline > +// images, out-of-line ftrace direct-call trampolines -- enters a Tasks > +// Trace RCU read-side critical section before calling out of itself and > +// leaves it before returning, so a task anywhere inside such a call-out, > +// preempted or not, is an ordinary rcu_read_lock_trace() reader and > +// synchronize_rcu_tasks_trace() waits for it. > +// > +// What that cannot cover is the handful of instructions in the trampoline > +// before the reader is entered and after it is left, and the one user that > +// has no trampoline at all: the bytes after a kprobe that the jump > +// optimizer is about to overwrite. A task can only linger in such > +// "unmarked" text by being interrupted there; unmarked text never calls > +// anything that could schedule. So a context switch on a CPU tells us that > +// whatever that CPU was running is out of unmarked text, with one > +// exception: a preemption from the irq-exit path, which can happen at any > +// instruction boundary. That path has the interrupted pt_regs in hand, so > +// just before it preempts it records the IP in the task and checks it > +// (rcu_tasks_trampoline_text()); if it is inside unmarked text the task > +// goes on a short holdout list first, and takes itself off again at its > +// next context switch outside such a preemption or its next irq-exit > +// check that finds it elsewhere. With that, every pass through > +// __schedule() is a per-CPU quiescent event, as are usermode and idle. > +// > +// A grace period is then: > +// > +// 1. Wait for every online, non-idle CPU to context switch, nudging > +// stragglers with resched_cpu(). Afterwards no task is in the leading > +// unmarked instructions of a dying trampoline unless it is on the > +// holdout list. > +// 2. Wait for the holdout list (as it stood) to drain. > +// 3. synchronize_rcu_tasks_trace(), for everything inside the readers. > +// 4. Repeat 1 and 2 for tasks that have since left the reader and are in > +// the trailing unmarked instructions. Alternatively the approach could be generalized to vanilla RCU, it could be possible to define a .text.rcu_no_qs section within which code running is considered as an RCU reader (with a pause while on the explicit RCU tasks section). It would be forbidden to voluntary sleep inside and to put explicit preemption points (CONFIG_PROVE_RCU could report misuses). Based on IP, RCU could consider those interrupted section as readers. This would require PREEMPT_RCU though. And then synchronize_rcu() would do the 1, 2, 4 jobs. -- Frederic Weisbecker SUSE Labs ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH RFC v3 03/13] rcu-tasks: Add a Tasks RCU implementation for reader-marked trampolines 2026-09-15 15:14 ` [PATCH RFC v3 03/13] rcu-tasks: Add a Tasks RCU implementation for reader-marked trampolines Frederic Weisbecker @ 2026-09-15 23:56 ` Paul E. McKenney 2026-09-16 12:40 ` Frederic Weisbecker 0 siblings, 1 reply; 20+ messages in thread From: Paul E. McKenney @ 2026-09-15 23:56 UTC (permalink / raw) To: Frederic Weisbecker Cc: Josef Bacik, Neeraj Upadhyay, Joel Fernandes, Boqun Feng, Thomas Gleixner, Peter Zijlstra, Steven Rostedt, Masami Hiramatsu, Mark Rutland, Jiri Olsa, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, x86, Catalin Marinas, Will Deacon, Puranjay Mohan, Xu Kuohai, Andy Lutomirski, Josh Triplett, Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang, Juergen Gross, Luis Chamberlain, Ihor Solodrai, linux-kernel, rcu, linux-trace-kernel, bpf, linux-arm-kernel, xen-devel On Tue, Sep 15, 2026 at 05:14:12PM +0200, Frederic Weisbecker wrote: > Le Tue, Sep 15, 2026 at 01:17:30PM +0000, Josef Bacik a écrit : > > Tasks RCU waits for every task to pass through a voluntary context > > switch, usermode or idle, because a preempted task might be sitting in a > > trampoline that is about to be freed and nothing marks it as such. With > > PREEMPT_LAZY that is a poor fit for servers: cond_resched() is a no-op, > > so a CPU-bound kthread only ever leaves the CPU by preemption, and one > > such kthread holds every synchronize_rcu_tasks() caller -- ftrace and > > BPF trampoline teardown under their mutexes, the kprobe jump optimizer > > under text_mutex and cpus_read_lock() -- hostage for as long as it runs. > > > > Following the discussion on v2, take the other road: let the > > architecture make its trampolines Tasks Trace RCU readers. When an > > architecture selects HAVE_RCU_TRAMPOLINE_READERS it promises that every > > trampoline whose lifetime Tasks RCU guards enters rcu_read_lock_trace() > > (or its assembly equivalent) before calling out and leaves it before > > returning, so a task anywhere inside such a call-out, preempted or not, > > is an ordinary Tasks Trace reader. > > > > That leaves the few instructions of trampoline text before the reader > > is entered and after it is left (plus, in a later patch, the bytes a > > kprobe jump optimization is about to overwrite). A task can only linger > > there by being interrupted there, and such text never calls anything > > that schedules, so instead of tracking tasks we track CPUs: every pass > > through __schedule() is a per-CPU quiescent event, except that the one > > context switch that can catch a task at an arbitrary instruction -- a > > preemption from irq exit -- first records the interrupted IP in the task > > and parks it on a per-CPU list for the duration (reusing the fields and > > lists the classic flavor keeps for its exit-path bookkeeping), and, if > > the IP is inside such "unmarked" text, puts the task on a short holdout > > list; the task takes itself off at its next context switch outside such > > a preemption or irq-exit check that finds it elsewhere. Usermode (the > > existing tick hook, or a nohz_full CPU in an RCU extended quiescent > > state) and idle count as well. rcu_tasks_trampoline_text() does the > > classification: anything outside core and module text, a new > > .text..rcu_tramp section for C glue that trampolines call before it has > > entered the reader (__rcu_trampoline), and an arch hook for things like > > static ftrace stubs and return thunks. > > > > The grace period, run by the existing rcu_tasks kthread so that > > call_rcu_tasks(), synchronize_rcu_tasks() and rcu_barrier_tasks() keep > > their names and callers, is: wait for every online non-idle CPU to > > context switch (nudging stragglers with resched_cpu() after a jiffy), > > drain the holdout list as it stood, synchronize_rcu_tasks_trace() for > > everything inside the readers, then one more CPU pass and drain for > > tasks that have since left the reader into the trailing instructions. > > That is bounded by a few jiffies, preempt-off latency and an SRCU grace > > period rather than by the longest stretch any task runs without > > sleeping, needs no per-task scan, and makes cond_resched_tasks_rcu_qs() > > unnecessary on such architectures. As before, idle tasks are not > > waited for. rcu_tasks_wait_irq_preempted() walks the parked lists for > > the one caller (the kprobe jump optimizer, later in the series) that > > makes ordinary text unsafe to be parked in and so has to wait out tasks > > that were preempted there before it said so. > > > > The classic implementation is untouched and remains the default; the > > new one is built only as CONFIG_TASKS_RCU_TRAMPOLINE_READERS when the > > architecture opts in and uses the generic irq entry code, whose > > reschedule check gains the rcu_tasks_irq_resched() call. Nothing > > selects it yet. > > > > Suggested-by: Paul E. McKenney <paulmck@kernel.org> > > Suggested-by: Alexei Starovoitov <ast@kernel.org> > > Assisted-by: LLM > > Signed-off-by: Josef Bacik <josef@toxicpanda.com> > > --- > > include/asm-generic/vmlinux.lds.h | 11 + > > include/linux/rcupdate.h | 32 ++- > > include/linux/sched.h | 1 + > > kernel/entry/common.c | 8 +- > > kernel/fork.c | 1 + > > kernel/rcu/Kconfig | 22 ++ > > kernel/rcu/tasks.h | 460 +++++++++++++++++++++++++++++++++++++- > > kernel/rcu/update.c | 2 + > > 8 files changed, 528 insertions(+), 9 deletions(-) > > > > diff --git a/include/asm-generic/vmlinux.lds.h b/include/asm-generic/vmlinux.lds.h > > index b2988aa12f66..86e58c4fe370 100644 > > --- a/include/asm-generic/vmlinux.lds.h > > +++ b/include/asm-generic/vmlinux.lds.h > > @@ -571,6 +571,16 @@ > > __cpuidle_text_end = .; \ > > __noinstr_text_end = .; > > > > +/* > > + * C glue called directly from Tasks-RCU-protected trampolines, bounded so > > + * that rcu_tasks_trampoline_text() can recognise it; see __rcu_trampoline. > > + */ > > +#define RCU_TRAMP_TEXT \ > > + ALIGN_FUNCTION(); \ > > + __rcu_tramp_text_start = .; \ > > + *(.text..rcu_tramp) \ > > + __rcu_tramp_text_end = .; > > + > > #define TEXT_SPLIT \ > > __split_text_start = .; \ > > *(.text.split .text.split.[0-9a-zA-Z_]*) \ > > @@ -607,6 +617,7 @@ > > TEXT_HOT \ > > *(TEXT_MAIN .text.fixup) \ > > NOINSTR_TEXT \ > > + RCU_TRAMP_TEXT \ > > *(.ref.text) > > > > /* sched.text is aling to function alignment to secure we have same > > diff --git a/include/linux/rcupdate.h b/include/linux/rcupdate.h > > index 44c07a66edff..fb2a3889a696 100644 > > --- a/include/linux/rcupdate.h > > +++ b/include/linux/rcupdate.h > > @@ -50,6 +50,31 @@ token_context_lock_instance(RCU, RCU_BH); > > /* Exported common interfaces */ > > void call_rcu(struct rcu_head *head, rcu_callback_t func); > > void rcu_barrier_tasks(void); > > + > > +/* > > + * Trampoline-reader Tasks RCU (CONFIG_TASKS_RCU_TRAMPOLINE_READERS), see > > + * kernel/rcu/tasks.h. rcu_tasks_irq_resched_enter()/_exit() bracket the > > + * irq-exit preemption; rcu_tasks_trampoline_text() and the arch_ override > > + * classify an interrupted IP; rcu_tasks_wait_irq_preempted() lets a caller > > + * wait out tasks already preempted somewhere it is about to make unsafe. > > + * __rcu_trampoline places C code that such trampolines call directly, before > > + * it has entered its Tasks Trace reader, where that classification can see it. > > + */ > > +void rcu_tasks_irq_resched_enter(unsigned long ip); > > +void rcu_tasks_irq_resched_exit(void); > > +bool rcu_tasks_trampoline_text(unsigned long ip); > > +bool arch_rcu_tasks_trampoline_text(unsigned long ip); > > +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS > > +void rcu_tasks_wait_irq_preempted(bool (*inside)(unsigned long ip)); > > +#else > > +static inline void rcu_tasks_wait_irq_preempted(bool (*inside)(unsigned long ip)) { } > > +#endif > > +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS > > +/* Also keeps instrumentation calls out of the prologue, ahead of the reader. */ > > +#define __rcu_trampoline __noinstr_section(".text..rcu_tramp") > > +#else > > +#define __rcu_trampoline > > +#endif > > void synchronize_rcu(void); > > > > /* > > @@ -180,11 +205,16 @@ static inline void rcu_nocb_flush_deferred_wakeup(void) { } > > #ifdef CONFIG_TASKS_RCU_GENERIC > > > > # ifdef CONFIG_TASKS_RCU > > -# define rcu_tasks_classic_qs(t, preempt) \ > > +# ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS > > +void rcu_tasks_note_qs(struct task_struct *t, bool preempt); > > +# define rcu_tasks_classic_qs(t, preempt) rcu_tasks_note_qs((t), (preempt)) > > +# else > > +# define rcu_tasks_classic_qs(t, preempt) \ > > do { \ > > if (!(preempt) && READ_ONCE((t)->rcu_tasks_holdout)) \ > > WRITE_ONCE((t)->rcu_tasks_holdout, false); \ > > } while (0) > > +# endif > > void call_rcu_tasks(struct rcu_head *head, rcu_callback_t func); > > void synchronize_rcu_tasks(void); > > void rcu_tasks_torture_stats_print(char *tt, char *tf); > > diff --git a/include/linux/sched.h b/include/linux/sched.h > > index 8b3d47a325cc..15beb44caa2c 100644 > > --- a/include/linux/sched.h > > +++ b/include/linux/sched.h > > @@ -957,6 +957,7 @@ struct task_struct { > > u8 rcu_tasks_holdout; > > u8 rcu_tasks_idx; > > int rcu_tasks_idle_cpu; > > + unsigned long rcu_tasks_irq_ip; > > struct list_head rcu_tasks_holdout_list; > > int rcu_tasks_exit_cpu; > > struct list_head rcu_tasks_exit_list; > > diff --git a/kernel/entry/common.c b/kernel/entry/common.c > > index e4acd50bd81a..94318519998c 100644 > > --- a/kernel/entry/common.c > > +++ b/kernel/entry/common.c > > @@ -6,6 +6,7 @@ > > #include <linux/jump_label.h> > > #include <linux/kmsan.h> > > #include <linux/livepatch.h> > > +#include <linux/rcupdate.h> > > #include <linux/resume_user_mode.h> > > #include <linux/tick.h> > > > > @@ -141,8 +142,13 @@ void raw_irqentry_exit_cond_resched(struct pt_regs *regs) > > rcu_irq_exit_check_preempt(); > > if (IS_ENABLED(CONFIG_DEBUG_ENTRY)) > > WARN_ON_ONCE(!on_thread_stack()); > > - if (need_resched() && arch_irqentry_exit_need_resched()) > > + if (need_resched() && arch_irqentry_exit_need_resched()) { > > + if (IS_ENABLED(CONFIG_TASKS_RCU_TRAMPOLINE_READERS)) > > + rcu_tasks_irq_resched_enter(instruction_pointer(regs)); > > preempt_schedule_irq(); > > + if (IS_ENABLED(CONFIG_TASKS_RCU_TRAMPOLINE_READERS)) > > + rcu_tasks_irq_resched_exit(); > > + } > > } > > } > > #ifdef CONFIG_PREEMPT_DYNAMIC > > diff --git a/kernel/fork.c b/kernel/fork.c > > index 416758c8a3d4..8077336bb136 100644 > > --- a/kernel/fork.c > > +++ b/kernel/fork.c > > @@ -1871,6 +1871,7 @@ static inline void rcu_copy_process(struct task_struct *p) > > p->rcu_tasks_holdout = false; > > INIT_LIST_HEAD(&p->rcu_tasks_holdout_list); > > p->rcu_tasks_idle_cpu = -1; > > + p->rcu_tasks_irq_ip = 0; > > INIT_LIST_HEAD(&p->rcu_tasks_exit_list); > > #endif /* #ifdef CONFIG_TASKS_RCU */ > > #ifdef CONFIG_TASKS_TRACE_RCU > > diff --git a/kernel/rcu/Kconfig b/kernel/rcu/Kconfig > > index 332df7a7a634..bbab14bc14c3 100644 > > --- a/kernel/rcu/Kconfig > > +++ b/kernel/rcu/Kconfig > > @@ -107,6 +107,28 @@ config TASKS_RCU > > default NEED_TASKS_RCU && PREEMPTION > > select IRQ_WORK > > > > +config HAVE_RCU_TRAMPOLINE_READERS > > + bool > > + help > > + Select this if the architecture uses the generic irq entry code and > > + every trampoline whose lifetime Tasks RCU guards on it (ftrace > > + trampolines, kprobe out-of-line and optimized-probe slots, BPF > > + trampolines, out-of-line ftrace direct-call trampolines) enters a > > + Tasks Trace RCU read-side critical section before calling out of > > + the trampoline and leaves it before returning, and any core text > > + that runs on behalf of such a trampoline outside that reader is > > + reported by arch_rcu_tasks_trampoline_text(). The assembly readers > > + use the this_cpu_inc() form of SRCU-fast, hence !NEED_SRCU_NMI_SAFE. > > + > > +config TASKS_RCU_TRAMPOLINE_READERS > > + def_bool TASKS_RCU && HAVE_RCU_TRAMPOLINE_READERS && GENERIC_IRQ_ENTRY && !NEED_SRCU_NMI_SAFE > > + select TASKS_TRACE_RCU > > + help > > + Implement the Tasks RCU grace period as a per-CPU pass over > > + context switches and irq-exit reschedules outside trampoline text > > + plus a Tasks Trace RCU grace period, instead of waiting for every > > + task to voluntarily context switch. See kernel/rcu/tasks.h. > > + > > config FORCE_TASKS_RUDE_RCU > > bool "Force selection of Tasks Rude RCU" > > depends on RCU_EXPERT > > diff --git a/kernel/rcu/tasks.h b/kernel/rcu/tasks.h > > index 627295396cd9..3a7c092361a6 100644 > > --- a/kernel/rcu/tasks.h > > +++ b/kernel/rcu/tasks.h > > @@ -152,7 +152,7 @@ static struct rcu_tasks rt_name = \ > > .kname = #rt_name, \ > > } > > > > -#ifdef CONFIG_TASKS_RCU > > +#if defined(CONFIG_TASKS_RCU) && !defined(CONFIG_TASKS_RCU_TRAMPOLINE_READERS) > > > > /* Report delay of scan exiting tasklist in rcu_tasks_postscan(). */ > > static void tasks_rcu_exit_stall(struct timer_list *unused); > > @@ -802,7 +802,7 @@ static void rcu_tasks_torture_stats_print_generic(struct rcu_tasks *rtp, char *t > > > > #endif // #ifndef CONFIG_TINY_RCU > > > > -#if defined(CONFIG_TASKS_RCU) > > +#if defined(CONFIG_TASKS_RCU) && !defined(CONFIG_TASKS_RCU_TRAMPOLINE_READERS) > > > > //////////////////////////////////////////////////////////////////////// > > // > > @@ -897,10 +897,445 @@ static void rcu_tasks_wait_gp(struct rcu_tasks *rtp) > > rtp->postgp_func(rtp); > > } > > > > -#endif /* #if defined(CONFIG_TASKS_RCU) */ > > +#endif /* #if defined(CONFIG_TASKS_RCU) && !defined(CONFIG_TASKS_RCU_TRAMPOLINE_READERS) */ > > > > #ifdef CONFIG_TASKS_RCU > > > > +static int rcu_tasks_lazy_ms = -1; > > +module_param(rcu_tasks_lazy_ms, int, 0444); > > + > > +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS > > + > > +//////////////////////////////////////////////////////////////////////// > > +// > > +// Tasks RCU for architectures whose trampolines are Tasks Trace RCU > > +// readers (CONFIG_HAVE_RCU_TRAMPOLINE_READERS). > > +// > > +// On these architectures every piece of text whose lifetime Tasks RCU > > +// guards -- ftrace trampolines, kprobe optinsn slots, BPF trampoline > > +// images, out-of-line ftrace direct-call trampolines -- enters a Tasks > > +// Trace RCU read-side critical section before calling out of itself and > > +// leaves it before returning, so a task anywhere inside such a call-out, > > +// preempted or not, is an ordinary rcu_read_lock_trace() reader and > > +// synchronize_rcu_tasks_trace() waits for it. > > +// > > +// What that cannot cover is the handful of instructions in the trampoline > > +// before the reader is entered and after it is left, and the one user that > > +// has no trampoline at all: the bytes after a kprobe that the jump > > +// optimizer is about to overwrite. A task can only linger in such > > +// "unmarked" text by being interrupted there; unmarked text never calls > > +// anything that could schedule. So a context switch on a CPU tells us that > > +// whatever that CPU was running is out of unmarked text, with one > > +// exception: a preemption from the irq-exit path, which can happen at any > > +// instruction boundary. That path has the interrupted pt_regs in hand, so > > +// just before it preempts it records the IP in the task and checks it > > +// (rcu_tasks_trampoline_text()); if it is inside unmarked text the task > > +// goes on a short holdout list first, and takes itself off again at its > > +// next context switch outside such a preemption or its next irq-exit > > +// check that finds it elsewhere. With that, every pass through > > +// __schedule() is a per-CPU quiescent event, as are usermode and idle. > > +// > > +// A grace period is then: > > +// > > +// 1. Wait for every online, non-idle CPU to context switch, nudging > > +// stragglers with resched_cpu(). Afterwards no task is in the leading > > +// unmarked instructions of a dying trampoline unless it is on the > > +// holdout list. > > +// 2. Wait for the holdout list (as it stood) to drain. > > +// 3. synchronize_rcu_tasks_trace(), for everything inside the readers. > > +// 4. Repeat 1 and 2 for tasks that have since left the reader and are in > > +// the trailing unmarked instructions. > > Alternatively the approach could be generalized to vanilla RCU, it could be > possible to define a .text.rcu_no_qs section within which code running is > considered as an RCU reader (with a pause while on the explicit RCU tasks > section). It would be forbidden to voluntary sleep inside > and to put explicit preemption points (CONFIG_PROVE_RCU could report misuses). > > Based on IP, RCU could consider those interrupted section as readers. This would > require PREEMPT_RCU though. > > And then synchronize_rcu() would do the 1, 2, 4 jobs. If I am following correctly (ha!), sleepable BPF programs rule out use of RCU in this manner. But your point is nevertheless valid, in that SRCU could be used. And because rcu_read_lock_trace() is a thin wrapper around SRCU-fast, we *might* be able to instead use rcu_read_lock_tasks_trace(), which would skip the task-struct increment and decrement, saving a few instructions. Then, instead of waiting for each task's counter to go to zero, instead just invoke synchronize_rcu_tasks_trace(). Which is pretty close to what Josef is proposing, just with the new RCU Tasks Trace read-side primitives. I think. ;-) This assumes that we do not need to flatten partially overlapping RCU Tasks Trace readers into one big reader. Or am I missing something here? Thanx, Paul ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH RFC v3 03/13] rcu-tasks: Add a Tasks RCU implementation for reader-marked trampolines 2026-09-15 23:56 ` Paul E. McKenney @ 2026-09-16 12:40 ` Frederic Weisbecker 2026-09-16 14:26 ` Paul E. McKenney 0 siblings, 1 reply; 20+ messages in thread From: Frederic Weisbecker @ 2026-09-16 12:40 UTC (permalink / raw) To: Paul E. McKenney Cc: Josef Bacik, Neeraj Upadhyay, Joel Fernandes, Boqun Feng, Thomas Gleixner, Peter Zijlstra, Steven Rostedt, Masami Hiramatsu, Mark Rutland, Jiri Olsa, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, x86, Catalin Marinas, Will Deacon, Puranjay Mohan, Xu Kuohai, Andy Lutomirski, Josh Triplett, Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang, Juergen Gross, Luis Chamberlain, Ihor Solodrai, linux-kernel, rcu, linux-trace-kernel, bpf, linux-arm-kernel, xen-devel Le Tue, Sep 15, 2026 at 04:56:41PM -0700, Paul E. McKenney a écrit : > > Alternatively the approach could be generalized to vanilla RCU, it could be > > possible to define a .text.rcu_no_qs section within which code running is > > considered as an RCU reader (with a pause while on the explicit RCU tasks > > section). It would be forbidden to voluntary sleep inside > > and to put explicit preemption points (CONFIG_PROVE_RCU could report misuses). > > > > Based on IP, RCU could consider those interrupted section as readers. This would > > require PREEMPT_RCU though. > > > > And then synchronize_rcu() would do the 1, 2, 4 jobs. > > If I am following correctly (ha!), sleepable BPF programs rule out use > of RCU in this manner. > > But your point is nevertheless valid, in that SRCU could be used. > And because rcu_read_lock_trace() is a thin wrapper around SRCU-fast, we > *might* be able to instead use rcu_read_lock_tasks_trace(), which would > skip the task-struct increment and decrement, saving a few instructions. > Then, instead of waiting for each task's counter to go to zero, instead > just invoke synchronize_rcu_tasks_trace(). > > Which is pretty close to what Josef is proposing, just with the new RCU > Tasks Trace read-side primitives. I think. ;-) > > This assumes that we do not need to flatten partially overlapping RCU > Tasks Trace readers into one big reader. > > Or am I missing something here? Yes I think that's what Josef does in this patchset. The problem is about handling the few instructions: 1) between the begining of the trampoline and the call to rcu_read_lock_trace() 2) between the call to rcu_read_unlock_trace() and the end of the trampoline So what I'm proposing is to make those two parts implicit RCU read lock sections. So the whole trampoline would be .text.rcu_no_qs: .text.rcu_no_qs trampoline: __________________________________________________________________________________________ |Few instructions 1 | rcu_read_lock_trace() .... rcu_read_unlock_trace | Few instructions 2| ___________________________________________________________________________________________ Then when a tick fires, rcu_flavor_sched_clock_irq() discards the interrupted code as QS if the IP was within .text.rcu_no_qs _unless_ it is in the rcu_read_lock_trace. Both are easy and quick to verify. Also preempt_schedule_irq() would make sure to verify the same condition and enqueue the task as a GP blocker if preempting inside "Few instructions 1" or "Few instructions 2". And since RCU tasks already does a synchronize RCU before and after the scan, that's all we would have to do. Thanks. -- Frederic Weisbecker SUSE Labs ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH RFC v3 03/13] rcu-tasks: Add a Tasks RCU implementation for reader-marked trampolines 2026-09-16 12:40 ` Frederic Weisbecker @ 2026-09-16 14:26 ` Paul E. McKenney 2026-09-16 14:35 ` Frederic Weisbecker 0 siblings, 1 reply; 20+ messages in thread From: Paul E. McKenney @ 2026-09-16 14:26 UTC (permalink / raw) To: Frederic Weisbecker Cc: Josef Bacik, Neeraj Upadhyay, Joel Fernandes, Boqun Feng, Thomas Gleixner, Peter Zijlstra, Steven Rostedt, Masami Hiramatsu, Mark Rutland, Jiri Olsa, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, x86, Catalin Marinas, Will Deacon, Puranjay Mohan, Xu Kuohai, Andy Lutomirski, Josh Triplett, Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang, Juergen Gross, Luis Chamberlain, Ihor Solodrai, linux-kernel, rcu, linux-trace-kernel, bpf, linux-arm-kernel, xen-devel On Wed, Sep 16, 2026 at 02:40:20PM +0200, Frederic Weisbecker wrote: > Le Tue, Sep 15, 2026 at 04:56:41PM -0700, Paul E. McKenney a écrit : > > > Alternatively the approach could be generalized to vanilla RCU, it could be > > > possible to define a .text.rcu_no_qs section within which code running is > > > considered as an RCU reader (with a pause while on the explicit RCU tasks > > > section). It would be forbidden to voluntary sleep inside > > > and to put explicit preemption points (CONFIG_PROVE_RCU could report misuses). > > > > > > Based on IP, RCU could consider those interrupted section as readers. This would > > > require PREEMPT_RCU though. > > > > > > And then synchronize_rcu() would do the 1, 2, 4 jobs. > > > > If I am following correctly (ha!), sleepable BPF programs rule out use > > of RCU in this manner. > > > > But your point is nevertheless valid, in that SRCU could be used. > > And because rcu_read_lock_trace() is a thin wrapper around SRCU-fast, we > > *might* be able to instead use rcu_read_lock_tasks_trace(), which would > > skip the task-struct increment and decrement, saving a few instructions. > > Then, instead of waiting for each task's counter to go to zero, instead > > just invoke synchronize_rcu_tasks_trace(). > > > > Which is pretty close to what Josef is proposing, just with the new RCU > > Tasks Trace read-side primitives. I think. ;-) > > > > This assumes that we do not need to flatten partially overlapping RCU > > Tasks Trace readers into one big reader. > > > > Or am I missing something here? > > Yes I think that's what Josef does in this patchset. The problem is about > handling the few instructions: > > 1) between the begining of the trampoline and the call to rcu_read_lock_trace() > > 2) between the call to rcu_read_unlock_trace() and the end of the trampoline > > So what I'm proposing is to make those two parts implicit RCU read lock sections. > > So the whole trampoline would be .text.rcu_no_qs: > > .text.rcu_no_qs trampoline: > __________________________________________________________________________________________ > |Few instructions 1 | rcu_read_lock_trace() .... rcu_read_unlock_trace | Few instructions 2| > ___________________________________________________________________________________________ > > Then when a tick fires, rcu_flavor_sched_clock_irq() discards the interrupted > code as QS if the IP was within .text.rcu_no_qs _unless_ it is in the > rcu_read_lock_trace. Both are easy and quick to verify. > > Also preempt_schedule_irq() would make sure to verify the same condition and > enqueue the task as a GP blocker if preempting inside "Few instructions 1" > or "Few instructions 2". Ah, OK, I might be following now. ;-) We also need both versions of rcu_exp_handler() to check the IP as well, given that sooner or later someone is going to want trampoline removal to go faster. Or am I still missing a turn in here somewhere? > And since RCU tasks already does a synchronize RCU before and after the scan, > that's all we would have to do. This is going to need some *serious* documentation. Also, what would be a good way to add tests for this to rcutorture? Designate some new rcutorture function as being in .text.rcu_no_qs and add this as another type of rcutorture reader? Or is there a better way? Thanx, Paul ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH RFC v3 03/13] rcu-tasks: Add a Tasks RCU implementation for reader-marked trampolines 2026-09-16 14:26 ` Paul E. McKenney @ 2026-09-16 14:35 ` Frederic Weisbecker 2026-09-16 14:47 ` Frederic Weisbecker 0 siblings, 1 reply; 20+ messages in thread From: Frederic Weisbecker @ 2026-09-16 14:35 UTC (permalink / raw) To: Paul E. McKenney Cc: Josef Bacik, Neeraj Upadhyay, Joel Fernandes, Boqun Feng, Thomas Gleixner, Peter Zijlstra, Steven Rostedt, Masami Hiramatsu, Mark Rutland, Jiri Olsa, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, x86, Catalin Marinas, Will Deacon, Puranjay Mohan, Xu Kuohai, Andy Lutomirski, Josh Triplett, Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang, Juergen Gross, Luis Chamberlain, Ihor Solodrai, linux-kernel, rcu, linux-trace-kernel, bpf, linux-arm-kernel, xen-devel Le Wed, Sep 16, 2026 at 07:26:35AM -0700, Paul E. McKenney a écrit : > On Wed, Sep 16, 2026 at 02:40:20PM +0200, Frederic Weisbecker wrote: > > Le Tue, Sep 15, 2026 at 04:56:41PM -0700, Paul E. McKenney a écrit : > > > > Alternatively the approach could be generalized to vanilla RCU, it could be > > > > possible to define a .text.rcu_no_qs section within which code running is > > > > considered as an RCU reader (with a pause while on the explicit RCU tasks > > > > section). It would be forbidden to voluntary sleep inside > > > > and to put explicit preemption points (CONFIG_PROVE_RCU could report misuses). > > > > > > > > Based on IP, RCU could consider those interrupted section as readers. This would > > > > require PREEMPT_RCU though. > > > > > > > > And then synchronize_rcu() would do the 1, 2, 4 jobs. > > > > > > If I am following correctly (ha!), sleepable BPF programs rule out use > > > of RCU in this manner. > > > > > > But your point is nevertheless valid, in that SRCU could be used. > > > And because rcu_read_lock_trace() is a thin wrapper around SRCU-fast, we > > > *might* be able to instead use rcu_read_lock_tasks_trace(), which would > > > skip the task-struct increment and decrement, saving a few instructions. > > > Then, instead of waiting for each task's counter to go to zero, instead > > > just invoke synchronize_rcu_tasks_trace(). > > > > > > Which is pretty close to what Josef is proposing, just with the new RCU > > > Tasks Trace read-side primitives. I think. ;-) > > > > > > This assumes that we do not need to flatten partially overlapping RCU > > > Tasks Trace readers into one big reader. > > > > > > Or am I missing something here? > > > > Yes I think that's what Josef does in this patchset. The problem is about > > handling the few instructions: > > > > 1) between the begining of the trampoline and the call to rcu_read_lock_trace() > > > > 2) between the call to rcu_read_unlock_trace() and the end of the trampoline > > > > So what I'm proposing is to make those two parts implicit RCU read lock sections. > > > > So the whole trampoline would be .text.rcu_no_qs: > > > > .text.rcu_no_qs trampoline: > > __________________________________________________________________________________________ > > |Few instructions 1 | rcu_read_lock_trace() .... rcu_read_unlock_trace | Few instructions 2| > > ___________________________________________________________________________________________ > > > > Then when a tick fires, rcu_flavor_sched_clock_irq() discards the interrupted > > code as QS if the IP was within .text.rcu_no_qs _unless_ it is in the > > rcu_read_lock_trace. Both are easy and quick to verify. > > > > Also preempt_schedule_irq() would make sure to verify the same condition and > > enqueue the task as a GP blocker if preempting inside "Few instructions 1" > > or "Few instructions 2". > > Ah, OK, I might be following now. ;-) > > We also need both versions of rcu_exp_handler() to check the IP as well, > given that sooner or later someone is going to want trampoline removal > to go faster. Or am I still missing a turn in here somewhere? Yes indeed, missed the exp part! > > > And since RCU tasks already does a synchronize RCU before and after the scan, > > that's all we would have to do. > > This is going to need some *serious* documentation. Yes :-) > Also, what would be a good way to add tests for this to rcutorture? > Designate some new rcutorture function as being in .text.rcu_no_qs and > add this as another type of rcutorture reader? Or is there a better way? Yes that sounds good! > > Thanx, Paul -- Frederic Weisbecker SUSE Labs ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH RFC v3 03/13] rcu-tasks: Add a Tasks RCU implementation for reader-marked trampolines 2026-09-16 14:35 ` Frederic Weisbecker @ 2026-09-16 14:47 ` Frederic Weisbecker 2026-09-16 14:55 ` Paul E. McKenney 0 siblings, 1 reply; 20+ messages in thread From: Frederic Weisbecker @ 2026-09-16 14:47 UTC (permalink / raw) To: Paul E. McKenney Cc: Josef Bacik, Neeraj Upadhyay, Joel Fernandes, Boqun Feng, Thomas Gleixner, Peter Zijlstra, Steven Rostedt, Masami Hiramatsu, Mark Rutland, Jiri Olsa, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, x86, Catalin Marinas, Will Deacon, Puranjay Mohan, Xu Kuohai, Andy Lutomirski, Josh Triplett, Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang, Juergen Gross, Luis Chamberlain, Ihor Solodrai, linux-kernel, rcu, linux-trace-kernel, bpf, linux-arm-kernel, xen-devel Le Wed, Sep 16, 2026 at 04:35:50PM +0200, Frederic Weisbecker a écrit : > Le Wed, Sep 16, 2026 at 07:26:35AM -0700, Paul E. McKenney a écrit : > > On Wed, Sep 16, 2026 at 02:40:20PM +0200, Frederic Weisbecker wrote: > > > Le Tue, Sep 15, 2026 at 04:56:41PM -0700, Paul E. McKenney a écrit : > > > > > Alternatively the approach could be generalized to vanilla RCU, it could be > > > > > possible to define a .text.rcu_no_qs section within which code running is > > > > > considered as an RCU reader (with a pause while on the explicit RCU tasks > > > > > section). It would be forbidden to voluntary sleep inside > > > > > and to put explicit preemption points (CONFIG_PROVE_RCU could report misuses). > > > > > > > > > > Based on IP, RCU could consider those interrupted section as readers. This would > > > > > require PREEMPT_RCU though. > > > > > > > > > > And then synchronize_rcu() would do the 1, 2, 4 jobs. > > > > > > > > If I am following correctly (ha!), sleepable BPF programs rule out use > > > > of RCU in this manner. > > > > > > > > But your point is nevertheless valid, in that SRCU could be used. > > > > And because rcu_read_lock_trace() is a thin wrapper around SRCU-fast, we > > > > *might* be able to instead use rcu_read_lock_tasks_trace(), which would > > > > skip the task-struct increment and decrement, saving a few instructions. > > > > Then, instead of waiting for each task's counter to go to zero, instead > > > > just invoke synchronize_rcu_tasks_trace(). > > > > > > > > Which is pretty close to what Josef is proposing, just with the new RCU > > > > Tasks Trace read-side primitives. I think. ;-) > > > > > > > > This assumes that we do not need to flatten partially overlapping RCU > > > > Tasks Trace readers into one big reader. > > > > > > > > Or am I missing something here? > > > > > > Yes I think that's what Josef does in this patchset. The problem is about > > > handling the few instructions: > > > > > > 1) between the begining of the trampoline and the call to rcu_read_lock_trace() > > > > > > 2) between the call to rcu_read_unlock_trace() and the end of the trampoline > > > > > > So what I'm proposing is to make those two parts implicit RCU read lock sections. > > > > > > So the whole trampoline would be .text.rcu_no_qs: > > > > > > .text.rcu_no_qs trampoline: > > > __________________________________________________________________________________________ > > > |Few instructions 1 | rcu_read_lock_trace() .... rcu_read_unlock_trace | Few instructions 2| > > > ___________________________________________________________________________________________ > > > > > > Then when a tick fires, rcu_flavor_sched_clock_irq() discards the interrupted > > > code as QS if the IP was within .text.rcu_no_qs _unless_ it is in the > > > rcu_read_lock_trace. Both are easy and quick to verify. > > > > > > Also preempt_schedule_irq() would make sure to verify the same condition and > > > enqueue the task as a GP blocker if preempting inside "Few instructions 1" > > > or "Few instructions 2". > > > > Ah, OK, I might be following now. ;-) > > > > We also need both versions of rcu_exp_handler() to check the IP as well, > > given that sooner or later someone is going to want trampoline removal > > to go faster. Or am I still missing a turn in here somewhere? > > Yes indeed, missed the exp part! What remains to handle also is non-preemptible RCU because if the task is preempted by an IRQ while in the .text.rcu_no_qs, we may still need to keep track of that somewhere. -- Frederic Weisbecker SUSE Labs ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH RFC v3 03/13] rcu-tasks: Add a Tasks RCU implementation for reader-marked trampolines 2026-09-16 14:47 ` Frederic Weisbecker @ 2026-09-16 14:55 ` Paul E. McKenney 2026-09-16 15:23 ` Frederic Weisbecker 0 siblings, 1 reply; 20+ messages in thread From: Paul E. McKenney @ 2026-09-16 14:55 UTC (permalink / raw) To: Frederic Weisbecker Cc: Josef Bacik, Neeraj Upadhyay, Joel Fernandes, Boqun Feng, Thomas Gleixner, Peter Zijlstra, Steven Rostedt, Masami Hiramatsu, Mark Rutland, Jiri Olsa, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, x86, Catalin Marinas, Will Deacon, Puranjay Mohan, Xu Kuohai, Andy Lutomirski, Josh Triplett, Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang, Juergen Gross, Luis Chamberlain, Ihor Solodrai, linux-kernel, rcu, linux-trace-kernel, bpf, linux-arm-kernel, xen-devel On Wed, Sep 16, 2026 at 04:47:29PM +0200, Frederic Weisbecker wrote: > Le Wed, Sep 16, 2026 at 04:35:50PM +0200, Frederic Weisbecker a écrit : > > Le Wed, Sep 16, 2026 at 07:26:35AM -0700, Paul E. McKenney a écrit : > > > On Wed, Sep 16, 2026 at 02:40:20PM +0200, Frederic Weisbecker wrote: > > > > Le Tue, Sep 15, 2026 at 04:56:41PM -0700, Paul E. McKenney a écrit : > > > > > > Alternatively the approach could be generalized to vanilla RCU, it could be > > > > > > possible to define a .text.rcu_no_qs section within which code running is > > > > > > considered as an RCU reader (with a pause while on the explicit RCU tasks > > > > > > section). It would be forbidden to voluntary sleep inside > > > > > > and to put explicit preemption points (CONFIG_PROVE_RCU could report misuses). > > > > > > > > > > > > Based on IP, RCU could consider those interrupted section as readers. This would > > > > > > require PREEMPT_RCU though. > > > > > > > > > > > > And then synchronize_rcu() would do the 1, 2, 4 jobs. > > > > > > > > > > If I am following correctly (ha!), sleepable BPF programs rule out use > > > > > of RCU in this manner. > > > > > > > > > > But your point is nevertheless valid, in that SRCU could be used. > > > > > And because rcu_read_lock_trace() is a thin wrapper around SRCU-fast, we > > > > > *might* be able to instead use rcu_read_lock_tasks_trace(), which would > > > > > skip the task-struct increment and decrement, saving a few instructions. > > > > > Then, instead of waiting for each task's counter to go to zero, instead > > > > > just invoke synchronize_rcu_tasks_trace(). > > > > > > > > > > Which is pretty close to what Josef is proposing, just with the new RCU > > > > > Tasks Trace read-side primitives. I think. ;-) > > > > > > > > > > This assumes that we do not need to flatten partially overlapping RCU > > > > > Tasks Trace readers into one big reader. > > > > > > > > > > Or am I missing something here? > > > > > > > > Yes I think that's what Josef does in this patchset. The problem is about > > > > handling the few instructions: > > > > > > > > 1) between the begining of the trampoline and the call to rcu_read_lock_trace() > > > > > > > > 2) between the call to rcu_read_unlock_trace() and the end of the trampoline > > > > > > > > So what I'm proposing is to make those two parts implicit RCU read lock sections. > > > > > > > > So the whole trampoline would be .text.rcu_no_qs: > > > > > > > > .text.rcu_no_qs trampoline: > > > > __________________________________________________________________________________________ > > > > |Few instructions 1 | rcu_read_lock_trace() .... rcu_read_unlock_trace | Few instructions 2| > > > > ___________________________________________________________________________________________ > > > > > > > > Then when a tick fires, rcu_flavor_sched_clock_irq() discards the interrupted > > > > code as QS if the IP was within .text.rcu_no_qs _unless_ it is in the > > > > rcu_read_lock_trace. Both are easy and quick to verify. > > > > > > > > Also preempt_schedule_irq() would make sure to verify the same condition and > > > > enqueue the task as a GP blocker if preempting inside "Few instructions 1" > > > > or "Few instructions 2". > > > > > > Ah, OK, I might be following now. ;-) > > > > > > We also need both versions of rcu_exp_handler() to check the IP as well, > > > given that sooner or later someone is going to want trampoline removal > > > to go faster. Or am I still missing a turn in here somewhere? I should add that the thing that I really like about Frederic's approach is that avoids the task-list scan. Or at least has the potential to do so. Such scans have proven problematic in the past. > > Yes indeed, missed the exp part! > > What remains to handle also is non-preemptible RCU because if the task is > preempted by an IRQ while in the .text.rcu_no_qs, we may still need to keep > track of that somewhere. Perhaps in rcu_core() in kernels booted with use_softirq? I am thinking specifically of the checks for deferred quiescent states. I don't (yet) see a need to modify rcu_check_quiescent_state(). Maybe other places as well. ;-) Thanx, Paul ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH RFC v3 03/13] rcu-tasks: Add a Tasks RCU implementation for reader-marked trampolines 2026-09-16 14:55 ` Paul E. McKenney @ 2026-09-16 15:23 ` Frederic Weisbecker 2026-09-16 15:41 ` Paul E. McKenney 0 siblings, 1 reply; 20+ messages in thread From: Frederic Weisbecker @ 2026-09-16 15:23 UTC (permalink / raw) To: Paul E. McKenney Cc: Josef Bacik, Neeraj Upadhyay, Joel Fernandes, Boqun Feng, Thomas Gleixner, Peter Zijlstra, Steven Rostedt, Masami Hiramatsu, Mark Rutland, Jiri Olsa, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, x86, Catalin Marinas, Will Deacon, Puranjay Mohan, Xu Kuohai, Andy Lutomirski, Josh Triplett, Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang, Juergen Gross, Luis Chamberlain, Ihor Solodrai, linux-kernel, rcu, linux-trace-kernel, bpf, linux-arm-kernel, xen-devel Le Wed, Sep 16, 2026 at 07:55:02AM -0700, Paul E. McKenney a écrit : > On Wed, Sep 16, 2026 at 04:47:29PM +0200, Frederic Weisbecker wrote: > > Le Wed, Sep 16, 2026 at 04:35:50PM +0200, Frederic Weisbecker a écrit : > > > Le Wed, Sep 16, 2026 at 07:26:35AM -0700, Paul E. McKenney a écrit : > > > > On Wed, Sep 16, 2026 at 02:40:20PM +0200, Frederic Weisbecker wrote: > > > > > Le Tue, Sep 15, 2026 at 04:56:41PM -0700, Paul E. McKenney a écrit : > > > > > > > Alternatively the approach could be generalized to vanilla RCU, it could be > > > > > > > possible to define a .text.rcu_no_qs section within which code running is > > > > > > > considered as an RCU reader (with a pause while on the explicit RCU tasks > > > > > > > section). It would be forbidden to voluntary sleep inside > > > > > > > and to put explicit preemption points (CONFIG_PROVE_RCU could report misuses). > > > > > > > > > > > > > > Based on IP, RCU could consider those interrupted section as readers. This would > > > > > > > require PREEMPT_RCU though. > > > > > > > > > > > > > > And then synchronize_rcu() would do the 1, 2, 4 jobs. > > > > > > > > > > > > If I am following correctly (ha!), sleepable BPF programs rule out use > > > > > > of RCU in this manner. > > > > > > > > > > > > But your point is nevertheless valid, in that SRCU could be used. > > > > > > And because rcu_read_lock_trace() is a thin wrapper around SRCU-fast, we > > > > > > *might* be able to instead use rcu_read_lock_tasks_trace(), which would > > > > > > skip the task-struct increment and decrement, saving a few instructions. > > > > > > Then, instead of waiting for each task's counter to go to zero, instead > > > > > > just invoke synchronize_rcu_tasks_trace(). > > > > > > > > > > > > Which is pretty close to what Josef is proposing, just with the new RCU > > > > > > Tasks Trace read-side primitives. I think. ;-) > > > > > > > > > > > > This assumes that we do not need to flatten partially overlapping RCU > > > > > > Tasks Trace readers into one big reader. > > > > > > > > > > > > Or am I missing something here? > > > > > > > > > > Yes I think that's what Josef does in this patchset. The problem is about > > > > > handling the few instructions: > > > > > > > > > > 1) between the begining of the trampoline and the call to rcu_read_lock_trace() > > > > > > > > > > 2) between the call to rcu_read_unlock_trace() and the end of the trampoline > > > > > > > > > > So what I'm proposing is to make those two parts implicit RCU read lock sections. > > > > > > > > > > So the whole trampoline would be .text.rcu_no_qs: > > > > > > > > > > .text.rcu_no_qs trampoline: > > > > > __________________________________________________________________________________________ > > > > > |Few instructions 1 | rcu_read_lock_trace() .... rcu_read_unlock_trace | Few instructions 2| > > > > > ___________________________________________________________________________________________ > > > > > > > > > > Then when a tick fires, rcu_flavor_sched_clock_irq() discards the interrupted > > > > > code as QS if the IP was within .text.rcu_no_qs _unless_ it is in the > > > > > rcu_read_lock_trace. Both are easy and quick to verify. > > > > > > > > > > Also preempt_schedule_irq() would make sure to verify the same condition and > > > > > enqueue the task as a GP blocker if preempting inside "Few instructions 1" > > > > > or "Few instructions 2". > > > > > > > > Ah, OK, I might be following now. ;-) > > > > > > > > We also need both versions of rcu_exp_handler() to check the IP as well, > > > > given that sooner or later someone is going to want trampoline removal > > > > to go faster. Or am I still missing a turn in here somewhere? > > I should add that the thing that I really like about Frederic's approach > is that avoids the task-list scan. Or at least has the potential to > do so. Such scans have proven problematic in the past. > > > > Yes indeed, missed the exp part! > > > > What remains to handle also is non-preemptible RCU because if the task is > > preempted by an IRQ while in the .text.rcu_no_qs, we may still need to keep > > track of that somewhere. > > Perhaps in rcu_core() in kernels booted with use_softirq? I am thinking > specifically of the checks for deferred quiescent states. I don't (yet) > see a need to modify rcu_check_quiescent_state(). > > Maybe other places as well. ;-) Hmm this tracking would have to happen on preempt_schedule() just like we do for PREEMPT_RCU. Or am I missing something? And then we would need a list scan of those tasks. Or we can build the blocked task list handling, that we already have for PREEMPT_RCU, when CONFIG_RCU_TASKS && !CONFIG_PREEMPT_RCU. We would just only add tasks when preempted in .text.rcu_no_qs since rcu_read_lock() would still disable preemption on normal explicit readers. So I wouldn't expect more overhead due to that blocked list tracking built since it would rarely track tasks. Thanks. -- Frederic Weisbecker SUSE Labs ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH RFC v3 03/13] rcu-tasks: Add a Tasks RCU implementation for reader-marked trampolines 2026-09-16 15:23 ` Frederic Weisbecker @ 2026-09-16 15:41 ` Paul E. McKenney 2026-09-17 12:14 ` Frederic Weisbecker 0 siblings, 1 reply; 20+ messages in thread From: Paul E. McKenney @ 2026-09-16 15:41 UTC (permalink / raw) To: Frederic Weisbecker Cc: Josef Bacik, Neeraj Upadhyay, Joel Fernandes, Boqun Feng, Thomas Gleixner, Peter Zijlstra, Steven Rostedt, Masami Hiramatsu, Mark Rutland, Jiri Olsa, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, x86, Catalin Marinas, Will Deacon, Puranjay Mohan, Xu Kuohai, Andy Lutomirski, Josh Triplett, Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang, Juergen Gross, Luis Chamberlain, Ihor Solodrai, linux-kernel, rcu, linux-trace-kernel, bpf, linux-arm-kernel, xen-devel On Wed, Sep 16, 2026 at 05:23:31PM +0200, Frederic Weisbecker wrote: > Le Wed, Sep 16, 2026 at 07:55:02AM -0700, Paul E. McKenney a écrit : > > On Wed, Sep 16, 2026 at 04:47:29PM +0200, Frederic Weisbecker wrote: > > > Le Wed, Sep 16, 2026 at 04:35:50PM +0200, Frederic Weisbecker a écrit : > > > > Le Wed, Sep 16, 2026 at 07:26:35AM -0700, Paul E. McKenney a écrit : > > > > > On Wed, Sep 16, 2026 at 02:40:20PM +0200, Frederic Weisbecker wrote: > > > > > > Le Tue, Sep 15, 2026 at 04:56:41PM -0700, Paul E. McKenney a écrit : > > > > > > > > Alternatively the approach could be generalized to vanilla RCU, it could be > > > > > > > > possible to define a .text.rcu_no_qs section within which code running is > > > > > > > > considered as an RCU reader (with a pause while on the explicit RCU tasks > > > > > > > > section). It would be forbidden to voluntary sleep inside > > > > > > > > and to put explicit preemption points (CONFIG_PROVE_RCU could report misuses). > > > > > > > > > > > > > > > > Based on IP, RCU could consider those interrupted section as readers. This would > > > > > > > > require PREEMPT_RCU though. > > > > > > > > > > > > > > > > And then synchronize_rcu() would do the 1, 2, 4 jobs. > > > > > > > > > > > > > > If I am following correctly (ha!), sleepable BPF programs rule out use > > > > > > > of RCU in this manner. > > > > > > > > > > > > > > But your point is nevertheless valid, in that SRCU could be used. > > > > > > > And because rcu_read_lock_trace() is a thin wrapper around SRCU-fast, we > > > > > > > *might* be able to instead use rcu_read_lock_tasks_trace(), which would > > > > > > > skip the task-struct increment and decrement, saving a few instructions. > > > > > > > Then, instead of waiting for each task's counter to go to zero, instead > > > > > > > just invoke synchronize_rcu_tasks_trace(). > > > > > > > > > > > > > > Which is pretty close to what Josef is proposing, just with the new RCU > > > > > > > Tasks Trace read-side primitives. I think. ;-) > > > > > > > > > > > > > > This assumes that we do not need to flatten partially overlapping RCU > > > > > > > Tasks Trace readers into one big reader. > > > > > > > > > > > > > > Or am I missing something here? > > > > > > > > > > > > Yes I think that's what Josef does in this patchset. The problem is about > > > > > > handling the few instructions: > > > > > > > > > > > > 1) between the begining of the trampoline and the call to rcu_read_lock_trace() > > > > > > > > > > > > 2) between the call to rcu_read_unlock_trace() and the end of the trampoline > > > > > > > > > > > > So what I'm proposing is to make those two parts implicit RCU read lock sections. > > > > > > > > > > > > So the whole trampoline would be .text.rcu_no_qs: > > > > > > > > > > > > .text.rcu_no_qs trampoline: > > > > > > __________________________________________________________________________________________ > > > > > > |Few instructions 1 | rcu_read_lock_trace() .... rcu_read_unlock_trace | Few instructions 2| > > > > > > ___________________________________________________________________________________________ > > > > > > > > > > > > Then when a tick fires, rcu_flavor_sched_clock_irq() discards the interrupted > > > > > > code as QS if the IP was within .text.rcu_no_qs _unless_ it is in the > > > > > > rcu_read_lock_trace. Both are easy and quick to verify. > > > > > > > > > > > > Also preempt_schedule_irq() would make sure to verify the same condition and > > > > > > enqueue the task as a GP blocker if preempting inside "Few instructions 1" > > > > > > or "Few instructions 2". > > > > > > > > > > Ah, OK, I might be following now. ;-) > > > > > > > > > > We also need both versions of rcu_exp_handler() to check the IP as well, > > > > > given that sooner or later someone is going to want trampoline removal > > > > > to go faster. Or am I still missing a turn in here somewhere? > > > > I should add that the thing that I really like about Frederic's approach > > is that avoids the task-list scan. Or at least has the potential to > > do so. Such scans have proven problematic in the past. > > > > > > Yes indeed, missed the exp part! > > > > > > What remains to handle also is non-preemptible RCU because if the task is > > > preempted by an IRQ while in the .text.rcu_no_qs, we may still need to keep > > > track of that somewhere. > > > > Perhaps in rcu_core() in kernels booted with use_softirq? I am thinking > > specifically of the checks for deferred quiescent states. I don't (yet) > > see a need to modify rcu_check_quiescent_state(). > > > > Maybe other places as well. ;-) > > Hmm this tracking would have to happen on preempt_schedule() just like we > do for PREEMPT_RCU. Or am I missing something? And then we would need a > list scan of those tasks. I am thinking of the case where a trampoline is interrupted before entering (or after leaving) its RCU Tasks Trace read-side critical section. Then there is a softirq handler on the back of that interrupt handler, and RCU_SOFTIRQ is invoked, calling rcu_core(). Specifically: /* Report any deferred quiescent states if preemption enabled. */ if (IS_ENABLED(CONFIG_PREEMPT_COUNT) && (!(preempt_count() & PREEMPT_MASK))) { rcu_preempt_deferred_qs(current); } else if (rcu_preempt_need_deferred_qs(current)) { guard(irqsave)(); set_need_resched_current(); } Preemption is enabled, but we should not report a quiescent state because we have interrupted a trampoline. Correct? > Or we can build the blocked task list handling, that we already have for PREEMPT_RCU, > when CONFIG_RCU_TASKS && !CONFIG_PREEMPT_RCU. We would just only add tasks when > preempted in .text.rcu_no_qs since rcu_read_lock() would still disable > preemption on normal explicit readers. So I wouldn't expect more overhead due to > that blocked list tracking built since it would rarely track tasks. Yes, we could avoid the list of tasks by treating the preemption within the trampoline the same as preemption within an RCU read-side critical section, but there might not be an rcu_read_unlock() to clean up. Which could be a problem. Trampolines that transfer control to tracing code could supply the needed cleanup call. But last I checked, there were trampolines that transferred directly back to the original code, with no opportunity for cleaning up. Or am I still missing a trick here? Thanx, Paul ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH RFC v3 03/13] rcu-tasks: Add a Tasks RCU implementation for reader-marked trampolines 2026-09-16 15:41 ` Paul E. McKenney @ 2026-09-17 12:14 ` Frederic Weisbecker 2026-09-17 15:40 ` Paul E. McKenney 0 siblings, 1 reply; 20+ messages in thread From: Frederic Weisbecker @ 2026-09-17 12:14 UTC (permalink / raw) To: Paul E. McKenney Cc: Josef Bacik, Neeraj Upadhyay, Joel Fernandes, Boqun Feng, Thomas Gleixner, Peter Zijlstra, Steven Rostedt, Masami Hiramatsu, Mark Rutland, Jiri Olsa, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, x86, Catalin Marinas, Will Deacon, Puranjay Mohan, Xu Kuohai, Andy Lutomirski, Josh Triplett, Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang, Juergen Gross, Luis Chamberlain, Ihor Solodrai, linux-kernel, rcu, linux-trace-kernel, bpf, linux-arm-kernel, xen-devel Le Wed, Sep 16, 2026 at 08:41:36AM -0700, Paul E. McKenney a écrit : > On Wed, Sep 16, 2026 at 05:23:31PM +0200, Frederic Weisbecker wrote: > > Le Wed, Sep 16, 2026 at 07:55:02AM -0700, Paul E. McKenney a écrit : > > > Perhaps in rcu_core() in kernels booted with use_softirq? I am thinking > > > specifically of the checks for deferred quiescent states. I don't (yet) > > > see a need to modify rcu_check_quiescent_state(). > > > > > > Maybe other places as well. ;-) > > > > Hmm this tracking would have to happen on preempt_schedule() just like we > > do for PREEMPT_RCU. Or am I missing something? And then we would need a > > list scan of those tasks. > > I am thinking of the case where a trampoline is interrupted before entering > (or after leaving) its RCU Tasks Trace read-side critical section. Then > there is a softirq handler on the back of that interrupt handler, and > RCU_SOFTIRQ is invoked, calling rcu_core(). Specifically: > > /* Report any deferred quiescent states if preemption enabled. */ > if (IS_ENABLED(CONFIG_PREEMPT_COUNT) && (!(preempt_count() & PREEMPT_MASK))) { > rcu_preempt_deferred_qs(current); > } else if (rcu_preempt_need_deferred_qs(current)) { > guard(irqsave)(); > set_need_resched_current(); > } > > Preemption is enabled, but we should not report a quiescent state because > we have interrupted a trampoline. Correct? Right! > > Or we can build the blocked task list handling, that we already have for PREEMPT_RCU, > > when CONFIG_RCU_TASKS && !CONFIG_PREEMPT_RCU. We would just only add tasks when > > preempted in .text.rcu_no_qs since rcu_read_lock() would still disable > > preemption on normal explicit readers. So I wouldn't expect more overhead due to > > that blocked list tracking built since it would rarely track tasks. > > Yes, we could avoid the list of tasks by treating the preemption within > the trampoline the same as preemption within an RCU read-side critical > section, but there might not be an rcu_read_unlock() to clean up. > Which could be a problem. Ah yes, good point. > > Trampolines that transfer control to tracing code could supply the needed > cleanup call. But last I checked, there were trampolines that transferred > directly back to the original code, with no opportunity for cleaning up. > > Or am I still missing a trick here? You're right. So we'll indeed need to reuse the deferred qs points here. Thanks. -- Frederic Weisbecker SUSE Labs ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH RFC v3 03/13] rcu-tasks: Add a Tasks RCU implementation for reader-marked trampolines 2026-09-17 12:14 ` Frederic Weisbecker @ 2026-09-17 15:40 ` Paul E. McKenney 2026-09-17 16:35 ` Josef Bacik 2026-09-17 18:45 ` Frederic Weisbecker 0 siblings, 2 replies; 20+ messages in thread From: Paul E. McKenney @ 2026-09-17 15:40 UTC (permalink / raw) To: Frederic Weisbecker Cc: Josef Bacik, Neeraj Upadhyay, Joel Fernandes, Boqun Feng, Thomas Gleixner, Peter Zijlstra, Steven Rostedt, Masami Hiramatsu, Mark Rutland, Jiri Olsa, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, x86, Catalin Marinas, Will Deacon, Puranjay Mohan, Xu Kuohai, Andy Lutomirski, Josh Triplett, Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang, Juergen Gross, Luis Chamberlain, Ihor Solodrai, linux-kernel, rcu, linux-trace-kernel, bpf, linux-arm-kernel, xen-devel On Thu, Sep 17, 2026 at 02:14:57PM +0200, Frederic Weisbecker wrote: > Le Wed, Sep 16, 2026 at 08:41:36AM -0700, Paul E. McKenney a écrit : > > On Wed, Sep 16, 2026 at 05:23:31PM +0200, Frederic Weisbecker wrote: > > > Le Wed, Sep 16, 2026 at 07:55:02AM -0700, Paul E. McKenney a écrit : > > > > Perhaps in rcu_core() in kernels booted with use_softirq? I am thinking > > > > specifically of the checks for deferred quiescent states. I don't (yet) > > > > see a need to modify rcu_check_quiescent_state(). > > > > > > > > Maybe other places as well. ;-) > > > > > > Hmm this tracking would have to happen on preempt_schedule() just like we > > > do for PREEMPT_RCU. Or am I missing something? And then we would need a > > > list scan of those tasks. > > > > I am thinking of the case where a trampoline is interrupted before entering > > (or after leaving) its RCU Tasks Trace read-side critical section. Then > > there is a softirq handler on the back of that interrupt handler, and > > RCU_SOFTIRQ is invoked, calling rcu_core(). Specifically: > > > > /* Report any deferred quiescent states if preemption enabled. */ > > if (IS_ENABLED(CONFIG_PREEMPT_COUNT) && (!(preempt_count() & PREEMPT_MASK))) { > > rcu_preempt_deferred_qs(current); > > } else if (rcu_preempt_need_deferred_qs(current)) { > > guard(irqsave)(); > > set_need_resched_current(); > > } > > > > Preemption is enabled, but we should not report a quiescent state because > > we have interrupted a trampoline. Correct? > > Right! > > > > Or we can build the blocked task list handling, that we already have for PREEMPT_RCU, > > > when CONFIG_RCU_TASKS && !CONFIG_PREEMPT_RCU. We would just only add tasks when > > > preempted in .text.rcu_no_qs since rcu_read_lock() would still disable > > > preemption on normal explicit readers. So I wouldn't expect more overhead due to > > > that blocked list tracking built since it would rarely track tasks. > > > > Yes, we could avoid the list of tasks by treating the preemption within > > the trampoline the same as preemption within an RCU read-side critical > > section, but there might not be an rcu_read_unlock() to clean up. > > Which could be a problem. > > Ah yes, good point. > > > > > Trampolines that transfer control to tracing code could supply the needed > > cleanup call. But last I checked, there were trampolines that transferred > > directly back to the original code, with no opportunity for cleaning up. > > > > Or am I still missing a trick here? > > You're right. So we'll indeed need to reuse the deferred qs points here. Except this is getting a bit involved. Don't get me wrong, if Josef is happy to take this on, far be it from me to stand in his way. But if not, we should be willing to treat this optimization as a follow-on effort, whether by Josef or someone else. Thanx, Paul ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH RFC v3 03/13] rcu-tasks: Add a Tasks RCU implementation for reader-marked trampolines 2026-09-17 15:40 ` Paul E. McKenney @ 2026-09-17 16:35 ` Josef Bacik 2026-09-17 16:55 ` Paul E. McKenney 2026-09-17 18:45 ` Frederic Weisbecker 1 sibling, 1 reply; 20+ messages in thread From: Josef Bacik @ 2026-09-17 16:35 UTC (permalink / raw) To: Paul E. McKenney, Frederic Weisbecker Cc: Neeraj Upadhyay, Joel Fernandes, Boqun Feng, Thomas Gleixner, Peter Zijlstra, Steven Rostedt, Masami Hiramatsu, Mark Rutland, Jiri Olsa, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, x86, Catalin Marinas, Will Deacon, Puranjay Mohan, Xu Kuohai, Andy Lutomirski, Josh Triplett, Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang, Juergen Gross, Luis Chamberlain, Ihor Solodrai, linux-kernel, rcu, linux-trace-kernel, bpf, linux-arm-kernel, xen-devel On Thu, 17 Sep 2026 08:40:10 -0700, Paul E. McKenney wrote: > On Thu, Sep 17, 2026 at 02:14:57PM +0200, Frederic Weisbecker wrote: > > You're right. So we'll indeed need to reuse the deferred qs points here. > > Except this is getting a bit involved. > > Don't get me wrong, if Josef is happy to take this on, far be it from me > to stand in his way. But if not, we should be willing to treat this > optimization as a follow-on effort, whether by Josef or someone else. Follow-on works for me. For what it's worth 03/13 is already fairly close to what Frederic describes, just outside the core flavor: no task-list scan (the GP waits per CPU for a pass through __schedule() or an EQS), the irq-exit preemption path checks the interrupted IP and queues the task as a holdout before the switch, and the holdout is keyed on where the task was interrupted so nothing is needed from the trampoline tail. Moving that IP check into the tick / rcu_exp_handler() / deferred-QS paths and reusing the blocked-tasks list is something I'm happy to look at once this has settled. v4 will pick up Alexei's ask (reader emitted by the BPF JIT around the fentry and fexit regions rather than in the glue) and the idle-CPU hole Sashiko found. Thanks, Josef ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH RFC v3 03/13] rcu-tasks: Add a Tasks RCU implementation for reader-marked trampolines 2026-09-17 16:35 ` Josef Bacik @ 2026-09-17 16:55 ` Paul E. McKenney 0 siblings, 0 replies; 20+ messages in thread From: Paul E. McKenney @ 2026-09-17 16:55 UTC (permalink / raw) To: Josef Bacik Cc: Frederic Weisbecker, Neeraj Upadhyay, Joel Fernandes, Boqun Feng, Thomas Gleixner, Peter Zijlstra, Steven Rostedt, Masami Hiramatsu, Mark Rutland, Jiri Olsa, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, x86, Catalin Marinas, Will Deacon, Puranjay Mohan, Xu Kuohai, Andy Lutomirski, Josh Triplett, Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang, Juergen Gross, Luis Chamberlain, Ihor Solodrai, linux-kernel, rcu, linux-trace-kernel, bpf, linux-arm-kernel, xen-devel On Thu, Sep 17, 2026 at 04:35:19PM +0000, Josef Bacik wrote: > On Thu, 17 Sep 2026 08:40:10 -0700, Paul E. McKenney wrote: > > On Thu, Sep 17, 2026 at 02:14:57PM +0200, Frederic Weisbecker wrote: > > > You're right. So we'll indeed need to reuse the deferred qs points here. > > > > Except this is getting a bit involved. > > > > Don't get me wrong, if Josef is happy to take this on, far be it from me > > to stand in his way. But if not, we should be willing to treat this > > optimization as a follow-on effort, whether by Josef or someone else. > > Follow-on works for me. For what it's worth 03/13 is already fairly > close to what Frederic describes, just outside the core flavor: no > task-list scan (the GP waits per CPU for a pass through __schedule() or > an EQS), the irq-exit preemption path checks the interrupted IP and > queues the task as a holdout before the switch, and the holdout is keyed > on where the task was interrupted so nothing is needed from the > trampoline tail. Moving that IP check into the tick / rcu_exp_handler() / > deferred-QS paths and reusing the blocked-tasks list is something I'm > happy to look at once this has settled. > > v4 will pick up Alexei's ask (reader emitted by the BPF JIT around the > fentry and fexit regions rather than in the glue) and the idle-CPU hole > Sashiko found. Even better! ;-) Thanx, Paul ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH RFC v3 03/13] rcu-tasks: Add a Tasks RCU implementation for reader-marked trampolines 2026-09-17 15:40 ` Paul E. McKenney 2026-09-17 16:35 ` Josef Bacik @ 2026-09-17 18:45 ` Frederic Weisbecker 2026-09-17 19:25 ` Paul E. McKenney 1 sibling, 1 reply; 20+ messages in thread From: Frederic Weisbecker @ 2026-09-17 18:45 UTC (permalink / raw) To: Paul E. McKenney Cc: Josef Bacik, Neeraj Upadhyay, Joel Fernandes, Boqun Feng, Thomas Gleixner, Peter Zijlstra, Steven Rostedt, Masami Hiramatsu, Mark Rutland, Jiri Olsa, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, x86, Catalin Marinas, Will Deacon, Puranjay Mohan, Xu Kuohai, Andy Lutomirski, Josh Triplett, Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang, Juergen Gross, Luis Chamberlain, Ihor Solodrai, linux-kernel, rcu, linux-trace-kernel, bpf, linux-arm-kernel, xen-devel Le Thu, Sep 17, 2026 at 08:40:10AM -0700, Paul E. McKenney a écrit : > > > Trampolines that transfer control to tracing code could supply the needed > > > cleanup call. But last I checked, there were trampolines that transferred > > > directly back to the original code, with no opportunity for cleaning up. > > > > > > Or am I still missing a trick here? > > > > You're right. So we'll indeed need to reuse the deferred qs points here. > > Except this is getting a bit involved. > > Don't get me wrong, if Josef is happy to take this on, far be it from me > to stand in his way. But if not, we should be willing to treat this > optimization as a follow-on effort, whether by Josef or someone else. Sure, I guess I can try the follow-on, especially if it leads to removing all this RCU tasks black magic. Thanks. -- Frederic Weisbecker SUSE Labs ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH RFC v3 03/13] rcu-tasks: Add a Tasks RCU implementation for reader-marked trampolines 2026-09-17 18:45 ` Frederic Weisbecker @ 2026-09-17 19:25 ` Paul E. McKenney 2026-09-17 20:31 ` Paul E. McKenney 0 siblings, 1 reply; 20+ messages in thread From: Paul E. McKenney @ 2026-09-17 19:25 UTC (permalink / raw) To: Frederic Weisbecker Cc: Josef Bacik, Neeraj Upadhyay, Joel Fernandes, Boqun Feng, Thomas Gleixner, Peter Zijlstra, Steven Rostedt, Masami Hiramatsu, Mark Rutland, Jiri Olsa, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, x86, Catalin Marinas, Will Deacon, Puranjay Mohan, Xu Kuohai, Andy Lutomirski, Josh Triplett, Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang, Juergen Gross, Luis Chamberlain, Ihor Solodrai, linux-kernel, rcu, linux-trace-kernel, bpf, linux-arm-kernel, xen-devel On Thu, Sep 17, 2026 at 08:45:36PM +0200, Frederic Weisbecker wrote: > Le Thu, Sep 17, 2026 at 08:40:10AM -0700, Paul E. McKenney a écrit : > > > > Trampolines that transfer control to tracing code could supply the needed > > > > cleanup call. But last I checked, there were trampolines that transferred > > > > directly back to the original code, with no opportunity for cleaning up. > > > > > > > > Or am I still missing a trick here? > > > > > > You're right. So we'll indeed need to reuse the deferred qs points here. > > > > Except this is getting a bit involved. > > > > Don't get me wrong, if Josef is happy to take this on, far be it from me > > to stand in his way. But if not, we should be willing to treat this > > optimization as a follow-on effort, whether by Josef or someone else. > > Sure, I guess I can try the follow-on, especially if it leads to removing > all this RCU tasks black magic. That sounds most excellent, thank you! Thanx, Paul ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH RFC v3 03/13] rcu-tasks: Add a Tasks RCU implementation for reader-marked trampolines 2026-09-17 19:25 ` Paul E. McKenney @ 2026-09-17 20:31 ` Paul E. McKenney 0 siblings, 0 replies; 20+ messages in thread From: Paul E. McKenney @ 2026-09-17 20:31 UTC (permalink / raw) To: Frederic Weisbecker Cc: Josef Bacik, Neeraj Upadhyay, Joel Fernandes, Boqun Feng, Thomas Gleixner, Peter Zijlstra, Steven Rostedt, Masami Hiramatsu, Mark Rutland, Jiri Olsa, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, x86, Catalin Marinas, Will Deacon, Puranjay Mohan, Xu Kuohai, Andy Lutomirski, Josh Triplett, Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang, Juergen Gross, Luis Chamberlain, Ihor Solodrai, linux-kernel, rcu, linux-trace-kernel, bpf, linux-arm-kernel, xen-devel On Thu, Sep 17, 2026 at 12:25:40PM -0700, Paul E. McKenney wrote: > On Thu, Sep 17, 2026 at 08:45:36PM +0200, Frederic Weisbecker wrote: > > Le Thu, Sep 17, 2026 at 08:40:10AM -0700, Paul E. McKenney a écrit : > > > > > Trampolines that transfer control to tracing code could supply the needed > > > > > cleanup call. But last I checked, there were trampolines that transferred > > > > > directly back to the original code, with no opportunity for cleaning up. > > > > > > > > > > Or am I still missing a trick here? > > > > > > > > You're right. So we'll indeed need to reuse the deferred qs points here. > > > > > > Except this is getting a bit involved. > > > > > > Don't get me wrong, if Josef is happy to take this on, far be it from me > > > to stand in his way. But if not, we should be willing to treat this > > > optimization as a follow-on effort, whether by Josef or someone else. > > > > Sure, I guess I can try the follow-on, especially if it leads to removing > > all this RCU tasks black magic. > > That sounds most excellent, thank you! Ah, and in case anyone (especially Josef) is wondering, one big advantage of the more elaborate approach is that it allowed the real-time guys to avoid yet another source of IPIs messing with their latencies. Thanx, Paul ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH RFC v3 03/13] rcu-tasks: Add a Tasks RCU implementation for reader-marked trampolines [not found] ` <20260915-b4-rcu-tasks-preempt-qs-v3-3-0ad30c4c5ee7@toxicpanda.com> 2026-09-15 15:14 ` [PATCH RFC v3 03/13] rcu-tasks: Add a Tasks RCU implementation for reader-marked trampolines Frederic Weisbecker @ 2026-09-17 20:20 ` Frederic Weisbecker 2026-09-22 15:42 ` Josef Bacik 1 sibling, 1 reply; 20+ messages in thread From: Frederic Weisbecker @ 2026-09-17 20:20 UTC (permalink / raw) To: Josef Bacik Cc: Paul E. McKenney, Neeraj Upadhyay, Joel Fernandes, Boqun Feng, Thomas Gleixner, Peter Zijlstra, Steven Rostedt, Masami Hiramatsu, Mark Rutland, Jiri Olsa, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, x86, Catalin Marinas, Will Deacon, Puranjay Mohan, Xu Kuohai, Andy Lutomirski, Josh Triplett, Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang, Juergen Gross, Luis Chamberlain, Ihor Solodrai, linux-kernel, rcu, linux-trace-kernel, bpf, linux-arm-kernel, xen-devel Le Tue, Sep 15, 2026 at 01:17:30PM +0000, Josef Bacik a écrit : > +static void rcu_tasks_tramp_hold(struct task_struct *t) > +{ > + unsigned long flags; > + > + if (t->rcu_tasks_holdout) > + return; > + raw_spin_lock_irqsave(&rcu_tasks_tramp_lock, flags); > + list_add_tail(&t->rcu_tasks_holdout_list, &rcu_tasks_tramp_holdouts); > + WRITE_ONCE(t->rcu_tasks_holdout, true); > + raw_spin_unlock_irqrestore(&rcu_tasks_tramp_lock, flags); > +} > + > +static void rcu_tasks_tramp_release(struct task_struct *t) > +{ > + unsigned long flags; > + > + if (likely(!t->rcu_tasks_holdout)) > + return; > + raw_spin_lock_irqsave(&rcu_tasks_tramp_lock, flags); > + list_del_init(&t->rcu_tasks_holdout_list); > + WRITE_ONCE(t->rcu_tasks_holdout, false); > + raw_spin_unlock_irqrestore(&rcu_tasks_tramp_lock, flags); > +} > + > +/** > + * rcu_tasks_irq_resched_enter - Tasks RCU hook for the irq-exit reschedule check > + * @ip: instruction pointer of the interrupted (task-level) context > + * > + * Called with interrupts disabled when an interrupt returning to kernel > + * mode is about to preempt_schedule_irq(), the one context switch that can > + * catch a task inside unmarked trampoline text. Record where the task is > + * parked for as long as it is (rcu_tasks_wait_irq_preempted() looks at > + * that), and if it is inside such text make it a holdout before > + * __schedule() reports the quiescent event; if it is not, this is as good > + * as a voluntary switch for ending an earlier hold. > + */ > +void rcu_tasks_irq_resched_enter(unsigned long ip) > +{ > + struct task_struct *t = current; > + struct rcu_tasks_percpu *rtpcp = this_cpu_ptr(rcu_tasks.rtpcpu); > + > + lockdep_assert_irqs_disabled(); > + WRITE_ONCE(t->rcu_tasks_irq_ip, ip); > + t->rcu_tasks_exit_cpu = smp_processor_id(); > + raw_spin_lock_rcu_node(rtpcp); > + list_add(&t->rcu_tasks_exit_list, &rtpcp->rtp_exit_list); > + raw_spin_unlock_rcu_node(rtpcp); I don't think we can do that. This is too much unconditional overhead on the hot preemption path. rcu_tasks_trampoline_text() should be a condition here. And do we really need to maintain both lists? I understand that they have different purposes. ->rcu_tasks_exit_list is to track preempted tasks on trampoline ->rcu_tasks_holdout_list is to track preempted tasks on trampoline until they ever voluntary schedule() Can the latter replace the former? I see it's used on kprobes and others but I haven't checked the details yet. Thanks. -- Frederic Weisbecker SUSE Labs ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH RFC v3 03/13] rcu-tasks: Add a Tasks RCU implementation for reader-marked trampolines 2026-09-17 20:20 ` Frederic Weisbecker @ 2026-09-22 15:42 ` Josef Bacik 0 siblings, 0 replies; 20+ messages in thread From: Josef Bacik @ 2026-09-22 15:42 UTC (permalink / raw) To: Frederic Weisbecker Cc: Paul E. McKenney, Boqun Feng, Thomas Gleixner, Peter Zijlstra, Steven Rostedt, Masami Hiramatsu, Mark Rutland, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, Puranjay Mohan, linux-kernel, rcu, linux-trace-kernel, bpf, linux-arm-kernel On Thu, 17 Sep 2026 22:20:17 +0200, Frederic Weisbecker wrote: > > + lockdep_assert_irqs_disabled(); > > + WRITE_ONCE(t->rcu_tasks_irq_ip, ip); > > + t->rcu_tasks_exit_cpu = smp_processor_id(); > > + raw_spin_lock_rcu_node(rtpcp); > > + list_add(&t->rcu_tasks_exit_list, &rtpcp->rtp_exit_list); > > + raw_spin_unlock_rcu_node(rtpcp); > > I don't think we can do that. This is too much unconditional overhead > on the hot preemption path. rcu_tasks_trampoline_text() should be > a condition here. Sorry, I missed this one before sending v4/v5. Agreed, and it is gone for v6: the hook is now just the WRITE_ONCE() of the IP plus the rcu_tasks_trampoline_text() check, and the exit side a single store. The IP store itself has to stay unconditional because of the kprobe jump optimizer: its window is ordinary text, so a task parked there before the optimizer decided to patch was not "trampoline text" when it was preempted, and the optimizer needs to find it afterwards. > And do we really need to maintain both lists? I understand that they > have different purposes. [...] > Can the latter replace the former? With the above there is only the holdout list left. The optimizer's rcu_tasks_wait_irq_preempted() now does what classic does for its scan: walk the task list plus the per-CPU exit lists (exit_tasks_rcu_start() and friends stay shared between the two flavors for that), checking each task's recorded IP. That is a slow path that only kprobe optimization hits. And thanks for picking up the core-RCU follow-on. Josef ^ permalink raw reply [flat|nested] 20+ messages in thread
[parent not found: <20260915-b4-rcu-tasks-preempt-qs-v3-6-0ad30c4c5ee7@toxicpanda.com>]
[parent not found: <DLGFJY2ZSY5M.11K7HK2CLM6I9@gmail.com>]
* Re: [PATCH RFC v3 06/13] bpf: Take a Tasks Trace reader in the trampoline glue [not found] ` <DLGFJY2ZSY5M.11K7HK2CLM6I9@gmail.com> @ 2026-09-17 1:16 ` Josef Bacik 2026-09-17 2:24 ` Alexei Starovoitov 0 siblings, 1 reply; 20+ messages in thread From: Josef Bacik @ 2026-09-17 1:16 UTC (permalink / raw) To: Alexei Starovoitov Cc: Paul E. McKenney, Frederic Weisbecker, Neeraj Upadhyay, Joel Fernandes, Boqun Feng, Thomas Gleixner, Peter Zijlstra, Steven Rostedt, Masami Hiramatsu, Mark Rutland, Jiri Olsa, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, x86, Catalin Marinas, Will Deacon, Puranjay Mohan, Xu Kuohai, Andy Lutomirski, Josh Triplett, Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang, Juergen Gross, Luis Chamberlain, Ihor Solodrai, linux-kernel, rcu, linux-trace-kernel, bpf, linux-arm-kernel, xen-devel On Wed, 16 Sep 2026 03:45:16 +0000, Alexei Starovoitov wrote: > On Tue Sep 15, 2026 at 1:17 PM UTC, Josef Bacik wrote: > > __acquires(RCU) > > { > > + bpf_tramp_read_lock_trace(); > > rcu_read_lock_dont_migrate(); > > This is double increment. rcu_read_lock_dont_migrate() includes > rcu_read_lock_trace(). Unless I'm looking at the wrong tree it doesn't, on Linus' master and on bpf-next it is static __always_inline void rcu_read_lock_dont_migrate(void) { if (IS_ENABLED(CONFIG_PREEMPT_RCU)) migrate_disable(); rcu_read_lock(); } so plain RCU plus migrate_disable(), no Tasks Trace reader. That is why the non-sleepable glue needs one added here: on these architectures the trampoline image the glue returns into is only kept alive by Tasks RCU while the task is a rcu_read_lock_trace() reader, and rcu_read_lock() does not give us that. It is two counters for a non-sleepable prog on x86-64/arm64 though, rcu_read_lock()'s and trc_reader_nesting plus the SRCU-fast percpu one, if that is what you meant. I don't see a way around it short of not using Tasks Trace as the trampoline reader: the prog still needs plain RCU for everything it dereferences, and the image needs something that survives preemption. It is compiled out on every other configuration and nothing changes in the JITed image. If you would rather the reader be taken once around the whole image in the JIT instead of per prog in the glue (which would also let the fentry-only teardown stay a single grace period), I can do that for x86 and arm64, it is what v2 did with the private counter. Separately, Junseo's "bpf: keep trampoline progs alive until image release" also adds bpf_tramp_image::nr_progs; if that lands first I will just use it here. Thanks, Josef ^ permalink raw reply [flat|nested] 20+ messages in thread
* Re: [PATCH RFC v3 06/13] bpf: Take a Tasks Trace reader in the trampoline glue 2026-09-17 1:16 ` [PATCH RFC v3 06/13] bpf: Take a Tasks Trace reader in the trampoline glue Josef Bacik @ 2026-09-17 2:24 ` Alexei Starovoitov 0 siblings, 0 replies; 20+ messages in thread From: Alexei Starovoitov @ 2026-09-17 2:24 UTC (permalink / raw) To: Josef Bacik Cc: Paul E. McKenney, Frederic Weisbecker, Neeraj Upadhyay, Joel Fernandes, Boqun Feng, Thomas Gleixner, Peter Zijlstra, Steven Rostedt, Masami Hiramatsu, Mark Rutland, Jiri Olsa, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, x86, Catalin Marinas, Will Deacon, Puranjay Mohan, Xu Kuohai, Andy Lutomirski, Josh Triplett, Uladzislau Rezki, Mathieu Desnoyers, Lai Jiangshan, Zqiang, Juergen Gross, Luis Chamberlain, Ihor Solodrai, linux-kernel, rcu, linux-trace-kernel, bpf, linux-arm-kernel, xen-devel On Thu Sep 17, 2026 at 1:16 AM UTC, Josef Bacik wrote: > On Wed, 16 Sep 2026 03:45:16 +0000, Alexei Starovoitov wrote: > > On Tue Sep 15, 2026 at 1:17 PM UTC, Josef Bacik wrote: > > > __acquires(RCU) > > > { > > > + bpf_tramp_read_lock_trace(); > > > rcu_read_lock_dont_migrate(); > > > > This is double increment. rcu_read_lock_dont_migrate() includes > > rcu_read_lock_trace(). > > Unless I'm looking at the wrong tree it doesn't, on Linus' master and on > bpf-next it is > > static __always_inline void rcu_read_lock_dont_migrate(void) > { > if (IS_ENABLED(CONFIG_PREEMPT_RCU)) > migrate_disable(); > rcu_read_lock(); > } > > so plain RCU plus migrate_disable(), no Tasks Trace reader. That is why > the non-sleepable glue needs one added here: on these architectures the > trampoline image the glue returns into is only kept alive by Tasks RCU > while the task is a rcu_read_lock_trace() reader, and rcu_read_lock() > does not give us that. Right. I got confused. Since rcu_read_lock_trace() CS will cover both sleepable and non-sleepable prog types let's do it once per fentry+fmod_ret region and 2nd time for fexit region. We probably don't want to hold it for the whole trampoline, since orig_call will delay freeing of progs. ^ permalink raw reply [flat|nested] 20+ messages in thread
end of thread, other threads:[~2026-09-22 15:43 UTC | newest]
Thread overview: 20+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
[not found] <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com>
[not found] ` <20260915-b4-rcu-tasks-preempt-qs-v3-3-0ad30c4c5ee7@toxicpanda.com>
2026-09-15 15:14 ` [PATCH RFC v3 03/13] rcu-tasks: Add a Tasks RCU implementation for reader-marked trampolines Frederic Weisbecker
2026-09-15 23:56 ` Paul E. McKenney
2026-09-16 12:40 ` Frederic Weisbecker
2026-09-16 14:26 ` Paul E. McKenney
2026-09-16 14:35 ` Frederic Weisbecker
2026-09-16 14:47 ` Frederic Weisbecker
2026-09-16 14:55 ` Paul E. McKenney
2026-09-16 15:23 ` Frederic Weisbecker
2026-09-16 15:41 ` Paul E. McKenney
2026-09-17 12:14 ` Frederic Weisbecker
2026-09-17 15:40 ` Paul E. McKenney
2026-09-17 16:35 ` Josef Bacik
2026-09-17 16:55 ` Paul E. McKenney
2026-09-17 18:45 ` Frederic Weisbecker
2026-09-17 19:25 ` Paul E. McKenney
2026-09-17 20:31 ` Paul E. McKenney
2026-09-17 20:20 ` Frederic Weisbecker
2026-09-22 15:42 ` Josef Bacik
[not found] ` <20260915-b4-rcu-tasks-preempt-qs-v3-6-0ad30c4c5ee7@toxicpanda.com>
[not found] ` <DLGFJY2ZSY5M.11K7HK2CLM6I9@gmail.com>
2026-09-17 1:16 ` [PATCH RFC v3 06/13] bpf: Take a Tasks Trace reader in the trampoline glue Josef Bacik
2026-09-17 2:24 ` Alexei Starovoitov
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox