The Linux Kernel Mailing List
 help / color / mirror / Atom feed
* [PATCH v4 0/6] rcu,srcu: Make call_rcu()/call_srcu() safe from any context
@ 2026-08-10 12:27 Puranjay Mohan
  2026-08-10 12:27 ` [PATCH v4 1/6] rcu: Make call_rcu() safe to call " Puranjay Mohan
                   ` (5 more replies)
  0 siblings, 6 replies; 11+ messages in thread
From: Puranjay Mohan @ 2026-08-10 12:27 UTC (permalink / raw)
  To: Lai Jiangshan, Paul E. McKenney, Josh Triplett, Onur Özkan,
	Frederic Weisbecker, Neeraj Upadhyay, Joel Fernandes, Boqun Feng,
	Uladzislau Rezki, Davidlohr Bueso, Andrii Nakryiko,
	Eduard Zingerman, Alexei Starovoitov, Daniel Borkmann,
	Kumar Kartikeya Dwivedi
  Cc: Puranjay Mohan, Steven Rostedt, Mathieu Desnoyers, Zqiang,
	Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
	Emil Tsalapatis, Matt Fleming, Harry Yoo (Oracle), linux-kernel,
	rcu, bpf, linux-rt-devel

call_rcu() and call_srcu() only ever touch their per-CPU callback lists
with interrupts disabled: the enqueue runs under local_irq_save() (and the
nocb locks when offloaded), and so do callback invocation and grace-period
work.  That is fine as long as call_rcu() itself is invoked with
interrupts enabled, but it is not always.  An NMI handler can call
call_rcu(), and instrumentation can reenter it.  The case that prompted
this is a BPF program attached to rcu_segcblist_enqueue() that frees an
object: the free reaches call_rcu_tasks_trace(), which is call_srcu()
under the hood, back on the same CPU with the srcu_data lock already held,
and it deadlocks on that lock.  Either way, enqueuing directly can corrupt
the list or deadlock.

Rather than scatter context checks through the enqueue, make it defer
whenever interrupts are disabled: stage the callback on a per-CPU lockless
list and re-issue it from an irq_work once interrupts are back on, going
straight to the enqueue helper so the re-issue cannot defer again.  Only
the drain side takes a lock; the staging is a bare llist_add() and stays
safe from NMI.  This is behind a new hidden CONFIG_RCU_DEFER, which is set
wherever a reentrant enqueue is possible (HAVE_NMI, KPROBES,
FUNCTION_TRACER or TRACEPOINTS); without it call_rcu() enqueues exactly as
before.  The gate is bare irqs_disabled(), so callers that merely hold
interrupts off are deferred too and pay one irq_work hop.

CPU offline is the awkward part.  A callback can be deferred very late in
the outgoing CPU's teardown -- from do_idle() or cpuhp_ap_report_dead(),
past the CPUHP_AP_SMPCFD_DYING flush that would otherwise run the irq_work
-- so the irq_work can no longer run there to re-issue it.  rcu_barrier()
and srcu_barrier() therefore drain every CPU's deferred list themselves
before they wait.  They drain rather than wait the irq_work out because
irq_work_sync() parks on an rcuwait, which holds a single waiter, so two
concurrent barriers would clobber each other's wakeup.
rcutree_migrate_callbacks() drains the outgoing CPU's list
too, so a late deferral still lands on a callback list even when nobody
calls a barrier.  To keep those drainers from stepping on each other, the
drain holds a per-CPU raw lock across the llist_del_all() and the
re-issue, so a drainer never returns having pulled callbacks off the
deferred list but not yet put them on a callback list.  Every lock the
re-issue touches (nocb, rcu_node, srcu_data) is already raw, so the
nesting is fine.

The drain re-issues with interrupts disabled, so instrumentation on the
enqueue path can re-enter call_rcu()/call_srcu() from inside it, stage
another callback, re-raise the irq_work, and the drain never finishes.  A
per-CPU flag catches that: a deferral that arrives while that CPU is
inside its own irq_work drain, and is not from an NMI, is dropped rather
than staged, with a WARN_ONCE() under CONFIG_PROVE_RCU whose backtrace
names the instrumentation responsible.  Dropping leaks that callback, and
can strand state its caller tied to it, but the alternative is a CPU that
never leaves the drain, and the producer is a BPF program that emits one
callback per enqueue, so there is nothing finite to wait for.  Only the
irq_work drain sets the flag: a direct drain from a barrier or from
CPU-offline re-issues onto the current CPU, so anything staged during it
is picked up by that CPU's own irq_work rather than feeding the drain in
progress, and no legitimate callback is dropped. Instrumenting the
irq_work machinery itself can still loop, as it can for any irq_work user,
and is not something this series can fix.

The irq_work is IRQ_WORK_INIT_HARD in all four flavors.  It is not needed
for correctness, but a non-HARD irq_work runs from a kthread on
PREEMPT_RT and can be delayed under load, letting deferred callbacks pile
up; running the re-issue in hard-irq context keeps that from turning into
an OOM.

Patches 1 and 2 do Tree and Tiny RCU, 3 and 4 Tree and Tiny SRCU.  Patch 5
teaches rcutorture to issue ->call() from a perf-overflow NMI -- the
nmi_calls parameter, on by default -- on the flavors that advertise it,
and checks that every callback issued from NMI is later invoked.  Patch 6
adds the BPF reentry reproducer described above.

Changelog:
v3: https://lore.kernel.org/rcu/20260805122346.269445-1-puranjay@kernel.org/
Changes in v4:
- rcu_barrier() and srcu_barrier() no longer skip a CPU whose deferred
  list looks empty.  The lockless llist_empty() test added in v3 skipped
  the lock that makes a drain conclusive: a concurrent drainer can already
  have emptied the list without having re-issued yet, so the barrier could
  return before those callbacks reached a callback list.
- Tree SRCU: drain before cleanup_srcu_struct()'s "just leak it" early
  returns, and in srcu_module_going() before it frees any ->sda.  Deferral
  skips check_init_srcu_struct(), so srcu_module_going() could skip
  cleanup_srcu_struct() and free ->sda with a staged srcu_data still
  chained on a per-CPU list.
- Dropped ->defer_exp: a deferred expedited call_srcu() now completes as a
  normal grace period.  The flag lived on the srcu_data rather than on the
  callback, so it expedited whole batches; only srcu_expedite_current()
  with interrupts already disabled can reach this, so the downgrade costs
  a delay and nothing else.
- rcutorture: use a fixed perf sample period rather than .freq, which sets
  TICK_DEP_BIT_PERF_EVENTS for every CPU and pinned the tick for whole
  runs of TREE04 and TREE07.  Also document nmi_calls, and say on the
  console when there is no PMU, since the issued==invoked check then
  compares zero to zero.
- kasan_record_aux_stack() moved to __call_rcu_common(), so a
  use-after-free report names the caller rather than the irq_work.
- Known gap, documented in patch 1: a callback deferred past the
  CPUHP_AP_SMPCFD_DYING flush leaves ->defer_work claimed with its
  self-IPI lost.  rcutree_migrate_callbacks() still re-issues the
  callback, but the first deferral after that CPU comes back raises no IPI
  and waits for the next irq_work there, or for rcu_barrier().

v2: https://lore.kernel.org/rcu/20260803134839.2103051-1-puranjay@kernel.org/
Changes in v3:
- Barriers no longer call irq_work_sync() on an online CPU's ->defer_work.
  irq_work_sync() waits on an rcuwait, which holds exactly one task, so two
  concurrent rcu_barrier()s syncing the same per-CPU irq_work could lose a
  wakeup and hang.  Both flushes now drain every CPU directly, guarded by a
  lockless llist_empty() test so the no-deferrals case stays cheap.
- The re-entry guard is now set only by the irq_work drain, which always
  runs on the CPU owning the list it drains.  In v2 a barrier draining a
  remote CPU set the flag on the draining CPU, so an unrelated irqs-off
  call_rcu() there was dropped and leaked even though it could not have fed
  the drain.
- The drain clears ->next before re-issuing.  A double call_rcu() on a head
  that is already debug-object-active makes llist_add() self-link it, and
  rcu_do_enqueue()'s double-free path returns without clearing ->next, so
  the drain span looped forever with interrupts disabled.
- An expedited call_srcu() is no longer silently downgraded: ->defer_exp
  records it per srcu_data and the batch is re-issued expedited.  Only
  srcu_expedite_current() is affected, __synchronize_srcu() sleeps and so
  is never deferred.
- The drop is now WARN_ONCE() under CONFIG_PROVE_RCU rather than an
  unconditional WARN, so instrumentation cannot reboot a panic_on_warn
  kernel; the backtrace is what identifies the offending program.
- Tiny RCU and Tiny SRCU use IRQ_WORK_INIT_HARD like the Tree flavors, and
  READ_ONCE()/WRITE_ONCE() on their draining flags.  TINY_SRCU is
  "default y if !SMP" with no PREEMPT_RT dependency, so it really can be
  built on RT where a non-HARD irq_work waits on the irq_workd kthread.
- Tiny SRCU: cleanup_srcu_struct() drains *and* irq_work_sync()s
  ->defer_iw.  That irq_work is embedded in the srcu_struct the caller is
  about to free, unlike the Tree flavors' static per-CPU ones.
- rcutorture: drive the perf counters from CPU-hotplug callbacks.  The
  one-shot for_each_online_cpu() loop lost them at the first CPU offline,
  after which the end-of-test issued==invoked check compared 0 == 0.
- rcutorture: per-CPU rcu_head instead of one global, so several CPUs can
  race the drain; count the call before issuing it so mid-run stats cannot
  show nmi-cbs > nmi-calls; release the perf events on the
  torture_cleanup_begin() early-return path; print nmi_calls in the module
  banner; report when nmi_calls is set but no NMI ->call() ever happened;
  and drop sample_freq to 100, since 1000 made perf lower the system-wide
  perf_event_max_sample_rate tenfold.
- selftests/bpf: the old ASSERT_EQ(reentered, 1) could not fail, and an
  atomic allocation failure in the nested task-storage delete made the test
  pass without ever re-entering call_srcu().  It now records and asserts
  the helper return values, probes for rcu_segcblist_enqueue() up front so
  a Tiny kernel skips instead of failing, and restores the CPU affinity it
  changes.

Testing: rcutorture rcu/srcu/srcud/tasks-tracing with nmi_calls, hotplug,
barriers and nocb toggling issued ~81000 callbacks from NMI with none lost,
under PROVE_LOCKING, PROVE_RAW_LOCK_NESTING, DEBUG_OBJECTS_RCU_HEAD and
RCU_LAZY, with no lockdep reports; SRCU-T, SRCU-U, TINY01 and TINY02 pass;
the BPF reproducer passes.

v1: https://lore.kernel.org/all/20260729162207.1567770-1-puranjay@kernel.org/
Changes in v2:
- Fixed the re-entry livelock Zqiang spotted: a BPF program on the enqueue
  path re-enters call_srcu() from inside srcu_defer_drain(), stages another
  callback and re-raises the irq_work, so the drain never finishes.  A
  per-CPU flag now drops such a deferral, with a warning, unless it comes
  from an NMI.
- cleanup_srcu_struct(): drain the deferred callbacks before syncing
  ->irq_work rather than after, since re-issuing one can start a grace
  period and re-queue that irq_work (Zqiang).
- Tiny SRCU: sync ->defer_iw in cleanup_srcu_struct() as well, so a
  deferred callback is re-issued onto ->srcu_cb_head where the leak checks
  can see it instead of being stranded on a soon-to-be-freed srcu_struct.
- Added Kumar's ack to the BPF selftest patch.

Testing: rcutorture rcu, srcu, srcud and tasks-tracing, each with
nmi_calls, CPU hotplug and barriers, and nocb toggling for rcu, all
End of test: SUCCESS with the count issued from NMI equal to the count
invoked (17987, 2924, 2661 and 4412 respectively) and clean consoles,
under PROVE_LOCKING, PROVE_RAW_LOCK_NESTING and DEBUG_OBJECTS_RCU_HEAD;
SRCU-T, SRCU-U, TINY01 and TINY02 pass; the BPF reproducer passes.
Build-tested x86_64 with TREE_RCU+TREE_SRCU, TINY_RCU+TINY_SRCU and
TREE_RCU+TINY_SRCU, each with CONFIG_RCU_DEFER=y and =n, plus arm64
TREE_RCU+TREE_SRCU with RCU_NOCB_CPU and RCU_LAZY.

Puranjay Mohan (6):
  rcu: Make call_rcu() safe to call from any context
  rcu: Make Tiny call_rcu() safe to call from any context
  srcu: Make call_srcu() safe to call from any context
  srcu: Make Tiny call_srcu() safe to call from any context
  rcutorture: Exercise ->call() from NMI context
  selftests/bpf: Add a call_srcu() re-entry reproducer

 .../admin-guide/kernel-parameters.txt         |   7 +
 include/linux/srcutiny.h                      |  12 +-
 include/linux/srcutree.h                      |   4 +
 kernel/rcu/Kconfig                            |   6 +
 kernel/rcu/rcu.h                              |  14 ++
 kernel/rcu/rcutorture.c                       | 150 +++++++++++++++-
 kernel/rcu/srcutiny.c                         |  93 +++++++++-
 kernel/rcu/srcutree.c                         | 170 +++++++++++++++++-
 kernel/rcu/tiny.c                             | 127 ++++++++++---
 kernel/rcu/tree.c                             | 133 ++++++++++++--
 kernel/rcu/tree.h                             |   6 +
 .../selftests/bpf/prog_tests/rcu_reentry.c    |  93 ++++++++++
 .../testing/selftests/bpf/progs/rcu_reentry.c |  51 ++++++
 13 files changed, 818 insertions(+), 48 deletions(-)
 create mode 100644 tools/testing/selftests/bpf/prog_tests/rcu_reentry.c
 create mode 100644 tools/testing/selftests/bpf/progs/rcu_reentry.c

base-commit: 9dc303e69bcd49f9668ca090ae45325269531fbb
-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 11+ messages in thread

* [PATCH v4 1/6] rcu: Make call_rcu() safe to call from any context
  2026-08-10 12:27 [PATCH v4 0/6] rcu,srcu: Make call_rcu()/call_srcu() safe from any context Puranjay Mohan
@ 2026-08-10 12:27 ` Puranjay Mohan
  2026-08-10 12:43   ` sashiko-bot
  2026-08-10 12:27 ` [PATCH v4 2/6] rcu: Make Tiny " Puranjay Mohan
                   ` (4 subsequent siblings)
  5 siblings, 1 reply; 11+ messages in thread
From: Puranjay Mohan @ 2026-08-10 12:27 UTC (permalink / raw)
  To: Lai Jiangshan, Paul E. McKenney, Josh Triplett, Onur Özkan,
	Frederic Weisbecker, Neeraj Upadhyay, Joel Fernandes, Boqun Feng,
	Uladzislau Rezki, Davidlohr Bueso, Andrii Nakryiko,
	Eduard Zingerman, Alexei Starovoitov, Daniel Borkmann,
	Kumar Kartikeya Dwivedi
  Cc: Puranjay Mohan, Steven Rostedt, Mathieu Desnoyers, Zqiang,
	Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
	Emil Tsalapatis, Matt Fleming, Harry Yoo (Oracle), linux-kernel,
	rcu, bpf, linux-rt-devel

RCU's per-CPU callback list is only touched with interrupts disabled: the
enqueue runs under local_irq_save() (and the nocb locks when offloaded),
as do callback invocation and grace-period work.  A call_rcu() that
arrives with interrupts already disabled, whether from an NMI or from
instrumentation that re-enters RCU, can interrupt one of those and corrupt
the list or deadlock.

Defer instead: stage the callback on a per-CPU llist and raise an irq_work
that re-issues it once interrupts are on, straight to the enqueue so it
cannot defer again.  The gate is bare irqs_disabled(), so callers that
merely hold interrupts off are deferred too and pay one irq_work hop.
Skip it while the scheduler is down (RCU_SCHEDULER_INACTIVE): irq_work is
not usable that early, rcu_init() already calls call_rcu(), and the per-CPU
deferral state is not initialised until rcu_init_one() runs later in it.

rcu_barrier() drains every CPU's ->defer_head before it scans the lists,
and rcutree_migrate_callbacks() drains an outgoing CPU's.  A drain
re-issues onto the draining CPU, so a barrier moves other CPUs' staged
callbacks onto
its own ->cblist; call_rcu() promises no CPU affinity for invocation.
->defer_lock is held across llist_del_all() and the whole re-issue so the
drainers
serialize: one that finds the list empty can conclude that everything
staged before it is already on a callback list.  Interrupts stay off for
the batch.  Where the arch has an irq_work self-IPI that is what one
interrupts-disabled region could stage, normally a single callback; where
arch_irq_work_has_interrupt() is false the drain waits for the tick, so
several regions can accumulate first.

The drain clears ->next before re-issuing.  A double call_rcu() on a head
that is already debug-object-active self-links the staged node, and
rcu_do_enqueue()'s duplicate path returns without clearing it, so the
drain would spin.  A re-add behind other staged callbacks makes a longer
cycle, which that does not bound; a double call_rcu() stays undefined.
llist_del_all() yields newest-first, so a batch is re-issued in reverse
call order; nothing depends on call_rcu() ordering.  The re-issue drops
the lazy hint, since staging records only ->func, so a deferred callback
loses its batching on CONFIG_RCU_LAZY.  kasan_record_aux_stack() moves to
__call_rcu_common() so a use-after-free report names the caller rather
than the irq_work.

The re-issue runs with interrupts disabled, so instrumentation on the
enqueue path can re-enter call_rcu(), stage another callback and re-raise
the irq_work, livelocking the drain.  A per-CPU flag guards it: a deferral
that arrives while this CPU is draining, and is not from an NMI, is
dropped.  The WARN_ONCE() is under CONFIG_PROVE_RCU, so a production kernel
drops it silently.  That leaks the callback and can strand state
the caller tied to it, since a one-shot flag only the callback clears
never resets, but the alternative is an unbounded loop.

A callback deferred past the CPUHP_AP_SMPCFD_DYING irq_work flush leaves
->defer_work claimed with its self-IPI lost.  rcutree_migrate_callbacks()
still re-issues the callback, but the first deferral after that CPU comes
back raises no IPI and waits for the next irq_work there, or for
rcu_barrier().  Unqueueing an irq_work is not something the API offers.

The irq_work is IRQ_WORK_INIT_HARD so the re-issue stays prompt on
PREEMPT_RT, where a non-HARD irq_work runs in a kthread that can be
delayed under load.  A hidden CONFIG_RCU_DEFER gates the deferral code and
its IRQ_WORK dependency, though the rcu_data members are unconditional;
without it call_rcu() enqueues directly as before.  Under
CONFIG_PROVE_RCU, warn if the direct path is reached from an NMI.

Suggested-by: Paul E. McKenney <paulmck@kernel.org>
Signed-off-by: Puranjay Mohan <puranjay@kernel.org>
---
 kernel/rcu/Kconfig |   6 +++
 kernel/rcu/rcu.h   |  11 ++++
 kernel/rcu/tree.c  | 131 +++++++++++++++++++++++++++++++++++++++++----
 kernel/rcu/tree.h  |   6 +++
 4 files changed, 143 insertions(+), 11 deletions(-)

diff --git a/kernel/rcu/Kconfig b/kernel/rcu/Kconfig
index f15da8038d0ba..1a5fb3156c062 100644
--- a/kernel/rcu/Kconfig
+++ b/kernel/rcu/Kconfig
@@ -175,6 +175,12 @@ config RCU_STALL_COMMON
 config RCU_NEED_SEGCBLIST
 	def_bool ( TREE_RCU || TREE_SRCU || TASKS_RCU_GENERIC )
 
+# The deferral (and the IRQ_WORK it uses) is only needed where call_rcu() /
+# call_srcu() can be invoked while a callback-list operation is in flight.
+config RCU_DEFER
+	def_bool HAVE_NMI || KPROBES || FUNCTION_TRACER || TRACEPOINTS
+	select IRQ_WORK
+
 config RCU_FANOUT
 	int "Tree-based hierarchical RCU fanout value"
 	range 2 64 if 64BIT
diff --git a/kernel/rcu/rcu.h b/kernel/rcu/rcu.h
index 39a9f6fa9a7b2..91e33571a554d 100644
--- a/kernel/rcu/rcu.h
+++ b/kernel/rcu/rcu.h
@@ -572,6 +572,17 @@ static inline void tasks_cblist_init_generic(void) { }
 #define RCU_SCHEDULER_INIT	1
 #define RCU_SCHEDULER_RUNNING	2
 
+/*
+ * Defer whenever interrupts are disabled, since a callback-list operation may
+ * be in flight on this CPU.  Not before the scheduler is up: irq_work is not
+ * usable that early, and rcu_init() itself calls call_rcu().
+ */
+static inline bool should_rcu_defer(void)
+{
+	return IS_ENABLED(CONFIG_RCU_DEFER) && irqs_disabled() &&
+	       rcu_scheduler_active != RCU_SCHEDULER_INACTIVE;
+}
+
 enum rcutorture_type {
 	RCU_FLAVOR,
 	RCU_TASKS_FLAVOR,
diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c
index 21b6ce1dffb63..3bf3a250f9de8 100644
--- a/kernel/rcu/tree.c
+++ b/kernel/rcu/tree.c
@@ -24,6 +24,7 @@
 #include <linux/smp.h>
 #include <linux/rcupdate_wait.h>
 #include <linux/interrupt.h>
+#include <linux/llist.h>
 #include <linux/sched.h>
 #include <linux/sched/debug.h>
 #include <linux/nmi.h>
@@ -3148,21 +3149,19 @@ static void check_cb_ovld(struct rcu_data *rdp)
 	raw_spin_unlock_rcu_node(rnp);
 }
 
-static void
-__call_rcu_common(struct rcu_head *head, rcu_callback_t func, bool lazy_in)
+/*
+ * Also called by __rcu_defer_drain() to re-issue a deferred callback, so it
+ * must not re-check the deferral condition.  Either caller may have interrupts
+ * already disabled, and a drain of a remote CPU re-issues onto the draining
+ * CPU.
+ */
+static void rcu_do_enqueue(struct rcu_head *head, rcu_callback_t func, bool lazy_in)
 {
 	static atomic_t doublefrees;
 	unsigned long flags;
 	bool lazy;
 	struct rcu_data *rdp;
 
-	/* Misaligned rcu_head! */
-	WARN_ON_ONCE((unsigned long)head & (sizeof(void *) - 1));
-
-	/* Avoid NULL dereference if callback is NULL. */
-	if (WARN_ON_ONCE(!func))
-		return;
-
 	if (debug_rcu_head_queue(head)) {
 		/*
 		 * Probable double call_rcu(), so leak the callback.
@@ -3178,7 +3177,6 @@ __call_rcu_common(struct rcu_head *head, rcu_callback_t func, bool lazy_in)
 	}
 	head->func = func;
 	head->next = NULL;
-	kasan_record_aux_stack(head);
 
 	local_irq_save(flags);
 	rdp = this_cpu_ptr(&rcu_data);
@@ -3206,6 +3204,103 @@ __call_rcu_common(struct rcu_head *head, rcu_callback_t func, bool lazy_in)
 	local_irq_restore(flags);
 }
 
+/*
+ * Re-issue deferred callbacks straight to the enqueue so they cannot defer
+ * again.  ->defer_lock serializes the drainers: this CPU's irq_work,
+ * rcu_defer_flush() and rcutree_migrate_callbacks().
+ */
+static void __rcu_defer_drain(struct rcu_data *rdp)
+{
+	struct llist_node *node, *next;
+	unsigned long flags;
+
+	if (!IS_ENABLED(CONFIG_RCU_DEFER))
+		return;
+
+	raw_spin_lock_irqsave(&rdp->defer_lock, flags);
+	llist_for_each_safe(node, next, llist_del_all(&rdp->defer_head)) {
+		struct rcu_head *head = (struct rcu_head *)node;
+
+		/* Bounds a node self-linked by a double call_rcu(). */
+		head->next = NULL;
+		rcu_do_enqueue(head, head->func, false);
+	}
+	raw_spin_unlock_irqrestore(&rdp->defer_lock, flags);
+}
+
+/*
+ * Only the irq_work drain can be re-fed by its own re-issue, so only it sets
+ * ->defer_draining.  Anything staged during a direct drain is picked up by the
+ * staging CPU's own irq_work.  Every caller of irq_work_run_list() has
+ * interrupts disabled, so the flag is never visible with them enabled.
+ */
+static void rcu_defer_drain(struct irq_work *iw)
+{
+	struct rcu_data *rdp = container_of(iw, struct rcu_data, defer_work);
+
+	WRITE_ONCE(rdp->defer_draining, true);
+	__rcu_defer_drain(rdp);
+	WRITE_ONCE(rdp->defer_draining, false);
+}
+
+/*
+ * Stage @head for this CPU's irq_work to re-issue once interrupts are on.  Only
+ * the drain side takes a lock, so this stays safe from NMI.
+ */
+static void call_rcu_defer(struct rcu_head *head, rcu_callback_t func)
+{
+	struct rcu_data *rdp = this_cpu_ptr(&rcu_data);
+
+	/*
+	 * Instrumentation on the enqueue path can re-enter here from inside the
+	 * drain.  Re-queuing would livelock it, so drop the callback; an NMI
+	 * cannot loop, so let it through.
+	 */
+	if (READ_ONCE(rdp->defer_draining) && !in_nmi()) {
+		WARN_ONCE(IS_ENABLED(CONFIG_PROVE_RCU),
+			  "call_rcu() re-entered during callback drain; leaking callback\n");
+		return;
+	}
+	head->func = func;
+	if (llist_add((struct llist_node *)head, &rdp->defer_head))
+		irq_work_queue(&rdp->defer_work);
+}
+
+static void rcu_defer_flush(void)
+{
+	int cpu;
+
+	for_each_possible_cpu(cpu)
+		__rcu_defer_drain(per_cpu_ptr(&rcu_data, cpu));
+}
+
+static void
+__call_rcu_common(struct rcu_head *head, rcu_callback_t func, bool lazy_in)
+{
+	/* Misaligned rcu_head! */
+	WARN_ON_ONCE((unsigned long)head & (sizeof(void *) - 1));
+
+	/* Avoid NULL dereference if callback is NULL. */
+	if (WARN_ON_ONCE(!func))
+		return;
+
+	/* Record the caller: the irq_work's stack says nothing about it. */
+	kasan_record_aux_stack(head);
+
+	if (should_rcu_defer()) {
+		call_rcu_defer(head, func);
+		return;
+	}
+
+	/*
+	 * Only reachable from an NMI when deferral is off: before the scheduler
+	 * is up, or with CONFIG_RCU_DEFER=n.  The enqueue can then race.
+	 */
+	WARN_ON_ONCE(IS_ENABLED(CONFIG_PROVE_RCU) && in_nmi());
+
+	rcu_do_enqueue(head, func, lazy_in);
+}
+
 #ifdef CONFIG_RCU_LAZY
 static bool enable_rcu_lazy __read_mostly = !IS_ENABLED(CONFIG_RCU_LAZY_DEFAULT_OFF);
 module_param(enable_rcu_lazy, bool, 0444);
@@ -3896,8 +3991,12 @@ void rcu_barrier(void)
 	unsigned long flags;
 	unsigned long gseq;
 	struct rcu_data *rdp;
-	unsigned long s = rcu_seq_snap(&rcu_state.barrier_sequence);
+	unsigned long s;
 
+	/* Register any deferred callbacks before snapshotting the sequence. */
+	rcu_defer_flush();
+
+	s = rcu_seq_snap(&rcu_state.barrier_sequence);
 	rcu_barrier_trace(TPS("Begin"), -1, s);
 
 	/* Take mutex to serialize concurrent rcu_barrier() requests. */
@@ -4231,6 +4330,9 @@ rcu_boot_init_percpu_data(int cpu)
 	rdp->rcu_onl_gp_state = RCU_GP_CLEANED;
 	rdp->last_sched_clock = jiffies;
 	rdp->cpu = cpu;
+	init_llist_head(&rdp->defer_head);
+	raw_spin_lock_init(&rdp->defer_lock);
+	rdp->defer_work = IRQ_WORK_INIT_HARD(rcu_defer_drain);
 	rcu_boot_init_nocb_percpu_data(rdp);
 }
 
@@ -4528,6 +4630,13 @@ void rcutree_migrate_callbacks(int cpu)
 	struct rcu_data *rdp = per_cpu_ptr(&rcu_data, cpu);
 	bool needwake;
 
+	/*
+	 * Callbacks deferred past the point the outgoing CPU's irq_work can run
+	 * sit on ->defer_head, which the ->cblist migration below does not
+	 * cover.  Drain them here, before the early returns.
+	 */
+	__rcu_defer_drain(rdp);
+
 	if (rcu_rdp_is_offloaded(rdp))
 		return;
 
diff --git a/kernel/rcu/tree.h b/kernel/rcu/tree.h
index eedfa43059e80..b7cac7a13b4f2 100644
--- a/kernel/rcu/tree.h
+++ b/kernel/rcu/tree.h
@@ -229,6 +229,12 @@ struct rcu_data {
 	struct rcu_head barrier_head;
 	int exp_watching_snap;		/* Double-check need for IPI. */
 
+	/* Deferral of an NMI/reentrant call_rcu(); see __call_rcu_common(). */
+	struct llist_head defer_head;
+	struct irq_work defer_work;
+	raw_spinlock_t defer_lock;
+	bool defer_draining;
+
 	/* 5) Callback offloading. */
 #ifdef CONFIG_RCU_NOCB_CPU
 	struct swait_queue_head nocb_cb_wq; /* For nocb kthreads to sleep on. */
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* [PATCH v4 2/6] rcu: Make Tiny call_rcu() safe to call from any context
  2026-08-10 12:27 [PATCH v4 0/6] rcu,srcu: Make call_rcu()/call_srcu() safe from any context Puranjay Mohan
  2026-08-10 12:27 ` [PATCH v4 1/6] rcu: Make call_rcu() safe to call " Puranjay Mohan
@ 2026-08-10 12:27 ` Puranjay Mohan
  2026-08-10 12:42   ` sashiko-bot
  2026-08-10 12:27 ` [PATCH v4 3/6] srcu: Make call_srcu() " Puranjay Mohan
                   ` (3 subsequent siblings)
  5 siblings, 1 reply; 11+ messages in thread
From: Puranjay Mohan @ 2026-08-10 12:27 UTC (permalink / raw)
  To: Lai Jiangshan, Paul E. McKenney, Josh Triplett, Onur Özkan,
	Frederic Weisbecker, Neeraj Upadhyay, Joel Fernandes, Boqun Feng,
	Uladzislau Rezki, Davidlohr Bueso, Andrii Nakryiko,
	Eduard Zingerman, Alexei Starovoitov, Daniel Borkmann,
	Kumar Kartikeya Dwivedi
  Cc: Puranjay Mohan, Steven Rostedt, Mathieu Desnoyers, Zqiang,
	Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
	Emil Tsalapatis, Matt Fleming, Harry Yoo (Oracle), linux-kernel,
	rcu, bpf, linux-rt-devel

Give Tiny call_rcu() the same treatment as Tree RCU.  When interrupts are
disabled and the scheduler is up, stage the callback on a lockless list
that an irq_work re-issues later.  One global list and irq_work suffice
since Tiny RCU is uniprocessor, and there is no CPU-offline drain.

The re-issue runs with interrupts disabled and can be re-entered by
instrumentation, so a draining flag drops a deferring call_rcu() seen
mid-drain (unless from an NMI), as in Tree RCU.  Gated by CONFIG_RCU_DEFER,
though the deferral state is unconditional.

Interrupts stay off for the whole batch, but the re-issue is a tail append
with no locks.  TINY_RCU implies !SMP, where arch_irq_work_has_interrupt()
is false, so the drain always waits for the tick and a batch is whatever
one tick's worth of interrupts-disabled call_rcu()s staged.  As in Tree
RCU the drain clears ->next before re-issuing, which bounds a node
self-linked by a double call_rcu(): rcu_do_enqueue()'s duplicate path
returns without clearing it.  A longer cycle is not bounded; a double
call_rcu() stays undefined.

The idle-task reschedule moves out of the enqueue helper so that a drain
does it once for the batch rather than once per callback, which would
otherwise take the runqueue lock N times with interrupts disabled.

Suggested-by: Paul E. McKenney <paulmck@kernel.org>
Signed-off-by: Puranjay Mohan <puranjay@kernel.org>
---
 kernel/rcu/tiny.c | 127 +++++++++++++++++++++++++++++++++++++---------
 1 file changed, 104 insertions(+), 23 deletions(-)

diff --git a/kernel/rcu/tiny.c b/kernel/rcu/tiny.c
index dccccd6be9411..656b6a682e31a 100644
--- a/kernel/rcu/tiny.c
+++ b/kernel/rcu/tiny.c
@@ -11,6 +11,8 @@
  */
 #include <linux/completion.h>
 #include <linux/interrupt.h>
+#include <linux/irq_work.h>
+#include <linux/llist.h>
 #include <linux/notifier.h>
 #include <linux/rcupdate_wait.h>
 #include <linux/kernel.h>
@@ -42,8 +44,100 @@ static struct rcu_ctrlblk rcu_ctrlblk = {
 	.gp_seq		= 0 - 300UL,
 };
 
+/*
+ * The callback list is only accessed with interrupts disabled, so a call_rcu()
+ * that arrives with interrupts off stages the callback on a lockless list that
+ * an irq_work re-issues later.  One global list and irq_work suffice, as Tiny
+ * RCU is uniprocessor.
+ */
+static void rcu_defer_drain(struct irq_work *iw);
+static LLIST_HEAD(rcu_defer_list);
+static struct irq_work rcu_defer_iw = IRQ_WORK_INIT_HARD(rcu_defer_drain);
+static bool rcu_defer_draining;
+
+/*
+ * Also called by __rcu_defer_drain() to re-issue a deferred callback, so it
+ * must not re-check the deferral condition.
+ */
+static void rcu_do_enqueue(struct rcu_head *head, rcu_callback_t func)
+{
+	static atomic_t doublefrees;
+	unsigned long flags;
+
+	if (debug_rcu_head_queue(head)) {
+		if (atomic_inc_return(&doublefrees) < 4) {
+			pr_err("%s(): Double-freed CB %p->%pS()!!!  ", __func__, head, head->func);
+			mem_dump_obj(head);
+		}
+		return;
+	}
+
+	head->func = func;
+	head->next = NULL;
+
+	local_irq_save(flags);
+	*rcu_ctrlblk.curtail = head;
+	rcu_ctrlblk.curtail = &head->next;
+	local_irq_restore(flags);
+}
+
+/* Force scheduling for rcu_qs() when enqueuing from the idle task. */
+static void rcu_resched_if_idle(void)
+{
+	if (unlikely(is_idle_task(current)))
+		resched_cpu(0);
+}
+
+static void __rcu_defer_drain(void)
+{
+	struct llist_node *node, *next;
+	bool drained = false;
+	unsigned long flags;
+
+	if (!IS_ENABLED(CONFIG_RCU_DEFER))
+		return;
+
+	/* Re-issued newest-first; nothing depends on call_rcu() ordering. */
+	local_irq_save(flags);
+	llist_for_each_safe(node, next, llist_del_all(&rcu_defer_list)) {
+		struct rcu_head *head = (struct rcu_head *)node;
+
+		/* Bounds a node self-linked by a double call_rcu(). */
+		head->next = NULL;
+		rcu_do_enqueue(head, head->func);
+		drained = true;
+	}
+	local_irq_restore(flags);
+
+	if (drained)
+		rcu_resched_if_idle();
+}
+
+/* Only the irq_work drain can be re-fed by its own re-issue; see Tree RCU. */
+static void rcu_defer_drain(struct irq_work *iw)
+{
+	WRITE_ONCE(rcu_defer_draining, true);
+	__rcu_defer_drain();
+	WRITE_ONCE(rcu_defer_draining, false);
+}
+
+static void call_rcu_defer(struct rcu_head *head, rcu_callback_t func)
+{
+	/* A re-entrant call_rcu() during the drain would livelock it; drop it. */
+	if (READ_ONCE(rcu_defer_draining) && !in_nmi()) {
+		WARN_ONCE(IS_ENABLED(CONFIG_PROVE_RCU),
+			  "call_rcu() re-entered during callback drain; leaking callback\n");
+		return;
+	}
+	head->func = func;
+	if (llist_add((struct llist_node *)head, &rcu_defer_list))
+		irq_work_queue(&rcu_defer_iw);
+}
+
 void rcu_barrier(void)
 {
+	/* Register any deferred callbacks so the wait below covers them. */
+	__rcu_defer_drain();
 	wait_rcu_gp(call_rcu_hurry);
 }
 EXPORT_SYMBOL(rcu_barrier);
@@ -157,29 +251,19 @@ EXPORT_SYMBOL_GPL(synchronize_rcu);
  */
 void call_rcu(struct rcu_head *head, rcu_callback_t func)
 {
-	static atomic_t doublefrees;
-	unsigned long flags;
-
-	if (debug_rcu_head_queue(head)) {
-		if (atomic_inc_return(&doublefrees) < 4) {
-			pr_err("%s(): Double-freed CB %p->%pS()!!!  ", __func__, head, head->func);
-			mem_dump_obj(head);
-		}
+	if (should_rcu_defer()) {
+		call_rcu_defer(head, func);
 		return;
 	}
 
-	head->func = func;
-	head->next = NULL;
+	/*
+	 * Only reachable from an NMI when deferral is off: before the scheduler
+	 * is up, or with CONFIG_RCU_DEFER=n.  The enqueue can then race.
+	 */
+	WARN_ON_ONCE(IS_ENABLED(CONFIG_PROVE_RCU) && in_nmi());
 
-	local_irq_save(flags);
-	*rcu_ctrlblk.curtail = head;
-	rcu_ctrlblk.curtail = &head->next;
-	local_irq_restore(flags);
-
-	if (unlikely(is_idle_task(current))) {
-		/* force scheduling for rcu_qs() */
-		resched_cpu(0);
-	}
+	rcu_do_enqueue(head, func);
+	rcu_resched_if_idle();
 }
 EXPORT_SYMBOL_GPL(call_rcu);
 
@@ -211,10 +295,7 @@ unsigned long start_poll_synchronize_rcu(void)
 {
 	unsigned long gp_seq = get_state_synchronize_rcu();
 
-	if (unlikely(is_idle_task(current))) {
-		/* force scheduling for rcu_qs() */
-		resched_cpu(0);
-	}
+	rcu_resched_if_idle();
 	return gp_seq;
 }
 EXPORT_SYMBOL_GPL(start_poll_synchronize_rcu);
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* [PATCH v4 3/6] srcu: Make call_srcu() safe to call from any context
  2026-08-10 12:27 [PATCH v4 0/6] rcu,srcu: Make call_rcu()/call_srcu() safe from any context Puranjay Mohan
  2026-08-10 12:27 ` [PATCH v4 1/6] rcu: Make call_rcu() safe to call " Puranjay Mohan
  2026-08-10 12:27 ` [PATCH v4 2/6] rcu: Make Tiny " Puranjay Mohan
@ 2026-08-10 12:27 ` Puranjay Mohan
  2026-08-10 12:27 ` [PATCH v4 4/6] srcu: Make Tiny " Puranjay Mohan
                   ` (2 subsequent siblings)
  5 siblings, 0 replies; 11+ messages in thread
From: Puranjay Mohan @ 2026-08-10 12:27 UTC (permalink / raw)
  To: Lai Jiangshan, Paul E. McKenney, Josh Triplett, Onur Özkan,
	Frederic Weisbecker, Neeraj Upadhyay, Joel Fernandes, Boqun Feng,
	Uladzislau Rezki, Davidlohr Bueso, Andrii Nakryiko,
	Eduard Zingerman, Alexei Starovoitov, Daniel Borkmann,
	Kumar Kartikeya Dwivedi
  Cc: Puranjay Mohan, Steven Rostedt, Mathieu Desnoyers, Zqiang,
	Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
	Emil Tsalapatis, Matt Fleming, Harry Yoo (Oracle), linux-kernel,
	rcu, bpf, linux-rt-devel

call_srcu() has the same constraint as call_rcu():
srcu_gp_start_if_needed() enqueues under raw_spin_lock_irqsave() and may
walk the srcu_node tree, as do callback invocation and grace-period work,
so a call_srcu() with interrupts already disabled can race an operation in
flight on this CPU.  call_rcu_tasks_trace() is call_srcu() under the hood,
so a sleepable BPF program freeing an object can reach this.

Defer as call_rcu() does: stage the callback on the srcu_data's
->defer_cbs, chain that srcu_data onto a per-CPU list, and raise a per-CPU
irq_work that re-issues it straight to the enqueue helper, never back
through __call_srcu().  The irqs-enabled path is unchanged; as for
call_rcu() the gate is bare irqs_disabled(), so callers that merely hold
interrupts off are deferred too and pay one irq_work hop, including
call_rcu_tasks_trace() from the BPF memalloc irq_work.

The irq_work is per-CPU rather than per-srcu_struct and statically
initialized, so deferral never runs check_init_srcu_struct(); it is
IRQ_WORK_INIT_HARD as for call_rcu().  srcu_barrier() flushes it first, and
rcutree_migrate_callbacks() calls srcu_offline_drain() for an outgoing CPU.
cleanup_srcu_struct() drains before its "just leak it" early returns, and
srcu_module_going() before freeing any ->sda, since a staged srcu_data left
chained on a per-CPU list would dangle.  Staging is two steps, the callback
onto
->defer_cbs and then the srcu_data onto the per-CPU list, so a flusher can
find the per-CPU list empty while a callback whose call_srcu() has not
returned sits on ->defer_cbs; the staging CPU's own irq_work takes that
one.

The per-CPU srcu_defer ->lock, not any srcu_data's, is held with interrupts
off across the whole nested drain: the chain of srcu_datas staged on that
CPU and, for each, its callbacks, with srcu_do_enqueue() taking that
srcu_data's ->lock and possibly starting a grace period for every one.
That is what serializes the drainers.  The bound is as for call_rcu(): what
one interrupts-disabled region could stage, normally a single callback, or
whatever accumulates before the tick where arch_irq_work_has_interrupt() is
false.  Both lists are drained newest-first; nothing depends on call_srcu()
ordering.

The drain clears ->next before re-issuing, which bounds a node self-linked
by a double call_srcu(); a longer cycle is not bounded, and a double
call_srcu() stays undefined, as for call_rcu().

A callback deferred past the CPUHP_AP_SMPCFD_DYING irq_work flush leaves
that CPU's srcu_defer ->iw claimed with its self-IPI lost, as for
call_rcu().  srcu_offline_drain() still re-issues the callback, but the
irq_work cannot be un-queued, and here the claim is shared by every
srcu_struct on the CPU.

As in call_rcu(), the re-issue runs with interrupts disabled and can be
re-entered by instrumentation, so a per-CPU flag, set only while that CPU
is inside its own irq_work drain, drops a deferring call_srcu() seen
mid-drain unless it comes from an NMI.  Such a drop can strand state the
caller associated with the callback, not just the callback itself.

Staging records only the callback, so a deferred expedited call_srcu()
completes as a normal grace period.  Only srcu_expedite_current() can hit
that, and only when invoked with interrupts already disabled.

Gated by CONFIG_RCU_DEFER, though the srcu_data members and the per-CPU
srcu_defer are unconditional.  Under CONFIG_PROVE_RCU, warn if the direct
path is reached from an NMI.

Suggested-by: Paul E. McKenney <paulmck@kernel.org>
Signed-off-by: Puranjay Mohan <puranjay@kernel.org>
---
 include/linux/srcutree.h |   4 +
 kernel/rcu/rcu.h         |   3 +
 kernel/rcu/srcutree.c    | 170 ++++++++++++++++++++++++++++++++++++++-
 kernel/rcu/tree.c        |   2 +
 4 files changed, 175 insertions(+), 4 deletions(-)

diff --git a/include/linux/srcutree.h b/include/linux/srcutree.h
index 75e54e4f963fa..1ce759fb70948 100644
--- a/include/linux/srcutree.h
+++ b/include/linux/srcutree.h
@@ -13,6 +13,8 @@
 
 #include <linux/rcu_node_tree.h>
 #include <linux/completion.h>
+#include <linux/irq_work_types.h>
+#include <linux/llist.h>
 
 struct srcu_node;
 struct srcu_struct;
@@ -41,6 +43,8 @@ struct srcu_data {
 	bool srcu_cblist_invoking;		/* Invoking these CBs? */
 	struct timer_list delay_work;		/* Delay for CB invoking */
 	struct work_struct work;		/* Context for CB invoking. */
+	struct llist_head defer_cbs;		/* Callbacks deferred on re-entry. */
+	struct llist_node defer_link;		/* Links onto the per-CPU deferral drain list */
 	struct rcu_head srcu_barrier_head;	/* For srcu_barrier() use. */
 	struct rcu_head srcu_ec_head;		/* For srcu_expedite_current() use. */
 	int srcu_ec_state;			/*  State for srcu_expedite_current(). */
diff --git a/kernel/rcu/rcu.h b/kernel/rcu/rcu.h
index 91e33571a554d..d60444bf3a02d 100644
--- a/kernel/rcu/rcu.h
+++ b/kernel/rcu/rcu.h
@@ -583,6 +583,9 @@ static inline bool should_rcu_defer(void)
 	       rcu_scheduler_active != RCU_SCHEDULER_INACTIVE;
 }
 
+/* Drain an outgoing CPU's deferred SRCU callbacks; see rcutree_migrate_callbacks(). */
+void srcu_offline_drain(int cpu);
+
 enum rcutorture_type {
 	RCU_FLAVOR,
 	RCU_TASKS_FLAVOR,
diff --git a/kernel/rcu/srcutree.c b/kernel/rcu/srcutree.c
index 304112674e8a2..10c89f34920c8 100644
--- a/kernel/rcu/srcutree.c
+++ b/kernel/rcu/srcutree.c
@@ -20,6 +20,7 @@
 #include <linux/percpu.h>
 #include <linux/preempt.h>
 #include <linux/irq_work.h>
+#include <linux/llist.h>
 #include <linux/rcupdate_wait.h>
 #include <linux/sched.h>
 #include <linux/smp.h>
@@ -79,6 +80,38 @@ static void process_srcu(struct work_struct *work);
 static void srcu_irq_work(struct irq_work *work);
 static void srcu_delay_timer(struct timer_list *t);
 
+struct srcu_defer;
+static void srcu_defer_drain(struct irq_work *iw);
+static void __srcu_defer_drain(struct srcu_defer *sndp);
+
+/*
+ * Per-CPU call_srcu() deferral state, shared by every srcu_struct.  A deferred
+ * callback is staged on its srcu_data's ->defer_cbs; that srcu_data is chained
+ * via ->defer_link onto ->list, which the irq_work walks.
+ */
+struct srcu_defer {
+	struct llist_head	list;
+	struct irq_work		iw;
+	raw_spinlock_t		lock;
+	bool			draining;
+};
+
+static DEFINE_PER_CPU(struct srcu_defer, srcu_defer) = {
+	.lock = __RAW_SPIN_LOCK_UNLOCKED(srcu_defer.lock),
+	.iw = IRQ_WORK_INIT_HARD(srcu_defer_drain),
+};
+
+/*
+ * Flush pending deferred callbacks so a following srcu_barrier() waits for them.
+ */
+static void srcu_defer_flush(void)
+{
+	int cpu;
+
+	for_each_possible_cpu(cpu)
+		__srcu_defer_drain(&per_cpu(srcu_defer, cpu));
+}
+
 /*
  * Initialize SRCU per-CPU data.  Note that statically allocated
  * srcu_struct structures might already have srcu_read_lock() and
@@ -107,6 +140,11 @@ static void init_srcu_struct_data(struct srcu_struct *ssp)
 		sdp->cpu = cpu;
 		INIT_WORK(&sdp->work, srcu_invoke_callbacks);
 		timer_setup(&sdp->delay_work, srcu_delay_timer, 0);
+		/*
+		 * ->defer_cbs and ->defer_link are valid when zeroed and are not
+		 * reinitialized here: that would clobber callbacks a reentrant
+		 * call_srcu() already staged.  See __call_srcu().
+		 */
 		sdp->ssp = ssp;
 	}
 }
@@ -688,6 +726,14 @@ void cleanup_srcu_struct(struct srcu_struct *ssp)
 	unsigned long delay;
 	struct srcu_usage *sup = ssp->srcu_sup;
 
+	/*
+	 * Drain before the early returns below: they leak the srcu_struct, but
+	 * srcu_module_going() frees ->sda regardless, and a staged srcu_data
+	 * left chained on a per-CPU list would then dangle.  Draining first also
+	 * has to precede the ->irq_work sync, since re-issuing a callback can
+	 * start a grace period and re-queue ->irq_work, which schedules ->work.
+	 */
+	srcu_defer_flush();
 	raw_spin_lock_irq_rcu_node(ssp->srcu_sup);
 	delay = srcu_get_delay(ssp);
 	raw_spin_unlock_irq_rcu_node(ssp->srcu_sup);
@@ -695,7 +741,6 @@ void cleanup_srcu_struct(struct srcu_struct *ssp)
 		return; /* Just leak it! */
 	if (WARN_ON(srcu_readers_active(ssp)))
 		return; /* Just leak it! */
-	/* Wait for irq_work to finish first as it may queue a new work. */
 	irq_work_sync(&sup->irq_work);
 	flush_delayed_work(&sup->work);
 	for_each_possible_cpu(cpu) {
@@ -1410,8 +1455,8 @@ static unsigned long srcu_gp_start_if_needed(struct srcu_struct *ssp,
  * srcu_read_lock(), and srcu_read_unlock() that are all passed the same
  * srcu_struct structure.
  */
-static void __call_srcu(struct srcu_struct *ssp, struct rcu_head *rhp,
-			rcu_callback_t func, bool do_norm)
+static void srcu_do_enqueue(struct srcu_struct *ssp, struct rcu_head *rhp,
+			    rcu_callback_t func, bool do_norm)
 {
 	if (debug_rcu_head_queue(rhp)) {
 		/* Probable double call_srcu(), so leak the callback. */
@@ -1423,6 +1468,108 @@ static void __call_srcu(struct srcu_struct *ssp, struct rcu_head *rhp,
 	(void)srcu_gp_start_if_needed(ssp, rhp, do_norm);
 }
 
+/*
+ * The srcu_cblist and srcu_node tree are only accessed with interrupts
+ * disabled, so defer when interrupts are already off rather than enqueue into
+ * an operation that may be in flight on this CPU.
+ */
+static void __call_srcu(struct srcu_struct *ssp, struct rcu_head *rhp,
+			rcu_callback_t func, bool do_norm)
+{
+	if (should_rcu_defer()) {
+		struct srcu_defer *sndp = this_cpu_ptr(&srcu_defer);
+		struct srcu_data *sdp;
+
+		/*
+		 * Instrumentation on the enqueue path can re-enter here from
+		 * inside the drain.  Re-queuing would livelock it, so drop the
+		 * callback; an NMI cannot loop, so let it through.
+		 */
+		if (READ_ONCE(sndp->draining) && !in_nmi()) {
+			WARN_ONCE(IS_ENABLED(CONFIG_PROVE_RCU),
+				  "call_srcu() re-entered during callback drain; leaking callback\n");
+			return;
+		}
+		sdp = this_cpu_ptr(ssp->sda);
+		rhp->func = func;
+		if (llist_add((struct llist_node *)rhp, &sdp->defer_cbs)) {
+			/*
+			 * Chain this srcu_data for the drain.  ->ssp must be
+			 * published here: deferral skips
+			 * check_init_srcu_struct(), so on a never-initialized
+			 * static srcu_struct the srcu_data are still zeroed and
+			 * the drain would read a NULL ->ssp.
+			 */
+			sdp->ssp = ssp;
+			if (llist_add(&sdp->defer_link, &sndp->list))
+				irq_work_queue(&sndp->iw);
+		}
+		return;
+	}
+
+	/*
+	 * Only reachable from an NMI when deferral is off: before the scheduler
+	 * is up, or with CONFIG_RCU_DEFER=n.  The enqueue can then race.
+	 */
+	WARN_ON_ONCE(IS_ENABLED(CONFIG_PROVE_RCU) && in_nmi());
+
+	srcu_do_enqueue(ssp, rhp, func, do_norm);
+}
+
+/*
+ * Re-issue deferred callbacks straight to srcu_do_enqueue() so they cannot defer
+ * again.  ->lock serializes the drainers: the irq_work, srcu_defer_flush() and
+ * srcu_offline_drain().
+ */
+static void __srcu_defer_drain(struct srcu_defer *sndp)
+{
+	struct llist_node *snode, *snext;
+	unsigned long flags;
+
+	if (!IS_ENABLED(CONFIG_RCU_DEFER))
+		return;
+
+	raw_spin_lock_irqsave(&sndp->lock, flags);
+	llist_for_each_safe(snode, snext, llist_del_all(&sndp->list)) {
+		struct srcu_data *sdp = container_of(snode, struct srcu_data, defer_link);
+		struct srcu_struct *ssp = sdp->ssp;
+		struct llist_node *cnode, *cnext;
+
+		cnode = llist_del_all(&sdp->defer_cbs);
+		llist_for_each_safe(cnode, cnext, cnode) {
+			struct rcu_head *rhp = (struct rcu_head *)cnode;
+
+			/* Bounds a node self-linked by a double call_srcu(). */
+			rhp->next = NULL;
+			srcu_do_enqueue(ssp, rhp, rhp->func, true);
+		}
+	}
+	raw_spin_unlock_irqrestore(&sndp->lock, flags);
+}
+
+/*
+ * Only the irq_work drain can be re-fed by its own re-issue, so only it sets
+ * ->draining.  A direct drain re-issues onto this CPU, and anything staged
+ * during it is picked up by that CPU's own irq_work.
+ */
+static void srcu_defer_drain(struct irq_work *iw)
+{
+	struct srcu_defer *sndp = container_of(iw, struct srcu_defer, iw);
+
+	WRITE_ONCE(sndp->draining, true);
+	__srcu_defer_drain(sndp);
+	WRITE_ONCE(sndp->draining, false);
+}
+
+/*
+ * Drain @cpu's deferred call_srcu() callbacks once @cpu is dead.  One pass
+ * covers every srcu_struct; the re-issue lands on the current CPU.
+ */
+void srcu_offline_drain(int cpu)
+{
+	__srcu_defer_drain(&per_cpu(srcu_defer, cpu));
+}
+
 /**
  * call_srcu() - Queue a callback for invocation after an SRCU grace period
  * @ssp: srcu_struct in queue the callback
@@ -1677,9 +1824,18 @@ void srcu_barrier(struct srcu_struct *ssp)
 {
 	int cpu;
 	int idx;
-	unsigned long s = rcu_seq_snap(&ssp->srcu_sup->srcu_barrier_seq);
+	unsigned long s;
 
 	check_init_srcu_struct(ssp);
+
+	/*
+	 * Register any deferred callbacks before snapshotting the sequence.  The
+	 * staging list is per-CPU, not per-srcu_struct, so this also drains
+	 * other srcu_structs'.
+	 */
+	srcu_defer_flush();
+
+	s = rcu_seq_snap(&ssp->srcu_sup->srcu_barrier_seq);
 	mutex_lock(&ssp->srcu_sup->srcu_barrier_mutex);
 	if (rcu_seq_done(&ssp->srcu_sup->srcu_barrier_seq, s)) {
 		smp_mb(); /* Force ordering following return. */
@@ -2134,6 +2290,12 @@ static void srcu_module_going(struct module *mod)
 	struct srcu_struct *ssp;
 	struct srcu_struct **sspp = mod->srcu_struct_ptrs;
 
+	/*
+	 * Deferral skips check_init_srcu_struct(), so cleanup_srcu_struct()
+	 * below can be skipped for an srcu_struct that has staged callbacks.
+	 * Drain them before any ->sda is freed.
+	 */
+	srcu_defer_flush();
 	for (i = 0; i < mod->num_srcu_structs; i++) {
 		ssp = *(sspp++);
 		if (!rcu_seq_state(smp_load_acquire(&ssp->srcu_sup->srcu_gp_seq_needed)) &&
diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c
index 3bf3a250f9de8..7a9b18825e029 100644
--- a/kernel/rcu/tree.c
+++ b/kernel/rcu/tree.c
@@ -4636,6 +4636,8 @@ void rcutree_migrate_callbacks(int cpu)
 	 * cover.  Drain them here, before the early returns.
 	 */
 	__rcu_defer_drain(rdp);
+	/* Likewise for the outgoing CPU's deferred call_srcu() callbacks. */
+	srcu_offline_drain(cpu);
 
 	if (rcu_rdp_is_offloaded(rdp))
 		return;
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* [PATCH v4 4/6] srcu: Make Tiny call_srcu() safe to call from any context
  2026-08-10 12:27 [PATCH v4 0/6] rcu,srcu: Make call_rcu()/call_srcu() safe from any context Puranjay Mohan
                   ` (2 preceding siblings ...)
  2026-08-10 12:27 ` [PATCH v4 3/6] srcu: Make call_srcu() " Puranjay Mohan
@ 2026-08-10 12:27 ` Puranjay Mohan
  2026-08-10 12:27 ` [PATCH v4 5/6] rcutorture: Exercise ->call() from NMI context Puranjay Mohan
  2026-08-10 12:27 ` [PATCH v4 6/6] selftests/bpf: Add a call_srcu() re-entry reproducer Puranjay Mohan
  5 siblings, 0 replies; 11+ messages in thread
From: Puranjay Mohan @ 2026-08-10 12:27 UTC (permalink / raw)
  To: Lai Jiangshan, Paul E. McKenney, Josh Triplett, Onur Özkan,
	Frederic Weisbecker, Neeraj Upadhyay, Joel Fernandes, Boqun Feng,
	Uladzislau Rezki, Davidlohr Bueso, Andrii Nakryiko,
	Eduard Zingerman, Alexei Starovoitov, Daniel Borkmann,
	Kumar Kartikeya Dwivedi
  Cc: Puranjay Mohan, Steven Rostedt, Mathieu Desnoyers, Zqiang,
	Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
	Emil Tsalapatis, Matt Fleming, Harry Yoo (Oracle), linux-kernel,
	rcu, bpf, linux-rt-devel

Give Tiny call_srcu() the same treatment as Tree SRCU.  When interrupts
are disabled and the scheduler is up, stage the callback on the
srcu_struct's lockless list for an irq_work to re-issue later.  Tiny SRCU
is uniprocessor, so there is no CPU-offline drain.  A draining flag drops
a deferring call_srcu() that re-enters mid-drain (unless from an NMI), as
in Tree SRCU; such a drop can strand state the caller tied to the callback,
not just the callback itself.

Interrupts stay off for the whole batch.  TINY_SRCU implies !SMP, where
arch_irq_work_has_interrupt() is false, so the drain always waits for the
tick and a batch is whatever one tick's worth of interrupts-disabled
call_srcu()s staged.  Unlike the other three flavors srcu_do_enqueue() here
has no debug_rcu_head_queue(), so nothing reports a double call_srcu();
termination of the drain rests on srcu_do_enqueue() clearing ->next, and
the callback list self-links at the tail exactly as a double call_srcu()
made it before.

srcu_barrier() (now out of line) and cleanup_srcu_struct() drain the
deferred list first, so a deferred callback is re-issued onto the callback
list and invoked by the grace-period work that cleanup_srcu_struct()
flushes, rather than stranded on a soon-to-be-freed srcu_struct.
cleanup_srcu_struct() also syncs ->defer_iw, since that irq_work is
embedded in the srcu_struct the caller is about to free.

The draining flag is global rather than per-srcu_struct: a re-entrant
call_srcu(B) inside a drain of A raises B's own ->defer_iw, whose drain can
stage back onto A, so a per-srcu_struct flag would not break the chain.
The cost is that a drain of A also drops a non-NMI call_srcu() to any
other srcu_struct for its duration.

Gated by CONFIG_RCU_DEFER like Tree SRCU, though the srcu_struct members
are unconditional.

Suggested-by: Paul E. McKenney <paulmck@kernel.org>
Signed-off-by: Puranjay Mohan <puranjay@kernel.org>
---
 include/linux/srcutiny.h | 12 ++++--
 kernel/rcu/srcutiny.c    | 93 ++++++++++++++++++++++++++++++++++++++--
 2 files changed, 97 insertions(+), 8 deletions(-)

diff --git a/include/linux/srcutiny.h b/include/linux/srcutiny.h
index fbcf13bc12d15..85b5de438450b 100644
--- a/include/linux/srcutiny.h
+++ b/include/linux/srcutiny.h
@@ -12,6 +12,7 @@
 #define _LINUX_SRCU_TINY_H
 
 #include <linux/irq_work_types.h>
+#include <linux/llist.h>
 #include <linux/swait.h>
 
 struct srcu_struct {
@@ -26,6 +27,8 @@ struct srcu_struct {
 	struct rcu_head **srcu_cb_tail;	/* Pending callbacks: Tail. */
 	struct work_struct srcu_work;	/* For driving grace periods. */
 	struct irq_work srcu_irq_work;	/* Defer schedule_work() to irq work. */
+	struct llist_head defer_cbs;	/* Callbacks deferred on re-entry. */
+	struct irq_work defer_iw;	/* Re-issues defer_cbs later. */
 #ifdef CONFIG_DEBUG_LOCK_ALLOC
 	struct lockdep_map dep_map;
 #endif /* #ifdef CONFIG_DEBUG_LOCK_ALLOC */
@@ -33,6 +36,7 @@ struct srcu_struct {
 
 void srcu_drive_gp(struct work_struct *wp);
 void srcu_tiny_irq_work(struct irq_work *irq_work);
+void srcu_defer_drain(struct irq_work *irq_work);
 
 #define __SRCU_STRUCT_INIT(name, __ignored, ___ignored, ____ignored)	\
 {									\
@@ -40,6 +44,9 @@ void srcu_tiny_irq_work(struct irq_work *irq_work);
 	.srcu_cb_tail = &name.srcu_cb_head,				\
 	.srcu_work = __WORK_INITIALIZER(name.srcu_work, srcu_drive_gp),	\
 	.srcu_irq_work = { .func = srcu_tiny_irq_work },		\
+	.defer_cbs = LLIST_HEAD_INIT(name.defer_cbs),			\
+	.defer_iw = { .node = { .u_flags = IRQ_WORK_HARD_IRQ },		\
+		      .func = srcu_defer_drain },			\
 	__SRCU_DEP_MAP_INIT(name)					\
 }
 
@@ -131,10 +138,7 @@ static inline void synchronize_srcu_expedited(struct srcu_struct *ssp)
 	synchronize_srcu(ssp);
 }
 
-static inline void srcu_barrier(struct srcu_struct *ssp)
-{
-	synchronize_srcu(ssp);
-}
+void srcu_barrier(struct srcu_struct *ssp);
 
 static inline void srcu_expedite_current(struct srcu_struct *ssp) { }
 #define srcu_check_read_flavor(ssp, read_flavor) do { } while (0)
diff --git a/kernel/rcu/srcutiny.c b/kernel/rcu/srcutiny.c
index f9c498ae75df2..32b37d63d58aa 100644
--- a/kernel/rcu/srcutiny.c
+++ b/kernel/rcu/srcutiny.c
@@ -10,6 +10,7 @@
 
 #include <linux/export.h>
 #include <linux/irq_work.h>
+#include <linux/llist.h>
 #include <linux/mutex.h>
 #include <linux/preempt.h>
 #include <linux/rcupdate_wait.h>
@@ -29,6 +30,8 @@ extern int rcu_scheduler_active;
 static LIST_HEAD(srcu_boot_list);
 static bool srcu_init_done;
 
+static void __srcu_defer_drain(struct srcu_struct *ssp);
+
 static int init_srcu_struct_fields(struct srcu_struct *ssp)
 {
 	ssp->srcu_lock_nesting[0] = 0;
@@ -43,6 +46,8 @@ static int init_srcu_struct_fields(struct srcu_struct *ssp)
 	INIT_WORK(&ssp->srcu_work, srcu_drive_gp);
 	INIT_LIST_HEAD(&ssp->srcu_work.entry);
 	init_irq_work(&ssp->srcu_irq_work, srcu_tiny_irq_work);
+	init_llist_head(&ssp->defer_cbs);
+	ssp->defer_iw = IRQ_WORK_INIT_HARD(srcu_defer_drain);
 	return 0;
 }
 
@@ -86,6 +91,16 @@ EXPORT_SYMBOL_GPL(init_srcu_struct_generic);
 void cleanup_srcu_struct(struct srcu_struct *ssp)
 {
 	WARN_ON(srcu_readers_active(ssp));
+	/*
+	 * Re-issue any deferred callbacks, then wait out ->defer_iw before it is
+	 * freed.  Skipped entirely with CONFIG_RCU_DEFER=n: irq_work_sync() ends
+	 * in an unconditional synchronize_rcu() wherever
+	 * arch_irq_work_has_interrupt() is false, which is every !SMP target.
+	 */
+	if (IS_ENABLED(CONFIG_RCU_DEFER)) {
+		__srcu_defer_drain(ssp);
+		irq_work_sync(&ssp->defer_iw);
+	}
 	irq_work_sync(&ssp->srcu_irq_work);
 	flush_work(&ssp->srcu_work);
 	WARN_ON(ssp->srcu_gp_running);
@@ -215,11 +230,11 @@ static void srcu_gp_start_if_needed(struct srcu_struct *ssp)
 }
 
 /*
- * Enqueue an SRCU callback on the specified srcu_struct structure,
- * initiating grace-period processing if it is not already running.
+ * Also called by __srcu_defer_drain() to re-issue a deferred callback, so it
+ * must not re-check the deferral condition.
  */
-void call_srcu(struct srcu_struct *ssp, struct rcu_head *rhp,
-	       rcu_callback_t func)
+static void srcu_do_enqueue(struct srcu_struct *ssp, struct rcu_head *rhp,
+			    rcu_callback_t func)
 {
 	unsigned long flags;
 
@@ -233,6 +248,68 @@ void call_srcu(struct srcu_struct *ssp, struct rcu_head *rhp,
 	srcu_gp_start_if_needed(ssp);
 	preempt_enable();
 }
+
+/*
+ * Set only by the irq_work drain, the one drain its own re-issue can re-feed;
+ * a callback staged during a direct drain is taken by ->defer_iw afterwards.
+ * Global rather than per-srcu_struct: a re-entrant call_srcu(B) inside a drain
+ * of A raises B's own ->defer_iw, whose drain can stage back onto A.
+ */
+static bool srcu_defer_draining;
+
+static void __srcu_defer_drain(struct srcu_struct *ssp)
+{
+	struct llist_node *node, *next;
+	unsigned long flags;
+
+	if (!IS_ENABLED(CONFIG_RCU_DEFER))
+		return;
+
+	/* Re-issued newest-first; nothing depends on call_srcu() ordering. */
+	local_irq_save(flags);
+	llist_for_each_safe(node, next, llist_del_all(&ssp->defer_cbs)) {
+		struct rcu_head *rhp = (struct rcu_head *)node;
+
+		srcu_do_enqueue(ssp, rhp, rhp->func);
+	}
+	local_irq_restore(flags);
+}
+
+/* Only the irq_work drain can be re-fed by its own re-issue; see Tree SRCU. */
+void srcu_defer_drain(struct irq_work *iw)
+{
+	struct srcu_struct *ssp = container_of(iw, struct srcu_struct, defer_iw);
+
+	WRITE_ONCE(srcu_defer_draining, true);
+	__srcu_defer_drain(ssp);
+	WRITE_ONCE(srcu_defer_draining, false);
+}
+EXPORT_SYMBOL_GPL(srcu_defer_drain);
+
+void call_srcu(struct srcu_struct *ssp, struct rcu_head *rhp,
+	       rcu_callback_t func)
+{
+	if (should_rcu_defer()) {
+		/* A re-entrant call_srcu() during the drain would livelock it. */
+		if (READ_ONCE(srcu_defer_draining) && !in_nmi()) {
+			WARN_ONCE(IS_ENABLED(CONFIG_PROVE_RCU),
+				  "call_srcu() re-entered during callback drain; leaking callback\n");
+			return;
+		}
+		rhp->func = func;
+		if (llist_add((struct llist_node *)rhp, &ssp->defer_cbs))
+			irq_work_queue(&ssp->defer_iw);
+		return;
+	}
+
+	/*
+	 * Only reachable from an NMI when deferral is off: before the scheduler
+	 * is up, or with CONFIG_RCU_DEFER=n.  The enqueue can then race.
+	 */
+	WARN_ON_ONCE(IS_ENABLED(CONFIG_PROVE_RCU) && in_nmi());
+
+	srcu_do_enqueue(ssp, rhp, func);
+}
 EXPORT_SYMBOL_GPL(call_srcu);
 
 /*
@@ -262,6 +339,14 @@ void synchronize_srcu(struct srcu_struct *ssp)
 }
 EXPORT_SYMBOL_GPL(synchronize_srcu);
 
+/* Register any deferred callbacks, then wait for all in-flight ones. */
+void srcu_barrier(struct srcu_struct *ssp)
+{
+	__srcu_defer_drain(ssp);
+	synchronize_srcu(ssp);
+}
+EXPORT_SYMBOL_GPL(srcu_barrier);
+
 /*
  * get_state_synchronize_srcu - Provide an end-of-grace-period cookie
  */
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* [PATCH v4 5/6] rcutorture: Exercise ->call() from NMI context
  2026-08-10 12:27 [PATCH v4 0/6] rcu,srcu: Make call_rcu()/call_srcu() safe from any context Puranjay Mohan
                   ` (3 preceding siblings ...)
  2026-08-10 12:27 ` [PATCH v4 4/6] srcu: Make Tiny " Puranjay Mohan
@ 2026-08-10 12:27 ` Puranjay Mohan
  2026-08-10 12:27 ` [PATCH v4 6/6] selftests/bpf: Add a call_srcu() re-entry reproducer Puranjay Mohan
  5 siblings, 0 replies; 11+ messages in thread
From: Puranjay Mohan @ 2026-08-10 12:27 UTC (permalink / raw)
  To: Lai Jiangshan, Paul E. McKenney, Josh Triplett, Onur Özkan,
	Frederic Weisbecker, Neeraj Upadhyay, Joel Fernandes, Boqun Feng,
	Uladzislau Rezki, Davidlohr Bueso, Andrii Nakryiko,
	Eduard Zingerman, Alexei Starovoitov, Daniel Borkmann,
	Kumar Kartikeya Dwivedi
  Cc: Puranjay Mohan, Steven Rostedt, Mathieu Desnoyers, Zqiang,
	Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
	Emil Tsalapatis, Matt Fleming, Harry Yoo (Oracle), linux-kernel,
	rcu, bpf, linux-rt-devel

call_rcu() and call_srcu() are now safe to invoke from NMI, but
rcutorture never does, leaving the deferral path untested.

Add an ->nmi_capable flag to rcu_torture_ops.  For flavors that set it,
arm a per-CPU hardware perf counter whose overflow handler submits a
callback via ->call().  The handler acts only when in_nmi(), so only a
genuine NMI exercises the deferral path.  One preallocated callback per
CPU is kept in flight, guarded by an atomic, to avoid allocating in NMI.
The counter uses a fixed sample period rather than a frequency: a
frequency-based event sets TICK_DEP_BIT_PERF_EVENTS and would pin the tick
for the whole run on NO_HZ_FULL kernels.

Report the count issued from NMI ("nmi-calls:") and the count invoked
("nmi-cbs:").  rcu_torture_cleanup() disables the counters and then calls
cb_barrier(), which drains every deferred callback, so the two counts must
then match; a mismatch fails the test.  This relies on
rcu_barrier()/srcu_barrier() flushing deferred callbacks, as added earlier
in the series.

Set ->nmi_capable on the NMI-safe flavors: rcu, srcu, srcud, and
tasks-tracing (call_srcu() under the hood).  Tasks and Tasks Rude are left
alone, as call_rcu_tasks_generic() is not yet NMI-safe.

Enabled by default; the nmi_calls parameter disables it, which helps rule
NMI handling in or out when triaging a failure.  Requires
CONFIG_PERF_EVENTS and a hardware PMU: without one nothing is issued from
NMI and the end-of-test check compares zero against zero, so a pass does
not by itself mean the path ran.  That is the case under kvm.sh, which
boots qemu with -cpu kvm64 and no vPMU; rcu_torture_nmi_cleanup() says so
on the console.

Signed-off-by: Puranjay Mohan <puranjay@kernel.org>
---
 .../admin-guide/kernel-parameters.txt         |   7 +
 kernel/rcu/rcutorture.c                       | 150 +++++++++++++++++-
 2 files changed, 155 insertions(+), 2 deletions(-)

diff --git a/Documentation/admin-guide/kernel-parameters.txt b/Documentation/admin-guide/kernel-parameters.txt
index 6ebc1cd26c7c8..a3d7a4f8bca33 100644
--- a/Documentation/admin-guide/kernel-parameters.txt
+++ b/Documentation/admin-guide/kernel-parameters.txt
@@ -6198,6 +6198,13 @@ Kernel parameters
 			stress RCU, they don't participate in the actual
 			test, hence the "fake".
 
+	rcutorture.nmi_calls= [KNL]
+			Enable issuing RCU callbacks from an NMI, on the
+			RCU flavors that support it, to exercise the
+			any-context callback path.  Requires
+			CONFIG_PERF_EVENTS and a hardware PMU; without
+			both, nothing is issued.  Defaults to enabled.
+
 	rcutorture.nocbs_nthreads= [KNL]
 			Set number of RCU callback-offload togglers.
 			Zero (the default) disables toggling.
diff --git a/kernel/rcu/rcutorture.c b/kernel/rcu/rcutorture.c
index 39426a8718fe9..0e8b88368bbfc 100644
--- a/kernel/rcu/rcutorture.c
+++ b/kernel/rcu/rcutorture.c
@@ -48,6 +48,7 @@
 #include <linux/tick.h>
 #include <linux/rcupdate_trace.h>
 #include <linux/nmi.h>
+#include <linux/perf_event.h>
 
 #include "rcu.h"
 
@@ -115,6 +116,7 @@ torture_param(int, leakpointer, 0, "Leak pointer dereferences from readers");
 torture_param(int, n_barrier_cbs, 0, "# of callbacks/kthreads for barrier testing");
 torture_param(int, n_up_down, 32, "# of concurrent up/down hrtimer-based RCU readers");
 torture_param(int, nfakewriters, 4, "Number of RCU fake writer threads");
+torture_param(bool, nmi_calls, true, "Exercise ->call() from NMI on nmi_capable flavors");
 torture_param(int, nreaders, -1, "Number of RCU reader threads");
 torture_param(bool, nwriters, 1, "Number of RCU writer threads (0 or 1)");
 torture_param(int, object_debug, 0, "Enable debug-object double call_rcu() testing");
@@ -216,6 +218,8 @@ static long n_rcu_torture_boost_failure;
 static long n_rcu_torture_boosts;
 static atomic_long_t n_rcu_torture_timers;
 static atomic_long_t n_rcu_torture_irqs;
+static atomic_long_t n_rcu_torture_nmi_call;
+static atomic_long_t n_rcu_torture_nmi_cb;
 static long n_barrier_attempts;
 static long n_barrier_successes; /* did rcu_barrier test succeed? */
 static unsigned long n_read_exits;
@@ -433,6 +437,7 @@ struct rcu_torture_ops {
 	bool (*is_task_rcu_boosted)(void);
 	long cbflood_max;
 	int irq_capable;
+	int nmi_capable;
 	int can_boost;
 	int extendables;
 	int slow_gps;
@@ -648,6 +653,7 @@ static struct rcu_torture_ops rcu_ops = {
 	.extendables		= RCUTORTURE_MAX_EXTEND,
 	.debug_objects		= 1,
 	.start_poll_irqsoff	= 1,
+	.nmi_capable		= 1,
 	.name			= "rcu"
 };
 
@@ -942,6 +948,7 @@ static struct rcu_torture_ops srcu_ops = {
 	.debug_objects	= 1,
 	.have_up_down	= IS_ENABLED(CONFIG_TINY_SRCU)
 				? 0 : SRCU_READ_FLAVOR_NORMAL | SRCU_READ_FLAVOR_FAST_UPDOWN,
+	.nmi_capable	= 1,
 	.name		= "srcu"
 };
 
@@ -1005,6 +1012,7 @@ static struct rcu_torture_ops srcud_ops = {
 	.debug_objects	= 1,
 	.have_up_down	= IS_ENABLED(CONFIG_TINY_SRCU)
 				? 0 : SRCU_READ_FLAVOR_NORMAL | SRCU_READ_FLAVOR_FAST_UPDOWN,
+	.nmi_capable	= 1,
 	.name		= "srcud"
 };
 
@@ -1269,6 +1277,7 @@ static struct rcu_torture_ops tasks_tracing_ops = {
 	.cbflood_max	= 50000,
 	.irq_capable	= 1,
 	.slow_gps	= 1,
+	.nmi_capable	= 1,
 	.name		= "tasks-tracing"
 };
 
@@ -2659,6 +2668,124 @@ static bool rcu_torture_one_read(struct torture_random_state *trsp, long myid)
 
 static DEFINE_TORTURE_RANDOM_PERCPU(rcu_torture_timer_rand);
 
+/*
+ * Exercise ->call() from NMI context for flavors that set ->nmi_capable.  A
+ * per-CPU hardware perf counter overflows into an NMI, and its handler submits
+ * a preallocated callback via ->call().  One callback per CPU is in flight at a
+ * time (guarded by an atomic) to avoid allocating in NMI.
+ */
+#ifdef CONFIG_PERF_EVENTS
+static struct perf_event_attr rcu_torture_nmi_attr = {
+	.type		= PERF_TYPE_HARDWARE,
+	.config		= PERF_COUNT_HW_CPU_CYCLES,
+	.size		= sizeof(struct perf_event_attr),
+	.pinned		= 1,
+	.disabled	= 1,
+	/*
+	 * A fixed period rather than .freq: a frequency-based event bumps
+	 * nr_freq_events, which sets TICK_DEP_BIT_PERF_EVENTS and would pin the
+	 * tick for the whole run on NO_HZ_FULL kernels.
+	 */
+	.sample_period	= 20 * 1000 * 1000,
+};
+
+/* One in-flight callback per CPU; ->inuse is released by the callback. */
+struct rcu_torture_nmi_cb {
+	struct rcu_head rh;
+	atomic_t inuse;
+};
+
+static struct perf_event **rcu_torture_nmi_events;
+static int rcu_torture_nmi_hp_state;
+static DEFINE_PER_CPU(struct rcu_torture_nmi_cb, rcu_torture_nmi_cb);
+
+static void rcu_torture_nmi_invoked(struct rcu_head *rhp)
+{
+	struct rcu_torture_nmi_cb *rtncp = container_of(rhp, struct rcu_torture_nmi_cb, rh);
+
+	atomic_long_inc(&n_rcu_torture_nmi_cb);
+	atomic_set(&rtncp->inuse, 0);
+}
+
+static void rcu_torture_nmi_overflow(struct perf_event *event,
+				     struct perf_sample_data *data,
+				     struct pt_regs *regs)
+{
+	struct rcu_torture_nmi_cb *rtncp = this_cpu_ptr(&rcu_torture_nmi_cb);
+
+	if (!in_nmi())
+		return;
+	if (cur_ops->call && !atomic_xchg(&rtncp->inuse, 1)) {
+		atomic_long_inc(&n_rcu_torture_nmi_call);
+		cur_ops->call(&rtncp->rh, rcu_torture_nmi_invoked);
+	}
+}
+
+static int rcu_torture_nmi_online(unsigned int cpu)
+{
+	struct perf_event *event;
+
+	event = perf_event_create_kernel_counter(&rcu_torture_nmi_attr, cpu, NULL,
+						 rcu_torture_nmi_overflow, NULL);
+	if (IS_ERR(event))
+		return 0;
+	rcu_torture_nmi_events[cpu] = event;
+	perf_event_enable(event);
+	return 0;
+}
+
+static int rcu_torture_nmi_offline(unsigned int cpu)
+{
+	struct perf_event *event = rcu_torture_nmi_events[cpu];
+
+	if (event) {
+		rcu_torture_nmi_events[cpu] = NULL;
+		perf_event_disable(event);
+		perf_event_release_kernel(event);
+	}
+	return 0;
+}
+
+/* Drive the counters from hotplug callbacks so coverage survives onoff. */
+static void rcu_torture_nmi_init(void)
+{
+	int ret;
+
+	if (!nmi_calls || !cur_ops->nmi_capable || !cur_ops->call)
+		return;
+	rcu_torture_nmi_events = kcalloc(nr_cpu_ids, sizeof(*rcu_torture_nmi_events),
+					 GFP_KERNEL);
+	if (!rcu_torture_nmi_events)
+		return;
+	ret = cpuhp_setup_state(CPUHP_AP_ONLINE_DYN, "rcutorture/nmi:online",
+				rcu_torture_nmi_online, rcu_torture_nmi_offline);
+	if (ret < 0) {
+		kfree(rcu_torture_nmi_events);
+		rcu_torture_nmi_events = NULL;
+		return;
+	}
+	rcu_torture_nmi_hp_state = ret;
+}
+
+static void rcu_torture_nmi_cleanup(void)
+{
+	if (!rcu_torture_nmi_events)
+		return;
+	if (rcu_torture_nmi_hp_state > 0) {
+		cpuhp_remove_state(rcu_torture_nmi_hp_state);
+		rcu_torture_nmi_hp_state = 0;
+	}
+	kfree(rcu_torture_nmi_events);
+	rcu_torture_nmi_events = NULL;
+	if (!atomic_long_read(&n_rcu_torture_nmi_call))
+		pr_alert("%s: nmi_calls set but no ->call() ever issued from NMI, so NMI ->call() went untested (no PMU, or NMIs unavailable here).\n",
+			 __func__);
+}
+#else /* #ifdef CONFIG_PERF_EVENTS */
+static void rcu_torture_nmi_init(void) { }
+static void rcu_torture_nmi_cleanup(void) { }
+#endif /* #else #ifdef CONFIG_PERF_EVENTS */
+
 /*
  * RCU torture reader from timer handler.  Dereferences rcu_torture_current,
  * incrementing the corresponding element of the pipeline array.  The
@@ -3047,6 +3174,9 @@ rcu_torture_stats_print(void)
 		data_race(n_barrier_attempts),
 		data_race(n_rcu_torture_barrier_error));
 	pr_cont("read-exits: %ld ", data_race(n_read_exits)); // Statistic.
+	pr_cont("nmi-calls: %ld nmi-cbs: %ld ",
+		atomic_long_read(&n_rcu_torture_nmi_call),
+		atomic_long_read(&n_rcu_torture_nmi_cb));
 	pr_cont("nocb-toggles: %ld:%ld ",
 		atomic_long_read(&n_nocb_offload), atomic_long_read(&n_nocb_deoffload));
 	pr_cont("gpwraps: %ld\n", n_gpwraps);
@@ -3195,7 +3325,7 @@ rcu_torture_print_module_parms(struct rcu_torture_ops *cur_ops, const char *tag)
 		 "read_exit_delay=%d read_exit_burst=%d "
 		 "reader_flavor=%x "
 		 "nocbs_nthreads=%d nocbs_toggle=%d "
-		 "test_nmis=%d "
+		 "test_nmis=%d nmi_calls=%d "
 		 "preempt_duration=%d preempt_interval=%d n_up_down=%d\n",
 		 torture_type, tag, nrealreaders, nwriters, nrealfakewriters,
 		 stat_interval, verbose, test_no_idle_hz, shuffle_interval,
@@ -3209,7 +3339,7 @@ rcu_torture_print_module_parms(struct rcu_torture_ops *cur_ops, const char *tag)
 		 read_exit_delay, read_exit_burst,
 		 reader_flavor,
 		 nocbs_nthreads, nocbs_toggle,
-		 test_nmis,
+		 test_nmis, nmi_calls,
 		 preempt_duration, preempt_interval, n_up_down);
 }
 
@@ -4283,6 +4413,7 @@ rcu_torture_cleanup(void)
 	int i;
 
 	if (torture_cleanup_begin()) {
+		rcu_torture_nmi_cleanup();
 		if (cur_ops->cb_barrier != NULL) {
 			pr_info("%s: Invoking %pS().\n", __func__, cur_ops->cb_barrier);
 			cur_ops->cb_barrier();
@@ -4325,6 +4456,7 @@ rcu_torture_cleanup(void)
 		kfree(reader_tasks);
 		reader_tasks = NULL;
 	}
+	rcu_torture_nmi_cleanup();
 	kfree(rcu_torture_reader_mbchk);
 	rcu_torture_reader_mbchk = NULL;
 
@@ -4354,6 +4486,19 @@ rcu_torture_cleanup(void)
 		pr_info("%s: Invoking %pS().\n", __func__, cur_ops->cb_barrier);
 		cur_ops->cb_barrier();
 	}
+
+	/*
+	 * cb_barrier() above drained every deferred callback, so the count
+	 * issued from NMI must equal the count invoked.
+	 */
+	if (atomic_long_read(&n_rcu_torture_nmi_call) !=
+	    atomic_long_read(&n_rcu_torture_nmi_cb)) {
+		pr_alert("%s: NMI ->call() lost a callback: issued %ld invoked %ld\n",
+			 __func__, atomic_long_read(&n_rcu_torture_nmi_call),
+			 atomic_long_read(&n_rcu_torture_nmi_cb));
+		atomic_inc(&n_rcu_torture_error);
+	}
+
 	if (cur_ops->cleanup != NULL)
 		cur_ops->cleanup();
 
@@ -4786,6 +4931,7 @@ rcu_torture_init(void)
 		firsterr = -ENOMEM;
 		goto unwind;
 	}
+	rcu_torture_nmi_init();
 	for (i = 0; i < nrealreaders; i++) {
 		rcu_torture_reader_mbchk[i].rtc_chkrdr = -1;
 		firsterr = torture_create_kthread(rcu_torture_reader, (void *)i,
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* [PATCH v4 6/6] selftests/bpf: Add a call_srcu() re-entry reproducer
  2026-08-10 12:27 [PATCH v4 0/6] rcu,srcu: Make call_rcu()/call_srcu() safe from any context Puranjay Mohan
                   ` (4 preceding siblings ...)
  2026-08-10 12:27 ` [PATCH v4 5/6] rcutorture: Exercise ->call() from NMI context Puranjay Mohan
@ 2026-08-10 12:27 ` Puranjay Mohan
  5 siblings, 0 replies; 11+ messages in thread
From: Puranjay Mohan @ 2026-08-10 12:27 UTC (permalink / raw)
  To: Lai Jiangshan, Paul E. McKenney, Josh Triplett, Onur Özkan,
	Frederic Weisbecker, Neeraj Upadhyay, Joel Fernandes, Boqun Feng,
	Uladzislau Rezki, Davidlohr Bueso, Andrii Nakryiko,
	Eduard Zingerman, Alexei Starovoitov, Daniel Borkmann,
	Kumar Kartikeya Dwivedi
  Cc: Puranjay Mohan, Steven Rostedt, Mathieu Desnoyers, Zqiang,
	Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
	Emil Tsalapatis, Matt Fleming, Harry Yoo (Oracle), linux-kernel,
	rcu, bpf, linux-rt-devel

Re-enter call_srcu() from a BPF program to exercise its any-context
safety, via call_rcu_tasks_trace(), which is call_srcu() on
rcu_tasks_trace_srcu_struct.

An fentry program on rcu_segcblist_enqueue() fires mid-enqueue: that
function is reached from srcu_gp_start_if_needed() with the srcu_data
->lock held.  The program does a task-storage delete, whose only deferred
work is call_rcu_tasks_trace(), re-entering the enqueue on the same CPU.
The handler matches on TID and fires once; pinning the thread removes the
migration window between picking the srcu_data and taking its lock.

Without the fix the nested call re-takes the same sdp lock and
self-deadlocks; with it the nested __call_srcu() sees interrupts disabled
and defers via irq_work, so the delete returns and the test passes.

The test skips where it does not apply: Tiny RCU has no
rcu_segcblist_enqueue() to attach to, and a UP+PREEMPT kernel pairs Tree
RCU with Tiny SRCU, so the attach succeeds but call_srcu() never reaches
the enqueue.  Tiny SRCU is told apart by srcu_expedite_current(), which it
stubs out, so on Tree SRCU a zero hit count fails rather than skips and the
reproducer cannot quietly stop reproducing.

Signed-off-by: Puranjay Mohan <puranjay@kernel.org>
Acked-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
 .../selftests/bpf/prog_tests/rcu_reentry.c    | 93 +++++++++++++++++++
 .../testing/selftests/bpf/progs/rcu_reentry.c | 51 ++++++++++
 2 files changed, 144 insertions(+)
 create mode 100644 tools/testing/selftests/bpf/prog_tests/rcu_reentry.c
 create mode 100644 tools/testing/selftests/bpf/progs/rcu_reentry.c

diff --git a/tools/testing/selftests/bpf/prog_tests/rcu_reentry.c b/tools/testing/selftests/bpf/prog_tests/rcu_reentry.c
new file mode 100644
index 0000000000000..de23a14b3d408
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/rcu_reentry.c
@@ -0,0 +1,93 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Exercise re-entry into call_srcu() from BPF; see progs/rcu_reentry.c. */
+#define _GNU_SOURCE
+#include <sched.h>
+#include <test_progs.h>
+#include "task_local_storage_helpers.h"
+#include "trace_helpers.h"
+#include "rcu_reentry.skel.h"
+
+/* Tiny RCU has no rcu_segcblist_enqueue() to attach to. */
+static bool have_attach_target(void)
+{
+	unsigned long long addr;
+
+	return kallsyms_find("rcu_segcblist_enqueue", &addr) == 0;
+}
+
+/* Tiny SRCU stubs out srcu_expedite_current(); Tree SRCU exports it. */
+static bool have_tree_srcu(void)
+{
+	unsigned long long addr;
+
+	return kallsyms_find("srcu_expedite_current", &addr) == 0;
+}
+
+void test_rcu_reentry(void)
+{
+	struct rcu_reentry *skel;
+	int err, pidfd = -1, map_fd;
+	cpu_set_t set, old_set;
+	bool affinity_saved;
+	__u64 val = 1;
+	int cpu;
+
+	if (!have_attach_target()) {
+		test__skip();
+		return;
+	}
+
+	skel = rcu_reentry__open_and_load();
+	if (!ASSERT_OK_PTR(skel, "skel_open_and_load"))
+		return;
+
+	err = rcu_reentry__attach(skel);
+	if (!ASSERT_OK(err, "skel_attach"))
+		goto out;
+
+	/* Keep the re-entry on a single CPU; a cpuset may exclude CPU 0. */
+	affinity_saved = !sched_getaffinity(0, sizeof(old_set), &old_set);
+	cpu = sched_getcpu();
+	if (!ASSERT_GE(cpu, 0, "getcpu"))
+		goto out;
+	CPU_ZERO(&set);
+	CPU_SET(cpu, &set);
+	if (!ASSERT_OK(sched_setaffinity(0, sizeof(set), &set), "setaffinity"))
+		goto out;
+
+	pidfd = sys_pidfd_open(getpid(), 0);
+	if (!ASSERT_GE(pidfd, 0, "pidfd_open"))
+		goto restore;
+	map_fd = bpf_map__fd(skel->maps.task_stg);
+	err = bpf_map_update_elem(map_fd, &pidfd, &val, BPF_NOEXIST);
+	if (!ASSERT_OK(err, "boot_create"))
+		goto restore;
+
+	/* Arm the handler for this thread, then trigger call_rcu_tasks_trace(). */
+	skel->bss->target_pid = syscall(__NR_gettid);
+	err = bpf_map_delete_elem(map_fd, &pidfd);
+	if (!ASSERT_OK(err, "boot_delete"))
+		goto restore;
+
+	/*
+	 * Only Tree SRCU reaches rcu_segcblist_enqueue() from call_srcu(); a
+	 * UP+PREEMPT kernel pairs Tree RCU with Tiny SRCU, so the attach
+	 * succeeds but nothing fires.  On Tree SRCU it must fire.
+	 */
+	if (!skel->bss->hits) {
+		if (have_tree_srcu())
+			ASSERT_GT(skel->bss->hits, 0, "prog_fired");
+		else
+			test__skip();
+		goto restore;
+	}
+	ASSERT_EQ(skel->bss->get_errs, 0, "nested_storage_get");
+	ASSERT_EQ(skel->bss->del_errs, 0, "nested_storage_delete");
+restore:
+	if (affinity_saved)
+		sched_setaffinity(0, sizeof(old_set), &old_set);
+out:
+	if (pidfd >= 0)
+		close(pidfd);
+	rcu_reentry__destroy(skel);
+}
diff --git a/tools/testing/selftests/bpf/progs/rcu_reentry.c b/tools/testing/selftests/bpf/progs/rcu_reentry.c
new file mode 100644
index 0000000000000..47a36f704cf3e
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/rcu_reentry.c
@@ -0,0 +1,51 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Re-enter call_srcu() from a BPF program.  fentry on rcu_segcblist_enqueue()
+ * fires inside call_srcu()'s enqueue (reached from srcu_gp_start_if_needed()
+ * with the srcu_data ->lock held); the handler then calls call_rcu_tasks_trace()
+ * -- itself call_srcu() on rcu_tasks_trace_srcu_struct -- re-entering the same
+ * srcu_data on the same CPU.
+ */
+#include "vmlinux.h"
+#include <bpf/bpf_helpers.h>
+#include <bpf/bpf_tracing.h>
+
+char _license[] SEC("license") = "GPL";
+
+struct {
+	__uint(type, BPF_MAP_TYPE_TASK_STORAGE);
+	__uint(map_flags, BPF_F_NO_PREALLOC);
+	__type(key, int);
+	__type(value, __u64);
+} task_stg SEC(".maps");
+
+int target_pid;
+int hits;
+int get_errs;
+int del_errs;
+int done;
+
+SEC("fentry/rcu_segcblist_enqueue")
+int BPF_PROG(reenter)
+{
+	struct task_struct *cur;
+
+	if (done || !target_pid)
+		return 0;
+
+	cur = bpf_get_current_task_btf();
+	if (cur->pid != target_pid)
+		return 0;
+
+	/* Issue the nested call exactly once, so the test is deterministic. */
+	done = 1;
+	__sync_fetch_and_add(&hits, 1);
+
+	/* Re-enter via a task-storage delete, which calls call_rcu_tasks_trace(). */
+	if (!bpf_task_storage_get(&task_stg, cur, 0, BPF_LOCAL_STORAGE_GET_F_CREATE))
+		__sync_fetch_and_add(&get_errs, 1);
+	else if (bpf_task_storage_delete(&task_stg, cur))
+		__sync_fetch_and_add(&del_errs, 1);
+
+	return 0;
+}
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* Re: [PATCH v4 2/6] rcu: Make Tiny call_rcu() safe to call from any context
  2026-08-10 12:27 ` [PATCH v4 2/6] rcu: Make Tiny " Puranjay Mohan
@ 2026-08-10 12:42   ` sashiko-bot
  2026-08-10 12:45     ` Puranjay Mohan
  0 siblings, 1 reply; 11+ messages in thread
From: sashiko-bot @ 2026-08-10 12:42 UTC (permalink / raw)
  To: Puranjay Mohan; +Cc: linux-kernel, bpf, linux-rt-devel

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] Moving `debug_rcu_head_queue()` into `rcu_do_enqueue()` bypasses debugobjects tracking for deferred RCU callbacks, masking use-after-free bugs and causing list corruption on double frees.
--

--- Patch [2]: [PATCH v4 2/6] rcu: Make Tiny call_rcu() safe to call from any context ---
commit b9d1d9e1aa870c19c3d9b30019cdb1bc509b6703
Author: Puranjay Mohan <puranjay@kernel.org>

rcu: Make Tiny call_rcu() safe to call from any context

This commit updates Tiny call_rcu() to handle interrupts-disabled contexts
when the scheduler is active. It defers the callback on a lockless list
which is re-issued later via an irq_work, similar to Tree RCU.

> diff --git a/kernel/rcu/tiny.c b/kernel/rcu/tiny.c
> index dccccd6be9411..656b6a682e31a 100644
> --- a/kernel/rcu/tiny.c
> +++ b/kernel/rcu/tiny.c
[ ... ]
> @@ -157,29 +251,19 @@ EXPORT_SYMBOL_GPL(synchronize_rcu);
>   */
>  void call_rcu(struct rcu_head *head, rcu_callback_t func)
>  {
> -	static atomic_t doublefrees;
> -	unsigned long flags;
> -
> -	if (debug_rcu_head_queue(head)) {
> -		if (atomic_inc_return(&doublefrees) < 4) {
> -			pr_err("%s(): Double-freed CB %p->%pS()!!!  ", __func__, head, head->func);
> -			mem_dump_obj(head);
> -		}
> +	if (should_rcu_defer()) {
> +		call_rcu_defer(head, func);
>  		return;
>  	}

[Severity: High]
Does moving debug_rcu_head_queue() into rcu_do_enqueue() bypass debugobjects
tracking for deferred RCU callbacks?

Execution is routed to call_rcu_defer() before debug_rcu_head_queue() can
activate the object in the debugobjects tracker. Adding to the lockless list
in call_rcu_defer() overwrites head->next without verifying if the RCU head
is already enqueued.

If a caller erroneously frees the object or double-queues it during the
deferral window before the timer tick issues the irq_work drain, could this
mask use-after-free bugs or create cyclic structures on double frees?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260810122758.183765-1-puranjay@kernel.org?part=2

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH v4 1/6] rcu: Make call_rcu() safe to call from any context
  2026-08-10 12:27 ` [PATCH v4 1/6] rcu: Make call_rcu() safe to call " Puranjay Mohan
@ 2026-08-10 12:43   ` sashiko-bot
  2026-08-10 12:48     ` Puranjay Mohan
  0 siblings, 1 reply; 11+ messages in thread
From: sashiko-bot @ 2026-08-10 12:43 UTC (permalink / raw)
  To: Puranjay Mohan; +Cc: bpf, linux-rt-devel, linux-kernel

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] Deferred callbacks bypass `debug_rcu_head_queue()` check, causing immediate list corruption on double `call_rcu()` and defeating `CONFIG_DEBUG_OBJECTS_RCU_HEAD`.
--

--- Patch [1]: [PATCH v4 1/6] rcu: Make call_rcu() safe to call from any context ---
commit 5205f0afe602a3f93ee865a26d94981f93553b7d
Author: Puranjay Mohan <puranjay@kernel.org>

rcu: Make call_rcu() safe to call from any context

This patch defers call_rcu() via a per-CPU llist and irq_work when called
with interrupts disabled. This avoids deadlocks or list corruption if
call_rcu() is invoked from NMI or instrumentation.

> diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c
> index 96848fc1f02b8..ff9a2395c9e8e 100644
> --- a/kernel/rcu/tree.c
> +++ b/kernel/rcu/tree.c

[ ... ]

> -static void
> -__call_rcu_common(struct rcu_head *head, rcu_callback_t func, bool lazy_in)
> +/*
> + * Also called by __rcu_defer_drain() to re-issue a deferred callback, so it
> + * must not re-check the deferral condition.  Either caller may have interrupts
> + * already disabled, and a drain of a remote CPU re-issues onto the draining
> + * CPU.
> + */
> +static void rcu_do_enqueue(struct rcu_head *head, rcu_callback_t func, bool lazy_in)
>  {
>  	static atomic_t doublefrees;
>  	unsigned long flags;
>  	bool lazy;
>  	struct rcu_data *rdp;
>  
> -	/* Misaligned rcu_head! */
> -	WARN_ON_ONCE((unsigned long)head & (sizeof(void *) - 1));
> -
> -	/* Avoid NULL dereference if callback is NULL. */
> -	if (WARN_ON_ONCE(!func))
> -		return;
> -
>  	if (debug_rcu_head_queue(head)) {

[ ... ]

> +static void call_rcu_defer(struct rcu_head *head, rcu_callback_t func)
> +{
> +	struct rcu_data *rdp = this_cpu_ptr(&rcu_data);
> +
> +	/*
> +	 * Instrumentation on the enqueue path can re-enter here from inside the
> +	 * drain.  Re-queuing would livelock it, so drop the callback; an NMI
> +	 * cannot loop, so let it through.
> +	 */
> +	if (READ_ONCE(rdp->defer_draining) && !in_nmi()) {
> +		WARN_ONCE(IS_ENABLED(CONFIG_PROVE_RCU),
> +			  "call_rcu() re-entered during callback drain; leaking callback\n");
> +		return;
> +	}
> +	head->func = func;
> +	if (llist_add((struct llist_node *)head, &rdp->defer_head))
> +		irq_work_queue(&rdp->defer_work);
> +}

[ ... ]

> +static void
> +__call_rcu_common(struct rcu_head *head, rcu_callback_t func, bool lazy_in)
> +{
> +	/* Misaligned rcu_head! */
> +	WARN_ON_ONCE((unsigned long)head & (sizeof(void *) - 1));
> +
> +	/* Avoid NULL dereference if callback is NULL. */
> +	if (WARN_ON_ONCE(!func))
> +		return;
> +
> +	/* Record the caller: the irq_work's stack says nothing about it. */
> +	kasan_record_aux_stack(head);
> +
> +	if (should_rcu_defer()) {
> +		call_rcu_defer(head, func);

[Severity: High]
Does this bypass the debug check for double call_rcu(), potentially corrupting
the callback list?

By moving debug_rcu_head_queue() from __call_rcu_common() into rcu_do_enqueue(),
deferred callbacks are added to the defer_head list before being validated.

If a double call_rcu() occurs, llist_add() in call_rcu_defer() will overwrite
head->next. If the callback was already in a normal cblist, its next pointer is
overwritten with defer_head->first, immediately splitting the cblist and
splicing it into the deferred list.

Could we check debug_rcu_head_queue() here before deferring, so that
CONFIG_DEBUG_OBJECTS_RCU_HEAD can safely catch the bug without list corruption?

> +		return;
> +	}
> +
> +	/*
> +	 * Only reachable from an NMI when deferral is off: before the scheduler
> +	 * is up, or with CONFIG_RCU_DEFER=n.  The enqueue can then race.
> +	 */
> +	WARN_ON_ONCE(IS_ENABLED(CONFIG_PROVE_RCU) && in_nmi());
> +
> +	rcu_do_enqueue(head, func, lazy_in);
> +}

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260810122758.183765-1-puranjay@kernel.org?part=1

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH v4 2/6] rcu: Make Tiny call_rcu() safe to call from any context
  2026-08-10 12:42   ` sashiko-bot
@ 2026-08-10 12:45     ` Puranjay Mohan
  0 siblings, 0 replies; 11+ messages in thread
From: Puranjay Mohan @ 2026-08-10 12:45 UTC (permalink / raw)
  To: sashiko-reviews; +Cc: linux-kernel, bpf, linux-rt-devel

On Mon, Aug 10, 2026 at 1:43 PM <sashiko-bot@kernel.org> wrote:
>
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] Moving `debug_rcu_head_queue()` into `rcu_do_enqueue()` bypasses debugobjects tracking for deferred RCU callbacks, masking use-after-free bugs and causing list corruption on double frees.
> --
>
> --- Patch [2]: [PATCH v4 2/6] rcu: Make Tiny call_rcu() safe to call from any context ---
> commit b9d1d9e1aa870c19c3d9b30019cdb1bc509b6703
> Author: Puranjay Mohan <puranjay@kernel.org>
>
> rcu: Make Tiny call_rcu() safe to call from any context
>
> This commit updates Tiny call_rcu() to handle interrupts-disabled contexts
> when the scheduler is active. It defers the callback on a lockless list
> which is re-issued later via an irq_work, similar to Tree RCU.
>
> > diff --git a/kernel/rcu/tiny.c b/kernel/rcu/tiny.c
> > index dccccd6be9411..656b6a682e31a 100644
> > --- a/kernel/rcu/tiny.c
> > +++ b/kernel/rcu/tiny.c
> [ ... ]
> > @@ -157,29 +251,19 @@ EXPORT_SYMBOL_GPL(synchronize_rcu);
> >   */
> >  void call_rcu(struct rcu_head *head, rcu_callback_t func)
> >  {
> > -     static atomic_t doublefrees;
> > -     unsigned long flags;
> > -
> > -     if (debug_rcu_head_queue(head)) {
> > -             if (atomic_inc_return(&doublefrees) < 4) {
> > -                     pr_err("%s(): Double-freed CB %p->%pS()!!!  ", __func__, head, head->func);
> > -                     mem_dump_obj(head);
> > -             }
> > +     if (should_rcu_defer()) {
> > +             call_rcu_defer(head, func);
> >               return;
> >       }
>
> [Severity: High]
> Does moving debug_rcu_head_queue() into rcu_do_enqueue() bypass debugobjects
> tracking for deferred RCU callbacks?
>
> Execution is routed to call_rcu_defer() before debug_rcu_head_queue() can
> activate the object in the debugobjects tracker. Adding to the lockless list
> in call_rcu_defer() overwrites head->next without verifying if the RCU head
> is already enqueued.
>
> If a caller erroneously frees the object or double-queues it during the
> deferral window before the timer tick issues the irq_work drain, could this
> mask use-after-free bugs or create cyclic structures on double frees?

double calling call_rcu() is already a BUG and doing it from NMI is
one level more. debug_rcu_head_queue() can't be called from NMI, until
we have something that can work from NMI we have to go with the
current implementation.

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH v4 1/6] rcu: Make call_rcu() safe to call from any context
  2026-08-10 12:43   ` sashiko-bot
@ 2026-08-10 12:48     ` Puranjay Mohan
  0 siblings, 0 replies; 11+ messages in thread
From: Puranjay Mohan @ 2026-08-10 12:48 UTC (permalink / raw)
  To: sashiko-reviews; +Cc: bpf, linux-rt-devel, linux-kernel

On Mon, Aug 10, 2026 at 1:43 PM <sashiko-bot@kernel.org> wrote:
>
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] Deferred callbacks bypass `debug_rcu_head_queue()` check, causing immediate list corruption on double `call_rcu()` and defeating `CONFIG_DEBUG_OBJECTS_RCU_HEAD`.
> --
>
> --- Patch [1]: [PATCH v4 1/6] rcu: Make call_rcu() safe to call from any context ---
> commit 5205f0afe602a3f93ee865a26d94981f93553b7d
> Author: Puranjay Mohan <puranjay@kernel.org>
>
> rcu: Make call_rcu() safe to call from any context
>
> This patch defers call_rcu() via a per-CPU llist and irq_work when called
> with interrupts disabled. This avoids deadlocks or list corruption if
> call_rcu() is invoked from NMI or instrumentation.
>
> > diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c
> > index 96848fc1f02b8..ff9a2395c9e8e 100644
> > --- a/kernel/rcu/tree.c
> > +++ b/kernel/rcu/tree.c
>
> [ ... ]
>
> > -static void
> > -__call_rcu_common(struct rcu_head *head, rcu_callback_t func, bool lazy_in)
> > +/*
> > + * Also called by __rcu_defer_drain() to re-issue a deferred callback, so it
> > + * must not re-check the deferral condition.  Either caller may have interrupts
> > + * already disabled, and a drain of a remote CPU re-issues onto the draining
> > + * CPU.
> > + */
> > +static void rcu_do_enqueue(struct rcu_head *head, rcu_callback_t func, bool lazy_in)
> >  {
> >       static atomic_t doublefrees;
> >       unsigned long flags;
> >       bool lazy;
> >       struct rcu_data *rdp;
> >
> > -     /* Misaligned rcu_head! */
> > -     WARN_ON_ONCE((unsigned long)head & (sizeof(void *) - 1));
> > -
> > -     /* Avoid NULL dereference if callback is NULL. */
> > -     if (WARN_ON_ONCE(!func))
> > -             return;
> > -
> >       if (debug_rcu_head_queue(head)) {
>
> [ ... ]
>
> > +static void call_rcu_defer(struct rcu_head *head, rcu_callback_t func)
> > +{
> > +     struct rcu_data *rdp = this_cpu_ptr(&rcu_data);
> > +
> > +     /*
> > +      * Instrumentation on the enqueue path can re-enter here from inside the
> > +      * drain.  Re-queuing would livelock it, so drop the callback; an NMI
> > +      * cannot loop, so let it through.
> > +      */
> > +     if (READ_ONCE(rdp->defer_draining) && !in_nmi()) {
> > +             WARN_ONCE(IS_ENABLED(CONFIG_PROVE_RCU),
> > +                       "call_rcu() re-entered during callback drain; leaking callback\n");
> > +             return;
> > +     }
> > +     head->func = func;
> > +     if (llist_add((struct llist_node *)head, &rdp->defer_head))
> > +             irq_work_queue(&rdp->defer_work);
> > +}
>
> [ ... ]
>
> > +static void
> > +__call_rcu_common(struct rcu_head *head, rcu_callback_t func, bool lazy_in)
> > +{
> > +     /* Misaligned rcu_head! */
> > +     WARN_ON_ONCE((unsigned long)head & (sizeof(void *) - 1));
> > +
> > +     /* Avoid NULL dereference if callback is NULL. */
> > +     if (WARN_ON_ONCE(!func))
> > +             return;
> > +
> > +     /* Record the caller: the irq_work's stack says nothing about it. */
> > +     kasan_record_aux_stack(head);
> > +
> > +     if (should_rcu_defer()) {
> > +             call_rcu_defer(head, func);
>
> [Severity: High]
> Does this bypass the debug check for double call_rcu(), potentially corrupting
> the callback list?
>
> By moving debug_rcu_head_queue() from __call_rcu_common() into rcu_do_enqueue(),
> deferred callbacks are added to the defer_head list before being validated.
>
> If a double call_rcu() occurs, llist_add() in call_rcu_defer() will overwrite
> head->next. If the callback was already in a normal cblist, its next pointer is
> overwritten with defer_head->first, immediately splitting the cblist and
> splicing it into the deferred list.
>
> Could we check debug_rcu_head_queue() here before deferring, so that
> CONFIG_DEBUG_OBJECTS_RCU_HEAD can safely catch the bug without list corruption?

No we can't debug_rcu_head_queue() can't be done from NMI and doubling
calling call_rcu() from NMI or interrupts disabled will result in
leaking some callbacks but it will not cause a deadlock (validated in
testing using a kernel module that does multiple call_rcu() from NMI).

^ permalink raw reply	[flat|nested] 11+ messages in thread

end of thread, other threads:[~2026-08-10 12:49 UTC | newest]

Thread overview: 11+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-10 12:27 [PATCH v4 0/6] rcu,srcu: Make call_rcu()/call_srcu() safe from any context Puranjay Mohan
2026-08-10 12:27 ` [PATCH v4 1/6] rcu: Make call_rcu() safe to call " Puranjay Mohan
2026-08-10 12:43   ` sashiko-bot
2026-08-10 12:48     ` Puranjay Mohan
2026-08-10 12:27 ` [PATCH v4 2/6] rcu: Make Tiny " Puranjay Mohan
2026-08-10 12:42   ` sashiko-bot
2026-08-10 12:45     ` Puranjay Mohan
2026-08-10 12:27 ` [PATCH v4 3/6] srcu: Make call_srcu() " Puranjay Mohan
2026-08-10 12:27 ` [PATCH v4 4/6] srcu: Make Tiny " Puranjay Mohan
2026-08-10 12:27 ` [PATCH v4 5/6] rcutorture: Exercise ->call() from NMI context Puranjay Mohan
2026-08-10 12:27 ` [PATCH v4 6/6] selftests/bpf: Add a call_srcu() re-entry reproducer Puranjay Mohan

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox