Linux RCU subsystem development
 help / color / mirror / Atom feed
From: Boqun Feng <boqun@kernel.org>
To: Puranjay Mohan <puranjay12@gmail.com>
Cc: "Paul E. McKenney" <paulmck@kernel.org>,
	rcu@vger.kernel.org, linux-kernel@vger.kernel.org,
	kernel-team@meta.com, rostedt@goodmis.org
Subject: Re: [PATCH 1/7] rcu: Make call_rcu() safe to call from any context
Date: Mon, 21 Sep 2026 20:30:42 +0200	[thread overview]
Message-ID: <arF30s1AAAkmq0DB@MacBook-0RXW5> (raw)
In-Reply-To: <CANk7y0g5EU8T8cLiU=+tLiZCcMEhZOjhnvs4uzD8Z=8GwN+7LA@mail.gmail.com>

On Mon, Sep 21, 2026 at 07:12:02PM +0100, Puranjay Mohan wrote:
> On Sat, Sep 19, 2026 at 3:07 PM Boqun Feng <boqun@kernel.org> wrote:
> >
> > On Fri, Sep 18, 2026 at 05:32:52PM -0700, Paul E. McKenney wrote:
> > > From: Puranjay Mohan <puranjay@kernel.org>
> > >
> >
> > Hi,
> >
> > Sorry for a bit late repsonse.
> >
> > > RCU's per-CPU callback list is only touched with interrupts disabled: the
> > > enqueue runs under local_irq_save() (and the nocb locks when offloaded),
> > > as do callback invocation and grace-period work.  A call_rcu() that
> > > arrives with interrupts already disabled, whether from an NMI or from
> > > instrumentation that re-enters RCU, can interrupt one of those and corrupt
> > > the list or deadlock.
> > >
> > > Defer instead: stage the callback on a per-CPU llist and raise an irq_work
> > > that re-issues it once interrupts are on, straight to the enqueue so it
> > > cannot defer again.  The gate is bare irqs_disabled(), so callers that
> > > merely hold interrupts off are deferred too and pay one irq_work hop.
> > > Skip it while the scheduler is down (RCU_SCHEDULER_INACTIVE): irq_work is
> > > not usable that early, rcu_init() already calls call_rcu(), and the per-CPU
> > > deferral state is not initialised until rcu_init_one() runs later in it.
> > >
> > > rcu_barrier() drains every CPU's ->defer_head before it scans the lists,
> > > and rcutree_migrate_callbacks() drains an outgoing CPU's.  A drain
> > > re-issues onto the draining CPU, so a barrier moves other CPUs' staged
> > > callbacks onto
> > > its own ->cblist; call_rcu() promises no CPU affinity for invocation.
> > > ->defer_lock is held across llist_del_all() and the whole re-issue so the
> > > drainers
> > > serialize: one that finds the list empty can conclude that everything
> > > staged before it is already on a callback list.  Interrupts stay off for
> > > the batch.  Where the arch has an irq_work self-IPI that is what one
> > > interrupts-disabled region could stage, normally a single callback; where
> > > arch_irq_work_has_interrupt() is false the drain waits for the tick, so
> > > several regions can accumulate first.
> > >
> > > The drain clears ->next before re-issuing.  A double call_rcu() on a head
> > > that is already debug-object-active self-links the staged node, and
> > > rcu_do_enqueue()'s duplicate path returns without clearing it, so the
> > > drain would spin.  A re-add behind other staged callbacks makes a longer
> > > cycle, which that does not bound; a double call_rcu() stays undefined.
> > > llist_del_all() yields newest-first, so a batch is re-issued in reverse
> > > call order; nothing depends on call_rcu() ordering.  The re-issue drops
> > > the lazy hint, since staging records only ->func, so a deferred callback
> > > loses its batching on CONFIG_RCU_LAZY.  kasan_record_aux_stack() moves to
> >
> > I'm not sure this is a good idea, because it effectively remove LAZY
> > support when DEFER is enabled. Since the goal of this patchset supports
> > BPF and NMI, would it be nicer that we skip the whole defer logic if the
> > callback is LAZY? Alternatively, you can have two llist (one for hurry
> > and one for lazy).
> 
> I had made this trade-off of removing the Lazy tag as I thought it is
> not necessary to support lazy when call_rcu() is called from nmi and
> bpf based instrumentation as they should not be frequent. But I like

Ah, I missed that you only defer if irqs_disabled() is true, so this is
much better than I used to think. So..

> the idea of two lists (skipping the defer logic is not possible as it
> could lead to deadlocks/corruption). I will also investigate if we can
> put the Lazy tag on the ->next pointer. But will it be acceptable if I
> do that as a follow up? I want to get the base support fully validated

.. definitely a follow-up would do, thank you!

Regards,
Boqun

> with the BPF side changes. I also have more optimizations planned as
> suggested by Sebastian.
> 
> Thanks,
> Puranjay
> 

  reply	other threads:[~2026-09-21 18:30 UTC|newest]

Thread overview: 12+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-19  0:32 [PATCH 0/7] Allow call_{s,}rcu() from NMI and BPF environments Paul E. McKenney
2026-09-19  0:32 ` [PATCH 1/7] rcu: Make call_rcu() safe to call from any context Paul E. McKenney
2026-09-19 14:07   ` Boqun Feng
2026-09-21 18:12     ` Puranjay Mohan
2026-09-21 18:30       ` Boqun Feng [this message]
2026-09-21 18:50         ` Paul E. McKenney
2026-09-19  0:32 ` [PATCH 2/7] rcu: Make Tiny " Paul E. McKenney
2026-09-19  0:32 ` [PATCH 3/7] srcu: Make call_srcu() " Paul E. McKenney
2026-09-19  0:32 ` [PATCH 4/7] srcu: Make Tiny " Paul E. McKenney
2026-09-19  0:32 ` [PATCH 5/7] rcutorture: Disable fragile readers during overload testing Paul E. McKenney
2026-09-19  0:32 ` [PATCH 6/7] rcutorture: Exercise ->call() from NMI context Paul E. McKenney
2026-09-19  0:32 ` [PATCH 7/7] selftests/bpf: Add a call_srcu() re-entry reproducer Paul E. McKenney

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=arF30s1AAAkmq0DB@MacBook-0RXW5 \
    --to=boqun@kernel.org \
    --cc=kernel-team@meta.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=paulmck@kernel.org \
    --cc=puranjay12@gmail.com \
    --cc=rcu@vger.kernel.org \
    --cc=rostedt@goodmis.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox