Linux RCU subsystem development
 help / color / mirror / Atom feed
From: Breno Leitao <leitao@debian.org>
To: Zqiang <qiang.zhang@linux.dev>
Cc: paulmck@kernel.org, frederic@kernel.org,
	neeraj.upadhyay@kernel.org,  joelagnelf@nvidia.com,
	urezki@gmail.com, boqun@kernel.org, rcu@vger.kernel.org,
	 linux-kernel@vger.kernel.org
Subject: Re: [PATCH] srcu: Queue sdp->work when the delay timer is successfully deleted
Date: Mon, 3 Aug 2026 03:49:03 -0700	[thread overview]
Message-ID: <anBxXePfyHY2Uz0d@gmail.com> (raw)
In-Reply-To: <0bba3fa18b116a3b08e3e83310b3197bcf780000@linux.dev>

On Sat, Aug 01, 2026 at 05:19:25AM +0000, Zqiang wrote:
> > 
> > Hello Zqiang,
> > 
> > On Thu, Jul 09, 2026 at 06:06:02PM +0800, Zqiang wrote:
> > 
> > > 
> > > In the cleanup_srcu_struct(), when iterating over per-cpu's srcu_data,
> > >  the timer_delete_sync(&sdp->delay_work) is called to cancel the delay
> > >  timer before flush_work(&sdp->work).
> > >  
> > >  However, if the timer_delete_sync() returns 1 means that it successfully
> > >  deleted an pending timer before it had a chance to fire, also means that
> > >  the sdp->work cannot be queued, the subsequent flush_work(&sdp->work)
> > >  will returns immediately without waiting for anything, this causes SRCU
> > >  callbacks to not be processed.
> > >  
> > >  Fix this by checking the return value of timer_delete_sync(), if it
> > >  returns 1, explicitly queue sdp->work so that the following flush_work()
> > >  can correctly wait for the work to complete.
> > >  
> > >  Signed-off-by: Zqiang <qiang.zhang@linux.dev>
> > >  ---
> > >  kernel/rcu/srcutree.c | 6 +++++-
> > >  1 file changed, 5 insertions(+), 1 deletion(-)
> > >  
> > >  diff --git a/kernel/rcu/srcutree.c b/kernel/rcu/srcutree.c
> > >  index 7c2f7cc131f7..02c322b7c6f1 100644
> > >  --- a/kernel/rcu/srcutree.c
> > >  +++ b/kernel/rcu/srcutree.c
> > >  @@ -725,7 +725,11 @@ void cleanup_srcu_struct(struct srcu_struct *ssp)
> > >  for_each_possible_cpu(cpu) {
> > >  struct srcu_data *sdp = per_cpu_ptr(ssp->sda, cpu);
> > >  
> > >  - timer_delete_sync(&sdp->delay_work);
> > >  + //In most scenarios, calling srcu_barrier before cleanup
> > >  + //will not trigger WARN_ON().
> > >  + if (WARN_ON(timer_delete_sync(&sdp->delay_work)) &&
> > >  + rcu_cpu_beenfullyonline(sdp->cpu))
> > >  + queue_work_on(sdp->cpu, rcu_gp_wq, &sdp->work);
> > > 
> > I started seeing this on my tests, it is not trivial to decode this one,
> > but, I can try harder if _really_ needed.
> >
> 
> The scenario I can think of is that we missed the call to srcu_barrier()
> before cleanup_srcu_struct():
> 
> loop_add()
> ->blk_mq_alloc_tag_set
>     init_srcu_struct(&set->tags_srcu)
> 
>     blk_mq_alloc_set_map_and_rqs() {
>      ->__blk_mq_alloc_rq_maps()
>        ->__blk_mq_alloc_map_and_rqs() return error
>          goto out_unwind: __blk_mq_free_map_and_rqs()
>                           ->blk_mq_free_rq_map()
>                             ->blk_mq_free_tags()
>                               ->call_srcu(&set->tags_srcu, &tags->rcu_head, blk_mq_free_tags_callback);
>     } return error 
>      
>     goto out_free_mq_map:
>          ....
>          cleanup_srcu_struct(&set->tags_srcu)
>          -> trigger WARN_ON(timer_delete_sync(&sdp->delay_work)
> 
> Can you try the following patch?

I tried it, and it does not help -- the WARN still fires at the same rate. I think the analysis points at the wrong call site.

Setup: linux-next-20260731 (arm64), 32 vCPU VM, HZ=1000, PROVE_LOCKING
and DEBUG_OBJECTS_TIMERS enabled, reproducer stress-ng --loop 32
--timeout 60s.

    baseline               3 x WARN srcutree.c:706 in 60s
    + your blk-mq patch    4 x WARN srcutree.c:706 in 60s

  (3 vs 4 is just jitter on a one-jiffy race, not a regression.)

The reason it cannot help is that the splat comes from
blk_mq_free_tag_set(), not from the blk_mq_alloc_tag_set() error path
your patch touches. All four splats in the patched run have the same
trace:

    cleanup_srcu_struct+0x274/0x450 (P)
    blk_mq_free_tag_set+0x1a4/0x1e0
    loop_remove+0x2c/0x78
    loop_control_ioctl+0x248/0x2a0
    __arm64_sys_ioctl+0x9c0/0xb00

I also put a pr_warn() at out_cleanup_tags_srcu: to be sure -- it fired
zero times over the whole run, so loop_add() never takes that error path
in this workload.

Why do you wangt to have this
WARN_ON(timer_delete_sync(&sdp->delay_work)) ?

timer_delete_sync() != 0 means "a timer was armed", not "callbacks are
pending".

  reply	other threads:[~2026-08-03 10:49 UTC|newest]

Thread overview: 10+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-09 10:06 [PATCH] srcu: Queue sdp->work when the delay timer is successfully deleted Zqiang
2026-07-10 18:34 ` Paul E. McKenney
2026-07-31 17:45 ` Breno Leitao
2026-08-01  5:19   ` Zqiang
2026-08-03 10:49     ` Breno Leitao [this message]
2026-08-03 14:11       ` Zqiang
2026-08-03 14:37         ` Zqiang
2026-08-03 23:40           ` Zqiang
2026-08-04  8:04             ` Breno Leitao
2026-08-04 21:58               ` Zqiang

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=anBxXePfyHY2Uz0d@gmail.com \
    --to=leitao@debian.org \
    --cc=boqun@kernel.org \
    --cc=frederic@kernel.org \
    --cc=joelagnelf@nvidia.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=neeraj.upadhyay@kernel.org \
    --cc=paulmck@kernel.org \
    --cc=qiang.zhang@linux.dev \
    --cc=rcu@vger.kernel.org \
    --cc=urezki@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox