From: "Zqiang" <qiang.zhang@linux.dev>
To: "Breno Leitao" <leitao@debian.org>
Cc: paulmck@kernel.org, frederic@kernel.org,
neeraj.upadhyay@kernel.org, joelagnelf@nvidia.com,
urezki@gmail.com, boqun@kernel.org, rcu@vger.kernel.org,
linux-kernel@vger.kernel.org
Subject: Re: [PATCH] srcu: Queue sdp->work when the delay timer is successfully deleted
Date: Mon, 03 Aug 2026 23:40:14 +0000 [thread overview]
Message-ID: <96d6edfce3ac43e153a21b23994251ede0fbd7f4@linux.dev> (raw)
In-Reply-To: <e3738457818909e56c8aa5aef99d8b3e26a328d3@linux.dev>
>
> >
> > On Sat, Aug 01, 2026 at 05:19:25AM +0000, Zqiang wrote:
> >
> >
> > Hello Zqiang,
> >
> > On Thu, Jul 09, 2026 at 06:06:02PM +0800, Zqiang wrote:
> >
> > >
> > > In the cleanup_srcu_struct(), when iterating over per-cpu's srcu_data,
> > > the timer_delete_sync(&sdp->delay_work) is called to cancel the delay
> > > timer before flush_work(&sdp->work).
> > >
> > > However, if the timer_delete_sync() returns 1 means that it successfully
> > > deleted an pending timer before it had a chance to fire, also means that
> > > the sdp->work cannot be queued, the subsequent flush_work(&sdp->work)
> > > will returns immediately without waiting for anything, this causes SRCU
> > > callbacks to not be processed.
> > >
> > > Fix this by checking the return value of timer_delete_sync(), if it
> > > returns 1, explicitly queue sdp->work so that the following flush_work()
> > > can correctly wait for the work to complete.
> > >
> > > Signed-off-by: Zqiang <qiang.zhang@linux.dev>
> > > ---
> > > kernel/rcu/srcutree.c | 6 +++++-
> > > 1 file changed, 5 insertions(+), 1 deletion(-)
> > >
> > > diff --git a/kernel/rcu/srcutree.c b/kernel/rcu/srcutree.c
> > > index 7c2f7cc131f7..02c322b7c6f1 100644
> > > --- a/kernel/rcu/srcutree.c
> > > +++ b/kernel/rcu/srcutree.c
> > > @@ -725,7 +725,11 @@ void cleanup_srcu_struct(struct srcu_struct *ssp)
> > > for_each_possible_cpu(cpu) {
> > > struct srcu_data *sdp = per_cpu_ptr(ssp->sda, cpu);
> > >
> > > - timer_delete_sync(&sdp->delay_work);
> > > + //In most scenarios, calling srcu_barrier before cleanup
> > > + //will not trigger WARN_ON().
> > > + if (WARN_ON(timer_delete_sync(&sdp->delay_work)) &&
> > > + rcu_cpu_beenfullyonline(sdp->cpu))
> > > + queue_work_on(sdp->cpu, rcu_gp_wq, &sdp->work);
> > >
> > I started seeing this on my tests, it is not trivial to decode this one,
> > but, I can try harder if _really_ needed.
> >
> >
> > The scenario I can think of is that we missed the call to srcu_barrier()
> > before cleanup_srcu_struct():
> >
> > loop_add()
> > ->blk_mq_alloc_tag_set
> > init_srcu_struct(&set->tags_srcu)
> >
> > blk_mq_alloc_set_map_and_rqs() {
> > ->__blk_mq_alloc_rq_maps()
> > ->__blk_mq_alloc_map_and_rqs() return error
> > goto out_unwind: __blk_mq_free_map_and_rqs()
> > ->blk_mq_free_rq_map()
> > ->blk_mq_free_tags()
> > ->call_srcu(&set->tags_srcu, &tags->rcu_head, blk_mq_free_tags_callback);
> > } return error
> >
> > goto out_free_mq_map:
> > ....
> > cleanup_srcu_struct(&set->tags_srcu)
> > -> trigger WARN_ON(timer_delete_sync(&sdp->delay_work)
> >
> > Can you try the following patch?
> >
> > I tried it, and it does not help -- the WARN still fires at the same rate. I think the analysis points at the wrong call site.
> >
> > Setup: linux-next-20260731 (arm64), 32 vCPU VM, HZ=1000, PROVE_LOCKING
> > and DEBUG_OBJECTS_TIMERS enabled, reproducer stress-ng --loop 32
> > --timeout 60s.
> >
> > Thanks for provide testing methods, I will also testing it.
> >
> >
> > baseline 3 x WARN srcutree.c:706 in 60s
> > + your blk-mq patch 4 x WARN srcutree.c:706 in 60s
> >
> > (3 vs 4 is just jitter on a one-jiffy race, not a regression.)
> >
> > The reason it cannot help is that the splat comes from
> > blk_mq_free_tag_set(), not from the blk_mq_alloc_tag_set() error path
> > your patch touches. All four splats in the patched run have the same
> > trace:
> >
> > cleanup_srcu_struct+0x274/0x450 (P)
> > blk_mq_free_tag_set+0x1a4/0x1e0
> > loop_remove+0x2c/0x78
> > loop_control_ioctl+0x248/0x2a0
> > __arm64_sys_ioctl+0x9c0/0xb00
> >
> > This may trigger a new srcu grace period again during the window period between
> > srcu-barrier() and cleanup_srcu_struct().
> >
> >
> > I also put a pr_warn() at out_cleanup_tags_srcu: to be sure -- it fired
> > zero times over the whole run, so loop_add() never takes that error path
> > in this workload.
> >
> > Why do you wangt to have this
> > WARN_ON(timer_delete_sync(&sdp->delay_work)) ?
> >
> > timer_delete_sync() != 0 "a timer was armed", not "callbacks are
> > pending".
> >
> > There are only two types of return values for timer_delete_sync(),
> > return 0 or 1, the timer_delete_sync() != 0 means that this timer
> > was pending and has been deactivated, right?
> >
> How about this?
>
> diff --git a/kernel/rcu/srcutree.c b/kernel/rcu/srcutree.c
> index b9fe57ff9100..62da31ec2bd0 100644
> --- a/kernel/rcu/srcutree.c
> +++ b/kernel/rcu/srcutree.c
> @@ -751,7 +751,8 @@ void cleanup_srcu_struct(struct srcu_struct *ssp)
>
> // Call srcu_barrier() before this cleanup_srcu_struct()
> // to avoid triggering this WARN_ON().
> - if (WARN_ON(timer_delete_sync(&sdp->delay_work)) &&
> + if (WARN_ON(rcu_segcblist_n_cbs(&sdp->srcu_cblist) &&
> + timer_delete_sync(&sdp->delay_work)) &&
> rcu_cpu_beenfullyonline(sdp->cpu))
> queue_work_on(sdp->cpu, rcu_gp_wq, &sdp->work);
> flush_work(&sdp->work);
Please ignore this change, this is my mistake.
There is a scenario where the 'rcu_segcblist_n_cbs(&sdp->srcu_cblist) == 0'
srcu_barrier() cannot intercept.
however, for sup->srcu_size_state being between SRCU_SIZE_WAIT_BARRIER and
SRCU_SIZE_BIG, we will queue sdp->delay_work for all CPUs belonging to the
leaf SRCU node, regardless of whether there is a callback on the current
CPU's sdp->srcu_cblist in srcu_gp_end().
so the timer_delete_sync(&sdp->delay_work) should be called unconditionally.
to ensure that the timer callback has ended or the timer which in pending
status has been successfully deleted.
diff --git a/kernel/rcu/srcutree.c b/kernel/rcu/srcutree.c
index b9fe57ff9100..d13c12f61150 100644
--- a/kernel/rcu/srcutree.c
+++ b/kernel/rcu/srcutree.c
@@ -751,8 +751,9 @@ void cleanup_srcu_struct(struct srcu_struct *ssp)
// Call srcu_barrier() before this cleanup_srcu_struct()
// to avoid triggering this WARN_ON().
- if (WARN_ON(timer_delete_sync(&sdp->delay_work)) &&
- rcu_cpu_beenfullyonline(sdp->cpu))
+ if (WARN_ON(timer_delete_sync(&sdp->delay_work) &&
+ rcu_segcblist_n_cbs(&sdp->srcu_cblist)) &&
+ rcu_cpu_beenfullyonline(sdp->cpu))
queue_work_on(sdp->cpu, rcu_gp_wq, &sdp->work);
flush_work(&sdp->work);
if (WARN_ON(rcu_segcblist_n_cbs(&sdp->srcu_cblist)))
Thanks
Zqiang
>
> Thanks
> Zqiang
>
> >
> > Thanks
> > Zqiang
> >
>
next prev parent reply other threads:[~2026-08-04 4:22 UTC|newest]
Thread overview: 9+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-09 10:06 [PATCH] srcu: Queue sdp->work when the delay timer is successfully deleted Zqiang
2026-07-10 18:34 ` Paul E. McKenney
2026-07-31 17:45 ` Breno Leitao
2026-08-01 5:19 ` Zqiang
2026-08-03 10:49 ` Breno Leitao
2026-08-03 14:11 ` Zqiang
2026-08-03 14:37 ` Zqiang
2026-08-03 23:40 ` Zqiang [this message]
2026-08-04 8:04 ` Breno Leitao
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=96d6edfce3ac43e153a21b23994251ede0fbd7f4@linux.dev \
--to=qiang.zhang@linux.dev \
--cc=boqun@kernel.org \
--cc=frederic@kernel.org \
--cc=joelagnelf@nvidia.com \
--cc=leitao@debian.org \
--cc=linux-kernel@vger.kernel.org \
--cc=neeraj.upadhyay@kernel.org \
--cc=paulmck@kernel.org \
--cc=rcu@vger.kernel.org \
--cc=urezki@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox