* [PATCH] md/raid10: fix missing wakeup in wait_barrier_nolock
@ 2026-08-10 13:55 Zizhi Wo
2026-08-10 14:12 ` sashiko-bot
0 siblings, 1 reply; 2+ messages in thread
From: Zizhi Wo @ 2026-08-10 13:55 UTC (permalink / raw)
To: song, yukuai, magiclinan, xiao, linux-raid
Cc: linux-kernel, yangerkun, chengzhihao1, wozizhi
From: Zizhi Wo <wozizhi@huawei.com>
[Bug]
Recently, our fuzz testing triggered a hungtask issue in RAID10:
INFO: task md0_raid10:1273 blocked for more than 120 seconds.
Not tainted 7.2.0-rc6+ #94
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
task:md0_raid10 state:D stack:0 pid:1273 tgid:1273 ppid:2
Call Trace:
<TASK>
__schedule+0xdf9/0x5c90
? _raw_spin_unlock_irqrestore+0xe/0x40
schedule+0x74/0x1f0
raid10d.cold+0x7db/0x1992
md_thread+0x1ce/0x3e0
kthread+0x327/0x410
......
[Cause]
The root cause of the issue is as follows:
[read process1] [read process2] [raid10d]
raid10_make_request
...
// nr_pending == 1
atomic_inc(&conf->nr_pending)
...
raid10_end_read_request
reschedule_retry
md_wakeup_thread(mddev->thread)
raid10_read_request
regular_request_wait
wait_barrier
wait_barrier_nolock
seq = read_seqbegin(&conf->resync_lock)
// nr_pending == 2
atomic_inc(&conf->nr_pending)
raid10d
handle_read_error
freeze_array
write_seqlock_irq(&conf->resync_lock)
conf->array_freeze_pending++
WRITE_ONCE(conf->barrier, conf->barrier + 1)
conf->nr_waiting++
// nr_pending == 2, nr_queued == 0
wait_event_barrier_cmd
write_sequnlock_irq(&(conf)->resync_lock)
schedule // hungtask!!
read_seqretry(&conf->resync_lock, seq)
// not wakeup because conf->nr_pending == 1
atomic_dec_and_test(&conf->nr_pending)
/* Hungtask will also occur here: since the barrier is non-zero,
* this I/O can never complete, so it can't call allow_barrier()
* to wake up the pending freeze_array(). */
wait_event_barrier(conf, stop_waiting_barrier(conf))
wait_barrier_nolock() speculatively increments nr_pending and, on a seqlock
retry, rolls it back with atomic_dec_and_test(). The wake fires only when
nr_pending reaches 0. This causes the freeze_array() in the aforementioned
raid10d flow to never be woken up.
[Fix]
Referring to allow_barrier(), this issue can be fixed by adding a wake-up
condition for "conf->array_freeze_pending" in wait_barrier_nolock().
Fixes: b9b083f9044a ("md/raid10: convert resync_lock to use seqlock")
Signed-off-by: Zizhi Wo <wozizhi@huawei.com>
Signed-off-by: Zizhi Wo <wozizhi@huaweicloud.com>
---
drivers/md/raid10.c | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
diff --git a/drivers/md/raid10.c b/drivers/md/raid10.c
index 0a3cfdd3f5df..fad634a444f8 100644
--- a/drivers/md/raid10.c
+++ b/drivers/md/raid10.c
@@ -996,7 +996,8 @@ static bool wait_barrier_nolock(struct r10conf *conf)
if (!read_seqretry(&conf->resync_lock, seq))
return true;
- if (atomic_dec_and_test(&conf->nr_pending))
+ if (atomic_dec_and_test(&conf->nr_pending) ||
+ conf->array_freeze_pending)
wake_up_barrier(conf);
return false;
--
2.52.0
^ permalink raw reply related [flat|nested] 2+ messages in thread
* Re: [PATCH] md/raid10: fix missing wakeup in wait_barrier_nolock
2026-08-10 13:55 [PATCH] md/raid10: fix missing wakeup in wait_barrier_nolock Zizhi Wo
@ 2026-08-10 14:12 ` sashiko-bot
0 siblings, 0 replies; 2+ messages in thread
From: sashiko-bot @ 2026-08-10 14:12 UTC (permalink / raw)
To: Zizhi Wo; +Cc: linux-raid, yukuai
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Low] Lockless read of `conf->array_freeze_pending` lacks `READ_ONCE()`, violating kernel concurrency guidelines and potentially triggering KCSAN warnings.
--
--- Patch [1]: [PATCH] md/raid10: fix missing wakeup in wait_barrier_nolock ---
commit 3470079c0f4271a77d161b3a70575176ea1d322b
Author: Zizhi Wo <wozizhi@huawei.com>
md/raid10: fix missing wakeup in wait_barrier_nolock
This commit fixes a hung task issue in RAID10 by ensuring that
wait_barrier_nolock() properly wakes up a pending freeze_array(). It adds a
wake-up condition for conf->array_freeze_pending when the pending counter is
decremented.
> diff --git a/drivers/md/raid10.c b/drivers/md/raid10.c
> index 0a3cfdd3f5df8..fad634a444f89 100644
> --- a/drivers/md/raid10.c
> +++ b/drivers/md/raid10.c
[ ... ]
> @@ -996,7 +996,8 @@ static bool wait_barrier_nolock(struct r10conf *conf)
> if (!read_seqretry(&conf->resync_lock, seq))
> return true;
>
> - if (atomic_dec_and_test(&conf->nr_pending))
> + if (atomic_dec_and_test(&conf->nr_pending) ||
> + conf->array_freeze_pending)
[Severity: Low]
Does this lockless read of conf->array_freeze_pending need a READ_ONCE()
wrapper?
Since freeze_array() modifies this variable under the resync_lock seqlock,
reading it here locklessly might trigger KCSAN data race warnings or be
subject to compiler optimizations.
While atomic_dec_and_test() provides a memory barrier, kernel concurrency
guidelines generally require READ_ONCE() for plain lockless reads of shared
variables to prevent compiler optimizations.
It looks like this pattern was copied from allow_barrier(), which also has a
lockless read without READ_ONCE():
drivers/md/raid10.c:allow_barrier() {
...
if ((atomic_dec_and_test(&conf->nr_pending)) ||
(conf->array_freeze_pending))
wake_up_barrier(conf);
...
}
Should both of these locations be updated to use
READ_ONCE(conf->array_freeze_pending)?
> wake_up_barrier(conf);
>
> return false;
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260810135521.3470422-1-wozizhi@huaweicloud.com?part=1
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-08-10 14:12 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-10 13:55 [PATCH] md/raid10: fix missing wakeup in wait_barrier_nolock Zizhi Wo
2026-08-10 14:12 ` sashiko-bot
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.