* [PATCH] ublk: return from ublk_wait_dev_ready_and_lock() on group exit
@ 2026-09-30 4:23 Nathan Gao
0 siblings, 0 replies; only message in thread
From: Nathan Gao @ 2026-09-30 4:23 UTC (permalink / raw)
To: Ming Lei, Jens Axboe; +Cc: linux-block, linux-kernel, stable, Nathan Gao
START_DEV and END_USER_RECOVERY wait in ublk_wait_dev_ready_and_lock()
until the server has issued a FETCH for every I/O tag. ublk defers both
commands to io-wq, so the wait runs on one of the server's io-wq
workers.
If the server is killed after fetching only some of the tags, the device
can never become ready, and the wait can only end on a signal. The
worker receives the SIGKILL too, but io-wq workers consume it without
exiting and then run the queued command anyway, so the wait starts with
no signal pending and never wakes up. Meanwhile the dying server waits
for that worker in io_wq_put_and_exit(), because io_uring_files_cancel()
does not cancel a uring_cmd on a non-io_uring file.
The server is left unkillable in D state, and a warning shows up in
dmesg from WARN_ON_ONCE(time_after(jiffies, warn_timeout)) in
io_wq_exit_workers():
WARNING: io_uring/io-wq.c:1373 at io_wq_put_and_exit+0x10f/0x2d0, CPU#5: kublk/215924
Call Trace:
io_uring_clean_tctx+0x85/0xb0
io_uring_cancel_generic+0x171/0x340
do_exit+0xb7/0x470
do_group_exit+0x2c/0x80
get_signal+0x8f8/0x940
arch_do_signal_or_restart+0x24/0xf0
exit_to_user_mode_loop+0xed/0x440
do_syscall_64+0x22d/0x500
entry_SYSCALL_64_after_hwframe+0x76/0x7e
tools/testing/selftests/ublk/test_generic_17.sh, which kills the server
part-way through a recovery fetch, hits this in a few percent of runs.
Fix it by returning -EINTR at the top of the wait loop once
SIGNAL_GROUP_EXIT is set. The flag lives in the signal_struct shared by
all threads, is set before SIGKILL is sent to each of them, and is never
cleared, so the worker still sees it after its SIGKILL is gone. If the
kill arrives after the check, the worker is woken with SIGKILL pending
and the wait is interrupted as before.
The hang goes back to commit fa8e442e832a ("ublk: honor
IO_URING_F_NONBLOCK for handling control command"), which moved these
commands onto io-wq; before it, the wait ran in the thread that receives
the fatal signal.
Fixes: fa8e442e832a ("ublk: honor IO_URING_F_NONBLOCK for handling control command")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Nathan Gao <zcgao@amazon.com>
---
Tested on v7.3-rc5: unpatched, test_generic_17.sh hangs within 3-45
iterations; patched, 150 iterations with no hang.
drivers/block/ublk_drv.c | 8 ++++++++
1 file changed, 8 insertions(+)
diff --git a/drivers/block/ublk_drv.c b/drivers/block/ublk_drv.c
index 66eb55e7162e..dac3084cb86a 100644
--- a/drivers/block/ublk_drv.c
+++ b/drivers/block/ublk_drv.c
@@ -4440,6 +4440,14 @@ static bool ublk_validate_user_pid(struct ublk_device *ub, pid_t ublksrv_pid)
static int ublk_wait_dev_ready_and_lock(struct ublk_device *ub)
{
while (true) {
+ /*
+ * A dying server can never make the device ready. On an io-wq
+ * worker the SIGKILL may already have been consumed, so check
+ * the group exit flag rather than signal_pending().
+ */
+ if (READ_ONCE(current->signal->flags) & SIGNAL_GROUP_EXIT)
+ return -EINTR;
+
if (wait_var_event_interruptible(&ub->nr_queue_ready,
ublk_dev_ready(ub)))
return -EINTR;
base-commit: 551c722f40809618230001baccf219193e22fc5a
--
2.50.1
^ permalink raw reply related [flat|nested] only message in thread
only message in thread, other threads:[~2026-09-30 4:23 UTC | newest]
Thread overview: (only message) (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-30 4:23 [PATCH] ublk: return from ublk_wait_dev_ready_and_lock() on group exit Nathan Gao
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox