From: huhai <15815827059@163.com>
To: "John Garry" <john.garry@linux.dev>
Cc: axboe@kernel.dk, linux-block@vger.kernel.org,
"Henry Hu" <huhai@kylinos.cn>
Subject: Re:Re: [PATCH] blk-mq: check passthrough state before reusing cached requests
Date: Thu, 17 Sep 2026 13:52:15 +0800 (CST) [thread overview]
Message-ID: <4deafc69.49ba.1a0adec3e0f.Coremail.15815827059@163.com> (raw)
In-Reply-To: <18a634dd-35fc-4f33-a7c2-4c72d210b0bb@linux.dev>
At 2026-09-17 03:19:58, "John Garry" <john.garry@linux.dev> wrote:
>On 9/16/26 16:16, huhai wrote:
>> From: Henry Hu <huhai@kylinos.cn>
>>
>> blk_mq_alloc_request() and blk_mq_submit_bio() share plug->cached_rqs,
>> but neither lookup checks the passthrough state. This allows an io_uring
>> batch to reuse a cached NVMe passthrough request for an ordinary bio.
>>
>> When an I/O scheduler is enabled, passthrough requests have RQF_SCHED_TAGS
>> set but RQF_USE_SCHED clear, and skip ->prepare_request(). Reuse updates
>> rq->cmd_flags without reinitializing the scheduler state, so the request
>> is inserted into the scheduler on plug flush, while completion skips
>> ->finish_request(). With kyber, this leaks the domain token acquired at
>> dispatch and can stall further I/O in that domain.
>>
>> This was reproduced with kyber and request merging disabled on a QEMU
>> NVMe device by submitting a passthrough read followed by O_DIRECT writes
>> in one io_uring batch. Tokens leak until the kyber write-domain token pool
>> is exhausted, blocking ext4 journal I/O and fsync():
>>
>> INFO: task jbd2/nvme0n1-8:97 blocked in I/O wait for more than 241 seconds.
>> Call Trace:
>> schedule+0xe9/0x300
>> io_schedule+0xca/0x150
>> bit_wait_io+0x1b/0x140
>> __wait_on_bit+0x63/0x170
>> jbd2_write_superblock+0x3d7/0x600
>> jbd2_journal_update_sb_log_tail+0x1e3/0x2d0
>> jbd2_journal_commit_transaction+0x11d1/0x5b50
>>
>> INFO: task fsynctest:107 blocked in I/O wait for more than 241 seconds.
>> Call Trace:
>> schedule+0xe9/0x300
>> io_schedule+0xca/0x150
>> folio_wait_bit_common+0x2ed/0x780
>> folio_wait_bit+0x18/0x30
>> folio_wait_writeback+0x4b/0x1e0
>> __filemap_fdatawait_range+0x120/0x1e0
>> file_write_and_wait_range+0xeb/0x130
>> ext4_sync_file+0x368/0xa80
>> do_fsync+0xa2/0x200
>>
>> Require the passthrough state to match before reusing a cached request.
>> With the fix, the workload completes normally.
>>
>> Assisted-by: LLM
>> Signed-off-by: Henry Hu <huhai@kylinos.cn>
>> ---
>> block/blk-mq.c | 4 ++++
>> 1 file changed, 4 insertions(+)
>>
>> diff --git a/block/blk-mq.c b/block/blk-mq.c
>> index a26a11c73ee3..3a9d57f6a569 100644
>> --- a/block/blk-mq.c
>> +++ b/block/blk-mq.c
>> @@ -648,6 +648,8 @@ static struct request *blk_mq_alloc_cached_request(struct request_queue *q,
>> return NULL;
>> if (op_is_flush(rq->cmd_flags) != op_is_flush(opf))
>> return NULL;
>> + if (blk_rq_is_passthrough(rq) != blk_op_is_passthrough(opf))
>> + return NULL;
>>
>> rq_list_pop(&plug->cached_rqs);
>> blk_mq_rq_time_init(rq, blk_time_get_ns());
>> @@ -3062,6 +3064,8 @@ static struct request *blk_mq_get_cached_request(struct blk_plug *plug,
>> return NULL;
>> if (op_is_flush(rq->cmd_flags) != op_is_flush(opf))
>> return NULL;
>> + if (blk_rq_is_passthrough(rq) != blk_op_is_passthrough(opf))
>> + return NULL;
>
>It seems like some code which code be factored out (between
>blk_mq_alloc_cached_request() and blk_mq_get_cached_request()).
>
>And maybe 7746564793978fe2f43b18a302b22dca0ad3a0e8 could to be repeated
>for blk_mq_alloc_cached_request(), which could mean even more factoring out.
Thanks for the review and suggestions.
I've confirmed dd6216bb16e8 is the patch that introduced the issue. I'll therefore
add the missing tag in v2:
Fixes: dd6216bb16e8 ("blk-mq: make sure elevator callbacks aren't called for passthrough request")
Before dd6216bb16e8 the reproducer is clean. At dd6216bb16e8 it triggers a
KASAN slab-out-of-bounds write in nvme_queue_rqs() during the io_uring plug
flush:
BUG: KASAN: slab-out-of-bounds in nvme_queue_rqs+0x8f3/0xa50
Write of size 8 by task repro/94
Call Trace:
nvme_queue_rqs+0x8f3/0xa50
__blk_mq_flush_plug_list+0x93/0xc0
blk_mq_flush_plug_list+0x152d/0x1d20
blk_finish_plug+0x5c/0xb0
io_submit_sqes+0x1142/0x1db0
__do_sys_io_uring_enter+0x7aa/0x1dc0
With this fix applied, the reproducer completes without the KASAN report.
I will continue to evaluate the suggested refactoring and applying 774656479397
to blk_mq_alloc_cached_request(), and confirm in v2..
Thanks,
Henry Hu.
>
>> rq_list_pop(&plug->cached_rqs);
>> return rq;
>> }
next prev parent reply other threads:[~2026-09-17 5:52 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-16 15:16 [PATCH] blk-mq: check passthrough state before reusing cached requests huhai
2026-09-16 19:19 ` John Garry
2026-09-17 5:52 ` huhai [this message]
2026-09-17 14:21 ` Keith Busch
2026-09-18 5:16 ` huhai
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=4deafc69.49ba.1a0adec3e0f.Coremail.15815827059@163.com \
--to=15815827059@163.com \
--cc=axboe@kernel.dk \
--cc=huhai@kylinos.cn \
--cc=john.garry@linux.dev \
--cc=linux-block@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox