Linux block layer
 help / color / mirror / Atom feed
From: huhai  <15815827059@163.com>
To: "John Garry" <john.garry@linux.dev>
Cc: axboe@kernel.dk, linux-block@vger.kernel.org,
	"Henry Hu" <huhai@kylinos.cn>
Subject: Re:Re: [PATCH] blk-mq: check passthrough state before reusing cached requests
Date: Thu, 17 Sep 2026 13:52:15 +0800 (CST)	[thread overview]
Message-ID: <4deafc69.49ba.1a0adec3e0f.Coremail.15815827059@163.com> (raw)
In-Reply-To: <18a634dd-35fc-4f33-a7c2-4c72d210b0bb@linux.dev>

At 2026-09-17 03:19:58, "John Garry" <john.garry@linux.dev> wrote:
>On 9/16/26 16:16, huhai wrote:
>> From: Henry Hu <huhai@kylinos.cn>
>> 
>> blk_mq_alloc_request() and blk_mq_submit_bio() share plug->cached_rqs,
>> but neither lookup checks the passthrough state. This allows an io_uring
>> batch to reuse a cached NVMe passthrough request for an ordinary bio.
>> 
>> When an I/O scheduler is enabled, passthrough requests have RQF_SCHED_TAGS
>> set but RQF_USE_SCHED clear, and skip ->prepare_request(). Reuse updates
>> rq->cmd_flags without reinitializing the scheduler state, so the request
>> is inserted into the scheduler on plug flush, while completion skips
>> ->finish_request(). With kyber, this leaks the domain token acquired at
>> dispatch and can stall further I/O in that domain.
>> 
>> This was reproduced with kyber and request merging disabled on a QEMU
>> NVMe device by submitting a passthrough read followed by O_DIRECT writes
>> in one io_uring batch. Tokens leak until the kyber write-domain token pool
>> is exhausted, blocking ext4 journal I/O and fsync():
>> 
>>    INFO: task jbd2/nvme0n1-8:97 blocked in I/O wait for more than 241 seconds.
>>      Call Trace:
>>        schedule+0xe9/0x300
>>        io_schedule+0xca/0x150
>>        bit_wait_io+0x1b/0x140
>>        __wait_on_bit+0x63/0x170
>>        jbd2_write_superblock+0x3d7/0x600
>>        jbd2_journal_update_sb_log_tail+0x1e3/0x2d0
>>        jbd2_journal_commit_transaction+0x11d1/0x5b50
>> 
>>    INFO: task fsynctest:107 blocked in I/O wait for more than 241 seconds.
>>      Call Trace:
>>        schedule+0xe9/0x300
>>        io_schedule+0xca/0x150
>>        folio_wait_bit_common+0x2ed/0x780
>>        folio_wait_bit+0x18/0x30
>>        folio_wait_writeback+0x4b/0x1e0
>>        __filemap_fdatawait_range+0x120/0x1e0
>>        file_write_and_wait_range+0xeb/0x130
>>        ext4_sync_file+0x368/0xa80
>>        do_fsync+0xa2/0x200
>> 
>> Require the passthrough state to match before reusing a cached request.
>> With the fix, the workload completes normally.
>> 
>> Assisted-by: LLM
>> Signed-off-by: Henry Hu <huhai@kylinos.cn>
>> ---
>>   block/blk-mq.c | 4 ++++
>>   1 file changed, 4 insertions(+)
>> 
>> diff --git a/block/blk-mq.c b/block/blk-mq.c
>> index a26a11c73ee3..3a9d57f6a569 100644
>> --- a/block/blk-mq.c
>> +++ b/block/blk-mq.c
>> @@ -648,6 +648,8 @@ static struct request *blk_mq_alloc_cached_request(struct request_queue *q,
>>   			return NULL;
>>   		if (op_is_flush(rq->cmd_flags) != op_is_flush(opf))
>>   			return NULL;
>> +		if (blk_rq_is_passthrough(rq) != blk_op_is_passthrough(opf))
>> +			return NULL;
>>   
>>   		rq_list_pop(&plug->cached_rqs);
>>   		blk_mq_rq_time_init(rq, blk_time_get_ns());
>> @@ -3062,6 +3064,8 @@ static struct request *blk_mq_get_cached_request(struct blk_plug *plug,
>>   		return NULL;
>>   	if (op_is_flush(rq->cmd_flags) != op_is_flush(opf))
>>   		return NULL;
>> +	if (blk_rq_is_passthrough(rq) != blk_op_is_passthrough(opf))
>> +		return NULL;
>
>It seems like some code which code be factored out (between 
>blk_mq_alloc_cached_request() and blk_mq_get_cached_request()).
>
>And maybe 7746564793978fe2f43b18a302b22dca0ad3a0e8 could to be repeated 
>for blk_mq_alloc_cached_request(), which could mean even more factoring out.

Thanks for the review and suggestions.

I've confirmed dd6216bb16e8 is the patch that introduced the issue. I'll therefore
add the missing tag in v2:

Fixes: dd6216bb16e8 ("blk-mq: make sure elevator callbacks aren't called for passthrough request")

Before dd6216bb16e8 the reproducer is clean. At dd6216bb16e8 it triggers a
KASAN slab-out-of-bounds write in nvme_queue_rqs() during the io_uring plug
flush:

  BUG: KASAN: slab-out-of-bounds in nvme_queue_rqs+0x8f3/0xa50
  Write of size 8 by task repro/94
  Call Trace:
    nvme_queue_rqs+0x8f3/0xa50
    __blk_mq_flush_plug_list+0x93/0xc0
    blk_mq_flush_plug_list+0x152d/0x1d20
    blk_finish_plug+0x5c/0xb0
    io_submit_sqes+0x1142/0x1db0
    __do_sys_io_uring_enter+0x7aa/0x1dc0

With this fix applied, the reproducer completes without the KASAN report.

I will continue to evaluate the suggested refactoring and applying 774656479397
to blk_mq_alloc_cached_request(), and confirm in v2..

Thanks,
Henry Hu.

>
>>   	rq_list_pop(&plug->cached_rqs);
>>   	return rq;
>>   }

  reply	other threads:[~2026-09-17  5:52 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-16 15:16 [PATCH] blk-mq: check passthrough state before reusing cached requests huhai
2026-09-16 19:19 ` John Garry
2026-09-17  5:52   ` huhai [this message]
2026-09-17 14:21 ` Keith Busch
2026-09-18  5:16   ` huhai

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=4deafc69.49ba.1a0adec3e0f.Coremail.15815827059@163.com \
    --to=15815827059@163.com \
    --cc=axboe@kernel.dk \
    --cc=huhai@kylinos.cn \
    --cc=john.garry@linux.dev \
    --cc=linux-block@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox