From: John Garry <john.garry@linux.dev>
To: huhai <15815827059@163.com>, axboe@kernel.dk
Cc: linux-block@vger.kernel.org, Henry Hu <huhai@kylinos.cn>
Subject: Re: [PATCH] blk-mq: check passthrough state before reusing cached requests
Date: Wed, 16 Sep 2026 20:19:58 +0100 [thread overview]
Message-ID: <18a634dd-35fc-4f33-a7c2-4c72d210b0bb@linux.dev> (raw)
In-Reply-To: <20260916151655.2588-1-15815827059@163.com>
On 9/16/26 16:16, huhai wrote:
> From: Henry Hu <huhai@kylinos.cn>
>
> blk_mq_alloc_request() and blk_mq_submit_bio() share plug->cached_rqs,
> but neither lookup checks the passthrough state. This allows an io_uring
> batch to reuse a cached NVMe passthrough request for an ordinary bio.
>
> When an I/O scheduler is enabled, passthrough requests have RQF_SCHED_TAGS
> set but RQF_USE_SCHED clear, and skip ->prepare_request(). Reuse updates
> rq->cmd_flags without reinitializing the scheduler state, so the request
> is inserted into the scheduler on plug flush, while completion skips
> ->finish_request(). With kyber, this leaks the domain token acquired at
> dispatch and can stall further I/O in that domain.
>
> This was reproduced with kyber and request merging disabled on a QEMU
> NVMe device by submitting a passthrough read followed by O_DIRECT writes
> in one io_uring batch. Tokens leak until the kyber write-domain token pool
> is exhausted, blocking ext4 journal I/O and fsync():
>
> INFO: task jbd2/nvme0n1-8:97 blocked in I/O wait for more than 241 seconds.
> Call Trace:
> schedule+0xe9/0x300
> io_schedule+0xca/0x150
> bit_wait_io+0x1b/0x140
> __wait_on_bit+0x63/0x170
> jbd2_write_superblock+0x3d7/0x600
> jbd2_journal_update_sb_log_tail+0x1e3/0x2d0
> jbd2_journal_commit_transaction+0x11d1/0x5b50
>
> INFO: task fsynctest:107 blocked in I/O wait for more than 241 seconds.
> Call Trace:
> schedule+0xe9/0x300
> io_schedule+0xca/0x150
> folio_wait_bit_common+0x2ed/0x780
> folio_wait_bit+0x18/0x30
> folio_wait_writeback+0x4b/0x1e0
> __filemap_fdatawait_range+0x120/0x1e0
> file_write_and_wait_range+0xeb/0x130
> ext4_sync_file+0x368/0xa80
> do_fsync+0xa2/0x200
>
> Require the passthrough state to match before reusing a cached request.
> With the fix, the workload completes normally.
>
> Assisted-by: LLM
> Signed-off-by: Henry Hu <huhai@kylinos.cn>
> ---
> block/blk-mq.c | 4 ++++
> 1 file changed, 4 insertions(+)
>
> diff --git a/block/blk-mq.c b/block/blk-mq.c
> index a26a11c73ee3..3a9d57f6a569 100644
> --- a/block/blk-mq.c
> +++ b/block/blk-mq.c
> @@ -648,6 +648,8 @@ static struct request *blk_mq_alloc_cached_request(struct request_queue *q,
> return NULL;
> if (op_is_flush(rq->cmd_flags) != op_is_flush(opf))
> return NULL;
> + if (blk_rq_is_passthrough(rq) != blk_op_is_passthrough(opf))
> + return NULL;
>
> rq_list_pop(&plug->cached_rqs);
> blk_mq_rq_time_init(rq, blk_time_get_ns());
> @@ -3062,6 +3064,8 @@ static struct request *blk_mq_get_cached_request(struct blk_plug *plug,
> return NULL;
> if (op_is_flush(rq->cmd_flags) != op_is_flush(opf))
> return NULL;
> + if (blk_rq_is_passthrough(rq) != blk_op_is_passthrough(opf))
> + return NULL;
It seems like some code which code be factored out (between
blk_mq_alloc_cached_request() and blk_mq_get_cached_request()).
And maybe 7746564793978fe2f43b18a302b22dca0ad3a0e8 could to be repeated
for blk_mq_alloc_cached_request(), which could mean even more factoring out.
> rq_list_pop(&plug->cached_rqs);
> return rq;
> }
next prev parent reply other threads:[~2026-09-16 19:20 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-16 15:16 [PATCH] blk-mq: check passthrough state before reusing cached requests huhai
2026-09-16 19:19 ` John Garry [this message]
2026-09-17 5:52 ` huhai
2026-09-17 14:21 ` Keith Busch
2026-09-18 5:16 ` huhai
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=18a634dd-35fc-4f33-a7c2-4c72d210b0bb@linux.dev \
--to=john.garry@linux.dev \
--cc=15815827059@163.com \
--cc=axboe@kernel.dk \
--cc=huhai@kylinos.cn \
--cc=linux-block@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox