From: Keith Busch <kbusch@kernel.org>
To: Hannes Reinecke <hare@suse.de>
Cc: Keith Busch <kbusch@meta.com>,
linux-nvme@lists.infradead.org, sagi@grimberg.me, hch@lst.de,
axboe@kernel.dk
Subject: Re: [PATCH RFC] nvme-tcp: allow multiple queues per hctx
Date: Mon, 21 Sep 2026 11:29:40 -0600 [thread overview]
Message-ID: <arFphFDlyHpnCv0J@kbusch-mbp> (raw)
In-Reply-To: <db8c19f8-d087-4e3c-8424-6a90760cdf16@suse.de>
On Mon, Sep 21, 2026 at 02:21:28PM +0200, Hannes Reinecke wrote:
> Interesting.
> We have so far refrained from similar attempts as the idea was that
> one CPU would be able to saturate the available bandwidth; additionally
> nvme_tcp_io_work() would be gobbling up any available data, so scheduling
> between two instances on the same queue would be pointless.
PCIe and RDMA can saturate the link with a single queue, but not TCP.
Note, this RFC specifically schedules multiple queues, not the same
queue. It is the same blk-mq hctx, but that fans out to different nvme
queues.
> So question would be: where does the speed up come from?
I have the workqueue unbounded, so if a batch of 4 requests comes in on
one thread, they get worked on in parallel on 4 different CPUs utilizing
different sockets. This also exploits NIC parallelisms via multiple RSS
buckets that wouldn't happen with single socket usage.
> In general not a bad idea. Especially if it helps to up our performance.
> But this really points to the same problem we're having with HW RAID
> HBAs: we have an issue if the per-queue I/O performance is below what
> the cpu can drive. Then it _does_ make sense to have more queues than\ CPUs,
> but the layout of which is beyond what blk-mq can handle.
> Ideally we should be able to handle that via blk-mq, too.
I'm still trying to see if we can achieve a similar result just by
making duplicate connections and letting the multipath policy handle the
spread. That does improve things significantly, but I'm getting worse
performance than this RFC, and I still don't know why yet. I didn't get
to work on this last week, so it's still on me to work out where it's
lacking.
> Nice topic for ALPSS ...
Thanks! We'll definitely touch on this next week.
next prev parent reply other threads:[~2026-09-21 17:29 UTC|newest]
Thread overview: 9+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-03 15:26 [PATCH RFC] nvme-tcp: allow multiple queues per hctx Keith Busch
2026-09-06 0:04 ` Sagi Grimberg
2026-09-08 14:21 ` Keith Busch
2026-09-08 17:10 ` Keith Busch
2026-09-09 5:21 ` Saravanan D
2026-09-11 21:44 ` Sagi Grimberg
2026-09-21 12:21 ` Hannes Reinecke
2026-09-21 17:29 ` Keith Busch [this message]
2026-10-06 11:03 ` Keith Busch
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=arFphFDlyHpnCv0J@kbusch-mbp \
--to=kbusch@kernel.org \
--cc=axboe@kernel.dk \
--cc=hare@suse.de \
--cc=hch@lst.de \
--cc=kbusch@meta.com \
--cc=linux-nvme@lists.infradead.org \
--cc=sagi@grimberg.me \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.