From: Saravanan D <saravanand@crusoe.ai>
To: Keith Busch <kbusch@kernel.org>
Cc: Saravanan D <saravanand@crusoe.ai>,
Sagi Grimberg <sagi@grimberg.me>, Keith Busch <kbusch@meta.com>,
linux-nvme@lists.infradead.org, Christoph Hellwig <hch@lst.de>,
Jens Axboe <axboe@kernel.dk>
Subject: Re: [PATCH RFC] nvme-tcp: allow multiple queues per hctx
Date: Tue, 8 Sep 2026 22:21:09 -0700 [thread overview]
Message-ID: <20260909052111.59944-1-saravanand@crusoe.ai> (raw)
In-Reply-To: <aqBBgzV1Q9pP7Z44@kbusch-mbp>
On Tue, 8 Sep 2026 11:10:27 -0600 Keith Busch <kbusch@kernel.org> wrote:
> Initial results show there's some promise to this suggestion, however, I
> think in addition to the wq_unbound, it looks like I also need to adjust
> the "io_cpu" to change to the submitter's CPU. I incorporated that in
> this RFC, but there's also a proposal specifically for that here:
>
> https://lore.kernel.org/linux-nvme/20260820083634.71689-1-saravanand@crusoe.ai/
>
> I need to catch up on the discussion there, but from what I can tell,
> the cpu hint provided to the queue work appears to be important.
A short summary to save you reading the whole thread. v1 and v2 adopted
the submitting CPU as io_cpu for every command except the fabrics
Connect, opt in per controller. The motivation is a multi tenant host
with more CPUs than controller queues, where the connect time pick can
land a queue's socket work on CPUs owned by a different tenant.
On a 384 cpu host whose controllers expose 128 io queues, blk-mq folds
three cpus into every map group, 9% of nvme_tcp_io_work executions ran
outside the submitting VM's cpuset by default, and adoption brought
99.99% back, so the cpu hint matters in our case as well.
Sagi suggested replacing the heuristic in the driver with a per queue
writable sysfs attribute so a control plane can set io_cpu exactly, and
my v3 submission will implement that. The user assignment is kept across
reconnects and writing -1 reverts to the connect time selection.
I will post it shortly.
Thanks,
Saravanan D.
next prev parent reply other threads:[~2026-09-09 5:21 UTC|newest]
Thread overview: 9+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-03 15:26 [PATCH RFC] nvme-tcp: allow multiple queues per hctx Keith Busch
2026-09-06 0:04 ` Sagi Grimberg
2026-09-08 14:21 ` Keith Busch
2026-09-08 17:10 ` Keith Busch
2026-09-09 5:21 ` Saravanan D [this message]
2026-09-11 21:44 ` Sagi Grimberg
2026-09-21 12:21 ` Hannes Reinecke
2026-09-21 17:29 ` Keith Busch
2026-10-06 11:03 ` Keith Busch
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260909052111.59944-1-saravanand@crusoe.ai \
--to=saravanand@crusoe.ai \
--cc=axboe@kernel.dk \
--cc=hch@lst.de \
--cc=kbusch@kernel.org \
--cc=kbusch@meta.com \
--cc=linux-nvme@lists.infradead.org \
--cc=sagi@grimberg.me \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.