From: Thomas Gleixner <tglx@linutronix.de>
To: Christoph Hellwig <hch@lst.de>
Cc: Keith Busch <keith.busch@intel.com>,
axboe@fb.com, linux-block@vger.kernel.org,
linux-nvme@lists.infradead.org
Subject: Re: [PATCH 4/7] blk-mq: allow the driver to pass in an affinity mask
Date: Wed, 7 Sep 2016 17:38:40 +0200 (CEST) [thread overview]
Message-ID: <alpine.DEB.2.20.1609071025360.5647@nanos> (raw)
In-Reply-To: <20160906165056.GB26214@lst.de>
On Tue, 6 Sep 2016, Christoph Hellwig wrote:
> [adding Thomas as it's about the affinity_mask he (we) added to the
> IRQ core]
>
> On Tue, Sep 06, 2016 at 10:39:28AM -0400, Keith Busch wrote:
> > > Always the previous one. Below is a patch to get us back to the
> > > previous behavior:
> >
> > No, that's not right.
> >
> > Here's my topology info:
> >
> > # numactl --hardware
> > available: 2 nodes (0-1)
> > node 0 cpus: 0 1 2 3 4 5 6 7 16 17 18 19 20 21 22 23
> > node 0 size: 15745 MB
> > node 0 free: 15319 MB
> > node 1 cpus: 8 9 10 11 12 13 14 15 24 25 26 27 28 29 30 31
> > node 1 size: 16150 MB
> > node 1 free: 15758 MB
> > node distances:
> > node 0 1
> > 0: 10 21
> > 1: 21 10
>
> How do you get that mapping? Does this CPU use Hyperthreading and
> thus expose siblings using topology_sibling_cpumask? As that's the
> only thing the old code used for any sort of special casing.
That's a normal Intel mapping with two sockets and HT enabled. The cpu
enumeration is
Socket0 - physical cores
Socket1 - physical cores
Socket0 - HT siblings
Socket1 - HT siblings
> I'll need to see if I can find a system with such a mapping to reproduce.
Any 2 socket Intel with HT enabled will do. If you need access to one let
me know.
> > If I have 16 vectors, the affinity_mask generated by what you're doing
> > looks like 0000ffff, CPU's 0-15. So the first 16 bits are set since each
> > of those are the first unique CPU, getting a unique vector just like you
> > wanted. If an unset bit just means share with the previous, then all of
> > my thread siblings (CPU's 16-31) get to share with CPU 15. That's awful!
> >
> > What we want for my CPU topology is the 16th CPU to pair with CPU 0,
> > 17 pairs with 1, 18 with 2, and so on. You can't convey that information
> > with this scheme. We need affinity_masks per vector.
>
> We actually have per-vector masks, but they are hidden inside the IRQ
> core and awkward to use. We could to the get_first_sibling magic
> in the block-mq queue mapping (and in fact with the current code I guess
> we need to). Or take a step back from trying to emulate the old code
> and instead look at NUMA nodes instead of siblings which some folks
> suggested a while ago.
I think you want both.
NUMA nodes are certainly the first decision factor. You split the number of
vectors to the nodes:
vecs_per_node = num_vector / num_nodes;
Then you spread the number of vectors per node by the number of cpus per
node.
cpus_per_vec = cpus_on(node) / vecs_per_node;
If the number of cpus per vector is <= 1 you just use a round robin
scheme. If not, you need to look at siblings.
Looking at the whole thing, I think we need to be more clever when setting
up the msi descriptor affinity masks.
I'll send a RFC series soon.
Thanks,
tglx
next prev parent reply other threads:[~2016-09-07 15:41 UTC|newest]
Thread overview: 19+ messages / expand[flat|nested] mbox.gz Atom feed top
2016-08-29 10:53 blk-mq: allow passing in an external queue mapping V2 Christoph Hellwig
2016-08-29 10:53 ` [PATCH 1/7] blk-mq: don't redistribute hardware queues on a CPU hotplug event Christoph Hellwig
2016-08-29 10:53 ` [PATCH 2/7] blk-mq: only allocate a single mq_map per tag_set Christoph Hellwig
2016-08-29 10:53 ` [PATCH 3/7] blk-mq: remove ->map_queue Christoph Hellwig
2016-08-29 10:53 ` [PATCH 4/7] blk-mq: allow the driver to pass in an affinity mask Christoph Hellwig
2016-08-31 16:38 ` Keith Busch
2016-09-01 8:46 ` Christoph Hellwig
2016-09-01 14:24 ` Keith Busch
2016-09-01 23:30 ` Keith Busch
2016-09-05 19:48 ` Christoph Hellwig
2016-09-06 14:39 ` Keith Busch
2016-09-06 16:50 ` Christoph Hellwig
2016-09-06 17:30 ` Keith Busch
2016-09-07 15:38 ` Thomas Gleixner [this message]
2016-08-29 10:53 ` [PATCH 5/7] nvme: switch to use pci_alloc_irq_vectors Christoph Hellwig
2016-08-29 10:53 ` [PATCH 6/7] nvme: remove the post_scan callout Christoph Hellwig
2016-08-29 10:53 ` [PATCH 7/7] blk-mq: get rid of the cpumask in struct blk_mq_tags Christoph Hellwig
2016-08-30 23:28 ` blk-mq: allow passing in an external queue mapping V2 Keith Busch
2016-09-01 8:45 ` Christoph Hellwig
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=alpine.DEB.2.20.1609071025360.5647@nanos \
--to=tglx@linutronix.de \
--cc=axboe@fb.com \
--cc=hch@lst.de \
--cc=keith.busch@intel.com \
--cc=linux-block@vger.kernel.org \
--cc=linux-nvme@lists.infradead.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox