* [RFC] PCI_IRQ_AFFINITY limits MSI-X allocation on 384 CPU / 1000+ NVMe system
@ 2026-07-28 6:42 santhosh kumar
2026-07-28 22:52 ` Thomas Gleixner
0 siblings, 1 reply; 5+ messages in thread
From: santhosh kumar @ 2026-07-28 6:42 UTC (permalink / raw)
To: linux-pci; +Cc: linux-kernel, tglx, akpm, mingo, bp, dave.hansen
Observed:
- Linux 6.13
- 384 CPUs
- 1000+ NVMe devices
With PCI_IRQ_AFFINITY:
- some NVMe devices fail to obtain dedicated MSI-X vectors
Without PCI_IRQ_AFFINITY:
- all NVMe devices obtain 2 MSI-X vectors
Investigation suggests an interaction between:
- group_cpus_evenly()
- irq_create_affinity_masks()
- x86 vector allocation
Looking for feedback on whether affinity-constrained vector allocation
could explain this behavior.
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [RFC] PCI_IRQ_AFFINITY limits MSI-X allocation on 384 CPU / 1000+ NVMe system
2026-07-28 6:42 [RFC] PCI_IRQ_AFFINITY limits MSI-X allocation on 384 CPU / 1000+ NVMe system santhosh kumar
@ 2026-07-28 22:52 ` Thomas Gleixner
2026-07-29 6:29 ` santhosh kumar
0 siblings, 1 reply; 5+ messages in thread
From: Thomas Gleixner @ 2026-07-28 22:52 UTC (permalink / raw)
To: santhosh kumar, linux-pci; +Cc: linux-kernel, akpm, mingo, bp, dave.hansen
On Tue, Jul 28 2026 at 12:12, santhosh kumar wrote:
> Observed:
> - Linux 6.13
> - 384 CPUs
> - 1000+ NVMe devices
>
>
> With PCI_IRQ_AFFINITY:
> - some NVMe devices fail to obtain dedicated MSI-X vectors
>
>
> Without PCI_IRQ_AFFINITY:
> - all NVMe devices obtain 2 MSI-X vectors
>
> Investigation suggests an interaction between:
> - group_cpus_evenly()
> - irq_create_affinity_masks()
> - x86 vector allocation
>
> Looking for feedback on whether affinity-constrained vector allocation
> could explain this behavior.
Math perhaps?
Number of CPUs: 384
Total number of available device vectors: ~ (200 * 384) = 76800
NVMe devices try to allocate min(nr_queues, NR_CPUS) queues where each
queue requires a dedicated interrupt. Add the managament queue to it and
then it's obvious that the total number of required vectors is larger
than the number of available vectors in the system when the number of
devices gets large enough.
Nothing to see here. It's simply resource exhaustion.
If you want that odd setup to be supported you have to talk to the NVME
people and not to a random list of folks which have absolutely nothing
to do with NVME.
Thanks,
tglx
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [RFC] PCI_IRQ_AFFINITY limits MSI-X allocation on 384 CPU / 1000+ NVMe system
2026-07-28 22:52 ` Thomas Gleixner
@ 2026-07-29 6:29 ` santhosh kumar
2026-07-30 13:16 ` Thomas Gleixner
0 siblings, 1 reply; 5+ messages in thread
From: santhosh kumar @ 2026-07-29 6:29 UTC (permalink / raw)
To: Thomas Gleixner; +Cc: linux-pci, linux-kernel, akpm, mingo, bp, dave.hansen
On Wed, Jul 29, 2026 at 4:22 AM Thomas Gleixner <tglx@linutronix.de> wrote:
>
> On Tue, Jul 28 2026 at 12:12, santhosh kumar wrote:
> > Observed:
> > - Linux 6.13
> > - 384 CPUs
> > - 1000+ NVMe devices
> >
> >
> > With PCI_IRQ_AFFINITY:
> > - some NVMe devices fail to obtain dedicated MSI-X vectors
> >
> >
> > Without PCI_IRQ_AFFINITY:
> > - all NVMe devices obtain 2 MSI-X vectors
> >
> > Investigation suggests an interaction between:
> > - group_cpus_evenly()
> > - irq_create_affinity_masks()
> > - x86 vector allocation
> >
> > Looking for feedback on whether affinity-constrained vector allocation
> > could explain this behavior.
>
> Math perhaps?
>
> Number of CPUs: 384
>
> Total number of available device vectors: ~ (200 * 384) = 76800
>
> NVMe devices try to allocate min(nr_queues, NR_CPUS) queues where each
> queue requires a dedicated interrupt. Add the managament queue to it and
> then it's obvious that the total number of required vectors is larger
> than the number of available vectors in the system when the number of
> devices gets large enough.
>
> Nothing to see here. It's simply resource exhaustion.
>
> If you want that odd setup to be supported you have to talk to the NVME
> people and not to a random list of folks which have absolutely nothing
> to do with NVME.
>
I am trying to create 1024 nvme devices , each device requesting 2 msix vectors.
Total we need 1024*2=2048 msix vectors
Total number of available device vectors: ~ (200 * 384) = 76800
But still seeing following messages:
[ 9166.370948] Interrupt reservation exceeds available resources
[ 9167.455029] Interrupt reservation exceeds available resources
[ 9167.455572] Interrupt reservation exceeds available resources
[ 9167.455888] Interrupt reservation exceeds available resources
● Summary: Affinity Mask Generation in Linux 6.13
The affinity mask is generated in multiple layers, here's the exact flow:
Key Files:
1. lib/group_cpus.c:347 - group_cpus_evenly()
- This function creates per-interrupt CPU masks
- Spreads interrupts across CPUs in a NUMA-aware manner
- Each interrupt gets assigned to 1 or a small set of specific CPUs
2. lib/group_cpus.c:14 - grp_spread_init_one()
- Actually picks which specific CPUs go into each mask
- Tries to use sibling CPUs (same core) for cache locality
3. kernel/irq/affinity.c:26 - irq_create_affinity_masks()
- Calls group_cpus_evenly() to get CPU assignments
- Creates affinity descriptor array with masks
- Marks interrupts as "managed" (cannot be changed from userspace)
4. drivers/pci/msi/msi.c:669 - msix_setup_interrupts()
- Calls irq_create_affinity_masks() when PCI_IRQ_AFFINITY is set
- Attaches the masks to MSI-X descriptors
5. arch/x86/kernel/apic/vector.c:295 - assign_irq_vector_any_locked()
- Reads the affinity mask created above
- ONLY allocates from CPUs in that specific mask
- No fallback to other CPUs!
The Problem Code:
// lib/group_cpus.c - Creates restricted CPU masks
grp_spread_init_one()
{
cpu = cpumask_first(nmsk); // Pick CPU 68 for nvme0q0
cpumask_set_cpu(cpu, irqmsk); // Mask = {68} only!
}
// arch/x86/kernel/apic/vector.c - Can only use that mask
assign_irq_vector_any_locked()
{
affmsk = irq_data_get_affinity_mask(irqd); // Gets {68}
assign_vector_locked(irqd, affmsk); // ONLY tries CPU 68!
// ❌ No fallback to CPU 200-383 even if they have free vectors
}
This is why disabling PCI_IRQ_AFFINITY fixes the issue - it prevents
the restrictive masks from being created in the first place!
In Linux 6.13, PCI_IRQ_AFFINITY causes MSI-X vectors to receive
affinity masks generated through irq_create_affinity_masks() and
group_cpus_evenly(). On very large systems with hundreds of CPUs and
thousands of devices, these masks can become sufficiently
restrictive that vector allocation is limited to a relatively small
subset of CPUs. When those CPUs become vector-constrained, some NVMe
devices receive fewer MSI-X vectors or fall back to shared interrupt
usage. Disabling PCI_IRQ_AFFINITY removes these affinity constraints
and allows the allocator to place vectors more flexibly, which
explains why all NVMe devices successfully obtain two MSI-X vectors in
the same configuration
> Thanks,
>
> tglx
>
>
>
>
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [RFC] PCI_IRQ_AFFINITY limits MSI-X allocation on 384 CPU / 1000+ NVMe system
2026-07-29 6:29 ` santhosh kumar
@ 2026-07-30 13:16 ` Thomas Gleixner
2026-07-30 13:28 ` Keith Busch
0 siblings, 1 reply; 5+ messages in thread
From: Thomas Gleixner @ 2026-07-30 13:16 UTC (permalink / raw)
To: santhosh kumar
Cc: linux-kernel, Christoph Hellwig, Keith Busch, Ming Lei, x86
On Wed, Jul 29 2026 at 11:59, santhosh kumar wrote:
> On Wed, Jul 29, 2026 at 4:22 AM Thomas Gleixner <tglx@linutronix.de> wrote:
>> Nothing to see here. It's simply resource exhaustion.
>>
>> If you want that odd setup to be supported you have to talk to the NVME
>> people and not to a random list of folks which have absolutely nothing
>> to do with NVME.
I've cleaned up the CC list for you because your AI buddy ignored that
request.
> I am trying to create 1024 nvme devices , each device requesting 2 msix vectors.
> Total we need 1024*2=2048 msix vectors
> Total number of available device vectors: ~ (200 * 384) = 76800
> But still seeing following messages:
> [ 9166.370948] Interrupt reservation exceeds available resources
> [ 9167.455029] Interrupt reservation exceeds available resources
> [ 9167.455572] Interrupt reservation exceeds available resources
> [ 9167.455888] Interrupt reservation exceeds available resources
>
> ● Summary: Affinity Mask Generation in Linux 6.13
Copying and pasting the output of your AI buddy is a pretty useless
exercise simply because your AI buddy does not understand how all of
this works. Neither did you actually validate that his slop makes any
sense.
Aside of that reports want to be against the latest upstream kernel and
not against something which is 8 revisions behind and eventually lacks a
ton of updates.
But let me explain you why your AI buddy gets it wrong.
> The affinity mask is generated in multiple layers, here's the exact flow:
We all know how that works even without the wisdom of your AI buddy.
> The Problem Code:
> // lib/group_cpus.c - Creates restricted CPU masks
> grp_spread_init_one()
> {
> cpu = cpumask_first(nmsk); // Pick CPU 68 for nvme0q0
> cpumask_set_cpu(cpu, irqmsk); // Mask = {68} only!
> }
See below.
> // arch/x86/kernel/apic/vector.c - Can only use that mask
> assign_irq_vector_any_locked()
> {
> affmsk = irq_data_get_affinity_mask(irqd); // Gets {68}
> assign_vector_locked(irqd, affmsk); // ONLY tries CPU 68!
> // ❌ No fallback to CPU 200-383 even if they have free vectors
Again your AI buddy is wrong.
assign_irq_vector_any_locked() has a fallback if the mask does not
result in an allocatable interrupt. It tries to get a close vector and
falls through to the very end of the function, which does:
/* Try the full online mask */
return assign_vector_locked(irqd, cpu_online_mask);
That's why the function has _any_ in the name. This function has a full
fallback. and that function is irrelevant for managed interrupts. It
only is invoked for non-managed interrupts.
Managed interrupts go through a different code path and that has no
fallback by design because that's the fundamental property of managed
interrupts.
So if group_spread_one() inits the mask for the managed interrupt with
only CPU 68 set then there is no fallback because there is no other
choice. But that also means that the driver did allocate more than two
vectors. Why?
If it really only allocates only two, then one is non-managed and the
other one is managed. In that case both end up with the a cpumask of
0-383, because spreading _ONE_ interrupt over 384 CPUs simply results in
a affinity mask with all CPUs set.
Just for illustration with a system with 64 CPUs and two nodes, where
CPU 0-15 and 32-47 are on node 0, CPU 16-31 and 48-63 are on node 1, the
spreading results in:
nr interrupts | nr_masks | masks
-----------------------------------------------------------------------
1 | 1 | [0-63]
2 | 2 | [0-15,32-47] [16-31,48-63]
4 | 4 | [0-7,32-39] [8-15,40-47] [16-23,48-55] ...
...
32 | 32 | [0,32] [1,33] ... [16,48] [17,49] ...
64 | 64 | [0] [1] [2] ....
That's the basic principle of the spreading mechanism. It's pretty
obvious, no?
Can you now explain me how that mechanism ends up creating a affinity
mask with a single CPU set if there is only _ONE_ interrupt to be
spreaded out?
I doubt it, but I can explain to you what happens in principle with NVME
and managed interrupts independent of the number of queues per device.
Managed interrupts ensure on startup, that for each CPU in the affinity
mask of each managed interrupt there is a vector guaranteed
available. That means:
nr interrupts | nr_masks | CPUs per mask | vectors | vectors
| | | per CPU | total
------------------------------------------------------------------------
1 | 1 | 64 | 1 | 64
2 | 2 | 32 | 1 | 64
4 | 4 | 16 | 1 | 64
...
32 | 32 | 2 | 1 | 64
64 | 64 | 1 | 1 | 64
So depending on the other interrupts allocated this will fail somewhere
around 200 devices for sure.
I'm sure it's the same on your side because that's independent of the
number of CPUs, even if you failed to provide that information.
As I told you before it's simple math and resource exhaustion due to the
way how NVME is implemented and the x86 vector space limitations.
If you want that to behave differently, then you have to talk to the
NVME people as I told you before.
Thanks,
tglx
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [RFC] PCI_IRQ_AFFINITY limits MSI-X allocation on 384 CPU / 1000+ NVMe system
2026-07-30 13:16 ` Thomas Gleixner
@ 2026-07-30 13:28 ` Keith Busch
0 siblings, 0 replies; 5+ messages in thread
From: Keith Busch @ 2026-07-30 13:28 UTC (permalink / raw)
To: Thomas Gleixner
Cc: santhosh kumar, linux-kernel, Christoph Hellwig, Ming Lei, x86
On Thu, Jul 30, 2026 at 03:16:46PM +0200, Thomas Gleixner wrote:
> As I told you before it's simple math and resource exhaustion due to the
> way how NVME is implemented and the x86 vector space limitations.
>
> If you want that to behave differently, then you have to talk to the
> NVME people as I told you before.
We can introduce a module parameter to throttle down the maximum number
of IO queues to allocate per controller. I don't think the driver can
automatically reason out what the correct number should be because it
doesn't know how many devices it's going to see.
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-07-30 13:28 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-28 6:42 [RFC] PCI_IRQ_AFFINITY limits MSI-X allocation on 384 CPU / 1000+ NVMe system santhosh kumar
2026-07-28 22:52 ` Thomas Gleixner
2026-07-29 6:29 ` santhosh kumar
2026-07-30 13:16 ` Thomas Gleixner
2026-07-30 13:28 ` Keith Busch
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.