From: Naman Jain <namjain@linux.microsoft.com>
To: Michael Kelley <mhklinux@outlook.com>,
Andrew Morton <akpm@linux-foundation.org>,
Thomas Gleixner <tglx@kernel.org>,
Ming Lei <tom.leiming@gmail.com>, Ming Lei <ming.lei@redhat.com>
Cc: Wangyang Guo <wangyang.guo@intel.com>,
Tianyou Li <tianyou.li@intel.com>,
Tim Chen <tim.c.chen@linux.intel.com>,
Long Li <longli@microsoft.com>,
"linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>,
"linux-hyperv@vger.kernel.org" <linux-hyperv@vger.kernel.org>
Subject: Re: [PATCH v2] lib/group_cpus: rotate extra groups to avoid IRQ stacking
Date: Mon, 31 Aug 2026 11:01:03 +0530 [thread overview]
Message-ID: <fa1b65fb-3f77-4089-8928-e4b8c656759f@linux.microsoft.com> (raw)
In-Reply-To: <SN6PR02MB4157F84374BFA68158E0D85AD4AA2@SN6PR02MB4157.namprd02.prod.outlook.com>
On 8/31/2026 5:27 AM, Michael Kelley wrote:
> From: Naman Jain <namjain@linux.microsoft.com> Sent: Thursday, August 27, 2026 2:11 AM
>>
>> On 8/26/2026 7:31 AM, Michael Kelley wrote:
>>> From: Naman Jain <namjain@linux.microsoft.com> Sent: Sunday, August 9, 2026 11:22 PM
>>>
>>> [snip]
>>>
>>>> @@ -510,6 +633,8 @@ struct cpumask *group_cpus_evenly(unsigned int numgrps, unsigned int *nummasks)
>>>> if (!masks)
>>>> goto fail_node_to_cpumask;
>>>>
>>>> + spread_offset = (unsigned int)atomic_fetch_inc(&group_spread_cnt);
>>>> +
>>>> build_node_to_cpumask(node_to_cpumask);
>>>>
>>>
>>> One additional observation: In my testing, group_cpus_evenly() is
>>> often called with numgrps set to 1. This happens in the block "loop"
>>> devices (drivers/block/loop.c) and for the NVMe admin queue. In
>>> these cases, the spread_offset is never used, but group_spread_cnt
>>> gets incremented anyway. Incrementing for NVMe admin queues
>>> tends to dirty the spreading for multiple NVMe devices with the
>>> same configuration because it is usually interleaved with the
>>> spreading of the main NVMe I/O queues.
>>>
>>> To improve this, I changed the above code to this:
>>>
>>> + if (numgrps == 1)
>>> + spread_offset = 0;
>>> + else
>>> + spread_offset = (unsigned int)atomic_fetch_inc(&group_spread_cnt);
>>>
>>> With this change, my configuration #1 (Azure L48s v2 VM) is noticeably
>>> better. All CPUs in NUMA node 1 have either 3 or 4 IRQs assigned. NUMA
>>> node 0 ranges from 3 to 5 IRQs, but that's partly because the NUMA
>>> nodes themselves aren't balanced, as previously discussed. With your
>>> change to apply group_spread_cnt to the NUMA nodes, and my change
>>> above, my config #1 is likely to work out very near optimal. Of course,
>>> there's no guarantee that some other device won't increment
>>> group_spread_cnt and dirty things, but for the typical case it probably
>>> works very well.
>>>
>>
>> Thanks for the suggestion, I tried this and it works fine.
>>
>>> This change to skip incrementing group_spread_cnt when numgrps == 1
>>> doesn't help my arm64 configs. I'm still thinking about ways to do better
>>> when there aren't any clusters. I have an idea that I'm experimenting
>>> with, but it may be a few more days before I reach any conclusions.
>>>
>
> I finally figured out why arm64 is different from x86. By adding a
> kernel boot line parameter for controlling the number of NVMe queues
> per controller (for experimentation only), I could construct identical
> configs on x86 and arm64. And even though the clustering is
> different, group_cpus_evenly() returns exactly the same set of
> cpumasks on the two architectures. So the clustering difference
> isn't the reason for the poor spreading on arm64.
>
> My experiments have been mostly cases where the number of NVMe
> queues is small compared with the number of CPUs -- e.g., 6 NVMe
> queues on each controller in a VM with 96 CPUs. In this example,
> there are 16 CPUs in the cpumask for each queue. That set of 16
> CPUs is the smp_affinity for the IRQ and is the same for both
> architectures. But determining the single CPU that is the
> effective_affinity is different. The x86 APIC vector code must load
> balance assignments across CPUs because each x86 CPU has a
> limited number of vectors available. At a result, x86 spreads out
> which CPU in the set of 16 becomes the effective_affinity. But arm64
> does not. The GICv3 function gic_set_affinity() always picks the
> 1st CPU in the set of 16 CPUs. So the NVMe IRQs get stacked on
> the same 6 CPUs and the other 90 CPUs get none.
>
> If the number of queues doesn't evenly divide into the number of
> CPUs, then your patch provides a modest amount of spreading in
> how the cpumasks are constructed, and the interrupt load gets
> slightly more spread on the arm64 CPUs, but not nearly as well
> as on x86.
>
> I was previously aware of the vector load balancing done by
> x86, but had never compared with what arm64 does. I had
> to run a few experiments to have the light bulb come on in my
> head. :-) But I learned something so it was time well spent.
>
> In any case, improving the spreading in such configurations
> on arm64 probably isn't the job of group_cpus_evenly(). The
> arm64 GIC code would need to do some spreading when
> picking the effective_affinity from the CPUs in the affinity mask.
>
> Michael
Thank you for looking into this and sharing this valuable information
Michael, it means a lot.
So for now, we have two enhancements - NUMA node level balancing and
additional (numgrps == 1) check.
I'll make these changes, evaluate them and send them in next version.
Regards,
Naman
next prev parent reply other threads:[~2026-08-31 5:31 UTC|newest]
Thread overview: 13+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-10 6:21 [PATCH v2] lib/group_cpus: rotate extra groups to avoid IRQ stacking Naman Jain
2026-08-10 6:32 ` sashiko-bot
2026-08-10 8:47 ` Naman Jain
2026-08-23 15:47 ` Michael Kelley
2026-08-24 14:35 ` Naman Jain
2026-08-26 2:01 ` Michael Kelley
2026-08-27 9:11 ` Naman Jain
2026-08-30 23:57 ` Michael Kelley
2026-08-31 5:31 ` Naman Jain [this message]
2026-09-01 0:02 ` Michael Kelley
2026-09-02 5:02 ` Naman Jain
2026-09-02 16:57 ` Michael Kelley
2026-09-03 3:49 ` Naman Jain
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=fa1b65fb-3f77-4089-8928-e4b8c656759f@linux.microsoft.com \
--to=namjain@linux.microsoft.com \
--cc=akpm@linux-foundation.org \
--cc=linux-hyperv@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=longli@microsoft.com \
--cc=mhklinux@outlook.com \
--cc=ming.lei@redhat.com \
--cc=tglx@kernel.org \
--cc=tianyou.li@intel.com \
--cc=tim.c.chen@linux.intel.com \
--cc=tom.leiming@gmail.com \
--cc=wangyang.guo@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.