* [RFC PATCH v3 01/13] lib/sbm: Introduce helpers for architectures to configure LLC properties
2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
2026-10-02 9:13 ` sashiko-bot
2026-10-07 5:43 ` Shrikanth Hegde
2026-10-01 19:28 ` [RFC PATCH v3 02/13] drivers/base/arch_topology: Add support for initializing sbm topology K Prateek Nayak
` (12 subsequent siblings)
13 siblings, 2 replies; 36+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
driver-core, Sudeep Holla, Greg Kroah-Hartman, Rafael J. Wysocki,
Danilo Krummrich, Huacai Chen, Thomas Bogendoerfer, Jiaxun Yang,
Madhavan Srinivasan, Heiko Carstens, Vasily Gorbik,
Alexander Gordeev, David S. Miller, Andreas Larsson,
Thomas Gleixner, Borislav Petkov, Dave Hansen, x86
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, Shrikanth Hegde, K Prateek Nayak, WANG Xuerui,
Michael Ellerman, Nicholas Piggin, Christophe Leroy,
Christian Borntraeger, Sven Schnelle, H. Peter Anvin
Subsequent commits will add arch/ hooks to configure the sparsebitmask
(sbm) topology - number of sbm instance and the maximum CPUs that a
single instance can represent.
Introduce sbm_set_topology() that helps architectures configure the sbm
topology. Also add arch_sbm_cpu_instance_id() which will be used to map
CPUs to the sbm instances during init to keep CPUs on the same instance
on the same sbm leaf.
Individual arch/ users will be added one at a time in the subsequent
commits.
Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
include/linux/sbm.h | 8 ++++++++
lib/Makefile | 2 +-
lib/sbm.c | 27 +++++++++++++++++++++++++++
3 files changed, 36 insertions(+), 1 deletion(-)
create mode 100644 include/linux/sbm.h
create mode 100644 lib/sbm.c
diff --git a/include/linux/sbm.h b/include/linux/sbm.h
new file mode 100644
index 000000000000..adac12ed233a
--- /dev/null
+++ b/include/linux/sbm.h
@@ -0,0 +1,8 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+#ifndef _LINUX_SBM_H
+#define _LINUX_SBM_H
+
+int arch_sbm_cpu_instance_id(int cpu);
+void sbm_set_topology(int num_instances, int max_threads_per_instance);
+
+#endif /* _LINUX_SBM_H */
diff --git a/lib/Makefile b/lib/Makefile
index dfab958327c5..7776330f5b25 100644
--- a/lib/Makefile
+++ b/lib/Makefile
@@ -40,7 +40,7 @@ lib-y := ctype.o string.o vsprintf.o cmdline.o \
is_single_threaded.o plist.o decompress.o kobject_uevent.o \
earlycpio.o seq_buf.o siphash.o dec_and_lock.o \
nmi_backtrace.o win_minmax.o memcat_p.o \
- buildid.o objpool.o iomem_copy.o sys_info.o
+ buildid.o objpool.o iomem_copy.o sys_info.o sbm.o
lib-$(CONFIG_UNION_FIND) += union_find.o
lib-$(CONFIG_PRINTK) += dump_stack.o
diff --git a/lib/sbm.c b/lib/sbm.c
new file mode 100644
index 000000000000..82280e3306de
--- /dev/null
+++ b/lib/sbm.c
@@ -0,0 +1,27 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+#include <linux/sbm.h>
+#include <linux/init.h>
+#include <linux/printk.h>
+
+static int sbm_max_threads_per_instance = -1;
+static int sbm_num_instance = -1;
+
+/*
+ * In absence of an arch definition, consider all CPUs to
+ * belong to the same sbm instance.
+ */
+int __weak arch_sbm_cpu_instance_id(int cpu)
+{
+ return 0;
+}
+
+void __init sbm_set_topology(int num_instances, int max_threads_per_instance)
+{
+ sbm_max_threads_per_instance = max_threads_per_instance;
+ sbm_num_instance = num_instances;
+
+ pr_info("sbm topology set to %d instance(s) with a maximum of %d thread(s) per instance\n",
+ sbm_num_instance,
+ sbm_max_threads_per_instance);
+}
+
--
2.34.1
^ permalink raw reply related [flat|nested] 36+ messages in thread* Re: [RFC PATCH v3 01/13] lib/sbm: Introduce helpers for architectures to configure LLC properties
2026-10-01 19:28 ` [RFC PATCH v3 01/13] lib/sbm: Introduce helpers for architectures to configure LLC properties K Prateek Nayak
@ 2026-10-02 9:13 ` sashiko-bot
2026-10-07 5:43 ` Shrikanth Hegde
1 sibling, 0 replies; 36+ messages in thread
From: sashiko-bot @ 2026-10-02 9:13 UTC (permalink / raw)
To: K Prateek Nayak
Cc: Alexander Gordeev, Heiko Carstens, Vasily Gorbik,
Christian Borntraeger, linux-s390
> Subsequent commits will add arch/ hooks to configure the sparsebitmask
> (sbm) topology - number of sbm instance and the maximum CPUs that a
> single instance can represent.
>
> Introduce sbm_set_topology() that helps architectures configure the sbm
> topology. Also add arch_sbm_cpu_instance_id() which will be used to map
> CPUs to the sbm instances during init to keep CPUs on the same instance
> on the same sbm leaf.
>
> Individual arch/ users will be added one at a time in the subsequent
> commits.
>
> Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
Sashiko has reviewed this patch and found no issues. It looks great!
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261001192849.74788-1-kprateek.nayak@amd.com?part=1
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [RFC PATCH v3 01/13] lib/sbm: Introduce helpers for architectures to configure LLC properties
2026-10-01 19:28 ` [RFC PATCH v3 01/13] lib/sbm: Introduce helpers for architectures to configure LLC properties K Prateek Nayak
2026-10-02 9:13 ` sashiko-bot
@ 2026-10-07 5:43 ` Shrikanth Hegde
1 sibling, 0 replies; 36+ messages in thread
From: Shrikanth Hegde @ 2026-10-07 5:43 UTC (permalink / raw)
To: K Prateek Nayak
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, WANG Xuerui, Michael Ellerman,
Nicholas Piggin, Christophe Leroy, Christian Borntraeger,
Sven Schnelle, H. Peter Anvin, Peter Zijlstra, Chen Yu, Tim Chen,
Ingo Molnar, Juri Lelli, Vincent Guittot, Andrew Morton,
Arnd Bergmann, linux-kernel, linux-arch, linux-s390, linuxppc-dev,
linux-mips, loongarch, driver-core, Sudeep Holla,
Greg Kroah-Hartman, Rafael J. Wysocki, Danilo Krummrich,
Huacai Chen, Thomas Bogendoerfer, Jiaxun Yang,
Madhavan Srinivasan, Heiko Carstens, Vasily Gorbik,
Alexander Gordeev, David S. Miller, Andreas Larsson,
Thomas Gleixner, Borislav Petkov, Dave Hansen, x86
Hi Prateek.
I ran the series on 320 CPU power11 system today.
hackbench shows high regression across the different number of groups and
it is more severe at lower number of groups. I haven't run anything else yet.
I haven't taken a look at the patches yet. Post your talk today will try to
go through and see if i can spot something.
On 10/2/26 12:58 AM, K Prateek Nayak wrote:
> Subsequent commits will add arch/ hooks to configure the sparsebitmask
> (sbm) topology - number of sbm instance and the maximum CPUs that a
> single instance can represent.
>
> Introduce sbm_set_topology() that helps architectures configure the sbm
> topology. Also add arch_sbm_cpu_instance_id() which will be used to map
> CPUs to the sbm instances during init to keep CPUs on the same instance
> on the same sbm leaf.
>
> Individual arch/ users will be added one at a time in the subsequent
> commits.
>
> Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
> ---
> include/linux/sbm.h | 8 ++++++++
> lib/Makefile | 2 +-
> lib/sbm.c | 27 +++++++++++++++++++++++++++
> 3 files changed, 36 insertions(+), 1 deletion(-)
> create mode 100644 include/linux/sbm.h
> create mode 100644 lib/sbm.c
>
> diff --git a/include/linux/sbm.h b/include/linux/sbm.h
> new file mode 100644
> index 000000000000..adac12ed233a
> --- /dev/null
> +++ b/include/linux/sbm.h
> @@ -0,0 +1,8 @@
> +/* SPDX-License-Identifier: GPL-2.0 */
> +#ifndef _LINUX_SBM_H
> +#define _LINUX_SBM_H
> +
> +int arch_sbm_cpu_instance_id(int cpu);
> +void sbm_set_topology(int num_instances, int max_threads_per_instance);
> +
> +#endif /* _LINUX_SBM_H */
> diff --git a/lib/Makefile b/lib/Makefile
> index dfab958327c5..7776330f5b25 100644
> --- a/lib/Makefile
> +++ b/lib/Makefile
> @@ -40,7 +40,7 @@ lib-y := ctype.o string.o vsprintf.o cmdline.o \
> is_single_threaded.o plist.o decompress.o kobject_uevent.o \
> earlycpio.o seq_buf.o siphash.o dec_and_lock.o \
> nmi_backtrace.o win_minmax.o memcat_p.o \
> - buildid.o objpool.o iomem_copy.o sys_info.o
> + buildid.o objpool.o iomem_copy.o sys_info.o sbm.o
>
> lib-$(CONFIG_UNION_FIND) += union_find.o
> lib-$(CONFIG_PRINTK) += dump_stack.o
> diff --git a/lib/sbm.c b/lib/sbm.c
> new file mode 100644
> index 000000000000..82280e3306de
> --- /dev/null
> +++ b/lib/sbm.c
> @@ -0,0 +1,27 @@
> +/* SPDX-License-Identifier: GPL-2.0 */
> +#include <linux/sbm.h>
> +#include <linux/init.h>
> +#include <linux/printk.h>
> +
> +static int sbm_max_threads_per_instance = -1;
> +static int sbm_num_instance = -1;
> +
> +/*
> + * In absence of an arch definition, consider all CPUs to
> + * belong to the same sbm instance.
> + */
> +int __weak arch_sbm_cpu_instance_id(int cpu)
> +{
> + return 0;
> +}
> +
> +void __init sbm_set_topology(int num_instances, int max_threads_per_instance)
> +{
> + sbm_max_threads_per_instance = max_threads_per_instance;
> + sbm_num_instance = num_instances;
> +
> + pr_info("sbm topology set to %d instance(s) with a maximum of %d thread(s) per instance\n",
> + sbm_num_instance,
> + sbm_max_threads_per_instance);
> +}
nit: Stray empty line at the end.
> +
^ permalink raw reply [flat|nested] 36+ messages in thread
* [RFC PATCH v3 02/13] drivers/base/arch_topology: Add support for initializing sbm topology
2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 01/13] lib/sbm: Introduce helpers for architectures to configure LLC properties K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
2026-10-02 9:13 ` sashiko-bot
2026-10-01 19:28 ` [RFC PATCH v3 03/13] LoongArch: Initialize CPU _PXM relation for disabled CPUs from SRAT K Prateek Nayak
` (11 subsequent siblings)
13 siblings, 1 reply; 36+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
driver-core, Sudeep Holla, Greg Kroah-Hartman, Rafael J. Wysocki,
Danilo Krummrich
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, Shrikanth Hegde, K Prateek Nayak
Add (fragile) support for initializing sparsebitmap topology for
architectures that support GENERIC_ARCH_TOPOLOGY.
Similar to setting cpu_smt_set_num_threads(), count the number of unique
LLCs (if last_level_cache_is_valid() true, otherwise) / packages, and
the maximum threads that exist in those instance and set
sbm_set_proc_config().
Any failure on the path will default to considering the entire processor
as a single sbm domain and the default logic will divide the system into
BITS_PER_LONG chunks. In case of a failure, sbm core will skip using
arch_sbm_cpu_instance_id() and assume all CPUs belong to the same
instance with the same instance id.
Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
drivers/base/arch_topology.c | 122 ++++++++++++++++++++++++++++++++++-
1 file changed, 121 insertions(+), 1 deletion(-)
diff --git a/drivers/base/arch_topology.c b/drivers/base/arch_topology.c
index 8c5e47c28d9a..f55745a93298 100644
--- a/drivers/base/arch_topology.c
+++ b/drivers/base/arch_topology.c
@@ -20,6 +20,7 @@
#include <linux/cpumask.h>
#include <linux/init.h>
#include <linux/rcupdate.h>
+#include <linux/sbm.h>
#include <linux/sched.h>
#include <linux/units.h>
@@ -930,6 +931,114 @@ __weak int __init parse_acpi_topology(void)
return 0;
}
+/*
+ * Note: arch_sbm_cpu_instance_id() is only used if init_sbm_topology()
+ * below succeeds at setting the sbm topology. Either the LLC
+ * information is valid or package_id is dependable.
+ */
+int arch_sbm_cpu_instance_id(int cpu)
+{
+ if (last_level_cache_is_valid(cpu)) {
+ struct cacheinfo *llc_info = get_cpu_cacheinfo_llc(cpu);
+
+ if (!llc_info)
+ goto out;
+
+ if (llc_info->attributes & CACHE_ID)
+ return llc_info->id;
+
+ /*
+ * XXX: fw_token be truncated from cast
+ * when the value is returned as an int.
+ */
+ return (int)((long)llc_info->fw_token);
+ }
+out:
+ return cpu_topology[cpu].package_id;
+}
+
+static void __init init_sbm_topology(void)
+{
+ int num_sbm_instances = 0, max_threads_per_instance = -1;
+ bool has_cache = false, has_package = false;
+ cpumask_var_t unique_cpus;
+ struct xarray instances;
+ unsigned long cpu, *count;
+
+ /*
+ * If the allocation fails, the sbm core will use
+ * the default logic of splitting CPUs evenly in
+ * BITS_PER_LONG chunk.
+ */
+ if (!zalloc_cpumask_var(&unique_cpus, GFP_KERNEL))
+ return;
+
+ xa_init(&instances);
+
+ for_each_possible_cpu(cpu) {
+ bool found = false;
+ int unique_cpu;
+
+ for_each_cpu(unique_cpu, unique_cpus) {
+ /*
+ * XXX: Assumes last_level_cache_is_valid() is uniformly true
+ * across the entire system if it is true for one CPU.
+ */
+ if (last_level_cache_is_valid(cpu)) {
+ has_cache = true;
+ if (last_level_cache_is_shared(cpu, unique_cpu)) {
+ found = true;
+ break;
+ }
+ } else {
+ /* Go by package_id if no LLC information is found. */
+ has_package = true;
+ if (cpu_topology[unique_cpu].package_id ==
+ cpu_topology[cpu].package_id) {
+ found = true;
+ break;
+ }
+ }
+ }
+
+ /* XXX: CPUs should not suddenly switch IDs mid way. */
+ if (has_cache && has_package)
+ goto out;
+
+ if (!found) {
+ count = kzalloc_obj(*count);
+ if (!count)
+ goto out;
+
+ cpumask_set_cpu(cpu, unique_cpus);
+ *count += 1;
+
+ xa_store(&instances, cpu, count, GFP_KERNEL);
+ continue;
+ }
+
+ count = xa_load(&instances, unique_cpu);
+ if (!count)
+ goto out;
+
+ *count += 1;
+ }
+
+ xa_for_each(&instances, cpu, count) {
+ max_threads_per_instance = max_t(int, max_threads_per_instance, *count);
+ num_sbm_instances++;
+ }
+
+ sbm_set_topology(num_sbm_instances, max_threads_per_instance);
+out:
+ xa_for_each(&instances, cpu, count) {
+ xa_erase(&instances, cpu);
+ kfree(count);
+ }
+ xa_destroy(&instances);
+ free_cpumask_var(unique_cpus);
+}
+
void __init init_cpu_topology(void)
{
int cpu, ret;
@@ -954,8 +1063,19 @@ void __init init_cpu_topology(void)
continue;
else if (ret != -ENOENT)
pr_err("Early cacheinfo failed, ret = %d\n", ret);
- return;
+ break;
}
+
+ /*
+ * If fetch_cache_info() fails for first CPU,
+ * init_cpu_sbm_topology() will use pacakge_id instead.
+ *
+ * Uniform cache topology is a necessary since implementation
+ * assumes last_level_cache_is_valid() gives same result for
+ * all possible CPUs
+ */
+ if (!ret || (cpu == 0 && ret == -ENOENT))
+ init_sbm_topology();
}
void store_cpu_topology(unsigned int cpuid)
--
2.34.1
^ permalink raw reply related [flat|nested] 36+ messages in thread* Re: [RFC PATCH v3 02/13] drivers/base/arch_topology: Add support for initializing sbm topology
2026-10-01 19:28 ` [RFC PATCH v3 02/13] drivers/base/arch_topology: Add support for initializing sbm topology K Prateek Nayak
@ 2026-10-02 9:13 ` sashiko-bot
0 siblings, 0 replies; 36+ messages in thread
From: sashiko-bot @ 2026-10-02 9:13 UTC (permalink / raw)
To: K Prateek Nayak
Cc: Alexander Gordeev, Christian Borntraeger, linux-s390,
Heiko Carstens, Vasily Gorbik
Thank you for your contribution! Sashiko AI review found 5 potential issue(s) to consider:
- [High] topology: 64-bit fw_token truncated to 32-bit int causing ID collisions
- [High] topology: hardcoded 'cpu == 0' fails initialization when CPU 0 is not possible
- [Medium] topology: memory leak of 'count' on xa_store() failure
- [Medium] topology: topology uniformity check silently ignores the first CPU
- [Medium] topology: mismatched cache type filtering causes inconsistent SBM domains
--
commit 3748179aae956b072baf4016333e144175bba72f
Author: K Prateek Nayak <kprateek.nayak@amd.com>
Subject: drivers/base/arch_topology: Add support for initializing sbm topology
This commit introduces support for initializing sparsebitmap (sbm) topology
for architectures supporting GENERIC_ARCH_TOPOLOGY. It counts the number of
unique LLCs or packages and the maximum threads per instance to set the sbm
processor configuration.
> diff --git a/drivers/base/arch_topology.c b/drivers/base/arch_topology.c
> index 8c5e47c28d9a3..f55745a932986 100644
> --- a/drivers/base/arch_topology.c
> +++ b/drivers/base/arch_topology.c
[ ... ]
> @@ -930,6 +931,114 @@ __weak int __init parse_acpi_topology(void)
> int arch_sbm_cpu_instance_id(int cpu)
> {
> if (last_level_cache_is_valid(cpu)) {
> struct cacheinfo *llc_info = get_cpu_cacheinfo_llc(cpu);
>
> if (!llc_info)
> goto out;
>
> if (llc_info->attributes & CACHE_ID)
> return llc_info->id;
>
> /*
> * XXX: fw_token be truncated from cast
> * when the value is returned as an int.
> */
> return (int)((long)llc_info->fw_token);
[Severity: High]
Does this explicit truncation risk SBM topology corruption?
On 64-bit architectures using ACPI PPTT for cache topology, fw_token contains
a 64-bit pointer. By casting it down to a 32-bit int, the upper 32 bits are
dropped. If two different cache structures collide in their lower 32 bits,
CPUs that do not share an LLC might be incorrectly grouped into the same
instance ID.
> }
> out:
> return cpu_topology[cpu].package_id;
> }
[ ... ]
> static void __init init_sbm_topology(void)
> {
> int num_sbm_instances = 0, max_threads_per_instance = -1;
> bool has_cache = false, has_package = false;
> cpumask_var_t unique_cpus;
> struct xarray instances;
> unsigned long cpu, *count;
[ ... ]
> for_each_possible_cpu(cpu) {
> bool found = false;
> int unique_cpu;
>
> for_each_cpu(unique_cpu, unique_cpus) {
[Severity: Medium]
Will this silently ignore the first CPU for the topology uniformity check?
When the first CPU is processed, unique_cpus is empty, so this inner
for_each_cpu loop is completely skipped. Since has_cache and has_package
are only updated inside this loop, the cache state for the first CPU is
never recorded.
If a subsequent CPU has a conflicting topology, won't the has_cache &&
has_package check fail to detect the mismatch?
> /*
> * XXX: Assumes last_level_cache_is_valid() is uniformly true
> * across the entire system if it is true for one CPU.
> */
> if (last_level_cache_is_valid(cpu)) {
[Severity: Medium]
Are we using mismatched cache type filtering rules?
Here in init_sbm_topology(), last_level_cache_is_valid() is used, which does
not filter out instruction caches. However, arch_sbm_cpu_instance_id() uses
get_cpu_cacheinfo_llc(), which strictly filters out instruction caches and
falls back to package_id.
If the LLC is an instruction cache, won't the domain counting diverge from ID
assignment, leading to inconsistent SBM domains?
> has_cache = true;
> if (last_level_cache_is_shared(cpu, unique_cpu)) {
> found = true;
> break;
> }
[ ... ]
> if (!found) {
> count = kzalloc_obj(*count);
> if (!count)
> goto out;
>
> cpumask_set_cpu(cpu, unique_cpus);
> *count += 1;
>
> xa_store(&instances, cpu, count, GFP_KERNEL);
> continue;
[Severity: Medium]
Does this code leak the newly allocated count object if xa_store() fails?
The return value of xa_store() can be an error pointer if internal memory
allocation fails. If xa_store() fails here, count is not added to the xarray,
but we continue the loop without freeing it, resulting in a memory leak.
> }
[ ... ]
> void __init init_cpu_topology(void)
> {
> int cpu, ret;
[ ... ]
> /*
> * If fetch_cache_info() fails for first CPU,
> * init_cpu_sbm_topology() will use pacakge_id instead.
> *
> * Uniform cache topology is a necessary since implementation
> * assumes last_level_cache_is_valid() gives same result for
> * all possible CPUs
> */
> if (!ret || (cpu == 0 && ret == -ENOENT))
> init_sbm_topology();
[Severity: High]
Can this hardcoded check for cpu == 0 silently bypass SBM initialization if
CPU 0 is offline or not possible?
If fetch_cache_info() fails with -ENOENT for the first possible CPU, the
loop preceding this block breaks and leaves that CPU in the cpu variable.
If the first possible CPU is greater than 0, cpu == 0 evaluates to false,
and init_sbm_topology() will never be called.
Should this dynamically resolve the first valid CPU via
cpumask_first(cpu_possible_mask) instead?
> }
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261001192849.74788-1-kprateek.nayak@amd.com?part=2
^ permalink raw reply [flat|nested] 36+ messages in thread
* [RFC PATCH v3 03/13] LoongArch: Initialize CPU _PXM relation for disabled CPUs from SRAT
2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 01/13] lib/sbm: Introduce helpers for architectures to configure LLC properties K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 02/13] drivers/base/arch_topology: Add support for initializing sbm topology K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
2026-10-02 9:13 ` sashiko-bot
2026-10-01 19:28 ` [RFC PATCH v3 04/13] LoongArch: Configure sbm topology during SMP preparation K Prateek Nayak
` (10 subsequent siblings)
13 siblings, 1 reply; 36+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
driver-core, Huacai Chen
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, Shrikanth Hegde, K Prateek Nayak, WANG Xuerui
Use the CPU _PXM (Node) relation for disabled CPUs too when parsing SRAT
table. This detail will be used for early topology parsing to count the
LLC domains and set sparsebitmask (sbm) properties correctly.
In addition to sbm enablement, this also helps mapping per-CPU areas
correctly for disabled CPUs by initializing the node relations via
set_cpuid_to_node() correctly and avoiding the "rr_node" based fallback
in smp_prepare_boot_cpu().
Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
arch/loongarch/kernel/acpi.c | 20 ++++++++++++--------
1 file changed, 12 insertions(+), 8 deletions(-)
diff --git a/arch/loongarch/kernel/acpi.c b/arch/loongarch/kernel/acpi.c
index cb454ee92b20..1f15f78713b7 100644
--- a/arch/loongarch/kernel/acpi.c
+++ b/arch/loongarch/kernel/acpi.c
@@ -311,8 +311,7 @@ acpi_numa_processor_affinity_init(struct acpi_srat_cpu_affinity *pa)
bad_srat();
return;
}
- if ((pa->flags & ACPI_SRAT_CPU_ENABLED) == 0)
- return;
+
pxm = pa->proximity_domain_lo;
if (acpi_srat_revision >= 2) {
pxm |= (pa->proximity_domain_hi[0] << 8);
@@ -332,9 +331,12 @@ acpi_numa_processor_affinity_init(struct acpi_srat_cpu_affinity *pa)
return;
}
- early_numa_add_cpu(pa->apic_id, node);
-
set_cpuid_to_node(pa->apic_id, node);
+
+ if ((pa->flags & ACPI_SRAT_CPU_ENABLED) == 0)
+ return;
+
+ early_numa_add_cpu(pa->apic_id, node);
node_set(node, numa_nodes_parsed);
pr_info("SRAT: PXM %u -> CPU 0x%02x -> Node %u\n", pxm, pa->apic_id, node);
}
@@ -350,8 +352,7 @@ acpi_numa_x2apic_affinity_init(struct acpi_srat_x2apic_cpu_affinity *pa)
bad_srat();
return;
}
- if ((pa->flags & ACPI_SRAT_CPU_ENABLED) == 0)
- return;
+
pxm = pa->proximity_domain;
node = acpi_map_pxm_to_node(pxm);
if (node < 0) {
@@ -366,9 +367,12 @@ acpi_numa_x2apic_affinity_init(struct acpi_srat_x2apic_cpu_affinity *pa)
return;
}
- early_numa_add_cpu(pa->apic_id, node);
-
set_cpuid_to_node(pa->apic_id, node);
+
+ if ((pa->flags & ACPI_SRAT_CPU_ENABLED) == 0)
+ return;
+
+ early_numa_add_cpu(pa->apic_id, node);
node_set(node, numa_nodes_parsed);
pr_info("SRAT: PXM %u -> CPU 0x%02x -> Node %u\n", pxm, pa->apic_id, node);
}
--
2.34.1
^ permalink raw reply related [flat|nested] 36+ messages in thread* Re: [RFC PATCH v3 03/13] LoongArch: Initialize CPU _PXM relation for disabled CPUs from SRAT
2026-10-01 19:28 ` [RFC PATCH v3 03/13] LoongArch: Initialize CPU _PXM relation for disabled CPUs from SRAT K Prateek Nayak
@ 2026-10-02 9:13 ` sashiko-bot
0 siblings, 0 replies; 36+ messages in thread
From: sashiko-bot @ 2026-10-02 9:13 UTC (permalink / raw)
To: K Prateek Nayak
Cc: Heiko Carstens, linux-s390, Alexander Gordeev,
Christian Borntraeger, Vasily Gorbik
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] loongarch/numa: skip adding disabled CPU node to numa_nodes_parsed
--
commit 9393724eb4ba65d5f1a0cdd15d60f325d36ffc28
Author: K Prateek Nayak <kprateek.nayak@amd.com>
Subject: LoongArch: Initialize CPU _PXM relation for disabled CPUs from SRAT
This commit updates SRAT parsing on LoongArch to initialize the CPU _PXM
(Node) relation even for disabled CPUs. This is intended to help early
topology parsing count LLC domains and set sparsebitmask properties, while
mapping per-CPU areas correctly.
> diff --git a/arch/loongarch/kernel/acpi.c b/arch/loongarch/kernel/acpi.c
> index cb454ee92b204..1f15f78713b7e 100644
> --- a/arch/loongarch/kernel/acpi.c
> +++ b/arch/loongarch/kernel/acpi.c
[ ... ]
> @@ -332,9 +331,12 @@ acpi_numa_processor_affinity_init(struct acpi_srat_cpu_affinity *pa)
> return;
> }
>
> - early_numa_add_cpu(pa->apic_id, node);
> -
> set_cpuid_to_node(pa->apic_id, node);
> +
> + if ((pa->flags & ACPI_SRAT_CPU_ENABLED) == 0)
> + return;
> +
> + early_numa_add_cpu(pa->apic_id, node);
> node_set(node, numa_nodes_parsed);
[Severity: High]
Does skipping node_set() leave the NUMA node uninitialized?
If the node lacks memory and other enabled CPUs, it won't be added to
numa_nodes_parsed. As a result, alloc_node_data() will not be called,
leaving NODE_DATA(node) as NULL.
Since set_cpuid_to_node() was already called for this disabled CPU,
early_cpu_to_node() will return this uninitialized node instead of
NUMA_NO_NODE.
This can cause smp_prepare_boot_cpu() to bypass the safe rr_node fallback
and assign the uninitialized node to the CPU:
arch/loongarch/kernel/smp.c:smp_prepare_boot_cpu() {
...
if (node != NUMA_NO_NODE)
set_cpu_numa_node(cpu, node);
else {
set_cpu_numa_node(cpu, rr_node);
rr_node = next_node_in(rr_node, node_online_map);
}
...
}
Could this lead to generic per-CPU allocations or CPU hotplug actions blindly
dereferencing the NULL NODE_DATA(node) pointer? For example, when hotplugging
the disabled CPU:
mm/memory_hotplug.c:hotadd_init_pgdat() {
...
pgdat = NODE_DATA(nid);
/* init node's zones as empty zones, we don't have any present pages.*/
if (free_area_init_core_hotplug(pgdat))
...
}
[ ... ]
> @@ -366,9 +367,12 @@ acpi_numa_x2apic_affinity_init(struct acpi_srat_x2apic_cpu_affinity *pa)
> return;
> }
>
> - early_numa_add_cpu(pa->apic_id, node);
> -
> set_cpuid_to_node(pa->apic_id, node);
> +
> + if ((pa->flags & ACPI_SRAT_CPU_ENABLED) == 0)
> + return;
> +
> + early_numa_add_cpu(pa->apic_id, node);
> node_set(node, numa_nodes_parsed);
[Severity: High]
Does this expose hotplugged CPUs in x2apic topologies to the same NULL
pointer dereference?
The exact same early return logic is present here for x2apic, skipping
the node_set() initialization for disabled CPUs.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261001192849.74788-1-kprateek.nayak@amd.com?part=3
^ permalink raw reply [flat|nested] 36+ messages in thread
* [RFC PATCH v3 04/13] LoongArch: Configure sbm topology during SMP preparation
2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
` (2 preceding siblings ...)
2026-10-01 19:28 ` [RFC PATCH v3 03/13] LoongArch: Initialize CPU _PXM relation for disabled CPUs from SRAT K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
2026-10-02 9:13 ` sashiko-bot
2026-10-07 3:51 ` [RFC PATCH v3.1 " K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 05/13] MIPS: Initialize sbm topology on multi-node systems K Prateek Nayak
` (9 subsequent siblings)
13 siblings, 2 replies; 36+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
driver-core, Huacai Chen
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, Shrikanth Hegde, K Prateek Nayak, WANG Xuerui
Configure the sparsebitmask (sbm) topology for multi-node processors
once SRAT is parsed and cpu_to_node() mappings are stable.
Former commit enabled parsing node mappings for disabled CPUs during
SRAT parsing and every CPU will have a valid mapping before the sbm
topology parsing is attempted ensuring correct bounds for
max_threads_per_instance.
XXX: If _PXM mappings cannot be trusted for disabled CPUs, an alternate
logic needs to be added. See PowerPC enablement later in the series.
Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
Only build tested with config from
https://lore.kernel.org/lkml/202609190905.OZTcO07l-lkp@intel.com/
---
arch/loongarch/kernel/smp.c | 36 ++++++++++++++++++++++++++++++++++++
1 file changed, 36 insertions(+)
diff --git a/arch/loongarch/kernel/smp.c b/arch/loongarch/kernel/smp.c
index d4b5d1b6bb01..d8155b26c4df 100644
--- a/arch/loongarch/kernel/smp.c
+++ b/arch/loongarch/kernel/smp.c
@@ -15,6 +15,7 @@
#include <linux/interrupt.h>
#include <linux/irq_work.h>
#include <linux/profile.h>
+#include <linux/sbm.h>
#include <linux/seq_file.h>
#include <linux/smp.h>
#include <linux/threads.h>
@@ -73,6 +74,10 @@ static cpumask_t cpu_llc_shared_setup_map;
/* representing cpus for which core maps can be computed */
static cpumask_t cpu_core_setup_map;
+/* sbm setup data - only needed during init */
+static cpumask_t cpu_sbm_setup_map __initdata;
+static int __node_thread_count[NR_CPUS] __initdata;
+
struct secondary_data cpuboot_data;
static DEFINE_PER_CPU(int, cpu_state);
@@ -362,10 +367,16 @@ void __init loongson_smp_setup(void)
pr_info("Detected %i available CPU(s)\n", loongson_sysconf.nr_cpus);
}
+int arch_sbm_cpu_instance_id(int cpu)
+{
+ return cpu_to_node(cpu);
+}
+
void __init loongson_prepare_cpus(unsigned int max_cpus)
{
int i = 0;
int threads_per_core = 0;
+ int num_sbm_instances, max_threads_per_instance = 1;
parse_acpi_topology();
cpu_data[0].global_id = cpu_logical_map(0);
@@ -388,6 +399,31 @@ void __init loongson_prepare_cpus(unsigned int max_cpus)
per_cpu(cpu_state, smp_processor_id()) = CPU_ONLINE;
cpu_smt_set_num_threads(threads_per_core, threads_per_core);
+
+ for_each_possible_cpu(i) {
+ bool found = false;
+ int j;
+
+ for_each_cpu(j, &cpu_sbm_setup_map) {
+ if (cpu_to_node(i) == cpu_to_node(j)) {
+ found = true;
+ break;
+ }
+ }
+
+ if (!found) {
+ cpumask_set_cpu(i, &cpu_sbm_setup_map);
+ __node_thread_count[i] = 1;
+ continue;
+ }
+
+ __node_thread_count[j] += 1;
+ max_threads_per_instance = max(max_threads_per_instance,
+ __node_thread_count[j]);
+ }
+
+ num_sbm_instances = cpumask_weight(&cpu_sbm_setup_map);
+ sbm_set_topology(num_sbm_instances, max_threads_per_instance);
}
/*
--
2.34.1
^ permalink raw reply related [flat|nested] 36+ messages in thread* Re: [RFC PATCH v3 04/13] LoongArch: Configure sbm topology during SMP preparation
2026-10-01 19:28 ` [RFC PATCH v3 04/13] LoongArch: Configure sbm topology during SMP preparation K Prateek Nayak
@ 2026-10-02 9:13 ` sashiko-bot
2026-10-07 3:51 ` [RFC PATCH v3.1 " K Prateek Nayak
1 sibling, 0 replies; 36+ messages in thread
From: sashiko-bot @ 2026-10-02 9:13 UTC (permalink / raw)
To: K Prateek Nayak
Cc: Christian Borntraeger, Alexander Gordeev, linux-s390,
Vasily Gorbik, Heiko Carstens
> Configure the sparsebitmask (sbm) topology for multi-node processors
> once SRAT is parsed and cpu_to_node() mappings are stable.
>
> Former commit enabled parsing node mappings for disabled CPUs during
> SRAT parsing and every CPU will have a valid mapping before the sbm
> topology parsing is attempted ensuring correct bounds for
> max_threads_per_instance.
>
> XXX: If _PXM mappings cannot be trusted for disabled CPUs, an alternate
> logic needs to be added. See PowerPC enablement later in the series.
>
> Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
Sashiko has reviewed this patch and found no issues. It looks great!
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261001192849.74788-1-kprateek.nayak@amd.com?part=4
^ permalink raw reply [flat|nested] 36+ messages in thread
* [RFC PATCH v3.1 04/13] LoongArch: Configure sbm topology during SMP preparation
2026-10-01 19:28 ` [RFC PATCH v3 04/13] LoongArch: Configure sbm topology during SMP preparation K Prateek Nayak
2026-10-02 9:13 ` sashiko-bot
@ 2026-10-07 3:51 ` K Prateek Nayak
1 sibling, 0 replies; 36+ messages in thread
From: K Prateek Nayak @ 2026-10-07 3:51 UTC (permalink / raw)
To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
driver-core, Huacai Chen
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, Shrikanth Hegde, K Prateek Nayak, WANG Xuerui
Configure the sparsebitmask (sbm) topology for multi-node processors
once SRAT is parsed and early_cpu_to_node() mappings are stable.
The _PXM mappings are not reliable for disabled CPUs, and the worst case
scenario is considered where the disabled CPUs can either form their own
NUMA node (increases num_instances) or joins the node with the largest
CPU count (increases max_cpus_per_instance).
Since the final topology cannot be predicted at boot time, both are
incremented with the count of disabled CPUs to account for topology
going either ways.
The sbm core will limit traversals to the nodes that are onlined and it
is acceptable to overshoot the final limits.
Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
Changelog rfc v3 .. rfc v3.1:
o Alternate scheme to use disabled CPUs to find the worst case topology
instead of relying of incorrect changes to SRAT parsing.
This was cross-compiled and tested on QEMU with:
qemu-system-loongarch64 \
-machine virt \
-m 4G \
-cpu la464 \
-smp sockets=2,cores=8 \
-bios QEMU_EFI.fd \
-kernel arch/loongarch/boot/vmlinuz.efi \
-initrd ramdisk \
-serial stdio \
-append "root=/dev/ram rdinit=/sbin/init console=ttyS0,115200" \
-object memory-backend-ram,size=2G,id=mem0 \
-object memory-backend-ram,size=2G,id=mem1 \
-numa node,nodeid=0,memdev=mem0,cpus=0-7 \
-numa node,nodeid=1,memdev=mem1,cpus=8-15
...
The incorrect SRAT parsing changes from v3 Patch 03/13 is no longer
required with this alternate fallback.
---
arch/loongarch/kernel/smp.c | 52 +++++++++++++++++++++++++++++++++++++
1 file changed, 52 insertions(+)
diff --git a/arch/loongarch/kernel/smp.c b/arch/loongarch/kernel/smp.c
index d4b5d1b6bb01..6cf1f5ada038 100644
--- a/arch/loongarch/kernel/smp.c
+++ b/arch/loongarch/kernel/smp.c
@@ -15,6 +15,7 @@
#include <linux/interrupt.h>
#include <linux/irq_work.h>
#include <linux/profile.h>
+#include <linux/sbm.h>
#include <linux/seq_file.h>
#include <linux/smp.h>
#include <linux/threads.h>
@@ -73,6 +74,10 @@ static cpumask_t cpu_llc_shared_setup_map;
/* representing cpus for which core maps can be computed */
static cpumask_t cpu_core_setup_map;
+/* sbm setup data - only needed during init */
+static cpumask_t cpu_sbm_setup_map __initdata;
+static int __node_thread_count[NR_CPUS] __initdata;
+
struct secondary_data cpuboot_data;
static DEFINE_PER_CPU(int, cpu_state);
@@ -362,10 +367,17 @@ void __init loongson_smp_setup(void)
pr_info("Detected %i available CPU(s)\n", loongson_sysconf.nr_cpus);
}
+int arch_sbm_cpu_instance_id(int cpu)
+{
+ return cpu_to_node(cpu);
+}
+
void __init loongson_prepare_cpus(unsigned int max_cpus)
{
int i = 0;
+ int disabled_cpus = 0;
int threads_per_core = 0;
+ int num_sbm_instances, max_threads_per_instance = 1;
parse_acpi_topology();
cpu_data[0].global_id = cpu_logical_map(0);
@@ -388,6 +400,46 @@ void __init loongson_prepare_cpus(unsigned int max_cpus)
per_cpu(cpu_state, smp_processor_id()) = CPU_ONLINE;
cpu_smt_set_num_threads(threads_per_core, threads_per_core);
+
+ for_each_possible_cpu(i) {
+ unsigned int node = early_cpu_to_node(i);
+ bool found = false;
+ int j;
+
+ if (node == NUMA_NO_NODE) {
+ disabled_cpus++;
+ continue;
+ }
+
+ for_each_cpu(j, &cpu_sbm_setup_map) {
+ if (node == early_cpu_to_node(j)) {
+ found = true;
+ break;
+ }
+ }
+
+ if (!found) {
+ cpumask_set_cpu(i, &cpu_sbm_setup_map);
+ __node_thread_count[i] = 1;
+ continue;
+ }
+
+ __node_thread_count[j] += 1;
+ max_threads_per_instance = max(max_threads_per_instance,
+ __node_thread_count[j]);
+ }
+
+ /*
+ * In case of disabled CPUs, expect them to either show up as a
+ * new node or get added to the largest NUMA node.
+ *
+ * XXX: Is there a way to establish the correct CPU <-> node
+ * relation for CPUs that have disabled APIC?
+ */
+ num_sbm_instances = cpumask_weight(&cpu_sbm_setup_map) + disabled_cpus;
+ max_threads_per_instance += disabled_cpus;
+
+ sbm_set_topology(num_sbm_instances, max_threads_per_instance);
}
/*
--
2.34.1
^ permalink raw reply related [flat|nested] 36+ messages in thread
* [RFC PATCH v3 05/13] MIPS: Initialize sbm topology on multi-node systems
2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
` (3 preceding siblings ...)
2026-10-01 19:28 ` [RFC PATCH v3 04/13] LoongArch: Configure sbm topology during SMP preparation K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
2026-10-02 9:13 ` sashiko-bot
2026-10-01 19:28 ` [RFC PATCH v3 06/13] powerpc/setup: Initialize sbm topology based on coregroup / NUMA topology K Prateek Nayak
` (8 subsequent siblings)
13 siblings, 1 reply; 36+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
driver-core, Thomas Bogendoerfer, Huacai Chen, Jiaxun Yang
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, Shrikanth Hegde, K Prateek Nayak
For MIPS systems that support NUMA, configure the sparsebitmap (sbm)
topology.
The two systems that toggle NUMA - loongson64 and sgi-ip27, find the
number of nodes and the maximum threads present in an instance. The
data structures used to parse the topology are marked __initdata to
allow reclaim post init.
In both cases - loongson64 and sgi-ip27, loongson3_smp_setup() and
node_scan_cpus() sets node IDs for all possible CPUs ensuring
configure_sbm_topology() captures the full possible system.
arch_sbm_cpu_instance_id() returns the cpu_to_node() mapping to derive
the instance ID and group the CPUs sharing the node on the same sbm
instance (leaf).
Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
Tested on QEMU with loongson3_defconfig from
https://gitlab.postmarketos.org/royka1/linux/-/blob/apollo-7.1.0-r3/arch/mips/configs/loongson3_defconfig?ref_type=tags
QEMU cmdline:
qemu-system-mips64el \
-M loongson3-virt,accel=tcg \
-smp sockets=2,cores=4 \
-cpu Loongson-3A1000 \
-kernel vmlinux \
-append "console=ttyS0 root=/dev/ram" \
-nographic
---
arch/mips/include/asm/topology.h | 6 +++++
arch/mips/kernel/topology.c | 43 ++++++++++++++++++++++++++++++++
arch/mips/loongson64/smp.c | 3 +++
arch/mips/sgi-ip27/ip27-smp.c | 3 +++
4 files changed, 55 insertions(+)
diff --git a/arch/mips/include/asm/topology.h b/arch/mips/include/asm/topology.h
index 5158c802eb65..7a57cbd59819 100644
--- a/arch/mips/include/asm/topology.h
+++ b/arch/mips/include/asm/topology.h
@@ -21,4 +21,10 @@ extern struct cpumask __cpu_primary_thread_mask;
#define cpu_primary_thread_mask ((const struct cpumask *)&__cpu_primary_thread_mask)
#endif
+#ifdef CONFIG_NUMA
+extern void configure_sbm_topology(void);
+#else /* !CONFIG_NUMA */
+static void __maybe_unused configure_sbm_topology(void) { }
+#endif /* CONFIG_NUMA */
+
#endif /* __ASM_TOPOLOGY_H */
diff --git a/arch/mips/kernel/topology.c b/arch/mips/kernel/topology.c
index 9429d85a4703..729a811e8b99 100644
--- a/arch/mips/kernel/topology.c
+++ b/arch/mips/kernel/topology.c
@@ -5,9 +5,52 @@
#include <linux/node.h>
#include <linux/nodemask.h>
#include <linux/percpu.h>
+#include <linux/sbm.h>
static DEFINE_PER_CPU(struct cpu, cpu_devices);
+#ifdef CONFIG_NUMA
+/* sbm setup data - only needed during init */
+static struct cpumask cpu_sbm_setup_map __initdata;
+static int __node_thread_count[NR_CPUS] __initdata;
+
+int arch_sbm_cpu_instance_id(int cpu)
+{
+ return cpu_to_node(cpu);
+}
+
+void __init configure_sbm_topology(void)
+{
+ int num_sbm_instances, max_threads_per_instance = 1;
+ int i;
+
+ for_each_possible_cpu(i) {
+ bool found = false;
+ int j;
+
+ for_each_cpu(j, &cpu_sbm_setup_map) {
+ if (cpu_to_node(i) == cpu_to_node(j)) {
+ found = true;
+ break;
+ }
+ }
+
+ if (!found) {
+ cpumask_set_cpu(i, &cpu_sbm_setup_map);
+ __node_thread_count[i] = 1;
+ continue;
+ }
+
+ __node_thread_count[j] += 1;
+ max_threads_per_instance = max(max_threads_per_instance,
+ __node_thread_count[j]);
+ }
+
+ num_sbm_instances = cpumask_weight(&cpu_sbm_setup_map);
+ sbm_set_topology(num_sbm_instances, max_threads_per_instance);
+}
+#endif /* CONFIG_NUMA */
+
static int __init topology_init(void)
{
int i, ret;
diff --git a/arch/mips/loongson64/smp.c b/arch/mips/loongson64/smp.c
index e584299d0fde..22174ae2b574 100644
--- a/arch/mips/loongson64/smp.c
+++ b/arch/mips/loongson64/smp.c
@@ -17,6 +17,7 @@
#include <asm/smp.h>
#include <asm/time.h>
#include <asm/tlbflush.h>
+#include <asm/topology.h>
#include <asm/cacheflush.h>
#include <loongson.h>
#include <loongson_regs.h>
@@ -493,6 +494,8 @@ static void __init loongson3_smp_setup(void)
if (smp_group[0])
ipi_write_enable(0);
+ configure_sbm_topology();
+
cpu_set_core(&cpu_data[0],
cpu_logical_map(0) % loongson_sysconf.cores_per_package);
cpu_data[0].package = cpu_logical_map(0) / loongson_sysconf.cores_per_package;
diff --git a/arch/mips/sgi-ip27/ip27-smp.c b/arch/mips/sgi-ip27/ip27-smp.c
index 62733e049570..244f8d987f13 100644
--- a/arch/mips/sgi-ip27/ip27-smp.c
+++ b/arch/mips/sgi-ip27/ip27-smp.c
@@ -15,6 +15,7 @@
#include <asm/page.h>
#include <asm/processor.h>
#include <asm/ptrace.h>
+#include <asm/topology.h>
#include <asm/sn/agent.h>
#include <asm/sn/arch.h>
#include <asm/sn/gda.h>
@@ -80,6 +81,8 @@ void cpu_node_probe(void)
highest = node_scan_cpus(nasid, highest);
}
+ configure_sbm_topology();
+
printk("Discovered %d cpus on %d nodes\n", highest + 1, num_online_nodes());
}
--
2.34.1
^ permalink raw reply related [flat|nested] 36+ messages in thread* Re: [RFC PATCH v3 05/13] MIPS: Initialize sbm topology on multi-node systems
2026-10-01 19:28 ` [RFC PATCH v3 05/13] MIPS: Initialize sbm topology on multi-node systems K Prateek Nayak
@ 2026-10-02 9:13 ` sashiko-bot
0 siblings, 0 replies; 36+ messages in thread
From: sashiko-bot @ 2026-10-02 9:13 UTC (permalink / raw)
To: K Prateek Nayak
Cc: linux-s390, Alexander Gordeev, Christian Borntraeger,
Vasily Gorbik, Heiko Carstens
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Low] mips: modpost section mismatch from cpu_node_probe() calling __init function
--
commit a182a04b59c7d5076dc4fa89322bc392b7a56f93
Author: K Prateek Nayak <kprateek.nayak@amd.com>
MIPS: Initialize sbm topology on multi-node systems
For MIPS systems that support NUMA, configure the sparsebitmap (sbm)
topology.
> diff --git a/arch/mips/kernel/topology.c b/arch/mips/kernel/topology.c
> index 9429d85a4703c..bb6d14bfe5ebb 100644
> --- a/arch/mips/kernel/topology.c
> +++ b/arch/mips/kernel/topology.c
> [ ... ]
> @@ -5,9 +5,52 @@
> #include <linux/node.h>
> #include <linux/nodemask.h>
> #include <linux/percpu.h>
> +#include <linux/sbm.h>
>
> static DEFINE_PER_CPU(struct cpu, cpu_devices);
>
> +#ifdef CONFIG_NUMA
> +/* sbm setup data - only needed during init */
> +static struct cpumask cpu_sbm_setup_map __initdata;
> +static int __node_thread_count[NR_CPUS] __initdata;
> +
> +int arch_sbm_cpu_instance_id(int cpu)
> +{
> + return cpu_to_node(cpu);
> +}
> +
> +void __init configure_sbm_topology(void)
> +{
> [ ... ]
> diff --git a/arch/mips/sgi-ip27/ip27-smp.c b/arch/mips/sgi-ip27/ip27-smp.c
> index 62733e049570a..244f8d987f13b 100644
> --- a/arch/mips/sgi-ip27/ip27-smp.c
> +++ b/arch/mips/sgi-ip27/ip27-smp.c
> [ ... ]
> @@ -80,6 +81,8 @@ void cpu_node_probe(void)
> highest = node_scan_cpus(nasid, highest);
> }
>
> + configure_sbm_topology();
[Severity: Low]
Does this cause a modpost section mismatch warning?
Since configure_sbm_topology() is marked __init in
arch/mips/kernel/topology.c, calling it from cpu_node_probe() creates a
cross-section reference because cpu_node_probe() is not marked __init.
While cpu_node_probe() might only be called during initialization,
shouldn't it also be marked __init to prevent the .text section from
referencing the freed .init.text section?
> +
> printk("Discovered %d cpus on %d nodes\n", highest + 1, num_online_nodes());
> }
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261001192849.74788-1-kprateek.nayak@amd.com?part=5
^ permalink raw reply [flat|nested] 36+ messages in thread
* [RFC PATCH v3 06/13] powerpc/setup: Initialize sbm topology based on coregroup / NUMA topology
2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
` (4 preceding siblings ...)
2026-10-01 19:28 ` [RFC PATCH v3 05/13] MIPS: Initialize sbm topology on multi-node systems K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
2026-10-02 9:13 ` sashiko-bot
2026-10-07 14:43 ` Shrikanth Hegde
2026-10-01 19:28 ` [RFC PATCH v3 07/13] s390/topology: Initialize sbm topology during topology_init_early() K Prateek Nayak
` (7 subsequent siblings)
13 siblings, 2 replies; 36+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
driver-core, Madhavan Srinivasan
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, Shrikanth Hegde, K Prateek Nayak,
Michael Ellerman, Nicholas Piggin, Christophe Leroy
Initialize sparsebitmap (sbm) topology based on the coregroup
information. Each coregorup gets its own sparsemask leaf.
pSeries and memory hotplug are interesting since a hotplug can place a
newly added CPU on any online node. This requires special care to allow
estimating bitmask size considering the worst case scenarios - each CPU
onlined is on a separate node, and this is a new N_CPU node.
Platform may enforce a stricter standards for the CPUs being online and
what nodes they can be mapped to but the current implementations makes
no assumptions and considers each CPU can be onlined on a unique node.
pSeries systems that can hotplug CPUs (detected using smp_ops) use the
NUMA topology instead for sbm initialization. arch_sbm_cpu_instance_id()
on these systems use cpu_to_node() mappings to match CPUs to sbm
instances.
XXX: This requires further optimizations to shorten sparsemask
traversals by keeping the number of leaf nodes to a minimum. If there
are nuances I'm not aware of, please reach out.
Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
Tested on ppc64le_defconfig with:
qemu-system-ppc64 \
-M pseries \
-cpu power10 \
-smp sockets=2,cores=2,threads=4 \
-m 10G -nographic \
-kernel vmlinux \
-append "root=/dev/ram sched_debug"
and also on ppce500 VM based on instructions in
https://www.qemu.org/docs/master/system/ppc/ppce500.html
---
arch/powerpc/kernel/setup-common.c | 88 ++++++++++++++++++++
arch/powerpc/platforms/pseries/hotplug-cpu.c | 10 +++
2 files changed, 98 insertions(+)
diff --git a/arch/powerpc/kernel/setup-common.c b/arch/powerpc/kernel/setup-common.c
index 4afaba19b586..4b57ad553172 100644
--- a/arch/powerpc/kernel/setup-common.c
+++ b/arch/powerpc/kernel/setup-common.c
@@ -11,6 +11,7 @@
#include <linux/export.h>
#include <linux/panic_notifier.h>
#include <linux/string.h>
+#include <linux/sbm.h>
#include <linux/sched.h>
#include <linux/init.h>
#include <linux/kernel.h>
@@ -602,6 +603,87 @@ static __init int add_pcspkr(void)
device_initcall(add_pcspkr);
#endif /* CONFIG_PCSPKR_PLATFORM */
+int arch_sbm_cpu_instance_id(int cpu)
+{
+ /*
+ * In case of pSeries processors, sbm masks are
+ * grouped by nodes where the cpuhotplug
+ * operations can remove and re-add same logical
+ * CPUs on different nodes.
+ *
+ * See comment in pseries_cpu_hotplug_init().
+ */
+ if (smp_ops->cpu_disable)
+ return cpu_to_node(cpu);
+
+ return cpu_to_coregroup_id(cpu);
+}
+
+static void __init setup_sbm_topology(void)
+{
+ int num_sbm_instances, max_threads_per_instance = 1;
+ struct cpumask *cpu_sbm_setup_map;
+ int i, *__node_thread_count;
+ int disabled_cpus = 0;
+
+ cpu_sbm_setup_map = memblock_alloc_or_panic(cpumask_size(), __alignof__(long));
+ __node_thread_count = memblock_alloc_or_panic(nr_cpu_ids * sizeof(int),
+ __alignof__(int));
+
+ memset(__node_thread_count, 0, nr_cpu_ids * sizeof(int));
+ memset(cpu_sbm_setup_map, 0, cpumask_size());
+
+ for_each_possible_cpu(i) {
+ bool found = false;
+ int j;
+
+ if (!cpu_present(i)) {
+ disabled_cpus += 1;
+ continue;
+ }
+
+ for_each_cpu(j, cpu_sbm_setup_map) {
+ if (cpu_to_coregroup_id(i) == cpu_to_coregroup_id(j)) {
+ found = true;
+ break;
+ }
+ }
+
+ if (!found) {
+ cpumask_set_cpu(i, cpu_sbm_setup_map);
+ __node_thread_count[i] = 1;
+ continue;
+ }
+
+ __node_thread_count[j] += 1;
+ max_threads_per_instance = max(max_threads_per_instance,
+ __node_thread_count[j]);
+ }
+
+ /*
+ * If CPUs are disabled, they may pop up on any online node.
+ *
+ * XXX: Any implementation nuances that can help this?
+ * pSeries says only online nodes can be extended.
+ */
+ if (disabled_cpus) {
+ num_sbm_instances = num_sbm_instances + disabled_cpus;
+ } else {
+ num_sbm_instances = cpumask_weight(cpu_sbm_setup_map);
+ }
+
+ /*
+ * If disabled threads exists, assume the maximum threads per
+ * instance can extend by the number of disabled threads if they
+ * are all added to the same node.
+ */
+ sbm_set_topology(num_sbm_instances,
+ max_threads_per_instance + disabled_cpus);
+
+ memblock_free(__node_thread_count, nr_cpu_ids * sizeof(int));
+ memblock_free(cpu_sbm_setup_map, cpumask_size());
+}
+
static char ppc_hw_desc_buf[128] __initdata;
struct seq_buf ppc_hw_desc __initdata = {
@@ -1006,6 +1088,12 @@ void __init setup_arch(char **cmdline_p)
early_memtest(min_low_pfn << PAGE_SHIFT, max_low_pfn << PAGE_SHIFT);
+ /*
+ * setup_arch() below can override topology for
+ * pSeries platforms as a result of hotplug nuances.
+ */
+ setup_sbm_topology();
+
if (ppc_md.setup_arch)
ppc_md.setup_arch();
diff --git a/arch/powerpc/platforms/pseries/hotplug-cpu.c b/arch/powerpc/platforms/pseries/hotplug-cpu.c
index bc6926dbf148..7c1c1ac3efde 100644
--- a/arch/powerpc/platforms/pseries/hotplug-cpu.c
+++ b/arch/powerpc/platforms/pseries/hotplug-cpu.c
@@ -19,6 +19,7 @@
#include <linux/kernel.h>
#include <linux/interrupt.h>
#include <linux/delay.h>
+#include <linux/sbm.h>
#include <linux/sched.h> /* for idle_task_exit */
#include <linux/sched/hotplug.h>
#include <linux/cpu.h>
@@ -870,6 +871,15 @@ void __init pseries_cpu_hotplug_init(void)
return;
}
+ /*
+ * find_cpu_id_range() only looks at online nodes.
+ *
+ * XXX: Is it possible for a CPU attached memory node to come
+ * online after this point? May need num_possbile_nodes() then
+ * unless there are platform nuances that can help optimize.
+ */
+ sbm_set_topology(num_online_nodes(), num_possible_cpus());
+
smp_ops->cpu_offline_self = pseries_cpu_offline_self;
smp_ops->cpu_disable = pseries_cpu_disable;
smp_ops->cpu_die = pseries_cpu_die;
--
2.34.1
^ permalink raw reply related [flat|nested] 36+ messages in thread* Re: [RFC PATCH v3 06/13] powerpc/setup: Initialize sbm topology based on coregroup / NUMA topology
2026-10-01 19:28 ` [RFC PATCH v3 06/13] powerpc/setup: Initialize sbm topology based on coregroup / NUMA topology K Prateek Nayak
@ 2026-10-02 9:13 ` sashiko-bot
2026-10-07 14:43 ` Shrikanth Hegde
1 sibling, 0 replies; 36+ messages in thread
From: sashiko-bot @ 2026-10-02 9:13 UTC (permalink / raw)
To: K Prateek Nayak
Cc: Christian Borntraeger, Heiko Carstens, linux-s390, Vasily Gorbik,
Alexander Gordeev
Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
- [High] powerpc/setup: uninitialized num_sbm_instances in setup_sbm_topology()
- [High] powerpc/setup: unconditional dereference of smp_ops causes boot panic
- [Medium] powerpc/pseries: sizing SBM by num_online_nodes() breaks memory hotplug
--
commit a4d9dadc8cce1e3232c36f0087110489e579c9b4
Author: K Prateek Nayak <kprateek.nayak@amd.com>
powerpc/setup: Initialize sbm topology based on coregroup / NUMA topology
This commit initializes the sparsebitmap (sbm) topology based on coregroup
information, with each coregroup getting its own sparsemask leaf. For pSeries,
NUMA topology is used instead to account for CPU hotplug behavior.
> diff --git a/arch/powerpc/kernel/setup-common.c b/arch/powerpc/kernel/setup-common.c
> index 4afaba19b5867..4b57ad553172d 100644
> --- a/arch/powerpc/kernel/setup-common.c
> +++ b/arch/powerpc/kernel/setup-common.c
[ ... ]
> @@ -602,6 +603,87 @@ static __init int add_pcspkr(void)
> device_initcall(add_pcspkr);
> #endif /* CONFIG_PCSPKR_PLATFORM */
>
> +int arch_sbm_cpu_instance_id(int cpu)
> +{
> + /*
> + * In case of pSeries processors, sbm masks are
> + * grouped by nodes where the cpuhotplug
> + * operations can remove and re-add same logical
> + * CPUs on different nodes.
> + *
> + * See comment in pseries_cpu_hotplug_init().
> + */
> + if (smp_ops->cpu_disable)
> + return cpu_to_node(cpu);
[Severity: High]
Is it possible for smp_ops to be NULL here?
If the platform does not initialize SMP (such as on UP platforms), evaluating
smp_ops->cpu_disable unconditionally can cause an immediate boot panic.
Additionally, on !CONFIG_SMP builds, struct smp_ops_t might be an incomplete
type, which can result in compilation errors.
> +
> + return cpu_to_coregroup_id(cpu);
> +}
> +
> +static void __init setup_sbm_topology(void)
> +{
> + int num_sbm_instances, max_threads_per_instance = 1;
[ ... ]
> + /*
> + * If CPUs are disabled, they may pop up on any online node.
> + *
> + * XXX: Any implementation nuances that can help this?
> + * pSeries says only online nodes can be extended.
> + */
> + if (disabled_cpus) {
> + num_sbm_instances = num_sbm_instances + disabled_cpus;
[Severity: High]
Is num_sbm_instances uninitialized when it is used in this addition?
If disabled_cpus > 0, the calculation incorporates stack garbage. This
determines the SBM topology size during boot and can lead to uncontrolled
memory allocation sizing in sbm_init(), potentially causing out-of-bounds
array access or out-of-memory panics.
> + } else {
> + num_sbm_instances = cpumask_weight(cpu_sbm_setup_map);
> + }
> +
> + /*
> + * If disabled threads exists, assume the maximum threads per
> + * instance can extend by the number of disabled threads if they
> + * are all added to the same node.
> + */
> + sbm_set_topology(num_sbm_instances,
> + max_threads_per_instance + disabled_cpus);
> +
> + memblock_free(__node_thread_count, nr_cpu_ids * sizeof(int));
> + memblock_free(cpu_sbm_setup_map, cpumask_size());
> +}
[ ... ]
> diff --git a/arch/powerpc/platforms/pseries/hotplug-cpu.c b/arch/powerpc/platforms/pseries/hotplug-cpu.c
> index bc6926dbf1489..7c1c1ac3efde4 100644
> --- a/arch/powerpc/platforms/pseries/hotplug-cpu.c
> +++ b/arch/powerpc/platforms/pseries/hotplug-cpu.c
[ ... ]
> @@ -870,6 +871,15 @@ void __init pseries_cpu_hotplug_init(void)
> return;
> }
>
> + /*
> + * find_cpu_id_range() only looks at online nodes.
> + *
> + * XXX: Is it possible for a CPU attached memory node to come
> + * online after this point? May need num_possbile_nodes() then
> + * unless there are platform nuances that can help optimize.
> + */
> + sbm_set_topology(num_online_nodes(), num_possible_cpus());
[Severity: Medium]
Does sizing SBM by num_online_nodes() break dynamic node hotplug?
pSeries platforms support Dynamic LPAR, where entirely new NUMA nodes can be
onlined after boot. When a CPU is added to a new node, its requested topology
instance will exceed the boot-time pool bounded by num_online_nodes(). This
exhausts the SBM metadata index pool, causing alloc_sbm_index() to return
-ENOENT and potentially corrupting the scheduler's idle tracking mechanisms.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261001192849.74788-1-kprateek.nayak@amd.com?part=6
^ permalink raw reply [flat|nested] 36+ messages in thread* Re: [RFC PATCH v3 06/13] powerpc/setup: Initialize sbm topology based on coregroup / NUMA topology
2026-10-01 19:28 ` [RFC PATCH v3 06/13] powerpc/setup: Initialize sbm topology based on coregroup / NUMA topology K Prateek Nayak
2026-10-02 9:13 ` sashiko-bot
@ 2026-10-07 14:43 ` Shrikanth Hegde
1 sibling, 0 replies; 36+ messages in thread
From: Shrikanth Hegde @ 2026-10-07 14:43 UTC (permalink / raw)
To: K Prateek Nayak, Peter Zijlstra, Michael Ellerman, linuxppc-dev,
Madhavan Srinivasan, Thomas Gleixner
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, Nicholas Piggin, Christophe Leroy, Chen Yu,
Tim Chen, Ingo Molnar, Juri Lelli, Vincent Guittot, Andrew Morton,
Arnd Bergmann, linux-kernel, linux-arch, linux-s390, linux-mips,
loongarch, driver-core, Ritesh Harjani, Srikar Dronamraju
Hi Prateek.
On 10/2/26 12:58 AM, K Prateek Nayak wrote:
> Initialize sparsebitmap (sbm) topology based on the coregroup
> information. Each coregorup gets its own sparsemask leaf.
>
> pSeries and memory hotplug are interesting since a hotplug can place a
> newly added CPU on any online node. This requires special care to allow
> estimating bitmask size considering the worst case scenarios - each CPU
> onlined is on a separate node, and this is a new N_CPU node.
>
> Platform may enforce a stricter standards for the CPUs being online and
> what nodes they can be mapped to but the current implementations makes
> no assumptions and considers each CPU can be onlined on a unique node.
>
I think this suffers the same fate as structures which are allocated at boot
time such as runqueues.
So your fallback option of putting all the disabled into singleton node may be
sensible option. (If you are not doing that already)
But yhea, will see more into it, this changing node stuff is new for me too.
Also i need to read your patch series too :)
> pSeries systems that can hotplug CPUs (detected using smp_ops) use the
> NUMA topology instead for sbm initialization. arch_sbm_cpu_instance_id()
> on these systems use cpu_to_node() mappings to match CPUs to sbm
> instances.
>
> XXX: This requires further optimizations to shorten sparsemask
> traversals by keeping the number of leaf nodes to a minimum. If there
> are nuances I'm not aware of, please reach out.
>
> Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
> ---
> Tested on ppc64le_defconfig with:
>
> qemu-system-ppc64 \
> -M pseries \
> -cpu power10 \
> -smp sockets=2,cores=2,threads=4 \
> -m 10G -nographic \
> -kernel vmlinux \
> -append "root=/dev/ram sched_debug"
>
> and also on ppce500 VM based on instructions in
> https://www.qemu.org/docs/master/system/ppc/ppce500.html
> ---
> arch/powerpc/kernel/setup-common.c | 88 ++++++++++++++++++++
> arch/powerpc/platforms/pseries/hotplug-cpu.c | 10 +++
> 2 files changed, 98 insertions(+)
>
> diff --git a/arch/powerpc/kernel/setup-common.c b/arch/powerpc/kernel/setup-common.c
> index 4afaba19b586..4b57ad553172 100644
> --- a/arch/powerpc/kernel/setup-common.c
> +++ b/arch/powerpc/kernel/setup-common.c
> @@ -11,6 +11,7 @@
> #include <linux/export.h>
> #include <linux/panic_notifier.h>
> #include <linux/string.h>
> +#include <linux/sbm.h>
> #include <linux/sched.h>
> #include <linux/init.h>
> #include <linux/kernel.h>
> @@ -602,6 +603,87 @@ static __init int add_pcspkr(void)
> device_initcall(add_pcspkr);
> #endif /* CONFIG_PCSPKR_PLATFORM */
>
> +int arch_sbm_cpu_instance_id(int cpu)
> +{
> + /*
> + * In case of pSeries processors, sbm masks are
> + * grouped by nodes where the cpuhotplug
> + * operations can remove and re-add same logical
> + * CPUs on different nodes.
> + *
> + * See comment in pseries_cpu_hotplug_init().
> + */
> + if (smp_ops->cpu_disable)
> + return cpu_to_node(cpu);
> +
> + return cpu_to_coregroup_id(cpu);
> +}
> +
> +static void __init setup_sbm_topology(void)
> +{
> + int num_sbm_instances, max_threads_per_instance = 1;
> + struct cpumask *cpu_sbm_setup_map;
> + int i, *__node_thread_count;
> + int disabled_cpus = 0;
> +
> + cpu_sbm_setup_map = memblock_alloc_or_panic(cpumask_size(), __alignof__(long));
> + __node_thread_count = memblock_alloc_or_panic(nr_cpu_ids * sizeof(int),
> + __alignof__(int));
> +
> + memset(__node_thread_count, 0, nr_cpu_ids * sizeof(int));
> + memset(cpu_sbm_setup_map, 0, cpumask_size());
> +
> + for_each_possible_cpu(i) {
> + bool found = false;
> + int j;
> +
> + if (!cpu_present(i)) {
> + disabled_cpus += 1;
> + continue;
> + }
> +
> + for_each_cpu(j, cpu_sbm_setup_map) {
> + if (cpu_to_coregroup_id(i) == cpu_to_coregroup_id(j)) {
> + found = true;
> + break;
> + }
> + }
> +
> + if (!found) {
> + cpumask_set_cpu(i, cpu_sbm_setup_map);
> + __node_thread_count[i] = 1;
> + continue;
> + }
> +
> + __node_thread_count[j] += 1;
> + max_threads_per_instance = max(max_threads_per_instance,
> + __node_thread_count[j]);
> + }
> +
> + /*
> + * If CPUs are disabled, they may pop up on any online node.
> + *
> + * XXX: Any implementation nuances that can help this?
> + * pSeries says only online nodes can be extended.
> + */
> + if (disabled_cpus) {
> + num_sbm_instances = num_sbm_instances + disabled_cpus;
> + } else {
> + num_sbm_instances = cpumask_weight(cpu_sbm_setup_map);
> + }
> +
> + /*
> + * If disabled threads exists, assume the maximum threads per
> + * instance can extend by the number of disabled threads if they
> + * are all added to the same node.
> + */
> + sbm_set_topology(num_sbm_instances,
> + max_threads_per_instance + disabled_cpus);
> +
> + memblock_free(__node_thread_count, nr_cpu_ids * sizeof(int));
> + memblock_free(cpu_sbm_setup_map, cpumask_size());
> +}
> +
> static char ppc_hw_desc_buf[128] __initdata;
>
> struct seq_buf ppc_hw_desc __initdata = {
> @@ -1006,6 +1088,12 @@ void __init setup_arch(char **cmdline_p)
>
> early_memtest(min_low_pfn << PAGE_SHIFT, max_low_pfn << PAGE_SHIFT);
>
> + /*
> + * setup_arch() below can override topology for
> + * pSeries platforms as a result of hotplug nuances.
> + */
> + setup_sbm_topology();
> +
> if (ppc_md.setup_arch)
> ppc_md.setup_arch();
>
> diff --git a/arch/powerpc/platforms/pseries/hotplug-cpu.c b/arch/powerpc/platforms/pseries/hotplug-cpu.c
> index bc6926dbf148..7c1c1ac3efde 100644
> --- a/arch/powerpc/platforms/pseries/hotplug-cpu.c
> +++ b/arch/powerpc/platforms/pseries/hotplug-cpu.c
> @@ -19,6 +19,7 @@
> #include <linux/kernel.h>
> #include <linux/interrupt.h>
> #include <linux/delay.h>
> +#include <linux/sbm.h>
> #include <linux/sched.h> /* for idle_task_exit */
> #include <linux/sched/hotplug.h>
> #include <linux/cpu.h>
> @@ -870,6 +871,15 @@ void __init pseries_cpu_hotplug_init(void)
> return;
> }
>
> + /*
> + * find_cpu_id_range() only looks at online nodes.
> + *
> + * XXX: Is it possible for a CPU attached memory node to come
> + * online after this point? May need num_possbile_nodes() then
> + * unless there are platform nuances that can help optimize.
> + */
> + sbm_set_topology(num_online_nodes(), num_possible_cpus());
> +
> smp_ops->cpu_offline_self = pseries_cpu_offline_self;
> smp_ops->cpu_disable = pseries_cpu_disable;
> smp_ops->cpu_die = pseries_cpu_die;
^ permalink raw reply [flat|nested] 36+ messages in thread
* [RFC PATCH v3 07/13] s390/topology: Initialize sbm topology during topology_init_early()
2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
` (5 preceding siblings ...)
2026-10-01 19:28 ` [RFC PATCH v3 06/13] powerpc/setup: Initialize sbm topology based on coregroup / NUMA topology K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
2026-10-02 9:13 ` sashiko-bot
2026-10-07 10:19 ` Mete Durlu
2026-10-01 19:28 ` [RFC PATCH v3 08/13] sparc64: Initialize sbm topology on multi-LLC system K Prateek Nayak
` (6 subsequent siblings)
13 siblings, 2 replies; 36+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
driver-core, Heiko Carstens, Vasily Gorbik, Alexander Gordeev
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, Shrikanth Hegde, K Prateek Nayak,
Christian Borntraeger, Sven Schnelle
When the topology mode is TOPOLOGY_MODE_HW, topology_init_early()
already parses the set of socket present on the system.
Use the socket_info to configure the sparsebitmask (sbm) properties for
the system - namely the number of sockets and maximum threads in a
socket instance.
arch_sbm_cpu_instance_id() traverses all socket instances to find
a matching CPU which is not optimal but instance ID is only checked
when the CPU is coming online. Since this is a rare event, the
inefficiency is tolerable.
Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
Build tested with LKP's s390 randconfig with QEMU cmdline:
qemu-system-s390x \
-cpu max \
-smp 2 \
-m 2048 \
-kernel ./arch/s390/boot/vmlinux \
-append "console=ttyAMA0 earlycon ignore_loglevel log_buf_len=10M print_fatal_signals=1 LOGLEVEL=8 sched_debug" \
-nographic
Note: Since ctop needs KVM acceleration to emulate TOPOLOGY_MODE_HW, the
actual spasemask setting is only build tested at the moment.
XXX: Any way around this using cross-compile and QEMU?
---
arch/s390/kernel/topology.c | 52 +++++++++++++++++++++++++++++++++++++
1 file changed, 52 insertions(+)
diff --git a/arch/s390/kernel/topology.c b/arch/s390/kernel/topology.c
index 1377c6f3f670..80dc51f5ce6f 100644
--- a/arch/s390/kernel/topology.c
+++ b/arch/s390/kernel/topology.c
@@ -20,6 +20,7 @@
#include <linux/init.h>
#include <linux/slab.h>
#include <linux/cpu.h>
+#include <linux/sbm.h>
#include <linux/smp.h>
#include <linux/mm.h>
#include <linux/nodemask.h>
@@ -561,6 +562,55 @@ static int __init detect_polarization(union topology_entry *tle)
return tl_core->pp != POLARIZATION_HRZ;
}
+int arch_sbm_cpu_instance_id(int cpu)
+{
+ struct mask_info *info = &socket_info;
+
+ /*
+ * sbm core should not call this if sbm_set_topology() was
+ * skipped below due to lack of topology information.
+ */
+ if (WARN_ON_ONCE(topology_mode != TOPOLOGY_MODE_HW))
+ return -1;
+
+ while (info) {
+ if (cpumask_test_cpu(cpu, &info->mask))
+ return info->id;
+ info = info->next;
+ }
+
+ pr_warn_once("socket mapping for CPU%d not found! Mapping to ID 0", cpu);
+ return 0;
+}
+
+static void __configure_sbm_topology(void)
+{
+ int num_sbm_instances = 0, max_threads_per_instance = -1;
+ struct mask_info *info = &socket_info;
+
+ /*
+ * Consider single package if topology
+ * information is unavailable.
+ *
+ * sbm core will handle the initialization.
+ */
+ if (topology_mode != TOPOLOGY_MODE_HW)
+ return;
+
+ while (info) {
+ num_sbm_instances += 1;
+ max_threads_per_instance = max_t(int,
+ cpumask_weight(&info->mask),
+ max_threads_per_instance);
+ info = info->next;
+ }
+
+ if (WARN_ON_ONCE(!num_sbm_instances || max_threads_per_instance < 1))
+ return;
+
+ sbm_set_topology(num_sbm_instances, max_threads_per_instance);
+}
+
void __init topology_init_early(void)
{
struct sysinfo_15_1_x *info;
@@ -588,6 +638,8 @@ void __init topology_init_early(void)
cpumask_set_cpu(0, &cpu_setup_mask);
__arch_update_cpu_topology();
__arch_update_dedicated_flag(NULL);
+
+ __configure_sbm_topology();
}
static inline int topology_get_mode(int enabled)
--
2.34.1
^ permalink raw reply related [flat|nested] 36+ messages in thread* Re: [RFC PATCH v3 07/13] s390/topology: Initialize sbm topology during topology_init_early()
2026-10-01 19:28 ` [RFC PATCH v3 07/13] s390/topology: Initialize sbm topology during topology_init_early() K Prateek Nayak
@ 2026-10-02 9:13 ` sashiko-bot
2026-10-07 10:19 ` Mete Durlu
1 sibling, 0 replies; 36+ messages in thread
From: sashiko-bot @ 2026-10-02 9:13 UTC (permalink / raw)
To: K Prateek Nayak
Cc: Alexander Gordeev, linux-s390, Christian Borntraeger,
Heiko Carstens, Vasily Gorbik
Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
- [High] s390/topology: OOB write in SBM due to underestimating max_threads_per_instance
- [High] s390/topology: arch_sbm_cpu_instance_id() fails to map secondary CPUs
- [Low] s390/topology: __configure_sbm_topology() counts dummy head node
--
commit 7e4b5901a7b8d62cc7a0551c478b49d8ad5ae92b
Author: K Prateek Nayak <kprateek.nayak@amd.com>
s390/topology: Initialize sbm topology during topology_init_early()
This commit configures sparsebitmask (sbm) properties for the system using
socket_info when the topology mode is TOPOLOGY_MODE_HW. It also adds
arch_sbm_cpu_instance_id() to map CPUs to socket instance IDs during
CPU hotplug events.
> diff --git a/arch/s390/kernel/topology.c b/arch/s390/kernel/topology.c
> index 1377c6f3f6709..80dc51f5ce6f5 100644
> --- a/arch/s390/kernel/topology.c
> +++ b/arch/s390/kernel/topology.c
[ ... ]
> @@ -561,6 +562,55 @@ static int __init detect_polarization(union topology_entry *tle)
> return tl_core->pp != POLARIZATION_HRZ;
> }
>
> +int arch_sbm_cpu_instance_id(int cpu)
> +{
> + struct mask_info *info = &socket_info;
> +
> + /*
> + * sbm core should not call this if sbm_set_topology() was
> + * skipped below due to lack of topology information.
> + */
> + if (WARN_ON_ONCE(topology_mode != TOPOLOGY_MODE_HW))
> + return -1;
> +
> + while (info) {
> + if (cpumask_test_cpu(cpu, &info->mask))
> + return info->id;
[Severity: High]
Can this fail to map newly hotplugged CPUs?
Since hardware topology masks on s390 are updated asynchronously via a
workqueue, the mask might not be populated before the CPU online callback
executes during CPU hotplug.
> + info = info->next;
> + }
> +
> + pr_warn_once("socket mapping for CPU%d not found! Mapping to ID 0", cpu);
> + return 0;
[Severity: High]
If the mask check above fails due to the asynchronous update race, does
falling back to instance 0 break SBM topology isolation by grouping all
secondary CPUs into a single instance?
> +}
> +
> +static void __configure_sbm_topology(void)
> +{
> + int num_sbm_instances = 0, max_threads_per_instance = -1;
> + struct mask_info *info = &socket_info;
[Severity: Low]
Should this traversal start at info->next instead of the dummy head node?
socket_info serves as a static dummy head, and alloc_masks() attaches newly
allocated instances starting at socket_info.next.
> +
> + /*
> + * Consider single package if topology
> + * information is unavailable.
> + *
> + * sbm core will handle the initialization.
> + */
> + if (topology_mode != TOPOLOGY_MODE_HW)
> + return;
> +
> + while (info) {
> + num_sbm_instances += 1;
[Severity: Low]
Counting the dummy head node as a valid instance appears to pass an inflated
count to sbm_set_topology(), which would overallocate the SBM instance
metadata arrays by exactly one element.
> + max_threads_per_instance = max_t(int,
> + cpumask_weight(&info->mask),
> + max_threads_per_instance);
[Severity: High]
Does calculating the max threads per instance here underestimate the
capacity?
Because this is called during topology_init_early() when only the boot CPU
is marked as present, info->mask will only contain the boot CPU's threads.
This results in a drastically undersized max_threads_per_instance.
When secondary CPUs come online later, alloc_sbm_index() assigns them
local indices. Will returning an index that exceeds this undersized limit
trigger an out-of-bounds write in the SBM core when accessing its internal
arrays?
> + info = info->next;
> + }
> +
> + if (WARN_ON_ONCE(!num_sbm_instances || max_threads_per_instance < 1))
> + return;
> +
> + sbm_set_topology(num_sbm_instances, max_threads_per_instance);
> +}
[ ... ]
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261001192849.74788-1-kprateek.nayak@amd.com?part=7
^ permalink raw reply [flat|nested] 36+ messages in thread* Re: [RFC PATCH v3 07/13] s390/topology: Initialize sbm topology during topology_init_early()
2026-10-01 19:28 ` [RFC PATCH v3 07/13] s390/topology: Initialize sbm topology during topology_init_early() K Prateek Nayak
2026-10-02 9:13 ` sashiko-bot
@ 2026-10-07 10:19 ` Mete Durlu
1 sibling, 0 replies; 36+ messages in thread
From: Mete Durlu @ 2026-10-07 10:19 UTC (permalink / raw)
To: K Prateek Nayak, Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar,
Juri Lelli, Vincent Guittot, Andrew Morton, Arnd Bergmann,
linux-kernel, linux-arch, linux-s390, linuxppc-dev, linux-mips,
loongarch, driver-core, Heiko Carstens, Vasily Gorbik,
Alexander Gordeev
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, Shrikanth Hegde, Christian Borntraeger,
Sven Schnelle
On 01/10/2026 21:28, K Prateek Nayak wrote:
> When the topology mode is TOPOLOGY_MODE_HW, topology_init_early()
> already parses the set of socket present on the system.
>
> Use the socket_info to configure the sparsebitmask (sbm) properties for
> the system - namely the number of sockets and maximum threads in a
> socket instance.
>
> arch_sbm_cpu_instance_id() traverses all socket instances to find
> a matching CPU which is not optimal but instance ID is only checked
> when the CPU is coming online. Since this is a rare event, the
> inefficiency is tolerable.
>
> Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
> ---
> Build tested with LKP's s390 randconfig with QEMU cmdline:
>
> qemu-system-s390x \
> -cpu max \
> -smp 2 \
> -m 2048 \
> -kernel ./arch/s390/boot/vmlinux \
> -append "console=ttyAMA0 earlycon ignore_loglevel log_buf_len=10M print_fatal_signals=1 LOGLEVEL=8 sched_debug" \
> -nographic
>
> Note: Since ctop needs KVM acceleration to emulate TOPOLOGY_MODE_HW, the
> actual spasemask setting is only build tested at the moment.
>
> XXX: Any way around this using cross-compile and QEMU?
> ---
> arch/s390/kernel/topology.c | 52 +++++++++++++++++++++++++++++++++++++
> 1 file changed, 52 insertions(+)
>
> diff --git a/arch/s390/kernel/topology.c b/arch/s390/kernel/topology.c
> index 1377c6f3f670..80dc51f5ce6f 100644
> --- a/arch/s390/kernel/topology.c
> +++ b/arch/s390/kernel/topology.c
> @@ -20,6 +20,7 @@
> #include <linux/init.h>
> #include <linux/slab.h>
> #include <linux/cpu.h>
> +#include <linux/sbm.h>
> #include <linux/smp.h>
> #include <linux/mm.h>
> #include <linux/nodemask.h>
> @@ -561,6 +562,55 @@ static int __init detect_polarization(union topology_entry *tle)
> return tl_core->pp != POLARIZATION_HRZ;
> }
topology_init_early() is called once during early boot and never again.
Considering that s390's topology can change during runtime this
implementation won't work as sbm won't be able to adapt.
On s390 max_threads_per_instance can change during runtime.
Like sashiko pointed out, this approach also does not consider
any newly added cpus.
IMHO, there is a slight redesign required so that smb_set_topology()
and smb_init() can be called each time a topology rebuild is necessary.
^ permalink raw reply [flat|nested] 36+ messages in thread
* [RFC PATCH v3 08/13] sparc64: Initialize sbm topology on multi-LLC system
2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
` (6 preceding siblings ...)
2026-10-01 19:28 ` [RFC PATCH v3 07/13] s390/topology: Initialize sbm topology during topology_init_early() K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
2026-10-02 9:13 ` sashiko-bot
2026-10-01 19:28 ` [RFC PATCH v3 09/13] x86/cpu/topology: Initialize sbm topology after topology parsing K Prateek Nayak
` (5 subsequent siblings)
13 siblings, 1 reply; 36+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
driver-core, David S. Miller, Andreas Larsson
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, Shrikanth Hegde, K Prateek Nayak
Once the CPU topology is parsed, initialize the sparsebitmap (sbm)
topology by counting the number of LLC / NUMA instances in the system
and the maximum number of threads in these instance.
Similar to alloc_irqstack_bootmem(), init_sbm_topology() is done after
setup_arch() has finished mapping the cache, socket, and NUMA IDs for
all the present CPUs.
arch_sbm_cpu_instance_id() uses a combination of cpu_to_node() and
max_cache_id to determine the sparsemask instances.
Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
Tested with sparc64_defconfig with
qemu-system-sparc64 \
-M sun4u \
-m 512 \
-kernel vmlinux \
-nographic -append "console=ttyS0"
Similar to some of my previous attempts, I was not able to boot a SMP
configuration with qemu-system-sparc64. For all purposes, sbm
initialization is only build tested.
---
arch/sparc/kernel/setup_64.c | 45 ++++++++++++++++++++++++++++++++++++
1 file changed, 45 insertions(+)
diff --git a/arch/sparc/kernel/setup_64.c b/arch/sparc/kernel/setup_64.c
index 63615f5c99b4..7ab52d17e46f 100644
--- a/arch/sparc/kernel/setup_64.c
+++ b/arch/sparc/kernel/setup_64.c
@@ -28,6 +28,7 @@
#include <linux/root_dev.h>
#include <linux/interrupt.h>
#include <linux/cpu.h>
+#include <linux/sbm.h>
#include <linux/initrd.h>
#include <linux/module.h>
#include <linux/start_kernel.h>
@@ -53,6 +54,7 @@
#include <asm/cacheflush.h>
#include <asm/dma.h>
#include <asm/irq.h>
+#include <asm/cpudata.h>
#ifdef CONFIG_IP_PNP
#include <net/ipconfig.h>
@@ -619,6 +621,48 @@ static void __init alloc_irqstack_bootmem(void)
}
}
+static struct cpumask cpu_sbm_setup_map __initdata;
+static int __node_thread_count[NR_CPUS] __initdata;
+
+int arch_sbm_cpu_instance_id(int cpu)
+{
+ /* Combine the node and cache identifier. */
+ return ((u32)cpu_to_node(cpu) << 16) |
+ cpu_data(cpu).max_cache_id;
+}
+
+static void __init init_sbm_topology(void)
+{
+ int num_sbm_instances, max_threads_per_instance = 1;
+ int i;
+
+ for_each_possible_cpu(i) {
+ bool found = false;
+ int j;
+
+ for_each_cpu(j, &cpu_sbm_setup_map) {
+ if (cpu_to_node(i) == cpu_to_node(j) &&
+ cpu_data(i).max_cache_id == cpu_data(j).max_cache_id) {
+ found = true;
+ break;
+ }
+ }
+
+ if (!found) {
+ cpumask_set_cpu(i, &cpu_sbm_setup_map);
+ __node_thread_count[i] = 1;
+ continue;
+ }
+
+ __node_thread_count[j] += 1;
+ max_threads_per_instance = max(max_threads_per_instance,
+ __node_thread_count[j]);
+ }
+
+ num_sbm_instances = cpumask_weight(&cpu_sbm_setup_map);
+ sbm_set_topology(num_sbm_instances, max_threads_per_instance);
+}
+
void __init setup_arch(char **cmdline_p)
{
/* Initialize PROM console and command line. */
@@ -677,6 +721,7 @@ void __init setup_arch(char **cmdline_p)
* allocate the IRQ stacks.
*/
alloc_irqstack_bootmem();
+ init_sbm_topology();
}
extern int stop_a_enabled;
--
2.34.1
^ permalink raw reply related [flat|nested] 36+ messages in thread* Re: [RFC PATCH v3 08/13] sparc64: Initialize sbm topology on multi-LLC system
2026-10-01 19:28 ` [RFC PATCH v3 08/13] sparc64: Initialize sbm topology on multi-LLC system K Prateek Nayak
@ 2026-10-02 9:13 ` sashiko-bot
0 siblings, 0 replies; 36+ messages in thread
From: sashiko-bot @ 2026-10-02 9:13 UTC (permalink / raw)
To: K Prateek Nayak
Cc: Alexander Gordeev, Vasily Gorbik, linux-s390, Heiko Carstens,
Christian Borntraeger
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] sparc64: sbm topology initialization runs before CPU data is populated
- [High] sparc64: ghost SBM instances created by iterating uninitialized possible CPUs
--
commit e3accb8257522fa2bde71974a4d31cb5abc096f8
Author: K Prateek Nayak <kprateek.nayak@amd.com>
sparc64: Initialize sbm topology on multi-LLC system
This commit initializes the sparsebitmap (sbm) topology for SPARC64 after
the CPU topology is parsed. It introduces init_sbm_topology() which counts
LLC/NUMA instances and their maximum thread counts, calling it from
setup_arch().
> diff --git a/arch/sparc/kernel/setup_64.c b/arch/sparc/kernel/setup_64.c
> index 63615f5c99b47..7ab52d17e46f4 100644
> --- a/arch/sparc/kernel/setup_64.c
> +++ b/arch/sparc/kernel/setup_64.c
[ ... ]
> @@ -619,6 +621,48 @@ static void __init alloc_irqstack_bootmem(void)
> }
> }
>
> +static struct cpumask cpu_sbm_setup_map __initdata;
> +static int __node_thread_count[NR_CPUS] __initdata;
> +
> +int arch_sbm_cpu_instance_id(int cpu)
> +{
> + /* Combine the node and cache identifier. */
> + return ((u32)cpu_to_node(cpu) << 16) |
> + cpu_data(cpu).max_cache_id;
> +}
> +
> +static void __init init_sbm_topology(void)
> +{
> + int num_sbm_instances, max_threads_per_instance = 1;
> + int i;
> +
> + for_each_possible_cpu(i) {
[Severity: High]
Will this loop over uninitialized offline or absent CPUs?
If the possible mask exceeds the present mask, absent CPUs will evaluate
their uninitialized max_cache_id to 0.
They would then incorrectly match each other and be grouped into a single
phantom instance, which could heavily inflate the max_threads_per_instance
calculation below and misallocate the global sparsebitmap arrays.
Should this use for_each_present_cpu() instead?
> + bool found = false;
> + int j;
> +
> + for_each_cpu(j, &cpu_sbm_setup_map) {
> + if (cpu_to_node(i) == cpu_to_node(j) &&
> + cpu_data(i).max_cache_id == cpu_data(j).max_cache_id) {
[Severity: High]
Are the max_cache_id and node values actually initialized when this runs?
This function is called from setup_arch(), but max_cache_id is populated
later during SMP initialization:
arch/sparc/kernel/smp_64.c:setup_per_cpu_areas() {
...
of_fill_in_cpu_data();
if (tlb_type == hypervisor)
mdesc_fill_in_cpu_data(cpu_all_mask);
...
}
Because setup_per_cpu_areas() runs after setup_arch(),
cpu_data(i).max_cache_id and cpu_to_node(i) will evaluate to zero for all
CPUs here. This would group all CPUs into a single instance instead of
distributing them based on max_cache_id.
If alloc_sbm_index() later finds the true max_cache_id, could it fail to
allocate leaves and return -ENOENT? This might leave __sbm_cpu_to_idx[cpu]
as -1, potentially causing out-of-bounds array accesses in the scheduler
fast paths.
> + found = true;
> + break;
> + }
> + }
[ ... ]
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261001192849.74788-1-kprateek.nayak@amd.com?part=8
^ permalink raw reply [flat|nested] 36+ messages in thread
* [RFC PATCH v3 09/13] x86/cpu/topology: Initialize sbm topology after topology parsing
2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
` (7 preceding siblings ...)
2026-10-01 19:28 ` [RFC PATCH v3 08/13] sparc64: Initialize sbm topology on multi-LLC system K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
2026-10-02 9:13 ` sashiko-bot
` (2 more replies)
2026-10-01 19:28 ` [RFC PATCH v3 10/13] lib/sbm: Dynamically allocate sbm index when CPU is activated K Prateek Nayak
` (4 subsequent siblings)
13 siblings, 3 replies; 36+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
driver-core, Thomas Gleixner, Borislav Petkov, Dave Hansen, x86
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, Shrikanth Hegde, K Prateek Nayak,
H. Peter Anvin
From: Peter Zijlstra <peterz@infradead.org>
Once the topology is parsed, the APICIDs of all present CPUs and topology
shifts are available for use.
Use the max APICID encountered during traversal and TOPO_DIE_DOMAIN /
TOPO_TILE_DOMAIN shift to determine the maximum number of CPUs that can
be associated with a single LLC instance and the maximum number of LLC
instances that can exist on a the processor to initialize sparsebitmap
(sbm) topology.
On AMD heterogeneous processors, the Cache Identifier leaf CPUID
0x8000001d field "NumSharingCache" give true value of threads sharing
the cache instance and not the number of bits reserved for a cache
instance in APICID.
To avoid incorrect calculations, shifts from extended topology leaf
0x80000026 on AMD and Hygon systems are used to derive the domain
shifts.
[ tim.c.chen: TOPO_DIE_DOMAIN / TOPO_TILE_DOMAIN based
initialization. ]
[ kprateek: Commit message. Adapting to new sbm helper. ]
(Not-yet-)Signed-off-by: Peter Zijlstra <peterz@infradead.org>
(Not-yet-)Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
arch/x86/kernel/cpu/topology.c | 40 +++++++++++++++++++++++++++++++++-
1 file changed, 39 insertions(+), 1 deletion(-)
diff --git a/arch/x86/kernel/cpu/topology.c b/arch/x86/kernel/cpu/topology.c
index 4913b64ec592..150e098f0ef8 100644
--- a/arch/x86/kernel/cpu/topology.c
+++ b/arch/x86/kernel/cpu/topology.c
@@ -23,6 +23,7 @@
*/
#define pr_fmt(fmt) "CPU topo: " fmt
#include <linux/cpu.h>
+#include <linux/sbm.h>
#include <xen/xen.h>
@@ -449,13 +450,45 @@ static __init bool restrict_to_up(void)
return apic_is_disabled;
}
+int arch_sbm_cpu_instance_id(int cpu)
+{
+ u32 sbm_shift = x86_topo_system.dom_shifts[TOPO_DIE_DOMAIN] - 1;
+ u32 apicid = cpuid_to_apicid[cpu];
+
+ if (boot_cpu_data.x86_vendor == X86_VENDOR_AMD ||
+ boot_cpu_data.x86_vendor == X86_VENDOR_HYGON)
+ sbm_shift = x86_topo_system.dom_shifts[TOPO_TILE_DOMAIN] - 1;
+
+ return (apicid >> sbm_shift);
+}
+
+static __init void init_sbm_topology(u32 max_apicid)
+{
+ u32 sbm_shift = x86_topo_system.dom_shifts[TOPO_DIE_DOMAIN] - 1;
+ int num_sbm_instances, max_threads_per_instance;
+
+ /*
+ * On Intel systems, memory controllers are present at TOPO_DIE_DOMAIN.
+ * On newer AMD and Hygon systems, LLC is at TOPO_TILE_DOMAIN so use
+ * that instead.
+ */
+ if (boot_cpu_data.x86_vendor == X86_VENDOR_AMD ||
+ boot_cpu_data.x86_vendor == X86_VENDOR_HYGON)
+ sbm_shift = x86_topo_system.dom_shifts[TOPO_TILE_DOMAIN] - 1;
+
+ num_sbm_instances = 1 + (max_apicid >> sbm_shift);
+ max_threads_per_instance = (1 << sbm_shift);
+
+ sbm_set_topology(num_sbm_instances, max_threads_per_instance);
+}
+
void __init topology_init_possible_cpus(void)
{
unsigned int assigned = topo_info.nr_assigned_cpus;
unsigned int disabled = topo_info.nr_disabled_cpus;
unsigned int cnta, cntb, cpu, allowed = 1;
unsigned int total = assigned + disabled;
- u32 apicid, firstid;
+ u32 apicid, firstid, maxid;
/*
* If there was no APIC registered, then fake one so that the
@@ -536,10 +569,13 @@ void __init topology_init_possible_cpus(void)
cpuid_to_apicid[topo_info.nr_assigned_cpus++] = apicid;
}
+ maxid = topo_info.boot_cpu_apic_id;
+
for (cpu = 0; cpu < allowed; cpu++) {
apicid = cpuid_to_apicid[cpu];
set_cpu_possible(cpu, true);
+ maxid = max(maxid, apicid);
if (apicid == BAD_APICID)
continue;
@@ -547,6 +583,8 @@ void __init topology_init_possible_cpus(void)
cpu_mark_primary_thread(cpu, apicid);
set_cpu_present(cpu, test_bit(apicid, phys_cpu_present_map));
}
+
+ init_sbm_topology(maxid);
}
/*
--
2.34.1
^ permalink raw reply related [flat|nested] 36+ messages in thread* Re: [RFC PATCH v3 09/13] x86/cpu/topology: Initialize sbm topology after topology parsing
2026-10-01 19:28 ` [RFC PATCH v3 09/13] x86/cpu/topology: Initialize sbm topology after topology parsing K Prateek Nayak
@ 2026-10-02 9:13 ` sashiko-bot
2026-10-03 8:27 ` Chen Yu
2026-10-07 3:52 ` [RFC PATCH v3.1 " K Prateek Nayak
2 siblings, 0 replies; 36+ messages in thread
From: sashiko-bot @ 2026-10-02 9:13 UTC (permalink / raw)
To: K Prateek Nayak
Cc: Christian Borntraeger, Alexander Gordeev, linux-s390,
Heiko Carstens, Vasily Gorbik
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] x86/topology: UB and incorrect topology IDs due to out-of-bounds subtract in dom_shifts
--
commit c4055a2c9b6199a420263b5bc2ed2a987d466f5a
Author: Peter Zijlstra <peterz@infradead.org>
x86/cpu/topology: Initialize sbm topology after topology parsing
This commit initializes the sparsebitmap (sbm) topology on x86 after parsing
the CPU topology. It uses the maximum APICID and topology domain shifts to
determine the maximum number of threads per LLC instance and the maximum
number of instances.
> diff --git a/arch/x86/kernel/cpu/topology.c b/arch/x86/kernel/cpu/topology.c
> index 4913b64ec592d..150e098f0ef8f 100644
> --- a/arch/x86/kernel/cpu/topology.c
> +++ b/arch/x86/kernel/cpu/topology.c
[ ... ]
> @@ -449,13 +450,45 @@ static __init bool restrict_to_up(void)
> return apic_is_disabled;
> }
>
> +int arch_sbm_cpu_instance_id(int cpu)
> +{
> + u32 sbm_shift = x86_topo_system.dom_shifts[TOPO_DIE_DOMAIN] - 1;
[Severity: High]
Does this code compute the wrong shift value by subtracting 1 from the
shift amount instead of the domain index?
If x86_topo_system.dom_shifts[TOPO_DIE_DOMAIN] happens to be 0 (such as on
CPUs without SMT or cores), this evaluates to 0xFFFFFFFF since sbm_shift is
an unsigned 32-bit integer. This would lead to undefined behavior when
evaluating apicid >> sbm_shift later in the function.
Even if the shift value is non-zero, this subtracts 1 from the bit shift
value itself, which calculates the wrong shift, inappropriately halving the
max threads per instance and generating incorrect SBM instance IDs.
Should this instead be x86_topo_system.dom_shifts[TOPO_DIE_DOMAIN - 1]?
Please note that this bug appears to remain unfixed at the end of the patch
series in commit 0b0c43c6e4d9b31b812d8407b146c5dc13783dc8.
> + u32 apicid = cpuid_to_apicid[cpu];
> +
> + if (boot_cpu_data.x86_vendor == X86_VENDOR_AMD ||
> + boot_cpu_data.x86_vendor == X86_VENDOR_HYGON)
> + sbm_shift = x86_topo_system.dom_shifts[TOPO_TILE_DOMAIN] - 1;
[Severity: High]
Is the subtract from the value rather than the index also an issue here
for AMD and Hygon systems?
> +
> + return (apicid >> sbm_shift);
> +}
> +
> +static __init void init_sbm_topology(u32 max_apicid)
> +{
> + u32 sbm_shift = x86_topo_system.dom_shifts[TOPO_DIE_DOMAIN] - 1;
[Severity: High]
Could this also lead to undefined behavior or incorrect shift calculations
during early boot initialization, for the same reasons as above?
> + int num_sbm_instances, max_threads_per_instance;
> +
> + /*
> + * On Intel systems, memory controllers are present at TOPO_DIE_DOMAIN.
> + * On newer AMD and Hygon systems, LLC is at TOPO_TILE_DOMAIN so use
> + * that instead.
> + */
> + if (boot_cpu_data.x86_vendor == X86_VENDOR_AMD ||
> + boot_cpu_data.x86_vendor == X86_VENDOR_HYGON)
> + sbm_shift = x86_topo_system.dom_shifts[TOPO_TILE_DOMAIN] - 1;
[Severity: High]
Does this code also have the same issue with subtracting 1 from the array
value instead of the array index during boot initialization?
> +
> + num_sbm_instances = 1 + (max_apicid >> sbm_shift);
> + max_threads_per_instance = (1 << sbm_shift);
> +
> + sbm_set_topology(num_sbm_instances, max_threads_per_instance);
> +}
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261001192849.74788-1-kprateek.nayak@amd.com?part=9
^ permalink raw reply [flat|nested] 36+ messages in thread* Re: [RFC PATCH v3 09/13] x86/cpu/topology: Initialize sbm topology after topology parsing
2026-10-01 19:28 ` [RFC PATCH v3 09/13] x86/cpu/topology: Initialize sbm topology after topology parsing K Prateek Nayak
2026-10-02 9:13 ` sashiko-bot
@ 2026-10-03 8:27 ` Chen Yu
2026-10-04 6:17 ` K Prateek Nayak
2026-10-07 3:52 ` [RFC PATCH v3.1 " K Prateek Nayak
2 siblings, 1 reply; 36+ messages in thread
From: Chen Yu @ 2026-10-03 8:27 UTC (permalink / raw)
To: K Prateek Nayak
Cc: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
driver-core, Thomas Gleixner, Borislav Petkov, Dave Hansen, x86,
Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, Shrikanth Hegde, H. Peter Anvin
Hello Prateek,
On Thu, Oct 01, 2026 at 07:28:45PM +0000, K Prateek Nayak wrote:
[ snip ]
> +static __init void init_sbm_topology(u32 max_apicid)
> +{
> + u32 sbm_shift = x86_topo_system.dom_shifts[TOPO_DIE_DOMAIN] - 1;
As sashiko reported, and also mentioned here:
https://lore.kernel.org/lkml/20260510155920.2587431-2-yu.c.chen@intel.com/
Maybe x86_topo_system.dom_shifts[TOPO_DIE_DOMAIN - 1] ?
> + int num_sbm_instances, max_threads_per_instance;
> +
> + /*
> + * On Intel systems, memory controllers are present at TOPO_DIE_DOMAIN.
> + * On newer AMD and Hygon systems, LLC is at TOPO_TILE_DOMAIN so use
> + * that instead.
> + */
> + if (boot_cpu_data.x86_vendor == X86_VENDOR_AMD ||
> + boot_cpu_data.x86_vendor == X86_VENDOR_HYGON)
> + sbm_shift = x86_topo_system.dom_shifts[TOPO_TILE_DOMAIN] - 1;
Ditto.
thanks,
Chenyu
^ permalink raw reply [flat|nested] 36+ messages in thread* Re: [RFC PATCH v3 09/13] x86/cpu/topology: Initialize sbm topology after topology parsing
2026-10-03 8:27 ` Chen Yu
@ 2026-10-04 6:17 ` K Prateek Nayak
0 siblings, 0 replies; 36+ messages in thread
From: K Prateek Nayak @ 2026-10-04 6:17 UTC (permalink / raw)
To: Chen Yu
Cc: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
driver-core, Thomas Gleixner, Borislav Petkov, Dave Hansen, x86,
Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, Shrikanth Hegde, H. Peter Anvin
Hello Chenyu,
On 10/3/2026 1:57 PM, Chen Yu wrote:
> Hello Prateek,
>
> On Thu, Oct 01, 2026 at 07:28:45PM +0000, K Prateek Nayak wrote:
>
> [ snip ]
>
>> +static __init void init_sbm_topology(u32 max_apicid)
>> +{
>> + u32 sbm_shift = x86_topo_system.dom_shifts[TOPO_DIE_DOMAIN] - 1;
>
> As sashiko reported, and also mentioned here:
> https://lore.kernel.org/lkml/20260510155920.2587431-2-yu.c.chen@intel.com/
> Maybe x86_topo_system.dom_shifts[TOPO_DIE_DOMAIN - 1] ?
Ack!
I was testing this on a 2CCX (TILE) = 1 CCD (DIE) machine and I
failed to notice my mistake since:
x86_topo_system.dom_shifts[TOPO_DIE_DOMAIN] - 1
and
x86_topo_system.dom_shifts[TOPO_DIE_DOMAIN] - 1
were the same.
>
>> + int num_sbm_instances, max_threads_per_instance;
>> +
>> + /*
>> + * On Intel systems, memory controllers are present at TOPO_DIE_DOMAIN.
>> + * On newer AMD and Hygon systems, LLC is at TOPO_TILE_DOMAIN so use
>> + * that instead.
>> + */
>> + if (boot_cpu_data.x86_vendor == X86_VENDOR_AMD ||
>> + boot_cpu_data.x86_vendor == X86_VENDOR_HYGON)
>> + sbm_shift = x86_topo_system.dom_shifts[TOPO_TILE_DOMAIN] - 1;
>
> Ditto.
Thank you! Will make sure I pull a VM with weird topology next time
during my testing to see if everything is correct.
--
Thanks and Regards,
Prateek
^ permalink raw reply [flat|nested] 36+ messages in thread
* [RFC PATCH v3.1 09/13] x86/cpu/topology: Initialize sbm topology after topology parsing
2026-10-01 19:28 ` [RFC PATCH v3 09/13] x86/cpu/topology: Initialize sbm topology after topology parsing K Prateek Nayak
2026-10-02 9:13 ` sashiko-bot
2026-10-03 8:27 ` Chen Yu
@ 2026-10-07 3:52 ` K Prateek Nayak
2 siblings, 0 replies; 36+ messages in thread
From: K Prateek Nayak @ 2026-10-07 3:52 UTC (permalink / raw)
To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
driver-core, Thomas Gleixner, Borislav Petkov, Dave Hansen, x86
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, Shrikanth Hegde, K Prateek Nayak,
H. Peter Anvin
From: Peter Zijlstra <peterz@infradead.org>
Once the topology is parsed, the APICIDs of all present CPUs and topology
shifts are available for use.
Use the max APICID encountered during traversal and TOPO_TILE_DOMAIN /
TOPO_MODULE_DOMAIN shift to determine the maximum number of CPUs that
can be associated with a single LLC instance and the maximum number of
LLC instances that can exist on a the processor to initialize
sparsebitmap (sbm) topology.
On AMD heterogeneous processors, the Cache Identifier leaf CPUID
0x8000001d field "NumSharingCache" give true value of threads sharing
the cache instance and not the number of bits reserved for a cache
instance in APICID.
To avoid incorrect calculations, shifts from extended topology leaf
0x80000026 on AMD and Hygon systems are used to derive the domain
shifts.
[ tim.c.chen: TOPO_TILE_DOMAIN / TOPO_MODULE_DOMAIN based
initialization. ]
[ kprateek: Commit message. Adapting to new sbm helper. ]
(Not-yet-)Signed-off-by: Peter Zijlstra <peterz@infradead.org>
(Not-yet-)Signed-off-by: Tim Chen <tim.c.chen@linux.intel.com>
Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
Changelog rfc v3 .. rfc v3.1:
o Fix incorrect shifts derivation by using
"dom_shifts[TOPO_TILE_DOMAIN - 1]" instead of
"dom_shifts[TOPO_TILE_DOMAIN] - 1" (Chenyu, Sashiko).
---
arch/x86/kernel/cpu/topology.c | 40 +++++++++++++++++++++++++++++++++-
1 file changed, 39 insertions(+), 1 deletion(-)
diff --git a/arch/x86/kernel/cpu/topology.c b/arch/x86/kernel/cpu/topology.c
index 4913b64ec592..7467fa90b879 100644
--- a/arch/x86/kernel/cpu/topology.c
+++ b/arch/x86/kernel/cpu/topology.c
@@ -23,6 +23,7 @@
*/
#define pr_fmt(fmt) "CPU topo: " fmt
#include <linux/cpu.h>
+#include <linux/sbm.h>
#include <xen/xen.h>
@@ -449,13 +450,45 @@ static __init bool restrict_to_up(void)
return apic_is_disabled;
}
+int arch_sbm_cpu_instance_id(int cpu)
+{
+ u32 sbm_shift = x86_topo_system.dom_shifts[TOPO_DIE_DOMAIN - 1];
+ u32 apicid = cpuid_to_apicid[cpu];
+
+ if (boot_cpu_data.x86_vendor == X86_VENDOR_AMD ||
+ boot_cpu_data.x86_vendor == X86_VENDOR_HYGON)
+ sbm_shift = x86_topo_system.dom_shifts[TOPO_TILE_DOMAIN - 1];
+
+ return (apicid >> sbm_shift);
+}
+
+static __init void init_sbm_topology(u32 max_apicid)
+{
+ u32 sbm_shift = x86_topo_system.dom_shifts[TOPO_DIE_DOMAIN - 1];
+ int num_sbm_instances, max_threads_per_instance;
+
+ /*
+ * On Intel systems, memory controllers are present at TOPO_DIE_DOMAIN.
+ * On newer AMD and Hygon systems, LLC is at TOPO_TILE_DOMAIN so use
+ * that instead.
+ */
+ if (boot_cpu_data.x86_vendor == X86_VENDOR_AMD ||
+ boot_cpu_data.x86_vendor == X86_VENDOR_HYGON)
+ sbm_shift = x86_topo_system.dom_shifts[TOPO_TILE_DOMAIN - 1];
+
+ num_sbm_instances = 1 + (max_apicid >> sbm_shift);
+ max_threads_per_instance = (1 << sbm_shift);
+
+ sbm_set_topology(num_sbm_instances, max_threads_per_instance);
+}
+
void __init topology_init_possible_cpus(void)
{
unsigned int assigned = topo_info.nr_assigned_cpus;
unsigned int disabled = topo_info.nr_disabled_cpus;
unsigned int cnta, cntb, cpu, allowed = 1;
unsigned int total = assigned + disabled;
- u32 apicid, firstid;
+ u32 apicid, firstid, maxid;
/*
* If there was no APIC registered, then fake one so that the
@@ -536,10 +569,13 @@ void __init topology_init_possible_cpus(void)
cpuid_to_apicid[topo_info.nr_assigned_cpus++] = apicid;
}
+ maxid = topo_info.boot_cpu_apic_id;
+
for (cpu = 0; cpu < allowed; cpu++) {
apicid = cpuid_to_apicid[cpu];
set_cpu_possible(cpu, true);
+ maxid = max(maxid, apicid);
if (apicid == BAD_APICID)
continue;
@@ -547,6 +583,8 @@ void __init topology_init_possible_cpus(void)
cpu_mark_primary_thread(cpu, apicid);
set_cpu_present(cpu, test_bit(apicid, phys_cpu_present_map));
}
+
+ init_sbm_topology(maxid);
}
/*
--
2.34.1
^ permalink raw reply related [flat|nested] 36+ messages in thread
* [RFC PATCH v3 10/13] lib/sbm: Dynamically allocate sbm index when CPU is activated
2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
` (8 preceding siblings ...)
2026-10-01 19:28 ` [RFC PATCH v3 09/13] x86/cpu/topology: Initialize sbm topology after topology parsing K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
2026-10-02 9:13 ` sashiko-bot
2026-10-01 19:28 ` [RFC PATCH v3 11/13] lib/sbm: Add helpers to allocate, set, clear, and traverse the bits on sbm K Prateek Nayak
` (3 subsequent siblings)
13 siblings, 1 reply; 36+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
driver-core, Sudeep Holla, Greg Kroah-Hartman, Rafael J. Wysocki,
Danilo Krummrich, Huacai Chen, Thomas Bogendoerfer, Jiaxun Yang,
Madhavan Srinivasan, Heiko Carstens, Vasily Gorbik,
Alexander Gordeev, David S. Miller, Andreas Larsson,
Thomas Gleixner, Borislav Petkov, Dave Hansen, x86
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, Shrikanth Hegde, K Prateek Nayak, WANG Xuerui,
Michael Ellerman, Nicholas Piggin, Christophe Leroy,
Christian Borntraeger, Sven Schnelle, H. Peter Anvin
Add infrastructure to establish CPU to sparsebitmap (sbm) index relation
before CPU is turned active. CPU coming online looks for a free slot
based on its instance ID and acquires a free slot.
If slots are exhausted, a new leaf is allocated for an instance ID. New
CPUs activating with same instance ID can claim free slots on the same
leaf but must never exceed max_threads_per_instance.
For architectures that have not initialized the sbm topology, sbm core
overrides the arch_sbm_cpu_instance_id() in sbm_cpu_instance_id()
wrapper to always return 0 keeping all CPUs on same instance.
Data structures that require fast access have been runtime constified to
enable faster access in kernel hot paths.
Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
include/asm-generic/vmlinux.lds.h | 6 +-
include/linux/sbm.h | 14 +++
init/main.c | 6 +
kernel/sched/core.c | 17 +++
lib/sbm.c | 175 +++++++++++++++++++++++++++++-
5 files changed, 215 insertions(+), 3 deletions(-)
diff --git a/include/asm-generic/vmlinux.lds.h b/include/asm-generic/vmlinux.lds.h
index b2988aa12f66..b0346d382cea 100644
--- a/include/asm-generic/vmlinux.lds.h
+++ b/include/asm-generic/vmlinux.lds.h
@@ -981,7 +981,11 @@
RUNTIME_CONST(ptr, __bfilp_cache) \
RUNTIME_CONST(shift, __futex_shift) \
RUNTIME_CONST(mask, __futex_mask) \
- RUNTIME_CONST(ptr, __futex_queues)
+ RUNTIME_CONST(ptr, __futex_queues) \
+ RUNTIME_CONST(shift, __sbm_shift) \
+ RUNTIME_CONST(mask, __sbm_mask) \
+ RUNTIME_CONST(ptr, __sbm_cpu_to_idx) \
+ RUNTIME_CONST(ptr, __sbm_idx_to_cpu)
/* Alignment must be consistent with (kunit_suite *) in include/kunit/test.h */
#define KUNIT_TABLE() \
diff --git a/include/linux/sbm.h b/include/linux/sbm.h
index adac12ed233a..232b0076bb3f 100644
--- a/include/linux/sbm.h
+++ b/include/linux/sbm.h
@@ -2,7 +2,21 @@
#ifndef _LINUX_SBM_H
#define _LINUX_SBM_H
+/*
+ * Masks and shifts for sbm index to translate
+ * a sbm leaf to CPU.
+ */
+extern int __sbm_shift;
+extern int __sbm_mask;
+
int arch_sbm_cpu_instance_id(int cpu);
void sbm_set_topology(int num_instances, int max_threads_per_instance);
+int sbm_cpu_to_idx(int cpu);
+int sbm_idx_to_cpu(int idx);
+
+int alloc_sbm_index(int cpu);
+void free_sbm_index(int cpu);
+int sbm_init(void);
+
#endif /* _LINUX_SBM_H */
diff --git a/init/main.c b/init/main.c
index 2613d3f9b3ce..b7406bd3acc8 100644
--- a/init/main.c
+++ b/init/main.c
@@ -72,6 +72,7 @@
#include <linux/pid_namespace.h>
#include <linux/device/driver.h>
#include <linux/kthread.h>
+#include <linux/sbm.h>
#include <linux/sched.h>
#include <linux/sched/init.h>
#include <linux/signal.h>
@@ -1652,6 +1653,11 @@ static noinline void __init kernel_init_freeable(void)
smp_prepare_cpus(setup_max_cpus);
+ sbm_init();
+
+ /* Finish initializing boot CPU since it is already active. */
+ alloc_sbm_index(smp_processor_id());
+
workqueue_init();
init_mm_internals();
diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index 0bb86a43a592..977f579da410 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -61,6 +61,7 @@
#include <linux/rcuwait_api.h>
#include <linux/rseq.h>
#include <linux/sched/wake_q.h>
+#include <linux/sbm.h>
#include <linux/scs.h>
#include <linux/slab.h>
#include <linux/syscalls.h>
@@ -8636,6 +8637,14 @@ int sched_cpu_activate(unsigned int cpu)
*/
balance_push_set(cpu, false);
+ alloc_sbm_index(cpu);
+
+ /*
+ * Make sure sbm mappings are visible
+ * before CPU is toggled active.
+ */
+ smp_mb();
+
/*
* When going up, increment the number of cores with SMT present.
*/
@@ -8688,6 +8697,14 @@ int sched_cpu_deactivate(unsigned int cpu)
set_cpu_active(cpu, false);
+ /*
+ * Make sure CPU is inactive before
+ * sbm indices are reclaimed.
+ */
+ smp_mb();
+
+ free_sbm_index(cpu);
+
/*
* From this point forward, this CPU will refuse to run any task that
* is not: migrate_disable() or KTHREAD_IS_PER_CPU, and will actively
diff --git a/lib/sbm.c b/lib/sbm.c
index 82280e3306de..e5b0508b6825 100644
--- a/lib/sbm.c
+++ b/lib/sbm.c
@@ -1,10 +1,58 @@
/* SPDX-License-Identifier: GPL-2.0 */
#include <linux/sbm.h>
#include <linux/init.h>
+#include <linux/log2.h>
+#include <linux/slab.h>
+#include <linux/cache.h>
#include <linux/printk.h>
+#include <linux/cpumask.h>
+#include <linux/jump_label.h>
-static int sbm_max_threads_per_instance = -1;
-static int sbm_num_instance = -1;
+#include <asm/runtime-const.h>
+
+static int sbm_max_threads_per_instance __ro_after_init = -1;
+static int sbm_num_instance __ro_after_init = -1;
+
+int __sbm_shift __ro_after_init;
+int __sbm_mask __ro_after_init;
+
+static struct {
+ int instance_id; /* Instance ID linked to the leaf. */
+ unsigned long allocated_mask; /* Set of IDs that have been allocated. */
+} *__sbm_idx_metadata __ro_after_init;
+
+/* Translations between cpu <-> sbm leaf */
+static int *__sbm_cpu_to_idx __ro_after_init;
+static int *__sbm_idx_to_cpu __ro_after_init;
+
+static __always_inline int *_sbm_cpu_to_idx(void)
+{
+ return runtime_const_ptr(__sbm_cpu_to_idx);
+}
+
+static __always_inline int *_sbm_idx_to_cpu(void)
+{
+ return runtime_const_ptr(__sbm_idx_to_cpu);
+}
+
+int sbm_cpu_to_idx(int cpu)
+{
+ return _sbm_cpu_to_idx()[cpu];
+}
+
+int sbm_idx_to_cpu(int idx)
+{
+ return _sbm_idx_to_cpu()[idx];
+}
+
+/*
+ * Certain architectures may skip initializing sbm propoerties
+ * while having an arch_sbm_cpu_instance_id() definition.
+ *
+ * In such cases, don't trust the arch/ side redefine and use
+ * the default single instance mapping.
+ */
+static DEFINE_STATIC_KEY_FALSE(sbm_arch_initialized);
/*
* In absence of an arch definition, consider all CPUs to
@@ -15,6 +63,72 @@ int __weak arch_sbm_cpu_instance_id(int cpu)
return 0;
}
+static int sbm_cpu_to_instance(int cpu)
+{
+ if (static_branch_likely(&sbm_arch_initialized))
+ return arch_sbm_cpu_instance_id(cpu);
+
+ return 0;
+}
+
+int alloc_sbm_index(int cpu)
+{
+ int cpu_instance = sbm_cpu_to_instance(cpu);
+ int i, idx = BITS_PER_LONG, free_index = -1;
+
+ for (i = 0; i < sbm_num_instance; ++i) {
+ if (__sbm_idx_metadata[i].instance_id == cpu_instance) {
+ idx = find_first_zero_bit(&__sbm_idx_metadata[i].allocated_mask,
+ BITS_PER_LONG);
+
+ if (idx < BITS_PER_LONG)
+ break;
+ }
+ if (free_index == -1 && __sbm_idx_metadata[i].instance_id == -1)
+ free_index = i;
+ }
+
+ if (i == sbm_num_instance && free_index == -1)
+ return -ENOENT;
+
+ if (i == sbm_num_instance) {
+ __sbm_idx_metadata[free_index].instance_id = cpu_instance;
+ i = free_index;
+ idx = 0;
+ }
+
+ WARN_ON_ONCE(idx >= sbm_max_threads_per_instance);
+
+ __set_bit(idx, &__sbm_idx_metadata[i].allocated_mask);
+
+ idx = (i << __sbm_shift) + idx;
+ _sbm_idx_to_cpu()[idx] = cpu;
+ _sbm_cpu_to_idx()[cpu] = idx;
+
+ return 0;
+}
+
+void free_sbm_index(int cpu)
+{
+ int idx = sbm_cpu_to_idx(cpu);
+ u32 leaf;
+
+ if (idx < 0)
+ return;
+
+ _sbm_idx_to_cpu()[idx] = -1;
+ _sbm_cpu_to_idx()[cpu] = -1;
+
+ leaf = runtime_const_shift_right_32(idx, __sbm_shift);
+ idx = runtime_const_mask_32(idx, __sbm_mask);
+
+ __clear_bit(idx, &__sbm_idx_metadata[leaf].allocated_mask);
+
+ if (find_first_bit(&__sbm_idx_metadata[leaf].allocated_mask, BITS_PER_LONG) ==
+ BITS_PER_LONG)
+ __sbm_idx_metadata[leaf].instance_id = -1;
+}
+
void __init sbm_set_topology(int num_instances, int max_threads_per_instance)
{
sbm_max_threads_per_instance = max_threads_per_instance;
@@ -25,3 +139,60 @@ void __init sbm_set_topology(int num_instances, int max_threads_per_instance)
sbm_max_threads_per_instance);
}
+int __init sbm_init(void)
+{
+ int i;
+
+ if (sbm_max_threads_per_instance > 0 && sbm_num_instance > 0) {
+ static_branch_enable(&sbm_arch_initialized);
+ goto init_properties;
+ }
+
+ sbm_max_threads_per_instance = BITS_PER_LONG;
+ sbm_num_instance = (nr_cpumask_bits / BITS_PER_LONG) + 1;
+
+init_properties:
+ /*
+ * If the number of CPUs per instance cross bitmask word boundary,
+ * split the instances into samller chunks on BITS_PER_LONG and
+ * increase the order of leaves.
+ */
+ if (sbm_max_threads_per_instance > BITS_PER_LONG) {
+ int split = (sbm_max_threads_per_instance + BITS_PER_LONG - 1) / BITS_PER_LONG;
+
+ sbm_max_threads_per_instance = BITS_PER_LONG;
+ sbm_num_instance *= split;
+ }
+
+ sbm_max_threads_per_instance = roundup_pow_of_two(sbm_max_threads_per_instance);
+
+ __sbm_shift = ilog2(sbm_max_threads_per_instance);
+ __sbm_mask = sbm_max_threads_per_instance - 1;
+
+ __sbm_idx_metadata = kzalloc_objs(*__sbm_idx_metadata, sbm_num_instance);
+ __sbm_cpu_to_idx = kzalloc_objs(*__sbm_cpu_to_idx, nr_cpumask_bits);
+ __sbm_idx_to_cpu = kzalloc_objs(*__sbm_idx_to_cpu,
+ sbm_max_threads_per_instance * sbm_num_instance);
+
+ BUG_ON(!__sbm_cpu_to_idx || !__sbm_idx_to_cpu || !__sbm_idx_metadata);
+
+ for (i = 0; i < nr_cpumask_bits; ++i)
+ __sbm_cpu_to_idx[i] = -1;
+
+ /* Set all instance id to -1 to allow for future allocations to claim them. */
+ for (i = 0; i < sbm_num_instance; ++i)
+ __sbm_idx_metadata[i].instance_id = -1;
+
+ runtime_const_init(shift, __sbm_shift);
+ runtime_const_init(mask, __sbm_mask);
+ runtime_const_init(ptr, __sbm_cpu_to_idx);
+ runtime_const_init(ptr, __sbm_idx_to_cpu);
+
+ barrier();
+
+ pr_info("sbm instance count: %d (maximum threads per instance: %d)\n",
+ sbm_num_instance,
+ sbm_max_threads_per_instance);
+
+ return 0;
+}
--
2.34.1
^ permalink raw reply related [flat|nested] 36+ messages in thread* Re: [RFC PATCH v3 10/13] lib/sbm: Dynamically allocate sbm index when CPU is activated
2026-10-01 19:28 ` [RFC PATCH v3 10/13] lib/sbm: Dynamically allocate sbm index when CPU is activated K Prateek Nayak
@ 2026-10-02 9:13 ` sashiko-bot
0 siblings, 0 replies; 36+ messages in thread
From: sashiko-bot @ 2026-10-02 9:13 UTC (permalink / raw)
To: K Prateek Nayak
Cc: Christian Borntraeger, Vasily Gorbik, linux-s390, Heiko Carstens,
Alexander Gordeev
Thank you for your contribution! Sashiko AI review found 4 potential issue(s) to consider:
- [High] lib/sbm: array overflow in alloc_sbm_index() due to hardcoded BITS_PER_LONG
- [High] sched/core: premature SBM index free causes UAF and out-of-bounds access
- [High] sched/core: unhandled alloc_sbm_index() failure allows corrupting sbm[-1]
- [Medium] lib/sbm: sentinel value collision with architecture instance IDs corrupts topology
--
commit 525dc1ae45d6fc1670c637f040964c4d0eb2f14a
Author: K Prateek Nayak <kprateek.nayak@amd.com>
lib/sbm: Dynamically allocate sbm index when CPU is activated
Add infrastructure to establish CPU to sparsebitmap (sbm) index relation
before CPU is turned active. CPU coming online looks for a free slot
based on its instance ID and acquires a free slot.
> diff --git a/init/main.c b/init/main.c
> index 2613d3f9b3ce9..b7406bd3acc87 100644
> --- a/init/main.c
> +++ b/init/main.c
[ ... ]
> @@ -1652,6 +1653,11 @@ static noinline void __init kernel_init_freeable(void)
>
> smp_prepare_cpus(setup_max_cpus);
>
> + sbm_init();
> +
> + /* Finish initializing boot CPU since it is already active. */
> + alloc_sbm_index(smp_processor_id());
[Severity: High]
Does this code need to check the return value of alloc_sbm_index()? If it
fails and returns -ENOENT, the boot CPU might proceed without a valid mapping,
potentially exposing the -1 default value to subsystem bitwise operations.
> +
> workqueue_init();
>
> init_mm_internals();
> diff --git a/kernel/sched/core.c b/kernel/sched/core.c
> index 0bb86a43a5920..977f579da4108 100644
> --- a/kernel/sched/core.c
> +++ b/kernel/sched/core.c
[ ... ]
> @@ -8636,6 +8637,14 @@ int sched_cpu_activate(unsigned int cpu)
> */
> balance_push_set(cpu, false);
>
> + alloc_sbm_index(cpu);
[Severity: High]
Similar to the boot CPU initialization, should we handle potential errors from
alloc_sbm_index() here? If the slots are full and it returns -ENOENT, the CPU
is brought fully online without a valid mapping, which might cause memory
corruption when writing to sbm[-1].bitmap.
> +
> + /*
> + * Make sure sbm mappings are visible
> + * before CPU is toggled active.
> + */
> + smp_mb();
> +
> /*
> * When going up, increment the number of cores with SMT present.
> */
[ ... ]
> @@ -8688,6 +8697,14 @@ int sched_cpu_deactivate(unsigned int cpu)
>
> set_cpu_active(cpu, false);
>
> + /*
> + * Make sure CPU is inactive before
> + * sbm indices are reclaimed.
> + */
> + smp_mb();
> +
> + free_sbm_index(cpu);
[Severity: High]
Can reclaiming the index before synchronize_rcu() cause out-of-bounds array
accesses? If a concurrent lockless reader iterates the sbm and reads the index
before it is cleared, but calls sbm_idx_to_cpu() after this free_sbm_index()
call, it might receive -1.
> +
> /*
> * From this point forward, this CPU will refuse to run any task that
> * is not: migrate_disable() or KTHREAD_IS_PER_CPU, and will actively
> diff --git a/lib/sbm.c b/lib/sbm.c
> index 82280e3306de8..b80ed1d961379 100644
> --- a/lib/sbm.c
> +++ b/lib/sbm.c
[ ... ]
> @@ -15,6 +63,72 @@ int __weak arch_sbm_cpu_instance_id(int cpu)
> return 0;
> }
>
> +static int sbm_cpu_to_instance(int cpu)
> +{
[ ... ]
> +int alloc_sbm_index(int cpu)
> +{
> + int cpu_instance = sbm_cpu_to_instance(cpu);
> + int i, idx = BITS_PER_LONG, free_index = -1;
> +
> + for (i = 0; i < sbm_num_instance; ++i) {
> + if (__sbm_idx_metadata[i].instance_id == cpu_instance) {
> + idx = find_first_zero_bit(&__sbm_idx_metadata[i].allocated_mask,
> + BITS_PER_LONG);
> +
> + if (idx < BITS_PER_LONG)
> + break;
> + }
> + if (free_index == -1 && __sbm_idx_metadata[i].instance_id == -1)
> + free_index = i;
[Severity: Medium]
Could this condition falsely match if an architecture returns -1 for an
instance ID (like NUMA_NO_NODE)? If the first CPU sets the instance ID to -1,
subsequent CPUs with valid instance IDs might mistakenly see this leaf as
empty, and overwrite the previous CPU's mapping.
> + }
> +
> + if (i == sbm_num_instance && free_index == -1)
> + return -ENOENT;
> +
> + if (i == sbm_num_instance) {
> + __sbm_idx_metadata[free_index].instance_id = cpu_instance;
> + i = free_index;
> + idx = 0;
> + }
> +
> + WARN_ON_ONCE(idx >= sbm_max_threads_per_instance);
[Severity: High]
Does this code overflow into subsequent leaves when
sbm_max_threads_per_instance is smaller than BITS_PER_LONG? Since the
find_first_zero_bit() limit and the break condition above hardcode
BITS_PER_LONG instead of checking against sbm_max_threads_per_instance, the
global idx computed below might overflow the leaf's boundary and corrupt
mappings for other instances.
> +
> + __set_bit(idx, &__sbm_idx_metadata[i].allocated_mask);
> +
> + idx = (i << __sbm_shift) + idx;
> + _sbm_idx_to_cpu()[idx] = cpu;
> + _sbm_cpu_to_idx()[cpu] = idx;
> +
> + return 0;
> +}
> +
> +void free_sbm_index(int cpu)
> +{
> + int idx = sbm_cpu_to_idx(cpu);
> + u32 leaf;
> +
> + if (idx < 0)
> + return;
> +
> + _sbm_idx_to_cpu()[idx] = -1;
> + _sbm_cpu_to_idx()[cpu] = -1;
> +
> + leaf = runtime_const_shift_right_32(idx, __sbm_shift);
> + idx = runtime_const_mask_32(idx, __sbm_mask);
> +
> + __clear_bit(idx, &__sbm_idx_metadata[leaf].allocated_mask);
[Severity: High]
If the allocation allowed an index to overflow past
sbm_max_threads_per_instance, will the shifted leaf calculation here resolve
to the wrong leaf and clear an allocated bit belonging to another instance?
> +
> + if (find_first_bit(&__sbm_idx_metadata[leaf].allocated_mask, BITS_PER_LONG) ==
> + BITS_PER_LONG)
> + __sbm_idx_metadata[leaf].instance_id = -1;
> +}
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261001192849.74788-1-kprateek.nayak@amd.com?part=10
^ permalink raw reply [flat|nested] 36+ messages in thread
* [RFC PATCH v3 11/13] lib/sbm: Add helpers to allocate, set, clear, and traverse the bits on sbm
2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
` (9 preceding siblings ...)
2026-10-01 19:28 ` [RFC PATCH v3 10/13] lib/sbm: Dynamically allocate sbm index when CPU is activated K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
2026-10-02 9:13 ` sashiko-bot
2026-10-01 19:28 ` [RFC PATCH v3 12/13] sched/fair: Allocate nohz.idle_cpus_mask during sched_init_smp() K Prateek Nayak
` (2 subsequent siblings)
13 siblings, 1 reply; 36+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
driver-core, Sudeep Holla, Greg Kroah-Hartman, Rafael J. Wysocki,
Danilo Krummrich, Huacai Chen, Thomas Bogendoerfer, Jiaxun Yang,
Madhavan Srinivasan, Heiko Carstens, Vasily Gorbik,
Alexander Gordeev, David S. Miller, Andreas Larsson,
Thomas Gleixner, Borislav Petkov, Dave Hansen, x86
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, Shrikanth Hegde, K Prateek Nayak, WANG Xuerui,
Michael Ellerman, Nicholas Piggin, Christophe Leroy,
Christian Borntraeger, Sven Schnelle, H. Peter Anvin
From: Peter Zijlstra <peterz@infradead.org>
Introduce helpers to allocate a sparsebitmap (sbm) of arch configured
length, set a bit on the sbm, clear a bit from the sbm, and iterate all
the set indices on a sbm structure.
[ yu.c.chen: Fixes for sbm implementation. ]
[ kprateek: Adapting sbm implementation to a flat array implementation. ]
(Not-yet-)Signed-off-by: Peter Zijlstra <peterz@infradead.org>
(Not-yet-)Signed-off-by: Chen Yu <yu.c.chen@intel.com>
Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
include/linux/sbm.h | 85 +++++++++++++++++++++++++++++++++++++++++++++
lib/sbm.c | 55 +++++++++++++++++++++++++++++
2 files changed, 140 insertions(+)
diff --git a/include/linux/sbm.h b/include/linux/sbm.h
index 232b0076bb3f..63b116e52e6c 100644
--- a/include/linux/sbm.h
+++ b/include/linux/sbm.h
@@ -2,6 +2,8 @@
#ifndef _LINUX_SBM_H
#define _LINUX_SBM_H
+#include <linux/bitmap.h>
+
/*
* Masks and shifts for sbm index to translate
* a sbm leaf to CPU.
@@ -9,12 +11,95 @@
extern int __sbm_shift;
extern int __sbm_mask;
+struct sbm {
+ unsigned long bitmap;
+} ____cacheline_aligned;
+
int arch_sbm_cpu_instance_id(int cpu);
void sbm_set_topology(int num_instances, int max_threads_per_instance);
int sbm_cpu_to_idx(int cpu);
int sbm_idx_to_cpu(int idx);
+struct sbm *sbm_alloc(void);
+bool sbm_empty(struct sbm *sbm);
+int sbm_find_next_bit(struct sbm *sbm, int start);
+
+#define __sbm_op(sbm, func) \
+({ \
+ int idx = sbm_cpu_to_idx(cpu); \
+ int nr = idx >> __sbm_shift; \
+ int bit = idx & __sbm_mask; \
+ \
+ func(bit, &sbm[nr].bitmap); \
+})
+
+static inline void sbm_cpu_set(struct sbm *sbm, int cpu)
+{
+ __sbm_op(sbm, set_bit);
+}
+
+static inline void sbm_cpu_clear(struct sbm *sbm, int cpu)
+{
+ __sbm_op(sbm, clear_bit);
+}
+
+static inline void __sbm_cpu_set(struct sbm *sbm, int cpu)
+{
+ __sbm_op(sbm, __set_bit);
+}
+
+static inline void __sbm_cpu_clear(struct sbm *sbm, int cpu)
+{
+ __sbm_op(sbm, __clear_bit);
+}
+
+static inline bool sbm_cpu_test(struct sbm *sbm, int cpu)
+{
+ return __sbm_op(sbm, test_bit);
+}
+
+static __always_inline
+unsigned int sbm_find_next_bit_wrap(struct sbm *sbm, int start)
+{
+ int bit = sbm_find_next_bit(sbm, start);
+
+ if (bit >= 0 || start == 0)
+ return bit;
+
+ bit = sbm_find_next_bit(sbm, 0);
+ return bit < start ? bit : -1;
+}
+
+static __always_inline
+unsigned int __sbm_for_each_wrap(struct sbm *sbm, int start, int n)
+{
+ int bit;
+
+ /* If not wrapped around */
+ if (n > start) {
+ /* and have a bit, just return it. */
+ bit = sbm_find_next_bit(sbm, n);
+ if (bit >= 0)
+ return bit;
+
+ /* Otherwise, wrap around and ... */
+ n = 0;
+ }
+
+ /* Search the other part. */
+ bit = sbm_find_next_bit(sbm, n);
+ return bit < start ? bit : -1;
+}
+
+#define sbm_for_each_set_bit(sbm, idx) \
+ for (int idx = sbm_find_next_bit(sbm, 0); \
+ idx >= 0; idx = sbm_find_next_bit(sbm, idx+1))
+
+#define sbm_for_each_set_bit_wrap(sbm, idx, start) \
+ for (int idx = sbm_find_next_bit_wrap(sbm, start); \
+ idx >= 0; idx = __sbm_for_each_wrap(sbm, start, idx+1))
+
int alloc_sbm_index(int cpu);
void free_sbm_index(int cpu);
int sbm_init(void);
diff --git a/lib/sbm.c b/lib/sbm.c
index e5b0508b6825..eeca5ce06d50 100644
--- a/lib/sbm.c
+++ b/lib/sbm.c
@@ -12,6 +12,7 @@
static int sbm_max_threads_per_instance __ro_after_init = -1;
static int sbm_num_instance __ro_after_init = -1;
+static int sbm_max_populated_index;
int __sbm_shift __ro_after_init;
int __sbm_mask __ro_after_init;
@@ -35,6 +36,11 @@ static __always_inline int *_sbm_idx_to_cpu(void)
return runtime_const_ptr(__sbm_idx_to_cpu);
}
+static int sbm_max_index(void)
+{
+ return READ_ONCE(sbm_max_populated_index);
+}
+
int sbm_cpu_to_idx(int cpu)
{
return _sbm_cpu_to_idx()[cpu];
@@ -45,6 +51,44 @@ int sbm_idx_to_cpu(int idx)
return _sbm_idx_to_cpu()[idx];
}
+struct sbm *sbm_alloc(void)
+{
+ return kzalloc_objs(struct sbm, sbm_max_threads_per_instance * sbm_num_instance);
+}
+
+bool sbm_empty(struct sbm *sbm)
+{
+ int i;
+
+ for (i = 0; i <= sbm_max_index(); ++i) {
+ if (sbm[i].bitmap)
+ return false;
+ }
+
+ return true;
+}
+
+int sbm_find_next_bit(struct sbm *sbm, int start)
+{
+ u32 nr = runtime_const_shift_right_32(start, __sbm_shift);
+ u32 bit = runtime_const_mask_32(start, __sbm_mask);
+ unsigned long tmp = 0, mask = (~0UL) << bit;
+
+ for (; nr <= sbm_max_index(); nr++) {
+ tmp = sbm[nr].bitmap & mask;
+ if (tmp)
+ break;
+ /*
+ * Consider full bitmask from
+ * second iteration.
+ */
+ mask = ~0UL;
+ }
+ if (!tmp)
+ return -1;
+ return (nr << __sbm_shift) | __ffs(tmp);
+}
+
/*
* Certain architectures may skip initializing sbm propoerties
* while having an arch_sbm_cpu_instance_id() definition.
@@ -105,6 +149,8 @@ int alloc_sbm_index(int cpu)
_sbm_idx_to_cpu()[idx] = cpu;
_sbm_cpu_to_idx()[cpu] = idx;
+ WRITE_ONCE(sbm_max_populated_index, max(sbm_max_populated_index, i));
+
return 0;
}
@@ -127,6 +173,15 @@ void free_sbm_index(int cpu)
if (find_first_bit(&__sbm_idx_metadata[leaf].allocated_mask, BITS_PER_LONG) ==
BITS_PER_LONG)
__sbm_idx_metadata[leaf].instance_id = -1;
+
+ if (leaf == sbm_max_populated_index) {
+ for (idx = leaf - 1; idx > -1; idx--) {
+ if (__sbm_idx_metadata[idx].instance_id != -1)
+ break;
+ }
+
+ WRITE_ONCE(sbm_max_populated_index, max(idx, 0));
+ }
}
void __init sbm_set_topology(int num_instances, int max_threads_per_instance)
--
2.34.1
^ permalink raw reply related [flat|nested] 36+ messages in thread* Re: [RFC PATCH v3 11/13] lib/sbm: Add helpers to allocate, set, clear, and traverse the bits on sbm
2026-10-01 19:28 ` [RFC PATCH v3 11/13] lib/sbm: Add helpers to allocate, set, clear, and traverse the bits on sbm K Prateek Nayak
@ 2026-10-02 9:13 ` sashiko-bot
0 siblings, 0 replies; 36+ messages in thread
From: sashiko-bot @ 2026-10-02 9:13 UTC (permalink / raw)
To: K Prateek Nayak
Cc: Christian Borntraeger, Alexander Gordeev, linux-s390,
Vasily Gorbik, Heiko Carstens
Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
- [High] sched/fair: out-of-bounds memory corruption in __sbm_op() during CPU hotplug
- [High] lib/sbm: premature decrement of max index in free_sbm_index() hides active CPUs
- [Low] lib/sbm: sbm_alloc() massively over-allocates cacheline-aligned memory
--
commit 8d84c933216fc8041e9a487dfb8a2121359eb99f
Author: Peter Zijlstra <peterz@infradead.org>
lib/sbm: Add helpers to allocate, set, clear, and traverse the bits on sbm
This commit introduces a set of helpers to allocate a sparsebitmap (sbm)
structure based on architecture-configured lengths. It provides macros and
inline functions to set, clear, test, and iterate over bits in the sbm array.
> diff --git a/include/linux/sbm.h b/include/linux/sbm.h
> --- a/include/linux/sbm.h
> +++ b/include/linux/sbm.h
[ ... ]
> @@ -9,12 +11,95 @@
> +#define __sbm_op(sbm, func) \
> +({ \
> + int idx = sbm_cpu_to_idx(cpu); \
> + int nr = idx >> __sbm_shift; \
> + int bit = idx & __sbm_mask; \
> + \
> + func(bit, &sbm[nr].bitmap); \
> +})
[Severity: High]
Can this sequence lead to out-of-bounds memory corruption during CPU
hotplug?
When a CPU is brought offline, free_sbm_index() sets its index to -1.
If the dying CPU continues to run the idle task to tear down, it can
trigger tick_nohz_idle_stop_tick() and consequently call sbm_cpu_set().
Because __sbm_op lacks bounds checking on the mapped idx, computing
idx = -1 results in nr = -1, which then unconditionally executes func()
on &sbm[-1].bitmap.
> diff --git a/lib/sbm.c b/lib/sbm.c
> --- a/lib/sbm.c
> +++ b/lib/sbm.c
[ ... ]
> @@ -45,6 +51,44 @@ int sbm_idx_to_cpu(int idx)
[ ... ]
> +struct sbm *sbm_alloc(void)
> +{
> + return kzalloc_objs(struct sbm, sbm_max_threads_per_instance * sbm_num_instance);
> +}
[Severity: Low]
Does this code allocate significantly more memory than needed?
Each struct sbm encapsulates a full bitmask that accommodates
sbm_max_threads_per_instance CPUs. Since the index calculation
nr = idx >> __sbm_shift directly corresponds to the leaf index, the
array only seems to require exactly sbm_num_instance elements.
Allocating sbm_max_threads_per_instance * sbm_num_instance elements
appears to request considerably more cacheline-aligned memory than is
actually necessary to map all bits.
[ ... ]
> @@ -127,6 +173,15 @@ void free_sbm_index(int cpu)
[ ... ]
> if (find_first_bit(&__sbm_idx_metadata[leaf].allocated_mask, BITS_PER_LONG) ==
> BITS_PER_LONG)
> __sbm_idx_metadata[leaf].instance_id = -1;
> +
> + if (leaf == sbm_max_populated_index) {
> + for (idx = leaf - 1; idx > -1; idx--) {
> + if (__sbm_idx_metadata[idx].instance_id != -1)
> + break;
> + }
> +
> + WRITE_ONCE(sbm_max_populated_index, max(idx, 0));
> + }
> }
[Severity: High]
Could this prematurely shrink the maximum index and strand active CPUs?
When clearing a bit in a leaf's allocated_mask, if the leaf matches
sbm_max_populated_index, this loop unconditionally scans backward to find
the next non-empty leaf and overwrites sbm_max_populated_index.
However, it does this without verifying if the current maximum leaf
actually became completely empty (instance_id == -1). If other CPUs
remain actively allocated in that same leaf, they appear to be
immediately hidden since the global max index is truncated below them.
This would cause functions like sbm_empty() and sbm_find_next_bit() to
permanently ignore these active CPUs during NOHZ idle load balancing.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261001192849.74788-1-kprateek.nayak@amd.com?part=11
^ permalink raw reply [flat|nested] 36+ messages in thread
* [RFC PATCH v3 12/13] sched/fair: Allocate nohz.idle_cpus_mask during sched_init_smp()
2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
` (10 preceding siblings ...)
2026-10-01 19:28 ` [RFC PATCH v3 11/13] lib/sbm: Add helpers to allocate, set, clear, and traverse the bits on sbm K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
2026-10-02 9:13 ` sashiko-bot
2026-10-01 19:28 ` [RFC PATCH v3 13/13] sched/fair: Switch nohz.idle_cpus to use sbm K Prateek Nayak
2026-10-03 9:10 ` [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) Chen Yu
13 siblings, 1 reply; 36+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
driver-core
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, Shrikanth Hegde, K Prateek Nayak
nohz.idle_cpus_mask is not required soon after early initialization
since the boot CPU cannot go tickless idle before CPU is activated and
tick is started.
Defer allocation of nohz.idle_cpus_mask to sched_init_smp() via
init_sched_fair_class_smp(). This is preparation to convert
nohz.idle_cpus_mask to sparsebitmap (sbm) which requires architectures
to finish smp_prepare before initializing the core bits.
No functional changes intended.
Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
kernel/sched/core.c | 2 ++
kernel/sched/fair.c | 8 +++++++-
kernel/sched/sched.h | 1 +
3 files changed, 10 insertions(+), 1 deletion(-)
diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index 977f579da410..127aed81d316 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -8849,6 +8849,8 @@ int sched_cpu_dying(unsigned int cpu)
void __init sched_init_smp(void)
{
+ init_sched_fair_class_smp();
+
sched_init_numa(NUMA_NO_NODE);
prandom_init_once(&sched_rnd_state);
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 1c687c3c70f1..a0a659f4c3be 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -15645,6 +15645,13 @@ void show_numa_stats(struct task_struct *p, struct seq_file *m)
}
#endif /* CONFIG_NUMA_BALANCING */
+__init void init_sched_fair_class_smp(void)
+{
+#ifdef CONFIG_NO_HZ_COMMON
+ zalloc_cpumask_var(&nohz.idle_cpus_mask, GFP_NOWAIT);
+#endif
+}
+
__init void init_sched_fair_class(void)
{
int i;
@@ -15666,6 +15673,5 @@ __init void init_sched_fair_class(void)
#ifdef CONFIG_NO_HZ_COMMON
nohz.next_balance = jiffies;
nohz.next_blocked = jiffies;
- zalloc_cpumask_var(&nohz.idle_cpus_mask, GFP_NOWAIT);
#endif
}
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index 7d2ec527b8a2..0726e8fb5967 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -3024,6 +3024,7 @@ extern void update_max_interval(void);
extern void init_sched_dl_class(void);
extern void init_sched_rt_class(void);
extern void init_sched_fair_class(void);
+extern void init_sched_fair_class_smp(void);
extern void resched_curr(struct rq *rq);
extern void resched_curr_lazy(struct rq *rq);
--
2.34.1
^ permalink raw reply related [flat|nested] 36+ messages in thread* [RFC PATCH v3 13/13] sched/fair: Switch nohz.idle_cpus to use sbm
2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
` (11 preceding siblings ...)
2026-10-01 19:28 ` [RFC PATCH v3 12/13] sched/fair: Allocate nohz.idle_cpus_mask during sched_init_smp() K Prateek Nayak
@ 2026-10-01 19:28 ` K Prateek Nayak
2026-10-02 9:13 ` sashiko-bot
2026-10-03 9:10 ` [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) Chen Yu
13 siblings, 1 reply; 36+ messages in thread
From: K Prateek Nayak @ 2026-10-01 19:28 UTC (permalink / raw)
To: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
driver-core
Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, Shrikanth Hegde, K Prateek Nayak
From: Peter Zijlstra <peterz@infradead.org>
With sbm infrastructure in place, convert the global nohz.idle_cpus
cpumask to use sparsebitmap (sbm).
[ prateek: Used sbm_for_each_bit_wrap(), and adapted find_new_ilb() to
the current SMT aware scheme. ]
(Not-yet-)Signed-off-by: Peter Zijlstra <peterz@infradead.org>
Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com>
---
kernel/sched/fair.c | 66 +++++++++++++++++++--------------------------
1 file changed, 27 insertions(+), 39 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index a0a659f4c3be..0b1458c4360e 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -51,8 +51,9 @@
#include <linux/profile.h>
#include <linux/psi.h>
#include <linux/ratelimit.h>
-#include <linux/task_work.h>
#include <linux/rbtree_augmented.h>
+#include <linux/sbm.h>
+#include <linux/task_work.h>
#include <asm/switch_to.h>
@@ -8210,7 +8211,7 @@ static DEFINE_PER_CPU(cpumask_var_t, should_we_balance_tmpmask);
#ifdef CONFIG_NO_HZ_COMMON
static struct {
- cpumask_var_t idle_cpus_mask;
+ struct sbm *sbm;
int has_blocked_load; /* Idle CPUS has blocked load */
int needs_update; /* Newly idle CPUs need their next_balance collated */
unsigned long next_balance; /* in jiffy units */
@@ -14030,7 +14031,7 @@ static inline int on_null_domain(struct rq *rq)
*/
static inline int find_new_ilb(void)
{
- struct cpumask *ilb_cpus;
+ struct cpumask *skip_cpus;
int ilb_cpu, fallback = -1;
lockdep_assert_irqs_disabled();
@@ -14039,18 +14040,22 @@ static inline int find_new_ilb(void)
* Reuse the per-CPU select_rq_mask, which is protected from concurrent
* use on this CPU by having interrupts disabled.
*/
- ilb_cpus = this_cpu_cpumask_var_ptr(select_rq_mask);
- cpumask_and(ilb_cpus, nohz.idle_cpus_mask,
- housekeeping_cpumask(HK_TYPE_KERNEL_NOISE));
+ skip_cpus = this_cpu_cpumask_var_ptr(select_rq_mask);
+ cpumask_clear(skip_cpus);
+
+ sbm_for_each_set_bit(nohz.sbm, idx) {
+ ilb_cpu = sbm_idx_to_cpu(idx);
+
+ if (cpumask_test_cpu(ilb_cpu, skip_cpus))
+ continue;
- for_each_cpu(ilb_cpu, ilb_cpus) {
if (!idle_cpu(ilb_cpu)) {
/*
* Once an idle fallback exists, a busy CPU proves that
* this core cannot be fully idle. Skip its siblings.
*/
if (sched_smt_active() && fallback >= 0)
- cpumask_andnot(ilb_cpus, ilb_cpus, cpu_smt_mask(ilb_cpu));
+ cpumask_or(skip_cpus, skip_cpus, cpu_smt_mask(ilb_cpu));
continue;
}
@@ -14069,8 +14074,7 @@ static inline int find_new_ilb(void)
* The core is not idle, so there is no need to check
* any of its other SMT siblings.
*/
- cpumask_andnot(ilb_cpus, ilb_cpus,
- cpu_smt_mask(ilb_cpu));
+ cpumask_or(skip_cpus, skip_cpus, cpu_smt_mask(ilb_cpu));
continue;
}
@@ -14134,7 +14138,7 @@ static void nohz_balancer_kick(struct rq *rq)
unsigned long now = jiffies;
struct sched_domain_shared *sds;
struct sched_domain *sd;
- int nr_busy, i, cpu = rq->cpu;
+ int nr_busy, cpu = rq->cpu;
unsigned int flags = 0;
if (unlikely(rq->idle_balance))
@@ -14162,11 +14166,7 @@ static void nohz_balancer_kick(struct rq *rq)
if (time_before(now, nohz.next_balance))
goto out;
- /*
- * None are in tickless mode and hence no need for NOHZ idle load
- * balancing
- */
- if (unlikely(cpumask_empty(nohz.idle_cpus_mask)))
+ if (unlikely(sbm_empty(nohz.sbm)))
return;
if (rq->nr_running >= 2) {
@@ -14186,24 +14186,6 @@ static void nohz_balancer_kick(struct rq *rq)
}
}
- sd = rcu_dereference_all(per_cpu(sd_asym_packing, cpu));
- if (sd) {
- /*
- * When ASYM_PACKING; see if there's a more preferred CPU
- * currently idle; in which case, kick the ILB to move tasks
- * around.
- *
- * When balancing between cores, all the SMT siblings of the
- * preferred CPU must be idle.
- */
- for_each_cpu_and(i, sched_domain_span(sd), nohz.idle_cpus_mask) {
- if (sched_asym(sd, i, cpu)) {
- flags |= NOHZ_STATS_KICK | NOHZ_BALANCE_KICK;
- goto out;
- }
- }
- }
-
sd = rcu_dereference_all(per_cpu(sd_asym_cpucapacity, cpu));
if (sd) {
/*
@@ -14270,7 +14252,8 @@ void nohz_balance_exit_idle(struct rq *rq)
return;
rq->nohz_tick_stopped = 0;
- cpumask_clear_cpu(rq->cpu, nohz.idle_cpus_mask);
+ if (cpumask_test_cpu(rq->cpu, housekeeping_cpumask(HK_TYPE_KERNEL_NOISE)))
+ sbm_cpu_clear(nohz.sbm, rq->cpu);
set_cpu_sd_state_busy(rq->cpu);
}
@@ -14324,7 +14307,8 @@ void nohz_balance_enter_idle(int cpu)
rq->nohz_tick_stopped = 1;
- cpumask_set_cpu(cpu, nohz.idle_cpus_mask);
+ if (cpumask_test_cpu(rq->cpu, housekeeping_cpumask(HK_TYPE_KERNEL_NOISE)))
+ sbm_cpu_set(nohz.sbm, rq->cpu);
/*
* Ensures that if nohz_idle_balance() fails to observe our
@@ -14351,7 +14335,7 @@ static bool update_nohz_stats(struct rq *rq)
if (!rq->has_blocked_load)
return false;
- if (!cpumask_test_cpu(cpu, nohz.idle_cpus_mask))
+ if (!sbm_cpu_test(nohz.sbm, cpu))
return false;
if (!time_after(jiffies, READ_ONCE(rq->last_blocked_load_update_tick)))
@@ -14377,6 +14361,7 @@ static void _nohz_idle_balance(struct rq *this_rq, unsigned int flags)
int this_cpu = this_rq->cpu;
int balance_cpu;
struct rq *rq;
+ int start;
WARN_ON_ONCE((flags & NOHZ_KICK_MASK) == NOHZ_BALANCE_KICK);
@@ -14405,7 +14390,10 @@ static void _nohz_idle_balance(struct rq *this_rq, unsigned int flags)
* Start with the next CPU after this_cpu so we will end with this_cpu and let a
* chance for other idle cpu to pull load.
*/
- for_each_cpu_wrap(balance_cpu, nohz.idle_cpus_mask, this_cpu+1) {
+ start = sbm_cpu_to_idx(cpumask_next_wrap(this_cpu, cpu_online_mask));
+ sbm_for_each_set_bit_wrap(nohz.sbm, idx, start) {
+ balance_cpu = sbm_idx_to_cpu(idx);
+
if (!idle_cpu(balance_cpu))
continue;
@@ -15648,7 +15636,7 @@ void show_numa_stats(struct task_struct *p, struct seq_file *m)
__init void init_sched_fair_class_smp(void)
{
#ifdef CONFIG_NO_HZ_COMMON
- zalloc_cpumask_var(&nohz.idle_cpus_mask, GFP_NOWAIT);
+ nohz.sbm = sbm_alloc();
#endif
}
--
2.34.1
^ permalink raw reply related [flat|nested] 36+ messages in thread* Re: [RFC PATCH v3 13/13] sched/fair: Switch nohz.idle_cpus to use sbm
2026-10-01 19:28 ` [RFC PATCH v3 13/13] sched/fair: Switch nohz.idle_cpus to use sbm K Prateek Nayak
@ 2026-10-02 9:13 ` sashiko-bot
0 siblings, 0 replies; 36+ messages in thread
From: sashiko-bot @ 2026-10-02 9:13 UTC (permalink / raw)
To: K Prateek Nayak
Cc: Christian Borntraeger, Heiko Carstens, Vasily Gorbik, linux-s390,
Alexander Gordeev
Thank you for your contribution! Sashiko AI review found 4 potential issue(s) to consider:
- [High] sched/fair: _nohz_idle_balance() aborts during CPU hotplug
- [Medium] sched/fair: NOHZ blocked load decay broken for NO_HZ_FULL CPUs
- [Medium] sched/fair: silent deletion of ASYM_PACKING idle balance kick
- [Low] sched/fair: missing atomic operation before smp_mb__after_atomic()
--
commit 0b0c43c6e4d9b31b812d8407b146c5dc13783dc8
Author: Peter Zijlstra <peterz@infradead.org>
sched/fair: Switch nohz.idle_cpus to use sbm
With sbm infrastructure in place, convert the global nohz.idle_cpus cpumask
to use sparsebitmap (sbm).
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index a0a659f4c3be6..0b1458c4360e3 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -14186,24 +14186,6 @@ static void nohz_balancer_kick(struct rq *rq)
> }
> }
>
> - sd = rcu_dereference_all(per_cpu(sd_asym_packing, cpu));
> - if (sd) {
> - /*
> - * When ASYM_PACKING; see if there's a more preferred CPU
> - * currently idle; in which case, kick the ILB to move tasks
> - * around.
> - *
> - * When balancing between cores, all the SMT siblings of the
> - * preferred CPU must be idle.
> - */
> - for_each_cpu_and(i, sched_domain_span(sd), nohz.idle_cpus_mask) {
> - if (sched_asym(sd, i, cpu)) {
> - flags |= NOHZ_STATS_KICK | NOHZ_BALANCE_KICK;
> - goto out;
> - }
> - }
> - }
> -
[Severity: Medium]
Was it intentional to completely remove the ASYM_PACKING handling block here
without a replacement?
This appears to mean tasks will no longer be proactively migrated to favored
cores during NOHZ idle load balancing on asymmetric capacity systems.
> sd = rcu_dereference_all(per_cpu(sd_asym_cpucapacity, cpu));
> if (sd) {
> /*
[ ... ]
> @@ -14324,7 +14307,8 @@ void nohz_balance_enter_idle(int cpu)
>
> rq->nohz_tick_stopped = 1;
>
> - cpumask_set_cpu(cpu, nohz.idle_cpus_mask);
> + if (cpumask_test_cpu(rq->cpu, housekeeping_cpumask(HK_TYPE_KERNEL_NOISE)))
> + sbm_cpu_set(nohz.sbm, rq->cpu);
>
> /*
> * Ensures that if nohz_idle_balance() fails to observe our
> * @idle_cpus_mask store, it must observe the @has_blocked_load
> * and @needs_update stores.
> */
> smp_mb__after_atomic();
[Severity: Low]
Does this conditional atomic operation violate the smp_mb__after_atomic() API
contract when the CPU is not in the housekeeping mask?
If a CPU is not in the housekeeping mask, sbm_cpu_set() is skipped. The
subsequent smp_mb__after_atomic() then executes without a preceding atomic
operation, which may fail to provide hardware memory ordering on architectures
where it is just a compiler barrier.
[ ... ]
> @@ -14351,7 +14335,7 @@ static bool update_nohz_stats(struct rq *rq)
> if (!rq->has_blocked_load)
> return false;
>
> - if (!cpumask_test_cpu(cpu, nohz.idle_cpus_mask))
> + if (!sbm_cpu_test(nohz.sbm, cpu))
> return false;
[Severity: Medium]
Will this prevent NO_HZ_FULL CPUs from ever decaying their blocked load when
idle?
When an isolated CPU enters idle, it is excluded from nohz.sbm due to the
housekeeping_cpumask() check in nohz_balance_enter_idle(). Consequently,
update_nohz_stats() entirely skips it because it checks sbm_cpu_test(), which
could lead to stale load averages and incorrect scheduling decisions for those
CPUs.
>
> if (!time_after(jiffies, READ_ONCE(rq->last_blocked_load_update_tick)))
[ ... ]
> @@ -14405,7 +14390,10 @@ static void _nohz_idle_balance(struct rq *this_rq, unsigned int flags)
> * Start with the next CPU after this_cpu so we will end with this_cpu and let a
> * chance for other idle cpu to pull load.
> */
> - for_each_cpu_wrap(balance_cpu, nohz.idle_cpus_mask, this_cpu+1) {
> + start = sbm_cpu_to_idx(cpumask_next_wrap(this_cpu, cpu_online_mask));
> + sbm_for_each_set_bit_wrap(nohz.sbm, idx, start) {
> + balance_cpu = sbm_idx_to_cpu(idx);
> +
[Severity: High]
Can this iteration abort entirely during CPU hotplug?
_nohz_idle_balance() uses cpumask_next_wrap() on cpu_online_mask. During CPU
teardown, sched_cpu_deactivate() calls free_sbm_index(cpu) which sets the CPU's
SBM index to -1, but the CPU remains in cpu_online_mask until later in the
teardown process.
If the ILB samples this offlining CPU, sbm_cpu_to_idx() returns -1, which is
passed as `start` to sbm_for_each_set_bit_wrap().
Since sbm_find_next_bit_wrap() returns -1 when start is -1:
include/linux/sbm.h:sbm_find_next_bit_wrap() {
...
bit = sbm_find_next_bit(sbm, 0);
return bit < start ? bit : -1;
}
won't this cause the loop to terminate immediately without balancing any CPUs?
> if (!idle_cpu(balance_cpu))
> continue;
>
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261001192849.74788-1-kprateek.nayak@amd.com?part=13
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm)
2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
` (12 preceding siblings ...)
2026-10-01 19:28 ` [RFC PATCH v3 13/13] sched/fair: Switch nohz.idle_cpus to use sbm K Prateek Nayak
@ 2026-10-03 9:10 ` Chen Yu
2026-10-04 6:13 ` K Prateek Nayak
13 siblings, 1 reply; 36+ messages in thread
From: Chen Yu @ 2026-10-03 9:10 UTC (permalink / raw)
To: K Prateek Nayak
Cc: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
driver-core, Sudeep Holla, Greg Kroah-Hartman, Rafael J. Wysocki,
Danilo Krummrich, Huacai Chen, Thomas Bogendoerfer, Jiaxun Yang,
Madhavan Srinivasan, Heiko Carstens, Vasily Gorbik,
Alexander Gordeev, David S. Miller, Andreas Larsson,
Thomas Gleixner, Borislav Petkov, Dave Hansen, x86,
Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, Shrikanth Hegde, WANG Xuerui,
Michael Ellerman, Nicholas Piggin, Christophe Leroy,
Christian Borntraeger, Sven Schnelle, H. Peter Anvin
Hello Prateek!
Thanks for bringing this interesting topic,
On Thu, Oct 01, 2026 at 07:28:36PM +0000, K Prateek Nayak wrote:
>
> Problem
> =======
>
> %cycles vs global mask operation
>
> global mask : 100.0000% (var: 3.28%)
> per-NUMA mask : 32.9209% (var: 7.77%)
> per-LLC mask : 1.2977% (var: 4.85%)
> per-LLC mask (u8 operation; no LOCK prefix) : 0.4930% (var: 0.83%)
>
This shows a significant latency improvement, especially in the per-LLC (u64, u8) case.
May I know if schbench / sched-messaging were used?
>
> Future work
> ===========
>
> o Interoperability with cpumaks since sbm lose the crucial optimizations
> that come naturally from for_each_cpu_and() iterations.
>
> o Different data representation - using the u8 variant for updates and
> then perform a "gather" operation to build a dense mask.
>
> o Extending sbm work to help in wakeup (and possibly resurrect Mel's
> optimization from [4] in some form). The current sbm is still far away
> from being used for wakeups since updates to sbm leaf, even on a
> 16CPUs per LLC system is visible in benchmark performance (~8-10%).
>
If we leverage sbm for the wakeup path, it is a per-LLC mask, there seems to be
no much difference from Mel Gorman's proposal of allocating per sd_share
unsigned long idle_cpus_span[]? The frequent update to this mask might still cause
c2c latency within 1 LLC. A wild guess is that maybe the u8 version is more suitable,
because it has only max-to-8 CPUs touching the mask at the same time?
My understanding is that the major case that sbm could fit is turning the global bitmask
into a per-LLC bitmask, because it mainly avoids CPUs on different LLC/node writing the
same cache line frequently (nohz.idle_cpus_mask set via nohz_balance_enter_idle() on
many CPUs, etc), which might cause a costly cache RFO event storm. Meanwhile, with sbm,
at the reader side, _nohz_idle_balance() could start scanning from the current CPU to find
an idle CPU, so as to avoid the costly HITM event - the reader is on LLC1, while the writer
is on LLC0 - so maybe:
for_each_cpu_wrap(balance_cpu, nohz.idle_cpus_mask, this_cpu+1)
could start from this_cpu's LLC sibling first, rather than this_cpu + 1, because this_cpu+1
might not always be the LLC sibling of this_cpu.
I found that in the current code, there are also other global mask:
rd->rto_mask(mentioned by Pan Deng when running ffmpeg[1])
rd->dlo_mask
tick_broadcast_**mask
maybe they can also be converted into sbm.
[1] https://lore.kernel.org/lkml/a3207ebf537bbe5605ff5454f63b5604d83a04a0.1753076363.git.pan.deng@intel.com/
thanks,
Chenyu
^ permalink raw reply [flat|nested] 36+ messages in thread* Re: [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm)
2026-10-03 9:10 ` [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) Chen Yu
@ 2026-10-04 6:13 ` K Prateek Nayak
0 siblings, 0 replies; 36+ messages in thread
From: K Prateek Nayak @ 2026-10-04 6:13 UTC (permalink / raw)
To: Chen Yu
Cc: Peter Zijlstra, Chen Yu, Tim Chen, Ingo Molnar, Juri Lelli,
Vincent Guittot, Andrew Morton, Arnd Bergmann, linux-kernel,
linux-arch, linux-s390, linuxppc-dev, linux-mips, loongarch,
driver-core, Sudeep Holla, Greg Kroah-Hartman, Rafael J. Wysocki,
Danilo Krummrich, Huacai Chen, Thomas Bogendoerfer, Jiaxun Yang,
Madhavan Srinivasan, Heiko Carstens, Vasily Gorbik,
Alexander Gordeev, David S. Miller, Andreas Larsson,
Thomas Gleixner, Borislav Petkov, Dave Hansen, x86,
Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, Shrikanth Hegde, WANG Xuerui,
Michael Ellerman, Nicholas Piggin, Christophe Leroy,
Christian Borntraeger, Sven Schnelle, H. Peter Anvin
Hello Chenyu,
On 10/3/2026 2:40 PM, Chen Yu wrote:
> Hello Prateek!
>
> Thanks for bringing this interesting topic,
>
> On Thu, Oct 01, 2026 at 07:28:36PM +0000, K Prateek Nayak wrote:
>>
>> Problem
>> =======
>>
>> %cycles vs global mask operation
>>
>> global mask : 100.0000% (var: 3.28%)
>> per-NUMA mask : 32.9209% (var: 7.77%)
>> per-LLC mask : 1.2977% (var: 4.85%)
>> per-LLC mask (u8 operation; no LOCK prefix) : 0.4930% (var: 0.83%)
>>
>
> This shows a significant latency improvement, especially in the per-LLC (u64, u8) case.
> May I know if schbench / sched-messaging were used?
No, this was a custom benchmark with two threads per
CPU - one setting the CPU on bitmask, and other clearing it,
continuously yielding to each other.
It gives an idea of what the worst case looks like when there
may be short idling followed by a short runtime going in
cycles.
>
>>
>> Future work
>> ===========
>>
>> o Interoperability with cpumaks since sbm lose the crucial optimizations
>> that come naturally from for_each_cpu_and() iterations.
>>
>> o Different data representation - using the u8 variant for updates and
>> then perform a "gather" operation to build a dense mask.
>>
>> o Extending sbm work to help in wakeup (and possibly resurrect Mel's
>> optimization from [4] in some form). The current sbm is still far away
>> from being used for wakeups since updates to sbm leaf, even on a
>> 16CPUs per LLC system is visible in benchmark performance (~8-10%).
>>
>
> If we leverage sbm for the wakeup path, it is a per-LLC mask, there seems to be
> no much difference from Mel Gorman's proposal of allocating per sd_share
> unsigned long idle_cpus_span[]?
Ack! But the current form is still pretty expensive. I see about a
10% overhead of just maintaining that mask which is what I'm trying
to reduce.
> The frequent update to this mask might still cause
> c2c latency within 1 LLC. A wild guess is that maybe the u8 version is more suitable,
> because it has only max-to-8 CPUs touching the mask at the same time?
With a 64B cacheline, single cacheline can contain data for up to
64 CPUs - unlike current sbm that uses first 8-bytes, this used
the whole 64 bytes.
P.S. All versions were tested with 16 CPUs per LLC on my systems.
Going from atomic u64 to plain u8 writes probably avoids an
expensive atomic path in the H/W making them faster.
> My understanding is that the major case that sbm could fit is turning the global bitmask
> into a per-LLC bitmask, because it mainly avoids CPUs on different LLC/node writing the
> same cache line frequently (nohz.idle_cpus_mask set via nohz_balance_enter_idle() on
> many CPUs, etc), which might cause a costly cache RFO event storm. Meanwhile, with sbm,
> at the reader side, _nohz_idle_balance() could start scanning from the current CPU to find
> an idle CPU, so as to avoid the costly HITM event - the reader is on LLC1, while the writer
> is on LLC0 - so maybe:
>
> for_each_cpu_wrap(balance_cpu, nohz.idle_cpus_mask, this_cpu+1)
>
> could start from this_cpu's LLC sibling first, rather than this_cpu + 1, because this_cpu+1
> might not always be the LLC sibling of this_cpu.
Ah! Good point. Let me see if wrapping within the bitmask leaf
first and then going out makes any difference.
>
> I found that in the current code, there are also other global mask:
> rd->rto_mask(mentioned by Pan Deng when running ffmpeg[1])
> rd->dlo_mask
> tick_broadcast_**mask
> maybe they can also be converted into sbm.
Ack! I was juts getting started somewhere to see if there is an
appetite for sbm :-)
>
> [1] https://lore.kernel.org/lkml/a3207ebf537bbe5605ff5454f63b5604d83a04a0.1753076363.git.pan.deng@intel.com/
>
> thanks,
> Chenyu
--
Thanks and Regards,
Prateek
^ permalink raw reply [flat|nested] 36+ messages in thread