Linux s390 Architecture development
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "K Prateek Nayak" <kprateek.nayak@amd.com>
Cc: "Christian Borntraeger" <borntraeger@linux.ibm.com>,
	"Vasily Gorbik" <gor@linux.ibm.com>,
	linux-s390@vger.kernel.org, "Heiko Carstens" <hca@linux.ibm.com>,
	"Alexander Gordeev" <agordeev@linux.ibm.com>
Subject: Re: [RFC PATCH v3 10/13] lib/sbm: Dynamically allocate sbm index when CPU is activated
Date: Fri, 02 Oct 2026 09:13:20 +0000	[thread overview]
Message-ID: <20261002091321.4E3121F00899@smtp.kernel.org> (raw)
In-Reply-To: <20261001192849.74788-11-kprateek.nayak@amd.com>

Thank you for your contribution! Sashiko AI review found 4 potential issue(s) to consider:
- [High] lib/sbm: array overflow in alloc_sbm_index() due to hardcoded BITS_PER_LONG
- [High] sched/core: premature SBM index free causes UAF and out-of-bounds access
- [High] sched/core: unhandled alloc_sbm_index() failure allows corrupting sbm[-1]
- [Medium] lib/sbm: sentinel value collision with architecture instance IDs corrupts topology
--

commit 525dc1ae45d6fc1670c637f040964c4d0eb2f14a
Author: K Prateek Nayak <kprateek.nayak@amd.com>

lib/sbm: Dynamically allocate sbm index when CPU is activated
    
Add infrastructure to establish CPU to sparsebitmap (sbm) index relation
before CPU is turned active. CPU coming online looks for a free slot
based on its instance ID and acquires a free slot.

> diff --git a/init/main.c b/init/main.c
> index 2613d3f9b3ce9..b7406bd3acc87 100644
> --- a/init/main.c
> +++ b/init/main.c
[ ... ]
> @@ -1652,6 +1653,11 @@ static noinline void __init kernel_init_freeable(void)
>  
>  	smp_prepare_cpus(setup_max_cpus);
>  
> +	sbm_init();
> +
> +	/* Finish initializing boot CPU since it is already active. */
> +	alloc_sbm_index(smp_processor_id());

[Severity: High]
Does this code need to check the return value of alloc_sbm_index()? If it
fails and returns -ENOENT, the boot CPU might proceed without a valid mapping,
potentially exposing the -1 default value to subsystem bitwise operations.

> +
>  	workqueue_init();
>  
>  	init_mm_internals();
> diff --git a/kernel/sched/core.c b/kernel/sched/core.c
> index 0bb86a43a5920..977f579da4108 100644
> --- a/kernel/sched/core.c
> +++ b/kernel/sched/core.c
[ ... ]
> @@ -8636,6 +8637,14 @@ int sched_cpu_activate(unsigned int cpu)
>  	 */
>  	balance_push_set(cpu, false);
>  
> +	alloc_sbm_index(cpu);

[Severity: High]
Similar to the boot CPU initialization, should we handle potential errors from
alloc_sbm_index() here? If the slots are full and it returns -ENOENT, the CPU
is brought fully online without a valid mapping, which might cause memory
corruption when writing to sbm[-1].bitmap.

> +
> +	/*
> +	 * Make sure sbm mappings are visible
> +	 * before CPU is toggled active.
> +	 */
> +	smp_mb();
> +
>  	/*
>  	 * When going up, increment the number of cores with SMT present.
>  	 */
[ ... ]
> @@ -8688,6 +8697,14 @@ int sched_cpu_deactivate(unsigned int cpu)
>  
>  	set_cpu_active(cpu, false);
>  
> +	/*
> +	 * Make sure CPU is inactive before
> +	 * sbm indices are reclaimed.
> +	 */
> +	smp_mb();
> +
> +	free_sbm_index(cpu);

[Severity: High]
Can reclaiming the index before synchronize_rcu() cause out-of-bounds array
accesses? If a concurrent lockless reader iterates the sbm and reads the index
before it is cleared, but calls sbm_idx_to_cpu() after this free_sbm_index()
call, it might receive -1.

> +
>  	/*
>  	 * From this point forward, this CPU will refuse to run any task that
>  	 * is not: migrate_disable() or KTHREAD_IS_PER_CPU, and will actively
> diff --git a/lib/sbm.c b/lib/sbm.c
> index 82280e3306de8..b80ed1d961379 100644
> --- a/lib/sbm.c
> +++ b/lib/sbm.c
[ ... ]
> @@ -15,6 +63,72 @@ int __weak arch_sbm_cpu_instance_id(int cpu)
>  	return 0;
>  }
>  
> +static int sbm_cpu_to_instance(int cpu)
> +{
[ ... ]
> +int alloc_sbm_index(int cpu)
> +{
> +	int cpu_instance = sbm_cpu_to_instance(cpu);
> +	int i, idx = BITS_PER_LONG, free_index = -1;
> +
> +	for (i = 0; i < sbm_num_instance; ++i) {
> +		if (__sbm_idx_metadata[i].instance_id == cpu_instance) {
> +			idx = find_first_zero_bit(&__sbm_idx_metadata[i].allocated_mask,
> +						  BITS_PER_LONG);
> +
> +			if (idx < BITS_PER_LONG)
> +				break;
> +		}
> +		if (free_index == -1 && __sbm_idx_metadata[i].instance_id == -1)
> +			free_index = i;

[Severity: Medium]
Could this condition falsely match if an architecture returns -1 for an
instance ID (like NUMA_NO_NODE)? If the first CPU sets the instance ID to -1,
subsequent CPUs with valid instance IDs might mistakenly see this leaf as
empty, and overwrite the previous CPU's mapping.

> +	}
> +
> +	if (i == sbm_num_instance && free_index == -1)
> +		return -ENOENT;
> +
> +	if (i == sbm_num_instance) {
> +		__sbm_idx_metadata[free_index].instance_id = cpu_instance;
> +		i = free_index;
> +		idx = 0;
> +	}
> +
> +	WARN_ON_ONCE(idx >= sbm_max_threads_per_instance);

[Severity: High]
Does this code overflow into subsequent leaves when
sbm_max_threads_per_instance is smaller than BITS_PER_LONG? Since the
find_first_zero_bit() limit and the break condition above hardcode
BITS_PER_LONG instead of checking against sbm_max_threads_per_instance, the
global idx computed below might overflow the leaf's boundary and corrupt
mappings for other instances.

> +
> +	__set_bit(idx, &__sbm_idx_metadata[i].allocated_mask);
> +
> +	idx = (i << __sbm_shift) + idx;
> +	_sbm_idx_to_cpu()[idx] = cpu;
> +	_sbm_cpu_to_idx()[cpu] = idx;
> +
> +	return 0;
> +}
> +
> +void free_sbm_index(int cpu)
> +{
> +	int idx = sbm_cpu_to_idx(cpu);
> +	u32 leaf;
> +
> +	if (idx < 0)
> +		return;
> +
> +	_sbm_idx_to_cpu()[idx] = -1;
> +	_sbm_cpu_to_idx()[cpu] = -1;
> +
> +	leaf = runtime_const_shift_right_32(idx, __sbm_shift);
> +	idx = runtime_const_mask_32(idx, __sbm_mask);
> +
> +	__clear_bit(idx, &__sbm_idx_metadata[leaf].allocated_mask);

[Severity: High]
If the allocation allowed an index to overflow past
sbm_max_threads_per_instance, will the shifted leaf calculation here resolve
to the wrong leaf and clear an allocated bit belonging to another instance?

> +
> +	if (find_first_bit(&__sbm_idx_metadata[leaf].allocated_mask, BITS_PER_LONG) ==
> +	    BITS_PER_LONG)
> +		__sbm_idx_metadata[leaf].instance_id = -1;
> +}

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261001192849.74788-1-kprateek.nayak@amd.com?part=10

  reply	other threads:[~2026-10-02  9:13 UTC|newest]

Thread overview: 36+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-01 19:28 [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 01/13] lib/sbm: Introduce helpers for architectures to configure LLC properties K Prateek Nayak
2026-10-02  9:13   ` sashiko-bot
2026-10-07  5:43   ` Shrikanth Hegde
2026-10-01 19:28 ` [RFC PATCH v3 02/13] drivers/base/arch_topology: Add support for initializing sbm topology K Prateek Nayak
2026-10-02  9:13   ` sashiko-bot
2026-10-01 19:28 ` [RFC PATCH v3 03/13] LoongArch: Initialize CPU _PXM relation for disabled CPUs from SRAT K Prateek Nayak
2026-10-02  9:13   ` sashiko-bot
2026-10-01 19:28 ` [RFC PATCH v3 04/13] LoongArch: Configure sbm topology during SMP preparation K Prateek Nayak
2026-10-02  9:13   ` sashiko-bot
2026-10-07  3:51   ` [RFC PATCH v3.1 " K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 05/13] MIPS: Initialize sbm topology on multi-node systems K Prateek Nayak
2026-10-02  9:13   ` sashiko-bot
2026-10-01 19:28 ` [RFC PATCH v3 06/13] powerpc/setup: Initialize sbm topology based on coregroup / NUMA topology K Prateek Nayak
2026-10-02  9:13   ` sashiko-bot
2026-10-07 14:43   ` Shrikanth Hegde
2026-10-01 19:28 ` [RFC PATCH v3 07/13] s390/topology: Initialize sbm topology during topology_init_early() K Prateek Nayak
2026-10-02  9:13   ` sashiko-bot
2026-10-07 10:19   ` Mete Durlu
2026-10-01 19:28 ` [RFC PATCH v3 08/13] sparc64: Initialize sbm topology on multi-LLC system K Prateek Nayak
2026-10-02  9:13   ` sashiko-bot
2026-10-01 19:28 ` [RFC PATCH v3 09/13] x86/cpu/topology: Initialize sbm topology after topology parsing K Prateek Nayak
2026-10-02  9:13   ` sashiko-bot
2026-10-03  8:27   ` Chen Yu
2026-10-04  6:17     ` K Prateek Nayak
2026-10-07  3:52   ` [RFC PATCH v3.1 " K Prateek Nayak
2026-10-01 19:28 ` [RFC PATCH v3 10/13] lib/sbm: Dynamically allocate sbm index when CPU is activated K Prateek Nayak
2026-10-02  9:13   ` sashiko-bot [this message]
2026-10-01 19:28 ` [RFC PATCH v3 11/13] lib/sbm: Add helpers to allocate, set, clear, and traverse the bits on sbm K Prateek Nayak
2026-10-02  9:13   ` sashiko-bot
2026-10-01 19:28 ` [RFC PATCH v3 12/13] sched/fair: Allocate nohz.idle_cpus_mask during sched_init_smp() K Prateek Nayak
2026-10-02  9:13   ` sashiko-bot
2026-10-01 19:28 ` [RFC PATCH v3 13/13] sched/fair: Switch nohz.idle_cpus to use sbm K Prateek Nayak
2026-10-02  9:13   ` sashiko-bot
2026-10-03  9:10 ` [RFC PATCH v3 00/13] lib, sched: Introduce sparsebitmap (sbm) Chen Yu
2026-10-04  6:13   ` K Prateek Nayak

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261002091321.4E3121F00899@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=agordeev@linux.ibm.com \
    --cc=borntraeger@linux.ibm.com \
    --cc=gor@linux.ibm.com \
    --cc=hca@linux.ibm.com \
    --cc=kprateek.nayak@amd.com \
    --cc=linux-s390@vger.kernel.org \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox