* [PATCH] riscv: smp: fix non-SPINWAIT secondary hart rendezvous for fw_dynamic platforms
@ 2026-07-29 8:34 Paul Sherman
2026-07-29 22:13 ` Samuel Holland
2026-07-30 10:16 ` Conor Dooley
0 siblings, 2 replies; 4+ messages in thread
From: Paul Sherman @ 2026-07-29 8:34 UTC (permalink / raw)
To: palmer, pjw, aou; +Cc: alex, linux-riscv, linux-kernel, Paul Sherman
On platforms where firmware (e.g. fw_dynamic) releases all harts to the
same Linux entry point simultaneously, the hart selected by OpenSBI as
Domain0 Boot HART is already in SBI_HSM_STATE_STARTED when Linux later
attempts to bring it online as a secondary CPU via SBI HART_START.
OpenSBI correctly returns SBI_ERR_ALREADY_STARTED, but Linux has no
recovery path: with CONFIG_RISCV_BOOT_SPINWAIT=n, there is no secondary
wait path in _start_kernel for harts that entered Linux directly from
firmware, and sbi_cpu_start() has no handler for SBI_ERR_ALREADY_STARTED.
This causes one CPU to be permanently dropped per boot. The missing CPU
is always the OpenSBI Domain0 Boot HART, which varies between boots on
Sophgo SG2042 (hart 1, 2, or 3), explaining the apparent 'moving victim'.
Fix with three cooperating changes:
1. Initialize boot_cpu_hartid to INVALID_HARTID instead of relying on
BSS zero-initialization. Without this, boot_cpu_hartid aliases with
hart 0, causing hart 0 to always appear to win the boot CPU race
regardless of which hart actually stored its hartid first.
2. Add a non-SPINWAIT secondary wait path in _start_kernel. When
firmware releases multiple harts to the same entry point, non-primary
harts divert into the existing spinwait rendezvous arrays (previously
used only by CONFIG_RISCV_BOOT_SPINWAIT) and wait for cpu_start() to
provide boot data before proceeding to secondary startup.
3. Handle SBI_ERR_ALREADY_STARTED (-EALREADY) in sbi_cpu_start(). When
HART_START returns -EALREADY, the hart is already executing in Linux
and spinning in .Lwait_for_cpu_up_sbi. Write the spinwait rendezvous
arrays to release it into secondary startup, matching the approach
used by cpu_ops_spinwait.c.
The arrays __cpu_spinwait_stack_pointer and __cpu_spinwait_task_pointer
are defined unconditionally in cpu_ops_spinwait.c but their extern
declarations in head.h were guarded by CONFIG_RISCV_BOOT_SPINWAIT.
Move the declarations outside the guard since the arrays are always
present and now used by both boot paths.
Note: The NR_CPUS bound check mirrors the identical pattern in
cpu_ops_spinwait.c:32 which guards the same arrays against out-of-range
hartids on platforms with discontiguous hart numbering.
Link: https://lore.kernel.org/linux-riscv/20260727221508.5179-1-shermanpauldylan@gmail.com/
Tested-on: Milk-V Pioneer (Sophgo SG2042, 64-hart RISC-V, 4-NUMA nodes,
128GB DDR4, OpenSBI v1.5, Linux v7.2-rc5)
Result: boot_cpu_hartid correctly reflects Domain0 Boot HART, all 64
CPUs online in 2.7 seconds (was 63 CPUs, boot always on hart 0)
Signed-off-by: Paul Sherman <shermanpauldylan@gmail.com>
---
arch/riscv/kernel/cpu_ops_sbi.c | 41 ++++++++++++++++++++++++++++++++-
arch/riscv/kernel/head.S | 25 ++++++++++++++++++++
arch/riscv/kernel/head.h | 2 --
arch/riscv/kernel/setup.c | 9 +++++++-
4 files changed, 73 insertions(+), 4 deletions(-)
diff --git a/arch/riscv/kernel/cpu_ops_sbi.c b/arch/riscv/kernel/cpu_ops_sbi.c
index ee6e4b5cc39e9..3268fda5be35f 100644
--- a/arch/riscv/kernel/cpu_ops_sbi.c
+++ b/arch/riscv/kernel/cpu_ops_sbi.c
@@ -12,6 +12,7 @@
#include <asm/cpu_ops_sbi.h>
#include <asm/sbi.h>
#include <asm/smp.h>
+#include "head.h"
extern char secondary_start_sbi[];
const struct cpu_operations cpu_ops_sbi;
@@ -23,6 +24,18 @@ const struct cpu_operations cpu_ops_sbi;
*/
static struct sbi_hart_boot_data boot_data[NR_CPUS];
+#ifndef CONFIG_RISCV_BOOT_SPINWAIT
+/*
+ * Secondary hart rendezvous arrays, shared with head.S.
+ * These arrays are named for historical reasons after the spinwait
+ * boot protocol, but serve a generic purpose: holding per-hart boot
+ * data until a secondary hart is ready to proceed. Defined here when
+ * CONFIG_RISCV_BOOT_SPINWAIT=n; otherwise defined in cpu_ops_spinwait.c.
+ */
+void *__cpu_spinwait_stack_pointer[NR_CPUS] __section(".data");
+void *__cpu_spinwait_task_pointer[NR_CPUS] __section(".data");
+#endif
+
static int sbi_hsm_hart_start(unsigned long hartid, unsigned long saddr,
unsigned long priv)
{
@@ -62,6 +75,7 @@ static int sbi_cpu_start(unsigned int cpuid, struct task_struct *tidle)
unsigned long boot_addr = __pa_symbol(secondary_start_sbi);
unsigned long hartid = cpuid_to_hartid_map(cpuid);
unsigned long hsm_data;
+ int ret;
struct sbi_hart_boot_data *bdata = &boot_data[cpuid];
/* Make sure tidle is updated */
@@ -71,7 +85,32 @@ static int sbi_cpu_start(unsigned int cpuid, struct task_struct *tidle)
/* Make sure boot data is updated */
smp_mb();
hsm_data = __pa(bdata);
- return sbi_hsm_hart_start(hartid, boot_addr, hsm_data);
+
+ ret = sbi_hsm_hart_start(hartid, boot_addr, hsm_data);
+
+ /*
+ * The firmware boot hart enters Linux directly from the bootloader
+ * and is already in SBI_HSM_STATE_STARTED when Linux attempts to
+ * bring it online as a secondary CPU. HART_START correctly returns
+ * SBI_ERR_ALREADY_STARTED in this case. The hart is spinning in
+ * .Lwait_for_cpu_up_sbi waiting for boot data - write the spinwait
+ * rendezvous arrays to release it into secondary startup.
+ *
+ * Guard against invalid or out-of-range hartids, matching the
+ * same constraint enforced in cpu_ops_spinwait.c.
+ */
+ if (ret == -EALREADY) {
+ if (hartid != INVALID_HARTID &&
+ hartid < (unsigned long)NR_CPUS) { /* array bound */
+ /* Ensure bdata writes visible before spinwait arrays */
+ smp_wmb();
+ WRITE_ONCE(__cpu_spinwait_stack_pointer[hartid],
+ task_pt_regs(tidle));
+ WRITE_ONCE(__cpu_spinwait_task_pointer[hartid], tidle);
+ }
+ ret = 0;
+ }
+ return ret;
}
#ifdef CONFIG_HOTPLUG_CPU
diff --git a/arch/riscv/kernel/head.S b/arch/riscv/kernel/head.S
index f6a8ca49e6277..da48f5efd928b 100644
--- a/arch/riscv/kernel/head.S
+++ b/arch/riscv/kernel/head.S
@@ -281,6 +281,31 @@ SYM_CODE_START(_start_kernel)
la a2, boot_cpu_hartid
REG_S a0, (a2)
+#ifndef CONFIG_RISCV_BOOT_SPINWAIT
+ /*
+ * On platforms where firmware releases all harts to the same
+ * entry point (e.g. fw_dynamic), non-primary harts must divert
+ * here before MMU setup. Wait in the spinwait rendezvous arrays
+ * until cpu_start() provides boot data. a0 = hartid.
+ */
+ REG_L a3, (a2)
+ beq a0, a3, .Lprimary_hart
+ slli a3, a0, LGREG
+ la a1, __cpu_spinwait_stack_pointer
+ la a2, __cpu_spinwait_task_pointer
+ add a1, a3, a1
+ add a2, a3, a2
+.Lwait_for_cpu_up_sbi:
+ fence r, r
+ REG_L sp, (a1)
+ REG_L tp, (a2)
+ beqz sp, .Lwait_for_cpu_up_sbi
+ beqz tp, .Lwait_for_cpu_up_sbi
+ fence
+ tail .Lsecondary_start_common
+.Lprimary_hart:
+#endif /* !CONFIG_RISCV_BOOT_SPINWAIT */
+
/* Initialize page tables and relocate to virtual addresses */
la tp, init_task
la sp, init_thread_union + THREAD_SIZE
diff --git a/arch/riscv/kernel/head.h b/arch/riscv/kernel/head.h
index 05a04bef442b1..1b34f8e655581 100644
--- a/arch/riscv/kernel/head.h
+++ b/arch/riscv/kernel/head.h
@@ -12,9 +12,7 @@ extern atomic_t hart_lottery;
asmlinkage void __init setup_vm(uintptr_t dtb_pa);
-#ifdef CONFIG_RISCV_BOOT_SPINWAIT
extern void *__cpu_spinwait_stack_pointer[];
extern void *__cpu_spinwait_task_pointer[];
-#endif
#endif /* __ASM_HEAD_H */
diff --git a/arch/riscv/kernel/setup.c b/arch/riscv/kernel/setup.c
index 52d1d2b8f338b..d2eecb7ad4965 100644
--- a/arch/riscv/kernel/setup.c
+++ b/arch/riscv/kernel/setup.c
@@ -47,7 +47,14 @@
* BSS.
*/
atomic_t hart_lottery __section(".sdata");
-unsigned long boot_cpu_hartid;
+/*
+ * Initialize to INVALID_HARTID so that the first hart to store its
+ * hartid in head.S is unambiguously the boot CPU. Without this,
+ * boot_cpu_hartid starts as 0 (BSS), which aliases with hart 0 and
+ * causes hart 0 to always appear to win the boot CPU race regardless
+ * of which hart actually wrote first.
+ */
+unsigned long boot_cpu_hartid = INVALID_HARTID;
EXPORT_SYMBOL_GPL(boot_cpu_hartid);
/*
--
2.53.0
_______________________________________________
linux-riscv mailing list
linux-riscv@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-riscv
^ permalink raw reply related [flat|nested] 4+ messages in thread
* Re: [PATCH] riscv: smp: fix non-SPINWAIT secondary hart rendezvous for fw_dynamic platforms
2026-07-29 8:34 [PATCH] riscv: smp: fix non-SPINWAIT secondary hart rendezvous for fw_dynamic platforms Paul Sherman
@ 2026-07-29 22:13 ` Samuel Holland
2026-07-30 1:43 ` Paul Sherman
2026-07-30 10:16 ` Conor Dooley
1 sibling, 1 reply; 4+ messages in thread
From: Samuel Holland @ 2026-07-29 22:13 UTC (permalink / raw)
To: Paul Sherman, palmer, pjw; +Cc: alex, linux-riscv, linux-kernel
Hi,
On 2026-07-29 3:34 AM, Paul Sherman wrote:
> On platforms where firmware (e.g. fw_dynamic) releases all harts to the
> same Linux entry point simultaneously, the hart selected by OpenSBI as
> Domain0 Boot HART is already in SBI_HSM_STATE_STARTED when Linux later
> attempts to bring it online as a secondary CPU via SBI HART_START.
> OpenSBI correctly returns SBI_ERR_ALREADY_STARTED, but Linux has no
> recovery path: with CONFIG_RISCV_BOOT_SPINWAIT=n, there is no secondary
> wait path in _start_kernel for harts that entered Linux directly from
> firmware, and sbi_cpu_start() has no handler for SBI_ERR_ALREADY_STARTED.
There is no need to handle SBI_ERR_ALREADY_STARTED because the contract for the
"ordered booting" entry method (see
https://kernel.org/doc/html/latest/arch/riscv/boot.html#kernel-entry) requires
that only one hart is in the STARTED state when the kernel begins executing.
It sounds like your firmware implements something closer to the
RISCV_BOOT_SPINWAIT entry method, so if you disable CONFIG_RISCV_BOOT_SPINWAIT,
it is unsurprising that there would be problems.
Now, the kernel assumes that if the SBI HSM extension is present, the firmware
implements the "ordered booting" entry method. This seems like a reasonable
assumption. Why doesn't your firmware do that? Can you fix your firmware?
Regards,
Samuel
> This causes one CPU to be permanently dropped per boot. The missing CPU
> is always the OpenSBI Domain0 Boot HART, which varies between boots on
> Sophgo SG2042 (hart 1, 2, or 3), explaining the apparent 'moving victim'.
>
> Fix with three cooperating changes:
>
> 1. Initialize boot_cpu_hartid to INVALID_HARTID instead of relying on
> BSS zero-initialization. Without this, boot_cpu_hartid aliases with
> hart 0, causing hart 0 to always appear to win the boot CPU race
> regardless of which hart actually stored its hartid first.
>
> 2. Add a non-SPINWAIT secondary wait path in _start_kernel. When
> firmware releases multiple harts to the same entry point, non-primary
> harts divert into the existing spinwait rendezvous arrays (previously
> used only by CONFIG_RISCV_BOOT_SPINWAIT) and wait for cpu_start() to
> provide boot data before proceeding to secondary startup.
>
> 3. Handle SBI_ERR_ALREADY_STARTED (-EALREADY) in sbi_cpu_start(). When
> HART_START returns -EALREADY, the hart is already executing in Linux
> and spinning in .Lwait_for_cpu_up_sbi. Write the spinwait rendezvous
> arrays to release it into secondary startup, matching the approach
> used by cpu_ops_spinwait.c.
>
> The arrays __cpu_spinwait_stack_pointer and __cpu_spinwait_task_pointer
> are defined unconditionally in cpu_ops_spinwait.c but their extern
> declarations in head.h were guarded by CONFIG_RISCV_BOOT_SPINWAIT.
> Move the declarations outside the guard since the arrays are always
> present and now used by both boot paths.
>
> Note: The NR_CPUS bound check mirrors the identical pattern in
> cpu_ops_spinwait.c:32 which guards the same arrays against out-of-range
> hartids on platforms with discontiguous hart numbering.
>
> Link: https://lore.kernel.org/linux-riscv/20260727221508.5179-1-shermanpauldylan@gmail.com/
> Tested-on: Milk-V Pioneer (Sophgo SG2042, 64-hart RISC-V, 4-NUMA nodes,
> 128GB DDR4, OpenSBI v1.5, Linux v7.2-rc5)
> Result: boot_cpu_hartid correctly reflects Domain0 Boot HART, all 64
> CPUs online in 2.7 seconds (was 63 CPUs, boot always on hart 0)
> Signed-off-by: Paul Sherman <shermanpauldylan@gmail.com>
> ---
> arch/riscv/kernel/cpu_ops_sbi.c | 41 ++++++++++++++++++++++++++++++++-
> arch/riscv/kernel/head.S | 25 ++++++++++++++++++++
> arch/riscv/kernel/head.h | 2 --
> arch/riscv/kernel/setup.c | 9 +++++++-
> 4 files changed, 73 insertions(+), 4 deletions(-)
>
> diff --git a/arch/riscv/kernel/cpu_ops_sbi.c b/arch/riscv/kernel/cpu_ops_sbi.c
> index ee6e4b5cc39e9..3268fda5be35f 100644
> --- a/arch/riscv/kernel/cpu_ops_sbi.c
> +++ b/arch/riscv/kernel/cpu_ops_sbi.c
> @@ -12,6 +12,7 @@
> #include <asm/cpu_ops_sbi.h>
> #include <asm/sbi.h>
> #include <asm/smp.h>
> +#include "head.h"
>
> extern char secondary_start_sbi[];
> const struct cpu_operations cpu_ops_sbi;
> @@ -23,6 +24,18 @@ const struct cpu_operations cpu_ops_sbi;
> */
> static struct sbi_hart_boot_data boot_data[NR_CPUS];
>
> +#ifndef CONFIG_RISCV_BOOT_SPINWAIT
> +/*
> + * Secondary hart rendezvous arrays, shared with head.S.
> + * These arrays are named for historical reasons after the spinwait
> + * boot protocol, but serve a generic purpose: holding per-hart boot
> + * data until a secondary hart is ready to proceed. Defined here when
> + * CONFIG_RISCV_BOOT_SPINWAIT=n; otherwise defined in cpu_ops_spinwait.c.
> + */
> +void *__cpu_spinwait_stack_pointer[NR_CPUS] __section(".data");
> +void *__cpu_spinwait_task_pointer[NR_CPUS] __section(".data");
> +#endif
> +
> static int sbi_hsm_hart_start(unsigned long hartid, unsigned long saddr,
> unsigned long priv)
> {
> @@ -62,6 +75,7 @@ static int sbi_cpu_start(unsigned int cpuid, struct task_struct *tidle)
> unsigned long boot_addr = __pa_symbol(secondary_start_sbi);
> unsigned long hartid = cpuid_to_hartid_map(cpuid);
> unsigned long hsm_data;
> + int ret;
> struct sbi_hart_boot_data *bdata = &boot_data[cpuid];
>
> /* Make sure tidle is updated */
> @@ -71,7 +85,32 @@ static int sbi_cpu_start(unsigned int cpuid, struct task_struct *tidle)
> /* Make sure boot data is updated */
> smp_mb();
> hsm_data = __pa(bdata);
> - return sbi_hsm_hart_start(hartid, boot_addr, hsm_data);
> +
> + ret = sbi_hsm_hart_start(hartid, boot_addr, hsm_data);
> +
> + /*
> + * The firmware boot hart enters Linux directly from the bootloader
> + * and is already in SBI_HSM_STATE_STARTED when Linux attempts to
> + * bring it online as a secondary CPU. HART_START correctly returns
> + * SBI_ERR_ALREADY_STARTED in this case. The hart is spinning in
> + * .Lwait_for_cpu_up_sbi waiting for boot data - write the spinwait
> + * rendezvous arrays to release it into secondary startup.
> + *
> + * Guard against invalid or out-of-range hartids, matching the
> + * same constraint enforced in cpu_ops_spinwait.c.
> + */
> + if (ret == -EALREADY) {
> + if (hartid != INVALID_HARTID &&
> + hartid < (unsigned long)NR_CPUS) { /* array bound */
> + /* Ensure bdata writes visible before spinwait arrays */
> + smp_wmb();
> + WRITE_ONCE(__cpu_spinwait_stack_pointer[hartid],
> + task_pt_regs(tidle));
> + WRITE_ONCE(__cpu_spinwait_task_pointer[hartid], tidle);
> + }
> + ret = 0;
> + }
> + return ret;
> }
>
> #ifdef CONFIG_HOTPLUG_CPU
> diff --git a/arch/riscv/kernel/head.S b/arch/riscv/kernel/head.S
> index f6a8ca49e6277..da48f5efd928b 100644
> --- a/arch/riscv/kernel/head.S
> +++ b/arch/riscv/kernel/head.S
> @@ -281,6 +281,31 @@ SYM_CODE_START(_start_kernel)
> la a2, boot_cpu_hartid
> REG_S a0, (a2)
>
> +#ifndef CONFIG_RISCV_BOOT_SPINWAIT
> + /*
> + * On platforms where firmware releases all harts to the same
> + * entry point (e.g. fw_dynamic), non-primary harts must divert
> + * here before MMU setup. Wait in the spinwait rendezvous arrays
> + * until cpu_start() provides boot data. a0 = hartid.
> + */
> + REG_L a3, (a2)
> + beq a0, a3, .Lprimary_hart
> + slli a3, a0, LGREG
> + la a1, __cpu_spinwait_stack_pointer
> + la a2, __cpu_spinwait_task_pointer
> + add a1, a3, a1
> + add a2, a3, a2
> +.Lwait_for_cpu_up_sbi:
> + fence r, r
> + REG_L sp, (a1)
> + REG_L tp, (a2)
> + beqz sp, .Lwait_for_cpu_up_sbi
> + beqz tp, .Lwait_for_cpu_up_sbi
> + fence
> + tail .Lsecondary_start_common
> +.Lprimary_hart:
> +#endif /* !CONFIG_RISCV_BOOT_SPINWAIT */
> +
> /* Initialize page tables and relocate to virtual addresses */
> la tp, init_task
> la sp, init_thread_union + THREAD_SIZE
> diff --git a/arch/riscv/kernel/head.h b/arch/riscv/kernel/head.h
> index 05a04bef442b1..1b34f8e655581 100644
> --- a/arch/riscv/kernel/head.h
> +++ b/arch/riscv/kernel/head.h
> @@ -12,9 +12,7 @@ extern atomic_t hart_lottery;
>
> asmlinkage void __init setup_vm(uintptr_t dtb_pa);
>
> -#ifdef CONFIG_RISCV_BOOT_SPINWAIT
> extern void *__cpu_spinwait_stack_pointer[];
> extern void *__cpu_spinwait_task_pointer[];
> -#endif
>
> #endif /* __ASM_HEAD_H */
> diff --git a/arch/riscv/kernel/setup.c b/arch/riscv/kernel/setup.c
> index 52d1d2b8f338b..d2eecb7ad4965 100644
> --- a/arch/riscv/kernel/setup.c
> +++ b/arch/riscv/kernel/setup.c
> @@ -47,7 +47,14 @@
> * BSS.
> */
> atomic_t hart_lottery __section(".sdata");
> -unsigned long boot_cpu_hartid;
> +/*
> + * Initialize to INVALID_HARTID so that the first hart to store its
> + * hartid in head.S is unambiguously the boot CPU. Without this,
> + * boot_cpu_hartid starts as 0 (BSS), which aliases with hart 0 and
> + * causes hart 0 to always appear to win the boot CPU race regardless
> + * of which hart actually wrote first.
> + */
> +unsigned long boot_cpu_hartid = INVALID_HARTID;
> EXPORT_SYMBOL_GPL(boot_cpu_hartid);
>
> /*
_______________________________________________
linux-riscv mailing list
linux-riscv@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-riscv
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH] riscv: smp: fix non-SPINWAIT secondary hart rendezvous for fw_dynamic platforms
2026-07-29 22:13 ` Samuel Holland
@ 2026-07-30 1:43 ` Paul Sherman
0 siblings, 0 replies; 4+ messages in thread
From: Paul Sherman @ 2026-07-30 1:43 UTC (permalink / raw)
To: Samuel Holland; +Cc: palmer, pjw, alex, linux-riscv, linux-kernel
Hi Samuel,
Thank you for the explanation and the pointer to the boot protocol
documentation.
On Wed, Jul 29, 2026 at 05:13:56PM -0500, Samuel Holland wrote:
>
> There is no need to handle SBI_ERR_ALREADY_STARTED because the contract for the
> "ordered booting" entry method (see
> https://kernel.org/doc/html/latest/arch/riscv/boot.html#kernel-entry) requires
> that only one hart is in the STARTED state when the kernel begins executing.
I agree that the documented ordered booting contract requires only
one hart to enter the kernel initially, and that HART_START should
therefore not return SBI_ERR_ALREADY_STARTED.
> It sounds like your firmware implements something closer to the
> RISCV_BOOT_SPINWAIT entry method, so if you disable CONFIG_RISCV_BOOT_SPINWAIT,
> it is unsurprising that there would be problems.
On the SG2042 platform I'm investigating, however, that is the
behavior I'm observing on shipping hardware. My investigation so far
has focused on making Linux recover from that situation. Based on
your reply, I'm now wondering whether the more appropriate fix is
in the platform firmware or OpenSBI instead.
> Why doesn't your firmware do that? Can you fix your firmware?
Is your expectation that platforms exhibiting this behavior should
be corrected to satisfy the ordered booting contract, rather than
Linux handling SBI_ERR_ALREADY_STARTED defensively?
Thanks,
Paul
_______________________________________________
linux-riscv mailing list
linux-riscv@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-riscv
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH] riscv: smp: fix non-SPINWAIT secondary hart rendezvous for fw_dynamic platforms
2026-07-29 8:34 [PATCH] riscv: smp: fix non-SPINWAIT secondary hart rendezvous for fw_dynamic platforms Paul Sherman
2026-07-29 22:13 ` Samuel Holland
@ 2026-07-30 10:16 ` Conor Dooley
1 sibling, 0 replies; 4+ messages in thread
From: Conor Dooley @ 2026-07-30 10:16 UTC (permalink / raw)
To: Paul Sherman; +Cc: palmer, pjw, aou, alex, linux-riscv, linux-kernel
[-- Attachment #1.1: Type: text/plain, Size: 2988 bytes --]
On Wed, Jul 29, 2026 at 01:34:14AM -0700, Paul Sherman wrote:
> On platforms where firmware (e.g. fw_dynamic) releases all harts to the
> same Linux entry point simultaneously, the hart selected by OpenSBI as
> Domain0 Boot HART is already in SBI_HSM_STATE_STARTED when Linux later
> attempts to bring it online as a secondary CPU via SBI HART_START.
> OpenSBI correctly returns SBI_ERR_ALREADY_STARTED, but Linux has no
> recovery path: with CONFIG_RISCV_BOOT_SPINWAIT=n, there is no secondary
> wait path in _start_kernel for harts that entered Linux directly from
> firmware, and sbi_cpu_start() has no handler for SBI_ERR_ALREADY_STARTED.
>
> This causes one CPU to be permanently dropped per boot. The missing CPU
> is always the OpenSBI Domain0 Boot HART, which varies between boots on
> Sophgo SG2042 (hart 1, 2, or 3), explaining the apparent 'moving victim'.
>
> Fix with three cooperating changes:
>
> 1. Initialize boot_cpu_hartid to INVALID_HARTID instead of relying on
> BSS zero-initialization. Without this, boot_cpu_hartid aliases with
> hart 0, causing hart 0 to always appear to win the boot CPU race
> regardless of which hart actually stored its hartid first.
>
> 2. Add a non-SPINWAIT secondary wait path in _start_kernel. When
> firmware releases multiple harts to the same entry point, non-primary
> harts divert into the existing spinwait rendezvous arrays (previously
> used only by CONFIG_RISCV_BOOT_SPINWAIT) and wait for cpu_start() to
> provide boot data before proceeding to secondary startup.
>
> 3. Handle SBI_ERR_ALREADY_STARTED (-EALREADY) in sbi_cpu_start(). When
> HART_START returns -EALREADY, the hart is already executing in Linux
> and spinning in .Lwait_for_cpu_up_sbi. Write the spinwait rendezvous
> arrays to release it into secondary startup, matching the approach
> used by cpu_ops_spinwait.c.
>
> The arrays __cpu_spinwait_stack_pointer and __cpu_spinwait_task_pointer
> are defined unconditionally in cpu_ops_spinwait.c but their extern
> declarations in head.h were guarded by CONFIG_RISCV_BOOT_SPINWAIT.
> Move the declarations outside the guard since the arrays are always
> present and now used by both boot paths.
>
> Note: The NR_CPUS bound check mirrors the identical pattern in
> cpu_ops_spinwait.c:32 which guards the same arrays against out-of-range
> hartids on platforms with discontiguous hart numbering.
>
> Link: https://lore.kernel.org/linux-riscv/20260727221508.5179-1-shermanpauldylan@gmail.com/
> Tested-on: Milk-V Pioneer (Sophgo SG2042, 64-hart RISC-V, 4-NUMA nodes,
> 128GB DDR4, OpenSBI v1.5, Linux v7.2-rc5)
> Result: boot_cpu_hartid correctly reflects Domain0 Boot HART, all 64
> CPUs online in 2.7 seconds (was 63 CPUs, boot always on hart 0)
FWIW, neither of these are valid tags, just put this information in the
commit message proper.
> Signed-off-by: Paul Sherman <shermanpauldylan@gmail.com>
[-- Attachment #1.2: signature.asc --]
[-- Type: application/pgp-signature, Size: 228 bytes --]
[-- Attachment #2: Type: text/plain, Size: 161 bytes --]
_______________________________________________
linux-riscv mailing list
linux-riscv@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-riscv
^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-07-30 10:16 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-29 8:34 [PATCH] riscv: smp: fix non-SPINWAIT secondary hart rendezvous for fw_dynamic platforms Paul Sherman
2026-07-29 22:13 ` Samuel Holland
2026-07-30 1:43 ` Paul Sherman
2026-07-30 10:16 ` Conor Dooley
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox