* [PATCH] arm64: archrandom: avoid trapping ID register read in __cpu_has_rng()
@ 2026-07-20 14:06 Aman Priyadarshi
2026-07-31 14:43 ` Will Deacon
0 siblings, 1 reply; 3+ messages in thread
From: Aman Priyadarshi @ 2026-07-20 14:06 UTC (permalink / raw)
To: catalin.marinas, will
Cc: Jason, linux-arm-kernel, linux-kernel, Aman Priyadarshi,
Ard Biesheuvel
__cpu_has_rng() has an early-boot fallback, taken before the
ARM64_HAS_RNG alternative is patched, that calls
this_cpu_has_cap(ARM64_HAS_RNG). With SCOPE_LOCAL_CPU that resolves the
capability by reading ID_AA64ISAR0_EL1 directly from hardware via
__read_sysreg_by_encoding(), on every invocation.
Until the CRNG is seeded, crng_make_state() routes every get_random_*()
through extract_entropy(), which drains architectural entropy via
arch_get_random_seed_longs()/arch_get_random_longs() and so calls
__cpu_has_rng() several times per request. On a direct (non-EFI) boot
there is no bootloader seed, so the CRNG stays unseeded for much of
boot and essentially every early randomness consumer takes this path.
Under virtualization this is costly: the hypervisor traps guest
accesses to the ID registers (HCR_EL2.TID3), making each read a vmexit,
producing ~200k trapped ID_AA64ISAR0_EL1 reads during boot.
The register value is invariant, so read the sanitised feature register
instead. read_sanitised_ftr_reg() returns the cached value from
arm64_ftr_regs[] with no sysreg access, and hence no trap. That array
is populated by cpuinfo_store_boot_cpu() in smp_prepare_boot_cpu(),
before the first early RNG use in random_init_early(), so it is always
valid here.
With this change the trapped reads drop from ~200k to handful number of
times, and the boot time drops roughly by 6.3% in the test environment.
Fixes: 2c03e16f4499 ("random: remove early archrandom abstraction")
Cc: Jason A. Donenfeld <Jason@zx2c4.com>
Cc: Ard Biesheuvel <ardb@kernel.org>
Cc: Will Deacon <will@kernel.org>
Signed-off-by: Aman Priyadarshi <amanp@apple.com>
---
arch/arm64/include/asm/archrandom.h | 18 ++++++++++++++++--
1 file changed, 16 insertions(+), 2 deletions(-)
diff --git a/arch/arm64/include/asm/archrandom.h b/arch/arm64/include/asm/archrandom.h
index 8babfbe31f95..8067e9a35641 100644
--- a/arch/arm64/include/asm/archrandom.h
+++ b/arch/arm64/include/asm/archrandom.h
@@ -61,8 +61,22 @@ static inline bool __arm64_rndrrs(unsigned long *v)
static __always_inline bool __cpu_has_rng(void)
{
- if (unlikely(!system_capabilities_finalized() && !preemptible()))
- return this_cpu_has_cap(ARM64_HAS_RNG);
+ if (unlikely(!system_capabilities_finalized() && !preemptible())) {
+ /*
+ * Until the ARM64_HAS_RNG alternative is patched we can't use
+ * the static-branch form, so consult the feature register
+ * directly. Don't use this_cpu_has_cap() here: it reads
+ * ID_AA64ISAR0_EL1 from hardware on every call, under
+ * virtualization each ID register read traps to the hypervisor
+ * (HCR_EL2.TID3) -- producing a storm of vmexits during boot.
+ * The sanitised value is cached in memory.
+ */
+ u64 isar0 = read_sanitised_ftr_reg(SYS_ID_AA64ISAR0_EL1);
+
+ return cpuid_feature_extract_unsigned_field(isar0,
+ ID_AA64ISAR0_EL1_RNDR_SHIFT) >=
+ ID_AA64ISAR0_EL1_RNDR_IMP;
+ }
return alternative_has_cap_unlikely(ARM64_HAS_RNG);
}
--
2.54.0 (Apple Git-156)
^ permalink raw reply related [flat|nested] 3+ messages in thread
* Re: [PATCH] arm64: archrandom: avoid trapping ID register read in __cpu_has_rng()
2026-07-20 14:06 [PATCH] arm64: archrandom: avoid trapping ID register read in __cpu_has_rng() Aman Priyadarshi
@ 2026-07-31 14:43 ` Will Deacon
2026-07-31 17:26 ` Aman Priyadarshi
0 siblings, 1 reply; 3+ messages in thread
From: Will Deacon @ 2026-07-31 14:43 UTC (permalink / raw)
To: Aman Priyadarshi
Cc: catalin.marinas, Jason, linux-arm-kernel, linux-kernel,
Ard Biesheuvel
Hi Aman,
On Mon, Jul 20, 2026 at 03:06:15PM +0100, Aman Priyadarshi wrote:
> __cpu_has_rng() has an early-boot fallback, taken before the
> ARM64_HAS_RNG alternative is patched, that calls
> this_cpu_has_cap(ARM64_HAS_RNG). With SCOPE_LOCAL_CPU that resolves the
> capability by reading ID_AA64ISAR0_EL1 directly from hardware via
> __read_sysreg_by_encoding(), on every invocation.
>
> Until the CRNG is seeded, crng_make_state() routes every get_random_*()
> through extract_entropy(), which drains architectural entropy via
> arch_get_random_seed_longs()/arch_get_random_longs() and so calls
> __cpu_has_rng() several times per request. On a direct (non-EFI) boot
> there is no bootloader seed, so the CRNG stays unseeded for much of
> boot and essentially every early randomness consumer takes this path.
>
> Under virtualization this is costly: the hypervisor traps guest
> accesses to the ID registers (HCR_EL2.TID3), making each read a vmexit,
> producing ~200k trapped ID_AA64ISAR0_EL1 reads during boot.
>
> The register value is invariant, so read the sanitised feature register
> instead. read_sanitised_ftr_reg() returns the cached value from
> arm64_ftr_regs[] with no sysreg access, and hence no trap. That array
> is populated by cpuinfo_store_boot_cpu() in smp_prepare_boot_cpu(),
> before the first early RNG use in random_init_early(), so it is always
> valid here.
>
> With this change the trapped reads drop from ~200k to handful number of
> times, and the boot time drops roughly by 6.3% in the test environment.
Yikes, that's quite a compelling performance improvement.
> diff --git a/arch/arm64/include/asm/archrandom.h b/arch/arm64/include/asm/archrandom.h
> index 8babfbe31f95..8067e9a35641 100644
> --- a/arch/arm64/include/asm/archrandom.h
> +++ b/arch/arm64/include/asm/archrandom.h
> @@ -61,8 +61,22 @@ static inline bool __arm64_rndrrs(unsigned long *v)
>
> static __always_inline bool __cpu_has_rng(void)
> {
> - if (unlikely(!system_capabilities_finalized() && !preemptible()))
> - return this_cpu_has_cap(ARM64_HAS_RNG);
> + if (unlikely(!system_capabilities_finalized() && !preemptible())) {
> + /*
> + * Until the ARM64_HAS_RNG alternative is patched we can't use
> + * the static-branch form, so consult the feature register
> + * directly. Don't use this_cpu_has_cap() here: it reads
> + * ID_AA64ISAR0_EL1 from hardware on every call, under
> + * virtualization each ID register read traps to the hypervisor
> + * (HCR_EL2.TID3) -- producing a storm of vmexits during boot.
> + * The sanitised value is cached in memory.
> + */
> + u64 isar0 = read_sanitised_ftr_reg(SYS_ID_AA64ISAR0_EL1);
> +
> + return cpuid_feature_extract_unsigned_field(isar0,
> + ID_AA64ISAR0_EL1_RNDR_SHIFT) >=
> + ID_AA64ISAR0_EL1_RNDR_IMP;
> + }
You can probably rewrite this a little more cleanly along the lines of
the (not even compile-tested) diff below. I was about to do that, but
then I got a bit confused by the whole thing. The preemptible() check is
presumably not needed if we're accessing the in-memory feature registers
rather than the per-CPU id registers, but then how do you handle races
with concurrent updates to the "safe value" made by CPUs concurrently
coming online?
Will
--->8
diff --git a/arch/arm64/include/asm/archrandom.h b/arch/arm64/include/asm/archrandom.h
index 8babfbe31f95..1c6ccc1776cd 100644
--- a/arch/arm64/include/asm/archrandom.h
+++ b/arch/arm64/include/asm/archrandom.h
@@ -61,8 +61,17 @@ static inline bool __arm64_rndrrs(unsigned long *v)
static __always_inline bool __cpu_has_rng(void)
{
- if (unlikely(!system_capabilities_finalized() && !preemptible()))
- return this_cpu_has_cap(ARM64_HAS_RNG);
+ if (unlikely(!system_capabilities_finalized() && !preemptible())) {
+ /*
+ * Query the in-memory sanitised value to avoid a potential
+ * trap when accessing the ID register under a hypervisor.
+ */
+ u64 isar0 = read_sanitised_ftr_reg(SYS_ID_AA64ISAR0_EL1);
+ u64 rndr = SYS_FIELD_GET(ID_AA64ISAR0_EL1, RNDR, isar0);
+
+ return rndr >= ID_AA64ISAR0_EL1_RNDR_IMP;
+ }
+
return alternative_has_cap_unlikely(ARM64_HAS_RNG);
}
^ permalink raw reply related [flat|nested] 3+ messages in thread
* Re: [PATCH] arm64: archrandom: avoid trapping ID register read in __cpu_has_rng()
2026-07-31 14:43 ` Will Deacon
@ 2026-07-31 17:26 ` Aman Priyadarshi
0 siblings, 0 replies; 3+ messages in thread
From: Aman Priyadarshi @ 2026-07-31 17:26 UTC (permalink / raw)
To: Will Deacon
Cc: catalin.marinas, Jason, linux-arm-kernel, linux-kernel,
Ard Biesheuvel
Hi Will,
Thank you for reviewing the patch!
> On 31 Jul 2026, at 15:43, Will Deacon <will@kernel.org> wrote:
>
> Hi Aman,
>
> On Mon, Jul 20, 2026 at 03:06:15PM +0100, Aman Priyadarshi wrote:
>> __cpu_has_rng() has an early-boot fallback, taken before the
>> ARM64_HAS_RNG alternative is patched, that calls
>> this_cpu_has_cap(ARM64_HAS_RNG). With SCOPE_LOCAL_CPU that resolves the
>> capability by reading ID_AA64ISAR0_EL1 directly from hardware via
>> __read_sysreg_by_encoding(), on every invocation.
>>
>> Until the CRNG is seeded, crng_make_state() routes every get_random_*()
>> through extract_entropy(), which drains architectural entropy via
>> arch_get_random_seed_longs()/arch_get_random_longs() and so calls
>> __cpu_has_rng() several times per request. On a direct (non-EFI) boot
>> there is no bootloader seed, so the CRNG stays unseeded for much of
>> boot and essentially every early randomness consumer takes this path.
>>
>> Under virtualization this is costly: the hypervisor traps guest
>> accesses to the ID registers (HCR_EL2.TID3), making each read a vmexit,
>> producing ~200k trapped ID_AA64ISAR0_EL1 reads during boot.
>>
>> The register value is invariant, so read the sanitised feature register
>> instead. read_sanitised_ftr_reg() returns the cached value from
>> arm64_ftr_regs[] with no sysreg access, and hence no trap. That array
>> is populated by cpuinfo_store_boot_cpu() in smp_prepare_boot_cpu(),
>> before the first early RNG use in random_init_early(), so it is always
>> valid here.
>>
>> With this change the trapped reads drop from ~200k to handful number of
>> times, and the boot time drops roughly by 6.3% in the test environment.
>
> Yikes, that's quite a compelling performance improvement.
>
>> diff --git a/arch/arm64/include/asm/archrandom.h b/arch/arm64/include/asm/archrandom.h
>> index 8babfbe31f95..8067e9a35641 100644
>> --- a/arch/arm64/include/asm/archrandom.h
>> +++ b/arch/arm64/include/asm/archrandom.h
>> @@ -61,8 +61,22 @@ static inline bool __arm64_rndrrs(unsigned long *v)
>>
>> static __always_inline bool __cpu_has_rng(void)
>> {
>> - if (unlikely(!system_capabilities_finalized() && !preemptible()))
>> - return this_cpu_has_cap(ARM64_HAS_RNG);
>> + if (unlikely(!system_capabilities_finalized() && !preemptible())) {
>> + /*
>> + * Until the ARM64_HAS_RNG alternative is patched we can't use
>> + * the static-branch form, so consult the feature register
>> + * directly. Don't use this_cpu_has_cap() here: it reads
>> + * ID_AA64ISAR0_EL1 from hardware on every call, under
>> + * virtualization each ID register read traps to the hypervisor
>> + * (HCR_EL2.TID3) -- producing a storm of vmexits during boot.
>> + * The sanitised value is cached in memory.
>> + */
>> + u64 isar0 = read_sanitised_ftr_reg(SYS_ID_AA64ISAR0_EL1);
>> +
>> + return cpuid_feature_extract_unsigned_field(isar0,
>> + ID_AA64ISAR0_EL1_RNDR_SHIFT) >=
>> + ID_AA64ISAR0_EL1_RNDR_IMP;
>> + }
>
> You can probably rewrite this a little more cleanly along the lines of
> the (not even compile-tested) diff below. I was about to do that, but
> then I got a bit confused by the whole thing. The preemptible() check is
> presumably not needed if we're accessing the in-memory feature registers
> rather than the per-CPU id registers, but then how do you handle races
> with concurrent updates to the "safe value" made by CPUs concurrently
> coming online?
I kept preemptible() check for this exact reason: the updates made by CPUs
concurrently coming online will always take the downgrade path (a secondary
CPU can clear RNDR, never set it), and therefore by taking the non-preemptible
branch I can guarantee that the pinned CPU supports the said feature.
I agree, this assumes that a secondary CPU folds its own ID registers into sys_val
before it can ever be a randomness consumer, but looking at the code that seems
the case, please feel free to correct me.
Besides, in my opinion, it's hard to argue correctness of this code without
preemptible() check.
>
> Will
>
> --->8
>
> diff --git a/arch/arm64/include/asm/archrandom.h b/arch/arm64/include/asm/archrandom.h
> index 8babfbe31f95..1c6ccc1776cd 100644
> --- a/arch/arm64/include/asm/archrandom.h
> +++ b/arch/arm64/include/asm/archrandom.h
> @@ -61,8 +61,17 @@ static inline bool __arm64_rndrrs(unsigned long *v)
>
> static __always_inline bool __cpu_has_rng(void)
> {
> - if (unlikely(!system_capabilities_finalized() && !preemptible()))
> - return this_cpu_has_cap(ARM64_HAS_RNG);
> + if (unlikely(!system_capabilities_finalized() && !preemptible())) {
> + /*
> + * Query the in-memory sanitised value to avoid a potential
> + * trap when accessing the ID register under a hypervisor.
> + */
> + u64 isar0 = read_sanitised_ftr_reg(SYS_ID_AA64ISAR0_EL1);
> + u64 rndr = SYS_FIELD_GET(ID_AA64ISAR0_EL1, RNDR, isar0);
> +
> + return rndr >= ID_AA64ISAR0_EL1_RNDR_IMP;
> + }
> +
> return alternative_has_cap_unlikely(ARM64_HAS_RNG);
> }
Agreed, this looks cleaner. Thanks!
- Aman Priyadarshi
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-07-31 17:28 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-20 14:06 [PATCH] arm64: archrandom: avoid trapping ID register read in __cpu_has_rng() Aman Priyadarshi
2026-07-31 14:43 ` Will Deacon
2026-07-31 17:26 ` Aman Priyadarshi
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.