Linux-ARM-Kernel Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH RFC] arm64: entry: PSTATE_I_SET is leaking on pseudo NMI mode
@ 2026-08-07 11:45 Breno Leitao
  2026-08-07 13:40 ` Will Deacon
  2026-08-07 13:57 ` Mark Rutland
  0 siblings, 2 replies; 6+ messages in thread
From: Breno Leitao @ 2026-08-07 11:45 UTC (permalink / raw)
  To: Catalin Marinas, Will Deacon, Peter Zijlstra (Intel),
	Mark Rutland, Jinjie Ruan
  Cc: linux-arm-kernel, linux-kernel, bpf, rmikey, kernel-team,
	Breno Leitao

Running retsnoop on a kernel with GIC priority masking and
CONFIG_ARM64_DEBUG_PRIORITY_MASKING=y trips the ICC_PMR_EL1 sanity check
in __pmr_local_irq_disable():

     WARNING: ./arch/arm64/include/asm/irqflags.h:63 at arm64_exit_to_kernel_mode+0xb8/0xc0, CPU#40: retsnoop/31805
     CPU: 40 UID: 0 PID: 31805 Comm: retsnoop Not tainted 7.2.0-rc6-next-20260805 #7 PREEMPTLAZY
     pstate: 234013c9 (nzCv DAIF +PAN -UAO +TCO +DIT +SSBS BTYPE=--)
     pc : arm64_exit_to_kernel_mode (arch/arm64/kernel/entry-common.c:63)
     lr : el1_abort (arch/arm64/kernel/entry-common.c:323)
     pmr: 000000f0
     Call trace:
D)    arm64_exit_to_kernel_mode (arch/arm64/kernel/entry-common.c:63) (P)
      el1_abort (arch/arm64/kernel/entry-common.c:323)
      el1h_64_sync_handler (arch/arm64/kernel/entry-common.c:449)
C)    el1h_64_sync (arch/arm64/kernel/entry.S:589)
      copy_from_kernel_nofault (mm/maccess.c:52) (P)
      bpf_probe_read_kernel (kernel/trace/bpf_trace.c:268)
      bpf_prog_db21a1730c2407e5_calib_exit+0xf0/0x160
      trace_call_bpf (kernel/trace/bpf_trace.c:147)
      kretprobe_perf_func (kernel/trace/trace_kprobe.c:1750)
      kretprobe_dispatcher (kernel/trace/trace_kprobe.c:1875)
      __kretprobe_trampoline_handler (kernel/kprobes.c:2116)
      kretprobe_brk_handler (arch/arm64/kernel/probes/kprobes.c:422)
      call_el1_break_hook (arch/arm64/kernel/debug-monitors.c:244)
      do_el1_brk64 (arch/arm64/kernel/debug-monitors.c:266)
B)    el1_brk64 (arch/arm64/kernel/entry-common.c:427)
      el1h_64_sync_handler (arch/arm64/kernel/entry-common.c:481)
      el1h_64_sync (arch/arm64/kernel/entry.S:589)
      invoke_syscall (arch/arm64/kernel/syscall.c:49) (P)
      do_el0_svc (arch/arm64/kernel/syscall.c:140)
      el0_svc (arch/arm64/kernel/entry-common.c:736)
      el0t_64_sync_handler (arch/arm64/kernel/entry-common.c:755)
A)    el0t_64_sync (arch/arm64/kernel/entry.S:594)

This is my understand of the current situation:

A) A task enters the kernel via a syscall.
	* PSTATE_I_SET becomes set on the live PMR
	* regs->PMR doesn't have PSR_I_SET set

B) A BRK fires at EL1 and that is what leaves the live PMR
   with PSR_I_SET.
	* At this stage PSTATE_I_SET is set on both on PMR and regs->PMR

C) A nested synchronous exception happens (Not sure why -- BPF related)
	* Now both the live PMR and regs->pmr have PSR_I_SET.

D) On the way out, local_irq_disable() → __pmr_local_irq_disable() warns.
   It detects that PMR is different than GIC_PRIO_IRQON and GIC_PRIO_IRQOFF,
   given live PMR and regs->PMR have PSTATE_I_SET ORed.

How to fix it? I don't know very well.

I am not certain that we want to have PSTATE_I_SET ever be sent to
regs->pstate. Do we ever need PSTATE_I_SET in regs->pstate?

In the current patch, I found that disabling IRQ in case it is disabled,
would solve the warning, but, this seems more a hack than a proper fix,
perhaps.

Fixes: ae654112eac0 ("arm64: entry: Use split preemption logic")
Signed-off-by: Breno Leitao <leitao@debian.org>
---
 arch/arm64/kernel/entry-common.c | 8 +++++++-
 1 file changed, 7 insertions(+), 1 deletion(-)

diff --git a/arch/arm64/kernel/entry-common.c b/arch/arm64/kernel/entry-common.c
index ceb4eb11232a6..fcce9ccd37108 100644
--- a/arch/arm64/kernel/entry-common.c
+++ b/arch/arm64/kernel/entry-common.c
@@ -55,7 +55,13 @@ static noinstr irqentry_state_t arm64_enter_from_kernel_mode(struct pt_regs *reg
 static void noinstr arm64_exit_to_kernel_mode(struct pt_regs *regs,
 					      irqentry_state_t state)
 {
-	local_irq_disable();
+	/*
+	 * Only irqentry_exit_to_kernel_mode_preempt() needs interrupts masked,
+	 * and it returns early when regs had them disabled. Skipping the
+	 * disable avoids clobbering a PMR the irqflags API does not expect.
+	 */
+	if (!regs_irqs_disabled(regs))
+		local_irq_disable();
 	irqentry_exit_to_kernel_mode_preempt(regs, state);
 	local_daif_mask();
 	mte_check_tfsr_exit();

---
base-commit: ea2bff00da89d7767d677bb68470130ba96f4928
change-id: 20260807-arm64_fix-47cad8fb6323

Best regards,
--  
Breno Leitao <leitao@debian.org>



^ permalink raw reply related	[flat|nested] 6+ messages in thread

end of thread, other threads:[~2026-08-07 16:29 UTC | newest]

Thread overview: 6+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-07 11:45 [PATCH RFC] arm64: entry: PSTATE_I_SET is leaking on pseudo NMI mode Breno Leitao
2026-08-07 13:40 ` Will Deacon
2026-08-07 13:57 ` Mark Rutland
2026-08-07 14:49   ` Vladimir Murzin
2026-08-07 14:58   ` Breno Leitao
2026-08-07 16:29     ` Breno Leitao

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox