* [PATCH v2 5.10.y/5.15.y 1/3] ARM: fix hash_name() fault
2026-09-17 2:30 [PATCH v2 5.10.y/5.15.y 0/3] ARM: fix might_sleep() WARNING for mmap_write_lock() around show_pte() Xie Yuanbin
@ 2026-09-17 2:30 ` Xie Yuanbin
2026-09-17 2:30 ` [PATCH v2 5.10.y/5.15.y 2/3] ARM: fix branch predictor hardening Xie Yuanbin
` (2 subsequent siblings)
3 siblings, 0 replies; 5+ messages in thread
From: Xie Yuanbin @ 2026-09-17 2:30 UTC (permalink / raw)
To: sashal, gregkh, linux, bigeasy, clrkwllms, rostedt, rmk+kernel,
linusw, kuninori.morimoto.gx, arnd, xiqi2, ljs, wozizhi
Cc: linux-arm-kernel, linux-kernel, linux-rt-devel, stable, patches,
lisongze2, wangbing6, Xie Yuanbin
From: "Russell King (Oracle)" <rmk+kernel@armlinux.org.uk>
[ Upstream commit 7733bc7d299d682f2723dc38fc7f370b9bf973e9 ]
Zizhi Wo reports:
"During the execution of hash_name()->load_unaligned_zeropad(), a
potential memory access beyond the PAGE boundary may occur. For
example, when the filename length is near the PAGE_SIZE boundary.
This triggers a page fault, which leads to a call to
do_page_fault()->mmap_read_trylock(). If we can't acquire the lock,
we have to fall back to the mmap_read_lock() path, which calls
might_sleep(). This breaks RCU semantics because path lookup occurs
under an RCU read-side critical section."
This is seen with CONFIG_DEBUG_ATOMIC_SLEEP=y and CONFIG_KFENCE=y.
Kernel addresses (with the exception of the vectors/kuser helper
page) do not have VMAs associated with them. If the vectors/kuser
helper page faults, then there are two possibilities:
1. if the fault happened while in kernel mode, then we're basically
dead, because the CPU won't be able to vector through this page
to handle the fault.
2. if the fault happened while in user mode, that means the page was
protected from user access, and we want to fault anyway.
Thus, we can handle kernel addresses from any context entirely
separately without going anywhere near the mmap lock. This gives us
an entirely non-sleeping path for all kernel mode kernel address
faults.
As we handle the kernel address faults before interrupts are enabled,
this change has the side effect of improving the branch predictor
hardening, but does not completely solve the issue.
[ Xie Yuanbin: At the upstream, the following patches are a patch set:
1. commit dea20281ac8822661576 ("ARM: group is_permission_fault() with
is_translation_fault()")
2. commit 40b466db1dffb41f0529 ("ARM: allow __do_kernel_fault() to
report execution of memory faults")
3. commit 7733bc7d299d682f2723 ("ARM: fix hash_name() fault")
4. commit fd2dee1c6e2256f726ba ("ARM: fix branch predictor hardening")
patch 1. and 2. is unneeded for 5.10.y and 5.15.y . This patch backports
patch 3. and simply adapts to the context differences. ]
Reported-by: Zizhi Wo <wozizhi@huaweicloud.com>
Reported-by: Xie Yuanbin <xieyuanbin1@huawei.com>
Link: https://lore.kernel.org/r/20251126090505.3057219-1-wozizhi@huaweicloud.com
Reviewed-by: Xie Yuanbin <xieyuanbin1@huawei.com>
Tested-by: Xie Yuanbin <xieyuanbin1@huawei.com>
Signed-off-by: Russell King (Oracle) <rmk+kernel@armlinux.org.uk>
Signed-off-by: Xie Yuanbin <xieyuanbin1@huawei.com>
---
arch/arm/mm/fault.c | 35 +++++++++++++++++++++++++++++++++++
1 file changed, 35 insertions(+)
diff --git a/arch/arm/mm/fault.c b/arch/arm/mm/fault.c
index c16d6a293b97..094137cd29c8 100644
--- a/arch/arm/mm/fault.c
+++ b/arch/arm/mm/fault.c
@@ -247,6 +247,35 @@ __do_page_fault(struct mm_struct *mm, unsigned long addr, unsigned int fsr,
return fault;
}
+static int __kprobes
+do_kernel_address_page_fault(struct mm_struct *mm, unsigned long addr,
+ unsigned int fsr, struct pt_regs *regs)
+{
+ if (user_mode(regs)) {
+ /*
+ * Fault from user mode for a kernel space address. User mode
+ * should not be faulting in kernel space, which includes the
+ * vector/khelper page. Send a SIGSEGV.
+ */
+ __do_user_fault(addr, fsr, SIGSEGV, SEGV_MAPERR, regs);
+ } else {
+ /*
+ * Fault from kernel mode. Enable interrupts if they were
+ * enabled in the parent context. Section (upper page table)
+ * translation faults are handled via do_translation_fault(),
+ * so we will only get here for a non-present kernel space
+ * PTE or PTE permission fault. This may happen in exceptional
+ * circumstances and need the fixup tables to be walked.
+ */
+ if (interrupts_enabled(regs))
+ local_irq_enable();
+
+ __do_kernel_fault(mm, addr, fsr, regs);
+ }
+
+ return 0;
+}
+
static int __kprobes
do_page_fault(unsigned long addr, unsigned int fsr, struct pt_regs *regs)
{
@@ -261,6 +290,12 @@ do_page_fault(unsigned long addr, unsigned int fsr, struct pt_regs *regs)
tsk = current;
mm = tsk->mm;
+ /*
+ * Handle kernel addresses faults separately, which avoids touching
+ * the mmap lock from contexts that are not able to sleep.
+ */
+ if (addr >= TASK_SIZE)
+ return do_kernel_address_page_fault(mm, addr, fsr, regs);
/* Enable interrupts if they were enabled in the parent context. */
if (interrupts_enabled(regs))
--
2.55.0
^ permalink raw reply related [flat|nested] 5+ messages in thread* [PATCH v2 5.10.y/5.15.y 2/3] ARM: fix branch predictor hardening
2026-09-17 2:30 [PATCH v2 5.10.y/5.15.y 0/3] ARM: fix might_sleep() WARNING for mmap_write_lock() around show_pte() Xie Yuanbin
2026-09-17 2:30 ` [PATCH v2 5.10.y/5.15.y 1/3] ARM: fix hash_name() fault Xie Yuanbin
@ 2026-09-17 2:30 ` Xie Yuanbin
2026-09-17 2:30 ` [PATCH v2 5.10.y/5.15.y 3/3] ARM: ensure interrupts are enabled in __do_user_fault() Xie Yuanbin
2026-09-18 0:52 ` [PATCH v2 5.10.y/5.15.y 0/3] ARM: fix might_sleep() WARNING for mmap_write_lock() around show_pte() Sasha Levin
3 siblings, 0 replies; 5+ messages in thread
From: Xie Yuanbin @ 2026-09-17 2:30 UTC (permalink / raw)
To: sashal, gregkh, linux, bigeasy, clrkwllms, rostedt, rmk+kernel,
linusw, kuninori.morimoto.gx, arnd, xiqi2, ljs, wozizhi
Cc: linux-arm-kernel, linux-kernel, linux-rt-devel, stable, patches,
lisongze2, wangbing6, Xie Yuanbin
From: "Russell King (Oracle)" <rmk+kernel@armlinux.org.uk>
[ Upstream commit fd2dee1c6e2256f726ba33fd3083a7be0efc80d3 ]
__do_user_fault() may be called with indeterminent interrupt enable
state, which means we may be preemptive at this point. This causes
problems when calling harden_branch_predictor(). For example, when
called from a data abort, do_alignment_fault()->do_bad_area().
Move harden_branch_predictor() out of __do_user_fault() and into the
calling contexts.
Moving it into do_kernel_address_page_fault(), we can be sure that
interrupts will be disabled here.
Converting do_translation_fault() to use do_kernel_address_page_fault()
rather than do_bad_area() means that we keep branch predictor handling
for translation faults. Interrupts will also be disabled at this call
site.
do_sect_fault() needs special handling, so detect user mode accesses
to kernel-addresses, and add an explicit call to branch predictor
hardening.
Finally, add branch predictor hardening to do_alignment() for the
faulting case (user mode accessing kernel addresses) before interrupts
are enabled.
This should cover all cases where harden_branch_predictor() is called,
ensuring that it is always has interrupts disabled, also ensuring that
it is called early in each call path.
[ Xie Yuanbin: At the upstream, the following patches are a patch set:
1. commit dea20281ac8822661576 ("ARM: group is_permission_fault() with
is_translation_fault()")
2. commit 40b466db1dffb41f0529 ("ARM: allow __do_kernel_fault() to
report execution of memory faults")
3. commit 7733bc7d299d682f2723 ("ARM: fix hash_name() fault")
4. commit fd2dee1c6e2256f726ba ("ARM: fix branch predictor hardening")
patch 1. and 2. is unneeded for 5.10.y and 5.15.y . This patch backports
patch 4. and simply adapts to the context differences. ]
Reviewed-by: Xie Yuanbin <xieyuanbin1@huawei.com>
Tested-by: Xie Yuanbin <xieyuanbin1@huawei.com>
Signed-off-by: Russell King (Oracle) <rmk+kernel@armlinux.org.uk>
Signed-off-by: Xie Yuanbin <xieyuanbin1@huawei.com>
---
arch/arm/mm/alignment.c | 4 ++++
arch/arm/mm/fault.c | 39 ++++++++++++++++++++++++++-------------
2 files changed, 30 insertions(+), 13 deletions(-)
diff --git a/arch/arm/mm/alignment.c b/arch/arm/mm/alignment.c
index bcefe3f51744..758504a28c13 100644
--- a/arch/arm/mm/alignment.c
+++ b/arch/arm/mm/alignment.c
@@ -23,6 +23,7 @@
#include <asm/cp15.h>
#include <asm/system_info.h>
#include <asm/unaligned.h>
+#include <asm/system_misc.h>
#include <asm/opcodes.h>
#include "fault.h"
@@ -809,6 +810,9 @@ do_alignment(unsigned long addr, unsigned int fsr, struct pt_regs *regs)
int thumb2_32b = 0;
int fault;
+ if (addr >= TASK_SIZE && user_mode(regs))
+ harden_branch_predictor();
+
if (interrupts_enabled(regs))
local_irq_enable();
diff --git a/arch/arm/mm/fault.c b/arch/arm/mm/fault.c
index 094137cd29c8..4afb076383a8 100644
--- a/arch/arm/mm/fault.c
+++ b/arch/arm/mm/fault.c
@@ -145,9 +145,6 @@ __do_user_fault(unsigned long addr, unsigned int fsr, unsigned int sig,
{
struct task_struct *tsk = current;
- if (addr > TASK_SIZE)
- harden_branch_predictor();
-
#ifdef CONFIG_DEBUG_USER
if (((user_debug & UDBG_SEGV) && (sig == SIGSEGV)) ||
((user_debug & UDBG_BUS) && (sig == SIGBUS))) {
@@ -255,8 +252,10 @@ do_kernel_address_page_fault(struct mm_struct *mm, unsigned long addr,
/*
* Fault from user mode for a kernel space address. User mode
* should not be faulting in kernel space, which includes the
- * vector/khelper page. Send a SIGSEGV.
+ * vector/khelper page. Handle the branch predictor hardening
+ * while interrupts are still disabled, then send a SIGSEGV.
*/
+ harden_branch_predictor();
__do_user_fault(addr, fsr, SIGSEGV, SEGV_MAPERR, regs);
} else {
/*
@@ -421,16 +420,20 @@ do_page_fault(unsigned long addr, unsigned int fsr, struct pt_regs *regs)
* We enter here because the first level page table doesn't contain
* a valid entry for the address.
*
- * If the address is in kernel space (>= TASK_SIZE), then we are
- * probably faulting in the vmalloc() area.
+ * If this is a user address (addr < TASK_SIZE), we handle this as a
+ * normal page fault. This leaves the remainder of the function to handle
+ * kernel address translation faults.
*
- * If the init_task's first level page tables contains the relevant
- * entry, we copy the it to this task. If not, we send the process
- * a signal, fixup the exception, or oops the kernel.
+ * Since user mode is not permitted to access kernel addresses, pass these
+ * directly to do_kernel_address_page_fault() to handle.
*
- * NOTE! We MUST NOT take any locks for this case. We may be in an
- * interrupt or a critical region, and should only copy the information
- * from the master page table, nothing more.
+ * Otherwise, we're probably faulting in the vmalloc() area, so try to fix
+ * that up. Note that we must not take any locks or enable interrupts in
+ * this case.
+ *
+ * If vmalloc() fixup fails, that means the non-leaf page tables did not
+ * contain an entry for this address, so handle this via
+ * do_kernel_address_page_fault().
*/
#ifdef CONFIG_MMU
static int __kprobes
@@ -496,7 +499,8 @@ do_translation_fault(unsigned long addr, unsigned int fsr,
return 0;
bad_area:
- do_bad_area(addr, fsr, regs);
+ do_kernel_address_page_fault(current->mm, addr, fsr, regs);
+
return 0;
}
#else /* CONFIG_MMU */
@@ -516,7 +520,16 @@ do_translation_fault(unsigned long addr, unsigned int fsr,
static int
do_sect_fault(unsigned long addr, unsigned int fsr, struct pt_regs *regs)
{
+ /*
+ * If this is a kernel address, but from user mode, then userspace
+ * is trying bad stuff. Invoke the branch predictor handling.
+ * Interrupts are disabled here.
+ */
+ if (addr >= TASK_SIZE && user_mode(regs))
+ harden_branch_predictor();
+
do_bad_area(addr, fsr, regs);
+
return 0;
}
#endif /* CONFIG_ARM_LPAE */
--
2.55.0
^ permalink raw reply related [flat|nested] 5+ messages in thread* [PATCH v2 5.10.y/5.15.y 3/3] ARM: ensure interrupts are enabled in __do_user_fault()
2026-09-17 2:30 [PATCH v2 5.10.y/5.15.y 0/3] ARM: fix might_sleep() WARNING for mmap_write_lock() around show_pte() Xie Yuanbin
2026-09-17 2:30 ` [PATCH v2 5.10.y/5.15.y 1/3] ARM: fix hash_name() fault Xie Yuanbin
2026-09-17 2:30 ` [PATCH v2 5.10.y/5.15.y 2/3] ARM: fix branch predictor hardening Xie Yuanbin
@ 2026-09-17 2:30 ` Xie Yuanbin
2026-09-18 0:52 ` [PATCH v2 5.10.y/5.15.y 0/3] ARM: fix might_sleep() WARNING for mmap_write_lock() around show_pte() Sasha Levin
3 siblings, 0 replies; 5+ messages in thread
From: Xie Yuanbin @ 2026-09-17 2:30 UTC (permalink / raw)
To: sashal, gregkh, linux, bigeasy, clrkwllms, rostedt, rmk+kernel,
linusw, kuninori.morimoto.gx, arnd, xiqi2, ljs, wozizhi
Cc: linux-arm-kernel, linux-kernel, linux-rt-devel, stable, patches,
lisongze2, wangbing6, Yadi.hu, Xie Yuanbin
From: "Russell King (Oracle)" <rmk+kernel@armlinux.org.uk>
[ Upstream commit 59e4f3b45b96a24fc9b7a89e5f8a2168b30f95af ]
__do_user_fault() may be called from fault handling paths where the
interrupts are enabled or disabled. E.g. do_page_fault() calls this
with interrupts enabled, whereas do_sect_fault()->do_bad_area()
will call this with interrupts disabled. Since this is a userspace
fault, we know that interrupts were enabled in the parent context,
so call local_irq_enable() here to give a consistent interrupt state.
This is necessary for force_sig_info() when PREEMPT_RT is enabled.
Reported-by: Yadi.hu <yadi.hu@windriver.com>
Reviewed-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
Signed-off-by: Russell King (Oracle) <rmk+kernel@armlinux.org.uk>
Signed-off-by: Xie Yuanbin <xieyuanbin1@huawei.com>
---
arch/arm/mm/fault.c | 6 +++++-
1 file changed, 5 insertions(+), 1 deletion(-)
diff --git a/arch/arm/mm/fault.c b/arch/arm/mm/fault.c
index 4afb076383a8..8b31d7704626 100644
--- a/arch/arm/mm/fault.c
+++ b/arch/arm/mm/fault.c
@@ -137,7 +137,8 @@ __do_kernel_fault(struct mm_struct *mm, unsigned long addr, unsigned int fsr,
/*
* Something tried to access memory that isn't in our memory map..
- * User mode accesses just cause a SIGSEGV
+ * User mode accesses just cause a SIGSEGV. Ensure interrupts are enabled
+ * for preempt RT.
*/
static void
__do_user_fault(unsigned long addr, unsigned int fsr, unsigned int sig,
@@ -145,6 +146,8 @@ __do_user_fault(unsigned long addr, unsigned int fsr, unsigned int sig,
{
struct task_struct *tsk = current;
+ local_irq_enable();
+
#ifdef CONFIG_DEBUG_USER
if (((user_debug & UDBG_SEGV) && (sig == SIGSEGV)) ||
((user_debug & UDBG_BUS) && (sig == SIGBUS))) {
@@ -254,6 +257,7 @@ do_kernel_address_page_fault(struct mm_struct *mm, unsigned long addr,
* should not be faulting in kernel space, which includes the
* vector/khelper page. Handle the branch predictor hardening
* while interrupts are still disabled, then send a SIGSEGV.
+ * Note that __do_user_fault() will enable interrupts.
*/
harden_branch_predictor();
__do_user_fault(addr, fsr, SIGSEGV, SEGV_MAPERR, regs);
--
2.55.0
^ permalink raw reply related [flat|nested] 5+ messages in thread* Re: [PATCH v2 5.10.y/5.15.y 0/3] ARM: fix might_sleep() WARNING for mmap_write_lock() around show_pte()
2026-09-17 2:30 [PATCH v2 5.10.y/5.15.y 0/3] ARM: fix might_sleep() WARNING for mmap_write_lock() around show_pte() Xie Yuanbin
` (2 preceding siblings ...)
2026-09-17 2:30 ` [PATCH v2 5.10.y/5.15.y 3/3] ARM: ensure interrupts are enabled in __do_user_fault() Xie Yuanbin
@ 2026-09-18 0:52 ` Sasha Levin
3 siblings, 0 replies; 5+ messages in thread
From: Sasha Levin @ 2026-09-18 0:52 UTC (permalink / raw)
To: gregkh, linux, bigeasy, clrkwllms, rostedt, rmk+kernel, linusw,
kuninori.morimoto.gx, arnd, xiqi2, ljs, wozizhi
Cc: Sasha Levin, linux-arm-kernel, linux-kernel, linux-rt-devel,
stable, patches, lisongze2, wangbing6, Xie Yuanbin
> This series backports all these upstream patches into the linux-5.10.y:
> 1. commit 7733bc7d299d682f2723 ("ARM: fix hash_name() fault")
> 2. commit fd2dee1c6e2256f726ba ("ARM: fix branch predictor hardening")
> 3. commit 59e4f3b45b96a24fc9b7 ("ARM: ensure interrupts are enabled in
> __do_user_fault()")
Queued the series for 5.15 and 5.10, thanks.
--
Thanks,
Sasha
^ permalink raw reply [flat|nested] 5+ messages in thread