Linux-ARM-Kernel Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [RFC PATCH v1 0/2] mm: use VMA lock for kernel faults on user addresses
@ 2026-08-02  7:40 Barry Song (Xiaomi)
  2026-08-02  7:40 ` [RFC PATCH v1 1/2] x86/mm: " Barry Song (Xiaomi)
  2026-08-02  7:40 ` [RFC PATCH v1 2/2] arm64/mm: " Barry Song (Xiaomi)
  0 siblings, 2 replies; 4+ messages in thread
From: Barry Song (Xiaomi) @ 2026-08-02  7:40 UTC (permalink / raw)
  To: akpm, linux-mm
  Cc: x86, linux-arm-kernel, surenb, liam, ljs, vbabka, shakeel.butt,
	david, linux-kernel, zhanghongru, willy, zhangbo56,
	Barry Song (Xiaomi)

Right now, kernel faults on user addresses, such as those from
copy_from_user() and copy_to_user(), unconditionally fall back
to the mmap_lock path.

This patchset switches them to the per-VMA lock path for three
reasons:

1. These faults are common. On a typical Ubuntu system, hundreds
   of kernel faults on user addresses occur every second in
   python3, apt-esm-hook, package-data-do, teamviewerd, bash,
   scudo, cscope, gnome-shell, systemd, and many other
   applications. Handling these faults under mmap_lock
   unnecessarily increases lock contention.

2. It removes one obstacle to simplifying filemap_fault().
   Matthew has proposed removing the page fault retry path and
   performing I/O while holding locks [1]. Based on that
   approach, Hongru has already reported regressions caused by
   performing I/O under mmap_lock for kernel faults on user
   addresses [2]. This patchset removes that obstacle.

   Another source of I/O under mmap_lock is GUP, which also
   relies on mmap_lock today and has been reported by Hongru to
   exhibit similar regressions [2]. It appears Rik van Riel may
   be addressing this separately [3].

3. The current implementation is inconsistent. Although kernel
   faults on user addresses always fall back to mmap_lock,
   arch/*/mm/fault.c still performs !user_mode checks in the
   per-VMA lock path, making that code effectively dead.

As an RFC, this patchset demonstrates the approach on x86 and
arm64 only. Other architectures will need to be updated as well.

[1] https://lore.kernel.org/linux-mm/20260625195040.2508362-1-willy@infradead.org/
[2] https://lore.kernel.org/linux-mm/20260712132759.2030823-1-zhanghongru@xiaomi.com/
[3] https://lore.kernel.org/linux-mm/20260724222934.1463812-1-riel@surriel.com/

Barry Song (Xiaomi) (2):
  x86/mm: use VMA lock for kernel faults on user addresses
  arm64/mm: use VMA lock for kernel faults on user addresses

 arch/arm64/mm/fault.c | 4 +++-
 arch/x86/mm/fault.c   | 3 ---
 2 files changed, 3 insertions(+), 4 deletions(-)

-- 
2.39.3 (Apple Git-146)



^ permalink raw reply	[flat|nested] 4+ messages in thread

* [RFC PATCH v1 1/2] x86/mm: use VMA lock for kernel faults on user addresses
  2026-08-02  7:40 [RFC PATCH v1 0/2] mm: use VMA lock for kernel faults on user addresses Barry Song (Xiaomi)
@ 2026-08-02  7:40 ` Barry Song (Xiaomi)
  2026-08-02  7:40 ` [RFC PATCH v1 2/2] arm64/mm: " Barry Song (Xiaomi)
  1 sibling, 0 replies; 4+ messages in thread
From: Barry Song (Xiaomi) @ 2026-08-02  7:40 UTC (permalink / raw)
  To: akpm, linux-mm
  Cc: x86, linux-arm-kernel, surenb, liam, ljs, vbabka, shakeel.butt,
	david, linux-kernel, zhanghongru, willy, zhangbo56,
	Barry Song (Xiaomi)

Use the VMA lock for kernel faults on user addresses. This also
makes the existing code below meaningful:

  /* Quick path to respond to signals */
  if (fault_signal_pending(fault, regs)) {
          if (!user_mode(regs))
                  kernelmode_fixup_or_oops(regs, error_code, address,
                                           SIGBUS, BUS_ADRERR,
                                           ARCH_DEFAULT_PKEY);
          return;
  }

Right now, the code above is dead because !user_mode always
takes the mmap_lock path.

Co-developed-by: Bo Zhang <zhangbo56@xiaomi.com>
Signed-off-by: Bo Zhang <zhangbo56@xiaomi.com>
Signed-off-by: Barry Song (Xiaomi) <baohua@kernel.org>
---
 arch/x86/mm/fault.c | 3 ---
 1 file changed, 3 deletions(-)

diff --git a/arch/x86/mm/fault.c b/arch/x86/mm/fault.c
index 45b99c3b1442..c22b74e0eeaf 100644
--- a/arch/x86/mm/fault.c
+++ b/arch/x86/mm/fault.c
@@ -1328,9 +1328,6 @@ void do_user_addr_fault(struct pt_regs *regs,
 	}
 #endif
 
-	if (!(flags & FAULT_FLAG_USER))
-		goto lock_mmap;
-
 	vma = lock_vma_under_rcu(mm, address);
 	if (!vma)
 		goto lock_mmap;
-- 
2.39.3 (Apple Git-146)



^ permalink raw reply related	[flat|nested] 4+ messages in thread

* [RFC PATCH v1 2/2] arm64/mm: use VMA lock for kernel faults on user addresses
  2026-08-02  7:40 [RFC PATCH v1 0/2] mm: use VMA lock for kernel faults on user addresses Barry Song (Xiaomi)
  2026-08-02  7:40 ` [RFC PATCH v1 1/2] x86/mm: " Barry Song (Xiaomi)
@ 2026-08-02  7:40 ` Barry Song (Xiaomi)
  2026-08-02  8:49   ` Barry Song
  1 sibling, 1 reply; 4+ messages in thread
From: Barry Song (Xiaomi) @ 2026-08-02  7:40 UTC (permalink / raw)
  To: akpm, linux-mm
  Cc: x86, linux-arm-kernel, surenb, liam, ljs, vbabka, shakeel.butt,
	david, linux-kernel, zhanghongru, willy, zhangbo56,
	Barry Song (Xiaomi)

Use the VMA lock for kernel faults on user addresses. This also
makes the existing code below meaningful:

/* Quick path to respond to signals */
if (fault_signal_pending(fault, regs)) {
	if (!user_mode(regs))
		goto no_context;
	return 0;
}

Right now, the code above is dead because !user_mode always
takes the mmap_lock path.

Co-developed-by: Bo Zhang <zhangbo56@xiaomi.com>
Signed-off-by: Bo Zhang <zhangbo56@xiaomi.com>
Signed-off-by: Barry Song (Xiaomi) <baohua@kernel.org>
---
 arch/arm64/mm/fault.c | 4 +++-
 1 file changed, 3 insertions(+), 1 deletion(-)

diff --git a/arch/arm64/mm/fault.c b/arch/arm64/mm/fault.c
index 85e23388f9bb..241f1ab07ab3 100644
--- a/arch/arm64/mm/fault.c
+++ b/arch/arm64/mm/fault.c
@@ -607,6 +607,7 @@ static int __kprobes do_page_fault(unsigned long far, unsigned long esr,
 	unsigned int mm_flags = FAULT_FLAG_DEFAULT;
 	unsigned long addr = untagged_addr(far);
 	struct vm_area_struct *vma;
+	bool uaccess = false;
 	int si_code;
 	int pkey = -1;
 
@@ -663,6 +664,7 @@ static int __kprobes do_page_fault(unsigned long far, unsigned long esr,
 		if (!insn_may_access_user(regs->pc, esr))
 			die_kernel_fault("access to user memory outside uaccess routines",
 					 addr, esr, regs);
+		uaccess = true;
 	}
 
 	if (is_pkvm_stage2_abort(esr)) {
@@ -674,7 +676,7 @@ static int __kprobes do_page_fault(unsigned long far, unsigned long esr,
 
 	perf_sw_event(PERF_COUNT_SW_PAGE_FAULTS, 1, regs, addr);
 
-	if (!(mm_flags & FAULT_FLAG_USER))
+	if (!(mm_flags & FAULT_FLAG_USER) && !uaccess)
 		goto lock_mmap;
 
 	vma = lock_vma_under_rcu(mm, addr);
-- 
2.39.3 (Apple Git-146)



^ permalink raw reply related	[flat|nested] 4+ messages in thread

* Re: [RFC PATCH v1 2/2] arm64/mm: use VMA lock for kernel faults on user addresses
  2026-08-02  7:40 ` [RFC PATCH v1 2/2] arm64/mm: " Barry Song (Xiaomi)
@ 2026-08-02  8:49   ` Barry Song
  0 siblings, 0 replies; 4+ messages in thread
From: Barry Song @ 2026-08-02  8:49 UTC (permalink / raw)
  To: akpm, linux-mm
  Cc: x86, linux-arm-kernel, surenb, liam, ljs, vbabka, shakeel.butt,
	david, linux-kernel, zhanghongru, willy, zhangbo56

On Sun, Aug 2, 2026 at 3:40 PM Barry Song (Xiaomi) <baohua@kernel.org> wrote:
>
> Use the VMA lock for kernel faults on user addresses. This also
> makes the existing code below meaningful:
>
> /* Quick path to respond to signals */
> if (fault_signal_pending(fault, regs)) {
>         if (!user_mode(regs))
>                 goto no_context;
>         return 0;
> }
>
> Right now, the code above is dead because !user_mode always
> takes the mmap_lock path.
>
> Co-developed-by: Bo Zhang <zhangbo56@xiaomi.com>
> Signed-off-by: Bo Zhang <zhangbo56@xiaomi.com>
> Signed-off-by: Barry Song (Xiaomi) <baohua@kernel.org>
> ---
>  arch/arm64/mm/fault.c | 4 +++-
>  1 file changed, 3 insertions(+), 1 deletion(-)
>
> diff --git a/arch/arm64/mm/fault.c b/arch/arm64/mm/fault.c
> index 85e23388f9bb..241f1ab07ab3 100644
> --- a/arch/arm64/mm/fault.c
> +++ b/arch/arm64/mm/fault.c
> @@ -607,6 +607,7 @@ static int __kprobes do_page_fault(unsigned long far, unsigned long esr,
>         unsigned int mm_flags = FAULT_FLAG_DEFAULT;
>         unsigned long addr = untagged_addr(far);
>         struct vm_area_struct *vma;
> +       bool uaccess = false;
>         int si_code;
>         int pkey = -1;
>
> @@ -663,6 +664,7 @@ static int __kprobes do_page_fault(unsigned long far, unsigned long esr,
>                 if (!insn_may_access_user(regs->pc, esr))
>                         die_kernel_fault("access to user memory outside uaccess routines",
>                                          addr, esr, regs);
> +               uaccess = true;

https://sashiko.dev/#/patchset/20260802074018.73887-1-baohua%40kernel.org

"Does this conditional bypass translation faults? Because this block is
guarded by:
if (is_ttbr0_addr(addr) && is_el1_permission_fault(addr, esr, regs)) {
it appears uaccess is only set for permission faults like CoW or PAN
violations.
If a kernel uaccess routine accesses an unmapped user address, a translation
fault occurs, making is_el1_permission_fault() evaluate to false.
Would this cause demand paging in uaccess routines to fall back to the slow
lock_mmap path, missing the intended optimization?"

Good catch! I should have only modified a single line. Then
unmapped PTEs would also benefit from the VMA lock:

diff --git a/arch/arm64/mm/fault.c b/arch/arm64/mm/fault.c
index 85e23388f9bb..e1b406d667aa 100644
--- a/arch/arm64/mm/fault.c
+++ b/arch/arm64/mm/fault.c
@@ -674,7 +674,7 @@ static int __kprobes do_page_fault(unsigned long
far, unsigned long esr,

        perf_sw_event(PERF_COUNT_SW_PAGE_FAULTS, 1, regs, addr);

-       if (!(mm_flags & FAULT_FLAG_USER))
+       if (!(mm_flags & FAULT_FLAG_USER) && !is_ttbr0_addr(addr))
                goto lock_mmap;

        vma = lock_vma_under_rcu(mm, addr);


^ permalink raw reply related	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2026-08-02  8:49 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-02  7:40 [RFC PATCH v1 0/2] mm: use VMA lock for kernel faults on user addresses Barry Song (Xiaomi)
2026-08-02  7:40 ` [RFC PATCH v1 1/2] x86/mm: " Barry Song (Xiaomi)
2026-08-02  7:40 ` [RFC PATCH v1 2/2] arm64/mm: " Barry Song (Xiaomi)
2026-08-02  8:49   ` Barry Song

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox