* [patch 00/18] entry: Consolidate and rework syscall entry handling
@ 2026-07-07 19:05 Thomas Gleixner
2026-07-07 19:05 ` [patch 01/18] powerpc: Move stack randomization after syscall_enter_from_user_mode() Thomas Gleixner
` (21 more replies)
0 siblings, 22 replies; 107+ messages in thread
From: Thomas Gleixner @ 2026-07-07 19:05 UTC (permalink / raw)
To: LKML
Cc: Peter Zijlstra, Michael Ellerman, Shrikanth Hegde, linuxppc-dev,
Kees Cook, Huacai Chen, loongarch, Paul Walmsley, Palmer Dabbelt,
linux-riscv, Sven Schnelle, linux-s390, x86, Mark Rutland,
Jinjie Ruan, Andy Lutomirski, Oleg Nesterov, Richard Henderson,
Russell King, Catalin Marinas, Guo Ren, Geert Uytterhoeven,
Thomas Bogendoerfer, Helge Deller, Yoshinori Sato,
Richard Weinberger, Chris Zankel, linux-arm-kernel, linux-alpha,
linux-csky, linux-m68k, linux-mips, linux-parisc, linux-sh,
linux-um, Arnd Bergmann, Vineet Gupta, Will Deacon, Brian Cain,
Michal Simek, Dinh Nguyen, David S. Miller, Andreas Larsson,
linux-snps-arc, linux-hexagon, linux-openrisc, sparclinux,
linux-arch, Michal Suchánek, Jonathan Corbet, linux-doc
Sorry for the long CC list, but this is a treewide change.
Michal recently posted a RFC patch to separate the potential syscall number
modifications in syscall_enter_user_mode_work() from the information
whether the syscall should be processed and the return value modified:
https://lore.kernel.org/lkml/CE1qW@kunlun.suse.cz
The existing logic is:
arch_syscall()
regs->result = -ENOSYS;
syscallnr = syscall_enter_from_user_mode(regs, syscall);
if (syscallnr != -1L)
regs->result = invoke_syscall(regs, syscall;
syscall_enter_from_user_mode() invokes ptrace, seccomp and
tracing/BPF/Probes. All of them can modify the syscall number.
ptrace and seccomp explicitly set the syscall number to -1L to indicate
that the syscall invocation needs to be skipped and the result has not to
be modified as it might have been modified by ptrace or seccomp. The
tracer/BPF/Probes mechanism can modify the syscall number as well and
relies implicitly on the -1L logic.
This can obviously not be differentiated from a syscall invocation where
userspace provided -1 as syscall number.
The general agreement of the discussion was that the current mechanism,
while functionally correct is non-intuitive and something like Michals
proposal would make that code clearer and easier to handle on the
architecture side:
arch_syscall()
regs->result = -ENOSYS;
if (syscall_enter_from_user_mode(regs, &syscall))
regs->result = invoke_syscall(regs, syscall;
That discussion made me look deeper into the related code and as usual
there were a lot of other things to discover.
1) Stack randomization
add_random_kstack_offset() can only be invoked after
enter_from_user_mode() established proper state as it calls into
instrumentable code.
PowerPC got that wrong and the other architectures either invoke it
after enter_from_user_mode() or after syscall_enter_from_user_mode().
The latter is suboptimal as the randomization takes place after all
the user mode entry work. Aside of that add_random_kstack_offset()
uses get/put_cpu_var(), which makes it usable in preemptible code, but
when invoked in the interrupt disabled region that's pointless
overhead.
2) As discussed in the above thread just changing the function signature
of syscall_enter_from_user_mode[_work]() so they take a pointer
argument for the syscall and then return 0 on success is not really
intuitive either. Aside of that this breaks the implicit assumption of
the tracer when setting the syscall number to -1.
3) The x86 entry code has some historically accumulated oddities
The following series addresses this by:
1) Providing new [syscall_]enter_from_user_mode() variants, which include
stack randomization and utilize a new add_random_kstack_offset_irqsoff()
variant, which avoids the get/put_cpu_var() overhead and converting all
usage sites over
2) Picking up Jinjie's seccomp patch from:
https://lore.kernel.org/lkml/20260629130616.642022-2-ruanjinjie@huawei.com
and addressing the feedback (renaming the seccomp functions)
3) Making the ptrace and tracer related functions return a boolean value
to indicate syscall permission
4) Addressing the x86 oddities
5) Converting the tree over to the new scheme
With that all architectures using the generic syscall entry code follow the
same scheme, apply stack randomization at the correct and earliest possible
place and skip syscall processing depending on the boolean return value of
syscall_enter_from_user_mode[_work]().
There should be no functional changes, at least there are none intended.
The resulting text size for the syscall entry code on x8664 is slightly
smaller than before these changes.
Testing syscall heavy workloads and micro benchmarks shows a small
performance gain for the general rework, but the last patch, which changes
the logic to be more understandable has no measurable impact in either
direction.
The series applies on Linus tree and is also available from git:
git://git.kernel.org/pub/scm/linux/kernel/git/tglx/devel.git entry-rework-v1
Thanks,
tglx
---
Documentation/core-api/entry.rst | 33 +++++---
arch/alpha/kernel/ptrace.c | 4 -
arch/arc/kernel/ptrace.c | 2
arch/arm/kernel/ptrace.c | 4 -
arch/arm64/kernel/ptrace.c | 4 -
arch/csky/kernel/ptrace.c | 4 -
arch/hexagon/kernel/traps.c | 2
arch/loongarch/kernel/syscall.c | 17 +---
arch/m68k/kernel/ptrace.c | 4 -
arch/microblaze/kernel/ptrace.c | 2
arch/mips/kernel/ptrace.c | 4 -
arch/nios2/kernel/ptrace.c | 2
arch/openrisc/kernel/ptrace.c | 2
arch/parisc/kernel/ptrace.c | 12 +--
arch/powerpc/kernel/syscall.c | 5 -
arch/riscv/kernel/traps.c | 14 +--
arch/s390/kernel/syscall.c | 11 +-
arch/sh/kernel/ptrace_32.c | 4 -
arch/sparc/kernel/ptrace_32.c | 2
arch/sparc/kernel/ptrace_64.c | 2
arch/um/kernel/ptrace.c | 2
arch/um/kernel/skas/syscall.c | 2
arch/x86/entry/syscall_32.c | 70 +++++++------------
arch/x86/entry/syscall_64.c | 61 ++++++----------
arch/x86/entry/vsyscall/vsyscall_64.c | 14 +--
arch/x86/include/asm/entry-common.h | 1
arch/x86/include/asm/syscall.h | 10 --
arch/xtensa/kernel/ptrace.c | 5 -
include/asm-generic/syscall.h | 4 -
include/linux/entry-common.h | 125 ++++++++++++++++++++--------------
include/linux/irq-entry-common.h | 6 -
include/linux/ptrace.h | 13 +--
include/linux/randomize_kstack.h | 19 +++++
include/linux/seccomp.h | 12 +--
kernel/entry/syscall-common.c | 7 +
kernel/seccomp.c | 35 ++++-----
36 files changed, 264 insertions(+), 256 deletions(-)
^ permalink raw reply [flat|nested] 107+ messages in thread
* [patch 01/18] powerpc: Move stack randomization after syscall_enter_from_user_mode()
2026-07-07 19:05 [patch 00/18] entry: Consolidate and rework syscall entry handling Thomas Gleixner
@ 2026-07-07 19:05 ` Thomas Gleixner
2026-07-08 14:07 ` Shrikanth Hegde
` (4 more replies)
2026-07-07 19:06 ` [patch 02/18] randomize_kstack: Provide add_random_kstack_offset_irqsoff() Thomas Gleixner
` (20 subsequent siblings)
21 siblings, 5 replies; 107+ messages in thread
From: Thomas Gleixner @ 2026-07-07 19:05 UTC (permalink / raw)
To: LKML
Cc: Peter Zijlstra, Michael Ellerman, Shrikanth Hegde, linuxppc-dev,
Kees Cook, Huacai Chen, loongarch, Paul Walmsley, Palmer Dabbelt,
linux-riscv, Sven Schnelle, linux-s390, x86, Mark Rutland,
Jinjie Ruan, Andy Lutomirski, Oleg Nesterov, Richard Henderson,
Russell King, Catalin Marinas, Guo Ren, Geert Uytterhoeven,
Thomas Bogendoerfer, Helge Deller, Yoshinori Sato,
Richard Weinberger, Chris Zankel, linux-arm-kernel, linux-alpha,
linux-csky, linux-m68k, linux-mips, linux-parisc, linux-sh,
linux-um, Arnd Bergmann, Vineet Gupta, Will Deacon, Brian Cain,
Michal Simek, Dinh Nguyen, David S. Miller, Andreas Larsson,
linux-snps-arc, linux-hexagon, linux-openrisc, sparclinux,
linux-arch, Michal Suchánek, Jonathan Corbet, linux-doc
add_random_kstack_offset() is invoked before syscall_enter_from_user_mode()
establishes state. That's wrong because add_random_kstack_offset() calls
into instrumentable code.
Move it after syscall_enter_from_user_mode() to ensure that state is
correctly established.
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Cc: Michael Ellerman <mpe@ellerman.id.au>
Cc: Shrikanth Hegde <sshegde@linux.ibm.com>
Cc: linuxppc-dev@lists.ozlabs.org
---
arch/powerpc/kernel/syscall.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
--- a/arch/powerpc/kernel/syscall.c
+++ b/arch/powerpc/kernel/syscall.c
@@ -19,8 +19,8 @@ notrace long system_call_exception(struc
long ret;
syscall_fn f;
- add_random_kstack_offset();
r0 = syscall_enter_from_user_mode(regs, r0);
+ add_random_kstack_offset();
if (unlikely(r0 >= NR_syscalls)) {
if (unlikely(trap_is_unsupported_scv(regs))) {
^ permalink raw reply [flat|nested] 107+ messages in thread
* [patch 02/18] randomize_kstack: Provide add_random_kstack_offset_irqsoff()
2026-07-07 19:05 [patch 00/18] entry: Consolidate and rework syscall entry handling Thomas Gleixner
2026-07-07 19:05 ` [patch 01/18] powerpc: Move stack randomization after syscall_enter_from_user_mode() Thomas Gleixner
@ 2026-07-07 19:06 ` Thomas Gleixner
2026-07-08 17:24 ` Radu Rendec
` (3 more replies)
2026-07-07 19:06 ` [patch 03/18] entry: Provide [syscall_]enter_from_user_mode_randomize_stack() Thomas Gleixner
` (19 subsequent siblings)
21 siblings, 4 replies; 107+ messages in thread
From: Thomas Gleixner @ 2026-07-07 19:06 UTC (permalink / raw)
To: LKML
Cc: Peter Zijlstra, Kees Cook, Michael Ellerman, Shrikanth Hegde,
linuxppc-dev, Huacai Chen, loongarch, Paul Walmsley,
Palmer Dabbelt, linux-riscv, Sven Schnelle, linux-s390, x86,
Mark Rutland, Jinjie Ruan, Andy Lutomirski, Oleg Nesterov,
Richard Henderson, Russell King, Catalin Marinas, Guo Ren,
Geert Uytterhoeven, Thomas Bogendoerfer, Helge Deller,
Yoshinori Sato, Richard Weinberger, Chris Zankel,
linux-arm-kernel, linux-alpha, linux-csky, linux-m68k, linux-mips,
linux-parisc, linux-sh, linux-um, Arnd Bergmann, Vineet Gupta,
Will Deacon, Brian Cain, Michal Simek, Dinh Nguyen,
David S. Miller, Andreas Larsson, linux-snps-arc, linux-hexagon,
linux-openrisc, sparclinux, linux-arch, Michal Suchánek,
Jonathan Corbet, linux-doc
add_random_kstack_offset() uses get/put_cpu_var() which is pointless
overhead when it is invoked from low level entry code with interrupts
disabled.
Provide a irqsoff() variant, which avoids that.
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Cc: Kees Cook <kees@kernel.org>
---
include/linux/randomize_kstack.h | 19 +++++++++++++++++++
1 file changed, 19 insertions(+)
--- a/include/linux/randomize_kstack.h
+++ b/include/linux/randomize_kstack.h
@@ -77,8 +77,27 @@ static __always_inline u32 get_kstack_of
} \
} while (0)
+/**
+ * add_random_kstack_offset_irqsoff - Increase stack utilization by a random offset.
+ *
+ * This should be used in the syscall entry path after user registers have been
+ * stored to the stack. Interrupts must be still disabled.
+ */
+#define add_random_kstack_offset_irqsoff() \
+do { \
+ lockdep_assert_irqs_disabled(); \
+ if (static_branch_maybe(CONFIG_RANDOMIZE_KSTACK_OFFSET_DEFAULT, \
+ &randomize_kstack_offset)) { \
+ u32 offset = prandom_u32_state(raw_cpu_ptr(&kstack_rnd_state)); \
+ u8 *ptr = __kstack_alloca(KSTACK_OFFSET_MAX(offset)); \
+ /* Keep allocation even after "ptr" loses scope. */ \
+ asm volatile("" :: "r"(ptr) : "memory"); \
+ } \
+} while (0)
+
#else /* CONFIG_RANDOMIZE_KSTACK_OFFSET */
#define add_random_kstack_offset() do { } while (0)
+#define add_random_kstack_offset_irqsoff() do { } while (0)
#endif /* CONFIG_RANDOMIZE_KSTACK_OFFSET */
#endif
^ permalink raw reply [flat|nested] 107+ messages in thread
* [patch 03/18] entry: Provide [syscall_]enter_from_user_mode_randomize_stack()
2026-07-07 19:05 [patch 00/18] entry: Consolidate and rework syscall entry handling Thomas Gleixner
2026-07-07 19:05 ` [patch 01/18] powerpc: Move stack randomization after syscall_enter_from_user_mode() Thomas Gleixner
2026-07-07 19:06 ` [patch 02/18] randomize_kstack: Provide add_random_kstack_offset_irqsoff() Thomas Gleixner
@ 2026-07-07 19:06 ` Thomas Gleixner
2026-07-08 17:26 ` Radu Rendec
` (5 more replies)
2026-07-07 19:06 ` [patch 04/18] loongarch/syscall: Use syscall_enter_from_user_mode_randomize_stack() Thomas Gleixner
` (18 subsequent siblings)
21 siblings, 6 replies; 107+ messages in thread
From: Thomas Gleixner @ 2026-07-07 19:06 UTC (permalink / raw)
To: LKML
Cc: Peter Zijlstra, Michael Ellerman, Shrikanth Hegde, linuxppc-dev,
Kees Cook, Huacai Chen, loongarch, Paul Walmsley, Palmer Dabbelt,
linux-riscv, Sven Schnelle, linux-s390, x86, Mark Rutland,
Jinjie Ruan, Andy Lutomirski, Oleg Nesterov, Richard Henderson,
Russell King, Catalin Marinas, Guo Ren, Geert Uytterhoeven,
Thomas Bogendoerfer, Helge Deller, Yoshinori Sato,
Richard Weinberger, Chris Zankel, linux-arm-kernel, linux-alpha,
linux-csky, linux-m68k, linux-mips, linux-parisc, linux-sh,
linux-um, Arnd Bergmann, Vineet Gupta, Will Deacon, Brian Cain,
Michal Simek, Dinh Nguyen, David S. Miller, Andreas Larsson,
linux-snps-arc, linux-hexagon, linux-openrisc, sparclinux,
linux-arch, Michal Suchánek, Jonathan Corbet, linux-doc
Randomizing the syscall stack can only happen after state is established
via enter_from_user_mode() or syscall_enter_from_user_mode(). The earlier
it happens the better.
Provide two new macros to consolidate that:
- enter_from_user_mode_randomize_stack()
enter_from_user_mode();
add_random_kstack_offset_irqsoff();
- syscall_enter_from_user_mode_randomize_stack()
enter_from_user_mode_randomize_stack();
syscall_enter_from_user_mode_work();
to reduce boiler plate code.
Those are macros and not inline functions as the latter would limit the
stack randomization scope to the inline function itself.
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
---
include/linux/entry-common.h | 56 +++++++++++++++++++++++++++++++++++++++++++
1 file changed, 56 insertions(+)
--- a/include/linux/entry-common.h
+++ b/include/linux/entry-common.h
@@ -6,6 +6,7 @@
#include <linux/irq-entry-common.h>
#include <linux/livepatch.h>
#include <linux/ptrace.h>
+#include <linux/randomize_kstack.h>
#include <linux/resume_user_mode.h>
#include <linux/seccomp.h>
#include <linux/sched.h>
@@ -150,6 +151,61 @@ static __always_inline long syscall_ente
}
/**
+ * enter_from_user_mode_randomize_stack - Establish state and add stack randomization
+ * before invoking syscall_enter_from_user_mode_work()
+ * @regs: Pointer to currents pt_regs
+ *
+ * Invoked from architecture specific syscall entry code with interrupts
+ * disabled. The calling code has to be non-instrumentable. When the function
+ * returns all state is correct, interrupts are still disabled and the
+ * subsequent functions can be instrumented.
+ *
+ * Implemented as a macro so that the stack randomization is effective
+ * throughout the function in which it is invoked. An inline would only make it
+ * effective in the scope of the inline function.
+ */
+#define enter_from_user_mode_randomize_stack(regs) \
+do { \
+ enter_from_user_mode(regs); \
+ instrumentation_begin(); \
+ add_random_kstack_offset_irqsoff(); \
+ instrumentation_end(); \
+} while (0)
+
+/**
+ * syscall_enter_from_user_mode_randomize_stack - Establish state and check and handle work
+ * before invoking a syscall
+ * @regs: Pointer to currents pt_regs
+ * @syscall: The syscall number
+ *
+ * Invoked from architecture specific syscall entry code with interrupts
+ * disabled. The calling code has to be non-instrumentable. When the
+ * function returns all state is correct, interrupts are enabled and the
+ * subsequent functions can be instrumented.
+ *
+ * This is the combination of enter_from_user_mode_randomize_stack() and
+ * syscall_enter_from_user_mode_work() to be used when there is no
+ * architecture specific work to be done between the two.
+ *
+ * Returns: The original or a modified syscall number. See
+ * syscall_enter_from_user_mode_work() for further explanation.
+ *
+ * Implemented as a macro to make stack randomization effective in the calling
+ * scope.
+ */
+#define syscall_enter_from_user_mode_randomize_stack(regs, syscall) \
+({ \
+ enter_from_user_mode_randomize_stack(regs); \
+ \
+ instrumentation_begin(); \
+ local_irq_enable(); \
+ long _ret = syscall_enter_from_user_mode_work(regs, syscall); \
+ instrumentation_end(); \
+ \
+ _ret; \
+})
+
+/**
* syscall_enter_from_user_mode - Establish state and check and handle work
* before invoking a syscall
* @regs: Pointer to currents pt_regs
^ permalink raw reply [flat|nested] 107+ messages in thread
* [patch 04/18] loongarch/syscall: Use syscall_enter_from_user_mode_randomize_stack()
2026-07-07 19:05 [patch 00/18] entry: Consolidate and rework syscall entry handling Thomas Gleixner
` (2 preceding siblings ...)
2026-07-07 19:06 ` [patch 03/18] entry: Provide [syscall_]enter_from_user_mode_randomize_stack() Thomas Gleixner
@ 2026-07-07 19:06 ` Thomas Gleixner
2026-07-08 18:37 ` Radu Rendec
` (3 more replies)
2026-07-07 19:06 ` [patch 05/18] powerpc/syscall: " Thomas Gleixner
` (17 subsequent siblings)
21 siblings, 4 replies; 107+ messages in thread
From: Thomas Gleixner @ 2026-07-07 19:06 UTC (permalink / raw)
To: LKML
Cc: Peter Zijlstra, Huacai Chen, loongarch, Michael Ellerman,
Shrikanth Hegde, linuxppc-dev, Kees Cook, Paul Walmsley,
Palmer Dabbelt, linux-riscv, Sven Schnelle, linux-s390, x86,
Mark Rutland, Jinjie Ruan, Andy Lutomirski, Oleg Nesterov,
Richard Henderson, Russell King, Catalin Marinas, Guo Ren,
Geert Uytterhoeven, Thomas Bogendoerfer, Helge Deller,
Yoshinori Sato, Richard Weinberger, Chris Zankel,
linux-arm-kernel, linux-alpha, linux-csky, linux-m68k, linux-mips,
linux-parisc, linux-sh, linux-um, Arnd Bergmann, Vineet Gupta,
Will Deacon, Brian Cain, Michal Simek, Dinh Nguyen,
David S. Miller, Andreas Larsson, linux-snps-arc, linux-hexagon,
linux-openrisc, sparclinux, linux-arch, Michal Suchánek,
Jonathan Corbet, linux-doc
syscall_enter_from_user_mode_randomize_stack() replaces
syscall_enter_from_user_mode() and the subsequent invocation of
add_random_kstack_offset().
The advantage is that it applies the stack randomization right after
enter_from_user_mode() and thereby avoids the overhead of get/put_cpu_var()
as that code is invoked with interrupts disabled.
No functional change.
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Cc: Huacai Chen <chenhuacai@kernel.org>
Cc: loongarch@lists.linux.dev
---
arch/loongarch/kernel/syscall.c | 5 +----
1 file changed, 1 insertion(+), 4 deletions(-)
--- a/arch/loongarch/kernel/syscall.c
+++ b/arch/loongarch/kernel/syscall.c
@@ -11,7 +11,6 @@
#include <linux/linkage.h>
#include <linux/nospec.h>
#include <linux/objtool.h>
-#include <linux/randomize_kstack.h>
#include <linux/syscalls.h>
#include <linux/unistd.h>
@@ -70,9 +69,7 @@ void noinstr __no_stack_protector do_sys
regs->orig_a0 = regs->regs[4];
regs->regs[4] = -ENOSYS;
- nr = syscall_enter_from_user_mode(regs, nr);
-
- add_random_kstack_offset();
+ nr = syscall_enter_from_user_mode_randomize_stack(regs, nr);
if (nr < NR_syscalls) {
syscall_fn = sys_call_table[array_index_nospec(nr, NR_syscalls)];
^ permalink raw reply [flat|nested] 107+ messages in thread
* [patch 05/18] powerpc/syscall: Use syscall_enter_from_user_mode_randomize_stack()
2026-07-07 19:05 [patch 00/18] entry: Consolidate and rework syscall entry handling Thomas Gleixner
` (3 preceding siblings ...)
2026-07-07 19:06 ` [patch 04/18] loongarch/syscall: Use syscall_enter_from_user_mode_randomize_stack() Thomas Gleixner
@ 2026-07-07 19:06 ` Thomas Gleixner
2026-07-08 18:35 ` Radu Rendec
` (3 more replies)
2026-07-07 19:06 ` [patch 06/18] riscv/syscall: " Thomas Gleixner
` (16 subsequent siblings)
21 siblings, 4 replies; 107+ messages in thread
From: Thomas Gleixner @ 2026-07-07 19:06 UTC (permalink / raw)
To: LKML
Cc: Peter Zijlstra, Michael Ellerman, Shrikanth Hegde, linuxppc-dev,
Kees Cook, Huacai Chen, loongarch, Paul Walmsley, Palmer Dabbelt,
linux-riscv, Sven Schnelle, linux-s390, x86, Mark Rutland,
Jinjie Ruan, Andy Lutomirski, Oleg Nesterov, Richard Henderson,
Russell King, Catalin Marinas, Guo Ren, Geert Uytterhoeven,
Thomas Bogendoerfer, Helge Deller, Yoshinori Sato,
Richard Weinberger, Chris Zankel, linux-arm-kernel, linux-alpha,
linux-csky, linux-m68k, linux-mips, linux-parisc, linux-sh,
linux-um, Arnd Bergmann, Vineet Gupta, Will Deacon, Brian Cain,
Michal Simek, Dinh Nguyen, David S. Miller, Andreas Larsson,
linux-snps-arc, linux-hexagon, linux-openrisc, sparclinux,
linux-arch, Michal Suchánek, Jonathan Corbet, linux-doc
syscall_enter_from_user_mode_randomize_stack() replaces
syscall_enter_from_user_mode() and the subsequent invocation of
add_random_kstack_offset().
The advantage is that it applies the stack randomization right after
enter_from_user_mode() and thereby avoids the overhead of get/put_cpu_var()
as that code is invoked with interrupts disabled.
No functional change.
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Cc: Michael Ellerman <mpe@ellerman.id.au>
Cc: Shrikanth Hegde <sshegde@linux.ibm.com>
Cc: linuxppc-dev@lists.ozlabs.org
---
arch/powerpc/kernel/syscall.c | 4 +---
1 file changed, 1 insertion(+), 3 deletions(-)
--- a/arch/powerpc/kernel/syscall.c
+++ b/arch/powerpc/kernel/syscall.c
@@ -2,7 +2,6 @@
#include <linux/compat.h>
#include <linux/context_tracking.h>
-#include <linux/randomize_kstack.h>
#include <linux/entry-common.h>
#include <asm/interrupt.h>
@@ -19,8 +18,7 @@ notrace long system_call_exception(struc
long ret;
syscall_fn f;
- r0 = syscall_enter_from_user_mode(regs, r0);
- add_random_kstack_offset();
+ r0 = syscall_enter_from_user_mode_randomize_stack(regs, r0);
if (unlikely(r0 >= NR_syscalls)) {
if (unlikely(trap_is_unsupported_scv(regs))) {
^ permalink raw reply [flat|nested] 107+ messages in thread
* [patch 06/18] riscv/syscall: Use syscall_enter_from_user_mode_randomize_stack()
2026-07-07 19:05 [patch 00/18] entry: Consolidate and rework syscall entry handling Thomas Gleixner
` (4 preceding siblings ...)
2026-07-07 19:06 ` [patch 05/18] powerpc/syscall: " Thomas Gleixner
@ 2026-07-07 19:06 ` Thomas Gleixner
2026-07-08 20:57 ` Radu Rendec
` (4 more replies)
2026-07-07 19:06 ` [patch 07/18] s390/syscall: Use enter_from_user_mode_randomize_stack() Thomas Gleixner
` (15 subsequent siblings)
21 siblings, 5 replies; 107+ messages in thread
From: Thomas Gleixner @ 2026-07-07 19:06 UTC (permalink / raw)
To: LKML
Cc: Peter Zijlstra, Paul Walmsley, Palmer Dabbelt, linux-riscv,
Michael Ellerman, Shrikanth Hegde, linuxppc-dev, Kees Cook,
Huacai Chen, loongarch, Sven Schnelle, linux-s390, x86,
Mark Rutland, Jinjie Ruan, Andy Lutomirski, Oleg Nesterov,
Richard Henderson, Russell King, Catalin Marinas, Guo Ren,
Geert Uytterhoeven, Thomas Bogendoerfer, Helge Deller,
Yoshinori Sato, Richard Weinberger, Chris Zankel,
linux-arm-kernel, linux-alpha, linux-csky, linux-m68k, linux-mips,
linux-parisc, linux-sh, linux-um, Arnd Bergmann, Vineet Gupta,
Will Deacon, Brian Cain, Michal Simek, Dinh Nguyen,
David S. Miller, Andreas Larsson, linux-snps-arc, linux-hexagon,
linux-openrisc, sparclinux, linux-arch, Michal Suchánek,
Jonathan Corbet, linux-doc
syscall_enter_from_user_mode_randomize_stack() replaces
syscall_enter_from_user_mode() and the subsequent invocation of
add_random_kstack_offset().
The advantage is that it applies the stack randomization right after
enter_from_user_mode() and thereby avoids the overhead of get/put_cpu_var()
as that code is invoked with interrupts disabled.
No functional change.
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Cc: Paul Walmsley <pjw@kernel.org>
Cc: Palmer Dabbelt <palmer@dabbelt.com>
Cc: linux-riscv@lists.infradead.org
---
arch/riscv/kernel/traps.c | 5 +----
1 file changed, 1 insertion(+), 4 deletions(-)
--- a/arch/riscv/kernel/traps.c
+++ b/arch/riscv/kernel/traps.c
@@ -7,7 +7,6 @@
#include <linux/kernel.h>
#include <linux/init.h>
#include <linux/irqflags.h>
-#include <linux/randomize_kstack.h>
#include <linux/sched.h>
#include <linux/sched/debug.h>
#include <linux/sched/signal.h>
@@ -333,9 +332,7 @@ void do_trap_ecall_u(struct pt_regs *reg
riscv_v_vstate_discard(regs);
- syscall = syscall_enter_from_user_mode(regs, syscall);
-
- add_random_kstack_offset();
+ syscall = syscall_enter_from_user_mode_randomize_stack(regs, syscall);
if (syscall >= 0 && syscall < NR_syscalls) {
syscall = array_index_nospec(syscall, NR_syscalls);
^ permalink raw reply [flat|nested] 107+ messages in thread
* [patch 07/18] s390/syscall: Use enter_from_user_mode_randomize_stack()
2026-07-07 19:05 [patch 00/18] entry: Consolidate and rework syscall entry handling Thomas Gleixner
` (5 preceding siblings ...)
2026-07-07 19:06 ` [patch 06/18] riscv/syscall: " Thomas Gleixner
@ 2026-07-07 19:06 ` Thomas Gleixner
2026-07-08 6:47 ` Sven Schnelle
` (4 more replies)
2026-07-07 19:06 ` [patch 08/18] x86/syscall: Use [syscall_]enter_from_user_mode_randomize_stack() Thomas Gleixner
` (14 subsequent siblings)
21 siblings, 5 replies; 107+ messages in thread
From: Thomas Gleixner @ 2026-07-07 19:06 UTC (permalink / raw)
To: LKML
Cc: Peter Zijlstra, Sven Schnelle, linux-s390, Michael Ellerman,
Shrikanth Hegde, linuxppc-dev, Kees Cook, Huacai Chen, loongarch,
Paul Walmsley, Palmer Dabbelt, linux-riscv, x86, Mark Rutland,
Jinjie Ruan, Andy Lutomirski, Oleg Nesterov, Richard Henderson,
Russell King, Catalin Marinas, Guo Ren, Geert Uytterhoeven,
Thomas Bogendoerfer, Helge Deller, Yoshinori Sato,
Richard Weinberger, Chris Zankel, linux-arm-kernel, linux-alpha,
linux-csky, linux-m68k, linux-mips, linux-parisc, linux-sh,
linux-um, Arnd Bergmann, Vineet Gupta, Will Deacon, Brian Cain,
Michal Simek, Dinh Nguyen, David S. Miller, Andreas Larsson,
linux-snps-arc, linux-hexagon, linux-openrisc, sparclinux,
linux-arch, Michal Suchánek, Jonathan Corbet, linux-doc
enter_from_user_mode_randomize_stack() replaces enter_from_user_mode() and
the subsequent invocation of add_random_kstack_offset_irqsoff().
As a bonus this avoids the overhead of get/put_cpu_var() in
add_random_kstack_offset().
No functional change.
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Cc: Sven Schnelle <svens@linux.ibm.com>
Cc: linux-s390@vger.kernel.org
---
arch/s390/kernel/syscall.c | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
--- a/arch/s390/kernel/syscall.c
+++ b/arch/s390/kernel/syscall.c
@@ -97,8 +97,8 @@ void noinstr __do_syscall(struct pt_regs
{
unsigned long nr;
- enter_from_user_mode(regs);
- add_random_kstack_offset();
+ enter_from_user_mode_randomize_stack(regs);
+
regs->psw = get_lowcore()->svc_old_psw;
regs->int_code = get_lowcore()->svc_int_code;
update_timer_sys();
^ permalink raw reply [flat|nested] 107+ messages in thread
* [patch 08/18] x86/syscall: Use [syscall_]enter_from_user_mode_randomize_stack()
2026-07-07 19:05 [patch 00/18] entry: Consolidate and rework syscall entry handling Thomas Gleixner
` (6 preceding siblings ...)
2026-07-07 19:06 ` [patch 07/18] s390/syscall: Use enter_from_user_mode_randomize_stack() Thomas Gleixner
@ 2026-07-07 19:06 ` Thomas Gleixner
2026-07-08 20:59 ` Radu Rendec
` (3 more replies)
2026-07-07 19:06 ` [patch 09/18] entry: Remove syscall_enter_from_user_mode() Thomas Gleixner
` (13 subsequent siblings)
21 siblings, 4 replies; 107+ messages in thread
From: Thomas Gleixner @ 2026-07-07 19:06 UTC (permalink / raw)
To: LKML
Cc: Peter Zijlstra, x86, Michael Ellerman, Shrikanth Hegde,
linuxppc-dev, Kees Cook, Huacai Chen, loongarch, Paul Walmsley,
Palmer Dabbelt, linux-riscv, Sven Schnelle, linux-s390,
Mark Rutland, Jinjie Ruan, Andy Lutomirski, Oleg Nesterov,
Richard Henderson, Russell King, Catalin Marinas, Guo Ren,
Geert Uytterhoeven, Thomas Bogendoerfer, Helge Deller,
Yoshinori Sato, Richard Weinberger, Chris Zankel,
linux-arm-kernel, linux-alpha, linux-csky, linux-m68k, linux-mips,
linux-parisc, linux-sh, linux-um, Arnd Bergmann, Vineet Gupta,
Will Deacon, Brian Cain, Michal Simek, Dinh Nguyen,
David S. Miller, Andreas Larsson, linux-snps-arc, linux-hexagon,
linux-openrisc, sparclinux, linux-arch, Michal Suchánek,
Jonathan Corbet, linux-doc
These functions integrate the stack randomization.
syscall_enter_from_user_mode_randomize_stack() has the advantage that the
randomization happens early right after enter_from_user_mode().
In both cases also the overhead of get/put_cpu_var() in
add_random_kstack_offset() is avoided.
No functional change.
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Cc: x86@kernel.org
---
arch/x86/entry/syscall_32.c | 19 +++++--------------
arch/x86/entry/syscall_64.c | 3 +--
arch/x86/include/asm/entry-common.h | 1 -
3 files changed, 6 insertions(+), 17 deletions(-)
--- a/arch/x86/entry/syscall_32.c
+++ b/arch/x86/entry/syscall_32.c
@@ -142,10 +142,9 @@ static __always_inline bool int80_is_ext
* int80_is_external() below which calls into the APIC driver.
* Identical for soft and external interrupts.
*/
- enter_from_user_mode(regs);
+ enter_from_user_mode_randomize_stack(regs);
instrumentation_begin();
- add_random_kstack_offset();
/* Validate that this is a soft interrupt to the extent possible */
if (unlikely(int80_is_external()))
@@ -210,11 +209,9 @@ DEFINE_FREDENTRY_RAW(int80_emulation)
{
int nr;
- enter_from_user_mode(regs);
+ enter_from_user_mode_randomize_stack(regs);
instrumentation_begin();
- add_random_kstack_offset();
-
/*
* FRED pushed 0 into regs::orig_ax and regs::ax contains the
* syscall number.
@@ -252,10 +249,10 @@ DEFINE_FREDENTRY_RAW(int80_emulation)
* orig_ax, the int return value truncates it. This matches
* the semantics of syscall_get_nr().
*/
- nr = syscall_enter_from_user_mode(regs, nr);
+ nr = syscall_enter_from_user_mode_randomize_stack(regs, nr);
+
instrumentation_begin();
- add_random_kstack_offset();
do_syscall_32_irqs_on(regs, nr);
instrumentation_end();
@@ -268,15 +265,9 @@ static noinstr bool __do_fast_syscall_32
int nr = syscall_32_enter(regs);
int res;
- /*
- * This cannot use syscall_enter_from_user_mode() as it has to
- * fetch EBP before invoking any of the syscall entry work
- * functions.
- */
- enter_from_user_mode(regs);
+ enter_from_user_mode_randomize_stack(regs);
instrumentation_begin();
- add_random_kstack_offset();
local_irq_enable();
/* Fetch EBP from where the vDSO stashed it. */
if (IS_ENABLED(CONFIG_X86_64)) {
--- a/arch/x86/entry/syscall_64.c
+++ b/arch/x86/entry/syscall_64.c
@@ -86,10 +86,9 @@ static __always_inline bool do_syscall_x
/* Returns true to return using SYSRET, or false to use IRET */
__visible noinstr bool do_syscall_64(struct pt_regs *regs, int nr)
{
- nr = syscall_enter_from_user_mode(regs, nr);
+ nr = syscall_enter_from_user_mode_randomize_stack(regs, nr);
instrumentation_begin();
- add_random_kstack_offset();
if (!do_syscall_x64(regs, nr) && !do_syscall_x32(regs, nr) && nr != -1) {
/* Invalid system call, but still a system call. */
--- a/arch/x86/include/asm/entry-common.h
+++ b/arch/x86/include/asm/entry-common.h
@@ -2,7 +2,6 @@
#ifndef _ASM_X86_ENTRY_COMMON_H
#define _ASM_X86_ENTRY_COMMON_H
-#include <linux/randomize_kstack.h>
#include <linux/user-return-notifier.h>
#include <asm/nospec-branch.h>
^ permalink raw reply [flat|nested] 107+ messages in thread
* [patch 09/18] entry: Remove syscall_enter_from_user_mode()
2026-07-07 19:05 [patch 00/18] entry: Consolidate and rework syscall entry handling Thomas Gleixner
` (7 preceding siblings ...)
2026-07-07 19:06 ` [patch 08/18] x86/syscall: Use [syscall_]enter_from_user_mode_randomize_stack() Thomas Gleixner
@ 2026-07-07 19:06 ` Thomas Gleixner
2026-07-08 21:21 ` Radu Rendec
` (2 more replies)
2026-07-07 19:06 ` [patch 10/18] entry: Use syscall number instead of rereading it Thomas Gleixner
` (12 subsequent siblings)
21 siblings, 3 replies; 107+ messages in thread
From: Thomas Gleixner @ 2026-07-07 19:06 UTC (permalink / raw)
To: LKML
Cc: Peter Zijlstra, Michael Ellerman, Shrikanth Hegde, linuxppc-dev,
Kees Cook, Huacai Chen, loongarch, Paul Walmsley, Palmer Dabbelt,
linux-riscv, Sven Schnelle, linux-s390, x86, Mark Rutland,
Jinjie Ruan, Andy Lutomirski, Oleg Nesterov, Richard Henderson,
Russell King, Catalin Marinas, Guo Ren, Geert Uytterhoeven,
Thomas Bogendoerfer, Helge Deller, Yoshinori Sato,
Richard Weinberger, Chris Zankel, linux-arm-kernel, linux-alpha,
linux-csky, linux-m68k, linux-mips, linux-parisc, linux-sh,
linux-um, Arnd Bergmann, Vineet Gupta, Will Deacon, Brian Cain,
Michal Simek, Dinh Nguyen, David S. Miller, Andreas Larsson,
linux-snps-arc, linux-hexagon, linux-openrisc, sparclinux,
linux-arch, Michal Suchánek, Jonathan Corbet, linux-doc
All architecture use either:
nr = enter_from_user_mode_randomize_stack(regs, nr);
or
enter_from_user_mode_randomize_stack(regs);
nr = syscall_enter_from_user_mode_work(regs, nr);
Remove the now unused function.
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
---
Documentation/core-api/entry.rst | 17 +++++++++-------
include/linux/entry-common.h | 40 +++------------------------------------
include/linux/irq-entry-common.h | 6 ++---
3 files changed, 17 insertions(+), 46 deletions(-)
--- a/Documentation/core-api/entry.rst
+++ b/Documentation/core-api/entry.rst
@@ -68,7 +68,7 @@ low-level C code must not be instrumente
noinstr void syscall(struct pt_regs *regs, int nr)
{
arch_syscall_enter(regs);
- nr = syscall_enter_from_user_mode(regs, nr);
+ nr = syscall_enter_from_user_mode_randomize_stack(regs, nr);
instrumentation_begin();
if (!invoke_syscall(regs, nr) && nr != -1)
@@ -78,12 +78,14 @@ low-level C code must not be instrumente
syscall_exit_to_user_mode(regs);
}
-syscall_enter_from_user_mode() first invokes enter_from_user_mode() which
-establishes state in the following order:
+syscall_enter_from_user_mode_randomize_stack() first invokes
+enter_from_user_mode_randomize_stack() which establishes state in the
+following order:
* Lockdep
* RCU / Context tracking
* Tracing
+ * Apply stack randomization
and then invokes the various entry work functions like ptrace, seccomp, audit,
syscall tracing, etc. After all that is done, the instrumentable invoke_syscall
@@ -99,10 +101,11 @@ that it invokes exit_to_user_mode() whic
* RCU / Context tracking
* Lockdep
-syscall_enter_from_user_mode() and syscall_exit_to_user_mode() are also
-available as fine grained subfunctions in cases where the architecture code
-has to do extra work between the various steps. In such cases it has to
-ensure that enter_from_user_mode() is called first on entry and
+syscall_enter_from_user_mode_randomize_stack() and
+syscall_exit_to_user_mode() are also available as fine grained subfunctions
+in cases where the architecture code has to do extra work between the
+various steps. In such cases it has to ensure that
+enter_from_user_mode_randomize_stack() is called first on entry and
exit_to_user_mode() is called last on exit.
Do not nest syscalls. Nested syscalls will cause RCU and/or context tracking
--- a/include/linux/entry-common.h
+++ b/include/linux/entry-common.h
@@ -19,7 +19,7 @@
#endif
/*
- * SYSCALL_WORK flags handled in syscall_enter_from_user_mode()
+ * SYSCALL_WORK flags handled in syscall_enter_from_user_mode_work()
*/
#define SYSCALL_WORK_ENTER (SYSCALL_WORK_SECCOMP | \
SYSCALL_WORK_SYSCALL_TRACEPOINT | \
@@ -205,42 +205,10 @@ do { \
_ret; \
})
-/**
- * syscall_enter_from_user_mode - Establish state and check and handle work
- * before invoking a syscall
- * @regs: Pointer to currents pt_regs
- * @syscall: The syscall number
- *
- * Invoked from architecture specific syscall entry code with interrupts
- * disabled. The calling code has to be non-instrumentable. When the
- * function returns all state is correct, interrupts are enabled and the
- * subsequent functions can be instrumented.
- *
- * This is the combination of enter_from_user_mode() and
- * syscall_enter_from_user_mode_work() to be used when there is no
- * architecture specific work to be done between the two.
- *
- * Returns: The original or a modified syscall number. See
- * syscall_enter_from_user_mode_work() for further explanation.
- */
-static __always_inline long syscall_enter_from_user_mode(struct pt_regs *regs, long syscall)
-{
- long ret;
-
- enter_from_user_mode(regs);
-
- instrumentation_begin();
- local_irq_enable();
- ret = syscall_enter_from_user_mode_work(regs, syscall);
- instrumentation_end();
-
- return ret;
-}
-
/*
- * If SYSCALL_EMU is set, then the only reason to report is when
- * SINGLESTEP is set (i.e. PTRACE_SYSEMU_SINGLESTEP). This syscall
- * instruction has been already reported in syscall_enter_from_user_mode().
+ * If SYSCALL_EMU is set, then the only reason to report is when SINGLESTEP is
+ * set (i.e. PTRACE_SYSEMU_SINGLESTEP). This syscall instruction has been
+ * already reported in syscall_enter_from_user_mode_work().
*/
static __always_inline bool report_single_step(unsigned long work)
{
--- a/include/linux/irq-entry-common.h
+++ b/include/linux/irq-entry-common.h
@@ -49,9 +49,9 @@
* Defaults to an empty implementation. Can be replaced by architecture
* specific code.
*
- * Invoked from syscall_enter_from_user_mode() in the non-instrumentable
- * section. Use __always_inline so the compiler cannot push it out of line
- * and make it instrumentable.
+ * Invoked from enter_from_user_mode_syscall_and_randomize_stack() in the
+ * non-instrumentable section. Use __always_inline so the compiler cannot push
+ * it out of line and make it instrumentable.
*/
static __always_inline void arch_enter_from_user_mode(struct pt_regs *regs);
^ permalink raw reply [flat|nested] 107+ messages in thread
* [patch 10/18] entry: Use syscall number instead of rereading it
2026-07-07 19:05 [patch 00/18] entry: Consolidate and rework syscall entry handling Thomas Gleixner
` (8 preceding siblings ...)
2026-07-07 19:06 ` [patch 09/18] entry: Remove syscall_enter_from_user_mode() Thomas Gleixner
@ 2026-07-07 19:06 ` Thomas Gleixner
2026-07-08 21:39 ` Radu Rendec
` (3 more replies)
2026-07-07 19:06 ` [patch 11/18] seccomp, treewide: Rename and convert __secure_computing() to return boolean Thomas Gleixner
` (11 subsequent siblings)
21 siblings, 4 replies; 107+ messages in thread
From: Thomas Gleixner @ 2026-07-07 19:06 UTC (permalink / raw)
To: LKML
Cc: Peter Zijlstra, Michael Ellerman, Shrikanth Hegde, linuxppc-dev,
Kees Cook, Huacai Chen, loongarch, Paul Walmsley, Palmer Dabbelt,
linux-riscv, Sven Schnelle, linux-s390, x86, Mark Rutland,
Jinjie Ruan, Andy Lutomirski, Oleg Nesterov, Richard Henderson,
Russell King, Catalin Marinas, Guo Ren, Geert Uytterhoeven,
Thomas Bogendoerfer, Helge Deller, Yoshinori Sato,
Richard Weinberger, Chris Zankel, linux-arm-kernel, linux-alpha,
linux-csky, linux-m68k, linux-mips, linux-parisc, linux-sh,
linux-um, Arnd Bergmann, Vineet Gupta, Will Deacon, Brian Cain,
Michal Simek, Dinh Nguyen, David S. Miller, Andreas Larsson,
linux-snps-arc, linux-hexagon, linux-openrisc, sparclinux,
linux-arch, Michal Suchánek, Jonathan Corbet, linux-doc
rseq_syscall_enter_work() is invoked before the syscall number can be
modified. So there is no point in rereading it from pt_regs.
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
---
include/linux/entry-common.h | 9 +++++----
1 file changed, 5 insertions(+), 4 deletions(-)
--- a/include/linux/entry-common.h
+++ b/include/linux/entry-common.h
@@ -70,9 +70,10 @@ static inline void syscall_enter_audit(s
}
}
-static __always_inline long syscall_trace_enter(struct pt_regs *regs, unsigned long work)
+static __always_inline long syscall_trace_enter(struct pt_regs *regs, unsigned long work,
+ long syscall)
{
- long syscall, ret = 0;
+ long ret = 0;
/*
* Handle Syscall User Dispatch. This must comes first, since
@@ -90,7 +91,7 @@ static __always_inline long syscall_trac
* through hrtimer_interrupt().
*/
if (work & SYSCALL_WORK_SYSCALL_RSEQ_SLICE)
- rseq_syscall_enter_work(syscall_get_nr(current, regs));
+ rseq_syscall_enter_work(syscall);
/* Handle ptrace */
if (work & (SYSCALL_WORK_SYSCALL_TRACE | SYSCALL_WORK_SYSCALL_EMU)) {
@@ -145,7 +146,7 @@ static __always_inline long syscall_ente
unsigned long work = READ_ONCE(current_thread_info()->syscall_work);
if (work & SYSCALL_WORK_ENTER)
- syscall = syscall_trace_enter(regs, work);
+ syscall = syscall_trace_enter(regs, work, syscall);
return syscall;
}
^ permalink raw reply [flat|nested] 107+ messages in thread
* [patch 11/18] seccomp, treewide: Rename and convert __secure_computing() to return boolean
2026-07-07 19:05 [patch 00/18] entry: Consolidate and rework syscall entry handling Thomas Gleixner
` (9 preceding siblings ...)
2026-07-07 19:06 ` [patch 10/18] entry: Use syscall number instead of rereading it Thomas Gleixner
@ 2026-07-07 19:06 ` Thomas Gleixner
2026-07-08 1:43 ` Jinjie Ruan
` (2 more replies)
2026-07-07 19:06 ` [patch 12/18] ptrace, treewide: Rename ptrace_report_syscall_entry() to ptrace_report_syscall_permit_entry() Thomas Gleixner
` (10 subsequent siblings)
21 siblings, 3 replies; 107+ messages in thread
From: Thomas Gleixner @ 2026-07-07 19:06 UTC (permalink / raw)
To: LKML
Cc: Peter Zijlstra, Mark Rutland, Jinjie Ruan, Kees Cook,
Andy Lutomirski, Oleg Nesterov, Richard Henderson, Russell King,
Catalin Marinas, Guo Ren, Geert Uytterhoeven, Thomas Bogendoerfer,
Helge Deller, Yoshinori Sato, Richard Weinberger, Chris Zankel,
linux-arm-kernel, linux-alpha, linux-csky, linux-m68k, linux-mips,
linux-parisc, linux-sh, linux-um, Michael Ellerman,
Shrikanth Hegde, linuxppc-dev, Huacai Chen, loongarch,
Paul Walmsley, Palmer Dabbelt, linux-riscv, Sven Schnelle,
linux-s390, x86, Arnd Bergmann, Vineet Gupta, Will Deacon,
Brian Cain, Michal Simek, Dinh Nguyen, David S. Miller,
Andreas Larsson, linux-snps-arc, linux-hexagon, linux-openrisc,
sparclinux, linux-arch, Michal Suchánek, Jonathan Corbet,
linux-doc
From: Jinjie Ruan <ruanjinjie@huawei.com>
The return value of __secure_computing() currently uses 0 to indicate
that a system call should be allowed, and -1 to indicate that it should
be blocked/killed. This 0/-1 pattern is non-intuitive for a security
check function and makes the control flow at the call sites less readable.
Furthermore, any potential future changes to these return values would
require a high-risk, error-prone audit of all its users across different
architectures.
Sanitize this logic by converting the return type of __secure_computing()
to a proper boolean, where 'true' explicitly means 'allow' and 'false'
means 'fail/deny'.
Update all the two dozen or so call sites across the tree to align with
this new boolean semantic. No functional changes are intended, as the
callers still return -1 to the lower-level assembly entry code upon
seccomp denial.
Rename the function to __seccomp_permit_syscall() so that the purpose is
entirely clear.
[ tglx: Rename the function ]
Suggested-by: Thomas Gleixner <tglx@kernel.org>
Suggested-by: Mark Rutland <mark.rutland@arm.com>
Signed-off-by: Jinjie Ruan <ruanjinjie@huawei.com>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Cc: Kees Cook <kees@kernel.org>
Cc: Andy Lutomirski <luto@kernel.org>
Cc: Oleg Nesterov <oleg@redhat.com>
Cc: Richard Henderson <richard.henderson@linaro.org>
Cc: Russell King <linux@armlinux.org.uk>
Cc: Catalin Marinas <catalin.marinas@arm.com>
Cc: Guo Ren <guoren@kernel.org>
Cc: Geert Uytterhoeven <geert@linux-m68k.org>
Cc: Thomas Bogendoerfer <tsbogend@alpha.franken.de>
Cc: Helge Deller <deller@gmx.de>
Cc: Yoshinori Sato <ysato@users.sourceforge.jp>
Cc: Richard Weinberger <richard@nod.at>
Cc: Chris Zankel <chris@zankel.net>
Cc: linux-arm-kernel@lists.infradead.org
Cc: linux-alpha@vger.kernel.org
Cc: linux-csky@vger.kernel.org
Cc: linux-m68k@lists.linux-m68k.org
Cc: linux-mips@vger.kernel.org
Cc: linux-parisc@vger.kernel.org
Cc: linux-sh@vger.kernel.org
Cc: linux-um@lists.infradead.org
---
arch/alpha/kernel/ptrace.c | 2 -
arch/arm/kernel/ptrace.c | 2 -
arch/arm64/kernel/ptrace.c | 2 -
arch/csky/kernel/ptrace.c | 2 -
arch/m68k/kernel/ptrace.c | 2 -
arch/mips/kernel/ptrace.c | 2 -
arch/parisc/kernel/ptrace.c | 2 -
arch/sh/kernel/ptrace_32.c | 2 -
arch/um/kernel/skas/syscall.c | 2 -
arch/x86/entry/vsyscall/vsyscall_64.c | 14 ++++++-------
arch/xtensa/kernel/ptrace.c | 3 --
include/linux/entry-common.h | 9 +++-----
include/linux/seccomp.h | 12 +++++------
kernel/seccomp.c | 35 +++++++++++++++++-----------------
14 files changed, 45 insertions(+), 46 deletions(-)
--- a/arch/alpha/kernel/ptrace.c
+++ b/arch/alpha/kernel/ptrace.c
@@ -387,7 +387,7 @@ asmlinkage unsigned long syscall_trace_e
* If this fails, seccomp may already have set up the return value
* (e.g. SECCOMP_RET_ERRNO / TRACE).
*/
- if (secure_computing() == -1) {
+ if (!seccomp_permit_syscall()) {
if (regs->r19 == 0 && regs->r0 == (unsigned long)-1)
syscall_set_return_value(current, regs, -ENOSYS, 0);
syscall_set_nr(current, regs, -1);
--- a/arch/arm/kernel/ptrace.c
+++ b/arch/arm/kernel/ptrace.c
@@ -855,7 +855,7 @@ asmlinkage int syscall_trace_enter(struc
/* Do seccomp after ptrace; syscall may have changed. */
#ifdef CONFIG_HAVE_ARCH_SECCOMP_FILTER
- if (secure_computing() == -1)
+ if (!seccomp_permit_syscall())
return -1;
#else
/* XXX: remove this once OABI gets fixed */
--- a/arch/arm64/kernel/ptrace.c
+++ b/arch/arm64/kernel/ptrace.c
@@ -2420,7 +2420,7 @@ int syscall_trace_enter(struct pt_regs *
}
/* Do the secure computing after ptrace; failures should be fast. */
- if (secure_computing() == -1)
+ if (!seccomp_permit_syscall())
return NO_SYSCALL;
if (test_thread_flag(TIF_SYSCALL_TRACEPOINT))
--- a/arch/csky/kernel/ptrace.c
+++ b/arch/csky/kernel/ptrace.c
@@ -323,7 +323,7 @@ asmlinkage int syscall_trace_enter(struc
if (ptrace_report_syscall_entry(regs))
return -1;
- if (secure_computing() == -1)
+ if (!seccomp_permit_syscall())
return -1;
if (test_thread_flag(TIF_SYSCALL_TRACEPOINT))
--- a/arch/m68k/kernel/ptrace.c
+++ b/arch/m68k/kernel/ptrace.c
@@ -281,7 +281,7 @@ asmlinkage int syscall_trace_enter(void)
if (test_thread_flag(TIF_SYSCALL_TRACE))
ret = ptrace_report_syscall_entry(task_pt_regs(current));
- if (secure_computing() == -1)
+ if (!seccomp_permit_syscall())
return -1;
return ret;
--- a/arch/mips/kernel/ptrace.c
+++ b/arch/mips/kernel/ptrace.c
@@ -1328,7 +1328,7 @@ asmlinkage long syscall_trace_enter(stru
return -1;
}
- if (secure_computing())
+ if (!seccomp_permit_syscall())
return -1;
if (unlikely(test_thread_flag(TIF_SYSCALL_TRACEPOINT)))
--- a/arch/parisc/kernel/ptrace.c
+++ b/arch/parisc/kernel/ptrace.c
@@ -351,7 +351,7 @@ long do_syscall_trace_enter(struct pt_re
}
/* Do the secure computing check after ptrace. */
- if (secure_computing() == -1)
+ if (!seccomp_permit_syscall())
return -1;
#ifdef CONFIG_HAVE_SYSCALL_TRACEPOINTS
--- a/arch/sh/kernel/ptrace_32.c
+++ b/arch/sh/kernel/ptrace_32.c
@@ -460,7 +460,7 @@ asmlinkage long do_syscall_trace_enter(s
return -1;
}
- if (secure_computing() == -1)
+ if (!seccomp_permit_syscall())
return -1;
if (unlikely(test_thread_flag(TIF_SYSCALL_TRACEPOINT)))
--- a/arch/um/kernel/skas/syscall.c
+++ b/arch/um/kernel/skas/syscall.c
@@ -27,7 +27,7 @@ void handle_syscall(struct uml_pt_regs *
goto out;
/* Do the seccomp check after ptrace; failures should be fast. */
- if (secure_computing() == -1)
+ if (!seccomp_permit_syscall())
goto out;
syscall = UPT_SYSCALL_NR(r);
--- a/arch/x86/entry/vsyscall/vsyscall_64.c
+++ b/arch/x86/entry/vsyscall/vsyscall_64.c
@@ -118,10 +118,10 @@ static bool write_ok_or_segv(unsigned lo
static bool __emulate_vsyscall(struct pt_regs *regs, unsigned long address)
{
- unsigned long caller;
- int vsyscall_nr, syscall_nr, tmp;
+ unsigned long caller, orig_dx;
+ int vsyscall_nr, syscall_nr;
+ bool skip;
long ret;
- unsigned long orig_dx;
/* Confirm that the fault happened in 64-bit user mode */
if (!user_64bit_mode(regs))
@@ -197,16 +197,16 @@ static bool __emulate_vsyscall(struct pt
*/
regs->orig_ax = syscall_nr;
regs->ax = -ENOSYS;
- tmp = secure_computing();
- if ((!tmp && regs->orig_ax != syscall_nr) || regs->ip != address) {
+ skip = !seccomp_permit_syscall();
+ if ((!skip && regs->orig_ax != syscall_nr) || regs->ip != address) {
warn_bad_vsyscall(KERN_DEBUG, regs,
"seccomp tried to change syscall nr or ip");
force_exit_sig(SIGSYS);
return true;
}
regs->orig_ax = -1;
- if (tmp)
- goto do_ret; /* skip requested */
+ if (skip)
+ goto do_ret;
/*
* With a real vsyscall, page faults cause SIGSEGV.
--- a/arch/xtensa/kernel/ptrace.c
+++ b/arch/xtensa/kernel/ptrace.c
@@ -553,8 +553,7 @@ int do_syscall_trace_enter(struct pt_reg
return 0;
}
- if (regs->syscall == NO_SYSCALL ||
- secure_computing() == -1) {
+ if (regs->syscall == NO_SYSCALL || !seccomp_permit_syscall()) {
do_syscall_trace_leave(regs);
return 0;
}
--- a/include/linux/entry-common.h
+++ b/include/linux/entry-common.h
@@ -102,9 +102,8 @@ static __always_inline long syscall_trac
/* Do seccomp after ptrace, to catch any tracer changes. */
if (work & SYSCALL_WORK_SECCOMP) {
- ret = __secure_computing();
- if (ret == -1L)
- return ret;
+ if (!__seccomp_permit_syscall())
+ return -1L;
}
/* Either of the above might have changed the syscall number */
@@ -115,7 +114,7 @@ static __always_inline long syscall_trac
syscall_enter_audit(regs, syscall);
- return ret ? : syscall;
+ return syscall;
}
/**
@@ -138,7 +137,7 @@ static __always_inline long syscall_trac
* It handles the following work items:
*
* 1) syscall_work flag dependent invocations of
- * ptrace_report_syscall_entry(), __secure_computing(), trace_sys_enter()
+ * ptrace_report_syscall_entry(), __seccomp_permit_syscall(), trace_sys_enter()
* 2) Invocation of audit_syscall_entry()
*/
static __always_inline long syscall_enter_from_user_mode_work(struct pt_regs *regs, long syscall)
--- a/include/linux/seccomp.h
+++ b/include/linux/seccomp.h
@@ -22,14 +22,14 @@
#include <linux/atomic.h>
#include <asm/seccomp.h>
-extern int __secure_computing(void);
+extern bool __seccomp_permit_syscall(void);
#ifdef CONFIG_HAVE_ARCH_SECCOMP_FILTER
-static inline int secure_computing(void)
+static __always_inline bool seccomp_permit_syscall(void)
{
if (unlikely(test_syscall_work(SECCOMP)))
- return __secure_computing();
- return 0;
+ return __seccomp_permit_syscall();
+ return true;
}
#else
extern void secure_computing_strict(int this_syscall);
@@ -50,11 +50,11 @@ static inline int seccomp_mode(struct se
struct seccomp_data;
#ifdef CONFIG_HAVE_ARCH_SECCOMP_FILTER
-static inline int secure_computing(void) { return 0; }
+static inline bool seccomp_permit_syscall(void) { return true; }
#else
static inline void secure_computing_strict(int this_syscall) { return; }
#endif
-static inline int __secure_computing(void) { return 0; }
+static inline bool __seccomp_permit_syscall(void) { return true; }
static inline long prctl_get_seccomp(void)
{
--- a/kernel/seccomp.c
+++ b/kernel/seccomp.c
@@ -1100,12 +1100,13 @@ void secure_computing_strict(int this_sy
else
BUG();
}
-int __secure_computing(void)
+
+bool __seccomp_permit_syscall(void)
{
int this_syscall = syscall_get_nr(current, current_pt_regs());
secure_computing_strict(this_syscall);
- return 0;
+ return true;
}
#else
@@ -1256,7 +1257,7 @@ static int seccomp_do_user_notification(
return -1;
}
-static int __seccomp_filter(int this_syscall, const bool recheck_after_trace)
+static bool __seccomp_filter(int this_syscall, const bool recheck_after_trace)
{
u32 filter_ret, action;
struct seccomp_data sd;
@@ -1294,7 +1295,7 @@ static int __seccomp_filter(int this_sys
case SECCOMP_RET_TRACE:
/* We've been put in this state by the ptracer already. */
if (recheck_after_trace)
- return 0;
+ return true;
/* ENOSYS these calls if there is no tracer attached. */
if (!ptrace_event_enabled(current, PTRACE_EVENT_SECCOMP)) {
@@ -1330,19 +1331,19 @@ static int __seccomp_filter(int this_sys
* a skip would have already been reported.
*/
if (__seccomp_filter(this_syscall, true))
- return -1;
+ return false;
- return 0;
+ return true;
case SECCOMP_RET_USER_NOTIF:
if (seccomp_do_user_notification(this_syscall, match, &sd))
goto skip;
- return 0;
+ return true;
case SECCOMP_RET_LOG:
seccomp_log(this_syscall, 0, action, true);
- return 0;
+ return true;
case SECCOMP_RET_ALLOW:
/*
@@ -1350,7 +1351,7 @@ static int __seccomp_filter(int this_sys
* this action since SECCOMP_RET_ALLOW is the starting
* state in seccomp_run_filters().
*/
- return 0;
+ return true;
case SECCOMP_RET_KILL_THREAD:
case SECCOMP_RET_KILL_PROCESS:
@@ -1367,46 +1368,46 @@ static int __seccomp_filter(int this_sys
} else {
do_exit(SIGSYS);
}
- return -1; /* skip the syscall go directly to signal handling */
+ return false; /* skip the syscall go directly to signal handling */
}
unreachable();
skip:
seccomp_log(this_syscall, 0, action, match ? match->log : false);
- return -1;
+ return false;
}
#else
-static int __seccomp_filter(int this_syscall, const bool recheck_after_trace)
+static bool __seccomp_filter(int this_syscall, const bool recheck_after_trace)
{
BUG();
- return -1;
+ return false;
}
#endif
-int __secure_computing(void)
+bool __seccomp_permit_syscall(void)
{
int mode = current->seccomp.mode;
int this_syscall;
if (IS_ENABLED(CONFIG_CHECKPOINT_RESTORE) &&
unlikely(current->ptrace & PT_SUSPEND_SECCOMP))
- return 0;
+ return true;
this_syscall = syscall_get_nr(current, current_pt_regs());
switch (mode) {
case SECCOMP_MODE_STRICT:
__secure_computing_strict(this_syscall); /* may call do_exit */
- return 0;
+ return true;
case SECCOMP_MODE_FILTER:
return __seccomp_filter(this_syscall, false);
/* Surviving SECCOMP_RET_KILL_* must be proactively impossible. */
case SECCOMP_MODE_DEAD:
WARN_ON_ONCE(1);
do_exit(SIGKILL);
- return -1;
+ return false;
default:
BUG();
}
^ permalink raw reply [flat|nested] 107+ messages in thread
* [patch 12/18] ptrace, treewide: Rename ptrace_report_syscall_entry() to ptrace_report_syscall_permit_entry()
2026-07-07 19:05 [patch 00/18] entry: Consolidate and rework syscall entry handling Thomas Gleixner
` (10 preceding siblings ...)
2026-07-07 19:06 ` [patch 11/18] seccomp, treewide: Rename and convert __secure_computing() to return boolean Thomas Gleixner
@ 2026-07-07 19:06 ` Thomas Gleixner
2026-07-08 15:46 ` Oleg Nesterov
` (5 more replies)
2026-07-07 19:06 ` [patch 13/18] entry: Make trace_syscall_enter() return type bool Thomas Gleixner
` (9 subsequent siblings)
21 siblings, 6 replies; 107+ messages in thread
From: Thomas Gleixner @ 2026-07-07 19:06 UTC (permalink / raw)
To: LKML
Cc: Peter Zijlstra, Arnd Bergmann, Oleg Nesterov, Richard Henderson,
Vineet Gupta, Russell King, Catalin Marinas, Will Deacon, Guo Ren,
Brian Cain, Geert Uytterhoeven, Michal Simek, Thomas Bogendoerfer,
Dinh Nguyen, Helge Deller, Yoshinori Sato, David S. Miller,
Andreas Larsson, Chris Zankel, linux-alpha, linux-snps-arc,
linux-arm-kernel, linux-csky, linux-hexagon, linux-m68k,
linux-mips, linux-openrisc, linux-parisc, linux-sh, sparclinux,
linux-um, linux-arch, Michael Ellerman, Shrikanth Hegde,
linuxppc-dev, Kees Cook, Huacai Chen, loongarch, Paul Walmsley,
Palmer Dabbelt, linux-riscv, Sven Schnelle, linux-s390, x86,
Mark Rutland, Jinjie Ruan, Andy Lutomirski, Richard Weinberger,
Michal Suchánek, Jonathan Corbet, linux-doc
The return value of that function is boolean and tells the caller whether
to permit the syscall processing or not.
Rename the function so the purpose is clear and make the return type bool.
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Cc: Arnd Bergmann <arnd@arndb.de>
Cc: Oleg Nesterov <oleg@redhat.com>
Cc: Richard Henderson <richard.henderson@linaro.org>
Cc: Vineet Gupta <vgupta@kernel.org>
Cc: Russell King <linux@armlinux.org.uk>
Cc: Catalin Marinas <catalin.marinas@arm.com>
Cc: Will Deacon <will@kernel.org>
Cc: Guo Ren <guoren@kernel.org>
Cc: Brian Cain <bcain@kernel.org>
Cc: Geert Uytterhoeven <geert@linux-m68k.org>
Cc: Michal Simek <monstr@monstr.eu>
Cc: Thomas Bogendoerfer <tsbogend@alpha.franken.de>
Cc: Dinh Nguyen <dinguyen@kernel.org>
Cc: Helge Deller <deller@gmx.de>
Cc: Yoshinori Sato <ysato@users.sourceforge.jp>
Cc: "David S. Miller" <davem@davemloft.net>
Cc: Andreas Larsson <andreas@gaisler.com>
Cc: Chris Zankel <chris@zankel.net>
Cc: linux-alpha@vger.kernel.org
Cc: linux-snps-arc@lists.infradead.org
Cc: linux-arm-kernel@lists.infradead.org
Cc: linux-csky@vger.kernel.org
Cc: linux-hexagon@vger.kernel.org
Cc: linux-m68k@lists.linux-m68k.org
Cc: linux-mips@vger.kernel.org
Cc: linux-openrisc@vger.kernel.org
Cc: linux-parisc@vger.kernel.org
Cc: linux-sh@vger.kernel.org
Cc: sparclinux@vger.kernel.org
Cc: linux-um@lists.infradead.org
Cc: linux-arch@vger.kernel.org
---
arch/alpha/kernel/ptrace.c | 2 +-
arch/arc/kernel/ptrace.c | 2 +-
arch/arm/kernel/ptrace.c | 2 +-
arch/arm64/kernel/ptrace.c | 2 +-
arch/csky/kernel/ptrace.c | 2 +-
arch/hexagon/kernel/traps.c | 2 +-
arch/m68k/kernel/ptrace.c | 2 +-
arch/microblaze/kernel/ptrace.c | 2 +-
arch/mips/kernel/ptrace.c | 2 +-
arch/nios2/kernel/ptrace.c | 2 +-
arch/openrisc/kernel/ptrace.c | 2 +-
arch/parisc/kernel/ptrace.c | 10 ++++------
arch/sh/kernel/ptrace_32.c | 2 +-
arch/sparc/kernel/ptrace_32.c | 2 +-
arch/sparc/kernel/ptrace_64.c | 2 +-
arch/um/kernel/ptrace.c | 2 +-
arch/xtensa/kernel/ptrace.c | 2 +-
include/asm-generic/syscall.h | 4 ++--
include/linux/entry-common.h | 25 ++++++++++++-------------
include/linux/ptrace.h | 13 ++++++-------
20 files changed, 40 insertions(+), 44 deletions(-)
--- a/arch/alpha/kernel/ptrace.c
+++ b/arch/alpha/kernel/ptrace.c
@@ -375,7 +375,7 @@ asmlinkage unsigned long syscall_trace_e
struct pt_regs *regs = current_pt_regs();
if (test_thread_flag(TIF_SYSCALL_TRACE) &&
- ptrace_report_syscall_entry(regs)) {
+ !ptrace_report_syscall_permit_entry(regs)) {
syscall_set_nr(current, regs, -1);
if (regs->r19 == 0 && regs->r0 == (unsigned long)-1)
syscall_set_return_value(current, regs, -ENOSYS, 0);
--- a/arch/arc/kernel/ptrace.c
+++ b/arch/arc/kernel/ptrace.c
@@ -342,7 +342,7 @@ long arch_ptrace(struct task_struct *chi
asmlinkage int syscall_trace_enter(struct pt_regs *regs)
{
if (test_thread_flag(TIF_SYSCALL_TRACE))
- if (ptrace_report_syscall_entry(regs))
+ if (!ptrace_report_syscall_permit_entry(regs))
return ULONG_MAX;
#ifdef CONFIG_HAVE_SYSCALL_TRACEPOINTS
--- a/arch/arm/kernel/ptrace.c
+++ b/arch/arm/kernel/ptrace.c
@@ -840,7 +840,7 @@ static void report_syscall(struct pt_reg
if (dir == PTRACE_SYSCALL_EXIT)
ptrace_report_syscall_exit(regs, 0);
- else if (ptrace_report_syscall_entry(regs))
+ else if (!ptrace_report_syscall_permit_entry(regs))
current_thread_info()->abi_syscall = -1;
regs->ARM_ip = ip;
--- a/arch/arm64/kernel/ptrace.c
+++ b/arch/arm64/kernel/ptrace.c
@@ -2379,7 +2379,7 @@ static int report_syscall_entry(struct p
int regno, ret;
saved_reg = ptrace_save_reg(regs, PTRACE_SYSCALL_ENTER, ®no);
- ret = ptrace_report_syscall_entry(regs);
+ ret = !ptrace_report_syscall_permit_entry(regs);
if (ret)
forget_syscall(regs);
regs->regs[regno] = saved_reg;
--- a/arch/csky/kernel/ptrace.c
+++ b/arch/csky/kernel/ptrace.c
@@ -320,7 +320,7 @@ long arch_ptrace(struct task_struct *chi
asmlinkage int syscall_trace_enter(struct pt_regs *regs)
{
if (test_thread_flag(TIF_SYSCALL_TRACE))
- if (ptrace_report_syscall_entry(regs))
+ if (!ptrace_report_syscall_permit_entry(regs))
return -1;
if (!seccomp_permit_syscall())
--- a/arch/hexagon/kernel/traps.c
+++ b/arch/hexagon/kernel/traps.c
@@ -345,7 +345,7 @@ void do_trap0(struct pt_regs *regs)
/* allow strace to catch syscall args */
if (unlikely(test_thread_flag(TIF_SYSCALL_TRACE) &&
- ptrace_report_syscall_entry(regs)))
+ !ptrace_report_syscall_permit_entry(regs)))
return; /* return -ENOSYS somewhere? */
/* Interrupts should be re-enabled for syscall processing */
--- a/arch/m68k/kernel/ptrace.c
+++ b/arch/m68k/kernel/ptrace.c
@@ -279,7 +279,7 @@ asmlinkage int syscall_trace_enter(void)
int ret = 0;
if (test_thread_flag(TIF_SYSCALL_TRACE))
- ret = ptrace_report_syscall_entry(task_pt_regs(current));
+ ret = !ptrace_report_syscall_permit_entry(task_pt_regs(current));
if (!seccomp_permit_syscall())
return -1;
--- a/arch/microblaze/kernel/ptrace.c
+++ b/arch/microblaze/kernel/ptrace.c
@@ -139,7 +139,7 @@ asmlinkage unsigned long do_syscall_trac
secure_computing_strict(regs->r12);
if (test_thread_flag(TIF_SYSCALL_TRACE) &&
- ptrace_report_syscall_entry(regs))
+ !ptrace_report_syscall_permit_entry(regs))
/*
* Tracing decided this syscall should not happen.
* We'll return a bogus call number to get an ENOSYS
--- a/arch/mips/kernel/ptrace.c
+++ b/arch/mips/kernel/ptrace.c
@@ -1324,7 +1324,7 @@ asmlinkage long syscall_trace_enter(stru
user_exit();
if (test_thread_flag(TIF_SYSCALL_TRACE)) {
- if (ptrace_report_syscall_entry(regs))
+ if (!ptrace_report_syscall_permit_entry(regs))
return -1;
}
--- a/arch/nios2/kernel/ptrace.c
+++ b/arch/nios2/kernel/ptrace.c
@@ -133,7 +133,7 @@ asmlinkage int do_syscall_trace_enter(vo
int ret = 0;
if (test_thread_flag(TIF_SYSCALL_TRACE))
- ret = ptrace_report_syscall_entry(task_pt_regs(current));
+ ret = !ptrace_report_syscall_permit_entry(task_pt_regs(current));
return ret;
}
--- a/arch/openrisc/kernel/ptrace.c
+++ b/arch/openrisc/kernel/ptrace.c
@@ -293,7 +293,7 @@ asmlinkage long do_syscall_trace_enter(s
long ret = 0;
if (test_thread_flag(TIF_SYSCALL_TRACE) &&
- ptrace_report_syscall_entry(regs))
+ !ptrace_report_syscall_permit_entry(regs))
/*
* Tracing decided this syscall should not happen.
* We'll return a bogus call number to get an ENOSYS
--- a/arch/parisc/kernel/ptrace.c
+++ b/arch/parisc/kernel/ptrace.c
@@ -326,7 +326,7 @@ long compat_arch_ptrace(struct task_stru
long do_syscall_trace_enter(struct pt_regs *regs)
{
if (test_thread_flag(TIF_SYSCALL_TRACE)) {
- int rc = ptrace_report_syscall_entry(regs);
+ bool permit = ptrace_report_syscall_permit_entry(regs);
/*
* As tracesys_next does not set %r28 to -ENOSYS
@@ -334,12 +334,10 @@ long do_syscall_trace_enter(struct pt_re
*/
regs->gr[28] = -ENOSYS;
- if (rc) {
+ if (!permit) {
/*
- * A nonzero return code from
- * ptrace_report_syscall_entry() tells us
- * to prevent the syscall execution. Skip
- * the syscall call and the syscall restart handling.
+ * Skip the syscall call and the syscall restart
+ * handling.
*
* Note that the tracer may also just change
* regs->gr[20] to an invalid syscall number,
--- a/arch/sh/kernel/ptrace_32.c
+++ b/arch/sh/kernel/ptrace_32.c
@@ -455,7 +455,7 @@ long arch_ptrace(struct task_struct *chi
asmlinkage long do_syscall_trace_enter(struct pt_regs *regs)
{
if (test_thread_flag(TIF_SYSCALL_TRACE) &&
- ptrace_report_syscall_entry(regs)) {
+ !ptrace_report_syscall_permit_entry(regs)) {
regs->regs[0] = -ENOSYS;
return -1;
}
--- a/arch/sparc/kernel/ptrace_32.c
+++ b/arch/sparc/kernel/ptrace_32.c
@@ -441,7 +441,7 @@ asmlinkage int syscall_trace(struct pt_r
if (syscall_exit_p)
ptrace_report_syscall_exit(regs, 0);
else
- ret = ptrace_report_syscall_entry(regs);
+ ret = !ptrace_report_syscall_permit_entry(regs);
}
return ret;
--- a/arch/sparc/kernel/ptrace_64.c
+++ b/arch/sparc/kernel/ptrace_64.c
@@ -1093,7 +1093,7 @@ asmlinkage int syscall_trace_enter(struc
user_exit();
if (test_thread_flag(TIF_SYSCALL_TRACE))
- ret = ptrace_report_syscall_entry(regs);
+ ret = !ptrace_report_syscall_permit_entry(regs);
if (unlikely(test_thread_flag(TIF_SYSCALL_TRACEPOINT)))
trace_sys_enter(regs, regs->u_regs[UREG_G1]);
--- a/arch/um/kernel/ptrace.c
+++ b/arch/um/kernel/ptrace.c
@@ -135,7 +135,7 @@ int syscall_trace_enter(struct pt_regs *
if (!test_thread_flag(TIF_SYSCALL_TRACE))
return 0;
- return ptrace_report_syscall_entry(regs);
+ return !ptrace_report_syscall_permit_entry(regs);
}
void syscall_trace_leave(struct pt_regs *regs)
--- a/arch/xtensa/kernel/ptrace.c
+++ b/arch/xtensa/kernel/ptrace.c
@@ -547,7 +547,7 @@ int do_syscall_trace_enter(struct pt_reg
regs->areg[2] = -ENOSYS;
if (test_thread_flag(TIF_SYSCALL_TRACE) &&
- ptrace_report_syscall_entry(regs)) {
+ !ptrace_report_syscall_permit_entry(regs)) {
regs->areg[2] = -ENOSYS;
regs->syscall = NO_SYSCALL;
return 0;
--- a/include/asm-generic/syscall.h
+++ b/include/asm-generic/syscall.h
@@ -58,8 +58,8 @@ void syscall_set_nr(struct task_struct *
*
* It's only valid to call this when @task is stopped for system
* call exit tracing (due to %SYSCALL_WORK_SYSCALL_TRACE or
- * %SYSCALL_WORK_SYSCALL_AUDIT), after ptrace_report_syscall_entry()
- * returned nonzero to prevent the system call from taking place.
+ * %SYSCALL_WORK_SYSCALL_AUDIT), after ptrace_report_syscall_permit_entry()
+ * returned False to prevent the system call from taking place.
*
* This rolls back the register state in @regs so it's as if the
* system call instruction was a no-op. The registers containing
--- a/include/linux/entry-common.h
+++ b/include/linux/entry-common.h
@@ -38,21 +38,22 @@
SYSCALL_WORK_SYSCALL_EXIT_TRAP)
/**
- * arch_ptrace_report_syscall_entry - Architecture specific ptrace_report_syscall_entry() wrapper
+ * arch_ptrace_report_syscall_permit_entry - Architecture specific wrapper for
+ * ptrace_report_syscall_permit_entry()
* @regs: Pointer to the register state at syscall entry
*
- * Invoked from syscall_trace_enter() to wrap ptrace_report_syscall_entry().
+ * Invoked from syscall_trace_enter() to wrap ptrace_report_syscall_permit_entry().
*
- * This allows architecture specific ptrace_report_syscall_entry()
+ * This allows architecture specific ptrace_report_syscall_permit_entry()
* implementations. If not defined by the architecture this falls back to
- * to ptrace_report_syscall_entry().
+ * to ptrace_report_syscall_permit_entry().
*/
-static __always_inline int arch_ptrace_report_syscall_entry(struct pt_regs *regs);
+static __always_inline bool arch_ptrace_report_syscall_permit_entry(struct pt_regs *regs);
-#ifndef arch_ptrace_report_syscall_entry
-static __always_inline int arch_ptrace_report_syscall_entry(struct pt_regs *regs)
+#ifndef arch_ptrace_report_syscall_permit_entry
+static __always_inline bool arch_ptrace_report_syscall_permit_entry(struct pt_regs *regs)
{
- return ptrace_report_syscall_entry(regs);
+ return ptrace_report_syscall_permit_entry(regs);
}
#endif
@@ -73,8 +74,6 @@ static inline void syscall_enter_audit(s
static __always_inline long syscall_trace_enter(struct pt_regs *regs, unsigned long work,
long syscall)
{
- long ret = 0;
-
/*
* Handle Syscall User Dispatch. This must comes first, since
* the ABI here can be something that doesn't make sense for
@@ -95,8 +94,8 @@ static __always_inline long syscall_trac
/* Handle ptrace */
if (work & (SYSCALL_WORK_SYSCALL_TRACE | SYSCALL_WORK_SYSCALL_EMU)) {
- ret = arch_ptrace_report_syscall_entry(regs);
- if (ret || (work & SYSCALL_WORK_SYSCALL_EMU))
+ if (!arch_ptrace_report_syscall_permit_entry(regs) ||
+ (work & SYSCALL_WORK_SYSCALL_EMU))
return -1L;
}
@@ -137,7 +136,7 @@ static __always_inline long syscall_trac
* It handles the following work items:
*
* 1) syscall_work flag dependent invocations of
- * ptrace_report_syscall_entry(), __seccomp_permit_syscall(), trace_sys_enter()
+ * ptrace_report_syscall_permit_entry(), __seccomp_permit_syscall(), trace_sys_enter()
* 2) Invocation of audit_syscall_entry()
*/
static __always_inline long syscall_enter_from_user_mode_work(struct pt_regs *regs, long syscall)
--- a/include/linux/ptrace.h
+++ b/include/linux/ptrace.h
@@ -405,13 +405,13 @@ extern void sigaction_compat_abi(struct
/*
* ptrace report for syscall entry and exit looks identical.
*/
-static inline int ptrace_report_syscall(unsigned long message)
+static inline bool ptrace_report_syscall(unsigned long message)
{
int ptrace = current->ptrace;
int signr;
if (!(ptrace & PT_PTRACED))
- return 0;
+ return true;
signr = ptrace_notify(SIGTRAP | ((ptrace & PT_TRACESYSGOOD) ? 0x80 : 0),
message);
@@ -424,11 +424,11 @@ static inline int ptrace_report_syscall(
if (signr)
send_sig(signr, current, 1);
- return fatal_signal_pending(current);
+ return !fatal_signal_pending(current);
}
/**
- * ptrace_report_syscall_entry - task is about to attempt a system call
+ * ptrace_report_syscall_permit_entry - task is about to attempt a system call
* @regs: user register state of current task
*
* This will be called if %SYSCALL_WORK_SYSCALL_TRACE or
@@ -438,7 +438,7 @@ static inline int ptrace_report_syscall(
* call number and arguments to be tried. It is safe to block here,
* preventing the system call from beginning.
*
- * Returns zero normally, or nonzero if the calling arch code should abort
+ * Returns True normally, or False if the calling architecture code should abort
* the system call. That must prevent normal entry so no system call is
* made. If @task ever returns to user mode after this, its register state
* is unspecified, but should be something harmless like an %ENOSYS error
@@ -447,8 +447,7 @@ static inline int ptrace_report_syscall(
*
* Called without locks, just after entering kernel mode.
*/
-static inline __must_check int ptrace_report_syscall_entry(
- struct pt_regs *regs)
+static inline __must_check bool ptrace_report_syscall_permit_entry(struct pt_regs *regs)
{
return ptrace_report_syscall(PTRACE_EVENTMSG_SYSCALL_ENTRY);
}
^ permalink raw reply [flat|nested] 107+ messages in thread
* [patch 13/18] entry: Make trace_syscall_enter() return type bool
2026-07-07 19:05 [patch 00/18] entry: Consolidate and rework syscall entry handling Thomas Gleixner
` (11 preceding siblings ...)
2026-07-07 19:06 ` [patch 12/18] ptrace, treewide: Rename ptrace_report_syscall_entry() to ptrace_report_syscall_permit_entry() Thomas Gleixner
@ 2026-07-07 19:06 ` Thomas Gleixner
2026-07-08 15:52 ` Michal Suchánek
2026-07-07 19:06 ` [patch 14/18] entry: Make return type of syscall_trace_enter() bool Thomas Gleixner
` (8 subsequent siblings)
21 siblings, 1 reply; 107+ messages in thread
From: Thomas Gleixner @ 2026-07-07 19:06 UTC (permalink / raw)
To: LKML
Cc: Peter Zijlstra, Michael Ellerman, Shrikanth Hegde, linuxppc-dev,
Kees Cook, Huacai Chen, loongarch, Paul Walmsley, Palmer Dabbelt,
linux-riscv, Sven Schnelle, linux-s390, x86, Mark Rutland,
Jinjie Ruan, Andy Lutomirski, Oleg Nesterov, Richard Henderson,
Russell King, Catalin Marinas, Guo Ren, Geert Uytterhoeven,
Thomas Bogendoerfer, Helge Deller, Yoshinori Sato,
Richard Weinberger, Chris Zankel, linux-arm-kernel, linux-alpha,
linux-csky, linux-m68k, linux-mips, linux-parisc, linux-sh,
linux-um, Arnd Bergmann, Vineet Gupta, Will Deacon, Brian Cain,
Michal Simek, Dinh Nguyen, David S. Miller, Andreas Larsson,
linux-snps-arc, linux-hexagon, linux-openrisc, sparclinux,
linux-arch, Michal Suchánek, Jonathan Corbet, linux-doc
In preparation of converting the return value of
syscall_enter_from_user_mode[_work]() bool, rework trace_syscall_enter() to
- update the syscall number via a pointer argument
- Return True if the syscall number is != -1, False otherwise
That aligns with ptrace_report_syscall_permit_enter() and
seccomp_permit_syscall().
The only difference is that this also returns False, when the syscall
number was already -1 to begin with, but there is not much which can be
done about that. As the architecture has to preset the return value to
-ENOSYS anyway, that results in the correct return value for such an
invalid syscall.
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
---
include/linux/entry-common.h | 8 +++++---
kernel/entry/syscall-common.c | 7 ++++---
2 files changed, 9 insertions(+), 6 deletions(-)
--- a/include/linux/entry-common.h
+++ b/include/linux/entry-common.h
@@ -58,7 +58,7 @@ static __always_inline bool arch_ptrace_
#endif
bool syscall_user_dispatch(struct pt_regs *regs);
-long trace_syscall_enter(struct pt_regs *regs, long syscall);
+bool trace_syscall_enter(struct pt_regs *regs, long *syscall);
void trace_syscall_exit(struct pt_regs *regs, long ret);
static inline void syscall_enter_audit(struct pt_regs *regs, long syscall)
@@ -108,8 +108,10 @@ static __always_inline long syscall_trac
/* Either of the above might have changed the syscall number */
syscall = syscall_get_nr(current, regs);
- if (unlikely(work & SYSCALL_WORK_SYSCALL_TRACEPOINT))
- syscall = trace_syscall_enter(regs, syscall);
+ if (unlikely(work & SYSCALL_WORK_SYSCALL_TRACEPOINT)) {
+ if (!trace_syscall_enter(regs, &syscall))
+ return -1L;
+ }
syscall_enter_audit(regs, syscall);
--- a/kernel/entry/syscall-common.c
+++ b/kernel/entry/syscall-common.c
@@ -7,14 +7,15 @@
/* Out of line to prevent tracepoint code duplication */
-long trace_syscall_enter(struct pt_regs *regs, long syscall)
+bool trace_syscall_enter(struct pt_regs *regs, long *syscall)
{
- trace_sys_enter(regs, syscall);
+ trace_sys_enter(regs, *syscall);
/*
* Probes or BPF hooks in the tracepoint may have changed the
* system call number. Reread it.
*/
- return syscall_get_nr(current, regs);
+ *syscall = syscall_get_nr(current, regs);
+ return *syscall != -1L;
}
void trace_syscall_exit(struct pt_regs *regs, long ret)
^ permalink raw reply [flat|nested] 107+ messages in thread
* [patch 14/18] entry: Make return type of syscall_trace_enter() bool
2026-07-07 19:05 [patch 00/18] entry: Consolidate and rework syscall entry handling Thomas Gleixner
` (12 preceding siblings ...)
2026-07-07 19:06 ` [patch 13/18] entry: Make trace_syscall_enter() return type bool Thomas Gleixner
@ 2026-07-07 19:06 ` Thomas Gleixner
2026-07-09 19:36 ` Mukesh Kumar Chaurasiya
2026-07-07 19:06 ` [patch 15/18] x86/entry: Make syscall functions static Thomas Gleixner
` (7 subsequent siblings)
21 siblings, 1 reply; 107+ messages in thread
From: Thomas Gleixner @ 2026-07-07 19:06 UTC (permalink / raw)
To: LKML
Cc: Peter Zijlstra, Michael Ellerman, Shrikanth Hegde, linuxppc-dev,
Kees Cook, Huacai Chen, loongarch, Paul Walmsley, Palmer Dabbelt,
linux-riscv, Sven Schnelle, linux-s390, x86, Mark Rutland,
Jinjie Ruan, Andy Lutomirski, Oleg Nesterov, Richard Henderson,
Russell King, Catalin Marinas, Guo Ren, Geert Uytterhoeven,
Thomas Bogendoerfer, Helge Deller, Yoshinori Sato,
Richard Weinberger, Chris Zankel, linux-arm-kernel, linux-alpha,
linux-csky, linux-m68k, linux-mips, linux-parisc, linux-sh,
linux-um, Arnd Bergmann, Vineet Gupta, Will Deacon, Brian Cain,
Michal Simek, Dinh Nguyen, David S. Miller, Andreas Larsson,
linux-snps-arc, linux-hexagon, linux-openrisc, sparclinux,
linux-arch, Michal Suchánek, Jonathan Corbet, linux-doc
This prepares for changing the return types of
syscall_enter_from_user_mode[_work]() to bool, which in turn separates the
decision of invoking the syscall from the syscall number, which might have
been changed in the call by ptrace, seccomp, tracing.
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
---
include/linux/entry-common.h | 28 +++++++++++++++-------------
1 file changed, 15 insertions(+), 13 deletions(-)
--- a/include/linux/entry-common.h
+++ b/include/linux/entry-common.h
@@ -71,8 +71,8 @@ static inline void syscall_enter_audit(s
}
}
-static __always_inline long syscall_trace_enter(struct pt_regs *regs, unsigned long work,
- long syscall)
+static __always_inline bool syscall_trace_enter(struct pt_regs *regs, unsigned long work,
+ long *syscall)
{
/*
* Handle Syscall User Dispatch. This must comes first, since
@@ -81,7 +81,7 @@ static __always_inline long syscall_trac
*/
if (work & SYSCALL_WORK_SYSCALL_USER_DISPATCH) {
if (syscall_user_dispatch(regs))
- return -1L;
+ return false;
}
/*
@@ -90,32 +90,32 @@ static __always_inline long syscall_trac
* through hrtimer_interrupt().
*/
if (work & SYSCALL_WORK_SYSCALL_RSEQ_SLICE)
- rseq_syscall_enter_work(syscall);
+ rseq_syscall_enter_work(*syscall);
/* Handle ptrace */
if (work & (SYSCALL_WORK_SYSCALL_TRACE | SYSCALL_WORK_SYSCALL_EMU)) {
if (!arch_ptrace_report_syscall_permit_entry(regs) ||
(work & SYSCALL_WORK_SYSCALL_EMU))
- return -1L;
+ return false;
}
/* Do seccomp after ptrace, to catch any tracer changes. */
if (work & SYSCALL_WORK_SECCOMP) {
if (!__seccomp_permit_syscall())
- return -1L;
+ return false;
}
/* Either of the above might have changed the syscall number */
- syscall = syscall_get_nr(current, regs);
+ *syscall = syscall_get_nr(current, regs);
if (unlikely(work & SYSCALL_WORK_SYSCALL_TRACEPOINT)) {
- if (!trace_syscall_enter(regs, &syscall))
- return -1L;
+ if (!trace_syscall_enter(regs, syscall))
+ return false;
}
- syscall_enter_audit(regs, syscall);
+ syscall_enter_audit(regs, *syscall);
- return syscall;
+ return true;
}
/**
@@ -145,8 +145,10 @@ static __always_inline long syscall_ente
{
unsigned long work = READ_ONCE(current_thread_info()->syscall_work);
- if (work & SYSCALL_WORK_ENTER)
- syscall = syscall_trace_enter(regs, work, syscall);
+ if (work & SYSCALL_WORK_ENTER) {
+ if (!syscall_trace_enter(regs, work, &syscall))
+ return -1L;
+ }
return syscall;
}
^ permalink raw reply [flat|nested] 107+ messages in thread
* [patch 15/18] x86/entry: Make syscall functions static
2026-07-07 19:05 [patch 00/18] entry: Consolidate and rework syscall entry handling Thomas Gleixner
` (13 preceding siblings ...)
2026-07-07 19:06 ` [patch 14/18] entry: Make return type of syscall_trace_enter() bool Thomas Gleixner
@ 2026-07-07 19:06 ` Thomas Gleixner
2026-07-09 1:47 ` Jinjie Ruan
2026-07-09 19:43 ` Mukesh Kumar Chaurasiya
2026-07-07 19:07 ` [patch 16/18] x86/entry: Get rid of the sys_ni_syscall() indirection Thomas Gleixner
` (6 subsequent siblings)
21 siblings, 2 replies; 107+ messages in thread
From: Thomas Gleixner @ 2026-07-07 19:06 UTC (permalink / raw)
To: LKML
Cc: Peter Zijlstra, Michael Ellerman, Shrikanth Hegde, linuxppc-dev,
Kees Cook, Huacai Chen, loongarch, Paul Walmsley, Palmer Dabbelt,
linux-riscv, Sven Schnelle, linux-s390, x86, Mark Rutland,
Jinjie Ruan, Andy Lutomirski, Oleg Nesterov, Richard Henderson,
Russell King, Catalin Marinas, Guo Ren, Geert Uytterhoeven,
Thomas Bogendoerfer, Helge Deller, Yoshinori Sato,
Richard Weinberger, Chris Zankel, linux-arm-kernel, linux-alpha,
linux-csky, linux-m68k, linux-mips, linux-parisc, linux-sh,
linux-um, Arnd Bergmann, Vineet Gupta, Will Deacon, Brian Cain,
Michal Simek, Dinh Nguyen, David S. Miller, Andreas Larsson,
linux-snps-arc, linux-hexagon, linux-openrisc, sparclinux,
linux-arch, Michal Suchánek, Jonathan Corbet, linux-doc
They are only used in the respective source files. No point in exposing
them.
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
---
arch/x86/entry/syscall_32.c | 2 +-
arch/x86/entry/syscall_64.c | 10 ++++++----
arch/x86/include/asm/syscall.h | 8 --------
3 files changed, 7 insertions(+), 13 deletions(-)
--- a/arch/x86/entry/syscall_32.c
+++ b/arch/x86/entry/syscall_32.c
@@ -41,7 +41,7 @@ const sys_call_ptr_t sys_call_table[] =
#endif
#define __SYSCALL(nr, sym) case nr: return __ia32_##sym(regs);
-long ia32_sys_call(const struct pt_regs *regs, unsigned int nr)
+static noinline long ia32_sys_call(const struct pt_regs *regs, unsigned int nr)
{
switch (nr) {
#include <asm/syscalls_32.h>
--- a/arch/x86/entry/syscall_64.c
+++ b/arch/x86/entry/syscall_64.c
@@ -32,7 +32,7 @@ const sys_call_ptr_t sys_call_table[] =
#undef __SYSCALL
#define __SYSCALL(nr, sym) case nr: return __x64_##sym(regs);
-long x64_sys_call(const struct pt_regs *regs, unsigned int nr)
+static noinline long x64_sys_call(const struct pt_regs *regs, unsigned int nr)
{
switch (nr) {
#include <asm/syscalls_64.h>
@@ -40,15 +40,17 @@ long x64_sys_call(const struct pt_regs *
}
}
-#ifdef CONFIG_X86_X32_ABI
-long x32_sys_call(const struct pt_regs *regs, unsigned int nr)
+static noinline long x32_sys_call(const struct pt_regs *regs, unsigned int nr)
{
+#ifdef CONFIG_X86_X32_ABI
switch (nr) {
#include <asm/syscalls_x32.h>
default: return __x64_sys_ni_syscall(regs);
}
-}
+#else
+ return -ENOSYS;
#endif
+}
static __always_inline bool do_syscall_x64(struct pt_regs *regs, int nr)
{
--- a/arch/x86/include/asm/syscall.h
+++ b/arch/x86/include/asm/syscall.h
@@ -21,14 +21,6 @@ typedef long (*sys_call_ptr_t)(const str
extern const sys_call_ptr_t sys_call_table[];
/*
- * These may not exist, but still put the prototypes in so we
- * can use IS_ENABLED().
- */
-extern long ia32_sys_call(const struct pt_regs *, unsigned int nr);
-extern long x32_sys_call(const struct pt_regs *, unsigned int nr);
-extern long x64_sys_call(const struct pt_regs *, unsigned int nr);
-
-/*
* Only the low 32 bits of orig_ax are meaningful, so we return int.
* This importantly ignores the high bits on 64-bit, so comparisons
* sign-extend the low 32 bits.
^ permalink raw reply [flat|nested] 107+ messages in thread
* [patch 16/18] x86/entry: Get rid of the sys_ni_syscall() indirection
2026-07-07 19:05 [patch 00/18] entry: Consolidate and rework syscall entry handling Thomas Gleixner
` (14 preceding siblings ...)
2026-07-07 19:06 ` [patch 15/18] x86/entry: Make syscall functions static Thomas Gleixner
@ 2026-07-07 19:07 ` Thomas Gleixner
2026-07-09 2:03 ` Jinjie Ruan
2026-07-07 19:07 ` [patch 17/18] x86/entry: Simplify the syscall number logic Thomas Gleixner
` (5 subsequent siblings)
21 siblings, 1 reply; 107+ messages in thread
From: Thomas Gleixner @ 2026-07-07 19:07 UTC (permalink / raw)
To: LKML
Cc: Peter Zijlstra, Michael Ellerman, Shrikanth Hegde, linuxppc-dev,
Kees Cook, Huacai Chen, loongarch, Paul Walmsley, Palmer Dabbelt,
linux-riscv, Sven Schnelle, linux-s390, x86, Mark Rutland,
Jinjie Ruan, Andy Lutomirski, Oleg Nesterov, Richard Henderson,
Russell King, Catalin Marinas, Guo Ren, Geert Uytterhoeven,
Thomas Bogendoerfer, Helge Deller, Yoshinori Sato,
Richard Weinberger, Chris Zankel, linux-arm-kernel, linux-alpha,
linux-csky, linux-m68k, linux-mips, linux-parisc, linux-sh,
linux-um, Arnd Bergmann, Vineet Gupta, Will Deacon, Brian Cain,
Michal Simek, Dinh Nguyen, David S. Miller, Andreas Larsson,
linux-snps-arc, linux-hexagon, linux-openrisc, sparclinux,
linux-arch, Michal Suchánek, Jonathan Corbet, linux-doc
Invoking sys_ni_syscall() from a code path, which already knows that the
syscall number is invalid just to assign -ENOSYS to regs->ax is a pointless
exercise. It's even redundant as the low level entry code already has set
regs->ax to -ENOSYS on entry.
Remove the extra conditionals and the function calls.
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
---
arch/x86/entry/syscall_32.c | 2 --
arch/x86/entry/syscall_64.c | 10 +++-------
2 files changed, 3 insertions(+), 9 deletions(-)
--- a/arch/x86/entry/syscall_32.c
+++ b/arch/x86/entry/syscall_32.c
@@ -81,8 +81,6 @@ static __always_inline void do_syscall_3
if (likely(unr < IA32_NR_syscalls)) {
unr = array_index_nospec(unr, IA32_NR_syscalls);
regs->ax = ia32_sys_call(regs, unr);
- } else if (nr != -1) {
- regs->ax = __ia32_sys_ni_syscall(regs);
}
}
--- a/arch/x86/entry/syscall_64.c
+++ b/arch/x86/entry/syscall_64.c
@@ -68,7 +68,7 @@ static __always_inline bool do_syscall_x
return false;
}
-static __always_inline bool do_syscall_x32(struct pt_regs *regs, int nr)
+static __always_inline void do_syscall_x32(struct pt_regs *regs, int nr)
{
/*
* Adjust the starting offset of the table, and convert numbers
@@ -80,9 +80,7 @@ static __always_inline bool do_syscall_x
if (IS_ENABLED(CONFIG_X86_X32_ABI) && likely(xnr < X32_NR_syscalls)) {
xnr = array_index_nospec(xnr, X32_NR_syscalls);
regs->ax = x32_sys_call(regs, xnr);
- return true;
}
- return false;
}
/* Returns true to return using SYSRET, or false to use IRET */
@@ -92,10 +90,8 @@ static __always_inline bool do_syscall_x
instrumentation_begin();
- if (!do_syscall_x64(regs, nr) && !do_syscall_x32(regs, nr) && nr != -1) {
- /* Invalid system call, but still a system call. */
- regs->ax = __x64_sys_ni_syscall(regs);
- }
+ if (!do_syscall_x64(regs, nr))
+ do_syscall_x32(regs, nr);
instrumentation_end();
syscall_exit_to_user_mode(regs);
^ permalink raw reply [flat|nested] 107+ messages in thread
* [patch 17/18] x86/entry: Simplify the syscall number logic
2026-07-07 19:05 [patch 00/18] entry: Consolidate and rework syscall entry handling Thomas Gleixner
` (15 preceding siblings ...)
2026-07-07 19:07 ` [patch 16/18] x86/entry: Get rid of the sys_ni_syscall() indirection Thomas Gleixner
@ 2026-07-07 19:07 ` Thomas Gleixner
2026-07-07 19:07 ` [patch 18/18] entry, treewide: Make syscall_enter_from_user_mode[_work]() indicate syscall execution Thomas Gleixner
` (4 subsequent siblings)
21 siblings, 0 replies; 107+ messages in thread
From: Thomas Gleixner @ 2026-07-07 19:07 UTC (permalink / raw)
To: LKML
Cc: Peter Zijlstra, Michael Ellerman, Shrikanth Hegde, linuxppc-dev,
Kees Cook, Huacai Chen, loongarch, Paul Walmsley, Palmer Dabbelt,
linux-riscv, Sven Schnelle, linux-s390, x86, Mark Rutland,
Jinjie Ruan, Andy Lutomirski, Oleg Nesterov, Richard Henderson,
Russell King, Catalin Marinas, Guo Ren, Geert Uytterhoeven,
Thomas Bogendoerfer, Helge Deller, Yoshinori Sato,
Richard Weinberger, Chris Zankel, linux-arm-kernel, linux-alpha,
linux-csky, linux-m68k, linux-mips, linux-parisc, linux-sh,
linux-um, Arnd Bergmann, Vineet Gupta, Will Deacon, Brian Cain,
Michal Simek, Dinh Nguyen, David S. Miller, Andreas Larsson,
linux-snps-arc, linux-hexagon, linux-openrisc, sparclinux,
linux-arch, Michal Suchánek, Jonathan Corbet, linux-doc
Converting from int to long, back to int and then to unsigned int is
confusing at best.
None of this voodoo is required. Negative syscall numbers including -1
don't need any of this treat