* [PATCH bpf v3 0/2] bpf, x86: fix per-CPU address resolution into an extended register @ 2026-08-14 22:02 Vineet Gupta 2026-08-14 22:02 ` [PATCH bpf v3 1/2] bpf, x86: Fix " Vineet Gupta 2026-08-14 22:02 ` [PATCH bpf v3 2/2] selftests/bpf: Check per-CPU address resolution per register Vineet Gupta 0 siblings, 2 replies; 8+ messages in thread From: Vineet Gupta @ 2026-08-14 22:02 UTC (permalink / raw) To: bpf; +Cc: ast, daniel, andrii, x86, stable, Vineet Gupta The JIT resolves a per-CPU address with add <dst>, gs:[this_cpu_off] but builds the REX prefix with add_1mod(), which sets REX.B. The destination is encoded in ModRM.reg, which REX.R extends, and the memory operand is disp32 with no base, so REX.B does nothing and the high register bit is dropped. Every extended destination therefore resolves into whichever register shares the low three bits: R5 -> RAX R7 -> RBP R8 -> RSI R9 -> RDI The address is left unadjusted and an unrelated register is clobbered. Patch 1 switches to add_2mod() so the bit goes through REX.R. Clang reloads the address into R1 before each per-CPU access, so the destination is never an extended register and the bug has been dormant since v6.10. GCC keeps several per-CPU addresses live at once, which is how it turned up: test_progs-bpf_gcc panics the kernel in global_percpu_data/init, with the address of a .percpu variable in R5. Patch 2 covers every register. A functional test only catches this if the address happens to land in an extended register, so the test matches the JITed add instead. Changes in v3: - Fold the five per-register programs into one that loads every register, and drop the comment explaining the register choice (Eduard Zingerman). - Move the percpu_data declaration inside the arch guard, so other targets no longer carry a .percpu section and an unused map (bpf-ci). - Match the movabsq of each address as well as the add, so the matchers stay on consecutive lines and the pair is checked to use the same register. - Restore the Reviewed-by on patch 1, dropped by mistake in v2. Changes in v2: - Add the selftest, patch 2/2 (Eduard Zingerman). It uses __jited() rather than __xlated(): the xlated stream is identical for every register, and the wrong prefix is only visible in the native encoding. - No functional change to patch 1. Vineet Gupta (2): bpf, x86: Fix per-CPU address resolution into an extended register selftests/bpf: Check per-CPU address resolution per register arch/x86/net/bpf_jit_comp.c | 2 +- .../selftests/bpf/prog_tests/verifier.c | 2 + .../bpf/progs/verifier_percpu_addr.c | 72 +++++++++++++++++++ 3 files changed, 75 insertions(+), 1 deletion(-) create mode 100644 tools/testing/selftests/bpf/progs/verifier_percpu_addr.c -- 2.53.0-Meta ^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH bpf v3 1/2] bpf, x86: Fix per-CPU address resolution into an extended register 2026-08-14 22:02 [PATCH bpf v3 0/2] bpf, x86: fix per-CPU address resolution into an extended register Vineet Gupta @ 2026-08-14 22:02 ` Vineet Gupta 2026-08-14 22:02 ` [PATCH bpf v3 2/2] selftests/bpf: Check per-CPU address resolution per register Vineet Gupta 1 sibling, 0 replies; 8+ messages in thread From: Vineet Gupta @ 2026-08-14 22:02 UTC (permalink / raw) To: bpf; +Cc: ast, daniel, andrii, x86, stable, Vineet Gupta, Eduard Zingerman The destination of the per-CPU address MOV is encoded in ModRM.reg, which is extended by REX.R, but the REX prefix is built with add_1mod(), which sets REX.B. REX.B extends ModRM.rm and SIB.base, and this instruction addresses memory as disp32 with no base, so the bit has no effect at all and the high register bit is simply lost. Every is_ereg() destination therefore resolves to the wrong register, picking whichever one shares the low three bits: R5 -> RAX R7 -> RBP R8 -> RSI R9 -> RDI With BPF_REG_5, whose reg2hex is 0, the emitted 65 49 03 04 25 <off> add %gs:<off>,%rax adds the per-CPU offset to RAX rather than R8. The destination keeps the unadjusted address and RAX is clobbered, so the program goes on to dereference a pointer that was never made per-CPU: BUG: unable to handle page fault for address: 0000607e386a8894 RIP: bpf_prog_707837aafd2aa9ae_update_percpu_data+0x93/0xc9 Call Trace: __bpf_prog_test_run_raw_tp+0x2dc/0x7d0 __flush_smp_call_function_queue+0x1e9/0xc80 Kernel panic - not syncing: Fatal exception in interrupt R5 is the mildest of the four, aliasing a scratch register and faulting at the store. R7 aliases RBP and would corrupt the frame pointer, R8 and R9 alias the argument registers. Use add_2mod() so the register goes through REX.R, matching how add_2reg() places it in ModRM.reg and how emit_priv_frame_ptr() hardcodes 0x4c for the same instruction with R9. Encodings for the non-extended registers are unchanged. Problem showed up when trying to resurrect BPF_GCC CI (selftests built with BPF_GCC). This has gone unnoticed because clang reloads the address into R1 before each per-CPU access, so the destination is never an extended register. GCC keeps several per-CPU addresses live at once, and test_progs-bpf_gcc panics the kernel in global_percpu_data/init, where the address of a .percpu variable ends up in R5. Fixes: 7bdbf7446305 ("bpf: add special internal-only MOV instruction to resolve per-CPU addrs") Cc: stable@vger.kernel.org Reviewed-by: Eduard Zingerman <eddyz87@gmail.com> Signed-off-by: Vineet Gupta <vineet.gupta@linux.dev> --- arch/x86/net/bpf_jit_comp.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c index d920772af7d5..1a9fb530adc3 100644 --- a/arch/x86/net/bpf_jit_comp.c +++ b/arch/x86/net/bpf_jit_comp.c @@ -1935,7 +1935,7 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int * EMIT_mov(dst_reg, src_reg); #ifdef CONFIG_SMP /* add <dst>, gs:[<off>] */ - EMIT2(0x65, add_1mod(0x48, dst_reg)); + EMIT2(0x65, add_2mod(0x48, 0, dst_reg)); EMIT3(0x03, add_2reg(0x04, 0, dst_reg), 0x25); EMIT((u32)(unsigned long)&this_cpu_off, 4); #endif -- 2.53.0-Meta ^ permalink raw reply related [flat|nested] 8+ messages in thread
* [PATCH bpf v3 2/2] selftests/bpf: Check per-CPU address resolution per register 2026-08-14 22:02 [PATCH bpf v3 0/2] bpf, x86: fix per-CPU address resolution into an extended register Vineet Gupta 2026-08-14 22:02 ` [PATCH bpf v3 1/2] bpf, x86: Fix " Vineet Gupta @ 2026-08-14 22:02 ` Vineet Gupta 2026-08-14 22:26 ` sashiko-bot 2026-08-14 22:39 ` bot+bpf-ci 1 sibling, 2 replies; 8+ messages in thread From: Vineet Gupta @ 2026-08-14 22:02 UTC (permalink / raw) To: bpf; +Cc: ast, daniel, andrii, x86, stable, Vineet Gupta An ld_imm64 of a per-CPU map value is followed by a mov_percpu_addr that reuses the same register, so which register the address lands in decides how the JIT encodes the add. Getting the REX prefix wrong there is invisible to a functional test unless the address happens to land in an extended register, which is why this went unnoticed. Load a .percpu variable into every register in one program and match the JITed add against the register each one must resolve into. Signed-off-by: Vineet Gupta <vineet.gupta@linux.dev> --- .../selftests/bpf/prog_tests/verifier.c | 2 + .../bpf/progs/verifier_percpu_addr.c | 72 +++++++++++++++++++ 2 files changed, 74 insertions(+) create mode 100644 tools/testing/selftests/bpf/progs/verifier_percpu_addr.c diff --git a/tools/testing/selftests/bpf/prog_tests/verifier.c b/tools/testing/selftests/bpf/prog_tests/verifier.c index 8113fea7ba86..64ac49ad67e6 100644 --- a/tools/testing/selftests/bpf/prog_tests/verifier.c +++ b/tools/testing/selftests/bpf/prog_tests/verifier.c @@ -79,6 +79,7 @@ #include "verifier_netfilter_retcode.skel.h" #include "verifier_bpf_fastcall.skel.h" #include "verifier_or_jmp32_k.skel.h" +#include "verifier_percpu_addr.skel.h" #include "verifier_precision.skel.h" #include "verifier_prevent_map_lookup.skel.h" #include "verifier_private_stack.skel.h" @@ -240,6 +241,7 @@ void test_verifier_netfilter_ctx(void) { RUN(verifier_netfilter_ctx); } void test_verifier_netfilter_retcode(void) { RUN(verifier_netfilter_retcode); } void test_verifier_bpf_fastcall(void) { RUN(verifier_bpf_fastcall); } void test_verifier_or_jmp32_k(void) { RUN(verifier_or_jmp32_k); } +void test_verifier_percpu_addr(void) { RUN(verifier_percpu_addr); } void test_verifier_precision(void) { RUN(verifier_precision); } void test_verifier_prevent_map_lookup(void) { RUN(verifier_prevent_map_lookup); } void test_verifier_private_stack(void) { RUN(verifier_private_stack); } diff --git a/tools/testing/selftests/bpf/progs/verifier_percpu_addr.c b/tools/testing/selftests/bpf/progs/verifier_percpu_addr.c new file mode 100644 index 000000000000..967f4e6e3a49 --- /dev/null +++ b/tools/testing/selftests/bpf/progs/verifier_percpu_addr.c @@ -0,0 +1,72 @@ +// SPDX-License-Identifier: GPL-2.0 + +#include <vmlinux.h> +#include <bpf/bpf_helpers.h> +#include "bpf_misc.h" + +#if defined(__TARGET_ARCH_x86) + +int percpu_data SEC(".percpu"); + +/* + * An ld_imm64 of a per-CPU map value is followed by a mov_percpu_addr that + * reuses the same register, so check that the add resolves into the register + * the address was loaded into, for every register. + */ +SEC("raw_tp") +__description("per-CPU address resolution") +__success +__arch_x86_64 +__jited(" movabsq $0x{{.*}}, %rax") +__jited(" addq %gs:{{.*}}, %rax") +__jited(" movabsq $0x{{.*}}, %rdi") +__jited(" addq %gs:{{.*}}, %rdi") +__jited(" movabsq $0x{{.*}}, %rsi") +__jited(" addq %gs:{{.*}}, %rsi") +__jited(" movabsq $0x{{.*}}, %rdx") +__jited(" addq %gs:{{.*}}, %rdx") +__jited(" movabsq $0x{{.*}}, %rcx") +__jited(" addq %gs:{{.*}}, %rcx") +__jited(" movabsq $0x{{.*}}, %r8") +__jited(" addq %gs:{{.*}}, %r8") +__jited(" movabsq $0x{{.*}}, %rbx") +__jited(" addq %gs:{{.*}}, %rbx") +__jited(" movabsq $0x{{.*}}, %r13") +__jited(" addq %gs:{{.*}}, %r13") +__jited(" movabsq $0x{{.*}}, %r14") +__jited(" addq %gs:{{.*}}, %r14") +__jited(" movabsq $0x{{.*}}, %r15") +__jited(" addq %gs:{{.*}}, %r15") +__naked void percpu_addr(void) +{ + asm volatile (" \ + r0 = %[percpu_data] ll; \ + r1 = %[percpu_data] ll; \ + r2 = %[percpu_data] ll; \ + r3 = %[percpu_data] ll; \ + r4 = %[percpu_data] ll; \ + r5 = %[percpu_data] ll; \ + r6 = %[percpu_data] ll; \ + r7 = %[percpu_data] ll; \ + r8 = %[percpu_data] ll; \ + r9 = %[percpu_data] ll; \ + r0 = 0; \ + exit; \ +" : + : __imm_addr(percpu_data) + : __clobber_all); +} + +#else + +SEC("raw_tp") +__description("percpu addr dummy") +__success +int dummy_test(void) +{ + return 0; +} + +#endif + +char _license[] SEC("license") = "GPL"; -- 2.53.0-Meta ^ permalink raw reply related [flat|nested] 8+ messages in thread
* Re: [PATCH bpf v3 2/2] selftests/bpf: Check per-CPU address resolution per register 2026-08-14 22:02 ` [PATCH bpf v3 2/2] selftests/bpf: Check per-CPU address resolution per register Vineet Gupta @ 2026-08-14 22:26 ` sashiko-bot 2026-08-14 22:39 ` Vineet Gupta 2026-08-14 22:39 ` bot+bpf-ci 1 sibling, 1 reply; 8+ messages in thread From: sashiko-bot @ 2026-08-14 22:26 UTC (permalink / raw) To: Vineet Gupta; +Cc: bpf Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider: - [High] The test uses an unrecognized section name `SEC(".percpu")` for its global variable, preventing libbpf from loading the test. - [High] The test's inline assembly is missing the `mov_percpu_addr` instruction, causing the `__jited` assertions to always fail. -- --- Patch [2]: [PATCH bpf v3 2/2] selftests/bpf: Check per-CPU address resolution per register --- commit 94ba39b00c54962b463da6e18c62f89978f29a3f Author: Vineet Gupta <vineet.gupta@linux.dev> selftests/bpf: Check per-CPU address resolution per register This commit adds a BPF selftest to verify that the x86 JIT correctly encodes the per-CPU address resolution. It loads a .percpu variable into all registers and matches the JITed add instruction against the expected register. > diff --git a/tools/testing/selftests/bpf/progs/verifier_percpu_addr.c b/tools/testing/selftests/bpf/progs/verifier_percpu_addr.c > new file mode 100644 > index 0000000000000..967f4e6e3a490 > --- /dev/null > +++ b/tools/testing/selftests/bpf/progs/verifier_percpu_addr.c > @@ -0,0 +1,72 @@ > +// SPDX-License-Identifier: GPL-2.0 > + > +#include <vmlinux.h> > +#include <bpf/bpf_helpers.h> > +#include "bpf_misc.h" > + > +#if defined(__TARGET_ARCH_x86) > + > +int percpu_data SEC(".percpu"); [Severity: High] Does libbpf support the SEC(".percpu") section name? When bpf_object__elf_collect runs, it skips the .percpu section as it doesn't match known data section prefixes like .data. Later, when bpf_object__relocate_data processes the R_BPF_64_64 relocation for the percpu_data load, bpf_object__section_to_libbpf_map_type returns LIBBPF_MAP_UNSPEC. Since the section is not the BTF maps section, this results in a bad map relo warning and returns -LIBBPF_ERRNO__RELOC, causing the test to fail to load. [ ... ] > +__naked void percpu_addr(void) > +{ > + asm volatile (" \ > + r0 = %[percpu_data] ll; \ > + r1 = %[percpu_data] ll; \ [Severity: High] Are the mov_percpu_addr instructions missing from this inline assembly? The commit message notes that an ld_imm64 of a per-CPU map value is followed by a mov_percpu_addr, but the assembly here only emits a BPF_LD_IMM64 instruction. The verifier processes this as BPF_PSEUDO_MAP_VALUE and converts it to a standard 64-bit immediate load. Without the BPF_MOV64_PERCPU_REG (mov_percpu_addr) instruction, the x86 JIT emits movabsq but never emits the addq %gs:..., %reg instruction. This causes the __jited assertions to fail. > + r2 = %[percpu_data] ll; \ > + r3 = %[percpu_data] ll; \ > + r4 = %[percpu_data] ll; \ [ ... ] -- Sashiko AI review · https://sashiko.dev/#/patchset/20260814220254.3797467-1-vineet.gupta@linux.dev?part=2 ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH bpf v3 2/2] selftests/bpf: Check per-CPU address resolution per register 2026-08-14 22:26 ` sashiko-bot @ 2026-08-14 22:39 ` Vineet Gupta 0 siblings, 0 replies; 8+ messages in thread From: Vineet Gupta @ 2026-08-14 22:39 UTC (permalink / raw) To: sashiko-reviews; +Cc: bpf On 8/14/26 3:26 PM, sashiko-bot@kernel.org wrote: > Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider: > - [High] The test uses an unrecognized section name `SEC(".percpu")` for its global variable, preventing libbpf from loading the test. > - [High] The test's inline assembly is missing the `mov_percpu_addr` instruction, causing the `__jited` assertions to always fail. > -- > > --- Patch [2]: [PATCH bpf v3 2/2] selftests/bpf: Check per-CPU address resolution per register --- > commit 94ba39b00c54962b463da6e18c62f89978f29a3f > Author: Vineet Gupta <vineet.gupta@linux.dev> > > selftests/bpf: Check per-CPU address resolution per register > > This commit adds a BPF selftest to verify that the x86 JIT correctly encodes > the per-CPU address resolution. It loads a .percpu variable into all registers > and matches the JITed add instruction against the expected register. > >> diff --git a/tools/testing/selftests/bpf/progs/verifier_percpu_addr.c b/tools/testing/selftests/bpf/progs/verifier_percpu_addr.c >> new file mode 100644 >> index 0000000000000..967f4e6e3a490 >> --- /dev/null >> +++ b/tools/testing/selftests/bpf/progs/verifier_percpu_addr.c >> @@ -0,0 +1,72 @@ >> +// SPDX-License-Identifier: GPL-2.0 >> + >> +#include <vmlinux.h> >> +#include <bpf/bpf_helpers.h> >> +#include "bpf_misc.h" >> + >> +#if defined(__TARGET_ARCH_x86) >> + >> +int percpu_data SEC(".percpu"); > [Severity: High] > Does libbpf support the SEC(".percpu") section name? Yes it does: - tools/lib/bpf/libbpf.c:544 — #define PERCPU_SEC ".percpu" > > When bpf_object__elf_collect runs, it skips the .percpu section as it doesn't > match known data section prefixes like .data. Later, when > bpf_object__relocate_data processes the R_BPF_64_64 relocation for the > percpu_data load, bpf_object__section_to_libbpf_map_type returns > LIBBPF_MAP_UNSPEC. > > Since the section is not the BTF maps section, this results in a bad map relo > warning and returns -LIBBPF_ERRNO__RELOC, causing the test to fail to load. > > [ ... ] > >> +__naked void percpu_addr(void) >> +{ >> + asm volatile (" \ >> + r0 = %[percpu_data] ll; \ >> + r1 = %[percpu_data] ll; \ > [Severity: High] > Are the mov_percpu_addr instructions missing from this inline assembly? They are inserted by the verifier. Thx, -Vineet > The commit message notes that an ld_imm64 of a per-CPU map value is followed > by a mov_percpu_addr, but the assembly here only emits a BPF_LD_IMM64 > instruction. > > The verifier processes this as BPF_PSEUDO_MAP_VALUE and converts it to a > standard 64-bit immediate load. Without the BPF_MOV64_PERCPU_REG > (mov_percpu_addr) instruction, the x86 JIT emits movabsq but never emits the > addq %gs:..., %reg instruction. This causes the __jited assertions to fail. > >> + r2 = %[percpu_data] ll; \ >> + r3 = %[percpu_data] ll; \ >> + r4 = %[percpu_data] ll; \ > [ ... ] > ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH bpf v3 2/2] selftests/bpf: Check per-CPU address resolution per register 2026-08-14 22:02 ` [PATCH bpf v3 2/2] selftests/bpf: Check per-CPU address resolution per register Vineet Gupta 2026-08-14 22:26 ` sashiko-bot @ 2026-08-14 22:39 ` bot+bpf-ci 2026-08-14 23:55 ` Vineet Gupta 1 sibling, 1 reply; 8+ messages in thread From: bot+bpf-ci @ 2026-08-14 22:39 UTC (permalink / raw) To: vineet.gupta, bpf Cc: ast, daniel, andrii, x86, stable, vineet.gupta, ast, andrii, daniel, martin.lau, eddyz87, yonghong.song, clm, ihor.solodrai [-- Attachment #1: Type: text/plain, Size: 4541 bytes --] > commit fe4e8b7aecfff8c04bc3c2f472f9243928effc6e > Author: Vineet Gupta <vineet.gupta@linux.dev> > Date: Fri Aug 14 15:02:54 2026 -0700 > > selftests/bpf: Check per-CPU address resolution per register > > An ld_imm64 of a per-CPU map value is followed by a mov_percpu_addr that > reuses the same register, so which register the address lands in decides > how the JIT encodes the add. Getting the REX prefix wrong there is > invisible to a functional test unless the address happens to land in an > extended register, which is why this went unnoticed. Does this mechanism actually exist in the tree? There is no verifier or JIT path that follows an ld_imm64 of a per-CPU map value with a mov_percpu_addr on the same register. BPF_MOV64_PERCPU_REG has exactly four call sites: kernel/bpf/arraymap.c:288 *insn++ = BPF_MOV64_PERCPU_REG(BPF_REG_0, BPF_REG_0); kernel/bpf/hashtab.c:2429 *insn++ = BPF_MOV64_PERCPU_REG(BPF_REG_0, BPF_REG_0); kernel/bpf/fixups.c:2115 insn_buf[1] = BPF_MOV64_PERCPU_REG(BPF_REG_0, BPF_REG_0); kernel/bpf/fixups.c:2136 insn_buf[1] = BPF_MOV64_PERCPU_REG(BPF_REG_0, BPF_REG_0); All four use BPF_REG_0 as both source and destination. BPF_REG_0 maps to RAX, so is_ereg() is always false and the REX.R encoding path is unreachable from any loadable program. An ld_imm64 of a percpu ksym goes through check_pseudo_btf_id() in the verifier, which only sets aux->btf_var.reg_type = PTR_TO_BTF_ID | MEM_PERCPU. No mov_percpu_addr insn is inserted after it. The per-CPU base is added later by bpf_per_cpu_ptr() or bpf_this_cpu_ptr() kfuncs, which return in R0. > diff --git a/tools/testing/selftests/bpf/progs/verifier_percpu_addr.c b/tools/testing/selftests/bpf/progs/verifier_percpu_addr.c > new file mode 100644 > index 000000000000..967f4e6e3a49 > --- /dev/null > +++ b/tools/testing/selftests/bpf/progs/verifier_percpu_addr.c > @@ -0,0 +1,72 @@ [ ... ] > +SEC("raw_tp") > +__description("per-CPU address resolution") > +__success > +__arch_x86_64 > +__jited(" movabsq $0x{{.*}}, %rax") > +__jited(" addq %gs:{{.*}}, %rax") > +__jited(" movabsq $0x{{.*}}, %rdi") > +__jited(" addq %gs:{{.*}}, %rdi") > +__jited(" movabsq $0x{{.*}}, %rsi") > +__jited(" addq %gs:{{.*}}, %rsi") > +__jited(" movabsq $0x{{.*}}, %rdx") > +__jited(" addq %gs:{{.*}}, %rdx") > +__jited(" movabsq $0x{{.*}}, %rcx") > +__jited(" addq %gs:{{.*}}, %rcx") > +__jited(" movabsq $0x{{.*}}, %r8") > +__jited(" addq %gs:{{.*}}, %r8") > +__jited(" movabsq $0x{{.*}}, %rbx") > +__jited(" addq %gs:{{.*}}, %rbx") > +__jited(" movabsq $0x{{.*}}, %r13") > +__jited(" addq %gs:{{.*}}, %r13") > +__jited(" movabsq $0x{{.*}}, %r14") > +__jited(" addq %gs:{{.*}}, %r14") > +__jited(" movabsq $0x{{.*}}, %r15") > +__jited(" addq %gs:{{.*}}, %r15") > +__naked void percpu_addr(void) > +{ > + asm volatile (" \ > + r0 = %[percpu_data] ll; \ > + r1 = %[percpu_data] ll; \ > + r2 = %[percpu_data] ll; \ > + r3 = %[percpu_data] ll; \ > + r4 = %[percpu_data] ll; \ > + r5 = %[percpu_data] ll; \ > + r6 = %[percpu_data] ll; \ > + r7 = %[percpu_data] ll; \ > + r8 = %[percpu_data] ll; \ > + r9 = %[percpu_data] ll; \ > + r0 = 0; \ > + exit; \ > +" : > + : __imm_addr(percpu_data) > + : __clobber_all); > +} Can these 'addq %gs:..., %<reg>' assertions ever match? A bare sequence of ten ld_imm64s produces ten 'movabsq' lines and zero 'addq %gs:' lines. The verifier does not insert a mov_percpu_addr insn after an ld_imm64 of a percpu ksym. The per-CPU base is added by the bpf_per_cpu_ptr() or bpf_this_cpu_ptr() kfuncs, which return in R0. So the JIT output would be ten consecutive 'movabsq' instructions with no 'addq' between them. The test expects alternating 'movabsq' and 'addq' lines, and consecutive __jited tags must match consecutive disassembly lines unless a '...' tag is inserted. Even if a mov_percpu_addr could be generated, all four call sites hardcode BPF_REG_0, so only the %rax form would ever be reachable. This means the test cannot cover the REX.R fix in acdf39f8be46 ("bpf, x86: Fix per-CPU address resolution into an extended register"). Should there be a Link: or Fixes: tag referencing acdf39f8be46, and does the commit message need to explain what new path makes the destination register reachable? --- AI reviewed your patch. Please fix the bug or email reply why it's not a bug. See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md CI run summary: https://github.com/kernel-patches/bpf/actions/runs/31845422431 ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH bpf v3 2/2] selftests/bpf: Check per-CPU address resolution per register 2026-08-14 22:39 ` bot+bpf-ci @ 2026-08-14 23:55 ` Vineet Gupta 2026-08-15 1:12 ` Borislav Petkov 0 siblings, 1 reply; 8+ messages in thread From: Vineet Gupta @ 2026-08-14 23:55 UTC (permalink / raw) To: bot+bpf-ci, bpf Cc: ast, daniel, andrii, x86, stable, martin.lau, eddyz87, yonghong.song, clm, ihor.solodrai On 8/14/26 3:39 PM, bot+bpf-ci@kernel.org wrote: >> commit fe4e8b7aecfff8c04bc3c2f472f9243928effc6e >> Author: Vineet Gupta <vineet.gupta@linux.dev> >> Date: Fri Aug 14 15:02:54 2026 -0700 >> >> selftests/bpf: Check per-CPU address resolution per register >> >> An ld_imm64 of a per-CPU map value is followed by a mov_percpu_addr that >> reuses the same register, so which register the address lands in decides >> how the JIT encodes the add. Getting the REX prefix wrong there is >> invisible to a functional test unless the address happens to land in an >> extended register, which is why this went unnoticed. > Does this mechanism actually exist in the tree? Indeed these changes are introduced in bpf-next not on *bpf* which the patches were targeting - noob mistake on my part. I'll respin v4 targeting bpf-next and drop stable etc. Thx, -Vineet > > There is no verifier or JIT path that follows an ld_imm64 of a per-CPU map > value with a mov_percpu_addr on the same register. BPF_MOV64_PERCPU_REG has > exactly four call sites: > > kernel/bpf/arraymap.c:288 *insn++ = BPF_MOV64_PERCPU_REG(BPF_REG_0, BPF_REG_0); > kernel/bpf/hashtab.c:2429 *insn++ = BPF_MOV64_PERCPU_REG(BPF_REG_0, BPF_REG_0); > kernel/bpf/fixups.c:2115 insn_buf[1] = BPF_MOV64_PERCPU_REG(BPF_REG_0, BPF_REG_0); > kernel/bpf/fixups.c:2136 insn_buf[1] = BPF_MOV64_PERCPU_REG(BPF_REG_0, BPF_REG_0); > > All four use BPF_REG_0 as both source and destination. BPF_REG_0 maps to > RAX, so is_ereg() is always false and the REX.R encoding path is > unreachable from any loadable program. > > An ld_imm64 of a percpu ksym goes through check_pseudo_btf_id() in the > verifier, which only sets aux->btf_var.reg_type = PTR_TO_BTF_ID | > MEM_PERCPU. No mov_percpu_addr insn is inserted after it. The per-CPU > base is added later by bpf_per_cpu_ptr() or bpf_this_cpu_ptr() kfuncs, > which return in R0. > >> diff --git a/tools/testing/selftests/bpf/progs/verifier_percpu_addr.c b/tools/testing/selftests/bpf/progs/verifier_percpu_addr.c >> new file mode 100644 >> index 000000000000..967f4e6e3a49 >> --- /dev/null >> +++ b/tools/testing/selftests/bpf/progs/verifier_percpu_addr.c >> @@ -0,0 +1,72 @@ > [ ... ] > >> +SEC("raw_tp") >> +__description("per-CPU address resolution") >> +__success >> +__arch_x86_64 >> +__jited(" movabsq $0x{{.*}}, %rax") >> +__jited(" addq %gs:{{.*}}, %rax") >> +__jited(" movabsq $0x{{.*}}, %rdi") >> +__jited(" addq %gs:{{.*}}, %rdi") >> +__jited(" movabsq $0x{{.*}}, %rsi") >> +__jited(" addq %gs:{{.*}}, %rsi") >> +__jited(" movabsq $0x{{.*}}, %rdx") >> +__jited(" addq %gs:{{.*}}, %rdx") >> +__jited(" movabsq $0x{{.*}}, %rcx") >> +__jited(" addq %gs:{{.*}}, %rcx") >> +__jited(" movabsq $0x{{.*}}, %r8") >> +__jited(" addq %gs:{{.*}}, %r8") >> +__jited(" movabsq $0x{{.*}}, %rbx") >> +__jited(" addq %gs:{{.*}}, %rbx") >> +__jited(" movabsq $0x{{.*}}, %r13") >> +__jited(" addq %gs:{{.*}}, %r13") >> +__jited(" movabsq $0x{{.*}}, %r14") >> +__jited(" addq %gs:{{.*}}, %r14") >> +__jited(" movabsq $0x{{.*}}, %r15") >> +__jited(" addq %gs:{{.*}}, %r15") >> +__naked void percpu_addr(void) >> +{ >> + asm volatile (" \ >> + r0 = %[percpu_data] ll; \ >> + r1 = %[percpu_data] ll; \ >> + r2 = %[percpu_data] ll; \ >> + r3 = %[percpu_data] ll; \ >> + r4 = %[percpu_data] ll; \ >> + r5 = %[percpu_data] ll; \ >> + r6 = %[percpu_data] ll; \ >> + r7 = %[percpu_data] ll; \ >> + r8 = %[percpu_data] ll; \ >> + r9 = %[percpu_data] ll; \ >> + r0 = 0; \ >> + exit; \ >> +" : >> + : __imm_addr(percpu_data) >> + : __clobber_all); >> +} > Can these 'addq %gs:..., %<reg>' assertions ever match? > > A bare sequence of ten ld_imm64s produces ten 'movabsq' lines and zero > 'addq %gs:' lines. The verifier does not insert a mov_percpu_addr insn > after an ld_imm64 of a percpu ksym. The per-CPU base is added by the > bpf_per_cpu_ptr() or bpf_this_cpu_ptr() kfuncs, which return in R0. > > So the JIT output would be ten consecutive 'movabsq' instructions with no > 'addq' between them. The test expects alternating 'movabsq' and 'addq' > lines, and consecutive __jited tags must match consecutive disassembly > lines unless a '...' tag is inserted. > > Even if a mov_percpu_addr could be generated, all four call sites hardcode > BPF_REG_0, so only the %rax form would ever be reachable. > > This means the test cannot cover the REX.R fix in acdf39f8be46 ("bpf, x86: > Fix per-CPU address resolution into an extended register"). > > Should there be a Link: or Fixes: tag referencing acdf39f8be46, and does > the commit message need to explain what new path makes the destination > register reachable? > > > --- > AI reviewed your patch. Please fix the bug or email reply why it's not a bug. > See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md > > CI run summary: https://github.com/kernel-patches/bpf/actions/runs/31845422431 ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH bpf v3 2/2] selftests/bpf: Check per-CPU address resolution per register 2026-08-14 23:55 ` Vineet Gupta @ 2026-08-15 1:12 ` Borislav Petkov 0 siblings, 0 replies; 8+ messages in thread From: Borislav Petkov @ 2026-08-15 1:12 UTC (permalink / raw) To: Vineet Gupta Cc: bot+bpf-ci, bpf, ast, daniel, andrii, x86, stable, martin.lau, eddyz87, yonghong.song, clm, ihor.solodrai On Fri, Aug 14, 2026 at 04:55:00PM -0700, Vineet Gupta wrote: > I'll respin v4 targeting bpf-next and drop stable etc. You've sent three revisions today. Slow down and stop spamming people! -- Regards/Gruss, Boris. https://people.kernel.org/tglx/notes-about-netiquette ^ permalink raw reply [flat|nested] 8+ messages in thread
end of thread, other threads:[~2026-08-15 1:13 UTC | newest] Thread overview: 8+ messages (download: mbox.gz follow: Atom feed -- links below jump to the message on this page -- 2026-08-14 22:02 [PATCH bpf v3 0/2] bpf, x86: fix per-CPU address resolution into an extended register Vineet Gupta 2026-08-14 22:02 ` [PATCH bpf v3 1/2] bpf, x86: Fix " Vineet Gupta 2026-08-14 22:02 ` [PATCH bpf v3 2/2] selftests/bpf: Check per-CPU address resolution per register Vineet Gupta 2026-08-14 22:26 ` sashiko-bot 2026-08-14 22:39 ` Vineet Gupta 2026-08-14 22:39 ` bot+bpf-ci 2026-08-14 23:55 ` Vineet Gupta 2026-08-15 1:12 ` Borislav Petkov
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox