From: Pedro Falcato <pfalcato@suse.de>
To: Nikunj A Dadhania <nikunj@amd.com>
Cc: rppt@kernel.org, akpm@linux-foundation.org,
anton.ivanov@cambridgegreys.com, bp@alien8.de,
brendan.higgins@linux.dev, da.gomez@samsung.com,
danielt@kernel.org, dave.hansen@linux.intel.com,
davidgow@google.com, dianders@chromium.org, hpa@zytor.com,
jason.wessel@windriver.com, jikos@kernel.org,
joe.lawrence@redhat.com, johannes@sipsolutions.net,
jpoimboe@kernel.org, kgdb-bugreport@lists.sourceforge.net,
kirill.shutemov@linux.intel.com, kunit-dev@googlegroups.com,
linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org,
linux-mm@kvack.org, linux-modules@vger.kernel.org,
linux-trace-kernel@vger.kernel.org, linux-um@lists.infradead.org,
live-patching@vger.kernel.org, lorenzo.stoakes@oracle.com,
luto@kernel.org, mark.rutland@arm.com, mbenes@suse.cz,
mcgrof@kernel.org, mhiramat@kernel.org, mingo@redhat.com,
peterz@infradead.org, petr.pavlu@suse.com, pmladek@suse.com,
richard@nod.at, rmoar@google.com, rostedt@goodmis.org,
samitolvanen@google.com, shuah@kernel.org, song@kernel.org,
tglx@linutronix.de, x86@kernel.org, bharata@amd.com,
ananth.narayan@amd.com
Subject: Re: [PATCH v3 9/9] x86: re-enable EXECMEM_ROX support
Date: Wed, 12 Aug 2026 12:38:08 +0100 [thread overview]
Message-ID: <anxa3VjtukAkY4Cl@pedro-suse> (raw)
In-Reply-To: <20260812063316.21371-1-nikunj@amd.com>
On Wed, Aug 12, 2026 at 06:33:15AM +0000, Nikunj A Dadhania wrote:
> Hi,
>
> A kernel BUG in __text_poke() is observed inside a Ubuntu 26.04 guest
> (v7.0.0.28 and also with upstream v7.2-4c5). Host is running on v7.2-rc5.
>
> SETUP
> -----
>
> 512 QEMU Q35 guests (2 vCPUs, 1G RAM each) boot concurrently across 200
> iterations. The crash is seen during module loading in early boot.
> 2-7 guests out of ~100,000 boot attempts crash with the kernel BUG.
>
> Reproduced on:
> - AMD EPYC, QEMU Q35 guests
> - Intel Icelake, QEMU Q35 guests
>
> Apparently it was seen once in the non-guest Ubuntu 26.04 environment as well
> https://discourse.ubuntu.com/t/failed-boot-one-time-on-ubuntu-unity-26-04-educational-mission/85300
>
> SYMPTOMS
> --------
>
> Either of the below two crash signatures is observed:
>
> Signature 1 — BUG_ON(!pages[0]) in __text_poke():
>
> kernel BUG at arch/x86/kernel/alternative.c:2564!
> Oops: invalid opcode: 0000 [#1] SMP NOPTI
> CPU: 0 UID: 0 PID: 741 Comm: (udev-worker) Not tainted 7.0.0-28-generic #28-Ubuntu PREEMPT(lazy)
> Hardware name: QEMU Standard PC (Q35 + ICH9, 2009)
> RIP: 0010:__text_poke+0x2a4/0x400
> Call Trace:
> <TASK>
> smp_text_poke_batch_finish+0x25f/0x5b0
> smp_text_poke_single+0x13/0x30
> __static_call_transform+0x10e/0x210
> arch_static_call_transform+0x5c/0xc0
> __static_call_init+0x154/0x300
> static_call_module_notify+0x144/0x170
> notifier_call_chain+0x62/0xf0
> blocking_notifier_call_chain_robust+0x62/0xc0
> load_module+0x5ef/0x960
> Modules linked in: ppdev(+) kvm(+) irqbypass ...
>
> Signature 2 — BUG_ON(memcmp()) in __text_poke():
>
> kernel BUG at arch/x86/kernel/alternative.c:2632!
> Oops: invalid opcode: 0000 [#1] SMP NOPTI
> CPU: 0 UID: 0 PID: 742 Comm: (udev-worker) Not tainted 7.0.0-28-generic #28-Ubuntu PREEMPT(lazy)
> Hardware name: QEMU Standard PC (Q35 + ICH9, 2009)
> RIP: 0010:__text_poke+0x380/0x400
> Call Trace:
> <TASK>
> smp_text_poke_batch_finish+0x323/0x5b0
> smp_text_poke_single+0x13/0x30
> __static_call_transform+0x10e/0x210
> arch_static_call_transform+0x5c/0xc0
> __static_call_init+0x154/0x300
> static_call_module_notify+0x144/0x170
> notifier_call_chain+0x62/0xf0
> blocking_notifier_call_chain_robust+0x62/0xc0
> load_module+0x5ef/0x960
> Modules linked in: kvm(+) irqbypass ...
>
> INSTRUMENTATION
> ---------------
>
> The following instrumentation was applied to capture the page table state at
> the moment of failure:
>
> In mm/vmalloc.c::vmalloc_to_page()
>
> ptep = pte_offset_kernel(pmd, addr);
> pte = ptep_get(ptep);
> if (pte_present(pte))
> page = pte_page(pte);
> else
> pr_err("vmalloc_to_page: pte_not_present for addr=%lx "
> "PGD=%lx P4D=%lx PUD=%lx PMD=%lx PTE=%lx\n",
> addr, pgd_val(*pgd), p4d_val(*p4d), pud_val(*pud),
> pmd_val(*pmd), pte_val(pte));
>
> In arch/x86/kernel/alternative.c::__text_poke()
>
> /* BUG 1: pages[0] NULL */
> if (!pages[0] || ...) {
> pr_err("%s: vmalloc_to_page failed for addr=%pK pages[0]=%pK pages[1]=%pK cross=%d\n", ...);
> BUG();
> }
>
> ...
>
> /* BUG 2: memcmp mismatch */
> if (memcmp(addr, src, len)) {
> struct page *page_now = vmalloc_to_page(addr);
> pr_err("%s: memcmp failed at addr=%pK len=%zu\n", __func__, addr, len);
> pr_err("%s: expected (src): %*ph\n", __func__, (int)len, src);
> pr_err("%s: found (dst): %*ph\n", __func__, (int)len, (u8 *)addr);
> pr_err("%s: pfn at poke time (vmalloc_to_page): %lx\n", __func__, page_to_pfn(pages[0]));
> pr_err("%s: pfn at readback time (vmalloc_to_page): %lx\n", __func__, page_to_pfn(page_now));
> text_poke_dump_pagetable((unsigned long)addr); /* reads CR3 */
> BUG();
> }
>
> Signature 1 — vmalloc_to_page() returned NULL:
>
> vmalloc_to_page: pte_not_present for addr=ffffffffc0697979 PGD=28242067 P4D=28243067 PUD=28245067 PMD=3be001a1 PTE=0
> SMP alternatives: __text_poke: vmalloc_to_page failed for addr=ffffffffc0697979 pages[0]=0000000000000000 pages[1]=0000000000000000 cross=0
>
> PMD was not a leaf, and the code reached a point to check pte_present() and
> finds the PTE to be not present. But in the dump in else case, pmd_val(*pmd)
> shows that it is a PMD leaf (PSE bit set).
>
> Signature 2 — vmalloc_to_page() returned a wrong page:
>
> SMP alternatives: __text_poke: memcmp failed at addr=ffffffffc06fc4bd len=1
> SMP alternatives: __text_poke: expected (src): cc
> SMP alternatives: __text_poke: found (dst): e8
> SMP alternatives: __text_poke: pfn at poke time (vmalloc_to_page): 3884a
> SMP alternatives: __text_poke: pfn at readback time (vmalloc_to_page): 52fc
> SMP alternatives: __text_poke: page table dump for addr:
> PGD 1ba42067 P4D 1ba43067 PUD 1ba45067 PMD 52001a1
>
> vmalloc_to_page() returned pfn=0x3884a when pages[0] was queried. The write
> of 0xcc landed there. When vmalloc_to_page() was called again during the
> memcmp readback in the instrumentation code, the vmap mapping resolved through
> PMD=52001a1 to a different pfn=0x52fc, which contains 0xe8 (existing
> instruction bytes). This is silent memory corruption — 0xcc was written to
> the wrong physical page (pfn=0x3884a).
>
> BISECTION
> ---------
>
> The following experiments all ran 200 iterations of 512 concurrent guest boots
> with zero crashes.
>
> 1. Remove "select ARCH_HAS_EXECMEM_ROX" from arch/x86/Kconfig
>
> 2. Boot with "nohugevmalloc"
>
> 3. Remove VM_ALLOW_HUGE_VMAP from vm_flags in execmem_cache_populate()
>
> - unsigned long vm_flags = VM_ALLOW_HUGE_VMAP;
> + unsigned long vm_flags = 0;
>
> 4. Remove CPA_COLLAPSE from set_memory_rox()
>
> return change_page_attr_set_clr(&addr, numpages, __pgprot(0), clr, 0,
> - CPA_COLLAPSE, NULL);
> + 0, NULL);
>
> Without CPA_COLLAPSE, set_memory_rox() changes 4K PTE permissions
> in-place and the PMD stays permanently as a PTE-page pointer after
> the first split — no re-collapse, no future race window. This
> confirms that CPA_COLLAPSE is a necessary part of the race sequence.
>
> 5. Hold pgd_lock around vmalloc_to_page() in __text_poke()
>
> if (!core_kernel_text((unsigned long)addr)) {
> + spin_lock(&pgd_lock);
> pages[0] = vmalloc_to_page(addr);
> if (cross_page_boundary)
> pages[1] = vmalloc_to_page(addr + PAGE_SIZE);
> + spin_unlock(&pgd_lock);
> } else {
>
> __split_large_page() holds pgd_lock while installing the new PMD
> and populating all 512 4K PTEs. Holding pgd_lock around
> vmalloc_to_page() prevents CPA from modifying the PMD concurrently
> with the page table walk. This confirms the race is precisely
> between vmalloc_to_page() reading the PMD/PTE and
> __split_large_page() modifying it under pgd_lock.
>
> Note: this is a diagnostic experiment, not a production fix.
>
> The experiments point to a race between vmalloc_to_page() in
> __text_poke() and __split_large_page() modifying the shared PMD under
> pgd_lock. __split_large_page() holds pgd_lock while modifying the
> PMD; vmalloc_to_page() holds no lock and can observe the PMD in an
> intermediate state.
>
> What would be the correct fix for this issue?
>
> Happy to test any proposed fix or run further experiments.
https://lore.kernel.org/linux-mm/anCK3eWFMwZqq5ka@pedro-suse/
--
Pedro
next prev parent reply other threads:[~2026-08-12 11:38 UTC|newest]
Thread overview: 18+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-01-26 7:47 [PATCH v3 0/9] x86/module: rework ROX cache to avoid writable copy Mike Rapoport
2025-01-26 7:47 ` [PATCH v3 1/9] x86/mm/pat: cpa-test: fix length for CPA_ARRAY test Mike Rapoport
2025-01-26 7:47 ` [PATCH v3 2/9] x86/mm/pat: drop duplicate variable in cpa_flush() Mike Rapoport
2025-01-26 7:47 ` [PATCH v3 3/9] x86/mm/pat: restore large ROX pages after fragmentation Mike Rapoport
2025-01-26 7:47 ` [PATCH v3 4/9] execmem: don't remove ROX cache from the direct map Mike Rapoport
2025-01-26 7:47 ` [PATCH v3 5/9] execmem: add API for temporal remapping as RW and restoring ROX afterwards Mike Rapoport
2025-01-26 7:47 ` [PATCH v3 6/9] module: switch to execmem API for remapping as RW and restoring ROX Mike Rapoport
2025-01-27 12:50 ` Petr Pavlu
2025-01-28 10:00 ` Mike Rapoport
2025-01-28 10:31 ` Petr Pavlu
2025-01-26 7:47 ` [PATCH v3 7/9] Revert "x86/module: prepare module loading for ROX allocations of text" Mike Rapoport
2025-01-26 7:47 ` [PATCH v3 8/9] module: drop unused module_writable_address() Mike Rapoport
2025-01-26 7:47 ` [PATCH v3 9/9] x86: re-enable EXECMEM_ROX support Mike Rapoport
2026-08-12 6:33 ` Nikunj A Dadhania
2026-08-12 11:38 ` Pedro Falcato [this message]
2026-08-12 13:20 ` Nikunj A. Dadhania
2025-01-27 11:19 ` [PATCH v3 0/9] x86/module: rework ROX cache to avoid writable copy Peter Zijlstra
2025-01-29 17:28 ` Lorenzo Stoakes
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=anxa3VjtukAkY4Cl@pedro-suse \
--to=pfalcato@suse.de \
--cc=akpm@linux-foundation.org \
--cc=ananth.narayan@amd.com \
--cc=anton.ivanov@cambridgegreys.com \
--cc=bharata@amd.com \
--cc=bp@alien8.de \
--cc=brendan.higgins@linux.dev \
--cc=da.gomez@samsung.com \
--cc=danielt@kernel.org \
--cc=dave.hansen@linux.intel.com \
--cc=davidgow@google.com \
--cc=dianders@chromium.org \
--cc=hpa@zytor.com \
--cc=jason.wessel@windriver.com \
--cc=jikos@kernel.org \
--cc=joe.lawrence@redhat.com \
--cc=johannes@sipsolutions.net \
--cc=jpoimboe@kernel.org \
--cc=kgdb-bugreport@lists.sourceforge.net \
--cc=kirill.shutemov@linux.intel.com \
--cc=kunit-dev@googlegroups.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux-modules@vger.kernel.org \
--cc=linux-trace-kernel@vger.kernel.org \
--cc=linux-um@lists.infradead.org \
--cc=live-patching@vger.kernel.org \
--cc=lorenzo.stoakes@oracle.com \
--cc=luto@kernel.org \
--cc=mark.rutland@arm.com \
--cc=mbenes@suse.cz \
--cc=mcgrof@kernel.org \
--cc=mhiramat@kernel.org \
--cc=mingo@redhat.com \
--cc=nikunj@amd.com \
--cc=peterz@infradead.org \
--cc=petr.pavlu@suse.com \
--cc=pmladek@suse.com \
--cc=richard@nod.at \
--cc=rmoar@google.com \
--cc=rostedt@goodmis.org \
--cc=rppt@kernel.org \
--cc=samitolvanen@google.com \
--cc=shuah@kernel.org \
--cc=song@kernel.org \
--cc=tglx@linutronix.de \
--cc=x86@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).