From: Shuai Xue <xueshuai@linux.alibaba.com>
To: Wei-Lin Chang <weilin.chang@arm.com>,
Marc Zyngier <maz@kernel.org>,
Wang Han <wanghan@linux.alibaba.com>
Cc: linux-arm-kernel@lists.infradead.org, kvmarm@lists.linux.dev,
linux-kernel@vger.kernel.org, oupton@kernel.org,
tabba@google.com, joey.gouly@arm.com, seiden@linux.ibm.com,
suzuki.poulose@arm.com, catalin.marinas@arm.com, will@kernel.org,
ljs@kernel.org, itaru.kitayama@fujitsu.com
Subject: Re: [PATCH v5 0/6] KVM: arm64: nv: Implement nested stage-2 reverse map (new data structure)
Date: Fri, 4 Sep 2026 15:01:24 +0800 [thread overview]
Message-ID: <46342b48-550e-42c1-9d9f-e800c72269c2@linux.alibaba.com> (raw)
In-Reply-To: <wsnf276xrun3g4hjr3t6alprapctg5z4kpj422v3hujqdjnrsa@swq4blkqvvj5>
On 9/3/26 9:28 PM, Wei-Lin Chang wrote:
> On Thu, Sep 03, 2026 at 08:43:35AM +0100, Marc Zyngier wrote:
>> On Wed, 02 Sep 2026 17:35:00 +0100,
>> Wang Han <wanghan@linux.alibaba.com> wrote:
>>>
>>> Hi Wei-Lin,
>>>
>>> I tested this series on a Yitian 710 system with an ARM Neoverse-N2 CPU
>>> (128 CPUs, 2 NUMA nodes).
>>>
>>> Test environment
>>> ----------------
>>>
>>> L0 kernel: Linux v7.2-rc6
>>> L1 guest: Ubuntu 26.04 LTS, kernel 7.0.0-27-generic (aarch64)
>>> QEMU: 10.2.3
>>>
>>> L0 NUMA balancing was enabled (`/proc/sys/kernel/numa_balancing=1`).
>>> The host was booted with `kvm_arm.mode=nested`.
>>>
>>> This series fixes a functional hang that is exposed when NUMA balancing is
>>> enabled. The previous nested stage-2 unmap path is too slow for this
>>> workload, making the performance problem user-visible: NUMA balancing can
>>> leave the L1 guest unable to make progress and eventually hang during boot.
>>>
>>> The L1 was started with 8 vCPUs and 32 GiB of RAM using:
>>>
>>> qemu-system-aarch64 -smp 8 -m 32G \
>>> -machine virt,accel=kvm,gic-version=3,virtualization=on \
>>> -cpu host -nographic -enable-kvm \
>>> -drive if=pflash,format=raw,readonly=on,file=pflash0_bak.img \
>>> -drive if=pflash,format=raw,file=pflash1_bak.img \
>>> -drive file=./ubuntu-vm.qcow2,format=qcow2,if=virtio,cache=none,aio=native \
>>> -nic user,model=virtio-net-pci,hostfwd=tcp::11234-:22 \
>>> -serial mon:stdio
>>>
>>
>> Puzzling. If you are only running an L1 in VHE mode, there is no
>> shadow S2, and therefore nothing to unmap. For shadow S2s to be built
>> and affect the MMU notifiers, you need to run an L2.
>
> I was thinking the same at first, but realized even with L1 in VHE mode
> there is a small period of time where L1 runs in its EL1 during boot, so
> one nested MMU will become valid for each vCPU. That causes
> kvm_nested_s2_unmap() to iterate through the entire IPA space 8 times
> (-smp 8).
>
> What I am curious about is whether one single notifier unmap is enough
> to hang L1, or were there multiple notifier unmaps.
>
> QEMU with -machine virt uses 40 IPA bits only, unmapping that takes:
> 1024 (4KB pages, unmapping 1GB per iteration)
> 32768 (16KB pages, unmapping 32MB per iteration)
> 2048 (64KB pages, unmapping 512MB per iteration)
> iterations for each page size. There aren't many mappings in each
> iteration too. Does this really take that long on real hardware (even if
> this must be done 8 times)?
>
> Thanks,
> Wei-Lin Chang
>
>>
>> So what are your actual test conditions?
>>
>> M.
>>
Hi, Wei-Lin and Marc,
I was able to reproduce this issue and capture ftrace evidence that confirms
the root cause. Below is the analysis, trace log, and timing data.
## Problem
Environment:
- Host (L0): ARM64, KVM with virtualization=on (nested virtualization)
- Guest (L1): Ubuntu 26.04, 8 vCPUs / 32 GB
- Host NUMA balancing enabled, numad active
When booting the L1 QEMU guest, the L1 kernel hits a soft lockup during
early boot (~45 s):
[ 45.646468] watchdog: BUG: soft lockup - CPU#0 stuck for 32s!
[kworker/0:2:330]
[ 45.646882] watchdog: BUG: soft lockup - CPU#2 stuck for 29s!
[snap:1146]
[ 45.647093] watchdog: BUG: soft lockup - CPU#3 stuck for 29s!
[snap:1141]
[ 45.647242] watchdog: BUG: soft lockup - CPU#7 stuck for 29s!
[snap:1144]
At the same time, L0 dmesg reports the QEMU main thread blocked in D-state
for more than 120 s:
[11239.341817] INFO: task qemu-system-aar:170799 blocked in I/O wait for
more than 120 seconds.
...
softleaf_entry_wait_on_locked+0x280/0x2d0
migration_entry_wait+0xdc/0x140
do_swap_page+0x834/0xd80
handle_pte_fault+0x208/0x2b8
__handle_mm_fault+0x228/0x528
handle_mm_fault+0xdc/0x2d8
do_page_fault+0x244/0x790
do_translation_fault+0x4c/0x88
do_mem_abort+0x4c/0xa0
el1_abort+0x50/0x80
el1h_64_sync_handler+0x50/0x108
el1h_64_sync+0x80/0x88
do_sys_poll+0x224/0x290
__arm64_sys_ppoll+0xa4/0x130
Note: This is not a 100% reproducible failure. In my automated loop the hang
reproduced on the 2nd boot attempt, but other attempts ran for 6–9 minutes
without hitting it. The bug is clearly timing-dependent on NUMA migration
activity during the L1 boot window.
## Root cause
The issue is in arch/arm64/kvm/nested.c:
void kvm_nested_s2_unmap(struct kvm *kvm, bool may_block)
{
int i;
lockdep_assert_held_write(&kvm->mmu_lock);
if (!kvm->arch.nested_mmus_size)
return;
for (i = 0; i < kvm->arch.nested_mmus_size; i++) {
struct kvm_s2_mmu *mmu = &kvm->arch.nested_mmus[i];
if (kvm_s2_mmu_valid(mmu))
kvm_stage2_unmap_range(mmu, 0, kvm_phys_size(mmu), may_block);
}
kvm_invalidate_vncr_ipa(kvm, 0, BIT(kvm->arch.mmu.pgt->ia_bits));
}
When L0 NUMA balancing migrates a page belonging to the QEMU process, the
MMU notifier path calls kvm_unmap_gfn_range():
bool kvm_unmap_gfn_range(struct kvm *kvm, struct kvm_gfn_range *range)
{
...
__unmap_stage2_range(&kvm->arch.mmu, range->start << PAGE_SHIFT,
(range->end - range->start) << PAGE_SHIFT,
range->may_block);
kvm_nested_s2_unmap(kvm, range->may_block); /* full unmap */
return false;
}
kvm_handle_hva_range() takes kvm->mmu_lock for writing before invoking the
handler and releases it only after the handler returns
(virt/kvm/kvm_main.c:622-642). Therefore, the entire kvm_nested_s2_unmap()
runs under kvm->mmu_lock.
The problem is that kvm_nested_s2_unmap() does not unmap only the affected
GPA range. Instead, it unmaps the entire IPA space (0 to kvm_phys_size(mmu))
for every nested S2 MMU. With nested virtualization enabled, this is very
expensive.
## Trace evidence
I captured function_graph traces for kvm_unmap_gfn_range,
kvm_nested_s2_unmap, and kvm_stage2_unmap_range, plus mm_migrate_pages
tracepoints.
1. numad-triggered full unmap
64) numad-1888 | | /* set_migration_pte:
addr=fff790579000, pte=3040c319d9680 order=0 */
64) numad-1888 | | kvm_unmap_gfn_range() {
64) numad-1888 | | kvm_nested_s2_unmap() {
64) numad-1888 | @ 877415.7 us | kvm_stage2_unmap_range();
64) numad-1888 | @ 877418.2 us | }
64) numad-1888 | @ 877422.0 us | }
Each set_migration_pte line is a single-page NUMA migration. Yet each
migration triggers one full kvm_nested_s2_unmap() that takes 877 ms.
Subsequent calls show per-page unmap durations between 841 ms and 1.16 s.
2. QEMU threads blocked as well
23) qemu-sy-227844 | | kvm_unmap_gfn_range() {
23) qemu-sy-227844 | | kvm_nested_s2_unmap() {
23) qemu-sy-227844 | $ 1170521 us | kvm_stage2_unmap_range();
23) qemu-sy-227844 | $ 1170529 us | }
23) qemu-sy-227844 | $ 1170542 us | }
QEMU's own threads also get stuck in the same full unmap, with one call
reaching 6.67 s.
3. Statistics
┌──────────────┬───────┬──────────────────────────────────────┐
│ Thread │ Calls │ kvm_nested_s2_unmap duration │
├──────────────┼───────┼──────────────────────────────────────┤
│ numad-1888 │ 59 │ min 0.84 s / avg 1.02 s / max 1.16 s │
├──────────────┼───────┼──────────────────────────────────────┤
│ QEMU threads │ 34 │ min 1.14 s / avg 5.14 s / max 6.67 s │
└──────────────┴───────┴──────────────────────────────────────┘
numad migrates pages one after another; each page holds kvm->mmu_lock for
about one second. QEMU vCPU threads cannot acquire mmu_lock and also block
on migration_entry_wait. The L1 guest vCPUs make no forward progress, and
the watchdog fires.
Soft lockup threshold check
Read directly inside the L1 guest:
root@ubuntu-vm:~# cat /proc/sys/kernel/watchdog_thresh
10
So the soft lockup threshold is 2 * watchdog_thresh = 20 s.
The L1 guest reported stuck times of 29~32 s, which exceeds the 20 s
threshold.
## Conclusion
The root cause is confirmed: kvm_nested_s2_unmap() performs a full IPA space
unmap in the MMU notifier path instead of unmapping only the affected
GPA/CPAI range. The interval-tree-based precise range unmap approach is the
right fix.
Please consider applying the patch that replaces the full unmap with
kvm_nested_unmap_cipa_range() to avoid scanning the entire nested stage-2
page table on every NUMA migration.
Thanks,
Shuai
next prev parent reply other threads:[~2026-09-04 7:01 UTC|newest]
Thread overview: 18+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-10 20:50 [PATCH v5 0/6] KVM: arm64: nv: Implement nested stage-2 reverse map (new data structure) Wei-Lin Chang
2026-08-10 20:50 ` [PATCH v5 1/6] KVM: arm64: Use a variable for the canonical IPA in kvm_s2_fault_map() Wei-Lin Chang
2026-08-10 20:50 ` [PATCH v5 2/6] KVM: arm64: nv: Introduce guest stage-2 tracking structures Wei-Lin Chang
2026-08-14 1:04 ` Itaru Kitayama
2026-08-14 10:42 ` Wei-Lin Chang
2026-08-16 22:01 ` Itaru Kitayama
2026-08-10 20:50 ` [PATCH v5 3/6] KVM: arm64: nv: Track guest stage-2 mapping creation Wei-Lin Chang
2026-08-10 20:50 ` [PATCH v5 4/6] KVM: arm64: nv: Track guest stage-2 mapping removal Wei-Lin Chang
2026-08-10 20:50 ` [PATCH v5 5/6] KVM: arm64: nv: Avoid full shadow stage-2 unmap Wei-Lin Chang
2026-08-10 20:50 ` [PATCH v5 6/6] KVM: arm64: Refactor kvm_unmap_gfn_range() with common variables Wei-Lin Chang
2026-08-12 2:12 ` [PATCH v5 0/6] KVM: arm64: nv: Implement nested stage-2 reverse map (new data structure) Itaru Kitayama
2026-09-02 16:35 ` Wang Han
2026-09-03 7:43 ` Marc Zyngier
2026-09-03 13:28 ` Wei-Lin Chang
2026-09-04 7:01 ` Shuai Xue [this message]
2026-09-04 7:54 ` Marc Zyngier
2026-09-04 7:49 ` Marc Zyngier
2026-09-04 11:37 ` Wei-Lin Chang
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=46342b48-550e-42c1-9d9f-e800c72269c2@linux.alibaba.com \
--to=xueshuai@linux.alibaba.com \
--cc=catalin.marinas@arm.com \
--cc=itaru.kitayama@fujitsu.com \
--cc=joey.gouly@arm.com \
--cc=kvmarm@lists.linux.dev \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=ljs@kernel.org \
--cc=maz@kernel.org \
--cc=oupton@kernel.org \
--cc=seiden@linux.ibm.com \
--cc=suzuki.poulose@arm.com \
--cc=tabba@google.com \
--cc=wanghan@linux.alibaba.com \
--cc=weilin.chang@arm.com \
--cc=will@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).