linux-arm-kernel.lists.infradead.org archive mirror
 help / color / mirror / Atom feed
From: Shuai Xue <xueshuai@linux.alibaba.com>
To: Wei-Lin Chang <weilin.chang@arm.com>,
	Marc Zyngier <maz@kernel.org>,
	Wang Han <wanghan@linux.alibaba.com>
Cc: linux-arm-kernel@lists.infradead.org, kvmarm@lists.linux.dev,
	linux-kernel@vger.kernel.org, oupton@kernel.org,
	tabba@google.com, joey.gouly@arm.com, seiden@linux.ibm.com,
	suzuki.poulose@arm.com, catalin.marinas@arm.com, will@kernel.org,
	ljs@kernel.org, itaru.kitayama@fujitsu.com
Subject: Re: [PATCH v5 0/6] KVM: arm64: nv: Implement nested stage-2 reverse map (new data structure)
Date: Fri, 4 Sep 2026 15:01:24 +0800	[thread overview]
Message-ID: <46342b48-550e-42c1-9d9f-e800c72269c2@linux.alibaba.com> (raw)
In-Reply-To: <wsnf276xrun3g4hjr3t6alprapctg5z4kpj422v3hujqdjnrsa@swq4blkqvvj5>



On 9/3/26 9:28 PM, Wei-Lin Chang wrote:
> On Thu, Sep 03, 2026 at 08:43:35AM +0100, Marc Zyngier wrote:
>> On Wed, 02 Sep 2026 17:35:00 +0100,
>> Wang Han <wanghan@linux.alibaba.com> wrote:
>>>
>>> Hi Wei-Lin,
>>>
>>> I tested this series on a Yitian 710 system with an ARM Neoverse-N2 CPU
>>> (128 CPUs, 2 NUMA nodes).
>>>
>>> Test environment
>>> ----------------
>>>
>>>    L0 kernel: Linux v7.2-rc6
>>>    L1 guest: Ubuntu 26.04 LTS, kernel 7.0.0-27-generic (aarch64)
>>>    QEMU: 10.2.3
>>>
>>> L0 NUMA balancing was enabled (`/proc/sys/kernel/numa_balancing=1`).
>>> The host was booted with `kvm_arm.mode=nested`.
>>>
>>> This series fixes a functional hang that is exposed when NUMA balancing is
>>> enabled.  The previous nested stage-2 unmap path is too slow for this
>>> workload, making the performance problem user-visible: NUMA balancing can
>>> leave the L1 guest unable to make progress and eventually hang during boot.
>>>
>>> The L1 was started with 8 vCPUs and 32 GiB of RAM using:
>>>
>>>    qemu-system-aarch64 -smp 8 -m 32G \
>>>      -machine virt,accel=kvm,gic-version=3,virtualization=on \
>>>      -cpu host -nographic -enable-kvm \
>>>      -drive if=pflash,format=raw,readonly=on,file=pflash0_bak.img \
>>>      -drive if=pflash,format=raw,file=pflash1_bak.img \
>>>      -drive file=./ubuntu-vm.qcow2,format=qcow2,if=virtio,cache=none,aio=native \
>>>      -nic user,model=virtio-net-pci,hostfwd=tcp::11234-:22 \
>>>      -serial mon:stdio
>>>
>>
>> Puzzling. If you are only running an L1 in VHE mode, there is no
>> shadow S2, and therefore nothing to unmap. For shadow S2s to be built
>> and affect the MMU notifiers, you need to run an L2.
> 
> I was thinking the same at first, but realized even with L1 in VHE mode
> there is a small period of time where L1 runs in its EL1 during boot, so
> one nested MMU will become valid for each vCPU. That causes
> kvm_nested_s2_unmap() to iterate through the entire IPA space 8 times
> (-smp 8).
> 
> What I am curious about is whether one single notifier unmap is enough
> to hang L1, or were there multiple notifier unmaps.
> 
> QEMU with -machine virt uses 40 IPA bits only, unmapping that takes:
> 1024  (4KB pages,  unmapping 1GB per iteration)
> 32768 (16KB pages, unmapping 32MB per iteration)
> 2048  (64KB pages, unmapping 512MB per iteration)
> iterations for each page size. There aren't many mappings in each
> iteration too. Does this really take that long on real hardware (even if
> this must be done 8 times)?
> 
> Thanks,
> Wei-Lin Chang
> 
>>
>> So what are your actual test conditions?
>>
>> 	M.
>>

Hi, Wei-Lin and Marc,

I was able to reproduce this issue and capture ftrace evidence that confirms
the root cause. Below is the analysis, trace log, and timing data.

## Problem

Environment:
- Host (L0): ARM64, KVM with virtualization=on (nested virtualization)
- Guest (L1): Ubuntu 26.04, 8 vCPUs / 32 GB
- Host NUMA balancing enabled, numad active

When booting the L1 QEMU guest, the L1 kernel hits a soft lockup during
early boot (~45 s):

  [   45.646468] watchdog: BUG: soft lockup - CPU#0 stuck for 32s!
  [kworker/0:2:330]
  [   45.646882] watchdog: BUG: soft lockup - CPU#2 stuck for 29s!
  [snap:1146]
  [   45.647093] watchdog: BUG: soft lockup - CPU#3 stuck for 29s!
  [snap:1141]
  [   45.647242] watchdog: BUG: soft lockup - CPU#7 stuck for 29s!
  [snap:1144]

At the same time, L0 dmesg reports the QEMU main thread blocked in D-state
for more than 120 s:

  [11239.341817] INFO: task qemu-system-aar:170799 blocked in I/O wait for
  more than 120 seconds.
  ...
    softleaf_entry_wait_on_locked+0x280/0x2d0
    migration_entry_wait+0xdc/0x140
    do_swap_page+0x834/0xd80
    handle_pte_fault+0x208/0x2b8
    __handle_mm_fault+0x228/0x528
    handle_mm_fault+0xdc/0x2d8
    do_page_fault+0x244/0x790
    do_translation_fault+0x4c/0x88
    do_mem_abort+0x4c/0xa0
    el1_abort+0x50/0x80
    el1h_64_sync_handler+0x50/0x108
    el1h_64_sync+0x80/0x88
    do_sys_poll+0x224/0x290
    __arm64_sys_ppoll+0xa4/0x130

Note: This is not a 100% reproducible failure. In my automated loop the hang
reproduced on the 2nd boot attempt, but other attempts ran for 6–9 minutes
without hitting it. The bug is clearly timing-dependent on NUMA migration
activity during the L1 boot window.

## Root cause

The issue is in arch/arm64/kvm/nested.c:
  void kvm_nested_s2_unmap(struct kvm *kvm, bool may_block)
  {
      int i;

      lockdep_assert_held_write(&kvm->mmu_lock);

      if (!kvm->arch.nested_mmus_size)
          return;

      for (i = 0; i < kvm->arch.nested_mmus_size; i++) {
          struct kvm_s2_mmu *mmu = &kvm->arch.nested_mmus[i];

          if (kvm_s2_mmu_valid(mmu))
              kvm_stage2_unmap_range(mmu, 0, kvm_phys_size(mmu), may_block);
      }

      kvm_invalidate_vncr_ipa(kvm, 0, BIT(kvm->arch.mmu.pgt->ia_bits));
  }

When L0 NUMA balancing migrates a page belonging to the QEMU process, the
MMU notifier path calls kvm_unmap_gfn_range():

  bool kvm_unmap_gfn_range(struct kvm *kvm, struct kvm_gfn_range *range)
  {
      ...
      __unmap_stage2_range(&kvm->arch.mmu, range->start << PAGE_SHIFT,
                           (range->end - range->start) << PAGE_SHIFT,
                           range->may_block);

      kvm_nested_s2_unmap(kvm, range->may_block);   /* full unmap */
      return false;
  }

kvm_handle_hva_range() takes kvm->mmu_lock for writing before invoking the
handler and releases it only after the handler returns
(virt/kvm/kvm_main.c:622-642). Therefore, the entire kvm_nested_s2_unmap()
runs under kvm->mmu_lock.

The problem is that kvm_nested_s2_unmap() does not unmap only the affected
GPA range. Instead, it unmaps the entire IPA space (0 to kvm_phys_size(mmu))
for every nested S2 MMU. With nested virtualization enabled, this is very
expensive.

## Trace evidence

I captured function_graph traces for kvm_unmap_gfn_range,
kvm_nested_s2_unmap, and kvm_stage2_unmap_range, plus mm_migrate_pages
tracepoints.

1. numad-triggered full unmap
    64)   numad-1888   |               |  /* set_migration_pte:
  addr=fff790579000, pte=3040c319d9680 order=0 */
    64)   numad-1888   |               |  kvm_unmap_gfn_range() {
    64)   numad-1888   |               |    kvm_nested_s2_unmap() {
    64)   numad-1888   | @ 877415.7 us |      kvm_stage2_unmap_range();
    64)   numad-1888   | @ 877418.2 us |    }
    64)   numad-1888   | @ 877422.0 us |  }

Each set_migration_pte line is a single-page NUMA migration. Yet each
migration triggers one full kvm_nested_s2_unmap() that takes 877 ms.

Subsequent calls show per-page unmap durations between 841 ms and 1.16 s.

2. QEMU threads blocked as well
    23) qemu-sy-227844 |               |  kvm_unmap_gfn_range() {
    23) qemu-sy-227844 |               |    kvm_nested_s2_unmap() {
    23) qemu-sy-227844 | $ 1170521 us  |      kvm_stage2_unmap_range();
    23) qemu-sy-227844 | $ 1170529 us |    }
    23) qemu-sy-227844 | $ 1170542 us |  }

QEMU's own threads also get stuck in the same full unmap, with one call
reaching 6.67 s.

3. Statistics
┌──────────────┬───────┬──────────────────────────────────────┐
│ Thread       │ Calls │ kvm_nested_s2_unmap duration         │
├──────────────┼───────┼──────────────────────────────────────┤
│ numad-1888   │ 59    │ min 0.84 s / avg 1.02 s / max 1.16 s │
├──────────────┼───────┼──────────────────────────────────────┤
│ QEMU threads │ 34    │ min 1.14 s / avg 5.14 s / max 6.67 s │
└──────────────┴───────┴──────────────────────────────────────┘

numad migrates pages one after another; each page holds kvm->mmu_lock for
about one second. QEMU vCPU threads cannot acquire mmu_lock and also block
on migration_entry_wait. The L1 guest vCPUs make no forward progress, and
the watchdog fires.

Soft lockup threshold check

Read directly inside the L1 guest:

  root@ubuntu-vm:~# cat /proc/sys/kernel/watchdog_thresh
  10

So the soft lockup threshold is 2 * watchdog_thresh = 20 s.

The L1 guest reported stuck times of 29~32 s, which exceeds the 20 s
threshold.

## Conclusion

The root cause is confirmed: kvm_nested_s2_unmap() performs a full IPA space
unmap in the MMU notifier path instead of unmapping only the affected
GPA/CPAI range. The interval-tree-based precise range unmap approach is the
right fix.

Please consider applying the patch that replaces the full unmap with
kvm_nested_unmap_cipa_range() to avoid scanning the entire nested stage-2
page table on every NUMA migration.

Thanks,
Shuai



  reply	other threads:[~2026-09-04  7:01 UTC|newest]

Thread overview: 18+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-10 20:50 [PATCH v5 0/6] KVM: arm64: nv: Implement nested stage-2 reverse map (new data structure) Wei-Lin Chang
2026-08-10 20:50 ` [PATCH v5 1/6] KVM: arm64: Use a variable for the canonical IPA in kvm_s2_fault_map() Wei-Lin Chang
2026-08-10 20:50 ` [PATCH v5 2/6] KVM: arm64: nv: Introduce guest stage-2 tracking structures Wei-Lin Chang
2026-08-14  1:04   ` Itaru Kitayama
2026-08-14 10:42     ` Wei-Lin Chang
2026-08-16 22:01       ` Itaru Kitayama
2026-08-10 20:50 ` [PATCH v5 3/6] KVM: arm64: nv: Track guest stage-2 mapping creation Wei-Lin Chang
2026-08-10 20:50 ` [PATCH v5 4/6] KVM: arm64: nv: Track guest stage-2 mapping removal Wei-Lin Chang
2026-08-10 20:50 ` [PATCH v5 5/6] KVM: arm64: nv: Avoid full shadow stage-2 unmap Wei-Lin Chang
2026-08-10 20:50 ` [PATCH v5 6/6] KVM: arm64: Refactor kvm_unmap_gfn_range() with common variables Wei-Lin Chang
2026-08-12  2:12 ` [PATCH v5 0/6] KVM: arm64: nv: Implement nested stage-2 reverse map (new data structure) Itaru Kitayama
2026-09-02 16:35 ` Wang Han
2026-09-03  7:43   ` Marc Zyngier
2026-09-03 13:28     ` Wei-Lin Chang
2026-09-04  7:01       ` Shuai Xue [this message]
2026-09-04  7:54         ` Marc Zyngier
2026-09-04  7:49       ` Marc Zyngier
2026-09-04 11:37         ` Wei-Lin Chang

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=46342b48-550e-42c1-9d9f-e800c72269c2@linux.alibaba.com \
    --to=xueshuai@linux.alibaba.com \
    --cc=catalin.marinas@arm.com \
    --cc=itaru.kitayama@fujitsu.com \
    --cc=joey.gouly@arm.com \
    --cc=kvmarm@lists.linux.dev \
    --cc=linux-arm-kernel@lists.infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=ljs@kernel.org \
    --cc=maz@kernel.org \
    --cc=oupton@kernel.org \
    --cc=seiden@linux.ibm.com \
    --cc=suzuki.poulose@arm.com \
    --cc=tabba@google.com \
    --cc=wanghan@linux.alibaba.com \
    --cc=weilin.chang@arm.com \
    --cc=will@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).