From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from bombadil.infradead.org (bombadil.infradead.org [198.137.202.133]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id A0113C79F85 for ; Sun, 6 Sep 2026 10:39:50 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=lists.infradead.org; s=bombadil.20210309; h=Sender:List-Subscribe:List-Help :List-Post:List-Archive:List-Unsubscribe:List-Id:Content-Type:MIME-Version: References:In-Reply-To:Subject:Cc:To:From:Message-ID:Date:Reply-To: Content-Transfer-Encoding:Content-ID:Content-Description:Resent-Date: Resent-From:Resent-Sender:Resent-To:Resent-Cc:Resent-Message-ID:List-Owner; bh=OxZ+7J7RcHfxj4BJowNDu3kn/mniDmEyOVm3o0nqITc=; b=g777zOCa18R9mhjDfaaowgQYvC OnbPNcQNlLkEcTxO5z4FLN82lHxEomMqcMRpKlLyeqfnrf/Fd6S+2+e3tPA0BKVogh4l6aNJC1ZuW iHOo+r0ylVMadE6lDa70mde5CwJR/ySJeM5mKED7C/gsm6Rl+iaz1knTXy/Ij4u+wi/OsJdMzNj4U t+sJHWbIW7F9DZ3/+j7sQpWaJTDuoqYV7IcKAxJLrwbQWwX0YDZQbilp6IAE9mCwPvSvKPCnA81Az 86PxdnfR5CvaHSQglLJAxNsxncaW0JEBah3yU344cfYIkRbvDBw9BgcakP8wvnkdehE1+3j5xBTB8 l4oIwIMg==; Received: from localhost ([::1] helo=bombadil.infradead.org) by bombadil.infradead.org with esmtp (Exim 4.99.1 #2 (Red Hat Linux)) id 1x3AHi-00000004wsi-36NV; Sun, 06 Sep 2026 10:39:38 +0000 Received: from sea.source.kernel.org ([172.234.252.31]) by bombadil.infradead.org with esmtps (Exim 4.99.1 #2 (Red Hat Linux)) id 1x3AHh-00000004wsQ-2XZ8 for linux-arm-kernel@lists.infradead.org; Sun, 06 Sep 2026 10:39:37 +0000 Received: from smtp.kernel.org (quasi.space.kernel.org [100.103.45.18]) by sea.source.kernel.org (Postfix) with ESMTP id EAF0F40495; Sun, 6 Sep 2026 10:39:36 +0000 (UTC) Received: by smtp.kernel.org (Postfix) with ESMTPSA id A79BB1F00A3A; Sun, 6 Sep 2026 10:39:36 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788691176; bh=OxZ+7J7RcHfxj4BJowNDu3kn/mniDmEyOVm3o0nqITc=; h=Date:From:To:Cc:Subject:In-Reply-To:References; b=ab3JilaTIX5Wyh0ADurJ8C4Tqhs3MVFg3dsoCs3Ece/XbxlhgfcqlDplQur1wqUmw f35vDNtUU+vnAYkep2RF+sUl/bOPDNeYdE0X5Y5SMgeL/28lZzNzES84CsOC6BWvMX TjEM/Du+uEEIyPRtNOD2g1Ee5eyC9dYqo0pUeZWFoj7iz3KXJdK5bB2S3CTVEqs4x5 eQoe1L6E8FRpHRiC5caSNOP+88mjSSW9k8zI+YRM3HI7lLbuyTGKNQwOMRpsY6CZn6 Oxzkow4rRU7M43SQHTWWw56E+dU5G09g5rOC9oBdN8YSHq0T/qazsdBESKBSpsFdwv nzVDb4EKijxIg== Received: from sofa.misterjones.org ([185.219.108.64] helo=lobster-girl.misterjones.org) by disco-boy.misterjones.org with esmtpsa (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.98.2) (envelope-from ) id 1x3AHe-00000005Nkm-0cn5; Sun, 06 Sep 2026 10:39:34 +0000 Date: Sun, 06 Sep 2026 11:42:08 +0100 Message-ID: <87o6ea4q67.wl-maz@kernel.org> From: Marc Zyngier To: Shuai Xue Cc: Wei-Lin Chang , Wang Han , linux-arm-kernel@lists.infradead.org, kvmarm@lists.linux.dev, linux-kernel@vger.kernel.org, oupton@kernel.org, tabba@google.com, joey.gouly@arm.com, seiden@linux.ibm.com, suzuki.poulose@arm.com, catalin.marinas@arm.com, will@kernel.org, ljs@kernel.org, itaru.kitayama@fujitsu.com Subject: Re: [PATCH v5 0/6] KVM: arm64: nv: Implement nested stage-2 reverse map (new data structure) In-Reply-To: References: <20260810205038.118843-1-weilin.chang@arm.com> <20260902163500.1841671-1-wanghan@linux.alibaba.com> <877bl26aqg.wl-maz@kernel.org> <46342b48-550e-42c1-9d9f-e800c72269c2@linux.alibaba.com> <874ig55u5h.wl-maz@kernel.org> User-Agent: Wanderlust/2.15.9 (Almost Unreal) SEMI-EPG/1.14.7 (Harue) FLIM-LB/1.14.9 (=?UTF-8?B?R29qxY0=?=) APEL-LB/10.8 EasyPG/1.0.0 Emacs/30.1 (aarch64-unknown-linux-gnu) MULE/6.0 (HANACHIRUSATO) MIME-Version: 1.0 (generated by SEMI-EPG 1.14.7 - "Harue") Content-Type: text/plain; charset=US-ASCII X-SA-Exim-Connect-IP: 185.219.108.64 X-SA-Exim-Rcpt-To: xueshuai@linux.alibaba.com, weilin.chang@arm.com, wanghan@linux.alibaba.com, linux-arm-kernel@lists.infradead.org, kvmarm@lists.linux.dev, linux-kernel@vger.kernel.org, oupton@kernel.org, tabba@google.com, joey.gouly@arm.com, seiden@linux.ibm.com, suzuki.poulose@arm.com, catalin.marinas@arm.com, will@kernel.org, ljs@kernel.org, itaru.kitayama@fujitsu.com X-SA-Exim-Mail-From: maz@kernel.org X-SA-Exim-Scanned: No (on disco-boy.misterjones.org); SAEximRunCond expanded to false X-BeenThere: linux-arm-kernel@lists.infradead.org X-Mailman-Version: 2.1.34 Precedence: list List-Id: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Sender: "linux-arm-kernel" Errors-To: linux-arm-kernel-bounces+linux-arm-kernel=archiver.kernel.org@lists.infradead.org On Sat, 05 Sep 2026 16:35:01 +0100, Shuai Xue wrote: > > > > On 9/4/26 3:54 PM, Marc Zyngier wrote: > > On Fri, 04 Sep 2026 08:01:24 +0100, > > Shuai Xue wrote: > >> > >> > >> > >> On 9/3/26 9:28 PM, Wei-Lin Chang wrote: > >>> On Thu, Sep 03, 2026 at 08:43:35AM +0100, Marc Zyngier wrote: > >>>> On Wed, 02 Sep 2026 17:35:00 +0100, > >>>> Wang Han wrote: > >>>>> > >>>>> Hi Wei-Lin, > >>>>> > >>>>> I tested this series on a Yitian 710 system with an ARM Neoverse-N2 CPU > >>>>> (128 CPUs, 2 NUMA nodes). > >>>>> > >>>>> Test environment > >>>>> ---------------- > >>>>> > >>>>> L0 kernel: Linux v7.2-rc6 > >>>>> L1 guest: Ubuntu 26.04 LTS, kernel 7.0.0-27-generic (aarch64) > >>>>> QEMU: 10.2.3 > >>>>> > >>>>> L0 NUMA balancing was enabled (`/proc/sys/kernel/numa_balancing=1`). > >>>>> The host was booted with `kvm_arm.mode=nested`. > >>>>> > >>>>> This series fixes a functional hang that is exposed when NUMA balancing is > >>>>> enabled. The previous nested stage-2 unmap path is too slow for this > >>>>> workload, making the performance problem user-visible: NUMA balancing can > >>>>> leave the L1 guest unable to make progress and eventually hang during boot. > >>>>> > >>>>> The L1 was started with 8 vCPUs and 32 GiB of RAM using: > >>>>> > >>>>> qemu-system-aarch64 -smp 8 -m 32G \ > >>>>> -machine virt,accel=kvm,gic-version=3,virtualization=on \ > >>>>> -cpu host -nographic -enable-kvm \ > >>>>> -drive if=pflash,format=raw,readonly=on,file=pflash0_bak.img \ > >>>>> -drive if=pflash,format=raw,file=pflash1_bak.img \ > >>>>> -drive file=./ubuntu-vm.qcow2,format=qcow2,if=virtio,cache=none,aio=native \ > >>>>> -nic user,model=virtio-net-pci,hostfwd=tcp::11234-:22 \ > >>>>> -serial mon:stdio > >>>>> > >>>> > >>>> Puzzling. If you are only running an L1 in VHE mode, there is no > >>>> shadow S2, and therefore nothing to unmap. For shadow S2s to be built > >>>> and affect the MMU notifiers, you need to run an L2. > >>> > >>> I was thinking the same at first, but realized even with L1 in VHE mode > >>> there is a small period of time where L1 runs in its EL1 during boot, so > >>> one nested MMU will become valid for each vCPU. That causes > >>> kvm_nested_s2_unmap() to iterate through the entire IPA space 8 times > >>> (-smp 8). > >>> > >>> What I am curious about is whether one single notifier unmap is enough > >>> to hang L1, or were there multiple notifier unmaps. > >>> > >>> QEMU with -machine virt uses 40 IPA bits only, unmapping that takes: > >>> 1024 (4KB pages, unmapping 1GB per iteration) > >>> 32768 (16KB pages, unmapping 32MB per iteration) > >>> 2048 (64KB pages, unmapping 512MB per iteration) > >>> iterations for each page size. There aren't many mappings in each > >>> iteration too. Does this really take that long on real hardware (even if > >>> this must be done 8 times)? > >>> > >>> Thanks, > >>> Wei-Lin Chang > >>> > >>>> > >>>> So what are your actual test conditions? > >>>> > >>>> M. > >>>> > >> > >> Hi, Wei-Lin and Marc, > >> > >> I was able to reproduce this issue and capture ftrace evidence that confirms > >> the root cause. Below is the analysis, trace log, and timing data. > > > > [...] > > > >> Each set_migration_pte line is a single-page NUMA migration. Yet each > >> migration triggers one full kvm_nested_s2_unmap() that takes 877 ms. > > > > And why is it taking so long? It should be *empty* after the first > > iteration. > > Good question. > > After dive into the details trace, let to try to answer the question. > > Yes, it is empty -- and that is exactly the point: the 877ms is paid *for* > an empty table. > > The cost is not in the walk and not in clearing PTEs; > **it is 262,144 broadcast TLB invalidations**, one at the end of every > 1GB chunk, each ~3.3us. Since v6.6 the cost of an unmap is > proportional to the size of the IPA range, not to the number of > mappings; an empty table pays in full. > > > > > [...] > > > >> ## Conclusion > >> > >> The root cause is confirmed: kvm_nested_s2_unmap() performs a full IPA space > >> unmap in the MMU notifier path instead of unmapping only the affected > >> GPA/CPAI range. The interval-tree-based precise range unmap approach is the > >> right fix. > > > > No. This just indicates that this is papering over a bigger problem, > > and your AI is jumping to conclusions. > > Sorry for the jumping up. > > The culprit is 7657ea920c54 ("KVM: arm64: Use TLBI range-based > instructions for unmap", v6.6). kvm_pgtable_stage2_unmap() ends > *every* call with an unconditional kvm_tlb_flush_vmid_range(), whether > or not the walk cleared a single PTE: > > ret = kvm_pgtable_walk(pgt, addr, size, &walker); > if (stage2_unmap_defer_tlb_flush(pgt)) > /* Perform the deferred TLB invalidations */ > kvm_tlb_flush_vmid_range(pgt->mmu, addr, size); > > stage2_apply_range() calls it once per 1GB chunk, and the nested MMU > covers the guest PARange -- 48 bits here, so 262,144 calls per > kvm_nested_s2_unmap(). Each broadcast is one IPAS2E1IS (range) plus one VMALLE1IS > plus two DSB(ish), ~3.2-3.4us without ftrace: > > 877ms / 262,144 chunks = 3.35us per chunk > > which is exactly the per-broadcast cost. Right. That's pretty compelling, thanks for digging into this. Your proposed approach (counting the invalidated regions) is interesting, but I don't think it is the correct one. The real issue here is that we treat a full S2 unmap as if it was a set of ranges. This is what needs fixing, because we can invalidate the whole thing with exactly *ONE* TLBI. This is even more important once you run an L2, as L1 will also perform its own TLB invalidation, and we want to avoid having trapping pointlessly. This also propagates in the way we handle TLB emulation. [...] > The candidate fix is below. Table entries > (KVM_PGTABLE_WALK_TABLE_POST) are unaffected: stage2_unmap_put_pte() > keeps issuing the immediate __kvm_tlb_flush_vmid_ipa() for them, and > their child leaves are counted as leaves within the same walk, so any > call that clears something still flushes a superset of what it > cleared. This restores the pre-v6.6 semantics -- cost proportional to > what is mapped. With it, an empty nested unmap costs ~110ms > instrumented (the pure walk); getting to "exactly zero" would > additionally require kvm_nested_s2_unmap() to skip nested MMUs that > are valid but empty. That's an interesting remark. I guess we could add some extra tracking for that, but let's see what we can do about the above first. I've hacked something together and pushed the result at [1] (compile-tested only). I'd appreciate it if you could put it to the test with your setup. Thanks, M. [1] https://web.git.kernel.org/pub/scm/linux/kernel/git/maz/arm-platforms.git/log/?h=kvm-arm64/unmap-vmall -- Jazz isn't dead. It just smells funny.