From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from bombadil.infradead.org (bombadil.infradead.org [198.137.202.133]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id 1500AC88E50 for ; Fri, 11 Sep 2026 15:27:19 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=lists.infradead.org; s=bombadil.20210309; h=Sender:List-Subscribe:List-Help :List-Post:List-Archive:List-Unsubscribe:List-Id:Content-Transfer-Encoding: Content-Type:MIME-Version:References:In-Reply-To:Subject:Cc:To:From: Message-ID:Date:Reply-To:Content-ID:Content-Description:Resent-Date: Resent-From:Resent-Sender:Resent-To:Resent-Cc:Resent-Message-ID:List-Owner; bh=dL5SIEsYkL5CFvwQMkla31xLNVufHe8l4HQ4/Pcc7zE=; b=DBbfcFTV7Kt+fwVf+TrNhUITzn t8Lnk5Po4NwEyB7lkE/DobgzQ8Jbw6Pd27mIY35C0CXEe/H4MfifYQPfuCRHEbB4sOqz378ow+KCw PhAVzVu5m2PuhZl6oWaCuDvWRAGk+n7Ux0omwtxf1cVmycJdXgA1zOUt2YXHvjd+Shx4hWfmRUI03 k3cSVBl7q2oUHBnEraIyOgcIzBJffhfhtNn4AshwGNTEA3fQdvgrtlQCxnJLSa1/qg8CKBDOBezaz u14B6+YcLjnWB4q1eIIwGL/Act9wN1Igs5ycs0NWdWVBVHYm0CXFhN2aFJuWBCNVu0Ud/5qqEPT3E HBkI53RQ==; Received: from localhost ([::1] helo=bombadil.infradead.org) by bombadil.infradead.org with esmtp (Exim 4.99.1 #2 (Red Hat Linux)) id 1x539i-0000000H38H-3Cgv; Fri, 11 Sep 2026 15:27:10 +0000 Received: from tor.source.kernel.org ([172.105.4.254]) by bombadil.infradead.org with esmtps (Exim 4.99.1 #2 (Red Hat Linux)) id 1x539e-0000000H37x-1kgA for linux-arm-kernel@lists.infradead.org; Fri, 11 Sep 2026 15:27:09 +0000 Received: from smtp.kernel.org (quasi.space.kernel.org [100.103.45.18]) by tor.source.kernel.org (Postfix) with ESMTP id 85CEE60202; Fri, 11 Sep 2026 15:27:05 +0000 (UTC) Received: by smtp.kernel.org (Postfix) with ESMTPSA id 1B92A1F000FF; Fri, 11 Sep 2026 15:27:05 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1789140425; bh=dL5SIEsYkL5CFvwQMkla31xLNVufHe8l4HQ4/Pcc7zE=; h=Date:From:To:Cc:Subject:In-Reply-To:References; b=agMogg7MfwE6qQ+hq1iJshNCB1YrISD9q+Xy8K8HxRCjrJ0szmZjoDYNFODVMnW4+ GMbyqli2WHZJOyRSEiXb6zMLnfw318oPj2C20GU4HuQP9i0KVdc9WYACfa9qfLN8GA wovSL8neofGZ6d76PawnIDdKHOdlQyMWQA7UFCL4frFJQZ2LGQxpKoofh+ZkgN5+sg SlhL9ITp/qb3mDs2QrXyx9OHVX6JM/tlmRMZr67pEvvhhTG+04sAdJU45vqiQjKHqd +GE7M9+x+EL5iTJzZq81pNnu8aPqGBTMqROfNZXnMoolXtRvC2r5yiX8PB2K4y0W67 +SqV6khmMSqeA== Received: from sofa.misterjones.org ([185.219.108.64] helo=goblin-girl.misterjones.org) by disco-boy.misterjones.org with esmtpsa (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.98.2) (envelope-from ) id 1x539b-00000007p44-0ROi; Fri, 11 Sep 2026 15:27:03 +0000 Date: Fri, 11 Sep 2026 16:27:02 +0100 Message-ID: <86h5jv7qrd.wl-maz@kernel.org> From: Marc Zyngier To: Wei-Lin Chang Cc: linux-arm-kernel@lists.infradead.org, kvmarm@lists.linux.dev, linux-kernel@vger.kernel.org, Oliver Upton , Fuad Tabba , Joey Gouly , Steffen Eiden , Suzuki K Poulose , Zenghui Yu , Catalin Marinas , Will Deacon , Lorenzo Stoakes , Itaru Kitayama Subject: Re: [PATCH v5 2/6] KVM: arm64: nv: Introduce guest stage-2 tracking structures In-Reply-To: References: <20260810205038.118843-1-weilin.chang@arm.com> <20260810205038.118843-3-weilin.chang@arm.com> <87mrtu4c21.wl-maz@kernel.org> User-Agent: Wanderlust/2.15.9 (Almost Unreal) SEMI-EPG/1.14.7 (Harue) FLIM-LB/1.14.9 (=?UTF-8?B?R29qxY0=?=) APEL-LB/10.8 EasyPG/1.0.0 Emacs/30.1 (aarch64-unknown-linux-gnu) MULE/6.0 (HANACHIRUSATO) MIME-Version: 1.0 (generated by SEMI-EPG 1.14.7 - "Harue") Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: quoted-printable X-SA-Exim-Connect-IP: 185.219.108.64 X-SA-Exim-Rcpt-To: weilin.chang@arm.com, linux-arm-kernel@lists.infradead.org, kvmarm@lists.linux.dev, linux-kernel@vger.kernel.org, oupton@kernel.org, tabba@google.com, joey.gouly@arm.com, seiden@linux.ibm.com, suzuki.poulose@arm.com, yuzenghui@huawei.com, catalin.marinas@arm.com, will@kernel.org, ljs@kernel.org, itaru.kitayama@fujitsu.com X-SA-Exim-Mail-From: maz@kernel.org X-SA-Exim-Scanned: No (on disco-boy.misterjones.org); SAEximRunCond expanded to false X-BeenThere: linux-arm-kernel@lists.infradead.org X-Mailman-Version: 2.1.34 Precedence: list List-Id: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Sender: "linux-arm-kernel" Errors-To: linux-arm-kernel-bounces+linux-arm-kernel=archiver.kernel.org@lists.infradead.org On Sun, 06 Sep 2026 20:56:34 +0100, Wei-Lin Chang wrote: >=20 > On Sun, Sep 06, 2026 at 04:47:02PM +0100, Marc Zyngier wrote: >=20 > [...] >=20 > > > +/* > > > + * Record of a guest stage-2 mapping, storing canonical and nested I= PA > > > + * ranges. Both ranges have the same size. > > > + */ > > > +struct kvm_guest_s2_mapping { > > > + struct interval_tree_node canonical; > > > + struct interval_tree_node nested; > > > + struct kvm_s2_mmu *nested_mmu; > > > +}; > > > + > >=20 > > I'm trying hard to find a way to reduce the size of this structure, > > because this is IMO the only real problem with this approach. > >=20 > > Obviously, the only thing we could kill is this nested_mmu field, as > > everything else is used by the interval trees. I can see two ugly ways > > to do that: > >=20 > > - either we iterate over all shadow MMUs to find the corresponding > > 'nested' node: really costly if we have a lot of mappings and/or a > > lot of shadow MMUs > >=20 > > - or we steal bits from the interval_tree_node to encode extra > > information. One realisation is that all addresses are PAGE_SIZE > > aligned, meaning that we have at least 12 bits that are always 0. We > > could, for example, encode an index in the bottom bits of the > > nested.start field. Probably easy enough, but may require some > > careful masking (and the addition of an index in the s2_mmu > > structure). > >=20 > > The result would be significant, as we could then use a 96 byte slab, > > which means the cost of a 1GB @4k granularity could fall to 24MB. >=20 > One way of meeting in the middle is to keep the nested_mmu pointer, and > use a dedicated kmem cache. That cache will then hand out exact 104-byte > objects. >=20 > That's 26MB for 1GB @4K. >=20 > This is a nice improvement from the current 128-byte allocations. > Although it's not as good as 96-byte allocations, the code logic can > remain simple. Would you take this approach? I don't think the logic has to be complicated to go all the way to 96 bytes, see the patch below. I hope we can further trim it over time, but this would be a good start. Thanks, M. =46rom 39e3bdaa1b02c460a1001a23a20dccca5501eb6d Mon Sep 17 00:00:00 2001 From: Marc Zyngier Date: Tue, 8 Sep 2026 17:07:50 +0100 Subject: [PATCH] KVM: arm64: Drop kvm_s2_mmu pointer from kvm_guest_s2_mapp= ing As it appears that the kvm_guest_s2_mapping structure is quite large, and results in a 128 byte slab allocation, there is some insentive to shrink a bit. For this, replace the S2 MMU back-pointer with an index tucked into the low bits of the nested.start field. This allows us to shrink the structure by 8 bytes, and therefore to fit in a 96 byte slab. The index is also stored in the S2 MMU structure itself, which comes for free as it fits in an existing hole. The structure is also repacked to avoid other disgracious holes. Signed-off-by: Marc Zyngier --- arch/arm64/include/asm/kvm_host.h | 16 ++++++----- arch/arm64/kvm/nested.c | 44 +++++++++++++++++++++++++------ 2 files changed, 45 insertions(+), 15 deletions(-) diff --git a/arch/arm64/include/asm/kvm_host.h b/arch/arm64/include/asm/kvm= _host.h index f12883a420817..4b3fc4d4a61f2 100644 --- a/arch/arm64/include/asm/kvm_host.h +++ b/arch/arm64/include/asm/kvm_host.h @@ -158,7 +158,6 @@ struct kvm_vmid { struct kvm_guest_s2_mapping { struct interval_tree_node canonical; struct interval_tree_node nested; - struct kvm_s2_mmu *nested_mmu; }; =20 struct kvm_s2_mmu { @@ -222,30 +221,33 @@ struct kvm_s2_mmu { u64 tlb_vttbr; u64 tlb_vtcr; =20 + /* Guest s2 mapping records indexed in this MMU's IPA space. */ + struct rb_root_cached guest_s2_mappings; + /* * true when this represents a nested context where virtual * HCR_EL2.VM =3D=3D 1 */ bool nested_stage2_enabled; =20 -#ifdef CONFIG_PTDUMP_STAGE2_DEBUGFS - struct dentry *shadow_pt_debugfs_dentry; -#endif - /* * true when this MMU needs to be unmapped before being used for a new * purpose. */ bool pending_unmap; =20 - /* Guest s2 mapping records indexed in this MMU's IPA space. */ - struct rb_root_cached guest_s2_mappings; + /* Index in the S2 MMU array, only valid for a shadow S2 */ + u16 s2_mmu_idx; =20 /* * 0: Nobody is currently using this, check vttbr for validity * >0: Somebody is actively using this. */ atomic_t refcnt; + +#ifdef CONFIG_PTDUMP_STAGE2_DEBUGFS + struct dentry *shadow_pt_debugfs_dentry; +#endif }; =20 struct kvm_arch_memory_slot { diff --git a/arch/arm64/kvm/nested.c b/arch/arm64/kvm/nested.c index e0941a261787c..540e52db9f180 100644 --- a/arch/arm64/kvm/nested.c +++ b/arch/arm64/kvm/nested.c @@ -45,11 +45,12 @@ struct vncr_tlb { * will invalidate them more often). */ #define S2_MMU_PER_VCPU 2 +#define S2_MMU_PER_VM (KVM_MAX_VCPUS * S2_MMU_PER_VCPU) =20 int kvm_init_nested(struct kvm *kvm) { kvm->arch.nested_mmus =3D kvmalloc_objs(struct kvm_s2_mmu *, - KVM_MAX_VCPUS * S2_MMU_PER_VCPU, + S2_MMU_PER_VM, GFP_KERNEL_ACCOUNT); kvm->arch.nested_mmus_size =3D 0; atomic_set(&kvm->arch.vncr_tlb_count, 0); @@ -128,8 +129,10 @@ int kvm_vcpu_init_nested(struct kvm_vcpu *vcpu) =20 guard(write_lock)(&kvm->mmu_lock); =20 - for (i =3D 0; i < S2_MMU_PER_VCPU; i++) + for (i =3D 0; i < S2_MMU_PER_VCPU; i++) { + tmp[i].s2_mmu_idx =3D i + kvm->arch.nested_mmus_size; kvm->arch.nested_mmus[i + kvm->arch.nested_mmus_size] =3D &tmp[i]; + } =20 kvm->arch.nested_mmus_size +=3D S2_MMU_PER_VCPU; } @@ -873,6 +876,27 @@ static struct kvm_s2_mmu *get_s2_mmu_nested(struct kvm= _vcpu *vcpu) return s2_mmu; } =20 +#define S2_MMU_IDX_MASK GENMASK_ULL(const_ilog2(S2_MMU_PER_VM) - 1, 0) + +static void tag_s2_mapping_mmu(struct kvm_guest_s2_mapping *mapping, + struct kvm_s2_mmu *mmu) +{ + BUILD_BUG_ON(const_ilog2(S2_MMU_PER_VM) >=3D 12); + mapping->nested.start &=3D ~S2_MMU_IDX_MASK; + mapping->nested.start |=3D mmu->s2_mmu_idx; +} + +static struct kvm_s2_mmu *s2_mapping_to_mmu(struct kvm *kvm, + struct kvm_guest_s2_mapping *mapping) +{ + return kvm->arch.nested_mmus[mapping->nested.start & S2_MMU_IDX_MASK]; +} + +static unsigned long s2_mapping_to_nested_start(struct kvm_guest_s2_mappin= g *mapping) +{ + return mapping->nested.start & ~S2_MMU_IDX_MASK; +} + void kvm_record_guest_s2_mapping(struct kvm_s2_mmu *mmu, gpa_t canonical_i= pa, gpa_t nested_ipa, size_t map_size, struct kvm_guest_s2_mapping *mapping) @@ -890,7 +914,7 @@ void kvm_record_guest_s2_mapping(struct kvm_s2_mmu *mmu= , gpa_t canonical_ipa, mapping->nested.start =3D nested_ipa; mapping->nested.last =3D nested_ipa + map_size - 1; =20 - mapping->nested_mmu =3D mmu; + tag_s2_mapping_mmu(mapping, mmu); =20 guard(spinlock)(&kvm->arch.guest_s2_tracking_lock); interval_tree_insert(&mapping->nested, &mmu->guest_s2_mappings); @@ -1357,17 +1381,21 @@ void kvm_nested_unmap_cipa_range(struct kvm *kvm, g= pa_t cipa, size_t unmap_size, =20 while ((node =3D interval_tree_iter_first(&kvm->arch.mmu.guest_s2_mapping= s, cipa, cipa_end))) { + unsigned long nested_start; + struct kvm_s2_mmu *mmu; + mapping =3D container_of(node, struct kvm_guest_s2_mapping, canonical); - mapping_size =3D mapping->nested.last - mapping->nested.start + 1; + nested_start =3D s2_mapping_to_nested_start(mapping); + mmu =3D s2_mapping_to_mmu(kvm, mapping); =20 - if (WARN_ON_ONCE(kvm_pgtable_stage2_unmap(mapping->nested_mmu->pgt, - mapping->nested.start, - mapping_size))) + mapping_size =3D mapping->nested.last - nested_start + 1; + + if (WARN_ON_ONCE(kvm_pgtable_stage2_unmap(mmu->pgt, nested_start, mappin= g_size))) return; =20 interval_tree_remove(node, &kvm->arch.mmu.guest_s2_mappings); - interval_tree_remove(&mapping->nested, &mapping->nested_mmu->guest_s2_ma= ppings); + interval_tree_remove(&mapping->nested, &mmu->guest_s2_mappings); kfree(mapping); =20 if (may_block) --=20 2.47.3 --=20 Without deviation from the norm, progress is not possible.