* [PATCH 0/4] KVM: arm64: Reduce overhead of full S2 teardown
@ 2026-09-12 10:48 Marc Zyngier
2026-09-12 10:48 ` [PATCH 1/4] KVM: arm64: pgtable: Add Stage-2 unmap without TLBI primitive Marc Zyngier
` (5 more replies)
0 siblings, 6 replies; 12+ messages in thread
From: Marc Zyngier @ 2026-09-12 10:48 UTC (permalink / raw)
To: kvmarm, linux-arm-kernel
Cc: Wei-Lin Chang, Wang Han, Shuai Xue, Steffen Eiden, Joey Gouly,
Suzuki K Poulose, Oliver Upton, Zenghui Yu, Fuad Tabba
Tearing down a full S2 is a pretty involved process, resulting in a
lot of TLB invalidation. These TLBIs are either on a per leaf basis if
the HW doesn't support range invalidation, or by top-level range if it
does. Amusingly, the latter occurs even when nothing has been
unmapped.
Things are made worse with NV, as we have a bucket-load of shadow S2s,
and the need to invalidate them all on the back of an MMU notifier.
The latter will eventually be solved by the reverse-map tracking that
Wei-Lin is working on, but we need to be better at full-S2 teardown.
This small series adds a "no TLBI" unmapping primitive, which allows
the caller to then whack the TLBs using a VMID-wide invalidation. This
results in far fewer TLBIs, and a better recursive virtualisation as
we get far fewer traps as a consequence.
This applies on top of my shadow-s2 lifetime fixes, and is expected to
be a prefix to Wei-Lin's series.
Marc Zyngier (4):
KVM: arm64: pgtable: Add Stage-2 unmap without TLBI primitive
KVM: arm64: MMU: Add kvm_stage2_unmap_all() helper
KVM: arm64: nv: Move full s2_mmu unmap over to kvm_stage2_unmap_all()
KVM: arm64: nv: Move TLBI VMALLS12E1* emulation over to
kvm_stage2_unmap_all()
arch/arm64/include/asm/kvm_mmu.h | 1 +
arch/arm64/include/asm/kvm_pgtable.h | 17 +++++++++++++
arch/arm64/include/asm/kvm_pkvm.h | 1 +
arch/arm64/kvm/hyp/pgtable.c | 38 +++++++++++++++++++++-------
arch/arm64/kvm/mmu.c | 15 +++++++++++
arch/arm64/kvm/nested.c | 4 +--
arch/arm64/kvm/pkvm.c | 2 ++
arch/arm64/kvm/sys_regs.c | 18 ++++++-------
8 files changed, 75 insertions(+), 21 deletions(-)
--
2.47.3
^ permalink raw reply [flat|nested] 12+ messages in thread
* [PATCH 1/4] KVM: arm64: pgtable: Add Stage-2 unmap without TLBI primitive
2026-09-12 10:48 [PATCH 0/4] KVM: arm64: Reduce overhead of full S2 teardown Marc Zyngier
@ 2026-09-12 10:48 ` Marc Zyngier
2026-09-14 8:48 ` Mark Rutland
2026-09-12 10:48 ` [PATCH 2/4] KVM: arm64: MMU: Add kvm_stage2_unmap_all() helper Marc Zyngier
` (4 subsequent siblings)
5 siblings, 1 reply; 12+ messages in thread
From: Marc Zyngier @ 2026-09-12 10:48 UTC (permalink / raw)
To: kvmarm, linux-arm-kernel
Cc: Wei-Lin Chang, Wang Han, Shuai Xue, Steffen Eiden, Joey Gouly,
Suzuki K Poulose, Oliver Upton, Zenghui Yu, Fuad Tabba
kvm_pgtable_stage2_unmap() iterates over a range, unmapping whatever is
within the range, and always guarantees that that the corresponding TLBs
are invalidated when the function returns.
While this is safe, it means that iterating over empty range on a system
that supports range invalidation results in a TLBI per largest block
mapping size (1GB, 32MB or 512MB, depending on the base granule size).
This can be pretty expensive in situation where the whole address space
is being torn down, as it happens with NV (where S2 MMUs are recycled
regularly), and it would be more efficient to elide the per-subrange
TLBIs to solely rely on a VMID-wide TLBI.
For this, provide a kvm_pgtable_stage2_unmap_notlbi() helper that elides
all TLBIs, and relies on the caller to do the work.
Note that for pKVM case, no additional helper is provided, and we
fallback on the TLBI-aware version.
Signed-off-by: Marc Zyngier <maz@kernel.org>
---
arch/arm64/include/asm/kvm_pgtable.h | 17 +++++++++++++
arch/arm64/include/asm/kvm_pkvm.h | 1 +
arch/arm64/kvm/hyp/pgtable.c | 38 +++++++++++++++++++++-------
arch/arm64/kvm/pkvm.c | 2 ++
4 files changed, 49 insertions(+), 9 deletions(-)
diff --git a/arch/arm64/include/asm/kvm_pgtable.h b/arch/arm64/include/asm/kvm_pgtable.h
index 41a8687938eb6..c370196888d1d 100644
--- a/arch/arm64/include/asm/kvm_pgtable.h
+++ b/arch/arm64/include/asm/kvm_pgtable.h
@@ -318,6 +318,8 @@ typedef bool (*kvm_pgtable_force_pte_cb_t)(u64 addr, u64 end,
* @KVM_PGTABLE_WALK_SKIP_CMO: Visit and update table entries
* without Cache maintenance
* operations required.
+ * @KVM_PGTABLE_WALK_SKIP_S2_TLBI: Visit and update table entries
+ * without Stage-2 TLB invalidation.
*/
enum kvm_pgtable_walk_flags {
KVM_PGTABLE_WALK_LEAF = BIT(0),
@@ -327,6 +329,7 @@ enum kvm_pgtable_walk_flags {
KVM_PGTABLE_WALK_IGNORE_EAGAIN = BIT(4),
KVM_PGTABLE_WALK_SKIP_BBM_TLBI = BIT(5),
KVM_PGTABLE_WALK_SKIP_CMO = BIT(6),
+ KVM_PGTABLE_WALK_SKIP_S2_TLBI = BIT(7),
};
struct kvm_pgtable_visit_ctx {
@@ -717,6 +720,20 @@ int kvm_pgtable_stage2_annotate(struct kvm_pgtable *pgt, u64 addr, u64 size,
*/
int kvm_pgtable_stage2_unmap(struct kvm_pgtable *pgt, u64 addr, u64 size);
+/**
+ * kvm_pgtable_stage2_unmap_notlbi() - Remove a mapping from a guest stage-2 page-table
+ * without TLB invalidation.
+ * @pgt: Page-table structure initialised by kvm_pgtable_stage2_init*().
+ * @addr: Intermediate physical address from which to remove the mapping.
+ * @size: Size of the mapping.
+ *
+ * Same as kvm_pgtable_stage2_unmap(), but does not invalidate the
+ * TLBs, which is the responsibility of the caller. Use with caution!
+ *
+ * Return: 0 on success, negative error code on failure.
+ */
+int kvm_pgtable_stage2_unmap_notlbi(struct kvm_pgtable *pgt, u64 addr, u64 size);
+
/**
* kvm_pgtable_stage2_wrprotect() - Write-protect guest stage-2 address range
* without TLB invalidation.
diff --git a/arch/arm64/include/asm/kvm_pkvm.h b/arch/arm64/include/asm/kvm_pkvm.h
index beea00e693a0a..273013c98ff17 100644
--- a/arch/arm64/include/asm/kvm_pkvm.h
+++ b/arch/arm64/include/asm/kvm_pkvm.h
@@ -214,6 +214,7 @@ int pkvm_pgtable_stage2_map(struct kvm_pgtable *pgt, u64 addr, u64 size, u64 phy
enum kvm_pgtable_prot prot, void *mc,
enum kvm_pgtable_walk_flags flags);
int pkvm_pgtable_stage2_unmap(struct kvm_pgtable *pgt, u64 addr, u64 size);
+int pkvm_pgtable_stage2_unmap_notlbi(struct kvm_pgtable *pgt, u64 addr, u64 size);
int pkvm_pgtable_stage2_wrprotect(struct kvm_pgtable *pgt, u64 addr, u64 size);
int pkvm_pgtable_stage2_flush(struct kvm_pgtable *pgt, u64 addr, u64 size);
bool pkvm_pgtable_stage2_test_clear_young(struct kvm_pgtable *pgt, u64 addr, u64 size, bool mkold);
diff --git a/arch/arm64/kvm/hyp/pgtable.c b/arch/arm64/kvm/hyp/pgtable.c
index b74dd5ce1efd3..6603fc236daa2 100644
--- a/arch/arm64/kvm/hyp/pgtable.c
+++ b/arch/arm64/kvm/hyp/pgtable.c
@@ -29,6 +29,11 @@ static bool kvm_pgtable_walk_skip_cmo(const struct kvm_pgtable_visit_ctx *ctx)
return unlikely(ctx->flags & KVM_PGTABLE_WALK_SKIP_CMO);
}
+static bool kvm_pgtable_walk_skip_s2_tlbi(const struct kvm_pgtable_visit_ctx *ctx)
+{
+ return unlikely(ctx->flags & KVM_PGTABLE_WALK_SKIP_S2_TLBI);
+}
+
static bool kvm_block_mapping_supported(const struct kvm_pgtable_visit_ctx *ctx, u64 phys)
{
u64 granule = kvm_granule_size(ctx->level);
@@ -905,12 +910,14 @@ static void stage2_unmap_put_pte(const struct kvm_pgtable_visit_ctx *ctx,
if (kvm_pte_valid(ctx->old)) {
kvm_clear_pte(ctx->ptep);
- if (kvm_pte_table(ctx->old, ctx->level)) {
- kvm_call_hyp(__kvm_tlb_flush_vmid_ipa, mmu, ctx->addr,
- TLBI_TTL_UNKNOWN);
- } else if (!stage2_unmap_defer_tlb_flush(pgt)) {
- kvm_call_hyp(__kvm_tlb_flush_vmid_ipa, mmu, ctx->addr,
- ctx->level);
+ if (!kvm_pgtable_walk_skip_s2_tlbi(ctx)) {
+ if (kvm_pte_table(ctx->old, ctx->level)) {
+ kvm_call_hyp(__kvm_tlb_flush_vmid_ipa, mmu, ctx->addr,
+ TLBI_TTL_UNKNOWN);
+ } else if (!stage2_unmap_defer_tlb_flush(pgt)) {
+ kvm_call_hyp(__kvm_tlb_flush_vmid_ipa, mmu, ctx->addr,
+ ctx->level);
+ }
}
}
@@ -1195,23 +1202,36 @@ static int stage2_unmap_walker(const struct kvm_pgtable_visit_ctx *ctx,
return 0;
}
-int kvm_pgtable_stage2_unmap(struct kvm_pgtable *pgt, u64 addr, u64 size)
+static int __kvm_pgtable_stage2_unmap(struct kvm_pgtable *pgt,
+ enum kvm_pgtable_walk_flags flags,
+ u64 addr, u64 size)
{
int ret;
struct kvm_pgtable_walker walker = {
.cb = stage2_unmap_walker,
.arg = pgt,
- .flags = KVM_PGTABLE_WALK_LEAF | KVM_PGTABLE_WALK_TABLE_POST,
+ .flags = KVM_PGTABLE_WALK_LEAF | KVM_PGTABLE_WALK_TABLE_POST | flags,
};
ret = kvm_pgtable_walk(pgt, addr, size, &walker);
- if (stage2_unmap_defer_tlb_flush(pgt))
+ if (stage2_unmap_defer_tlb_flush(pgt) &&
+ !(flags & KVM_PGTABLE_WALK_SKIP_S2_TLBI))
/* Perform the deferred TLB invalidations */
kvm_tlb_flush_vmid_range(pgt->mmu, addr, size);
return ret;
}
+int kvm_pgtable_stage2_unmap(struct kvm_pgtable *pgt, u64 addr, u64 size)
+{
+ return __kvm_pgtable_stage2_unmap(pgt, 0, addr, size);
+}
+
+int kvm_pgtable_stage2_unmap_notlbi(struct kvm_pgtable *pgt, u64 addr, u64 size)
+{
+ return __kvm_pgtable_stage2_unmap(pgt, KVM_PGTABLE_WALK_SKIP_S2_TLBI, addr, size);
+}
+
struct stage2_attr_data {
kvm_pte_t attr_set;
kvm_pte_t attr_clr;
diff --git a/arch/arm64/kvm/pkvm.c b/arch/arm64/kvm/pkvm.c
index 8e4c6e4bec123..ec151005fbe4d 100644
--- a/arch/arm64/kvm/pkvm.c
+++ b/arch/arm64/kvm/pkvm.c
@@ -488,6 +488,8 @@ int pkvm_pgtable_stage2_unmap(struct kvm_pgtable *pgt, u64 addr, u64 size)
return __pkvm_pgtable_stage2_unshare(pgt, addr, addr + size);
}
+int pkvm_pgtable_stage2_unmap_notlbi(struct kvm_pgtable *pgt, u64 addr, u64 size) __alias(pkvm_pgtable_stage2_unmap);
+
int pkvm_pgtable_stage2_wrprotect(struct kvm_pgtable *pgt, u64 addr, u64 size)
{
struct kvm *kvm = kvm_s2_mmu_to_kvm(pgt->mmu);
--
2.47.3
^ permalink raw reply related [flat|nested] 12+ messages in thread
* [PATCH 2/4] KVM: arm64: MMU: Add kvm_stage2_unmap_all() helper
2026-09-12 10:48 [PATCH 0/4] KVM: arm64: Reduce overhead of full S2 teardown Marc Zyngier
2026-09-12 10:48 ` [PATCH 1/4] KVM: arm64: pgtable: Add Stage-2 unmap without TLBI primitive Marc Zyngier
@ 2026-09-12 10:48 ` Marc Zyngier
2026-09-12 10:48 ` [PATCH 3/4] KVM: arm64: nv: Move full s2_mmu unmap over to kvm_stage2_unmap_all() Marc Zyngier
` (3 subsequent siblings)
5 siblings, 0 replies; 12+ messages in thread
From: Marc Zyngier @ 2026-09-12 10:48 UTC (permalink / raw)
To: kvmarm, linux-arm-kernel
Cc: Wei-Lin Chang, Wang Han, Shuai Xue, Steffen Eiden, Joey Gouly,
Suzuki K Poulose, Oliver Upton, Zenghui Yu, Fuad Tabba
Build on top of kvm_pgtable_stage2_unmap_notlbi() to provide a primitive
tearing down a whole Stage-2 MMU and invalidating the corresponding TLBs
by VMID, which is far more efficient than performing range invalidation
(or even worse, single mappings).
Signed-off-by: Marc Zyngier <maz@kernel.org>
---
arch/arm64/include/asm/kvm_mmu.h | 1 +
arch/arm64/kvm/mmu.c | 15 +++++++++++++++
2 files changed, 16 insertions(+)
diff --git a/arch/arm64/include/asm/kvm_mmu.h b/arch/arm64/include/asm/kvm_mmu.h
index 6eae7e7e2a684..f450557d4b7b3 100644
--- a/arch/arm64/include/asm/kvm_mmu.h
+++ b/arch/arm64/include/asm/kvm_mmu.h
@@ -171,6 +171,7 @@ void __init free_hyp_pgds(void);
void kvm_stage2_unmap_range(struct kvm_s2_mmu *mmu, phys_addr_t start,
u64 size, bool may_block);
+void kvm_stage2_unmap_all(struct kvm_s2_mmu *mmu, bool may_block);
void kvm_stage2_flush_range(struct kvm_s2_mmu *mmu, phys_addr_t addr, phys_addr_t end);
void kvm_stage2_wp_range(struct kvm_s2_mmu *mmu, phys_addr_t addr, phys_addr_t end);
diff --git a/arch/arm64/kvm/mmu.c b/arch/arm64/kvm/mmu.c
index 9ba86450fe4af..68bfd09870b76 100644
--- a/arch/arm64/kvm/mmu.c
+++ b/arch/arm64/kvm/mmu.c
@@ -346,6 +346,21 @@ void kvm_stage2_unmap_range(struct kvm_s2_mmu *mmu, phys_addr_t start,
__unmap_stage2_range(mmu, start, size, may_block);
}
+void kvm_stage2_unmap_all(struct kvm_s2_mmu *mmu, bool may_block)
+{
+ struct kvm *kvm = kvm_s2_mmu_to_kvm(mmu);
+
+ if (kvm_vm_is_protected(kvm))
+ return;
+
+ lockdep_assert_held_write(&kvm->mmu_lock);
+ WARN_ON(stage2_apply_range(mmu, 0, kvm_phys_size(mmu),
+ KVM_PGT_FN(kvm_pgtable_stage2_unmap_notlbi),
+ may_block));
+
+ kvm_call_hyp(__kvm_tlb_flush_vmid, mmu);
+}
+
void kvm_stage2_flush_range(struct kvm_s2_mmu *mmu, phys_addr_t addr, phys_addr_t end)
{
stage2_apply_range_resched(mmu, addr, end, KVM_PGT_FN(kvm_pgtable_stage2_flush));
--
2.47.3
^ permalink raw reply related [flat|nested] 12+ messages in thread
* [PATCH 3/4] KVM: arm64: nv: Move full s2_mmu unmap over to kvm_stage2_unmap_all()
2026-09-12 10:48 [PATCH 0/4] KVM: arm64: Reduce overhead of full S2 teardown Marc Zyngier
2026-09-12 10:48 ` [PATCH 1/4] KVM: arm64: pgtable: Add Stage-2 unmap without TLBI primitive Marc Zyngier
2026-09-12 10:48 ` [PATCH 2/4] KVM: arm64: MMU: Add kvm_stage2_unmap_all() helper Marc Zyngier
@ 2026-09-12 10:48 ` Marc Zyngier
2026-09-12 10:48 ` [PATCH 4/4] KVM: arm64: nv: Move TLBI VMALLS12E1* emulation " Marc Zyngier
` (2 subsequent siblings)
5 siblings, 0 replies; 12+ messages in thread
From: Marc Zyngier @ 2026-09-12 10:48 UTC (permalink / raw)
To: kvmarm, linux-arm-kernel
Cc: Wei-Lin Chang, Wang Han, Shuai Xue, Steffen Eiden, Joey Gouly,
Suzuki K Poulose, Oliver Upton, Zenghui Yu, Fuad Tabba
Now that we have kvm_stage2_unmap_all(), use it when getting rid of a
shadow S2 MMU.
Signed-off-by: Marc Zyngier <maz@kernel.org>
---
arch/arm64/kvm/nested.c | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
diff --git a/arch/arm64/kvm/nested.c b/arch/arm64/kvm/nested.c
index d60f6f69e293d..ebd8d23df0757 100644
--- a/arch/arm64/kvm/nested.c
+++ b/arch/arm64/kvm/nested.c
@@ -1295,7 +1295,7 @@ void kvm_nested_s2_unmap(struct kvm *kvm, bool may_block)
struct kvm_s2_mmu *mmu = kvm->arch.nested_mmus[i];
if (kvm_s2_mmu_valid(mmu))
- kvm_stage2_unmap_range(mmu, 0, kvm_phys_size(mmu), may_block);
+ kvm_stage2_unmap_all(mmu, may_block);
}
kvm_invalidate_vncr_ipa(kvm, 0, BIT(kvm->arch.mmu.pgt->ia_bits));
@@ -2014,7 +2014,7 @@ void check_nested_vcpu_requests(struct kvm_vcpu *vcpu)
write_lock(&vcpu->kvm->mmu_lock);
if (mmu->pending_unmap) {
- kvm_stage2_unmap_range(mmu, 0, kvm_phys_size(mmu), true);
+ kvm_stage2_unmap_all(mmu, true);
mmu->pending_unmap = false;
}
write_unlock(&vcpu->kvm->mmu_lock);
--
2.47.3
^ permalink raw reply related [flat|nested] 12+ messages in thread
* [PATCH 4/4] KVM: arm64: nv: Move TLBI VMALLS12E1* emulation over to kvm_stage2_unmap_all()
2026-09-12 10:48 [PATCH 0/4] KVM: arm64: Reduce overhead of full S2 teardown Marc Zyngier
` (2 preceding siblings ...)
2026-09-12 10:48 ` [PATCH 3/4] KVM: arm64: nv: Move full s2_mmu unmap over to kvm_stage2_unmap_all() Marc Zyngier
@ 2026-09-12 10:48 ` Marc Zyngier
2026-09-13 23:27 ` [PATCH 0/4] KVM: arm64: Reduce overhead of full S2 teardown Itaru Kitayama
2026-09-14 6:44 ` Shuai Xue
5 siblings, 0 replies; 12+ messages in thread
From: Marc Zyngier @ 2026-09-12 10:48 UTC (permalink / raw)
To: kvmarm, linux-arm-kernel
Cc: Wei-Lin Chang, Wang Han, Shuai Xue, Steffen Eiden, Joey Gouly,
Suzuki K Poulose, Oliver Upton, Zenghui Yu, Fuad Tabba
Instead of emulating TLBI VMALLS12E1* with a range invalidation, use
kvm_stage2_unmap_all() which is more efficient, specially in deeply
nested cases.
Signed-off-by: Marc Zyngier <maz@kernel.org>
---
arch/arm64/kvm/sys_regs.c | 18 ++++++++----------
1 file changed, 8 insertions(+), 10 deletions(-)
diff --git a/arch/arm64/kvm/sys_regs.c b/arch/arm64/kvm/sys_regs.c
index 44aae52c473d7..2c90e185c6e8e 100644
--- a/arch/arm64/kvm/sys_regs.c
+++ b/arch/arm64/kvm/sys_regs.c
@@ -4124,26 +4124,24 @@ static void s2_mmu_unmap_range(struct kvm_s2_mmu *mmu,
kvm_stage2_unmap_range(mmu, info->range.start, info->range.size, true);
}
+static void s2_mmu_unmap_all(struct kvm_s2_mmu *mmu,
+ const union tlbi_info *info)
+{
+ kvm_stage2_unmap_all(mmu, true);
+}
+
static bool handle_vmalls12e1is(struct kvm_vcpu *vcpu, struct sys_reg_params *p,
const struct sys_reg_desc *r)
{
u32 sys_encoding = sys_insn(p->Op0, p->Op1, p->CRn, p->CRm, p->Op2);
- u64 limit, vttbr;
+ u64 vttbr;
if (!kvm_supported_tlbi_s12_op(vcpu, sys_encoding))
return undef_access(vcpu, p, r);
vttbr = vcpu_read_sys_reg(vcpu, VTTBR_EL2);
- limit = BIT_ULL(kvm_get_pa_bits(vcpu->kvm));
- kvm_s2_mmu_iterate_by_vmid(vcpu->kvm, get_vmid(vttbr),
- &(union tlbi_info) {
- .range = {
- .start = 0,
- .size = limit,
- },
- },
- s2_mmu_unmap_range);
+ kvm_s2_mmu_iterate_by_vmid(vcpu->kvm, get_vmid(vttbr), NULL, s2_mmu_unmap_all);
return true;
}
--
2.47.3
^ permalink raw reply related [flat|nested] 12+ messages in thread
* Re: [PATCH 0/4] KVM: arm64: Reduce overhead of full S2 teardown
2026-09-12 10:48 [PATCH 0/4] KVM: arm64: Reduce overhead of full S2 teardown Marc Zyngier
` (3 preceding siblings ...)
2026-09-12 10:48 ` [PATCH 4/4] KVM: arm64: nv: Move TLBI VMALLS12E1* emulation " Marc Zyngier
@ 2026-09-13 23:27 ` Itaru Kitayama
2026-09-14 6:44 ` Shuai Xue
5 siblings, 0 replies; 12+ messages in thread
From: Itaru Kitayama @ 2026-09-13 23:27 UTC (permalink / raw)
To: Marc Zyngier
Cc: kvmarm, linux-arm-kernel, Wei-Lin Chang, Wang Han, Shuai Xue,
Steffen Eiden, Joey Gouly, Suzuki K Poulose, Oliver Upton,
Zenghui Yu, Fuad Tabba
Hi Marc,
On Sat, Sep 12, 2026 at 11:48:30AM +0100, Marc Zyngier wrote:
> Tearing down a full S2 is a pretty involved process, resulting in a
> lot of TLB invalidation. These TLBIs are either on a per leaf basis if
> the HW doesn't support range invalidation, or by top-level range if it
> does. Amusingly, the latter occurs even when nothing has been
> unmapped.
>
> Things are made worse with NV, as we have a bucket-load of shadow S2s,
> and the need to invalidate them all on the back of an MMU notifier.
> The latter will eventually be solved by the reverse-map tracking that
> Wei-Lin is working on, but we need to be better at full-S2 teardown.
>
> This small series adds a "no TLBI" unmapping primitive, which allows
> the caller to then whack the TLBs using a VMID-wide invalidation. This
> results in far fewer TLBIs, and a better recursive virtualisation as
> we get far fewer traps as a consequence.
>
> This applies on top of my shadow-s2 lifetime fixes, and is expected to
> be a prefix to Wei-Lin's series.
I've tested this series on QEMU with TCG mode, and it boots.
Tested-by: Itaru Kitayama <itaru.kitayama@fujitsu.com>
Thanks,
Itaru.
>
> Marc Zyngier (4):
> KVM: arm64: pgtable: Add Stage-2 unmap without TLBI primitive
> KVM: arm64: MMU: Add kvm_stage2_unmap_all() helper
> KVM: arm64: nv: Move full s2_mmu unmap over to kvm_stage2_unmap_all()
> KVM: arm64: nv: Move TLBI VMALLS12E1* emulation over to
> kvm_stage2_unmap_all()
>
> arch/arm64/include/asm/kvm_mmu.h | 1 +
> arch/arm64/include/asm/kvm_pgtable.h | 17 +++++++++++++
> arch/arm64/include/asm/kvm_pkvm.h | 1 +
> arch/arm64/kvm/hyp/pgtable.c | 38 +++++++++++++++++++++-------
> arch/arm64/kvm/mmu.c | 15 +++++++++++
> arch/arm64/kvm/nested.c | 4 +--
> arch/arm64/kvm/pkvm.c | 2 ++
> arch/arm64/kvm/sys_regs.c | 18 ++++++-------
> 8 files changed, 75 insertions(+), 21 deletions(-)
>
> --
> 2.47.3
>
>
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [PATCH 0/4] KVM: arm64: Reduce overhead of full S2 teardown
2026-09-12 10:48 [PATCH 0/4] KVM: arm64: Reduce overhead of full S2 teardown Marc Zyngier
` (4 preceding siblings ...)
2026-09-13 23:27 ` [PATCH 0/4] KVM: arm64: Reduce overhead of full S2 teardown Itaru Kitayama
@ 2026-09-14 6:44 ` Shuai Xue
2026-09-14 8:14 ` Marc Zyngier
5 siblings, 1 reply; 12+ messages in thread
From: Shuai Xue @ 2026-09-14 6:44 UTC (permalink / raw)
To: Marc Zyngier, kvmarm, linux-arm-kernel
Cc: Wei-Lin Chang, Wang Han, Steffen Eiden, Joey Gouly,
Suzuki K Poulose, Oliver Upton, Zenghui Yu, Fuad Tabba
On 9/12/26 6:48 PM, Marc Zyngier wrote:
> Tearing down a full S2 is a pretty involved process, resulting in a
> lot of TLB invalidation. These TLBIs are either on a per leaf basis if
> the HW doesn't support range invalidation, or by top-level range if it
> does. Amusingly, the latter occurs even when nothing has been
> unmapped.
>
> Things are made worse with NV, as we have a bucket-load of shadow S2s,
> and the need to invalidate them all on the back of an MMU notifier.
> The latter will eventually be solved by the reverse-map tracking that
> Wei-Lin is working on, but we need to be better at full-S2 teardown.
>
> This small series adds a "no TLBI" unmapping primitive, which allows
> the caller to then whack the TLBs using a VMID-wide invalidation. This
> results in far fewer TLBIs, and a better recursive virtualisation as
> we get far fewer traps as a consequence.
>
> This applies on top of my shadow-s2 lifetime fixes, and is expected to
> be a prefix to Wei-Lin's series.
>
> Marc Zyngier (4):
> KVM: arm64: pgtable: Add Stage-2 unmap without TLBI primitive
> KVM: arm64: MMU: Add kvm_stage2_unmap_all() helper
> KVM: arm64: nv: Move full s2_mmu unmap over to kvm_stage2_unmap_all()
> KVM: arm64: nv: Move TLBI VMALLS12E1* emulation over to
> kvm_stage2_unmap_all()
>
> arch/arm64/include/asm/kvm_mmu.h | 1 +
> arch/arm64/include/asm/kvm_pgtable.h | 17 +++++++++++++
> arch/arm64/include/asm/kvm_pkvm.h | 1 +
> arch/arm64/kvm/hyp/pgtable.c | 38 +++++++++++++++++++++-------
> arch/arm64/kvm/mmu.c | 15 +++++++++++
> arch/arm64/kvm/nested.c | 4 +--
> arch/arm64/kvm/pkvm.c | 2 ++
> arch/arm64/kvm/sys_regs.c | 18 ++++++-------
> 8 files changed, 75 insertions(+), 21 deletions(-)
>
Hi Marc,
Thanks for putting this series together. I reviewed the four patches
and revisited the traces from my earlier Marc-only tests.
I have a correctness concern about child page-table reclamation.
1. Child page-table reclamation before the final TLBI
In patch 1, SKIP_S2_TLBI suppresses invalidation for both leaf and table
descriptors, while stage2_unmap_walker() still immediately releases
empty child tables.
For a child table with page_count(childp) == 1, the sequence is:
stage2_unmap_walker()
stage2_unmap_put_pte()
clear the parent table descriptor
skip its TLBI
mm_ops->put_page(childp)
kvm_s2_put_page()
put_page() /* drop the child's last reference */
... process the remaining address ranges ...
__kvm_tlb_flush_vmid() /* final invalidation in patch 2 */
This path does not use the free_unlinked_table()/call_rcu() deferred
reclamation mechanism. The existing deferred-range-TLBI path still
invalidates table descriptors immediately; the new flag skips that
invalidation too.
Patch 4 provides a caller operating on active shadow MMUs:
kvm_s2_mmu_iterate_by_vmid() holds mmu_lock for write and visits valid
matching shadow MMUs, but does not require refcnt == 0 or wait for
other vCPUs using the MMU to exit.
Another vCPU can therefore still use that shadow S2. The write lock
excludes software page-table updates, not hardware table walks.
The following interleaving is allowed:
vCPU B / hardware walker vCPU A
------------------------ ----------------------------
Holds an old reference to T
Clears parent, skips TLBI
Drops T's last reference
T is reused by the allocator
Accesses T via the old reference
Performs final VMID-wide TLBI
The final flush barriers cannot retroactively protect a table that
has already been freed and reused.
Is there a lifetime guarantee for these active-MMU callers that I
have missed? Otherwise, the corresponding invalidation needs to
complete before the child table's memory can be reused.
Retaining table-descriptor TLBI looks like a smaller correction
that would still remove empty-range TLBI amplification. Batching
those invalidations as well would require deferred child-table
reclamation, accounting for the lock drop/reacquisition boundaries
in the may_block path.
Thanks.
Shuai
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [PATCH 0/4] KVM: arm64: Reduce overhead of full S2 teardown
2026-09-14 6:44 ` Shuai Xue
@ 2026-09-14 8:14 ` Marc Zyngier
2026-09-14 9:06 ` Shuai Xue
2026-09-15 23:19 ` Oliver Upton
0 siblings, 2 replies; 12+ messages in thread
From: Marc Zyngier @ 2026-09-14 8:14 UTC (permalink / raw)
To: Shuai Xue
Cc: kvmarm, linux-arm-kernel, Wei-Lin Chang, Wang Han, Steffen Eiden,
Joey Gouly, Suzuki K Poulose, Oliver Upton, Zenghui Yu,
Fuad Tabba
On Mon, 14 Sep 2026 07:44:15 +0100,
Shuai Xue <xueshuai@linux.alibaba.com> wrote:
>
>
>
> On 9/12/26 6:48 PM, Marc Zyngier wrote:
> > Tearing down a full S2 is a pretty involved process, resulting in a
> > lot of TLB invalidation. These TLBIs are either on a per leaf basis if
> > the HW doesn't support range invalidation, or by top-level range if it
> > does. Amusingly, the latter occurs even when nothing has been
> > unmapped.
> >
> > Things are made worse with NV, as we have a bucket-load of shadow S2s,
> > and the need to invalidate them all on the back of an MMU notifier.
> > The latter will eventually be solved by the reverse-map tracking that
> > Wei-Lin is working on, but we need to be better at full-S2 teardown.
> >
> > This small series adds a "no TLBI" unmapping primitive, which allows
> > the caller to then whack the TLBs using a VMID-wide invalidation. This
> > results in far fewer TLBIs, and a better recursive virtualisation as
> > we get far fewer traps as a consequence.
> >
> > This applies on top of my shadow-s2 lifetime fixes, and is expected to
> > be a prefix to Wei-Lin's series.
> >
> > Marc Zyngier (4):
> > KVM: arm64: pgtable: Add Stage-2 unmap without TLBI primitive
> > KVM: arm64: MMU: Add kvm_stage2_unmap_all() helper
> > KVM: arm64: nv: Move full s2_mmu unmap over to kvm_stage2_unmap_all()
> > KVM: arm64: nv: Move TLBI VMALLS12E1* emulation over to
> > kvm_stage2_unmap_all()
> >
> > arch/arm64/include/asm/kvm_mmu.h | 1 +
> > arch/arm64/include/asm/kvm_pgtable.h | 17 +++++++++++++
> > arch/arm64/include/asm/kvm_pkvm.h | 1 +
> > arch/arm64/kvm/hyp/pgtable.c | 38 +++++++++++++++++++++-------
> > arch/arm64/kvm/mmu.c | 15 +++++++++++
> > arch/arm64/kvm/nested.c | 4 +--
> > arch/arm64/kvm/pkvm.c | 2 ++
> > arch/arm64/kvm/sys_regs.c | 18 ++++++-------
> > 8 files changed, 75 insertions(+), 21 deletions(-)
> >
>
>
> Hi Marc,
>
> Thanks for putting this series together. I reviewed the four patches
> and revisited the traces from my earlier Marc-only tests.
>
> I have a correctness concern about child page-table reclamation.
>
You? Or your AI model? I'd really expect you to explain *your*
perception of the problem rather than dumping the result of your AI in
an email.
> 1. Child page-table reclamation before the final TLBI
>
> In patch 1, SKIP_S2_TLBI suppresses invalidation for both leaf and table
> descriptors, while stage2_unmap_walker() still immediately releases
> empty child tables.
What is a "child" table?
>
> For a child table with page_count(childp) == 1, the sequence is:
>
> stage2_unmap_walker()
> stage2_unmap_put_pte()
> clear the parent table descriptor
> skip its TLBI
> mm_ops->put_page(childp)
> kvm_s2_put_page()
> put_page() /* drop the child's last reference */
>
> ... process the remaining address ranges ...
>
> __kvm_tlb_flush_vmid() /* final invalidation in patch 2 */
>
> This path does not use the free_unlinked_table()/call_rcu() deferred
> reclamation mechanism. The existing deferred-range-TLBI path still
> invalidates table descriptors immediately; the new flag skips that
> invalidation too.
And? What is the actual problem here?
>
> Patch 4 provides a caller operating on active shadow MMUs:
> kvm_s2_mmu_iterate_by_vmid() holds mmu_lock for write and visits valid
> matching shadow MMUs, but does not require refcnt == 0 or wait for
> other vCPUs using the MMU to exit.
Of course it doesn't, since this is simply emulating an instruction
local to that vcpu. How would the actual HW "wait" for another CPU to
stop using a set of translation?
>
> Another vCPU can therefore still use that shadow S2. The write lock
> excludes software page-table updates, not hardware table walks.
> The following interleaving is allowed:
>
> vCPU B / hardware walker vCPU A
> ------------------------ ----------------------------
> Holds an old reference to T
> Clears parent, skips TLBI
> Drops T's last reference
> T is reused by the allocator
> Accesses T via the old reference
> Performs final VMID-wide TLBI
>
> The final flush barriers cannot retroactively protect a table that
> has already been freed and reused.
T is a shadow page-table page on the host. How can vcpu B hold a
reference on that page? It isn't even in the same address space.
Now, I can see that B could have a VA that *translate through* T, and
that's rather annoying.
I'll have a think.
M.
--
Without deviation from the norm, progress is not possible.
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [PATCH 1/4] KVM: arm64: pgtable: Add Stage-2 unmap without TLBI primitive
2026-09-12 10:48 ` [PATCH 1/4] KVM: arm64: pgtable: Add Stage-2 unmap without TLBI primitive Marc Zyngier
@ 2026-09-14 8:48 ` Mark Rutland
2026-09-14 9:23 ` Mark Rutland
0 siblings, 1 reply; 12+ messages in thread
From: Mark Rutland @ 2026-09-14 8:48 UTC (permalink / raw)
To: Marc Zyngier
Cc: kvmarm, linux-arm-kernel, Wei-Lin Chang, Wang Han, Shuai Xue,
Steffen Eiden, Joey Gouly, Suzuki K Poulose, Oliver Upton,
Zenghui Yu, Fuad Tabba
On Sat, Sep 12, 2026 at 11:48:31AM +0100, Marc Zyngier wrote:
> kvm_pgtable_stage2_unmap() iterates over a range, unmapping whatever is
> within the range, and always guarantees that that the corresponding TLBs
> are invalidated when the function returns.
>
> While this is safe, it means that iterating over empty range on a system
> that supports range invalidation results in a TLBI per largest block
> mapping size (1GB, 32MB or 512MB, depending on the base granule size).
>
> This can be pretty expensive in situation where the whole address space
> is being torn down, as it happens with NV (where S2 MMUs are recycled
> regularly), and it would be more efficient to elide the per-subrange
> TLBIs to solely rely on a VMID-wide TLBI.
>
> For this, provide a kvm_pgtable_stage2_unmap_notlbi() helper that elides
> all TLBIs, and relies on the caller to do the work.
>
> Note that for pKVM case, no additional helper is provided, and we
> fallback on the TLBI-aware version.
Just to check: I assume that before this is called, we have somehow
ensured that the S2 being torn down isn't live on any PE, and cannot
become live on any PE? I asssume that's a natural part of S2 lifetime
management, but I couldn't figure that out from a quick skim of the hyp
pgtable code.
Assuming so, it might be worth mentioning that in the commit message,
since it explains why it's safe to invalidate *after* intermediate
tables are freed by stage2_unmap_walker() calling mm_ops->put_page(). We
might also be able to add some test/assertion in
kvm_pgtable_stage2_unmap_notlbi() to ensure it is not called where the
tables could be live on a PE.
Otherwise, this all looks sensible to me!
Mark.
>
> Signed-off-by: Marc Zyngier <maz@kernel.org>
> ---
> arch/arm64/include/asm/kvm_pgtable.h | 17 +++++++++++++
> arch/arm64/include/asm/kvm_pkvm.h | 1 +
> arch/arm64/kvm/hyp/pgtable.c | 38 +++++++++++++++++++++-------
> arch/arm64/kvm/pkvm.c | 2 ++
> 4 files changed, 49 insertions(+), 9 deletions(-)
>
> diff --git a/arch/arm64/include/asm/kvm_pgtable.h b/arch/arm64/include/asm/kvm_pgtable.h
> index 41a8687938eb6..c370196888d1d 100644
> --- a/arch/arm64/include/asm/kvm_pgtable.h
> +++ b/arch/arm64/include/asm/kvm_pgtable.h
> @@ -318,6 +318,8 @@ typedef bool (*kvm_pgtable_force_pte_cb_t)(u64 addr, u64 end,
> * @KVM_PGTABLE_WALK_SKIP_CMO: Visit and update table entries
> * without Cache maintenance
> * operations required.
> + * @KVM_PGTABLE_WALK_SKIP_S2_TLBI: Visit and update table entries
> + * without Stage-2 TLB invalidation.
> */
> enum kvm_pgtable_walk_flags {
> KVM_PGTABLE_WALK_LEAF = BIT(0),
> @@ -327,6 +329,7 @@ enum kvm_pgtable_walk_flags {
> KVM_PGTABLE_WALK_IGNORE_EAGAIN = BIT(4),
> KVM_PGTABLE_WALK_SKIP_BBM_TLBI = BIT(5),
> KVM_PGTABLE_WALK_SKIP_CMO = BIT(6),
> + KVM_PGTABLE_WALK_SKIP_S2_TLBI = BIT(7),
> };
>
> struct kvm_pgtable_visit_ctx {
> @@ -717,6 +720,20 @@ int kvm_pgtable_stage2_annotate(struct kvm_pgtable *pgt, u64 addr, u64 size,
> */
> int kvm_pgtable_stage2_unmap(struct kvm_pgtable *pgt, u64 addr, u64 size);
>
> +/**
> + * kvm_pgtable_stage2_unmap_notlbi() - Remove a mapping from a guest stage-2 page-table
> + * without TLB invalidation.
> + * @pgt: Page-table structure initialised by kvm_pgtable_stage2_init*().
> + * @addr: Intermediate physical address from which to remove the mapping.
> + * @size: Size of the mapping.
> + *
> + * Same as kvm_pgtable_stage2_unmap(), but does not invalidate the
> + * TLBs, which is the responsibility of the caller. Use with caution!
> + *
> + * Return: 0 on success, negative error code on failure.
> + */
> +int kvm_pgtable_stage2_unmap_notlbi(struct kvm_pgtable *pgt, u64 addr, u64 size);
> +
> /**
> * kvm_pgtable_stage2_wrprotect() - Write-protect guest stage-2 address range
> * without TLB invalidation.
> diff --git a/arch/arm64/include/asm/kvm_pkvm.h b/arch/arm64/include/asm/kvm_pkvm.h
> index beea00e693a0a..273013c98ff17 100644
> --- a/arch/arm64/include/asm/kvm_pkvm.h
> +++ b/arch/arm64/include/asm/kvm_pkvm.h
> @@ -214,6 +214,7 @@ int pkvm_pgtable_stage2_map(struct kvm_pgtable *pgt, u64 addr, u64 size, u64 phy
> enum kvm_pgtable_prot prot, void *mc,
> enum kvm_pgtable_walk_flags flags);
> int pkvm_pgtable_stage2_unmap(struct kvm_pgtable *pgt, u64 addr, u64 size);
> +int pkvm_pgtable_stage2_unmap_notlbi(struct kvm_pgtable *pgt, u64 addr, u64 size);
> int pkvm_pgtable_stage2_wrprotect(struct kvm_pgtable *pgt, u64 addr, u64 size);
> int pkvm_pgtable_stage2_flush(struct kvm_pgtable *pgt, u64 addr, u64 size);
> bool pkvm_pgtable_stage2_test_clear_young(struct kvm_pgtable *pgt, u64 addr, u64 size, bool mkold);
> diff --git a/arch/arm64/kvm/hyp/pgtable.c b/arch/arm64/kvm/hyp/pgtable.c
> index b74dd5ce1efd3..6603fc236daa2 100644
> --- a/arch/arm64/kvm/hyp/pgtable.c
> +++ b/arch/arm64/kvm/hyp/pgtable.c
> @@ -29,6 +29,11 @@ static bool kvm_pgtable_walk_skip_cmo(const struct kvm_pgtable_visit_ctx *ctx)
> return unlikely(ctx->flags & KVM_PGTABLE_WALK_SKIP_CMO);
> }
>
> +static bool kvm_pgtable_walk_skip_s2_tlbi(const struct kvm_pgtable_visit_ctx *ctx)
> +{
> + return unlikely(ctx->flags & KVM_PGTABLE_WALK_SKIP_S2_TLBI);
> +}
> +
> static bool kvm_block_mapping_supported(const struct kvm_pgtable_visit_ctx *ctx, u64 phys)
> {
> u64 granule = kvm_granule_size(ctx->level);
> @@ -905,12 +910,14 @@ static void stage2_unmap_put_pte(const struct kvm_pgtable_visit_ctx *ctx,
> if (kvm_pte_valid(ctx->old)) {
> kvm_clear_pte(ctx->ptep);
>
> - if (kvm_pte_table(ctx->old, ctx->level)) {
> - kvm_call_hyp(__kvm_tlb_flush_vmid_ipa, mmu, ctx->addr,
> - TLBI_TTL_UNKNOWN);
> - } else if (!stage2_unmap_defer_tlb_flush(pgt)) {
> - kvm_call_hyp(__kvm_tlb_flush_vmid_ipa, mmu, ctx->addr,
> - ctx->level);
> + if (!kvm_pgtable_walk_skip_s2_tlbi(ctx)) {
> + if (kvm_pte_table(ctx->old, ctx->level)) {
> + kvm_call_hyp(__kvm_tlb_flush_vmid_ipa, mmu, ctx->addr,
> + TLBI_TTL_UNKNOWN);
> + } else if (!stage2_unmap_defer_tlb_flush(pgt)) {
> + kvm_call_hyp(__kvm_tlb_flush_vmid_ipa, mmu, ctx->addr,
> + ctx->level);
> + }
> }
> }
>
> @@ -1195,23 +1202,36 @@ static int stage2_unmap_walker(const struct kvm_pgtable_visit_ctx *ctx,
> return 0;
> }
>
> -int kvm_pgtable_stage2_unmap(struct kvm_pgtable *pgt, u64 addr, u64 size)
> +static int __kvm_pgtable_stage2_unmap(struct kvm_pgtable *pgt,
> + enum kvm_pgtable_walk_flags flags,
> + u64 addr, u64 size)
> {
> int ret;
> struct kvm_pgtable_walker walker = {
> .cb = stage2_unmap_walker,
> .arg = pgt,
> - .flags = KVM_PGTABLE_WALK_LEAF | KVM_PGTABLE_WALK_TABLE_POST,
> + .flags = KVM_PGTABLE_WALK_LEAF | KVM_PGTABLE_WALK_TABLE_POST | flags,
> };
>
> ret = kvm_pgtable_walk(pgt, addr, size, &walker);
> - if (stage2_unmap_defer_tlb_flush(pgt))
> + if (stage2_unmap_defer_tlb_flush(pgt) &&
> + !(flags & KVM_PGTABLE_WALK_SKIP_S2_TLBI))
> /* Perform the deferred TLB invalidations */
> kvm_tlb_flush_vmid_range(pgt->mmu, addr, size);
>
> return ret;
> }
>
> +int kvm_pgtable_stage2_unmap(struct kvm_pgtable *pgt, u64 addr, u64 size)
> +{
> + return __kvm_pgtable_stage2_unmap(pgt, 0, addr, size);
> +}
> +
> +int kvm_pgtable_stage2_unmap_notlbi(struct kvm_pgtable *pgt, u64 addr, u64 size)
> +{
> + return __kvm_pgtable_stage2_unmap(pgt, KVM_PGTABLE_WALK_SKIP_S2_TLBI, addr, size);
> +}
> +
> struct stage2_attr_data {
> kvm_pte_t attr_set;
> kvm_pte_t attr_clr;
> diff --git a/arch/arm64/kvm/pkvm.c b/arch/arm64/kvm/pkvm.c
> index 8e4c6e4bec123..ec151005fbe4d 100644
> --- a/arch/arm64/kvm/pkvm.c
> +++ b/arch/arm64/kvm/pkvm.c
> @@ -488,6 +488,8 @@ int pkvm_pgtable_stage2_unmap(struct kvm_pgtable *pgt, u64 addr, u64 size)
> return __pkvm_pgtable_stage2_unshare(pgt, addr, addr + size);
> }
>
> +int pkvm_pgtable_stage2_unmap_notlbi(struct kvm_pgtable *pgt, u64 addr, u64 size) __alias(pkvm_pgtable_stage2_unmap);
> +
> int pkvm_pgtable_stage2_wrprotect(struct kvm_pgtable *pgt, u64 addr, u64 size)
> {
> struct kvm *kvm = kvm_s2_mmu_to_kvm(pgt->mmu);
> --
> 2.47.3
>
>
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [PATCH 0/4] KVM: arm64: Reduce overhead of full S2 teardown
2026-09-14 8:14 ` Marc Zyngier
@ 2026-09-14 9:06 ` Shuai Xue
2026-09-15 23:19 ` Oliver Upton
1 sibling, 0 replies; 12+ messages in thread
From: Shuai Xue @ 2026-09-14 9:06 UTC (permalink / raw)
To: Marc Zyngier
Cc: kvmarm, linux-arm-kernel, Wei-Lin Chang, Wang Han, Steffen Eiden,
Joey Gouly, Suzuki K Poulose, Oliver Upton, Zenghui Yu,
Fuad Tabba
On 9/14/26 4:14 PM, Marc Zyngier wrote:
> On Mon, 14 Sep 2026 07:44:15 +0100,
> Shuai Xue <xueshuai@linux.alibaba.com> wrote:
>>
>>
>>
>> On 9/12/26 6:48 PM, Marc Zyngier wrote:
>>> Tearing down a full S2 is a pretty involved process, resulting in a
>>> lot of TLB invalidation. These TLBIs are either on a per leaf basis if
>>> the HW doesn't support range invalidation, or by top-level range if it
>>> does. Amusingly, the latter occurs even when nothing has been
>>> unmapped.
>>>
>>> Things are made worse with NV, as we have a bucket-load of shadow S2s,
>>> and the need to invalidate them all on the back of an MMU notifier.
>>> The latter will eventually be solved by the reverse-map tracking that
>>> Wei-Lin is working on, but we need to be better at full-S2 teardown.
>>>
>>> This small series adds a "no TLBI" unmapping primitive, which allows
>>> the caller to then whack the TLBs using a VMID-wide invalidation. This
>>> results in far fewer TLBIs, and a better recursive virtualisation as
>>> we get far fewer traps as a consequence.
>>>
>>> This applies on top of my shadow-s2 lifetime fixes, and is expected to
>>> be a prefix to Wei-Lin's series.
>>>
>>> Marc Zyngier (4):
>>> KVM: arm64: pgtable: Add Stage-2 unmap without TLBI primitive
>>> KVM: arm64: MMU: Add kvm_stage2_unmap_all() helper
>>> KVM: arm64: nv: Move full s2_mmu unmap over to kvm_stage2_unmap_all()
>>> KVM: arm64: nv: Move TLBI VMALLS12E1* emulation over to
>>> kvm_stage2_unmap_all()
>>>
>>> arch/arm64/include/asm/kvm_mmu.h | 1 +
>>> arch/arm64/include/asm/kvm_pgtable.h | 17 +++++++++++++
>>> arch/arm64/include/asm/kvm_pkvm.h | 1 +
>>> arch/arm64/kvm/hyp/pgtable.c | 38 +++++++++++++++++++++-------
>>> arch/arm64/kvm/mmu.c | 15 +++++++++++
>>> arch/arm64/kvm/nested.c | 4 +--
>>> arch/arm64/kvm/pkvm.c | 2 ++
>>> arch/arm64/kvm/sys_regs.c | 18 ++++++-------
>>> 8 files changed, 75 insertions(+), 21 deletions(-)
>>>
>>
>>
>> Hi Marc,
>>
>> Thanks for putting this series together. I reviewed the four patches
>> and revisited the traces from my earlier Marc-only tests.
>>
>> I have a correctness concern about child page-table reclamation.
>>
>
> You? Or your AI model? I'd really expect you to explain *your*
> perception of the problem rather than dumping the result of your AI in
> an email.
Hi Marc,
You're right. I used AI assistance (GPT-6 Astra) for the analysis.
Let me clarify the technical points.
>
>> 1. Child page-table reclamation before the final TLBI
>>
>> In patch 1, SKIP_S2_TLBI suppresses invalidation for both leaf and table
>> descriptors, while stage2_unmap_walker() still immediately releases
>> empty child tables.
>
> What is a "child" table?
I meant the next-level shadow S2 table obtained here in
stage2_unmap_walker():
if (kvm_pte_table(ctx->old, ctx->level)) {
childp = kvm_pte_follow(ctx->old, mm_ops);
For example, if ctx->old is an L2 table descriptor, childp points
to the L3 table that it describes. This is a host-allocated shadow
page-table page, not a guest data page.
>
>>
>> For a child table with page_count(childp) == 1, the sequence is:
>>
>> stage2_unmap_walker()
>> stage2_unmap_put_pte()
>> clear the parent table descriptor
>> skip its TLBI
>> mm_ops->put_page(childp)
>> kvm_s2_put_page()
>> put_page() /* drop the child's last reference */
>>
>> ... process the remaining address ranges ...
>>
>> __kvm_tlb_flush_vmid() /* final invalidation in patch 2 */
>>
>> This path does not use the free_unlinked_table()/call_rcu() deferred
>> reclamation mechanism. The existing deferred-range-TLBI path still
>> invalidates table descriptors immediately; the new flag skips that
>> invalidation too.
>
> And? What is the actual problem here?
The concern is that this page becomes available for reuse before
the corresponding invalidation completes.
With page_count(childp) == 1, the walker clears the parent descriptor
through stage2_unmap_put_pte(), then releases the next-level table
with mm_ops->put_page(childp). SKIP_S2_TLBI suppresses the table-
descriptor TLBI that previously happened before this release.
The final VMID-wide TLBI happens after the full walk.
If hardware can still walk through that page using old intermediate
translation state, it could interpret contents written by a new
owner as S2 descriptors. That is the failure scenario I intended
to describe.
>
>>
>> Patch 4 provides a caller operating on active shadow MMUs:
>> kvm_s2_mmu_iterate_by_vmid() holds mmu_lock for write and visits valid
>> matching shadow MMUs, but does not require refcnt == 0 or wait for
>> other vCPUs using the MMU to exit.
>
> Of course it doesn't, since this is simply emulating an instruction
> local to that vcpu. How would the actual HW "wait" for another CPU to
> stop using a set of translation?
Agreed. I should not have presented the absence of a vCPU-stop
mechanism as a problem. Stopping another vCPU and waiting for
invalidation to complete are different things.
My concern is the lifetime of the host shadow page-table memory:
whether it can be reclaimed before the final TLBI completes.
>
>>
>> Another vCPU can therefore still use that shadow S2. The write lock
>> excludes software page-table updates, not hardware table walks.
>> The following interleaving is allowed:
>>
>> vCPU B / hardware walker vCPU A
>> ------------------------ ----------------------------
>> Holds an old reference to T
>> Clears parent, skips TLBI
>> Drops T's last reference
>> T is reused by the allocator
>> Accesses T via the old reference
>> Performs final VMID-wide TLBI
>>
>> The final flush barriers cannot retroactively protect a table that
>> has already been freed and reused.
>
> T is a shadow page-table page on the host. How can vcpu B hold a
> reference on that page? It isn't even in the same address space.
Sorry for the misleading. I meant a hardware table walk
on the PE running B going through T, not a software reference held
by B or a guest mapping of T.
>
> Now, I can see that B could have a VA that *translate through* T, and
> that's rather annoying.
>
> I'll have a think.
>
> M.
Yes, translating through T is what I meant.
Would retaining the table-descriptor TLBI be a reasonable minimal
fix? That would preserve invalidation before releasing the next-level
table, while still avoiding the per-range flushes for empty chunks.
Thanks for looking into it.
Shuai
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [PATCH 1/4] KVM: arm64: pgtable: Add Stage-2 unmap without TLBI primitive
2026-09-14 8:48 ` Mark Rutland
@ 2026-09-14 9:23 ` Mark Rutland
0 siblings, 0 replies; 12+ messages in thread
From: Mark Rutland @ 2026-09-14 9:23 UTC (permalink / raw)
To: Marc Zyngier
Cc: kvmarm, linux-arm-kernel, Wei-Lin Chang, Wang Han, Shuai Xue,
Steffen Eiden, Joey Gouly, Suzuki K Poulose, Oliver Upton,
Zenghui Yu, Fuad Tabba
On Mon, Sep 14, 2026 at 09:48:03AM +0100, Mark Rutland wrote:
> On Sat, Sep 12, 2026 at 11:48:31AM +0100, Marc Zyngier wrote:
> > kvm_pgtable_stage2_unmap() iterates over a range, unmapping whatever is
> > within the range, and always guarantees that that the corresponding TLBs
> > are invalidated when the function returns.
> >
> > While this is safe, it means that iterating over empty range on a system
> > that supports range invalidation results in a TLBI per largest block
> > mapping size (1GB, 32MB or 512MB, depending on the base granule size).
> >
> > This can be pretty expensive in situation where the whole address space
> > is being torn down, as it happens with NV (where S2 MMUs are recycled
> > regularly), and it would be more efficient to elide the per-subrange
> > TLBIs to solely rely on a VMID-wide TLBI.
> >
> > For this, provide a kvm_pgtable_stage2_unmap_notlbi() helper that elides
> > all TLBIs, and relies on the caller to do the work.
> >
> > Note that for pKVM case, no additional helper is provided, and we
> > fallback on the TLBI-aware version.
>
> Just to check: I assume that before this is called, we have somehow
> ensured that the S2 being torn down isn't live on any PE, and cannot
> become live on any PE? I asssume that's a natural part of S2 lifetime
> management, but I couldn't figure that out from a quick skim of the hyp
> pgtable code.
>
> Assuming so, it might be worth mentioning that in the commit message,
> since it explains why it's safe to invalidate *after* intermediate
> tables are freed by stage2_unmap_walker() calling mm_ops->put_page(). We
> might also be able to add some test/assertion in
> kvm_pgtable_stage2_unmap_notlbi() to ensure it is not called where the
> tables could be live on a PE.
I see Shuai Xue said something in this area, but just to elaborate:
If it's possible that some PE is performing a translation table walk of
the tables being freed, and if a page for an intermediate table gets
freed and reallocated/reused before the TLBI is executed, then HW might
read a garbage value when trying to read a decriptor from that page.
That could lead to a variety of problems (e.g. permit a guest to access
an arbitrary PA, or cause a HW walk to access an arbitrary PA).
Above I had assumed that the S2 table being unmapped+freed weren't live
(e.g. not programmed into any PE's VTTBR) at the time we performed this
invalidation, and hence kvm_pgtable_stage2_unmap_notlbi() could check
some refcount or something to verify that.
If that is the case, all's good, and you can ignore the rest of this
mail. If that is not the case, read on.
If the tables are potentially live on some PE, then it is not safe to
defer the TLB invalidation. However, we can still reduce the TLBIs with
a slightly more elaborate sequence:
(1) Make a temporary copy of the root table (whatever VTTBR points to)
for the S2 being invalidated.
(2) Clear the entries from the root table (leaving any child
intermediate/leaf tables as-is).
(3) Perform TLB invalidation for the entire VMID. This will ensure HW
can't walk to any of the child tables.
(4) Walk the temporary copy of the root table to free the child tables,
without any TLB invalidation.
(5) Free the temporary copy of the root table.
Mark.
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [PATCH 0/4] KVM: arm64: Reduce overhead of full S2 teardown
2026-09-14 8:14 ` Marc Zyngier
2026-09-14 9:06 ` Shuai Xue
@ 2026-09-15 23:19 ` Oliver Upton
1 sibling, 0 replies; 12+ messages in thread
From: Oliver Upton @ 2026-09-15 23:19 UTC (permalink / raw)
To: Marc Zyngier
Cc: Shuai Xue, kvmarm, linux-arm-kernel, Wei-Lin Chang, Wang Han,
Steffen Eiden, Joey Gouly, Suzuki K Poulose, Zenghui Yu,
Fuad Tabba
Hey,
On Mon, Sep 14, 2026 at 09:14:24AM +0100, Marc Zyngier wrote:
> > Another vCPU can therefore still use that shadow S2. The write lock
> > excludes software page-table updates, not hardware table walks.
> > The following interleaving is allowed:
> >
> > vCPU B / hardware walker vCPU A
> > ------------------------ ----------------------------
> > Holds an old reference to T
> > Clears parent, skips TLBI
> > Drops T's last reference
> > T is reused by the allocator
> > Accesses T via the old reference
> > Performs final VMID-wide TLBI
> >
> > The final flush barriers cannot retroactively protect a table that
> > has already been freed and reused.
>
> T is a shadow page-table page on the host. How can vcpu B hold a
> reference on that page? It isn't even in the same address space.
>
> Now, I can see that B could have a VA that *translate through* T, and
> that's rather annoying.
>
> I'll have a think.
One option that may be worth considering is reusing the KVM_REQ_NESTED_S2_UNMAP
request for emulating VMALLS12E1*. All vCPUs that share the S2 MMU would
be forced to take EL1&0 out of context until the unmap operation
completes, at the expense of an unnecessary kick to vCPUs in a different
MMU context.
A bit hacky but seems more straightforward than wiring up some batching
mechanism for freeing table pages.
Thanks,
Oliver
^ permalink raw reply [flat|nested] 12+ messages in thread
end of thread, other threads:[~2026-09-15 23:19 UTC | newest]
Thread overview: 12+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-12 10:48 [PATCH 0/4] KVM: arm64: Reduce overhead of full S2 teardown Marc Zyngier
2026-09-12 10:48 ` [PATCH 1/4] KVM: arm64: pgtable: Add Stage-2 unmap without TLBI primitive Marc Zyngier
2026-09-14 8:48 ` Mark Rutland
2026-09-14 9:23 ` Mark Rutland
2026-09-12 10:48 ` [PATCH 2/4] KVM: arm64: MMU: Add kvm_stage2_unmap_all() helper Marc Zyngier
2026-09-12 10:48 ` [PATCH 3/4] KVM: arm64: nv: Move full s2_mmu unmap over to kvm_stage2_unmap_all() Marc Zyngier
2026-09-12 10:48 ` [PATCH 4/4] KVM: arm64: nv: Move TLBI VMALLS12E1* emulation " Marc Zyngier
2026-09-13 23:27 ` [PATCH 0/4] KVM: arm64: Reduce overhead of full S2 teardown Itaru Kitayama
2026-09-14 6:44 ` Shuai Xue
2026-09-14 8:14 ` Marc Zyngier
2026-09-14 9:06 ` Shuai Xue
2026-09-15 23:19 ` Oliver Upton
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.