From: Mark Rutland <mark.rutland@arm.com>
To: Marc Zyngier <maz@kernel.org>
Cc: kvmarm@lists.linux.dev, linux-arm-kernel@lists.infradead.org,
Wei-Lin Chang <weilin.chang@arm.com>,
Wang Han <wanghan@linux.alibaba.com>,
Shuai Xue <xueshuai@linux.alibaba.com>,
Steffen Eiden <seiden@linux.ibm.com>,
Joey Gouly <joey.gouly@arm.com>,
Suzuki K Poulose <suzuki.poulose@arm.com>,
Oliver Upton <oupton@kernel.org>,
Zenghui Yu <yuzenghui@huawei.com>,
Fuad Tabba <fuad.tabba@linux.dev>
Subject: Re: [PATCH 1/4] KVM: arm64: pgtable: Add Stage-2 unmap without TLBI primitive
Date: Mon, 14 Sep 2026 09:48:03 +0100 [thread overview]
Message-ID: <aqe0w0myu7iIT_tR@J2N7QTR9R3> (raw)
In-Reply-To: <20260912104834.3093878-2-maz@kernel.org>
On Sat, Sep 12, 2026 at 11:48:31AM +0100, Marc Zyngier wrote:
> kvm_pgtable_stage2_unmap() iterates over a range, unmapping whatever is
> within the range, and always guarantees that that the corresponding TLBs
> are invalidated when the function returns.
>
> While this is safe, it means that iterating over empty range on a system
> that supports range invalidation results in a TLBI per largest block
> mapping size (1GB, 32MB or 512MB, depending on the base granule size).
>
> This can be pretty expensive in situation where the whole address space
> is being torn down, as it happens with NV (where S2 MMUs are recycled
> regularly), and it would be more efficient to elide the per-subrange
> TLBIs to solely rely on a VMID-wide TLBI.
>
> For this, provide a kvm_pgtable_stage2_unmap_notlbi() helper that elides
> all TLBIs, and relies on the caller to do the work.
>
> Note that for pKVM case, no additional helper is provided, and we
> fallback on the TLBI-aware version.
Just to check: I assume that before this is called, we have somehow
ensured that the S2 being torn down isn't live on any PE, and cannot
become live on any PE? I asssume that's a natural part of S2 lifetime
management, but I couldn't figure that out from a quick skim of the hyp
pgtable code.
Assuming so, it might be worth mentioning that in the commit message,
since it explains why it's safe to invalidate *after* intermediate
tables are freed by stage2_unmap_walker() calling mm_ops->put_page(). We
might also be able to add some test/assertion in
kvm_pgtable_stage2_unmap_notlbi() to ensure it is not called where the
tables could be live on a PE.
Otherwise, this all looks sensible to me!
Mark.
>
> Signed-off-by: Marc Zyngier <maz@kernel.org>
> ---
> arch/arm64/include/asm/kvm_pgtable.h | 17 +++++++++++++
> arch/arm64/include/asm/kvm_pkvm.h | 1 +
> arch/arm64/kvm/hyp/pgtable.c | 38 +++++++++++++++++++++-------
> arch/arm64/kvm/pkvm.c | 2 ++
> 4 files changed, 49 insertions(+), 9 deletions(-)
>
> diff --git a/arch/arm64/include/asm/kvm_pgtable.h b/arch/arm64/include/asm/kvm_pgtable.h
> index 41a8687938eb6..c370196888d1d 100644
> --- a/arch/arm64/include/asm/kvm_pgtable.h
> +++ b/arch/arm64/include/asm/kvm_pgtable.h
> @@ -318,6 +318,8 @@ typedef bool (*kvm_pgtable_force_pte_cb_t)(u64 addr, u64 end,
> * @KVM_PGTABLE_WALK_SKIP_CMO: Visit and update table entries
> * without Cache maintenance
> * operations required.
> + * @KVM_PGTABLE_WALK_SKIP_S2_TLBI: Visit and update table entries
> + * without Stage-2 TLB invalidation.
> */
> enum kvm_pgtable_walk_flags {
> KVM_PGTABLE_WALK_LEAF = BIT(0),
> @@ -327,6 +329,7 @@ enum kvm_pgtable_walk_flags {
> KVM_PGTABLE_WALK_IGNORE_EAGAIN = BIT(4),
> KVM_PGTABLE_WALK_SKIP_BBM_TLBI = BIT(5),
> KVM_PGTABLE_WALK_SKIP_CMO = BIT(6),
> + KVM_PGTABLE_WALK_SKIP_S2_TLBI = BIT(7),
> };
>
> struct kvm_pgtable_visit_ctx {
> @@ -717,6 +720,20 @@ int kvm_pgtable_stage2_annotate(struct kvm_pgtable *pgt, u64 addr, u64 size,
> */
> int kvm_pgtable_stage2_unmap(struct kvm_pgtable *pgt, u64 addr, u64 size);
>
> +/**
> + * kvm_pgtable_stage2_unmap_notlbi() - Remove a mapping from a guest stage-2 page-table
> + * without TLB invalidation.
> + * @pgt: Page-table structure initialised by kvm_pgtable_stage2_init*().
> + * @addr: Intermediate physical address from which to remove the mapping.
> + * @size: Size of the mapping.
> + *
> + * Same as kvm_pgtable_stage2_unmap(), but does not invalidate the
> + * TLBs, which is the responsibility of the caller. Use with caution!
> + *
> + * Return: 0 on success, negative error code on failure.
> + */
> +int kvm_pgtable_stage2_unmap_notlbi(struct kvm_pgtable *pgt, u64 addr, u64 size);
> +
> /**
> * kvm_pgtable_stage2_wrprotect() - Write-protect guest stage-2 address range
> * without TLB invalidation.
> diff --git a/arch/arm64/include/asm/kvm_pkvm.h b/arch/arm64/include/asm/kvm_pkvm.h
> index beea00e693a0a..273013c98ff17 100644
> --- a/arch/arm64/include/asm/kvm_pkvm.h
> +++ b/arch/arm64/include/asm/kvm_pkvm.h
> @@ -214,6 +214,7 @@ int pkvm_pgtable_stage2_map(struct kvm_pgtable *pgt, u64 addr, u64 size, u64 phy
> enum kvm_pgtable_prot prot, void *mc,
> enum kvm_pgtable_walk_flags flags);
> int pkvm_pgtable_stage2_unmap(struct kvm_pgtable *pgt, u64 addr, u64 size);
> +int pkvm_pgtable_stage2_unmap_notlbi(struct kvm_pgtable *pgt, u64 addr, u64 size);
> int pkvm_pgtable_stage2_wrprotect(struct kvm_pgtable *pgt, u64 addr, u64 size);
> int pkvm_pgtable_stage2_flush(struct kvm_pgtable *pgt, u64 addr, u64 size);
> bool pkvm_pgtable_stage2_test_clear_young(struct kvm_pgtable *pgt, u64 addr, u64 size, bool mkold);
> diff --git a/arch/arm64/kvm/hyp/pgtable.c b/arch/arm64/kvm/hyp/pgtable.c
> index b74dd5ce1efd3..6603fc236daa2 100644
> --- a/arch/arm64/kvm/hyp/pgtable.c
> +++ b/arch/arm64/kvm/hyp/pgtable.c
> @@ -29,6 +29,11 @@ static bool kvm_pgtable_walk_skip_cmo(const struct kvm_pgtable_visit_ctx *ctx)
> return unlikely(ctx->flags & KVM_PGTABLE_WALK_SKIP_CMO);
> }
>
> +static bool kvm_pgtable_walk_skip_s2_tlbi(const struct kvm_pgtable_visit_ctx *ctx)
> +{
> + return unlikely(ctx->flags & KVM_PGTABLE_WALK_SKIP_S2_TLBI);
> +}
> +
> static bool kvm_block_mapping_supported(const struct kvm_pgtable_visit_ctx *ctx, u64 phys)
> {
> u64 granule = kvm_granule_size(ctx->level);
> @@ -905,12 +910,14 @@ static void stage2_unmap_put_pte(const struct kvm_pgtable_visit_ctx *ctx,
> if (kvm_pte_valid(ctx->old)) {
> kvm_clear_pte(ctx->ptep);
>
> - if (kvm_pte_table(ctx->old, ctx->level)) {
> - kvm_call_hyp(__kvm_tlb_flush_vmid_ipa, mmu, ctx->addr,
> - TLBI_TTL_UNKNOWN);
> - } else if (!stage2_unmap_defer_tlb_flush(pgt)) {
> - kvm_call_hyp(__kvm_tlb_flush_vmid_ipa, mmu, ctx->addr,
> - ctx->level);
> + if (!kvm_pgtable_walk_skip_s2_tlbi(ctx)) {
> + if (kvm_pte_table(ctx->old, ctx->level)) {
> + kvm_call_hyp(__kvm_tlb_flush_vmid_ipa, mmu, ctx->addr,
> + TLBI_TTL_UNKNOWN);
> + } else if (!stage2_unmap_defer_tlb_flush(pgt)) {
> + kvm_call_hyp(__kvm_tlb_flush_vmid_ipa, mmu, ctx->addr,
> + ctx->level);
> + }
> }
> }
>
> @@ -1195,23 +1202,36 @@ static int stage2_unmap_walker(const struct kvm_pgtable_visit_ctx *ctx,
> return 0;
> }
>
> -int kvm_pgtable_stage2_unmap(struct kvm_pgtable *pgt, u64 addr, u64 size)
> +static int __kvm_pgtable_stage2_unmap(struct kvm_pgtable *pgt,
> + enum kvm_pgtable_walk_flags flags,
> + u64 addr, u64 size)
> {
> int ret;
> struct kvm_pgtable_walker walker = {
> .cb = stage2_unmap_walker,
> .arg = pgt,
> - .flags = KVM_PGTABLE_WALK_LEAF | KVM_PGTABLE_WALK_TABLE_POST,
> + .flags = KVM_PGTABLE_WALK_LEAF | KVM_PGTABLE_WALK_TABLE_POST | flags,
> };
>
> ret = kvm_pgtable_walk(pgt, addr, size, &walker);
> - if (stage2_unmap_defer_tlb_flush(pgt))
> + if (stage2_unmap_defer_tlb_flush(pgt) &&
> + !(flags & KVM_PGTABLE_WALK_SKIP_S2_TLBI))
> /* Perform the deferred TLB invalidations */
> kvm_tlb_flush_vmid_range(pgt->mmu, addr, size);
>
> return ret;
> }
>
> +int kvm_pgtable_stage2_unmap(struct kvm_pgtable *pgt, u64 addr, u64 size)
> +{
> + return __kvm_pgtable_stage2_unmap(pgt, 0, addr, size);
> +}
> +
> +int kvm_pgtable_stage2_unmap_notlbi(struct kvm_pgtable *pgt, u64 addr, u64 size)
> +{
> + return __kvm_pgtable_stage2_unmap(pgt, KVM_PGTABLE_WALK_SKIP_S2_TLBI, addr, size);
> +}
> +
> struct stage2_attr_data {
> kvm_pte_t attr_set;
> kvm_pte_t attr_clr;
> diff --git a/arch/arm64/kvm/pkvm.c b/arch/arm64/kvm/pkvm.c
> index 8e4c6e4bec123..ec151005fbe4d 100644
> --- a/arch/arm64/kvm/pkvm.c
> +++ b/arch/arm64/kvm/pkvm.c
> @@ -488,6 +488,8 @@ int pkvm_pgtable_stage2_unmap(struct kvm_pgtable *pgt, u64 addr, u64 size)
> return __pkvm_pgtable_stage2_unshare(pgt, addr, addr + size);
> }
>
> +int pkvm_pgtable_stage2_unmap_notlbi(struct kvm_pgtable *pgt, u64 addr, u64 size) __alias(pkvm_pgtable_stage2_unmap);
> +
> int pkvm_pgtable_stage2_wrprotect(struct kvm_pgtable *pgt, u64 addr, u64 size)
> {
> struct kvm *kvm = kvm_s2_mmu_to_kvm(pgt->mmu);
> --
> 2.47.3
>
>
next prev parent reply other threads:[~2026-09-14 8:48 UTC|newest]
Thread overview: 12+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-12 10:48 [PATCH 0/4] KVM: arm64: Reduce overhead of full S2 teardown Marc Zyngier
2026-09-12 10:48 ` [PATCH 1/4] KVM: arm64: pgtable: Add Stage-2 unmap without TLBI primitive Marc Zyngier
2026-09-14 8:48 ` Mark Rutland [this message]
2026-09-14 9:23 ` Mark Rutland
2026-09-12 10:48 ` [PATCH 2/4] KVM: arm64: MMU: Add kvm_stage2_unmap_all() helper Marc Zyngier
2026-09-12 10:48 ` [PATCH 3/4] KVM: arm64: nv: Move full s2_mmu unmap over to kvm_stage2_unmap_all() Marc Zyngier
2026-09-12 10:48 ` [PATCH 4/4] KVM: arm64: nv: Move TLBI VMALLS12E1* emulation " Marc Zyngier
2026-09-13 23:27 ` [PATCH 0/4] KVM: arm64: Reduce overhead of full S2 teardown Itaru Kitayama
2026-09-14 6:44 ` Shuai Xue
2026-09-14 8:14 ` Marc Zyngier
2026-09-14 9:06 ` Shuai Xue
2026-09-15 23:19 ` Oliver Upton
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aqe0w0myu7iIT_tR@J2N7QTR9R3 \
--to=mark.rutland@arm.com \
--cc=fuad.tabba@linux.dev \
--cc=joey.gouly@arm.com \
--cc=kvmarm@lists.linux.dev \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=maz@kernel.org \
--cc=oupton@kernel.org \
--cc=seiden@linux.ibm.com \
--cc=suzuki.poulose@arm.com \
--cc=wanghan@linux.alibaba.com \
--cc=weilin.chang@arm.com \
--cc=xueshuai@linux.alibaba.com \
--cc=yuzenghui@huawei.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox