From: Tian Zheng <zhengtian10@huawei.com>
To: <maz@kernel.org>, <oupton@kernel.org>, <catalin.marinas@arm.com>,
<will@kernel.org>, <corbet@lwn.net>, <pbonzini@redhat.com>,
<zhengtian10@huawei.com>, <leo.bras@arm.com>
Cc: <yuzenghui@huawei.com>, <wangzhou1@hisilicon.com>,
<yangjinqian1@huawei.com>, <caijian11@h-partners.com>,
<liuyonglong@huawei.com>, <tangchengchang@huawei.com>,
<yezhenyu2@huawei.com>, <yubihong@huawei.com>,
<linuxarm@huawei.com>, <joey.gouly@arm.com>,
<kvmarm@lists.linux.dev>, <kvm@vger.kernel.org>,
<linux-arm-kernel@lists.infradead.org>,
<linux-kernel@vger.kernel.org>, <seiden@linux.ibm.com>,
<suzuki.poulose@arm.com>, <fuad.tabba@linux.dev>,
<mark.rutland@arm.com>, <seanjc@google.com>,
<rdunlap@infradead.org>, <linux-doc@vger.kernel.org>,
<linux-kselftest@vger.kernel.org>, <skhan@linuxfoundation.org>
Subject: [PATCH v5 05/15] KVM: arm64: Harvest stage-2 dirty state into the host folio account
Date: Tue, 29 Sep 2026 18:36:45 +0800 [thread overview]
Message-ID: <20260929103655.85107-6-zhengtian10@huawei.com> (raw)
In-Reply-To: <20260929103655.85107-1-zhengtian10@huawei.com>
With VTCR_EL2.HD set, hardware promotes writable-clean descriptors to
writable-dirty without any VM exit, and the only record of the write
is the S2AP[1] bit in the stage-2 PTE. The write never passed through
the host stage-1, and no unmap path reads the bit, so the record dies
with the PTE and reclaim may discard written guest data (silent
corruption).
The fault paths already mark the folio speculatively at fault-in, but
the mapping lifecycle still needs exact harvesting, mirroring
zap_present_folio_ptes() in the generic mm:
- stage2_unmap_walker(): harvest the output address of a valid leaf
with S2AP[1] set before tearing it down.
- stage2_wrprotect_walker(): a WD -> WC transition drops the dirty
state, so harvest before the clear.
Both hooks go through a new kvm_pgtable_mm_ops::mark_page_dirty
callback, as the walkers are also compiled into the nVHE hypervisor,
where SetPageDirty() is unavailable. pKVM leaves the callback NULL.
The folio is dirtied at its head, so one harvest covers a whole block
mapping. MMIO/PFNMAP ranges and reserved pages are skipped, mirroring
kvm_is_ad_tracked_page().
Signed-off-by: Tian Zheng <zhengtian10@huawei.com>
---
arch/arm64/include/asm/kvm_pgtable.h | 2 ++
arch/arm64/kvm/hyp/pgtable.c | 14 ++++++++++++--
arch/arm64/kvm/mmu.c | 16 ++++++++++++++++
3 files changed, 30 insertions(+), 2 deletions(-)
diff --git a/arch/arm64/include/asm/kvm_pgtable.h b/arch/arm64/include/asm/kvm_pgtable.h
index 379031c74cbc..ea71c13615f7 100644
--- a/arch/arm64/include/asm/kvm_pgtable.h
+++ b/arch/arm64/include/asm/kvm_pgtable.h
@@ -246,6 +246,8 @@ struct kvm_pgtable_mm_ops {
phys_addr_t (*virt_to_phys)(void *addr);
void (*dcache_clean_inval_poc)(void *addr, size_t size);
void (*icache_inval_pou)(void *addr, size_t size);
+ /* NULL where folios are not tracked. */
+ void (*mark_page_dirty)(u64 pa);
};
/**
diff --git a/arch/arm64/kvm/hyp/pgtable.c b/arch/arm64/kvm/hyp/pgtable.c
index aa0448d3a6a4..9dde7e779699 100644
--- a/arch/arm64/kvm/hyp/pgtable.c
+++ b/arch/arm64/kvm/hyp/pgtable.c
@@ -1182,8 +1182,13 @@ static int stage2_unmap_walker(const struct kvm_pgtable_visit_ctx *ctx,
if (mm_ops->page_count(childp) != 1)
return 0;
- } else if (stage2_pte_cacheable(pgt, ctx->old)) {
- need_flush = !cpus_have_final_cap(ARM64_HAS_STAGE2_FWB);
+ } else {
+ if (stage2_pte_cacheable(pgt, ctx->old))
+ need_flush = !cpus_have_final_cap(ARM64_HAS_STAGE2_FWB);
+
+ if ((ctx->old & KVM_PTE_LEAF_ATTR_LO_S2_S2AP_W) &&
+ mm_ops->mark_page_dirty)
+ mm_ops->mark_page_dirty(kvm_pte_to_phys(ctx->old));
}
/*
@@ -1302,6 +1307,11 @@ static int stage2_wrprotect_walker(const struct kvm_pgtable_visit_ctx *ctx,
if (ctx->level < KVM_PGTABLE_LAST_LEVEL)
new &= ~KVM_PTE_LEAF_ATTR_HI_S2_DBM;
+ if (kvm_pte_valid(ctx->old) && ctx->old != new &&
+ (ctx->old & KVM_PTE_LEAF_ATTR_LO_S2_S2AP_W) &&
+ ctx->mm_ops->mark_page_dirty)
+ ctx->mm_ops->mark_page_dirty(kvm_pte_to_phys(ctx->old));
+
/*
* We may race with the CPU trying to set the access flag here,
* but worst-case the access flag update gets lost and will be
diff --git a/arch/arm64/kvm/mmu.c b/arch/arm64/kvm/mmu.c
index 698a87e85a6d..85a98d2c23a9 100644
--- a/arch/arm64/kvm/mmu.c
+++ b/arch/arm64/kvm/mmu.c
@@ -897,6 +897,21 @@ static int get_user_mapping_size(struct kvm *kvm, u64 addr)
return BIT(ARM64_HW_PGTABLE_LEVEL_SHIFT(level));
}
+static void kvm_s2_mark_page_dirty(u64 pa)
+{
+ unsigned long pfn = pa >> PAGE_SHIFT;
+ struct page *page;
+
+ if (!pfn_valid(pfn))
+ return;
+
+ page = pfn_to_page(pfn);
+ if (PageReserved(page))
+ return;
+
+ SetPageDirty(page);
+}
+
static struct kvm_pgtable_mm_ops kvm_s2_mm_ops = {
.zalloc_page = stage2_memcache_zalloc_page,
.zalloc_pages_exact = kvm_s2_zalloc_pages_exact,
@@ -909,6 +924,7 @@ static struct kvm_pgtable_mm_ops kvm_s2_mm_ops = {
.virt_to_phys = kvm_host_pa,
.dcache_clean_inval_poc = clean_dcache_guest_page,
.icache_inval_pou = invalidate_icache_guest_page,
+ .mark_page_dirty = kvm_s2_mark_page_dirty,
};
static int kvm_init_ipa_range(struct kvm_s2_mmu *mmu, unsigned long type)
--
2.43.0
next prev parent reply other threads:[~2026-09-29 12:24 UTC|newest]
Thread overview: 20+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-29 10:36 [PATCH v5 00/15] KVM: arm64: FEAT_HDBSS support for stage-2 dirty tracking Tian Zheng
2026-09-29 10:36 ` [PATCH v5 01/15] KVM: arm64: pgtables: Change write bit from S2AP_W to DBM Tian Zheng
2026-09-30 0:25 ` Oliver Upton
2026-09-30 2:44 ` Tian Zheng
2026-09-29 10:36 ` [PATCH v5 02/15] KVM: arm64: Add KVM_PGTABLE_PROT_DIRTY Tian Zheng
2026-09-30 0:35 ` Oliver Upton
2026-09-30 2:57 ` Tian Zheng
2026-09-29 10:36 ` [PATCH v5 03/15] KVM: arm64: Introduce a dedicated walker for stage2 write-protect Tian Zheng
2026-09-29 10:36 ` [PATCH v5 04/15] KVM: arm64: Add KVM_REQ_RELOAD_STAGE2 Tian Zheng
2026-09-29 10:36 ` Tian Zheng [this message]
2026-09-29 10:36 ` [PATCH v5 06/15] KVM: arm64: Add support for FEAT_HDBSS Tian Zheng
2026-09-29 10:36 ` [PATCH v5 07/15] KVM: arm64: Add HDBSS per-vCPU buffer management Tian Zheng
2026-09-29 10:36 ` [PATCH v5 08/15] KVM: arm64: Flush the HDBSS buffer on VM exit Tian Zheng
2026-09-29 10:36 ` [PATCH v5 09/15] KVM: arm64: Handle HDBSS faults Tian Zheng
2026-09-29 10:36 ` [PATCH v5 10/15] KVM: Add kvm_arch_dirty_ring_size_updated() hook Tian Zheng
2026-09-29 10:36 ` [PATCH v5 11/15] KVM: arm64: Reserve dirty ring space for the HDBSS buffer Tian Zheng
2026-09-29 10:36 ` [PATCH v5 12/15] KVM: arm64: Derive the VM hardware dirty mode from dirty logging Tian Zheng
2026-09-29 10:36 ` [PATCH v5 13/15] KVM: arm64: Add HDBSS buffer size ioctl for dirty-bitmap mode Tian Zheng
2026-09-29 10:36 ` [PATCH v5 14/15] KVM: arm64: Document HDBSS buffer size ioctl Tian Zheng
2026-09-29 10:36 ` [PATCH v5 15/15] KVM: arm64: selftests: Add HDBSS buffer size ioctl interface test Tian Zheng
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260929103655.85107-6-zhengtian10@huawei.com \
--to=zhengtian10@huawei.com \
--cc=caijian11@h-partners.com \
--cc=catalin.marinas@arm.com \
--cc=corbet@lwn.net \
--cc=fuad.tabba@linux.dev \
--cc=joey.gouly@arm.com \
--cc=kvm@vger.kernel.org \
--cc=kvmarm@lists.linux.dev \
--cc=leo.bras@arm.com \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linuxarm@huawei.com \
--cc=liuyonglong@huawei.com \
--cc=mark.rutland@arm.com \
--cc=maz@kernel.org \
--cc=oupton@kernel.org \
--cc=pbonzini@redhat.com \
--cc=rdunlap@infradead.org \
--cc=seanjc@google.com \
--cc=seiden@linux.ibm.com \
--cc=skhan@linuxfoundation.org \
--cc=suzuki.poulose@arm.com \
--cc=tangchengchang@huawei.com \
--cc=wangzhou1@hisilicon.com \
--cc=will@kernel.org \
--cc=yangjinqian1@huawei.com \
--cc=yezhenyu2@huawei.com \
--cc=yubihong@huawei.com \
--cc=yuzenghui@huawei.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox