From: sashiko-bot@kernel.org
To: "Vincent Donnefort" <vdonnefort@google.com>
Cc: kvmarm@lists.linux.dev, "Oliver Upton" <oupton@kernel.org>,
"Marc Zyngier" <maz@kernel.org>
Subject: Re: [PATCH 18/20] KVM: arm64: Add PKVM_HYP_REQ_SPLIT
Date: Mon, 03 Aug 2026 10:54:11 +0000 [thread overview]
Message-ID: <20260803105411.B6ACE1F000E9@smtp.kernel.org> (raw)
In-Reply-To: <20260803100904.3563942-19-vdonnefort@google.com>
Thank you for your contribution! Sashiko AI review found 5 potential issue(s) to consider:
- [Critical] Time-of-Check to Time-of-Use (TOCTOU) race in `__pkvm_pgtable_stage2_split` due to missing MMU notifier retry protocol.
- [Critical] Using `pin_user_pages()` introduces a `FOLL_PIN` counting bias that causes massive memory leaks during KVM teardown.
- [High] Memory cgroup accounting bypass when allocating spare `pkvm_mapping` structures.
- [Medium] Splitting read-only guest mappings is incorrectly forbidden by requiring write access.
- [Low] `__pkvm_pgtable_stage2_split()` mixes `goto`-based error handling with scope-based `__free(kfree)` cleanup.
--
--- Patch [18]: [PATCH 18/20] KVM: arm64: Add PKVM_HYP_REQ_SPLIT ---
commit eb20478193ba08806cd19923ba4b65163db37c40
Author: Vincent Donnefort <vdonnefort@google.com>
KVM: arm64: Add PKVM_HYP_REQ_SPLIT
This commit introduces a new pKVM hypervisor request, PKVM_HYP_REQ_SPLIT,
enabling the hypervisor to request the host to split stage-2 huge mappings.
This ensures synchronous splitting of both the guest stage-2 page tables
and the host's pkvm_mappings red-black tree, preventing concurrent
teardowns from observing inconsistent page sizes.
> diff --git a/arch/arm64/kvm/pkvm.c b/arch/arm64/kvm/pkvm.c
> index 379bdc2b258a6..089b77cf2f6a9 100644
> --- a/arch/arm64/kvm/pkvm.c
> +++ b/arch/arm64/kvm/pkvm.c
[ ... ]
> +static int pkvm_mapping_alloc_spares(struct list_head *head, u64 nr_spares)
> +{
> + struct pkvm_mapping *m;
> +
> + while (nr_spares--) {
> + m = kzalloc_obj(*m);
[Severity: High]
Does this bypass memory cgroup limits?
This allocation uses kzalloc_obj() which defaults to GFP_KERNEL without
the __GFP_ACCOUNT flag.
Because each PMD split allocates 511 structures, a guest could repeatedly
trigger splits across its memory footprint.
Can this allow the guest to consume large amounts of uncharged host memory,
potentially leading to host memory exhaustion?
> + if (!m) {
> + pkvm_mapping_free_spares(head);
> + return -ENOMEM;
> + }
> +
> + list_add(&m->list, head);
> + }
> +
> + return 0;
> +}
[ ... ]
> +static int __pkvm_pgtable_stage2_split(struct kvm_vcpu *vcpu, phys_addr_t ipa, u64 size)
> +{
> + struct kvm_hyp_memcache *mc = &vcpu->arch.pkvm_memcache;
> + struct kvm_pgtable *pgt = vcpu->arch.hw_mmu->pgt;
> + struct page **pages __free(kfree) = NULL;
[Severity: Low]
Is it safe to mix scope-based cleanup with goto error handling in this
function?
The pages array is declared with __free(kfree), but the function uses goto
labels like goto unlock_srcu and goto unlock_mmu for unwinding locks and
freeing the spares list.
The kernel cleanup guidelines note that functions should use either
scope-based cleanup for all resources or goto labels for all resources,
but not mix them, as it creates confusing ownership semantics.
> + struct mm_struct *mm = current->mm;
> + struct kvm_memory_slot *memslot;
[ ... ]
> + idx = srcu_read_lock(&kvm->srcu);
> + memslot = gfn_to_memslot(kvm, gfn);
> + hva = gfn_to_hva_memslot_prot(memslot, gfn, &writable);
> + if (kvm_is_error_hva(hva) || !writable) {
[Severity: Medium]
Should read-only mappings be allowed to split?
This explicitly checks that the memslot is writable and fails with -EFAULT
otherwise. It also passes FOLL_WRITE to pin_user_pages() later on.
However, splitting a huge page block into smaller PTEs changes the mapping
granularity without modifying the underlying memory contents.
Does requiring write access prevent the hypervisor from splitting read-only
stage-2 block mappings, such as guest ROMs?
> + ret = -EFAULT;
> + goto unlock_srcu;
> + }
> +
> + mmap_read_lock(mm);
> + ret = pin_user_pages(hva, nr_pages, FOLL_HWPOISON | FOLL_LONGTERM | FOLL_WRITE, pages);
[Severity: Critical]
Does using pin_user_pages() here cause a permanent memory leak during KVM
teardown?
Calling pin_user_pages() adds a GUP_PIN_COUNTING_BIAS to the page refcount.
However, normal KVM stage-2 teardown drops references using standard
put_page() semantics which only subtracts one from the refcount.
When the VM is destroyed, will these tail pages retain phantom references
and be permanently leaked?
> + mmap_read_unlock(mm);
> + if (ret != nr_pages) {
> + if (ret > 0)
> + unpin_user_pages(pages, ret);
> + ret = -EFAULT;
> + goto unlock_srcu;
> + }
> +
> + write_lock(&kvm->mmu_lock);
> + mapping = pkvm_mapping_iter_first(&pgt->pkvm_mappings, ipa, ipa + size - 1);
[Severity: Critical]
Is there a Time-of-Check to Time-of-Use race here since the MMU
notifier retry protocol is missing?
The resolution of the HVA to PFNs via pin_user_pages() happens outside
kvm->mmu_lock. A concurrent host MMU invalidation could unmap the region
and remove the pkvm_mapping, and a subsequent guest fault could establish
a new huge page mapping at the same IPA.
When this thread acquires kvm->mmu_lock, it finds the new mapping in the
tree and proceeds to split it using the PFNs derived from the old mapping.
Without capturing kvm->mmu_invalidate_seq before GUP and checking
mmu_invalidate_retry() under the lock, can this lead to a refcount
underflow on the new huge page and a permanent leak of the old pages?
> + if (!pkvm_mapping_can_split(mapping)) {
> + ret = -EINVAL;
> + goto unlock_mmu;
> + }
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260803100904.3563942-1-vdonnefort@google.com?part=18
next prev parent reply other threads:[~2026-08-03 10:54 UTC|newest]
Thread overview: 32+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-03 10:08 [PATCH 00/20] Huge mapping support for protected VMs Vincent Donnefort
2026-08-03 10:08 ` [PATCH 01/20] KVM: arm64: Prefault host stage-2 entries on block split Vincent Donnefort
2026-08-03 10:31 ` sashiko-bot
2026-08-03 10:08 ` [PATCH 02/20] KVM: arm64: Propagate host stage-2 annotated " Vincent Donnefort
2026-08-03 10:08 ` [PATCH 03/20] KVM: arm64: Allow block-level stage-2 annotation Vincent Donnefort
2026-08-03 10:40 ` sashiko-bot
2026-08-03 10:08 ` [PATCH 04/20] KVM: arm64: Use block-level annotations when setting up the host stage-2 Vincent Donnefort
2026-08-03 10:08 ` [PATCH 05/20] KVM: arm64: Make pKVM ownership selftest an HVC Vincent Donnefort
2026-08-03 10:08 ` [PATCH 06/20] KVM: arm64: Add a range to __pkvm_host_share/unshare_hyp() Vincent Donnefort
2026-08-03 10:08 ` [PATCH 07/20] KVM: arm64: Add a range to __pkvm_host_donate_guest() Vincent Donnefort
2026-08-03 10:26 ` sashiko-bot
2026-08-03 10:08 ` [PATCH 08/20] KVM: arm64: Add a range to __pkvm_host_reclaim_page_guest() Vincent Donnefort
2026-08-03 10:28 ` sashiko-bot
2026-08-03 10:08 ` [PATCH 09/20] KVM: arm64: Add a range to __pkvm_guest_share_host() Vincent Donnefort
2026-08-03 10:08 ` [PATCH 10/20] KVM: arm64: Add a range to __pkvm_guest_unshare_host() Vincent Donnefort
2026-08-03 10:08 ` [PATCH 11/20] KVM: arm64: Add a range to pKVM ownership selftest Vincent Donnefort
2026-08-03 10:48 ` sashiko-bot
2026-08-03 10:08 ` [PATCH 12/20] KVM: arm64: Handle huge mappings in __pkvm_host_force_reclaim_page_guest() Vincent Donnefort
2026-08-03 10:08 ` [PATCH 13/20] KVM: arm64: Handle huge mappings in __pkvm_vcpu_in_poison_fault() Vincent Donnefort
2026-08-03 10:08 ` [PATCH 14/20] KVM: arm64: pkvm: Warn on guest stage-2 block collapse Vincent Donnefort
2026-08-03 10:08 ` [PATCH 15/20] KVM: arm64: Add pkvm_hyp_req infrastructure Vincent Donnefort
2026-08-03 10:46 ` sashiko-bot
2026-08-03 10:09 ` [PATCH 16/20] KVM: arm64: Add __pkvm_host_split_guest HVC Vincent Donnefort
2026-08-03 10:54 ` sashiko-bot
2026-08-03 10:09 ` [PATCH 17/20] KVM: arm64: Extend pKVM page ownership selftests to cover guest block split Vincent Donnefort
2026-08-03 11:01 ` sashiko-bot
2026-08-03 10:09 ` [PATCH 18/20] KVM: arm64: Add PKVM_HYP_REQ_SPLIT Vincent Donnefort
2026-08-03 10:54 ` sashiko-bot [this message]
2026-08-03 10:09 ` [PATCH 19/20] KVM: arm64: Raise PKVM_HYP_REQ_SPLIT on guest to host sharing Vincent Donnefort
2026-08-03 11:02 ` sashiko-bot
2026-08-03 10:09 ` [PATCH 20/20] KVM: arm64: Stage-2 huge mappings for protected VMs Vincent Donnefort
2026-08-03 11:04 ` sashiko-bot
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260803105411.B6ACE1F000E9@smtp.kernel.org \
--to=sashiko-bot@kernel.org \
--cc=kvmarm@lists.linux.dev \
--cc=maz@kernel.org \
--cc=oupton@kernel.org \
--cc=sashiko-reviews@lists.linux.dev \
--cc=vdonnefort@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.