From: Vincent Donnefort <vdonnefort@google.com>
To: sashiko-reviews@lists.linux.dev
Cc: kvmarm@lists.linux.dev, Oliver Upton <oupton@kernel.org>,
Marc Zyngier <maz@kernel.org>
Subject: Re: [PATCH 18/20] KVM: arm64: Add PKVM_HYP_REQ_SPLIT
Date: Mon, 7 Sep 2026 11:23:10 +0100 [thread overview]
Message-ID: <ap6QjjJxysWVYCVM@google.com> (raw)
In-Reply-To: <20260803105411.B6ACE1F000E9@smtp.kernel.org>
On Mon, Aug 03, 2026 at 10:54:11AM +0000, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 5 potential issue(s) to consider:
> - [Critical] Time-of-Check to Time-of-Use (TOCTOU) race in `__pkvm_pgtable_stage2_split` due to missing MMU notifier retry protocol.
> - [Critical] Using `pin_user_pages()` introduces a `FOLL_PIN` counting bias that causes massive memory leaks during KVM teardown.
> - [High] Memory cgroup accounting bypass when allocating spare `pkvm_mapping` structures.
> - [Medium] Splitting read-only guest mappings is incorrectly forbidden by requiring write access.
> - [Low] `__pkvm_pgtable_stage2_split()` mixes `goto`-based error handling with scope-based `__free(kfree)` cleanup.
> --
>
> --- Patch [18]: [PATCH 18/20] KVM: arm64: Add PKVM_HYP_REQ_SPLIT ---
> commit eb20478193ba08806cd19923ba4b65163db37c40
> Author: Vincent Donnefort <vdonnefort@google.com>
>
> KVM: arm64: Add PKVM_HYP_REQ_SPLIT
>
> This commit introduces a new pKVM hypervisor request, PKVM_HYP_REQ_SPLIT,
> enabling the hypervisor to request the host to split stage-2 huge mappings.
> This ensures synchronous splitting of both the guest stage-2 page tables
> and the host's pkvm_mappings red-black tree, preventing concurrent
> teardowns from observing inconsistent page sizes.
>
> > diff --git a/arch/arm64/kvm/pkvm.c b/arch/arm64/kvm/pkvm.c
> > index 379bdc2b258a6..089b77cf2f6a9 100644
> > --- a/arch/arm64/kvm/pkvm.c
> > +++ b/arch/arm64/kvm/pkvm.c
>
> [ ... ]
>
> > +static int pkvm_mapping_alloc_spares(struct list_head *head, u64 nr_spares)
> > +{
> > + struct pkvm_mapping *m;
> > +
> > + while (nr_spares--) {
> > + m = kzalloc_obj(*m);
>
> [Severity: High]
> Does this bypass memory cgroup limits?
>
> This allocation uses kzalloc_obj() which defaults to GFP_KERNEL without
> the __GFP_ACCOUNT flag.
>
> Because each PMD split allocates 511 structures, a guest could repeatedly
> trigger splits across its memory footprint.
>
> Can this allow the guest to consume large amounts of uncharged host memory,
> potentially leading to host memory exhaustion?
>
> > + if (!m) {
> > + pkvm_mapping_free_spares(head);
> > + return -ENOMEM;
> > + }
> > +
> > + list_add(&m->list, head);
> > + }
> > +
> > + return 0;
> > +}
>
> [ ... ]
>
> > +static int __pkvm_pgtable_stage2_split(struct kvm_vcpu *vcpu, phys_addr_t ipa, u64 size)
> > +{
> > + struct kvm_hyp_memcache *mc = &vcpu->arch.pkvm_memcache;
> > + struct kvm_pgtable *pgt = vcpu->arch.hw_mmu->pgt;
> > + struct page **pages __free(kfree) = NULL;
>
> [Severity: Low]
> Is it safe to mix scope-based cleanup with goto error handling in this
> function?
>
> The pages array is declared with __free(kfree), but the function uses goto
> labels like goto unlock_srcu and goto unlock_mmu for unwinding locks and
> freeing the spares list.
>
> The kernel cleanup guidelines note that functions should use either
> scope-based cleanup for all resources or goto labels for all resources,
> but not mix them, as it creates confusing ownership semantics.
>
> > + struct mm_struct *mm = current->mm;
> > + struct kvm_memory_slot *memslot;
>
> [ ... ]
>
> > + idx = srcu_read_lock(&kvm->srcu);
> > + memslot = gfn_to_memslot(kvm, gfn);
> > + hva = gfn_to_hva_memslot_prot(memslot, gfn, &writable);
> > + if (kvm_is_error_hva(hva) || !writable) {
>
> [Severity: Medium]
> Should read-only mappings be allowed to split?
>
> This explicitly checks that the memslot is writable and fails with -EFAULT
> otherwise. It also passes FOLL_WRITE to pin_user_pages() later on.
>
> However, splitting a huge page block into smaller PTEs changes the mapping
> granularity without modifying the underlying memory contents.
>
> Does requiring write access prevent the hypervisor from splitting read-only
> stage-2 block mappings, such as guest ROMs?
>
> > + ret = -EFAULT;
> > + goto unlock_srcu;
> > + }
> > +
> > + mmap_read_lock(mm);
> > + ret = pin_user_pages(hva, nr_pages, FOLL_HWPOISON | FOLL_LONGTERM | FOLL_WRITE, pages);
>
> [Severity: Critical]
> Does using pin_user_pages() here cause a permanent memory leak during KVM
> teardown?
>
> Calling pin_user_pages() adds a GUP_PIN_COUNTING_BIAS to the page refcount.
> However, normal KVM stage-2 teardown drops references using standard
> put_page() semantics which only subtracts one from the refcount.
>
> When the VM is destroyed, will these tail pages retain phantom references
> and be permanently leaked?
>
> > + mmap_read_unlock(mm);
> > + if (ret != nr_pages) {
> > + if (ret > 0)
> > + unpin_user_pages(pages, ret);
> > + ret = -EFAULT;
> > + goto unlock_srcu;
> > + }
> > +
> > + write_lock(&kvm->mmu_lock);
> > + mapping = pkvm_mapping_iter_first(&pgt->pkvm_mappings, ipa, ipa + size - 1);
>
> [Severity: Critical]
> Is there a Time-of-Check to Time-of-Use race here since the MMU
> notifier retry protocol is missing?
>
> The resolution of the HVA to PFNs via pin_user_pages() happens outside
> kvm->mmu_lock. A concurrent host MMU invalidation could unmap the region
> and remove the pkvm_mapping, and a subsequent guest fault could establish
> a new huge page mapping at the same IPA.
This is called on the back of a hypervisor request. This is only for
protected-VM and we do have an existing pin on the huge-page, preventing
any unmap.
We do not use MMU notifier for protected VMs.
--
Vincent
>
> When this thread acquires kvm->mmu_lock, it finds the new mapping in the
> tree and proceeds to split it using the PFNs derived from the old mapping.
>
> Without capturing kvm->mmu_invalidate_seq before GUP and checking
> mmu_invalidate_retry() under the lock, can this lead to a refcount
> underflow on the new huge page and a permanent leak of the old pages?
>
> > + if (!pkvm_mapping_can_split(mapping)) {
> > + ret = -EINVAL;
> > + goto unlock_mmu;
> > + }
>
> --
> Sashiko AI review · https://sashiko.dev/#/patchset/20260803100904.3563942-1-vdonnefort@google.com?part=18
next prev parent reply other threads:[~2026-09-07 10:23 UTC|newest]
Thread overview: 55+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-03 10:08 [PATCH 00/20] Huge mapping support for protected VMs Vincent Donnefort
2026-08-03 10:08 ` [PATCH 01/20] KVM: arm64: Prefault host stage-2 entries on block split Vincent Donnefort
2026-08-03 10:31 ` sashiko-bot
2026-09-07 18:34 ` Wei-Lin Chang
2026-08-03 10:08 ` [PATCH 02/20] KVM: arm64: Propagate host stage-2 annotated " Vincent Donnefort
2026-09-07 13:58 ` Wei-Lin Chang
2026-09-08 10:07 ` Vincent Donnefort
2026-08-03 10:08 ` [PATCH 03/20] KVM: arm64: Allow block-level stage-2 annotation Vincent Donnefort
2026-08-03 10:40 ` sashiko-bot
2026-08-03 10:08 ` [PATCH 04/20] KVM: arm64: Use block-level annotations when setting up the host stage-2 Vincent Donnefort
2026-09-07 14:37 ` Wei-Lin Chang
2026-09-08 10:18 ` Vincent Donnefort
2026-08-03 10:08 ` [PATCH 05/20] KVM: arm64: Make pKVM ownership selftest an HVC Vincent Donnefort
2026-08-03 10:08 ` [PATCH 06/20] KVM: arm64: Add a range to __pkvm_host_share/unshare_hyp() Vincent Donnefort
2026-08-03 10:08 ` [PATCH 07/20] KVM: arm64: Add a range to __pkvm_host_donate_guest() Vincent Donnefort
2026-08-03 10:26 ` sashiko-bot
2026-08-03 10:08 ` [PATCH 08/20] KVM: arm64: Add a range to __pkvm_host_reclaim_page_guest() Vincent Donnefort
2026-08-03 10:28 ` sashiko-bot
2026-09-07 14:52 ` Wei-Lin Chang
2026-09-08 10:22 ` Vincent Donnefort
2026-08-03 10:08 ` [PATCH 09/20] KVM: arm64: Add a range to __pkvm_guest_share_host() Vincent Donnefort
2026-08-03 10:08 ` [PATCH 10/20] KVM: arm64: Add a range to __pkvm_guest_unshare_host() Vincent Donnefort
2026-08-03 10:08 ` [PATCH 11/20] KVM: arm64: Add a range to pKVM ownership selftest Vincent Donnefort
2026-08-03 10:48 ` sashiko-bot
2026-09-07 10:13 ` Vincent Donnefort
2026-09-07 14:58 ` Wei-Lin Chang
2026-09-08 10:25 ` Vincent Donnefort
2026-09-10 10:04 ` Vincent Donnefort
2026-08-03 10:08 ` [PATCH 12/20] KVM: arm64: Handle huge mappings in __pkvm_host_force_reclaim_page_guest() Vincent Donnefort
2026-09-07 16:52 ` Wei-Lin Chang
2026-09-08 10:26 ` Vincent Donnefort
2026-08-03 10:08 ` [PATCH 13/20] KVM: arm64: Handle huge mappings in __pkvm_vcpu_in_poison_fault() Vincent Donnefort
2026-08-03 10:08 ` [PATCH 14/20] KVM: arm64: pkvm: Warn on guest stage-2 block collapse Vincent Donnefort
2026-08-03 10:08 ` [PATCH 15/20] KVM: arm64: Add pkvm_hyp_req infrastructure Vincent Donnefort
2026-08-03 10:46 ` sashiko-bot
2026-08-03 10:09 ` [PATCH 16/20] KVM: arm64: Add __pkvm_host_split_guest HVC Vincent Donnefort
2026-08-03 10:54 ` sashiko-bot
2026-09-07 10:14 ` Vincent Donnefort
2026-09-07 17:39 ` Wei-Lin Chang
2026-09-08 14:46 ` Vincent Donnefort
2026-08-03 10:09 ` [PATCH 17/20] KVM: arm64: Extend pKVM page ownership selftests to cover guest block split Vincent Donnefort
2026-08-03 11:01 ` sashiko-bot
2026-08-03 10:09 ` [PATCH 18/20] KVM: arm64: Add PKVM_HYP_REQ_SPLIT Vincent Donnefort
2026-08-03 10:54 ` sashiko-bot
2026-09-07 10:23 ` Vincent Donnefort [this message]
2026-09-08 10:35 ` Wei-Lin Chang
2026-09-07 18:15 ` Wei-Lin Chang
2026-09-08 14:56 ` Vincent Donnefort
2026-08-03 10:09 ` [PATCH 19/20] KVM: arm64: Raise PKVM_HYP_REQ_SPLIT on guest to host sharing Vincent Donnefort
2026-08-03 11:02 ` sashiko-bot
2026-09-07 10:26 ` Vincent Donnefort
2026-08-03 10:09 ` [PATCH 20/20] KVM: arm64: Stage-2 huge mappings for protected VMs Vincent Donnefort
2026-08-03 11:04 ` sashiko-bot
2026-09-07 18:31 ` Wei-Lin Chang
2026-09-07 18:41 ` [PATCH 00/20] Huge mapping support " Wei-Lin Chang
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=ap6QjjJxysWVYCVM@google.com \
--to=vdonnefort@google.com \
--cc=kvmarm@lists.linux.dev \
--cc=maz@kernel.org \
--cc=oupton@kernel.org \
--cc=sashiko-reviews@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox