All of lore.kernel.org
 help / color / mirror / Atom feed
From: Vincent Donnefort <vdonnefort@google.com>
To: sashiko-reviews@lists.linux.dev
Cc: kvmarm@lists.linux.dev, Oliver Upton <oupton@kernel.org>,
	Marc Zyngier <maz@kernel.org>
Subject: Re: [PATCH 18/20] KVM: arm64: Add PKVM_HYP_REQ_SPLIT
Date: Mon, 7 Sep 2026 11:23:10 +0100	[thread overview]
Message-ID: <ap6QjjJxysWVYCVM@google.com> (raw)
In-Reply-To: <20260803105411.B6ACE1F000E9@smtp.kernel.org>

On Mon, Aug 03, 2026 at 10:54:11AM +0000, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 5 potential issue(s) to consider:
> - [Critical] Time-of-Check to Time-of-Use (TOCTOU) race in `__pkvm_pgtable_stage2_split` due to missing MMU notifier retry protocol.
> - [Critical] Using `pin_user_pages()` introduces a `FOLL_PIN` counting bias that causes massive memory leaks during KVM teardown.
> - [High] Memory cgroup accounting bypass when allocating spare `pkvm_mapping` structures.
> - [Medium] Splitting read-only guest mappings is incorrectly forbidden by requiring write access.
> - [Low] `__pkvm_pgtable_stage2_split()` mixes `goto`-based error handling with scope-based `__free(kfree)` cleanup.
> --
> 
> --- Patch [18]: [PATCH 18/20] KVM: arm64: Add PKVM_HYP_REQ_SPLIT ---
> commit eb20478193ba08806cd19923ba4b65163db37c40
> Author: Vincent Donnefort <vdonnefort@google.com>
> 
> KVM: arm64: Add PKVM_HYP_REQ_SPLIT
> 
> This commit introduces a new pKVM hypervisor request, PKVM_HYP_REQ_SPLIT,
> enabling the hypervisor to request the host to split stage-2 huge mappings.
> This ensures synchronous splitting of both the guest stage-2 page tables
> and the host's pkvm_mappings red-black tree, preventing concurrent
> teardowns from observing inconsistent page sizes.
> 
> > diff --git a/arch/arm64/kvm/pkvm.c b/arch/arm64/kvm/pkvm.c
> > index 379bdc2b258a6..089b77cf2f6a9 100644
> > --- a/arch/arm64/kvm/pkvm.c
> > +++ b/arch/arm64/kvm/pkvm.c
> 
> [ ... ]
> 
> > +static int pkvm_mapping_alloc_spares(struct list_head *head, u64 nr_spares)
> > +{
> > +	struct pkvm_mapping *m;
> > +
> > +	while (nr_spares--) {
> > +		m = kzalloc_obj(*m);
> 
> [Severity: High]
> Does this bypass memory cgroup limits?
> 
> This allocation uses kzalloc_obj() which defaults to GFP_KERNEL without
> the __GFP_ACCOUNT flag.
> 
> Because each PMD split allocates 511 structures, a guest could repeatedly
> trigger splits across its memory footprint.
> 
> Can this allow the guest to consume large amounts of uncharged host memory,
> potentially leading to host memory exhaustion?
> 
> > +		if (!m) {
> > +			pkvm_mapping_free_spares(head);
> > +			return -ENOMEM;
> > +		}
> > +
> > +		list_add(&m->list, head);
> > +	}
> > +
> > +	return 0;
> > +}
> 
> [ ... ]
> 
> > +static int __pkvm_pgtable_stage2_split(struct kvm_vcpu *vcpu, phys_addr_t ipa, u64 size)
> > +{
> > +	struct kvm_hyp_memcache *mc = &vcpu->arch.pkvm_memcache;
> > +	struct kvm_pgtable *pgt = vcpu->arch.hw_mmu->pgt;
> > +	struct page **pages __free(kfree) = NULL;
> 
> [Severity: Low]
> Is it safe to mix scope-based cleanup with goto error handling in this
> function?
> 
> The pages array is declared with __free(kfree), but the function uses goto
> labels like goto unlock_srcu and goto unlock_mmu for unwinding locks and
> freeing the spares list.
> 
> The kernel cleanup guidelines note that functions should use either
> scope-based cleanup for all resources or goto labels for all resources,
> but not mix them, as it creates confusing ownership semantics.
> 
> > +	struct mm_struct *mm = current->mm;
> > +	struct kvm_memory_slot *memslot;
> 
> [ ... ]
> 
> > +	idx = srcu_read_lock(&kvm->srcu);
> > +	memslot = gfn_to_memslot(kvm, gfn);
> > +	hva = gfn_to_hva_memslot_prot(memslot, gfn, &writable);
> > +	if (kvm_is_error_hva(hva) || !writable) {
> 
> [Severity: Medium]
> Should read-only mappings be allowed to split?
> 
> This explicitly checks that the memslot is writable and fails with -EFAULT
> otherwise. It also passes FOLL_WRITE to pin_user_pages() later on.
> 
> However, splitting a huge page block into smaller PTEs changes the mapping
> granularity without modifying the underlying memory contents.
> 
> Does requiring write access prevent the hypervisor from splitting read-only
> stage-2 block mappings, such as guest ROMs?
> 
> > +		ret = -EFAULT;
> > +		goto unlock_srcu;
> > +	}
> > +
> > +	mmap_read_lock(mm);
> > +	ret = pin_user_pages(hva, nr_pages, FOLL_HWPOISON | FOLL_LONGTERM | FOLL_WRITE, pages);
> 
> [Severity: Critical]
> Does using pin_user_pages() here cause a permanent memory leak during KVM
> teardown?
> 
> Calling pin_user_pages() adds a GUP_PIN_COUNTING_BIAS to the page refcount.
> However, normal KVM stage-2 teardown drops references using standard
> put_page() semantics which only subtracts one from the refcount.
> 
> When the VM is destroyed, will these tail pages retain phantom references
> and be permanently leaked?
> 
> > +	mmap_read_unlock(mm);
> > +	if (ret != nr_pages) {
> > +		if (ret > 0)
> > +			unpin_user_pages(pages, ret);
> > +		ret = -EFAULT;
> > +		goto unlock_srcu;
> > +	}
> > +
> > +	write_lock(&kvm->mmu_lock);
> > +	mapping = pkvm_mapping_iter_first(&pgt->pkvm_mappings, ipa, ipa + size - 1);
> 
> [Severity: Critical]
> Is there a Time-of-Check to Time-of-Use race here since the MMU
> notifier retry protocol is missing?
> 
> The resolution of the HVA to PFNs via pin_user_pages() happens outside
> kvm->mmu_lock. A concurrent host MMU invalidation could unmap the region
> and remove the pkvm_mapping, and a subsequent guest fault could establish
> a new huge page mapping at the same IPA.

This is called on the back of a hypervisor request. This is only for
protected-VM and we do have an existing pin on the huge-page, preventing
any unmap. 

We do not use MMU notifier for protected VMs.

-- 
Vincent

> 
> When this thread acquires kvm->mmu_lock, it finds the new mapping in the
> tree and proceeds to split it using the PFNs derived from the old mapping.
> 
> Without capturing kvm->mmu_invalidate_seq before GUP and checking
> mmu_invalidate_retry() under the lock, can this lead to a refcount
> underflow on the new huge page and a permanent leak of the old pages?
> 
> > +	if (!pkvm_mapping_can_split(mapping)) {
> > +		ret = -EINVAL;
> > +		goto unlock_mmu;
> > +	}
> 
> -- 
> Sashiko AI review · https://sashiko.dev/#/patchset/20260803100904.3563942-1-vdonnefort@google.com?part=18

  reply	other threads:[~2026-09-07 10:23 UTC|newest]

Thread overview: 55+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-03 10:08 [PATCH 00/20] Huge mapping support for protected VMs Vincent Donnefort
2026-08-03 10:08 ` [PATCH 01/20] KVM: arm64: Prefault host stage-2 entries on block split Vincent Donnefort
2026-08-03 10:31   ` sashiko-bot
2026-09-07 18:34   ` Wei-Lin Chang
2026-08-03 10:08 ` [PATCH 02/20] KVM: arm64: Propagate host stage-2 annotated " Vincent Donnefort
2026-09-07 13:58   ` Wei-Lin Chang
2026-09-08 10:07     ` Vincent Donnefort
2026-08-03 10:08 ` [PATCH 03/20] KVM: arm64: Allow block-level stage-2 annotation Vincent Donnefort
2026-08-03 10:40   ` sashiko-bot
2026-08-03 10:08 ` [PATCH 04/20] KVM: arm64: Use block-level annotations when setting up the host stage-2 Vincent Donnefort
2026-09-07 14:37   ` Wei-Lin Chang
2026-09-08 10:18     ` Vincent Donnefort
2026-08-03 10:08 ` [PATCH 05/20] KVM: arm64: Make pKVM ownership selftest an HVC Vincent Donnefort
2026-08-03 10:08 ` [PATCH 06/20] KVM: arm64: Add a range to __pkvm_host_share/unshare_hyp() Vincent Donnefort
2026-08-03 10:08 ` [PATCH 07/20] KVM: arm64: Add a range to __pkvm_host_donate_guest() Vincent Donnefort
2026-08-03 10:26   ` sashiko-bot
2026-08-03 10:08 ` [PATCH 08/20] KVM: arm64: Add a range to __pkvm_host_reclaim_page_guest() Vincent Donnefort
2026-08-03 10:28   ` sashiko-bot
2026-09-07 14:52   ` Wei-Lin Chang
2026-09-08 10:22     ` Vincent Donnefort
2026-08-03 10:08 ` [PATCH 09/20] KVM: arm64: Add a range to __pkvm_guest_share_host() Vincent Donnefort
2026-08-03 10:08 ` [PATCH 10/20] KVM: arm64: Add a range to __pkvm_guest_unshare_host() Vincent Donnefort
2026-08-03 10:08 ` [PATCH 11/20] KVM: arm64: Add a range to pKVM ownership selftest Vincent Donnefort
2026-08-03 10:48   ` sashiko-bot
2026-09-07 10:13     ` Vincent Donnefort
2026-09-07 14:58   ` Wei-Lin Chang
2026-09-08 10:25     ` Vincent Donnefort
2026-09-10 10:04       ` Vincent Donnefort
2026-08-03 10:08 ` [PATCH 12/20] KVM: arm64: Handle huge mappings in __pkvm_host_force_reclaim_page_guest() Vincent Donnefort
2026-09-07 16:52   ` Wei-Lin Chang
2026-09-08 10:26     ` Vincent Donnefort
2026-08-03 10:08 ` [PATCH 13/20] KVM: arm64: Handle huge mappings in __pkvm_vcpu_in_poison_fault() Vincent Donnefort
2026-08-03 10:08 ` [PATCH 14/20] KVM: arm64: pkvm: Warn on guest stage-2 block collapse Vincent Donnefort
2026-08-03 10:08 ` [PATCH 15/20] KVM: arm64: Add pkvm_hyp_req infrastructure Vincent Donnefort
2026-08-03 10:46   ` sashiko-bot
2026-08-03 10:09 ` [PATCH 16/20] KVM: arm64: Add __pkvm_host_split_guest HVC Vincent Donnefort
2026-08-03 10:54   ` sashiko-bot
2026-09-07 10:14     ` Vincent Donnefort
2026-09-07 17:39   ` Wei-Lin Chang
2026-09-08 14:46     ` Vincent Donnefort
2026-08-03 10:09 ` [PATCH 17/20] KVM: arm64: Extend pKVM page ownership selftests to cover guest block split Vincent Donnefort
2026-08-03 11:01   ` sashiko-bot
2026-08-03 10:09 ` [PATCH 18/20] KVM: arm64: Add PKVM_HYP_REQ_SPLIT Vincent Donnefort
2026-08-03 10:54   ` sashiko-bot
2026-09-07 10:23     ` Vincent Donnefort [this message]
2026-09-08 10:35       ` Wei-Lin Chang
2026-09-07 18:15   ` Wei-Lin Chang
2026-09-08 14:56     ` Vincent Donnefort
2026-08-03 10:09 ` [PATCH 19/20] KVM: arm64: Raise PKVM_HYP_REQ_SPLIT on guest to host sharing Vincent Donnefort
2026-08-03 11:02   ` sashiko-bot
2026-09-07 10:26     ` Vincent Donnefort
2026-08-03 10:09 ` [PATCH 20/20] KVM: arm64: Stage-2 huge mappings for protected VMs Vincent Donnefort
2026-08-03 11:04   ` sashiko-bot
2026-09-07 18:31   ` Wei-Lin Chang
2026-09-07 18:41 ` [PATCH 00/20] Huge mapping support " Wei-Lin Chang

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=ap6QjjJxysWVYCVM@google.com \
    --to=vdonnefort@google.com \
    --cc=kvmarm@lists.linux.dev \
    --cc=maz@kernel.org \
    --cc=oupton@kernel.org \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.