Kernel KVM virtualization development
 help / color / mirror / Atom feed
From: Suzuki K Poulose <suzuki.poulose@arm.com>
To: Ackerley Tng <ackerleytng@google.com>,
	Steven Price <steven.price@arm.com>,
	kvm@vger.kernel.org, kvmarm@lists.linux.dev
Cc: Catalin Marinas <catalin.marinas@arm.com>,
	Marc Zyngier <maz@kernel.org>, Will Deacon <will@kernel.org>,
	James Morse <james.morse@arm.com>,
	Oliver Upton <oliver.upton@linux.dev>,
	Zenghui Yu <yuzenghui@huawei.com>,
	linux-arm-kernel@lists.infradead.org,
	linux-kernel@vger.kernel.org, Joey Gouly <joey.gouly@arm.com>,
	Alexandru Elisei <alexandru.elisei@arm.com>,
	Christoffer Dall <christoffer.dall@arm.com>,
	Fuad Tabba <tabba@google.com>,
	linux-coco@lists.linux.dev,
	Ganapatrao Kulkarni <gankulkarni@os.amperecomputing.com>,
	Gavin Shan <gshan@redhat.com>,
	Shanker Donthineni <sdonthineni@nvidia.com>,
	Alper Gun <alpergun@google.com>,
	"Aneesh Kumar K . V" <aneesh.kumar@kernel.org>,
	Emi Kisanuki <fj0570is@fujitsu.com>,
	Vishal Annapurve <vannapurve@google.com>,
	WeiLin.Chang@arm.com, Lorenzo Pieralisi <lpieralisi@kernel.org>,
	rick.p.edgecombe@intel.com, yan.y.zhao@intel.com
Subject: Re: [PATCH v16 27/45] KVM: arm64: CCA: Allow populating initial contents
Date: Fri, 7 Aug 2026 11:58:53 +0100	[thread overview]
Message-ID: <7ff49972-a9bf-4d99-9088-3d5e1eaa1608@arm.com> (raw)
In-Reply-To: <CAEvNRgHF-30EwJEprM466+Ag7GSuJWrzMGNAw7kCvXyT7YVucg@mail.gmail.com>

On 06/08/2026 23:43, Ackerley Tng wrote:
> Steven Price <steven.price@arm.com> writes:
> 
>> The VMM needs to populate the realm with some data before starting (e.g.
>> a kernel and initrd). This is measured by the RMM and used as part of
>> the attestation later on.
>>
>> Signed-off-by: Steven Price <steven.price@arm.com>

...

>> +static int realm_data_map_init(struct kvm *kvm, unsigned long ipa,
>> +			       kvm_pfn_t dst_pfn, kvm_pfn_t src_pfn,
>> +			       unsigned long flags)
>> +{
>> +	struct realm *realm = &kvm->arch.realm;
>> +	phys_addr_t rd = virt_to_phys(realm->rd);
>> +	phys_addr_t dst_phys, src_phys;
>> +	long ret;
>> +
>> +	lockdep_assert_held(&kvm->slots_lock);
>> +	lockdep_assert_held(&kvm->arch.config_lock);
>> +
>> +	dst_phys = __pfn_to_phys(dst_pfn);
>> +	src_phys = __pfn_to_phys(src_pfn);
>> +
>> +	if (rmi_delegate_page(dst_phys))
>> +		return -ENXIO;
>> +
>> +retry:
>> +	ret = rmi_rtt_data_map_init(rd, dst_phys, ipa, src_phys, flags);
>> +	if (ret >= 0 && RMI_RETURN_STATUS(ret) == RMI_ERROR_RTT) {
>> +		/* Create missing RTTs and retry */
>> +		int level = RMI_RETURN_INDEX(ret);
>> +
>> +		KVM_BUG_ON(level >= KVM_PGTABLE_LAST_LEVEL, kvm);
>> +
>> +		ret = realm_create_rtt_levels(realm, ipa, level,
>> +					      level + 1, NULL);
>> +		if (!ret)
>> +			goto retry;
>> +	}
>> +
>> +	if (ret && WARN_ON(rmi_undelegate_page(dst_phys))) {
>> +		/* Leak the page if the undelegate fails */
>> +		get_page(pfn_to_page(dst_pfn));
> 
> Is there some way to avoid taking a reference on the page? This would
> interfere with conversions. There was a similar discussion for TDX as
> well [1].

Unfortunately, no. The page was transferred to the Realm world (with
rmi_delegate_page() above the retry: ). If we fail to bring it back,
that page is still in the Realm PAS and any access to it by the normal
world would result in a GPF and eventually bring down the system
if it happens from the kernel.

We don't expect that undelegate to fail. The granule_delegate()
should fail if the page was already in use by the RMM for some
purpose (e.g., already mapped at the IPA, because VMM issued
DATA_MAP_INIT twice. Even with the relaxation coming in the
RMM, we will mandate that the "populate" cases will request
strict conditions for granule delegate).

Please note that this is NOT the "unmap" failure, but it is
"Bring the page back to the NS world" failure that causes
the WARN_ON and the leaking.

> 
> TDX originally incremented folio refounts for these:
> 
> + when mapping folios into the Secure EPTs. This one was easier to agree
>    to remove, since TDX can trust guest_memfd to keep pages around on
>    behalf of the guest.
> + To indicate unmapping failure (IIUC this is the same situation as
>    above). This interferes with conversions.
>      + An alternative discussed was to mark these pages as HWPOISON, but
>        that was eventually rejected as adding unnecessary complexity to
>        make TDX special for code paths that only occur on kernel
>        bugs. (In TDX's case the unmap failures would probably only be for
>        kernel bugs.)
>      + I later worked a bit more on memory failure for guest_memfd
>        HugeTLB and found that because we will need to restructure huge
>        pages for conversions, using the HWPOISON flag would be hard to
>        handle. For TDX since the conclusion was not to use a HWPOISON
>        flag to indicate unmap failures anyway, this turned out to be a
>        non-issue. Nobody wanted to use the HWPOISON flag. (I hope you
>        won't need to either)
> 
> So for TDX, on an unmap failure we do a KVM_BUG_ON() and mark the VM as
> dead, and do nothing about the page, it still gets returned to the
> system as if nothing happened.
> 
> Here's my understanding of why this is okay for TDX (Rick and Yan, could
> you please help me here):
> 
> + For unmap failures, the page remains in TDX's Physical Address
>    Metadata Table (PAMT), and the page is still assigned to some TD.
> + If the page was assigned to some other TD, it would be blocked, since
>    the PAMT shows it as already assigned.
> + If the page was used by something completely unrelated to TDX, then in
>    the TDX model the host is free to write and read pages. Nothing goes
>    bad until the TD tries to use that same page, but that TD would never
>    use the page again, that TD is already dead and the HKID for the TD
>    was leaked.

This is not true for CCA. Like I said above, touching the page in Realm
PAS is going to be disastrous for the Host.


> 
> [1] https://lore.kernel.org/all/diqz34bolnta.fsf@ackerleytng-ctop.c.googlers.com/
> 
>> +	}
>> +
>> +	return ret <= 0 ? ret : -ENXIO;
>> +}
>> +
>> +static int populate_region_cb(struct kvm *kvm, gfn_t gfn, kvm_pfn_t pfn,
>> +			      struct page *src_page, void *opaque)
>> +{
>> +	unsigned long data_flags = *(unsigned long *)opaque;
>> +	phys_addr_t ipa = gfn_to_gpa(gfn);
>> +
>> +	return realm_data_map_init(kvm, ipa, pfn, page_to_pfn(src_page),
>> +				   data_flags);
>> +}
>> +
>> +static long populate_region(struct kvm *kvm,
>> +			    gfn_t base_gfn,
>> +			    unsigned long pages,
>> +			    u64 uaddr,
>> +			    unsigned long data_flags)
>> +{
>> +	long ret = 0;
>> +
>> +	lockdep_assert_held(&kvm->slots_lock);
>> +	lockdep_assert_held(&kvm->arch.config_lock);
>> +
>> +	if (!uaddr)
>> +		return -EINVAL;
>> +
> 
> Why not check for !uaddr together with the other checks in
> kvm_arm_rmi_populate?

Yep, we could move it there.

> 
> Also would it be okay to inline populate_region into
> kvm_arm_rmi_populate below?

Ack.

Suzuki

  reply	other threads:[~2026-08-07 10:58 UTC|newest]

Thread overview: 65+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-03 13:43 [PATCH v16 00/45] arm64: Support for Arm CCA in KVM Steven Price
2026-08-03 13:43 ` [PATCH v16 01/45] firmware: arm_rmm: Add SMC definitions for calling the RMM Steven Price
2026-08-03 13:43 ` [PATCH v16 02/45] firmware: arm_rmm: Add wrappers for direct RMI calls Steven Price
2026-08-03 13:43 ` [PATCH v16 03/45] firmware: arm_rmm: Check for RMI support at init Steven Price
2026-08-03 13:43 ` [PATCH v16 04/45] firmware: arm_rmm: Configure the RMM with the host's page size Steven Price
2026-08-03 13:43 ` [PATCH v16 05/45] firmware: arm_rmm: Add support for SRO Steven Price
2026-08-03 13:43 ` [PATCH v16 06/45] firmware: arm_rmm: Ensure the RMM has GPT entries for memory Steven Price
2026-08-03 13:43 ` [PATCH v16 07/45] arm64: mm: Handle Granule Protection Faults (GPFs) Steven Price
2026-08-03 13:43 ` [PATCH v16 08/45] KVM: arm64: Include kvm_emulate.h in kvm/arm_psci.h Steven Price
2026-08-03 13:43 ` [PATCH v16 09/45] KVM: arm64: Avoid including linux/kvm_host.h in kvm_pgtable.h Steven Price
2026-08-03 13:43 ` [PATCH v16 10/45] KVM: arm64: CCA: Add wrappers for realm related RMIs Steven Price
2026-08-03 13:43 ` [PATCH v16 11/45] KVM: arm64: CCA: Check for RMI support at KVM init Steven Price
2026-08-04 14:55   ` Fuad Tabba
2026-08-04 14:59     ` Suzuki K Poulose
2026-08-03 13:43 ` [PATCH v16 12/45] KVM: arm64: CCA: Check for LPA2 support Steven Price
2026-08-03 13:43 ` [PATCH v16 13/45] KVM: arm64: CCA: Define the user ABI Steven Price
2026-08-03 13:43 ` [PATCH v16 14/45] KVM: arm64: CCA: Add basic infrastructure for creating a realm Steven Price
2026-08-03 13:43 ` [PATCH v16 15/45] KVM: arm64: CCA: Don't expose unsupported capabilities for realm guests Steven Price
2026-08-03 13:43 ` [PATCH v16 16/45] KVM: arm64: CCA: Allow passing the machine type in KVM creation Steven Price
2026-08-03 13:43 ` [PATCH v16 17/45] KVM: arm64: CCA: Tear down RTTs Steven Price
2026-08-03 22:29   ` Alper Gun
2026-08-04 12:16     ` Suzuki K Poulose
2026-08-03 13:43 ` [PATCH v16 18/45] KVM: arm64: CCA: Allocate and free RECs to match vCPUs Steven Price
2026-08-03 13:43 ` [PATCH v16 19/45] KVM: arm64: CCA: Support the VGIC in realms Steven Price
2026-08-03 13:43 ` [PATCH v16 20/45] KVM: arm64: CCA: Support timers in realm RECs Steven Price
2026-08-03 13:43 ` [PATCH v16 21/45] KVM: arm64: CCA: Handle realm enter/exit Steven Price
2026-08-04  8:57   ` Aneesh Kumar K.V
2026-08-04 13:36   ` Aneesh Kumar K.V
2026-08-03 13:43 ` [PATCH v16 22/45] KVM: arm64: CCA: Handle RMI_EXIT_RIPAS_CHANGE Steven Price
2026-08-05 15:59   ` Ackerley Tng
2026-08-06  8:42     ` Suzuki K Poulose
2026-08-03 13:43 ` [PATCH v16 23/45] KVM: arm64: CCA: Handle realm MMIO emulation Steven Price
2026-08-03 13:43 ` [PATCH v16 24/45] KVM: arm64: Expose support for private memory Steven Price
2026-08-05 16:02   ` Ackerley Tng
2026-08-07 10:12     ` Suzuki K Poulose
2026-08-03 13:43 ` [PATCH v16 25/45] KVM: arm64: CCA: Create the realm descriptor Steven Price
2026-08-03 13:43 ` [PATCH v16 26/45] KVM: arm64: CCA: Activate realms on first vCPU run Steven Price
2026-08-03 13:43 ` [PATCH v16 27/45] KVM: arm64: CCA: Allow populating initial contents Steven Price
2026-08-06 22:43   ` Ackerley Tng
2026-08-07 10:58     ` Suzuki K Poulose [this message]
2026-08-03 13:43 ` [PATCH v16 28/45] KVM: arm64: CCA: Set RIPAS of initial memslots Steven Price
2026-08-03 13:43 ` [PATCH v16 29/45] KVM: arm64: CCA: Support runtime faulting of memory Steven Price
2026-08-06 23:11   ` Ackerley Tng
2026-08-03 13:43 ` [PATCH v16 30/45] KVM: arm64: CCA: Handle realm vCPU load Steven Price
2026-08-03 13:43 ` [PATCH v16 31/45] KVM: arm64: CCA: Validate register access for Realm VMs Steven Price
2026-08-03 13:43 ` [PATCH v16 32/45] KVM: arm64: CCA: Handle Realm PSCI requests Steven Price
2026-08-03 13:43 ` [PATCH v16 33/45] KVM: arm64: WARN on injected undef exceptions Steven Price
2026-08-03 13:43 ` [PATCH v16 34/45] KVM: arm64: CCA: Allow userspace to inject aborts Steven Price
2026-08-03 13:43 ` [PATCH v16 35/45] KVM: arm64: CCA: Support RSI_HOST_CALL Steven Price
2026-08-03 13:43 ` [PATCH v16 36/45] KVM: arm64: CCA: Allow checking SVE on VM instance Steven Price
2026-08-03 13:43 ` [PATCH v16 37/45] KVM: arm64: CCA: Prevent Device mappings for realms Steven Price
2026-08-03 13:43 ` [PATCH v16 38/45] KVM: arm64: CCA: Propagate breakpoint and watchpoint counts to userspace Steven Price
2026-08-03 13:43 ` [PATCH v16 39/45] KVM: arm64: CCA: Set breakpoint parameters through SET_ONE_REG Steven Price
2026-08-03 13:43 ` [PATCH v16 40/45] KVM: arm64: CCA: Propagate max SVE vector length from the RMM Steven Price
2026-08-03 13:43 ` [PATCH v16 41/45] KVM: arm64: CCA: Configure max SVE vector length for a Realm Steven Price
2026-08-03 13:43 ` [PATCH v16 42/45] KVM: arm64: CCA: Provide register list for unfinalized RECs Steven Price
2026-08-03 13:43 ` [PATCH v16 43/45] KVM: arm64: CCA: Provide an accurate register list Steven Price
2026-08-03 13:44 ` [PATCH v16 44/45] KVM: arm64: CCA: Require ICH_HCR_EL2.TDIR for realms Steven Price
2026-08-03 13:44 ` [PATCH v16 45/45] KVM: arm64: CCA: Enable realms to be created Steven Price
2026-08-03 15:01 ` [PATCH v16 00/45] arm64: Support for Arm CCA in KVM Marc Zyngier
2026-08-03 15:06   ` Steven Price
2026-08-03 15:20     ` Marc Zyngier
2026-08-03 15:45       ` Steven Price
2026-08-04 14:24 ` Fuad Tabba
2026-08-06 22:05 ` Suzuki K Poulose

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=7ff49972-a9bf-4d99-9088-3d5e1eaa1608@arm.com \
    --to=suzuki.poulose@arm.com \
    --cc=WeiLin.Chang@arm.com \
    --cc=ackerleytng@google.com \
    --cc=alexandru.elisei@arm.com \
    --cc=alpergun@google.com \
    --cc=aneesh.kumar@kernel.org \
    --cc=catalin.marinas@arm.com \
    --cc=christoffer.dall@arm.com \
    --cc=fj0570is@fujitsu.com \
    --cc=gankulkarni@os.amperecomputing.com \
    --cc=gshan@redhat.com \
    --cc=james.morse@arm.com \
    --cc=joey.gouly@arm.com \
    --cc=kvm@vger.kernel.org \
    --cc=kvmarm@lists.linux.dev \
    --cc=linux-arm-kernel@lists.infradead.org \
    --cc=linux-coco@lists.linux.dev \
    --cc=linux-kernel@vger.kernel.org \
    --cc=lpieralisi@kernel.org \
    --cc=maz@kernel.org \
    --cc=oliver.upton@linux.dev \
    --cc=rick.p.edgecombe@intel.com \
    --cc=sdonthineni@nvidia.com \
    --cc=steven.price@arm.com \
    --cc=tabba@google.com \
    --cc=vannapurve@google.com \
    --cc=will@kernel.org \
    --cc=yan.y.zhao@intel.com \
    --cc=yuzenghui@huawei.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox