From: Suzuki K Poulose <suzuki.poulose@arm.com>
To: Ackerley Tng <ackerleytng@google.com>,
"David Hildenbrand (Arm)" <david@kernel.org>,
"linux-coco@lists.linux.dev" <linux-coco@lists.linux.dev>,
"linux-mm@kvack.org" <linux-mm@kvack.org>,
KVM <kvm@vger.kernel.org>
Cc: amit@infradead.org, aneeshkumar.kizhakeveetil@arm.com,
ashish.kalra@amd.com, dwmw2@infradead.org, eberman@quicinc.com,
fvdl@google.com, gshan@redhat.com, jackmanb@google.com,
jackyli@google.com, jthoughton@google.com, kalyazin@amazon.com,
kas@kernel.org, kevinloughlin@google.com, liruxin@google.com,
michael.day@amd.com, michael.roth@amd.com,
mike.rapoport@gmail.com, mvaralar@redhat.com,
pankaj.gupta@amd.com, papaluri@amd.com, patrick.roy@linux.dev,
Peter Xu <peterx@redhat.com>,
pheragu@quicinc.com, pkondeti@qti.qualcomm.com, prty@google.com,
psalian@google.com, qinkun@google.com, seanjc@google.com,
shan.gavin@gmail.com, shivankg@amd.com, sidtelang@google.com,
tabba@google.com, tatashin@google.com, vannapurve@google.com,
vbabka@suse.com, wyihan@google.com
Subject: Re: [Invitation] bi-weekly guest_memfd upstream call on 2026-09-17
Date: Fri, 18 Sep 2026 10:28:20 +0100 [thread overview]
Message-ID: <e4bcc676-34d3-45ed-9eda-51c9a60e977f@arm.com> (raw)
In-Reply-To: <CAEvNRgGW_VHmT82q+bUqHh_5G2Z9sUF1jJOTiVN4Z8_cA9A50g@mail.gmail.com>
Hi Ackerley
Thanks for the writeup, my responses in-line.
On 17/09/2026 19:41, Ackerley Tng wrote:
> "David Hildenbrand (Arm)" <david@kernel.org> writes:
>
>> On 9/16/26 17:11, David Hildenbrand (Arm) wrote:
>>> Hi,
>>>
>>> Our next guest_memfd upstream call is scheduled for tomorrow, Thursday,
>>> 2026-09-16 8:00 - 9:00am (GMT-07:00) Pacific Time - Vancouver.
>>
>> Sorry, tomorrow (17) of course :(
>>
>> --
>> Cheers,
>>
>> David
>
> I'd like to try and restate the conversion problem we discussed today to
> understand better :)
>
> CCA's conversion protocol is:
>
> 1. Guest tells RMM to convert a GPA range
> 2. RMM notes down, in a vCPU object within the RMM, the conversion range
> 3. On re-entering the guest, RMM tells the guest if the conversion
> progress, something like:
> + success, GPA start to GPA end was converted or
> + failed, (with some error?)
> 4. Guest can
> + Be happy that whatever it requested is fulfilled
> + Retry to finish the parts that wasn't yet converted or
> + Be sad that it failed and figure it out.
>
Correct.
> I looked more into it and I'm surprised that what I was thinking of as
> "tell RMM to mark shared" isn't even correct.
>
> RMI_RTT_SET_RIPAS() takes these parameters: rd (the realm), rec_ptr (the
> vCPU), base (GPA start) and top (GPA after, or base + size).
>
> RMI_RTT_SET_RIPAS doesn't even take anything about shared or private!
This was deliberately removed because :
1. Host knows the original request from the VCPU exit.
2. RMM caches the request (range, ripas) in the VCPU object
3. Host confirms to the RMM, complete the RIPAS transition by
RMI_RTT_SET_RIPAS(), with the values it received from the VCPU exit.
4. RMM matches the range provided by the host to match with the VCPU
cached request and performs the "ripas" transition as per the
request from Realm.
>
> RMI_RTT_SET_RIPAS() is actually saying "host permits the conversion from
> base to top", it's not telling the RMM what to set it to, this is also
> different from SNP.
As mentioned above, host knows the requested "RIPAS" from VCPU exit.
RMM knows the "RIPAS" from the VCPU object.
Now: Privates vs Shared is translated to RIPAS_RAM vs RIPAS_EMPTY
in the CCA (well, roughly)
RIPAS_RAM implies, the GPA can be mapped into the private address
space of the Realm and it is integrity protected. Host cannot
replace the GPA with another content (it can unmap and DESTROY
the GPA mapping. But unless the Realm consents to replace the
GPA, again via SET_RIPAS request).
>
> Interestingly RMI_RTT_SET_RIPAS() also errors out if the current
> shared/private state is different from the one tracked in the vCPU
> object in the RMM? Did I get that right? I'm looking at the base_align
Correct.
> Failure condition, where it says ripas_pre != rec.ripas_value.
Please note that, in such cases, RMM needs a deeper page table level
and the error is RMI_ERROR_RTT, indicating to the host that:
Look I need a deeper level table to satisfy the request.
e.g., base = 4K, but walk.level = 2, and ripas_pre="private"
i.e. the requested base is mapped at L2 as block with "private" ripas.
If you want to convert the "base" to shared, it needs L3 table. The host
would follow up with RMI_RTT_CREATE and then retry the request.
> If two vCPUs race to convert the same range to shared, both vCPUs would
> have rec.ripas_value = private. The first conversion would be fine, but
nit: rec.ripas_value = empty (shared)
> the second one woul see ripas_pre = shared but rec.ripas_value = private
> and would definitely get an error?
rec.ripas_value == empty (shared) as per the guest request. And the RMM
will find the state is already "shared" and would confirm the success
back to host.
>
> And there's no "accept" step in the guest after conversions, which makes
> it different from TDX and SNP.
Correct, the guest "permitted" the GPA to be made private with unknown
contents anyway. RMM guarantees that the "data" is scrubbed when the
GPA is made valid. The advantage with this approach is, once the
Guest sets the RIPAS_RAM, the host can lazily donate pages at
fault time
>
> Difficulty in using current gmem hooks that SNP uses:
>
> * .gmem_make_shared() called from conversions doesn't have the vCPU
> context, finding the right vCPU context is expensive.
> * Sean, don't we already iterate vCPUs to find VMSA pages to kick
> the right vCPUs?
We could keep a list of VCPUs with pending set-ripas request for e.g.
Or before the vCPU enter, we check the state of the region and
do the sync with RMM.
> * .gmem_make_private at fault time is too late
> * At the next vCPU enter, the RMM would already read it's state to
> report success/failure, and there's no fault in-between for the
> .gmem_make_private to happen.
>
> At the call Sean suggested mirroring the RMM's tracking in KVM, but that
> sounds quite arch-specific and it's like doing arch-specific validation
> within KVM.
Suzuki
>
> p.s. Fuad, for pKVM you also mentioned that you'll need to check if the
> guest had requested for conversion first? This might be the same/similar
> problem.
next prev parent reply other threads:[~2026-09-18 9:28 UTC|newest]
Thread overview: 9+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-16 15:11 [Invitation] bi-weekly guest_memfd upstream call on 2026-09-16 David Hildenbrand (Arm)
2026-09-16 15:16 ` [Invitation] bi-weekly guest_memfd upstream call on 2026-09-17 David Hildenbrand (Arm)
2026-09-17 18:41 ` Ackerley Tng
2026-09-17 19:30 ` Sean Christopherson
2026-09-18 16:56 ` Ackerley Tng
2026-09-18 9:28 ` Suzuki K Poulose [this message]
2026-09-18 16:21 ` Ackerley Tng
2026-09-20 11:22 ` Suzuki K Poulose
2026-09-25 8:40 ` Fuad Tabba
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=e4bcc676-34d3-45ed-9eda-51c9a60e977f@arm.com \
--to=suzuki.poulose@arm.com \
--cc=ackerleytng@google.com \
--cc=amit@infradead.org \
--cc=aneeshkumar.kizhakeveetil@arm.com \
--cc=ashish.kalra@amd.com \
--cc=david@kernel.org \
--cc=dwmw2@infradead.org \
--cc=eberman@quicinc.com \
--cc=fvdl@google.com \
--cc=gshan@redhat.com \
--cc=jackmanb@google.com \
--cc=jackyli@google.com \
--cc=jthoughton@google.com \
--cc=kalyazin@amazon.com \
--cc=kas@kernel.org \
--cc=kevinloughlin@google.com \
--cc=kvm@vger.kernel.org \
--cc=linux-coco@lists.linux.dev \
--cc=linux-mm@kvack.org \
--cc=liruxin@google.com \
--cc=michael.day@amd.com \
--cc=michael.roth@amd.com \
--cc=mike.rapoport@gmail.com \
--cc=mvaralar@redhat.com \
--cc=pankaj.gupta@amd.com \
--cc=papaluri@amd.com \
--cc=patrick.roy@linux.dev \
--cc=peterx@redhat.com \
--cc=pheragu@quicinc.com \
--cc=pkondeti@qti.qualcomm.com \
--cc=prty@google.com \
--cc=psalian@google.com \
--cc=qinkun@google.com \
--cc=seanjc@google.com \
--cc=shan.gavin@gmail.com \
--cc=shivankg@amd.com \
--cc=sidtelang@google.com \
--cc=tabba@google.com \
--cc=tatashin@google.com \
--cc=vannapurve@google.com \
--cc=vbabka@suse.com \
--cc=wyihan@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox