Kernel KVM virtualization development
 help / color / mirror / Atom feed
* [Invitation] bi-weekly guest_memfd upstream call on 2026-09-16
@ 2026-09-16 15:11 David Hildenbrand (Arm)
  2026-09-16 15:16 ` [Invitation] bi-weekly guest_memfd upstream call on 2026-09-17 David Hildenbrand (Arm)
  0 siblings, 1 reply; 9+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-16 15:11 UTC (permalink / raw)
  To: linux-coco@lists.linux.dev, linux-mm@kvack.org, KVM
  Cc: ackerleytng, amit, aneeshkumar.kizhakeveetil, ashish.kalra, dwmw2,
	eberman, fvdl, gshan, jackmanb, jackyli, jthoughton, kalyazin,
	kas, kevinloughlin, liruxin, michael.day, michael.roth,
	mike.rapoport, mvaralar, pankaj.gupta, papaluri, patrick.roy,
	Peter Xu, pheragu, pkondeti, prty, psalian, qinkun, seanjc,
	shan.gavin, shivankg, sidtelang, suzuki.poulose, tabba, tatashin,
	vannapurve, vbabka, wyihan

Hi,

Our next guest_memfd upstream call is scheduled for tomorrow, Thursday,
2026-09-16 8:00 - 9:00am (GMT-07:00) Pacific Time - Vancouver.

So far we don't have any topics, so I assume we'll just briefly sync on current
upstream work and discuss whatever comes up.

We'll be using the following Google meet:
http://meet.google.com/wxp-wtju-jzw

The meeting notes can be found at [1], where we also link recordings and
collect current guest_memfd upstream proposals. If you want an google
calendar invitation that also covers all future meetings, just write me
or Ackerley a mail.

To put something to discuss onto the agenda, reply to this mail or add
them to the "Topics/questions for next meeting(s)" section in the
meeting notes as a comment.

[1]
https://docs.google.com/document/d/1M6766BzdY1Lhk7LiR5IqVR8B8mG3cr-cxTxOrAosPOk/edit?usp=sharing

-- 
Cheers,

David


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [Invitation] bi-weekly guest_memfd upstream call on 2026-09-17
  2026-09-16 15:11 [Invitation] bi-weekly guest_memfd upstream call on 2026-09-16 David Hildenbrand (Arm)
@ 2026-09-16 15:16 ` David Hildenbrand (Arm)
  2026-09-17 18:41   ` Ackerley Tng
  0 siblings, 1 reply; 9+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-16 15:16 UTC (permalink / raw)
  To: linux-coco@lists.linux.dev, linux-mm@kvack.org, KVM
  Cc: ackerleytng, amit, aneeshkumar.kizhakeveetil, ashish.kalra, dwmw2,
	eberman, fvdl, gshan, jackmanb, jackyli, jthoughton, kalyazin,
	kas, kevinloughlin, liruxin, michael.day, michael.roth,
	mike.rapoport, mvaralar, pankaj.gupta, papaluri, patrick.roy,
	Peter Xu, pheragu, pkondeti, prty, psalian, qinkun, seanjc,
	shan.gavin, shivankg, sidtelang, suzuki.poulose, tabba, tatashin,
	vannapurve, vbabka, wyihan

On 9/16/26 17:11, David Hildenbrand (Arm) wrote:
> Hi,
> 
> Our next guest_memfd upstream call is scheduled for tomorrow, Thursday,
> 2026-09-16 8:00 - 9:00am (GMT-07:00) Pacific Time - Vancouver.

Sorry, tomorrow (17) of course :(

-- 
Cheers,

David

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [Invitation] bi-weekly guest_memfd upstream call on 2026-09-17
  2026-09-16 15:16 ` [Invitation] bi-weekly guest_memfd upstream call on 2026-09-17 David Hildenbrand (Arm)
@ 2026-09-17 18:41   ` Ackerley Tng
  2026-09-17 19:30     ` Sean Christopherson
                       ` (2 more replies)
  0 siblings, 3 replies; 9+ messages in thread
From: Ackerley Tng @ 2026-09-17 18:41 UTC (permalink / raw)
  To: David Hildenbrand (Arm), linux-coco@lists.linux.dev,
	linux-mm@kvack.org, KVM
  Cc: amit, aneeshkumar.kizhakeveetil, ashish.kalra, dwmw2, eberman,
	fvdl, gshan, jackmanb, jackyli, jthoughton, kalyazin, kas,
	kevinloughlin, liruxin, michael.day, michael.roth, mike.rapoport,
	mvaralar, pankaj.gupta, papaluri, patrick.roy, Peter Xu, pheragu,
	pkondeti, prty, psalian, qinkun, seanjc, shan.gavin, shivankg,
	sidtelang, suzuki.poulose, tabba, tatashin, vannapurve, vbabka,
	wyihan

"David Hildenbrand (Arm)" <david@kernel.org> writes:

> On 9/16/26 17:11, David Hildenbrand (Arm) wrote:
>> Hi,
>>
>> Our next guest_memfd upstream call is scheduled for tomorrow, Thursday,
>> 2026-09-16 8:00 - 9:00am (GMT-07:00) Pacific Time - Vancouver.
>
> Sorry, tomorrow (17) of course :(
>
> --
> Cheers,
>
> David

I'd like to try and restate the conversion problem we discussed today to
understand better :)

CCA's conversion protocol is:

1. Guest tells RMM to convert a GPA range
2. RMM notes down, in a vCPU object within the RMM, the conversion range
3. On re-entering the guest, RMM tells the guest if the conversion
   progress, something like:
     + success, GPA start to GPA end was converted or
     + failed, (with some error?)
4. Guest can
     + Be happy that whatever it requested is fulfilled
     + Retry to finish the parts that wasn't yet converted or
     + Be sad that it failed and figure it out.

I looked more into it and I'm surprised that what I was thinking of as
"tell RMM to mark shared" isn't even correct.

RMI_RTT_SET_RIPAS() takes these parameters: rd (the realm), rec_ptr (the
vCPU), base (GPA start) and top (GPA after, or base + size).

RMI_RTT_SET_RIPAS doesn't even take anything about shared or private!

RMI_RTT_SET_RIPAS() is actually saying "host permits the conversion from
base to top", it's not telling the RMM what to set it to, this is also
different from SNP.

Interestingly RMI_RTT_SET_RIPAS() also errors out if the current
shared/private state is different from the one tracked in the vCPU
object in the RMM? Did I get that right? I'm looking at the base_align
Failure condition, where it says ripas_pre != rec.ripas_value.

If two vCPUs race to convert the same range to shared, both vCPUs would
have rec.ripas_value = private. The first conversion would be fine, but
the second one woul see ripas_pre = shared but rec.ripas_value = private
and would definitely get an error?

And there's no "accept" step in the guest after conversions, which makes
it different from TDX and SNP.

Difficulty in using current gmem hooks that SNP uses:

* .gmem_make_shared() called from conversions doesn't have the vCPU
  context, finding the right vCPU context is expensive.
    * Sean, don't we already iterate vCPUs to find VMSA pages to kick
      the right vCPUs?
* .gmem_make_private at fault time is too late
    * At the next vCPU enter, the RMM would already read it's state to
      report success/failure, and there's no fault in-between for the
      .gmem_make_private to happen.

At the call Sean suggested mirroring the RMM's tracking in KVM, but that
sounds quite arch-specific and it's like doing arch-specific validation
within KVM.

p.s. Fuad, for pKVM you also mentioned that you'll need to check if the
guest had requested for conversion first? This might be the same/similar
problem.

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [Invitation] bi-weekly guest_memfd upstream call on 2026-09-17
  2026-09-17 18:41   ` Ackerley Tng
@ 2026-09-17 19:30     ` Sean Christopherson
  2026-09-18 16:56       ` Ackerley Tng
  2026-09-18  9:28     ` Suzuki K Poulose
  2026-09-25  8:40     ` Fuad Tabba
  2 siblings, 1 reply; 9+ messages in thread
From: Sean Christopherson @ 2026-09-17 19:30 UTC (permalink / raw)
  To: Ackerley Tng
  Cc: David Hildenbrand (Arm), linux-coco@lists.linux.dev,
	linux-mm@kvack.org, KVM, amit, aneeshkumar.kizhakeveetil,
	ashish.kalra, dwmw2, eberman, fvdl, gshan, jackmanb, jackyli,
	jthoughton, kalyazin, kas, kevinloughlin, liruxin, michael.day,
	michael.roth, mike.rapoport, mvaralar, pankaj.gupta, papaluri,
	patrick.roy, Peter Xu, pheragu, pkondeti, prty, psalian, qinkun,
	shan.gavin, shivankg, sidtelang, suzuki.poulose, tabba, tatashin,
	vannapurve, vbabka, wyihan

On Thu, Sep 17, 2026, Ackerley Tng wrote:
> At the call Sean suggested mirroring the RMM's tracking in KVM, but that
> sounds quite arch-specific and it's like doing arch-specific validation
> within KVM.

I suggested "mirroring" the tracking in KVM arm64, not in guest_memfd.

E.g. put a structure in arm64's kvm_vcpu_arch with a list_head object, then insert
into a per-VM list store in kvm_arch on a conversion request.  But before inserting,
walk the list to see if there are conflicting requests, and if so, reject the new
request.

Then on KVM_RUN, if a vCPU has an outstanding request, verify the gmem page exists,
is in the correct state, and is mapped into the guest (or at least, is known to the
RMM?).  If any steps fail, exit to userspace with -EFAULT + KVM_EXIT_MEMORY_FAULT.

The downside to such a simplistic implementation is that it'll require a per-VM
lock to serialize the requests.  If the contention ends up being too painful, e.g.
because there are usage patterns where guests due batched conversions across many
vCPUs, then KVM could use a more sophisticated approach as needed.  E.g. shard the
tracking+locking at 1GiB or something?

> p.s. Fuad, for pKVM you also mentioned that you'll need to check if the
> guest had requested for conversion first? This might be the same/similar
> problem.

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [Invitation] bi-weekly guest_memfd upstream call on 2026-09-17
  2026-09-17 18:41   ` Ackerley Tng
  2026-09-17 19:30     ` Sean Christopherson
@ 2026-09-18  9:28     ` Suzuki K Poulose
  2026-09-18 16:21       ` Ackerley Tng
  2026-09-25  8:40     ` Fuad Tabba
  2 siblings, 1 reply; 9+ messages in thread
From: Suzuki K Poulose @ 2026-09-18  9:28 UTC (permalink / raw)
  To: Ackerley Tng, David Hildenbrand (Arm), linux-coco@lists.linux.dev,
	linux-mm@kvack.org, KVM
  Cc: amit, aneeshkumar.kizhakeveetil, ashish.kalra, dwmw2, eberman,
	fvdl, gshan, jackmanb, jackyli, jthoughton, kalyazin, kas,
	kevinloughlin, liruxin, michael.day, michael.roth, mike.rapoport,
	mvaralar, pankaj.gupta, papaluri, patrick.roy, Peter Xu, pheragu,
	pkondeti, prty, psalian, qinkun, seanjc, shan.gavin, shivankg,
	sidtelang, tabba, tatashin, vannapurve, vbabka, wyihan

Hi Ackerley

Thanks for the writeup, my responses in-line.

On 17/09/2026 19:41, Ackerley Tng wrote:
> "David Hildenbrand (Arm)" <david@kernel.org> writes:
> 
>> On 9/16/26 17:11, David Hildenbrand (Arm) wrote:
>>> Hi,
>>>
>>> Our next guest_memfd upstream call is scheduled for tomorrow, Thursday,
>>> 2026-09-16 8:00 - 9:00am (GMT-07:00) Pacific Time - Vancouver.
>>
>> Sorry, tomorrow (17) of course :(
>>
>> --
>> Cheers,
>>
>> David
> 
> I'd like to try and restate the conversion problem we discussed today to
> understand better :)
> 
> CCA's conversion protocol is:
> 
> 1. Guest tells RMM to convert a GPA range
> 2. RMM notes down, in a vCPU object within the RMM, the conversion range
> 3. On re-entering the guest, RMM tells the guest if the conversion
>     progress, something like:
>       + success, GPA start to GPA end was converted or
>       + failed, (with some error?)
> 4. Guest can
>       + Be happy that whatever it requested is fulfilled
>       + Retry to finish the parts that wasn't yet converted or
>       + Be sad that it failed and figure it out.
> 

Correct.

> I looked more into it and I'm surprised that what I was thinking of as
> "tell RMM to mark shared" isn't even correct.
> 
> RMI_RTT_SET_RIPAS() takes these parameters: rd (the realm), rec_ptr (the
> vCPU), base (GPA start) and top (GPA after, or base + size).
> 
> RMI_RTT_SET_RIPAS doesn't even take anything about shared or private!

This was deliberately removed because :

1. Host knows the original request from the VCPU exit.
2. RMM caches the request (range, ripas) in the VCPU object
3. Host confirms to the RMM, complete the RIPAS transition by
    RMI_RTT_SET_RIPAS(), with the values it received from the VCPU exit.
4. RMM matches the range provided by the host to match with the VCPU
    cached request and performs the "ripas" transition as per the
    request from Realm.


> 
> RMI_RTT_SET_RIPAS() is actually saying "host permits the conversion from
> base to top", it's not telling the RMM what to set it to, this is also
> different from SNP.

As mentioned above, host knows the requested "RIPAS" from VCPU exit.
RMM knows the "RIPAS" from the VCPU object.

Now: Privates vs Shared is translated to RIPAS_RAM vs RIPAS_EMPTY
in the CCA (well, roughly)

RIPAS_RAM implies, the GPA can be mapped into the private address
space of the Realm and it is integrity protected. Host cannot
replace the GPA with another content (it can unmap and DESTROY
the GPA mapping. But unless the Realm consents to replace the
GPA, again via SET_RIPAS request).


> 
> Interestingly RMI_RTT_SET_RIPAS() also errors out if the current
> shared/private state is different from the one tracked in the vCPU
> object in the RMM? Did I get that right? I'm looking at the base_align

Correct.

> Failure condition, where it says ripas_pre != rec.ripas_value.

Please note that, in such cases, RMM needs a deeper page table level
and the error is RMI_ERROR_RTT, indicating to the host that:
Look I need a deeper level table to satisfy the request.

e.g., base = 4K, but walk.level = 2, and ripas_pre="private"

i.e. the requested base is mapped at L2 as block with "private" ripas.
If you want to convert the "base" to shared, it needs L3 table. The host
would follow up with RMI_RTT_CREATE and then retry the request.


> If two vCPUs race to convert the same range to shared, both vCPUs would
> have rec.ripas_value = private. The first conversion would be fine, but

nit: rec.ripas_value = empty (shared)


> the second one woul see ripas_pre = shared but rec.ripas_value = private
> and would definitely get an error?

rec.ripas_value == empty (shared) as per the guest request. And the RMM
will find the state is already "shared" and would confirm the success
back to host.


> 
> And there's no "accept" step in the guest after conversions, which makes
> it different from TDX and SNP.

Correct, the guest "permitted" the GPA to be made private with unknown
contents anyway. RMM guarantees that the "data" is scrubbed when the
GPA is made valid. The advantage with this approach is, once the
Guest sets the RIPAS_RAM, the host can lazily donate pages at
fault time

> 
> Difficulty in using current gmem hooks that SNP uses:
> 
> * .gmem_make_shared() called from conversions doesn't have the vCPU
>    context, finding the right vCPU context is expensive.
>      * Sean, don't we already iterate vCPUs to find VMSA pages to kick
>        the right vCPUs?

We could keep a list of VCPUs with pending set-ripas request for e.g.
Or before the vCPU enter, we check the state of the region and
do the sync with RMM.

> * .gmem_make_private at fault time is too late
>      * At the next vCPU enter, the RMM would already read it's state to
>        report success/failure, and there's no fault in-between for the
>        .gmem_make_private to happen.
> 
> At the call Sean suggested mirroring the RMM's tracking in KVM, but that
> sounds quite arch-specific and it's like doing arch-specific validation
> within KVM.

Suzuki


> 
> p.s. Fuad, for pKVM you also mentioned that you'll need to check if the
> guest had requested for conversion first? This might be the same/similar
> problem.


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [Invitation] bi-weekly guest_memfd upstream call on 2026-09-17
  2026-09-18  9:28     ` Suzuki K Poulose
@ 2026-09-18 16:21       ` Ackerley Tng
  2026-09-20 11:22         ` Suzuki K Poulose
  0 siblings, 1 reply; 9+ messages in thread
From: Ackerley Tng @ 2026-09-18 16:21 UTC (permalink / raw)
  To: Suzuki K Poulose, David Hildenbrand (Arm),
	linux-coco@lists.linux.dev, linux-mm@kvack.org, KVM
  Cc: amit, aneeshkumar.kizhakeveetil, ashish.kalra, dwmw2, eberman,
	fvdl, gshan, jackmanb, jackyli, jthoughton, kalyazin, kas,
	kevinloughlin, liruxin, michael.day, michael.roth, mike.rapoport,
	mvaralar, pankaj.gupta, papaluri, patrick.roy, Peter Xu, pheragu,
	pkondeti, prty, psalian, qinkun, seanjc, shan.gavin, shivankg,
	sidtelang, tabba, tatashin, vannapurve, vbabka, wyihan

Suzuki K Poulose <suzuki.poulose@arm.com> writes:

>
> [...snip...]
>
>>
>> RMI_RTT_SET_RIPAS() is actually saying "host permits the conversion from
>> base to top", it's not telling the RMM what to set it to, this is also
>> different from SNP.
>
> As mentioned above, host knows the requested "RIPAS" from VCPU exit.
> RMM knows the "RIPAS" from the VCPU object.
>
> Now: Privates vs Shared is translated to RIPAS_RAM vs RIPAS_EMPTY
> in the CCA (well, roughly)
>
> RIPAS_RAM implies, the GPA can be mapped into the private address
> space of the Realm and it is integrity protected. Host cannot
> replace the GPA with another content (it can unmap and DESTROY
> the GPA mapping. But unless the Realm consents to replace the
> GPA, again via SET_RIPAS request).
>
>
>>
>> Interestingly RMI_RTT_SET_RIPAS() also errors out if the current
>> shared/private state is different from the one tracked in the vCPU
>> object in the RMM? Did I get that right? I'm looking at the base_align
>
> Correct.
>
>> Failure condition, where it says ripas_pre != rec.ripas_value.
>
> Please note that, in such cases, RMM needs a deeper page table level
> and the error is RMI_ERROR_RTT, indicating to the host that:
> Look I need a deeper level table to satisfy the request.
>
> e.g., base = 4K, but walk.level = 2, and ripas_pre="private"
>
> i.e. the requested base is mapped at L2 as block with "private" ripas.
> If you want to convert the "base" to shared, it needs L3 table. The host
> would follow up with RMI_RTT_CREATE and then retry the request.
>
>
>> If two vCPUs race to convert the same range to shared, both vCPUs would
>> have rec.ripas_value = private. The first conversion would be fine, but
>
> nit: rec.ripas_value = empty (shared)
>

Sorry that was a bad mixture of pseudocode and constants on my part.

>> the second one woul see ripas_pre = shared but rec.ripas_value = private
>> and would definitely get an error?
>
> rec.ripas_value == empty (shared) as per the guest request. And the RMM
> will find the state is already "shared" and would confirm the success
> back to host.
>

Oh yes, my bad.

base_align says (rephrased)

  if (!AddrIsRttLevelAligned(base, walk.level) &&
      ripas_pre != rec.ripas_value)
          return RMI_ERROR_RTT

Is the alignment check to do with another vCPU splitting the record in
the table?

Why compare ripas_pre and rec.ripas_value when the level is not aligned?

>
>>
>> And there's no "accept" step in the guest after conversions, which makes
>> it different from TDX and SNP.
>
> Correct, the guest "permitted" the GPA to be made private with unknown
> contents anyway. RMM guarantees that the "data" is scrubbed when the
> GPA is made valid. The advantage with this approach is, once the
> Guest sets the RIPAS_RAM, the host can lazily donate pages at
> fault time
>
>>
>> Difficulty in using current gmem hooks that SNP uses:
>>
>> * .gmem_make_shared() called from conversions doesn't have the vCPU
>>    context, finding the right vCPU context is expensive.
>>      * Sean, don't we already iterate vCPUs to find VMSA pages to kick
>>        the right vCPUs?
>
> We could keep a list of VCPUs with pending set-ripas request for e.g.
> Or before the vCPU enter, we check the state of the region and
> do the sync with RMM.
>
>> * .gmem_make_private at fault time is too late
>>      * At the next vCPU enter, the RMM would already read it's state to
>>        report success/failure, and there's no fault in-between for the
>>        .gmem_make_private to happen.
>>
>> At the call Sean suggested mirroring the RMM's tracking in KVM, but that
>> sounds quite arch-specific and it's like doing arch-specific validation
>> within KVM.
>
> Suzuki
>
>
>>
>> p.s. Fuad, for pKVM you also mentioned that you'll need to check if the
>> guest had requested for conversion first? This might be the same/similar
>> problem.

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [Invitation] bi-weekly guest_memfd upstream call on 2026-09-17
  2026-09-17 19:30     ` Sean Christopherson
@ 2026-09-18 16:56       ` Ackerley Tng
  0 siblings, 0 replies; 9+ messages in thread
From: Ackerley Tng @ 2026-09-18 16:56 UTC (permalink / raw)
  To: Sean Christopherson
  Cc: David Hildenbrand (Arm), linux-coco@lists.linux.dev,
	linux-mm@kvack.org, KVM, amit, aneeshkumar.kizhakeveetil,
	ashish.kalra, dwmw2, eberman, fvdl, gshan, jackmanb, jackyli,
	jthoughton, kalyazin, kas, kevinloughlin, liruxin, michael.day,
	michael.roth, mike.rapoport, mvaralar, pankaj.gupta, papaluri,
	patrick.roy, Peter Xu, pheragu, pkondeti, prty, psalian, qinkun,
	shan.gavin, shivankg, sidtelang, suzuki.poulose, tabba, tatashin,
	vannapurve, vbabka, wyihan

Sean Christopherson <seanjc@google.com> writes:

> On Thu, Sep 17, 2026, Ackerley Tng wrote:
>> At the call Sean suggested mirroring the RMM's tracking in KVM, but that
>> sounds quite arch-specific and it's like doing arch-specific validation
>> within KVM.
>
> I suggested "mirroring" the tracking in KVM arm64, not in guest_memfd.
>
> E.g. put a structure in arm64's kvm_vcpu_arch with a list_head object, then insert
> into a per-VM list store in kvm_arch on a conversion request.  But before inserting,
> walk the list to see if there are conflicting requests, and if so, reject the new
> request.
>
> Then on KVM_RUN, if a vCPU has an outstanding request, verify the gmem page exists,
> is in the correct state, and is mapped into the guest (or at least, is known to the
> RMM?).  If any steps fail, exit to userspace with -EFAULT + KVM_EXIT_MEMORY_FAULT.
>
> The downside to such a simplistic implementation is that it'll require a per-VM
> lock to serialize the requests.  If the contention ends up being too painful, e.g.
> because there are usage patterns where guests due batched conversions across many
> vCPUs, then KVM could use a more sophisticated approach as needed.  E.g. shard the
> tracking+locking at 1GiB or something?
>

Perhaps naive question: What if, in the complete_userspace_io handler,
RMI_RTT_SET_RIPAS reads the vcpu_run struct and passes on exactly what
userspace said was handled? (Make the vcpu_run struct for
KVM_EXIT_MEMORY_FAULT bidirectional.)

Is the issue context switches? If the vcpu continuing the vcpu_run is
not the same vcpu that exited?

Perhaps userspace can track conversion requests, like on
KVM_EXIT_MEMORY_FAULT, save the request as <TOKEN>. Then any CPU can
take the token and do the conversion with guest_memfd, associate results
with <TOKEN>. Then on continuing the given vCPU associated with <TOKEN>,
pick up the results for that token and continue.

With that, all races and errors are for the guest/RMM/userspace to
report and handle.


>> p.s. Fuad, for pKVM you also mentioned that you'll need to check if the
>> guest had requested for conversion first? This might be the same/similar
>> problem.

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [Invitation] bi-weekly guest_memfd upstream call on 2026-09-17
  2026-09-18 16:21       ` Ackerley Tng
@ 2026-09-20 11:22         ` Suzuki K Poulose
  0 siblings, 0 replies; 9+ messages in thread
From: Suzuki K Poulose @ 2026-09-20 11:22 UTC (permalink / raw)
  To: Ackerley Tng, David Hildenbrand (Arm), linux-coco@lists.linux.dev,
	linux-mm@kvack.org, KVM
  Cc: amit, aneeshkumar.kizhakeveetil, ashish.kalra, dwmw2, eberman,
	fvdl, gshan, jackmanb, jackyli, jthoughton, kalyazin, kas,
	kevinloughlin, liruxin, michael.day, michael.roth, mike.rapoport,
	mvaralar, pankaj.gupta, papaluri, patrick.roy, Peter Xu, pheragu,
	pkondeti, prty, psalian, qinkun, seanjc, shan.gavin, shivankg,
	sidtelang, tabba, tatashin, vannapurve, vbabka, wyihan

Hi Ackerley,

On 18/09/2026 17:21, Ackerley Tng wrote:
> Suzuki K Poulose <suzuki.poulose@arm.com> writes:
> 
>>
>> [...snip...]
>>
>>>
>>> RMI_RTT_SET_RIPAS() is actually saying "host permits the conversion from
>>> base to top", it's not telling the RMM what to set it to, this is also
>>> different from SNP.
>>
>> As mentioned above, host knows the requested "RIPAS" from VCPU exit.
>> RMM knows the "RIPAS" from the VCPU object.
>>
>> Now: Privates vs Shared is translated to RIPAS_RAM vs RIPAS_EMPTY
>> in the CCA (well, roughly)
>>
>> RIPAS_RAM implies, the GPA can be mapped into the private address
>> space of the Realm and it is integrity protected. Host cannot
>> replace the GPA with another content (it can unmap and DESTROY
>> the GPA mapping. But unless the Realm consents to replace the
>> GPA, again via SET_RIPAS request).
>>
>>
>>>
>>> Interestingly RMI_RTT_SET_RIPAS() also errors out if the current
>>> shared/private state is different from the one tracked in the vCPU
>>> object in the RMM? Did I get that right? I'm looking at the base_align
>>
>> Correct.
>>
>>> Failure condition, where it says ripas_pre != rec.ripas_value.
>>
>> Please note that, in such cases, RMM needs a deeper page table level
>> and the error is RMI_ERROR_RTT, indicating to the host that:
>> Look I need a deeper level table to satisfy the request.
>>
>> e.g., base = 4K, but walk.level = 2, and ripas_pre="private"
>>
>> i.e. the requested base is mapped at L2 as block with "private" ripas.
>> If you want to convert the "base" to shared, it needs L3 table. The host
>> would follow up with RMI_RTT_CREATE and then retry the request.
>>
>>
>>> If two vCPUs race to convert the same range to shared, both vCPUs would
>>> have rec.ripas_value = private. The first conversion would be fine, but
>>
>> nit: rec.ripas_value = empty (shared)
>>
> 
> Sorry that was a bad mixture of pseudocode and constants on my part.
> 
>>> the second one woul see ripas_pre = shared but rec.ripas_value = private
>>> and would definitely get an error?
>>
>> rec.ripas_value == empty (shared) as per the guest request. And the RMM
>> will find the state is already "shared" and would confirm the success
>> back to host.
>>
> 
> Oh yes, my bad.
> 
> base_align says (rephrased)
> 
>    if (!AddrIsRttLevelAligned(base, walk.level) &&
>        ripas_pre != rec.ripas_value)
>            return RMI_ERROR_RTT
> 
> Is the alignment check to do with another vCPU splitting the record in
> the table?

Nope. The Host doesn't keep track of the levels at which the S2 mapping
is created nor does it track the "RIPAS" of the regions. That check is
making sure that we don't unnecessarily break down "block" mappings.

e.g., Before boot, the host might set RIPAS for certain regions (e.g.,
populated contents.) The Guest (firmware) could at boot then convert
all of the "DRAM" regions to RIPAS_RAM without checking what was
populated (this is anyway available in the Measurement). Thus a
conversion request may encounter regions that are already in the
requested state (RIPAS_RAM in the above case). And thus RMM
can report SUCCESS if it is already in the state.

Otherwise, we now need to split the block mapping to a deeper level
to apply the "state" change to a subset of the "block" range. This
is why we return RMI_ERROR_RTT when the requested RIPAS doesn't
match the "ripas_pre"

Does that help ? Remember that we don't have the concept of "ACCEPT"
in CCA.

Cheers
Suzuki


> 
> Why compare ripas_pre and rec.ripas_value when the level is not aligned?
> 
>>
>>>
>>> And there's no "accept" step in the guest after conversions, which makes
>>> it different from TDX and SNP.
>>
>> Correct, the guest "permitted" the GPA to be made private with unknown
>> contents anyway. RMM guarantees that the "data" is scrubbed when the
>> GPA is made valid. The advantage with this approach is, once the
>> Guest sets the RIPAS_RAM, the host can lazily donate pages at
>> fault time
>>
>>>
>>> Difficulty in using current gmem hooks that SNP uses:
>>>
>>> * .gmem_make_shared() called from conversions doesn't have the vCPU
>>>     context, finding the right vCPU context is expensive.
>>>       * Sean, don't we already iterate vCPUs to find VMSA pages to kick
>>>         the right vCPUs?
>>
>> We could keep a list of VCPUs with pending set-ripas request for e.g.
>> Or before the vCPU enter, we check the state of the region and
>> do the sync with RMM.
>>
>>> * .gmem_make_private at fault time is too late
>>>       * At the next vCPU enter, the RMM would already read it's state to
>>>         report success/failure, and there's no fault in-between for the
>>>         .gmem_make_private to happen.
>>>
>>> At the call Sean suggested mirroring the RMM's tracking in KVM, but that
>>> sounds quite arch-specific and it's like doing arch-specific validation
>>> within KVM.
>>
>> Suzuki
>>
>>
>>>
>>> p.s. Fuad, for pKVM you also mentioned that you'll need to check if the
>>> guest had requested for conversion first? This might be the same/similar
>>> problem.


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [Invitation] bi-weekly guest_memfd upstream call on 2026-09-17
  2026-09-17 18:41   ` Ackerley Tng
  2026-09-17 19:30     ` Sean Christopherson
  2026-09-18  9:28     ` Suzuki K Poulose
@ 2026-09-25  8:40     ` Fuad Tabba
  2 siblings, 0 replies; 9+ messages in thread
From: Fuad Tabba @ 2026-09-25  8:40 UTC (permalink / raw)
  To: Ackerley Tng
  Cc: David Hildenbrand (Arm), linux-coco@lists.linux.dev,
	linux-mm@kvack.org, KVM, amit, aneeshkumar.kizhakeveetil,
	ashish.kalra, dwmw2, eberman, fvdl, gshan, jackmanb, jackyli,
	jthoughton, kalyazin, kas, kevinloughlin, liruxin, michael.day,
	michael.roth, mike.rapoport, mvaralar, pankaj.gupta, papaluri,
	patrick.roy, Peter Xu, pheragu, pkondeti, prty, psalian, qinkun,
	seanjc, shan.gavin, shivankg, sidtelang, suzuki.poulose, tatashin,
	vannapurve, vbabka, wyihan

Hi Ackerley,

On Thu, 17 Sep 2026 19:41:14 +0100, Ackerley Tng <ackerleytng@google.com> wrote:
[...]
> p.s. Fuad, for pKVM you also mentioned that you'll need to check if the
> guest had requested for conversion first? This might be the same/similar
> problem.

Sorry, I missed this.

Yes, for unsharing. I'd said EL2 does the transition first and the
host then accepts the conversion [1]. That holds for sharing, but for
unsharing with guest_memfd the host would have to convert to private
first, and only then could EL2 complete the unshare. Otherwise a
failed conversion (-EAGAIN) leaves the host holding a page it can no
longer access.

As with CCA, the unshare would stay pending until the host checks on
KVM_RUN that the range is private, like the check Sean described [2].
I don't think the per-VM list in the host is required, though: EL2
already takes the VM's lock for the transition, so it can track the
pending request and reject conflicts there.

Cheers,
/fuad

[1] https://lore.kernel.org/all/CA+EHjTwLRzrq0vY0-OosENJ2Gkknerux53aVBj0zGRQUkRQP=A@mail.gmail.com/
[2] https://lore.kernel.org/all/aqw_uBIYig_uo5Eq@google.com/

^ permalink raw reply	[flat|nested] 9+ messages in thread

end of thread, other threads:[~2026-09-25  8:41 UTC | newest]

Thread overview: 9+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-16 15:11 [Invitation] bi-weekly guest_memfd upstream call on 2026-09-16 David Hildenbrand (Arm)
2026-09-16 15:16 ` [Invitation] bi-weekly guest_memfd upstream call on 2026-09-17 David Hildenbrand (Arm)
2026-09-17 18:41   ` Ackerley Tng
2026-09-17 19:30     ` Sean Christopherson
2026-09-18 16:56       ` Ackerley Tng
2026-09-18  9:28     ` Suzuki K Poulose
2026-09-18 16:21       ` Ackerley Tng
2026-09-20 11:22         ` Suzuki K Poulose
2026-09-25  8:40     ` Fuad Tabba

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox