Linux-ARM-Kernel Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: David Hildenbrand <david.hildenbrand@arm.com>
To: Sean Christopherson <seanjc@google.com>
Cc: Alexandru Elisei <alexandru.elisei@arm.com>,
	Mark Rutland <mark.rutland@arm.com>,
	pbonzini@redhat.com, kvm@vger.kernel.org, maz@kernel.org,
	oupton@kernel.org, joey.gouly@arm.com, seiden@linux.ibm.com,
	suzuki.poulose@arm.com, yuzenghui@huawei.com,
	linux-arm-kernel@lists.infradead.org, kvmarm@lists.linux.dev,
	fuad.tabba@linux.dev
Subject: Re: [RFC PATCH 0/3] KVM: Dirty page logging for guest_memfd-only memslots
Date: Fri, 14 Aug 2026 09:51:17 +0200	[thread overview]
Message-ID: <487aa57c-72ad-453b-971d-ca24a8380429@arm.com> (raw)
In-Reply-To: <an5UJ4S_P1LUFmzD@google.com>

On 8/14/26 01:32, Sean Christopherson wrote:
> On Thu, Aug 13, 2026, David Hildenbrand wrote:
>> On 8/13/26 18:09, Alexandru Elisei wrote:
>>>
>>> There's already a proposal for how to do this in the migratable guest_memfd
>>> series [1] - it's a new guest_memfd creation flag that disables migration.
>>
>> In the guest_memfd call I was arguing against the flag in the first version, and
>> instead adding it when actually required.
> 
> Ya, right now a flag is meaningless.  Telling guest_memfd not to do something it
> doesn't ever do...

Right.

>>>
>>> I'm a slightly concerned though that all of this will work by chance, and
>>> not by design, and in the future the behaviour might change to allow
>>> guest_memfd memory to be unmapped from stage 2 without the VMM or KVM
>>> explicitly allowing it or initiating it.
> 
> Meh, TDX on x86 already has the same requirement.  Unmapping a page from the S-EPT
> kills the VM unless the VM was expecting the page to be lost. 

Right.

> 
>> Thus my idea of the explicit contract between KVM and guest_memfd. Instead of
>> being a "this doesn't support migration" it would be a "S2 always mapped"
>> kind-of contract.
> 
> Who would that contract be between though?  KVM can tell a guest_memfd instance
> that page migration is/isn't supported, but telling guest_memfd that the VM will
> always keep the entirety of the guest_memfd mapped in stage-2 is nonsensical.
> guest_memfd simply doesn't care if the page is mapped or not, it only cares if
> migration is supported.

I think there are more details to this. Let me try to express what my mind was
up to last night.

We have a hardware feature that requires pages to always be mapped into S2. Some
things I had in mind:

1) Page migration would not be a problem as long as hardware could be paused
   while migrating (e.g., kick all vCPUs). I doubt someone would implement that
   right now,  but you could consider it an implementation detail that page
   migration cannot be supported right now.

2) Newer hardware could mitigate this problem, allowing the feature to support
   pages temporarily being unmapped from S2.

3) Disallowing page migration is really just one implication of "pages must
   always be mapped into S2".

So what we really want is "if feature X is enabled and hardware requires it,
always keep pages mapped into S2, which currently implies that page migration
cannot be supported."

Which isn't all that different to "if a confidential VM is run on current TDX
hardware, always keep pages mapped into S2, which currently implies that page
migration cannot be supported."

So I was wondering whether the flow could be:

User space enabled CPU feature for VM -> KVM knows that current hardware
requires for that CPU feature to have S2 always mapped -> KVM tells guest_memfd
that S2 must be always mapped / disables page migration.

That would be in contrast to user space having to guess that page migration on
the current hardware with the current guest_memfd implementation does not
support page migration, to then disable exactly that.

Does that explanation makes sense? I don't know the exact mechanism to do that,
but that's just my high-level thinking.

> 
>>> My understanding from the conversation so far is that the plan for the
>>> future of guest_memfd is to support an option/mode where the memory is
>>> effectively "pinned" at stage 2 (but which allows userspace to explicitly
>>> free/unmap it, of course). Is that correct, or am I being overly optimistic
>>> in my interpretation?
>>
>> We could then even disallow fallocate() to punch holes if that contract is
>> negotiated.
> 
> Why?  If userspace pulls a stupid and kills its guest, that's userspace's problem.

It was late, agreed. Unplugging memory would just work, so that's not a concern.
It's a userspace's problem.

> 
> KVM would also have to block memslot changes, and probably other things in the
> future that would unmap stage-2 in response to userspace syscalls/ioctls.


-- 
Cheers,

David


  reply	other threads:[~2026-08-14  7:52 UTC|newest]

Thread overview: 31+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-02 14:29 [RFC PATCH 0/3] KVM: Dirty page logging for guest_memfd-only memslots Alexandru Elisei
2026-07-02 14:29 ` [RFC PATCH 1/3] KVM: guest_memfd: Use memslot id to keep track of associated memslots Alexandru Elisei
2026-07-06  7:14   ` David Hildenbrand
2026-07-06 13:45     ` Alexandru Elisei
2026-07-06 21:46       ` Sean Christopherson
2026-07-07 17:05         ` Alexandru Elisei
2026-07-06 21:43   ` Sean Christopherson
2026-07-07 17:05     ` Alexandru Elisei
2026-07-13 14:03       ` David Hildenbrand
2026-07-13 16:13         ` Sean Christopherson
2026-07-13 13:42     ` David Hildenbrand
2026-07-02 14:29 ` [RFC PATCH 2/3] KVM: Implement dirty page logging for guest_memfd-only memslots Alexandru Elisei
2026-07-07  1:29   ` Sean Christopherson
2026-07-07 17:12     ` Alexandru Elisei
2026-07-14  5:39       ` Kishen Maloor
2026-07-02 14:29 ` [RFC PATCH 3/3] KVM: arm64: Allow " Alexandru Elisei
2026-07-07  0:56 ` [RFC PATCH 0/3] KVM: Dirty " Sean Christopherson
2026-07-07 16:58   ` Alexandru Elisei
2026-07-07 17:12     ` Sean Christopherson
2026-07-09 11:21       ` Mark Rutland
2026-07-09 19:01         ` Sean Christopherson
2026-07-10 10:26           ` Mark Rutland
2026-07-13 14:11             ` David Hildenbrand
2026-08-13 16:09               ` Alexandru Elisei
2026-08-13 20:01                 ` David Hildenbrand
2026-08-13 23:32                   ` Sean Christopherson
2026-08-14  7:51                     ` David Hildenbrand [this message]
2026-07-09 20:33     ` Oliver Upton
2026-07-10 10:44       ` Alexandru Elisei
2026-07-10 18:18         ` Oliver Upton
2026-07-13 13:48           ` Alexandru Elisei

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=487aa57c-72ad-453b-971d-ca24a8380429@arm.com \
    --to=david.hildenbrand@arm.com \
    --cc=alexandru.elisei@arm.com \
    --cc=fuad.tabba@linux.dev \
    --cc=joey.gouly@arm.com \
    --cc=kvm@vger.kernel.org \
    --cc=kvmarm@lists.linux.dev \
    --cc=linux-arm-kernel@lists.infradead.org \
    --cc=mark.rutland@arm.com \
    --cc=maz@kernel.org \
    --cc=oupton@kernel.org \
    --cc=pbonzini@redhat.com \
    --cc=seanjc@google.com \
    --cc=seiden@linux.ibm.com \
    --cc=suzuki.poulose@arm.com \
    --cc=yuzenghui@huawei.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox