* [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
@ 2026-08-31 7:13 Tony Lindgren
2026-08-31 7:13 ` [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests Tony Lindgren
` (6 more replies)
0 siblings, 7 replies; 84+ messages in thread
From: Tony Lindgren @ 2026-08-31 7:13 UTC (permalink / raw)
To: Paolo Bonzini, Sean Christopherson
Cc: Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Jörg Rödel ,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
Hi all,
As discussed in a recent PUCK call, Sean suggested we post what Intel is
using for the KVM live migration API for TDX as an example to see if we
can come up with APIs that are not vendor specific.
The goal of this patch series is to start a discussion about live migration
kernel APIs for CoCo guests.
Tom, since you mentioned that AMD SEV-SNP and Intel TDX live migration
sound similar, can you please take a look how the API might work for
SEV-SNP?
For CoCo VMs, the guest memory and vCPU states are not accessible to the
userspace or KVM for live migration. The memory and vCPU states need to be
extracted into encrypted blobs on the source, and decrypted on the
destination. Before live migration, an encryption key needs to be
negotiated between the source and destination.
The layer handling the encryption for live migration is implementation
specific. It can be the TDX module or Coconut-SVSM for example.
For CoCo VMs, the KVM_MEMORY_ENCRYPT_OP ioctl() has been used with vendor
specific sub-commands. Adding more vendor specific sub-commands is an
option also for live migration. However, depending on how similar the KVM
needs are, it may be possible to have a common API.
For TDX, we're using a group of ioctl()s that might be possible to adapt
also for other CoCo implementations.
Artem has put together a brief description below of the example API and the
migration flow:
Example API
===========
- KVM_CAP_LIVE_MIGRATION - if a VM supports live migration through this
uAPI.
- KVM_MIGRATE_CMD - the main ioctl that drives the migration phases. Each
command takes vendor-specific flags and a buffer for the blob that travels
between the hosts.
- KVM_MIGRATE_SETUP - establish the migration session and transfer the
immutable VM state.
- KVM_MIGRATE_ITERATION - close a memory copy round.
- KVM_MIGRATE_STOP_AND_COPY - pause the VM and transfer the remaining VM
state.
- KVM_MIGRATE_END - complete the migration, or abort it.
- KVM_EXPORT_MEMORY - export memory pages on the source host.
- KVM_IMPORT_MEMORY - import memory pages on the destination host.
- KVM_EXPORT_VCPU - export vCPU state on the source host.
- KVM_IMPORT_VCPU - import vCPU state on the destination host.
Dirty page tracking does not add a new uAPI. Userspace keeps using
KVM_GET_DIRTY_LOG and KVM_CLEAR_DIRTY_LOG.
Migration flow
==============
Source host Destination host
=========== ================
CMD(SETUP/SESSION) <--- setup msgs ---> CMD(SETUP/SESSION)
| (repeated) |
CMD(SETUP/IMMUTABLE_STATE) - immutable state -> CMD(SETUP/IMMUTABLE_STATE)
| |
KVM_GET_DIRTY_LOG |
KVM_EXPORT_MEMORY --- memory data ---> KVM_IMPORT_MEMORY
CMD(ITERATION) --- epoch token ---> CMD(ITERATION)
| (repeat until convergence) |
CMD(STOP_AND_COPY/PAUSE) |
CMD(STOP_AND_COPY/TD_STATE) --- VM state ------> CMD(STOP_AND_COPY/TD_STATE)
KVM_EXPORT_VCPU --- vCPU state ----> KVM_IMPORT_VCPU
KVM_EXPORT_MEMORY -- final memory ---> KVM_IMPORT_MEMORY
CMD(ITERATION/DONE) --- start token ---> CMD(ITERATION)
| |
CMD(END) CMD(END)
For the TDX implementation, the above map to the TDX module SEAMCALLs.
Regards,
Tony
Changes since v1 at [0] below:
- Drop KVM_MIGRATE_CMD sub-command PREPARE, SETUP sub-command has been
enough for TDX at least
- Rename KVM_MIGRATE_CMD sub-command KVM_MIGRATE_TOKEN to
KVM_MIGRATE_ITERATION
- Rename KVM_MIGRATE_CMD sub-command KVM_MIGRATE_SOURCE_BLACKOUT to
KVM_MIGRATE_STOP_AND_COPY
- Add x86 ioctl handling
[0] https://lore.kernel.org/kvm/20251006113524.1573116-1-tony.lindgren@linux.intel.com/
Tony Lindgren (4):
Documentation: KVM: Add live migration API for confidential guests
KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY
KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU
Documentation/virt/kvm/api.rst | 205 +++++++++++++++++++++++++++++
arch/x86/include/asm/kvm-x86-ops.h | 6 +
arch/x86/include/asm/kvm_host.h | 6 +
arch/x86/kvm/x86.c | 102 ++++++++++++++
include/uapi/linux/kvm.h | 43 ++++++
5 files changed, 362 insertions(+)
base-commit: dc59e4fea9d83f03bad6bddf3fa2e52491777482
--
2.43.0
^ permalink raw reply [flat|nested] 84+ messages in thread* [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests 2026-08-31 7:13 [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Tony Lindgren @ 2026-08-31 7:13 ` Tony Lindgren 2026-08-31 7:20 ` sashiko-bot 2026-09-18 11:35 ` Peter Xu 2026-08-31 7:13 ` [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD Tony Lindgren ` (5 subsequent siblings) 6 siblings, 2 replies; 84+ messages in thread From: Tony Lindgren @ 2026-08-31 7:13 UTC (permalink / raw) To: Paolo Bonzini, Sean Christopherson Cc: Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel , Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm For CoCo VMs, the guest memory and vCPU states are not accessible to the userspace or KVM for live migration. The memory and vCPU states need to be extracted into encrypted blobs on the source, and decrypted on the destination. Before live migration, an encryption key needs to be negotiated between the source and destination. KVM help is needed to talk to the layer exporting and importing the encrypted state. Document the KVM live migration API for confidential guests. Co-developed-by: Kishen Maloor <kishen.maloor@intel.com> Signed-off-by: Kishen Maloor <kishen.maloor@intel.com> Signed-off-by: Tony Lindgren <tony.lindgren@linux.intel.com> --- Documentation/virt/kvm/api.rst | 205 +++++++++++++++++++++++++++++++++ 1 file changed, 205 insertions(+) diff --git a/Documentation/virt/kvm/api.rst b/Documentation/virt/kvm/api.rst index a5f9ee92f43e8..9d546d288af5f 100644 --- a/Documentation/virt/kvm/api.rst +++ b/Documentation/virt/kvm/api.rst @@ -6566,6 +6566,200 @@ KVM_S390_KEYOP_SSKE Sets the storage key for the guest address ``guest_addr`` to the key specified in ``key``, returning the previous value in ``key``. +.. _KVM_MIGRATE_CMD: + +4.145 KVM_MIGRATE_CMD +--------------------- + +:Capability: KVM_CAP_LIVE_MIGRATION +:Architectures: x86 +:Type: vm ioctl +:Parameters: struct kvm_migrate_cmd (in/out) +:Returns: 0 on success, < 0 on error + +Allows userspace to send live migration related commands to KVM for vendor +specific handling. + +For confidential computing, live migration related commands may be needed. +The commands typically use encrypted data that needs to be passed between the +source and destination hosts. The hosts may also require specific coordination +steps during migration that must be triggered at precise points in the +migration process. + +The vendor specific implementation handles locking and checks the valid flags +bits. If KVM_CAP_LIVE_MIGRATION is not available for the VM, -ENOTTY is +returned. + +The KVM_MIGRATE_CMD subcommand passed in struct kvm_migrate_cmd is one of:: + + #define KVM_MIGRATE_SETUP 0 + #define KVM_MIGRATE_ITERATION 1 + #define KVM_MIGRATE_STOP_AND_COPY 2 + #define KVM_MIGRATE_ABORT 3 + #define KVM_MIGRATE_END 4 + +The kvm_transfer_buffer is:: + + /** + * @address: Userspace buffer address + * @size: Size of the userspace buffer + * @reserved: Reserved for future use + */ + struct kvm_transfer_buffer { + __u64 address; + __u32 size; + __u32 reserved; + }; + +The kvm_migrate_cmd is:: + + /** + * @command: One of the defined KVM_MIGRATE commands + * @flags: Hardware specific flags + * @reserved: Reserved for future use + * @buf: Userspace buffer for hardware specific data + */ + struct kvm_migrate_cmd { + __u16 command; + __u16 flags; + __u32 reserved; + struct kvm_transfer_buffer buf; + }; + +.. _KVM_EXPORT_MEMORY: + +4.146 KVM_EXPORT_MEMORY +----------------------- + +:Capability: KVM_CAP_LIVE_MIGRATION +:Architectures: x86 +:Type: vm ioctl +:Parameters: struct kvm_memory_transfer (in/out) +:Returns: 0 on success, < 0 on error + +Allows userspace to request the host to export an array of memory pages to a +userspace buffer. + +The private memory may not be accessible to KVM because of encryption. For +confidential computing, the guest memory is encrypted and only accessible to +the guest. + +If KVM_CAP_LIVE_MIGRATION is not available for the VM, -ENOTTY is returned. + +The vendor specific ID is used at least for TDX for the migration thread +index. + +The kvm_memory_transfer is:: + + /** + * @gfns: Userspace address of an array of nr_gfns __u64 GFNs to export + * @nr_gfns: Number of GFNs in the @gfns array + * @id: Optional vendor specific transfer ID + * @flags: Vendor specific flags + * @reserved: Reserved for future use + * @buf: Userspace buffer to export memory to + */ + struct kvm_memory_transfer { + __u64 gfns; + __u32 nr_gfns; + __u16 id; + __u16 flags; + __u64 reserved; + struct kvm_transfer_buffer buf; + }; + +The transfer buffer size is vendor specific. + +For the transfer buffer, seeo :ref:`KVM_MIGRATE_CMD <KVM_MIGRATE_CMD>`. + +For memory import, see also :ref:`KVM_IMPORT_MEMORY <KVM_IMPORT_MEMORY>`. + + +.. _KVM_IMPORT_MEMORY: + +4.147 KVM_IMPORT_MEMORY +----------------------- + +:Capability: KVM_CAP_LIVE_MIGRATION +:Architectures: x86 +:Type: vm ioctl +:Parameters: struct kvm_memory_transfer (in/out) +:Returns: 0 on success, < 0 on error + +Allows userspace to request the host to import an array of memory pages from a +userspace buffer. + +The private memory may not be accessible to KVM because of encryption. For +confidential computing, the guest memory is encrypted and only accessible to +the guest. + +If KVM_CAP_LIVE_MIGRATION is not available for the VM, -ENOTTY is returned. + +The vendor specific ID is used at least for TDX for the migration thread +index. + +The transfer buffer size is vendor specific. + +For kvm_memory_transfer, see :ref:`KVM_EXPORT_MEMORY <KVM_EXPORT_MEMORY>`. + +For the transfer buffer, seeo :ref:`KVM_MIGRATE_CMD <KVM_MIGRATE_CMD>`. + +.. _KVM_EXPORT_VCPU: + +4.149 KVM_EXPORT_VCPU +--------------------- +:Capability: KVM_CAP_LIVE_MIGRATION +:Architectures: arm64, x86 +:Type: vcpu ioctl +:Parameters: struct kvm_vcpu_transfer (in/out) +:Returns: 0 on success, < 0 on error + +Allows userspace to request the host to export a VCPU state to a userspace +buffer. + +The VCPU state may not be directly accessible to KVM because of encryption. For +confidential computing, the VCPU state is encrypted and only accessible to the +guest. + +The vcpu_transfer is:: + + /** + * @flags: Hardware specific flags + * @reserved: Reserved for future use + * @buf: Userspace buffer to export VCPU state to + */ + struct kvm_vcpu_transfer { + __u32 flags; + __u32 reserved; + struct kvm_transfer_buffer buf; + }; + +For the transfer buffer, see :ref:`KVM_MIGRATE_CMD <KVM_MIGRATE_CMD>`. + +For vCPU import, see also :ref:`KVM_IMPORT_VCPU <KVM_IMPORT_VCPU>`. + +.. _KVM_IMPORT_VCPU: + +4.148 KVM_IMPORT_VCPU +--------------------- + +:Capability: KVM_CAP_LIVE_MIGRATION +:Architectures: arm64, x86 +:Type: vcpu ioctl +:Parameters: struct kvm_vcpu_transfer (in/out) +:Returns: 0 on success, < 0 on error + +Allows userspace to request the host to import a VCPU state from a userspace +buffer. + +The VCPU state may not be directly accessible to KVM because of encryption. For +confidential computing, the VCPU state is encrypted and only accessible to the +guest. + +For vcpu_transfer and vCPU import, see :ref:`KVM_IMPORT_VCPU <KVM_IMPORT_VCPU>`. + +For the transfer buffer, see also :ref:`KVM_MIGRATE_CMD <KVM_MIGRATE_CMD>`. + .. _kvm_run: 5. The kvm_run structure @@ -9493,6 +9687,17 @@ take care to differentiate between these cases. The presence of this capability indicates that the nested KVM guest can start in ESA mode. +8.48 KVM_CAP_LIVE_MIGRATION +--------------------------- + +:Architectures: x86 +:Target: VM +:Parameters: None + +Indicates that the VM needs to use KVM calls for live migration, and that the +KVM_MIGRATE_CMD ioctl and the KVM_EXPORT_MEMORY, KVM_IMPORT_MEMORY, +KVM_EXPORT_VCPU and KVM_IMPORT_VCPU ioctls are available. + 9. Known KVM API problems ========================= -- 2.43.0 ^ permalink raw reply related [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests 2026-08-31 7:13 ` [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests Tony Lindgren @ 2026-08-31 7:20 ` sashiko-bot 2026-09-18 11:35 ` Peter Xu 1 sibling, 0 replies; 84+ messages in thread From: sashiko-bot @ 2026-08-31 7:20 UTC (permalink / raw) To: Tony Lindgren; +Cc: kvm Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider: - [Low] The newly added documentation for the KVM live migration API contains architectural contradictions, backwards numbering, self-referencing links, and typographical errors. -- commit 3a6c0c9c47ef3277e964b009e8dd06e7aeab932d Author: Tony Lindgren <tony.lindgren@linux.intel.com> Documentation: KVM: Add live migration API for confidential guests For CoCo VMs, the guest memory and vCPU states are not accessible to the userspace or KVM for live migration. The memory and vCPU states need to be extracted into encrypted blobs on the source, and decrypted on the destination. Before live migration, an encryption key needs to be negotiated between the source and destination. > diff --git a/Documentation/virt/kvm/api.rst b/Documentation/virt/kvm/api.rst > index a5f9ee92f43e8..9d546d288af5f 100644 > --- a/Documentation/virt/kvm/api.rst > +++ b/Documentation/virt/kvm/api.rst [ ... ] > +The transfer buffer size is vendor specific. > + > +For the transfer buffer, seeo :ref:`KVM_MIGRATE_CMD <KVM_MIGRATE_CMD>`. [Severity: Low] Is this "seeo" a typo for "see"? > + > +For memory import, see also :ref:`KVM_IMPORT_MEMORY <KVM_IMPORT_MEMORY>`. > + > + > +.. _KVM_IMPORT_MEMORY: > + > +4.147 KVM_IMPORT_MEMORY > +----------------------- [ ... ] > +The transfer buffer size is vendor specific. > + > +For kvm_memory_transfer, see :ref:`KVM_EXPORT_MEMORY <KVM_EXPORT_MEMORY>`. > + > +For the transfer buffer, seeo :ref:`KVM_MIGRATE_CMD <KVM_MIGRATE_CMD>`. [Severity: Low] Is there another "seeo" typo here? > + > +.. _KVM_EXPORT_VCPU: > + > +4.149 KVM_EXPORT_VCPU > +--------------------- > +:Capability: KVM_CAP_LIVE_MIGRATION > +:Architectures: arm64, x86 [Severity: Low] Is it intentional that KVM_EXPORT_VCPU is numbered 4.149 before KVM_IMPORT_VCPU which is numbered 4.148? Also, KVM_EXPORT_VCPU lists arm64 as a supported architecture, but relies on KVM_CAP_LIVE_MIGRATION which is documented as x86 only below. Should arm64 be removed here, or added to the capability documentation? > +:Type: vcpu ioctl > +:Parameters: struct kvm_vcpu_transfer (in/out) > +:Returns: 0 on success, < 0 on error [ ... ] > +.. _KVM_IMPORT_VCPU: > + > +4.148 KVM_IMPORT_VCPU > +--------------------- > + > +:Capability: KVM_CAP_LIVE_MIGRATION > +:Architectures: arm64, x86 > +:Type: vcpu ioctl > +:Parameters: struct kvm_vcpu_transfer (in/out) > +:Returns: 0 on success, < 0 on error > + > +Allows userspace to request the host to import a VCPU state from a userspace > +buffer. > + > +The VCPU state may not be directly accessible to KVM because of encryption. For > +confidential computing, the VCPU state is encrypted and only accessible to the > +guest. > + > +For vcpu_transfer and vCPU import, see :ref:`KVM_IMPORT_VCPU <KVM_IMPORT_VCPU>`. [Severity: Low] Does this circular reference point to itself? Should this point to KVM_EXPORT_VCPU instead to help the reader find the struct kvm_vcpu_transfer definition? > + > +For the transfer buffer, see also :ref:`KVM_MIGRATE_CMD <KVM_MIGRATE_CMD>`. > + > .. _kvm_run: > > 5. The kvm_run structure > @@ -9493,6 +9687,17 @@ take care to differentiate between these cases. > The presence of this capability indicates that the nested KVM guest can > start in ESA mode. > > +8.48 KVM_CAP_LIVE_MIGRATION > +--------------------------- > + > +:Architectures: x86 > +:Target: VM > +:Parameters: None [Severity: Low] The KVM_CAP_LIVE_MIGRATION capability is documented here as x86 only, but KVM_EXPORT_VCPU and KVM_IMPORT_VCPU state they are for arm64 and x86. Is this capability intended to support arm64 as well? -- Sashiko AI review · https://sashiko.dev/#/patchset/20260831071304.762939-1-tony.lindgren@linux.intel.com?part=1 ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests 2026-08-31 7:13 ` [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests Tony Lindgren 2026-08-31 7:20 ` sashiko-bot @ 2026-09-18 11:35 ` Peter Xu 2026-09-21 4:20 ` Tony Lindgren 2026-09-24 1:50 ` Wei Wang 1 sibling, 2 replies; 84+ messages in thread From: Peter Xu @ 2026-09-18 11:35 UTC (permalink / raw) To: Tony Lindgren Cc: Paolo Bonzini, Sean Christopherson, Artem Bityutskiy, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Mon, Aug 31, 2026 at 10:13:01AM +0300, Tony Lindgren wrote: > +:Capability: KVM_CAP_LIVE_MIGRATION IMHO this is slightly misleading, some "CONFIDENTIAL_" or other prefix would be nice. Thanks, -- Peter Xu ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests 2026-09-18 11:35 ` Peter Xu @ 2026-09-21 4:20 ` Tony Lindgren 2026-09-24 1:50 ` Wei Wang 1 sibling, 0 replies; 84+ messages in thread From: Tony Lindgren @ 2026-09-21 4:20 UTC (permalink / raw) To: Peter Xu Cc: Paolo Bonzini, Sean Christopherson, Artem Bityutskiy, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm Hi, On Fri, Sep 18, 2026 at 07:35:39AM -0400, Peter Xu wrote: > On Mon, Aug 31, 2026 at 10:13:01AM +0300, Tony Lindgren wrote: > > +:Capability: KVM_CAP_LIVE_MIGRATION > > IMHO this is slightly misleading, some "CONFIDENTIAL_" or other prefix > would be nice. OK. For the possible prefixes to consider, also HW for hardware specific ops might work. I'm not aware of kernel side migration needs for this other than CoCo though. Regards, Tony ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests 2026-09-18 11:35 ` Peter Xu 2026-09-21 4:20 ` Tony Lindgren @ 2026-09-24 1:50 ` Wei Wang 2026-09-24 4:51 ` Tony Lindgren 1 sibling, 1 reply; 84+ messages in thread From: Wei Wang @ 2026-09-24 1:50 UTC (permalink / raw) To: Peter Xu, Tony Lindgren Cc: Paolo Bonzini, Sean Christopherson, Artem Bityutskiy, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On 9/18/26 7:35 PM, Peter Xu wrote: > On Mon, Aug 31, 2026 at 10:13:01AM +0300, Tony Lindgren wrote: >> +:Capability: KVM_CAP_LIVE_MIGRATION > > IMHO this is slightly misleading, some "CONFIDENTIAL_" or other prefix > would be nice. > An earlier version of this series used KVM_CAP_CGM (CGM stands for Confidential Guest Migration). I think it's cleaner to use a short CGM prefix for the related uAPIs, so they're easy to identify as a group. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests 2026-09-24 1:50 ` Wei Wang @ 2026-09-24 4:51 ` Tony Lindgren 0 siblings, 0 replies; 84+ messages in thread From: Tony Lindgren @ 2026-09-24 4:51 UTC (permalink / raw) To: Wei Wang Cc: Peter Xu, Paolo Bonzini, Sean Christopherson, Artem Bityutskiy, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm Hi Wei, On Thu, Sep 24, 2026 at 09:50:45AM +0800, Wei Wang wrote: > On 9/18/26 7:35 PM, Peter Xu wrote: > > On Mon, Aug 31, 2026 at 10:13:01AM +0300, Tony Lindgren wrote: > > > +:Capability: KVM_CAP_LIVE_MIGRATION > > > > IMHO this is slightly misleading, some "CONFIDENTIAL_" or other prefix > > would be nice. > > > An earlier version of this series used KVM_CAP_CGM (CGM stands for > Confidential Guest Migration). I think it's cleaner to use a short > CGM prefix for the related uAPIs, so they're easy to identify as a group. Using an abbreviation like CGM or MIG or HW has a problem where it's hard to decipher for anybody not familiar with the live migration. So my vote is currently on using Peter's suggestion for CONFIDENTIAL prefix. Regards, Tony ^ permalink raw reply [flat|nested] 84+ messages in thread
* [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-08-31 7:13 [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Tony Lindgren 2026-08-31 7:13 ` [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests Tony Lindgren @ 2026-08-31 7:13 ` Tony Lindgren 2026-08-31 7:23 ` sashiko-bot ` (2 more replies) 2026-08-31 7:13 ` [RFC PATCH v2 3/4] KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY Tony Lindgren ` (4 subsequent siblings) 6 siblings, 3 replies; 84+ messages in thread From: Tony Lindgren @ 2026-08-31 7:13 UTC (permalink / raw) To: Paolo Bonzini, Sean Christopherson Cc: Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel , Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm Live migration of confidential guests needs the help of KVM at least for TDX. Add KVM_CAP_LIVE_MIGRATION for when hardware specific live migration functions must be used. Add KVM_MIGRATE_CMD to configure the hardware for live migration. Assisted-by: Claude-Code:claude-opus-5 checkpatch [ used AI to review and simplify the code ] Signed-off-by: Tony Lindgren <tony.lindgren@linux.intel.com> --- arch/x86/include/asm/kvm-x86-ops.h | 2 ++ arch/x86/include/asm/kvm_host.h | 2 ++ arch/x86/kvm/x86.c | 25 +++++++++++++++++++++++++ include/uapi/linux/kvm.h | 22 ++++++++++++++++++++++ 4 files changed, 51 insertions(+) diff --git a/arch/x86/include/asm/kvm-x86-ops.h b/arch/x86/include/asm/kvm-x86-ops.h index 83dc5086138b3..ac080b556b0c8 100644 --- a/arch/x86/include/asm/kvm-x86-ops.h +++ b/arch/x86/include/asm/kvm-x86-ops.h @@ -148,6 +148,8 @@ KVM_X86_OP_OPTIONAL(alloc_apic_backing_page) KVM_X86_OP_OPTIONAL_RET0(gmem_prepare) KVM_X86_OP_OPTIONAL_RET0(gmem_max_mapping_level) KVM_X86_OP_OPTIONAL(gmem_invalidate) +KVM_X86_OP_OPTIONAL_RET0(cap_live_migration) +KVM_X86_OP_OPTIONAL(migrate_cmd) #endif #undef KVM_X86_OP diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h index 5f6c1ce9673b7..d9291a8a97bb1 100644 --- a/arch/x86/include/asm/kvm_host.h +++ b/arch/x86/include/asm/kvm_host.h @@ -2010,6 +2010,8 @@ struct kvm_x86_ops { int (*gmem_prepare)(struct kvm *kvm, kvm_pfn_t pfn, gfn_t gfn, int max_order); void (*gmem_invalidate)(kvm_pfn_t start, kvm_pfn_t end); int (*gmem_max_mapping_level)(struct kvm *kvm, kvm_pfn_t pfn, bool is_private); + bool (*cap_live_migration)(struct kvm *kvm); + int (*migrate_cmd)(struct kvm *kvm, struct kvm_migrate_cmd *cmd); }; struct kvm_x86_nested_ops { diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c index afcac1042947a..7064fd709e56d 100644 --- a/arch/x86/kvm/x86.c +++ b/arch/x86/kvm/x86.c @@ -4973,6 +4973,9 @@ int kvm_vm_ioctl_check_extension(struct kvm *kvm, long ext) case KVM_CAP_READONLY_MEM: r = kvm ? kvm_arch_has_readonly_mem(kvm) : 1; break; + case KVM_CAP_LIVE_MIGRATION: + r = kvm ? kvm_x86_call(cap_live_migration)(kvm) : 0; + break; default: break; } @@ -7614,6 +7617,28 @@ int kvm_arch_vm_ioctl(struct file *filp, unsigned int ioctl, unsigned long arg) r = kvm_vm_ioctl_set_msr_filter(kvm, &filter); break; } + case KVM_MIGRATE_CMD: { + struct kvm_migrate_cmd cmd; + + if (!kvm_x86_ops.migrate_cmd || + !kvm_x86_call(cap_live_migration)(kvm)) + return -ENOTTY; + + if (copy_from_user(&cmd, argp, sizeof(cmd))) + return -EFAULT; + + if (cmd.reserved || cmd.buf.reserved) + return -EINVAL; + + r = kvm_x86_call(migrate_cmd)(kvm, &cmd); + if (r > 0) + r = -EIO; + + /* Copy back also on an error to report a partially done command */ + if (copy_to_user(argp, &cmd, sizeof(cmd))) + return -EFAULT; + break; + } default: r = -ENOTTY; } diff --git a/include/uapi/linux/kvm.h b/include/uapi/linux/kvm.h index 419011097fa8e..e5b227d3a8e40 100644 --- a/include/uapi/linux/kvm.h +++ b/include/uapi/linux/kvm.h @@ -997,6 +997,7 @@ struct kvm_enable_cap { #define KVM_CAP_S390_KEYOP 247 #define KVM_CAP_S390_VSIE_ESAMODE 248 #define KVM_CAP_S390_HPAGE_2G 249 +#define KVM_CAP_LIVE_MIGRATION 250 struct kvm_irq_routing_irqchip { __u32 irqchip; @@ -1350,6 +1351,8 @@ struct kvm_s390_keyop { #define KVM_GET_DEVICE_ATTR _IOW(KVMIO, 0xe2, struct kvm_device_attr) #define KVM_HAS_DEVICE_ATTR _IOW(KVMIO, 0xe3, struct kvm_device_attr) +#define KVM_MIGRATE_CMD _IOWR(KVMIO, 0xe4, struct kvm_migrate_cmd) + /* * ioctls for vcpu fds */ @@ -1670,4 +1673,23 @@ struct kvm_pre_fault_memory { __u64 padding[5]; }; +#define KVM_MIGRATE_SETUP 0 +#define KVM_MIGRATE_ITERATION 1 +#define KVM_MIGRATE_STOP_AND_COPY 2 +#define KVM_MIGRATE_ABORT 3 +#define KVM_MIGRATE_END 4 + +struct kvm_transfer_buffer { + __u64 address; + __u32 size; + __u32 reserved; +}; + +struct kvm_migrate_cmd { + __u16 command; + __u16 flags; + __u32 reserved; + struct kvm_transfer_buffer buf; +}; + #endif /* __LINUX_KVM_H */ -- 2.43.0 ^ permalink raw reply related [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-08-31 7:13 ` [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD Tony Lindgren @ 2026-08-31 7:23 ` sashiko-bot 2026-09-01 6:03 ` Tony Lindgren 2026-09-07 11:53 ` Tony Lindgren 2026-09-18 4:33 ` Kishen Maloor 2 siblings, 1 reply; 84+ messages in thread From: sashiko-bot @ 2026-08-31 7:23 UTC (permalink / raw) To: Tony Lindgren; +Cc: kvm Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider: - [High] The 'flags' field in the new 'struct kvm_migrate_cmd' UAPI is not validated to be zero. -- commit b50bc10fe65ad9a54c9b979c445505e5baf76592 Author: Tony Lindgren <tony.lindgren@linux.intel.com> KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD This patch introduces the KVM_CAP_LIVE_MIGRATION capability and the KVM_MIGRATE_CMD ioctl. These additions allow KVM to configure hardware-specific live migration functions, which is necessary for confidential guests like TDX. > diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c > index afcac1042947a..7064fd709e56d 100644 > --- a/arch/x86/kvm/x86.c > +++ b/arch/x86/kvm/x86.c [ ... ] > @@ -7614,6 +7617,28 @@ int kvm_arch_vm_ioctl(struct file *filp, unsigned int ioctl, unsigned long arg) > r = kvm_vm_ioctl_set_msr_filter(kvm, &filter); > break; > } > + case KVM_MIGRATE_CMD: { > + struct kvm_migrate_cmd cmd; > + > + if (!kvm_x86_ops.migrate_cmd || > + !kvm_x86_call(cap_live_migration)(kvm)) > + return -ENOTTY; > + > + if (copy_from_user(&cmd, argp, sizeof(cmd))) > + return -EFAULT; > + > + if (cmd.reserved || cmd.buf.reserved) > + return -EINVAL; [Severity: High] Does this code also need to validate that the new flags field is zero? If uninitialized memory or garbage is passed by userspace in cmd.flags, it will currently be silently ignored. When a future kernel introduces a new flag, old userspace applications passing uninitialized flags will inadvertently enable it, which could permanently break backwards compatibility. > + > + r = kvm_x86_call(migrate_cmd)(kvm, &cmd); > + if (r > 0) > + r = -EIO; > + > + /* Copy back also on an error to report a partially done command */ > + if (copy_to_user(argp, &cmd, sizeof(cmd))) > + return -EFAULT; > + break; > + } -- Sashiko AI review · https://sashiko.dev/#/patchset/20260831071304.762939-1-tony.lindgren@linux.intel.com?part=2 ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-08-31 7:23 ` sashiko-bot @ 2026-09-01 6:03 ` Tony Lindgren 0 siblings, 0 replies; 84+ messages in thread From: Tony Lindgren @ 2026-09-01 6:03 UTC (permalink / raw) To: sashiko-reviews; +Cc: kvm On Mon, Aug 31, 2026 at 07:23:11AM +0000, sashiko-bot@kernel.org wrote: > Author: Tony Lindgren <tony.lindgren@linux.intel.com> > > --- a/arch/x86/kvm/x86.c > > +++ b/arch/x86/kvm/x86.c > > @@ -7614,6 +7617,28 @@ int kvm_arch_vm_ioctl(struct file *filp, unsigned int ioctl, unsigned long arg) > > r = kvm_vm_ioctl_set_msr_filter(kvm, &filter); > > break; > > } > > + case KVM_MIGRATE_CMD: { > > + struct kvm_migrate_cmd cmd; > > + > > + if (!kvm_x86_ops.migrate_cmd || > > + !kvm_x86_call(cap_live_migration)(kvm)) > > + return -ENOTTY; > > + > > + if (copy_from_user(&cmd, argp, sizeof(cmd))) > > + return -EFAULT; > > + > > + if (cmd.reserved || cmd.buf.reserved) > > + return -EINVAL; > > [Severity: High] > Does this code also need to validate that the new flags field is zero? > > If uninitialized memory or garbage is passed by userspace in cmd.flags, it > will currently be silently ignored. When a future kernel introduces a new > flag, old userspace applications passing uninitialized flags will inadvertently > enable it, which could permanently break backwards compatibility. Could be zero, or we could check for the known flags. We are currently using the following flags for KVM_MIGRATE_CMD sub-commands for TDX: 1. KVM_MIGRATE_SETUP ==================== TDX_MIGRATE_SETUP_SESSION exchange the migration keys TDX_MIGRATE_IMMUTABLE_STATE transfer the TDX immutable state 2. KVM_MIGRATE_ITERATION ======================== TDX_MIGRATE_IN_ORDER_DONE end the in-order migration 3. KVM_MIGRATE_STOP_AND_COPY ============================ TDX_MIGRATE_STOP_COPY_PAUSE pause the TD TDX_MIGRATE_STOP_COPY_TD_STATE transfer the mutable TD state 4. KVM_MIGRATE_END ================== TDX_MIGRATE_ABORT abort migration Maybe all the above TDX specific flags could be turned into generic flags for the sub-commands. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-08-31 7:13 ` [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD Tony Lindgren 2026-08-31 7:23 ` sashiko-bot @ 2026-09-07 11:53 ` Tony Lindgren 2026-09-07 13:15 ` Jörg Rödel 2026-09-18 4:33 ` Kishen Maloor 2 siblings, 1 reply; 84+ messages in thread From: Tony Lindgren @ 2026-09-07 11:53 UTC (permalink / raw) To: Paolo Bonzini, Sean Christopherson Cc: Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Mon, Aug 31, 2026 at 10:13:02AM +0300, Tony Lindgren wrote: > --- a/include/uapi/linux/kvm.h > +++ b/include/uapi/linux/kvm.h > @@ -1670,4 +1673,23 @@ struct kvm_pre_fault_memory { > __u64 padding[5]; > }; > > +#define KVM_MIGRATE_SETUP 0 > +#define KVM_MIGRATE_ITERATION 1 > +#define KVM_MIGRATE_STOP_AND_COPY 2 > +#define KVM_MIGRATE_ABORT 3 > +#define KVM_MIGRATE_END 4 > + > +struct kvm_transfer_buffer { > + __u64 address; > + __u32 size; > + __u32 reserved; > +}; > + > +struct kvm_migrate_cmd { > + __u16 command; > + __u16 flags; > + __u32 reserved; > + struct kvm_transfer_buffer buf; > +}; For the common flags, KVM_MIGRATE_CMD probably should have migration direction. Or maybe we could have KVM_EXPORT_CMD and KVM_IMPORT_CMD. The migration direction is needed early for TDX. We currently pass a flag for delayed init for the migration destination in KVM_TDX_INIT_VM. This is to prevent the TD and vCPU init SEAMCALLs on the destination. The TD and vCPU are initialized only later on with the import calls. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-07 11:53 ` Tony Lindgren @ 2026-09-07 13:15 ` Jörg Rödel 2026-09-07 13:32 ` Artem Bityutskiy 0 siblings, 1 reply; 84+ messages in thread From: Jörg Rödel @ 2026-09-07 13:15 UTC (permalink / raw) To: Tony Lindgren Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Mon, Sep 07, 2026 at 02:53:05PM +0300, Tony Lindgren wrote: > On Mon, Aug 31, 2026 at 10:13:02AM +0300, Tony Lindgren wrote: > > --- a/include/uapi/linux/kvm.h > > +++ b/include/uapi/linux/kvm.h > > @@ -1670,4 +1673,23 @@ struct kvm_pre_fault_memory { > > __u64 padding[5]; > > }; > > > > +#define KVM_MIGRATE_SETUP 0 > > +#define KVM_MIGRATE_ITERATION 1 > > +#define KVM_MIGRATE_STOP_AND_COPY 2 > > +#define KVM_MIGRATE_ABORT 3 > > +#define KVM_MIGRATE_END 4 > > + > > +struct kvm_transfer_buffer { > > + __u64 address; > > + __u32 size; > > + __u32 reserved; > > +}; > > + > > +struct kvm_migrate_cmd { > > + __u16 command; > > + __u16 flags; > > + __u32 reserved; > > + struct kvm_transfer_buffer buf; > > +}; > > For the common flags, KVM_MIGRATE_CMD probably should have migration > direction. Or maybe we could have KVM_EXPORT_CMD and KVM_IMPORT_CMD. The direction is always the same over a single live migration session, right? So it could be a setup flag, on the other hand having separate KVM_EXPORT_CMD and KVM_IMPORT_CMD seems to be a cleaner ABI. -Joerg ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-07 13:15 ` Jörg Rödel @ 2026-09-07 13:32 ` Artem Bityutskiy 2026-09-08 4:15 ` Tony Lindgren 2026-09-08 4:43 ` Tony Lindgren 0 siblings, 2 replies; 84+ messages in thread From: Artem Bityutskiy @ 2026-09-07 13:32 UTC (permalink / raw) To: Jörg Rödel, Tony Lindgren Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Mon, 2026-09-07 at 15:15 +0200, Jörg Rödel wrote: > The direction is always the same over a single live migration session, right? > So it could be a setup flag, on the other hand having separate KVM_EXPORT_CMD > and KVM_IMPORT_CMD seems to be a cleaner ABI. Yes, the source stays the source, and the destination stays the destination for the entire session. This is a special case of a broader question I keep coming back to: ** How much state should KVM keep about a migration session? ** For direction specifically, we could pass it to KVM on every call and let KVM stay stateless about it, or KVM could record it once and remember it for the rest of the session. But in general. And this is addressed not just to Jörg, but community. Traditional VM migration is driven by QEMU. KVM provides building blocks such as dirty page tracking and vCPU state get/set APIs, but it does not track the overall migration session. The migration session state lives in QEMU. Our TDX live migration prototype keeps some per-migration state in KVM, for example the direction, the migration phase (setup done, started, paused, and so on). This lets use validate inputs and issue the correct TDX module seamcalls from KVM. A different uAPI could shift this balance either way. In the **extreme** case, we could expose a uAPI for each migration seamcall. KVM would then just pass inputs and outputs between the TDX module and QEMU, staying a thin layer with no per-session state. We have not tried this, but it illustrates the trade-off. So where is the right boundary for migration state in KVM? Should KVM manage none of it, or is some state acceptable? Would be interesting to know what KVM community thinks on this. Thanks, Artem. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-07 13:32 ` Artem Bityutskiy @ 2026-09-08 4:15 ` Tony Lindgren 2026-09-08 4:43 ` Tony Lindgren 1 sibling, 0 replies; 84+ messages in thread From: Tony Lindgren @ 2026-09-08 4:15 UTC (permalink / raw) To: Artem Bityutskiy Cc: Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Mon, Sep 07, 2026 at 04:32:33PM +0300, Artem Bityutskiy wrote: > Our TDX live migration prototype keeps some per-migration state in KVM, > for example the direction, the migration phase (setup done, started, > paused, and so on). This lets use validate inputs and issue the correct TDX > module seamcalls from KVM. Yeah for TDX, we currently keep track of some of the TDX module state for migration. For most part it can be done with the existing kvm_tdx->state. The paused state is additional TD_STATE_PAUSED. Some states are trickier though, the setup done state means the TDX module has migration keys configured. If the keys are not installed, the migration SEAMCALLs return errors. Does the kernel need to keep track of this? With proper errors returned, maybe not. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-07 13:32 ` Artem Bityutskiy 2026-09-08 4:15 ` Tony Lindgren @ 2026-09-08 4:43 ` Tony Lindgren 2026-09-09 0:22 ` Kishen Maloor 1 sibling, 1 reply; 84+ messages in thread From: Tony Lindgren @ 2026-09-08 4:43 UTC (permalink / raw) To: Artem Bityutskiy Cc: Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Mon, Sep 07, 2026 at 04:32:33PM +0300, Artem Bityutskiy wrote: > On Mon, 2026-09-07 at 15:15 +0200, Jörg Rödel wrote: > > The direction is always the same over a single live migration session, right? > > So it could be a setup flag, on the other hand having separate KVM_EXPORT_CMD > > and KVM_IMPORT_CMD seems to be a cleaner ABI. OK > Yes, the source stays the source, and the destination stays the destination > for the entire session. > > This is a special case of a broader question I keep coming back to: > > ** How much state should KVM keep about a migration session? ** > > For direction specifically, we could pass it to KVM on every call and let > KVM stay stateless about it, or KVM could record it once and remember it for > the rest of the session. Yes the "record and remember" is another option, it could be a sub-command something like KVM_MIGRATE_DIRECTION. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-08 4:43 ` Tony Lindgren @ 2026-09-09 0:22 ` Kishen Maloor 2026-09-09 6:57 ` Tony Lindgren 0 siblings, 1 reply; 84+ messages in thread From: Kishen Maloor @ 2026-09-09 0:22 UTC (permalink / raw) To: Tony Lindgren, Artem Bityutskiy Cc: Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On 9/7/26 9:43 PM, Tony Lindgren wrote: > On Mon, Sep 07, 2026 at 04:32:33PM +0300, Artem Bityutskiy wrote: >> On Mon, 2026-09-07 at 15:15 +0200, Jörg Rödel wrote: >>> The direction is always the same over a single live migration session, right? >>> So it could be a setup flag, on the other hand having separate KVM_EXPORT_CMD >>> and KVM_IMPORT_CMD seems to be a cleaner ABI. > > OK Just sharing an alternate point of view: The roles are fixed over a migration session. A split ABI is certainly more self-describing, but it restates that invariant on every call, and therefore also permits it to be contradicted -- a failure mode that does not otherwise exist. With a single KVM_MIGRATE_CMD and a per-session role recorded once (more on that below), there is no need for per-call policing: the role could be checked once when the session is established. Along these lines: it raises a question of whether the MEMORY and VCPU calls should be coalesced as well into KVM_MIGRATE_MEMORY and KVM_MIGRATE_VCPU. As posted, direction is implicit for KVM_MIGRATE_CMD but encoded in the ioctl number for those transfers, so collapsing them would at least make the uAPI consistent about where direction comes from, and free two ioctls. > >> ... >> >> For direction specifically, we could pass it to KVM on every call and let >> KVM stay stateless about it, or KVM could record it once and remember it for >> the rest of the session. > > Yes the "record and remember" is another option, it could be a sub-command > something like KVM_MIGRATE_DIRECTION. Agreed on record-and-remember, with a refinement on scope. A VM that was migrated in can later be migrated out, so the role is not a property of the VM -- it has to be recorded per session. SETUP is the call that starts a migration and runs ahead of all other migration calls, so its arguments look like the natural place for userspace to state the role; a separate sub-command would need its own scope rules relative to SETUP. KVM can then ask the vendor layer whether the requested role is permitted for this VM -- only it knows the confidential-VM state -- and on success record it in generic KVM state, where it then selects the export or import callbacks. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-09 0:22 ` Kishen Maloor @ 2026-09-09 6:57 ` Tony Lindgren 2026-09-10 1:11 ` Kishen Maloor 0 siblings, 1 reply; 84+ messages in thread From: Tony Lindgren @ 2026-09-09 6:57 UTC (permalink / raw) To: Kishen Maloor Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Tue, Sep 08, 2026 at 05:22:30PM -0700, Kishen Maloor wrote: > On 9/7/26 9:43 PM, Tony Lindgren wrote: > > On Mon, Sep 07, 2026 at 04:32:33PM +0300, Artem Bityutskiy wrote: > >> On Mon, 2026-09-07 at 15:15 +0200, Jörg Rödel wrote: > >>> The direction is always the same over a single live migration session, right? > >>> So it could be a setup flag, on the other hand having separate KVM_EXPORT_CMD > >>> and KVM_IMPORT_CMD seems to be a cleaner ABI. > > > > OK > > > Just sharing an alternate point of view: > > The roles are fixed over a migration session. A split ABI is certainly more > self-describing, but it restates that invariant on every call, and therefore > also permits it to be contradicted -- a failure mode that does not otherwise > exist. With a single KVM_MIGRATE_CMD and a per-session role recorded once (more > on that below), there is no need for per-call policing: the role could be > checked once when the session is established. There are two occasions the role is set or changed. On starting the destination the incoming role needs to configured at least for TDX. And then after the migration, the role changes if re-migrated. I don't think there are other cases for role change, maybe cancelled migration could require that for some hardware possibly. > Along these lines: it raises a question of whether the MEMORY and VCPU > calls should be coalesced as well into KVM_MIGRATE_MEMORY and KVM_MIGRATE_VCPU. > As posted, direction is implicit for KVM_MIGRATE_CMD but encoded in the ioctl > number for those transfers, so collapsing them would at least make the uAPI > consistent about where direction comes from, and free two ioctls. Using naming KVM_TRANSFER_MEMORY and KVM_TRANSFER_VCPU might be more descriptive? Eventually these same commands could be used to save the state to disk for power management use. And going back to the dmaengine like analogy of what is being done.. The transfer direction flags could be KVM_TRANSFER_FROM_GUEST and KVM_TRANSFER_TO_GUEST? > >> For direction specifically, we could pass it to KVM on every call and let > >> KVM stay stateless about it, or KVM could record it once and remember it for > >> the rest of the session. > > > > Yes the "record and remember" is another option, it could be a sub-command > > something like KVM_MIGRATE_DIRECTION. > > > Agreed on record-and-remember, with a refinement on scope. > > A VM that was migrated in can later be migrated out, so the role is not a > property of the VM -- it has to be recorded per session. SETUP is the call > that starts a migration and runs ahead of all other migration calls, so its > arguments look like the natural place for userspace to state the role; a > separate sub-command would need its own scope rules relative to SETUP. > KVM can then ask the vendor layer whether the requested role is permitted for > this VM -- only it knows the confidential-VM state -- and on success record it > in generic KVM state, where it then selects the export or import callbacks. The role can change, but it can be VM specific for starting the migration destination even before migration is started. At least for TDX we need to specify direction for migration destination on init to prevent fully initializing the TD and vCPUs. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-09 6:57 ` Tony Lindgren @ 2026-09-10 1:11 ` Kishen Maloor 2026-09-10 6:33 ` Tony Lindgren 0 siblings, 1 reply; 84+ messages in thread From: Kishen Maloor @ 2026-09-10 1:11 UTC (permalink / raw) To: Tony Lindgren Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On 9/8/26 11:57 PM, Tony Lindgren wrote: > On Tue, Sep 08, 2026 at 05:22:30PM -0700, Kishen Maloor wrote: >> On 9/7/26 9:43 PM, Tony Lindgren wrote: >>> On Mon, Sep 07, 2026 at 04:32:33PM +0300, Artem Bityutskiy wrote: >>>> On Mon, 2026-09-07 at 15:15 +0200, Jörg Rödel wrote: >>>>> The direction is always the same over a single live migration session, right? >>>>> So it could be a setup flag, on the other hand having separate KVM_EXPORT_CMD >>>>> and KVM_IMPORT_CMD seems to be a cleaner ABI. >>> >>> OK >> >> >> Just sharing an alternate point of view: >> >> The roles are fixed over a migration session. A split ABI is certainly more >> self-describing, but it restates that invariant on every call, and therefore >> also permits it to be contradicted -- a failure mode that does not otherwise >> exist. With a single KVM_MIGRATE_CMD and a per-session role recorded once (more >> on that below), there is no need for per-call policing: the role could be >> checked once when the session is established. > > There are two occasions the role is set or changed. On starting the > destination the incoming role needs to configured at least for TDX. > And then after the migration, the role changes if re-migrated. A destination TD needs a directive to not initialize the TD and its vCPUs. It comes from its launch parameters (e.g. QEMU cmdline) which selects the delayed_init path. That is a construction directive though, and it applies only to a destination -- a source needs nothing at init. A session role is symmetric and is what the transfer calls consume, so I don't think the two need to be the same thing. > I don't think there are other cases for role change, maybe cancelled > migration could require that for some hardware possibly. Architecturally, per-session scoping of migration roles should be straightforward with any platform: each side asserts a role at the start of every migration session and vendor code will either accept or reject the stated role. >> Along these lines: it raises a question of whether the MEMORY and VCPU >> calls should be coalesced as well into KVM_MIGRATE_MEMORY and KVM_MIGRATE_VCPU. >> As posted, direction is implicit for KVM_MIGRATE_CMD but encoded in the ioctl >> number for those transfers, so collapsing them would at least make the uAPI >> consistent about where direction comes from, and free two ioctls. > > Using naming KVM_TRANSFER_MEMORY and KVM_TRANSFER_VCPU might be more > descriptive? > > Eventually these same commands could be used to save the state to disk > for power management use. Sure, and the save-to-disk case is a good argument for a more generic name. > And going back to the dmaengine like analogy of what is being done.. > > The transfer direction flags could be KVM_TRANSFER_FROM_GUEST and > KVM_TRANSFER_TO_GUEST? FROM_GUEST/TO_GUEST still encodes direction per call, which is the open question above. If direction is a per-session property, then KVM_TRANSFER_MEMORY and KVM_TRANSFER_VCPU are sufficient on their own -- no direction flag, and no separate export/import ioctls. And if we settle on a per-session property, then SETUP could conceivably state a role for a non-migration transfer session as well. > >>>> For direction specifically, we could pass it to KVM on every call and let >>>> KVM stay stateless about it, or KVM could record it once and remember it for >>>> the rest of the session. >>> >>> Yes the "record and remember" is another option, it could be a sub-command >>> something like KVM_MIGRATE_DIRECTION. >> >> >> Agreed on record-and-remember, with a refinement on scope. >> >> A VM that was migrated in can later be migrated out, so the role is not a >> property of the VM -- it has to be recorded per session. SETUP is the call >> that starts a migration and runs ahead of all other migration calls, so its >> arguments look like the natural place for userspace to state the role; a >> separate sub-command would need its own scope rules relative to SETUP. >> KVM can then ask the vendor layer whether the requested role is permitted for >> this VM -- only it knows the confidential-VM state -- and on success record it >> in generic KVM state, where it then selects the export or import callbacks. > > The role can change, but it can be VM specific for starting the migration > destination even before migration is started. > > At least for TDX we need to specify direction for migration destination on > init to prevent fully initializing the TD and vCPUs. As mentioned above, that is a destination launch time directive that we needn't conflate with a migration/transfer session role. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-10 1:11 ` Kishen Maloor @ 2026-09-10 6:33 ` Tony Lindgren 2026-09-11 1:40 ` Kishen Maloor 0 siblings, 1 reply; 84+ messages in thread From: Tony Lindgren @ 2026-09-10 6:33 UTC (permalink / raw) To: Kishen Maloor Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Wed, Sep 09, 2026 at 06:11:11PM -0700, Kishen Maloor wrote: > On 9/8/26 11:57 PM, Tony Lindgren wrote: > > On Tue, Sep 08, 2026 at 05:22:30PM -0700, Kishen Maloor wrote: > >> On 9/7/26 9:43 PM, Tony Lindgren wrote: > >>> On Mon, Sep 07, 2026 at 04:32:33PM +0300, Artem Bityutskiy wrote: > >>>> On Mon, 2026-09-07 at 15:15 +0200, Jörg Rödel wrote: > >>>>> The direction is always the same over a single live migration session, right? > >>>>> So it could be a setup flag, on the other hand having separate KVM_EXPORT_CMD > >>>>> and KVM_IMPORT_CMD seems to be a cleaner ABI. > >>> > >>> OK > >> > >> > >> Just sharing an alternate point of view: > >> > >> The roles are fixed over a migration session. A split ABI is certainly more > >> self-describing, but it restates that invariant on every call, and therefore > >> also permits it to be contradicted -- a failure mode that does not otherwise > >> exist. With a single KVM_MIGRATE_CMD and a per-session role recorded once (more > >> on that below), there is no need for per-call policing: the role could be > >> checked once when the session is established. > > > > There are two occasions the role is set or changed. On starting the > > destination the incoming role needs to configured at least for TDX. > > And then after the migration, the role changes if re-migrated. > > A destination TD needs a directive to not initialize the TD and its vCPUs. > It comes from its launch parameters (e.g. QEMU cmdline) which selects the > delayed_init path. That is a construction directive though, and it applies > only to a destination -- a source needs nothing at init. A session role is > symmetric and is what the transfer calls consume, so I don't think the two > need to be the same thing. Yes the source vs destination role is there from the start for sure. And changes on re-migration. Could be set in different ways. > > I don't think there are other cases for role change, maybe cancelled > > migration could require that for some hardware possibly. > > Architecturally, per-session scoping of migration roles should be > straightforward with any platform: each side asserts a role at the start of every > migration session and vendor code will either accept or reject the stated role. Agreed. > >> Along these lines: it raises a question of whether the MEMORY and VCPU > >> calls should be coalesced as well into KVM_MIGRATE_MEMORY and KVM_MIGRATE_VCPU. > >> As posted, direction is implicit for KVM_MIGRATE_CMD but encoded in the ioctl > >> number for those transfers, so collapsing them would at least make the uAPI > >> consistent about where direction comes from, and free two ioctls. > > > > Using naming KVM_TRANSFER_MEMORY and KVM_TRANSFER_VCPU might be more > > descriptive? > > > > Eventually these same commands could be used to save the state to disk > > for power management use. > > Sure, and the save-to-disk case is a good argument for a more generic name. > > > And going back to the dmaengine like analogy of what is being done.. > > > > The transfer direction flags could be KVM_TRANSFER_FROM_GUEST and > > KVM_TRANSFER_TO_GUEST? > > FROM_GUEST/TO_GUEST still encodes direction per call, which is the open > question above. If direction is a per-session property, then KVM_TRANSFER_MEMORY > and KVM_TRANSFER_VCPU are sufficient on their own -- no direction flag, and no > separate export/import ioctls. > > And if we settle on a per-session property, then SETUP could conceivably state a > role for a non-migration transfer session as well. > > > > >>>> For direction specifically, we could pass it to KVM on every call and let > >>>> KVM stay stateless about it, or KVM could record it once and remember it for > >>>> the rest of the session. > >>> > >>> Yes the "record and remember" is another option, it could be a sub-command > >>> something like KVM_MIGRATE_DIRECTION. > >> > >> > >> Agreed on record-and-remember, with a refinement on scope. > >> > >> A VM that was migrated in can later be migrated out, so the role is not a > >> property of the VM -- it has to be recorded per session. SETUP is the call > >> that starts a migration and runs ahead of all other migration calls, so its > >> arguments look like the natural place for userspace to state the role; a > >> separate sub-command would need its own scope rules relative to SETUP. > >> KVM can then ask the vendor layer whether the requested role is permitted for > >> this VM -- only it knows the confidential-VM state -- and on success record it > >> in generic KVM state, where it then selects the export or import callbacks. > > > > The role can change, but it can be VM specific for starting the migration > > destination even before migration is started. > > > > At least for TDX we need to specify direction for migration destination on > > init to prevent fully initializing the TD and vCPUs. > > As mentioned above, that is a destination launch time directive that we needn't > conflate with a migration/transfer session role. It's still the same role though. Yes we can set it on init, but would be nice to have some generic way to do it for qemu -incoming. Just brainstorming.. I wonder if we need two things though. A source vs destination role. And then at some point possibly later on also a data transfer direction enumeration similar to what Linux has in include/linux/dma-direction.h. We already need to make use of the QEMU return-path for the migration key exchange. What if some hardware needs to make use of KVM_TRANSFER_MEMORY from source to destination, and after that back from destination to source to ack the transfer? Sure this is just speculation, I'm not aware of this need right now. In any case with handling the source vs destination role, the enumeration for data direction can be added to the transfer flags later on as needed. No need to try to stuff the data direction flag there until really needed. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-10 6:33 ` Tony Lindgren @ 2026-09-11 1:40 ` Kishen Maloor 2026-09-11 4:23 ` Tony Lindgren 0 siblings, 1 reply; 84+ messages in thread From: Kishen Maloor @ 2026-09-11 1:40 UTC (permalink / raw) To: Tony Lindgren Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On 9/9/26 11:33 PM, Tony Lindgren wrote: > On Wed, Sep 09, 2026 at 06:11:11PM -0700, Kishen Maloor wrote: > ... >> >> As mentioned above, that is a destination launch time directive that we needn't >> conflate with a migration/transfer session role. > > It's still the same role though. Yes we can set it on init, but would > be nice to have some generic way to do it for qemu -incoming. I'd separate these. On the role: it seems we agree on recording roles per-session. My only point then is that a VM created through the delayed_init flow doesn't additionally need a destination role recorded for it if SETUP will assert one when the migration is kicked off. On generic plumbing for -incoming: is this about KVM_TDX_INIT_VM_F_DELAY_INIT? A generic mechanism would make sense to me if the flag were consumed by generic KVM code, but that isn't the case here. Userspace has to make a vendor-specific VM-init call like KVM_TDX_INIT_VM anyway, with the flag passed on that call. So I'm not sure what a generic version would add, unless you have something else in mind. > Just brainstorming.. I wonder if we need two things though. A source vs > destination role. And then at some point possibly later on also a data > transfer direction enumeration similar to what Linux has in > include/linux/dma-direction.h. > > We already need to make use of the QEMU return-path for the migration key > exchange. What if some hardware needs to make use of KVM_TRANSFER_MEMORY > from source to destination, and after that back from destination to source > to ack the transfer? Sure this is just speculation, I'm not aware of this > need right now. > > In any case with handling the source vs destination role, the enumeration > for data direction can be added to the transfer flags later on as needed. > No need to try to stuff the data direction flag there until really needed. Agree on deferring such an enumeration. More generally though, the session role determines which operations are permitted, and it seems like vendor code could be expected to handle those in context. The TDX architecture already has a destination-to-source example: TDH.IMPORT.ABORT emits an abort token on the destination that TDH.EXPORT.ABORT consumes on the source. So KVM on the source would dispatch to its export-abort path and let the vendor layer decide what to do with the buffer. This doesn't require a direction flag. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-11 1:40 ` Kishen Maloor @ 2026-09-11 4:23 ` Tony Lindgren 2026-09-15 0:14 ` Kishen Maloor 0 siblings, 1 reply; 84+ messages in thread From: Tony Lindgren @ 2026-09-11 4:23 UTC (permalink / raw) To: Kishen Maloor Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Thu, Sep 10, 2026 at 06:40:04PM -0700, Kishen Maloor wrote: > On 9/9/26 11:33 PM, Tony Lindgren wrote: > > On Wed, Sep 09, 2026 at 06:11:11PM -0700, Kishen Maloor wrote: > > ... > >> > >> As mentioned above, that is a destination launch time directive that we needn't > >> conflate with a migration/transfer session role. > > > > It's still the same role though. Yes we can set it on init, but would > > be nice to have some generic way to do it for qemu -incoming. > > I'd separate these. > > On the role: it seems we agree on recording roles per-session. My only point > then is that a VM created through the delayed_init flow doesn't additionally > need a destination role recorded for it if SETUP will assert one when the > migration is kicked off. > > On generic plumbing for -incoming: is this about KVM_TDX_INIT_VM_F_DELAY_INIT? > A generic mechanism would make sense to me if the flag were consumed by generic > KVM code, but that isn't the case here. Userspace has to make a > vendor-specific VM-init call like KVM_TDX_INIT_VM anyway, with the flag passed > on that call. So I'm not sure what a generic version would add, unless you have > something else in mind. So we could add a SETUP subcommand SET_ROLE or SET_INCOMING. The implementation could store the role at least initially. And if we want to set the role with KVM_TDX_INIT_VM, we could recycle the role bit there. > > Just brainstorming.. I wonder if we need two things though. A source vs > > destination role. And then at some point possibly later on also a data > > transfer direction enumeration similar to what Linux has in > > include/linux/dma-direction.h. > > > > We already need to make use of the QEMU return-path for the migration key > > exchange. What if some hardware needs to make use of KVM_TRANSFER_MEMORY > > from source to destination, and after that back from destination to source > > to ack the transfer? Sure this is just speculation, I'm not aware of this > > need right now. > > > > In any case with handling the source vs destination role, the enumeration > > for data direction can be added to the transfer flags later on as needed. > > No need to try to stuff the data direction flag there until really needed. > > Agree on deferring such an enumeration. More generally though, the session > role determines which operations are permitted, and it seems like vendor code > could be expected to handle those in context. > > The TDX architecture already has a destination-to-source example: > TDH.IMPORT.ABORT emits an abort token on the destination that > TDH.EXPORT.ABORT consumes on the source. So KVM on the source would dispatch > to its export-abort path and let the vendor layer decide what to do with > the buffer. This doesn't require a direction flag. Yes agreed with the role we can handle migration related transfers both ways at least for TDX. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-11 4:23 ` Tony Lindgren @ 2026-09-15 0:14 ` Kishen Maloor 2026-09-15 4:44 ` Tony Lindgren 0 siblings, 1 reply; 84+ messages in thread From: Kishen Maloor @ 2026-09-15 0:14 UTC (permalink / raw) To: Tony Lindgren Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On 9/10/26 9:23 PM, Tony Lindgren wrote: > On Thu, Sep 10, 2026 at 06:40:04PM -0700, Kishen Maloor wrote: >> On 9/9/26 11:33 PM, Tony Lindgren wrote: >>> On Wed, Sep 09, 2026 at 06:11:11PM -0700, Kishen Maloor wrote: >>> ... >>>> >>>> As mentioned above, that is a destination launch time directive that we needn't >>>> conflate with a migration/transfer session role. >>> >>> It's still the same role though. Yes we can set it on init, but would >>> be nice to have some generic way to do it for qemu -incoming. >> >> I'd separate these. >> >> On the role: it seems we agree on recording roles per-session. My only point >> then is that a VM created through the delayed_init flow doesn't additionally >> need a destination role recorded for it if SETUP will assert one when the >> migration is kicked off. >> >> On generic plumbing for -incoming: is this about KVM_TDX_INIT_VM_F_DELAY_INIT? >> A generic mechanism would make sense to me if the flag were consumed by generic >> KVM code, but that isn't the case here. Userspace has to make a >> vendor-specific VM-init call like KVM_TDX_INIT_VM anyway, with the flag passed >> on that call. So I'm not sure what a generic version would add, unless you have >> something else in mind. > > So we could add a SETUP subcommand SET_ROLE or SET_INCOMING. It might be better to pass the role as a parameter of the SETUP call rather than a separate SET_ROLE call. A separate SET_ROLE would bring its own ordering rules relative to the other SETUP sub-commands. The specific role (src or dst) still needs someplace to go, and flags is already the sub-command selector. An option is to carve out room in the __u32 reserved field to carry a role argument, something like 0=unset, 1=src, 2=dest so KVM can verify that a role was indeed set. It could further be written into a generic KVM struct (kvm_arch or kvm) which could be queried on KVM_TRANSFER_MEMORY, etc. to identify the relevant callback. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-15 0:14 ` Kishen Maloor @ 2026-09-15 4:44 ` Tony Lindgren 2026-09-15 15:53 ` Kishen Maloor 0 siblings, 1 reply; 84+ messages in thread From: Tony Lindgren @ 2026-09-15 4:44 UTC (permalink / raw) To: Kishen Maloor Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Mon, Sep 14, 2026 at 05:14:32PM -0700, Kishen Maloor wrote: > On 9/10/26 9:23 PM, Tony Lindgren wrote: > > On Thu, Sep 10, 2026 at 06:40:04PM -0700, Kishen Maloor wrote: > >> On 9/9/26 11:33 PM, Tony Lindgren wrote: > >>> On Wed, Sep 09, 2026 at 06:11:11PM -0700, Kishen Maloor wrote: > >>> ... > >>>> > >>>> As mentioned above, that is a destination launch time directive that we needn't > >>>> conflate with a migration/transfer session role. > >>> > >>> It's still the same role though. Yes we can set it on init, but would > >>> be nice to have some generic way to do it for qemu -incoming. > >> > >> I'd separate these. > >> > >> On the role: it seems we agree on recording roles per-session. My only point > >> then is that a VM created through the delayed_init flow doesn't additionally > >> need a destination role recorded for it if SETUP will assert one when the > >> migration is kicked off. > >> > >> On generic plumbing for -incoming: is this about KVM_TDX_INIT_VM_F_DELAY_INIT? > >> A generic mechanism would make sense to me if the flag were consumed by generic > >> KVM code, but that isn't the case here. Userspace has to make a > >> vendor-specific VM-init call like KVM_TDX_INIT_VM anyway, with the flag passed > >> on that call. So I'm not sure what a generic version would add, unless you have > >> something else in mind. > > > > So we could add a SETUP subcommand SET_ROLE or SET_INCOMING. > It might be better to pass the role as a parameter of the SETUP call > rather than a separate SET_ROLE call. > A separate SET_ROLE would bring its own ordering rules relative to the > other SETUP sub-commands. OK a flag for SETUP sounds good to me. SETUP is needed anyways for each migration session. > The specific role (src or dst) still needs someplace to go, and flags is > already the sub-command selector. An option is to carve out room in > the __u32 reserved field to carry a role argument, something > like 0=unset, 1=src, 2=dest so KVM can verify that a role was indeed set. To me it seems that 0=src can be the natural default starting point, I don't think we need 0=unset. > It could further be written into a generic KVM struct (kvm_arch or kvm) which > could be queried on KVM_TRANSFER_MEMORY, etc. to identify the relevant > callback. Yeah eventually some generic place for it would be nice. But that's easy to add later on too. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-15 4:44 ` Tony Lindgren @ 2026-09-15 15:53 ` Kishen Maloor 2026-09-16 5:09 ` Tony Lindgren 0 siblings, 1 reply; 84+ messages in thread From: Kishen Maloor @ 2026-09-15 15:53 UTC (permalink / raw) To: Tony Lindgren Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On 9/14/26 9:44 PM, Tony Lindgren wrote: > On Mon, Sep 14, 2026 at 05:14:32PM -0700, Kishen Maloor wrote: >> On 9/10/26 9:23 PM, Tony Lindgren wrote: >>> On Thu, Sep 10, 2026 at 06:40:04PM -0700, Kishen Maloor wrote: >>>> On 9/9/26 11:33 PM, Tony Lindgren wrote: >>>>> On Wed, Sep 09, 2026 at 06:11:11PM -0700, Kishen Maloor wrote: >>>>> ... >>>>>> >>>>>> As mentioned above, that is a destination launch time directive that we needn't >>>>>> conflate with a migration/transfer session role. >>>>> >>>>> It's still the same role though. Yes we can set it on init, but would >>>>> be nice to have some generic way to do it for qemu -incoming. >>>> >>>> I'd separate these. >>>> >>>> On the role: it seems we agree on recording roles per-session. My only point >>>> then is that a VM created through the delayed_init flow doesn't additionally >>>> need a destination role recorded for it if SETUP will assert one when the >>>> migration is kicked off. >>>> >>>> On generic plumbing for -incoming: is this about KVM_TDX_INIT_VM_F_DELAY_INIT? >>>> A generic mechanism would make sense to me if the flag were consumed by generic >>>> KVM code, but that isn't the case here. Userspace has to make a >>>> vendor-specific VM-init call like KVM_TDX_INIT_VM anyway, with the flag passed >>>> on that call. So I'm not sure what a generic version would add, unless you have >>>> something else in mind. >>> >>> So we could add a SETUP subcommand SET_ROLE or SET_INCOMING. >> It might be better to pass the role as a parameter of the SETUP call >> rather than a separate SET_ROLE call. >> A separate SET_ROLE would bring its own ordering rules relative to the >> other SETUP sub-commands. > > OK a flag for SETUP sounds good to me. SETUP is needed anyways for each > migration session. To be clear, I was suggesting a field in kvm_migrate_cmd and not a flag to pass the role as an argument to SETUP. As I mentioned in my last comments (right below), 'flags' in the current proposal carry the vendor-defined sub-command values, so a generic role argument wouldn't belong there. > >> The specific role (src or dst) still needs someplace to go, and flags is >> already the sub-command selector. An option is to carve out room in >> the __u32 reserved field to carry a role argument, something >> like 0=unset, 1=src, 2=dest so KVM can verify that a role was indeed set. > > To me it seems that 0=src can be the natural default starting point, I > don't think we need 0=unset. With 0=src, the field would only carry information when it's a destination, thereby making it an is_dest boolean rather than a role. It would also mean any VM that never established a session still reads as a source, so a KVM_TRANSFER_MEMORY aimed at the wrong VM by buggy or rogue userspace would get dispatched to the export path instead of rejected outright. The generic dispatcher shouldn't have to rely on the vendor layer to catch that. Reserving 0 for 'unset' costs nothing and lets KVM reject a session that never stated a role. > >> It could further be written into a generic KVM struct (kvm_arch or kvm) which >> could be queried on KVM_TRANSFER_MEMORY, etc. to identify the relevant >> callback. > > Yeah eventually some generic place for it would be nice. But that's easy > to add later on too. Sure, we don't have to decide now as we're still discussing the UAPI. But we'd want this detail also settled sometime before we call the UAPI complete as it determines whether the dispatch is generic. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-15 15:53 ` Kishen Maloor @ 2026-09-16 5:09 ` Tony Lindgren 2026-09-17 3:31 ` Kishen Maloor 0 siblings, 1 reply; 84+ messages in thread From: Tony Lindgren @ 2026-09-16 5:09 UTC (permalink / raw) To: Kishen Maloor Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Tue, Sep 15, 2026 at 08:53:18AM -0700, Kishen Maloor wrote: > On 9/14/26 9:44 PM, Tony Lindgren wrote: > > On Mon, Sep 14, 2026 at 05:14:32PM -0700, Kishen Maloor wrote: > >> On 9/10/26 9:23 PM, Tony Lindgren wrote: > >>> On Thu, Sep 10, 2026 at 06:40:04PM -0700, Kishen Maloor wrote: > >>>> On 9/9/26 11:33 PM, Tony Lindgren wrote: > >>>>> On Wed, Sep 09, 2026 at 06:11:11PM -0700, Kishen Maloor wrote: > >>>>> ... > >>>>>> > >>>>>> As mentioned above, that is a destination launch time directive that we needn't > >>>>>> conflate with a migration/transfer session role. > >>>>> > >>>>> It's still the same role though. Yes we can set it on init, but would > >>>>> be nice to have some generic way to do it for qemu -incoming. > >>>> > >>>> I'd separate these. > >>>> > >>>> On the role: it seems we agree on recording roles per-session. My only point > >>>> then is that a VM created through the delayed_init flow doesn't additionally > >>>> need a destination role recorded for it if SETUP will assert one when the > >>>> migration is kicked off. > >>>> > >>>> On generic plumbing for -incoming: is this about KVM_TDX_INIT_VM_F_DELAY_INIT? > >>>> A generic mechanism would make sense to me if the flag were consumed by generic > >>>> KVM code, but that isn't the case here. Userspace has to make a > >>>> vendor-specific VM-init call like KVM_TDX_INIT_VM anyway, with the flag passed > >>>> on that call. So I'm not sure what a generic version would add, unless you have > >>>> something else in mind. > >>> > >>> So we could add a SETUP subcommand SET_ROLE or SET_INCOMING. > >> It might be better to pass the role as a parameter of the SETUP call > >> rather than a separate SET_ROLE call. > >> A separate SET_ROLE would bring its own ordering rules relative to the > >> other SETUP sub-commands. > > > > OK a flag for SETUP sounds good to me. SETUP is needed anyways for each > > migration session. > > To be clear, I was suggesting a field in kvm_migrate_cmd and not a flag to pass > the role as an argument to SETUP. As I mentioned in my last comments (right > below), 'flags' in the current proposal carry the vendor-defined sub-command values, > so a generic role argument wouldn't belong there. I was thinking 8 bits for common flags and 8 bits for vendor flags but yeah that can be a bit tight. Sorry if the SETUP above caused extra confusion. So trying to summarize the common flags for the role and separate vendor flags: struct kvm_migrate_cmd { __u16 command; __u16 flags; __u16 vflags; __u16 reserved; __u32 reserved; struct kvm_transfer_buffer buf; }; Is the above along the lines what you were thinking? > >> The specific role (src or dst) still needs someplace to go, and flags is > >> already the sub-command selector. An option is to carve out room in > >> the __u32 reserved field to carry a role argument, something > >> like 0=unset, 1=src, 2=dest so KVM can verify that a role was indeed set. > > > > To me it seems that 0=src can be the natural default starting point, I > > don't think we need 0=unset. > > With 0=src, the field would only carry information when it's a destination, > thereby making it an is_dest boolean rather than a role. It would also mean any > VM that never established a session still reads as a source, so a > KVM_TRANSFER_MEMORY aimed at the wrong VM by buggy or rogue userspace would get > dispatched to the export path instead of rejected outright. The generic > dispatcher shouldn't have to rely on the vendor layer to catch that. Reserving 0 > for 'unset' costs nothing and lets KVM reject a session that never stated a > role. Ah OK, yes that would also tell "the hardware has been initialized to a certain migration role". That seems like a usable common feature. > >> It could further be written into a generic KVM struct (kvm_arch or kvm) which > >> could be queried on KVM_TRANSFER_MEMORY, etc. to identify the relevant > >> callback. > > > > Yeah eventually some generic place for it would be nice. But that's easy > > to add later on too. > > Sure, we don't have to decide now as we're still discussing the UAPI. > But we'd want this detail also settled sometime before we call the UAPI > complete as it determines whether the dispatch is generic. Yes a shared place for the role would make some generic sanity checks easier. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-16 5:09 ` Tony Lindgren @ 2026-09-17 3:31 ` Kishen Maloor 2026-09-17 6:42 ` Tony Lindgren 0 siblings, 1 reply; 84+ messages in thread From: Kishen Maloor @ 2026-09-17 3:31 UTC (permalink / raw) To: Tony Lindgren Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On 9/15/26 10:09 PM, Tony Lindgren wrote: > On Tue, Sep 15, 2026 at 08:53:18AM -0700, Kishen Maloor wrote: >> On 9/14/26 9:44 PM, Tony Lindgren wrote: >>> On Mon, Sep 14, 2026 at 05:14:32PM -0700, Kishen Maloor wrote: >>>> On 9/10/26 9:23 PM, Tony Lindgren wrote: >>>>> On Thu, Sep 10, 2026 at 06:40:04PM -0700, Kishen Maloor wrote: >>>>>> On 9/9/26 11:33 PM, Tony Lindgren wrote: >>>>>>> On Wed, Sep 09, 2026 at 06:11:11PM -0700, Kishen Maloor wrote: >>>>>>> ... >>>>>>>> >>>>>>>> As mentioned above, that is a destination launch time directive that we needn't >>>>>>>> conflate with a migration/transfer session role. >>>>>>> >>>>>>> It's still the same role though. Yes we can set it on init, but would >>>>>>> be nice to have some generic way to do it for qemu -incoming. >>>>>> >>>>>> I'd separate these. >>>>>> >>>>>> On the role: it seems we agree on recording roles per-session. My only point >>>>>> then is that a VM created through the delayed_init flow doesn't additionally >>>>>> need a destination role recorded for it if SETUP will assert one when the >>>>>> migration is kicked off. >>>>>> >>>>>> On generic plumbing for -incoming: is this about KVM_TDX_INIT_VM_F_DELAY_INIT? >>>>>> A generic mechanism would make sense to me if the flag were consumed by generic >>>>>> KVM code, but that isn't the case here. Userspace has to make a >>>>>> vendor-specific VM-init call like KVM_TDX_INIT_VM anyway, with the flag passed >>>>>> on that call. So I'm not sure what a generic version would add, unless you have >>>>>> something else in mind. >>>>> >>>>> So we could add a SETUP subcommand SET_ROLE or SET_INCOMING. >>>> It might be better to pass the role as a parameter of the SETUP call >>>> rather than a separate SET_ROLE call. >>>> A separate SET_ROLE would bring its own ordering rules relative to the >>>> other SETUP sub-commands. >>> >>> OK a flag for SETUP sounds good to me. SETUP is needed anyways for each >>> migration session. >> >> To be clear, I was suggesting a field in kvm_migrate_cmd and not a flag to pass >> the role as an argument to SETUP. As I mentioned in my last comments (right >> below), 'flags' in the current proposal carry the vendor-defined sub-command values, >> so a generic role argument wouldn't belong there. > > I was thinking 8 bits for common flags and 8 bits for vendor flags but > yeah that can be a bit tight. Sorry if the SETUP above caused extra > confusion. > > So trying to summarize the common flags for the role and separate vendor > flags: > > struct kvm_migrate_cmd { > __u16 command; > __u16 flags; > __u16 vflags; > __u16 reserved; > __u32 reserved; > struct kvm_transfer_buffer buf; > }; > > Is the above along the lines what you were thinking? No, I was suggesting a 'role' field carved out of the 'reserved' space, like this: struct kvm_migrate_cmd { __u16 command; __u16 flags; __u8 role; /* 0 = unset, 1 = source, 2 = destination */ __u8 reserved[3]; struct kvm_transfer_buffer buf; }; We haven't defined any generic flags. Thus far in this proposal 'flags' contains only vendor-defined sub-command values. > >>>> The specific role (src or dst) still needs someplace to go, and flags is >>>> already the sub-command selector. An option is to carve out room in >>>> the __u32 reserved field to carry a role argument, something >>>> like 0=unset, 1=src, 2=dest so KVM can verify that a role was indeed set. >>> >>> To me it seems that 0=src can be the natural default starting point, I >>> don't think we need 0=unset. >> >> With 0=src, the field would only carry information when it's a destination, >> thereby making it an is_dest boolean rather than a role. It would also mean any >> VM that never established a session still reads as a source, so a >> KVM_TRANSFER_MEMORY aimed at the wrong VM by buggy or rogue userspace would get >> dispatched to the export path instead of rejected outright. The generic >> dispatcher shouldn't have to rely on the vendor layer to catch that. Reserving 0 >> for 'unset' costs nothing and lets KVM reject a session that never stated a >> role. > > Ah OK, yes that would also tell "the hardware has been initialized to a > certain migration role". That seems like a usable common feature. Not quite. It tells us that userspace asserted a role for this VM's migration session. Whether a TD was created for import is a separate, vendor-level detail. The generic layer only needs the role to reject a session that never stated one, and to pick the export or import callback. That callback then knows which side it's on and can reject an incorrect role (e.g., if SETUP asserted dst for a src TD). > >>>> It could further be written into a generic KVM struct (kvm_arch or kvm) which >>>> could be queried on KVM_TRANSFER_MEMORY, etc. to identify the relevant >>>> callback. >>> >>> Yeah eventually some generic place for it would be nice. But that's easy >>> to add later on too. >> >> Sure, we don't have to decide now as we're still discussing the UAPI. >> But we'd want this detail also settled sometime before we call the UAPI >> complete as it determines whether the dispatch is generic. > > Yes a shared place for the role would make some generic sanity checks > easier. Agreed. The field above is just an argument to SETUP. Where that role gets stored for KVM's top-level dispatcher to check is the detail we can settle later. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-17 3:31 ` Kishen Maloor @ 2026-09-17 6:42 ` Tony Lindgren 2026-09-18 4:32 ` Kishen Maloor 0 siblings, 1 reply; 84+ messages in thread From: Tony Lindgren @ 2026-09-17 6:42 UTC (permalink / raw) To: Kishen Maloor Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Wed, Sep 16, 2026 at 08:31:32PM -0700, Kishen Maloor wrote: > On 9/15/26 10:09 PM, Tony Lindgren wrote: > > So trying to summarize the common flags for the role and separate vendor > > flags: > > > > struct kvm_migrate_cmd { > > __u16 command; > > __u16 flags; > > __u16 vflags; > > __u16 reserved; > > __u32 reserved; > > struct kvm_transfer_buffer buf; > > }; > > > > Is the above along the lines what you were thinking? > > No, I was suggesting a 'role' field carved out of the 'reserved' space, > like this: > > struct kvm_migrate_cmd { > __u16 command; > __u16 flags; > __u8 role; /* 0 = unset, 1 = source, 2 = destination */ > __u8 reserved[3]; > struct kvm_transfer_buffer buf; > }; OK yes thanks for clarifying, that works for me. > > Ah OK, yes that would also tell "the hardware has been initialized to a > > certain migration role". That seems like a usable common feature. > > Not quite. It tells us that userspace asserted a role for this VM's migration > session. Whether a TD was created for import is a separate, vendor-level detail. > The generic layer only needs the role to reject a session that never stated one, > and to pick the export or import callback. That callback then knows which side > it's on and can reject an incorrect role (e.g., if SETUP asserted dst for a src TD). That's a good point, the hardware role may not be set yet. I'm still wondering if there is a need to stash the userspace set role in KVM though. Likely only the hardware specific code can properly track the state of the hardware and adjust to the userspace requests. Seems just being able to pass the role in struct kvm_migrate_cmd should be enough? ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-17 6:42 ` Tony Lindgren @ 2026-09-18 4:32 ` Kishen Maloor 2026-09-18 5:58 ` Tony Lindgren 0 siblings, 1 reply; 84+ messages in thread From: Kishen Maloor @ 2026-09-18 4:32 UTC (permalink / raw) To: Tony Lindgren Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On 9/16/26 11:42 PM, Tony Lindgren wrote: > On Wed, Sep 16, 2026 at 08:31:32PM -0700, Kishen Maloor wrote: >> On 9/15/26 10:09 PM, Tony Lindgren wrote: >>> So trying to summarize the common flags for the role and separate vendor >>> flags: >>> >>> struct kvm_migrate_cmd { >>> __u16 command; >>> __u16 flags; >>> __u16 vflags; >>> __u16 reserved; >>> __u32 reserved; >>> struct kvm_transfer_buffer buf; >>> }; >>> >>> Is the above along the lines what you were thinking? >> >> No, I was suggesting a 'role' field carved out of the 'reserved' space, >> like this: >> >> struct kvm_migrate_cmd { >> __u16 command; >> __u16 flags; >> __u8 role; /* 0 = unset, 1 = source, 2 = destination */ >> __u8 reserved[3]; >> struct kvm_transfer_buffer buf; >> }; > > OK yes thanks for clarifying, that works for me. > >>> Ah OK, yes that would also tell "the hardware has been initialized to a >>> certain migration role". That seems like a usable common feature. >> >> Not quite. It tells us that userspace asserted a role for this VM's migration >> session. Whether a TD was created for import is a separate, vendor-level detail. >> The generic layer only needs the role to reject a session that never stated one, >> and to pick the export or import callback. That callback then knows which side >> it's on and can reject an incorrect role (e.g., if SETUP asserted dst for a src TD). > > That's a good point, the hardware role may not be set yet. > > I'm still wondering if there is a need to stash the userspace set role in > KVM though. Likely only the hardware specific code can properly track the > state of the hardware and adjust to the userspace requests. Seems just > being able to pass the role in struct kvm_migrate_cmd should be enough? Passing it in kvm_migrate_cmd is enough for SETUP itself, but the commands after SETUP like memory/vcpu transfers still have to reach the right callback. So KVM would need to remember what was asserted so that the generic layer can dispatch to the export or import facing callbacks. We've been sketching (on this thread) an alternative UAPI set (3 vs 5 ioctls) for consideration which this stored role enables: Proposed in the RFC Alternative KVM_MIGRATE_CMD KVM_MIGRATE_CMD KVM_EXPORT_MEMORY KVM_IMPORT_MEMORY KVM_TRANSFER_MEMORY KVM_EXPORT_VCPU KVM_IMPORT_VCPU KVM_TRANSFER_VCPU It's just a record (1 byte) of what userspace asserted for the current session at SETUP. Vendor code still owns the hardware state and remains free to reject a role that doesn't match it. It is also what lets the generic layer reject a command to a VM that never set up a session. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-18 4:32 ` Kishen Maloor @ 2026-09-18 5:58 ` Tony Lindgren 2026-09-21 0:13 ` Kishen Maloor 0 siblings, 1 reply; 84+ messages in thread From: Tony Lindgren @ 2026-09-18 5:58 UTC (permalink / raw) To: Kishen Maloor Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Thu, Sep 17, 2026 at 09:32:23PM -0700, Kishen Maloor wrote: > On 9/16/26 11:42 PM, Tony Lindgren wrote: > > On Wed, Sep 16, 2026 at 08:31:32PM -0700, Kishen Maloor wrote: > >> On 9/15/26 10:09 PM, Tony Lindgren wrote: > >>> So trying to summarize the common flags for the role and separate vendor > >>> flags: > >>> > >>> struct kvm_migrate_cmd { > >>> __u16 command; > >>> __u16 flags; > >>> __u16 vflags; > >>> __u16 reserved; > >>> __u32 reserved; > >>> struct kvm_transfer_buffer buf; > >>> }; > >>> > >>> Is the above along the lines what you were thinking? > >> > >> No, I was suggesting a 'role' field carved out of the 'reserved' space, > >> like this: > >> > >> struct kvm_migrate_cmd { > >> __u16 command; > >> __u16 flags; > >> __u8 role; /* 0 = unset, 1 = source, 2 = destination */ > >> __u8 reserved[3]; > >> struct kvm_transfer_buffer buf; > >> }; > > > > OK yes thanks for clarifying, that works for me. > > > >>> Ah OK, yes that would also tell "the hardware has been initialized to a > >>> certain migration role". That seems like a usable common feature. > >> > >> Not quite. It tells us that userspace asserted a role for this VM's migration > >> session. Whether a TD was created for import is a separate, vendor-level detail. > >> The generic layer only needs the role to reject a session that never stated one, > >> and to pick the export or import callback. That callback then knows which side > >> it's on and can reject an incorrect role (e.g., if SETUP asserted dst for a src TD). > > > > That's a good point, the hardware role may not be set yet. > > > > I'm still wondering if there is a need to stash the userspace set role in > > KVM though. Likely only the hardware specific code can properly track the > > state of the hardware and adjust to the userspace requests. Seems just > > being able to pass the role in struct kvm_migrate_cmd should be enough? > > Passing it in kvm_migrate_cmd is enough for SETUP itself, but the commands > after SETUP like memory/vcpu transfers still have to reach the right > callback. So KVM would need to remember what was asserted so that the > generic layer can dispatch to the export or import facing callbacks. > We've been sketching (on this thread) an alternative UAPI set > (3 vs 5 ioctls) for consideration which this stored role enables: > > Proposed in the RFC Alternative > KVM_MIGRATE_CMD KVM_MIGRATE_CMD > KVM_EXPORT_MEMORY > KVM_IMPORT_MEMORY KVM_TRANSFER_MEMORY > KVM_EXPORT_VCPU > KVM_IMPORT_VCPU KVM_TRANSFER_VCPU > > It's just a record (1 byte) of what userspace asserted for the current session > at SETUP. Vendor code still owns the hardware state and remains free to reject a > role that doesn't match it. It is also what lets the generic layer reject a > command to a VM that never set up a session. For the TRANSFER style operations, I would assume the direction is passed for each transfer, just like the Linux does for the dmaengine. It's possible that there may be transfers going both directions without the role changing. So looks like we have tree things to consider: userspace set migration role, the hardware state, and transfer direction. What if userspace always passes the role and transfer direction where it makes sense? And then the hardware specific implementation tracks the hardware state? ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-18 5:58 ` Tony Lindgren @ 2026-09-21 0:13 ` Kishen Maloor 2026-09-21 6:52 ` Tony Lindgren 0 siblings, 1 reply; 84+ messages in thread From: Kishen Maloor @ 2026-09-21 0:13 UTC (permalink / raw) To: Tony Lindgren Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On 9/17/26 10:58 PM, Tony Lindgren wrote: > On Thu, Sep 17, 2026 at 09:32:23PM -0700, Kishen Maloor wrote: >> On 9/16/26 11:42 PM, Tony Lindgren wrote: >>> On Wed, Sep 16, 2026 at 08:31:32PM -0700, Kishen Maloor wrote: >>>> On 9/15/26 10:09 PM, Tony Lindgren wrote: >>>>> So trying to summarize the common flags for the role and separate vendor >>>>> flags: >>>>> >>>>> struct kvm_migrate_cmd { >>>>> __u16 command; >>>>> __u16 flags; >>>>> __u16 vflags; >>>>> __u16 reserved; >>>>> __u32 reserved; >>>>> struct kvm_transfer_buffer buf; >>>>> }; >>>>> >>>>> Is the above along the lines what you were thinking? >>>> >>>> No, I was suggesting a 'role' field carved out of the 'reserved' space, >>>> like this: >>>> >>>> struct kvm_migrate_cmd { >>>> __u16 command; >>>> __u16 flags; >>>> __u8 role; /* 0 = unset, 1 = source, 2 = destination */ >>>> __u8 reserved[3]; >>>> struct kvm_transfer_buffer buf; >>>> }; >>> >>> OK yes thanks for clarifying, that works for me. >>> >>>>> Ah OK, yes that would also tell "the hardware has been initialized to a >>>>> certain migration role". That seems like a usable common feature. >>>> >>>> Not quite. It tells us that userspace asserted a role for this VM's migration >>>> session. Whether a TD was created for import is a separate, vendor-level detail. >>>> The generic layer only needs the role to reject a session that never stated one, >>>> and to pick the export or import callback. That callback then knows which side >>>> it's on and can reject an incorrect role (e.g., if SETUP asserted dst for a src TD). >>> >>> That's a good point, the hardware role may not be set yet. >>> >>> I'm still wondering if there is a need to stash the userspace set role in >>> KVM though. Likely only the hardware specific code can properly track the >>> state of the hardware and adjust to the userspace requests. Seems just >>> being able to pass the role in struct kvm_migrate_cmd should be enough? >> >> Passing it in kvm_migrate_cmd is enough for SETUP itself, but the commands >> after SETUP like memory/vcpu transfers still have to reach the right >> callback. So KVM would need to remember what was asserted so that the >> generic layer can dispatch to the export or import facing callbacks. >> We've been sketching (on this thread) an alternative UAPI set >> (3 vs 5 ioctls) for consideration which this stored role enables: >> >> Proposed in the RFC Alternative >> KVM_MIGRATE_CMD KVM_MIGRATE_CMD >> KVM_EXPORT_MEMORY >> KVM_IMPORT_MEMORY KVM_TRANSFER_MEMORY >> KVM_EXPORT_VCPU >> KVM_IMPORT_VCPU KVM_TRANSFER_VCPU >> >> It's just a record (1 byte) of what userspace asserted for the current session >> at SETUP. Vendor code still owns the hardware state and remains free to reject a >> role that doesn't match it. It is also what lets the generic layer reject a >> command to a VM that never set up a session. > > For the TRANSFER style operations, I would assume the direction is passed > for each transfer, just like the Linux does for the dmaengine. It's > possible that there may be transfers going both directions without the > role changing. > > So looks like we have tree things to consider: userspace set migration > role, the hardware state, and transfer direction. Two of those three I agree with: userspace passes the role at SETUP, and the vendor implementation tracks the hardware state. It's the per-transfer direction I don't think we need. > What if userspace always passes the role and transfer direction where it > makes sense? And then the hardware specific implementation tracks the > hardware state? Unless there's another meaning, direction would indicate that an operation must produce or consume a blob into/from a buffer. Such an indication would be necessary if the layer below the API cannot know what to do with a buffer, which may well be the case in your dmaengine analogy. However, in this case a vendor implementation sits below the generic layer and could derive what it needs to do from the role, command, and any session state it maintains. The TDH.EXPORT.ABORT example I cited upthread expects a token only once the session has left its pre-copy phase, so what to do with the buffer follows from state the vendor layer already holds. The gap I do see is in how the buffer itself is described: we have no way to express a command that takes an input and produces an output. A direction flag doesn't help there either, since it can only say one thing. My comments on patch 2 suggest a new 'capacity' field alongside a reframing of 'size' in kvm_transfer_buffer, where capacity bounds what the kernel may write and size reports what is actually there. That covers the both-ways case, and it also leaves nothing for a direction field to convey. Maybe that addresses your concern? ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-21 0:13 ` Kishen Maloor @ 2026-09-21 6:52 ` Tony Lindgren 2026-09-21 9:24 ` Tony Lindgren 2026-09-22 3:57 ` Kishen Maloor 0 siblings, 2 replies; 84+ messages in thread From: Tony Lindgren @ 2026-09-21 6:52 UTC (permalink / raw) To: Kishen Maloor Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Sun, Sep 20, 2026 at 05:13:10PM -0700, Kishen Maloor wrote: > On 9/17/26 10:58 PM, Tony Lindgren wrote: > > On Thu, Sep 17, 2026 at 09:32:23PM -0700, Kishen Maloor wrote: > >> On 9/16/26 11:42 PM, Tony Lindgren wrote: > >>> On Wed, Sep 16, 2026 at 08:31:32PM -0700, Kishen Maloor wrote: > >>>> On 9/15/26 10:09 PM, Tony Lindgren wrote: > >>>>> So trying to summarize the common flags for the role and separate vendor > >>>>> flags: > >>>>> > >>>>> struct kvm_migrate_cmd { > >>>>> __u16 command; > >>>>> __u16 flags; > >>>>> __u16 vflags; > >>>>> __u16 reserved; > >>>>> __u32 reserved; > >>>>> struct kvm_transfer_buffer buf; > >>>>> }; > >>>>> > >>>>> Is the above along the lines what you were thinking? > >>>> > >>>> No, I was suggesting a 'role' field carved out of the 'reserved' space, > >>>> like this: > >>>> > >>>> struct kvm_migrate_cmd { > >>>> __u16 command; > >>>> __u16 flags; > >>>> __u8 role; /* 0 = unset, 1 = source, 2 = destination */ > >>>> __u8 reserved[3]; > >>>> struct kvm_transfer_buffer buf; > >>>> }; > >>> > >>> OK yes thanks for clarifying, that works for me. > >>> > >>>>> Ah OK, yes that would also tell "the hardware has been initialized to a > >>>>> certain migration role". That seems like a usable common feature. > >>>> > >>>> Not quite. It tells us that userspace asserted a role for this VM's migration > >>>> session. Whether a TD was created for import is a separate, vendor-level detail. > >>>> The generic layer only needs the role to reject a session that never stated one, > >>>> and to pick the export or import callback. That callback then knows which side > >>>> it's on and can reject an incorrect role (e.g., if SETUP asserted dst for a src TD). > >>> > >>> That's a good point, the hardware role may not be set yet. > >>> > >>> I'm still wondering if there is a need to stash the userspace set role in > >>> KVM though. Likely only the hardware specific code can properly track the > >>> state of the hardware and adjust to the userspace requests. Seems just > >>> being able to pass the role in struct kvm_migrate_cmd should be enough? > >> > >> Passing it in kvm_migrate_cmd is enough for SETUP itself, but the commands > >> after SETUP like memory/vcpu transfers still have to reach the right > >> callback. So KVM would need to remember what was asserted so that the > >> generic layer can dispatch to the export or import facing callbacks. > >> We've been sketching (on this thread) an alternative UAPI set > >> (3 vs 5 ioctls) for consideration which this stored role enables: > >> > >> Proposed in the RFC Alternative > >> KVM_MIGRATE_CMD KVM_MIGRATE_CMD > >> KVM_EXPORT_MEMORY > >> KVM_IMPORT_MEMORY KVM_TRANSFER_MEMORY > >> KVM_EXPORT_VCPU > >> KVM_IMPORT_VCPU KVM_TRANSFER_VCPU > >> > >> It's just a record (1 byte) of what userspace asserted for the current session > >> at SETUP. Vendor code still owns the hardware state and remains free to reject a > >> role that doesn't match it. It is also what lets the generic layer reject a > >> command to a VM that never set up a session. > > > > For the TRANSFER style operations, I would assume the direction is passed > > for each transfer, just like the Linux does for the dmaengine. It's > > possible that there may be transfers going both directions without the > > role changing. > > > > So looks like we have tree things to consider: userspace set migration > > role, the hardware state, and transfer direction. > > Two of those three I agree with: userspace passes the role at SETUP, and > the vendor implementation tracks the hardware state. It's the per-transfer > direction I don't think we need. Ack on the userspace passing the role at SETUP and vendor implementation tracking the hardware state. Then for KVM tracking the role, I don't think we need it with the two above. The role tracking can always be added if really needed. Any other opinions on this one? The transfer direction is there with the EXPORT/IMPORT naming. Maybe just let's keep that naming for easier readability rather than try to switch to TRANSFER style naming. No transfer direction flag needed. > > What if userspace always passes the role and transfer direction where it > > makes sense? And then the hardware specific implementation tracks the > > hardware state? > > Unless there's another meaning, direction would indicate that an operation > must produce or consume a blob into/from a buffer. Such an indication would be > necessary if the layer below the API cannot know what to do with a buffer, > which may well be the case in your dmaengine analogy. However, in this case a > vendor implementation sits below the generic layer and could derive what it > needs to do from the role, command, and any session state it maintains. The > TDH.EXPORT.ABORT example I cited upthread expects a token only once the > session has left its pre-copy phase, so what to do with the buffer follows > from state the vendor layer already holds. Yeah I don't think the dmaengine API ever expects to get back a blob as a result of an outgoing transfer.. That would be a separate DMA transfer. > The gap I do see is in how the buffer itself is described: we have no way to > express a command that takes an input and produces an output. A direction flag > doesn't help there either, since it can only say one thing. My comments on patch 2 > suggest a new 'capacity' field alongside a reframing of 'size' in kvm_transfer_buffer, > where capacity bounds what the kernel may write and size reports what is actually > there. That covers the both-ways case, and it also leaves nothing for a direction > field to convey. Maybe that addresses your concern? Yes that's a good point, a transfer command may also return data in the transfer buffer and the size of returned data needs to be known. Replied to your patch #2 comments with some ideas on it. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-21 6:52 ` Tony Lindgren @ 2026-09-21 9:24 ` Tony Lindgren 2026-09-21 10:58 ` Tony Lindgren 2026-09-22 3:57 ` Kishen Maloor 1 sibling, 1 reply; 84+ messages in thread From: Tony Lindgren @ 2026-09-21 9:24 UTC (permalink / raw) To: Kishen Maloor Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Mon, Sep 21, 2026 at 09:52:42AM +0300, Tony Lindgren wrote: > On Sun, Sep 20, 2026 at 05:13:10PM -0700, Kishen Maloor wrote: > > On 9/17/26 10:58 PM, Tony Lindgren wrote: > > > So looks like we have tree things to consider: userspace set migration > > > role, the hardware state, and transfer direction. > > > > Two of those three I agree with: userspace passes the role at SETUP, and > > the vendor implementation tracks the hardware state. It's the per-transfer > > direction I don't think we need. > > Ack on the userspace passing the role at SETUP and vendor implementation > tracking the hardware state. > > Then for KVM tracking the role, I don't think we need it with the two > above. The role tracking can always be added if really needed. Any other > opinions on this one? > > The transfer direction is there with the EXPORT/IMPORT naming. Maybe > just let's keep that naming for easier readability rather than try to > switch to TRANSFER style naming. No transfer direction flag needed. Looking at include/linux/dmaengine.h again, the transfer specific direction is depreated.. So to follow the the dmaengine analogy, I now agree that migration SETUP is the place to set the transfer direction like you're suggesting. And so it seems the role at SETUP and the vendor implementation tracking the hardware state is all we need. The EXPORT/IMPORT vs TRANFER naming both work for me. I find the EXPORT/IMPORT easier to read, and I think Jörg also seemed to like that naming better. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-21 9:24 ` Tony Lindgren @ 2026-09-21 10:58 ` Tony Lindgren 0 siblings, 0 replies; 84+ messages in thread From: Tony Lindgren @ 2026-09-21 10:58 UTC (permalink / raw) To: Kishen Maloor Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Mon, Sep 21, 2026 at 12:24:26PM +0300, Tony Lindgren wrote: > On Mon, Sep 21, 2026 at 09:52:42AM +0300, Tony Lindgren wrote: > > On Sun, Sep 20, 2026 at 05:13:10PM -0700, Kishen Maloor wrote: > > > On 9/17/26 10:58 PM, Tony Lindgren wrote: > > > > So looks like we have tree things to consider: userspace set migration > > > > role, the hardware state, and transfer direction. > > > > > > Two of those three I agree with: userspace passes the role at SETUP, and > > > the vendor implementation tracks the hardware state. It's the per-transfer > > > direction I don't think we need. > > > > Ack on the userspace passing the role at SETUP and vendor implementation > > tracking the hardware state. > > > > Then for KVM tracking the role, I don't think we need it with the two > > above. The role tracking can always be added if really needed. Any other > > opinions on this one? > > > > The transfer direction is there with the EXPORT/IMPORT naming. Maybe > > just let's keep that naming for easier readability rather than try to > > switch to TRANSFER style naming. No transfer direction flag needed. > > Looking at include/linux/dmaengine.h again, the transfer specific > direction is depreated.. So to follow the the dmaengine analogy, I now > agree that migration SETUP is the place to set the transfer direction > like you're suggesting. And so it seems the role at SETUP and the vendor > implementation tracking the hardware state is all we need. Sorry above I meant "SETUP is the place to set the role", not the direction. > The EXPORT/IMPORT vs TRANFER naming both work for me. I find the > EXPORT/IMPORT easier to read, and I think Jörg also seemed to like > that naming better. And KVM is using KVM_GET/SET_REGS type ioctls, not TRANSFER. Also, I think EXPORT/IMPORT is better naming than GET/SET here as we export into an encrypted blob. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-21 6:52 ` Tony Lindgren 2026-09-21 9:24 ` Tony Lindgren @ 2026-09-22 3:57 ` Kishen Maloor 2026-09-22 5:25 ` Tony Lindgren 1 sibling, 1 reply; 84+ messages in thread From: Kishen Maloor @ 2026-09-22 3:57 UTC (permalink / raw) To: Tony Lindgren Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On 9/20/26 11:52 PM, Tony Lindgren wrote: > On Sun, Sep 20, 2026 at 05:13:10PM -0700, Kishen Maloor wrote: >> On 9/17/26 10:58 PM, Tony Lindgren wrote: >>> On Thu, Sep 17, 2026 at 09:32:23PM -0700, Kishen Maloor wrote: >>>> On 9/16/26 11:42 PM, Tony Lindgren wrote: >>>>> On Wed, Sep 16, 2026 at 08:31:32PM -0700, Kishen Maloor wrote: >>>>>> On 9/15/26 10:09 PM, Tony Lindgren wrote: >>>>>>> So trying to summarize the common flags for the role and separate vendor >>>>>>> flags: >>>>>>> >>>>>>> struct kvm_migrate_cmd { >>>>>>> __u16 command; >>>>>>> __u16 flags; >>>>>>> __u16 vflags; >>>>>>> __u16 reserved; >>>>>>> __u32 reserved; >>>>>>> struct kvm_transfer_buffer buf; >>>>>>> }; >>>>>>> >>>>>>> Is the above along the lines what you were thinking? >>>>>> >>>>>> No, I was suggesting a 'role' field carved out of the 'reserved' space, >>>>>> like this: >>>>>> >>>>>> struct kvm_migrate_cmd { >>>>>> __u16 command; >>>>>> __u16 flags; >>>>>> __u8 role; /* 0 = unset, 1 = source, 2 = destination */ >>>>>> __u8 reserved[3]; >>>>>> struct kvm_transfer_buffer buf; >>>>>> }; >>>>> >>>>> OK yes thanks for clarifying, that works for me. >>>>> >>>>>>> Ah OK, yes that would also tell "the hardware has been initialized to a >>>>>>> certain migration role". That seems like a usable common feature. >>>>>> >>>>>> Not quite. It tells us that userspace asserted a role for this VM's migration >>>>>> session. Whether a TD was created for import is a separate, vendor-level detail. >>>>>> The generic layer only needs the role to reject a session that never stated one, >>>>>> and to pick the export or import callback. That callback then knows which side >>>>>> it's on and can reject an incorrect role (e.g., if SETUP asserted dst for a src TD). >>>>> >>>>> That's a good point, the hardware role may not be set yet. >>>>> >>>>> I'm still wondering if there is a need to stash the userspace set role in >>>>> KVM though. Likely only the hardware specific code can properly track the >>>>> state of the hardware and adjust to the userspace requests. Seems just >>>>> being able to pass the role in struct kvm_migrate_cmd should be enough? >>>> >>>> Passing it in kvm_migrate_cmd is enough for SETUP itself, but the commands >>>> after SETUP like memory/vcpu transfers still have to reach the right >>>> callback. So KVM would need to remember what was asserted so that the >>>> generic layer can dispatch to the export or import facing callbacks. >>>> We've been sketching (on this thread) an alternative UAPI set >>>> (3 vs 5 ioctls) for consideration which this stored role enables: >>>> >>>> Proposed in the RFC Alternative >>>> KVM_MIGRATE_CMD KVM_MIGRATE_CMD >>>> KVM_EXPORT_MEMORY >>>> KVM_IMPORT_MEMORY KVM_TRANSFER_MEMORY >>>> KVM_EXPORT_VCPU >>>> KVM_IMPORT_VCPU KVM_TRANSFER_VCPU >>>> >>>> It's just a record (1 byte) of what userspace asserted for the current session >>>> at SETUP. Vendor code still owns the hardware state and remains free to reject a >>>> role that doesn't match it. It is also what lets the generic layer reject a >>>> command to a VM that never set up a session. >>> >>> For the TRANSFER style operations, I would assume the direction is passed >>> for each transfer, just like the Linux does for the dmaengine. It's >>> possible that there may be transfers going both directions without the >>> role changing. >>> >>> So looks like we have tree things to consider: userspace set migration >>> role, the hardware state, and transfer direction. >> >> Two of those three I agree with: userspace passes the role at SETUP, and >> the vendor implementation tracks the hardware state. It's the per-transfer >> direction I don't think we need. > > Ack on the userspace passing the role at SETUP and vendor implementation > tracking the hardware state. > > Then for KVM tracking the role, I don't think we need it with the two > above. The role tracking can always be added if really needed. Any other > opinions on this one? I do think it's useful for generic KVM to track this role (1 byte). - It lets _TRANSFER_ style calls dispatch directly to import or export callbacks based on the role. - Even if we don't adopt the _TRANSFER_ style, it enables generic KVM to reject mismatched calls, e.g., KVM_EXPORT_MEMORY on a destination. > > The transfer direction is there with the EXPORT/IMPORT naming. Maybe > just let's keep that naming for easier readability rather than try to > switch to TRANSFER style naming. No transfer direction flag needed. I can't say I have a clear preference between EXPORT/IMPORT vs _TRANSFER_. If we keep the EXPORT/IMPORT naming, then consistency would arguably call for splitting MIGRATE_CMD too, which makes it 6 vs 3 (or 5 vs 3 against the RFC as posted): EXPORT/IMPORT style _TRANSFER_ style KVM_EXPORT_CMD KVM_MIGRATE_CMD KVM_IMPORT_CMD KVM_EXPORT_MEMORY KVM_TRANSFER_MEMORY KVM_IMPORT_MEMORY KVM_EXPORT_VCPU KVM_TRANSFER_VCPU KVM_IMPORT_VCPU I guess the main benefit of the _TRANSFER_ style is reduced duplication. Each EXPORT/IMPORT pair takes the same struct and differs only by the role that the session already knows. Merging them gives one entry point per call type and lets userspace drive both ends from the same call site which could be considered a win. It doesn't reduce kernel code though as the top-level handler still branches internally on the role. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-22 3:57 ` Kishen Maloor @ 2026-09-22 5:25 ` Tony Lindgren 2026-09-23 0:38 ` Kishen Maloor 0 siblings, 1 reply; 84+ messages in thread From: Tony Lindgren @ 2026-09-22 5:25 UTC (permalink / raw) To: Kishen Maloor Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Mon, Sep 21, 2026 at 08:57:45PM -0700, Kishen Maloor wrote: > On 9/20/26 11:52 PM, Tony Lindgren wrote: > > Then for KVM tracking the role, I don't think we need it with the two > > above. The role tracking can always be added if really needed. Any other > > opinions on this one? > > I do think it's useful for generic KVM to track this role (1 byte). > - It lets _TRANSFER_ style calls dispatch directly to import or export > callbacks based on the role. > - Even if we don't adopt the _TRANSFER_ style, it enables generic KVM to > reject mismatched calls, e.g., KVM_EXPORT_MEMORY on a destination. Having KVM do generic checks on the calls is a good idea. There might be a simpler way of handling it though. Rather than having KVM track the migration state, how about we add a function to check for the migration session state from the vendor code? So something like this for the states you suggested earlier: enum kvm_lmstate { KVM_LM_NONE, KVM_LM_SOURCE, KVM_LM_DESTINATION, }; With something like this to get the state from the vendor code: enum kmv_lmstate kvm_arch_get_lmstate(struct kvm *); For x86 it would end up calling kvm_x86_call(get_lmstate)(kvm) and for the TDX specific case tdx_get_lmstate(). It would allow KVM to do the generic checks for the migration related calls you're describing. And having KVM start tracking the state can be still added later on too if it is needed. > > The transfer direction is there with the EXPORT/IMPORT naming. Maybe > > just let's keep that naming for easier readability rather than try to > > switch to TRANSFER style naming. No transfer direction flag needed. > > I can't say I have a clear preference between EXPORT/IMPORT vs _TRANSFER_. > If we keep the EXPORT/IMPORT naming, then consistency would arguably call for > splitting MIGRATE_CMD too, which makes it 6 vs 3 (or 5 vs 3 against the RFC > as posted): > > EXPORT/IMPORT style _TRANSFER_ style > KVM_EXPORT_CMD KVM_MIGRATE_CMD > KVM_IMPORT_CMD > KVM_EXPORT_MEMORY KVM_TRANSFER_MEMORY > KVM_IMPORT_MEMORY > KVM_EXPORT_VCPU KVM_TRANSFER_VCPU > KVM_IMPORT_VCPU Agreed we should split the MIGRATE_CMD too. Probably the number of ioctls is not and issue compared to following the KVM style and better readability. So my vote is now on EXPORT/IMPORT style naming. > I guess the main benefit of the _TRANSFER_ style is reduced duplication. > Each EXPORT/IMPORT pair takes the same struct and differs only by the role > that the session already knows. Merging them gives one entry point per call > type and lets userspace drive both ends from the same call site which could > be considered a win. It doesn't reduce kernel code though as the top-level > handler still branches internally on the role. Yup not much of a win for the TRANSFER style naming. Trying to summarize again after we sorted out the direction flag issue in the transfer: role per migration session (cannot change during the migration) direction per command, EXPORT/IMPORT hardware state set and tracked by vendor specific code And checking again against the dmaengine analogy: Compared to dmaengine, the migration role is modeled similar to the dma channel configuration. The migration direction with EXPORT/IMPORT is modeled similar to dmaengine_prep_slave_sg(). The migration hardware state is modeled similar to the dmaengine driver managing the hardware state. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-22 5:25 ` Tony Lindgren @ 2026-09-23 0:38 ` Kishen Maloor 2026-09-23 6:04 ` Tony Lindgren 0 siblings, 1 reply; 84+ messages in thread From: Kishen Maloor @ 2026-09-23 0:38 UTC (permalink / raw) To: Tony Lindgren Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On 9/21/26 10:25 PM, Tony Lindgren wrote: > On Mon, Sep 21, 2026 at 08:57:45PM -0700, Kishen Maloor wrote: >> On 9/20/26 11:52 PM, Tony Lindgren wrote: >>> Then for KVM tracking the role, I don't think we need it with the two >>> above. The role tracking can always be added if really needed. Any other >>> opinions on this one? >> >> I do think it's useful for generic KVM to track this role (1 byte). >> - It lets _TRANSFER_ style calls dispatch directly to import or export >> callbacks based on the role. >> - Even if we don't adopt the _TRANSFER_ style, it enables generic KVM to >> reject mismatched calls, e.g., KVM_EXPORT_MEMORY on a destination. > > Having KVM do generic checks on the calls is a good idea. There might be > a simpler way of handling it though. Rather than having KVM track the > migration state, how about we add a function to check for the migration > session state from the vendor code? > > So something like this for the states you suggested earlier: > > enum kvm_lmstate { > KVM_LM_NONE, > KVM_LM_SOURCE, > KVM_LM_DESTINATION, > }; > > With something like this to get the state from the vendor code: > > enum kmv_lmstate kvm_arch_get_lmstate(struct kvm *); > > For x86 it would end up calling kvm_x86_call(get_lmstate)(kvm) and for > the TDX specific case tdx_get_lmstate(). How is this simpler? It trades one byte in a KVM struct for a new generic enum, a new kvm_arch_get_lmstate(), a new kvm_x86_ops entry, and a vendor implementation per vendor, plus a cross-layer call on every command just to learn the role. Directly checking a stored byte (0=unset/1=src/2=dst) seems simplest, no? > It would allow KVM to do the generic checks for the migration related > calls you're describing. And having KVM start tracking the state can be > still added later on too if it is needed. I think we'd need to pick one way or the other before the UAPI settles if we want generic KVM to reject mismatched calls. OTOH if we want to defer this generic KVM validation, then yeah, it could be settled later. > >>> The transfer direction is there with the EXPORT/IMPORT naming. Maybe >>> just let's keep that naming for easier readability rather than try to >>> switch to TRANSFER style naming. No transfer direction flag needed. >> >> I can't say I have a clear preference between EXPORT/IMPORT vs _TRANSFER_. >> If we keep the EXPORT/IMPORT naming, then consistency would arguably call for >> splitting MIGRATE_CMD too, which makes it 6 vs 3 (or 5 vs 3 against the RFC >> as posted): >> >> EXPORT/IMPORT style _TRANSFER_ style >> KVM_EXPORT_CMD KVM_MIGRATE_CMD >> KVM_IMPORT_CMD >> KVM_EXPORT_MEMORY KVM_TRANSFER_MEMORY >> KVM_IMPORT_MEMORY >> KVM_EXPORT_VCPU KVM_TRANSFER_VCPU >> KVM_IMPORT_VCPU > > Agreed we should split the MIGRATE_CMD too. Probably the number of ioctls > is not and issue compared to following the KVM style and better > readability. So my vote is now on EXPORT/IMPORT style naming. > >> I guess the main benefit of the _TRANSFER_ style is reduced duplication. >> Each EXPORT/IMPORT pair takes the same struct and differs only by the role >> that the session already knows. Merging them gives one entry point per call >> type and lets userspace drive both ends from the same call site which could >> be considered a win. It doesn't reduce kernel code though as the top-level >> handler still branches internally on the role. > > Yup not much of a win for the TRANSFER style naming. Yeah, my goal was just to enumerate alternatives for consideration. One other benefit of KVM_EXPORT_CMD/KVM_IMPORT_CMD is that the per-session role is implicitly conveyed - in other words, a successful KVM_EXPORT_CMD/SETUP (for e.g.) would indicate that this is a 'source'. So, we wouldn't require a 'role' field in struct kvm_migrate_cmd to explicitly assert one during SETUP. > > Trying to summarize again after we sorted out the direction flag issue in > the transfer: > > role per migration session (cannot change during the migration) > direction per command, EXPORT/IMPORT > hardware state set and tracked by vendor specific code The 'role' and 'direction' as you define it above are essentially saying the same thing - a source only invokes the vendor's EXPORT call, a destination only invokes the vendor's IMPORT call, and the role doesn't change over the session. If a destination needs to send data to be consumed by the source, then that still invokes the vendor's EXPORT call. > And checking again against the dmaengine analogy: > > Compared to dmaengine, the migration role is modeled similar to the dma > channel configuration. > > The migration direction with EXPORT/IMPORT is modeled similar to > dmaengine_prep_slave_sg(). I'll leave the dmaengine comparison to you, I don't know that API well enough to map it properly :) ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-23 0:38 ` Kishen Maloor @ 2026-09-23 6:04 ` Tony Lindgren 2026-09-24 5:53 ` Kishen Maloor 0 siblings, 1 reply; 84+ messages in thread From: Tony Lindgren @ 2026-09-23 6:04 UTC (permalink / raw) To: Kishen Maloor Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Tue, Sep 22, 2026 at 05:38:43PM -0700, Kishen Maloor wrote: > On 9/21/26 10:25 PM, Tony Lindgren wrote: > > On Mon, Sep 21, 2026 at 08:57:45PM -0700, Kishen Maloor wrote: > >> On 9/20/26 11:52 PM, Tony Lindgren wrote: > >>> Then for KVM tracking the role, I don't think we need it with the two > >>> above. The role tracking can always be added if really needed. Any other > >>> opinions on this one? > >> > >> I do think it's useful for generic KVM to track this role (1 byte). > >> - It lets _TRANSFER_ style calls dispatch directly to import or export > >> callbacks based on the role. > >> - Even if we don't adopt the _TRANSFER_ style, it enables generic KVM to > >> reject mismatched calls, e.g., KVM_EXPORT_MEMORY on a destination. > > > > Having KVM do generic checks on the calls is a good idea. There might be > > a simpler way of handling it though. Rather than having KVM track the > > migration state, how about we add a function to check for the migration > > session state from the vendor code? > > > > So something like this for the states you suggested earlier: > > > > enum kvm_lmstate { > > KVM_LM_NONE, > > KVM_LM_SOURCE, > > KVM_LM_DESTINATION, > > }; > > > > With something like this to get the state from the vendor code: > > > > enum kmv_lmstate kvm_arch_get_lmstate(struct kvm *); > > > > For x86 it would end up calling kvm_x86_call(get_lmstate)(kvm) and for > > the TDX specific case tdx_get_lmstate(). > > How is this simpler? It trades one byte in a KVM struct for a new generic > enum, a new kvm_arch_get_lmstate(), a new kvm_x86_ops entry, and a vendor > implementation per vendor, plus a cross-layer call on every command just > to learn the role. > Directly checking a stored byte (0=unset/1=src/2=dst) seems simplest, no? It would avoid dragging KVM into the "track the migration state" business at least for now. I guess the question in general is: What does KVM need to do with the migration role beyond generic checks on the migration related calls? > > It would allow KVM to do the generic checks for the migration related > > calls you're describing. And having KVM start tracking the state can be > > still added later on too if it is needed. > > I think we'd need to pick one way or the other before the UAPI settles > if we want generic KVM to reject mismatched calls. OTOH if we want to > defer this generic KVM validation, then yeah, it could be settled later. > > > > >>> The transfer direction is there with the EXPORT/IMPORT naming. Maybe > >>> just let's keep that naming for easier readability rather than try to > >>> switch to TRANSFER style naming. No transfer direction flag needed. > >> > >> I can't say I have a clear preference between EXPORT/IMPORT vs _TRANSFER_. > >> If we keep the EXPORT/IMPORT naming, then consistency would arguably call for > >> splitting MIGRATE_CMD too, which makes it 6 vs 3 (or 5 vs 3 against the RFC > >> as posted): > >> > >> EXPORT/IMPORT style _TRANSFER_ style > >> KVM_EXPORT_CMD KVM_MIGRATE_CMD > >> KVM_IMPORT_CMD > >> KVM_EXPORT_MEMORY KVM_TRANSFER_MEMORY > >> KVM_IMPORT_MEMORY > >> KVM_EXPORT_VCPU KVM_TRANSFER_VCPU > >> KVM_IMPORT_VCPU > > > > Agreed we should split the MIGRATE_CMD too. Probably the number of ioctls > > is not and issue compared to following the KVM style and better > > readability. So my vote is now on EXPORT/IMPORT style naming. > > > >> I guess the main benefit of the _TRANSFER_ style is reduced duplication. > >> Each EXPORT/IMPORT pair takes the same struct and differs only by the role > >> that the session already knows. Merging them gives one entry point per call > >> type and lets userspace drive both ends from the same call site which could > >> be considered a win. It doesn't reduce kernel code though as the top-level > >> handler still branches internally on the role. > > > > Yup not much of a win for the TRANSFER style naming. > > Yeah, my goal was just to enumerate alternatives for consideration. > > One other benefit of KVM_EXPORT_CMD/KVM_IMPORT_CMD is that the per-session > role is implicitly conveyed - in other words, a successful KVM_EXPORT_CMD/SETUP > (for e.g.) would indicate that this is a 'source'. So, we wouldn't require a 'role' > field in struct kvm_migrate_cmd to explicitly assert one during SETUP. Yes good point with the KVM_EXPORT/IMPORT_CMD, that sounds good to me. > > Trying to summarize again after we sorted out the direction flag issue in > > the transfer: > > > > role per migration session (cannot change during the migration) > > direction per command, EXPORT/IMPORT > > hardware state set and tracked by vendor specific code > > The 'role' and 'direction' as you define it above are essentially saying the > same thing - a source only invokes the vendor's EXPORT call, a destination only > invokes the vendor's IMPORT call, and the role doesn't change over the session. > If a destination needs to send data to be consumed by the source, then that > still invokes the vendor's EXPORT call. Yup. > > And checking again against the dmaengine analogy: > > > > Compared to dmaengine, the migration role is modeled similar to the dma > > channel configuration. > > > > The migration direction with EXPORT/IMPORT is modeled similar to > > dmaengine_prep_slave_sg(). > I'll leave the dmaengine comparison to you, I don't know that API well enough > to map it properly :) Heh just a sanity check for trying to relate this to something existing. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-23 6:04 ` Tony Lindgren @ 2026-09-24 5:53 ` Kishen Maloor 2026-09-24 6:59 ` Tony Lindgren 0 siblings, 1 reply; 84+ messages in thread From: Kishen Maloor @ 2026-09-24 5:53 UTC (permalink / raw) To: Tony Lindgren Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On 9/22/26 11:04 PM, Tony Lindgren wrote: > On Tue, Sep 22, 2026 at 05:38:43PM -0700, Kishen Maloor wrote: >> On 9/21/26 10:25 PM, Tony Lindgren wrote: >>> On Mon, Sep 21, 2026 at 08:57:45PM -0700, Kishen Maloor wrote: >>>> On 9/20/26 11:52 PM, Tony Lindgren wrote: >>>>> Then for KVM tracking the role, I don't think we need it with the two >>>>> above. The role tracking can always be added if really needed. Any other >>>>> opinions on this one? >>>> >>>> I do think it's useful for generic KVM to track this role (1 byte). >>>> - It lets _TRANSFER_ style calls dispatch directly to import or export >>>> callbacks based on the role. >>>> - Even if we don't adopt the _TRANSFER_ style, it enables generic KVM to >>>> reject mismatched calls, e.g., KVM_EXPORT_MEMORY on a destination. >>> >>> Having KVM do generic checks on the calls is a good idea. There might be >>> a simpler way of handling it though. Rather than having KVM track the >>> migration state, how about we add a function to check for the migration >>> session state from the vendor code? >>> >>> So something like this for the states you suggested earlier: >>> >>> enum kvm_lmstate { >>> KVM_LM_NONE, >>> KVM_LM_SOURCE, >>> KVM_LM_DESTINATION, >>> }; >>> >>> With something like this to get the state from the vendor code: >>> >>> enum kmv_lmstate kvm_arch_get_lmstate(struct kvm *); >>> >>> For x86 it would end up calling kvm_x86_call(get_lmstate)(kvm) and for >>> the TDX specific case tdx_get_lmstate(). >> >> How is this simpler? It trades one byte in a KVM struct for a new generic >> enum, a new kvm_arch_get_lmstate(), a new kvm_x86_ops entry, and a vendor >> implementation per vendor, plus a cross-layer call on every command just >> to learn the role. >> Directly checking a stored byte (0=unset/1=src/2=dst) seems simplest, no? > > It would avoid dragging KVM into the "track the migration state" business > at least for now. > > I guess the question in general is: What does KVM need to do with the > migration role beyond generic checks on the migration related calls? I haven't thought of any other uses for it. But if we do want those generic checks, I'd still lean toward the stored byte. > >>> It would allow KVM to do the generic checks for the migration related >>> calls you're describing. And having KVM start tracking the state can be >>> still added later on too if it is needed. >> >> I think we'd need to pick one way or the other before the UAPI settles >> if we want generic KVM to reject mismatched calls. OTOH if we want to >> defer this generic KVM validation, then yeah, it could be settled later. >> >>> >>>>> The transfer direction is there with the EXPORT/IMPORT naming. Maybe >>>>> just let's keep that naming for easier readability rather than try to >>>>> switch to TRANSFER style naming. No transfer direction flag needed. >>>> >>>> I can't say I have a clear preference between EXPORT/IMPORT vs _TRANSFER_. >>>> If we keep the EXPORT/IMPORT naming, then consistency would arguably call for >>>> splitting MIGRATE_CMD too, which makes it 6 vs 3 (or 5 vs 3 against the RFC >>>> as posted): >>>> >>>> EXPORT/IMPORT style _TRANSFER_ style >>>> KVM_EXPORT_CMD KVM_MIGRATE_CMD >>>> KVM_IMPORT_CMD >>>> KVM_EXPORT_MEMORY KVM_TRANSFER_MEMORY >>>> KVM_IMPORT_MEMORY >>>> KVM_EXPORT_VCPU KVM_TRANSFER_VCPU >>>> KVM_IMPORT_VCPU >>> >>> Agreed we should split the MIGRATE_CMD too. Probably the number of ioctls >>> is not and issue compared to following the KVM style and better >>> readability. So my vote is now on EXPORT/IMPORT style naming. >>> >>>> I guess the main benefit of the _TRANSFER_ style is reduced duplication. >>>> Each EXPORT/IMPORT pair takes the same struct and differs only by the role >>>> that the session already knows. Merging them gives one entry point per call >>>> type and lets userspace drive both ends from the same call site which could >>>> be considered a win. It doesn't reduce kernel code though as the top-level >>>> handler still branches internally on the role. >>> >>> Yup not much of a win for the TRANSFER style naming. >> >> Yeah, my goal was just to enumerate alternatives for consideration. >> >> One other benefit of KVM_EXPORT_CMD/KVM_IMPORT_CMD is that the per-session >> role is implicitly conveyed - in other words, a successful KVM_EXPORT_CMD/SETUP >> (for e.g.) would indicate that this is a 'source'. So, we wouldn't require a 'role' >> field in struct kvm_migrate_cmd to explicitly assert one during SETUP. > > Yes good point with the KVM_EXPORT/IMPORT_CMD, that sounds good to me. > >>> Trying to summarize again after we sorted out the direction flag issue in >>> the transfer: >>> >>> role per migration session (cannot change during the migration) >>> direction per command, EXPORT/IMPORT >>> hardware state set and tracked by vendor specific code >> >> The 'role' and 'direction' as you define it above are essentially saying the >> same thing - a source only invokes the vendor's EXPORT call, a destination only >> invokes the vendor's IMPORT call, and the role doesn't change over the session. >> If a destination needs to send data to be consumed by the source, then that >> still invokes the vendor's EXPORT call. > > Yup. > >>> And checking again against the dmaengine analogy: >>> >>> Compared to dmaengine, the migration role is modeled similar to the dma >>> channel configuration. >>> >>> The migration direction with EXPORT/IMPORT is modeled similar to >>> dmaengine_prep_slave_sg(). >> I'll leave the dmaengine comparison to you, I don't know that API well enough >> to map it properly :) > > Heh just a sanity check for trying to relate this to something existing. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-24 5:53 ` Kishen Maloor @ 2026-09-24 6:59 ` Tony Lindgren 0 siblings, 0 replies; 84+ messages in thread From: Tony Lindgren @ 2026-09-24 6:59 UTC (permalink / raw) To: Kishen Maloor Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Wed, Sep 23, 2026 at 10:53:28PM -0700, Kishen Maloor wrote: > On 9/22/26 11:04 PM, Tony Lindgren wrote: > > On Tue, Sep 22, 2026 at 05:38:43PM -0700, Kishen Maloor wrote: > >> On 9/21/26 10:25 PM, Tony Lindgren wrote: > >>> On Mon, Sep 21, 2026 at 08:57:45PM -0700, Kishen Maloor wrote: > >>>> On 9/20/26 11:52 PM, Tony Lindgren wrote: > >>>>> Then for KVM tracking the role, I don't think we need it with the two > >>>>> above. The role tracking can always be added if really needed. Any other > >>>>> opinions on this one? > >>>> > >>>> I do think it's useful for generic KVM to track this role (1 byte). > >>>> - It lets _TRANSFER_ style calls dispatch directly to import or export > >>>> callbacks based on the role. > >>>> - Even if we don't adopt the _TRANSFER_ style, it enables generic KVM to > >>>> reject mismatched calls, e.g., KVM_EXPORT_MEMORY on a destination. > >>> > >>> Having KVM do generic checks on the calls is a good idea. There might be > >>> a simpler way of handling it though. Rather than having KVM track the > >>> migration state, how about we add a function to check for the migration > >>> session state from the vendor code? > >>> > >>> So something like this for the states you suggested earlier: > >>> > >>> enum kvm_lmstate { > >>> KVM_LM_NONE, > >>> KVM_LM_SOURCE, > >>> KVM_LM_DESTINATION, > >>> }; > >>> > >>> With something like this to get the state from the vendor code: > >>> > >>> enum kmv_lmstate kvm_arch_get_lmstate(struct kvm *); > >>> > >>> For x86 it would end up calling kvm_x86_call(get_lmstate)(kvm) and for > >>> the TDX specific case tdx_get_lmstate(). > >> > >> How is this simpler? It trades one byte in a KVM struct for a new generic > >> enum, a new kvm_arch_get_lmstate(), a new kvm_x86_ops entry, and a vendor > >> implementation per vendor, plus a cross-layer call on every command just > >> to learn the role. > >> Directly checking a stored byte (0=unset/1=src/2=dst) seems simplest, no? > > > > It would avoid dragging KVM into the "track the migration state" business > > at least for now. > > > > I guess the question in general is: What does KVM need to do with the > > migration role beyond generic checks on the migration related calls? > > I haven't thought of any other uses for it. But if we do want those generic > checks, I'd still lean toward the stored byte. We can easily add the KVM role later on if real KVM generic need for carrying the migration role comes up. We can just have the vendor code do the checks for now. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-08-31 7:13 ` [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD Tony Lindgren 2026-08-31 7:23 ` sashiko-bot 2026-09-07 11:53 ` Tony Lindgren @ 2026-09-18 4:33 ` Kishen Maloor 2026-09-21 5:58 ` Tony Lindgren 2 siblings, 1 reply; 84+ messages in thread From: Kishen Maloor @ 2026-09-18 4:33 UTC (permalink / raw) To: Tony Lindgren, Paolo Bonzini, Sean Christopherson Cc: Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On 8/31/26 12:13 AM, Tony Lindgren wrote: I was wondering how a vendor would use kvm_transfer_buffer for a command that passes an input and also returns an output. size can describe the input length or the space available for output, not both, so the kernel has no way to know how much it may write. Two suggestions below. They're orthogonal. > ... > diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h > index 5f6c1ce9673b7..d9291a8a97bb1 100644 > --- a/arch/x86/include/asm/kvm_host.h > +++ b/arch/x86/include/asm/kvm_host.h > @@ -2010,6 +2010,8 @@ struct kvm_x86_ops { > int (*gmem_prepare)(struct kvm *kvm, kvm_pfn_t pfn, gfn_t gfn, int max_order); > void (*gmem_invalidate)(kvm_pfn_t start, kvm_pfn_t end); > int (*gmem_max_mapping_level)(struct kvm *kvm, kvm_pfn_t pfn, bool is_private); > + bool (*cap_live_migration)(struct kvm *kvm); Should this return an 'int' specifying the maximum buffer size that the vendor impl requires/uses? Userspace can then learn this once. > ... > + > +struct kvm_transfer_buffer { > + __u64 address; > + __u32 size; > + __u32 reserved; > +}; Should this struct include a 'capacity' field (u32) that is set on each command? It would be the number of bytes writable at address. size would be the input length on entry (0 if the command passes none), and the number of bytes produced on return (0 if none). ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-18 4:33 ` Kishen Maloor @ 2026-09-21 5:58 ` Tony Lindgren 2026-09-21 6:56 ` Tony Lindgren 0 siblings, 1 reply; 84+ messages in thread From: Tony Lindgren @ 2026-09-21 5:58 UTC (permalink / raw) To: Kishen Maloor Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Thu, Sep 17, 2026 at 09:33:18PM -0700, Kishen Maloor wrote: > On 8/31/26 12:13 AM, Tony Lindgren wrote: > > I was wondering how a vendor would use kvm_transfer_buffer for a command that > passes an input and also returns an output. size can describe the input > length or the space available for output, not both, so the kernel has no way to > know how much it may write. > > Two suggestions below. They're orthogonal. > > > ... > > diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h > > index 5f6c1ce9673b7..d9291a8a97bb1 100644 > > --- a/arch/x86/include/asm/kvm_host.h > > +++ b/arch/x86/include/asm/kvm_host.h > > @@ -2010,6 +2010,8 @@ struct kvm_x86_ops { > > int (*gmem_prepare)(struct kvm *kvm, kvm_pfn_t pfn, gfn_t gfn, int max_order); > > void (*gmem_invalidate)(kvm_pfn_t start, kvm_pfn_t end); > > int (*gmem_max_mapping_level)(struct kvm *kvm, kvm_pfn_t pfn, bool is_private); > > + bool (*cap_live_migration)(struct kvm *kvm); > > Should this return an 'int' specifying the maximum buffer size that the vendor impl > requires/uses? Userspace can then learn this once. OK that sounds good to me. I assume you are thinking the maximum hardware specific buffer size per thread? As in 512 * 4096 bytes for the TDX case? > > + > > +struct kvm_transfer_buffer { > > + __u64 address; > > + __u32 size; > > + __u32 reserved; > > +}; > > Should this struct include a 'capacity' field (u32) that is set on each command? > It would be the number of bytes writable at address. > size would be the input length on entry (0 if the command passes none), and the > number of bytes produced on return (0 if none). Hmm so the transfer command return value can return how many bytes were written of the input. But yeah we don't know how many bytes were written back to the transfer buffer as result of the transfer command. How about if we add the bytes returned to the transfer struct? Then the kvm_transfer_buffer can stay as just a buffer. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-21 5:58 ` Tony Lindgren @ 2026-09-21 6:56 ` Tony Lindgren 2026-09-22 3:56 ` Kishen Maloor 0 siblings, 1 reply; 84+ messages in thread From: Tony Lindgren @ 2026-09-21 6:56 UTC (permalink / raw) To: Kishen Maloor Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Mon, Sep 21, 2026 at 08:58:17AM +0300, Tony Lindgren wrote: > On Thu, Sep 17, 2026 at 09:33:18PM -0700, Kishen Maloor wrote: > > On 8/31/26 12:13 AM, Tony Lindgren wrote: > > > > I was wondering how a vendor would use kvm_transfer_buffer for a command that > > passes an input and also returns an output. size can describe the input > > length or the space available for output, not both, so the kernel has no way to > > know how much it may write. > > > > Two suggestions below. They're orthogonal. > > > > > ... > > > diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h > > > index 5f6c1ce9673b7..d9291a8a97bb1 100644 > > > --- a/arch/x86/include/asm/kvm_host.h > > > +++ b/arch/x86/include/asm/kvm_host.h > > > @@ -2010,6 +2010,8 @@ struct kvm_x86_ops { > > > int (*gmem_prepare)(struct kvm *kvm, kvm_pfn_t pfn, gfn_t gfn, int max_order); > > > void (*gmem_invalidate)(kvm_pfn_t start, kvm_pfn_t end); > > > int (*gmem_max_mapping_level)(struct kvm *kvm, kvm_pfn_t pfn, bool is_private); > > > + bool (*cap_live_migration)(struct kvm *kvm); > > > > Should this return an 'int' specifying the maximum buffer size that the vendor impl > > requires/uses? Userspace can then learn this once. > > OK that sounds good to me. I assume you are thinking the maximum hardware > specific buffer size per thread? As in 512 * 4096 bytes for the TDX case? > > > > + > > > +struct kvm_transfer_buffer { > > > + __u64 address; > > > + __u32 size; > > > + __u32 reserved; > > > +}; > > > > Should this struct include a 'capacity' field (u32) that is set on each command? > > It would be the number of bytes writable at address. > > size would be the input length on entry (0 if the command passes none), and the > > number of bytes produced on return (0 if none). > > Hmm so the transfer command return value can return how many bytes were > written of the input. But yeah we don't know how many bytes were written > back to the transfer buffer as result of the transfer command. > > How about if we add the bytes returned to the transfer struct? Then the > kvm_transfer_buffer can stay as just a buffer. Actually, for the possible cases with input+output, we could reserve space in the transfer struct for another struct kvm_transfer_buffer for the results? ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-21 6:56 ` Tony Lindgren @ 2026-09-22 3:56 ` Kishen Maloor 2026-09-22 6:27 ` Tony Lindgren 0 siblings, 1 reply; 84+ messages in thread From: Kishen Maloor @ 2026-09-22 3:56 UTC (permalink / raw) To: Tony Lindgren Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On 9/20/26 11:56 PM, Tony Lindgren wrote: > On Mon, Sep 21, 2026 at 08:58:17AM +0300, Tony Lindgren wrote: >> On Thu, Sep 17, 2026 at 09:33:18PM -0700, Kishen Maloor wrote: >>> On 8/31/26 12:13 AM, Tony Lindgren wrote: >>> >>> I was wondering how a vendor would use kvm_transfer_buffer for a command that >>> passes an input and also returns an output. size can describe the input >>> length or the space available for output, not both, so the kernel has no way to >>> know how much it may write. >>> >>> Two suggestions below. They're orthogonal. >>> >>>> ... >>>> diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h >>>> index 5f6c1ce9673b7..d9291a8a97bb1 100644 >>>> --- a/arch/x86/include/asm/kvm_host.h >>>> +++ b/arch/x86/include/asm/kvm_host.h >>>> @@ -2010,6 +2010,8 @@ struct kvm_x86_ops { >>>> int (*gmem_prepare)(struct kvm *kvm, kvm_pfn_t pfn, gfn_t gfn, int max_order); >>>> void (*gmem_invalidate)(kvm_pfn_t start, kvm_pfn_t end); >>>> int (*gmem_max_mapping_level)(struct kvm *kvm, kvm_pfn_t pfn, bool is_private); >>>> + bool (*cap_live_migration)(struct kvm *kvm); >>> >>> Should this return an 'int' specifying the maximum buffer size that the vendor impl >>> requires/uses? Userspace can then learn this once. >> >> OK that sounds good to me. I assume you are thinking the maximum hardware >> specific buffer size per thread? As in 512 * 4096 bytes for the TDX case? Yeah, an upper bound on the buffer size, so 516 * 4096 (4 for the GPA+MAC lists and MBMD) in TDX. It could be a hint to userspace to size its buffers at the start. Though it would be up to userspace to decide how to use that information. >> >>>> + >>>> +struct kvm_transfer_buffer { >>>> + __u64 address; >>>> + __u32 size; >>>> + __u32 reserved; >>>> +}; >>> >>> Should this struct include a 'capacity' field (u32) that is set on each command? >>> It would be the number of bytes writable at address. >>> size would be the input length on entry (0 if the command passes none), and the >>> number of bytes produced on return (0 if none). >> >> Hmm so the transfer command return value can return how many bytes were >> written of the input. But yeah we don't know how many bytes were written >> back to the transfer buffer as result of the transfer command. The transfer command return value could return how many bytes were written into the buffer. But in an input-output call, the kernel handler wouldn't know how many bytes it could write, or for that matter even how many pages to pin up front in case it needs to return an output because 'size' couldn't simultaneously convey the input length and buffer capacity. That was the gap that I thought a read-only 'capacity' field could bridge. Of course, this assumes that the output is written in-place. >> >> How about if we add the bytes returned to the transfer struct? Then the >> kvm_transfer_buffer can stay as just a buffer. > > Actually, for the possible cases with input+output, we could reserve space > in the transfer struct for another struct kvm_transfer_buffer for the > results? Yes, say an 'in' and 'out' kvm_transfer_buffer inside struct kvm_migrate_cmd should close this out and shouldn't require a 'capacity' field. Maybe we then establish this convention: - A non-zero 'size' on 'in' at call entry would signal that there is input. - A non-zero 'size' on 'out' at call entry would convey the buffer capacity. - A non-zero 'size' on 'out' at call exit would convey that there is output. - A zeroed 'size' on 'out' at call exit would convey that there is no output. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-22 3:56 ` Kishen Maloor @ 2026-09-22 6:27 ` Tony Lindgren 2026-09-23 0:37 ` Kishen Maloor 0 siblings, 1 reply; 84+ messages in thread From: Tony Lindgren @ 2026-09-22 6:27 UTC (permalink / raw) To: Kishen Maloor Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Mon, Sep 21, 2026 at 08:56:48PM -0700, Kishen Maloor wrote: > On 9/20/26 11:56 PM, Tony Lindgren wrote: > > On Mon, Sep 21, 2026 at 08:58:17AM +0300, Tony Lindgren wrote: > >> On Thu, Sep 17, 2026 at 09:33:18PM -0700, Kishen Maloor wrote: > >>> On 8/31/26 12:13 AM, Tony Lindgren wrote: > >>> > >>> I was wondering how a vendor would use kvm_transfer_buffer for a command that > >>> passes an input and also returns an output. size can describe the input > >>> length or the space available for output, not both, so the kernel has no way to > >>> know how much it may write. > >>> > >>> Two suggestions below. They're orthogonal. > >>> > >>>> ... > >>>> diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h > >>>> index 5f6c1ce9673b7..d9291a8a97bb1 100644 > >>>> --- a/arch/x86/include/asm/kvm_host.h > >>>> +++ b/arch/x86/include/asm/kvm_host.h > >>>> @@ -2010,6 +2010,8 @@ struct kvm_x86_ops { > >>>> int (*gmem_prepare)(struct kvm *kvm, kvm_pfn_t pfn, gfn_t gfn, int max_order); > >>>> void (*gmem_invalidate)(kvm_pfn_t start, kvm_pfn_t end); > >>>> int (*gmem_max_mapping_level)(struct kvm *kvm, kvm_pfn_t pfn, bool is_private); > >>>> + bool (*cap_live_migration)(struct kvm *kvm); > >>> > >>> Should this return an 'int' specifying the maximum buffer size that the vendor impl > >>> requires/uses? Userspace can then learn this once. > >> > >> OK that sounds good to me. I assume you are thinking the maximum hardware > >> specific buffer size per thread? As in 512 * 4096 bytes for the TDX case? > > Yeah, an upper bound on the buffer size, so 516 * 4096 (4 for the GPA+MAC lists > and MBMD) in TDX. It could be a hint to userspace to size its buffers at the > start. Though it would be up to userspace to decide how to use that information. Oh right thanks. I think this is really the maximum transfer buffer size Peter asked, not just a hint to userspace :) > >>>> +struct kvm_transfer_buffer { > >>>> + __u64 address; > >>>> + __u32 size; > >>>> + __u32 reserved; > >>>> +}; > >>> > >>> Should this struct include a 'capacity' field (u32) that is set on each command? > >>> It would be the number of bytes writable at address. > >>> size would be the input length on entry (0 if the command passes none), and the > >>> number of bytes produced on return (0 if none). > >> > >> Hmm so the transfer command return value can return how many bytes were > >> written of the input. But yeah we don't know how many bytes were written > >> back to the transfer buffer as result of the transfer command. > > The transfer command return value could return how many bytes were written into the buffer. > But in an input-output call, the kernel handler wouldn't know how many bytes it could write, > or for that matter even how many pages to pin up front in case it needs to return an output > because 'size' couldn't simultaneously convey the input length and buffer capacity. That was > the gap that I thought a read-only 'capacity' field could bridge. Of course, this > assumes that the output is written in-place. Hmm yeah this inplace capacity vs transferred issue remains still. So I agree we need to specify the capacity in struct kvm_transfer_buffer like you suggested. To me size is already the size of the buffer though. So instead of changing size to capacity, how about something like datasize or len for the input and output transfer length? > >> How about if we add the bytes returned to the transfer struct? Then the > >> kvm_transfer_buffer can stay as just a buffer. > > > > Actually, for the possible cases with input+output, we could reserve space > > in the transfer struct for another struct kvm_transfer_buffer for the > > results? > > Yes, say an 'in' and 'out' kvm_transfer_buffer inside struct kvm_migrate_cmd should > close this out and shouldn't require a 'capacity' field. Maybe we then establish this > convention: > - A non-zero 'size' on 'in' at call entry would signal that there is input. > - A non-zero 'size' on 'out' at call entry would convey the buffer capacity. > - A non-zero 'size' on 'out' at call exit would convey that there is output. > - A zeroed 'size' on 'out' at call exit would convey that there is no output. Looks doable to me but with the inplace issue as above.. Sounds like we just need to reserve space for a case with a separate output buffer though. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-22 6:27 ` Tony Lindgren @ 2026-09-23 0:37 ` Kishen Maloor 2026-09-23 6:50 ` Tony Lindgren 0 siblings, 1 reply; 84+ messages in thread From: Kishen Maloor @ 2026-09-23 0:37 UTC (permalink / raw) To: Tony Lindgren Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On 9/21/26 11:27 PM, Tony Lindgren wrote: > On Mon, Sep 21, 2026 at 08:56:48PM -0700, Kishen Maloor wrote: >> On 9/20/26 11:56 PM, Tony Lindgren wrote: >>> On Mon, Sep 21, 2026 at 08:58:17AM +0300, Tony Lindgren wrote: >>>> On Thu, Sep 17, 2026 at 09:33:18PM -0700, Kishen Maloor wrote: >>>>> On 8/31/26 12:13 AM, Tony Lindgren wrote: >>>>> >>>>> I was wondering how a vendor would use kvm_transfer_buffer for a command that >>>>> passes an input and also returns an output. size can describe the input >>>>> length or the space available for output, not both, so the kernel has no way to >>>>> know how much it may write. >>>>> >>>>> Two suggestions below. They're orthogonal. >>>>> >>>>>> ... >>>>>> diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h >>>>>> index 5f6c1ce9673b7..d9291a8a97bb1 100644 >>>>>> --- a/arch/x86/include/asm/kvm_host.h >>>>>> +++ b/arch/x86/include/asm/kvm_host.h >>>>>> @@ -2010,6 +2010,8 @@ struct kvm_x86_ops { >>>>>> int (*gmem_prepare)(struct kvm *kvm, kvm_pfn_t pfn, gfn_t gfn, int max_order); >>>>>> void (*gmem_invalidate)(kvm_pfn_t start, kvm_pfn_t end); >>>>>> int (*gmem_max_mapping_level)(struct kvm *kvm, kvm_pfn_t pfn, bool is_private); >>>>>> + bool (*cap_live_migration)(struct kvm *kvm); >>>>> >>>>> Should this return an 'int' specifying the maximum buffer size that the vendor impl >>>>> requires/uses? Userspace can then learn this once. >>>> >>>> OK that sounds good to me. I assume you are thinking the maximum hardware >>>> specific buffer size per thread? As in 512 * 4096 bytes for the TDX case? >> >> Yeah, an upper bound on the buffer size, so 516 * 4096 (4 for the GPA+MAC lists >> and MBMD) in TDX. It could be a hint to userspace to size its buffers at the >> start. Though it would be up to userspace to decide how to use that information. > > Oh right thanks. I think this is really the maximum transfer buffer size > Peter asked, not just a hint to userspace :) It is a strict upper bound. Maybe it's semantics, but I called it a "hint" because allocating for that entire size could be optional. If a userspace driver for say TDX wants to send smaller batches (say 128) then it can refer to the spec, do the math, and allocate 131 pages and the kernel should permit that; it's not wrong. If userspace ever allocates less room than a call requires, it would fail. A different userspace driver could simply allocate that max size and be done; no need to refer to the spec or do the math. It is for this second case where I thought returning the upper bound would be useful. Hence the earlier suggestion. > >>>>>> +struct kvm_transfer_buffer { >>>>>> + __u64 address; >>>>>> + __u32 size; >>>>>> + __u32 reserved; >>>>>> +}; >>>>> >>>>> Should this struct include a 'capacity' field (u32) that is set on each command? >>>>> It would be the number of bytes writable at address. >>>>> size would be the input length on entry (0 if the command passes none), and the >>>>> number of bytes produced on return (0 if none). >>>> >>>> Hmm so the transfer command return value can return how many bytes were >>>> written of the input. But yeah we don't know how many bytes were written >>>> back to the transfer buffer as result of the transfer command. >> >> The transfer command return value could return how many bytes were written into the buffer. >> But in an input-output call, the kernel handler wouldn't know how many bytes it could write, >> or for that matter even how many pages to pin up front in case it needs to return an output >> because 'size' couldn't simultaneously convey the input length and buffer capacity. That was >> the gap that I thought a read-only 'capacity' field could bridge. Of course, this >> assumes that the output is written in-place. > > Hmm yeah this inplace capacity vs transferred issue remains still. So I > agree we need to specify the capacity in struct kvm_transfer_buffer like > you suggested. To be clear, I think the in/out split for the buffers along with the convention I laid out closes that gap I saw without needing a 'capacity' field. Because out/size could now unambiguously convey capacity on entry and output length on return. > > To me size is already the size of the buffer though. So instead of changing > size to capacity, how about something like datasize or len for the input > and output transfer length? But I understand that (and please correct me if I'm wrong): a) You'd still prefer to not have 'size' serve that double duty. b) 'size' in your mental model already means buffer capacity. In that case, we could add a 'datasize' field to convey the length of valid data in the buffer, like this: struct kvm_transfer_buffer { __u64 address; __u32 size; __u32 datasize; __u64 reserved; }; The convention then becomes: - A non-zero 'datasize' on 'in' at call entry conveys that there is input. - A non-zero 'datasize' on 'out' at call exit conveys that there is output. - out/datasize on call entry is ignored. - in/size and out/size are seeded with the buffer capacity. > >>>> How about if we add the bytes returned to the transfer struct? Then the >>>> kvm_transfer_buffer can stay as just a buffer. >>> >>> Actually, for the possible cases with input+output, we could reserve space >>> in the transfer struct for another struct kvm_transfer_buffer for the >>> results? >> >> Yes, say an 'in' and 'out' kvm_transfer_buffer inside struct kvm_migrate_cmd should >> close this out and shouldn't require a 'capacity' field. Maybe we then establish this >> convention: >> - A non-zero 'size' on 'in' at call entry would signal that there is input. >> - A non-zero 'size' on 'out' at call entry would convey the buffer capacity. >> - A non-zero 'size' on 'out' at call exit would convey that there is output. >> - A zeroed 'size' on 'out' at call exit would convey that there is no output. > > Looks doable to me but with the inplace issue as above.. Sounds like we just > need to reserve space for a case with a separate output buffer though. To be clear, this is what I thought we were talking about :) To add a 2nd kvm_transfer_buffer to kvm_migrate_cmd, like this: struct kvm_migrate_cmd { __u16 command; __u16 flags; __u32 reserved; struct kvm_transfer_buffer in; struct kvm_transfer_buffer out; }; If there is agreement on this model, then yeah, we'd want to define these fields now, since the struct can't grow later without a new ioctl number. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-23 0:37 ` Kishen Maloor @ 2026-09-23 6:50 ` Tony Lindgren 2026-09-24 5:34 ` Kishen Maloor 0 siblings, 1 reply; 84+ messages in thread From: Tony Lindgren @ 2026-09-23 6:50 UTC (permalink / raw) To: Kishen Maloor Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Tue, Sep 22, 2026 at 05:37:33PM -0700, Kishen Maloor wrote: > On 9/21/26 11:27 PM, Tony Lindgren wrote: > > Oh right thanks. I think this is really the maximum transfer buffer size > > Peter asked, not just a hint to userspace :) > > It is a strict upper bound. Maybe it's semantics, but I called > it a "hint" because allocating for that entire size could be optional. > If a userspace driver for say TDX wants to send smaller batches (say 128) then > it can refer to the spec, do the math, and allocate 131 pages and the kernel > should permit that; it's not wrong. If userspace ever allocates less room than > a call requires, it would fail. A different userspace driver could simply > allocate that max size and be done; no need to refer to the spec or do the math. > It is for this second case where I thought returning the upper bound would be > useful. Hence the earlier suggestion. OK > >>>>>> +struct kvm_transfer_buffer { > >>>>>> + __u64 address; > >>>>>> + __u32 size; > >>>>>> + __u32 reserved; > >>>>>> +}; > >>>>> > >>>>> Should this struct include a 'capacity' field (u32) that is set on each command? > >>>>> It would be the number of bytes writable at address. > >>>>> size would be the input length on entry (0 if the command passes none), and the > >>>>> number of bytes produced on return (0 if none). > >>>> > >>>> Hmm so the transfer command return value can return how many bytes were > >>>> written of the input. But yeah we don't know how many bytes were written > >>>> back to the transfer buffer as result of the transfer command. > >> > >> The transfer command return value could return how many bytes were written into the buffer. > >> But in an input-output call, the kernel handler wouldn't know how many bytes it could write, > >> or for that matter even how many pages to pin up front in case it needs to return an output > >> because 'size' couldn't simultaneously convey the input length and buffer capacity. That was > >> the gap that I thought a read-only 'capacity' field could bridge. Of course, this > >> assumes that the output is written in-place. > > > > Hmm yeah this inplace capacity vs transferred issue remains still. So I > > agree we need to specify the capacity in struct kvm_transfer_buffer like > > you suggested. > > To be clear, I think the in/out split for the buffers along with the convention > I laid out closes that gap I saw without needing a 'capacity' field. > Because out/size could now unambiguously convey capacity on entry and output > length on return. But for an inplace buffer use with some input data smaller than the output data, would it work? To me it seems you need both buffer size and data size for that. > > To me size is already the size of the buffer though. So instead of changing > > size to capacity, how about something like datasize or len for the input > > and output transfer length? > > But I understand that (and please correct me if I'm wrong): > a) You'd still prefer to not have 'size' serve that double duty. > b) 'size' in your mental model already means buffer capacity. Heh yes correct for the above. > In that case, we could add a 'datasize' field to convey the length > of valid data in the buffer, like this: > > struct kvm_transfer_buffer { > __u64 address; > __u32 size; > __u32 datasize; > __u64 reserved; > }; Maybe bufsize and datasize? Then the difference would be obvious while reading the code. > The convention then becomes: > - A non-zero 'datasize' on 'in' at call entry conveys that there is input. > - A non-zero 'datasize' on 'out' at call exit conveys that there is output. > - out/datasize on call entry is ignored. > - in/size and out/size are seeded with the buffer capacity. > > > > >>>> How about if we add the bytes returned to the transfer struct? Then the > >>>> kvm_transfer_buffer can stay as just a buffer. > >>> > >>> Actually, for the possible cases with input+output, we could reserve space > >>> in the transfer struct for another struct kvm_transfer_buffer for the > >>> results? > >> > >> Yes, say an 'in' and 'out' kvm_transfer_buffer inside struct kvm_migrate_cmd should > >> close this out and shouldn't require a 'capacity' field. Maybe we then establish this > >> convention: > >> - A non-zero 'size' on 'in' at call entry would signal that there is input. > >> - A non-zero 'size' on 'out' at call entry would convey the buffer capacity. > >> - A non-zero 'size' on 'out' at call exit would convey that there is output. > >> - A zeroed 'size' on 'out' at call exit would convey that there is no output. > > > > Looks doable to me but with the inplace issue as above.. Sounds like we just > > need to reserve space for a case with a separate output buffer though. > > To be clear, this is what I thought we were talking about :) Heh yeah we're talking two things with the inplace use vs two buffers :) > To add a 2nd kvm_transfer_buffer to kvm_migrate_cmd, like this: > > struct kvm_migrate_cmd { > __u16 command; > __u16 flags; > __u32 reserved; > struct kvm_transfer_buffer in; > struct kvm_transfer_buffer out; > }; > > If there is agreement on this model, then yeah, we'd want to define > these fields now, since the struct can't grow later without a new ioctl > number. Based on what we've discussed, my preference is the following: Keep the current buf naming. For the EXPORT/IMPORT type functions the use should be obvious from the transfer type. Reserve enough space for a separate output buffer or results buffer or whatever it might get called if such a use case ever pops up. Add the datasize to struct kvm_transfer_buffer like you suggested and rename size to bufsize. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-23 6:50 ` Tony Lindgren @ 2026-09-24 5:34 ` Kishen Maloor 2026-09-24 7:15 ` Tony Lindgren 0 siblings, 1 reply; 84+ messages in thread From: Kishen Maloor @ 2026-09-24 5:34 UTC (permalink / raw) To: Tony Lindgren Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On 9/22/26 11:50 PM, Tony Lindgren wrote: > On Tue, Sep 22, 2026 at 05:37:33PM -0700, Kishen Maloor wrote: >> On 9/21/26 11:27 PM, Tony Lindgren wrote: >>> Oh right thanks. I think this is really the maximum transfer buffer size >>> Peter asked, not just a hint to userspace :) >> >> It is a strict upper bound. Maybe it's semantics, but I called >> it a "hint" because allocating for that entire size could be optional. >> If a userspace driver for say TDX wants to send smaller batches (say 128) then >> it can refer to the spec, do the math, and allocate 131 pages and the kernel >> should permit that; it's not wrong. If userspace ever allocates less room than >> a call requires, it would fail. A different userspace driver could simply >> allocate that max size and be done; no need to refer to the spec or do the math. >> It is for this second case where I thought returning the upper bound would be >> useful. Hence the earlier suggestion. > > OK > >>>>>>>> +struct kvm_transfer_buffer { >>>>>>>> + __u64 address; >>>>>>>> + __u32 size; >>>>>>>> + __u32 reserved; >>>>>>>> +}; >>>>>>> >>>>>>> Should this struct include a 'capacity' field (u32) that is set on each command? >>>>>>> It would be the number of bytes writable at address. >>>>>>> size would be the input length on entry (0 if the command passes none), and the >>>>>>> number of bytes produced on return (0 if none). >>>>>> >>>>>> Hmm so the transfer command return value can return how many bytes were >>>>>> written of the input. But yeah we don't know how many bytes were written >>>>>> back to the transfer buffer as result of the transfer command. >>>> >>>> The transfer command return value could return how many bytes were written into the buffer. >>>> But in an input-output call, the kernel handler wouldn't know how many bytes it could write, >>>> or for that matter even how many pages to pin up front in case it needs to return an output >>>> because 'size' couldn't simultaneously convey the input length and buffer capacity. That was >>>> the gap that I thought a read-only 'capacity' field could bridge. Of course, this >>>> assumes that the output is written in-place. >>> >>> Hmm yeah this inplace capacity vs transferred issue remains still. So I >>> agree we need to specify the capacity in struct kvm_transfer_buffer like >>> you suggested. >> >> To be clear, I think the in/out split for the buffers along with the convention >> I laid out closes that gap I saw without needing a 'capacity' field. >> Because out/size could now unambiguously convey capacity on entry and output >> length on return. > > But for an inplace buffer use with some input data smaller than the output > data, would it work? To me it seems you need both buffer size and data > size for that. Yes, with two kvm_transfer_buffers in the transfer struct, it would work, whether the call uses a single userspace buffer or two. > >>> To me size is already the size of the buffer though. So instead of changing >>> size to capacity, how about something like datasize or len for the input >>> and output transfer length? >> >> But I understand that (and please correct me if I'm wrong): >> a) You'd still prefer to not have 'size' serve that double duty. >> b) 'size' in your mental model already means buffer capacity. > > Heh yes correct for the above. > >> In that case, we could add a 'datasize' field to convey the length >> of valid data in the buffer, like this: >> >> struct kvm_transfer_buffer { >> __u64 address; >> __u32 size; >> __u32 datasize; >> __u64 reserved; >> }; > > Maybe bufsize and datasize? Then the difference would be obvious while > reading the code. Sure. > >> The convention then becomes: >> - A non-zero 'datasize' on 'in' at call entry conveys that there is input. >> - A non-zero 'datasize' on 'out' at call exit conveys that there is output. >> - out/datasize on call entry is ignored. >> - in/size and out/size are seeded with the buffer capacity. >> >>> >>>>>> How about if we add the bytes returned to the transfer struct? Then the >>>>>> kvm_transfer_buffer can stay as just a buffer. >>>>> >>>>> Actually, for the possible cases with input+output, we could reserve space >>>>> in the transfer struct for another struct kvm_transfer_buffer for the >>>>> results? >>>> >>>> Yes, say an 'in' and 'out' kvm_transfer_buffer inside struct kvm_migrate_cmd should >>>> close this out and shouldn't require a 'capacity' field. Maybe we then establish this >>>> convention: >>>> - A non-zero 'size' on 'in' at call entry would signal that there is input. >>>> - A non-zero 'size' on 'out' at call entry would convey the buffer capacity. >>>> - A non-zero 'size' on 'out' at call exit would convey that there is output. >>>> - A zeroed 'size' on 'out' at call exit would convey that there is no output. >>> >>> Looks doable to me but with the inplace issue as above.. Sounds like we just >>> need to reserve space for a case with a separate output buffer though. >> >> To be clear, this is what I thought we were talking about :) > > Heh yeah we're talking two things with the inplace use vs two buffers :) > >> To add a 2nd kvm_transfer_buffer to kvm_migrate_cmd, like this: >> >> struct kvm_migrate_cmd { >> __u16 command; >> __u16 flags; >> __u32 reserved; >> struct kvm_transfer_buffer in; >> struct kvm_transfer_buffer out; >> }; >> >> If there is agreement on this model, then yeah, we'd want to define >> these fields now, since the struct can't grow later without a new ioctl >> number. > > Based on what we've discussed, my preference is the following: > > Keep the current buf naming. For the EXPORT/IMPORT type functions the use > should be obvious from the transfer type. > > Reserve enough space for a separate output buffer or results buffer or > whatever it might get called if such a use case ever pops up. Reserved space can be named later without changing sizeof, so either way works. I'd mildly prefer to declare the second kvm_transfer_buffer just because declaring both would settle the second buffer's semantics now. Not something I'd push hard on though if you prefer reserving. > Add the datasize to struct kvm_transfer_buffer like you suggested and > rename size to bufsize. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-24 5:34 ` Kishen Maloor @ 2026-09-24 7:15 ` Tony Lindgren 2026-10-08 9:22 ` Tony Lindgren 0 siblings, 1 reply; 84+ messages in thread From: Tony Lindgren @ 2026-09-24 7:15 UTC (permalink / raw) To: Kishen Maloor Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Wed, Sep 23, 2026 at 10:34:14PM -0700, Kishen Maloor wrote: > On 9/22/26 11:50 PM, Tony Lindgren wrote: > > Based on what we've discussed, my preference is the following: > > > > Keep the current buf naming. For the EXPORT/IMPORT type functions the use > > should be obvious from the transfer type. > > > > Reserve enough space for a separate output buffer or results buffer or > > whatever it might get called if such a use case ever pops up. > > Reserved space can be named later without changing sizeof, so either way > works. I'd mildly prefer to declare the second kvm_transfer_buffer just > because declaring both would settle the second buffer's semantics now. Not > something I'd push hard on though if you prefer reserving. Let's just use buf and reserved space then. There is no known usecase needing separate in and out buffers for EXPORT or IMPORT. And the separate in and out buffers would have to be needed the same time rather than first in and then out.. > > Add the datasize to struct kvm_transfer_buffer like you suggested and > > rename size to bufsize. And to recap, with the bufsize and datasize in the buffer, the needs we discussed for separate in and out buffers went away. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD 2026-09-24 7:15 ` Tony Lindgren @ 2026-10-08 9:22 ` Tony Lindgren 0 siblings, 0 replies; 84+ messages in thread From: Tony Lindgren @ 2026-10-08 9:22 UTC (permalink / raw) To: Kishen Maloor Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm, mathias.brossard On Thu, Sep 24, 2026 at 10:15:17AM +0300, Tony Lindgren wrote: > On Wed, Sep 23, 2026 at 10:34:14PM -0700, Kishen Maloor wrote: > > On 9/22/26 11:50 PM, Tony Lindgren wrote: > > > Based on what we've discussed, my preference is the following: > > > > > > Keep the current buf naming. For the EXPORT/IMPORT type functions the use > > > should be obvious from the transfer type. > > > > > > Reserve enough space for a separate output buffer or results buffer or > > > whatever it might get called if such a use case ever pops up. > > > > Reserved space can be named later without changing sizeof, so either way > > works. I'd mildly prefer to declare the second kvm_transfer_buffer just > > because declaring both would settle the second buffer's semantics now. Not > > something I'd push hard on though if you prefer reserving. > > Let's just use buf and reserved space then. There is no known usecase > needing separate in and out buffers for EXPORT or IMPORT. And the separate > in and out buffers would have to be needed the same time rather than first > in and then out.. > > > > Add the datasize to struct kvm_transfer_buffer like you suggested and > > > rename size to bufsize. > > And to recap, with the bufsize and datasize in the buffer, the needs we > discussed for separate in and out buffers went away. Based on the ARM CCA live migration presentation at LPC by Mathias Brossard, we now have a user for a separate metadata buffer. Adding Mathias to Cc. Page 8 of the presentation at [0] below lists it under feedback. And see also page 10 for feedback. Probably best to make the buffer a userspace array with an arch specific define for the number of buffers. So far looks like two buffers for ARM, and one might be enough for x86. Looks like we should have a separate SETUP command to query the buffer sizes for cmd, memory and vCPU transfers. [0] https://lpc.events/event/20/contributions/2476/attachments/2235/4909/LPC2026_CCA_Live_Migration.pdf ^ permalink raw reply [flat|nested] 84+ messages in thread
* [RFC PATCH v2 3/4] KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY 2026-08-31 7:13 [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Tony Lindgren 2026-08-31 7:13 ` [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests Tony Lindgren 2026-08-31 7:13 ` [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD Tony Lindgren @ 2026-08-31 7:13 ` Tony Lindgren 2026-08-31 7:23 ` sashiko-bot 2026-08-31 7:13 ` [RFC PATCH v2 4/4] KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU Tony Lindgren ` (3 subsequent siblings) 6 siblings, 1 reply; 84+ messages in thread From: Tony Lindgren @ 2026-08-31 7:13 UTC (permalink / raw) To: Paolo Bonzini, Sean Christopherson Cc: Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel , Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm Add support to export and import KVM memory for cases where the memory is only accessible to the guest. Live migration of confidential computing needs help of KVM for the vendor specific calls at least for TDX. Introduce optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY. Based on earlier code by Wei Wang <wei.w.wang@intel.com>. Co-developed-by: Kishen Maloor <kishen.maloor@intel.com> Signed-off-by: Kishen Maloor <kishen.maloor@intel.com> Assisted-by: Claude-Code:claude-opus-5 checkpatch [ used AI to review and simplify the code ] Signed-off-by: Tony Lindgren <tony.lindgren@linux.intel.com> --- arch/x86/include/asm/kvm-x86-ops.h | 2 ++ arch/x86/include/asm/kvm_host.h | 2 ++ arch/x86/kvm/x86.c | 37 ++++++++++++++++++++++++++++++ include/uapi/linux/kvm.h | 13 +++++++++++ 4 files changed, 54 insertions(+) diff --git a/arch/x86/include/asm/kvm-x86-ops.h b/arch/x86/include/asm/kvm-x86-ops.h index ac080b556b0c8..173d0c4f1115e 100644 --- a/arch/x86/include/asm/kvm-x86-ops.h +++ b/arch/x86/include/asm/kvm-x86-ops.h @@ -150,6 +150,8 @@ KVM_X86_OP_OPTIONAL_RET0(gmem_max_mapping_level) KVM_X86_OP_OPTIONAL(gmem_invalidate) KVM_X86_OP_OPTIONAL_RET0(cap_live_migration) KVM_X86_OP_OPTIONAL(migrate_cmd) +KVM_X86_OP_OPTIONAL(export_memory) +KVM_X86_OP_OPTIONAL(import_memory) #endif #undef KVM_X86_OP diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h index d9291a8a97bb1..9a517bfc2f3a6 100644 --- a/arch/x86/include/asm/kvm_host.h +++ b/arch/x86/include/asm/kvm_host.h @@ -2012,6 +2012,8 @@ struct kvm_x86_ops { int (*gmem_max_mapping_level)(struct kvm *kvm, kvm_pfn_t pfn, bool is_private); bool (*cap_live_migration)(struct kvm *kvm); int (*migrate_cmd)(struct kvm *kvm, struct kvm_migrate_cmd *cmd); + int (*export_memory)(struct kvm *kvm, struct kvm_memory_transfer *mem); + int (*import_memory)(struct kvm *kvm, struct kvm_memory_transfer *mem); }; struct kvm_x86_nested_ops { diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c index 7064fd709e56d..8a99c665008a3 100644 --- a/arch/x86/kvm/x86.c +++ b/arch/x86/kvm/x86.c @@ -7258,6 +7258,37 @@ long kvm_arch_vcpu_unlocked_ioctl(struct file *filp, unsigned int ioctl, return -ENOIOCTLCMD; } +static int kvm_vm_ioctl_transfer_memory(struct kvm *kvm, bool import, + void __user *argp) +{ + struct kvm_memory_transfer mem; + int r; + + if (!kvm_x86_call(cap_live_migration)(kvm) || + (import && !kvm_x86_ops.import_memory) || + (!import && !kvm_x86_ops.export_memory)) + return -ENOTTY; + + if (copy_from_user(&mem, argp, sizeof(mem))) + return -EFAULT; + + if (mem.reserved || mem.buf.reserved || !mem.nr_gfns) + return -EINVAL; + + if (import) + r = kvm_x86_call(import_memory)(kvm, &mem); + else + r = kvm_x86_call(export_memory)(kvm, &mem); + if (r > 0) + r = -EIO; + + /* Copy back also on an error to report a partially done transfer */ + if (copy_to_user(argp, &mem, sizeof(mem))) + return -EFAULT; + + return r; +} + int kvm_arch_vm_ioctl(struct file *filp, unsigned int ioctl, unsigned long arg) { struct kvm *kvm = filp->private_data; @@ -7639,6 +7670,12 @@ int kvm_arch_vm_ioctl(struct file *filp, unsigned int ioctl, unsigned long arg) return -EFAULT; break; } + case KVM_EXPORT_MEMORY: + r = kvm_vm_ioctl_transfer_memory(kvm, false, argp); + break; + case KVM_IMPORT_MEMORY: + r = kvm_vm_ioctl_transfer_memory(kvm, true, argp); + break; default: r = -ENOTTY; } diff --git a/include/uapi/linux/kvm.h b/include/uapi/linux/kvm.h index e5b227d3a8e40..666bbdf220d65 100644 --- a/include/uapi/linux/kvm.h +++ b/include/uapi/linux/kvm.h @@ -1493,6 +1493,10 @@ struct kvm_enc_region { #define KVM_GET_SREGS2 _IOR(KVMIO, 0xcc, struct kvm_sregs2) #define KVM_SET_SREGS2 _IOW(KVMIO, 0xcd, struct kvm_sregs2) +/* Available with KVM_CAP_LIVE_MIGRATION */ +#define KVM_EXPORT_MEMORY _IOWR(KVMIO, 0xe5, struct kvm_memory_transfer) +#define KVM_IMPORT_MEMORY _IOWR(KVMIO, 0xe6, struct kvm_memory_transfer) + #define KVM_DIRTY_LOG_MANUAL_PROTECT_ENABLE (1 << 0) #define KVM_DIRTY_LOG_INITIALLY_SET (1 << 1) @@ -1692,4 +1696,13 @@ struct kvm_migrate_cmd { struct kvm_transfer_buffer buf; }; +struct kvm_memory_transfer { + __u64 gfns; + __u32 nr_gfns; + __u16 id; + __u16 flags; + __u64 reserved; + struct kvm_transfer_buffer buf; +}; + #endif /* __LINUX_KVM_H */ -- 2.43.0 ^ permalink raw reply related [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 3/4] KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY 2026-08-31 7:13 ` [RFC PATCH v2 3/4] KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY Tony Lindgren @ 2026-08-31 7:23 ` sashiko-bot 2026-09-01 6:10 ` Tony Lindgren 0 siblings, 1 reply; 84+ messages in thread From: sashiko-bot @ 2026-08-31 7:23 UTC (permalink / raw) To: Tony Lindgren; +Cc: kvm Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider: - [Medium] Unvalidated UAPI fields `flags` and `id` in `struct kvm_memory_transfer` can lead to ABI breakage. -- commit b6f49e80a58eebaaa2a6d123950fd36709972a10 Author: Tony Lindgren <tony.lindgren@linux.intel.com> KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY This commit adds support to export and import KVM memory for cases where the memory is only accessible to the guest. It introduces the optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY ioctls to assist with live migration of confidential computing. > diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c > index 7064fd709e56d..8a99c665008a3 100644 > --- a/arch/x86/kvm/x86.c > +++ b/arch/x86/kvm/x86.c [ ... ] > @@ -7258,6 +7258,37 @@ long kvm_arch_vcpu_unlocked_ioctl(struct file *filp, unsigned int ioctl, > return -ENOIOCTLCMD; > } > > +static int kvm_vm_ioctl_transfer_memory(struct kvm *kvm, bool import, > + void __user *argp) > +{ > + struct kvm_memory_transfer mem; > + int r; > + > + if (!kvm_x86_call(cap_live_migration)(kvm) || > + (import && !kvm_x86_ops.import_memory) || > + (!import && !kvm_x86_ops.export_memory)) > + return -ENOTTY; > + > + if (copy_from_user(&mem, argp, sizeof(mem))) > + return -EFAULT; > + > + if (mem.reserved || mem.buf.reserved || !mem.nr_gfns) > + return -EINVAL; [Severity: Medium] Does this input validation for the kvm_memory_transfer UAPI struct miss checking mem.flags == 0 and mem.id == 0? Failing to validate that unused UAPI fields are zero allows userspace to pass uninitialized garbage. If KVM later assigns specific meaning to these fields, older applications that unintentionally passed non-zero values could break, leading to an ABI compatibility issue. Standard Linux UAPI design usually requires enforcing these unused or reserved fields to be zero for future extensibility. -- Sashiko AI review · https://sashiko.dev/#/patchset/20260831071304.762939-1-tony.lindgren@linux.intel.com?part=3 ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 3/4] KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY 2026-08-31 7:23 ` sashiko-bot @ 2026-09-01 6:10 ` Tony Lindgren 0 siblings, 0 replies; 84+ messages in thread From: Tony Lindgren @ 2026-09-01 6:10 UTC (permalink / raw) To: sashiko-reviews; +Cc: kvm On Mon, Aug 31, 2026 at 07:23:47AM +0000, sashiko-bot@kernel.org wrote: > commit b6f49e80a58eebaaa2a6d123950fd36709972a10 > Author: Tony Lindgren <tony.lindgren@linux.intel.com> > > --- a/arch/x86/kvm/x86.c > > +++ b/arch/x86/kvm/x86.c > [ ... ] > > @@ -7258,6 +7258,37 @@ long kvm_arch_vcpu_unlocked_ioctl(struct file *filp, unsigned int ioctl, > > return -ENOIOCTLCMD; > > } > > > > +static int kvm_vm_ioctl_transfer_memory(struct kvm *kvm, bool import, > > + void __user *argp) > > +{ > > + struct kvm_memory_transfer mem; > > + int r; > > + > > + if (!kvm_x86_call(cap_live_migration)(kvm) || > > + (import && !kvm_x86_ops.import_memory) || > > + (!import && !kvm_x86_ops.export_memory)) > > + return -ENOTTY; > > + > > + if (copy_from_user(&mem, argp, sizeof(mem))) > > + return -EFAULT; > > + > > + if (mem.reserved || mem.buf.reserved || !mem.nr_gfns) > > + return -EINVAL; > > [Severity: Medium] > Does this input validation for the kvm_memory_transfer UAPI struct miss > checking mem.flags == 0 and mem.id == 0? Yes flags is unused for memory transfers at least for TDX. The id TDX can eventually use for the thread index for QEMU multifd transfers. ^ permalink raw reply [flat|nested] 84+ messages in thread
* [RFC PATCH v2 4/4] KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU 2026-08-31 7:13 [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Tony Lindgren ` (2 preceding siblings ...) 2026-08-31 7:13 ` [RFC PATCH v2 3/4] KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY Tony Lindgren @ 2026-08-31 7:13 ` Tony Lindgren 2026-08-31 7:23 ` sashiko-bot 2026-09-04 18:24 ` [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Artem Bityutskiy ` (2 subsequent siblings) 6 siblings, 1 reply; 84+ messages in thread From: Tony Lindgren @ 2026-08-31 7:13 UTC (permalink / raw) To: Paolo Bonzini, Sean Christopherson Cc: Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel , Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm Add support to export and import VCPU for cases where the VCPU state is only accessible to the guest. Live migration of confidential computing needs help of KVM for the firmware specific calls at least for TDX. Introduce optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU. Based on earlier code by Wei Wang <wei.w.wang@intel.com>. Co-developed-by: Kishen Maloor <kishen.maloor@intel.com> Signed-off-by: Kishen Maloor <kishen.maloor@intel.com> Signed-off-by: Tony Lindgren <tony.lindgren@linux.intel.com> --- arch/x86/include/asm/kvm-x86-ops.h | 2 ++ arch/x86/include/asm/kvm_host.h | 2 ++ arch/x86/kvm/x86.c | 40 ++++++++++++++++++++++++++++++ include/uapi/linux/kvm.h | 8 ++++++ 4 files changed, 52 insertions(+) diff --git a/arch/x86/include/asm/kvm-x86-ops.h b/arch/x86/include/asm/kvm-x86-ops.h index 173d0c4f1115e..7f110f80d6f82 100644 --- a/arch/x86/include/asm/kvm-x86-ops.h +++ b/arch/x86/include/asm/kvm-x86-ops.h @@ -152,6 +152,8 @@ KVM_X86_OP_OPTIONAL_RET0(cap_live_migration) KVM_X86_OP_OPTIONAL(migrate_cmd) KVM_X86_OP_OPTIONAL(export_memory) KVM_X86_OP_OPTIONAL(import_memory) +KVM_X86_OP_OPTIONAL(export_vcpu) +KVM_X86_OP_OPTIONAL(import_vcpu) #endif #undef KVM_X86_OP diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h index 9a517bfc2f3a6..b6362408dab80 100644 --- a/arch/x86/include/asm/kvm_host.h +++ b/arch/x86/include/asm/kvm_host.h @@ -2014,6 +2014,8 @@ struct kvm_x86_ops { int (*migrate_cmd)(struct kvm *kvm, struct kvm_migrate_cmd *cmd); int (*export_memory)(struct kvm *kvm, struct kvm_memory_transfer *mem); int (*import_memory)(struct kvm *kvm, struct kvm_memory_transfer *mem); + int (*export_vcpu)(struct kvm_vcpu *vcpu, struct kvm_vcpu_transfer *vcpu_state); + int (*import_vcpu)(struct kvm_vcpu *vcpu, struct kvm_vcpu_transfer *vcpu_state); }; struct kvm_x86_nested_ops { diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c index 8a99c665008a3..e8385326894b1 100644 --- a/arch/x86/kvm/x86.c +++ b/arch/x86/kvm/x86.c @@ -6189,6 +6189,38 @@ static int kvm_get_reg_list(struct kvm_vcpu *vcpu, return 0; } +static int kvm_vcpu_ioctl_transfer_vcpu(struct kvm_vcpu *vcpu, bool import, + void __user *argp) +{ + struct kvm_vcpu_transfer vcpu_state; + struct kvm *kvm = vcpu->kvm; + int r; + + if (!kvm_x86_call(cap_live_migration)(kvm) || + (import && !kvm_x86_ops.import_vcpu) || + (!import && !kvm_x86_ops.export_vcpu)) + return -ENOTTY; + + if (copy_from_user(&vcpu_state, argp, sizeof(vcpu_state))) + return -EFAULT; + + if (vcpu_state.reserved || vcpu_state.buf.reserved) + return -EINVAL; + + if (import) + r = kvm_x86_call(import_vcpu)(vcpu, &vcpu_state); + else + r = kvm_x86_call(export_vcpu)(vcpu, &vcpu_state); + if (r > 0) + r = -EIO; + + /* Copy back also on an error to report a partially done transfer */ + if (copy_to_user(argp, &vcpu_state, sizeof(vcpu_state))) + r = -EFAULT; + + return r; +} + long kvm_arch_vcpu_ioctl(struct file *filp, unsigned int ioctl, unsigned long arg) { @@ -6659,6 +6691,14 @@ long kvm_arch_vcpu_ioctl(struct file *filp, goto out; r = kvm_x86_ops.vcpu_mem_enc_ioctl(vcpu, argp); break; + case KVM_EXPORT_VCPU: { + r = kvm_vcpu_ioctl_transfer_vcpu(vcpu, false, argp); + break; + } + case KVM_IMPORT_VCPU: { + r = kvm_vcpu_ioctl_transfer_vcpu(vcpu, true, argp); + break; + } default: r = -EINVAL; } diff --git a/include/uapi/linux/kvm.h b/include/uapi/linux/kvm.h index 666bbdf220d65..0a9aa126daadb 100644 --- a/include/uapi/linux/kvm.h +++ b/include/uapi/linux/kvm.h @@ -1496,6 +1496,8 @@ struct kvm_enc_region { /* Available with KVM_CAP_LIVE_MIGRATION */ #define KVM_EXPORT_MEMORY _IOWR(KVMIO, 0xe5, struct kvm_memory_transfer) #define KVM_IMPORT_MEMORY _IOWR(KVMIO, 0xe6, struct kvm_memory_transfer) +#define KVM_EXPORT_VCPU _IOWR(KVMIO, 0xe7, struct kvm_vcpu_transfer) +#define KVM_IMPORT_VCPU _IOWR(KVMIO, 0xe8, struct kvm_vcpu_transfer) #define KVM_DIRTY_LOG_MANUAL_PROTECT_ENABLE (1 << 0) #define KVM_DIRTY_LOG_INITIALLY_SET (1 << 1) @@ -1705,4 +1707,10 @@ struct kvm_memory_transfer { struct kvm_transfer_buffer buf; }; +struct kvm_vcpu_transfer { + __u32 flags; + __u32 reserved; + struct kvm_transfer_buffer buf; +}; + #endif /* __LINUX_KVM_H */ -- 2.43.0 ^ permalink raw reply related [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 4/4] KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU 2026-08-31 7:13 ` [RFC PATCH v2 4/4] KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU Tony Lindgren @ 2026-08-31 7:23 ` sashiko-bot 2026-09-01 6:12 ` Tony Lindgren 0 siblings, 1 reply; 84+ messages in thread From: sashiko-bot @ 2026-08-31 7:23 UTC (permalink / raw) To: Tony Lindgren; +Cc: kvm Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider: - [High] Missing validation of the flags field in UAPI struct kvm_vcpu_transfer breaks forward compatibility. -- commit 4d7dccde0750ca7fd34158829ad4120d967ac2c3 Author: Tony Lindgren <tony.lindgren@linux.intel.com> KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU This commit adds support to export and import a VCPU for cases where the VCPU state is only accessible to the guest. It introduces the new KVM_EXPORT_VCPU and KVM_IMPORT_VCPU ioctls, along with the required kvm_vcpu_transfer structure. > diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c > index 8a99c665008a3..e8385326894b1 100644 > --- a/arch/x86/kvm/x86.c > +++ b/arch/x86/kvm/x86.c [ ... ] > @@ -6189,6 +6189,38 @@ static int kvm_get_reg_list(struct kvm_vcpu *vcpu, > return 0; > } > > +static int kvm_vcpu_ioctl_transfer_vcpu(struct kvm_vcpu *vcpu, bool import, > + void __user *argp) > +{ > + struct kvm_vcpu_transfer vcpu_state; > + struct kvm *kvm = vcpu->kvm; > + int r; [ ... ] > + if (copy_from_user(&vcpu_state, argp, sizeof(vcpu_state))) > + return -EFAULT; > + > + if (vcpu_state.reserved || vcpu_state.buf.reserved) > + return -EINVAL; [Severity: High] Should we also validate the flags field here to ensure forward compatibility? The new kvm_vcpu_transfer structure introduces a flags field, but this validation step only explicitly rejects non-zero reserved fields. If the kernel silently ignores non-zero flags, userspace might inadvertently pass uninitialized or arbitrary values without receiving an error. If KVM later assigns meaning to these flags, old userspace programs that have been unknowingly passing garbage could unexpectedly trigger new behaviors or break. Would it be appropriate to require that vcpu_state.flags is zero for now? > + > + if (import) > + r = kvm_x86_call(import_vcpu)(vcpu, &vcpu_state); [ ... ] -- Sashiko AI review · https://sashiko.dev/#/patchset/20260831071304.762939-1-tony.lindgren@linux.intel.com?part=4 ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 4/4] KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU 2026-08-31 7:23 ` sashiko-bot @ 2026-09-01 6:12 ` Tony Lindgren 0 siblings, 0 replies; 84+ messages in thread From: Tony Lindgren @ 2026-09-01 6:12 UTC (permalink / raw) To: sashiko-reviews; +Cc: kvm On Mon, Aug 31, 2026 at 07:23:53AM +0000, sashiko-bot@kernel.org wrote: > commit 4d7dccde0750ca7fd34158829ad4120d967ac2c3 > Author: Tony Lindgren <tony.lindgren@linux.intel.com> > > --- a/arch/x86/kvm/x86.c > > +++ b/arch/x86/kvm/x86.c > > [ ... ] > > > @@ -6189,6 +6189,38 @@ static int kvm_get_reg_list(struct kvm_vcpu *vcpu, > > return 0; > > } > > > > +static int kvm_vcpu_ioctl_transfer_vcpu(struct kvm_vcpu *vcpu, bool import, > > + void __user *argp) > > +{ > > + struct kvm_vcpu_transfer vcpu_state; > > + struct kvm *kvm = vcpu->kvm; > > + int r; > > [ ... ] > > > + if (copy_from_user(&vcpu_state, argp, sizeof(vcpu_state))) > > + return -EFAULT; > > + > > + if (vcpu_state.reserved || vcpu_state.buf.reserved) > > + return -EINVAL; > > [Severity: High] > Should we also validate the flags field here to ensure forward compatibility? > > The new kvm_vcpu_transfer structure introduces a flags field, but this > validation step only explicitly rejects non-zero reserved fields. > > If the kernel silently ignores non-zero flags, userspace might > inadvertently pass uninitialized or arbitrary values without receiving > an error. If KVM later assigns meaning to these flags, old userspace > programs that have been unknowingly passing garbage could unexpectedly > trigger new behaviors or break. > > Would it be appropriate to require that vcpu_state.flags is zero for now? Yes flags is unused for vCPU transfers at least for TDX and can be zero for now. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-08-31 7:13 [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Tony Lindgren ` (3 preceding siblings ...) 2026-08-31 7:13 ` [RFC PATCH v2 4/4] KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU Tony Lindgren @ 2026-09-04 18:24 ` Artem Bityutskiy 2026-09-17 21:27 ` Peter Xu 2026-09-18 18:36 ` Ionut Mihalcea 2026-09-25 16:03 ` Serge Hallyn (AMD) 6 siblings, 1 reply; 84+ messages in thread From: Artem Bityutskiy @ 2026-09-04 18:24 UTC (permalink / raw) To: Tony Lindgren, Paolo Bonzini, Sean Christopherson Cc: Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Mon, 2026-08-31 at 10:13 +0300, Tony Lindgren wrote: > Tom, since you mentioned that AMD SEV-SNP and Intel TDX live migration > sound similar, can you please take a look how the API might work for > SEV-SNP? The presented example uAPIs were designed to fit the TDX live migration flow, with the intent that they could also be used by other CoCo VMs. It would be super great if someone could commend on how suitable these uAPIs are for AMD and ARM flows. AFAIU, what is generic in the presented uAPI is more of usage pattern. - The order in which QEMU calls them. - The idea that QEMU/KVM is a transport for opaque blobs. The proposed container - 'struct kvm_transfer_buffer' - only has address and a size. The vendor defines the layout of the data in it. Some example high-level topics that would be nice to get feedback on: - Can we come up with a single set of generic migration uAPIs for different CoCo models? - Or should some uAPIs be generic while others are vendor-specific? - Or should each CoCo model have its own vendor-specific set of migration uAPIs? - Should the same uAPIs also support traditional VMs? But the only use-case I imagine here is "for testing purposes". > Artem has put together a brief description below of the example API and the > migration flow: .. snip ... > Migration flow > ============== > > Source host Destination host > =========== ================ > > CMD(SETUP/SESSION) <--- setup msgs ---> CMD(SETUP/SESSION) > | (repeated) | > CMD(SETUP/IMMUTABLE_STATE) - immutable state -> CMD(SETUP/IMMUTABLE_STATE) > | | > KVM_GET_DIRTY_LOG | > KVM_EXPORT_MEMORY --- memory data ---> KVM_IMPORT_MEMORY > CMD(ITERATION) --- epoch token ---> CMD(ITERATION) > | (repeat until convergence) | > CMD(STOP_AND_COPY/PAUSE) | > CMD(STOP_AND_COPY/TD_STATE) --- VM state ------> CMD(STOP_AND_COPY/TD_STATE) > KVM_EXPORT_VCPU --- vCPU state ----> KVM_IMPORT_VCPU > KVM_EXPORT_MEMORY -- final memory ---> KVM_IMPORT_MEMORY > CMD(ITERATION/DONE) --- start token ---> CMD(ITERATION) > | | > CMD(END) CMD(END) ... snip ... As I mentioned, the example uAPI is modeled around the TDX migration flow. In case it helps the reader, here is a summary of that flow that was presented in PUCK. It describes what the TDX module offers today and focuses on pre-copy migration. This is our interpretation of the TDX specifications, not a specification itself, and may contain errors or omissions. Please refer to the official TDX specifications for authoritative information. Migration Overview ------------------ In TDX, the TDX module implements the migration logic. The VMM drives migration by issuing seamcalls to the source and destination TDX modules and transports the resulting encrypted blobs between them. The VMM does not need to know the contents of those blobs. The TDX security model enforces two hard rules: - Only one instance of the TD may run at a time. Either the source or the destination may run, but never both. In other words, cloning a TD is not allowed. - When migration completes, the destination must have the same memory and vCPU state as the source. It must not end up with a partial or mixed state. Simple Overview --------------- Run the migration setup session | v Transfer immutable TD state | v Copy dirty memory pages while the TD runs <-------+ | | +-- iterate until convergence criteria is met --+ | v Pause the source TD and copy the remaining state | v Start the destination TD The migration process starts with a setup session. During the setup session, the source and destination TDX modules exchange encrypted blobs and establish migration encryption keys. For example, these keys protect memory contents transferred during migration. Next, the source transfers the immutable TD state to the destination and initializes the destination TD. This state includes the read-only VM data, such as its vCPU count and topology. The source and destination then perform iterative memory-copy rounds. In each round, the source scans for memory pages that need to be migrated and exports them in encrypted form. The destination imports the pages. The rounds continue until the number of pages that change between rounds is small enough to meet the convergence criteria. The source TD is then paused. The VMM transfers the remaining TD state, such as vCPU state, and the final dirty memory pages. Finally, the destination TD is started and the migration completes. Notice that the VMM's role in this process is to issue the required seamcalls and send encrypted blobs between the source and destination. The VMM does not need to know the contents of those blobs. More Detailed Overview ---------------------- The following diagram shows the TDX seamcalls used by the source and destination: Source host Destination host =========== ================ TDH.MIG.SETUP -- crypto keys, attestation -> TDH.MIG.SETUP | | TDH.EXPORT.STATE.IMMUTABLE -- read-only TD state -> TDH.IMPORT.STATE.IMMUTABLE | | TDH.MEM.SCAN.RANGE | TDH.MEM.TRACK + IPIs | TDH.EXPORT.MEM ------ memory data ---------> TDH.IMPORT.MEM TDH.EXPORT.TRACK ------ epoch token ---------> TDH.IMPORT.TRACK | | TDH.EXPORT.PAUSE | | | TDH.EXPORT.STATE.TD ------ global TD state -----> TDH.IMPORT.STATE.TD | | TDH.EXPORT.STATE.VP ------ vCPU state ----------> TDH.IMPORT.STATE.VP | | TDH.MEM.SCAN.COMP | TDH.MEM.TRACK + IPIs | TDH.EXPORT.MEM ------ final memory data ---> TDH.IMPORT.MEM TDH.EXPORT.TRACK ------ start token ---------> TDH.IMPORT.TRACK | TDH.IMPORT.END The source and destination go through the following stages. Setup ----- The VMM issues TDH.MIG.SETUP on the source and destination TDX modules iteratively to perform the setup session. The modules return status and may also return an encrypted blob, which the VMM passes between the two sides. In other words, the migration protocol is between the two TDX modules, while the VMM is simply the transport mechanism. The setup session performs mutual attestation, establishes trust between the modules, establishes migration encryption keys, and loads the migration policy. It completes when the TDX module returns success. Immutable State Transfer ------------------------ The source calls TDH.EXPORT.STATE.IMMUTABLE to export the immutable TD state. The destination VMM creates the destination TD skeleton and configures migration streams, then calls TDH.IMPORT.STATE.IMMUTABLE, which finalizes the destination TD initialization. Iterative Memory Copy --------------------- While the source TD continues running, the source and destination run iterative memory copy rounds: - The source calls TDH.MEM.SCAN.RANGE to find migration candidate pages. - The source calls TDH.MEM.TRACK and sends IPIs to the TD's vCPUs in order to ensure the dirty page scanning algorithm correctness. - The source calls TDH.EXPORT.MEM to export private memory pages, which the VMM sends to the destination. - The destination calls TDH.IMPORT.MEM to import the received memory pages. - The source calls TDH.EXPORT.TRACK to generate an epoch token. The VMM sends it to the destination, which calls TDH.IMPORT.TRACK to consume the token. This verifies that all data exported from the source was imported on the destination. The rounds continue until the number of pages that change between rounds is small enough to meet the convergence criteria. Stop and Copy ------------- The VMM pauses the source TD, which begins the downtime period, then calls TDH.EXPORT.PAUSE to start the TDX-enforced blackout period. The source then calls TDH.EXPORT.STATE.TD to export mutable TD-scope state and TDH.EXPORT.STATE.VP to export mutable state for each vCPU. It then performs the final dirty page scan, runs TDH.MEM.TRACK and sends IPIs, and calls TDH.EXPORT.MEM to export the final dirty pages as encrypted blobs. The destination calls TDH.IMPORT.STATE.TD to import mutable TD-scope state and TDH.IMPORT.STATE.VP to import the mutable state of each vCPU. It calls TDH.IMPORT.MEM to import the final dirty pages. After exporting the final dirty pages, the source's last migration seamcall is TDH.EXPORT.TRACK with IN_ORDER_DONE=1, which generates the start token. The VMM sends the start token to the destination. The destination calls TDH.IMPORT.TRACK with this token. This verifies that the mutable TD state has been imported and allows the destination to start the TD. Finally, the destination calls TDH.IMPORT.END, which ends the migration. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-09-04 18:24 ` [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Artem Bityutskiy @ 2026-09-17 21:27 ` Peter Xu 2026-09-18 12:46 ` Artem Bityutskiy 2026-09-20 23:56 ` Kishen Maloor 0 siblings, 2 replies; 84+ messages in thread From: Peter Xu @ 2026-09-17 21:27 UTC (permalink / raw) To: Artem Bityutskiy Cc: Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Fri, Sep 04, 2026 at 09:24:25PM +0300, Artem Bityutskiy wrote: > On Mon, 2026-08-31 at 10:13 +0300, Tony Lindgren wrote: > > Tom, since you mentioned that AMD SEV-SNP and Intel TDX live migration > > sound similar, can you please take a look how the API might work for > > SEV-SNP? > > The presented example uAPIs were designed to fit the TDX live migration > flow, with the intent that they could also be used by other CoCo VMs. > > It would be super great if someone could commend on how suitable these > uAPIs are for AMD and ARM flows. > > AFAIU, what is generic in the presented uAPI is more of usage pattern. > > - The order in which QEMU calls them. > - The idea that QEMU/KVM is a transport for opaque blobs. > > The proposed container - 'struct kvm_transfer_buffer' - only has address > and a size. The vendor defines the layout of the data in it. > > Some example high-level topics that would be nice to get feedback on: > > - Can we come up with a single set of generic migration uAPIs for different > CoCo models? > - Or should some uAPIs be generic while others are vendor-specific? > - Or should each CoCo model have its own vendor-specific set of migration > uAPIs? It's always good if we can put together as much function to be shared with generic ioctls as possible. At some point, IMHO we need to collect such information somehow, so when merging the generic API we know what vendor specific API will be needed. Hopefully this series is a good start. > - Should the same uAPIs also support traditional VMs? But the only use-case > I imagine here is "for testing purposes". This is an interesting idea, I think this could be useful. Especially, I wonder if you already have it done and PoC branches you can share, so that I can play with it. > > > Artem has put together a brief description below of the example API and the > > migration flow: > > .. snip ... > > > Migration flow > > ============== > > > > Source host Destination host > > =========== ================ > > > > CMD(SETUP/SESSION) <--- setup msgs ---> CMD(SETUP/SESSION) > > | (repeated) | > > CMD(SETUP/IMMUTABLE_STATE) - immutable state -> CMD(SETUP/IMMUTABLE_STATE) > > | | > > KVM_GET_DIRTY_LOG | > > KVM_EXPORT_MEMORY --- memory data ---> KVM_IMPORT_MEMORY > > CMD(ITERATION) --- epoch token ---> CMD(ITERATION) Could you elaborate this ITERATION operation? Is that something the userapp must do after full scan of a round of guest memory? > > | (repeat until convergence) | > > CMD(STOP_AND_COPY/PAUSE) | > > CMD(STOP_AND_COPY/TD_STATE) --- VM state ------> CMD(STOP_AND_COPY/TD_STATE) > > KVM_EXPORT_VCPU --- vCPU state ----> KVM_IMPORT_VCPU > > KVM_EXPORT_MEMORY -- final memory ---> KVM_IMPORT_MEMORY When read/write encrypted memories, two questions: - Is there an upper bound of the buffer size per-page? - Does this operation supports concurrency? If it supports, how well it scales per expectation (e.g. is there known big lock for that)? Similar question to the vCPU getter and setter. For now even without CoCo we serialize vCPU get/set, but I want to understand the potential of concurrent operations, and see if there's anything special for CoCo from that regard. > > CMD(ITERATION/DONE) --- start token ---> CMD(ITERATION) > > | | > > CMD(END) CMD(END) > > ... snip ... > > > As I mentioned, the example uAPI is modeled around the TDX migration flow. > In case it helps the reader, here is a summary of that flow that was > presented in PUCK. > > It describes what the TDX module offers today and focuses on pre-copy > migration. This is our interpretation of the TDX specifications, not a IMHO we should really take postcopy into account when designing the API and state machine. We don't need to implement it in the first version, even until merging, but we need to make sure postcopy will be new ioctls on top of existing and it should have no major loopholes that it'll need a new set of APIs. For example, I think we should consider KVM_EXPORT_MEMORY being usable after END on source, KVM_IMPORT_MEMORY while TD is in operation, etc. We should likely also need to still picture the rough process of postcopy, reserve those APIs since the start (but return -EINVAL or something). AFAIU, postcopy is so far still the best solution for extremely large or extremely busy VMs regarding user experience, and it will happen to CoCo VMs one day or another. > specification itself, and may contain errors or omissions. Please refer to > the official TDX specifications for authoritative information. > > Migration Overview > ------------------ > > In TDX, the TDX module implements the migration logic. The VMM drives > migration by issuing seamcalls to the source and destination TDX modules and > transports the resulting encrypted blobs between them. The VMM does not need > to know the contents of those blobs. > > The TDX security model enforces two hard rules: > > - Only one instance of the TD may run at a time. Either the source or the Just curious - could I ask why this limitation? > destination may run, but never both. In other words, cloning a TD is not > allowed. > - When migration completes, the destination must have the same memory and > vCPU state as the source. It must not end up with a partial or mixed > state. If such happens, it's definitely a bug, even without CoCo. Anything specific about CoCo? Like, whole-VM checksum? I recall QEMU could have some devices touching the memory during its post load process (after destination QEMU receive the device states and apply). I'm not sure how much it affects. > > Simple Overview > --------------- > > Run the migration setup session > | > v > Transfer immutable TD state > | > v > Copy dirty memory pages while the TD runs <-------+ > | | > +-- iterate until convergence criteria is met --+ > | > v > Pause the source TD and copy the remaining state > | > v > Start the destination TD > > The migration process starts with a setup session. During the setup session, > the source and destination TDX modules exchange encrypted blobs and > establish migration encryption keys. For example, these keys protect memory > contents transferred during migration. > > Next, the source transfers the immutable TD state to the destination and > initializes the destination TD. This state includes the read-only VM data, > such as its vCPU count and topology. > > The source and destination then perform iterative memory-copy rounds. In > each round, the source scans for memory pages that need to be migrated and > exports them in encrypted form. The destination imports the pages. The > rounds continue until the number of pages that change between rounds is > small enough to meet the convergence criteria. > > The source TD is then paused. The VMM transfers the remaining TD state, such I want to understand what is extra for a CoCo VM in terms of "pause", say, what's more than "stopping the vCPU threads". I saw there's mention of PRE_COPY_STOP state. One example question is, when reaching this state, can the guest memory still change? What happens if some emulated device are still DMAing to the guest memory (assuming flipped from private to shared)? In case of future IO zone support, what happens if in case of VFIO-PCI assigned doing encrypted DMA? From that regard, VFIO has the P2P state where it quiesce initiation of any DMA from this specific device, then another round to fully stop all devices into STOP_COPY phase. I wonder if CoCo VMs need similar treatment. > as vCPU state, and the final dirty memory pages. Finally, the destination TD > is started and the migration completes. > > Notice that the VMM's role in this process is to issue the required > seamcalls and send encrypted blobs between the source and destination. The > VMM does not need to know the contents of those blobs. > > More Detailed Overview > ---------------------- > > The following diagram shows the TDX seamcalls used by the source and > destination: > > Source host Destination host > =========== ================ > > TDH.MIG.SETUP -- crypto keys, attestation -> TDH.MIG.SETUP > | | > TDH.EXPORT.STATE.IMMUTABLE -- read-only TD state -> > TDH.IMPORT.STATE.IMMUTABLE > | | > TDH.MEM.SCAN.RANGE | Could you elaborate what's the relations between TDH.MEM.SCAN.RANGE and the GET_DIRTY_LOG ioctl? I recall above mentioned GET_DIRTY_LOG will be available even for CoCo, which makes sense assuming dirty information isn't confidential. However then I don't understand what TDH.MEM.SCAN.RANGE plays the role here. Thanks, > TDH.MEM.TRACK + IPIs | > TDH.EXPORT.MEM ------ memory data ---------> TDH.IMPORT.MEM > TDH.EXPORT.TRACK ------ epoch token ---------> TDH.IMPORT.TRACK > | | > TDH.EXPORT.PAUSE | > | | > TDH.EXPORT.STATE.TD ------ global TD state -----> TDH.IMPORT.STATE.TD > | | > TDH.EXPORT.STATE.VP ------ vCPU state ----------> TDH.IMPORT.STATE.VP > | | > TDH.MEM.SCAN.COMP | > TDH.MEM.TRACK + IPIs | > TDH.EXPORT.MEM ------ final memory data ---> TDH.IMPORT.MEM > TDH.EXPORT.TRACK ------ start token ---------> TDH.IMPORT.TRACK > | > TDH.IMPORT.END > > The source and destination go through the following stages. > > Setup > ----- > > The VMM issues TDH.MIG.SETUP on the source and destination TDX modules > iteratively to perform the setup session. The modules return status and may > also return an encrypted blob, which the VMM passes between the two sides. > In other words, the migration protocol is between the two TDX modules, while > the VMM is simply the transport mechanism. The setup session performs mutual > attestation, establishes trust between the modules, establishes migration > encryption keys, and loads the migration policy. It completes when the TDX > module returns success. > > Immutable State Transfer > ------------------------ > > The source calls TDH.EXPORT.STATE.IMMUTABLE to export the immutable TD > state. The destination VMM creates the destination TD skeleton and > configures migration streams, then calls TDH.IMPORT.STATE.IMMUTABLE, which > finalizes the destination TD initialization. > > Iterative Memory Copy > --------------------- > > While the source TD continues running, the source and destination run > iterative memory copy rounds: > > - The source calls TDH.MEM.SCAN.RANGE to find migration candidate pages. > - The source calls TDH.MEM.TRACK and sends IPIs to the TD's vCPUs in order > to ensure the dirty page scanning algorithm correctness. > - The source calls TDH.EXPORT.MEM to export private memory pages, which the > VMM sends to the destination. > - The destination calls TDH.IMPORT.MEM to import the received memory pages. > - The source calls TDH.EXPORT.TRACK to generate an epoch token. The VMM > sends it to the destination, which calls TDH.IMPORT.TRACK to consume the > token. This verifies that all data exported from the source was imported > on the destination. > > The rounds continue until the number of pages that change between rounds is > small enough to meet the convergence criteria. > > Stop and Copy > ------------- > > The VMM pauses the source TD, which begins the downtime period, then calls > TDH.EXPORT.PAUSE to start the TDX-enforced blackout period. The source then > calls TDH.EXPORT.STATE.TD to export mutable TD-scope state and > TDH.EXPORT.STATE.VP to export mutable state for each vCPU. It then performs > the final dirty page scan, runs TDH.MEM.TRACK and sends IPIs, and calls > TDH.EXPORT.MEM to export the final dirty pages as encrypted blobs. > > The destination calls TDH.IMPORT.STATE.TD to import mutable TD-scope state > and TDH.IMPORT.STATE.VP to import the mutable state of each vCPU. It calls > TDH.IMPORT.MEM to import the final dirty pages. > > After exporting the final dirty pages, the source's last migration seamcall > is TDH.EXPORT.TRACK with IN_ORDER_DONE=1, which generates the start token. > The VMM sends the start token to the destination. The destination calls > TDH.IMPORT.TRACK with this token. This verifies that the mutable TD state > has been imported and allows the destination to start the TD. Finally, the > destination calls TDH.IMPORT.END, which ends the migration. > -- Peter Xu ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-09-17 21:27 ` Peter Xu @ 2026-09-18 12:46 ` Artem Bityutskiy 2026-09-18 15:53 ` Peter Xu 2026-09-20 23:56 ` Kishen Maloor 1 sibling, 1 reply; 84+ messages in thread From: Artem Bityutskiy @ 2026-09-18 12:46 UTC (permalink / raw) To: Peter Xu Cc: Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm Hi Peter, thank you for your good comments and questions. A quick disclaimer before I address them. In my answers I try to keep three distinct things separate: 1. The CoCo migration uAPI - the generic interface we ultimately want to converge on. This is the end goal, not necessarily what these patches propose. What is good for this uAPI is my priority at this point. 2. The TDX migration model - the mechanics offered by Intel TDX module for TDX guest live migration today. 3. Our PoC - a concrete uAPI proposal and its example implementation for TDX guests. On Thu, 2026-09-17 at 17:27 -0400, Peter Xu wrote: > > Some example high-level topics that would be nice to get feedback on: > > > > - Can we come up with a single set of generic migration uAPIs for different > > CoCo models? > > - Or should some uAPIs be generic while others are vendor-specific? > > - Or should each CoCo model have its own vendor-specific set of migration > > uAPIs? > > It's always good if we can put together as much function to be shared with > generic ioctls as possible. At some point, IMHO we need to collect such > information somehow, so when merging the generic API we know what vendor > specific API will be needed. Hopefully this series is a good start. Thanks, agreed. On that note - does anyone know of a good doc describing the AMD, ARM, or other CoCo migration models? My knowledge is limited to TDX, so it is hard to tell what is common and what is TDX-specific. FYI, I am working on a TDX migration model document. It describes what the TDX module offers, but unlike the specs it is oriented towards software engineers: much easier to read and it does not require deep TDX knowledge. It is all based on public specs, just distilled into readable mental models. I plan to publish it publicly. I am about 80% done. > > - Should the same uAPIs also support traditional VMs? But the only use-case > > I imagine here is "for testing purposes". > > This is an interesting idea, I think this could be useful. Especially, I > wonder if you already have it done and PoC branches you can share, so that > I can play with it. We do not have code. But Kishen spent time playing with it, and I think he concluded not to proceed with this. But he might have evaluated it from the "unify all migration into a single generic API" perspective. But may be as a "this is a test framework" perspective is different, at least I feel it may be the case. I think Kishen can provide more insight if needed, he is in CC. > > > > > Migration flow > > > ============== > > > > > > Source host Destination host > > > =========== ================ > > > > > > CMD(SETUP/SESSION) <--- setup msgs ---> CMD(SETUP/SESSION) > > > | (repeated) | > > > CMD(SETUP/IMMUTABLE_STATE) - immutable state -> CMD(SETUP/IMMUTABLE_STATE) > > > | | > > > KVM_GET_DIRTY_LOG | > > > KVM_EXPORT_MEMORY --- memory data ---> KVM_IMPORT_MEMORY > > > CMD(ITERATION) --- epoch token ---> CMD(ITERATION) > > Could you elaborate this ITERATION operation? Is that something the > userapp must do after full scan of a round of guest memory? = Why no duplicate instances? = First, let me answer the "why" question you asked further below: "why does the TDX module require that only the source or only the destination runs, never both?". Answering it first makes the tokens easier to understand. And the tokens are why we proposed the "ITERATION" operation. I believe this is not TDX-specific, it is a confidential computing requirement. In short, cloning would give the VMM a very powerful primitive to attack confidential VMs. Here are a couple of example attack approaches: - When you attest a CoCo VM remotely, you get an assurance that you talk to this one specific instance. If duplication were allowed, many instances could exist, and that assurance is gone. - With a clone you can security-upgrade the state of one copy, attest the upgraded state, and then use the pre-upgraded copy: the user believes they are working with an up-to-date CoCo VM, but in fact use an older, possibly vulnerable version. This is also why TDX migration requires that not only must the two copies never run at the same time, but after migration the destination must be exactly the same as the source. For example, the destination must not end up using an older copy of a page. = ITERATION Operation = In the QEMU model, the pre-copy phase is a set of rounds: 1. Get the list of dirty pages. 2. Copy them to the destination. 3. Repeat until the convergence criteria are met. ITERATION is the explicit uAPI that ends the current pre-copy round. There is no equivalent uAPI for traditional VMs today. In the TDX model, ending a round needs an extra step: the source generates an epoch token, and the destination imports it. Two seamcalls do this: - TDH.EXPORT.TRACK generates the epoch token on the source. - TDH.IMPORT.TRACK imports it on the destination. The token enforces integrity and ordering, for example: - Every page exported on the source must be imported on the destination. - Once a newer version of a page is imported, an older version can no longer be imported. The final round is special. It uses the "done" flag, which is passed to the `TDH.EXPORT.TRACK` seamcall, and makes it export a special variant of the epoch token that is called the start token. On top of the integrity and ordering guarantees, the start token is what allows the destination to start: until the destination imports it, the TDX module will not let the destination TD run with partial state. > > > | (repeat until convergence) | > > > CMD(STOP_AND_COPY/PAUSE) | > > > CMD(STOP_AND_COPY/TD_STATE) --- VM state ------> CMD(STOP_AND_COPY/TD_STATE) > > > KVM_EXPORT_VCPU --- vCPU state ----> KVM_IMPORT_VCPU > > > KVM_EXPORT_MEMORY -- final memory ---> KVM_IMPORT_MEMORY > > When read/write encrypted memories, two questions: > > - Is there an upper bound of the buffer size per-page? So generally, the assumption is that exporting N pages requires M pages, M > N, because there may be some metadata (e.g., MACs for integrity checks). In case of TDX module, M is predictable and can be calculated in advance. I believe in our current PoC, the ioctl requires the buffer size to be large enough to hold all the requested pages. But this is specific to our current PoC implementation. In general, I feel like if buffer size is not enough, the uAPI could fill it with as much data as fits, and communicate back about what GPAs were exported. The caller could export the rest separately. Context: I am new in the Intel TDX live migration team, and did not participate in TDX PoC, that's why I use "I believe". I am catching up. But I assume others will (Tony, Kishen) will correct me if I am wrong. > - Does this operation supports concurrency? If it supports, how well it > scales per expectation (e.g. is there known big lock for that)? From the TDX module perspective, parallel exports of different GPAs can run on multiple CPUs, so I expect the QEMU multifd model to work and scale. In our current PoC the ioctl does not take a VM-wide lock, and concurrency is per-stream. I can expand on the stream concept if needed, but it is exactly about parallel import/export of memory and vCPU state. In our PoC we are focusing on the basics, but multifd support is definitely a goal too, just later. Kishen was already prototyping it in QEMU. From the uAPI point of view, I believe parallel export/import should be allowed. If a specific CoCo VM has issues with that, it would need to serialize the operations internally, I'd say. > - Does this operation supports concurrency? If it supports, how well it > scales per expectation (e.g. is there known big lock for that)? > > Similar question to the vCPU getter and setter. For now even without CoCo > we serialize vCPU get/set, but I want to understand the potential of > concurrent operations, and see if there's anything special for CoCo from > that regard. Similar to memory import/export: the TDX module explicitly allows vCPU state to be exported and imported in parallel. Our current PoC does not take advantage of this yet. vCPU export currently grabs the KVM MMU write lock, so vCPU exports are serialized today. I think Tony can comment more on the technical difficulties there. From the uAPI point of view, I'd propose to allow concurrent vCPU operations. > > > > It describes what the TDX module offers today and focuses on pre-copy > > migration. This is our interpretation of the TDX specifications, not a > > IMHO we should really take postcopy into account when designing the API and > state machine. We don't need to implement it in the first version, even > until merging, but we need to make sure postcopy will be new ioctls on top > of existing and it should have no major loopholes that it'll need a new set > of APIs. I totally agree. I have not yet dug into the TDX module implementation details for post-copy, but I know it is supported and I know the basics. I plan to study it in detail later. So far I have not noticed anything that would prevent adding post-copy on top later. > > For example, I think we should consider KVM_EXPORT_MEMORY being usable > after END on source, KVM_IMPORT_MEMORY while TD is in operation, etc. We > should likely also need to still picture the rough process of postcopy, > reserve those APIs since the start (but return -EINVAL or something). Yes, agreed. I will spend more time looking at this. But at this point, I just assumed that the proposed uAPIs can be used at the post-copy phase in parallel with on-demand page delivery. Just FYI, TDX module model allows for this, but we did not try it. > > AFAIU, postcopy is so far still the best solution for extremely large or > extremely busy VMs regarding user experience, and it will happen to CoCo > VMs one day or another. Sure, thanks for sharing. > > destination may run, but never both. In other words, cloning a TD is not > > allowed. > > - When migration completes, the destination must have the same memory and > > vCPU state as the source. It must not end up with a partial or mixed > > state. > > If such happens, it's definitely a bug, even without CoCo. Anything > specific about CoCo? Like, whole-VM checksum? Well, in CoCo VMs it is not just a bug, it is something the CoCo framework needs to make impossible, because VMM is considered to be untrusted, it can try to manipulate things and half-migrate, use it not as a bug but as attack vector. In TDX case, the TDX module will not allow you to run the TD - the TDH.VP.ENTER seamcall will fail. Regarding checksums: there is no single whole-VM checksum in TDX migration model. Instead integrity is enforced continuously - every exported blob carries a MACs that the destination TDX module verifies on import, and the epoch/start tokens guarantee that everything was imported, in order. > I want to understand what is extra for a CoCo VM in terms of "pause", say, > what's more than "stopping the vCPU threads". TDX module guarantees the source won't run, even if VMM tries, the TDH.VP.ENTER seamcall will fail. So the source TD state is effectively frozen and cannot be modified by the VMM. > > I saw there's mention of PRE_COPY_STOP state. One example question is, > when reaching this state, can the guest memory still change? What happens > if some emulated device are still DMAing to the guest memory (assuming > flipped from private to shared)? In case of future IO zone support, what > happens if in case of VFIO-PCI assigned doing encrypted DMA? > > From that regard, VFIO has the P2P state where it quiesce initiation of any > DMA from this specific device, then another round to fully stop all devices > into STOP_COPY phase. I wonder if CoCo VMs need similar treatment. Let me split this by device type, because TDX treats them very differently. Emulated (virtio-net, virtio-blk, etc.) only use shared memory - they cannot read or DMA into TD private memory. So full device state lives in shared memory, and QEMU migrates them exactly the same way as for a traditional VM. This is entirely outside the TDX module migration model and outside the proposed uAPI - the uAPI is only for TD private memory. Directly assigned devices are only possible with TDX Connect, where a physical PCIe/CXL function (a "TDI") is assigned to the TD and can DMA into private memory over a cryptographically protected link. This is not implemented in Linux yet. For migration, the TDX module requires all TDIs to be unassigned before the source TD is paused - the TDH.EXPORT.PAUSE seamcall actually checks this. Unassigning a TDI tears down its whole TD-private footprint (MMIO unmapped from the Secure EPT, trusted DMA mappings removed), so no device-specific state is left to migrate. From the TD's point of view it is a full hot-unplug on the source and a fresh hot-plug on the destination. Our current TDX guest migration PoC is built on this assumption. > Could you elaborate what's the relations between TDH.MEM.SCAN.RANGE and the > GET_DIRTY_LOG ioctl? I recall above mentioned GET_DIRTY_LOG will be > available even for CoCo, which makes sense assuming dirty information isn't > confidential. However then I don't understand what TDH.MEM.SCAN.RANGE > plays the role here. A few things here. First, our plan is that the standard GET_DIRTY_LOG ioctl is backed by the `TDH.MEM.SCAN.RANGE` seamcall - that is how dirty tracking is implemented for a TD. So we do not propose any special uAPI for dirty tracking. And I think you are right that the dirty information is not confidential - the TDX module exposes dirty page information to the VMM. Second, FYI, in current TDX migration model, dirty page scanning (`TDH.MEM.SCAN.RANGE`) is only allowed during a migration session. The TDX module returns an error if the seamcall is issued before the session is set up (i.e. before the source and destination TDX modules have exchanged the migration key and established trust - what the SETUP command does in the proposed uAPI). In other words, with our PoC, if someone tries to use GET_DIRTY_LOG without going through the migration setup - the ioctl will return an error. But I wish it were an independent feature instead. Then we could work on upstreaming it on its own - Tony estimates it is about 20% of the current TDX migration PoC code. I already raised this with the Intel TDX module architects, and they asked for use-cases. The only one we came up with is QEMU estimating the TD dirty rate before starting migration (the calc_dirty_rate command). I understand their position: without a use-case there is little reason to implement it, and they would also need to study the security implications - can it help an attacker in some way? So if you or anyone else can educate me about use-cases for independent dirty page scanning, I would really appreciate it - I could take them back to the TDX module architects. Thanks, Artem. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-09-18 12:46 ` Artem Bityutskiy @ 2026-09-18 15:53 ` Peter Xu 2026-09-22 8:09 ` Artem Bityutskiy 0 siblings, 1 reply; 84+ messages in thread From: Peter Xu @ 2026-09-18 15:53 UTC (permalink / raw) To: Artem Bityutskiy Cc: Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Fri, Sep 18, 2026 at 03:46:32PM +0300, Artem Bityutskiy wrote: > Hi Peter, Hi, Artem, > > thank you for your good comments and questions. A quick disclaimer before I > address them. In my answers I try to keep three distinct things separate: > > 1. The CoCo migration uAPI - the generic interface we ultimately want to > converge on. This is the end goal, not necessarily what these patches > propose. What is good for this uAPI is my priority at this point. > 2. The TDX migration model - the mechanics offered by Intel TDX module for > TDX guest live migration today. > 3. Our PoC - a concrete uAPI proposal and its example implementation for TDX > guests. > > On Thu, 2026-09-17 at 17:27 -0400, Peter Xu wrote: > > > Some example high-level topics that would be nice to get feedback on: > > > > > > - Can we come up with a single set of generic migration uAPIs for different > > > CoCo models? > > > - Or should some uAPIs be generic while others are vendor-specific? > > > - Or should each CoCo model have its own vendor-specific set of migration > > > uAPIs? > > > > It's always good if we can put together as much function to be shared with > > generic ioctls as possible. At some point, IMHO we need to collect such > > information somehow, so when merging the generic API we know what vendor > > specific API will be needed. Hopefully this series is a good start. > > Thanks, agreed. > > On that note - does anyone know of a good doc describing the AMD, ARM, or > other CoCo migration models? My knowledge is limited to TDX, so it is hard > to tell what is common and what is TDX-specific. > > FYI, I am working on a TDX migration model document. It describes what the > TDX module offers, but unlike the specs it is oriented towards software > engineers: much easier to read and it does not require deep TDX knowledge. > It is all based on public specs, just distilled into readable mental models. > I plan to publish it publicly. I am about 80% done. That will be very useful, thanks for doing this. I'll be more than happy to read it when it's done. I wonder if we can, after TDX bits done, use this doc as a base for others to add in theirs in separate tabs, keeping everything together for the unified migration API work. Out of pure curiosity, I also don't know where s390 stands; it's almost not mentioned in the current API plan. > > > > - Should the same uAPIs also support traditional VMs? But the only use-case > > > I imagine here is "for testing purposes". > > > > This is an interesting idea, I think this could be useful. Especially, I > > wonder if you already have it done and PoC branches you can share, so that > > I can play with it. > > We do not have code. But Kishen spent time playing with it, and I think he > concluded not to proceed with this. But he might have evaluated it from > the "unify all migration into a single generic API" perspective. But may > be as a "this is a test framework" perspective is different, at least I feel > it may be the case. I think Kishen can provide more insight if needed, he > is in CC. Yes, thanks. I'm willing to hear more, and I'm a bit surprised that non-coco migration didn't fit already well into it, because IIUC non-coco needs less in this case, not more. > > > > > > > > Migration flow > > > > ============== > > > > > > > > Source host Destination host > > > > =========== ================ > > > > > > > > CMD(SETUP/SESSION) <--- setup msgs ---> CMD(SETUP/SESSION) > > > > | (repeated) | > > > > CMD(SETUP/IMMUTABLE_STATE) - immutable state -> CMD(SETUP/IMMUTABLE_STATE) > > > > | | > > > > KVM_GET_DIRTY_LOG | > > > > KVM_EXPORT_MEMORY --- memory data ---> KVM_IMPORT_MEMORY > > > > CMD(ITERATION) --- epoch token ---> CMD(ITERATION) > > > > Could you elaborate this ITERATION operation? Is that something the > > userapp must do after full scan of a round of guest memory? > > = Why no duplicate instances? = > > First, let me answer the "why" question you asked further below: "why does > the TDX module require that only the source or only the destination runs, > never both?". Answering it first makes the tokens easier to understand. And > the tokens are why we proposed the "ITERATION" operation. > > I believe this is not TDX-specific, it is a confidential computing > requirement. In short, cloning would give the VMM a very powerful primitive > to attack confidential VMs. Here are a couple of example attack approaches: > > - When you attest a CoCo VM remotely, you get an assurance that you talk to > this one specific instance. If duplication were allowed, many instances > could exist, and that assurance is gone. > - With a clone you can security-upgrade the state of one copy, attest the > upgraded state, and then use the pre-upgraded copy: the user believes they > are working with an up-to-date CoCo VM, but in fact use an older, possibly > vulnerable version. > > This is also why TDX migration requires that not only must the two copies > never run at the same time, but after migration the destination must be > exactly the same as the source. For example, the destination must not end up > using an older copy of a page. > > = ITERATION Operation = > > In the QEMU model, the pre-copy phase is a set of rounds: > 1. Get the list of dirty pages. > 2. Copy them to the destination. > 3. Repeat until the convergence criteria are met. > > ITERATION is the explicit uAPI that ends the current pre-copy round. There > is no equivalent uAPI for traditional VMs today. > > In the TDX model, ending a round needs an extra step: the source generates > an epoch token, and the destination imports it. Two seamcalls do this: > > - TDH.EXPORT.TRACK generates the epoch token on the source. > - TDH.IMPORT.TRACK imports it on the destination. > > The token enforces integrity and ordering, for example: > > - Every page exported on the source must be imported on the destination. > - Once a newer version of a page is imported, an older version can no longer > be imported. > > The final round is special. It uses the "done" flag, which is passed to the > `TDH.EXPORT.TRACK` seamcall, and makes it export a special variant of the > epoch token that is called the start token. On top of the integrity and > ordering guarantees, the start token is what allows the destination to > start: until the destination imports it, the TDX module will not let the > destination TD run with partial state. The methodology on unique VM attestation sounds comlex, but I think I get it now, thank you. It'll be nice if some of these reasonings will also be there in the doc you're drafting. > > > > > | (repeat until convergence) | > > > > CMD(STOP_AND_COPY/PAUSE) | > > > > CMD(STOP_AND_COPY/TD_STATE) --- VM state ------> CMD(STOP_AND_COPY/TD_STATE) > > > > KVM_EXPORT_VCPU --- vCPU state ----> KVM_IMPORT_VCPU > > > > KVM_EXPORT_MEMORY -- final memory ---> KVM_IMPORT_MEMORY > > > > When read/write encrypted memories, two questions: > > > > - Is there an upper bound of the buffer size per-page? > > So generally, the assumption is that exporting N pages requires M pages, > M > N, because there may be some metadata (e.g., MACs for integrity > checks). In case of TDX module, M is predictable and can be calculated in > advance. I am just thinking out loud here: if the hardware is good enough to do encryption plus (some?) compression, that would be very nice. For "some", I meant minimum over zero pages. Because in this case even if host wants to play tricks with zero pages, it can't anymore when un-readable. Maybe the guest driver can play some trick, but I'm also not sure if in CoCo. M > N may imply it's not the case for now, but it's still sane as a start even if so. > > I believe in our current PoC, the ioctl requires the buffer size to be large > enough to hold all the requested pages. But this is specific to our current > PoC implementation. > > In general, I feel like if buffer size is not enough, the uAPI could fill it > with as much data as fits, and communicate back about what GPAs were > exported. The caller could export the rest separately. Yes, this will work. Or maybe it's simpler to be able to export an upper bound in another API that probes it (some KVM cap)? Any retry is a wasted round trip from perf perspective. The upper bound can be relatively large, IMHO, which should be non-issue. It should be simpler for both userapp and kernel if feasible. > > Context: I am new in the Intel TDX live migration team, and did not > participate in TDX PoC, that's why I use "I believe". I am catching up. But > I assume others will (Tony, Kishen) will correct me if I am wrong. No worries, thanks for the detailed answers whatever offered; they're already very helpful. > > > - Does this operation supports concurrency? If it supports, how well it > > scales per expectation (e.g. is there known big lock for that)? > > From the TDX module perspective, parallel exports of different GPAs can run > on multiple CPUs, so I expect the QEMU multifd model to work and scale. > > In our current PoC the ioctl does not take a VM-wide lock, and concurrency > is per-stream. I can expand on the stream concept if needed, but it is > exactly about parallel import/export of memory and vCPU state. > > In our PoC we are focusing on the basics, but multifd support is definitely > a goal too, just later. Kishen was already prototyping it in QEMU. Great. > > From the uAPI point of view, I believe parallel export/import should be > allowed. If a specific CoCo VM has issues with that, it would need to > serialize the operations internally, I'd say. Yes, it would be good to keep the critical section as small as possible in this case. As long as we are fully aware of the concurrent use model from userapp then it's good enough for now. > > > - Does this operation supports concurrency? If it supports, how well it > > scales per expectation (e.g. is there known big lock for that)? > > > > Similar question to the vCPU getter and setter. For now even without CoCo > > we serialize vCPU get/set, but I want to understand the potential of > > concurrent operations, and see if there's anything special for CoCo from > > that regard. > > Similar to memory import/export: the TDX module explicitly allows vCPU state > to be exported and imported in parallel. > > Our current PoC does not take advantage of this yet. vCPU export currently > grabs the KVM MMU write lock, so vCPU exports are serialized today. I think > Tony can comment more on the technical difficulties there. > > From the uAPI point of view, I'd propose to allow concurrent vCPU > operations. Sounds good. > > > > > > > It describes what the TDX module offers today and focuses on pre-copy > > > migration. This is our interpretation of the TDX specifications, not a > > > > IMHO we should really take postcopy into account when designing the API and > > state machine. We don't need to implement it in the first version, even > > until merging, but we need to make sure postcopy will be new ioctls on top > > of existing and it should have no major loopholes that it'll need a new set > > of APIs. > > I totally agree. I have not yet dug into the TDX module implementation > details for post-copy, but I know it is supported and I know the basics. I > plan to study it in detail later. So far I have not noticed anything that > would prevent adding post-copy on top later. As long as we have that in mind across working on this, that's good enough, thanks. > > > > > For example, I think we should consider KVM_EXPORT_MEMORY being usable > > after END on source, KVM_IMPORT_MEMORY while TD is in operation, etc. We > > should likely also need to still picture the rough process of postcopy, > > reserve those APIs since the start (but return -EINVAL or something). > > Yes, agreed. I will spend more time looking at this. But at this point, I > just assumed that the proposed uAPIs can be used at the post-copy phase in > parallel with on-demand page delivery. > > Just FYI, TDX module model allows for this, but we did not try it. > > > > > AFAIU, postcopy is so far still the best solution for extremely large or > > extremely busy VMs regarding user experience, and it will happen to CoCo > > VMs one day or another. > > Sure, thanks for sharing. > > > > destination may run, but never both. In other words, cloning a TD is not > > > allowed. > > > - When migration completes, the destination must have the same memory and > > > vCPU state as the source. It must not end up with a partial or mixed > > > state. > > > > If such happens, it's definitely a bug, even without CoCo. Anything > > specific about CoCo? Like, whole-VM checksum? > > Well, in CoCo VMs it is not just a bug, it is something the CoCo framework > needs to make impossible, because VMM is considered to be untrusted, it can > try to manipulate things and half-migrate, use it not as a bug but as attack > vector. In TDX case, the TDX module will not allow you to run the TD - the > TDH.VP.ENTER seamcall will fail. > > Regarding checksums: there is no single whole-VM checksum in TDX migration > model. Instead integrity is enforced continuously - every exported blob > carries a MACs that the destination TDX module verifies on import, and the > epoch/start tokens guarantee that everything was imported, in order. > > > I want to understand what is extra for a CoCo VM in terms of "pause", say, > > what's more than "stopping the vCPU threads". > > TDX module guarantees the source won't run, even if VMM tries, the > TDH.VP.ENTER seamcall will fail. So the source TD state is effectively > frozen and cannot be modified by the VMM. IIUC this should work for QEMU. Said that, we'll need to be careful then in case of migration fallbacks at the final stage. Nowadays, I believe QEMU can still fallback to source side at a very, very late stage after all things applied. If I'm not mistaken, the final handshake is done at migration_incoming_state_destroy() -> migrate_send_rp_shut() telling source to be gone. After reading above, one thing we may want to make sure is TDX ENTER on dest be exactly the last thing to do on destination, rather than dest QEMU ENTER done then something else seems wrong, then dest can't fallback anymore. I didn't check into details, though, more of a heads-up to whoever is working on QEMU for this in case useful. > > > > > I saw there's mention of PRE_COPY_STOP state. One example question is, > > when reaching this state, can the guest memory still change? What happens > > if some emulated device are still DMAing to the guest memory (assuming > > flipped from private to shared)? In case of future IO zone support, what > > happens if in case of VFIO-PCI assigned doing encrypted DMA? > > > > From that regard, VFIO has the P2P state where it quiesce initiation of any > > DMA from this specific device, then another round to fully stop all devices > > into STOP_COPY phase. I wonder if CoCo VMs need similar treatment. > > Let me split this by device type, because TDX treats them very differently. > > Emulated (virtio-net, virtio-blk, etc.) only use shared memory - they cannot > read or DMA into TD private memory. So full device state lives in shared > memory, and QEMU migrates them exactly the same way as for a traditional VM. > This is entirely outside the TDX module migration model and outside the > proposed uAPI - the uAPI is only for TD private memory. I hope I understand it right, that all shared pages are out of the secure zone TD manages, during migration or not (hence the same as some random page the VMM has allocated)? If so, anything about post_load operations shouldn't be any concern, and should work as usual. > > Directly assigned devices are only possible with TDX Connect, where a > physical PCIe/CXL function (a "TDI") is assigned to the TD and can DMA into > private memory over a cryptographically protected link. This is not > implemented in Linux yet. For migration, the TDX module requires all TDIs to > be unassigned before the source TD is paused - the TDH.EXPORT.PAUSE seamcall > actually checks this. Unassigning a TDI tears down its whole TD-private > footprint (MMIO unmapped from the Secure EPT, trusted DMA mappings removed), > so no device-specific state is left to migrate. From the TD's point of view > it is a full hot-unplug on the source and a fresh hot-plug on the > destination. Hmm, interesting. Sorry if this follow up question may be slightly off-topic of the new API, or maybe it matters, depends on the answer: do you know who is in charge of this unplug / plug operation? VMM or TD (hence, transparent to VMM)? If it's VMM, is QEMU involved? > > Our current TDX guest migration PoC is built on this assumption. > > > Could you elaborate what's the relations between TDH.MEM.SCAN.RANGE and the > > GET_DIRTY_LOG ioctl? I recall above mentioned GET_DIRTY_LOG will be > > available even for CoCo, which makes sense assuming dirty information isn't > > confidential. However then I don't understand what TDH.MEM.SCAN.RANGE > > plays the role here. > > A few things here. > > First, our plan is that the standard GET_DIRTY_LOG ioctl is backed by the > `TDH.MEM.SCAN.RANGE` seamcall - that is how dirty tracking is implemented > for a TD. So we do not propose any special uAPI for dirty tracking. > > And I think you are right that the dirty information is not confidential - > the TDX module exposes dirty page information to the VMM. > > Second, FYI, in current TDX migration model, dirty page scanning > (`TDH.MEM.SCAN.RANGE`) is only allowed during a migration session. The TDX > module returns an error if the seamcall is issued before the session is set > up (i.e. before the source and destination TDX modules have exchanged the > migration key and established trust - what the SETUP command does in the > proposed uAPI). > > In other words, with our PoC, if someone tries to use GET_DIRTY_LOG without > going through the migration setup - the ioctl will return an error. > > But I wish it were an independent feature instead. Then we could work on > upstreaming it on its own - Tony estimates it is about 20% of the current > TDX migration PoC code. Unexpected, that's a lot just for tracking. > > I already raised this with the Intel TDX module architects, and they asked > for use-cases. The only one we came up with is QEMU estimating the TD dirty > rate before starting migration (the calc_dirty_rate command). I understand > their position: without a use-case there is little reason to implement it, > and they would also need to study the security implications - can it help > an attacker in some way? > > So if you or anyone else can educate me about use-cases for independent > dirty page scanning, I would really appreciate it - I could take them back > to the TDX module architects. Yes, calc_dirty_rate will use it, I also can't think of another use case that needs it. RHEL supports it, so it would be nice this will be supported for CoCo too. QEMU also has another way to do it without KVM tracking, which is kind of simle page hash daemon fully done in userspace. But in case of CoCo it'll stop working too when most memory unreadable. So GET_DIRTY_LOG seems the only way to go. We can still "emulate" it by initiating a remote migration just to collect this info, I believe all attestations will simply pass and we do a fallback when the admin turns it off.. but it's awkward and all the rest (including trying to find another host suitable for migration, allocate resources..) are pure wastes. Btw, do you know if the private pages will be write-trappable by the host kernel, with things like userfaultfd-wp or soft-dirty? That'll be another way to do this without TD involvement, but I don't know well enough to say. It just sounds like if it's doable in KVM, it's still the best place, also since GET_DIRTY_LOG is available as long as memslot marked tracking, it sounds good to keep CoCo be compatible with that API if it is to be reused, hence GET_DIRTY_LOG API is consistent across coco / non-coco. Thanks, -- Peter Xu ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-09-18 15:53 ` Peter Xu @ 2026-09-22 8:09 ` Artem Bityutskiy 2026-09-22 9:42 ` Tony Lindgren ` (2 more replies) 0 siblings, 3 replies; 84+ messages in thread From: Artem Bityutskiy @ 2026-09-22 8:09 UTC (permalink / raw) To: Peter Xu Cc: Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm Hi Peter, thanks again for good comments and questions. On Fri, 2026-09-18 at 11:53 -0400, Peter Xu wrote: > > > > FYI, I am working on a TDX migration model document. It describes what the > > TDX module offers, but unlike the specs it is oriented towards software > > engineers: much easier to read and it does not require deep TDX knowledge. > > It is all based on public specs, just distilled into readable mental models. > > I plan to publish it publicly. I am about 80% done. > > That will be very useful, thanks for doing this. I'll be more than happy to > read it when it's done. I wonder if we can, after TDX bits done, use this > doc as a base for others to add in theirs in separate tabs, keeping > everything together for the unified migration API work. > > Out of pure curiosity, I also don't know where s390 stands; it's almost not > mentioned in the current API plan. In general, sounds like a good idea. Let me finish TDX and share it first, then we can think about the next step. But in general a master doc with high level overview and comparison of different CoCo migration models would be awesome. s390 - yes, would be curious to know. > > > > We do not have code. But Kishen spent time playing with it, and I think he > > concluded not to proceed with this. But he might have evaluated it from > > the "unify all migration into a single generic API" perspective. But may > > be as a "this is a test framework" perspective is different, at least I feel > > it may be the case. I think Kishen can provide more insight if needed, he > > is in CC. > > Yes, thanks. I'm willing to hear more, and I'm a bit surprised that > non-coco migration didn't fit already well into it, because IIUC non-coco > needs less in this case, not more. Yeah, just thinking aloud now. - Suppose QEMU and KVM support CoCo migration. - Does it make sense to treat traditional VM as a CoCo VM for migration purposes? - Not for performance reasons - mmapped memory I/O is going to be more efficient than any sorts of export/import uAPIs. - Yes for testing/experimental purposes. - Would it add a lot of code and complexity: would the benefits outweigh the costs? - QEMU: Should require minimum amount of code. Would need a flag like "treat me as CoCo VM", may be some QEMU operator commands or cmdline options. - KVM: Not sure, but intuitively it may add noticeable amount of churn. > > The final round is special. It uses the "done" flag, which is passed to the > > `TDH.EXPORT.TRACK` seamcall, and makes it export a special variant of the > > epoch token that is called the start token. On top of the integrity and > > ordering guarantees, the start token is what allows the destination to > > start: until the destination imports it, the TDX module will not let the > > destination TD run with partial state. > > The methodology on unique VM attestation sounds comlex, but I think I get > it now, thank you. It'll be nice if some of these reasonings will also be > there in the doc you're drafting. Yes, will add. > > > > > I am just thinking out loud here: if the hardware is good enough to do > encryption plus (some?) compression, that would be very nice. For "some", > I meant minimum over zero pages. Because in this case even if host wants > to play tricks with zero pages, it can't anymore when un-readable. Maybe > the guest driver can play some trick, but I'm also not sure if in CoCo. > > M > N may imply it's not the case for now, but it's still sane as a start > even if so. Yeah. In case of TDX, the TDX module does not do compression today, and there is no zero-page optimization either. But keep in mind that what pages in TD are zero pages is by itself a secret. Any sort of theoretical TDX module zero page compression optimization would need to be done in a way that VMM cannot figure out what pages were zero pages. So with a disclaimer that I am not really a security expert, I'd say it is far-fetched. > > I believe in our current PoC, the ioctl requires the buffer size to be large > > enough to hold all the requested pages. But this is specific to our current > > PoC implementation. > > > > In general, I feel like if buffer size is not enough, the uAPI could fill it > > with as much data as fits, and communicate back about what GPAs were > > exported. The caller could export the rest separately. > > Yes, this will work. > > Or maybe it's simpler to be able to export an upper bound in another API > that probes it (some KVM cap)? > > Any retry is a wasted round trip from perf perspective. The upper bound > can be relatively large, IMHO, which should be non-issue. It should be > simpler for both userapp and kernel if feasible. OK, so by upper bound here you mean the maximum amount of data that can be exported in a single memory export ioctl operation, right? And you mean that there should be some sort of command to query it, similar to KVM_CAP_XSAVE2 capability returning the XSAVE buffer size? If so - yes, I think exposing such an upper bound would be useful. Let's see. In case of TDX, the `TDH.EXPORT.MEM` seamcall today can export max. 512 pages at a time. Today only 4KiB pages are supported, so it is just a 2MiB buffer. So this is the upper bound for the seamcall. So in TDX case, 512 pages would be the absolute upper bound for the uAPI input buffer size. The alternative is to accept larger input buffers and internally do multiple seamcalls. For TDX case, I do not think the latter makes much sense though. But, what I do think is that the uAPI itself shouldn't mandate one approach over the other - a different CoCo implementation may prefer to aggregate multiple hardware calls internally. Also, today TDX can only migrate 4KiB pages - VMM must split larger pages into 4KiB pages before migration. But I expect that in the future TDX may implement larger pages migration too. So I think the cap should be expressed in page count rather than bytes. That said, I don't think exposing this upper bound cap should replace the flexibility we discussed earlier: the API itself should still allow exporting as much as fits the output buffer, regardless of what the cap reports. A specific CoCo implementation could still choose to just fail if the buffer isn't large enough to hold all requested pages - but that would be an implementation limitation, not an API limitation. And the other question is whether to allow exporting a mix of large and small pages, like N x 4KiB along with M x 2MiB and K x 1GiB pages. Or the uAPI would allow to only export one page size at a time? If large pages are added to TDX, I'd speculate that TDX seamcall would allow the mixing and matching - I can see that already today the seamcall API is provisioned for this. uAPI could require that the input buffer (set of pages) can contain a mix which should match the GPAs of the corresponding memory regions. But then race conditions - the GPA layout of the VM may change by the time the request reaches the hardware: large pages can be split or the other way round, I guess? Then uAPI could return a "retry" sort of exit code, may be? What do you think? I am just thinking aloud: looks like pages of different size add a degree of complexity - I did not take that into account until this conversation, thanks. > > Said that, we'll need to be careful then in case of migration fallbacks at > the final stage. Nowadays, I believe QEMU can still fallback to source side > at a very, very late stage after all things applied. If I'm not mistaken, > the final handshake is done at migration_incoming_state_destroy() -> > migrate_send_rp_shut() telling source to be gone. Right. In traditional pre-copy VM migration model, the fallback is possible at any point before the destination VM starts running and modifying its state. The same is true for TDX, but with more complexity and limitations. We have 2 points: 1. Before the source has exported the start token - fallback is similar to traditional VM migration - just abort the migration on source and continue running the source TD, and just destroy the destination TD. 2. After the source has exported the start token - fallback is still possible, but more complex - it requires the destination to first generate the abort token, which should be delivered to the source and consumed there. There are seamcalls for both generating and consuming the abort token. The token basically makes sure the destination cannot run, and the source can run again - same "only one of the TDs can run at a time" security rule that we discussed earlier. Now, the limitation here is that if the abort is because the network connectivity is lost, the abort token cannot be delivered. Then I'd guess migration should stop, without destroying the source and the destination, and the token should be delivered later manually, QEMU could even provide an infrastructure / commands for that. Would be very helpful to learn the abort protocol the other CoCo vendors offer - is it similar to TDX or not? In our PoC we did not implement the TDX abort token procedure. But the proposed uAPI is provisioned for the abort token export/import. > After reading above, one thing we may want to make sure is TDX ENTER on > dest be exactly the last thing to do on destination, rather than dest QEMU > ENTER done then something else seems wrong, then dest can't fallback > anymore. I didn't check into details, though, more of a heads-up to > whoever is working on QEMU for this in case useful. Well, first of all, as the above describes, in case of TDX late fallback is still possible, just requires a more complex protocol. But from recoverability point of view, I agree with your suggestion to make TDH.VP.ENTER the last thing done on the destination, rather than starting the destination early and doing more afterwards. However, the paramount performance goal is to minimize the downtime, and it conflicts with the suggested idea. In my mind, minimizing downtime should take precedence. > > > I saw there's mention of PRE_COPY_STOP state. One example question is, > > > when reaching this state, can the guest memory still change? What happens > > > if some emulated device are still DMAing to the guest memory (assuming > > > flipped from private to shared)? In case of future IO zone support, what > > > happens if in case of VFIO-PCI assigned doing encrypted DMA? > > > > > > From that regard, VFIO has the P2P state where it quiesce initiation of any > > > DMA from this specific device, then another round to fully stop all devices > > > into STOP_COPY phase. I wonder if CoCo VMs need similar treatment. > > > > Let me split this by device type, because TDX treats them very differently. > > > > Emulated (virtio-net, virtio-blk, etc.) only use shared memory - they cannot > > read or DMA into TD private memory. So full device state lives in shared > > memory, and QEMU migrates them exactly the same way as for a traditional VM. > > This is entirely outside the TDX module migration model and outside the > > proposed uAPI - the uAPI is only for TD private memory. > > I hope I understand it right, that all shared pages are out of the secure > zone TD manages, during migration or not (hence the same as some random > page the VMM has allocated)? If so, anything about post_load operations > shouldn't be any concern, and should work as usual. Yes, that is correct. Shared pages are fully managed by KVM, not by the TDX module, regardless of whether migration is in progress or not. They are migrated the standard QEMU way, same as for traditional VMs, so post_load and any other migration code dealing with shared memory should work exactly as it does today. Again, good to know how it is for other CoCo VMs. > > Directly assigned devices are only possible with TDX Connect, where a > > physical PCIe/CXL function (a "TDI") is assigned to the TD and can DMA into > > private memory over a cryptographically protected link. This is not > > implemented in Linux yet. For migration, the TDX module requires all TDIs to > > be unassigned before the source TD is paused - the TDH.EXPORT.PAUSE seamcall > > actually checks this. Unassigning a TDI tears down its whole TD-private > > footprint (MMIO unmapped from the Secure EPT, trusted DMA mappings removed), > > so no device-specific state is left to migrate. From the TD's point of view > > it is a full hot-unplug on the source and a fresh hot-plug on the > > destination. > > Hmm, interesting. Sorry if this follow up question may be slightly > off-topic of the new API, or maybe it matters, depends on the answer: do > you know who is in charge of this unplug / plug operation? VMM or TD > (hence, transparent to VMM)? > > If it's VMM, is QEMU involved? Specifically about TDX - the pause seamcall will return an error if TDIs are not unassigned. While this is something that is not implemented in Linux yet, I believe the model will be that there is some uAPI to unassign TDIs, and it is not related to migration. QEMU would just need to exercise this uAPI at the right time. > > > > But I wish it were an independent feature instead. Then we could work on > > upstreaming it on its own - Tony estimates it is about 20% of the current > > TDX migration PoC code. > > Unexpected, that's a lot just for tracking. I assume because it brings various seamcall wrappers and some shared infra code. But also Tony might have over-estimated - I'll let him comment on this. > > I already raised this with the Intel TDX module architects, and they asked > > for use-cases. The only one we came up with is QEMU estimating the TD dirty > > rate before starting migration (the calc_dirty_rate command). I understand > > their position: without a use-case there is little reason to implement it, > > and they would also need to study the security implications - can it help > > an attacker in some way? > > > > So if you or anyone else can educate me about use-cases for independent > > dirty page scanning, I would really appreciate it - I could take them back > > to the TDX module architects. > > Yes, calc_dirty_rate will use it, I also can't think of another use case > that needs it. RHEL supports it, so it would be nice this will be > supported for CoCo too. Do you or someone else know how important it is to have it working? Do actual customers use it in real setups and rely on it? How bad or painful is it if this feature is not available? The reason I am asking is that my assumption was that it is not important. But if it is, I will come back to the TDX module architects with a request to revise the design to support independent dirty page tracking. Of course they may have some reasons for not doing it, but I would try at least. > QEMU also has another way to do it without KVM tracking, which is kind of > simle page hash daemon fully done in userspace. But in case of CoCo it'll > stop working too when most memory unreadable. So GET_DIRTY_LOG seems the > only way to go. > > We can still "emulate" it by initiating a remote migration just to collect > this info, I believe all attestations will simply pass and we do a fallback > when the admin turns it off.. but it's awkward and all the rest (including > trying to find another host suitable for migration, allocate resources..) > are pure wastes. Understood, thanks - sounds like this workaround is doable, but wasteful, something to avoid if possible. That in itself is useful evidence I can bring back to the TDX module architects. Just to understand the real-world deployments better: do you know if anyone actually falls back to this workaround today, or is it more of a theoretical idea that nobody actually uses? > > Btw, do you know if the private pages will be write-trappable by the host > kernel, with things like userfaultfd-wp or soft-dirty? That'll be another > way to do this without TD involvement, but I don't know well enough to say. The short answer is - no, private pages are not write-trappable outside of a migration session, so trapping cannot be used as a workaround for independent dirty tracking today. But a bit longer answer is that the TDX module provides 2 alternative dirty page tracking mechanisms, and one of them is in fact a write-trapping mechanism. In the TDX specs, these mechanisms are called Write-Blocking Export and Non-Blocking Export. - Write-blocking: the VMM write-blocks pages using `TDH.EXPORT.BLOCKW`, and gets a VM exit (an EPT violation) when the TD attempts to write a blocked page. - Non-blocking: based on Secure EPT scanning, using `TDH.MEM.SCAN.RANGE`, which was mentioned earlier in this cover letter. I refer to it as dirty scanning. Today, both mechanisms are only available during migration and require the migration setup to be done first - neither can be used as a standalone. In our PoC we use the dirty scanning mechanism, as we expect it to be more performant. But our PoC actually started with the write-blocking method. According to Kishen and Tony, it required a lot of ugly code and was very intrusive into the MMU code. But I'll let Kishen and Tony provide more details. Anyway, if I can get the TDX module to allow dirty tracking independent of migration, we could pick either mechanism, or even implement both and add a TDX-specific ioctl to select which one to use. We could plug either method into the KVM_DIRTY_LOG ioctl - both would work, just differently. = Dirty Scanning Optimization in TDX Module = Now, I am diverging, but just in case: in the PUCK call where we presented TDX migration, Sean made an immediate observation that WP-based dirty tracking is not categorically worse than PML-based dirty ring tracking, it depends on the workload. Then he learned that TDX module's dirty scanning does not use PML, and was understandably surprised. Sean was concerned about dirty scanning performance. So what I learned then about dirty scanning is that it is optimized for minimizing the downtime. Before the source TD is paused, there are a couple of seamcalls to invoke, let me refer to them as prescan seamcalls. They scan the Secure EPT while the source is running, and mark the sub-trees that do not have migration candidates as "clean". Then during the actual final dirty scan in the downtime window, the scan can skip those entire sub-trees and be more efficient. Sean correctly pointed out that this would be problematic because on Intel CPUs the EPT dirty flag is set only at leaf level, and does not propagate all the way to the root level. However, what I learned is that the TDX module uses the "accessed" bit instead of the dirty bit in this prescan optimization - the accessed bit does propagate all the way to the root level. I thought this was a nifty trick, but obviously it would also mark sub-trees as "not-clean" on read access. Now, I personally did not benchmark this optimization, but I heard that some people did and the results showed acceptable performance. > It just sounds like if it's doable in KVM, it's still the best place, also > since GET_DIRTY_LOG is available as long as memslot marked tracking, it > sounds good to keep CoCo be compatible with that API if it is to be reused, > hence GET_DIRTY_LOG API is consistent across coco / non-coco. Yes, that's exactly the goal and the proposal: QEMU would use GET_DIRTY_LOG API regardless of the VM type. In TDX guest case, KVM would use TDX-specific dirty page scanning seamcalls to implement the GET_DIRTY_LOG API. Thanks, Artem. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-09-22 8:09 ` Artem Bityutskiy @ 2026-09-22 9:42 ` Tony Lindgren 2026-09-22 11:54 ` Artem Bityutskiy 2026-09-22 21:18 ` Peter Xu 2026-09-23 15:28 ` Serge Hallyn (AMD) 2 siblings, 1 reply; 84+ messages in thread From: Tony Lindgren @ 2026-09-22 9:42 UTC (permalink / raw) To: Artem Bityutskiy Cc: Peter Xu, Paolo Bonzini, Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Tue, Sep 22, 2026 at 11:09:42AM +0300, Artem Bityutskiy wrote: > On Fri, 2026-09-18 at 11:53 -0400, Peter Xu wrote: > > Unexpected, that's a lot just for tracking. > > I assume because it brings various seamcall wrappers and some shared infra > code. But also Tony might have over-estimated - I'll let him comment on > this. Yes dirty log support is currently about 20% of the number of patches for the POC. About 100 LOC changes, about 75% it is TDX specific code with SEAMCALLs and and the scanning functions. > > > I already raised this with the Intel TDX module architects, and they asked > > > for use-cases. The only one we came up with is QEMU estimating the TD dirty > > > rate before starting migration (the calc_dirty_rate command). I understand > > > their position: without a use-case there is little reason to implement it, > > > and they would also need to study the security implications - can it help > > > an attacker in some way? > > > > > > So if you or anyone else can educate me about use-cases for independent > > > dirty page scanning, I would really appreciate it - I could take them back > > > to the TDX module architects. > > > > Yes, calc_dirty_rate will use it, I also can't think of another use case > > that needs it. RHEL supports it, so it would be nice this will be > > supported for CoCo too. > > Do you or someone else know how important it is to have it working? > Do actual customers use it in real setups and rely on it? > How bad or painful is it if this feature is not available? > > The reason I am asking is that my assumption was that it is not important. > But if it is, I will come back to the TDX module architects with a > request to revise the design to support independent dirty page tracking. Of > course they may have some reasons for not doing it, but I would try at > least. > > > QEMU also has another way to do it without KVM tracking, which is kind of > > simle page hash daemon fully done in userspace. But in case of CoCo it'll > > stop working too when most memory unreadable. So GET_DIRTY_LOG seems the > > only way to go. > > > > We can still "emulate" it by initiating a remote migration just to collect > > this info, I believe all attestations will simply pass and we do a fallback > > when the admin turns it off.. but it's awkward and all the rest (including > > trying to find another host suitable for migration, allocate resources..) > > are pure wastes. > > Understood, thanks - sounds like this workaround is doable, but wasteful, > something to avoid if possible. That in itself is useful evidence I can > bring back to the TDX module architects. > > Just to understand the real-world deployments better: do you know if anyone > actually falls back to this workaround today, or is it more of a theoretical > idea that nobody actually uses? Interesting workaround :) I too agree that proper dirty log support is the way to go. To me it seems non-migration dirty log support can be added to the TDX module as an additional feature. And adding it probably would not even need KVM changes except a new flag for the scan SEAMCALL. > > Btw, do you know if the private pages will be write-trappable by the host > > kernel, with things like userfaultfd-wp or soft-dirty? That'll be another > > way to do this without TD involvement, but I don't know well enough to say. > > The short answer is - no, private pages are not write-trappable outside of > a migration session, so trapping cannot be used as a workaround for > independent dirty tracking today. > > But a bit longer answer is that the TDX module provides 2 alternative dirty > page tracking mechanisms, and one of them is in fact a write-trapping > mechanism. In the TDX specs, these mechanisms are called Write-Blocking > Export and Non-Blocking Export. > > - Write-blocking: the VMM write-blocks pages using `TDH.EXPORT.BLOCKW`, and > gets a VM exit (an EPT violation) when the TD attempts to write a > blocked page. > - Non-blocking: based on Secure EPT scanning, using `TDH.MEM.SCAN.RANGE`, > which was mentioned earlier in this cover letter. I refer to it as dirty > scanning. > > Today, both mechanisms are only available during migration and require the > migration setup to be done first - neither can be used as a standalone. In > our PoC we use the dirty scanning mechanism, as we expect it to be more > performant. > > But our PoC actually started with the write-blocking method. According to > Kishen and Tony, it required a lot of ugly code and was very intrusive into > the MMU code. But I'll let Kishen and Tony provide more details. Yes write-blocking has issues with being intrusive. Exported pages are locked by the TDX module and only cleared on import or after a cancel operation. KVM MMU error handling gets tricky. The non-blocking migration makes the tricky parts go away at the cost of adding TDX specific code to handle the GET_DIRTY_LOG scanning. > Anyway, if I can get the TDX module to allow dirty tracking independent of > migration, we could pick either mechanism, or even implement both and add a > TDX-specific ioctl to select which one to use. We could plug either method > into the KVM_DIRTY_LOG ioctl - both would work, just differently. Note that for TDX write-blocking is an older approach. The non-blocking SEAMCALLs were added because of the issues noticed. Both features are not usable the same time. If non-blocking is enabled for a TDX module write-blocking cannot be used. Also note that the dirty log scan features depend on non-blocking features being enabled. IMO no reason for Linux to try to support the write-blocking migration at all. > = Dirty Scanning Optimization in TDX Module = > > Now, I am diverging, but just in case: in the PUCK call where we presented > TDX migration, Sean made an immediate observation that WP-based dirty > tracking is not categorically worse than PML-based dirty ring tracking, it > depends on the workload. Then he learned that TDX module's dirty scanning > does not use PML, and was understandably surprised. Sean was concerned about > dirty scanning performance. > > So what I learned then about dirty scanning is that it is optimized for > minimizing the downtime. Before the source TD is paused, there are a couple > of seamcalls to invoke, let me refer to them as prescan seamcalls. > > They scan the Secure EPT while the source is running, and mark the > sub-trees that do not have migration candidates as "clean". Then during > the actual final dirty scan in the downtime window, the scan can skip > those entire sub-trees and be more efficient. > > Sean correctly pointed out that this would be problematic because on Intel > CPUs the EPT dirty flag is set only at leaf level, and does not propagate > all the way to the root level. However, what I learned is that the TDX > module uses the "accessed" bit instead of the dirty bit in this prescan > optimization - the accessed bit does propagate all the way to the root > level. I thought this was a nifty trick, but obviously it would also mark > sub-trees as "not-clean" on read access. > > Now, I personally did not benchmark this optimization, but I heard that > some people did and the results showed acceptable performance. > > > It just sounds like if it's doable in KVM, it's still the best place, also > > since GET_DIRTY_LOG is available as long as memslot marked tracking, it > > sounds good to keep CoCo be compatible with that API if it is to be reused, > > hence GET_DIRTY_LOG API is consistent across coco / non-coco. > > > Yes, that's exactly the goal and the proposal: QEMU would use GET_DIRTY_LOG > API regardless of the VM type. In TDX guest case, KVM would use TDX-specific > dirty page scanning seamcalls to implement the GET_DIRTY_LOG API. Yes. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-09-22 9:42 ` Tony Lindgren @ 2026-09-22 11:54 ` Artem Bityutskiy 2026-09-23 4:20 ` Tony Lindgren 0 siblings, 1 reply; 84+ messages in thread From: Artem Bityutskiy @ 2026-09-22 11:54 UTC (permalink / raw) To: Tony Lindgren Cc: Peter Xu, Paolo Bonzini, Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Tue, 2026-09-22 at 12:42 +0300, Tony Lindgren wrote: > > > Anyway, if I can get the TDX module to allow dirty tracking independent of > > migration, we could pick either mechanism, or even implement both and add a > > TDX-specific ioctl to select which one to use. We could plug either method > > into the KVM_DIRTY_LOG ioctl - both would work, just differently. > > Note that for TDX write-blocking is an older approach. The non-blocking > SEAMCALLs were added because of the issues noticed. Both features are not > usable the same time. But let me clarify one point: write-blocking (the WP-based mechanism) isn't broken - it works fine. For fairness: - MMU-intrusiveness - valid argument. - It is "older" - not on its own a strong argument. Older does not automatically make it categorically worse. Side note: although I always try to remember that if there is a real need, we can work with the TDX module architects on the ABI to reduce that intrusiveness. I do not see the need yet, though. > If non-blocking is enabled for a TDX module write-blocking cannot be used. > Also note that the dirty log scan features depend on non-blocking features > being enabled. To make it more pronounced: Tony says the blocking vs. non-blocking TDX export mode is: - Global for all TDs, not a per-TD choice. - Decided at TDX module initialization time, and cannot be changed without re-initializing or updating the module. So in practice, it's a kernel init time decision. So my suggestion that both could be implemented and selected via a uAPI is moot in practice. Good point, thanks! Artem. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-09-22 11:54 ` Artem Bityutskiy @ 2026-09-23 4:20 ` Tony Lindgren 0 siblings, 0 replies; 84+ messages in thread From: Tony Lindgren @ 2026-09-23 4:20 UTC (permalink / raw) To: Artem Bityutskiy Cc: Peter Xu, Paolo Bonzini, Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Tue, Sep 22, 2026 at 02:54:29PM +0300, Artem Bityutskiy wrote: > On Tue, 2026-09-22 at 12:42 +0300, Tony Lindgren wrote: > > > > > Anyway, if I can get the TDX module to allow dirty tracking independent of > > > migration, we could pick either mechanism, or even implement both and add a > > > TDX-specific ioctl to select which one to use. We could plug either method > > > into the KVM_DIRTY_LOG ioctl - both would work, just differently. > > > > Note that for TDX write-blocking is an older approach. The non-blocking > > SEAMCALLs were added because of the issues noticed. Both features are not > > usable the same time. > > But let me clarify one point: write-blocking (the WP-based mechanism) isn't > broken - it works fine. For fairness: > > - MMU-intrusiveness - valid argument. > - It is "older" - not on its own a strong argument. Older does not > automatically make it categorically worse. Correct. And I'm of course a bit biased on the write-blocking migration after tinkering with it some so please excuse me. For reference, there are earlier WIP patches from 2023 for TDX using the write-blocking at [0] and [1] below. While the write-blocking itself seems fairly straight forward, it's interaction with live migration is complicated to unblock pages. You may want to add a VM exit and TDX module unblock for each 4K page to the list for the TDX write-blocking migration. > Side note: although I always try to remember that if there is a real need, > we can work with the TDX module architects on the ABI to reduce that > intrusiveness. I do not see the need yet, though. > > > If non-blocking is enabled for a TDX module write-blocking cannot be used. > > Also note that the dirty log scan features depend on non-blocking features > > being enabled. > > To make it more pronounced: Tony says the blocking vs. non-blocking TDX > export mode is: > > - Global for all TDs, not a per-TD choice. > - Decided at TDX module initialization time, and cannot be changed without > re-initializing or updating the module. So in practice, it's a kernel > init time decision. > > So my suggestion that both could be implemented and selected via a uAPI is > moot in practice. Yup. If somebody really needs it, selecting write-blocking migration could be a module param. [0] https://github.com/intel-staging/tdx/commit/d6c63ebf463297a076984fe248aabeba0fa58864 [1] https://github.com/intel-staging/tdx/commit/782649ab74c1b3f05df7e38d5fc8e3f907a9394e ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-09-22 8:09 ` Artem Bityutskiy 2026-09-22 9:42 ` Tony Lindgren @ 2026-09-22 21:18 ` Peter Xu 2026-09-23 12:05 ` Artem Bityutskiy 2026-09-23 15:28 ` Serge Hallyn (AMD) 2 siblings, 1 reply; 84+ messages in thread From: Peter Xu @ 2026-09-22 21:18 UTC (permalink / raw) To: Artem Bityutskiy Cc: Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Tue, Sep 22, 2026 at 11:09:42AM +0300, Artem Bityutskiy wrote: > Hi Peter, Hello, Artem, > > thanks again for good comments and questions. My pleasure if I helped anything at all. [...] > > Yes, thanks. I'm willing to hear more, and I'm a bit surprised that > > non-coco migration didn't fit already well into it, because IIUC non-coco > > needs less in this case, not more. > > Yeah, just thinking aloud now. > > - Suppose QEMU and KVM support CoCo migration. > - Does it make sense to treat traditional VM as a CoCo VM for migration > purposes? > - Not for performance reasons - mmapped memory I/O is going to be more > efficient than any sorts of export/import uAPIs. > - Yes for testing/experimental purposes. > - Would it add a lot of code and complexity: would the benefits outweigh the > costs? > - QEMU: Should require minimum amount of code. Would need a flag like > "treat me as CoCo VM", may be some QEMU operator commands or cmdline > options. In QEMU, we could likely stick with calling them non-CoCo VMs even if the new migration API will work for them. We can have a qemu flag (kvm_coco_migration_api), which will be required / enforced by CoCo VMs to migrate. Then if the new API will support non-CoCo, we can export that kvm flag so non-CoCo can select it too. > - KVM: Not sure, but intuitively it may add noticeable amount of churn. True. Also because of this, personally I would not request for this to be supported, only if you all see fit or value. Said that, this PoC problem that Kishen described PAM in the other email helped me to notice something very important that I overlooked, that at least I want to understand. I'll reply there later separately. [...] > > I am just thinking out loud here: if the hardware is good enough to do > > encryption plus (some?) compression, that would be very nice. For "some", > > I meant minimum over zero pages. Because in this case even if host wants > > to play tricks with zero pages, it can't anymore when un-readable. Maybe > > the guest driver can play some trick, but I'm also not sure if in CoCo. > > > > M > N may imply it's not the case for now, but it's still sane as a start > > even if so. > > Yeah. In case of TDX, the TDX module does not do compression today, > and there is no zero-page optimization either. > > But keep in mind that what pages in TD are zero pages is by itself a secret. > Any sort of theoretical TDX module zero page compression optimization would > need to be done in a way that VMM cannot figure out what pages were zero > pages. So with a disclaimer that I am not really a security expert, I'd say > it is far-fetched. Agreed. That's why I think TDX, if able to compress zero pages in that buffer stream, will hide that information into the encrypted stream, so it's still not visible to VMM, but should be visible to destination TDX. But yeah, this is slightly off topic, we can skip this one for now regardless. [...] > OK, so by upper bound here you mean the maximum amount of data that can be > exported in a single memory export ioctl operation, right? And you mean that > there should be some sort of command to query it, similar to KVM_CAP_XSAVE2 > capability returning the XSAVE buffer size? If so - yes, I think exposing > such an upper bound would be useful. > > Let's see. In case of TDX, the `TDH.EXPORT.MEM` seamcall today can export > max. 512 pages at a time. Today only 4KiB pages are supported, so it is just > a 2MiB buffer. So this is the upper bound for the seamcall. > > So in TDX case, 512 pages would be the absolute upper bound for the uAPI > input buffer size. The alternative is to accept larger input buffers and > internally do multiple seamcalls. For TDX case, I do not think the latter > makes much sense though. But, what I do think is that the uAPI itself > shouldn't mandate one approach over the other - a different CoCo > implementation may prefer to aggregate multiple hardware calls internally. > > Also, today TDX can only migrate 4KiB pages - VMM must split larger pages > into 4KiB pages before migration. But I expect that in the future TDX may > implement larger pages migration too. So I think the cap should be expressed > in page count rather than bytes. > > That said, I don't think exposing this upper bound cap should replace the > flexibility we discussed earlier: the API itself should still allow > exporting as much as fits the output buffer, regardless of what the cap > reports. A specific CoCo implementation could still choose to just fail if > the buffer isn't large enough to hold all requested pages - but that would > be an implementation limitation, not an API limitation. > > And the other question is whether to allow exporting a mix of large and > small pages, like N x 4KiB along with M x 2MiB and K x 1GiB pages. Or the > uAPI would allow to only export one page size at a time? > > If large pages are added to TDX, I'd speculate that TDX seamcall would allow > the mixing and matching - I can see that already today the seamcall API > is provisioned for this. uAPI could require that the input buffer (set of > pages) can contain a mix which should match the GPAs of the corresponding > memory regions. But then race conditions - the GPA layout of the VM may > change by the time the request reaches the hardware: large pages can be > split or the other way round, I guess? Then uAPI could return a "retry" sort > of exit code, may be? > > What do you think? I am just thinking aloud: looks like pages of different > size add a degree of complexity - I did not take that into account until > this conversation, thanks. I apologize if my question was misleading. My question was more about the size of user buffer QEMU needs to prepare. I think I got the answer from the other email. Still, thanks for sharing your thoughts on huge page handling. Personally I'm not sure if we need to identify huge pages: we can always still define the GFNs to always represent PAGE_SIZE no matter the size of host pages, especially in migration context, we'll need to split in the first place. But that'll be a question to ask for later. > > > > > Said that, we'll need to be careful then in case of migration fallbacks at > > the final stage. Nowadays, I believe QEMU can still fallback to source side > > at a very, very late stage after all things applied. If I'm not mistaken, > > the final handshake is done at migration_incoming_state_destroy() -> > > migrate_send_rp_shut() telling source to be gone. > > Right. In traditional pre-copy VM migration model, the fallback is possible > at any point before the destination VM starts running and modifying its > state. > > The same is true for TDX, but with more complexity and limitations. > > We have 2 points: > 1. Before the source has exported the start token - fallback is similar to > traditional VM migration - just abort the migration on source and > continue running the source TD, and just destroy the destination TD. > 2. After the source has exported the start token - fallback is still > possible, but more complex - it requires the destination to first > generate the abort token, which should be delivered to the source and > consumed there. There are seamcalls for both generating and consuming the > abort token. The token basically makes sure the destination cannot run, > and the source can run again - same "only one of the TDs can run at a > time" security rule that we discussed earlier. > > Now, the limitation here is that if the abort is because the network > connectivity is lost, the abort token cannot be delivered. Then I'd guess > migration should stop, without destroying the source and the destination, > and the token should be delivered later manually, QEMU could even provide an > infrastructure / commands for that. Heh, this reminded me of postcopy recover that QEMU supported for years: not a trivial feature at all, people always want it to be there, but I'm not sure who is using it at all in production.. Especially, normally who cares about migration will provide dedicated and reliable networking for migration purpose, so the failure will be even more unlikely. Meanwhile who doesn't care enough, may simply crash the VM at a postcopy failure and reboot.. not bother to recover. I bet rarely people will hit a network failure happens exactly at forwarding the TDX START token to destination host, similarly for case (2) on ABORT token lost. But yes, likely this is required to make the whole thing complete. > > Would be very helpful to learn the abort protocol the other CoCo vendors > offer - is it similar to TDX or not? > > In our PoC we did not implement the TDX abort token procedure. But the > proposed uAPI is provisioned for the abort token export/import. [...] > > Hmm, interesting. Sorry if this follow up question may be slightly > > off-topic of the new API, or maybe it matters, depends on the answer: do > > you know who is in charge of this unplug / plug operation? VMM or TD > > (hence, transparent to VMM)? > > > > If it's VMM, is QEMU involved? > > Specifically about TDX - the pause seamcall will return an error if TDIs > are not unassigned. > > While this is something that is not implemented in Linux yet, I believe the > model will be that there is some uAPI to unassign TDIs, and it is not > related to migration. QEMU would just need to exercise this uAPI at the > right time. OK, this sounds working, but then it means the migration will be visible to the guest. I wonder whether there's any attempt to make it more transparent, but we can also leave this question for later. [...] > Do you or someone else know how important it is to have it working? > Do actual customers use it in real setups and rely on it? > How bad or painful is it if this feature is not available? > > The reason I am asking is that my assumption was that it is not important. > But if it is, I will come back to the TDX module architects with a > request to revise the design to support independent dirty page tracking. Of > course they may have some reasons for not doing it, but I would try at > least. Thanks, I'll talk to our team and revisit this after I collect answers. [...] > Understood, thanks - sounds like this workaround is doable, but wasteful, > something to avoid if possible. That in itself is useful evidence I can > bring back to the TDX module architects. > > Just to understand the real-world deployments better: do you know if anyone > actually falls back to this workaround today, or is it more of a theoretical > idea that nobody actually uses? Nop, I am not aware of any CoCo migration worked in any production deployments. So it's only a wild idea assuming GET_DIRTY_LOG will be gone without migration process. > > > > > Btw, do you know if the private pages will be write-trappable by the host > > kernel, with things like userfaultfd-wp or soft-dirty? That'll be another > > way to do this without TD involvement, but I don't know well enough to say. > > The short answer is - no, private pages are not write-trappable outside of > a migration session, so trapping cannot be used as a workaround for > independent dirty tracking today. > > But a bit longer answer is that the TDX module provides 2 alternative dirty > page tracking mechanisms, and one of them is in fact a write-trapping > mechanism. In the TDX specs, these mechanisms are called Write-Blocking > Export and Non-Blocking Export. > > - Write-blocking: the VMM write-blocks pages using `TDH.EXPORT.BLOCKW`, and > gets a VM exit (an EPT violation) when the TD attempts to write a > blocked page. > - Non-blocking: based on Secure EPT scanning, using `TDH.MEM.SCAN.RANGE`, > which was mentioned earlier in this cover letter. I refer to it as dirty > scanning. > > Today, both mechanisms are only available during migration and require the > migration setup to be done first - neither can be used as a standalone. In > our PoC we use the dirty scanning mechanism, as we expect it to be more > performant. > > But our PoC actually started with the write-blocking method. According to > Kishen and Tony, it required a lot of ugly code and was very intrusive into > the MMU code. But I'll let Kishen and Tony provide more details. > > Anyway, if I can get the TDX module to allow dirty tracking independent of > migration, we could pick either mechanism, or even implement both and add a > TDX-specific ioctl to select which one to use. We could plug either method > into the KVM_DIRTY_LOG ioctl - both would work, just differently. > > = Dirty Scanning Optimization in TDX Module = > > Now, I am diverging, but just in case: in the PUCK call where we presented > TDX migration, Sean made an immediate observation that WP-based dirty > tracking is not categorically worse than PML-based dirty ring tracking, it > depends on the workload. Hmm, I was expecting PML is still superior in most cases. For "depending on workloads", is that perhaps when (1) huge pages are used, and (2) the workload writes only a small portion of guest memory? Then maybe RO huge page will be a good hint to skip whole huge-page-range. But that's only my wild guess, could be wrong. > Then he learned that TDX module's dirty scanning does not use PML, and > was understandably surprised. Sean was concerned about dirty scanning > performance. I'm definitely surprised too that PML isn't used. Could I ask if there's any simple reason not to use it for TDX? Per my understanding, PML works with all kinds of loads, and I was expecting PML to be efficient and most ideal. Another approach is if TDX can take over the bitmap buffer from the relevant kvm memslots, update directly there alongside setting D bits in EPT PTEs; after all IIUC we assumed dirty info not part of confidential materials. But that sounds more complex than PML if it's already working for years. The current scan approach sounds like unpredictable in terms of downtime, in that even if with a scanner I don't see how TDX can guaratee the downtime for the last dirty sync from QEMU, which will completely be part of the blackout downtime. Whenever switchover decision made, QEMU stops the VM, do the last time sync, move anything left. So that last one will matter a lot if we care about downtime. Thanks, -- Peter Xu ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-09-22 21:18 ` Peter Xu @ 2026-09-23 12:05 ` Artem Bityutskiy 2026-09-24 21:19 ` Peter Xu 0 siblings, 1 reply; 84+ messages in thread From: Artem Bityutskiy @ 2026-09-23 12:05 UTC (permalink / raw) To: Peter Xu Cc: Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Tue, 2026-09-22 at 17:18 -0400, Peter Xu wrote: > > thanks again for good comments and questions. > > My pleasure if I helped anything at all. Honestly, yes. New insights, and even commenting on questions makes me research more and dig deeper. > > > > Now, the limitation here is that if the abort is because the network > > connectivity is lost, the abort token cannot be delivered. Then I'd guess > > migration should stop, without destroying the source and the destination, > > and the token should be delivered later manually, QEMU could even provide an > > infrastructure / commands for that. > > Heh, this reminded me of postcopy recover that QEMU supported for years: > not a trivial feature at all, people always want it to be there, but I'm > not sure who is using it at all in production.. > > Especially, normally who cares about migration will provide dedicated and > reliable networking for migration purpose, so the failure will be even more > unlikely. Meanwhile who doesn't care enough, may simply crash the VM at a > postcopy failure and reboot.. not bother to recover. > > I bet rarely people will hit a network failure happens exactly at > forwarding the TDX START token to destination host, similarly for case (2) > on ABORT token lost. But yes, likely this is required to make the whole > thing complete. Thanks for the insights about users who really care about having a reliable dedicated network. I would like to make sure I understand correctly though. You refer to QEMU post-copy recovery, which is basically all about restoring network connectivity and finishing the migration. Is this right? And to make sure we are on the same page, here is how I see things at a high level. 1. Post-copy recovery is about handling the state split between source and destination. I think this aspect is going to be similar, if not the same, for traditional and CoCo VM migration. Same problem, same recovery strategies, I'd guess. 2. The abort token stuff I discussed is about pre-copy. It is an artifact of the switchover: when src is paused and dst is allowed to start, an abort token is needed to reverse this process. And whether post-copy is used or not is orthogonal to the abort token stuff. Did I miss something? > > > Specifically about TDX - the pause seamcall will return an error if TDIs > > are not unassigned. > > > > While this is something that is not implemented in Linux yet, I believe the > > model will be that there is some uAPI to unassign TDIs, and it is not > > related to migration. QEMU would just need to exercise this uAPI at the > > right time. > > OK, this sounds working, but then it means the migration will be visible to > the guest. I wonder whether there's any attempt to make it more > transparent, but we can also leave this question for later. Correct. With a disclaimer that I am not a TDX Connect expert, I had the impression that this is more of a compromise solution. The VMM is an untrusted entity in the CoCo model, so the trust between the TD and the PCIe/CXL TDX Connect device is built by the TD itself. It is the TD establishing the cryptographic trust with a specific physical PCIe/CXL device. Directly, not via VMM. This is part of the industry standard SPDM protocol, which stands for Security Protocol and Data Model. As I understand it, with the current SPDM protocol, this trust cannot be transparently moved from one host to another: the destination host has a different physical device, with different cryptographic keys, and the TD would need to re-establish trust explicitly. VMM cannot do it on behalf of the TD in a transparent manner. I would only speculate that this means preserving the device state is a hard problem. Maybe future TDX Connect and SPDM revisions will solve this problem, but for now, this is the compromise solution we have. But again, take it with a grain of salt, it is more of my intuition than based on concrete knowledge. > > The reason I am asking is that my assumption was that it is not important. > > But if it is, I will come back to the TDX module architects with a > > request to revise the design to support independent dirty page tracking. Of > > course they may have some reasons for not doing it, but I would try at > > least. > > Thanks, I'll talk to our team and revisit this after I collect answers. Many thanks! > > Now, I am diverging, but just in case: in the PUCK call where we presented > > TDX migration, Sean made an immediate observation that WP-based dirty > > tracking is not categorically worse than PML-based dirty ring tracking, it > > depends on the workload. > > Hmm, I was expecting PML is still superior in most cases. For "depending > on workloads", is that perhaps when (1) huge pages are used, and (2) the > workload writes only a small portion of guest memory? Sorry, I did not communicate it correctly.Sean did not talk about huge pages. I need to be very careful here. What I think was Sean's point is that VM exits forced by the WP-based dirty tracking cause "back-pressure" as he put it, meaning they work as a natural way to slow down vCPUs and improve migration convergence. PML-based tracking does not cause as much back-pressure, so, depending on workload, they may require artificial vCPU throttling. But disclaimer, this is not a cite, this is my interpretation. > > Then he learned that TDX module's dirty scanning does not use PML, and > > was understandably surprised. Sean was concerned about dirty scanning > > performance. > > I'm definitely surprised too that PML isn't used. Could I ask if there's > any simple reason not to use it for TDX? Per my understanding, PML works > with all kinds of loads, and I was expecting PML to be efficient and most > ideal. Another point where I need to be careful to not miscommunicate. The honest answer is that I do not know for sure why PML specifically was not chosen. Please take my comments below with a grain of salt - I am a software person who tries to understand the design decisions made in the TDX module, but I am not a TDX module architect. Current Intel processors do not support PML for DMA - a TDI's DMA writes would go untracked by PML. But they do set the Secure EPT Dirty bit, so scanning works for catching both CPU and DMA writes. My speculation is that once non-blocking scanning had to be built to cover the TDX Connect case, it made sense to use it as the single mechanism for TD migration in general. AFAIU, the TDX guest migration implementation benchmarking results are satisfactory with the scanning approach, but I do not have hard numbers to share. I also feel that PML could offer better performance, at least for memory-intensive workloads. But this is intuition only. > Another approach is if TDX can take over the bitmap buffer from the > relevant kvm memslots, update directly there alongside setting D bits in > EPT PTEs; after all IIUC we assumed dirty info not part of confidential > materials. But that sounds more complex than PML if it's already working > for years. Well, then the dirty bitmap specifics would become the ABI - the hard contract between the Intel platform and the OS. But this is effectively what TDX module dirty scanning does already today: on input you give it an array of up to 512 GPAs, on the output it marks which entries are migration candidates and also for what reason. Keep in mind that dirty pages are the majority of migration candidates, but not all of them. Sometimes a migration candidate can be a page that was already exported, but then was, for example, converted from private to shared, or unaccepted by the TD (gone, in other words). In this case the TDX module flags it as a migration candidate too. The memory export seamcall treats it differently too - instead of exporting encrypted page data, it exports a small record indicating that the page has changed its status (gone). IOW, in the TDX migration case, it is not just dirty pages. In an abstract way, it is useful to think of it as the TDX module tracking both page data and metadata changes. We (me, Kishen, Tony) call them "dirty pages" for simplicity, but TDX specs use the term "migration candidates". But again, most of them are dirty pages. > The current scan approach sounds like unpredictable in terms of downtime, > in that even if with a scanner I don't see how TDX can guaratee the > downtime for the last dirty sync from QEMU, which will completely be part > of the blackout downtime. Could you help me understand exactly what you mean by "predictable" here? Let me walk through how I see it. Please correct me if I am wrong. For a traditional VM: at some point QEMU decides that pre-copy has converged. But the source VM keeps running until it is actually paused, and it can dirty more pages. QEMU has no way to know in advance how many more dirty pages there will be by the time the source VM is actually paused - could be a few, could be a lot. So the time it takes to find and copy them during the downtime is not fully predictable. The same logic applies to a TDX guest using dirty scanning. Suppose the final dirty scan is slower than a theoretical TDX PML-based approach would have been. The prescan optimization I described in the previous e-mail should help with this in an average case, but let's assume it does not help for some special case - when the TD touches most of its memory, so all EPT sub-trees end up touched. This should be rare, but let's assume it happens. In this case, whether that scan slowdown will actually matter for the overall downtime also depends on the network. If the network is very fast, the scan itself can become the dominant part of the downtime. If the network is the slower part, the scan slowdown may barely matter. So my understanding is that downtime is never fully predictable, for either type of VM. Intuitively, the TDX case does feel "less predictable". The open question for me is whether the degree of unpredictability is large enough to bother users. My attitude is to focus on getting something simple done first, learn from real-world behavior, and improve it later if needed, including exploring PML. The best is the enemy of the good sort of attitude. > Whenever switchover decision made, QEMU stops the VM, do the last time > sync, move anything left. So that last one will matter a lot if we care > about downtime. Could you please help me understand: in your experience, how much do users care about the overall migration time? My current assumption is that users mostly care about downtime, and care little about the overall migration time. I guess nobody wants migration to take days, but if it takes, say, 10 minutes, I assume no one would put in a lot of effort to improve it to 9 minutes. Thanks, Artem. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-09-23 12:05 ` Artem Bityutskiy @ 2026-09-24 21:19 ` Peter Xu 2026-09-28 14:15 ` Artem Bityutskiy 0 siblings, 1 reply; 84+ messages in thread From: Peter Xu @ 2026-09-24 21:19 UTC (permalink / raw) To: Artem Bityutskiy Cc: Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Wed, Sep 23, 2026 at 03:05:38PM +0300, Artem Bityutskiy wrote: [...] > I would like to make sure I understand correctly though. You refer to QEMU > post-copy recovery, which is basically all about restoring network > connectivity and finishing the migration. Is this right? Correct. > > And to make sure we are on the same page, here is how I see things at a > high level. > > 1. Post-copy recovery is about handling the state split between source and > destination. I think this aspect is going to be similar, if not the same, > for traditional and CoCo VM migration. Same problem, same recovery > strategies, I'd guess. > > 2. The abort token stuff I discussed is about pre-copy. It is an artifact of > the switchover: when src is paused and dst is allowed to start, an abort > token is needed to reverse this process. And whether post-copy is used or > not is orthogonal to the abort token stuff. > > Did I miss something? I believe we're on the same page. I mentioned the recovery feature because both of them (even if ABORT is part of precopy rather than postcopy) describe such an use case where network interruption can cause some form of split brain of the VM, causing neither side be able to continue. We used to not have such case with precopy, but then this ABORT / START message can make it happen similarly like postcopy. Said that, the window is much smaller than postcopy. > > > > > > Specifically about TDX - the pause seamcall will return an error if TDIs > > > are not unassigned. > > > > > > While this is something that is not implemented in Linux yet, I believe the > > > model will be that there is some uAPI to unassign TDIs, and it is not > > > related to migration. QEMU would just need to exercise this uAPI at the > > > right time. > > > > OK, this sounds working, but then it means the migration will be visible to > > the guest. I wonder whether there's any attempt to make it more > > transparent, but we can also leave this question for later. > > Correct. With a disclaimer that I am not a TDX Connect expert, I had the > impression that this is more of a compromise solution. The VMM is an > untrusted entity in the CoCo model, so the trust between the TD and the > PCIe/CXL TDX Connect device is built by the TD itself. It is the TD > establishing the cryptographic trust with a specific physical PCIe/CXL > device. Directly, not via VMM. This is part of the industry standard SPDM > protocol, which stands for Security Protocol and Data Model. > > As I understand it, with the current SPDM protocol, this trust cannot be > transparently moved from one host to another: the destination host has a > different physical device, with different cryptographic keys, and the TD > would need to re-establish trust explicitly. VMM cannot do it on behalf of > the TD in a transparent manner. I would only speculate that this means > preserving the device state is a hard problem. > > Maybe future TDX Connect and SPDM revisions will solve this problem, but for > now, this is the compromise solution we have. > > But again, take it with a grain of salt, it is more of my intuition than > based on concrete knowledge. AFAIU, preserving device states were a hard problem even on non-CoCo before, but then I guess people thought VFIO performs so good, after that people managed to work the problem out.. and now more people start to rely on VFIO precopy migrations working in the clusters, non-CoCo. I had a gut feeling it will happen too for CoCo some day, that unplug approach was exactly what happens before VFIO migration is implemented... But yes, let's leave this for later, thanks for sharing. > > > The reason I am asking is that my assumption was that it is not important. > > > But if it is, I will come back to the TDX module architects with a > > > request to revise the design to support independent dirty page tracking. Of > > > course they may have some reasons for not doing it, but I would try at > > > least. > > > > Thanks, I'll talk to our team and revisit this after I collect answers. > > Many thanks! I got some feedback on this, I'll try to provide a summary. So, first of all, calc_dirty_rate isn't seem to be widely used across our customers. However, we do have customer case using calc_dirty_rate to evaluate migrations of a VM fleet for cases like from one data centre to another. I think it makes sense because the normal "try to migrate and fallback otherwise" idea applies well to one VM, but perhaps not that good on a fleet. When a fleet is involved, we don't want to migrate 400 VMs then found there're 30 critical VMs too busy and can't migrate, then due to whatever reason (inter-VM communication / service locality ?) one is forced to migrate that 400 VMs backwards. IOW, it seems helpful to provide high-level evalutions of migration decisions over a full cluster, concurrently and efficiently. > > Hmm, I was expecting PML is still superior in most cases. For "depending > > on workloads", is that perhaps when (1) huge pages are used, and (2) the > > workload writes only a small portion of guest memory? > > Sorry, I did not communicate it correctly.Sean did not talk about huge > pages. I need to be very careful here. What I think was Sean's point is that > VM exits forced by the WP-based dirty tracking cause "back-pressure" as he > put it, meaning they work as a natural way to slow down vCPUs and improve > migration convergence. PML-based tracking does not cause as much > back-pressure, so, depending on workload, they may require artificial vCPU > throttling. But disclaimer, this is not a cite, this is my interpretation. No worires, thanks for sharing your thoughts. And I agree there is that back pressure effect. Migration performance is one of the most weird performance engineering topics I'm aware of for sure; sometimes, the better a work done, the less likely it converges.. > > > > Then he learned that TDX module's dirty scanning does not use PML, and > > > was understandably surprised. Sean was concerned about dirty scanning > > > performance. > > > > I'm definitely surprised too that PML isn't used. Could I ask if there's > > any simple reason not to use it for TDX? Per my understanding, PML works > > with all kinds of loads, and I was expecting PML to be efficient and most > > ideal. > > Another point where I need to be careful to not miscommunicate. The honest > answer is that I do not know for sure why PML specifically was not chosen. > Please take my comments below with a grain of salt - I am a software person > who tries to understand the design decisions made in the TDX module, but I > am not a TDX module architect. > > Current Intel processors do not support PML for DMA - a TDI's DMA writes > would go untracked by PML. But they do set the Secure EPT Dirty bit, so > scanning works for catching both CPU and DMA writes. > > My speculation is that once non-blocking scanning had to be built to cover > the TDX Connect case, it made sense to use it as the single mechanism for TD > migration in general. > > AFAIU, the TDX guest migration implementation benchmarking results are > satisfactory with the scanning approach, but I do not have hard numbers to > share. I also feel that PML could offer better performance, at least for > memory-intensive workloads. But this is intuition only. My gut feeling is DMA shouldn't be a blocker for PML: AFAIU we don't track DMA from KVM side. Assigned device should have its own dirty tracking for DMAs, either via device's own tracking facilities, or the IOMMU on the host. Feel free to refer to vfio_listener_log_sync() in QEMU. In all cases, it'll be great you could share the reason if you have more solid clues. > > > Another approach is if TDX can take over the bitmap buffer from the > > relevant kvm memslots, update directly there alongside setting D bits in > > EPT PTEs; after all IIUC we assumed dirty info not part of confidential > > materials. But that sounds more complex than PML if it's already working > > for years. > > Well, then the dirty bitmap specifics would become the ABI - the hard > contract between the Intel platform and the OS. > > But this is effectively what TDX module dirty scanning does already today: > on input you give it an array of up to 512 GPAs, on the output it marks > which entries are migration candidates and also for what reason. Are we talking about the memory export/import API or GET_DIRTY_LOG? IIUC, GET_DIRTY_LOG always applies to a whole memslot, struct kvm_dirty_log { __u32 slot; __u32 padding1; union { void *dirty_bitmap; /* one bit per page */ __u64 padding2; }; }; > > Keep in mind that dirty pages are the majority of migration candidates, but > not all of them. Sometimes a migration candidate can be a page that was > already exported, but then was, for example, converted from private to > shared, or unaccepted by the TD (gone, in other words). In this case the TDX > module flags it as a migration candidate too. The memory export seamcall > treats it differently too - instead of exporting encrypted page data, it > exports a small record indicating that the page has changed its status > (gone). Yes, it makes sense. So can I inteprete this as GET_DIRTY_LOG works seamlessly for both private and shared pages (or even, unaccepted pages)? Then I assume it means MEMORY.EXPORT should also be able to read shared or unaccepted pages too, am I right? Same to when apply with IMPORT. Another counter example is MEMORY.EXPORT returns a flag saying "this page is shared, go read it directly from HVA", but then QEMU reading it may race with a concurrent shared->private conversion crashing VMM. Looks to me MEMORY.EXPORT must support shared too, then. > > IOW, in the TDX migration case, it is not just dirty pages. In an abstract > way, it is useful to think of it as the TDX module tracking both page data > and metadata changes. > > We (me, Kishen, Tony) call them "dirty pages" for simplicity, but TDX specs > use the term "migration candidates". But again, most of them are dirty > pages. Yes, I didn't notice it before, but now I see that marking converted or unaccepted pages to be dirty makes sense. IIUC it's because that info (shared, or private, or unaccepted) is part of page [meta]data that needs to be migrated to reconstruct the whole VM on the other host. > > > The current scan approach sounds like unpredictable in terms of downtime, > > in that even if with a scanner I don't see how TDX can guaratee the > > downtime for the last dirty sync from QEMU, which will completely be part > > of the blackout downtime. > > Could you help me understand exactly what you mean by "predictable" here? > Let me walk through how I see it. Please correct me if I am wrong. > > For a traditional VM: at some point QEMU decides that pre-copy has > converged. But the source VM keeps running until it is actually paused, and > it can dirty more pages. QEMU has no way to know in advance how many more > dirty pages there will be by the time the source VM is actually paused - > could be a few, could be a lot. So the time it takes to find and copy them > during the downtime is not fully predictable. Correct. > > The same logic applies to a TDX guest using dirty scanning. Suppose the > final dirty scan is slower than a theoretical TDX PML-based approach would > have been. The prescan optimization I described in the previous e-mail > should help with this in an average case, but let's assume it does not > help for some special case - when the TD touches most of its memory, so > all EPT sub-trees end up touched. This should be rare, but let's assume it > happens. > > In this case, whether that scan slowdown will actually matter for the > overall downtime also depends on the network. If the network is very fast, > the scan itself can become the dominant part of the downtime. If the > network is the slower part, the scan slowdown may barely matter. > > So my understanding is that downtime is never fully predictable, for > either type of VM. Right, but IMHO background scan of EPT pgtable dirty bits adds a completely new reason to introduce downtime, and when I said "unpredictable", it is about that part. Also, I worry in some worst case this can be pretty large. So we have two overheads here at this stage, unpredictable: (a) Scanning EPT pgtable, when very unlucky, can take a lot of time to finally reports to a GET_DIRTY_LOG request, (b) Migrating of dirty pages during blackout phase, which should be roughly linear to how many dirty pages we just collected. (NOTE! I think we may have way to fix this (b) or optimize it.. but this is off-topic; let's focus on the difference of (a) and (b) first) When with PML, IIUC (a) is predictable: we have the bitmap on hand, plus a maximum of some (my memory is, 512?) PML entries to flush per vCPU. That'll be flushed automatically when we do vm_stop(), likely also concurrently, atomically updating the bitmaps. I never measured it, but it is bounded, and sounds pretty fast. When with scanning, (a) seems more unpredictable. That's the part I was slightly concerned. But now after thinking a bit more, it seems fine. Please read below. > > Intuitively, the TDX case does feel "less predictable". The open > question for me is whether the degree of unpredictability is large enough > to bother users. My attitude is to focus on getting something simple done > first, learn from real-world behavior, and improve it later if needed, > including exploring PML. The best is the enemy of the good sort of > attitude. Yes, I think it's always fine we start with whatever is most feasible. I think it actually may not be that bad. The last sync is special at least on how QEMU treats it, it should look like: - GET_DIRTY_LOG, to do last math, decide to switchover, <------ [1] - vm_stop() - GET_DIRTY_LOG, this collects all rest dirty bits <------ [2] - migrates the dirty pages, device states, etc. So I expect there should be normally very small window between two continuous GET_DIRTY_LOG across system. Only [2] will be part of downtime. Since you explained to me on how the background rescan roughly works, by relying on A bit in pgtable directory entries, I do feel like in this case most of the memory regions shouldn't be accessed during small window of [1]->[2], then the range to scan should be very much under control too. In reality, it will likely be even smaller, [1]->vm_stop(), because after that vCPUs are halted. So it may not really be an issue in practise, but it still depends. In all cases, some measurements after PoC ready would be nice on some large and relatively busy VMs. > > > Whenever switchover decision made, QEMU stops the VM, do the last time > > sync, move anything left. So that last one will matter a lot if we care > > about downtime. > > Could you please help me understand: in your experience, how much do users > care about the overall migration time? > > My current assumption is that users mostly care about downtime, and care > little about the overall migration time. I guess nobody wants migration to > take days, but if it takes, say, 10 minutes, I assume no one would put in a > lot of effort to improve it to 9 minutes. I agree. I think some use case may care about total migration time, but I would say in most cases, "10min or 9min total migration time" difference is less of a concern than downtime effects. Thanks, -- Peter Xu ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-09-24 21:19 ` Peter Xu @ 2026-09-28 14:15 ` Artem Bityutskiy 2026-09-29 21:05 ` Peter Xu 0 siblings, 1 reply; 84+ messages in thread From: Artem Bityutskiy @ 2026-09-28 14:15 UTC (permalink / raw) To: Peter Xu Cc: Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm Hi Peter, thanks for reply again. On Thu, 2026-09-24 at 17:19 -0400, Peter Xu wrote: > > > > The reason I am asking is that my assumption was that it is not important. > > > > But if it is, I will come back to the TDX module architects with a > > > > request to revise the design to support independent dirty page tracking. Of > > > > course they may have some reasons for not doing it, but I would try at > > > > least. > > > > > > Thanks, I'll talk to our team and revisit this after I collect answers. > > > > Many thanks! > > I got some feedback on this, I'll try to provide a summary. > > So, first of all, calc_dirty_rate isn't seem to be widely used across our > customers. > > However, we do have customer case using calc_dirty_rate to evaluate > migrations of a VM fleet for cases like from one data centre to another. > > I think it makes sense because the normal "try to migrate and fallback > otherwise" idea applies well to one VM, but perhaps not that good on a > fleet. Yes, that makes sense. > When a fleet is involved, we don't want to migrate 400 VMs then found > there're 30 critical VMs too busy and can't migrate, then due to whatever > reason (inter-VM communication / service locality ?) one is forced to > migrate that 400 VMs backwards. > > IOW, it seems helpful to provide high-level evalutions of migration > decisions over a full cluster, concurrently and efficiently. Thank you. I'll work on this internally. It will take time. Just to give wider context: Sean and Paolo gave us feedback regarding the entire SEPT scan approach - they believe it is too costly and won't scale, and suggested using PML instead. For now, dirty scanning is the best we have, but I continue to explore other options internally, and standalone dirty tracking is one of them. > > AFAIU, the TDX guest migration implementation benchmarking results are > > satisfactory with the scanning approach, but I do not have hard numbers to > > share. I also feel that PML could offer better performance, at least for > > memory-intensive workloads. But this is intuition only. > > My gut feeling is DMA shouldn't be a blocker for PML: AFAIU we don't track > DMA from KVM side. Assigned device should have its own dirty tracking for > DMAs, either via device's own tracking facilities, or the IOMMU on the > host. Feel free to refer to vfio_listener_log_sync() in QEMU. In all > cases, it'll be great you could share the reason if you have more solid > clues. Yeah, this is also part of the internal exploration I mentioned above too. > > > Are we talking about the memory export/import API or GET_DIRTY_LOG? IIUC, > GET_DIRTY_LOG always applies to a whole memslot, > > struct kvm_dirty_log { > __u32 slot; > __u32 padding1; > union { > void *dirty_bitmap; /* one bit per page */ > __u64 padding2; > }; > }; I apologize, I worte something unrelated to the context. 512 GPAs at a time is the limit for exporting the memory, not for dirty scanning. For the dirty scanning seamcall (TDH.MEM.SCAN.RANGE), the limit is 512 * 512 GPAs at a time, which is 262,144 GPAs, or 1GiB. So if a memslot is larger than 1GiB, multiple calls to the dirty scanning seamcall are needed. > > Keep in mind that dirty pages are the majority of migration candidates, but > > not all of them. Sometimes a migration candidate can be a page that was > > already exported, but then was, for example, converted from private to > > shared, or unaccepted by the TD (gone, in other words). In this case the TDX > > module flags it as a migration candidate too. The memory export seamcall > > treats it differently too - instead of exporting encrypted page data, it > > exports a small record indicating that the page has changed its status > > (gone). > > Yes, it makes sense. > > So can I inteprete this as GET_DIRTY_LOG works seamlessly for both private > and shared pages (or even, unaccepted pages)? > > Then I assume it means MEMORY.EXPORT should also be able to read shared or > unaccepted pages too, am I right? Same to when apply with IMPORT. Another > counter example is MEMORY.EXPORT returns a flag saying "this page is > shared, go read it directly from HVA", but then QEMU reading it may race > with a concurrent shared->private conversion crashing VMM. > > Looks to me MEMORY.EXPORT must support shared too, then. Hmm... First of all, it does sound like a possible approach. But it is not the approach we took in our PoC today. I hope Kishen will chime in to correct me. Here is how I saw this, but I may be missing something (my excuse is that I am still new to the team and still learning). 1. QEMU has a bitmap of shared pages in RAMBlockAttributes, so it can distinguish shared pages. 2. In general, QEMU does not distinguish private vs unaccepted pages, so unaccepted pages are treated as private pages. Dirty tracking: - QEMU uses the same KVM_DIRTY_LOG mechanism for tracking shared, private, and unaccepted pages. - For shared GFNs, KVM uses the normal VM dirty tracking mechanism. For private and unaccepted GFNs, KVM goes to the TDX-specific code. - But the final bitmap that QEMU sees covers all page types. - KVM calls TDH.MEM.SCAN.RANGE on both private and unaccepted GFNs. - For private GFNs, the TDX module reports it as a migration candidate if its data changed or its status changed (e.g., converted to shared or unaccepted). - For unaccepted GFNs, the TDX module reports it as a migration candidate in the first round (so it appears as dirty in KVM_DIRTY_LOG reply). Then it reports it as clean, unless its status changes - it becomes accepted. Page export: - QEMU migrates shared pages the old way - it does not try to use the proposed CoCo migration uAPI for that. - For private and unaccepted pages, QEMU uses the CoCo migration uAPI. - Our export uAPI PoC implementation does not try to check GFN type - it just feeds them all to the TDX module TDH.EXPORT.MEM seamcall. Here is what TDX module does depending on the page type: - Shared pages: just skip, no errors. - Private pages: export the data in encrypted form. - Unaccepted pages: export a small record telling that the page is unaccepted. This record should be delivered to the destination and imported there, just like private pages. But clearly this is part of the uAPI contract that must be discussed and made explicit. What I describe above is obviously our PoC implementation, plus TDX module behavior details. > > The same logic applies to a TDX guest using dirty scanning. Suppose the > > final dirty scan is slower than a theoretical TDX PML-based approach would > > have been. The prescan optimization I described in the previous e-mail > > should help with this in an average case, but let's assume it does not > > help for some special case - when the TD touches most of its memory, so > > all EPT sub-trees end up touched. This should be rare, but let's assume it > > happens. > > > > In this case, whether that scan slowdown will actually matter for the > > overall downtime also depends on the network. If the network is very fast, > > the scan itself can become the dominant part of the downtime. If the > > network is the slower part, the scan slowdown may barely matter. > > > > So my understanding is that downtime is never fully predictable, for > > either type of VM. > > Right, but IMHO background scan of EPT pgtable dirty bits adds a completely > new reason to introduce downtime, and when I said "unpredictable", it is > about that part. Also, I worry in some worst case this can be pretty large. Just to clarify on the "background" part. Yes, it is "background" relative to the TD - some CPU is running it, in parallel with vCPUs running on other CPUs. But there is no background activity in the TDX module itself. All the seamcalls run synchronously on the CPU that invokes them. This means, for example, that to speed up memory export, one can run the TDH.EXPORT.MEM seamcall for different GPAs in parallel on different CPUs. Same for dirty scanning - if one could run the TDH.MEM.SCAN.RANGE seamcall for different GPAs in parallel on different CPUs, that would make scanning much faster. Also a bit separately, one point to keep in mind is that in CoCo the page export and import are heavy, compute-intensive crypto operations. So when we talk about slower scanning, we need to keep in mind that it is not that slow in relation to the export crypto. But I understand that this is not an apples-to-apples comparison: - Scanning is potentially about a large SEPT in a VM with terabytes of memory. - Exporting is only about the pages found to be dirty. So this is only to remind that in the CoCo case, memory export also has a high price tag, compared to a traditional VM. > So we have two overheads here at this stage, unpredictable: > > (a) Scanning EPT pgtable, when very unlucky, can take a lot of time to > finally reports to a GET_DIRTY_LOG request, > > (b) Migrating of dirty pages during blackout phase, which should be > roughly linear to how many dirty pages we just collected. (NOTE! I > think we may have way to fix this (b) or optimize it.. but this is > off-topic; let's focus on the difference of (a) and (b) first) > > When with PML, IIUC (a) is predictable: we have the bitmap on hand, plus a > maximum of some (my memory is, 512?) PML entries to flush per vCPU. > That'll be flushed automatically when we do vm_stop(), likely also > concurrently, atomically updating the bitmaps. I never measured it, but it > is bounded, and sounds pretty fast. Yes, I agree. Now my secret desire is that TDX module can eventually plug PML under the hood, consider it a "hardware accelerator" without changing the ABI. But I do not know whether keeping the ABI unchanged is possible, or how soon it could happen. This is something I am working on internally with Intel TDX module team. > When with scanning, (a) seems more unpredictable. That's the part I was > slightly concerned. But now after thinking a bit more, it seems fine. > Please read below. Sure, thanks. > > Intuitively, the TDX case does feel "less predictable". The open > > question for me is whether the degree of unpredictability is large enough > > to bother users. My attitude is to focus on getting something simple done > > first, learn from real-world behavior, and improve it later if needed, > > including exploring PML. The best is the enemy of the good sort of > > attitude. > > Yes, I think it's always fine we start with whatever is most feasible. > > I think it actually may not be that bad. The last sync is special at least > on how QEMU treats it, it should look like: > > - GET_DIRTY_LOG, to do last math, decide to switchover, <------ [1] > - vm_stop() > - GET_DIRTY_LOG, this collects all rest dirty bits <------ [2] > - migrates the dirty pages, device states, etc. > > So I expect there should be normally very small window between two > continuous GET_DIRTY_LOG across system. Only [2] will be part of downtime. > > Since you explained to me on how the background rescan roughly works, by > relying on A bit in pgtable directory entries, I do feel like in this case > most of the memory regions shouldn't be accessed during small window of > [1]->[2], then the range to scan should be very much under control too. In > reality, it will likely be even smaller, [1]->vm_stop(), because after that > vCPUs are halted. Yes, I agree. Just to flag the word "background" again, and to make sure we are aligned - this "rescan" happens between [1] and [2]. The idea is that [2] will be very fast after the "rescan". But it does increase the time between [1] and [2], and the TD has time to dirty more pages. > So it may not really be an issue in practise, but it still depends. In all > cases, some measurements after PoC ready would be nice on some large and > relatively busy VMs. Yes, I agree. > Thanks, Artem. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-09-28 14:15 ` Artem Bityutskiy @ 2026-09-29 21:05 ` Peter Xu 2026-10-02 19:57 ` Artem Bityutskiy 0 siblings, 1 reply; 84+ messages in thread From: Peter Xu @ 2026-09-29 21:05 UTC (permalink / raw) To: Artem Bityutskiy Cc: Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Mon, Sep 28, 2026 at 05:15:00PM +0300, Artem Bityutskiy wrote: > Hi Peter, thanks for reply again. > > On Thu, 2026-09-24 at 17:19 -0400, Peter Xu wrote: > > > > > The reason I am asking is that my assumption was that it is not important. > > > > > But if it is, I will come back to the TDX module architects with a > > > > > request to revise the design to support independent dirty page tracking. Of > > > > > course they may have some reasons for not doing it, but I would try at > > > > > least. > > > > > > > > Thanks, I'll talk to our team and revisit this after I collect answers. > > > > > > Many thanks! > > > > I got some feedback on this, I'll try to provide a summary. > > > > So, first of all, calc_dirty_rate isn't seem to be widely used across our > > customers. > > > > However, we do have customer case using calc_dirty_rate to evaluate > > migrations of a VM fleet for cases like from one data centre to another. > > > > I think it makes sense because the normal "try to migrate and fallback > > otherwise" idea applies well to one VM, but perhaps not that good on a > > fleet. > > Yes, that makes sense. > > > When a fleet is involved, we don't want to migrate 400 VMs then found > > there're 30 critical VMs too busy and can't migrate, then due to whatever > > reason (inter-VM communication / service locality ?) one is forced to > > migrate that 400 VMs backwards. > > > > IOW, it seems helpful to provide high-level evalutions of migration > > decisions over a full cluster, concurrently and efficiently. > > Thank you. I'll work on this internally. It will take time. Yes, thanks. Nothing urgent, good to know it's on the map. > > Just to give wider context: Sean and Paolo gave us feedback regarding the > entire SEPT scan approach - they believe it is too costly and won't scale, > and suggested using PML instead. For now, dirty scanning is the best we > have, but I continue to explore other options internally, and standalone > dirty tracking is one of them. Ok. > > > > AFAIU, the TDX guest migration implementation benchmarking results are > > > satisfactory with the scanning approach, but I do not have hard numbers to > > > share. I also feel that PML could offer better performance, at least for > > > memory-intensive workloads. But this is intuition only. > > > > My gut feeling is DMA shouldn't be a blocker for PML: AFAIU we don't track > > DMA from KVM side. Assigned device should have its own dirty tracking for > > DMAs, either via device's own tracking facilities, or the IOMMU on the > > host. Feel free to refer to vfio_listener_log_sync() in QEMU. In all > > cases, it'll be great you could share the reason if you have more solid > > clues. > > Yeah, this is also part of the internal exploration I mentioned above too. > > > > > > Are we talking about the memory export/import API or GET_DIRTY_LOG? IIUC, > > GET_DIRTY_LOG always applies to a whole memslot, > > > > struct kvm_dirty_log { > > __u32 slot; > > __u32 padding1; > > union { > > void *dirty_bitmap; /* one bit per page */ > > __u64 padding2; > > }; > > }; > > I apologize, I worte something unrelated to the context. 512 GPAs at a time > is the limit for exporting the memory, not for dirty scanning. > > For the dirty scanning seamcall (TDH.MEM.SCAN.RANGE), the limit is 512 * 512 > GPAs at a time, which is 262,144 GPAs, or 1GiB. So if a memslot is larger > than 1GiB, multiple calls to the dirty scanning seamcall are needed. > > > > Keep in mind that dirty pages are the majority of migration candidates, but > > > not all of them. Sometimes a migration candidate can be a page that was > > > already exported, but then was, for example, converted from private to > > > shared, or unaccepted by the TD (gone, in other words). In this case the TDX > > > module flags it as a migration candidate too. The memory export seamcall > > > treats it differently too - instead of exporting encrypted page data, it > > > exports a small record indicating that the page has changed its status > > > (gone). > > > > Yes, it makes sense. > > > > So can I inteprete this as GET_DIRTY_LOG works seamlessly for both private > > and shared pages (or even, unaccepted pages)? > > > > Then I assume it means MEMORY.EXPORT should also be able to read shared or > > unaccepted pages too, am I right? Same to when apply with IMPORT. Another > > counter example is MEMORY.EXPORT returns a flag saying "this page is > > shared, go read it directly from HVA", but then QEMU reading it may race > > with a concurrent shared->private conversion crashing VMM. > > > > Looks to me MEMORY.EXPORT must support shared too, then. > > Hmm... First of all, it does sound like a possible approach. But it is not > the approach we took in our PoC today. I hope Kishen will chime in to > correct me. > > Here is how I saw this, but I may be missing something (my excuse is that I > am still new to the team and still learning). > > 1. QEMU has a bitmap of shared pages in RAMBlockAttributes, so it can > distinguish shared pages. > 2. In general, QEMU does not distinguish private vs unaccepted pages, so > unaccepted pages are treated as private pages. Yes, the latter seems uncontroversial. The 1st one is true, and it just reminded me if the conversion is synchronous and one step requires the hypercall to QEMU, then indeed background conversion can be avoided by some form of userspace locking. Perhaps, a rwlock suites, each vCPU takes it for write whenever page conversion requested from the guest (private <-> shared; nothing about "accepted" that matters). Then the migration threads, one or multiple, take the read lock, lookup the bit, do MEM.EXPORT, unlock. Then it seems fine in general, except that I donno if things can still go wrong when there are multiple versions of "if this page is private or shared". Say, minimum of three? (a) QEMU maintains the bitmap in RAMBlockAttributes, each bit represents if the page is shared or private (b) KVM should maintain one, looks to me, kvm->mem_attr_array (c) Hardware / Firmware may maintain its own, in case of TDX, is that one bit on the SEPT pgtable? They don't change together, AFAIU, they change in order, I believe (c)->(a)->(b) if my above understanding is correct. Then, what if they report different things? Say, during migration the guest wants to convert a page from shared to private. (c) can be already done saying one page "private" now for TDX, (a) tries to mark it "private" too, but now assuming page being accessed (read lock held), it may be trying to take a write lock and sleep, which means (b) will be "shared" so far. So what happens is, QEMU thinks this page "shared" because the conversion hasn't take place waiting for the write lock, however at least TDX may think it already "private" instead. Then QEMU logically can access HVA of that page, with (a)=shared, (b)=shared, (c)=private. Would it cause trouble? > > Dirty tracking: > > - QEMU uses the same KVM_DIRTY_LOG mechanism for tracking shared, private, > and unaccepted pages. > - For shared GFNs, KVM uses the normal VM dirty tracking mechanism. For > private and unaccepted GFNs, KVM goes to the TDX-specific code. > - But the final bitmap that QEMU sees covers all page types. > - KVM calls TDH.MEM.SCAN.RANGE on both private and unaccepted GFNs. > - For private GFNs, the TDX module reports it as a migration candidate if > its data changed or its status changed (e.g., converted to shared or > unaccepted). > - For unaccepted GFNs, the TDX module reports it as a migration candidate in > the first round (so it appears as dirty in KVM_DIRTY_LOG reply). Then it > reports it as clean, unless its status changes - it becomes accepted. > > Page export: > > - QEMU migrates shared pages the old way - it does not try to use the > proposed CoCo migration uAPI for that. > - For private and unaccepted pages, QEMU uses the CoCo migration uAPI. > - Our export uAPI PoC implementation does not try to check GFN type - it > just feeds them all to the TDX module TDH.EXPORT.MEM seamcall. Here > is what TDX module does depending on the page type: > - Shared pages: just skip, no errors. > - Private pages: export the data in encrypted form. > - Unaccepted pages: export a small record telling that the page is > unaccepted. This record should be delivered to the destination and > imported there, just like private pages. > > But clearly this is part of the uAPI contract that must be discussed and > made explicit. What I describe above is obviously our PoC implementation, > plus TDX module behavior details. > > > > The same logic applies to a TDX guest using dirty scanning. Suppose the > > > final dirty scan is slower than a theoretical TDX PML-based approach would > > > have been. The prescan optimization I described in the previous e-mail > > > should help with this in an average case, but let's assume it does not > > > help for some special case - when the TD touches most of its memory, so > > > all EPT sub-trees end up touched. This should be rare, but let's assume it > > > happens. > > > > > > In this case, whether that scan slowdown will actually matter for the > > > overall downtime also depends on the network. If the network is very fast, > > > the scan itself can become the dominant part of the downtime. If the > > > network is the slower part, the scan slowdown may barely matter. > > > > > > So my understanding is that downtime is never fully predictable, for > > > either type of VM. > > > > Right, but IMHO background scan of EPT pgtable dirty bits adds a completely > > new reason to introduce downtime, and when I said "unpredictable", it is > > about that part. Also, I worry in some worst case this can be pretty large. > > Just to clarify on the "background" part. Yes, it is "background" relative > to the TD - some CPU is running it, in parallel with vCPUs running on other > CPUs. > > But there is no background activity in the TDX module itself. All the > seamcalls run synchronously on the CPU that invokes them. > > This means, for example, that to speed up memory export, one can run the > TDH.EXPORT.MEM seamcall for different GPAs in parallel on different CPUs. > > Same for dirty scanning - if one could run the TDH.MEM.SCAN.RANGE seamcall > for different GPAs in parallel on different CPUs, that would make scanning > much faster. > > Also a bit separately, one point to keep in mind is that in CoCo the page > export and import are heavy, compute-intensive crypto operations. So when we > talk about slower scanning, we need to keep in mind that it is not that slow > in relation to the export crypto. But I understand that this is not an > apples-to-apples comparison: > > - Scanning is potentially about a large SEPT in a VM with terabytes of > memory. > - Exporting is only about the pages found to be dirty. > > So this is only to remind that in the CoCo case, memory export also has a > high price tag, compared to a traditional VM. Ok, I'll keep that in mind. > > > So we have two overheads here at this stage, unpredictable: > > > > (a) Scanning EPT pgtable, when very unlucky, can take a lot of time to > > finally reports to a GET_DIRTY_LOG request, > > > > (b) Migrating of dirty pages during blackout phase, which should be > > roughly linear to how many dirty pages we just collected. (NOTE! I > > think we may have way to fix this (b) or optimize it.. but this is > > off-topic; let's focus on the difference of (a) and (b) first) > > > > When with PML, IIUC (a) is predictable: we have the bitmap on hand, plus a > > maximum of some (my memory is, 512?) PML entries to flush per vCPU. > > That'll be flushed automatically when we do vm_stop(), likely also > > concurrently, atomically updating the bitmaps. I never measured it, but it > > is bounded, and sounds pretty fast. > > Yes, I agree. > > Now my secret desire is that TDX module can eventually plug PML under the > hood, consider it a "hardware accelerator" without changing the ABI. But I > do not know whether keeping the ABI unchanged is possible, or how soon it > could happen. This is something I am working on internally with Intel TDX > module team. > > > When with scanning, (a) seems more unpredictable. That's the part I was > > slightly concerned. But now after thinking a bit more, it seems fine. > > Please read below. > > Sure, thanks. > > > > Intuitively, the TDX case does feel "less predictable". The open > > > question for me is whether the degree of unpredictability is large enough > > > to bother users. My attitude is to focus on getting something simple done > > > first, learn from real-world behavior, and improve it later if needed, > > > including exploring PML. The best is the enemy of the good sort of > > > attitude. > > > > Yes, I think it's always fine we start with whatever is most feasible. > > > > I think it actually may not be that bad. The last sync is special at least > > on how QEMU treats it, it should look like: > > > > - GET_DIRTY_LOG, to do last math, decide to switchover, <------ [1] > > - vm_stop() > > - GET_DIRTY_LOG, this collects all rest dirty bits <------ [2] > > - migrates the dirty pages, device states, etc. > > > > So I expect there should be normally very small window between two > > continuous GET_DIRTY_LOG across system. Only [2] will be part of downtime. > > > > Since you explained to me on how the background rescan roughly works, by > > relying on A bit in pgtable directory entries, I do feel like in this case > > most of the memory regions shouldn't be accessed during small window of > > [1]->[2], then the range to scan should be very much under control too. In > > reality, it will likely be even smaller, [1]->vm_stop(), because after that > > vCPUs are halted. > > Yes, I agree. Just to flag the word "background" again, and to make sure we > are aligned - this "rescan" happens between [1] and [2]. The idea is that > [2] will be very fast after the "rescan". But it does increase the time > between [1] and [2], and the TD has time to dirty more pages. In general, looks like either way we can still stick with GET_DIRTY_LOG. PML would be better, otherwise from a review on the API level, that's still ok. Thanks, -- Peter Xu ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-09-29 21:05 ` Peter Xu @ 2026-10-02 19:57 ` Artem Bityutskiy 2026-10-07 20:00 ` Peter Xu 0 siblings, 1 reply; 84+ messages in thread From: Artem Bityutskiy @ 2026-10-02 19:57 UTC (permalink / raw) To: Peter Xu Cc: Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Tue, 2026-09-29 at 17:05 -0400, Peter Xu wrote: > > Here is how I saw this, but I may be missing something (my excuse is that I > > am still new to the team and still learning). > > > > 1. QEMU has a bitmap of shared pages in RAMBlockAttributes, so it can > > distinguish shared pages. > > 2. In general, QEMU does not distinguish private vs unaccepted pages, so > > unaccepted pages are treated as private pages. > > Yes, the latter seems uncontroversial. > > The 1st one is true, and it just reminded me if the conversion is > synchronous and one step requires the hypercall to QEMU, then indeed > background conversion can be avoided by some form of userspace locking. > > Perhaps, a rwlock suites, each vCPU takes it for write whenever page > conversion requested from the guest (private <-> shared; nothing about > "accepted" that matters). Then the migration threads, one or multiple, > take the read lock, lookup the bit, do MEM.EXPORT, unlock. > > Then it seems fine in general, except that I donno if things can still go > wrong when there are multiple versions of "if this page is private or > shared". Say, minimum of three? > > (a) QEMU maintains the bitmap in RAMBlockAttributes, each bit represents > if the page is shared or private > > (b) KVM should maintain one, looks to me, kvm->mem_attr_array > > (c) Hardware / Firmware may maintain its own, in case of TDX, is that one > bit on the SEPT pgtable? > > They don't change together, AFAIU, they change in order, I believe > (c)->(a)->(b) if my above understanding is correct. It looks like in terms of which layer saves the page type change first: - Private->Shared: TDX -> KVM -> QEMU - Shared->Private: KVM -> QEMU -> TDX In both cases the TD initiates the change. This ends up with a TD exit, followed by KVM exiting to QEMU. QEMU calls kvm_convert_memory(), which calls back into KVM (KVM_SET_MEMORY_ATTRIBUTES). At this point SEPT did not change yet. Private->Shared: - KVM first removes the page from SEPT, so TDX sees the change first. - KVM updates own data (kvm->mem_attr_array). So KVM "gets" the change second. - QEMU updates RAMBlockAttributes, so QEMU "gets" the change last. Shared->Private: - KVM removes the page from the shared EPT, updates own data (kvm->mem_attr_array). - QEMU updates RAMBlockAttributes, discards backing storage. - Back to TD, which accepts the page. This causes an EPT violation, and KVM adds the page to SEPT (TDH.MEM.PAGE.AUG). > Then, what if they report different things? > > Say, during migration the guest wants to convert a page from shared to > private. (c) can be already done saying one page "private" now for TDX, > (a) tries to mark it "private" too, but now assuming page being accessed > (read lock held), it may be trying to take a write lock and sleep, which > means (b) will be "shared" so far. > > So what happens is, QEMU thinks this page "shared" because the conversion > hasn't take place waiting for the write lock, however at least TDX may > think it already "private" instead. > > Then QEMU logically can access HVA of that page, with (a)=shared, > (b)=shared, (c)=private. > > Would it cause trouble? (c) should see it as private only after KVM and QEMU do. But I am not sure about the entire idea. Holding the read lock around the export ioctl means that a vCPU requesting a conversion waits until the export finishes. One export call may cover many MiB of crypto work, so the vCPU stalling may be significant, right? Let's check the 2 cases. Private -> Shared QEMU calls the export ioctl for a page that it thinks is private, but meanwhile it became shared. In this case, if the semantics of the export ioctl is that such pages are skipped, we should be fine, right? QEMU will just handle this page during the next round. Shared -> Private QEMU tries to migrate a shared page, which meanwhile became private. Reading it would result in zeros or some stale data, right? Would it help if QEMU used a lock-check_if_still_shared-copy-release, and the same lock around kvm_convert_memory()? Artem. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-10-02 19:57 ` Artem Bityutskiy @ 2026-10-07 20:00 ` Peter Xu 0 siblings, 0 replies; 84+ messages in thread From: Peter Xu @ 2026-10-07 20:00 UTC (permalink / raw) To: Artem Bityutskiy Cc: Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Fri, Oct 02, 2026 at 10:57:46PM +0300, Artem Bityutskiy wrote: > On Tue, 2026-09-29 at 17:05 -0400, Peter Xu wrote: > > > Here is how I saw this, but I may be missing something (my excuse is that I > > > am still new to the team and still learning). > > > > > > 1. QEMU has a bitmap of shared pages in RAMBlockAttributes, so it can > > > distinguish shared pages. > > > 2. In general, QEMU does not distinguish private vs unaccepted pages, so > > > unaccepted pages are treated as private pages. > > > > Yes, the latter seems uncontroversial. > > > > The 1st one is true, and it just reminded me if the conversion is > > synchronous and one step requires the hypercall to QEMU, then indeed > > background conversion can be avoided by some form of userspace locking. > > > > Perhaps, a rwlock suites, each vCPU takes it for write whenever page > > conversion requested from the guest (private <-> shared; nothing about > > "accepted" that matters). Then the migration threads, one or multiple, > > take the read lock, lookup the bit, do MEM.EXPORT, unlock. > > > > Then it seems fine in general, except that I donno if things can still go > > wrong when there are multiple versions of "if this page is private or > > shared". Say, minimum of three? > > > > (a) QEMU maintains the bitmap in RAMBlockAttributes, each bit represents > > if the page is shared or private > > > > (b) KVM should maintain one, looks to me, kvm->mem_attr_array > > > > (c) Hardware / Firmware may maintain its own, in case of TDX, is that one > > bit on the SEPT pgtable? > > > > They don't change together, AFAIU, they change in order, I believe > > (c)->(a)->(b) if my above understanding is correct. > > It looks like in terms of which layer saves the page type change first: > > - Private->Shared: TDX -> KVM -> QEMU > - Shared->Private: KVM -> QEMU -> TDX > > In both cases the TD initiates the change. This ends up with a TD exit, > followed by KVM exiting to QEMU. QEMU calls kvm_convert_memory(), which > calls back into KVM (KVM_SET_MEMORY_ATTRIBUTES). At this point SEPT did not > change yet. > > Private->Shared: > - KVM first removes the page from SEPT, so TDX sees the change first. > - KVM updates own data (kvm->mem_attr_array). So KVM "gets" the change > second. > - QEMU updates RAMBlockAttributes, so QEMU "gets" the change last. > > Shared->Private: > - KVM removes the page from the shared EPT, updates own data > (kvm->mem_attr_array). > - QEMU updates RAMBlockAttributes, discards backing storage. > - Back to TD, which accepts the page. This causes an EPT violation, and > KVM adds the page to SEPT (TDH.MEM.PAGE.AUG). Oh, this reminded me that ram_block_attributes_state_change() is done after the ioctl(KVM_SET_MEMORY_ATTRIBUTES); I believe I didn't notice this detail and assumed the other way round. In that case, yes, KVM's page status will always change before QEMU's. > > > Then, what if they report different things? > > > > Say, during migration the guest wants to convert a page from shared to > > private. (c) can be already done saying one page "private" now for TDX, > > (a) tries to mark it "private" too, but now assuming page being accessed > > (read lock held), it may be trying to take a write lock and sleep, which > > means (b) will be "shared" so far. > > > > So what happens is, QEMU thinks this page "shared" because the conversion > > hasn't take place waiting for the write lock, however at least TDX may > > think it already "private" instead. > > > > Then QEMU logically can access HVA of that page, with (a)=shared, > > (b)=shared, (c)=private. > > > > Would it cause trouble? > > (c) should see it as private only after KVM and QEMU do. > > But I am not sure about the entire idea. Holding the read lock around the > export ioctl means that a vCPU requesting a conversion waits until the > export finishes. One export call may cover many MiB of crypto work, so the > vCPU stalling may be significant, right? > > Let's check the 2 cases. > > Private -> Shared > > QEMU calls the export ioctl for a page that it thinks is private, but > meanwhile it became shared. In this case, if the semantics of the export > ioctl is that such pages are skipped, we should be fine, right? QEMU will > just handle this page during the next round. Yes, this should be benign as long as conversion of page status set the dirty bit; QEMU guarantees to clear the D-bit before reading the page, then it must read it again later. > > Shared -> Private > > QEMU tries to migrate a shared page, which meanwhile became private. Reading > it would result in zeros or some stale data, right? Would it help if QEMU > used a lock-check_if_still_shared-copy-release, and the same lock around > kvm_convert_memory()? AFAIU, this should SIGBUS QEMU, which is the major issue, that should be what happens in current linux tree, where now only INIT_SHARED is available for fault() processing. I also think that's the plan for afterwards, but still worth check kernel tree with TDX migration integrated: kvm_gmem_fault_user_mapping() should decide what happens.. -- Peter Xu ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-09-22 8:09 ` Artem Bityutskiy 2026-09-22 9:42 ` Tony Lindgren 2026-09-22 21:18 ` Peter Xu @ 2026-09-23 15:28 ` Serge Hallyn (AMD) 2 siblings, 0 replies; 84+ messages in thread From: Serge Hallyn (AMD) @ 2026-09-23 15:28 UTC (permalink / raw) To: Artem Bityutskiy Cc: Peter Xu, Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Tue, Sep 22, 2026 at 11:09:42AM +0300, Artem Bityutskiy wrote: > Hi Peter, ... > > Said that, we'll need to be careful then in case of migration fallbacks at > > the final stage. Nowadays, I believe QEMU can still fallback to source side > > at a very, very late stage after all things applied. If I'm not mistaken, > > the final handshake is done at migration_incoming_state_destroy() -> > > migrate_send_rp_shut() telling source to be gone. > > Right. In traditional pre-copy VM migration model, the fallback is possible > at any point before the destination VM starts running and modifying its > state. > > The same is true for TDX, but with more complexity and limitations. > > We have 2 points: > 1. Before the source has exported the start token - fallback is similar to > traditional VM migration - just abort the migration on source and > continue running the source TD, and just destroy the destination TD. > 2. After the source has exported the start token - fallback is still > possible, but more complex - it requires the destination to first > generate the abort token, which should be delivered to the source and > consumed there. There are seamcalls for both generating and consuming the > abort token. The token basically makes sure the destination cannot run, > and the source can run again - same "only one of the TDs can run at a > time" security rule that we discussed earlier. > > Now, the limitation here is that if the abort is because the network > connectivity is lost, the abort token cannot be delivered. Then I'd guess > migration should stop, without destroying the source and the destination, > and the token should be delivered later manually, QEMU could even provide an > infrastructure / commands for that. > > Would be very helpful to learn the abort protocol the other CoCo vendors > offer - is it similar to TDX or not? Yup, that sounds very similar to the expected SEV-SNP late abort semantics, -serge ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-09-17 21:27 ` Peter Xu 2026-09-18 12:46 ` Artem Bityutskiy @ 2026-09-20 23:56 ` Kishen Maloor 2026-09-23 21:36 ` Peter Xu 1 sibling, 1 reply; 84+ messages in thread From: Kishen Maloor @ 2026-09-20 23:56 UTC (permalink / raw) To: Peter Xu, Artem Bityutskiy Cc: Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm Hi Peter, Thank you for your comments. Just adding a few other details to complement Artem's response. On 9/17/26 2:27 PM, Peter Xu wrote: > On Fri, Sep 04, 2026 at 09:24:25PM +0300, Artem Bityutskiy wrote: >> On Mon, 2026-08-31 at 10:13 +0300, Tony Lindgren wrote: > ... >> - Should the same uAPIs also support traditional VMs? But the only use-case >> I imagine here is "for testing purposes". > > This is an interesting idea, I think this could be useful. Especially, I > wonder if you already have it done and PoC branches you can share, so that > I can play with it. I had once attempted this and hit a breaking case while migrating regular VMs. I didn't dig further since the primary use of this UAPI set is with CoCo VMs. Here's what I gathered: KVM-mediated transfer needs KVM to resolve a GFN to the HVA where the data lands, and on x86 QEMU mutates that mapping at runtime in a way that is opaque to KVM. For the PAM window QEMU overlays a separate region on top of pc.ram, and KVM sees only the flattened view, the winning memslot, with no way to address what that slot shadows. On the source, firmware reprograms PAM and QEMU drops the overlay. On the destination the firmware never runs, so it still has its boot-time layout and the overlay is still in place. So importing, say, GFN 0xc0 lands on the overlay's backing store, which is a read-only memslot, and the GFN->HVA translation fails. I suppose even if it were writable, the data would land there rather than in the pc.ram underneath where it belongs. Without mediation, QEMU is able to write to the HVA directly, so the problem doesn't arise. Generally, I think KVM mediation only works if KVM's view of memory is authoritative. TDX skirts this issue entirely because QEMU doesn't create overlays for TDX VMs. > ... > Could you elaborate this ITERATION operation? Is that something the > userapp must do after full scan of a round of guest memory? Essentially, yes. TDX migration architecture delimits such pre-copy rounds as "migration epochs" and emits an epoch token at each round boundary that needs to be consumed on the destination. It allows the TDX module to verify that everything from the prior round has been received at the destination and that two versions of the same GPA aren't sent in the same epoch. It is userspace that decides where a round ends, but any deviation from this model and migration would fail. > >>> | (repeat until convergence) | >>> CMD(STOP_AND_COPY/PAUSE) | >>> CMD(STOP_AND_COPY/TD_STATE) --- VM state ------> CMD(STOP_AND_COPY/TD_STATE) >>> KVM_EXPORT_VCPU --- vCPU state ----> KVM_IMPORT_VCPU >>> KVM_EXPORT_MEMORY -- final memory ---> KVM_IMPORT_MEMORY > > When read/write encrypted memories, two questions: > > - Is there an upper bound of the buffer size per-page? Yes. For TDX a 4KB guest page produces exactly 4KB of encrypted payload plus a small fixed amount of ancillary data. The bound can be pre-computed and userspace can size its buffers accordingly. The UAPI itself doesn't impose one. The vendor implementation decides how pages and ancillary data are laid out. > > - Does this operation supports concurrency? If it supports, how well it > scales per expectation (e.g. is there known big lock for that)? Yes, these operations can support multifd and kvm_memory_transfer carries the channel ID. The encryption/decryption work is per-channel, so it scales as multifd does. The one serialization point is the epoch boundary, i.e. the ITERATION call, which has to drain in-flight transfers for that round. That's inherent to the epoch model rather than a lock that could be dropped. > ... > IMHO we should really take postcopy into account when designing the API and > state machine. We don't need to implement it in the first version, even > until merging, but we need to make sure postcopy will be new ioctls on top > of existing and it should have no major loopholes that it'll need a new set > of APIs. Agreed. We honestly haven't looked at post-copy in depth, but supporting it thus far appears to be additive. One item that comes to mind is a new command in KVM_MIGRATE_CMD to mark the post-copy switchover point. On the source, vendor code could send all previously exported and dirty pages (deferring all un-exported pages to post-copy) and prepare to serve pages on demand. On the destination, vendor code could make the VM runnable with incomplete memory. This is based on a preliminary assessment though. > For example, I think we should consider KVM_EXPORT_MEMORY being usable > after END on source, KVM_IMPORT_MEMORY while TD is in operation, etc. We Yes, though I think some of these details could be handled in the vendor implementation -- e.g., KVM_EXPORT/IMPORT_MEMORY might not need to know that they're servicing post-copy. > should likely also need to still picture the rough process of postcopy, > reserve those APIs since the start (but return -EINVAL or something). Agreed. We shall attempt to sketch that flow. Your thoughts and feedback would be very helpful as we work through this. > ... > Could you elaborate what's the relations between TDH.MEM.SCAN.RANGE and the > GET_DIRTY_LOG ioctl? I recall above mentioned GET_DIRTY_LOG will be > available even for CoCo, which makes sense assuming dirty information isn't > confidential. However then I don't understand what TDH.MEM.SCAN.RANGE > plays the role here. GET_DIRTY_LOG is the UAPI, but the slot dirty bitmaps need to be populated somehow and TDH.MEM.SCAN.RANGE fills that gap for TDX. It is a SEAMCALL that scans the SEPT upon request to return the list of currently dirty pages. We wanted to avoid inventing a new UAPI for this and GET_DIRTY_LOG seemed to be a natural fit. In our current PoC, we've added a barebones kvm_x86_ops hook that is plugged into kvm_arch_sync_dirty_log(). The TDX implementation of that hook invokes the scans and records the dirty pages into KVM's slot dirty bitmaps that GET_DIRTY_LOG serves up to userspace (as usual). ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-09-20 23:56 ` Kishen Maloor @ 2026-09-23 21:36 ` Peter Xu 2026-09-24 4:27 ` Kishen Maloor 0 siblings, 1 reply; 84+ messages in thread From: Peter Xu @ 2026-09-23 21:36 UTC (permalink / raw) To: Kishen Maloor Cc: Artem Bityutskiy, Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Sun, Sep 20, 2026 at 04:56:20PM -0700, Kishen Maloor wrote: > Hi Peter, Hi, Kishen, > > Thank you for your comments. Just adding a few other details to > complement Artem's response. > > On 9/17/26 2:27 PM, Peter Xu wrote: > > On Fri, Sep 04, 2026 at 09:24:25PM +0300, Artem Bityutskiy wrote: > >> On Mon, 2026-08-31 at 10:13 +0300, Tony Lindgren wrote: > > ... > >> - Should the same uAPIs also support traditional VMs? But the only use-case > >> I imagine here is "for testing purposes". > > > > This is an interesting idea, I think this could be useful. Especially, I > > wonder if you already have it done and PoC branches you can share, so that > > I can play with it. > > I had once attempted this and hit a breaking case while migrating regular VMs. > I didn't dig further since the primary use of this UAPI set is with CoCo VMs. > > Here's what I gathered: KVM-mediated transfer needs KVM to resolve a GFN to the > HVA where the data lands, and on x86 QEMU mutates that mapping at runtime in a > way that is opaque to KVM. For the PAM window QEMU overlays a separate region on > top of pc.ram, and KVM sees only the flattened view, the winning memslot, with > no way to address what that slot shadows. On the source, firmware reprograms PAM > and QEMU drops the overlay. On the destination the firmware never runs, so it still > has its boot-time layout and the overlay is still in place. So importing, say, GFN 0xc0 > lands on the overlay's backing store, which is a read-only memslot, and the GFN->HVA > translation fails. I suppose even if it were writable, the data would land there > rather than in the pc.ram underneath where it belongs. Without mediation, QEMU is > able to write to the HVA directly, so the problem doesn't arise. > > Generally, I think KVM mediation only works if KVM's view of memory is authoritative. > TDX skirts this issue entirely because QEMU doesn't create overlays for TDX VMs. I see this one slightly differently, and I have a major question on the choice of GPA for KVM's memory access API. If reusing GET_DIRTY_LOG, it means slot_id works for CoCo like before because that's the old interface there. But the new KVM_EXPORT_MEM (vice versa) used GPA arrays. Could I ask why the change? IIUC this is the fundamental reason why you hit that PAM issue: at a specific GPA, QEMU/KVM can map different things, hence GPA is not yet an unified identifier for a physical page that guest uses. However, (slot_id, slot_offset) will be. If the migration memory access API will be using that instead of GPA, I think problems will be gone because then the PAM regions (likely, pc.ram underneath) will be migrated in form of KVM memslots, destination apply that to the PAM's (slot_id, slot_offset), then when switchover, that PAM memory region will be enabled by QEMU and it will be visible to destination QEMU when VM starts on destination. I'm actually not sure if such (e.g. PAM) will ever be supported at all in CoCo context, but it seems to me using (slot_id, slot_offset) (for each entry; it can also be an array) is more flexible than GPA arrays in that it allows GPA mapping to change on the fly if needed even during migration. I'm not sure if that's explicitly forbidden for CoCo, though. > > > ... > > Could you elaborate this ITERATION operation? Is that something the > > userapp must do after full scan of a round of guest memory? > > Essentially, yes. TDX migration architecture delimits such pre-copy rounds as > "migration epochs" and emits an epoch token at each round boundary that needs to be > consumed on the destination. It allows the TDX module to verify that everything from > the prior round has been received at the destination and that two versions of the > same GPA aren't sent in the same epoch. It is userspace that decides where a round > ends, but any deviation from this model and migration would fail. I'm a bit surprised that TDX will also monitor how many times the same GPA is updated per iteration. What if below happens: ... ITERATION sync n GET_DIRTY_LOG, see page P dirty migrate page P GET_DIRTY_LOG, see page P dirty again migrate page P again <--------------------- [a] ITERATION sync n+1 ... Would above crash on destination TDX at step [a] applying P 2nd time? Or does it mean GET_DIRTY_LOG can only be invoked once per iteration? > > > > >>> | (repeat until convergence) | > >>> CMD(STOP_AND_COPY/PAUSE) | > >>> CMD(STOP_AND_COPY/TD_STATE) --- VM state ------> CMD(STOP_AND_COPY/TD_STATE) > >>> KVM_EXPORT_VCPU --- vCPU state ----> KVM_IMPORT_VCPU > >>> KVM_EXPORT_MEMORY -- final memory ---> KVM_IMPORT_MEMORY > > > > When read/write encrypted memories, two questions: > > > > - Is there an upper bound of the buffer size per-page? > > Yes. For TDX a 4KB guest page produces exactly 4KB of encrypted payload plus a small > fixed amount of ancillary data. The bound can be pre-computed and userspace > can size its buffers accordingly. The UAPI itself doesn't impose one. The vendor > implementation decides how pages and ancillary data are laid out. Get it now. I still think if that varies across impl, it might be good to somehow notify the userapp on choosing the buffer size. But nah, it's not a big deal - when there's retry, userapp can always probe with any size then double it until it fits.. Thanks, -- Peter Xu ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-09-23 21:36 ` Peter Xu @ 2026-09-24 4:27 ` Kishen Maloor 2026-09-25 14:18 ` Peter Xu 0 siblings, 1 reply; 84+ messages in thread From: Kishen Maloor @ 2026-09-24 4:27 UTC (permalink / raw) To: Peter Xu Cc: Artem Bityutskiy, Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm Hi Peter, Thanks! It really helps to go over details. On 9/23/26 2:36 PM, Peter Xu wrote: > ... > > I see this one slightly differently, and I have a major question on the > choice of GPA for KVM's memory access API. > > If reusing GET_DIRTY_LOG, it means slot_id works for CoCo like before > because that's the old interface there. It works because we map backwards from GPAs in TDX's dirty scan results to a KVM memslot on the source, and a bit in its dirty bitmap, in keeping with the GET_DIRTY_LOG contract. > But the new KVM_EXPORT_MEM (vice versa) used GPA arrays. Could I ask why > the change? The migration ABI (at least in TDX, possibly others) is GPA shaped. Also, everything downstream like SEPT entries key off GPAs. The EXPORT and IMPORT ABIs take GPAs alongside the encrypted blob, so both the source and the destination need it at the call site that implements the UAPI. Passing GFNs in the UAPI is merely a convenience in that respect. If we passed in slot+offset instead, it would just be an indirection that needs to resolve back to the GPA anyway to issue the vendor migration call. But this is not the blocking issue for the PAM case, as I'll explain below. > IIUC this is the fundamental reason why you hit that PAM issue: at a > specific GPA, QEMU/KVM can map different things, hence GPA is not yet an > unified identifier for a physical page that guest uses. However, (slot_id, > slot_offset) will be. > > If the migration memory access API will be using that instead of GPA, I > think problems will be gone because then the PAM regions (likely, pc.ram > underneath) will be migrated in form of KVM memslots, destination apply > that to the PAM's (slot_id, slot_offset), then when switchover, that PAM > memory region will be enabled by QEMU and it will be visible to destination > QEMU when VM starts on destination. I think using slot+offset won't make the problem vanish, because the PAM window in the underlying pc.ram region is not addressable by KVM on the destination at that time. KVM's memslots come from QEMU's FlatView at any moment and QEMU reconfigures them on the fly. In the destination's boot-time setup, it seems like pc.ram's memslot is split around that window, and what covers the range instead is a separate, read-only memslot for the overlay, which is an allocation with its own backing, not the pc.ram bytes underneath. And since only the flattened view is ever registered, there's no slot+offset that refers to those bytes either. Only when the source's PAM configuration lands at the destination, which happens after memory migration concludes, do those pc.ram bytes become visible to KVM. In the TDX case, QEMU does not create that overlay at the start, so the destination has a memslot corresponding to the pc.ram region that stays put. KVM is able to do a GFN->PFN translation to hand to TDX's IMPORT ABI. So, it appears that a UAPI change will not address this corner case. > I'm actually not sure if such (e.g. PAM) will ever be supported at all in > CoCo context, but it seems to me using (slot_id, slot_offset) (for each > entry; it can also be an array) is more flexible than GPA arrays in that it > allows GPA mapping to change on the fly if needed even during migration. > > I'm not sure if that's explicitly forbidden for CoCo, though. A TD private page lives at a fixed GPA in the SEPT, and I believe the EXPORT/IMPORT bundle binds the GPA into the page's integrity check, so a page has to be imported at the GPA it was exported from. >>> Could you elaborate this ITERATION operation? Is that something the >>> userapp must do after full scan of a round of guest memory? >> >> Essentially, yes. TDX migration architecture delimits such pre-copy rounds as >> "migration epochs" and emits an epoch token at each round boundary that needs to be >> consumed on the destination. It allows the TDX module to verify that everything from >> the prior round has been received at the destination and that two versions of the >> same GPA aren't sent in the same epoch. It is userspace that decides where a round >> ends, but any deviation from this model and migration would fail. > > I'm a bit surprised that TDX will also monitor how many times the same GPA > is updated per iteration. What if below happens: Yes, and maybe that is because it won't know the ordering if the same GPA were caught at different times in the same epoch and fed to different migration streams (which one is newer?) On regular VMs, I believe QEMU doesn't migrate the same GPA twice in one round. As I understand it, a migrated GPA is revisited only in the next round. I suppose TDX just makes that behavior architectural. > ... > ITERATION sync n > GET_DIRTY_LOG, see page P dirty > migrate page P > GET_DIRTY_LOG, see page P dirty again > migrate page P again <--------------------- [a] QEMU shouldn't transfer P again in the same round, right? > ITERATION sync n+1 > ... > > Would above crash on destination TDX at step [a] applying P 2nd time? Or If P were migrated again, then I believe it would fail to import on the destination, and also abort the import session, so migration would fail. TDX records the current epoch against a GPA on every import, and I think that's how it catches a 2nd import attempt. > does it mean GET_DIRTY_LOG can only be invoked once per iteration? We don't touch QEMU's stock dirty logging/transfer flows so however it handles it currently stays. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-09-24 4:27 ` Kishen Maloor @ 2026-09-25 14:18 ` Peter Xu 2026-09-29 1:28 ` Kishen Maloor 0 siblings, 1 reply; 84+ messages in thread From: Peter Xu @ 2026-09-25 14:18 UTC (permalink / raw) To: Kishen Maloor Cc: Artem Bityutskiy, Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Wed, Sep 23, 2026 at 09:27:52PM -0700, Kishen Maloor wrote: > Hi Peter, Hi, Kishen, [...] > > IIUC this is the fundamental reason why you hit that PAM issue: at a > > specific GPA, QEMU/KVM can map different things, hence GPA is not yet an > > unified identifier for a physical page that guest uses. However, (slot_id, > > slot_offset) will be. > > > > If the migration memory access API will be using that instead of GPA, I > > think problems will be gone because then the PAM regions (likely, pc.ram > > underneath) will be migrated in form of KVM memslots, destination apply > > that to the PAM's (slot_id, slot_offset), then when switchover, that PAM > > memory region will be enabled by QEMU and it will be visible to destination > > QEMU when VM starts on destination. > > I think using slot+offset won't make the problem vanish, because the PAM window > in the underlying pc.ram region is not addressable by KVM on the destination at > that time. KVM's memslots come from QEMU's FlatView at any moment and QEMU > reconfigures them on the fly. In the destination's boot-time setup, it seems > like pc.ram's memslot is split around that window, and what covers the > range instead is a separate, read-only memslot for the overlay, which is an > allocation with its own backing, not the pc.ram bytes underneath. > And since only the flattened view is ever registered, there's no slot+offset > that refers to those bytes either. Only when the source's PAM configuration > lands at the destination, which happens after memory migration concludes, do > those pc.ram bytes become visible to KVM. > > In the TDX case, QEMU does not create that overlay at the start, so the > destination has a memslot corresponding to the pc.ram region that stays put. > KVM is able to do a GFN->PFN translation to hand to TDX's IMPORT ABI. > > So, it appears that a UAPI change will not address this corner case. Yes, I think you're right, the ABI isn't the major issue. If from KVM's perspective it's always convertable between GPA <-> slots, then it's the same. I believe my mindset when replying was pretty much in QEMU's perspective, where in qemu we can have two ramblocks plugged into the same GPA range, only one of them will be visible to KVM and guest (e.g. which one has higher MemoryRegion priority, but it's not the only factor). What migration module does right now is, it allows both ramblocks to be migrated with no issue, even if one is not visible, but if it used to be touched, since that ramblock will maintain its own dirty bitmap (GET_DIRTY_LOG on that kvm memslot when it was visible to KVM, or maybe set within QEMU userspace somehow). IOW, what QEMU could do here is, after MEMORY.EXPORT, convert the GPA address space into ramblock ranges in QEMU, migrate with that not GPA. In case of PAM, it's part of pc.ram. Dest QEMU sees it is pc.ram, it should "apply" those data to pc.ram ramblock only. For non-CoCo, it's as easy as writting to some host HVA pointer. Now, the question is, the new migration API only allows applying data in GPA ranges. I don't think with TDX there's a way to "apply" the data.. dest QEMU just booted, this specific portion of pc.ram may not be mapped at that GPA source fetched due to reset status of PAM registers. That (rather than the ABI interface), might be the real thing I wanted to point out. Maybe it means TDX just can't work with it by definition? I think it'll be fine, and now I wonder if it means PAM will be working for TDX only if PAM boots too early so it was before TDX initializes, then after TDX enabled anything like PAM will not work anymore? So it seems TDX will "lock" the memory footprint in place when enabled, allow accept/unaccept (or say, plug / unplug) memories, but anything like "flipping this to that" will not work. Then there's a very corner case question I want to double check, and I apologize if this is stupid only due to my ignorance on TDX knowledge: can someone migrate a VM too early so TDX is just hasn't been enabled at all? [...] > > I'm a bit surprised that TDX will also monitor how many times the same GPA > > is updated per iteration. What if below happens: > > Yes, and maybe that is because it won't know the ordering if the same > GPA were caught at different times in the same epoch and fed to different > migration streams (which one is newer?) > On regular VMs, I believe QEMU doesn't migrate the same GPA twice in one round. > As I understand it, a migrated GPA is revisited only in the next round. > I suppose TDX just makes that behavior architectural. > > > ... > > ITERATION sync n > > GET_DIRTY_LOG, see page P dirty > > migrate page P > > GET_DIRTY_LOG, see page P dirty again > > migrate page P again <--------------------- [a] > > QEMU shouldn't transfer P again in the same round, right? I believe yes with current QEMU, I can't think of anything otherwise. But still, this is very specific impl detail. There's definitely no issue migrating one page twice or more in non-CoCo. I can give one example to illustrate what could happen. In postcopy, we support preemption mode, which is simply a separate fast path for requested / urgent pages. It's possible while background thread transferring one page, the fast path saw a request on this same page. The current algorithm is simple, it will wait for that in progress background send to complete. But logically, we could do it the other way too: send the page again on fast path, in postcopy the page content is guaranteed to be identical and unchnaged, it means the fast path can land this page earlier, reducing fault latency. If so, a minimum cap we need is MEMORY.EXPORT be able to be done twice, so the fast path can read the 2nd time. IMPORT is more flexible, because QEMU can maintain what has been applied, so logically background loader should be able to skip the 2nd IMPORT. That is not a good example, at least because it's postcopy and doesn't happen with precopy. So far, I also don't think a major risk, but I confess I don't understand why TDX needs to add hard requirement on "only sample one page once per iteration": even if VMM sampled a page twice, it's still encrypted and confidential. I believe it has something to do with the whole attestation logic. I think it'll be more flexible with less restrictions, but no issue I see either, hence please only treat that a verbose FYI. > > > ITERATION sync n+1 > > ... > > > > Would above crash on destination TDX at step [a] applying P 2nd time? Or > > If P were migrated again, then I believe it would fail to import on the > destination, and also abort the import session, so migration would fail. I wonder if we can just fail the 2nd IMPORT without abort the whole process, then if userapp wants to detect it there's a way to (similar to an -EEXIST). Not a request, more like a pure question. > TDX records the current epoch against a GPA on every import, and I think > that's how it catches a 2nd import attempt. > > > does it mean GET_DIRTY_LOG can only be invoked once per iteration? > > We don't touch QEMU's stock dirty logging/transfer flows so however it > handles it currently stays. I see, then it should be good. Just to say, qemu may sync dirty bitmaps in the background nowadays, can refer to cpu_throttle_dirty_sync_timer_tick(). So it can be decoupled from iteration runs. Thanks, -- Peter Xu ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-09-25 14:18 ` Peter Xu @ 2026-09-29 1:28 ` Kishen Maloor 2026-09-30 20:42 ` Peter Xu 0 siblings, 1 reply; 84+ messages in thread From: Kishen Maloor @ 2026-09-29 1:28 UTC (permalink / raw) To: Peter Xu Cc: Artem Bityutskiy, Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm Hi Peter, Great comments! On 9/25/26 7:18 AM, Peter Xu wrote: > On Wed, Sep 23, 2026 at 09:27:52PM -0700, Kishen Maloor wrote: > ... > Yes, I think you're right, the ABI isn't the major issue. If from KVM's > perspective it's always convertable between GPA <-> slots, then it's the > same. Correct. It's just that when the bytes aren't in any memslot there's no way to reach pc.ram's HVA on the destination in KVM's view to write to it. > I believe my mindset when replying was pretty much in QEMU's perspective, > where in qemu we can have two ramblocks plugged into the same GPA range, > only one of them will be visible to KVM and guest (e.g. which one has > higher MemoryRegion priority, but it's not the only factor). What migration I understand, and it matches what I found. > module does right now is, it allows both ramblocks to be migrated with no > issue, even if one is not visible, but if it used to be touched, since that > ramblock will maintain its own dirty bitmap (GET_DIRTY_LOG on that kvm > memslot when it was visible to KVM, or maybe set within QEMU userspace > somehow). Correct, and regular VM migration bypasses any notion of GPAs/memslots entirely. It migrates RAMBlocks by block+offset. QEMU sets every RAMBlock's migration bitmap to all ones before the first pass, so everything is transferred in round 1 regardless of dirty state. > IOW, what QEMU could do here is, after MEMORY.EXPORT, convert the GPA > address space into ramblock ranges in QEMU, migrate with that not GPA. In > case of PAM, it's part of pc.ram. Dest QEMU sees it is pc.ram, it should > "apply" those data to pc.ram ramblock only. > > For non-CoCo, it's as easy as writting to some host HVA pointer. I understand that this is how it works for regular VMs as the destination is able to memcpy into the HVA obtained using the block+offset it receives over the stream. Also, when migration is kicked off after the source has fully booted, the PAM window already maps to the source's pc.ram area, so the export transfers the right bytes. The gap is purely on the destination because its layout is still in the boot-time configuration until device state lands. > Now, the question is, the new migration API only allows applying data in > GPA ranges. I don't think with TDX there's a way to "apply" the data.. > dest QEMU just booted, this specific portion of pc.ram may not be mapped at > that GPA source fetched due to reset status of PAM registers. Correct. In TDX, the page is placed by the import SEAMCALL at the GPA it was exported from. So QEMU has to turn the block+offset back into a GPA, and that GPA has to be covered by a memslot for KVM to resolve it, which it is unable to do at that point. > That (rather than the ABI interface), might be the real thing I wanted to > point out. Understood. I did also wonder whether other VMMs organize their guest memory hierarchy in similar ways. > Maybe it means TDX just can't work with it by definition? I think it'll be > fine, and now I wonder if it means PAM will be working for TDX only if PAM > boots too early so it was before TDX initializes, then after TDX enabled > anything like PAM will not work anymore? Upstream QEMU gates the separate pc.rom allocation on !is_tdx_vm(), so there's no second allocation to flip to. I don't see any other TDX gating around this, so I presume that PAM operates as usual otherwise. It doesn't appear to be a question of PAM coming up before TDX initializes. > So it seems TDX will "lock" the memory footprint in place when enabled, > allow accept/unaccept (or say, plug / unplug) memories, but anything like > "flipping this to that" will not work. Yes, I think so. The TDX module holds the GPA->page binding, so any change has to pass through it. Adding and removing pages plumb through the module. But a sort of content preserving remap of GPAs in the way QEMU does for PAM is not possible AFAIU. > Then there's a very corner case question I want to double check, and I > apologize if this is stupid only due to my ignorance on TDX knowledge: can > someone migrate a VM too early so TDX is just hasn't been enabled at all? Not a stupid question at all. I've actually tried this. TDX VMs appear to migrate fine beyond a certain point mid-boot of the source. Kicking off migration any earlier causes the destination not to resume, even though the migration itself reports success. I haven't yet confirmed what that point is to explain it. For now we're trying to settle on sound fundamentals for the UAPI and flow, but this is definitely an area to dig further into. >>> I'm a bit surprised that TDX will also monitor how many times the same GPA >>> is updated per iteration. What if below happens: >> >> Yes, and maybe that is because it won't know the ordering if the same >> GPA were caught at different times in the same epoch and fed to different >> migration streams (which one is newer?) >> On regular VMs, I believe QEMU doesn't migrate the same GPA twice in one round. >> As I understand it, a migrated GPA is revisited only in the next round. >> I suppose TDX just makes that behavior architectural. >> >>> ... >>> ITERATION sync n >>> GET_DIRTY_LOG, see page P dirty >>> migrate page P >>> GET_DIRTY_LOG, see page P dirty again >>> migrate page P again <--------------------- [a] >> >> QEMU shouldn't transfer P again in the same round, right? > > I believe yes with current QEMU, I can't think of anything otherwise. But > still, this is very specific impl detail. There's definitely no issue > migrating one page twice or more in non-CoCo. Thanks for confirming this. But do you think it would be a problem during pre-copy with multifd on regular VM migration if we didn't have this "transfer once per round" logic? If the same page got fed to different multifd queues, couldn't the older copy land after the newer one at the destination? [1] With a single channel it shouldn't matter, but it could with multiple channels. I think since the bitmap sweep is shared with multifd, QEMU ends up transferring a page only once per round either way. At least this was the rationale I conceived in trying to explain the TDX policy of not allowing multiple imports of the same page in one epoch. Though the underlying concern may not be TDX-specific. > I can give one example to illustrate what could happen. > > In postcopy, we support preemption mode, which is simply a separate fast > path for requested / urgent pages. It's possible while background thread > transferring one page, the fast path saw a request on this same page. The > current algorithm is simple, it will wait for that in progress background > send to complete. > > But logically, we could do it the other way too: send the page again on > fast path, in postcopy the page content is guaranteed to be identical and > unchnaged, it means the fast path can land this page earlier, reducing > fault latency. If so, a minimum cap we need is MEMORY.EXPORT be able to be > done twice, so the fast path can read the 2nd time. IMPORT is more We shall check about this. TDH.EXPORT.MEM (the export-side SEAMCALL on the source) may already permit a 2nd export of the same page during post-copy. If it doesn't, there's a good case for asking for it, since it would give both the background and preemption threads a shot at fulfilling a request ASAP. The import side might still reject it, so your suggestion below looks like the right place to handle that. > flexible, because QEMU can maintain what has been applied, so logically > background loader should be able to skip the 2nd IMPORT. This would make sense to do I suppose, because accepting a 2nd IMPORT would clobber a page that the destination previously received and has itself modified. > That is not a good example, at least because it's postcopy and doesn't > happen with precopy. So far, I also don't think a major risk, but I confess > I don't understand why TDX needs to add hard requirement on "only sample > one page once per iteration": even if VMM sampled a page twice, it's still > encrypted and confidential. I believe it has something to do with the > whole attestation logic. I think it'll be more flexible with less > restrictions, but no issue I see either, hence please only treat that a > verbose FYI. I think the transfer once per iteration might have everything to do with reason [1] I mentioned above. I've generally noticed that the TDX migration architecture reflects established practices of VMMs. That being said, we should look for instances where a restriction deviates and poses a problem. > >> >>> ITERATION sync n+1 >>> ... >>> >>> Would above crash on destination TDX at step [a] applying P 2nd time? Or >> >> If P were migrated again, then I believe it would fail to import on the >> destination, and also abort the import session, so migration would fail. > > I wonder if we can just fail the 2nd IMPORT without abort the whole > process, then if userapp wants to detect it there's a way to (similar to an > -EEXIST). Not a request, more like a pure question. Understood. I think it's TDX's way of assuring correctness in case a VMM transfers a page more than once in a round. QEMU doesn't. In theory it could relax that when only one migration stream is registered as the ordering would be unambiguous there, so a second import is provably newer and could just be accepted. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-09-29 1:28 ` Kishen Maloor @ 2026-09-30 20:42 ` Peter Xu 2026-10-07 4:27 ` Kishen Maloor 0 siblings, 1 reply; 84+ messages in thread From: Peter Xu @ 2026-09-30 20:42 UTC (permalink / raw) To: Kishen Maloor Cc: Artem Bityutskiy, Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Mon, Sep 28, 2026 at 06:28:39PM -0700, Kishen Maloor wrote: > Hi Peter, Hi, Kishen, [...] > > So it seems TDX will "lock" the memory footprint in place when enabled, > > allow accept/unaccept (or say, plug / unplug) memories, but anything like > > "flipping this to that" will not work. > > Yes, I think so. The TDX module holds the GPA->page binding, so any change has > to pass through it. Adding and removing pages plumb through the module. > But a sort of content preserving remap of GPAs in the way QEMU does for PAM > is not possible AFAIU. This still looks fine for TDX then, but I think we may want to make sure it works properly for all the rest consumers of this API that we're aware of, say, AMD, ARM CCA, etc. Especially, MEM.IMPORT seems trickier to me: MEM.EXPORT even if described in in GPA address space, can be converted to other address spaces at will, as long as QEMU has full knowledge of the mappings. MEM.IMPORT, OTOH, may not trivially work in GPA terms, if memory mappings can overlap. > > > Then there's a very corner case question I want to double check, and I > > apologize if this is stupid only due to my ignorance on TDX knowledge: can > > someone migrate a VM too early so TDX is just hasn't been enabled at all? > > Not a stupid question at all. I've actually tried this. TDX VMs appear to > migrate fine beyond a certain point mid-boot of the source. Kicking off > migration any earlier causes the destination not to resume, even though the > migration itself reports success. I haven't yet confirmed what that point is > to explain it. > > For now we're trying to settle on sound fundamentals for the UAPI and flow, > but this is definitely an area to dig further into. Ok. Before any fixing lands, perhaps one can add a temprorary migration blocker in QEMU and remove it after that mid-boot phase. > > >>> I'm a bit surprised that TDX will also monitor how many times the same GPA > >>> is updated per iteration. What if below happens: > >> > >> Yes, and maybe that is because it won't know the ordering if the same > >> GPA were caught at different times in the same epoch and fed to different > >> migration streams (which one is newer?) > >> On regular VMs, I believe QEMU doesn't migrate the same GPA twice in one round. > >> As I understand it, a migrated GPA is revisited only in the next round. > >> I suppose TDX just makes that behavior architectural. > >> > >>> ... > >>> ITERATION sync n > >>> GET_DIRTY_LOG, see page P dirty > >>> migrate page P > >>> GET_DIRTY_LOG, see page P dirty again > >>> migrate page P again <--------------------- [a] > >> > >> QEMU shouldn't transfer P again in the same round, right? > > > > I believe yes with current QEMU, I can't think of anything otherwise. But > > still, this is very specific impl detail. There's definitely no issue > > migrating one page twice or more in non-CoCo. > > Thanks for confirming this. But do you think it would be a problem during > pre-copy with multifd on regular VM migration if we didn't have this > "transfer once per round" logic? If the same page got fed to different > multifd queues, couldn't the older copy land after the newer one at the > destination? [1] Multifd has a flush logic per-iteration, so no chance an old version page lands after a new version. Can refer to multifd_ram_round_notify(). QEMU sometimes calls it "round" in the code or comment but it is the ITERATION term we're talking about here. NOTE, that function was merged very recently; you may need to pull latest QEMU, but the logic existed for a long time, so even old multifd behaves that way. > > With a single channel it shouldn't matter, but it could with multiple > channels. I think since the bitmap sweep is shared with multifd, QEMU ends up > transferring a page only once per round either way. Normally, yes, the scanner submits pages to multifd threads, the scanner, if scan one round of bitmap, can only submit one page once until the next round / ITERATION. But it may also be a matter of how that ioctl(ITERATION) will be done inside QEMU; I'll need to reference to the QEMU PoC branch when ready to be sure I have the correct understanding. E.g., IIUC that ioctl(ITERATION) needs to be done similarly "per-round" here to make sure one page appears once, and we should also always keep in mind QEMU can still have GET_DIRTY_LOG run in the background concurrently of the scanner, as I mentioned that too somewhere else (which means one page can be dirtied right after sending it, but as long as the scanner moves on with the next page it looks still OK). Thanks, -- Peter Xu ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-09-30 20:42 ` Peter Xu @ 2026-10-07 4:27 ` Kishen Maloor 2026-10-07 20:13 ` Peter Xu 0 siblings, 1 reply; 84+ messages in thread From: Kishen Maloor @ 2026-10-07 4:27 UTC (permalink / raw) To: Peter Xu Cc: Artem Bityutskiy, Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm Hi Peter, On 9/30/26 1:42 PM, Peter Xu wrote: > On Mon, Sep 28, 2026 at 06:28:39PM -0700, Kishen Maloor wrote: > ... >>> So it seems TDX will "lock" the memory footprint in place when enabled, >>> allow accept/unaccept (or say, plug / unplug) memories, but anything like >>> "flipping this to that" will not work. >> >> Yes, I think so. The TDX module holds the GPA->page binding, so any change has >> to pass through it. Adding and removing pages plumb through the module. >> But a sort of content preserving remap of GPAs in the way QEMU does for PAM >> is not possible AFAIU. > > This still looks fine for TDX then, but I think we may want to make sure it > works properly for all the rest consumers of this API that we're aware of, > say, AMD, ARM CCA, etc. Agreed. It really comes down to the shape of the other migration ABIs out there. > Especially, MEM.IMPORT seems trickier to me: MEM.EXPORT even if described > in in GPA address space, can be converted to other address spaces at will, > as long as QEMU has full knowledge of the mappings. MEM.IMPORT, OTOH, may > not trivially work in GPA terms, if memory mappings can overlap. I can appreciate that. Working in GPA terms assumes that a GPA reference is unambiguous over the duration of migration, which holds true in the TDX case. But it would break if another backend is able to point a mapped GPA at a different protected page while keeping both pages valid -- essentially the PAM case on regular VMs. One option that comes to mind is to add a 2nd optional identifier to kvm_memory_transfer consisting of a guest_memfd fd and an offsets array. That would be at parity with QEMU's block+offset and directly reference specific storage units. Then QEMU and KVM backends for a platform would decide what identifier to use. The value might be that a bundle could be accepted by a backend for storage that isn't currently mapped in a memslot. OTOH, slot+offset as an identifier might face the same problem as GPA in this case when a particular RAMBlock's storage isn't mapped at the time on the destination. Not sure if this idea is viable or even makes sense; this is mostly for discussion, but would appreciate your thoughts. >>> Then there's a very corner case question I want to double check, and I >>> apologize if this is stupid only due to my ignorance on TDX knowledge: can >>> someone migrate a VM too early so TDX is just hasn't been enabled at all? >> >> Not a stupid question at all. I've actually tried this. TDX VMs appear to >> migrate fine beyond a certain point mid-boot of the source. Kicking off >> migration any earlier causes the destination not to resume, even though the >> migration itself reports success. I haven't yet confirmed what that point is >> to explain it. >> >> For now we're trying to settle on sound fundamentals for the UAPI and flow, >> but this is definitely an area to dig further into. > > Ok. Before any fixing lands, perhaps one can add a temprorary migration > blocker in QEMU and remove it after that mid-boot phase. Yeah, possibly. But I haven't seen a TD boot progress hook in QEMU to place the blocker removal. We might just have to chase down the underlying issue for clarity. >> ... >> With a single channel it shouldn't matter, but it could with multiple >> channels. I think since the bitmap sweep is shared with multifd, QEMU ends up >> transferring a page only once per round either way. > > Normally, yes, the scanner submits pages to multifd threads, the scanner, > if scan one round of bitmap, can only submit one page once until the next > round / ITERATION. > > But it may also be a matter of how that ioctl(ITERATION) will be done > inside QEMU; I'll need to reference to the QEMU PoC branch when ready to be > sure I have the correct understanding. E.g., IIUC that ioctl(ITERATION) > needs to be done similarly "per-round" here to make sure one page appears Correct, ioctl(ITERATION) is also done at a round boundary. > once, and we should also always keep in mind QEMU can still have > GET_DIRTY_LOG run in the background concurrently of the scanner, as I > mentioned that too somewhere else (which means one page can be dirtied > right after sending it, but as long as the scanner moves on with the next > page it looks still OK). This should be covered. TDX requires a TLB flush following a scan so subsequent writes would be caught. If the write happens before the page is exported, the export fails for that entry with a PAGE_DIRTY status; if it happens after, the dirty bit is set (as usual). Either way it gets reported in a subsequent scan and sent. ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-10-07 4:27 ` Kishen Maloor @ 2026-10-07 20:13 ` Peter Xu 2026-10-08 6:23 ` Tony Lindgren 0 siblings, 1 reply; 84+ messages in thread From: Peter Xu @ 2026-10-07 20:13 UTC (permalink / raw) To: Kishen Maloor Cc: Artem Bityutskiy, Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Tue, Oct 06, 2026 at 09:27:47PM -0700, Kishen Maloor wrote: > Hi Peter, Hi, Kishen, [...] > > Especially, MEM.IMPORT seems trickier to me: MEM.EXPORT even if described > > in in GPA address space, can be converted to other address spaces at will, > > as long as QEMU has full knowledge of the mappings. MEM.IMPORT, OTOH, may > > not trivially work in GPA terms, if memory mappings can overlap. > > I can appreciate that. Working in GPA terms assumes that a GPA reference is > unambiguous over the duration of migration, which holds true in the TDX case. > > But it would break if another backend is able to point a mapped GPA at a > different protected page while keeping both pages valid -- essentially > the PAM case on regular VMs. > > One option that comes to mind is to add a 2nd optional identifier to > kvm_memory_transfer consisting of a guest_memfd fd and an offsets array. > That would be at parity with QEMU's block+offset and directly reference > specific storage units. Then QEMU and KVM backends for a platform would > decide what identifier to use. > > The value might be that a bundle could be accepted by a backend for > storage that isn't currently mapped in a memslot. > > OTOH, slot+offset as an identifier might face the same problem as GPA in > this case when a particular RAMBlock's storage isn't mapped at the time > on the destination. > > Not sure if this idea is viable or even makes sense; this is mostly > for discussion, but would appreciate your thoughts. Yes, I would think (gmemfd, [array of fd offsets]) a nice interface. If we say gmemfd is the container of data in KVM world, then if it can provide all data accessors it feels all natural and fit. And yes, it also works for overlapped. Said that, it's a matter of if, TDX or any other solution, will use GPA to be part of attestation process when loading the data. My intepretation is GPA was part of TDX MEM.IMPORT workflow and required, but now since you proposed gmem for indexing pages, maybe I was wrong. In this case, I'd be more than happy if I was wrong.. Thanks, -- Peter Xu ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-10-07 20:13 ` Peter Xu @ 2026-10-08 6:23 ` Tony Lindgren 0 siblings, 0 replies; 84+ messages in thread From: Tony Lindgren @ 2026-10-08 6:23 UTC (permalink / raw) To: Peter Xu Cc: Kishen Maloor, Artem Bityutskiy, Paolo Bonzini, Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Wed, Oct 07, 2026 at 04:13:32PM -0400, Peter Xu wrote: > On Tue, Oct 06, 2026 at 09:27:47PM -0700, Kishen Maloor wrote: > > Hi Peter, > > Hi, Kishen, > > [...] > > > > Especially, MEM.IMPORT seems trickier to me: MEM.EXPORT even if described > > > in in GPA address space, can be converted to other address spaces at will, > > > as long as QEMU has full knowledge of the mappings. MEM.IMPORT, OTOH, may > > > not trivially work in GPA terms, if memory mappings can overlap. > > > > I can appreciate that. Working in GPA terms assumes that a GPA reference is > > unambiguous over the duration of migration, which holds true in the TDX case. > > > > But it would break if another backend is able to point a mapped GPA at a > > different protected page while keeping both pages valid -- essentially > > the PAM case on regular VMs. > > > > One option that comes to mind is to add a 2nd optional identifier to > > kvm_memory_transfer consisting of a guest_memfd fd and an offsets array. > > That would be at parity with QEMU's block+offset and directly reference > > specific storage units. Then QEMU and KVM backends for a platform would > > decide what identifier to use. > > > > The value might be that a bundle could be accepted by a backend for > > storage that isn't currently mapped in a memslot. > > > > OTOH, slot+offset as an identifier might face the same problem as GPA in > > this case when a particular RAMBlock's storage isn't mapped at the time > > on the destination. > > > > Not sure if this idea is viable or even makes sense; this is mostly > > for discussion, but would appreciate your thoughts. > > Yes, I would think (gmemfd, [array of fd offsets]) a nice interface. > > If we say gmemfd is the container of data in KVM world, then if it can > provide all data accessors it feels all natural and fit. And yes, it also > works for overlapped. > > Said that, it's a matter of if, TDX or any other solution, will use GPA to > be part of attestation process when loading the data. My intepretation is > GPA was part of TDX MEM.IMPORT workflow and required, but now since you > proposed gmem for indexing pages, maybe I was wrong. In this case, I'd be > more than happy if I was wrong.. Yes GPAs are part of the MAC generated by tdh_export_mem() so cannot change for import. This can be seen at the TDX module source at [0] below. What if we add a flag for the memory functions to describe the type of addressing used? [0] https://github.com/intel/confidential-computing.tdx.tdx-module/blob/tdx_1.5/src/vmm_dispatcher/migration_api_calls/tdh_export_mem.c#L1083 ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-08-31 7:13 [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Tony Lindgren ` (4 preceding siblings ...) 2026-09-04 18:24 ` [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Artem Bityutskiy @ 2026-09-18 18:36 ` Ionut Mihalcea 2026-09-21 4:35 ` Tony Lindgren 2026-09-25 16:03 ` Serge Hallyn (AMD) 6 siblings, 1 reply; 84+ messages in thread From: Ionut Mihalcea @ 2026-09-18 18:36 UTC (permalink / raw) To: Tony Lindgren, Paolo Bonzini, Sean Christopherson Cc: Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel , Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm@vger.kernel.org, Mathias Brossard, nd On 31 Aug 2026, at 08:13, Tony Lindgren <tony.lindgren@linux.intel.com> wrote: > Tom, since you mentioned that AMD SEV-SNP and Intel TDX live migration > sound similar, can you please take a look how the API might work for > SEV-SNP? While the x86 implementation is outside our scope, we're also currently investigating how the Arm CCA ABIs for Live Migration ABIs would fit as a backend for this API. We will provide more feedback on the API shape and semantics as the analysis progresses. As an intro to our approach (and a potential touchpoint on the topic of this API), Mathias Brossard will be presenting Arm CCA Live Migration at LPC next month [0]. > Dirty page tracking does not add a new uAPI. Userspace keeps using > KVM_GET_DIRTY_LOG and KVM_CLEAR_DIRTY_LOG. One concern we currently have is about an impedance mismatch between the non-CoCo dirty log APIs and our approach to dirty reporting. This is just to highlight the concern at the moment, without any explicit proposal as an alternative. [0] https://lpc.events/event/20/contributions/2476/ ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-09-18 18:36 ` Ionut Mihalcea @ 2026-09-21 4:35 ` Tony Lindgren 0 siblings, 0 replies; 84+ messages in thread From: Tony Lindgren @ 2026-09-21 4:35 UTC (permalink / raw) To: Ionut Mihalcea Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm@vger.kernel.org, Mathias Brossard, nd Hi, On Fri, Sep 18, 2026 at 06:36:38PM +0000, Ionut Mihalcea wrote: > On 31 Aug 2026, at 08:13, Tony Lindgren <tony.lindgren@linux.intel.com> wrote: > > Tom, since you mentioned that AMD SEV-SNP and Intel TDX live migration > > sound similar, can you please take a look how the API might work for > > SEV-SNP? > > While the x86 implementation is outside our scope, we're also currently > investigating how the Arm CCA ABIs for Live Migration ABIs would fit as a > backend for this API. We will provide more feedback on the API shape and > semantics as the analysis progresses. OK great, yes great if we can make this work also for ARM too. > As an intro to our approach (and a potential touchpoint on the topic of this > API), Mathias Brossard will be presenting Arm CCA Live Migration at LPC next > month [0]. OK > > Dirty page tracking does not add a new uAPI. Userspace keeps using > > KVM_GET_DIRTY_LOG and KVM_CLEAR_DIRTY_LOG. > > One concern we currently have is about an impedance mismatch between the > non-CoCo dirty log APIs and our approach to dirty reporting. This is just > to highlight the concern at the moment, without any explicit proposal as > an alternative. Care to describe a bit more what issues you are seeing with dirty log? What we have noticed with TDX is that KVM_GET_DIRTY_LOG works with additional vendor ops for live migration. But it is not usable outside migration. We currently get a list of migration candidates from the TDX module and that status sticks until migration or cancel operation. So the dirty log stats are invalid outside the migration context. This means that for example qemu calc_dirty_rate is not available before the migration for stats. I guess we could eventually get preprocessed stats from the TDX module to for reporting the dirty rate. But it seems that initially we may want to return nothing for dirty log outside migration. > [0] https://lpc.events/event/20/contributions/2476/ > > ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-08-31 7:13 [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Tony Lindgren ` (5 preceding siblings ...) 2026-09-18 18:36 ` Ionut Mihalcea @ 2026-09-25 16:03 ` Serge Hallyn (AMD) 2026-09-28 3:24 ` Kishen Maloor 6 siblings, 1 reply; 84+ messages in thread From: Serge Hallyn (AMD) @ 2026-09-25 16:03 UTC (permalink / raw) To: Tony Lindgren Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On Mon, Aug 31, 2026 at 10:13:00AM +0300, Tony Lindgren wrote: > Hi all, > > As discussed in a recent PUCK call, Sean suggested we post what Intel is > using for the KVM live migration API for TDX as an example to see if we > can come up with APIs that are not vendor specific. > > The goal of this patch series is to start a discussion about live migration > kernel APIs for CoCo guests. > > Tom, since you mentioned that AMD SEV-SNP and Intel TDX live migration > sound similar, can you please take a look how the API might work for > SEV-SNP? > > For CoCo VMs, the guest memory and vCPU states are not accessible to the > userspace or KVM for live migration. The memory and vCPU states need to be > extracted into encrypted blobs on the source, and decrypted on the > destination. Before live migration, an encryption key needs to be > negotiated between the source and destination. > > The layer handling the encryption for live migration is implementation > specific. It can be the TDX module or Coconut-SVSM for example. > > For CoCo VMs, the KVM_MEMORY_ENCRYPT_OP ioctl() has been used with vendor > specific sub-commands. Adding more vendor specific sub-commands is an > option also for live migration. However, depending on how similar the KVM > needs are, it may be possible to have a common API. > > For TDX, we're using a group of ioctl()s that might be possible to adapt > also for other CoCo implementations. > > Artem has put together a brief description below of the example API and the > migration flow: Hi Tony, is there any pubically available qemu git branch or patchset to show how you currently are driving this? > Example API > =========== > > - KVM_CAP_LIVE_MIGRATION - if a VM supports live migration through this > uAPI. > - KVM_MIGRATE_CMD - the main ioctl that drives the migration phases. Each > command takes vendor-specific flags and a buffer for the blob that travels > between the hosts. > - KVM_MIGRATE_SETUP - establish the migration session and transfer the > immutable VM state. > - KVM_MIGRATE_ITERATION - close a memory copy round. > - KVM_MIGRATE_STOP_AND_COPY - pause the VM and transfer the remaining VM > state. > - KVM_MIGRATE_END - complete the migration, or abort it. > - KVM_EXPORT_MEMORY - export memory pages on the source host. > - KVM_IMPORT_MEMORY - import memory pages on the destination host. > - KVM_EXPORT_VCPU - export vCPU state on the source host. > - KVM_IMPORT_VCPU - import vCPU state on the destination host. > > Dirty page tracking does not add a new uAPI. Userspace keeps using > KVM_GET_DIRTY_LOG and KVM_CLEAR_DIRTY_LOG. > > Migration flow > ============== > > Source host Destination host > =========== ================ > > CMD(SETUP/SESSION) <--- setup msgs ---> CMD(SETUP/SESSION) > | (repeated) | > CMD(SETUP/IMMUTABLE_STATE) - immutable state -> CMD(SETUP/IMMUTABLE_STATE) > | | > KVM_GET_DIRTY_LOG | > KVM_EXPORT_MEMORY --- memory data ---> KVM_IMPORT_MEMORY > CMD(ITERATION) --- epoch token ---> CMD(ITERATION) > | (repeat until convergence) | > CMD(STOP_AND_COPY/PAUSE) | > CMD(STOP_AND_COPY/TD_STATE) --- VM state ------> CMD(STOP_AND_COPY/TD_STATE) > KVM_EXPORT_VCPU --- vCPU state ----> KVM_IMPORT_VCPU > KVM_EXPORT_MEMORY -- final memory ---> KVM_IMPORT_MEMORY > CMD(ITERATION/DONE) --- start token ---> CMD(ITERATION) > | | > CMD(END) CMD(END) > > For the TDX implementation, the above map to the TDX module SEAMCALLs. > > Regards, > > Tony > > Changes since v1 at [0] below: > > - Drop KVM_MIGRATE_CMD sub-command PREPARE, SETUP sub-command has been > enough for TDX at least > > - Rename KVM_MIGRATE_CMD sub-command KVM_MIGRATE_TOKEN to > KVM_MIGRATE_ITERATION > > - Rename KVM_MIGRATE_CMD sub-command KVM_MIGRATE_SOURCE_BLACKOUT to > KVM_MIGRATE_STOP_AND_COPY > > - Add x86 ioctl handling > > [0] https://lore.kernel.org/kvm/20251006113524.1573116-1-tony.lindgren@linux.intel.com/ > > > Tony Lindgren (4): > Documentation: KVM: Add live migration API for confidential guests > KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD > KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY > KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU > > Documentation/virt/kvm/api.rst | 205 +++++++++++++++++++++++++++++ > arch/x86/include/asm/kvm-x86-ops.h | 6 + > arch/x86/include/asm/kvm_host.h | 6 + > arch/x86/kvm/x86.c | 102 ++++++++++++++ > include/uapi/linux/kvm.h | 43 ++++++ > 5 files changed, 362 insertions(+) > > > base-commit: dc59e4fea9d83f03bad6bddf3fa2e52491777482 > -- > 2.43.0 > ^ permalink raw reply [flat|nested] 84+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration 2026-09-25 16:03 ` Serge Hallyn (AMD) @ 2026-09-28 3:24 ` Kishen Maloor 0 siblings, 0 replies; 84+ messages in thread From: Kishen Maloor @ 2026-09-28 3:24 UTC (permalink / raw) To: Serge Hallyn (AMD), Tony Lindgren Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price, Anup Patel, Samuel Ortiz, Jakub Růžička, Jörg Rödel, Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm On 9/25/26 9:03 AM, Serge Hallyn (AMD) wrote: > On Mon, Aug 31, 2026 at 10:13:00AM +0300, Tony Lindgren wrote: > ... > > is there any pubically available qemu git branch or patchset to show > how you currently are driving this? > Hi Serge, Nothing public yet. Our QEMU code is throwaway scaffolding for exercising the ioctls. We're working on something more suitable to post as an RFC. ^ permalink raw reply [flat|nested] 84+ messages in thread
end of thread, other threads:[~2026-10-08 9:22 UTC | newest] Thread overview: 84+ messages (download: mbox.gz follow: Atom feed -- links below jump to the message on this page -- 2026-08-31 7:13 [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Tony Lindgren 2026-08-31 7:13 ` [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests Tony Lindgren 2026-08-31 7:20 ` sashiko-bot 2026-09-18 11:35 ` Peter Xu 2026-09-21 4:20 ` Tony Lindgren 2026-09-24 1:50 ` Wei Wang 2026-09-24 4:51 ` Tony Lindgren 2026-08-31 7:13 ` [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD Tony Lindgren 2026-08-31 7:23 ` sashiko-bot 2026-09-01 6:03 ` Tony Lindgren 2026-09-07 11:53 ` Tony Lindgren 2026-09-07 13:15 ` Jörg Rödel 2026-09-07 13:32 ` Artem Bityutskiy 2026-09-08 4:15 ` Tony Lindgren 2026-09-08 4:43 ` Tony Lindgren 2026-09-09 0:22 ` Kishen Maloor 2026-09-09 6:57 ` Tony Lindgren 2026-09-10 1:11 ` Kishen Maloor 2026-09-10 6:33 ` Tony Lindgren 2026-09-11 1:40 ` Kishen Maloor 2026-09-11 4:23 ` Tony Lindgren 2026-09-15 0:14 ` Kishen Maloor 2026-09-15 4:44 ` Tony Lindgren 2026-09-15 15:53 ` Kishen Maloor 2026-09-16 5:09 ` Tony Lindgren 2026-09-17 3:31 ` Kishen Maloor 2026-09-17 6:42 ` Tony Lindgren 2026-09-18 4:32 ` Kishen Maloor 2026-09-18 5:58 ` Tony Lindgren 2026-09-21 0:13 ` Kishen Maloor 2026-09-21 6:52 ` Tony Lindgren 2026-09-21 9:24 ` Tony Lindgren 2026-09-21 10:58 ` Tony Lindgren 2026-09-22 3:57 ` Kishen Maloor 2026-09-22 5:25 ` Tony Lindgren 2026-09-23 0:38 ` Kishen Maloor 2026-09-23 6:04 ` Tony Lindgren 2026-09-24 5:53 ` Kishen Maloor 2026-09-24 6:59 ` Tony Lindgren 2026-09-18 4:33 ` Kishen Maloor 2026-09-21 5:58 ` Tony Lindgren 2026-09-21 6:56 ` Tony Lindgren 2026-09-22 3:56 ` Kishen Maloor 2026-09-22 6:27 ` Tony Lindgren 2026-09-23 0:37 ` Kishen Maloor 2026-09-23 6:50 ` Tony Lindgren 2026-09-24 5:34 ` Kishen Maloor 2026-09-24 7:15 ` Tony Lindgren 2026-10-08 9:22 ` Tony Lindgren 2026-08-31 7:13 ` [RFC PATCH v2 3/4] KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY Tony Lindgren 2026-08-31 7:23 ` sashiko-bot 2026-09-01 6:10 ` Tony Lindgren 2026-08-31 7:13 ` [RFC PATCH v2 4/4] KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU Tony Lindgren 2026-08-31 7:23 ` sashiko-bot 2026-09-01 6:12 ` Tony Lindgren 2026-09-04 18:24 ` [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Artem Bityutskiy 2026-09-17 21:27 ` Peter Xu 2026-09-18 12:46 ` Artem Bityutskiy 2026-09-18 15:53 ` Peter Xu 2026-09-22 8:09 ` Artem Bityutskiy 2026-09-22 9:42 ` Tony Lindgren 2026-09-22 11:54 ` Artem Bityutskiy 2026-09-23 4:20 ` Tony Lindgren 2026-09-22 21:18 ` Peter Xu 2026-09-23 12:05 ` Artem Bityutskiy 2026-09-24 21:19 ` Peter Xu 2026-09-28 14:15 ` Artem Bityutskiy 2026-09-29 21:05 ` Peter Xu 2026-10-02 19:57 ` Artem Bityutskiy 2026-10-07 20:00 ` Peter Xu 2026-09-23 15:28 ` Serge Hallyn (AMD) 2026-09-20 23:56 ` Kishen Maloor 2026-09-23 21:36 ` Peter Xu 2026-09-24 4:27 ` Kishen Maloor 2026-09-25 14:18 ` Peter Xu 2026-09-29 1:28 ` Kishen Maloor 2026-09-30 20:42 ` Peter Xu 2026-10-07 4:27 ` Kishen Maloor 2026-10-07 20:13 ` Peter Xu 2026-10-08 6:23 ` Tony Lindgren 2026-09-18 18:36 ` Ionut Mihalcea 2026-09-21 4:35 ` Tony Lindgren 2026-09-25 16:03 ` Serge Hallyn (AMD) 2026-09-28 3:24 ` Kishen Maloor
This is an external index of several public inboxes, see mirroring instructions on how to clone and mirror all data and code used by this external index.