* [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
@ 2026-08-31 7:13 Tony Lindgren
2026-08-31 7:13 ` [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests Tony Lindgren
` (6 more replies)
0 siblings, 7 replies; 78+ messages in thread
From: Tony Lindgren @ 2026-08-31 7:13 UTC (permalink / raw)
To: Paolo Bonzini, Sean Christopherson
Cc: Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Jörg Rödel ,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
Hi all,
As discussed in a recent PUCK call, Sean suggested we post what Intel is
using for the KVM live migration API for TDX as an example to see if we
can come up with APIs that are not vendor specific.
The goal of this patch series is to start a discussion about live migration
kernel APIs for CoCo guests.
Tom, since you mentioned that AMD SEV-SNP and Intel TDX live migration
sound similar, can you please take a look how the API might work for
SEV-SNP?
For CoCo VMs, the guest memory and vCPU states are not accessible to the
userspace or KVM for live migration. The memory and vCPU states need to be
extracted into encrypted blobs on the source, and decrypted on the
destination. Before live migration, an encryption key needs to be
negotiated between the source and destination.
The layer handling the encryption for live migration is implementation
specific. It can be the TDX module or Coconut-SVSM for example.
For CoCo VMs, the KVM_MEMORY_ENCRYPT_OP ioctl() has been used with vendor
specific sub-commands. Adding more vendor specific sub-commands is an
option also for live migration. However, depending on how similar the KVM
needs are, it may be possible to have a common API.
For TDX, we're using a group of ioctl()s that might be possible to adapt
also for other CoCo implementations.
Artem has put together a brief description below of the example API and the
migration flow:
Example API
===========
- KVM_CAP_LIVE_MIGRATION - if a VM supports live migration through this
uAPI.
- KVM_MIGRATE_CMD - the main ioctl that drives the migration phases. Each
command takes vendor-specific flags and a buffer for the blob that travels
between the hosts.
- KVM_MIGRATE_SETUP - establish the migration session and transfer the
immutable VM state.
- KVM_MIGRATE_ITERATION - close a memory copy round.
- KVM_MIGRATE_STOP_AND_COPY - pause the VM and transfer the remaining VM
state.
- KVM_MIGRATE_END - complete the migration, or abort it.
- KVM_EXPORT_MEMORY - export memory pages on the source host.
- KVM_IMPORT_MEMORY - import memory pages on the destination host.
- KVM_EXPORT_VCPU - export vCPU state on the source host.
- KVM_IMPORT_VCPU - import vCPU state on the destination host.
Dirty page tracking does not add a new uAPI. Userspace keeps using
KVM_GET_DIRTY_LOG and KVM_CLEAR_DIRTY_LOG.
Migration flow
==============
Source host Destination host
=========== ================
CMD(SETUP/SESSION) <--- setup msgs ---> CMD(SETUP/SESSION)
| (repeated) |
CMD(SETUP/IMMUTABLE_STATE) - immutable state -> CMD(SETUP/IMMUTABLE_STATE)
| |
KVM_GET_DIRTY_LOG |
KVM_EXPORT_MEMORY --- memory data ---> KVM_IMPORT_MEMORY
CMD(ITERATION) --- epoch token ---> CMD(ITERATION)
| (repeat until convergence) |
CMD(STOP_AND_COPY/PAUSE) |
CMD(STOP_AND_COPY/TD_STATE) --- VM state ------> CMD(STOP_AND_COPY/TD_STATE)
KVM_EXPORT_VCPU --- vCPU state ----> KVM_IMPORT_VCPU
KVM_EXPORT_MEMORY -- final memory ---> KVM_IMPORT_MEMORY
CMD(ITERATION/DONE) --- start token ---> CMD(ITERATION)
| |
CMD(END) CMD(END)
For the TDX implementation, the above map to the TDX module SEAMCALLs.
Regards,
Tony
Changes since v1 at [0] below:
- Drop KVM_MIGRATE_CMD sub-command PREPARE, SETUP sub-command has been
enough for TDX at least
- Rename KVM_MIGRATE_CMD sub-command KVM_MIGRATE_TOKEN to
KVM_MIGRATE_ITERATION
- Rename KVM_MIGRATE_CMD sub-command KVM_MIGRATE_SOURCE_BLACKOUT to
KVM_MIGRATE_STOP_AND_COPY
- Add x86 ioctl handling
[0] https://lore.kernel.org/kvm/20251006113524.1573116-1-tony.lindgren@linux.intel.com/
Tony Lindgren (4):
Documentation: KVM: Add live migration API for confidential guests
KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY
KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU
Documentation/virt/kvm/api.rst | 205 +++++++++++++++++++++++++++++
arch/x86/include/asm/kvm-x86-ops.h | 6 +
arch/x86/include/asm/kvm_host.h | 6 +
arch/x86/kvm/x86.c | 102 ++++++++++++++
include/uapi/linux/kvm.h | 43 ++++++
5 files changed, 362 insertions(+)
base-commit: dc59e4fea9d83f03bad6bddf3fa2e52491777482
--
2.43.0
^ permalink raw reply [flat|nested] 78+ messages in thread
* [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests
2026-08-31 7:13 [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Tony Lindgren
@ 2026-08-31 7:13 ` Tony Lindgren
2026-08-31 7:20 ` sashiko-bot
2026-09-18 11:35 ` Peter Xu
2026-08-31 7:13 ` [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD Tony Lindgren
` (5 subsequent siblings)
6 siblings, 2 replies; 78+ messages in thread
From: Tony Lindgren @ 2026-08-31 7:13 UTC (permalink / raw)
To: Paolo Bonzini, Sean Christopherson
Cc: Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Jörg Rödel ,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
For CoCo VMs, the guest memory and vCPU states are not accessible to the
userspace or KVM for live migration. The memory and vCPU states need to be
extracted into encrypted blobs on the source, and decrypted on the
destination. Before live migration, an encryption key needs to be
negotiated between the source and destination.
KVM help is needed to talk to the layer exporting and importing the
encrypted state. Document the KVM live migration API for confidential
guests.
Co-developed-by: Kishen Maloor <kishen.maloor@intel.com>
Signed-off-by: Kishen Maloor <kishen.maloor@intel.com>
Signed-off-by: Tony Lindgren <tony.lindgren@linux.intel.com>
---
Documentation/virt/kvm/api.rst | 205 +++++++++++++++++++++++++++++++++
1 file changed, 205 insertions(+)
diff --git a/Documentation/virt/kvm/api.rst b/Documentation/virt/kvm/api.rst
index a5f9ee92f43e8..9d546d288af5f 100644
--- a/Documentation/virt/kvm/api.rst
+++ b/Documentation/virt/kvm/api.rst
@@ -6566,6 +6566,200 @@ KVM_S390_KEYOP_SSKE
Sets the storage key for the guest address ``guest_addr`` to the key
specified in ``key``, returning the previous value in ``key``.
+.. _KVM_MIGRATE_CMD:
+
+4.145 KVM_MIGRATE_CMD
+---------------------
+
+:Capability: KVM_CAP_LIVE_MIGRATION
+:Architectures: x86
+:Type: vm ioctl
+:Parameters: struct kvm_migrate_cmd (in/out)
+:Returns: 0 on success, < 0 on error
+
+Allows userspace to send live migration related commands to KVM for vendor
+specific handling.
+
+For confidential computing, live migration related commands may be needed.
+The commands typically use encrypted data that needs to be passed between the
+source and destination hosts. The hosts may also require specific coordination
+steps during migration that must be triggered at precise points in the
+migration process.
+
+The vendor specific implementation handles locking and checks the valid flags
+bits. If KVM_CAP_LIVE_MIGRATION is not available for the VM, -ENOTTY is
+returned.
+
+The KVM_MIGRATE_CMD subcommand passed in struct kvm_migrate_cmd is one of::
+
+ #define KVM_MIGRATE_SETUP 0
+ #define KVM_MIGRATE_ITERATION 1
+ #define KVM_MIGRATE_STOP_AND_COPY 2
+ #define KVM_MIGRATE_ABORT 3
+ #define KVM_MIGRATE_END 4
+
+The kvm_transfer_buffer is::
+
+ /**
+ * @address: Userspace buffer address
+ * @size: Size of the userspace buffer
+ * @reserved: Reserved for future use
+ */
+ struct kvm_transfer_buffer {
+ __u64 address;
+ __u32 size;
+ __u32 reserved;
+ };
+
+The kvm_migrate_cmd is::
+
+ /**
+ * @command: One of the defined KVM_MIGRATE commands
+ * @flags: Hardware specific flags
+ * @reserved: Reserved for future use
+ * @buf: Userspace buffer for hardware specific data
+ */
+ struct kvm_migrate_cmd {
+ __u16 command;
+ __u16 flags;
+ __u32 reserved;
+ struct kvm_transfer_buffer buf;
+ };
+
+.. _KVM_EXPORT_MEMORY:
+
+4.146 KVM_EXPORT_MEMORY
+-----------------------
+
+:Capability: KVM_CAP_LIVE_MIGRATION
+:Architectures: x86
+:Type: vm ioctl
+:Parameters: struct kvm_memory_transfer (in/out)
+:Returns: 0 on success, < 0 on error
+
+Allows userspace to request the host to export an array of memory pages to a
+userspace buffer.
+
+The private memory may not be accessible to KVM because of encryption. For
+confidential computing, the guest memory is encrypted and only accessible to
+the guest.
+
+If KVM_CAP_LIVE_MIGRATION is not available for the VM, -ENOTTY is returned.
+
+The vendor specific ID is used at least for TDX for the migration thread
+index.
+
+The kvm_memory_transfer is::
+
+ /**
+ * @gfns: Userspace address of an array of nr_gfns __u64 GFNs to export
+ * @nr_gfns: Number of GFNs in the @gfns array
+ * @id: Optional vendor specific transfer ID
+ * @flags: Vendor specific flags
+ * @reserved: Reserved for future use
+ * @buf: Userspace buffer to export memory to
+ */
+ struct kvm_memory_transfer {
+ __u64 gfns;
+ __u32 nr_gfns;
+ __u16 id;
+ __u16 flags;
+ __u64 reserved;
+ struct kvm_transfer_buffer buf;
+ };
+
+The transfer buffer size is vendor specific.
+
+For the transfer buffer, seeo :ref:`KVM_MIGRATE_CMD <KVM_MIGRATE_CMD>`.
+
+For memory import, see also :ref:`KVM_IMPORT_MEMORY <KVM_IMPORT_MEMORY>`.
+
+
+.. _KVM_IMPORT_MEMORY:
+
+4.147 KVM_IMPORT_MEMORY
+-----------------------
+
+:Capability: KVM_CAP_LIVE_MIGRATION
+:Architectures: x86
+:Type: vm ioctl
+:Parameters: struct kvm_memory_transfer (in/out)
+:Returns: 0 on success, < 0 on error
+
+Allows userspace to request the host to import an array of memory pages from a
+userspace buffer.
+
+The private memory may not be accessible to KVM because of encryption. For
+confidential computing, the guest memory is encrypted and only accessible to
+the guest.
+
+If KVM_CAP_LIVE_MIGRATION is not available for the VM, -ENOTTY is returned.
+
+The vendor specific ID is used at least for TDX for the migration thread
+index.
+
+The transfer buffer size is vendor specific.
+
+For kvm_memory_transfer, see :ref:`KVM_EXPORT_MEMORY <KVM_EXPORT_MEMORY>`.
+
+For the transfer buffer, seeo :ref:`KVM_MIGRATE_CMD <KVM_MIGRATE_CMD>`.
+
+.. _KVM_EXPORT_VCPU:
+
+4.149 KVM_EXPORT_VCPU
+---------------------
+:Capability: KVM_CAP_LIVE_MIGRATION
+:Architectures: arm64, x86
+:Type: vcpu ioctl
+:Parameters: struct kvm_vcpu_transfer (in/out)
+:Returns: 0 on success, < 0 on error
+
+Allows userspace to request the host to export a VCPU state to a userspace
+buffer.
+
+The VCPU state may not be directly accessible to KVM because of encryption. For
+confidential computing, the VCPU state is encrypted and only accessible to the
+guest.
+
+The vcpu_transfer is::
+
+ /**
+ * @flags: Hardware specific flags
+ * @reserved: Reserved for future use
+ * @buf: Userspace buffer to export VCPU state to
+ */
+ struct kvm_vcpu_transfer {
+ __u32 flags;
+ __u32 reserved;
+ struct kvm_transfer_buffer buf;
+ };
+
+For the transfer buffer, see :ref:`KVM_MIGRATE_CMD <KVM_MIGRATE_CMD>`.
+
+For vCPU import, see also :ref:`KVM_IMPORT_VCPU <KVM_IMPORT_VCPU>`.
+
+.. _KVM_IMPORT_VCPU:
+
+4.148 KVM_IMPORT_VCPU
+---------------------
+
+:Capability: KVM_CAP_LIVE_MIGRATION
+:Architectures: arm64, x86
+:Type: vcpu ioctl
+:Parameters: struct kvm_vcpu_transfer (in/out)
+:Returns: 0 on success, < 0 on error
+
+Allows userspace to request the host to import a VCPU state from a userspace
+buffer.
+
+The VCPU state may not be directly accessible to KVM because of encryption. For
+confidential computing, the VCPU state is encrypted and only accessible to the
+guest.
+
+For vcpu_transfer and vCPU import, see :ref:`KVM_IMPORT_VCPU <KVM_IMPORT_VCPU>`.
+
+For the transfer buffer, see also :ref:`KVM_MIGRATE_CMD <KVM_MIGRATE_CMD>`.
+
.. _kvm_run:
5. The kvm_run structure
@@ -9493,6 +9687,17 @@ take care to differentiate between these cases.
The presence of this capability indicates that the nested KVM guest can
start in ESA mode.
+8.48 KVM_CAP_LIVE_MIGRATION
+---------------------------
+
+:Architectures: x86
+:Target: VM
+:Parameters: None
+
+Indicates that the VM needs to use KVM calls for live migration, and that the
+KVM_MIGRATE_CMD ioctl and the KVM_EXPORT_MEMORY, KVM_IMPORT_MEMORY,
+KVM_EXPORT_VCPU and KVM_IMPORT_VCPU ioctls are available.
+
9. Known KVM API problems
=========================
--
2.43.0
^ permalink raw reply related [flat|nested] 78+ messages in thread
* [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-08-31 7:13 [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Tony Lindgren
2026-08-31 7:13 ` [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests Tony Lindgren
@ 2026-08-31 7:13 ` Tony Lindgren
2026-08-31 7:23 ` sashiko-bot
` (2 more replies)
2026-08-31 7:13 ` [RFC PATCH v2 3/4] KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY Tony Lindgren
` (4 subsequent siblings)
6 siblings, 3 replies; 78+ messages in thread
From: Tony Lindgren @ 2026-08-31 7:13 UTC (permalink / raw)
To: Paolo Bonzini, Sean Christopherson
Cc: Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Jörg Rödel ,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
Live migration of confidential guests needs the help of KVM at least for
TDX.
Add KVM_CAP_LIVE_MIGRATION for when hardware specific live migration
functions must be used.
Add KVM_MIGRATE_CMD to configure the hardware for live migration.
Assisted-by: Claude-Code:claude-opus-5 checkpatch
[ used AI to review and simplify the code ]
Signed-off-by: Tony Lindgren <tony.lindgren@linux.intel.com>
---
arch/x86/include/asm/kvm-x86-ops.h | 2 ++
arch/x86/include/asm/kvm_host.h | 2 ++
arch/x86/kvm/x86.c | 25 +++++++++++++++++++++++++
include/uapi/linux/kvm.h | 22 ++++++++++++++++++++++
4 files changed, 51 insertions(+)
diff --git a/arch/x86/include/asm/kvm-x86-ops.h b/arch/x86/include/asm/kvm-x86-ops.h
index 83dc5086138b3..ac080b556b0c8 100644
--- a/arch/x86/include/asm/kvm-x86-ops.h
+++ b/arch/x86/include/asm/kvm-x86-ops.h
@@ -148,6 +148,8 @@ KVM_X86_OP_OPTIONAL(alloc_apic_backing_page)
KVM_X86_OP_OPTIONAL_RET0(gmem_prepare)
KVM_X86_OP_OPTIONAL_RET0(gmem_max_mapping_level)
KVM_X86_OP_OPTIONAL(gmem_invalidate)
+KVM_X86_OP_OPTIONAL_RET0(cap_live_migration)
+KVM_X86_OP_OPTIONAL(migrate_cmd)
#endif
#undef KVM_X86_OP
diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h
index 5f6c1ce9673b7..d9291a8a97bb1 100644
--- a/arch/x86/include/asm/kvm_host.h
+++ b/arch/x86/include/asm/kvm_host.h
@@ -2010,6 +2010,8 @@ struct kvm_x86_ops {
int (*gmem_prepare)(struct kvm *kvm, kvm_pfn_t pfn, gfn_t gfn, int max_order);
void (*gmem_invalidate)(kvm_pfn_t start, kvm_pfn_t end);
int (*gmem_max_mapping_level)(struct kvm *kvm, kvm_pfn_t pfn, bool is_private);
+ bool (*cap_live_migration)(struct kvm *kvm);
+ int (*migrate_cmd)(struct kvm *kvm, struct kvm_migrate_cmd *cmd);
};
struct kvm_x86_nested_ops {
diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c
index afcac1042947a..7064fd709e56d 100644
--- a/arch/x86/kvm/x86.c
+++ b/arch/x86/kvm/x86.c
@@ -4973,6 +4973,9 @@ int kvm_vm_ioctl_check_extension(struct kvm *kvm, long ext)
case KVM_CAP_READONLY_MEM:
r = kvm ? kvm_arch_has_readonly_mem(kvm) : 1;
break;
+ case KVM_CAP_LIVE_MIGRATION:
+ r = kvm ? kvm_x86_call(cap_live_migration)(kvm) : 0;
+ break;
default:
break;
}
@@ -7614,6 +7617,28 @@ int kvm_arch_vm_ioctl(struct file *filp, unsigned int ioctl, unsigned long arg)
r = kvm_vm_ioctl_set_msr_filter(kvm, &filter);
break;
}
+ case KVM_MIGRATE_CMD: {
+ struct kvm_migrate_cmd cmd;
+
+ if (!kvm_x86_ops.migrate_cmd ||
+ !kvm_x86_call(cap_live_migration)(kvm))
+ return -ENOTTY;
+
+ if (copy_from_user(&cmd, argp, sizeof(cmd)))
+ return -EFAULT;
+
+ if (cmd.reserved || cmd.buf.reserved)
+ return -EINVAL;
+
+ r = kvm_x86_call(migrate_cmd)(kvm, &cmd);
+ if (r > 0)
+ r = -EIO;
+
+ /* Copy back also on an error to report a partially done command */
+ if (copy_to_user(argp, &cmd, sizeof(cmd)))
+ return -EFAULT;
+ break;
+ }
default:
r = -ENOTTY;
}
diff --git a/include/uapi/linux/kvm.h b/include/uapi/linux/kvm.h
index 419011097fa8e..e5b227d3a8e40 100644
--- a/include/uapi/linux/kvm.h
+++ b/include/uapi/linux/kvm.h
@@ -997,6 +997,7 @@ struct kvm_enable_cap {
#define KVM_CAP_S390_KEYOP 247
#define KVM_CAP_S390_VSIE_ESAMODE 248
#define KVM_CAP_S390_HPAGE_2G 249
+#define KVM_CAP_LIVE_MIGRATION 250
struct kvm_irq_routing_irqchip {
__u32 irqchip;
@@ -1350,6 +1351,8 @@ struct kvm_s390_keyop {
#define KVM_GET_DEVICE_ATTR _IOW(KVMIO, 0xe2, struct kvm_device_attr)
#define KVM_HAS_DEVICE_ATTR _IOW(KVMIO, 0xe3, struct kvm_device_attr)
+#define KVM_MIGRATE_CMD _IOWR(KVMIO, 0xe4, struct kvm_migrate_cmd)
+
/*
* ioctls for vcpu fds
*/
@@ -1670,4 +1673,23 @@ struct kvm_pre_fault_memory {
__u64 padding[5];
};
+#define KVM_MIGRATE_SETUP 0
+#define KVM_MIGRATE_ITERATION 1
+#define KVM_MIGRATE_STOP_AND_COPY 2
+#define KVM_MIGRATE_ABORT 3
+#define KVM_MIGRATE_END 4
+
+struct kvm_transfer_buffer {
+ __u64 address;
+ __u32 size;
+ __u32 reserved;
+};
+
+struct kvm_migrate_cmd {
+ __u16 command;
+ __u16 flags;
+ __u32 reserved;
+ struct kvm_transfer_buffer buf;
+};
+
#endif /* __LINUX_KVM_H */
--
2.43.0
^ permalink raw reply related [flat|nested] 78+ messages in thread
* [RFC PATCH v2 3/4] KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY
2026-08-31 7:13 [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Tony Lindgren
2026-08-31 7:13 ` [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests Tony Lindgren
2026-08-31 7:13 ` [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD Tony Lindgren
@ 2026-08-31 7:13 ` Tony Lindgren
2026-08-31 7:23 ` sashiko-bot
2026-08-31 7:13 ` [RFC PATCH v2 4/4] KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU Tony Lindgren
` (3 subsequent siblings)
6 siblings, 1 reply; 78+ messages in thread
From: Tony Lindgren @ 2026-08-31 7:13 UTC (permalink / raw)
To: Paolo Bonzini, Sean Christopherson
Cc: Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Jörg Rödel ,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
Add support to export and import KVM memory for cases where the memory
is only accessible to the guest. Live migration of confidential computing
needs help of KVM for the vendor specific calls at least for TDX.
Introduce optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY.
Based on earlier code by Wei Wang <wei.w.wang@intel.com>.
Co-developed-by: Kishen Maloor <kishen.maloor@intel.com>
Signed-off-by: Kishen Maloor <kishen.maloor@intel.com>
Assisted-by: Claude-Code:claude-opus-5 checkpatch
[ used AI to review and simplify the code ]
Signed-off-by: Tony Lindgren <tony.lindgren@linux.intel.com>
---
arch/x86/include/asm/kvm-x86-ops.h | 2 ++
arch/x86/include/asm/kvm_host.h | 2 ++
arch/x86/kvm/x86.c | 37 ++++++++++++++++++++++++++++++
include/uapi/linux/kvm.h | 13 +++++++++++
4 files changed, 54 insertions(+)
diff --git a/arch/x86/include/asm/kvm-x86-ops.h b/arch/x86/include/asm/kvm-x86-ops.h
index ac080b556b0c8..173d0c4f1115e 100644
--- a/arch/x86/include/asm/kvm-x86-ops.h
+++ b/arch/x86/include/asm/kvm-x86-ops.h
@@ -150,6 +150,8 @@ KVM_X86_OP_OPTIONAL_RET0(gmem_max_mapping_level)
KVM_X86_OP_OPTIONAL(gmem_invalidate)
KVM_X86_OP_OPTIONAL_RET0(cap_live_migration)
KVM_X86_OP_OPTIONAL(migrate_cmd)
+KVM_X86_OP_OPTIONAL(export_memory)
+KVM_X86_OP_OPTIONAL(import_memory)
#endif
#undef KVM_X86_OP
diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h
index d9291a8a97bb1..9a517bfc2f3a6 100644
--- a/arch/x86/include/asm/kvm_host.h
+++ b/arch/x86/include/asm/kvm_host.h
@@ -2012,6 +2012,8 @@ struct kvm_x86_ops {
int (*gmem_max_mapping_level)(struct kvm *kvm, kvm_pfn_t pfn, bool is_private);
bool (*cap_live_migration)(struct kvm *kvm);
int (*migrate_cmd)(struct kvm *kvm, struct kvm_migrate_cmd *cmd);
+ int (*export_memory)(struct kvm *kvm, struct kvm_memory_transfer *mem);
+ int (*import_memory)(struct kvm *kvm, struct kvm_memory_transfer *mem);
};
struct kvm_x86_nested_ops {
diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c
index 7064fd709e56d..8a99c665008a3 100644
--- a/arch/x86/kvm/x86.c
+++ b/arch/x86/kvm/x86.c
@@ -7258,6 +7258,37 @@ long kvm_arch_vcpu_unlocked_ioctl(struct file *filp, unsigned int ioctl,
return -ENOIOCTLCMD;
}
+static int kvm_vm_ioctl_transfer_memory(struct kvm *kvm, bool import,
+ void __user *argp)
+{
+ struct kvm_memory_transfer mem;
+ int r;
+
+ if (!kvm_x86_call(cap_live_migration)(kvm) ||
+ (import && !kvm_x86_ops.import_memory) ||
+ (!import && !kvm_x86_ops.export_memory))
+ return -ENOTTY;
+
+ if (copy_from_user(&mem, argp, sizeof(mem)))
+ return -EFAULT;
+
+ if (mem.reserved || mem.buf.reserved || !mem.nr_gfns)
+ return -EINVAL;
+
+ if (import)
+ r = kvm_x86_call(import_memory)(kvm, &mem);
+ else
+ r = kvm_x86_call(export_memory)(kvm, &mem);
+ if (r > 0)
+ r = -EIO;
+
+ /* Copy back also on an error to report a partially done transfer */
+ if (copy_to_user(argp, &mem, sizeof(mem)))
+ return -EFAULT;
+
+ return r;
+}
+
int kvm_arch_vm_ioctl(struct file *filp, unsigned int ioctl, unsigned long arg)
{
struct kvm *kvm = filp->private_data;
@@ -7639,6 +7670,12 @@ int kvm_arch_vm_ioctl(struct file *filp, unsigned int ioctl, unsigned long arg)
return -EFAULT;
break;
}
+ case KVM_EXPORT_MEMORY:
+ r = kvm_vm_ioctl_transfer_memory(kvm, false, argp);
+ break;
+ case KVM_IMPORT_MEMORY:
+ r = kvm_vm_ioctl_transfer_memory(kvm, true, argp);
+ break;
default:
r = -ENOTTY;
}
diff --git a/include/uapi/linux/kvm.h b/include/uapi/linux/kvm.h
index e5b227d3a8e40..666bbdf220d65 100644
--- a/include/uapi/linux/kvm.h
+++ b/include/uapi/linux/kvm.h
@@ -1493,6 +1493,10 @@ struct kvm_enc_region {
#define KVM_GET_SREGS2 _IOR(KVMIO, 0xcc, struct kvm_sregs2)
#define KVM_SET_SREGS2 _IOW(KVMIO, 0xcd, struct kvm_sregs2)
+/* Available with KVM_CAP_LIVE_MIGRATION */
+#define KVM_EXPORT_MEMORY _IOWR(KVMIO, 0xe5, struct kvm_memory_transfer)
+#define KVM_IMPORT_MEMORY _IOWR(KVMIO, 0xe6, struct kvm_memory_transfer)
+
#define KVM_DIRTY_LOG_MANUAL_PROTECT_ENABLE (1 << 0)
#define KVM_DIRTY_LOG_INITIALLY_SET (1 << 1)
@@ -1692,4 +1696,13 @@ struct kvm_migrate_cmd {
struct kvm_transfer_buffer buf;
};
+struct kvm_memory_transfer {
+ __u64 gfns;
+ __u32 nr_gfns;
+ __u16 id;
+ __u16 flags;
+ __u64 reserved;
+ struct kvm_transfer_buffer buf;
+};
+
#endif /* __LINUX_KVM_H */
--
2.43.0
^ permalink raw reply related [flat|nested] 78+ messages in thread
* [RFC PATCH v2 4/4] KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU
2026-08-31 7:13 [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Tony Lindgren
` (2 preceding siblings ...)
2026-08-31 7:13 ` [RFC PATCH v2 3/4] KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY Tony Lindgren
@ 2026-08-31 7:13 ` Tony Lindgren
2026-08-31 7:23 ` sashiko-bot
2026-09-04 18:24 ` [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Artem Bityutskiy
` (2 subsequent siblings)
6 siblings, 1 reply; 78+ messages in thread
From: Tony Lindgren @ 2026-08-31 7:13 UTC (permalink / raw)
To: Paolo Bonzini, Sean Christopherson
Cc: Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Jörg Rödel ,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
Add support to export and import VCPU for cases where the VCPU state is
only accessible to the guest. Live migration of confidential computing
needs help of KVM for the firmware specific calls at least for TDX.
Introduce optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU.
Based on earlier code by Wei Wang <wei.w.wang@intel.com>.
Co-developed-by: Kishen Maloor <kishen.maloor@intel.com>
Signed-off-by: Kishen Maloor <kishen.maloor@intel.com>
Signed-off-by: Tony Lindgren <tony.lindgren@linux.intel.com>
---
arch/x86/include/asm/kvm-x86-ops.h | 2 ++
arch/x86/include/asm/kvm_host.h | 2 ++
arch/x86/kvm/x86.c | 40 ++++++++++++++++++++++++++++++
include/uapi/linux/kvm.h | 8 ++++++
4 files changed, 52 insertions(+)
diff --git a/arch/x86/include/asm/kvm-x86-ops.h b/arch/x86/include/asm/kvm-x86-ops.h
index 173d0c4f1115e..7f110f80d6f82 100644
--- a/arch/x86/include/asm/kvm-x86-ops.h
+++ b/arch/x86/include/asm/kvm-x86-ops.h
@@ -152,6 +152,8 @@ KVM_X86_OP_OPTIONAL_RET0(cap_live_migration)
KVM_X86_OP_OPTIONAL(migrate_cmd)
KVM_X86_OP_OPTIONAL(export_memory)
KVM_X86_OP_OPTIONAL(import_memory)
+KVM_X86_OP_OPTIONAL(export_vcpu)
+KVM_X86_OP_OPTIONAL(import_vcpu)
#endif
#undef KVM_X86_OP
diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h
index 9a517bfc2f3a6..b6362408dab80 100644
--- a/arch/x86/include/asm/kvm_host.h
+++ b/arch/x86/include/asm/kvm_host.h
@@ -2014,6 +2014,8 @@ struct kvm_x86_ops {
int (*migrate_cmd)(struct kvm *kvm, struct kvm_migrate_cmd *cmd);
int (*export_memory)(struct kvm *kvm, struct kvm_memory_transfer *mem);
int (*import_memory)(struct kvm *kvm, struct kvm_memory_transfer *mem);
+ int (*export_vcpu)(struct kvm_vcpu *vcpu, struct kvm_vcpu_transfer *vcpu_state);
+ int (*import_vcpu)(struct kvm_vcpu *vcpu, struct kvm_vcpu_transfer *vcpu_state);
};
struct kvm_x86_nested_ops {
diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c
index 8a99c665008a3..e8385326894b1 100644
--- a/arch/x86/kvm/x86.c
+++ b/arch/x86/kvm/x86.c
@@ -6189,6 +6189,38 @@ static int kvm_get_reg_list(struct kvm_vcpu *vcpu,
return 0;
}
+static int kvm_vcpu_ioctl_transfer_vcpu(struct kvm_vcpu *vcpu, bool import,
+ void __user *argp)
+{
+ struct kvm_vcpu_transfer vcpu_state;
+ struct kvm *kvm = vcpu->kvm;
+ int r;
+
+ if (!kvm_x86_call(cap_live_migration)(kvm) ||
+ (import && !kvm_x86_ops.import_vcpu) ||
+ (!import && !kvm_x86_ops.export_vcpu))
+ return -ENOTTY;
+
+ if (copy_from_user(&vcpu_state, argp, sizeof(vcpu_state)))
+ return -EFAULT;
+
+ if (vcpu_state.reserved || vcpu_state.buf.reserved)
+ return -EINVAL;
+
+ if (import)
+ r = kvm_x86_call(import_vcpu)(vcpu, &vcpu_state);
+ else
+ r = kvm_x86_call(export_vcpu)(vcpu, &vcpu_state);
+ if (r > 0)
+ r = -EIO;
+
+ /* Copy back also on an error to report a partially done transfer */
+ if (copy_to_user(argp, &vcpu_state, sizeof(vcpu_state)))
+ r = -EFAULT;
+
+ return r;
+}
+
long kvm_arch_vcpu_ioctl(struct file *filp,
unsigned int ioctl, unsigned long arg)
{
@@ -6659,6 +6691,14 @@ long kvm_arch_vcpu_ioctl(struct file *filp,
goto out;
r = kvm_x86_ops.vcpu_mem_enc_ioctl(vcpu, argp);
break;
+ case KVM_EXPORT_VCPU: {
+ r = kvm_vcpu_ioctl_transfer_vcpu(vcpu, false, argp);
+ break;
+ }
+ case KVM_IMPORT_VCPU: {
+ r = kvm_vcpu_ioctl_transfer_vcpu(vcpu, true, argp);
+ break;
+ }
default:
r = -EINVAL;
}
diff --git a/include/uapi/linux/kvm.h b/include/uapi/linux/kvm.h
index 666bbdf220d65..0a9aa126daadb 100644
--- a/include/uapi/linux/kvm.h
+++ b/include/uapi/linux/kvm.h
@@ -1496,6 +1496,8 @@ struct kvm_enc_region {
/* Available with KVM_CAP_LIVE_MIGRATION */
#define KVM_EXPORT_MEMORY _IOWR(KVMIO, 0xe5, struct kvm_memory_transfer)
#define KVM_IMPORT_MEMORY _IOWR(KVMIO, 0xe6, struct kvm_memory_transfer)
+#define KVM_EXPORT_VCPU _IOWR(KVMIO, 0xe7, struct kvm_vcpu_transfer)
+#define KVM_IMPORT_VCPU _IOWR(KVMIO, 0xe8, struct kvm_vcpu_transfer)
#define KVM_DIRTY_LOG_MANUAL_PROTECT_ENABLE (1 << 0)
#define KVM_DIRTY_LOG_INITIALLY_SET (1 << 1)
@@ -1705,4 +1707,10 @@ struct kvm_memory_transfer {
struct kvm_transfer_buffer buf;
};
+struct kvm_vcpu_transfer {
+ __u32 flags;
+ __u32 reserved;
+ struct kvm_transfer_buffer buf;
+};
+
#endif /* __LINUX_KVM_H */
--
2.43.0
^ permalink raw reply related [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests
2026-08-31 7:13 ` [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests Tony Lindgren
@ 2026-08-31 7:20 ` sashiko-bot
2026-09-18 11:35 ` Peter Xu
1 sibling, 0 replies; 78+ messages in thread
From: sashiko-bot @ 2026-08-31 7:20 UTC (permalink / raw)
To: Tony Lindgren; +Cc: kvm
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Low] The newly added documentation for the KVM live migration API contains architectural contradictions, backwards numbering, self-referencing links, and typographical errors.
--
commit 3a6c0c9c47ef3277e964b009e8dd06e7aeab932d
Author: Tony Lindgren <tony.lindgren@linux.intel.com>
Documentation: KVM: Add live migration API for confidential guests
For CoCo VMs, the guest memory and vCPU states are not accessible to the
userspace or KVM for live migration. The memory and vCPU states need to be
extracted into encrypted blobs on the source, and decrypted on the
destination. Before live migration, an encryption key needs to be
negotiated between the source and destination.
> diff --git a/Documentation/virt/kvm/api.rst b/Documentation/virt/kvm/api.rst
> index a5f9ee92f43e8..9d546d288af5f 100644
> --- a/Documentation/virt/kvm/api.rst
> +++ b/Documentation/virt/kvm/api.rst
[ ... ]
> +The transfer buffer size is vendor specific.
> +
> +For the transfer buffer, seeo :ref:`KVM_MIGRATE_CMD <KVM_MIGRATE_CMD>`.
[Severity: Low]
Is this "seeo" a typo for "see"?
> +
> +For memory import, see also :ref:`KVM_IMPORT_MEMORY <KVM_IMPORT_MEMORY>`.
> +
> +
> +.. _KVM_IMPORT_MEMORY:
> +
> +4.147 KVM_IMPORT_MEMORY
> +-----------------------
[ ... ]
> +The transfer buffer size is vendor specific.
> +
> +For kvm_memory_transfer, see :ref:`KVM_EXPORT_MEMORY <KVM_EXPORT_MEMORY>`.
> +
> +For the transfer buffer, seeo :ref:`KVM_MIGRATE_CMD <KVM_MIGRATE_CMD>`.
[Severity: Low]
Is there another "seeo" typo here?
> +
> +.. _KVM_EXPORT_VCPU:
> +
> +4.149 KVM_EXPORT_VCPU
> +---------------------
> +:Capability: KVM_CAP_LIVE_MIGRATION
> +:Architectures: arm64, x86
[Severity: Low]
Is it intentional that KVM_EXPORT_VCPU is numbered 4.149 before
KVM_IMPORT_VCPU which is numbered 4.148?
Also, KVM_EXPORT_VCPU lists arm64 as a supported architecture, but relies on
KVM_CAP_LIVE_MIGRATION which is documented as x86 only below. Should arm64 be
removed here, or added to the capability documentation?
> +:Type: vcpu ioctl
> +:Parameters: struct kvm_vcpu_transfer (in/out)
> +:Returns: 0 on success, < 0 on error
[ ... ]
> +.. _KVM_IMPORT_VCPU:
> +
> +4.148 KVM_IMPORT_VCPU
> +---------------------
> +
> +:Capability: KVM_CAP_LIVE_MIGRATION
> +:Architectures: arm64, x86
> +:Type: vcpu ioctl
> +:Parameters: struct kvm_vcpu_transfer (in/out)
> +:Returns: 0 on success, < 0 on error
> +
> +Allows userspace to request the host to import a VCPU state from a userspace
> +buffer.
> +
> +The VCPU state may not be directly accessible to KVM because of encryption. For
> +confidential computing, the VCPU state is encrypted and only accessible to the
> +guest.
> +
> +For vcpu_transfer and vCPU import, see :ref:`KVM_IMPORT_VCPU <KVM_IMPORT_VCPU>`.
[Severity: Low]
Does this circular reference point to itself? Should this point to
KVM_EXPORT_VCPU instead to help the reader find the struct kvm_vcpu_transfer
definition?
> +
> +For the transfer buffer, see also :ref:`KVM_MIGRATE_CMD <KVM_MIGRATE_CMD>`.
> +
> .. _kvm_run:
>
> 5. The kvm_run structure
> @@ -9493,6 +9687,17 @@ take care to differentiate between these cases.
> The presence of this capability indicates that the nested KVM guest can
> start in ESA mode.
>
> +8.48 KVM_CAP_LIVE_MIGRATION
> +---------------------------
> +
> +:Architectures: x86
> +:Target: VM
> +:Parameters: None
[Severity: Low]
The KVM_CAP_LIVE_MIGRATION capability is documented here as x86 only, but
KVM_EXPORT_VCPU and KVM_IMPORT_VCPU state they are for arm64 and x86. Is
this capability intended to support arm64 as well?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260831071304.762939-1-tony.lindgren@linux.intel.com?part=1
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-08-31 7:13 ` [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD Tony Lindgren
@ 2026-08-31 7:23 ` sashiko-bot
2026-09-01 6:03 ` Tony Lindgren
2026-09-07 11:53 ` Tony Lindgren
2026-09-18 4:33 ` Kishen Maloor
2 siblings, 1 reply; 78+ messages in thread
From: sashiko-bot @ 2026-08-31 7:23 UTC (permalink / raw)
To: Tony Lindgren; +Cc: kvm
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] The 'flags' field in the new 'struct kvm_migrate_cmd' UAPI is not validated to be zero.
--
commit b50bc10fe65ad9a54c9b979c445505e5baf76592
Author: Tony Lindgren <tony.lindgren@linux.intel.com>
KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
This patch introduces the KVM_CAP_LIVE_MIGRATION capability and the
KVM_MIGRATE_CMD ioctl. These additions allow KVM to configure hardware-specific
live migration functions, which is necessary for confidential guests like TDX.
> diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c
> index afcac1042947a..7064fd709e56d 100644
> --- a/arch/x86/kvm/x86.c
> +++ b/arch/x86/kvm/x86.c
[ ... ]
> @@ -7614,6 +7617,28 @@ int kvm_arch_vm_ioctl(struct file *filp, unsigned int ioctl, unsigned long arg)
> r = kvm_vm_ioctl_set_msr_filter(kvm, &filter);
> break;
> }
> + case KVM_MIGRATE_CMD: {
> + struct kvm_migrate_cmd cmd;
> +
> + if (!kvm_x86_ops.migrate_cmd ||
> + !kvm_x86_call(cap_live_migration)(kvm))
> + return -ENOTTY;
> +
> + if (copy_from_user(&cmd, argp, sizeof(cmd)))
> + return -EFAULT;
> +
> + if (cmd.reserved || cmd.buf.reserved)
> + return -EINVAL;
[Severity: High]
Does this code also need to validate that the new flags field is zero?
If uninitialized memory or garbage is passed by userspace in cmd.flags, it
will currently be silently ignored. When a future kernel introduces a new
flag, old userspace applications passing uninitialized flags will inadvertently
enable it, which could permanently break backwards compatibility.
> +
> + r = kvm_x86_call(migrate_cmd)(kvm, &cmd);
> + if (r > 0)
> + r = -EIO;
> +
> + /* Copy back also on an error to report a partially done command */
> + if (copy_to_user(argp, &cmd, sizeof(cmd)))
> + return -EFAULT;
> + break;
> + }
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260831071304.762939-1-tony.lindgren@linux.intel.com?part=2
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 3/4] KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY
2026-08-31 7:13 ` [RFC PATCH v2 3/4] KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY Tony Lindgren
@ 2026-08-31 7:23 ` sashiko-bot
2026-09-01 6:10 ` Tony Lindgren
0 siblings, 1 reply; 78+ messages in thread
From: sashiko-bot @ 2026-08-31 7:23 UTC (permalink / raw)
To: Tony Lindgren; +Cc: kvm
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] Unvalidated UAPI fields `flags` and `id` in `struct kvm_memory_transfer` can lead to ABI breakage.
--
commit b6f49e80a58eebaaa2a6d123950fd36709972a10
Author: Tony Lindgren <tony.lindgren@linux.intel.com>
KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY
This commit adds support to export and import KVM memory for cases where
the memory is only accessible to the guest. It introduces the optional
KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY ioctls to assist with live
migration of confidential computing.
> diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c
> index 7064fd709e56d..8a99c665008a3 100644
> --- a/arch/x86/kvm/x86.c
> +++ b/arch/x86/kvm/x86.c
[ ... ]
> @@ -7258,6 +7258,37 @@ long kvm_arch_vcpu_unlocked_ioctl(struct file *filp, unsigned int ioctl,
> return -ENOIOCTLCMD;
> }
>
> +static int kvm_vm_ioctl_transfer_memory(struct kvm *kvm, bool import,
> + void __user *argp)
> +{
> + struct kvm_memory_transfer mem;
> + int r;
> +
> + if (!kvm_x86_call(cap_live_migration)(kvm) ||
> + (import && !kvm_x86_ops.import_memory) ||
> + (!import && !kvm_x86_ops.export_memory))
> + return -ENOTTY;
> +
> + if (copy_from_user(&mem, argp, sizeof(mem)))
> + return -EFAULT;
> +
> + if (mem.reserved || mem.buf.reserved || !mem.nr_gfns)
> + return -EINVAL;
[Severity: Medium]
Does this input validation for the kvm_memory_transfer UAPI struct miss
checking mem.flags == 0 and mem.id == 0?
Failing to validate that unused UAPI fields are zero allows userspace to pass
uninitialized garbage. If KVM later assigns specific meaning to these fields,
older applications that unintentionally passed non-zero values could break,
leading to an ABI compatibility issue.
Standard Linux UAPI design usually requires enforcing these unused or reserved
fields to be zero for future extensibility.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260831071304.762939-1-tony.lindgren@linux.intel.com?part=3
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 4/4] KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU
2026-08-31 7:13 ` [RFC PATCH v2 4/4] KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU Tony Lindgren
@ 2026-08-31 7:23 ` sashiko-bot
2026-09-01 6:12 ` Tony Lindgren
0 siblings, 1 reply; 78+ messages in thread
From: sashiko-bot @ 2026-08-31 7:23 UTC (permalink / raw)
To: Tony Lindgren; +Cc: kvm
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] Missing validation of the flags field in UAPI struct kvm_vcpu_transfer breaks forward compatibility.
--
commit 4d7dccde0750ca7fd34158829ad4120d967ac2c3
Author: Tony Lindgren <tony.lindgren@linux.intel.com>
KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU
This commit adds support to export and import a VCPU for cases where the
VCPU state is only accessible to the guest. It introduces the new
KVM_EXPORT_VCPU and KVM_IMPORT_VCPU ioctls, along with the required
kvm_vcpu_transfer structure.
> diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c
> index 8a99c665008a3..e8385326894b1 100644
> --- a/arch/x86/kvm/x86.c
> +++ b/arch/x86/kvm/x86.c
[ ... ]
> @@ -6189,6 +6189,38 @@ static int kvm_get_reg_list(struct kvm_vcpu *vcpu,
> return 0;
> }
>
> +static int kvm_vcpu_ioctl_transfer_vcpu(struct kvm_vcpu *vcpu, bool import,
> + void __user *argp)
> +{
> + struct kvm_vcpu_transfer vcpu_state;
> + struct kvm *kvm = vcpu->kvm;
> + int r;
[ ... ]
> + if (copy_from_user(&vcpu_state, argp, sizeof(vcpu_state)))
> + return -EFAULT;
> +
> + if (vcpu_state.reserved || vcpu_state.buf.reserved)
> + return -EINVAL;
[Severity: High]
Should we also validate the flags field here to ensure forward compatibility?
The new kvm_vcpu_transfer structure introduces a flags field, but this
validation step only explicitly rejects non-zero reserved fields.
If the kernel silently ignores non-zero flags, userspace might
inadvertently pass uninitialized or arbitrary values without receiving
an error. If KVM later assigns meaning to these flags, old userspace
programs that have been unknowingly passing garbage could unexpectedly
trigger new behaviors or break.
Would it be appropriate to require that vcpu_state.flags is zero for now?
> +
> + if (import)
> + r = kvm_x86_call(import_vcpu)(vcpu, &vcpu_state);
[ ... ]
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260831071304.762939-1-tony.lindgren@linux.intel.com?part=4
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-08-31 7:23 ` sashiko-bot
@ 2026-09-01 6:03 ` Tony Lindgren
0 siblings, 0 replies; 78+ messages in thread
From: Tony Lindgren @ 2026-09-01 6:03 UTC (permalink / raw)
To: sashiko-reviews; +Cc: kvm
On Mon, Aug 31, 2026 at 07:23:11AM +0000, sashiko-bot@kernel.org wrote:
> Author: Tony Lindgren <tony.lindgren@linux.intel.com>
> > --- a/arch/x86/kvm/x86.c
> > +++ b/arch/x86/kvm/x86.c
> > @@ -7614,6 +7617,28 @@ int kvm_arch_vm_ioctl(struct file *filp, unsigned int ioctl, unsigned long arg)
> > r = kvm_vm_ioctl_set_msr_filter(kvm, &filter);
> > break;
> > }
> > + case KVM_MIGRATE_CMD: {
> > + struct kvm_migrate_cmd cmd;
> > +
> > + if (!kvm_x86_ops.migrate_cmd ||
> > + !kvm_x86_call(cap_live_migration)(kvm))
> > + return -ENOTTY;
> > +
> > + if (copy_from_user(&cmd, argp, sizeof(cmd)))
> > + return -EFAULT;
> > +
> > + if (cmd.reserved || cmd.buf.reserved)
> > + return -EINVAL;
>
> [Severity: High]
> Does this code also need to validate that the new flags field is zero?
>
> If uninitialized memory or garbage is passed by userspace in cmd.flags, it
> will currently be silently ignored. When a future kernel introduces a new
> flag, old userspace applications passing uninitialized flags will inadvertently
> enable it, which could permanently break backwards compatibility.
Could be zero, or we could check for the known flags.
We are currently using the following flags for KVM_MIGRATE_CMD
sub-commands for TDX:
1. KVM_MIGRATE_SETUP
====================
TDX_MIGRATE_SETUP_SESSION exchange the migration keys
TDX_MIGRATE_IMMUTABLE_STATE transfer the TDX immutable state
2. KVM_MIGRATE_ITERATION
========================
TDX_MIGRATE_IN_ORDER_DONE end the in-order migration
3. KVM_MIGRATE_STOP_AND_COPY
============================
TDX_MIGRATE_STOP_COPY_PAUSE pause the TD
TDX_MIGRATE_STOP_COPY_TD_STATE transfer the mutable TD state
4. KVM_MIGRATE_END
==================
TDX_MIGRATE_ABORT abort migration
Maybe all the above TDX specific flags could be turned into generic
flags for the sub-commands.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 3/4] KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY
2026-08-31 7:23 ` sashiko-bot
@ 2026-09-01 6:10 ` Tony Lindgren
0 siblings, 0 replies; 78+ messages in thread
From: Tony Lindgren @ 2026-09-01 6:10 UTC (permalink / raw)
To: sashiko-reviews; +Cc: kvm
On Mon, Aug 31, 2026 at 07:23:47AM +0000, sashiko-bot@kernel.org wrote:
> commit b6f49e80a58eebaaa2a6d123950fd36709972a10
> Author: Tony Lindgren <tony.lindgren@linux.intel.com>
> > --- a/arch/x86/kvm/x86.c
> > +++ b/arch/x86/kvm/x86.c
> [ ... ]
> > @@ -7258,6 +7258,37 @@ long kvm_arch_vcpu_unlocked_ioctl(struct file *filp, unsigned int ioctl,
> > return -ENOIOCTLCMD;
> > }
> >
> > +static int kvm_vm_ioctl_transfer_memory(struct kvm *kvm, bool import,
> > + void __user *argp)
> > +{
> > + struct kvm_memory_transfer mem;
> > + int r;
> > +
> > + if (!kvm_x86_call(cap_live_migration)(kvm) ||
> > + (import && !kvm_x86_ops.import_memory) ||
> > + (!import && !kvm_x86_ops.export_memory))
> > + return -ENOTTY;
> > +
> > + if (copy_from_user(&mem, argp, sizeof(mem)))
> > + return -EFAULT;
> > +
> > + if (mem.reserved || mem.buf.reserved || !mem.nr_gfns)
> > + return -EINVAL;
>
> [Severity: Medium]
> Does this input validation for the kvm_memory_transfer UAPI struct miss
> checking mem.flags == 0 and mem.id == 0?
Yes flags is unused for memory transfers at least for TDX. The id TDX can
eventually use for the thread index for QEMU multifd transfers.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 4/4] KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU
2026-08-31 7:23 ` sashiko-bot
@ 2026-09-01 6:12 ` Tony Lindgren
0 siblings, 0 replies; 78+ messages in thread
From: Tony Lindgren @ 2026-09-01 6:12 UTC (permalink / raw)
To: sashiko-reviews; +Cc: kvm
On Mon, Aug 31, 2026 at 07:23:53AM +0000, sashiko-bot@kernel.org wrote:
> commit 4d7dccde0750ca7fd34158829ad4120d967ac2c3
> Author: Tony Lindgren <tony.lindgren@linux.intel.com>
> > --- a/arch/x86/kvm/x86.c
> > +++ b/arch/x86/kvm/x86.c
>
> [ ... ]
>
> > @@ -6189,6 +6189,38 @@ static int kvm_get_reg_list(struct kvm_vcpu *vcpu,
> > return 0;
> > }
> >
> > +static int kvm_vcpu_ioctl_transfer_vcpu(struct kvm_vcpu *vcpu, bool import,
> > + void __user *argp)
> > +{
> > + struct kvm_vcpu_transfer vcpu_state;
> > + struct kvm *kvm = vcpu->kvm;
> > + int r;
>
> [ ... ]
>
> > + if (copy_from_user(&vcpu_state, argp, sizeof(vcpu_state)))
> > + return -EFAULT;
> > +
> > + if (vcpu_state.reserved || vcpu_state.buf.reserved)
> > + return -EINVAL;
>
> [Severity: High]
> Should we also validate the flags field here to ensure forward compatibility?
>
> The new kvm_vcpu_transfer structure introduces a flags field, but this
> validation step only explicitly rejects non-zero reserved fields.
>
> If the kernel silently ignores non-zero flags, userspace might
> inadvertently pass uninitialized or arbitrary values without receiving
> an error. If KVM later assigns meaning to these flags, old userspace
> programs that have been unknowingly passing garbage could unexpectedly
> trigger new behaviors or break.
>
> Would it be appropriate to require that vcpu_state.flags is zero for now?
Yes flags is unused for vCPU transfers at least for TDX and can be zero
for now.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
2026-08-31 7:13 [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Tony Lindgren
` (3 preceding siblings ...)
2026-08-31 7:13 ` [RFC PATCH v2 4/4] KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU Tony Lindgren
@ 2026-09-04 18:24 ` Artem Bityutskiy
2026-09-17 21:27 ` Peter Xu
2026-09-18 18:36 ` Ionut Mihalcea
2026-09-25 16:03 ` Serge Hallyn (AMD)
6 siblings, 1 reply; 78+ messages in thread
From: Artem Bityutskiy @ 2026-09-04 18:24 UTC (permalink / raw)
To: Tony Lindgren, Paolo Bonzini, Sean Christopherson
Cc: Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky,
Marc Zyngier, Oliver Upton, Steven Price, Anup Patel,
Samuel Ortiz, Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
On Mon, 2026-08-31 at 10:13 +0300, Tony Lindgren wrote:
> Tom, since you mentioned that AMD SEV-SNP and Intel TDX live migration
> sound similar, can you please take a look how the API might work for
> SEV-SNP?
The presented example uAPIs were designed to fit the TDX live migration
flow, with the intent that they could also be used by other CoCo VMs.
It would be super great if someone could commend on how suitable these
uAPIs are for AMD and ARM flows.
AFAIU, what is generic in the presented uAPI is more of usage pattern.
- The order in which QEMU calls them.
- The idea that QEMU/KVM is a transport for opaque blobs.
The proposed container - 'struct kvm_transfer_buffer' - only has address
and a size. The vendor defines the layout of the data in it.
Some example high-level topics that would be nice to get feedback on:
- Can we come up with a single set of generic migration uAPIs for different
CoCo models?
- Or should some uAPIs be generic while others are vendor-specific?
- Or should each CoCo model have its own vendor-specific set of migration
uAPIs?
- Should the same uAPIs also support traditional VMs? But the only use-case
I imagine here is "for testing purposes".
> Artem has put together a brief description below of the example API and the
> migration flow:
.. snip ...
> Migration flow
> ==============
>
> Source host Destination host
> =========== ================
>
> CMD(SETUP/SESSION) <--- setup msgs ---> CMD(SETUP/SESSION)
> | (repeated) |
> CMD(SETUP/IMMUTABLE_STATE) - immutable state -> CMD(SETUP/IMMUTABLE_STATE)
> | |
> KVM_GET_DIRTY_LOG |
> KVM_EXPORT_MEMORY --- memory data ---> KVM_IMPORT_MEMORY
> CMD(ITERATION) --- epoch token ---> CMD(ITERATION)
> | (repeat until convergence) |
> CMD(STOP_AND_COPY/PAUSE) |
> CMD(STOP_AND_COPY/TD_STATE) --- VM state ------> CMD(STOP_AND_COPY/TD_STATE)
> KVM_EXPORT_VCPU --- vCPU state ----> KVM_IMPORT_VCPU
> KVM_EXPORT_MEMORY -- final memory ---> KVM_IMPORT_MEMORY
> CMD(ITERATION/DONE) --- start token ---> CMD(ITERATION)
> | |
> CMD(END) CMD(END)
... snip ...
As I mentioned, the example uAPI is modeled around the TDX migration flow.
In case it helps the reader, here is a summary of that flow that was
presented in PUCK.
It describes what the TDX module offers today and focuses on pre-copy
migration. This is our interpretation of the TDX specifications, not a
specification itself, and may contain errors or omissions. Please refer to
the official TDX specifications for authoritative information.
Migration Overview
------------------
In TDX, the TDX module implements the migration logic. The VMM drives
migration by issuing seamcalls to the source and destination TDX modules and
transports the resulting encrypted blobs between them. The VMM does not need
to know the contents of those blobs.
The TDX security model enforces two hard rules:
- Only one instance of the TD may run at a time. Either the source or the
destination may run, but never both. In other words, cloning a TD is not
allowed.
- When migration completes, the destination must have the same memory and
vCPU state as the source. It must not end up with a partial or mixed
state.
Simple Overview
---------------
Run the migration setup session
|
v
Transfer immutable TD state
|
v
Copy dirty memory pages while the TD runs <-------+
| |
+-- iterate until convergence criteria is met --+
|
v
Pause the source TD and copy the remaining state
|
v
Start the destination TD
The migration process starts with a setup session. During the setup session,
the source and destination TDX modules exchange encrypted blobs and
establish migration encryption keys. For example, these keys protect memory
contents transferred during migration.
Next, the source transfers the immutable TD state to the destination and
initializes the destination TD. This state includes the read-only VM data,
such as its vCPU count and topology.
The source and destination then perform iterative memory-copy rounds. In
each round, the source scans for memory pages that need to be migrated and
exports them in encrypted form. The destination imports the pages. The
rounds continue until the number of pages that change between rounds is
small enough to meet the convergence criteria.
The source TD is then paused. The VMM transfers the remaining TD state, such
as vCPU state, and the final dirty memory pages. Finally, the destination TD
is started and the migration completes.
Notice that the VMM's role in this process is to issue the required
seamcalls and send encrypted blobs between the source and destination. The
VMM does not need to know the contents of those blobs.
More Detailed Overview
----------------------
The following diagram shows the TDX seamcalls used by the source and
destination:
Source host Destination host
=========== ================
TDH.MIG.SETUP -- crypto keys, attestation -> TDH.MIG.SETUP
| |
TDH.EXPORT.STATE.IMMUTABLE -- read-only TD state ->
TDH.IMPORT.STATE.IMMUTABLE
| |
TDH.MEM.SCAN.RANGE |
TDH.MEM.TRACK + IPIs |
TDH.EXPORT.MEM ------ memory data ---------> TDH.IMPORT.MEM
TDH.EXPORT.TRACK ------ epoch token ---------> TDH.IMPORT.TRACK
| |
TDH.EXPORT.PAUSE |
| |
TDH.EXPORT.STATE.TD ------ global TD state -----> TDH.IMPORT.STATE.TD
| |
TDH.EXPORT.STATE.VP ------ vCPU state ----------> TDH.IMPORT.STATE.VP
| |
TDH.MEM.SCAN.COMP |
TDH.MEM.TRACK + IPIs |
TDH.EXPORT.MEM ------ final memory data ---> TDH.IMPORT.MEM
TDH.EXPORT.TRACK ------ start token ---------> TDH.IMPORT.TRACK
|
TDH.IMPORT.END
The source and destination go through the following stages.
Setup
-----
The VMM issues TDH.MIG.SETUP on the source and destination TDX modules
iteratively to perform the setup session. The modules return status and may
also return an encrypted blob, which the VMM passes between the two sides.
In other words, the migration protocol is between the two TDX modules, while
the VMM is simply the transport mechanism. The setup session performs mutual
attestation, establishes trust between the modules, establishes migration
encryption keys, and loads the migration policy. It completes when the TDX
module returns success.
Immutable State Transfer
------------------------
The source calls TDH.EXPORT.STATE.IMMUTABLE to export the immutable TD
state. The destination VMM creates the destination TD skeleton and
configures migration streams, then calls TDH.IMPORT.STATE.IMMUTABLE, which
finalizes the destination TD initialization.
Iterative Memory Copy
---------------------
While the source TD continues running, the source and destination run
iterative memory copy rounds:
- The source calls TDH.MEM.SCAN.RANGE to find migration candidate pages.
- The source calls TDH.MEM.TRACK and sends IPIs to the TD's vCPUs in order
to ensure the dirty page scanning algorithm correctness.
- The source calls TDH.EXPORT.MEM to export private memory pages, which the
VMM sends to the destination.
- The destination calls TDH.IMPORT.MEM to import the received memory pages.
- The source calls TDH.EXPORT.TRACK to generate an epoch token. The VMM
sends it to the destination, which calls TDH.IMPORT.TRACK to consume the
token. This verifies that all data exported from the source was imported
on the destination.
The rounds continue until the number of pages that change between rounds is
small enough to meet the convergence criteria.
Stop and Copy
-------------
The VMM pauses the source TD, which begins the downtime period, then calls
TDH.EXPORT.PAUSE to start the TDX-enforced blackout period. The source then
calls TDH.EXPORT.STATE.TD to export mutable TD-scope state and
TDH.EXPORT.STATE.VP to export mutable state for each vCPU. It then performs
the final dirty page scan, runs TDH.MEM.TRACK and sends IPIs, and calls
TDH.EXPORT.MEM to export the final dirty pages as encrypted blobs.
The destination calls TDH.IMPORT.STATE.TD to import mutable TD-scope state
and TDH.IMPORT.STATE.VP to import the mutable state of each vCPU. It calls
TDH.IMPORT.MEM to import the final dirty pages.
After exporting the final dirty pages, the source's last migration seamcall
is TDH.EXPORT.TRACK with IN_ORDER_DONE=1, which generates the start token.
The VMM sends the start token to the destination. The destination calls
TDH.IMPORT.TRACK with this token. This verifies that the mutable TD state
has been imported and allows the destination to start the TD. Finally, the
destination calls TDH.IMPORT.END, which ends the migration.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-08-31 7:13 ` [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD Tony Lindgren
2026-08-31 7:23 ` sashiko-bot
@ 2026-09-07 11:53 ` Tony Lindgren
2026-09-07 13:15 ` Jörg Rödel
2026-09-18 4:33 ` Kishen Maloor
2 siblings, 1 reply; 78+ messages in thread
From: Tony Lindgren @ 2026-09-07 11:53 UTC (permalink / raw)
To: Paolo Bonzini, Sean Christopherson
Cc: Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
On Mon, Aug 31, 2026 at 10:13:02AM +0300, Tony Lindgren wrote:
> --- a/include/uapi/linux/kvm.h
> +++ b/include/uapi/linux/kvm.h
> @@ -1670,4 +1673,23 @@ struct kvm_pre_fault_memory {
> __u64 padding[5];
> };
>
> +#define KVM_MIGRATE_SETUP 0
> +#define KVM_MIGRATE_ITERATION 1
> +#define KVM_MIGRATE_STOP_AND_COPY 2
> +#define KVM_MIGRATE_ABORT 3
> +#define KVM_MIGRATE_END 4
> +
> +struct kvm_transfer_buffer {
> + __u64 address;
> + __u32 size;
> + __u32 reserved;
> +};
> +
> +struct kvm_migrate_cmd {
> + __u16 command;
> + __u16 flags;
> + __u32 reserved;
> + struct kvm_transfer_buffer buf;
> +};
For the common flags, KVM_MIGRATE_CMD probably should have migration
direction. Or maybe we could have KVM_EXPORT_CMD and KVM_IMPORT_CMD.
The migration direction is needed early for TDX. We currently pass a flag
for delayed init for the migration destination in KVM_TDX_INIT_VM. This
is to prevent the TD and vCPU init SEAMCALLs on the destination. The TD
and vCPU are initialized only later on with the import calls.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-07 11:53 ` Tony Lindgren
@ 2026-09-07 13:15 ` Jörg Rödel
2026-09-07 13:32 ` Artem Bityutskiy
0 siblings, 1 reply; 78+ messages in thread
From: Jörg Rödel @ 2026-09-07 13:15 UTC (permalink / raw)
To: Tony Lindgren
Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Artem Bityutskiy,
Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky,
Marc Zyngier, Oliver Upton, Steven Price, Anup Patel,
Samuel Ortiz, Jakub Růžička, Vishal Annapurve,
Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg,
Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm
On Mon, Sep 07, 2026 at 02:53:05PM +0300, Tony Lindgren wrote:
> On Mon, Aug 31, 2026 at 10:13:02AM +0300, Tony Lindgren wrote:
> > --- a/include/uapi/linux/kvm.h
> > +++ b/include/uapi/linux/kvm.h
> > @@ -1670,4 +1673,23 @@ struct kvm_pre_fault_memory {
> > __u64 padding[5];
> > };
> >
> > +#define KVM_MIGRATE_SETUP 0
> > +#define KVM_MIGRATE_ITERATION 1
> > +#define KVM_MIGRATE_STOP_AND_COPY 2
> > +#define KVM_MIGRATE_ABORT 3
> > +#define KVM_MIGRATE_END 4
> > +
> > +struct kvm_transfer_buffer {
> > + __u64 address;
> > + __u32 size;
> > + __u32 reserved;
> > +};
> > +
> > +struct kvm_migrate_cmd {
> > + __u16 command;
> > + __u16 flags;
> > + __u32 reserved;
> > + struct kvm_transfer_buffer buf;
> > +};
>
> For the common flags, KVM_MIGRATE_CMD probably should have migration
> direction. Or maybe we could have KVM_EXPORT_CMD and KVM_IMPORT_CMD.
The direction is always the same over a single live migration session, right?
So it could be a setup flag, on the other hand having separate KVM_EXPORT_CMD
and KVM_IMPORT_CMD seems to be a cleaner ABI.
-Joerg
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-07 13:15 ` Jörg Rödel
@ 2026-09-07 13:32 ` Artem Bityutskiy
2026-09-08 4:15 ` Tony Lindgren
2026-09-08 4:43 ` Tony Lindgren
0 siblings, 2 replies; 78+ messages in thread
From: Artem Bityutskiy @ 2026-09-07 13:32 UTC (permalink / raw)
To: Jörg Rödel, Tony Lindgren
Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Fabiano Rosas,
Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Vishal Annapurve, Elena Reshetova,
Kai Huang, Kishen Maloor, Mika Westerberg, Peter Fang,
Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm
On Mon, 2026-09-07 at 15:15 +0200, Jörg Rödel wrote:
> The direction is always the same over a single live migration session, right?
> So it could be a setup flag, on the other hand having separate KVM_EXPORT_CMD
> and KVM_IMPORT_CMD seems to be a cleaner ABI.
Yes, the source stays the source, and the destination stays the destination
for the entire session.
This is a special case of a broader question I keep coming back to:
** How much state should KVM keep about a migration session? **
For direction specifically, we could pass it to KVM on every call and let
KVM stay stateless about it, or KVM could record it once and remember it for
the rest of the session.
But in general. And this is addressed not just to Jörg, but community.
Traditional VM migration is driven by QEMU. KVM provides building blocks
such as dirty page tracking and vCPU state get/set APIs, but it does not
track the overall migration session. The migration session state lives in
QEMU.
Our TDX live migration prototype keeps some per-migration state in KVM,
for example the direction, the migration phase (setup done, started,
paused, and so on). This lets use validate inputs and issue the correct TDX
module seamcalls from KVM.
A different uAPI could shift this balance either way.
In the **extreme** case, we could expose a uAPI for each migration
seamcall. KVM would then just pass inputs and outputs between the TDX
module and QEMU, staying a thin layer with no per-session state. We have
not tried this, but it illustrates the trade-off.
So where is the right boundary for migration state in KVM?
Should KVM manage none of it, or is some state acceptable?
Would be interesting to know what KVM community thinks on this.
Thanks,
Artem.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-07 13:32 ` Artem Bityutskiy
@ 2026-09-08 4:15 ` Tony Lindgren
2026-09-08 4:43 ` Tony Lindgren
1 sibling, 0 replies; 78+ messages in thread
From: Tony Lindgren @ 2026-09-08 4:15 UTC (permalink / raw)
To: Artem Bityutskiy
Cc: Jörg Rödel, Paolo Bonzini, Sean Christopherson,
Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky,
Marc Zyngier, Oliver Upton, Steven Price, Anup Patel,
Samuel Ortiz, Jakub Růžička, Vishal Annapurve,
Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg,
Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm
On Mon, Sep 07, 2026 at 04:32:33PM +0300, Artem Bityutskiy wrote:
> Our TDX live migration prototype keeps some per-migration state in KVM,
> for example the direction, the migration phase (setup done, started,
> paused, and so on). This lets use validate inputs and issue the correct TDX
> module seamcalls from KVM.
Yeah for TDX, we currently keep track of some of the TDX module state for
migration. For most part it can be done with the existing kvm_tdx->state.
The paused state is additional TD_STATE_PAUSED.
Some states are trickier though, the setup done state means the TDX module
has migration keys configured. If the keys are not installed, the
migration SEAMCALLs return errors. Does the kernel need to keep track of
this? With proper errors returned, maybe not.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-07 13:32 ` Artem Bityutskiy
2026-09-08 4:15 ` Tony Lindgren
@ 2026-09-08 4:43 ` Tony Lindgren
2026-09-09 0:22 ` Kishen Maloor
1 sibling, 1 reply; 78+ messages in thread
From: Tony Lindgren @ 2026-09-08 4:43 UTC (permalink / raw)
To: Artem Bityutskiy
Cc: Jörg Rödel, Paolo Bonzini, Sean Christopherson,
Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky,
Marc Zyngier, Oliver Upton, Steven Price, Anup Patel,
Samuel Ortiz, Jakub Růžička, Vishal Annapurve,
Elena Reshetova, Kai Huang, Kishen Maloor, Mika Westerberg,
Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm
On Mon, Sep 07, 2026 at 04:32:33PM +0300, Artem Bityutskiy wrote:
> On Mon, 2026-09-07 at 15:15 +0200, Jörg Rödel wrote:
> > The direction is always the same over a single live migration session, right?
> > So it could be a setup flag, on the other hand having separate KVM_EXPORT_CMD
> > and KVM_IMPORT_CMD seems to be a cleaner ABI.
OK
> Yes, the source stays the source, and the destination stays the destination
> for the entire session.
>
> This is a special case of a broader question I keep coming back to:
>
> ** How much state should KVM keep about a migration session? **
>
> For direction specifically, we could pass it to KVM on every call and let
> KVM stay stateless about it, or KVM could record it once and remember it for
> the rest of the session.
Yes the "record and remember" is another option, it could be a sub-command
something like KVM_MIGRATE_DIRECTION.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-08 4:43 ` Tony Lindgren
@ 2026-09-09 0:22 ` Kishen Maloor
2026-09-09 6:57 ` Tony Lindgren
0 siblings, 1 reply; 78+ messages in thread
From: Kishen Maloor @ 2026-09-09 0:22 UTC (permalink / raw)
To: Tony Lindgren, Artem Bityutskiy
Cc: Jörg Rödel, Paolo Bonzini, Sean Christopherson,
Peter Xu, Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky,
Marc Zyngier, Oliver Upton, Steven Price, Anup Patel,
Samuel Ortiz, Jakub Růžička, Vishal Annapurve,
Elena Reshetova, Kai Huang, Mika Westerberg, Peter Fang,
Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm
On 9/7/26 9:43 PM, Tony Lindgren wrote:
> On Mon, Sep 07, 2026 at 04:32:33PM +0300, Artem Bityutskiy wrote:
>> On Mon, 2026-09-07 at 15:15 +0200, Jörg Rödel wrote:
>>> The direction is always the same over a single live migration session, right?
>>> So it could be a setup flag, on the other hand having separate KVM_EXPORT_CMD
>>> and KVM_IMPORT_CMD seems to be a cleaner ABI.
>
> OK
Just sharing an alternate point of view:
The roles are fixed over a migration session. A split ABI is certainly more
self-describing, but it restates that invariant on every call, and therefore
also permits it to be contradicted -- a failure mode that does not otherwise
exist. With a single KVM_MIGRATE_CMD and a per-session role recorded once (more
on that below), there is no need for per-call policing: the role could be
checked once when the session is established.
Along these lines: it raises a question of whether the MEMORY and VCPU
calls should be coalesced as well into KVM_MIGRATE_MEMORY and KVM_MIGRATE_VCPU.
As posted, direction is implicit for KVM_MIGRATE_CMD but encoded in the ioctl
number for those transfers, so collapsing them would at least make the uAPI
consistent about where direction comes from, and free two ioctls.
>
>> ...
>>
>> For direction specifically, we could pass it to KVM on every call and let
>> KVM stay stateless about it, or KVM could record it once and remember it for
>> the rest of the session.
>
> Yes the "record and remember" is another option, it could be a sub-command
> something like KVM_MIGRATE_DIRECTION.
Agreed on record-and-remember, with a refinement on scope.
A VM that was migrated in can later be migrated out, so the role is not a
property of the VM -- it has to be recorded per session. SETUP is the call
that starts a migration and runs ahead of all other migration calls, so its
arguments look like the natural place for userspace to state the role; a
separate sub-command would need its own scope rules relative to SETUP.
KVM can then ask the vendor layer whether the requested role is permitted for
this VM -- only it knows the confidential-VM state -- and on success record it
in generic KVM state, where it then selects the export or import callbacks.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-09 0:22 ` Kishen Maloor
@ 2026-09-09 6:57 ` Tony Lindgren
2026-09-10 1:11 ` Kishen Maloor
0 siblings, 1 reply; 78+ messages in thread
From: Tony Lindgren @ 2026-09-09 6:57 UTC (permalink / raw)
To: Kishen Maloor
Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini,
Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On Tue, Sep 08, 2026 at 05:22:30PM -0700, Kishen Maloor wrote:
> On 9/7/26 9:43 PM, Tony Lindgren wrote:
> > On Mon, Sep 07, 2026 at 04:32:33PM +0300, Artem Bityutskiy wrote:
> >> On Mon, 2026-09-07 at 15:15 +0200, Jörg Rödel wrote:
> >>> The direction is always the same over a single live migration session, right?
> >>> So it could be a setup flag, on the other hand having separate KVM_EXPORT_CMD
> >>> and KVM_IMPORT_CMD seems to be a cleaner ABI.
> >
> > OK
>
>
> Just sharing an alternate point of view:
>
> The roles are fixed over a migration session. A split ABI is certainly more
> self-describing, but it restates that invariant on every call, and therefore
> also permits it to be contradicted -- a failure mode that does not otherwise
> exist. With a single KVM_MIGRATE_CMD and a per-session role recorded once (more
> on that below), there is no need for per-call policing: the role could be
> checked once when the session is established.
There are two occasions the role is set or changed. On starting the
destination the incoming role needs to configured at least for TDX.
And then after the migration, the role changes if re-migrated.
I don't think there are other cases for role change, maybe cancelled
migration could require that for some hardware possibly.
> Along these lines: it raises a question of whether the MEMORY and VCPU
> calls should be coalesced as well into KVM_MIGRATE_MEMORY and KVM_MIGRATE_VCPU.
> As posted, direction is implicit for KVM_MIGRATE_CMD but encoded in the ioctl
> number for those transfers, so collapsing them would at least make the uAPI
> consistent about where direction comes from, and free two ioctls.
Using naming KVM_TRANSFER_MEMORY and KVM_TRANSFER_VCPU might be more
descriptive?
Eventually these same commands could be used to save the state to disk
for power management use.
And going back to the dmaengine like analogy of what is being done..
The transfer direction flags could be KVM_TRANSFER_FROM_GUEST and
KVM_TRANSFER_TO_GUEST?
> >> For direction specifically, we could pass it to KVM on every call and let
> >> KVM stay stateless about it, or KVM could record it once and remember it for
> >> the rest of the session.
> >
> > Yes the "record and remember" is another option, it could be a sub-command
> > something like KVM_MIGRATE_DIRECTION.
>
>
> Agreed on record-and-remember, with a refinement on scope.
>
> A VM that was migrated in can later be migrated out, so the role is not a
> property of the VM -- it has to be recorded per session. SETUP is the call
> that starts a migration and runs ahead of all other migration calls, so its
> arguments look like the natural place for userspace to state the role; a
> separate sub-command would need its own scope rules relative to SETUP.
> KVM can then ask the vendor layer whether the requested role is permitted for
> this VM -- only it knows the confidential-VM state -- and on success record it
> in generic KVM state, where it then selects the export or import callbacks.
The role can change, but it can be VM specific for starting the migration
destination even before migration is started.
At least for TDX we need to specify direction for migration destination on
init to prevent fully initializing the TD and vCPUs.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-09 6:57 ` Tony Lindgren
@ 2026-09-10 1:11 ` Kishen Maloor
2026-09-10 6:33 ` Tony Lindgren
0 siblings, 1 reply; 78+ messages in thread
From: Kishen Maloor @ 2026-09-10 1:11 UTC (permalink / raw)
To: Tony Lindgren
Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini,
Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On 9/8/26 11:57 PM, Tony Lindgren wrote:
> On Tue, Sep 08, 2026 at 05:22:30PM -0700, Kishen Maloor wrote:
>> On 9/7/26 9:43 PM, Tony Lindgren wrote:
>>> On Mon, Sep 07, 2026 at 04:32:33PM +0300, Artem Bityutskiy wrote:
>>>> On Mon, 2026-09-07 at 15:15 +0200, Jörg Rödel wrote:
>>>>> The direction is always the same over a single live migration session, right?
>>>>> So it could be a setup flag, on the other hand having separate KVM_EXPORT_CMD
>>>>> and KVM_IMPORT_CMD seems to be a cleaner ABI.
>>>
>>> OK
>>
>>
>> Just sharing an alternate point of view:
>>
>> The roles are fixed over a migration session. A split ABI is certainly more
>> self-describing, but it restates that invariant on every call, and therefore
>> also permits it to be contradicted -- a failure mode that does not otherwise
>> exist. With a single KVM_MIGRATE_CMD and a per-session role recorded once (more
>> on that below), there is no need for per-call policing: the role could be
>> checked once when the session is established.
>
> There are two occasions the role is set or changed. On starting the
> destination the incoming role needs to configured at least for TDX.
> And then after the migration, the role changes if re-migrated.
A destination TD needs a directive to not initialize the TD and its vCPUs.
It comes from its launch parameters (e.g. QEMU cmdline) which selects the
delayed_init path. That is a construction directive though, and it applies
only to a destination -- a source needs nothing at init. A session role is
symmetric and is what the transfer calls consume, so I don't think the two
need to be the same thing.
> I don't think there are other cases for role change, maybe cancelled
> migration could require that for some hardware possibly.
Architecturally, per-session scoping of migration roles should be
straightforward with any platform: each side asserts a role at the start of every
migration session and vendor code will either accept or reject the stated role.
>> Along these lines: it raises a question of whether the MEMORY and VCPU
>> calls should be coalesced as well into KVM_MIGRATE_MEMORY and KVM_MIGRATE_VCPU.
>> As posted, direction is implicit for KVM_MIGRATE_CMD but encoded in the ioctl
>> number for those transfers, so collapsing them would at least make the uAPI
>> consistent about where direction comes from, and free two ioctls.
>
> Using naming KVM_TRANSFER_MEMORY and KVM_TRANSFER_VCPU might be more
> descriptive?
>
> Eventually these same commands could be used to save the state to disk
> for power management use.
Sure, and the save-to-disk case is a good argument for a more generic name.
> And going back to the dmaengine like analogy of what is being done..
>
> The transfer direction flags could be KVM_TRANSFER_FROM_GUEST and
> KVM_TRANSFER_TO_GUEST?
FROM_GUEST/TO_GUEST still encodes direction per call, which is the open
question above. If direction is a per-session property, then KVM_TRANSFER_MEMORY
and KVM_TRANSFER_VCPU are sufficient on their own -- no direction flag, and no
separate export/import ioctls.
And if we settle on a per-session property, then SETUP could conceivably state a
role for a non-migration transfer session as well.
>
>>>> For direction specifically, we could pass it to KVM on every call and let
>>>> KVM stay stateless about it, or KVM could record it once and remember it for
>>>> the rest of the session.
>>>
>>> Yes the "record and remember" is another option, it could be a sub-command
>>> something like KVM_MIGRATE_DIRECTION.
>>
>>
>> Agreed on record-and-remember, with a refinement on scope.
>>
>> A VM that was migrated in can later be migrated out, so the role is not a
>> property of the VM -- it has to be recorded per session. SETUP is the call
>> that starts a migration and runs ahead of all other migration calls, so its
>> arguments look like the natural place for userspace to state the role; a
>> separate sub-command would need its own scope rules relative to SETUP.
>> KVM can then ask the vendor layer whether the requested role is permitted for
>> this VM -- only it knows the confidential-VM state -- and on success record it
>> in generic KVM state, where it then selects the export or import callbacks.
>
> The role can change, but it can be VM specific for starting the migration
> destination even before migration is started.
>
> At least for TDX we need to specify direction for migration destination on
> init to prevent fully initializing the TD and vCPUs.
As mentioned above, that is a destination launch time directive that we needn't
conflate with a migration/transfer session role.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-10 1:11 ` Kishen Maloor
@ 2026-09-10 6:33 ` Tony Lindgren
2026-09-11 1:40 ` Kishen Maloor
0 siblings, 1 reply; 78+ messages in thread
From: Tony Lindgren @ 2026-09-10 6:33 UTC (permalink / raw)
To: Kishen Maloor
Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini,
Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On Wed, Sep 09, 2026 at 06:11:11PM -0700, Kishen Maloor wrote:
> On 9/8/26 11:57 PM, Tony Lindgren wrote:
> > On Tue, Sep 08, 2026 at 05:22:30PM -0700, Kishen Maloor wrote:
> >> On 9/7/26 9:43 PM, Tony Lindgren wrote:
> >>> On Mon, Sep 07, 2026 at 04:32:33PM +0300, Artem Bityutskiy wrote:
> >>>> On Mon, 2026-09-07 at 15:15 +0200, Jörg Rödel wrote:
> >>>>> The direction is always the same over a single live migration session, right?
> >>>>> So it could be a setup flag, on the other hand having separate KVM_EXPORT_CMD
> >>>>> and KVM_IMPORT_CMD seems to be a cleaner ABI.
> >>>
> >>> OK
> >>
> >>
> >> Just sharing an alternate point of view:
> >>
> >> The roles are fixed over a migration session. A split ABI is certainly more
> >> self-describing, but it restates that invariant on every call, and therefore
> >> also permits it to be contradicted -- a failure mode that does not otherwise
> >> exist. With a single KVM_MIGRATE_CMD and a per-session role recorded once (more
> >> on that below), there is no need for per-call policing: the role could be
> >> checked once when the session is established.
> >
> > There are two occasions the role is set or changed. On starting the
> > destination the incoming role needs to configured at least for TDX.
> > And then after the migration, the role changes if re-migrated.
>
> A destination TD needs a directive to not initialize the TD and its vCPUs.
> It comes from its launch parameters (e.g. QEMU cmdline) which selects the
> delayed_init path. That is a construction directive though, and it applies
> only to a destination -- a source needs nothing at init. A session role is
> symmetric and is what the transfer calls consume, so I don't think the two
> need to be the same thing.
Yes the source vs destination role is there from the start for sure. And
changes on re-migration. Could be set in different ways.
> > I don't think there are other cases for role change, maybe cancelled
> > migration could require that for some hardware possibly.
>
> Architecturally, per-session scoping of migration roles should be
> straightforward with any platform: each side asserts a role at the start of every
> migration session and vendor code will either accept or reject the stated role.
Agreed.
> >> Along these lines: it raises a question of whether the MEMORY and VCPU
> >> calls should be coalesced as well into KVM_MIGRATE_MEMORY and KVM_MIGRATE_VCPU.
> >> As posted, direction is implicit for KVM_MIGRATE_CMD but encoded in the ioctl
> >> number for those transfers, so collapsing them would at least make the uAPI
> >> consistent about where direction comes from, and free two ioctls.
> >
> > Using naming KVM_TRANSFER_MEMORY and KVM_TRANSFER_VCPU might be more
> > descriptive?
> >
> > Eventually these same commands could be used to save the state to disk
> > for power management use.
>
> Sure, and the save-to-disk case is a good argument for a more generic name.
>
> > And going back to the dmaengine like analogy of what is being done..
> >
> > The transfer direction flags could be KVM_TRANSFER_FROM_GUEST and
> > KVM_TRANSFER_TO_GUEST?
>
> FROM_GUEST/TO_GUEST still encodes direction per call, which is the open
> question above. If direction is a per-session property, then KVM_TRANSFER_MEMORY
> and KVM_TRANSFER_VCPU are sufficient on their own -- no direction flag, and no
> separate export/import ioctls.
>
> And if we settle on a per-session property, then SETUP could conceivably state a
> role for a non-migration transfer session as well.
>
> >
> >>>> For direction specifically, we could pass it to KVM on every call and let
> >>>> KVM stay stateless about it, or KVM could record it once and remember it for
> >>>> the rest of the session.
> >>>
> >>> Yes the "record and remember" is another option, it could be a sub-command
> >>> something like KVM_MIGRATE_DIRECTION.
> >>
> >>
> >> Agreed on record-and-remember, with a refinement on scope.
> >>
> >> A VM that was migrated in can later be migrated out, so the role is not a
> >> property of the VM -- it has to be recorded per session. SETUP is the call
> >> that starts a migration and runs ahead of all other migration calls, so its
> >> arguments look like the natural place for userspace to state the role; a
> >> separate sub-command would need its own scope rules relative to SETUP.
> >> KVM can then ask the vendor layer whether the requested role is permitted for
> >> this VM -- only it knows the confidential-VM state -- and on success record it
> >> in generic KVM state, where it then selects the export or import callbacks.
> >
> > The role can change, but it can be VM specific for starting the migration
> > destination even before migration is started.
> >
> > At least for TDX we need to specify direction for migration destination on
> > init to prevent fully initializing the TD and vCPUs.
>
> As mentioned above, that is a destination launch time directive that we needn't
> conflate with a migration/transfer session role.
It's still the same role though. Yes we can set it on init, but would
be nice to have some generic way to do it for qemu -incoming.
Just brainstorming.. I wonder if we need two things though. A source vs
destination role. And then at some point possibly later on also a data
transfer direction enumeration similar to what Linux has in
include/linux/dma-direction.h.
We already need to make use of the QEMU return-path for the migration key
exchange. What if some hardware needs to make use of KVM_TRANSFER_MEMORY
from source to destination, and after that back from destination to source
to ack the transfer? Sure this is just speculation, I'm not aware of this
need right now.
In any case with handling the source vs destination role, the enumeration
for data direction can be added to the transfer flags later on as needed.
No need to try to stuff the data direction flag there until really needed.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-10 6:33 ` Tony Lindgren
@ 2026-09-11 1:40 ` Kishen Maloor
2026-09-11 4:23 ` Tony Lindgren
0 siblings, 1 reply; 78+ messages in thread
From: Kishen Maloor @ 2026-09-11 1:40 UTC (permalink / raw)
To: Tony Lindgren
Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini,
Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On 9/9/26 11:33 PM, Tony Lindgren wrote:
> On Wed, Sep 09, 2026 at 06:11:11PM -0700, Kishen Maloor wrote:
> ...
>>
>> As mentioned above, that is a destination launch time directive that we needn't
>> conflate with a migration/transfer session role.
>
> It's still the same role though. Yes we can set it on init, but would
> be nice to have some generic way to do it for qemu -incoming.
I'd separate these.
On the role: it seems we agree on recording roles per-session. My only point
then is that a VM created through the delayed_init flow doesn't additionally
need a destination role recorded for it if SETUP will assert one when the
migration is kicked off.
On generic plumbing for -incoming: is this about KVM_TDX_INIT_VM_F_DELAY_INIT?
A generic mechanism would make sense to me if the flag were consumed by generic
KVM code, but that isn't the case here. Userspace has to make a
vendor-specific VM-init call like KVM_TDX_INIT_VM anyway, with the flag passed
on that call. So I'm not sure what a generic version would add, unless you have
something else in mind.
> Just brainstorming.. I wonder if we need two things though. A source vs
> destination role. And then at some point possibly later on also a data
> transfer direction enumeration similar to what Linux has in
> include/linux/dma-direction.h.
>
> We already need to make use of the QEMU return-path for the migration key
> exchange. What if some hardware needs to make use of KVM_TRANSFER_MEMORY
> from source to destination, and after that back from destination to source
> to ack the transfer? Sure this is just speculation, I'm not aware of this
> need right now.
>
> In any case with handling the source vs destination role, the enumeration
> for data direction can be added to the transfer flags later on as needed.
> No need to try to stuff the data direction flag there until really needed.
Agree on deferring such an enumeration. More generally though, the session
role determines which operations are permitted, and it seems like vendor code
could be expected to handle those in context.
The TDX architecture already has a destination-to-source example:
TDH.IMPORT.ABORT emits an abort token on the destination that
TDH.EXPORT.ABORT consumes on the source. So KVM on the source would dispatch
to its export-abort path and let the vendor layer decide what to do with
the buffer. This doesn't require a direction flag.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-11 1:40 ` Kishen Maloor
@ 2026-09-11 4:23 ` Tony Lindgren
2026-09-15 0:14 ` Kishen Maloor
0 siblings, 1 reply; 78+ messages in thread
From: Tony Lindgren @ 2026-09-11 4:23 UTC (permalink / raw)
To: Kishen Maloor
Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini,
Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On Thu, Sep 10, 2026 at 06:40:04PM -0700, Kishen Maloor wrote:
> On 9/9/26 11:33 PM, Tony Lindgren wrote:
> > On Wed, Sep 09, 2026 at 06:11:11PM -0700, Kishen Maloor wrote:
> > ...
> >>
> >> As mentioned above, that is a destination launch time directive that we needn't
> >> conflate with a migration/transfer session role.
> >
> > It's still the same role though. Yes we can set it on init, but would
> > be nice to have some generic way to do it for qemu -incoming.
>
> I'd separate these.
>
> On the role: it seems we agree on recording roles per-session. My only point
> then is that a VM created through the delayed_init flow doesn't additionally
> need a destination role recorded for it if SETUP will assert one when the
> migration is kicked off.
>
> On generic plumbing for -incoming: is this about KVM_TDX_INIT_VM_F_DELAY_INIT?
> A generic mechanism would make sense to me if the flag were consumed by generic
> KVM code, but that isn't the case here. Userspace has to make a
> vendor-specific VM-init call like KVM_TDX_INIT_VM anyway, with the flag passed
> on that call. So I'm not sure what a generic version would add, unless you have
> something else in mind.
So we could add a SETUP subcommand SET_ROLE or SET_INCOMING. The
implementation could store the role at least initially. And if we want to
set the role with KVM_TDX_INIT_VM, we could recycle the role bit there.
> > Just brainstorming.. I wonder if we need two things though. A source vs
> > destination role. And then at some point possibly later on also a data
> > transfer direction enumeration similar to what Linux has in
> > include/linux/dma-direction.h.
> >
> > We already need to make use of the QEMU return-path for the migration key
> > exchange. What if some hardware needs to make use of KVM_TRANSFER_MEMORY
> > from source to destination, and after that back from destination to source
> > to ack the transfer? Sure this is just speculation, I'm not aware of this
> > need right now.
> >
> > In any case with handling the source vs destination role, the enumeration
> > for data direction can be added to the transfer flags later on as needed.
> > No need to try to stuff the data direction flag there until really needed.
>
> Agree on deferring such an enumeration. More generally though, the session
> role determines which operations are permitted, and it seems like vendor code
> could be expected to handle those in context.
>
> The TDX architecture already has a destination-to-source example:
> TDH.IMPORT.ABORT emits an abort token on the destination that
> TDH.EXPORT.ABORT consumes on the source. So KVM on the source would dispatch
> to its export-abort path and let the vendor layer decide what to do with
> the buffer. This doesn't require a direction flag.
Yes agreed with the role we can handle migration related transfers both
ways at least for TDX.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-11 4:23 ` Tony Lindgren
@ 2026-09-15 0:14 ` Kishen Maloor
2026-09-15 4:44 ` Tony Lindgren
0 siblings, 1 reply; 78+ messages in thread
From: Kishen Maloor @ 2026-09-15 0:14 UTC (permalink / raw)
To: Tony Lindgren
Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini,
Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On 9/10/26 9:23 PM, Tony Lindgren wrote:
> On Thu, Sep 10, 2026 at 06:40:04PM -0700, Kishen Maloor wrote:
>> On 9/9/26 11:33 PM, Tony Lindgren wrote:
>>> On Wed, Sep 09, 2026 at 06:11:11PM -0700, Kishen Maloor wrote:
>>> ...
>>>>
>>>> As mentioned above, that is a destination launch time directive that we needn't
>>>> conflate with a migration/transfer session role.
>>>
>>> It's still the same role though. Yes we can set it on init, but would
>>> be nice to have some generic way to do it for qemu -incoming.
>>
>> I'd separate these.
>>
>> On the role: it seems we agree on recording roles per-session. My only point
>> then is that a VM created through the delayed_init flow doesn't additionally
>> need a destination role recorded for it if SETUP will assert one when the
>> migration is kicked off.
>>
>> On generic plumbing for -incoming: is this about KVM_TDX_INIT_VM_F_DELAY_INIT?
>> A generic mechanism would make sense to me if the flag were consumed by generic
>> KVM code, but that isn't the case here. Userspace has to make a
>> vendor-specific VM-init call like KVM_TDX_INIT_VM anyway, with the flag passed
>> on that call. So I'm not sure what a generic version would add, unless you have
>> something else in mind.
>
> So we could add a SETUP subcommand SET_ROLE or SET_INCOMING.
It might be better to pass the role as a parameter of the SETUP call
rather than a separate SET_ROLE call.
A separate SET_ROLE would bring its own ordering rules relative to the
other SETUP sub-commands.
The specific role (src or dst) still needs someplace to go, and flags is
already the sub-command selector. An option is to carve out room in
the __u32 reserved field to carry a role argument, something
like 0=unset, 1=src, 2=dest so KVM can verify that a role was indeed set.
It could further be written into a generic KVM struct (kvm_arch or kvm) which
could be queried on KVM_TRANSFER_MEMORY, etc. to identify the relevant
callback.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-15 0:14 ` Kishen Maloor
@ 2026-09-15 4:44 ` Tony Lindgren
2026-09-15 15:53 ` Kishen Maloor
0 siblings, 1 reply; 78+ messages in thread
From: Tony Lindgren @ 2026-09-15 4:44 UTC (permalink / raw)
To: Kishen Maloor
Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini,
Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On Mon, Sep 14, 2026 at 05:14:32PM -0700, Kishen Maloor wrote:
> On 9/10/26 9:23 PM, Tony Lindgren wrote:
> > On Thu, Sep 10, 2026 at 06:40:04PM -0700, Kishen Maloor wrote:
> >> On 9/9/26 11:33 PM, Tony Lindgren wrote:
> >>> On Wed, Sep 09, 2026 at 06:11:11PM -0700, Kishen Maloor wrote:
> >>> ...
> >>>>
> >>>> As mentioned above, that is a destination launch time directive that we needn't
> >>>> conflate with a migration/transfer session role.
> >>>
> >>> It's still the same role though. Yes we can set it on init, but would
> >>> be nice to have some generic way to do it for qemu -incoming.
> >>
> >> I'd separate these.
> >>
> >> On the role: it seems we agree on recording roles per-session. My only point
> >> then is that a VM created through the delayed_init flow doesn't additionally
> >> need a destination role recorded for it if SETUP will assert one when the
> >> migration is kicked off.
> >>
> >> On generic plumbing for -incoming: is this about KVM_TDX_INIT_VM_F_DELAY_INIT?
> >> A generic mechanism would make sense to me if the flag were consumed by generic
> >> KVM code, but that isn't the case here. Userspace has to make a
> >> vendor-specific VM-init call like KVM_TDX_INIT_VM anyway, with the flag passed
> >> on that call. So I'm not sure what a generic version would add, unless you have
> >> something else in mind.
> >
> > So we could add a SETUP subcommand SET_ROLE or SET_INCOMING.
> It might be better to pass the role as a parameter of the SETUP call
> rather than a separate SET_ROLE call.
> A separate SET_ROLE would bring its own ordering rules relative to the
> other SETUP sub-commands.
OK a flag for SETUP sounds good to me. SETUP is needed anyways for each
migration session.
> The specific role (src or dst) still needs someplace to go, and flags is
> already the sub-command selector. An option is to carve out room in
> the __u32 reserved field to carry a role argument, something
> like 0=unset, 1=src, 2=dest so KVM can verify that a role was indeed set.
To me it seems that 0=src can be the natural default starting point, I
don't think we need 0=unset.
> It could further be written into a generic KVM struct (kvm_arch or kvm) which
> could be queried on KVM_TRANSFER_MEMORY, etc. to identify the relevant
> callback.
Yeah eventually some generic place for it would be nice. But that's easy
to add later on too.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-15 4:44 ` Tony Lindgren
@ 2026-09-15 15:53 ` Kishen Maloor
2026-09-16 5:09 ` Tony Lindgren
0 siblings, 1 reply; 78+ messages in thread
From: Kishen Maloor @ 2026-09-15 15:53 UTC (permalink / raw)
To: Tony Lindgren
Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini,
Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On 9/14/26 9:44 PM, Tony Lindgren wrote:
> On Mon, Sep 14, 2026 at 05:14:32PM -0700, Kishen Maloor wrote:
>> On 9/10/26 9:23 PM, Tony Lindgren wrote:
>>> On Thu, Sep 10, 2026 at 06:40:04PM -0700, Kishen Maloor wrote:
>>>> On 9/9/26 11:33 PM, Tony Lindgren wrote:
>>>>> On Wed, Sep 09, 2026 at 06:11:11PM -0700, Kishen Maloor wrote:
>>>>> ...
>>>>>>
>>>>>> As mentioned above, that is a destination launch time directive that we needn't
>>>>>> conflate with a migration/transfer session role.
>>>>>
>>>>> It's still the same role though. Yes we can set it on init, but would
>>>>> be nice to have some generic way to do it for qemu -incoming.
>>>>
>>>> I'd separate these.
>>>>
>>>> On the role: it seems we agree on recording roles per-session. My only point
>>>> then is that a VM created through the delayed_init flow doesn't additionally
>>>> need a destination role recorded for it if SETUP will assert one when the
>>>> migration is kicked off.
>>>>
>>>> On generic plumbing for -incoming: is this about KVM_TDX_INIT_VM_F_DELAY_INIT?
>>>> A generic mechanism would make sense to me if the flag were consumed by generic
>>>> KVM code, but that isn't the case here. Userspace has to make a
>>>> vendor-specific VM-init call like KVM_TDX_INIT_VM anyway, with the flag passed
>>>> on that call. So I'm not sure what a generic version would add, unless you have
>>>> something else in mind.
>>>
>>> So we could add a SETUP subcommand SET_ROLE or SET_INCOMING.
>> It might be better to pass the role as a parameter of the SETUP call
>> rather than a separate SET_ROLE call.
>> A separate SET_ROLE would bring its own ordering rules relative to the
>> other SETUP sub-commands.
>
> OK a flag for SETUP sounds good to me. SETUP is needed anyways for each
> migration session.
To be clear, I was suggesting a field in kvm_migrate_cmd and not a flag to pass
the role as an argument to SETUP. As I mentioned in my last comments (right
below), 'flags' in the current proposal carry the vendor-defined sub-command values,
so a generic role argument wouldn't belong there.
>
>> The specific role (src or dst) still needs someplace to go, and flags is
>> already the sub-command selector. An option is to carve out room in
>> the __u32 reserved field to carry a role argument, something
>> like 0=unset, 1=src, 2=dest so KVM can verify that a role was indeed set.
>
> To me it seems that 0=src can be the natural default starting point, I
> don't think we need 0=unset.
With 0=src, the field would only carry information when it's a destination,
thereby making it an is_dest boolean rather than a role. It would also mean any
VM that never established a session still reads as a source, so a
KVM_TRANSFER_MEMORY aimed at the wrong VM by buggy or rogue userspace would get
dispatched to the export path instead of rejected outright. The generic
dispatcher shouldn't have to rely on the vendor layer to catch that. Reserving 0
for 'unset' costs nothing and lets KVM reject a session that never stated a
role.
>
>> It could further be written into a generic KVM struct (kvm_arch or kvm) which
>> could be queried on KVM_TRANSFER_MEMORY, etc. to identify the relevant
>> callback.
>
> Yeah eventually some generic place for it would be nice. But that's easy
> to add later on too.
Sure, we don't have to decide now as we're still discussing the UAPI.
But we'd want this detail also settled sometime before we call the UAPI
complete as it determines whether the dispatch is generic.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-15 15:53 ` Kishen Maloor
@ 2026-09-16 5:09 ` Tony Lindgren
2026-09-17 3:31 ` Kishen Maloor
0 siblings, 1 reply; 78+ messages in thread
From: Tony Lindgren @ 2026-09-16 5:09 UTC (permalink / raw)
To: Kishen Maloor
Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini,
Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On Tue, Sep 15, 2026 at 08:53:18AM -0700, Kishen Maloor wrote:
> On 9/14/26 9:44 PM, Tony Lindgren wrote:
> > On Mon, Sep 14, 2026 at 05:14:32PM -0700, Kishen Maloor wrote:
> >> On 9/10/26 9:23 PM, Tony Lindgren wrote:
> >>> On Thu, Sep 10, 2026 at 06:40:04PM -0700, Kishen Maloor wrote:
> >>>> On 9/9/26 11:33 PM, Tony Lindgren wrote:
> >>>>> On Wed, Sep 09, 2026 at 06:11:11PM -0700, Kishen Maloor wrote:
> >>>>> ...
> >>>>>>
> >>>>>> As mentioned above, that is a destination launch time directive that we needn't
> >>>>>> conflate with a migration/transfer session role.
> >>>>>
> >>>>> It's still the same role though. Yes we can set it on init, but would
> >>>>> be nice to have some generic way to do it for qemu -incoming.
> >>>>
> >>>> I'd separate these.
> >>>>
> >>>> On the role: it seems we agree on recording roles per-session. My only point
> >>>> then is that a VM created through the delayed_init flow doesn't additionally
> >>>> need a destination role recorded for it if SETUP will assert one when the
> >>>> migration is kicked off.
> >>>>
> >>>> On generic plumbing for -incoming: is this about KVM_TDX_INIT_VM_F_DELAY_INIT?
> >>>> A generic mechanism would make sense to me if the flag were consumed by generic
> >>>> KVM code, but that isn't the case here. Userspace has to make a
> >>>> vendor-specific VM-init call like KVM_TDX_INIT_VM anyway, with the flag passed
> >>>> on that call. So I'm not sure what a generic version would add, unless you have
> >>>> something else in mind.
> >>>
> >>> So we could add a SETUP subcommand SET_ROLE or SET_INCOMING.
> >> It might be better to pass the role as a parameter of the SETUP call
> >> rather than a separate SET_ROLE call.
> >> A separate SET_ROLE would bring its own ordering rules relative to the
> >> other SETUP sub-commands.
> >
> > OK a flag for SETUP sounds good to me. SETUP is needed anyways for each
> > migration session.
>
> To be clear, I was suggesting a field in kvm_migrate_cmd and not a flag to pass
> the role as an argument to SETUP. As I mentioned in my last comments (right
> below), 'flags' in the current proposal carry the vendor-defined sub-command values,
> so a generic role argument wouldn't belong there.
I was thinking 8 bits for common flags and 8 bits for vendor flags but
yeah that can be a bit tight. Sorry if the SETUP above caused extra
confusion.
So trying to summarize the common flags for the role and separate vendor
flags:
struct kvm_migrate_cmd {
__u16 command;
__u16 flags;
__u16 vflags;
__u16 reserved;
__u32 reserved;
struct kvm_transfer_buffer buf;
};
Is the above along the lines what you were thinking?
> >> The specific role (src or dst) still needs someplace to go, and flags is
> >> already the sub-command selector. An option is to carve out room in
> >> the __u32 reserved field to carry a role argument, something
> >> like 0=unset, 1=src, 2=dest so KVM can verify that a role was indeed set.
> >
> > To me it seems that 0=src can be the natural default starting point, I
> > don't think we need 0=unset.
>
> With 0=src, the field would only carry information when it's a destination,
> thereby making it an is_dest boolean rather than a role. It would also mean any
> VM that never established a session still reads as a source, so a
> KVM_TRANSFER_MEMORY aimed at the wrong VM by buggy or rogue userspace would get
> dispatched to the export path instead of rejected outright. The generic
> dispatcher shouldn't have to rely on the vendor layer to catch that. Reserving 0
> for 'unset' costs nothing and lets KVM reject a session that never stated a
> role.
Ah OK, yes that would also tell "the hardware has been initialized to a
certain migration role". That seems like a usable common feature.
> >> It could further be written into a generic KVM struct (kvm_arch or kvm) which
> >> could be queried on KVM_TRANSFER_MEMORY, etc. to identify the relevant
> >> callback.
> >
> > Yeah eventually some generic place for it would be nice. But that's easy
> > to add later on too.
>
> Sure, we don't have to decide now as we're still discussing the UAPI.
> But we'd want this detail also settled sometime before we call the UAPI
> complete as it determines whether the dispatch is generic.
Yes a shared place for the role would make some generic sanity checks
easier.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-16 5:09 ` Tony Lindgren
@ 2026-09-17 3:31 ` Kishen Maloor
2026-09-17 6:42 ` Tony Lindgren
0 siblings, 1 reply; 78+ messages in thread
From: Kishen Maloor @ 2026-09-17 3:31 UTC (permalink / raw)
To: Tony Lindgren
Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini,
Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On 9/15/26 10:09 PM, Tony Lindgren wrote:
> On Tue, Sep 15, 2026 at 08:53:18AM -0700, Kishen Maloor wrote:
>> On 9/14/26 9:44 PM, Tony Lindgren wrote:
>>> On Mon, Sep 14, 2026 at 05:14:32PM -0700, Kishen Maloor wrote:
>>>> On 9/10/26 9:23 PM, Tony Lindgren wrote:
>>>>> On Thu, Sep 10, 2026 at 06:40:04PM -0700, Kishen Maloor wrote:
>>>>>> On 9/9/26 11:33 PM, Tony Lindgren wrote:
>>>>>>> On Wed, Sep 09, 2026 at 06:11:11PM -0700, Kishen Maloor wrote:
>>>>>>> ...
>>>>>>>>
>>>>>>>> As mentioned above, that is a destination launch time directive that we needn't
>>>>>>>> conflate with a migration/transfer session role.
>>>>>>>
>>>>>>> It's still the same role though. Yes we can set it on init, but would
>>>>>>> be nice to have some generic way to do it for qemu -incoming.
>>>>>>
>>>>>> I'd separate these.
>>>>>>
>>>>>> On the role: it seems we agree on recording roles per-session. My only point
>>>>>> then is that a VM created through the delayed_init flow doesn't additionally
>>>>>> need a destination role recorded for it if SETUP will assert one when the
>>>>>> migration is kicked off.
>>>>>>
>>>>>> On generic plumbing for -incoming: is this about KVM_TDX_INIT_VM_F_DELAY_INIT?
>>>>>> A generic mechanism would make sense to me if the flag were consumed by generic
>>>>>> KVM code, but that isn't the case here. Userspace has to make a
>>>>>> vendor-specific VM-init call like KVM_TDX_INIT_VM anyway, with the flag passed
>>>>>> on that call. So I'm not sure what a generic version would add, unless you have
>>>>>> something else in mind.
>>>>>
>>>>> So we could add a SETUP subcommand SET_ROLE or SET_INCOMING.
>>>> It might be better to pass the role as a parameter of the SETUP call
>>>> rather than a separate SET_ROLE call.
>>>> A separate SET_ROLE would bring its own ordering rules relative to the
>>>> other SETUP sub-commands.
>>>
>>> OK a flag for SETUP sounds good to me. SETUP is needed anyways for each
>>> migration session.
>>
>> To be clear, I was suggesting a field in kvm_migrate_cmd and not a flag to pass
>> the role as an argument to SETUP. As I mentioned in my last comments (right
>> below), 'flags' in the current proposal carry the vendor-defined sub-command values,
>> so a generic role argument wouldn't belong there.
>
> I was thinking 8 bits for common flags and 8 bits for vendor flags but
> yeah that can be a bit tight. Sorry if the SETUP above caused extra
> confusion.
>
> So trying to summarize the common flags for the role and separate vendor
> flags:
>
> struct kvm_migrate_cmd {
> __u16 command;
> __u16 flags;
> __u16 vflags;
> __u16 reserved;
> __u32 reserved;
> struct kvm_transfer_buffer buf;
> };
>
> Is the above along the lines what you were thinking?
No, I was suggesting a 'role' field carved out of the 'reserved' space,
like this:
struct kvm_migrate_cmd {
__u16 command;
__u16 flags;
__u8 role; /* 0 = unset, 1 = source, 2 = destination */
__u8 reserved[3];
struct kvm_transfer_buffer buf;
};
We haven't defined any generic flags. Thus far in this proposal 'flags'
contains only vendor-defined sub-command values.
>
>>>> The specific role (src or dst) still needs someplace to go, and flags is
>>>> already the sub-command selector. An option is to carve out room in
>>>> the __u32 reserved field to carry a role argument, something
>>>> like 0=unset, 1=src, 2=dest so KVM can verify that a role was indeed set.
>>>
>>> To me it seems that 0=src can be the natural default starting point, I
>>> don't think we need 0=unset.
>>
>> With 0=src, the field would only carry information when it's a destination,
>> thereby making it an is_dest boolean rather than a role. It would also mean any
>> VM that never established a session still reads as a source, so a
>> KVM_TRANSFER_MEMORY aimed at the wrong VM by buggy or rogue userspace would get
>> dispatched to the export path instead of rejected outright. The generic
>> dispatcher shouldn't have to rely on the vendor layer to catch that. Reserving 0
>> for 'unset' costs nothing and lets KVM reject a session that never stated a
>> role.
>
> Ah OK, yes that would also tell "the hardware has been initialized to a
> certain migration role". That seems like a usable common feature.
Not quite. It tells us that userspace asserted a role for this VM's migration
session. Whether a TD was created for import is a separate, vendor-level detail.
The generic layer only needs the role to reject a session that never stated one,
and to pick the export or import callback. That callback then knows which side
it's on and can reject an incorrect role (e.g., if SETUP asserted dst for a src TD).
>
>>>> It could further be written into a generic KVM struct (kvm_arch or kvm) which
>>>> could be queried on KVM_TRANSFER_MEMORY, etc. to identify the relevant
>>>> callback.
>>>
>>> Yeah eventually some generic place for it would be nice. But that's easy
>>> to add later on too.
>>
>> Sure, we don't have to decide now as we're still discussing the UAPI.
>> But we'd want this detail also settled sometime before we call the UAPI
>> complete as it determines whether the dispatch is generic.
>
> Yes a shared place for the role would make some generic sanity checks
> easier.
Agreed. The field above is just an argument to SETUP. Where that role gets stored
for KVM's top-level dispatcher to check is the detail we can settle later.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-17 3:31 ` Kishen Maloor
@ 2026-09-17 6:42 ` Tony Lindgren
2026-09-18 4:32 ` Kishen Maloor
0 siblings, 1 reply; 78+ messages in thread
From: Tony Lindgren @ 2026-09-17 6:42 UTC (permalink / raw)
To: Kishen Maloor
Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini,
Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On Wed, Sep 16, 2026 at 08:31:32PM -0700, Kishen Maloor wrote:
> On 9/15/26 10:09 PM, Tony Lindgren wrote:
> > So trying to summarize the common flags for the role and separate vendor
> > flags:
> >
> > struct kvm_migrate_cmd {
> > __u16 command;
> > __u16 flags;
> > __u16 vflags;
> > __u16 reserved;
> > __u32 reserved;
> > struct kvm_transfer_buffer buf;
> > };
> >
> > Is the above along the lines what you were thinking?
>
> No, I was suggesting a 'role' field carved out of the 'reserved' space,
> like this:
>
> struct kvm_migrate_cmd {
> __u16 command;
> __u16 flags;
> __u8 role; /* 0 = unset, 1 = source, 2 = destination */
> __u8 reserved[3];
> struct kvm_transfer_buffer buf;
> };
OK yes thanks for clarifying, that works for me.
> > Ah OK, yes that would also tell "the hardware has been initialized to a
> > certain migration role". That seems like a usable common feature.
>
> Not quite. It tells us that userspace asserted a role for this VM's migration
> session. Whether a TD was created for import is a separate, vendor-level detail.
> The generic layer only needs the role to reject a session that never stated one,
> and to pick the export or import callback. That callback then knows which side
> it's on and can reject an incorrect role (e.g., if SETUP asserted dst for a src TD).
That's a good point, the hardware role may not be set yet.
I'm still wondering if there is a need to stash the userspace set role in
KVM though. Likely only the hardware specific code can properly track the
state of the hardware and adjust to the userspace requests. Seems just
being able to pass the role in struct kvm_migrate_cmd should be enough?
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
2026-09-04 18:24 ` [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Artem Bityutskiy
@ 2026-09-17 21:27 ` Peter Xu
2026-09-18 12:46 ` Artem Bityutskiy
2026-09-20 23:56 ` Kishen Maloor
0 siblings, 2 replies; 78+ messages in thread
From: Peter Xu @ 2026-09-17 21:27 UTC (permalink / raw)
To: Artem Bityutskiy
Cc: Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas,
Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
On Fri, Sep 04, 2026 at 09:24:25PM +0300, Artem Bityutskiy wrote:
> On Mon, 2026-08-31 at 10:13 +0300, Tony Lindgren wrote:
> > Tom, since you mentioned that AMD SEV-SNP and Intel TDX live migration
> > sound similar, can you please take a look how the API might work for
> > SEV-SNP?
>
> The presented example uAPIs were designed to fit the TDX live migration
> flow, with the intent that they could also be used by other CoCo VMs.
>
> It would be super great if someone could commend on how suitable these
> uAPIs are for AMD and ARM flows.
>
> AFAIU, what is generic in the presented uAPI is more of usage pattern.
>
> - The order in which QEMU calls them.
> - The idea that QEMU/KVM is a transport for opaque blobs.
>
> The proposed container - 'struct kvm_transfer_buffer' - only has address
> and a size. The vendor defines the layout of the data in it.
>
> Some example high-level topics that would be nice to get feedback on:
>
> - Can we come up with a single set of generic migration uAPIs for different
> CoCo models?
> - Or should some uAPIs be generic while others are vendor-specific?
> - Or should each CoCo model have its own vendor-specific set of migration
> uAPIs?
It's always good if we can put together as much function to be shared with
generic ioctls as possible. At some point, IMHO we need to collect such
information somehow, so when merging the generic API we know what vendor
specific API will be needed. Hopefully this series is a good start.
> - Should the same uAPIs also support traditional VMs? But the only use-case
> I imagine here is "for testing purposes".
This is an interesting idea, I think this could be useful. Especially, I
wonder if you already have it done and PoC branches you can share, so that
I can play with it.
>
> > Artem has put together a brief description below of the example API and the
> > migration flow:
>
> .. snip ...
>
> > Migration flow
> > ==============
> >
> > Source host Destination host
> > =========== ================
> >
> > CMD(SETUP/SESSION) <--- setup msgs ---> CMD(SETUP/SESSION)
> > | (repeated) |
> > CMD(SETUP/IMMUTABLE_STATE) - immutable state -> CMD(SETUP/IMMUTABLE_STATE)
> > | |
> > KVM_GET_DIRTY_LOG |
> > KVM_EXPORT_MEMORY --- memory data ---> KVM_IMPORT_MEMORY
> > CMD(ITERATION) --- epoch token ---> CMD(ITERATION)
Could you elaborate this ITERATION operation? Is that something the
userapp must do after full scan of a round of guest memory?
> > | (repeat until convergence) |
> > CMD(STOP_AND_COPY/PAUSE) |
> > CMD(STOP_AND_COPY/TD_STATE) --- VM state ------> CMD(STOP_AND_COPY/TD_STATE)
> > KVM_EXPORT_VCPU --- vCPU state ----> KVM_IMPORT_VCPU
> > KVM_EXPORT_MEMORY -- final memory ---> KVM_IMPORT_MEMORY
When read/write encrypted memories, two questions:
- Is there an upper bound of the buffer size per-page?
- Does this operation supports concurrency? If it supports, how well it
scales per expectation (e.g. is there known big lock for that)?
Similar question to the vCPU getter and setter. For now even without CoCo
we serialize vCPU get/set, but I want to understand the potential of
concurrent operations, and see if there's anything special for CoCo from
that regard.
> > CMD(ITERATION/DONE) --- start token ---> CMD(ITERATION)
> > | |
> > CMD(END) CMD(END)
>
> ... snip ...
>
>
> As I mentioned, the example uAPI is modeled around the TDX migration flow.
> In case it helps the reader, here is a summary of that flow that was
> presented in PUCK.
>
> It describes what the TDX module offers today and focuses on pre-copy
> migration. This is our interpretation of the TDX specifications, not a
IMHO we should really take postcopy into account when designing the API and
state machine. We don't need to implement it in the first version, even
until merging, but we need to make sure postcopy will be new ioctls on top
of existing and it should have no major loopholes that it'll need a new set
of APIs.
For example, I think we should consider KVM_EXPORT_MEMORY being usable
after END on source, KVM_IMPORT_MEMORY while TD is in operation, etc. We
should likely also need to still picture the rough process of postcopy,
reserve those APIs since the start (but return -EINVAL or something).
AFAIU, postcopy is so far still the best solution for extremely large or
extremely busy VMs regarding user experience, and it will happen to CoCo
VMs one day or another.
> specification itself, and may contain errors or omissions. Please refer to
> the official TDX specifications for authoritative information.
>
> Migration Overview
> ------------------
>
> In TDX, the TDX module implements the migration logic. The VMM drives
> migration by issuing seamcalls to the source and destination TDX modules and
> transports the resulting encrypted blobs between them. The VMM does not need
> to know the contents of those blobs.
>
> The TDX security model enforces two hard rules:
>
> - Only one instance of the TD may run at a time. Either the source or the
Just curious - could I ask why this limitation?
> destination may run, but never both. In other words, cloning a TD is not
> allowed.
> - When migration completes, the destination must have the same memory and
> vCPU state as the source. It must not end up with a partial or mixed
> state.
If such happens, it's definitely a bug, even without CoCo. Anything
specific about CoCo? Like, whole-VM checksum?
I recall QEMU could have some devices touching the memory during its post
load process (after destination QEMU receive the device states and apply).
I'm not sure how much it affects.
>
> Simple Overview
> ---------------
>
> Run the migration setup session
> |
> v
> Transfer immutable TD state
> |
> v
> Copy dirty memory pages while the TD runs <-------+
> | |
> +-- iterate until convergence criteria is met --+
> |
> v
> Pause the source TD and copy the remaining state
> |
> v
> Start the destination TD
>
> The migration process starts with a setup session. During the setup session,
> the source and destination TDX modules exchange encrypted blobs and
> establish migration encryption keys. For example, these keys protect memory
> contents transferred during migration.
>
> Next, the source transfers the immutable TD state to the destination and
> initializes the destination TD. This state includes the read-only VM data,
> such as its vCPU count and topology.
>
> The source and destination then perform iterative memory-copy rounds. In
> each round, the source scans for memory pages that need to be migrated and
> exports them in encrypted form. The destination imports the pages. The
> rounds continue until the number of pages that change between rounds is
> small enough to meet the convergence criteria.
>
> The source TD is then paused. The VMM transfers the remaining TD state, such
I want to understand what is extra for a CoCo VM in terms of "pause", say,
what's more than "stopping the vCPU threads".
I saw there's mention of PRE_COPY_STOP state. One example question is,
when reaching this state, can the guest memory still change? What happens
if some emulated device are still DMAing to the guest memory (assuming
flipped from private to shared)? In case of future IO zone support, what
happens if in case of VFIO-PCI assigned doing encrypted DMA?
From that regard, VFIO has the P2P state where it quiesce initiation of any
DMA from this specific device, then another round to fully stop all devices
into STOP_COPY phase. I wonder if CoCo VMs need similar treatment.
> as vCPU state, and the final dirty memory pages. Finally, the destination TD
> is started and the migration completes.
>
> Notice that the VMM's role in this process is to issue the required
> seamcalls and send encrypted blobs between the source and destination. The
> VMM does not need to know the contents of those blobs.
>
> More Detailed Overview
> ----------------------
>
> The following diagram shows the TDX seamcalls used by the source and
> destination:
>
> Source host Destination host
> =========== ================
>
> TDH.MIG.SETUP -- crypto keys, attestation -> TDH.MIG.SETUP
> | |
> TDH.EXPORT.STATE.IMMUTABLE -- read-only TD state ->
> TDH.IMPORT.STATE.IMMUTABLE
> | |
> TDH.MEM.SCAN.RANGE |
Could you elaborate what's the relations between TDH.MEM.SCAN.RANGE and the
GET_DIRTY_LOG ioctl? I recall above mentioned GET_DIRTY_LOG will be
available even for CoCo, which makes sense assuming dirty information isn't
confidential. However then I don't understand what TDH.MEM.SCAN.RANGE
plays the role here.
Thanks,
> TDH.MEM.TRACK + IPIs |
> TDH.EXPORT.MEM ------ memory data ---------> TDH.IMPORT.MEM
> TDH.EXPORT.TRACK ------ epoch token ---------> TDH.IMPORT.TRACK
> | |
> TDH.EXPORT.PAUSE |
> | |
> TDH.EXPORT.STATE.TD ------ global TD state -----> TDH.IMPORT.STATE.TD
> | |
> TDH.EXPORT.STATE.VP ------ vCPU state ----------> TDH.IMPORT.STATE.VP
> | |
> TDH.MEM.SCAN.COMP |
> TDH.MEM.TRACK + IPIs |
> TDH.EXPORT.MEM ------ final memory data ---> TDH.IMPORT.MEM
> TDH.EXPORT.TRACK ------ start token ---------> TDH.IMPORT.TRACK
> |
> TDH.IMPORT.END
>
> The source and destination go through the following stages.
>
> Setup
> -----
>
> The VMM issues TDH.MIG.SETUP on the source and destination TDX modules
> iteratively to perform the setup session. The modules return status and may
> also return an encrypted blob, which the VMM passes between the two sides.
> In other words, the migration protocol is between the two TDX modules, while
> the VMM is simply the transport mechanism. The setup session performs mutual
> attestation, establishes trust between the modules, establishes migration
> encryption keys, and loads the migration policy. It completes when the TDX
> module returns success.
>
> Immutable State Transfer
> ------------------------
>
> The source calls TDH.EXPORT.STATE.IMMUTABLE to export the immutable TD
> state. The destination VMM creates the destination TD skeleton and
> configures migration streams, then calls TDH.IMPORT.STATE.IMMUTABLE, which
> finalizes the destination TD initialization.
>
> Iterative Memory Copy
> ---------------------
>
> While the source TD continues running, the source and destination run
> iterative memory copy rounds:
>
> - The source calls TDH.MEM.SCAN.RANGE to find migration candidate pages.
> - The source calls TDH.MEM.TRACK and sends IPIs to the TD's vCPUs in order
> to ensure the dirty page scanning algorithm correctness.
> - The source calls TDH.EXPORT.MEM to export private memory pages, which the
> VMM sends to the destination.
> - The destination calls TDH.IMPORT.MEM to import the received memory pages.
> - The source calls TDH.EXPORT.TRACK to generate an epoch token. The VMM
> sends it to the destination, which calls TDH.IMPORT.TRACK to consume the
> token. This verifies that all data exported from the source was imported
> on the destination.
>
> The rounds continue until the number of pages that change between rounds is
> small enough to meet the convergence criteria.
>
> Stop and Copy
> -------------
>
> The VMM pauses the source TD, which begins the downtime period, then calls
> TDH.EXPORT.PAUSE to start the TDX-enforced blackout period. The source then
> calls TDH.EXPORT.STATE.TD to export mutable TD-scope state and
> TDH.EXPORT.STATE.VP to export mutable state for each vCPU. It then performs
> the final dirty page scan, runs TDH.MEM.TRACK and sends IPIs, and calls
> TDH.EXPORT.MEM to export the final dirty pages as encrypted blobs.
>
> The destination calls TDH.IMPORT.STATE.TD to import mutable TD-scope state
> and TDH.IMPORT.STATE.VP to import the mutable state of each vCPU. It calls
> TDH.IMPORT.MEM to import the final dirty pages.
>
> After exporting the final dirty pages, the source's last migration seamcall
> is TDH.EXPORT.TRACK with IN_ORDER_DONE=1, which generates the start token.
> The VMM sends the start token to the destination. The destination calls
> TDH.IMPORT.TRACK with this token. This verifies that the mutable TD state
> has been imported and allows the destination to start the TD. Finally, the
> destination calls TDH.IMPORT.END, which ends the migration.
>
--
Peter Xu
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-17 6:42 ` Tony Lindgren
@ 2026-09-18 4:32 ` Kishen Maloor
2026-09-18 5:58 ` Tony Lindgren
0 siblings, 1 reply; 78+ messages in thread
From: Kishen Maloor @ 2026-09-18 4:32 UTC (permalink / raw)
To: Tony Lindgren
Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini,
Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On 9/16/26 11:42 PM, Tony Lindgren wrote:
> On Wed, Sep 16, 2026 at 08:31:32PM -0700, Kishen Maloor wrote:
>> On 9/15/26 10:09 PM, Tony Lindgren wrote:
>>> So trying to summarize the common flags for the role and separate vendor
>>> flags:
>>>
>>> struct kvm_migrate_cmd {
>>> __u16 command;
>>> __u16 flags;
>>> __u16 vflags;
>>> __u16 reserved;
>>> __u32 reserved;
>>> struct kvm_transfer_buffer buf;
>>> };
>>>
>>> Is the above along the lines what you were thinking?
>>
>> No, I was suggesting a 'role' field carved out of the 'reserved' space,
>> like this:
>>
>> struct kvm_migrate_cmd {
>> __u16 command;
>> __u16 flags;
>> __u8 role; /* 0 = unset, 1 = source, 2 = destination */
>> __u8 reserved[3];
>> struct kvm_transfer_buffer buf;
>> };
>
> OK yes thanks for clarifying, that works for me.
>
>>> Ah OK, yes that would also tell "the hardware has been initialized to a
>>> certain migration role". That seems like a usable common feature.
>>
>> Not quite. It tells us that userspace asserted a role for this VM's migration
>> session. Whether a TD was created for import is a separate, vendor-level detail.
>> The generic layer only needs the role to reject a session that never stated one,
>> and to pick the export or import callback. That callback then knows which side
>> it's on and can reject an incorrect role (e.g., if SETUP asserted dst for a src TD).
>
> That's a good point, the hardware role may not be set yet.
>
> I'm still wondering if there is a need to stash the userspace set role in
> KVM though. Likely only the hardware specific code can properly track the
> state of the hardware and adjust to the userspace requests. Seems just
> being able to pass the role in struct kvm_migrate_cmd should be enough?
Passing it in kvm_migrate_cmd is enough for SETUP itself, but the commands
after SETUP like memory/vcpu transfers still have to reach the right
callback. So KVM would need to remember what was asserted so that the
generic layer can dispatch to the export or import facing callbacks.
We've been sketching (on this thread) an alternative UAPI set
(3 vs 5 ioctls) for consideration which this stored role enables:
Proposed in the RFC Alternative
KVM_MIGRATE_CMD KVM_MIGRATE_CMD
KVM_EXPORT_MEMORY
KVM_IMPORT_MEMORY KVM_TRANSFER_MEMORY
KVM_EXPORT_VCPU
KVM_IMPORT_VCPU KVM_TRANSFER_VCPU
It's just a record (1 byte) of what userspace asserted for the current session
at SETUP. Vendor code still owns the hardware state and remains free to reject a
role that doesn't match it. It is also what lets the generic layer reject a
command to a VM that never set up a session.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-08-31 7:13 ` [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD Tony Lindgren
2026-08-31 7:23 ` sashiko-bot
2026-09-07 11:53 ` Tony Lindgren
@ 2026-09-18 4:33 ` Kishen Maloor
2026-09-21 5:58 ` Tony Lindgren
2 siblings, 1 reply; 78+ messages in thread
From: Kishen Maloor @ 2026-09-18 4:33 UTC (permalink / raw)
To: Tony Lindgren, Paolo Bonzini, Sean Christopherson
Cc: Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg,
Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm
On 8/31/26 12:13 AM, Tony Lindgren wrote:
I was wondering how a vendor would use kvm_transfer_buffer for a command that
passes an input and also returns an output. size can describe the input
length or the space available for output, not both, so the kernel has no way to
know how much it may write.
Two suggestions below. They're orthogonal.
> ...
> diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h
> index 5f6c1ce9673b7..d9291a8a97bb1 100644
> --- a/arch/x86/include/asm/kvm_host.h
> +++ b/arch/x86/include/asm/kvm_host.h
> @@ -2010,6 +2010,8 @@ struct kvm_x86_ops {
> int (*gmem_prepare)(struct kvm *kvm, kvm_pfn_t pfn, gfn_t gfn, int max_order);
> void (*gmem_invalidate)(kvm_pfn_t start, kvm_pfn_t end);
> int (*gmem_max_mapping_level)(struct kvm *kvm, kvm_pfn_t pfn, bool is_private);
> + bool (*cap_live_migration)(struct kvm *kvm);
Should this return an 'int' specifying the maximum buffer size that the vendor impl
requires/uses? Userspace can then learn this once.
> ...
> +
> +struct kvm_transfer_buffer {
> + __u64 address;
> + __u32 size;
> + __u32 reserved;
> +};
Should this struct include a 'capacity' field (u32) that is set on each command?
It would be the number of bytes writable at address.
size would be the input length on entry (0 if the command passes none), and the
number of bytes produced on return (0 if none).
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-18 4:32 ` Kishen Maloor
@ 2026-09-18 5:58 ` Tony Lindgren
2026-09-21 0:13 ` Kishen Maloor
0 siblings, 1 reply; 78+ messages in thread
From: Tony Lindgren @ 2026-09-18 5:58 UTC (permalink / raw)
To: Kishen Maloor
Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini,
Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On Thu, Sep 17, 2026 at 09:32:23PM -0700, Kishen Maloor wrote:
> On 9/16/26 11:42 PM, Tony Lindgren wrote:
> > On Wed, Sep 16, 2026 at 08:31:32PM -0700, Kishen Maloor wrote:
> >> On 9/15/26 10:09 PM, Tony Lindgren wrote:
> >>> So trying to summarize the common flags for the role and separate vendor
> >>> flags:
> >>>
> >>> struct kvm_migrate_cmd {
> >>> __u16 command;
> >>> __u16 flags;
> >>> __u16 vflags;
> >>> __u16 reserved;
> >>> __u32 reserved;
> >>> struct kvm_transfer_buffer buf;
> >>> };
> >>>
> >>> Is the above along the lines what you were thinking?
> >>
> >> No, I was suggesting a 'role' field carved out of the 'reserved' space,
> >> like this:
> >>
> >> struct kvm_migrate_cmd {
> >> __u16 command;
> >> __u16 flags;
> >> __u8 role; /* 0 = unset, 1 = source, 2 = destination */
> >> __u8 reserved[3];
> >> struct kvm_transfer_buffer buf;
> >> };
> >
> > OK yes thanks for clarifying, that works for me.
> >
> >>> Ah OK, yes that would also tell "the hardware has been initialized to a
> >>> certain migration role". That seems like a usable common feature.
> >>
> >> Not quite. It tells us that userspace asserted a role for this VM's migration
> >> session. Whether a TD was created for import is a separate, vendor-level detail.
> >> The generic layer only needs the role to reject a session that never stated one,
> >> and to pick the export or import callback. That callback then knows which side
> >> it's on and can reject an incorrect role (e.g., if SETUP asserted dst for a src TD).
> >
> > That's a good point, the hardware role may not be set yet.
> >
> > I'm still wondering if there is a need to stash the userspace set role in
> > KVM though. Likely only the hardware specific code can properly track the
> > state of the hardware and adjust to the userspace requests. Seems just
> > being able to pass the role in struct kvm_migrate_cmd should be enough?
>
> Passing it in kvm_migrate_cmd is enough for SETUP itself, but the commands
> after SETUP like memory/vcpu transfers still have to reach the right
> callback. So KVM would need to remember what was asserted so that the
> generic layer can dispatch to the export or import facing callbacks.
> We've been sketching (on this thread) an alternative UAPI set
> (3 vs 5 ioctls) for consideration which this stored role enables:
>
> Proposed in the RFC Alternative
> KVM_MIGRATE_CMD KVM_MIGRATE_CMD
> KVM_EXPORT_MEMORY
> KVM_IMPORT_MEMORY KVM_TRANSFER_MEMORY
> KVM_EXPORT_VCPU
> KVM_IMPORT_VCPU KVM_TRANSFER_VCPU
>
> It's just a record (1 byte) of what userspace asserted for the current session
> at SETUP. Vendor code still owns the hardware state and remains free to reject a
> role that doesn't match it. It is also what lets the generic layer reject a
> command to a VM that never set up a session.
For the TRANSFER style operations, I would assume the direction is passed
for each transfer, just like the Linux does for the dmaengine. It's
possible that there may be transfers going both directions without the
role changing.
So looks like we have tree things to consider: userspace set migration
role, the hardware state, and transfer direction.
What if userspace always passes the role and transfer direction where it
makes sense? And then the hardware specific implementation tracks the
hardware state?
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests
2026-08-31 7:13 ` [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests Tony Lindgren
2026-08-31 7:20 ` sashiko-bot
@ 2026-09-18 11:35 ` Peter Xu
2026-09-21 4:20 ` Tony Lindgren
2026-09-24 1:50 ` Wei Wang
1 sibling, 2 replies; 78+ messages in thread
From: Peter Xu @ 2026-09-18 11:35 UTC (permalink / raw)
To: Tony Lindgren
Cc: Paolo Bonzini, Sean Christopherson, Artem Bityutskiy,
Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky,
Marc Zyngier, Oliver Upton, Steven Price, Anup Patel,
Samuel Ortiz, Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
On Mon, Aug 31, 2026 at 10:13:01AM +0300, Tony Lindgren wrote:
> +:Capability: KVM_CAP_LIVE_MIGRATION
IMHO this is slightly misleading, some "CONFIDENTIAL_" or other prefix
would be nice.
Thanks,
--
Peter Xu
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
2026-09-17 21:27 ` Peter Xu
@ 2026-09-18 12:46 ` Artem Bityutskiy
2026-09-18 15:53 ` Peter Xu
2026-09-20 23:56 ` Kishen Maloor
1 sibling, 1 reply; 78+ messages in thread
From: Artem Bityutskiy @ 2026-09-18 12:46 UTC (permalink / raw)
To: Peter Xu
Cc: Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas,
Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
Hi Peter,
thank you for your good comments and questions. A quick disclaimer before I
address them. In my answers I try to keep three distinct things separate:
1. The CoCo migration uAPI - the generic interface we ultimately want to
converge on. This is the end goal, not necessarily what these patches
propose. What is good for this uAPI is my priority at this point.
2. The TDX migration model - the mechanics offered by Intel TDX module for
TDX guest live migration today.
3. Our PoC - a concrete uAPI proposal and its example implementation for TDX
guests.
On Thu, 2026-09-17 at 17:27 -0400, Peter Xu wrote:
> > Some example high-level topics that would be nice to get feedback on:
> >
> > - Can we come up with a single set of generic migration uAPIs for different
> > CoCo models?
> > - Or should some uAPIs be generic while others are vendor-specific?
> > - Or should each CoCo model have its own vendor-specific set of migration
> > uAPIs?
>
> It's always good if we can put together as much function to be shared with
> generic ioctls as possible. At some point, IMHO we need to collect such
> information somehow, so when merging the generic API we know what vendor
> specific API will be needed. Hopefully this series is a good start.
Thanks, agreed.
On that note - does anyone know of a good doc describing the AMD, ARM, or
other CoCo migration models? My knowledge is limited to TDX, so it is hard
to tell what is common and what is TDX-specific.
FYI, I am working on a TDX migration model document. It describes what the
TDX module offers, but unlike the specs it is oriented towards software
engineers: much easier to read and it does not require deep TDX knowledge.
It is all based on public specs, just distilled into readable mental models.
I plan to publish it publicly. I am about 80% done.
> > - Should the same uAPIs also support traditional VMs? But the only use-case
> > I imagine here is "for testing purposes".
>
> This is an interesting idea, I think this could be useful. Especially, I
> wonder if you already have it done and PoC branches you can share, so that
> I can play with it.
We do not have code. But Kishen spent time playing with it, and I think he
concluded not to proceed with this. But he might have evaluated it from
the "unify all migration into a single generic API" perspective. But may
be as a "this is a test framework" perspective is different, at least I feel
it may be the case. I think Kishen can provide more insight if needed, he
is in CC.
> >
> > > Migration flow
> > > ==============
> > >
> > > Source host Destination host
> > > =========== ================
> > >
> > > CMD(SETUP/SESSION) <--- setup msgs ---> CMD(SETUP/SESSION)
> > > | (repeated) |
> > > CMD(SETUP/IMMUTABLE_STATE) - immutable state -> CMD(SETUP/IMMUTABLE_STATE)
> > > | |
> > > KVM_GET_DIRTY_LOG |
> > > KVM_EXPORT_MEMORY --- memory data ---> KVM_IMPORT_MEMORY
> > > CMD(ITERATION) --- epoch token ---> CMD(ITERATION)
>
> Could you elaborate this ITERATION operation? Is that something the
> userapp must do after full scan of a round of guest memory?
= Why no duplicate instances? =
First, let me answer the "why" question you asked further below: "why does
the TDX module require that only the source or only the destination runs,
never both?". Answering it first makes the tokens easier to understand. And
the tokens are why we proposed the "ITERATION" operation.
I believe this is not TDX-specific, it is a confidential computing
requirement. In short, cloning would give the VMM a very powerful primitive
to attack confidential VMs. Here are a couple of example attack approaches:
- When you attest a CoCo VM remotely, you get an assurance that you talk to
this one specific instance. If duplication were allowed, many instances
could exist, and that assurance is gone.
- With a clone you can security-upgrade the state of one copy, attest the
upgraded state, and then use the pre-upgraded copy: the user believes they
are working with an up-to-date CoCo VM, but in fact use an older, possibly
vulnerable version.
This is also why TDX migration requires that not only must the two copies
never run at the same time, but after migration the destination must be
exactly the same as the source. For example, the destination must not end up
using an older copy of a page.
= ITERATION Operation =
In the QEMU model, the pre-copy phase is a set of rounds:
1. Get the list of dirty pages.
2. Copy them to the destination.
3. Repeat until the convergence criteria are met.
ITERATION is the explicit uAPI that ends the current pre-copy round. There
is no equivalent uAPI for traditional VMs today.
In the TDX model, ending a round needs an extra step: the source generates
an epoch token, and the destination imports it. Two seamcalls do this:
- TDH.EXPORT.TRACK generates the epoch token on the source.
- TDH.IMPORT.TRACK imports it on the destination.
The token enforces integrity and ordering, for example:
- Every page exported on the source must be imported on the destination.
- Once a newer version of a page is imported, an older version can no longer
be imported.
The final round is special. It uses the "done" flag, which is passed to the
`TDH.EXPORT.TRACK` seamcall, and makes it export a special variant of the
epoch token that is called the start token. On top of the integrity and
ordering guarantees, the start token is what allows the destination to
start: until the destination imports it, the TDX module will not let the
destination TD run with partial state.
> > > | (repeat until convergence) |
> > > CMD(STOP_AND_COPY/PAUSE) |
> > > CMD(STOP_AND_COPY/TD_STATE) --- VM state ------> CMD(STOP_AND_COPY/TD_STATE)
> > > KVM_EXPORT_VCPU --- vCPU state ----> KVM_IMPORT_VCPU
> > > KVM_EXPORT_MEMORY -- final memory ---> KVM_IMPORT_MEMORY
>
> When read/write encrypted memories, two questions:
>
> - Is there an upper bound of the buffer size per-page?
So generally, the assumption is that exporting N pages requires M pages,
M > N, because there may be some metadata (e.g., MACs for integrity
checks). In case of TDX module, M is predictable and can be calculated in
advance.
I believe in our current PoC, the ioctl requires the buffer size to be large
enough to hold all the requested pages. But this is specific to our current
PoC implementation.
In general, I feel like if buffer size is not enough, the uAPI could fill it
with as much data as fits, and communicate back about what GPAs were
exported. The caller could export the rest separately.
Context: I am new in the Intel TDX live migration team, and did not
participate in TDX PoC, that's why I use "I believe". I am catching up. But
I assume others will (Tony, Kishen) will correct me if I am wrong.
> - Does this operation supports concurrency? If it supports, how well it
> scales per expectation (e.g. is there known big lock for that)?
From the TDX module perspective, parallel exports of different GPAs can run
on multiple CPUs, so I expect the QEMU multifd model to work and scale.
In our current PoC the ioctl does not take a VM-wide lock, and concurrency
is per-stream. I can expand on the stream concept if needed, but it is
exactly about parallel import/export of memory and vCPU state.
In our PoC we are focusing on the basics, but multifd support is definitely
a goal too, just later. Kishen was already prototyping it in QEMU.
From the uAPI point of view, I believe parallel export/import should be
allowed. If a specific CoCo VM has issues with that, it would need to
serialize the operations internally, I'd say.
> - Does this operation supports concurrency? If it supports, how well it
> scales per expectation (e.g. is there known big lock for that)?
>
> Similar question to the vCPU getter and setter. For now even without CoCo
> we serialize vCPU get/set, but I want to understand the potential of
> concurrent operations, and see if there's anything special for CoCo from
> that regard.
Similar to memory import/export: the TDX module explicitly allows vCPU state
to be exported and imported in parallel.
Our current PoC does not take advantage of this yet. vCPU export currently
grabs the KVM MMU write lock, so vCPU exports are serialized today. I think
Tony can comment more on the technical difficulties there.
From the uAPI point of view, I'd propose to allow concurrent vCPU
operations.
> >
> > It describes what the TDX module offers today and focuses on pre-copy
> > migration. This is our interpretation of the TDX specifications, not a
>
> IMHO we should really take postcopy into account when designing the API and
> state machine. We don't need to implement it in the first version, even
> until merging, but we need to make sure postcopy will be new ioctls on top
> of existing and it should have no major loopholes that it'll need a new set
> of APIs.
I totally agree. I have not yet dug into the TDX module implementation
details for post-copy, but I know it is supported and I know the basics. I
plan to study it in detail later. So far I have not noticed anything that
would prevent adding post-copy on top later.
>
> For example, I think we should consider KVM_EXPORT_MEMORY being usable
> after END on source, KVM_IMPORT_MEMORY while TD is in operation, etc. We
> should likely also need to still picture the rough process of postcopy,
> reserve those APIs since the start (but return -EINVAL or something).
Yes, agreed. I will spend more time looking at this. But at this point, I
just assumed that the proposed uAPIs can be used at the post-copy phase in
parallel with on-demand page delivery.
Just FYI, TDX module model allows for this, but we did not try it.
>
> AFAIU, postcopy is so far still the best solution for extremely large or
> extremely busy VMs regarding user experience, and it will happen to CoCo
> VMs one day or another.
Sure, thanks for sharing.
> > destination may run, but never both. In other words, cloning a TD is not
> > allowed.
> > - When migration completes, the destination must have the same memory and
> > vCPU state as the source. It must not end up with a partial or mixed
> > state.
>
> If such happens, it's definitely a bug, even without CoCo. Anything
> specific about CoCo? Like, whole-VM checksum?
Well, in CoCo VMs it is not just a bug, it is something the CoCo framework
needs to make impossible, because VMM is considered to be untrusted, it can
try to manipulate things and half-migrate, use it not as a bug but as attack
vector. In TDX case, the TDX module will not allow you to run the TD - the
TDH.VP.ENTER seamcall will fail.
Regarding checksums: there is no single whole-VM checksum in TDX migration
model. Instead integrity is enforced continuously - every exported blob
carries a MACs that the destination TDX module verifies on import, and the
epoch/start tokens guarantee that everything was imported, in order.
> I want to understand what is extra for a CoCo VM in terms of "pause", say,
> what's more than "stopping the vCPU threads".
TDX module guarantees the source won't run, even if VMM tries, the
TDH.VP.ENTER seamcall will fail. So the source TD state is effectively
frozen and cannot be modified by the VMM.
>
> I saw there's mention of PRE_COPY_STOP state. One example question is,
> when reaching this state, can the guest memory still change? What happens
> if some emulated device are still DMAing to the guest memory (assuming
> flipped from private to shared)? In case of future IO zone support, what
> happens if in case of VFIO-PCI assigned doing encrypted DMA?
>
> From that regard, VFIO has the P2P state where it quiesce initiation of any
> DMA from this specific device, then another round to fully stop all devices
> into STOP_COPY phase. I wonder if CoCo VMs need similar treatment.
Let me split this by device type, because TDX treats them very differently.
Emulated (virtio-net, virtio-blk, etc.) only use shared memory - they cannot
read or DMA into TD private memory. So full device state lives in shared
memory, and QEMU migrates them exactly the same way as for a traditional VM.
This is entirely outside the TDX module migration model and outside the
proposed uAPI - the uAPI is only for TD private memory.
Directly assigned devices are only possible with TDX Connect, where a
physical PCIe/CXL function (a "TDI") is assigned to the TD and can DMA into
private memory over a cryptographically protected link. This is not
implemented in Linux yet. For migration, the TDX module requires all TDIs to
be unassigned before the source TD is paused - the TDH.EXPORT.PAUSE seamcall
actually checks this. Unassigning a TDI tears down its whole TD-private
footprint (MMIO unmapped from the Secure EPT, trusted DMA mappings removed),
so no device-specific state is left to migrate. From the TD's point of view
it is a full hot-unplug on the source and a fresh hot-plug on the
destination.
Our current TDX guest migration PoC is built on this assumption.
> Could you elaborate what's the relations between TDH.MEM.SCAN.RANGE and the
> GET_DIRTY_LOG ioctl? I recall above mentioned GET_DIRTY_LOG will be
> available even for CoCo, which makes sense assuming dirty information isn't
> confidential. However then I don't understand what TDH.MEM.SCAN.RANGE
> plays the role here.
A few things here.
First, our plan is that the standard GET_DIRTY_LOG ioctl is backed by the
`TDH.MEM.SCAN.RANGE` seamcall - that is how dirty tracking is implemented
for a TD. So we do not propose any special uAPI for dirty tracking.
And I think you are right that the dirty information is not confidential -
the TDX module exposes dirty page information to the VMM.
Second, FYI, in current TDX migration model, dirty page scanning
(`TDH.MEM.SCAN.RANGE`) is only allowed during a migration session. The TDX
module returns an error if the seamcall is issued before the session is set
up (i.e. before the source and destination TDX modules have exchanged the
migration key and established trust - what the SETUP command does in the
proposed uAPI).
In other words, with our PoC, if someone tries to use GET_DIRTY_LOG without
going through the migration setup - the ioctl will return an error.
But I wish it were an independent feature instead. Then we could work on
upstreaming it on its own - Tony estimates it is about 20% of the current
TDX migration PoC code.
I already raised this with the Intel TDX module architects, and they asked
for use-cases. The only one we came up with is QEMU estimating the TD dirty
rate before starting migration (the calc_dirty_rate command). I understand
their position: without a use-case there is little reason to implement it,
and they would also need to study the security implications - can it help
an attacker in some way?
So if you or anyone else can educate me about use-cases for independent
dirty page scanning, I would really appreciate it - I could take them back
to the TDX module architects.
Thanks,
Artem.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
2026-09-18 12:46 ` Artem Bityutskiy
@ 2026-09-18 15:53 ` Peter Xu
2026-09-22 8:09 ` Artem Bityutskiy
0 siblings, 1 reply; 78+ messages in thread
From: Peter Xu @ 2026-09-18 15:53 UTC (permalink / raw)
To: Artem Bityutskiy
Cc: Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas,
Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
On Fri, Sep 18, 2026 at 03:46:32PM +0300, Artem Bityutskiy wrote:
> Hi Peter,
Hi, Artem,
>
> thank you for your good comments and questions. A quick disclaimer before I
> address them. In my answers I try to keep three distinct things separate:
>
> 1. The CoCo migration uAPI - the generic interface we ultimately want to
> converge on. This is the end goal, not necessarily what these patches
> propose. What is good for this uAPI is my priority at this point.
> 2. The TDX migration model - the mechanics offered by Intel TDX module for
> TDX guest live migration today.
> 3. Our PoC - a concrete uAPI proposal and its example implementation for TDX
> guests.
>
> On Thu, 2026-09-17 at 17:27 -0400, Peter Xu wrote:
> > > Some example high-level topics that would be nice to get feedback on:
> > >
> > > - Can we come up with a single set of generic migration uAPIs for different
> > > CoCo models?
> > > - Or should some uAPIs be generic while others are vendor-specific?
> > > - Or should each CoCo model have its own vendor-specific set of migration
> > > uAPIs?
> >
> > It's always good if we can put together as much function to be shared with
> > generic ioctls as possible. At some point, IMHO we need to collect such
> > information somehow, so when merging the generic API we know what vendor
> > specific API will be needed. Hopefully this series is a good start.
>
> Thanks, agreed.
>
> On that note - does anyone know of a good doc describing the AMD, ARM, or
> other CoCo migration models? My knowledge is limited to TDX, so it is hard
> to tell what is common and what is TDX-specific.
>
> FYI, I am working on a TDX migration model document. It describes what the
> TDX module offers, but unlike the specs it is oriented towards software
> engineers: much easier to read and it does not require deep TDX knowledge.
> It is all based on public specs, just distilled into readable mental models.
> I plan to publish it publicly. I am about 80% done.
That will be very useful, thanks for doing this. I'll be more than happy to
read it when it's done. I wonder if we can, after TDX bits done, use this
doc as a base for others to add in theirs in separate tabs, keeping
everything together for the unified migration API work.
Out of pure curiosity, I also don't know where s390 stands; it's almost not
mentioned in the current API plan.
>
> > > - Should the same uAPIs also support traditional VMs? But the only use-case
> > > I imagine here is "for testing purposes".
> >
> > This is an interesting idea, I think this could be useful. Especially, I
> > wonder if you already have it done and PoC branches you can share, so that
> > I can play with it.
>
> We do not have code. But Kishen spent time playing with it, and I think he
> concluded not to proceed with this. But he might have evaluated it from
> the "unify all migration into a single generic API" perspective. But may
> be as a "this is a test framework" perspective is different, at least I feel
> it may be the case. I think Kishen can provide more insight if needed, he
> is in CC.
Yes, thanks. I'm willing to hear more, and I'm a bit surprised that
non-coco migration didn't fit already well into it, because IIUC non-coco
needs less in this case, not more.
>
> > >
> > > > Migration flow
> > > > ==============
> > > >
> > > > Source host Destination host
> > > > =========== ================
> > > >
> > > > CMD(SETUP/SESSION) <--- setup msgs ---> CMD(SETUP/SESSION)
> > > > | (repeated) |
> > > > CMD(SETUP/IMMUTABLE_STATE) - immutable state -> CMD(SETUP/IMMUTABLE_STATE)
> > > > | |
> > > > KVM_GET_DIRTY_LOG |
> > > > KVM_EXPORT_MEMORY --- memory data ---> KVM_IMPORT_MEMORY
> > > > CMD(ITERATION) --- epoch token ---> CMD(ITERATION)
> >
> > Could you elaborate this ITERATION operation? Is that something the
> > userapp must do after full scan of a round of guest memory?
>
> = Why no duplicate instances? =
>
> First, let me answer the "why" question you asked further below: "why does
> the TDX module require that only the source or only the destination runs,
> never both?". Answering it first makes the tokens easier to understand. And
> the tokens are why we proposed the "ITERATION" operation.
>
> I believe this is not TDX-specific, it is a confidential computing
> requirement. In short, cloning would give the VMM a very powerful primitive
> to attack confidential VMs. Here are a couple of example attack approaches:
>
> - When you attest a CoCo VM remotely, you get an assurance that you talk to
> this one specific instance. If duplication were allowed, many instances
> could exist, and that assurance is gone.
> - With a clone you can security-upgrade the state of one copy, attest the
> upgraded state, and then use the pre-upgraded copy: the user believes they
> are working with an up-to-date CoCo VM, but in fact use an older, possibly
> vulnerable version.
>
> This is also why TDX migration requires that not only must the two copies
> never run at the same time, but after migration the destination must be
> exactly the same as the source. For example, the destination must not end up
> using an older copy of a page.
>
> = ITERATION Operation =
>
> In the QEMU model, the pre-copy phase is a set of rounds:
> 1. Get the list of dirty pages.
> 2. Copy them to the destination.
> 3. Repeat until the convergence criteria are met.
>
> ITERATION is the explicit uAPI that ends the current pre-copy round. There
> is no equivalent uAPI for traditional VMs today.
>
> In the TDX model, ending a round needs an extra step: the source generates
> an epoch token, and the destination imports it. Two seamcalls do this:
>
> - TDH.EXPORT.TRACK generates the epoch token on the source.
> - TDH.IMPORT.TRACK imports it on the destination.
>
> The token enforces integrity and ordering, for example:
>
> - Every page exported on the source must be imported on the destination.
> - Once a newer version of a page is imported, an older version can no longer
> be imported.
>
> The final round is special. It uses the "done" flag, which is passed to the
> `TDH.EXPORT.TRACK` seamcall, and makes it export a special variant of the
> epoch token that is called the start token. On top of the integrity and
> ordering guarantees, the start token is what allows the destination to
> start: until the destination imports it, the TDX module will not let the
> destination TD run with partial state.
The methodology on unique VM attestation sounds comlex, but I think I get
it now, thank you. It'll be nice if some of these reasonings will also be
there in the doc you're drafting.
>
> > > > | (repeat until convergence) |
> > > > CMD(STOP_AND_COPY/PAUSE) |
> > > > CMD(STOP_AND_COPY/TD_STATE) --- VM state ------> CMD(STOP_AND_COPY/TD_STATE)
> > > > KVM_EXPORT_VCPU --- vCPU state ----> KVM_IMPORT_VCPU
> > > > KVM_EXPORT_MEMORY -- final memory ---> KVM_IMPORT_MEMORY
> >
> > When read/write encrypted memories, two questions:
> >
> > - Is there an upper bound of the buffer size per-page?
>
> So generally, the assumption is that exporting N pages requires M pages,
> M > N, because there may be some metadata (e.g., MACs for integrity
> checks). In case of TDX module, M is predictable and can be calculated in
> advance.
I am just thinking out loud here: if the hardware is good enough to do
encryption plus (some?) compression, that would be very nice. For "some",
I meant minimum over zero pages. Because in this case even if host wants
to play tricks with zero pages, it can't anymore when un-readable. Maybe
the guest driver can play some trick, but I'm also not sure if in CoCo.
M > N may imply it's not the case for now, but it's still sane as a start
even if so.
>
> I believe in our current PoC, the ioctl requires the buffer size to be large
> enough to hold all the requested pages. But this is specific to our current
> PoC implementation.
>
> In general, I feel like if buffer size is not enough, the uAPI could fill it
> with as much data as fits, and communicate back about what GPAs were
> exported. The caller could export the rest separately.
Yes, this will work.
Or maybe it's simpler to be able to export an upper bound in another API
that probes it (some KVM cap)?
Any retry is a wasted round trip from perf perspective. The upper bound
can be relatively large, IMHO, which should be non-issue. It should be
simpler for both userapp and kernel if feasible.
>
> Context: I am new in the Intel TDX live migration team, and did not
> participate in TDX PoC, that's why I use "I believe". I am catching up. But
> I assume others will (Tony, Kishen) will correct me if I am wrong.
No worries, thanks for the detailed answers whatever offered; they're
already very helpful.
>
> > - Does this operation supports concurrency? If it supports, how well it
> > scales per expectation (e.g. is there known big lock for that)?
>
> From the TDX module perspective, parallel exports of different GPAs can run
> on multiple CPUs, so I expect the QEMU multifd model to work and scale.
>
> In our current PoC the ioctl does not take a VM-wide lock, and concurrency
> is per-stream. I can expand on the stream concept if needed, but it is
> exactly about parallel import/export of memory and vCPU state.
>
> In our PoC we are focusing on the basics, but multifd support is definitely
> a goal too, just later. Kishen was already prototyping it in QEMU.
Great.
>
> From the uAPI point of view, I believe parallel export/import should be
> allowed. If a specific CoCo VM has issues with that, it would need to
> serialize the operations internally, I'd say.
Yes, it would be good to keep the critical section as small as possible in
this case. As long as we are fully aware of the concurrent use model from
userapp then it's good enough for now.
>
> > - Does this operation supports concurrency? If it supports, how well it
> > scales per expectation (e.g. is there known big lock for that)?
> >
> > Similar question to the vCPU getter and setter. For now even without CoCo
> > we serialize vCPU get/set, but I want to understand the potential of
> > concurrent operations, and see if there's anything special for CoCo from
> > that regard.
>
> Similar to memory import/export: the TDX module explicitly allows vCPU state
> to be exported and imported in parallel.
>
> Our current PoC does not take advantage of this yet. vCPU export currently
> grabs the KVM MMU write lock, so vCPU exports are serialized today. I think
> Tony can comment more on the technical difficulties there.
>
> From the uAPI point of view, I'd propose to allow concurrent vCPU
> operations.
Sounds good.
>
> > >
> > > It describes what the TDX module offers today and focuses on pre-copy
> > > migration. This is our interpretation of the TDX specifications, not a
> >
> > IMHO we should really take postcopy into account when designing the API and
> > state machine. We don't need to implement it in the first version, even
> > until merging, but we need to make sure postcopy will be new ioctls on top
> > of existing and it should have no major loopholes that it'll need a new set
> > of APIs.
>
> I totally agree. I have not yet dug into the TDX module implementation
> details for post-copy, but I know it is supported and I know the basics. I
> plan to study it in detail later. So far I have not noticed anything that
> would prevent adding post-copy on top later.
As long as we have that in mind across working on this, that's good enough,
thanks.
>
> >
> > For example, I think we should consider KVM_EXPORT_MEMORY being usable
> > after END on source, KVM_IMPORT_MEMORY while TD is in operation, etc. We
> > should likely also need to still picture the rough process of postcopy,
> > reserve those APIs since the start (but return -EINVAL or something).
>
> Yes, agreed. I will spend more time looking at this. But at this point, I
> just assumed that the proposed uAPIs can be used at the post-copy phase in
> parallel with on-demand page delivery.
>
> Just FYI, TDX module model allows for this, but we did not try it.
>
> >
> > AFAIU, postcopy is so far still the best solution for extremely large or
> > extremely busy VMs regarding user experience, and it will happen to CoCo
> > VMs one day or another.
>
> Sure, thanks for sharing.
>
> > > destination may run, but never both. In other words, cloning a TD is not
> > > allowed.
> > > - When migration completes, the destination must have the same memory and
> > > vCPU state as the source. It must not end up with a partial or mixed
> > > state.
> >
> > If such happens, it's definitely a bug, even without CoCo. Anything
> > specific about CoCo? Like, whole-VM checksum?
>
> Well, in CoCo VMs it is not just a bug, it is something the CoCo framework
> needs to make impossible, because VMM is considered to be untrusted, it can
> try to manipulate things and half-migrate, use it not as a bug but as attack
> vector. In TDX case, the TDX module will not allow you to run the TD - the
> TDH.VP.ENTER seamcall will fail.
>
> Regarding checksums: there is no single whole-VM checksum in TDX migration
> model. Instead integrity is enforced continuously - every exported blob
> carries a MACs that the destination TDX module verifies on import, and the
> epoch/start tokens guarantee that everything was imported, in order.
>
> > I want to understand what is extra for a CoCo VM in terms of "pause", say,
> > what's more than "stopping the vCPU threads".
>
> TDX module guarantees the source won't run, even if VMM tries, the
> TDH.VP.ENTER seamcall will fail. So the source TD state is effectively
> frozen and cannot be modified by the VMM.
IIUC this should work for QEMU.
Said that, we'll need to be careful then in case of migration fallbacks at
the final stage. Nowadays, I believe QEMU can still fallback to source side
at a very, very late stage after all things applied. If I'm not mistaken,
the final handshake is done at migration_incoming_state_destroy() ->
migrate_send_rp_shut() telling source to be gone.
After reading above, one thing we may want to make sure is TDX ENTER on
dest be exactly the last thing to do on destination, rather than dest QEMU
ENTER done then something else seems wrong, then dest can't fallback
anymore. I didn't check into details, though, more of a heads-up to
whoever is working on QEMU for this in case useful.
>
> >
> > I saw there's mention of PRE_COPY_STOP state. One example question is,
> > when reaching this state, can the guest memory still change? What happens
> > if some emulated device are still DMAing to the guest memory (assuming
> > flipped from private to shared)? In case of future IO zone support, what
> > happens if in case of VFIO-PCI assigned doing encrypted DMA?
> >
> > From that regard, VFIO has the P2P state where it quiesce initiation of any
> > DMA from this specific device, then another round to fully stop all devices
> > into STOP_COPY phase. I wonder if CoCo VMs need similar treatment.
>
> Let me split this by device type, because TDX treats them very differently.
>
> Emulated (virtio-net, virtio-blk, etc.) only use shared memory - they cannot
> read or DMA into TD private memory. So full device state lives in shared
> memory, and QEMU migrates them exactly the same way as for a traditional VM.
> This is entirely outside the TDX module migration model and outside the
> proposed uAPI - the uAPI is only for TD private memory.
I hope I understand it right, that all shared pages are out of the secure
zone TD manages, during migration or not (hence the same as some random
page the VMM has allocated)? If so, anything about post_load operations
shouldn't be any concern, and should work as usual.
>
> Directly assigned devices are only possible with TDX Connect, where a
> physical PCIe/CXL function (a "TDI") is assigned to the TD and can DMA into
> private memory over a cryptographically protected link. This is not
> implemented in Linux yet. For migration, the TDX module requires all TDIs to
> be unassigned before the source TD is paused - the TDH.EXPORT.PAUSE seamcall
> actually checks this. Unassigning a TDI tears down its whole TD-private
> footprint (MMIO unmapped from the Secure EPT, trusted DMA mappings removed),
> so no device-specific state is left to migrate. From the TD's point of view
> it is a full hot-unplug on the source and a fresh hot-plug on the
> destination.
Hmm, interesting. Sorry if this follow up question may be slightly
off-topic of the new API, or maybe it matters, depends on the answer: do
you know who is in charge of this unplug / plug operation? VMM or TD
(hence, transparent to VMM)?
If it's VMM, is QEMU involved?
>
> Our current TDX guest migration PoC is built on this assumption.
>
> > Could you elaborate what's the relations between TDH.MEM.SCAN.RANGE and the
> > GET_DIRTY_LOG ioctl? I recall above mentioned GET_DIRTY_LOG will be
> > available even for CoCo, which makes sense assuming dirty information isn't
> > confidential. However then I don't understand what TDH.MEM.SCAN.RANGE
> > plays the role here.
>
> A few things here.
>
> First, our plan is that the standard GET_DIRTY_LOG ioctl is backed by the
> `TDH.MEM.SCAN.RANGE` seamcall - that is how dirty tracking is implemented
> for a TD. So we do not propose any special uAPI for dirty tracking.
>
> And I think you are right that the dirty information is not confidential -
> the TDX module exposes dirty page information to the VMM.
>
> Second, FYI, in current TDX migration model, dirty page scanning
> (`TDH.MEM.SCAN.RANGE`) is only allowed during a migration session. The TDX
> module returns an error if the seamcall is issued before the session is set
> up (i.e. before the source and destination TDX modules have exchanged the
> migration key and established trust - what the SETUP command does in the
> proposed uAPI).
>
> In other words, with our PoC, if someone tries to use GET_DIRTY_LOG without
> going through the migration setup - the ioctl will return an error.
>
> But I wish it were an independent feature instead. Then we could work on
> upstreaming it on its own - Tony estimates it is about 20% of the current
> TDX migration PoC code.
Unexpected, that's a lot just for tracking.
>
> I already raised this with the Intel TDX module architects, and they asked
> for use-cases. The only one we came up with is QEMU estimating the TD dirty
> rate before starting migration (the calc_dirty_rate command). I understand
> their position: without a use-case there is little reason to implement it,
> and they would also need to study the security implications - can it help
> an attacker in some way?
>
> So if you or anyone else can educate me about use-cases for independent
> dirty page scanning, I would really appreciate it - I could take them back
> to the TDX module architects.
Yes, calc_dirty_rate will use it, I also can't think of another use case
that needs it. RHEL supports it, so it would be nice this will be
supported for CoCo too.
QEMU also has another way to do it without KVM tracking, which is kind of
simle page hash daemon fully done in userspace. But in case of CoCo it'll
stop working too when most memory unreadable. So GET_DIRTY_LOG seems the
only way to go.
We can still "emulate" it by initiating a remote migration just to collect
this info, I believe all attestations will simply pass and we do a fallback
when the admin turns it off.. but it's awkward and all the rest (including
trying to find another host suitable for migration, allocate resources..)
are pure wastes.
Btw, do you know if the private pages will be write-trappable by the host
kernel, with things like userfaultfd-wp or soft-dirty? That'll be another
way to do this without TD involvement, but I don't know well enough to say.
It just sounds like if it's doable in KVM, it's still the best place, also
since GET_DIRTY_LOG is available as long as memslot marked tracking, it
sounds good to keep CoCo be compatible with that API if it is to be reused,
hence GET_DIRTY_LOG API is consistent across coco / non-coco.
Thanks,
--
Peter Xu
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
2026-08-31 7:13 [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Tony Lindgren
` (4 preceding siblings ...)
2026-09-04 18:24 ` [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Artem Bityutskiy
@ 2026-09-18 18:36 ` Ionut Mihalcea
2026-09-21 4:35 ` Tony Lindgren
2026-09-25 16:03 ` Serge Hallyn (AMD)
6 siblings, 1 reply; 78+ messages in thread
From: Ionut Mihalcea @ 2026-09-18 18:36 UTC (permalink / raw)
To: Tony Lindgren, Paolo Bonzini, Sean Christopherson
Cc: Peter Xu, Artem Bityutskiy, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Jörg Rödel ,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm@vger.kernel.org, Mathias Brossard, nd
On 31 Aug 2026, at 08:13, Tony Lindgren <tony.lindgren@linux.intel.com> wrote:
> Tom, since you mentioned that AMD SEV-SNP and Intel TDX live migration
> sound similar, can you please take a look how the API might work for
> SEV-SNP?
While the x86 implementation is outside our scope, we're also currently
investigating how the Arm CCA ABIs for Live Migration ABIs would fit as a
backend for this API. We will provide more feedback on the API shape and
semantics as the analysis progresses.
As an intro to our approach (and a potential touchpoint on the topic of this
API), Mathias Brossard will be presenting Arm CCA Live Migration at LPC next
month [0].
> Dirty page tracking does not add a new uAPI. Userspace keeps using
> KVM_GET_DIRTY_LOG and KVM_CLEAR_DIRTY_LOG.
One concern we currently have is about an impedance mismatch between the
non-CoCo dirty log APIs and our approach to dirty reporting. This is just
to highlight the concern at the moment, without any explicit proposal as
an alternative.
[0] https://lpc.events/event/20/contributions/2476/
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
2026-09-17 21:27 ` Peter Xu
2026-09-18 12:46 ` Artem Bityutskiy
@ 2026-09-20 23:56 ` Kishen Maloor
2026-09-23 21:36 ` Peter Xu
1 sibling, 1 reply; 78+ messages in thread
From: Kishen Maloor @ 2026-09-20 23:56 UTC (permalink / raw)
To: Peter Xu, Artem Bityutskiy
Cc: Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas,
Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg,
Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm
Hi Peter,
Thank you for your comments. Just adding a few other details to
complement Artem's response.
On 9/17/26 2:27 PM, Peter Xu wrote:
> On Fri, Sep 04, 2026 at 09:24:25PM +0300, Artem Bityutskiy wrote:
>> On Mon, 2026-08-31 at 10:13 +0300, Tony Lindgren wrote:
> ...
>> - Should the same uAPIs also support traditional VMs? But the only use-case
>> I imagine here is "for testing purposes".
>
> This is an interesting idea, I think this could be useful. Especially, I
> wonder if you already have it done and PoC branches you can share, so that
> I can play with it.
I had once attempted this and hit a breaking case while migrating regular VMs.
I didn't dig further since the primary use of this UAPI set is with CoCo VMs.
Here's what I gathered: KVM-mediated transfer needs KVM to resolve a GFN to the
HVA where the data lands, and on x86 QEMU mutates that mapping at runtime in a
way that is opaque to KVM. For the PAM window QEMU overlays a separate region on
top of pc.ram, and KVM sees only the flattened view, the winning memslot, with
no way to address what that slot shadows. On the source, firmware reprograms PAM
and QEMU drops the overlay. On the destination the firmware never runs, so it still
has its boot-time layout and the overlay is still in place. So importing, say, GFN 0xc0
lands on the overlay's backing store, which is a read-only memslot, and the GFN->HVA
translation fails. I suppose even if it were writable, the data would land there
rather than in the pc.ram underneath where it belongs. Without mediation, QEMU is
able to write to the HVA directly, so the problem doesn't arise.
Generally, I think KVM mediation only works if KVM's view of memory is authoritative.
TDX skirts this issue entirely because QEMU doesn't create overlays for TDX VMs.
> ...
> Could you elaborate this ITERATION operation? Is that something the
> userapp must do after full scan of a round of guest memory?
Essentially, yes. TDX migration architecture delimits such pre-copy rounds as
"migration epochs" and emits an epoch token at each round boundary that needs to be
consumed on the destination. It allows the TDX module to verify that everything from
the prior round has been received at the destination and that two versions of the
same GPA aren't sent in the same epoch. It is userspace that decides where a round
ends, but any deviation from this model and migration would fail.
>
>>> | (repeat until convergence) |
>>> CMD(STOP_AND_COPY/PAUSE) |
>>> CMD(STOP_AND_COPY/TD_STATE) --- VM state ------> CMD(STOP_AND_COPY/TD_STATE)
>>> KVM_EXPORT_VCPU --- vCPU state ----> KVM_IMPORT_VCPU
>>> KVM_EXPORT_MEMORY -- final memory ---> KVM_IMPORT_MEMORY
>
> When read/write encrypted memories, two questions:
>
> - Is there an upper bound of the buffer size per-page?
Yes. For TDX a 4KB guest page produces exactly 4KB of encrypted payload plus a small
fixed amount of ancillary data. The bound can be pre-computed and userspace
can size its buffers accordingly. The UAPI itself doesn't impose one. The vendor
implementation decides how pages and ancillary data are laid out.
>
> - Does this operation supports concurrency? If it supports, how well it
> scales per expectation (e.g. is there known big lock for that)?
Yes, these operations can support multifd and kvm_memory_transfer carries the
channel ID. The encryption/decryption work is per-channel, so it scales as multifd
does. The one serialization point is the epoch boundary, i.e. the ITERATION call,
which has to drain in-flight transfers for that round. That's inherent to the epoch
model rather than a lock that could be dropped.
> ...
> IMHO we should really take postcopy into account when designing the API and
> state machine. We don't need to implement it in the first version, even
> until merging, but we need to make sure postcopy will be new ioctls on top
> of existing and it should have no major loopholes that it'll need a new set
> of APIs.
Agreed. We honestly haven't looked at post-copy in depth, but supporting it thus
far appears to be additive. One item that comes to mind is a new command in KVM_MIGRATE_CMD
to mark the post-copy switchover point. On the source, vendor code could send all
previously exported and dirty pages (deferring all un-exported pages to post-copy) and
prepare to serve pages on demand. On the destination, vendor code could make the VM
runnable with incomplete memory. This is based on a preliminary assessment though.
> For example, I think we should consider KVM_EXPORT_MEMORY being usable
> after END on source, KVM_IMPORT_MEMORY while TD is in operation, etc. We
Yes, though I think some of these details could be handled in the vendor
implementation -- e.g., KVM_EXPORT/IMPORT_MEMORY might not need to know that
they're servicing post-copy.
> should likely also need to still picture the rough process of postcopy,
> reserve those APIs since the start (but return -EINVAL or something).
Agreed. We shall attempt to sketch that flow. Your thoughts and feedback would
be very helpful as we work through this.
> ...
> Could you elaborate what's the relations between TDH.MEM.SCAN.RANGE and the
> GET_DIRTY_LOG ioctl? I recall above mentioned GET_DIRTY_LOG will be
> available even for CoCo, which makes sense assuming dirty information isn't
> confidential. However then I don't understand what TDH.MEM.SCAN.RANGE
> plays the role here.
GET_DIRTY_LOG is the UAPI, but the slot dirty bitmaps need to be populated somehow and
TDH.MEM.SCAN.RANGE fills that gap for TDX. It is a SEAMCALL that scans the SEPT upon
request to return the list of currently dirty pages. We wanted to avoid inventing a new
UAPI for this and GET_DIRTY_LOG seemed to be a natural fit. In our current PoC, we've added
a barebones kvm_x86_ops hook that is plugged into kvm_arch_sync_dirty_log(). The TDX
implementation of that hook invokes the scans and records the dirty pages into KVM's slot
dirty bitmaps that GET_DIRTY_LOG serves up to userspace (as usual).
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-18 5:58 ` Tony Lindgren
@ 2026-09-21 0:13 ` Kishen Maloor
2026-09-21 6:52 ` Tony Lindgren
0 siblings, 1 reply; 78+ messages in thread
From: Kishen Maloor @ 2026-09-21 0:13 UTC (permalink / raw)
To: Tony Lindgren
Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini,
Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On 9/17/26 10:58 PM, Tony Lindgren wrote:
> On Thu, Sep 17, 2026 at 09:32:23PM -0700, Kishen Maloor wrote:
>> On 9/16/26 11:42 PM, Tony Lindgren wrote:
>>> On Wed, Sep 16, 2026 at 08:31:32PM -0700, Kishen Maloor wrote:
>>>> On 9/15/26 10:09 PM, Tony Lindgren wrote:
>>>>> So trying to summarize the common flags for the role and separate vendor
>>>>> flags:
>>>>>
>>>>> struct kvm_migrate_cmd {
>>>>> __u16 command;
>>>>> __u16 flags;
>>>>> __u16 vflags;
>>>>> __u16 reserved;
>>>>> __u32 reserved;
>>>>> struct kvm_transfer_buffer buf;
>>>>> };
>>>>>
>>>>> Is the above along the lines what you were thinking?
>>>>
>>>> No, I was suggesting a 'role' field carved out of the 'reserved' space,
>>>> like this:
>>>>
>>>> struct kvm_migrate_cmd {
>>>> __u16 command;
>>>> __u16 flags;
>>>> __u8 role; /* 0 = unset, 1 = source, 2 = destination */
>>>> __u8 reserved[3];
>>>> struct kvm_transfer_buffer buf;
>>>> };
>>>
>>> OK yes thanks for clarifying, that works for me.
>>>
>>>>> Ah OK, yes that would also tell "the hardware has been initialized to a
>>>>> certain migration role". That seems like a usable common feature.
>>>>
>>>> Not quite. It tells us that userspace asserted a role for this VM's migration
>>>> session. Whether a TD was created for import is a separate, vendor-level detail.
>>>> The generic layer only needs the role to reject a session that never stated one,
>>>> and to pick the export or import callback. That callback then knows which side
>>>> it's on and can reject an incorrect role (e.g., if SETUP asserted dst for a src TD).
>>>
>>> That's a good point, the hardware role may not be set yet.
>>>
>>> I'm still wondering if there is a need to stash the userspace set role in
>>> KVM though. Likely only the hardware specific code can properly track the
>>> state of the hardware and adjust to the userspace requests. Seems just
>>> being able to pass the role in struct kvm_migrate_cmd should be enough?
>>
>> Passing it in kvm_migrate_cmd is enough for SETUP itself, but the commands
>> after SETUP like memory/vcpu transfers still have to reach the right
>> callback. So KVM would need to remember what was asserted so that the
>> generic layer can dispatch to the export or import facing callbacks.
>> We've been sketching (on this thread) an alternative UAPI set
>> (3 vs 5 ioctls) for consideration which this stored role enables:
>>
>> Proposed in the RFC Alternative
>> KVM_MIGRATE_CMD KVM_MIGRATE_CMD
>> KVM_EXPORT_MEMORY
>> KVM_IMPORT_MEMORY KVM_TRANSFER_MEMORY
>> KVM_EXPORT_VCPU
>> KVM_IMPORT_VCPU KVM_TRANSFER_VCPU
>>
>> It's just a record (1 byte) of what userspace asserted for the current session
>> at SETUP. Vendor code still owns the hardware state and remains free to reject a
>> role that doesn't match it. It is also what lets the generic layer reject a
>> command to a VM that never set up a session.
>
> For the TRANSFER style operations, I would assume the direction is passed
> for each transfer, just like the Linux does for the dmaengine. It's
> possible that there may be transfers going both directions without the
> role changing.
>
> So looks like we have tree things to consider: userspace set migration
> role, the hardware state, and transfer direction.
Two of those three I agree with: userspace passes the role at SETUP, and
the vendor implementation tracks the hardware state. It's the per-transfer
direction I don't think we need.
> What if userspace always passes the role and transfer direction where it
> makes sense? And then the hardware specific implementation tracks the
> hardware state?
Unless there's another meaning, direction would indicate that an operation
must produce or consume a blob into/from a buffer. Such an indication would be
necessary if the layer below the API cannot know what to do with a buffer,
which may well be the case in your dmaengine analogy. However, in this case a
vendor implementation sits below the generic layer and could derive what it
needs to do from the role, command, and any session state it maintains. The
TDH.EXPORT.ABORT example I cited upthread expects a token only once the
session has left its pre-copy phase, so what to do with the buffer follows
from state the vendor layer already holds.
The gap I do see is in how the buffer itself is described: we have no way to
express a command that takes an input and produces an output. A direction flag
doesn't help there either, since it can only say one thing. My comments on patch 2
suggest a new 'capacity' field alongside a reframing of 'size' in kvm_transfer_buffer,
where capacity bounds what the kernel may write and size reports what is actually
there. That covers the both-ways case, and it also leaves nothing for a direction
field to convey. Maybe that addresses your concern?
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests
2026-09-18 11:35 ` Peter Xu
@ 2026-09-21 4:20 ` Tony Lindgren
2026-09-24 1:50 ` Wei Wang
1 sibling, 0 replies; 78+ messages in thread
From: Tony Lindgren @ 2026-09-21 4:20 UTC (permalink / raw)
To: Peter Xu
Cc: Paolo Bonzini, Sean Christopherson, Artem Bityutskiy,
Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky,
Marc Zyngier, Oliver Upton, Steven Price, Anup Patel,
Samuel Ortiz, Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
Hi,
On Fri, Sep 18, 2026 at 07:35:39AM -0400, Peter Xu wrote:
> On Mon, Aug 31, 2026 at 10:13:01AM +0300, Tony Lindgren wrote:
> > +:Capability: KVM_CAP_LIVE_MIGRATION
>
> IMHO this is slightly misleading, some "CONFIDENTIAL_" or other prefix
> would be nice.
OK. For the possible prefixes to consider, also HW for hardware specific
ops might work. I'm not aware of kernel side migration needs for this
other than CoCo though.
Regards,
Tony
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
2026-09-18 18:36 ` Ionut Mihalcea
@ 2026-09-21 4:35 ` Tony Lindgren
0 siblings, 0 replies; 78+ messages in thread
From: Tony Lindgren @ 2026-09-21 4:35 UTC (permalink / raw)
To: Ionut Mihalcea
Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Artem Bityutskiy,
Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky,
Marc Zyngier, Oliver Upton, Steven Price, Anup Patel,
Samuel Ortiz, Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm@vger.kernel.org, Mathias Brossard, nd
Hi,
On Fri, Sep 18, 2026 at 06:36:38PM +0000, Ionut Mihalcea wrote:
> On 31 Aug 2026, at 08:13, Tony Lindgren <tony.lindgren@linux.intel.com> wrote:
> > Tom, since you mentioned that AMD SEV-SNP and Intel TDX live migration
> > sound similar, can you please take a look how the API might work for
> > SEV-SNP?
>
> While the x86 implementation is outside our scope, we're also currently
> investigating how the Arm CCA ABIs for Live Migration ABIs would fit as a
> backend for this API. We will provide more feedback on the API shape and
> semantics as the analysis progresses.
OK great, yes great if we can make this work also for ARM too.
> As an intro to our approach (and a potential touchpoint on the topic of this
> API), Mathias Brossard will be presenting Arm CCA Live Migration at LPC next
> month [0].
OK
> > Dirty page tracking does not add a new uAPI. Userspace keeps using
> > KVM_GET_DIRTY_LOG and KVM_CLEAR_DIRTY_LOG.
>
> One concern we currently have is about an impedance mismatch between the
> non-CoCo dirty log APIs and our approach to dirty reporting. This is just
> to highlight the concern at the moment, without any explicit proposal as
> an alternative.
Care to describe a bit more what issues you are seeing with dirty log?
What we have noticed with TDX is that KVM_GET_DIRTY_LOG works with
additional vendor ops for live migration. But it is not usable outside
migration. We currently get a list of migration candidates from the TDX
module and that status sticks until migration or cancel operation.
So the dirty log stats are invalid outside the migration context. This
means that for example qemu calc_dirty_rate is not available before the
migration for stats.
I guess we could eventually get preprocessed stats from the TDX module to
for reporting the dirty rate. But it seems that initially we may want to
return nothing for dirty log outside migration.
> [0] https://lpc.events/event/20/contributions/2476/
>
>
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-18 4:33 ` Kishen Maloor
@ 2026-09-21 5:58 ` Tony Lindgren
2026-09-21 6:56 ` Tony Lindgren
0 siblings, 1 reply; 78+ messages in thread
From: Tony Lindgren @ 2026-09-21 5:58 UTC (permalink / raw)
To: Kishen Maloor
Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Artem Bityutskiy,
Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky,
Marc Zyngier, Oliver Upton, Steven Price, Anup Patel,
Samuel Ortiz, Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg,
Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm
On Thu, Sep 17, 2026 at 09:33:18PM -0700, Kishen Maloor wrote:
> On 8/31/26 12:13 AM, Tony Lindgren wrote:
>
> I was wondering how a vendor would use kvm_transfer_buffer for a command that
> passes an input and also returns an output. size can describe the input
> length or the space available for output, not both, so the kernel has no way to
> know how much it may write.
>
> Two suggestions below. They're orthogonal.
>
> > ...
> > diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h
> > index 5f6c1ce9673b7..d9291a8a97bb1 100644
> > --- a/arch/x86/include/asm/kvm_host.h
> > +++ b/arch/x86/include/asm/kvm_host.h
> > @@ -2010,6 +2010,8 @@ struct kvm_x86_ops {
> > int (*gmem_prepare)(struct kvm *kvm, kvm_pfn_t pfn, gfn_t gfn, int max_order);
> > void (*gmem_invalidate)(kvm_pfn_t start, kvm_pfn_t end);
> > int (*gmem_max_mapping_level)(struct kvm *kvm, kvm_pfn_t pfn, bool is_private);
> > + bool (*cap_live_migration)(struct kvm *kvm);
>
> Should this return an 'int' specifying the maximum buffer size that the vendor impl
> requires/uses? Userspace can then learn this once.
OK that sounds good to me. I assume you are thinking the maximum hardware
specific buffer size per thread? As in 512 * 4096 bytes for the TDX case?
> > +
> > +struct kvm_transfer_buffer {
> > + __u64 address;
> > + __u32 size;
> > + __u32 reserved;
> > +};
>
> Should this struct include a 'capacity' field (u32) that is set on each command?
> It would be the number of bytes writable at address.
> size would be the input length on entry (0 if the command passes none), and the
> number of bytes produced on return (0 if none).
Hmm so the transfer command return value can return how many bytes were
written of the input. But yeah we don't know how many bytes were written
back to the transfer buffer as result of the transfer command.
How about if we add the bytes returned to the transfer struct? Then the
kvm_transfer_buffer can stay as just a buffer.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-21 0:13 ` Kishen Maloor
@ 2026-09-21 6:52 ` Tony Lindgren
2026-09-21 9:24 ` Tony Lindgren
2026-09-22 3:57 ` Kishen Maloor
0 siblings, 2 replies; 78+ messages in thread
From: Tony Lindgren @ 2026-09-21 6:52 UTC (permalink / raw)
To: Kishen Maloor
Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini,
Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On Sun, Sep 20, 2026 at 05:13:10PM -0700, Kishen Maloor wrote:
> On 9/17/26 10:58 PM, Tony Lindgren wrote:
> > On Thu, Sep 17, 2026 at 09:32:23PM -0700, Kishen Maloor wrote:
> >> On 9/16/26 11:42 PM, Tony Lindgren wrote:
> >>> On Wed, Sep 16, 2026 at 08:31:32PM -0700, Kishen Maloor wrote:
> >>>> On 9/15/26 10:09 PM, Tony Lindgren wrote:
> >>>>> So trying to summarize the common flags for the role and separate vendor
> >>>>> flags:
> >>>>>
> >>>>> struct kvm_migrate_cmd {
> >>>>> __u16 command;
> >>>>> __u16 flags;
> >>>>> __u16 vflags;
> >>>>> __u16 reserved;
> >>>>> __u32 reserved;
> >>>>> struct kvm_transfer_buffer buf;
> >>>>> };
> >>>>>
> >>>>> Is the above along the lines what you were thinking?
> >>>>
> >>>> No, I was suggesting a 'role' field carved out of the 'reserved' space,
> >>>> like this:
> >>>>
> >>>> struct kvm_migrate_cmd {
> >>>> __u16 command;
> >>>> __u16 flags;
> >>>> __u8 role; /* 0 = unset, 1 = source, 2 = destination */
> >>>> __u8 reserved[3];
> >>>> struct kvm_transfer_buffer buf;
> >>>> };
> >>>
> >>> OK yes thanks for clarifying, that works for me.
> >>>
> >>>>> Ah OK, yes that would also tell "the hardware has been initialized to a
> >>>>> certain migration role". That seems like a usable common feature.
> >>>>
> >>>> Not quite. It tells us that userspace asserted a role for this VM's migration
> >>>> session. Whether a TD was created for import is a separate, vendor-level detail.
> >>>> The generic layer only needs the role to reject a session that never stated one,
> >>>> and to pick the export or import callback. That callback then knows which side
> >>>> it's on and can reject an incorrect role (e.g., if SETUP asserted dst for a src TD).
> >>>
> >>> That's a good point, the hardware role may not be set yet.
> >>>
> >>> I'm still wondering if there is a need to stash the userspace set role in
> >>> KVM though. Likely only the hardware specific code can properly track the
> >>> state of the hardware and adjust to the userspace requests. Seems just
> >>> being able to pass the role in struct kvm_migrate_cmd should be enough?
> >>
> >> Passing it in kvm_migrate_cmd is enough for SETUP itself, but the commands
> >> after SETUP like memory/vcpu transfers still have to reach the right
> >> callback. So KVM would need to remember what was asserted so that the
> >> generic layer can dispatch to the export or import facing callbacks.
> >> We've been sketching (on this thread) an alternative UAPI set
> >> (3 vs 5 ioctls) for consideration which this stored role enables:
> >>
> >> Proposed in the RFC Alternative
> >> KVM_MIGRATE_CMD KVM_MIGRATE_CMD
> >> KVM_EXPORT_MEMORY
> >> KVM_IMPORT_MEMORY KVM_TRANSFER_MEMORY
> >> KVM_EXPORT_VCPU
> >> KVM_IMPORT_VCPU KVM_TRANSFER_VCPU
> >>
> >> It's just a record (1 byte) of what userspace asserted for the current session
> >> at SETUP. Vendor code still owns the hardware state and remains free to reject a
> >> role that doesn't match it. It is also what lets the generic layer reject a
> >> command to a VM that never set up a session.
> >
> > For the TRANSFER style operations, I would assume the direction is passed
> > for each transfer, just like the Linux does for the dmaengine. It's
> > possible that there may be transfers going both directions without the
> > role changing.
> >
> > So looks like we have tree things to consider: userspace set migration
> > role, the hardware state, and transfer direction.
>
> Two of those three I agree with: userspace passes the role at SETUP, and
> the vendor implementation tracks the hardware state. It's the per-transfer
> direction I don't think we need.
Ack on the userspace passing the role at SETUP and vendor implementation
tracking the hardware state.
Then for KVM tracking the role, I don't think we need it with the two
above. The role tracking can always be added if really needed. Any other
opinions on this one?
The transfer direction is there with the EXPORT/IMPORT naming. Maybe
just let's keep that naming for easier readability rather than try to
switch to TRANSFER style naming. No transfer direction flag needed.
> > What if userspace always passes the role and transfer direction where it
> > makes sense? And then the hardware specific implementation tracks the
> > hardware state?
>
> Unless there's another meaning, direction would indicate that an operation
> must produce or consume a blob into/from a buffer. Such an indication would be
> necessary if the layer below the API cannot know what to do with a buffer,
> which may well be the case in your dmaengine analogy. However, in this case a
> vendor implementation sits below the generic layer and could derive what it
> needs to do from the role, command, and any session state it maintains. The
> TDH.EXPORT.ABORT example I cited upthread expects a token only once the
> session has left its pre-copy phase, so what to do with the buffer follows
> from state the vendor layer already holds.
Yeah I don't think the dmaengine API ever expects to get back a blob as a
result of an outgoing transfer.. That would be a separate DMA transfer.
> The gap I do see is in how the buffer itself is described: we have no way to
> express a command that takes an input and produces an output. A direction flag
> doesn't help there either, since it can only say one thing. My comments on patch 2
> suggest a new 'capacity' field alongside a reframing of 'size' in kvm_transfer_buffer,
> where capacity bounds what the kernel may write and size reports what is actually
> there. That covers the both-ways case, and it also leaves nothing for a direction
> field to convey. Maybe that addresses your concern?
Yes that's a good point, a transfer command may also return data in the
transfer buffer and the size of returned data needs to be known. Replied
to your patch #2 comments with some ideas on it.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-21 5:58 ` Tony Lindgren
@ 2026-09-21 6:56 ` Tony Lindgren
2026-09-22 3:56 ` Kishen Maloor
0 siblings, 1 reply; 78+ messages in thread
From: Tony Lindgren @ 2026-09-21 6:56 UTC (permalink / raw)
To: Kishen Maloor
Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Artem Bityutskiy,
Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky,
Marc Zyngier, Oliver Upton, Steven Price, Anup Patel,
Samuel Ortiz, Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg,
Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm
On Mon, Sep 21, 2026 at 08:58:17AM +0300, Tony Lindgren wrote:
> On Thu, Sep 17, 2026 at 09:33:18PM -0700, Kishen Maloor wrote:
> > On 8/31/26 12:13 AM, Tony Lindgren wrote:
> >
> > I was wondering how a vendor would use kvm_transfer_buffer for a command that
> > passes an input and also returns an output. size can describe the input
> > length or the space available for output, not both, so the kernel has no way to
> > know how much it may write.
> >
> > Two suggestions below. They're orthogonal.
> >
> > > ...
> > > diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h
> > > index 5f6c1ce9673b7..d9291a8a97bb1 100644
> > > --- a/arch/x86/include/asm/kvm_host.h
> > > +++ b/arch/x86/include/asm/kvm_host.h
> > > @@ -2010,6 +2010,8 @@ struct kvm_x86_ops {
> > > int (*gmem_prepare)(struct kvm *kvm, kvm_pfn_t pfn, gfn_t gfn, int max_order);
> > > void (*gmem_invalidate)(kvm_pfn_t start, kvm_pfn_t end);
> > > int (*gmem_max_mapping_level)(struct kvm *kvm, kvm_pfn_t pfn, bool is_private);
> > > + bool (*cap_live_migration)(struct kvm *kvm);
> >
> > Should this return an 'int' specifying the maximum buffer size that the vendor impl
> > requires/uses? Userspace can then learn this once.
>
> OK that sounds good to me. I assume you are thinking the maximum hardware
> specific buffer size per thread? As in 512 * 4096 bytes for the TDX case?
>
> > > +
> > > +struct kvm_transfer_buffer {
> > > + __u64 address;
> > > + __u32 size;
> > > + __u32 reserved;
> > > +};
> >
> > Should this struct include a 'capacity' field (u32) that is set on each command?
> > It would be the number of bytes writable at address.
> > size would be the input length on entry (0 if the command passes none), and the
> > number of bytes produced on return (0 if none).
>
> Hmm so the transfer command return value can return how many bytes were
> written of the input. But yeah we don't know how many bytes were written
> back to the transfer buffer as result of the transfer command.
>
> How about if we add the bytes returned to the transfer struct? Then the
> kvm_transfer_buffer can stay as just a buffer.
Actually, for the possible cases with input+output, we could reserve space
in the transfer struct for another struct kvm_transfer_buffer for the
results?
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-21 6:52 ` Tony Lindgren
@ 2026-09-21 9:24 ` Tony Lindgren
2026-09-21 10:58 ` Tony Lindgren
2026-09-22 3:57 ` Kishen Maloor
1 sibling, 1 reply; 78+ messages in thread
From: Tony Lindgren @ 2026-09-21 9:24 UTC (permalink / raw)
To: Kishen Maloor
Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini,
Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On Mon, Sep 21, 2026 at 09:52:42AM +0300, Tony Lindgren wrote:
> On Sun, Sep 20, 2026 at 05:13:10PM -0700, Kishen Maloor wrote:
> > On 9/17/26 10:58 PM, Tony Lindgren wrote:
> > > So looks like we have tree things to consider: userspace set migration
> > > role, the hardware state, and transfer direction.
> >
> > Two of those three I agree with: userspace passes the role at SETUP, and
> > the vendor implementation tracks the hardware state. It's the per-transfer
> > direction I don't think we need.
>
> Ack on the userspace passing the role at SETUP and vendor implementation
> tracking the hardware state.
>
> Then for KVM tracking the role, I don't think we need it with the two
> above. The role tracking can always be added if really needed. Any other
> opinions on this one?
>
> The transfer direction is there with the EXPORT/IMPORT naming. Maybe
> just let's keep that naming for easier readability rather than try to
> switch to TRANSFER style naming. No transfer direction flag needed.
Looking at include/linux/dmaengine.h again, the transfer specific
direction is depreated.. So to follow the the dmaengine analogy, I now
agree that migration SETUP is the place to set the transfer direction
like you're suggesting. And so it seems the role at SETUP and the vendor
implementation tracking the hardware state is all we need.
The EXPORT/IMPORT vs TRANFER naming both work for me. I find the
EXPORT/IMPORT easier to read, and I think Jörg also seemed to like
that naming better.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-21 9:24 ` Tony Lindgren
@ 2026-09-21 10:58 ` Tony Lindgren
0 siblings, 0 replies; 78+ messages in thread
From: Tony Lindgren @ 2026-09-21 10:58 UTC (permalink / raw)
To: Kishen Maloor
Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini,
Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On Mon, Sep 21, 2026 at 12:24:26PM +0300, Tony Lindgren wrote:
> On Mon, Sep 21, 2026 at 09:52:42AM +0300, Tony Lindgren wrote:
> > On Sun, Sep 20, 2026 at 05:13:10PM -0700, Kishen Maloor wrote:
> > > On 9/17/26 10:58 PM, Tony Lindgren wrote:
> > > > So looks like we have tree things to consider: userspace set migration
> > > > role, the hardware state, and transfer direction.
> > >
> > > Two of those three I agree with: userspace passes the role at SETUP, and
> > > the vendor implementation tracks the hardware state. It's the per-transfer
> > > direction I don't think we need.
> >
> > Ack on the userspace passing the role at SETUP and vendor implementation
> > tracking the hardware state.
> >
> > Then for KVM tracking the role, I don't think we need it with the two
> > above. The role tracking can always be added if really needed. Any other
> > opinions on this one?
> >
> > The transfer direction is there with the EXPORT/IMPORT naming. Maybe
> > just let's keep that naming for easier readability rather than try to
> > switch to TRANSFER style naming. No transfer direction flag needed.
>
> Looking at include/linux/dmaengine.h again, the transfer specific
> direction is depreated.. So to follow the the dmaengine analogy, I now
> agree that migration SETUP is the place to set the transfer direction
> like you're suggesting. And so it seems the role at SETUP and the vendor
> implementation tracking the hardware state is all we need.
Sorry above I meant "SETUP is the place to set the role", not the
direction.
> The EXPORT/IMPORT vs TRANFER naming both work for me. I find the
> EXPORT/IMPORT easier to read, and I think Jörg also seemed to like
> that naming better.
And KVM is using KVM_GET/SET_REGS type ioctls, not TRANSFER. Also, I
think EXPORT/IMPORT is better naming than GET/SET here as we export
into an encrypted blob.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-21 6:56 ` Tony Lindgren
@ 2026-09-22 3:56 ` Kishen Maloor
2026-09-22 6:27 ` Tony Lindgren
0 siblings, 1 reply; 78+ messages in thread
From: Kishen Maloor @ 2026-09-22 3:56 UTC (permalink / raw)
To: Tony Lindgren
Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Artem Bityutskiy,
Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky,
Marc Zyngier, Oliver Upton, Steven Price, Anup Patel,
Samuel Ortiz, Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg,
Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm
On 9/20/26 11:56 PM, Tony Lindgren wrote:
> On Mon, Sep 21, 2026 at 08:58:17AM +0300, Tony Lindgren wrote:
>> On Thu, Sep 17, 2026 at 09:33:18PM -0700, Kishen Maloor wrote:
>>> On 8/31/26 12:13 AM, Tony Lindgren wrote:
>>>
>>> I was wondering how a vendor would use kvm_transfer_buffer for a command that
>>> passes an input and also returns an output. size can describe the input
>>> length or the space available for output, not both, so the kernel has no way to
>>> know how much it may write.
>>>
>>> Two suggestions below. They're orthogonal.
>>>
>>>> ...
>>>> diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h
>>>> index 5f6c1ce9673b7..d9291a8a97bb1 100644
>>>> --- a/arch/x86/include/asm/kvm_host.h
>>>> +++ b/arch/x86/include/asm/kvm_host.h
>>>> @@ -2010,6 +2010,8 @@ struct kvm_x86_ops {
>>>> int (*gmem_prepare)(struct kvm *kvm, kvm_pfn_t pfn, gfn_t gfn, int max_order);
>>>> void (*gmem_invalidate)(kvm_pfn_t start, kvm_pfn_t end);
>>>> int (*gmem_max_mapping_level)(struct kvm *kvm, kvm_pfn_t pfn, bool is_private);
>>>> + bool (*cap_live_migration)(struct kvm *kvm);
>>>
>>> Should this return an 'int' specifying the maximum buffer size that the vendor impl
>>> requires/uses? Userspace can then learn this once.
>>
>> OK that sounds good to me. I assume you are thinking the maximum hardware
>> specific buffer size per thread? As in 512 * 4096 bytes for the TDX case?
Yeah, an upper bound on the buffer size, so 516 * 4096 (4 for the GPA+MAC lists
and MBMD) in TDX. It could be a hint to userspace to size its buffers at the
start. Though it would be up to userspace to decide how to use that information.
>>
>>>> +
>>>> +struct kvm_transfer_buffer {
>>>> + __u64 address;
>>>> + __u32 size;
>>>> + __u32 reserved;
>>>> +};
>>>
>>> Should this struct include a 'capacity' field (u32) that is set on each command?
>>> It would be the number of bytes writable at address.
>>> size would be the input length on entry (0 if the command passes none), and the
>>> number of bytes produced on return (0 if none).
>>
>> Hmm so the transfer command return value can return how many bytes were
>> written of the input. But yeah we don't know how many bytes were written
>> back to the transfer buffer as result of the transfer command.
The transfer command return value could return how many bytes were written into the buffer.
But in an input-output call, the kernel handler wouldn't know how many bytes it could write,
or for that matter even how many pages to pin up front in case it needs to return an output
because 'size' couldn't simultaneously convey the input length and buffer capacity. That was
the gap that I thought a read-only 'capacity' field could bridge. Of course, this
assumes that the output is written in-place.
>>
>> How about if we add the bytes returned to the transfer struct? Then the
>> kvm_transfer_buffer can stay as just a buffer.
>
> Actually, for the possible cases with input+output, we could reserve space
> in the transfer struct for another struct kvm_transfer_buffer for the
> results?
Yes, say an 'in' and 'out' kvm_transfer_buffer inside struct kvm_migrate_cmd should
close this out and shouldn't require a 'capacity' field. Maybe we then establish this
convention:
- A non-zero 'size' on 'in' at call entry would signal that there is input.
- A non-zero 'size' on 'out' at call entry would convey the buffer capacity.
- A non-zero 'size' on 'out' at call exit would convey that there is output.
- A zeroed 'size' on 'out' at call exit would convey that there is no output.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-21 6:52 ` Tony Lindgren
2026-09-21 9:24 ` Tony Lindgren
@ 2026-09-22 3:57 ` Kishen Maloor
2026-09-22 5:25 ` Tony Lindgren
1 sibling, 1 reply; 78+ messages in thread
From: Kishen Maloor @ 2026-09-22 3:57 UTC (permalink / raw)
To: Tony Lindgren
Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini,
Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On 9/20/26 11:52 PM, Tony Lindgren wrote:
> On Sun, Sep 20, 2026 at 05:13:10PM -0700, Kishen Maloor wrote:
>> On 9/17/26 10:58 PM, Tony Lindgren wrote:
>>> On Thu, Sep 17, 2026 at 09:32:23PM -0700, Kishen Maloor wrote:
>>>> On 9/16/26 11:42 PM, Tony Lindgren wrote:
>>>>> On Wed, Sep 16, 2026 at 08:31:32PM -0700, Kishen Maloor wrote:
>>>>>> On 9/15/26 10:09 PM, Tony Lindgren wrote:
>>>>>>> So trying to summarize the common flags for the role and separate vendor
>>>>>>> flags:
>>>>>>>
>>>>>>> struct kvm_migrate_cmd {
>>>>>>> __u16 command;
>>>>>>> __u16 flags;
>>>>>>> __u16 vflags;
>>>>>>> __u16 reserved;
>>>>>>> __u32 reserved;
>>>>>>> struct kvm_transfer_buffer buf;
>>>>>>> };
>>>>>>>
>>>>>>> Is the above along the lines what you were thinking?
>>>>>>
>>>>>> No, I was suggesting a 'role' field carved out of the 'reserved' space,
>>>>>> like this:
>>>>>>
>>>>>> struct kvm_migrate_cmd {
>>>>>> __u16 command;
>>>>>> __u16 flags;
>>>>>> __u8 role; /* 0 = unset, 1 = source, 2 = destination */
>>>>>> __u8 reserved[3];
>>>>>> struct kvm_transfer_buffer buf;
>>>>>> };
>>>>>
>>>>> OK yes thanks for clarifying, that works for me.
>>>>>
>>>>>>> Ah OK, yes that would also tell "the hardware has been initialized to a
>>>>>>> certain migration role". That seems like a usable common feature.
>>>>>>
>>>>>> Not quite. It tells us that userspace asserted a role for this VM's migration
>>>>>> session. Whether a TD was created for import is a separate, vendor-level detail.
>>>>>> The generic layer only needs the role to reject a session that never stated one,
>>>>>> and to pick the export or import callback. That callback then knows which side
>>>>>> it's on and can reject an incorrect role (e.g., if SETUP asserted dst for a src TD).
>>>>>
>>>>> That's a good point, the hardware role may not be set yet.
>>>>>
>>>>> I'm still wondering if there is a need to stash the userspace set role in
>>>>> KVM though. Likely only the hardware specific code can properly track the
>>>>> state of the hardware and adjust to the userspace requests. Seems just
>>>>> being able to pass the role in struct kvm_migrate_cmd should be enough?
>>>>
>>>> Passing it in kvm_migrate_cmd is enough for SETUP itself, but the commands
>>>> after SETUP like memory/vcpu transfers still have to reach the right
>>>> callback. So KVM would need to remember what was asserted so that the
>>>> generic layer can dispatch to the export or import facing callbacks.
>>>> We've been sketching (on this thread) an alternative UAPI set
>>>> (3 vs 5 ioctls) for consideration which this stored role enables:
>>>>
>>>> Proposed in the RFC Alternative
>>>> KVM_MIGRATE_CMD KVM_MIGRATE_CMD
>>>> KVM_EXPORT_MEMORY
>>>> KVM_IMPORT_MEMORY KVM_TRANSFER_MEMORY
>>>> KVM_EXPORT_VCPU
>>>> KVM_IMPORT_VCPU KVM_TRANSFER_VCPU
>>>>
>>>> It's just a record (1 byte) of what userspace asserted for the current session
>>>> at SETUP. Vendor code still owns the hardware state and remains free to reject a
>>>> role that doesn't match it. It is also what lets the generic layer reject a
>>>> command to a VM that never set up a session.
>>>
>>> For the TRANSFER style operations, I would assume the direction is passed
>>> for each transfer, just like the Linux does for the dmaengine. It's
>>> possible that there may be transfers going both directions without the
>>> role changing.
>>>
>>> So looks like we have tree things to consider: userspace set migration
>>> role, the hardware state, and transfer direction.
>>
>> Two of those three I agree with: userspace passes the role at SETUP, and
>> the vendor implementation tracks the hardware state. It's the per-transfer
>> direction I don't think we need.
>
> Ack on the userspace passing the role at SETUP and vendor implementation
> tracking the hardware state.
>
> Then for KVM tracking the role, I don't think we need it with the two
> above. The role tracking can always be added if really needed. Any other
> opinions on this one?
I do think it's useful for generic KVM to track this role (1 byte).
- It lets _TRANSFER_ style calls dispatch directly to import or export
callbacks based on the role.
- Even if we don't adopt the _TRANSFER_ style, it enables generic KVM to
reject mismatched calls, e.g., KVM_EXPORT_MEMORY on a destination.
>
> The transfer direction is there with the EXPORT/IMPORT naming. Maybe
> just let's keep that naming for easier readability rather than try to
> switch to TRANSFER style naming. No transfer direction flag needed.
I can't say I have a clear preference between EXPORT/IMPORT vs _TRANSFER_.
If we keep the EXPORT/IMPORT naming, then consistency would arguably call for
splitting MIGRATE_CMD too, which makes it 6 vs 3 (or 5 vs 3 against the RFC
as posted):
EXPORT/IMPORT style _TRANSFER_ style
KVM_EXPORT_CMD KVM_MIGRATE_CMD
KVM_IMPORT_CMD
KVM_EXPORT_MEMORY KVM_TRANSFER_MEMORY
KVM_IMPORT_MEMORY
KVM_EXPORT_VCPU KVM_TRANSFER_VCPU
KVM_IMPORT_VCPU
I guess the main benefit of the _TRANSFER_ style is reduced duplication.
Each EXPORT/IMPORT pair takes the same struct and differs only by the role
that the session already knows. Merging them gives one entry point per call
type and lets userspace drive both ends from the same call site which could
be considered a win. It doesn't reduce kernel code though as the top-level
handler still branches internally on the role.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-22 3:57 ` Kishen Maloor
@ 2026-09-22 5:25 ` Tony Lindgren
2026-09-23 0:38 ` Kishen Maloor
0 siblings, 1 reply; 78+ messages in thread
From: Tony Lindgren @ 2026-09-22 5:25 UTC (permalink / raw)
To: Kishen Maloor
Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini,
Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On Mon, Sep 21, 2026 at 08:57:45PM -0700, Kishen Maloor wrote:
> On 9/20/26 11:52 PM, Tony Lindgren wrote:
> > Then for KVM tracking the role, I don't think we need it with the two
> > above. The role tracking can always be added if really needed. Any other
> > opinions on this one?
>
> I do think it's useful for generic KVM to track this role (1 byte).
> - It lets _TRANSFER_ style calls dispatch directly to import or export
> callbacks based on the role.
> - Even if we don't adopt the _TRANSFER_ style, it enables generic KVM to
> reject mismatched calls, e.g., KVM_EXPORT_MEMORY on a destination.
Having KVM do generic checks on the calls is a good idea. There might be
a simpler way of handling it though. Rather than having KVM track the
migration state, how about we add a function to check for the migration
session state from the vendor code?
So something like this for the states you suggested earlier:
enum kvm_lmstate {
KVM_LM_NONE,
KVM_LM_SOURCE,
KVM_LM_DESTINATION,
};
With something like this to get the state from the vendor code:
enum kmv_lmstate kvm_arch_get_lmstate(struct kvm *);
For x86 it would end up calling kvm_x86_call(get_lmstate)(kvm) and for
the TDX specific case tdx_get_lmstate().
It would allow KVM to do the generic checks for the migration related
calls you're describing. And having KVM start tracking the state can be
still added later on too if it is needed.
> > The transfer direction is there with the EXPORT/IMPORT naming. Maybe
> > just let's keep that naming for easier readability rather than try to
> > switch to TRANSFER style naming. No transfer direction flag needed.
>
> I can't say I have a clear preference between EXPORT/IMPORT vs _TRANSFER_.
> If we keep the EXPORT/IMPORT naming, then consistency would arguably call for
> splitting MIGRATE_CMD too, which makes it 6 vs 3 (or 5 vs 3 against the RFC
> as posted):
>
> EXPORT/IMPORT style _TRANSFER_ style
> KVM_EXPORT_CMD KVM_MIGRATE_CMD
> KVM_IMPORT_CMD
> KVM_EXPORT_MEMORY KVM_TRANSFER_MEMORY
> KVM_IMPORT_MEMORY
> KVM_EXPORT_VCPU KVM_TRANSFER_VCPU
> KVM_IMPORT_VCPU
Agreed we should split the MIGRATE_CMD too. Probably the number of ioctls
is not and issue compared to following the KVM style and better
readability. So my vote is now on EXPORT/IMPORT style naming.
> I guess the main benefit of the _TRANSFER_ style is reduced duplication.
> Each EXPORT/IMPORT pair takes the same struct and differs only by the role
> that the session already knows. Merging them gives one entry point per call
> type and lets userspace drive both ends from the same call site which could
> be considered a win. It doesn't reduce kernel code though as the top-level
> handler still branches internally on the role.
Yup not much of a win for the TRANSFER style naming.
Trying to summarize again after we sorted out the direction flag issue in
the transfer:
role per migration session (cannot change during the migration)
direction per command, EXPORT/IMPORT
hardware state set and tracked by vendor specific code
And checking again against the dmaengine analogy:
Compared to dmaengine, the migration role is modeled similar to the dma
channel configuration.
The migration direction with EXPORT/IMPORT is modeled similar to
dmaengine_prep_slave_sg().
The migration hardware state is modeled similar to the dmaengine driver
managing the hardware state.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-22 3:56 ` Kishen Maloor
@ 2026-09-22 6:27 ` Tony Lindgren
2026-09-23 0:37 ` Kishen Maloor
0 siblings, 1 reply; 78+ messages in thread
From: Tony Lindgren @ 2026-09-22 6:27 UTC (permalink / raw)
To: Kishen Maloor
Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Artem Bityutskiy,
Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky,
Marc Zyngier, Oliver Upton, Steven Price, Anup Patel,
Samuel Ortiz, Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg,
Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm
On Mon, Sep 21, 2026 at 08:56:48PM -0700, Kishen Maloor wrote:
> On 9/20/26 11:56 PM, Tony Lindgren wrote:
> > On Mon, Sep 21, 2026 at 08:58:17AM +0300, Tony Lindgren wrote:
> >> On Thu, Sep 17, 2026 at 09:33:18PM -0700, Kishen Maloor wrote:
> >>> On 8/31/26 12:13 AM, Tony Lindgren wrote:
> >>>
> >>> I was wondering how a vendor would use kvm_transfer_buffer for a command that
> >>> passes an input and also returns an output. size can describe the input
> >>> length or the space available for output, not both, so the kernel has no way to
> >>> know how much it may write.
> >>>
> >>> Two suggestions below. They're orthogonal.
> >>>
> >>>> ...
> >>>> diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h
> >>>> index 5f6c1ce9673b7..d9291a8a97bb1 100644
> >>>> --- a/arch/x86/include/asm/kvm_host.h
> >>>> +++ b/arch/x86/include/asm/kvm_host.h
> >>>> @@ -2010,6 +2010,8 @@ struct kvm_x86_ops {
> >>>> int (*gmem_prepare)(struct kvm *kvm, kvm_pfn_t pfn, gfn_t gfn, int max_order);
> >>>> void (*gmem_invalidate)(kvm_pfn_t start, kvm_pfn_t end);
> >>>> int (*gmem_max_mapping_level)(struct kvm *kvm, kvm_pfn_t pfn, bool is_private);
> >>>> + bool (*cap_live_migration)(struct kvm *kvm);
> >>>
> >>> Should this return an 'int' specifying the maximum buffer size that the vendor impl
> >>> requires/uses? Userspace can then learn this once.
> >>
> >> OK that sounds good to me. I assume you are thinking the maximum hardware
> >> specific buffer size per thread? As in 512 * 4096 bytes for the TDX case?
>
> Yeah, an upper bound on the buffer size, so 516 * 4096 (4 for the GPA+MAC lists
> and MBMD) in TDX. It could be a hint to userspace to size its buffers at the
> start. Though it would be up to userspace to decide how to use that information.
Oh right thanks. I think this is really the maximum transfer buffer size
Peter asked, not just a hint to userspace :)
> >>>> +struct kvm_transfer_buffer {
> >>>> + __u64 address;
> >>>> + __u32 size;
> >>>> + __u32 reserved;
> >>>> +};
> >>>
> >>> Should this struct include a 'capacity' field (u32) that is set on each command?
> >>> It would be the number of bytes writable at address.
> >>> size would be the input length on entry (0 if the command passes none), and the
> >>> number of bytes produced on return (0 if none).
> >>
> >> Hmm so the transfer command return value can return how many bytes were
> >> written of the input. But yeah we don't know how many bytes were written
> >> back to the transfer buffer as result of the transfer command.
>
> The transfer command return value could return how many bytes were written into the buffer.
> But in an input-output call, the kernel handler wouldn't know how many bytes it could write,
> or for that matter even how many pages to pin up front in case it needs to return an output
> because 'size' couldn't simultaneously convey the input length and buffer capacity. That was
> the gap that I thought a read-only 'capacity' field could bridge. Of course, this
> assumes that the output is written in-place.
Hmm yeah this inplace capacity vs transferred issue remains still. So I
agree we need to specify the capacity in struct kvm_transfer_buffer like
you suggested.
To me size is already the size of the buffer though. So instead of changing
size to capacity, how about something like datasize or len for the input
and output transfer length?
> >> How about if we add the bytes returned to the transfer struct? Then the
> >> kvm_transfer_buffer can stay as just a buffer.
> >
> > Actually, for the possible cases with input+output, we could reserve space
> > in the transfer struct for another struct kvm_transfer_buffer for the
> > results?
>
> Yes, say an 'in' and 'out' kvm_transfer_buffer inside struct kvm_migrate_cmd should
> close this out and shouldn't require a 'capacity' field. Maybe we then establish this
> convention:
> - A non-zero 'size' on 'in' at call entry would signal that there is input.
> - A non-zero 'size' on 'out' at call entry would convey the buffer capacity.
> - A non-zero 'size' on 'out' at call exit would convey that there is output.
> - A zeroed 'size' on 'out' at call exit would convey that there is no output.
Looks doable to me but with the inplace issue as above.. Sounds like we just
need to reserve space for a case with a separate output buffer though.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
2026-09-18 15:53 ` Peter Xu
@ 2026-09-22 8:09 ` Artem Bityutskiy
2026-09-22 9:42 ` Tony Lindgren
` (2 more replies)
0 siblings, 3 replies; 78+ messages in thread
From: Artem Bityutskiy @ 2026-09-22 8:09 UTC (permalink / raw)
To: Peter Xu
Cc: Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas,
Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
Hi Peter,
thanks again for good comments and questions.
On Fri, 2026-09-18 at 11:53 -0400, Peter Xu wrote:
> >
> > FYI, I am working on a TDX migration model document. It describes what the
> > TDX module offers, but unlike the specs it is oriented towards software
> > engineers: much easier to read and it does not require deep TDX knowledge.
> > It is all based on public specs, just distilled into readable mental models.
> > I plan to publish it publicly. I am about 80% done.
>
> That will be very useful, thanks for doing this. I'll be more than happy to
> read it when it's done. I wonder if we can, after TDX bits done, use this
> doc as a base for others to add in theirs in separate tabs, keeping
> everything together for the unified migration API work.
>
> Out of pure curiosity, I also don't know where s390 stands; it's almost not
> mentioned in the current API plan.
In general, sounds like a good idea. Let me finish TDX and share it first,
then we can think about the next step. But in general a master doc with high
level overview and comparison of different CoCo migration models would be
awesome.
s390 - yes, would be curious to know.
> >
> > We do not have code. But Kishen spent time playing with it, and I think he
> > concluded not to proceed with this. But he might have evaluated it from
> > the "unify all migration into a single generic API" perspective. But may
> > be as a "this is a test framework" perspective is different, at least I feel
> > it may be the case. I think Kishen can provide more insight if needed, he
> > is in CC.
>
> Yes, thanks. I'm willing to hear more, and I'm a bit surprised that
> non-coco migration didn't fit already well into it, because IIUC non-coco
> needs less in this case, not more.
Yeah, just thinking aloud now.
- Suppose QEMU and KVM support CoCo migration.
- Does it make sense to treat traditional VM as a CoCo VM for migration
purposes?
- Not for performance reasons - mmapped memory I/O is going to be more
efficient than any sorts of export/import uAPIs.
- Yes for testing/experimental purposes.
- Would it add a lot of code and complexity: would the benefits outweigh the
costs?
- QEMU: Should require minimum amount of code. Would need a flag like
"treat me as CoCo VM", may be some QEMU operator commands or cmdline
options.
- KVM: Not sure, but intuitively it may add noticeable amount of churn.
> > The final round is special. It uses the "done" flag, which is passed to the
> > `TDH.EXPORT.TRACK` seamcall, and makes it export a special variant of the
> > epoch token that is called the start token. On top of the integrity and
> > ordering guarantees, the start token is what allows the destination to
> > start: until the destination imports it, the TDX module will not let the
> > destination TD run with partial state.
>
> The methodology on unique VM attestation sounds comlex, but I think I get
> it now, thank you. It'll be nice if some of these reasonings will also be
> there in the doc you're drafting.
Yes, will add.
> > > >
> I am just thinking out loud here: if the hardware is good enough to do
> encryption plus (some?) compression, that would be very nice. For "some",
> I meant minimum over zero pages. Because in this case even if host wants
> to play tricks with zero pages, it can't anymore when un-readable. Maybe
> the guest driver can play some trick, but I'm also not sure if in CoCo.
>
> M > N may imply it's not the case for now, but it's still sane as a start
> even if so.
Yeah. In case of TDX, the TDX module does not do compression today,
and there is no zero-page optimization either.
But keep in mind that what pages in TD are zero pages is by itself a secret.
Any sort of theoretical TDX module zero page compression optimization would
need to be done in a way that VMM cannot figure out what pages were zero
pages. So with a disclaimer that I am not really a security expert, I'd say
it is far-fetched.
> > I believe in our current PoC, the ioctl requires the buffer size to be large
> > enough to hold all the requested pages. But this is specific to our current
> > PoC implementation.
> >
> > In general, I feel like if buffer size is not enough, the uAPI could fill it
> > with as much data as fits, and communicate back about what GPAs were
> > exported. The caller could export the rest separately.
>
> Yes, this will work.
>
> Or maybe it's simpler to be able to export an upper bound in another API
> that probes it (some KVM cap)?
>
> Any retry is a wasted round trip from perf perspective. The upper bound
> can be relatively large, IMHO, which should be non-issue. It should be
> simpler for both userapp and kernel if feasible.
OK, so by upper bound here you mean the maximum amount of data that can be
exported in a single memory export ioctl operation, right? And you mean that
there should be some sort of command to query it, similar to KVM_CAP_XSAVE2
capability returning the XSAVE buffer size? If so - yes, I think exposing
such an upper bound would be useful.
Let's see. In case of TDX, the `TDH.EXPORT.MEM` seamcall today can export
max. 512 pages at a time. Today only 4KiB pages are supported, so it is just
a 2MiB buffer. So this is the upper bound for the seamcall.
So in TDX case, 512 pages would be the absolute upper bound for the uAPI
input buffer size. The alternative is to accept larger input buffers and
internally do multiple seamcalls. For TDX case, I do not think the latter
makes much sense though. But, what I do think is that the uAPI itself
shouldn't mandate one approach over the other - a different CoCo
implementation may prefer to aggregate multiple hardware calls internally.
Also, today TDX can only migrate 4KiB pages - VMM must split larger pages
into 4KiB pages before migration. But I expect that in the future TDX may
implement larger pages migration too. So I think the cap should be expressed
in page count rather than bytes.
That said, I don't think exposing this upper bound cap should replace the
flexibility we discussed earlier: the API itself should still allow
exporting as much as fits the output buffer, regardless of what the cap
reports. A specific CoCo implementation could still choose to just fail if
the buffer isn't large enough to hold all requested pages - but that would
be an implementation limitation, not an API limitation.
And the other question is whether to allow exporting a mix of large and
small pages, like N x 4KiB along with M x 2MiB and K x 1GiB pages. Or the
uAPI would allow to only export one page size at a time?
If large pages are added to TDX, I'd speculate that TDX seamcall would allow
the mixing and matching - I can see that already today the seamcall API
is provisioned for this. uAPI could require that the input buffer (set of
pages) can contain a mix which should match the GPAs of the corresponding
memory regions. But then race conditions - the GPA layout of the VM may
change by the time the request reaches the hardware: large pages can be
split or the other way round, I guess? Then uAPI could return a "retry" sort
of exit code, may be?
What do you think? I am just thinking aloud: looks like pages of different
size add a degree of complexity - I did not take that into account until
this conversation, thanks.
>
> Said that, we'll need to be careful then in case of migration fallbacks at
> the final stage. Nowadays, I believe QEMU can still fallback to source side
> at a very, very late stage after all things applied. If I'm not mistaken,
> the final handshake is done at migration_incoming_state_destroy() ->
> migrate_send_rp_shut() telling source to be gone.
Right. In traditional pre-copy VM migration model, the fallback is possible
at any point before the destination VM starts running and modifying its
state.
The same is true for TDX, but with more complexity and limitations.
We have 2 points:
1. Before the source has exported the start token - fallback is similar to
traditional VM migration - just abort the migration on source and
continue running the source TD, and just destroy the destination TD.
2. After the source has exported the start token - fallback is still
possible, but more complex - it requires the destination to first
generate the abort token, which should be delivered to the source and
consumed there. There are seamcalls for both generating and consuming the
abort token. The token basically makes sure the destination cannot run,
and the source can run again - same "only one of the TDs can run at a
time" security rule that we discussed earlier.
Now, the limitation here is that if the abort is because the network
connectivity is lost, the abort token cannot be delivered. Then I'd guess
migration should stop, without destroying the source and the destination,
and the token should be delivered later manually, QEMU could even provide an
infrastructure / commands for that.
Would be very helpful to learn the abort protocol the other CoCo vendors
offer - is it similar to TDX or not?
In our PoC we did not implement the TDX abort token procedure. But the
proposed uAPI is provisioned for the abort token export/import.
> After reading above, one thing we may want to make sure is TDX ENTER on
> dest be exactly the last thing to do on destination, rather than dest QEMU
> ENTER done then something else seems wrong, then dest can't fallback
> anymore. I didn't check into details, though, more of a heads-up to
> whoever is working on QEMU for this in case useful.
Well, first of all, as the above describes, in case of TDX late fallback is
still possible, just requires a more complex protocol. But from
recoverability point of view, I agree with your suggestion to make
TDH.VP.ENTER the last thing done on the destination, rather than starting
the destination early and doing more afterwards. However, the paramount
performance goal is to minimize the downtime, and it conflicts with the
suggested idea. In my mind, minimizing downtime should take precedence.
> > > I saw there's mention of PRE_COPY_STOP state. One example question is,
> > > when reaching this state, can the guest memory still change? What happens
> > > if some emulated device are still DMAing to the guest memory (assuming
> > > flipped from private to shared)? In case of future IO zone support, what
> > > happens if in case of VFIO-PCI assigned doing encrypted DMA?
> > >
> > > From that regard, VFIO has the P2P state where it quiesce initiation of any
> > > DMA from this specific device, then another round to fully stop all devices
> > > into STOP_COPY phase. I wonder if CoCo VMs need similar treatment.
> >
> > Let me split this by device type, because TDX treats them very differently.
> >
> > Emulated (virtio-net, virtio-blk, etc.) only use shared memory - they cannot
> > read or DMA into TD private memory. So full device state lives in shared
> > memory, and QEMU migrates them exactly the same way as for a traditional VM.
> > This is entirely outside the TDX module migration model and outside the
> > proposed uAPI - the uAPI is only for TD private memory.
>
> I hope I understand it right, that all shared pages are out of the secure
> zone TD manages, during migration or not (hence the same as some random
> page the VMM has allocated)? If so, anything about post_load operations
> shouldn't be any concern, and should work as usual.
Yes, that is correct. Shared pages are fully managed by KVM, not by the TDX
module, regardless of whether migration is in progress or not. They are
migrated the standard QEMU way, same as for traditional VMs, so post_load
and any other migration code dealing with shared memory should work exactly
as it does today. Again, good to know how it is for other CoCo VMs.
> > Directly assigned devices are only possible with TDX Connect, where a
> > physical PCIe/CXL function (a "TDI") is assigned to the TD and can DMA into
> > private memory over a cryptographically protected link. This is not
> > implemented in Linux yet. For migration, the TDX module requires all TDIs to
> > be unassigned before the source TD is paused - the TDH.EXPORT.PAUSE seamcall
> > actually checks this. Unassigning a TDI tears down its whole TD-private
> > footprint (MMIO unmapped from the Secure EPT, trusted DMA mappings removed),
> > so no device-specific state is left to migrate. From the TD's point of view
> > it is a full hot-unplug on the source and a fresh hot-plug on the
> > destination.
>
> Hmm, interesting. Sorry if this follow up question may be slightly
> off-topic of the new API, or maybe it matters, depends on the answer: do
> you know who is in charge of this unplug / plug operation? VMM or TD
> (hence, transparent to VMM)?
>
> If it's VMM, is QEMU involved?
Specifically about TDX - the pause seamcall will return an error if TDIs
are not unassigned.
While this is something that is not implemented in Linux yet, I believe the
model will be that there is some uAPI to unassign TDIs, and it is not
related to migration. QEMU would just need to exercise this uAPI at the
right time.
> >
> > But I wish it were an independent feature instead. Then we could work on
> > upstreaming it on its own - Tony estimates it is about 20% of the current
> > TDX migration PoC code.
>
> Unexpected, that's a lot just for tracking.
I assume because it brings various seamcall wrappers and some shared infra
code. But also Tony might have over-estimated - I'll let him comment on
this.
> > I already raised this with the Intel TDX module architects, and they asked
> > for use-cases. The only one we came up with is QEMU estimating the TD dirty
> > rate before starting migration (the calc_dirty_rate command). I understand
> > their position: without a use-case there is little reason to implement it,
> > and they would also need to study the security implications - can it help
> > an attacker in some way?
> >
> > So if you or anyone else can educate me about use-cases for independent
> > dirty page scanning, I would really appreciate it - I could take them back
> > to the TDX module architects.
>
> Yes, calc_dirty_rate will use it, I also can't think of another use case
> that needs it. RHEL supports it, so it would be nice this will be
> supported for CoCo too.
Do you or someone else know how important it is to have it working?
Do actual customers use it in real setups and rely on it?
How bad or painful is it if this feature is not available?
The reason I am asking is that my assumption was that it is not important.
But if it is, I will come back to the TDX module architects with a
request to revise the design to support independent dirty page tracking. Of
course they may have some reasons for not doing it, but I would try at
least.
> QEMU also has another way to do it without KVM tracking, which is kind of
> simle page hash daemon fully done in userspace. But in case of CoCo it'll
> stop working too when most memory unreadable. So GET_DIRTY_LOG seems the
> only way to go.
>
> We can still "emulate" it by initiating a remote migration just to collect
> this info, I believe all attestations will simply pass and we do a fallback
> when the admin turns it off.. but it's awkward and all the rest (including
> trying to find another host suitable for migration, allocate resources..)
> are pure wastes.
Understood, thanks - sounds like this workaround is doable, but wasteful,
something to avoid if possible. That in itself is useful evidence I can
bring back to the TDX module architects.
Just to understand the real-world deployments better: do you know if anyone
actually falls back to this workaround today, or is it more of a theoretical
idea that nobody actually uses?
>
> Btw, do you know if the private pages will be write-trappable by the host
> kernel, with things like userfaultfd-wp or soft-dirty? That'll be another
> way to do this without TD involvement, but I don't know well enough to say.
The short answer is - no, private pages are not write-trappable outside of
a migration session, so trapping cannot be used as a workaround for
independent dirty tracking today.
But a bit longer answer is that the TDX module provides 2 alternative dirty
page tracking mechanisms, and one of them is in fact a write-trapping
mechanism. In the TDX specs, these mechanisms are called Write-Blocking
Export and Non-Blocking Export.
- Write-blocking: the VMM write-blocks pages using `TDH.EXPORT.BLOCKW`, and
gets a VM exit (an EPT violation) when the TD attempts to write a
blocked page.
- Non-blocking: based on Secure EPT scanning, using `TDH.MEM.SCAN.RANGE`,
which was mentioned earlier in this cover letter. I refer to it as dirty
scanning.
Today, both mechanisms are only available during migration and require the
migration setup to be done first - neither can be used as a standalone. In
our PoC we use the dirty scanning mechanism, as we expect it to be more
performant.
But our PoC actually started with the write-blocking method. According to
Kishen and Tony, it required a lot of ugly code and was very intrusive into
the MMU code. But I'll let Kishen and Tony provide more details.
Anyway, if I can get the TDX module to allow dirty tracking independent of
migration, we could pick either mechanism, or even implement both and add a
TDX-specific ioctl to select which one to use. We could plug either method
into the KVM_DIRTY_LOG ioctl - both would work, just differently.
= Dirty Scanning Optimization in TDX Module =
Now, I am diverging, but just in case: in the PUCK call where we presented
TDX migration, Sean made an immediate observation that WP-based dirty
tracking is not categorically worse than PML-based dirty ring tracking, it
depends on the workload. Then he learned that TDX module's dirty scanning
does not use PML, and was understandably surprised. Sean was concerned about
dirty scanning performance.
So what I learned then about dirty scanning is that it is optimized for
minimizing the downtime. Before the source TD is paused, there are a couple
of seamcalls to invoke, let me refer to them as prescan seamcalls.
They scan the Secure EPT while the source is running, and mark the
sub-trees that do not have migration candidates as "clean". Then during
the actual final dirty scan in the downtime window, the scan can skip
those entire sub-trees and be more efficient.
Sean correctly pointed out that this would be problematic because on Intel
CPUs the EPT dirty flag is set only at leaf level, and does not propagate
all the way to the root level. However, what I learned is that the TDX
module uses the "accessed" bit instead of the dirty bit in this prescan
optimization - the accessed bit does propagate all the way to the root
level. I thought this was a nifty trick, but obviously it would also mark
sub-trees as "not-clean" on read access.
Now, I personally did not benchmark this optimization, but I heard that
some people did and the results showed acceptable performance.
> It just sounds like if it's doable in KVM, it's still the best place, also
> since GET_DIRTY_LOG is available as long as memslot marked tracking, it
> sounds good to keep CoCo be compatible with that API if it is to be reused,
> hence GET_DIRTY_LOG API is consistent across coco / non-coco.
Yes, that's exactly the goal and the proposal: QEMU would use GET_DIRTY_LOG
API regardless of the VM type. In TDX guest case, KVM would use TDX-specific
dirty page scanning seamcalls to implement the GET_DIRTY_LOG API.
Thanks, Artem.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
2026-09-22 8:09 ` Artem Bityutskiy
@ 2026-09-22 9:42 ` Tony Lindgren
2026-09-22 11:54 ` Artem Bityutskiy
2026-09-22 21:18 ` Peter Xu
2026-09-23 15:28 ` Serge Hallyn (AMD)
2 siblings, 1 reply; 78+ messages in thread
From: Tony Lindgren @ 2026-09-22 9:42 UTC (permalink / raw)
To: Artem Bityutskiy
Cc: Peter Xu, Paolo Bonzini, Sean Christopherson, Fabiano Rosas,
Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
On Tue, Sep 22, 2026 at 11:09:42AM +0300, Artem Bityutskiy wrote:
> On Fri, 2026-09-18 at 11:53 -0400, Peter Xu wrote:
> > Unexpected, that's a lot just for tracking.
>
> I assume because it brings various seamcall wrappers and some shared infra
> code. But also Tony might have over-estimated - I'll let him comment on
> this.
Yes dirty log support is currently about 20% of the number of patches for
the POC. About 100 LOC changes, about 75% it is TDX specific code with
SEAMCALLs and and the scanning functions.
> > > I already raised this with the Intel TDX module architects, and they asked
> > > for use-cases. The only one we came up with is QEMU estimating the TD dirty
> > > rate before starting migration (the calc_dirty_rate command). I understand
> > > their position: without a use-case there is little reason to implement it,
> > > and they would also need to study the security implications - can it help
> > > an attacker in some way?
> > >
> > > So if you or anyone else can educate me about use-cases for independent
> > > dirty page scanning, I would really appreciate it - I could take them back
> > > to the TDX module architects.
> >
> > Yes, calc_dirty_rate will use it, I also can't think of another use case
> > that needs it. RHEL supports it, so it would be nice this will be
> > supported for CoCo too.
>
> Do you or someone else know how important it is to have it working?
> Do actual customers use it in real setups and rely on it?
> How bad or painful is it if this feature is not available?
>
> The reason I am asking is that my assumption was that it is not important.
> But if it is, I will come back to the TDX module architects with a
> request to revise the design to support independent dirty page tracking. Of
> course they may have some reasons for not doing it, but I would try at
> least.
>
> > QEMU also has another way to do it without KVM tracking, which is kind of
> > simle page hash daemon fully done in userspace. But in case of CoCo it'll
> > stop working too when most memory unreadable. So GET_DIRTY_LOG seems the
> > only way to go.
> >
> > We can still "emulate" it by initiating a remote migration just to collect
> > this info, I believe all attestations will simply pass and we do a fallback
> > when the admin turns it off.. but it's awkward and all the rest (including
> > trying to find another host suitable for migration, allocate resources..)
> > are pure wastes.
>
> Understood, thanks - sounds like this workaround is doable, but wasteful,
> something to avoid if possible. That in itself is useful evidence I can
> bring back to the TDX module architects.
>
> Just to understand the real-world deployments better: do you know if anyone
> actually falls back to this workaround today, or is it more of a theoretical
> idea that nobody actually uses?
Interesting workaround :) I too agree that proper dirty log support is the
way to go. To me it seems non-migration dirty log support can be added to
the TDX module as an additional feature. And adding it probably would not
even need KVM changes except a new flag for the scan SEAMCALL.
> > Btw, do you know if the private pages will be write-trappable by the host
> > kernel, with things like userfaultfd-wp or soft-dirty? That'll be another
> > way to do this without TD involvement, but I don't know well enough to say.
>
> The short answer is - no, private pages are not write-trappable outside of
> a migration session, so trapping cannot be used as a workaround for
> independent dirty tracking today.
>
> But a bit longer answer is that the TDX module provides 2 alternative dirty
> page tracking mechanisms, and one of them is in fact a write-trapping
> mechanism. In the TDX specs, these mechanisms are called Write-Blocking
> Export and Non-Blocking Export.
>
> - Write-blocking: the VMM write-blocks pages using `TDH.EXPORT.BLOCKW`, and
> gets a VM exit (an EPT violation) when the TD attempts to write a
> blocked page.
> - Non-blocking: based on Secure EPT scanning, using `TDH.MEM.SCAN.RANGE`,
> which was mentioned earlier in this cover letter. I refer to it as dirty
> scanning.
>
> Today, both mechanisms are only available during migration and require the
> migration setup to be done first - neither can be used as a standalone. In
> our PoC we use the dirty scanning mechanism, as we expect it to be more
> performant.
>
> But our PoC actually started with the write-blocking method. According to
> Kishen and Tony, it required a lot of ugly code and was very intrusive into
> the MMU code. But I'll let Kishen and Tony provide more details.
Yes write-blocking has issues with being intrusive. Exported pages are
locked by the TDX module and only cleared on import or after a cancel
operation. KVM MMU error handling gets tricky. The non-blocking migration
makes the tricky parts go away at the cost of adding TDX specific code
to handle the GET_DIRTY_LOG scanning.
> Anyway, if I can get the TDX module to allow dirty tracking independent of
> migration, we could pick either mechanism, or even implement both and add a
> TDX-specific ioctl to select which one to use. We could plug either method
> into the KVM_DIRTY_LOG ioctl - both would work, just differently.
Note that for TDX write-blocking is an older approach. The non-blocking
SEAMCALLs were added because of the issues noticed. Both features are not
usable the same time.
If non-blocking is enabled for a TDX module write-blocking cannot be used.
Also note that the dirty log scan features depend on non-blocking features
being enabled.
IMO no reason for Linux to try to support the write-blocking migration at
all.
> = Dirty Scanning Optimization in TDX Module =
>
> Now, I am diverging, but just in case: in the PUCK call where we presented
> TDX migration, Sean made an immediate observation that WP-based dirty
> tracking is not categorically worse than PML-based dirty ring tracking, it
> depends on the workload. Then he learned that TDX module's dirty scanning
> does not use PML, and was understandably surprised. Sean was concerned about
> dirty scanning performance.
>
> So what I learned then about dirty scanning is that it is optimized for
> minimizing the downtime. Before the source TD is paused, there are a couple
> of seamcalls to invoke, let me refer to them as prescan seamcalls.
>
> They scan the Secure EPT while the source is running, and mark the
> sub-trees that do not have migration candidates as "clean". Then during
> the actual final dirty scan in the downtime window, the scan can skip
> those entire sub-trees and be more efficient.
>
> Sean correctly pointed out that this would be problematic because on Intel
> CPUs the EPT dirty flag is set only at leaf level, and does not propagate
> all the way to the root level. However, what I learned is that the TDX
> module uses the "accessed" bit instead of the dirty bit in this prescan
> optimization - the accessed bit does propagate all the way to the root
> level. I thought this was a nifty trick, but obviously it would also mark
> sub-trees as "not-clean" on read access.
>
> Now, I personally did not benchmark this optimization, but I heard that
> some people did and the results showed acceptable performance.
>
> > It just sounds like if it's doable in KVM, it's still the best place, also
> > since GET_DIRTY_LOG is available as long as memslot marked tracking, it
> > sounds good to keep CoCo be compatible with that API if it is to be reused,
> > hence GET_DIRTY_LOG API is consistent across coco / non-coco.
>
>
> Yes, that's exactly the goal and the proposal: QEMU would use GET_DIRTY_LOG
> API regardless of the VM type. In TDX guest case, KVM would use TDX-specific
> dirty page scanning seamcalls to implement the GET_DIRTY_LOG API.
Yes.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
2026-09-22 9:42 ` Tony Lindgren
@ 2026-09-22 11:54 ` Artem Bityutskiy
2026-09-23 4:20 ` Tony Lindgren
0 siblings, 1 reply; 78+ messages in thread
From: Artem Bityutskiy @ 2026-09-22 11:54 UTC (permalink / raw)
To: Tony Lindgren
Cc: Peter Xu, Paolo Bonzini, Sean Christopherson, Fabiano Rosas,
Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
On Tue, 2026-09-22 at 12:42 +0300, Tony Lindgren wrote:
>
> > Anyway, if I can get the TDX module to allow dirty tracking independent of
> > migration, we could pick either mechanism, or even implement both and add a
> > TDX-specific ioctl to select which one to use. We could plug either method
> > into the KVM_DIRTY_LOG ioctl - both would work, just differently.
>
> Note that for TDX write-blocking is an older approach. The non-blocking
> SEAMCALLs were added because of the issues noticed. Both features are not
> usable the same time.
But let me clarify one point: write-blocking (the WP-based mechanism) isn't
broken - it works fine. For fairness:
- MMU-intrusiveness - valid argument.
- It is "older" - not on its own a strong argument. Older does not
automatically make it categorically worse.
Side note: although I always try to remember that if there is a real need,
we can work with the TDX module architects on the ABI to reduce that
intrusiveness. I do not see the need yet, though.
> If non-blocking is enabled for a TDX module write-blocking cannot be used.
> Also note that the dirty log scan features depend on non-blocking features
> being enabled.
To make it more pronounced: Tony says the blocking vs. non-blocking TDX
export mode is:
- Global for all TDs, not a per-TD choice.
- Decided at TDX module initialization time, and cannot be changed without
re-initializing or updating the module. So in practice, it's a kernel
init time decision.
So my suggestion that both could be implemented and selected via a uAPI is
moot in practice.
Good point, thanks!
Artem.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
2026-09-22 8:09 ` Artem Bityutskiy
2026-09-22 9:42 ` Tony Lindgren
@ 2026-09-22 21:18 ` Peter Xu
2026-09-23 12:05 ` Artem Bityutskiy
2026-09-23 15:28 ` Serge Hallyn (AMD)
2 siblings, 1 reply; 78+ messages in thread
From: Peter Xu @ 2026-09-22 21:18 UTC (permalink / raw)
To: Artem Bityutskiy
Cc: Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas,
Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
On Tue, Sep 22, 2026 at 11:09:42AM +0300, Artem Bityutskiy wrote:
> Hi Peter,
Hello, Artem,
>
> thanks again for good comments and questions.
My pleasure if I helped anything at all.
[...]
> > Yes, thanks. I'm willing to hear more, and I'm a bit surprised that
> > non-coco migration didn't fit already well into it, because IIUC non-coco
> > needs less in this case, not more.
>
> Yeah, just thinking aloud now.
>
> - Suppose QEMU and KVM support CoCo migration.
> - Does it make sense to treat traditional VM as a CoCo VM for migration
> purposes?
> - Not for performance reasons - mmapped memory I/O is going to be more
> efficient than any sorts of export/import uAPIs.
> - Yes for testing/experimental purposes.
> - Would it add a lot of code and complexity: would the benefits outweigh the
> costs?
> - QEMU: Should require minimum amount of code. Would need a flag like
> "treat me as CoCo VM", may be some QEMU operator commands or cmdline
> options.
In QEMU, we could likely stick with calling them non-CoCo VMs even if the
new migration API will work for them.
We can have a qemu flag (kvm_coco_migration_api), which will be required /
enforced by CoCo VMs to migrate. Then if the new API will support
non-CoCo, we can export that kvm flag so non-CoCo can select it too.
> - KVM: Not sure, but intuitively it may add noticeable amount of churn.
True. Also because of this, personally I would not request for this to be
supported, only if you all see fit or value.
Said that, this PoC problem that Kishen described PAM in the other email
helped me to notice something very important that I overlooked, that at
least I want to understand. I'll reply there later separately.
[...]
> > I am just thinking out loud here: if the hardware is good enough to do
> > encryption plus (some?) compression, that would be very nice. For "some",
> > I meant minimum over zero pages. Because in this case even if host wants
> > to play tricks with zero pages, it can't anymore when un-readable. Maybe
> > the guest driver can play some trick, but I'm also not sure if in CoCo.
> >
> > M > N may imply it's not the case for now, but it's still sane as a start
> > even if so.
>
> Yeah. In case of TDX, the TDX module does not do compression today,
> and there is no zero-page optimization either.
>
> But keep in mind that what pages in TD are zero pages is by itself a secret.
> Any sort of theoretical TDX module zero page compression optimization would
> need to be done in a way that VMM cannot figure out what pages were zero
> pages. So with a disclaimer that I am not really a security expert, I'd say
> it is far-fetched.
Agreed. That's why I think TDX, if able to compress zero pages in that
buffer stream, will hide that information into the encrypted stream, so
it's still not visible to VMM, but should be visible to destination TDX.
But yeah, this is slightly off topic, we can skip this one for now
regardless.
[...]
> OK, so by upper bound here you mean the maximum amount of data that can be
> exported in a single memory export ioctl operation, right? And you mean that
> there should be some sort of command to query it, similar to KVM_CAP_XSAVE2
> capability returning the XSAVE buffer size? If so - yes, I think exposing
> such an upper bound would be useful.
>
> Let's see. In case of TDX, the `TDH.EXPORT.MEM` seamcall today can export
> max. 512 pages at a time. Today only 4KiB pages are supported, so it is just
> a 2MiB buffer. So this is the upper bound for the seamcall.
>
> So in TDX case, 512 pages would be the absolute upper bound for the uAPI
> input buffer size. The alternative is to accept larger input buffers and
> internally do multiple seamcalls. For TDX case, I do not think the latter
> makes much sense though. But, what I do think is that the uAPI itself
> shouldn't mandate one approach over the other - a different CoCo
> implementation may prefer to aggregate multiple hardware calls internally.
>
> Also, today TDX can only migrate 4KiB pages - VMM must split larger pages
> into 4KiB pages before migration. But I expect that in the future TDX may
> implement larger pages migration too. So I think the cap should be expressed
> in page count rather than bytes.
>
> That said, I don't think exposing this upper bound cap should replace the
> flexibility we discussed earlier: the API itself should still allow
> exporting as much as fits the output buffer, regardless of what the cap
> reports. A specific CoCo implementation could still choose to just fail if
> the buffer isn't large enough to hold all requested pages - but that would
> be an implementation limitation, not an API limitation.
>
> And the other question is whether to allow exporting a mix of large and
> small pages, like N x 4KiB along with M x 2MiB and K x 1GiB pages. Or the
> uAPI would allow to only export one page size at a time?
>
> If large pages are added to TDX, I'd speculate that TDX seamcall would allow
> the mixing and matching - I can see that already today the seamcall API
> is provisioned for this. uAPI could require that the input buffer (set of
> pages) can contain a mix which should match the GPAs of the corresponding
> memory regions. But then race conditions - the GPA layout of the VM may
> change by the time the request reaches the hardware: large pages can be
> split or the other way round, I guess? Then uAPI could return a "retry" sort
> of exit code, may be?
>
> What do you think? I am just thinking aloud: looks like pages of different
> size add a degree of complexity - I did not take that into account until
> this conversation, thanks.
I apologize if my question was misleading. My question was more about the
size of user buffer QEMU needs to prepare. I think I got the answer from
the other email.
Still, thanks for sharing your thoughts on huge page handling. Personally
I'm not sure if we need to identify huge pages: we can always still define
the GFNs to always represent PAGE_SIZE no matter the size of host pages,
especially in migration context, we'll need to split in the first place.
But that'll be a question to ask for later.
>
> >
> > Said that, we'll need to be careful then in case of migration fallbacks at
> > the final stage. Nowadays, I believe QEMU can still fallback to source side
> > at a very, very late stage after all things applied. If I'm not mistaken,
> > the final handshake is done at migration_incoming_state_destroy() ->
> > migrate_send_rp_shut() telling source to be gone.
>
> Right. In traditional pre-copy VM migration model, the fallback is possible
> at any point before the destination VM starts running and modifying its
> state.
>
> The same is true for TDX, but with more complexity and limitations.
>
> We have 2 points:
> 1. Before the source has exported the start token - fallback is similar to
> traditional VM migration - just abort the migration on source and
> continue running the source TD, and just destroy the destination TD.
> 2. After the source has exported the start token - fallback is still
> possible, but more complex - it requires the destination to first
> generate the abort token, which should be delivered to the source and
> consumed there. There are seamcalls for both generating and consuming the
> abort token. The token basically makes sure the destination cannot run,
> and the source can run again - same "only one of the TDs can run at a
> time" security rule that we discussed earlier.
>
> Now, the limitation here is that if the abort is because the network
> connectivity is lost, the abort token cannot be delivered. Then I'd guess
> migration should stop, without destroying the source and the destination,
> and the token should be delivered later manually, QEMU could even provide an
> infrastructure / commands for that.
Heh, this reminded me of postcopy recover that QEMU supported for years:
not a trivial feature at all, people always want it to be there, but I'm
not sure who is using it at all in production..
Especially, normally who cares about migration will provide dedicated and
reliable networking for migration purpose, so the failure will be even more
unlikely. Meanwhile who doesn't care enough, may simply crash the VM at a
postcopy failure and reboot.. not bother to recover.
I bet rarely people will hit a network failure happens exactly at
forwarding the TDX START token to destination host, similarly for case (2)
on ABORT token lost. But yes, likely this is required to make the whole
thing complete.
>
> Would be very helpful to learn the abort protocol the other CoCo vendors
> offer - is it similar to TDX or not?
>
> In our PoC we did not implement the TDX abort token procedure. But the
> proposed uAPI is provisioned for the abort token export/import.
[...]
> > Hmm, interesting. Sorry if this follow up question may be slightly
> > off-topic of the new API, or maybe it matters, depends on the answer: do
> > you know who is in charge of this unplug / plug operation? VMM or TD
> > (hence, transparent to VMM)?
> >
> > If it's VMM, is QEMU involved?
>
> Specifically about TDX - the pause seamcall will return an error if TDIs
> are not unassigned.
>
> While this is something that is not implemented in Linux yet, I believe the
> model will be that there is some uAPI to unassign TDIs, and it is not
> related to migration. QEMU would just need to exercise this uAPI at the
> right time.
OK, this sounds working, but then it means the migration will be visible to
the guest. I wonder whether there's any attempt to make it more
transparent, but we can also leave this question for later.
[...]
> Do you or someone else know how important it is to have it working?
> Do actual customers use it in real setups and rely on it?
> How bad or painful is it if this feature is not available?
>
> The reason I am asking is that my assumption was that it is not important.
> But if it is, I will come back to the TDX module architects with a
> request to revise the design to support independent dirty page tracking. Of
> course they may have some reasons for not doing it, but I would try at
> least.
Thanks, I'll talk to our team and revisit this after I collect answers.
[...]
> Understood, thanks - sounds like this workaround is doable, but wasteful,
> something to avoid if possible. That in itself is useful evidence I can
> bring back to the TDX module architects.
>
> Just to understand the real-world deployments better: do you know if anyone
> actually falls back to this workaround today, or is it more of a theoretical
> idea that nobody actually uses?
Nop, I am not aware of any CoCo migration worked in any production
deployments. So it's only a wild idea assuming GET_DIRTY_LOG will be gone
without migration process.
>
> >
> > Btw, do you know if the private pages will be write-trappable by the host
> > kernel, with things like userfaultfd-wp or soft-dirty? That'll be another
> > way to do this without TD involvement, but I don't know well enough to say.
>
> The short answer is - no, private pages are not write-trappable outside of
> a migration session, so trapping cannot be used as a workaround for
> independent dirty tracking today.
>
> But a bit longer answer is that the TDX module provides 2 alternative dirty
> page tracking mechanisms, and one of them is in fact a write-trapping
> mechanism. In the TDX specs, these mechanisms are called Write-Blocking
> Export and Non-Blocking Export.
>
> - Write-blocking: the VMM write-blocks pages using `TDH.EXPORT.BLOCKW`, and
> gets a VM exit (an EPT violation) when the TD attempts to write a
> blocked page.
> - Non-blocking: based on Secure EPT scanning, using `TDH.MEM.SCAN.RANGE`,
> which was mentioned earlier in this cover letter. I refer to it as dirty
> scanning.
>
> Today, both mechanisms are only available during migration and require the
> migration setup to be done first - neither can be used as a standalone. In
> our PoC we use the dirty scanning mechanism, as we expect it to be more
> performant.
>
> But our PoC actually started with the write-blocking method. According to
> Kishen and Tony, it required a lot of ugly code and was very intrusive into
> the MMU code. But I'll let Kishen and Tony provide more details.
>
> Anyway, if I can get the TDX module to allow dirty tracking independent of
> migration, we could pick either mechanism, or even implement both and add a
> TDX-specific ioctl to select which one to use. We could plug either method
> into the KVM_DIRTY_LOG ioctl - both would work, just differently.
>
> = Dirty Scanning Optimization in TDX Module =
>
> Now, I am diverging, but just in case: in the PUCK call where we presented
> TDX migration, Sean made an immediate observation that WP-based dirty
> tracking is not categorically worse than PML-based dirty ring tracking, it
> depends on the workload.
Hmm, I was expecting PML is still superior in most cases. For "depending
on workloads", is that perhaps when (1) huge pages are used, and (2) the
workload writes only a small portion of guest memory?
Then maybe RO huge page will be a good hint to skip whole huge-page-range.
But that's only my wild guess, could be wrong.
> Then he learned that TDX module's dirty scanning does not use PML, and
> was understandably surprised. Sean was concerned about dirty scanning
> performance.
I'm definitely surprised too that PML isn't used. Could I ask if there's
any simple reason not to use it for TDX? Per my understanding, PML works
with all kinds of loads, and I was expecting PML to be efficient and most
ideal.
Another approach is if TDX can take over the bitmap buffer from the
relevant kvm memslots, update directly there alongside setting D bits in
EPT PTEs; after all IIUC we assumed dirty info not part of confidential
materials. But that sounds more complex than PML if it's already working
for years.
The current scan approach sounds like unpredictable in terms of downtime,
in that even if with a scanner I don't see how TDX can guaratee the
downtime for the last dirty sync from QEMU, which will completely be part
of the blackout downtime.
Whenever switchover decision made, QEMU stops the VM, do the last time
sync, move anything left. So that last one will matter a lot if we care
about downtime.
Thanks,
--
Peter Xu
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-22 6:27 ` Tony Lindgren
@ 2026-09-23 0:37 ` Kishen Maloor
2026-09-23 6:50 ` Tony Lindgren
0 siblings, 1 reply; 78+ messages in thread
From: Kishen Maloor @ 2026-09-23 0:37 UTC (permalink / raw)
To: Tony Lindgren
Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Artem Bityutskiy,
Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky,
Marc Zyngier, Oliver Upton, Steven Price, Anup Patel,
Samuel Ortiz, Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg,
Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm
On 9/21/26 11:27 PM, Tony Lindgren wrote:
> On Mon, Sep 21, 2026 at 08:56:48PM -0700, Kishen Maloor wrote:
>> On 9/20/26 11:56 PM, Tony Lindgren wrote:
>>> On Mon, Sep 21, 2026 at 08:58:17AM +0300, Tony Lindgren wrote:
>>>> On Thu, Sep 17, 2026 at 09:33:18PM -0700, Kishen Maloor wrote:
>>>>> On 8/31/26 12:13 AM, Tony Lindgren wrote:
>>>>>
>>>>> I was wondering how a vendor would use kvm_transfer_buffer for a command that
>>>>> passes an input and also returns an output. size can describe the input
>>>>> length or the space available for output, not both, so the kernel has no way to
>>>>> know how much it may write.
>>>>>
>>>>> Two suggestions below. They're orthogonal.
>>>>>
>>>>>> ...
>>>>>> diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h
>>>>>> index 5f6c1ce9673b7..d9291a8a97bb1 100644
>>>>>> --- a/arch/x86/include/asm/kvm_host.h
>>>>>> +++ b/arch/x86/include/asm/kvm_host.h
>>>>>> @@ -2010,6 +2010,8 @@ struct kvm_x86_ops {
>>>>>> int (*gmem_prepare)(struct kvm *kvm, kvm_pfn_t pfn, gfn_t gfn, int max_order);
>>>>>> void (*gmem_invalidate)(kvm_pfn_t start, kvm_pfn_t end);
>>>>>> int (*gmem_max_mapping_level)(struct kvm *kvm, kvm_pfn_t pfn, bool is_private);
>>>>>> + bool (*cap_live_migration)(struct kvm *kvm);
>>>>>
>>>>> Should this return an 'int' specifying the maximum buffer size that the vendor impl
>>>>> requires/uses? Userspace can then learn this once.
>>>>
>>>> OK that sounds good to me. I assume you are thinking the maximum hardware
>>>> specific buffer size per thread? As in 512 * 4096 bytes for the TDX case?
>>
>> Yeah, an upper bound on the buffer size, so 516 * 4096 (4 for the GPA+MAC lists
>> and MBMD) in TDX. It could be a hint to userspace to size its buffers at the
>> start. Though it would be up to userspace to decide how to use that information.
>
> Oh right thanks. I think this is really the maximum transfer buffer size
> Peter asked, not just a hint to userspace :)
It is a strict upper bound. Maybe it's semantics, but I called
it a "hint" because allocating for that entire size could be optional.
If a userspace driver for say TDX wants to send smaller batches (say 128) then
it can refer to the spec, do the math, and allocate 131 pages and the kernel
should permit that; it's not wrong. If userspace ever allocates less room than
a call requires, it would fail. A different userspace driver could simply
allocate that max size and be done; no need to refer to the spec or do the math.
It is for this second case where I thought returning the upper bound would be
useful. Hence the earlier suggestion.
>
>>>>>> +struct kvm_transfer_buffer {
>>>>>> + __u64 address;
>>>>>> + __u32 size;
>>>>>> + __u32 reserved;
>>>>>> +};
>>>>>
>>>>> Should this struct include a 'capacity' field (u32) that is set on each command?
>>>>> It would be the number of bytes writable at address.
>>>>> size would be the input length on entry (0 if the command passes none), and the
>>>>> number of bytes produced on return (0 if none).
>>>>
>>>> Hmm so the transfer command return value can return how many bytes were
>>>> written of the input. But yeah we don't know how many bytes were written
>>>> back to the transfer buffer as result of the transfer command.
>>
>> The transfer command return value could return how many bytes were written into the buffer.
>> But in an input-output call, the kernel handler wouldn't know how many bytes it could write,
>> or for that matter even how many pages to pin up front in case it needs to return an output
>> because 'size' couldn't simultaneously convey the input length and buffer capacity. That was
>> the gap that I thought a read-only 'capacity' field could bridge. Of course, this
>> assumes that the output is written in-place.
>
> Hmm yeah this inplace capacity vs transferred issue remains still. So I
> agree we need to specify the capacity in struct kvm_transfer_buffer like
> you suggested.
To be clear, I think the in/out split for the buffers along with the convention
I laid out closes that gap I saw without needing a 'capacity' field.
Because out/size could now unambiguously convey capacity on entry and output
length on return.
>
> To me size is already the size of the buffer though. So instead of changing
> size to capacity, how about something like datasize or len for the input
> and output transfer length?
But I understand that (and please correct me if I'm wrong):
a) You'd still prefer to not have 'size' serve that double duty.
b) 'size' in your mental model already means buffer capacity.
In that case, we could add a 'datasize' field to convey the length
of valid data in the buffer, like this:
struct kvm_transfer_buffer {
__u64 address;
__u32 size;
__u32 datasize;
__u64 reserved;
};
The convention then becomes:
- A non-zero 'datasize' on 'in' at call entry conveys that there is input.
- A non-zero 'datasize' on 'out' at call exit conveys that there is output.
- out/datasize on call entry is ignored.
- in/size and out/size are seeded with the buffer capacity.
>
>>>> How about if we add the bytes returned to the transfer struct? Then the
>>>> kvm_transfer_buffer can stay as just a buffer.
>>>
>>> Actually, for the possible cases with input+output, we could reserve space
>>> in the transfer struct for another struct kvm_transfer_buffer for the
>>> results?
>>
>> Yes, say an 'in' and 'out' kvm_transfer_buffer inside struct kvm_migrate_cmd should
>> close this out and shouldn't require a 'capacity' field. Maybe we then establish this
>> convention:
>> - A non-zero 'size' on 'in' at call entry would signal that there is input.
>> - A non-zero 'size' on 'out' at call entry would convey the buffer capacity.
>> - A non-zero 'size' on 'out' at call exit would convey that there is output.
>> - A zeroed 'size' on 'out' at call exit would convey that there is no output.
>
> Looks doable to me but with the inplace issue as above.. Sounds like we just
> need to reserve space for a case with a separate output buffer though.
To be clear, this is what I thought we were talking about :)
To add a 2nd kvm_transfer_buffer to kvm_migrate_cmd, like this:
struct kvm_migrate_cmd {
__u16 command;
__u16 flags;
__u32 reserved;
struct kvm_transfer_buffer in;
struct kvm_transfer_buffer out;
};
If there is agreement on this model, then yeah, we'd want to define
these fields now, since the struct can't grow later without a new ioctl
number.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-22 5:25 ` Tony Lindgren
@ 2026-09-23 0:38 ` Kishen Maloor
2026-09-23 6:04 ` Tony Lindgren
0 siblings, 1 reply; 78+ messages in thread
From: Kishen Maloor @ 2026-09-23 0:38 UTC (permalink / raw)
To: Tony Lindgren
Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini,
Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On 9/21/26 10:25 PM, Tony Lindgren wrote:
> On Mon, Sep 21, 2026 at 08:57:45PM -0700, Kishen Maloor wrote:
>> On 9/20/26 11:52 PM, Tony Lindgren wrote:
>>> Then for KVM tracking the role, I don't think we need it with the two
>>> above. The role tracking can always be added if really needed. Any other
>>> opinions on this one?
>>
>> I do think it's useful for generic KVM to track this role (1 byte).
>> - It lets _TRANSFER_ style calls dispatch directly to import or export
>> callbacks based on the role.
>> - Even if we don't adopt the _TRANSFER_ style, it enables generic KVM to
>> reject mismatched calls, e.g., KVM_EXPORT_MEMORY on a destination.
>
> Having KVM do generic checks on the calls is a good idea. There might be
> a simpler way of handling it though. Rather than having KVM track the
> migration state, how about we add a function to check for the migration
> session state from the vendor code?
>
> So something like this for the states you suggested earlier:
>
> enum kvm_lmstate {
> KVM_LM_NONE,
> KVM_LM_SOURCE,
> KVM_LM_DESTINATION,
> };
>
> With something like this to get the state from the vendor code:
>
> enum kmv_lmstate kvm_arch_get_lmstate(struct kvm *);
>
> For x86 it would end up calling kvm_x86_call(get_lmstate)(kvm) and for
> the TDX specific case tdx_get_lmstate().
How is this simpler? It trades one byte in a KVM struct for a new generic
enum, a new kvm_arch_get_lmstate(), a new kvm_x86_ops entry, and a vendor
implementation per vendor, plus a cross-layer call on every command just
to learn the role.
Directly checking a stored byte (0=unset/1=src/2=dst) seems simplest, no?
> It would allow KVM to do the generic checks for the migration related
> calls you're describing. And having KVM start tracking the state can be
> still added later on too if it is needed.
I think we'd need to pick one way or the other before the UAPI settles
if we want generic KVM to reject mismatched calls. OTOH if we want to
defer this generic KVM validation, then yeah, it could be settled later.
>
>>> The transfer direction is there with the EXPORT/IMPORT naming. Maybe
>>> just let's keep that naming for easier readability rather than try to
>>> switch to TRANSFER style naming. No transfer direction flag needed.
>>
>> I can't say I have a clear preference between EXPORT/IMPORT vs _TRANSFER_.
>> If we keep the EXPORT/IMPORT naming, then consistency would arguably call for
>> splitting MIGRATE_CMD too, which makes it 6 vs 3 (or 5 vs 3 against the RFC
>> as posted):
>>
>> EXPORT/IMPORT style _TRANSFER_ style
>> KVM_EXPORT_CMD KVM_MIGRATE_CMD
>> KVM_IMPORT_CMD
>> KVM_EXPORT_MEMORY KVM_TRANSFER_MEMORY
>> KVM_IMPORT_MEMORY
>> KVM_EXPORT_VCPU KVM_TRANSFER_VCPU
>> KVM_IMPORT_VCPU
>
> Agreed we should split the MIGRATE_CMD too. Probably the number of ioctls
> is not and issue compared to following the KVM style and better
> readability. So my vote is now on EXPORT/IMPORT style naming.
>
>> I guess the main benefit of the _TRANSFER_ style is reduced duplication.
>> Each EXPORT/IMPORT pair takes the same struct and differs only by the role
>> that the session already knows. Merging them gives one entry point per call
>> type and lets userspace drive both ends from the same call site which could
>> be considered a win. It doesn't reduce kernel code though as the top-level
>> handler still branches internally on the role.
>
> Yup not much of a win for the TRANSFER style naming.
Yeah, my goal was just to enumerate alternatives for consideration.
One other benefit of KVM_EXPORT_CMD/KVM_IMPORT_CMD is that the per-session
role is implicitly conveyed - in other words, a successful KVM_EXPORT_CMD/SETUP
(for e.g.) would indicate that this is a 'source'. So, we wouldn't require a 'role'
field in struct kvm_migrate_cmd to explicitly assert one during SETUP.
>
> Trying to summarize again after we sorted out the direction flag issue in
> the transfer:
>
> role per migration session (cannot change during the migration)
> direction per command, EXPORT/IMPORT
> hardware state set and tracked by vendor specific code
The 'role' and 'direction' as you define it above are essentially saying the
same thing - a source only invokes the vendor's EXPORT call, a destination only
invokes the vendor's IMPORT call, and the role doesn't change over the session.
If a destination needs to send data to be consumed by the source, then that
still invokes the vendor's EXPORT call.
> And checking again against the dmaengine analogy:
>
> Compared to dmaengine, the migration role is modeled similar to the dma
> channel configuration.
>
> The migration direction with EXPORT/IMPORT is modeled similar to
> dmaengine_prep_slave_sg().
I'll leave the dmaengine comparison to you, I don't know that API well enough
to map it properly :)
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
2026-09-22 11:54 ` Artem Bityutskiy
@ 2026-09-23 4:20 ` Tony Lindgren
0 siblings, 0 replies; 78+ messages in thread
From: Tony Lindgren @ 2026-09-23 4:20 UTC (permalink / raw)
To: Artem Bityutskiy
Cc: Peter Xu, Paolo Bonzini, Sean Christopherson, Fabiano Rosas,
Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
On Tue, Sep 22, 2026 at 02:54:29PM +0300, Artem Bityutskiy wrote:
> On Tue, 2026-09-22 at 12:42 +0300, Tony Lindgren wrote:
> >
> > > Anyway, if I can get the TDX module to allow dirty tracking independent of
> > > migration, we could pick either mechanism, or even implement both and add a
> > > TDX-specific ioctl to select which one to use. We could plug either method
> > > into the KVM_DIRTY_LOG ioctl - both would work, just differently.
> >
> > Note that for TDX write-blocking is an older approach. The non-blocking
> > SEAMCALLs were added because of the issues noticed. Both features are not
> > usable the same time.
>
> But let me clarify one point: write-blocking (the WP-based mechanism) isn't
> broken - it works fine. For fairness:
>
> - MMU-intrusiveness - valid argument.
> - It is "older" - not on its own a strong argument. Older does not
> automatically make it categorically worse.
Correct. And I'm of course a bit biased on the write-blocking migration
after tinkering with it some so please excuse me.
For reference, there are earlier WIP patches from 2023 for TDX using the
write-blocking at [0] and [1] below. While the write-blocking itself seems
fairly straight forward, it's interaction with live migration is complicated
to unblock pages.
You may want to add a VM exit and TDX module unblock for each 4K page to the
list for the TDX write-blocking migration.
> Side note: although I always try to remember that if there is a real need,
> we can work with the TDX module architects on the ABI to reduce that
> intrusiveness. I do not see the need yet, though.
>
> > If non-blocking is enabled for a TDX module write-blocking cannot be used.
> > Also note that the dirty log scan features depend on non-blocking features
> > being enabled.
>
> To make it more pronounced: Tony says the blocking vs. non-blocking TDX
> export mode is:
>
> - Global for all TDs, not a per-TD choice.
> - Decided at TDX module initialization time, and cannot be changed without
> re-initializing or updating the module. So in practice, it's a kernel
> init time decision.
>
> So my suggestion that both could be implemented and selected via a uAPI is
> moot in practice.
Yup. If somebody really needs it, selecting write-blocking migration could
be a module param.
[0] https://github.com/intel-staging/tdx/commit/d6c63ebf463297a076984fe248aabeba0fa58864
[1] https://github.com/intel-staging/tdx/commit/782649ab74c1b3f05df7e38d5fc8e3f907a9394e
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-23 0:38 ` Kishen Maloor
@ 2026-09-23 6:04 ` Tony Lindgren
2026-09-24 5:53 ` Kishen Maloor
0 siblings, 1 reply; 78+ messages in thread
From: Tony Lindgren @ 2026-09-23 6:04 UTC (permalink / raw)
To: Kishen Maloor
Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini,
Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On Tue, Sep 22, 2026 at 05:38:43PM -0700, Kishen Maloor wrote:
> On 9/21/26 10:25 PM, Tony Lindgren wrote:
> > On Mon, Sep 21, 2026 at 08:57:45PM -0700, Kishen Maloor wrote:
> >> On 9/20/26 11:52 PM, Tony Lindgren wrote:
> >>> Then for KVM tracking the role, I don't think we need it with the two
> >>> above. The role tracking can always be added if really needed. Any other
> >>> opinions on this one?
> >>
> >> I do think it's useful for generic KVM to track this role (1 byte).
> >> - It lets _TRANSFER_ style calls dispatch directly to import or export
> >> callbacks based on the role.
> >> - Even if we don't adopt the _TRANSFER_ style, it enables generic KVM to
> >> reject mismatched calls, e.g., KVM_EXPORT_MEMORY on a destination.
> >
> > Having KVM do generic checks on the calls is a good idea. There might be
> > a simpler way of handling it though. Rather than having KVM track the
> > migration state, how about we add a function to check for the migration
> > session state from the vendor code?
> >
> > So something like this for the states you suggested earlier:
> >
> > enum kvm_lmstate {
> > KVM_LM_NONE,
> > KVM_LM_SOURCE,
> > KVM_LM_DESTINATION,
> > };
> >
> > With something like this to get the state from the vendor code:
> >
> > enum kmv_lmstate kvm_arch_get_lmstate(struct kvm *);
> >
> > For x86 it would end up calling kvm_x86_call(get_lmstate)(kvm) and for
> > the TDX specific case tdx_get_lmstate().
>
> How is this simpler? It trades one byte in a KVM struct for a new generic
> enum, a new kvm_arch_get_lmstate(), a new kvm_x86_ops entry, and a vendor
> implementation per vendor, plus a cross-layer call on every command just
> to learn the role.
> Directly checking a stored byte (0=unset/1=src/2=dst) seems simplest, no?
It would avoid dragging KVM into the "track the migration state" business
at least for now.
I guess the question in general is: What does KVM need to do with the
migration role beyond generic checks on the migration related calls?
> > It would allow KVM to do the generic checks for the migration related
> > calls you're describing. And having KVM start tracking the state can be
> > still added later on too if it is needed.
>
> I think we'd need to pick one way or the other before the UAPI settles
> if we want generic KVM to reject mismatched calls. OTOH if we want to
> defer this generic KVM validation, then yeah, it could be settled later.
>
> >
> >>> The transfer direction is there with the EXPORT/IMPORT naming. Maybe
> >>> just let's keep that naming for easier readability rather than try to
> >>> switch to TRANSFER style naming. No transfer direction flag needed.
> >>
> >> I can't say I have a clear preference between EXPORT/IMPORT vs _TRANSFER_.
> >> If we keep the EXPORT/IMPORT naming, then consistency would arguably call for
> >> splitting MIGRATE_CMD too, which makes it 6 vs 3 (or 5 vs 3 against the RFC
> >> as posted):
> >>
> >> EXPORT/IMPORT style _TRANSFER_ style
> >> KVM_EXPORT_CMD KVM_MIGRATE_CMD
> >> KVM_IMPORT_CMD
> >> KVM_EXPORT_MEMORY KVM_TRANSFER_MEMORY
> >> KVM_IMPORT_MEMORY
> >> KVM_EXPORT_VCPU KVM_TRANSFER_VCPU
> >> KVM_IMPORT_VCPU
> >
> > Agreed we should split the MIGRATE_CMD too. Probably the number of ioctls
> > is not and issue compared to following the KVM style and better
> > readability. So my vote is now on EXPORT/IMPORT style naming.
> >
> >> I guess the main benefit of the _TRANSFER_ style is reduced duplication.
> >> Each EXPORT/IMPORT pair takes the same struct and differs only by the role
> >> that the session already knows. Merging them gives one entry point per call
> >> type and lets userspace drive both ends from the same call site which could
> >> be considered a win. It doesn't reduce kernel code though as the top-level
> >> handler still branches internally on the role.
> >
> > Yup not much of a win for the TRANSFER style naming.
>
> Yeah, my goal was just to enumerate alternatives for consideration.
>
> One other benefit of KVM_EXPORT_CMD/KVM_IMPORT_CMD is that the per-session
> role is implicitly conveyed - in other words, a successful KVM_EXPORT_CMD/SETUP
> (for e.g.) would indicate that this is a 'source'. So, we wouldn't require a 'role'
> field in struct kvm_migrate_cmd to explicitly assert one during SETUP.
Yes good point with the KVM_EXPORT/IMPORT_CMD, that sounds good to me.
> > Trying to summarize again after we sorted out the direction flag issue in
> > the transfer:
> >
> > role per migration session (cannot change during the migration)
> > direction per command, EXPORT/IMPORT
> > hardware state set and tracked by vendor specific code
>
> The 'role' and 'direction' as you define it above are essentially saying the
> same thing - a source only invokes the vendor's EXPORT call, a destination only
> invokes the vendor's IMPORT call, and the role doesn't change over the session.
> If a destination needs to send data to be consumed by the source, then that
> still invokes the vendor's EXPORT call.
Yup.
> > And checking again against the dmaengine analogy:
> >
> > Compared to dmaengine, the migration role is modeled similar to the dma
> > channel configuration.
> >
> > The migration direction with EXPORT/IMPORT is modeled similar to
> > dmaengine_prep_slave_sg().
> I'll leave the dmaengine comparison to you, I don't know that API well enough
> to map it properly :)
Heh just a sanity check for trying to relate this to something existing.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-23 0:37 ` Kishen Maloor
@ 2026-09-23 6:50 ` Tony Lindgren
2026-09-24 5:34 ` Kishen Maloor
0 siblings, 1 reply; 78+ messages in thread
From: Tony Lindgren @ 2026-09-23 6:50 UTC (permalink / raw)
To: Kishen Maloor
Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Artem Bityutskiy,
Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky,
Marc Zyngier, Oliver Upton, Steven Price, Anup Patel,
Samuel Ortiz, Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg,
Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm
On Tue, Sep 22, 2026 at 05:37:33PM -0700, Kishen Maloor wrote:
> On 9/21/26 11:27 PM, Tony Lindgren wrote:
> > Oh right thanks. I think this is really the maximum transfer buffer size
> > Peter asked, not just a hint to userspace :)
>
> It is a strict upper bound. Maybe it's semantics, but I called
> it a "hint" because allocating for that entire size could be optional.
> If a userspace driver for say TDX wants to send smaller batches (say 128) then
> it can refer to the spec, do the math, and allocate 131 pages and the kernel
> should permit that; it's not wrong. If userspace ever allocates less room than
> a call requires, it would fail. A different userspace driver could simply
> allocate that max size and be done; no need to refer to the spec or do the math.
> It is for this second case where I thought returning the upper bound would be
> useful. Hence the earlier suggestion.
OK
> >>>>>> +struct kvm_transfer_buffer {
> >>>>>> + __u64 address;
> >>>>>> + __u32 size;
> >>>>>> + __u32 reserved;
> >>>>>> +};
> >>>>>
> >>>>> Should this struct include a 'capacity' field (u32) that is set on each command?
> >>>>> It would be the number of bytes writable at address.
> >>>>> size would be the input length on entry (0 if the command passes none), and the
> >>>>> number of bytes produced on return (0 if none).
> >>>>
> >>>> Hmm so the transfer command return value can return how many bytes were
> >>>> written of the input. But yeah we don't know how many bytes were written
> >>>> back to the transfer buffer as result of the transfer command.
> >>
> >> The transfer command return value could return how many bytes were written into the buffer.
> >> But in an input-output call, the kernel handler wouldn't know how many bytes it could write,
> >> or for that matter even how many pages to pin up front in case it needs to return an output
> >> because 'size' couldn't simultaneously convey the input length and buffer capacity. That was
> >> the gap that I thought a read-only 'capacity' field could bridge. Of course, this
> >> assumes that the output is written in-place.
> >
> > Hmm yeah this inplace capacity vs transferred issue remains still. So I
> > agree we need to specify the capacity in struct kvm_transfer_buffer like
> > you suggested.
>
> To be clear, I think the in/out split for the buffers along with the convention
> I laid out closes that gap I saw without needing a 'capacity' field.
> Because out/size could now unambiguously convey capacity on entry and output
> length on return.
But for an inplace buffer use with some input data smaller than the output
data, would it work? To me it seems you need both buffer size and data
size for that.
> > To me size is already the size of the buffer though. So instead of changing
> > size to capacity, how about something like datasize or len for the input
> > and output transfer length?
>
> But I understand that (and please correct me if I'm wrong):
> a) You'd still prefer to not have 'size' serve that double duty.
> b) 'size' in your mental model already means buffer capacity.
Heh yes correct for the above.
> In that case, we could add a 'datasize' field to convey the length
> of valid data in the buffer, like this:
>
> struct kvm_transfer_buffer {
> __u64 address;
> __u32 size;
> __u32 datasize;
> __u64 reserved;
> };
Maybe bufsize and datasize? Then the difference would be obvious while
reading the code.
> The convention then becomes:
> - A non-zero 'datasize' on 'in' at call entry conveys that there is input.
> - A non-zero 'datasize' on 'out' at call exit conveys that there is output.
> - out/datasize on call entry is ignored.
> - in/size and out/size are seeded with the buffer capacity.
>
> >
> >>>> How about if we add the bytes returned to the transfer struct? Then the
> >>>> kvm_transfer_buffer can stay as just a buffer.
> >>>
> >>> Actually, for the possible cases with input+output, we could reserve space
> >>> in the transfer struct for another struct kvm_transfer_buffer for the
> >>> results?
> >>
> >> Yes, say an 'in' and 'out' kvm_transfer_buffer inside struct kvm_migrate_cmd should
> >> close this out and shouldn't require a 'capacity' field. Maybe we then establish this
> >> convention:
> >> - A non-zero 'size' on 'in' at call entry would signal that there is input.
> >> - A non-zero 'size' on 'out' at call entry would convey the buffer capacity.
> >> - A non-zero 'size' on 'out' at call exit would convey that there is output.
> >> - A zeroed 'size' on 'out' at call exit would convey that there is no output.
> >
> > Looks doable to me but with the inplace issue as above.. Sounds like we just
> > need to reserve space for a case with a separate output buffer though.
>
> To be clear, this is what I thought we were talking about :)
Heh yeah we're talking two things with the inplace use vs two buffers :)
> To add a 2nd kvm_transfer_buffer to kvm_migrate_cmd, like this:
>
> struct kvm_migrate_cmd {
> __u16 command;
> __u16 flags;
> __u32 reserved;
> struct kvm_transfer_buffer in;
> struct kvm_transfer_buffer out;
> };
>
> If there is agreement on this model, then yeah, we'd want to define
> these fields now, since the struct can't grow later without a new ioctl
> number.
Based on what we've discussed, my preference is the following:
Keep the current buf naming. For the EXPORT/IMPORT type functions the use
should be obvious from the transfer type.
Reserve enough space for a separate output buffer or results buffer or
whatever it might get called if such a use case ever pops up.
Add the datasize to struct kvm_transfer_buffer like you suggested and
rename size to bufsize.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
2026-09-22 21:18 ` Peter Xu
@ 2026-09-23 12:05 ` Artem Bityutskiy
2026-09-24 21:19 ` Peter Xu
0 siblings, 1 reply; 78+ messages in thread
From: Artem Bityutskiy @ 2026-09-23 12:05 UTC (permalink / raw)
To: Peter Xu
Cc: Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas,
Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
On Tue, 2026-09-22 at 17:18 -0400, Peter Xu wrote:
> > thanks again for good comments and questions.
>
> My pleasure if I helped anything at all.
Honestly, yes. New insights, and even commenting on questions makes me
research more and dig deeper.
> >
> > Now, the limitation here is that if the abort is because the network
> > connectivity is lost, the abort token cannot be delivered. Then I'd guess
> > migration should stop, without destroying the source and the destination,
> > and the token should be delivered later manually, QEMU could even provide an
> > infrastructure / commands for that.
>
> Heh, this reminded me of postcopy recover that QEMU supported for years:
> not a trivial feature at all, people always want it to be there, but I'm
> not sure who is using it at all in production..
>
> Especially, normally who cares about migration will provide dedicated and
> reliable networking for migration purpose, so the failure will be even more
> unlikely. Meanwhile who doesn't care enough, may simply crash the VM at a
> postcopy failure and reboot.. not bother to recover.
>
> I bet rarely people will hit a network failure happens exactly at
> forwarding the TDX START token to destination host, similarly for case (2)
> on ABORT token lost. But yes, likely this is required to make the whole
> thing complete.
Thanks for the insights about users who really care about having a reliable
dedicated network.
I would like to make sure I understand correctly though. You refer to QEMU
post-copy recovery, which is basically all about restoring network
connectivity and finishing the migration. Is this right?
And to make sure we are on the same page, here is how I see things at a
high level.
1. Post-copy recovery is about handling the state split between source and
destination. I think this aspect is going to be similar, if not the same,
for traditional and CoCo VM migration. Same problem, same recovery
strategies, I'd guess.
2. The abort token stuff I discussed is about pre-copy. It is an artifact of
the switchover: when src is paused and dst is allowed to start, an abort
token is needed to reverse this process. And whether post-copy is used or
not is orthogonal to the abort token stuff.
Did I miss something?
>
> > Specifically about TDX - the pause seamcall will return an error if TDIs
> > are not unassigned.
> >
> > While this is something that is not implemented in Linux yet, I believe the
> > model will be that there is some uAPI to unassign TDIs, and it is not
> > related to migration. QEMU would just need to exercise this uAPI at the
> > right time.
>
> OK, this sounds working, but then it means the migration will be visible to
> the guest. I wonder whether there's any attempt to make it more
> transparent, but we can also leave this question for later.
Correct. With a disclaimer that I am not a TDX Connect expert, I had the
impression that this is more of a compromise solution. The VMM is an
untrusted entity in the CoCo model, so the trust between the TD and the
PCIe/CXL TDX Connect device is built by the TD itself. It is the TD
establishing the cryptographic trust with a specific physical PCIe/CXL
device. Directly, not via VMM. This is part of the industry standard SPDM
protocol, which stands for Security Protocol and Data Model.
As I understand it, with the current SPDM protocol, this trust cannot be
transparently moved from one host to another: the destination host has a
different physical device, with different cryptographic keys, and the TD
would need to re-establish trust explicitly. VMM cannot do it on behalf of
the TD in a transparent manner. I would only speculate that this means
preserving the device state is a hard problem.
Maybe future TDX Connect and SPDM revisions will solve this problem, but for
now, this is the compromise solution we have.
But again, take it with a grain of salt, it is more of my intuition than
based on concrete knowledge.
> > The reason I am asking is that my assumption was that it is not important.
> > But if it is, I will come back to the TDX module architects with a
> > request to revise the design to support independent dirty page tracking. Of
> > course they may have some reasons for not doing it, but I would try at
> > least.
>
> Thanks, I'll talk to our team and revisit this after I collect answers.
Many thanks!
> > Now, I am diverging, but just in case: in the PUCK call where we presented
> > TDX migration, Sean made an immediate observation that WP-based dirty
> > tracking is not categorically worse than PML-based dirty ring tracking, it
> > depends on the workload.
>
> Hmm, I was expecting PML is still superior in most cases. For "depending
> on workloads", is that perhaps when (1) huge pages are used, and (2) the
> workload writes only a small portion of guest memory?
Sorry, I did not communicate it correctly.Sean did not talk about huge
pages. I need to be very careful here. What I think was Sean's point is that
VM exits forced by the WP-based dirty tracking cause "back-pressure" as he
put it, meaning they work as a natural way to slow down vCPUs and improve
migration convergence. PML-based tracking does not cause as much
back-pressure, so, depending on workload, they may require artificial vCPU
throttling. But disclaimer, this is not a cite, this is my interpretation.
> > Then he learned that TDX module's dirty scanning does not use PML, and
> > was understandably surprised. Sean was concerned about dirty scanning
> > performance.
>
> I'm definitely surprised too that PML isn't used. Could I ask if there's
> any simple reason not to use it for TDX? Per my understanding, PML works
> with all kinds of loads, and I was expecting PML to be efficient and most
> ideal.
Another point where I need to be careful to not miscommunicate. The honest
answer is that I do not know for sure why PML specifically was not chosen.
Please take my comments below with a grain of salt - I am a software person
who tries to understand the design decisions made in the TDX module, but I
am not a TDX module architect.
Current Intel processors do not support PML for DMA - a TDI's DMA writes
would go untracked by PML. But they do set the Secure EPT Dirty bit, so
scanning works for catching both CPU and DMA writes.
My speculation is that once non-blocking scanning had to be built to cover
the TDX Connect case, it made sense to use it as the single mechanism for TD
migration in general.
AFAIU, the TDX guest migration implementation benchmarking results are
satisfactory with the scanning approach, but I do not have hard numbers to
share. I also feel that PML could offer better performance, at least for
memory-intensive workloads. But this is intuition only.
> Another approach is if TDX can take over the bitmap buffer from the
> relevant kvm memslots, update directly there alongside setting D bits in
> EPT PTEs; after all IIUC we assumed dirty info not part of confidential
> materials. But that sounds more complex than PML if it's already working
> for years.
Well, then the dirty bitmap specifics would become the ABI - the hard
contract between the Intel platform and the OS.
But this is effectively what TDX module dirty scanning does already today:
on input you give it an array of up to 512 GPAs, on the output it marks
which entries are migration candidates and also for what reason.
Keep in mind that dirty pages are the majority of migration candidates, but
not all of them. Sometimes a migration candidate can be a page that was
already exported, but then was, for example, converted from private to
shared, or unaccepted by the TD (gone, in other words). In this case the TDX
module flags it as a migration candidate too. The memory export seamcall
treats it differently too - instead of exporting encrypted page data, it
exports a small record indicating that the page has changed its status
(gone).
IOW, in the TDX migration case, it is not just dirty pages. In an abstract
way, it is useful to think of it as the TDX module tracking both page data
and metadata changes.
We (me, Kishen, Tony) call them "dirty pages" for simplicity, but TDX specs
use the term "migration candidates". But again, most of them are dirty
pages.
> The current scan approach sounds like unpredictable in terms of downtime,
> in that even if with a scanner I don't see how TDX can guaratee the
> downtime for the last dirty sync from QEMU, which will completely be part
> of the blackout downtime.
Could you help me understand exactly what you mean by "predictable" here?
Let me walk through how I see it. Please correct me if I am wrong.
For a traditional VM: at some point QEMU decides that pre-copy has
converged. But the source VM keeps running until it is actually paused, and
it can dirty more pages. QEMU has no way to know in advance how many more
dirty pages there will be by the time the source VM is actually paused -
could be a few, could be a lot. So the time it takes to find and copy them
during the downtime is not fully predictable.
The same logic applies to a TDX guest using dirty scanning. Suppose the
final dirty scan is slower than a theoretical TDX PML-based approach would
have been. The prescan optimization I described in the previous e-mail
should help with this in an average case, but let's assume it does not
help for some special case - when the TD touches most of its memory, so
all EPT sub-trees end up touched. This should be rare, but let's assume it
happens.
In this case, whether that scan slowdown will actually matter for the
overall downtime also depends on the network. If the network is very fast,
the scan itself can become the dominant part of the downtime. If the
network is the slower part, the scan slowdown may barely matter.
So my understanding is that downtime is never fully predictable, for
either type of VM.
Intuitively, the TDX case does feel "less predictable". The open
question for me is whether the degree of unpredictability is large enough
to bother users. My attitude is to focus on getting something simple done
first, learn from real-world behavior, and improve it later if needed,
including exploring PML. The best is the enemy of the good sort of
attitude.
> Whenever switchover decision made, QEMU stops the VM, do the last time
> sync, move anything left. So that last one will matter a lot if we care
> about downtime.
Could you please help me understand: in your experience, how much do users
care about the overall migration time?
My current assumption is that users mostly care about downtime, and care
little about the overall migration time. I guess nobody wants migration to
take days, but if it takes, say, 10 minutes, I assume no one would put in a
lot of effort to improve it to 9 minutes.
Thanks, Artem.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
2026-09-22 8:09 ` Artem Bityutskiy
2026-09-22 9:42 ` Tony Lindgren
2026-09-22 21:18 ` Peter Xu
@ 2026-09-23 15:28 ` Serge Hallyn (AMD)
2 siblings, 0 replies; 78+ messages in thread
From: Serge Hallyn (AMD) @ 2026-09-23 15:28 UTC (permalink / raw)
To: Artem Bityutskiy
Cc: Peter Xu, Tony Lindgren, Paolo Bonzini, Sean Christopherson,
Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky,
Marc Zyngier, Oliver Upton, Steven Price, Anup Patel,
Samuel Ortiz, Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
On Tue, Sep 22, 2026 at 11:09:42AM +0300, Artem Bityutskiy wrote:
> Hi Peter,
...
> > Said that, we'll need to be careful then in case of migration fallbacks at
> > the final stage. Nowadays, I believe QEMU can still fallback to source side
> > at a very, very late stage after all things applied. If I'm not mistaken,
> > the final handshake is done at migration_incoming_state_destroy() ->
> > migrate_send_rp_shut() telling source to be gone.
>
> Right. In traditional pre-copy VM migration model, the fallback is possible
> at any point before the destination VM starts running and modifying its
> state.
>
> The same is true for TDX, but with more complexity and limitations.
>
> We have 2 points:
> 1. Before the source has exported the start token - fallback is similar to
> traditional VM migration - just abort the migration on source and
> continue running the source TD, and just destroy the destination TD.
> 2. After the source has exported the start token - fallback is still
> possible, but more complex - it requires the destination to first
> generate the abort token, which should be delivered to the source and
> consumed there. There are seamcalls for both generating and consuming the
> abort token. The token basically makes sure the destination cannot run,
> and the source can run again - same "only one of the TDs can run at a
> time" security rule that we discussed earlier.
>
> Now, the limitation here is that if the abort is because the network
> connectivity is lost, the abort token cannot be delivered. Then I'd guess
> migration should stop, without destroying the source and the destination,
> and the token should be delivered later manually, QEMU could even provide an
> infrastructure / commands for that.
>
> Would be very helpful to learn the abort protocol the other CoCo vendors
> offer - is it similar to TDX or not?
Yup, that sounds very similar to the expected SEV-SNP late abort semantics,
-serge
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
2026-09-20 23:56 ` Kishen Maloor
@ 2026-09-23 21:36 ` Peter Xu
2026-09-24 4:27 ` Kishen Maloor
0 siblings, 1 reply; 78+ messages in thread
From: Peter Xu @ 2026-09-23 21:36 UTC (permalink / raw)
To: Kishen Maloor
Cc: Artem Bityutskiy, Tony Lindgren, Paolo Bonzini,
Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta,
Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price,
Anup Patel, Samuel Ortiz, Jakub Růžička,
Jörg Rödel, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On Sun, Sep 20, 2026 at 04:56:20PM -0700, Kishen Maloor wrote:
> Hi Peter,
Hi, Kishen,
>
> Thank you for your comments. Just adding a few other details to
> complement Artem's response.
>
> On 9/17/26 2:27 PM, Peter Xu wrote:
> > On Fri, Sep 04, 2026 at 09:24:25PM +0300, Artem Bityutskiy wrote:
> >> On Mon, 2026-08-31 at 10:13 +0300, Tony Lindgren wrote:
> > ...
> >> - Should the same uAPIs also support traditional VMs? But the only use-case
> >> I imagine here is "for testing purposes".
> >
> > This is an interesting idea, I think this could be useful. Especially, I
> > wonder if you already have it done and PoC branches you can share, so that
> > I can play with it.
>
> I had once attempted this and hit a breaking case while migrating regular VMs.
> I didn't dig further since the primary use of this UAPI set is with CoCo VMs.
>
> Here's what I gathered: KVM-mediated transfer needs KVM to resolve a GFN to the
> HVA where the data lands, and on x86 QEMU mutates that mapping at runtime in a
> way that is opaque to KVM. For the PAM window QEMU overlays a separate region on
> top of pc.ram, and KVM sees only the flattened view, the winning memslot, with
> no way to address what that slot shadows. On the source, firmware reprograms PAM
> and QEMU drops the overlay. On the destination the firmware never runs, so it still
> has its boot-time layout and the overlay is still in place. So importing, say, GFN 0xc0
> lands on the overlay's backing store, which is a read-only memslot, and the GFN->HVA
> translation fails. I suppose even if it were writable, the data would land there
> rather than in the pc.ram underneath where it belongs. Without mediation, QEMU is
> able to write to the HVA directly, so the problem doesn't arise.
>
> Generally, I think KVM mediation only works if KVM's view of memory is authoritative.
> TDX skirts this issue entirely because QEMU doesn't create overlays for TDX VMs.
I see this one slightly differently, and I have a major question on the
choice of GPA for KVM's memory access API.
If reusing GET_DIRTY_LOG, it means slot_id works for CoCo like before
because that's the old interface there.
But the new KVM_EXPORT_MEM (vice versa) used GPA arrays. Could I ask why
the change?
IIUC this is the fundamental reason why you hit that PAM issue: at a
specific GPA, QEMU/KVM can map different things, hence GPA is not yet an
unified identifier for a physical page that guest uses. However, (slot_id,
slot_offset) will be.
If the migration memory access API will be using that instead of GPA, I
think problems will be gone because then the PAM regions (likely, pc.ram
underneath) will be migrated in form of KVM memslots, destination apply
that to the PAM's (slot_id, slot_offset), then when switchover, that PAM
memory region will be enabled by QEMU and it will be visible to destination
QEMU when VM starts on destination.
I'm actually not sure if such (e.g. PAM) will ever be supported at all in
CoCo context, but it seems to me using (slot_id, slot_offset) (for each
entry; it can also be an array) is more flexible than GPA arrays in that it
allows GPA mapping to change on the fly if needed even during migration.
I'm not sure if that's explicitly forbidden for CoCo, though.
>
> > ...
> > Could you elaborate this ITERATION operation? Is that something the
> > userapp must do after full scan of a round of guest memory?
>
> Essentially, yes. TDX migration architecture delimits such pre-copy rounds as
> "migration epochs" and emits an epoch token at each round boundary that needs to be
> consumed on the destination. It allows the TDX module to verify that everything from
> the prior round has been received at the destination and that two versions of the
> same GPA aren't sent in the same epoch. It is userspace that decides where a round
> ends, but any deviation from this model and migration would fail.
I'm a bit surprised that TDX will also monitor how many times the same GPA
is updated per iteration. What if below happens:
...
ITERATION sync n
GET_DIRTY_LOG, see page P dirty
migrate page P
GET_DIRTY_LOG, see page P dirty again
migrate page P again <--------------------- [a]
ITERATION sync n+1
...
Would above crash on destination TDX at step [a] applying P 2nd time? Or
does it mean GET_DIRTY_LOG can only be invoked once per iteration?
>
> >
> >>> | (repeat until convergence) |
> >>> CMD(STOP_AND_COPY/PAUSE) |
> >>> CMD(STOP_AND_COPY/TD_STATE) --- VM state ------> CMD(STOP_AND_COPY/TD_STATE)
> >>> KVM_EXPORT_VCPU --- vCPU state ----> KVM_IMPORT_VCPU
> >>> KVM_EXPORT_MEMORY -- final memory ---> KVM_IMPORT_MEMORY
> >
> > When read/write encrypted memories, two questions:
> >
> > - Is there an upper bound of the buffer size per-page?
>
> Yes. For TDX a 4KB guest page produces exactly 4KB of encrypted payload plus a small
> fixed amount of ancillary data. The bound can be pre-computed and userspace
> can size its buffers accordingly. The UAPI itself doesn't impose one. The vendor
> implementation decides how pages and ancillary data are laid out.
Get it now. I still think if that varies across impl, it might be good to
somehow notify the userapp on choosing the buffer size.
But nah, it's not a big deal - when there's retry, userapp can always probe
with any size then double it until it fits..
Thanks,
--
Peter Xu
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests
2026-09-18 11:35 ` Peter Xu
2026-09-21 4:20 ` Tony Lindgren
@ 2026-09-24 1:50 ` Wei Wang
2026-09-24 4:51 ` Tony Lindgren
1 sibling, 1 reply; 78+ messages in thread
From: Wei Wang @ 2026-09-24 1:50 UTC (permalink / raw)
To: Peter Xu, Tony Lindgren
Cc: Paolo Bonzini, Sean Christopherson, Artem Bityutskiy,
Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky,
Marc Zyngier, Oliver Upton, Steven Price, Anup Patel,
Samuel Ortiz, Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
On 9/18/26 7:35 PM, Peter Xu wrote:
> On Mon, Aug 31, 2026 at 10:13:01AM +0300, Tony Lindgren wrote:
>> +:Capability: KVM_CAP_LIVE_MIGRATION
>
> IMHO this is slightly misleading, some "CONFIDENTIAL_" or other prefix
> would be nice.
>
An earlier version of this series used KVM_CAP_CGM (CGM stands for
Confidential Guest Migration). I think it's cleaner to use a short
CGM prefix for the related uAPIs, so they're easy to identify as a group.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
2026-09-23 21:36 ` Peter Xu
@ 2026-09-24 4:27 ` Kishen Maloor
2026-09-25 14:18 ` Peter Xu
0 siblings, 1 reply; 78+ messages in thread
From: Kishen Maloor @ 2026-09-24 4:27 UTC (permalink / raw)
To: Peter Xu
Cc: Artem Bityutskiy, Tony Lindgren, Paolo Bonzini,
Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta,
Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price,
Anup Patel, Samuel Ortiz, Jakub Růžička,
Jörg Rödel, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
Hi Peter,
Thanks! It really helps to go over details.
On 9/23/26 2:36 PM, Peter Xu wrote:
> ...
>
> I see this one slightly differently, and I have a major question on the
> choice of GPA for KVM's memory access API.
>
> If reusing GET_DIRTY_LOG, it means slot_id works for CoCo like before
> because that's the old interface there.
It works because we map backwards from GPAs in TDX's dirty scan results to
a KVM memslot on the source, and a bit in its dirty bitmap, in keeping with
the GET_DIRTY_LOG contract.
> But the new KVM_EXPORT_MEM (vice versa) used GPA arrays. Could I ask why
> the change?
The migration ABI (at least in TDX, possibly others) is GPA shaped. Also,
everything downstream like SEPT entries key off GPAs. The EXPORT and IMPORT
ABIs take GPAs alongside the encrypted blob, so both the source and the
destination need it at the call site that implements the UAPI. Passing GFNs
in the UAPI is merely a convenience in that respect. If we passed in
slot+offset instead, it would just be an indirection that needs to resolve
back to the GPA anyway to issue the vendor migration call. But this is not
the blocking issue for the PAM case, as I'll explain below.
> IIUC this is the fundamental reason why you hit that PAM issue: at a
> specific GPA, QEMU/KVM can map different things, hence GPA is not yet an
> unified identifier for a physical page that guest uses. However, (slot_id,
> slot_offset) will be.
>
> If the migration memory access API will be using that instead of GPA, I
> think problems will be gone because then the PAM regions (likely, pc.ram
> underneath) will be migrated in form of KVM memslots, destination apply
> that to the PAM's (slot_id, slot_offset), then when switchover, that PAM
> memory region will be enabled by QEMU and it will be visible to destination
> QEMU when VM starts on destination.
I think using slot+offset won't make the problem vanish, because the PAM window
in the underlying pc.ram region is not addressable by KVM on the destination at
that time. KVM's memslots come from QEMU's FlatView at any moment and QEMU
reconfigures them on the fly. In the destination's boot-time setup, it seems
like pc.ram's memslot is split around that window, and what covers the
range instead is a separate, read-only memslot for the overlay, which is an
allocation with its own backing, not the pc.ram bytes underneath.
And since only the flattened view is ever registered, there's no slot+offset
that refers to those bytes either. Only when the source's PAM configuration
lands at the destination, which happens after memory migration concludes, do
those pc.ram bytes become visible to KVM.
In the TDX case, QEMU does not create that overlay at the start, so the
destination has a memslot corresponding to the pc.ram region that stays put.
KVM is able to do a GFN->PFN translation to hand to TDX's IMPORT ABI.
So, it appears that a UAPI change will not address this corner case.
> I'm actually not sure if such (e.g. PAM) will ever be supported at all in
> CoCo context, but it seems to me using (slot_id, slot_offset) (for each
> entry; it can also be an array) is more flexible than GPA arrays in that it
> allows GPA mapping to change on the fly if needed even during migration.
>
> I'm not sure if that's explicitly forbidden for CoCo, though.
A TD private page lives at a fixed GPA in the SEPT, and I believe the
EXPORT/IMPORT bundle binds the GPA into the page's integrity check, so a page
has to be imported at the GPA it was exported from.
>>> Could you elaborate this ITERATION operation? Is that something the
>>> userapp must do after full scan of a round of guest memory?
>>
>> Essentially, yes. TDX migration architecture delimits such pre-copy rounds as
>> "migration epochs" and emits an epoch token at each round boundary that needs to be
>> consumed on the destination. It allows the TDX module to verify that everything from
>> the prior round has been received at the destination and that two versions of the
>> same GPA aren't sent in the same epoch. It is userspace that decides where a round
>> ends, but any deviation from this model and migration would fail.
>
> I'm a bit surprised that TDX will also monitor how many times the same GPA
> is updated per iteration. What if below happens:
Yes, and maybe that is because it won't know the ordering if the same
GPA were caught at different times in the same epoch and fed to different
migration streams (which one is newer?)
On regular VMs, I believe QEMU doesn't migrate the same GPA twice in one round.
As I understand it, a migrated GPA is revisited only in the next round.
I suppose TDX just makes that behavior architectural.
> ...
> ITERATION sync n
> GET_DIRTY_LOG, see page P dirty
> migrate page P
> GET_DIRTY_LOG, see page P dirty again
> migrate page P again <--------------------- [a]
QEMU shouldn't transfer P again in the same round, right?
> ITERATION sync n+1
> ...
>
> Would above crash on destination TDX at step [a] applying P 2nd time? Or
If P were migrated again, then I believe it would fail to import on the
destination, and also abort the import session, so migration would fail.
TDX records the current epoch against a GPA on every import, and I think
that's how it catches a 2nd import attempt.
> does it mean GET_DIRTY_LOG can only be invoked once per iteration?
We don't touch QEMU's stock dirty logging/transfer flows so however it
handles it currently stays.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests
2026-09-24 1:50 ` Wei Wang
@ 2026-09-24 4:51 ` Tony Lindgren
0 siblings, 0 replies; 78+ messages in thread
From: Tony Lindgren @ 2026-09-24 4:51 UTC (permalink / raw)
To: Wei Wang
Cc: Peter Xu, Paolo Bonzini, Sean Christopherson, Artem Bityutskiy,
Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky,
Marc Zyngier, Oliver Upton, Steven Price, Anup Patel,
Samuel Ortiz, Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
Hi Wei,
On Thu, Sep 24, 2026 at 09:50:45AM +0800, Wei Wang wrote:
> On 9/18/26 7:35 PM, Peter Xu wrote:
> > On Mon, Aug 31, 2026 at 10:13:01AM +0300, Tony Lindgren wrote:
> > > +:Capability: KVM_CAP_LIVE_MIGRATION
> >
> > IMHO this is slightly misleading, some "CONFIDENTIAL_" or other prefix
> > would be nice.
> >
> An earlier version of this series used KVM_CAP_CGM (CGM stands for
> Confidential Guest Migration). I think it's cleaner to use a short
> CGM prefix for the related uAPIs, so they're easy to identify as a group.
Using an abbreviation like CGM or MIG or HW has a problem where it's hard
to decipher for anybody not familiar with the live migration. So my vote
is currently on using Peter's suggestion for CONFIDENTIAL prefix.
Regards,
Tony
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-23 6:50 ` Tony Lindgren
@ 2026-09-24 5:34 ` Kishen Maloor
2026-09-24 7:15 ` Tony Lindgren
0 siblings, 1 reply; 78+ messages in thread
From: Kishen Maloor @ 2026-09-24 5:34 UTC (permalink / raw)
To: Tony Lindgren
Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Artem Bityutskiy,
Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky,
Marc Zyngier, Oliver Upton, Steven Price, Anup Patel,
Samuel Ortiz, Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg,
Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm
On 9/22/26 11:50 PM, Tony Lindgren wrote:
> On Tue, Sep 22, 2026 at 05:37:33PM -0700, Kishen Maloor wrote:
>> On 9/21/26 11:27 PM, Tony Lindgren wrote:
>>> Oh right thanks. I think this is really the maximum transfer buffer size
>>> Peter asked, not just a hint to userspace :)
>>
>> It is a strict upper bound. Maybe it's semantics, but I called
>> it a "hint" because allocating for that entire size could be optional.
>> If a userspace driver for say TDX wants to send smaller batches (say 128) then
>> it can refer to the spec, do the math, and allocate 131 pages and the kernel
>> should permit that; it's not wrong. If userspace ever allocates less room than
>> a call requires, it would fail. A different userspace driver could simply
>> allocate that max size and be done; no need to refer to the spec or do the math.
>> It is for this second case where I thought returning the upper bound would be
>> useful. Hence the earlier suggestion.
>
> OK
>
>>>>>>>> +struct kvm_transfer_buffer {
>>>>>>>> + __u64 address;
>>>>>>>> + __u32 size;
>>>>>>>> + __u32 reserved;
>>>>>>>> +};
>>>>>>>
>>>>>>> Should this struct include a 'capacity' field (u32) that is set on each command?
>>>>>>> It would be the number of bytes writable at address.
>>>>>>> size would be the input length on entry (0 if the command passes none), and the
>>>>>>> number of bytes produced on return (0 if none).
>>>>>>
>>>>>> Hmm so the transfer command return value can return how many bytes were
>>>>>> written of the input. But yeah we don't know how many bytes were written
>>>>>> back to the transfer buffer as result of the transfer command.
>>>>
>>>> The transfer command return value could return how many bytes were written into the buffer.
>>>> But in an input-output call, the kernel handler wouldn't know how many bytes it could write,
>>>> or for that matter even how many pages to pin up front in case it needs to return an output
>>>> because 'size' couldn't simultaneously convey the input length and buffer capacity. That was
>>>> the gap that I thought a read-only 'capacity' field could bridge. Of course, this
>>>> assumes that the output is written in-place.
>>>
>>> Hmm yeah this inplace capacity vs transferred issue remains still. So I
>>> agree we need to specify the capacity in struct kvm_transfer_buffer like
>>> you suggested.
>>
>> To be clear, I think the in/out split for the buffers along with the convention
>> I laid out closes that gap I saw without needing a 'capacity' field.
>> Because out/size could now unambiguously convey capacity on entry and output
>> length on return.
>
> But for an inplace buffer use with some input data smaller than the output
> data, would it work? To me it seems you need both buffer size and data
> size for that.
Yes, with two kvm_transfer_buffers in the transfer struct, it would work, whether
the call uses a single userspace buffer or two.
>
>>> To me size is already the size of the buffer though. So instead of changing
>>> size to capacity, how about something like datasize or len for the input
>>> and output transfer length?
>>
>> But I understand that (and please correct me if I'm wrong):
>> a) You'd still prefer to not have 'size' serve that double duty.
>> b) 'size' in your mental model already means buffer capacity.
>
> Heh yes correct for the above.
>
>> In that case, we could add a 'datasize' field to convey the length
>> of valid data in the buffer, like this:
>>
>> struct kvm_transfer_buffer {
>> __u64 address;
>> __u32 size;
>> __u32 datasize;
>> __u64 reserved;
>> };
>
> Maybe bufsize and datasize? Then the difference would be obvious while
> reading the code.
Sure.
>
>> The convention then becomes:
>> - A non-zero 'datasize' on 'in' at call entry conveys that there is input.
>> - A non-zero 'datasize' on 'out' at call exit conveys that there is output.
>> - out/datasize on call entry is ignored.
>> - in/size and out/size are seeded with the buffer capacity.
>>
>>>
>>>>>> How about if we add the bytes returned to the transfer struct? Then the
>>>>>> kvm_transfer_buffer can stay as just a buffer.
>>>>>
>>>>> Actually, for the possible cases with input+output, we could reserve space
>>>>> in the transfer struct for another struct kvm_transfer_buffer for the
>>>>> results?
>>>>
>>>> Yes, say an 'in' and 'out' kvm_transfer_buffer inside struct kvm_migrate_cmd should
>>>> close this out and shouldn't require a 'capacity' field. Maybe we then establish this
>>>> convention:
>>>> - A non-zero 'size' on 'in' at call entry would signal that there is input.
>>>> - A non-zero 'size' on 'out' at call entry would convey the buffer capacity.
>>>> - A non-zero 'size' on 'out' at call exit would convey that there is output.
>>>> - A zeroed 'size' on 'out' at call exit would convey that there is no output.
>>>
>>> Looks doable to me but with the inplace issue as above.. Sounds like we just
>>> need to reserve space for a case with a separate output buffer though.
>>
>> To be clear, this is what I thought we were talking about :)
>
> Heh yeah we're talking two things with the inplace use vs two buffers :)
>
>> To add a 2nd kvm_transfer_buffer to kvm_migrate_cmd, like this:
>>
>> struct kvm_migrate_cmd {
>> __u16 command;
>> __u16 flags;
>> __u32 reserved;
>> struct kvm_transfer_buffer in;
>> struct kvm_transfer_buffer out;
>> };
>>
>> If there is agreement on this model, then yeah, we'd want to define
>> these fields now, since the struct can't grow later without a new ioctl
>> number.
>
> Based on what we've discussed, my preference is the following:
>
> Keep the current buf naming. For the EXPORT/IMPORT type functions the use
> should be obvious from the transfer type.
>
> Reserve enough space for a separate output buffer or results buffer or
> whatever it might get called if such a use case ever pops up.
Reserved space can be named later without changing sizeof, so either way
works. I'd mildly prefer to declare the second kvm_transfer_buffer just
because declaring both would settle the second buffer's semantics now. Not
something I'd push hard on though if you prefer reserving.
> Add the datasize to struct kvm_transfer_buffer like you suggested and
> rename size to bufsize.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-23 6:04 ` Tony Lindgren
@ 2026-09-24 5:53 ` Kishen Maloor
2026-09-24 6:59 ` Tony Lindgren
0 siblings, 1 reply; 78+ messages in thread
From: Kishen Maloor @ 2026-09-24 5:53 UTC (permalink / raw)
To: Tony Lindgren
Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini,
Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On 9/22/26 11:04 PM, Tony Lindgren wrote:
> On Tue, Sep 22, 2026 at 05:38:43PM -0700, Kishen Maloor wrote:
>> On 9/21/26 10:25 PM, Tony Lindgren wrote:
>>> On Mon, Sep 21, 2026 at 08:57:45PM -0700, Kishen Maloor wrote:
>>>> On 9/20/26 11:52 PM, Tony Lindgren wrote:
>>>>> Then for KVM tracking the role, I don't think we need it with the two
>>>>> above. The role tracking can always be added if really needed. Any other
>>>>> opinions on this one?
>>>>
>>>> I do think it's useful for generic KVM to track this role (1 byte).
>>>> - It lets _TRANSFER_ style calls dispatch directly to import or export
>>>> callbacks based on the role.
>>>> - Even if we don't adopt the _TRANSFER_ style, it enables generic KVM to
>>>> reject mismatched calls, e.g., KVM_EXPORT_MEMORY on a destination.
>>>
>>> Having KVM do generic checks on the calls is a good idea. There might be
>>> a simpler way of handling it though. Rather than having KVM track the
>>> migration state, how about we add a function to check for the migration
>>> session state from the vendor code?
>>>
>>> So something like this for the states you suggested earlier:
>>>
>>> enum kvm_lmstate {
>>> KVM_LM_NONE,
>>> KVM_LM_SOURCE,
>>> KVM_LM_DESTINATION,
>>> };
>>>
>>> With something like this to get the state from the vendor code:
>>>
>>> enum kmv_lmstate kvm_arch_get_lmstate(struct kvm *);
>>>
>>> For x86 it would end up calling kvm_x86_call(get_lmstate)(kvm) and for
>>> the TDX specific case tdx_get_lmstate().
>>
>> How is this simpler? It trades one byte in a KVM struct for a new generic
>> enum, a new kvm_arch_get_lmstate(), a new kvm_x86_ops entry, and a vendor
>> implementation per vendor, plus a cross-layer call on every command just
>> to learn the role.
>> Directly checking a stored byte (0=unset/1=src/2=dst) seems simplest, no?
>
> It would avoid dragging KVM into the "track the migration state" business
> at least for now.
>
> I guess the question in general is: What does KVM need to do with the
> migration role beyond generic checks on the migration related calls?
I haven't thought of any other uses for it. But if we do want those generic
checks, I'd still lean toward the stored byte.
>
>>> It would allow KVM to do the generic checks for the migration related
>>> calls you're describing. And having KVM start tracking the state can be
>>> still added later on too if it is needed.
>>
>> I think we'd need to pick one way or the other before the UAPI settles
>> if we want generic KVM to reject mismatched calls. OTOH if we want to
>> defer this generic KVM validation, then yeah, it could be settled later.
>>
>>>
>>>>> The transfer direction is there with the EXPORT/IMPORT naming. Maybe
>>>>> just let's keep that naming for easier readability rather than try to
>>>>> switch to TRANSFER style naming. No transfer direction flag needed.
>>>>
>>>> I can't say I have a clear preference between EXPORT/IMPORT vs _TRANSFER_.
>>>> If we keep the EXPORT/IMPORT naming, then consistency would arguably call for
>>>> splitting MIGRATE_CMD too, which makes it 6 vs 3 (or 5 vs 3 against the RFC
>>>> as posted):
>>>>
>>>> EXPORT/IMPORT style _TRANSFER_ style
>>>> KVM_EXPORT_CMD KVM_MIGRATE_CMD
>>>> KVM_IMPORT_CMD
>>>> KVM_EXPORT_MEMORY KVM_TRANSFER_MEMORY
>>>> KVM_IMPORT_MEMORY
>>>> KVM_EXPORT_VCPU KVM_TRANSFER_VCPU
>>>> KVM_IMPORT_VCPU
>>>
>>> Agreed we should split the MIGRATE_CMD too. Probably the number of ioctls
>>> is not and issue compared to following the KVM style and better
>>> readability. So my vote is now on EXPORT/IMPORT style naming.
>>>
>>>> I guess the main benefit of the _TRANSFER_ style is reduced duplication.
>>>> Each EXPORT/IMPORT pair takes the same struct and differs only by the role
>>>> that the session already knows. Merging them gives one entry point per call
>>>> type and lets userspace drive both ends from the same call site which could
>>>> be considered a win. It doesn't reduce kernel code though as the top-level
>>>> handler still branches internally on the role.
>>>
>>> Yup not much of a win for the TRANSFER style naming.
>>
>> Yeah, my goal was just to enumerate alternatives for consideration.
>>
>> One other benefit of KVM_EXPORT_CMD/KVM_IMPORT_CMD is that the per-session
>> role is implicitly conveyed - in other words, a successful KVM_EXPORT_CMD/SETUP
>> (for e.g.) would indicate that this is a 'source'. So, we wouldn't require a 'role'
>> field in struct kvm_migrate_cmd to explicitly assert one during SETUP.
>
> Yes good point with the KVM_EXPORT/IMPORT_CMD, that sounds good to me.
>
>>> Trying to summarize again after we sorted out the direction flag issue in
>>> the transfer:
>>>
>>> role per migration session (cannot change during the migration)
>>> direction per command, EXPORT/IMPORT
>>> hardware state set and tracked by vendor specific code
>>
>> The 'role' and 'direction' as you define it above are essentially saying the
>> same thing - a source only invokes the vendor's EXPORT call, a destination only
>> invokes the vendor's IMPORT call, and the role doesn't change over the session.
>> If a destination needs to send data to be consumed by the source, then that
>> still invokes the vendor's EXPORT call.
>
> Yup.
>
>>> And checking again against the dmaengine analogy:
>>>
>>> Compared to dmaengine, the migration role is modeled similar to the dma
>>> channel configuration.
>>>
>>> The migration direction with EXPORT/IMPORT is modeled similar to
>>> dmaengine_prep_slave_sg().
>> I'll leave the dmaengine comparison to you, I don't know that API well enough
>> to map it properly :)
>
> Heh just a sanity check for trying to relate this to something existing.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-24 5:53 ` Kishen Maloor
@ 2026-09-24 6:59 ` Tony Lindgren
0 siblings, 0 replies; 78+ messages in thread
From: Tony Lindgren @ 2026-09-24 6:59 UTC (permalink / raw)
To: Kishen Maloor
Cc: Artem Bityutskiy, Jörg Rödel, Paolo Bonzini,
Sean Christopherson, Peter Xu, Fabiano Rosas, Jon Grimm,
Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On Wed, Sep 23, 2026 at 10:53:28PM -0700, Kishen Maloor wrote:
> On 9/22/26 11:04 PM, Tony Lindgren wrote:
> > On Tue, Sep 22, 2026 at 05:38:43PM -0700, Kishen Maloor wrote:
> >> On 9/21/26 10:25 PM, Tony Lindgren wrote:
> >>> On Mon, Sep 21, 2026 at 08:57:45PM -0700, Kishen Maloor wrote:
> >>>> On 9/20/26 11:52 PM, Tony Lindgren wrote:
> >>>>> Then for KVM tracking the role, I don't think we need it with the two
> >>>>> above. The role tracking can always be added if really needed. Any other
> >>>>> opinions on this one?
> >>>>
> >>>> I do think it's useful for generic KVM to track this role (1 byte).
> >>>> - It lets _TRANSFER_ style calls dispatch directly to import or export
> >>>> callbacks based on the role.
> >>>> - Even if we don't adopt the _TRANSFER_ style, it enables generic KVM to
> >>>> reject mismatched calls, e.g., KVM_EXPORT_MEMORY on a destination.
> >>>
> >>> Having KVM do generic checks on the calls is a good idea. There might be
> >>> a simpler way of handling it though. Rather than having KVM track the
> >>> migration state, how about we add a function to check for the migration
> >>> session state from the vendor code?
> >>>
> >>> So something like this for the states you suggested earlier:
> >>>
> >>> enum kvm_lmstate {
> >>> KVM_LM_NONE,
> >>> KVM_LM_SOURCE,
> >>> KVM_LM_DESTINATION,
> >>> };
> >>>
> >>> With something like this to get the state from the vendor code:
> >>>
> >>> enum kmv_lmstate kvm_arch_get_lmstate(struct kvm *);
> >>>
> >>> For x86 it would end up calling kvm_x86_call(get_lmstate)(kvm) and for
> >>> the TDX specific case tdx_get_lmstate().
> >>
> >> How is this simpler? It trades one byte in a KVM struct for a new generic
> >> enum, a new kvm_arch_get_lmstate(), a new kvm_x86_ops entry, and a vendor
> >> implementation per vendor, plus a cross-layer call on every command just
> >> to learn the role.
> >> Directly checking a stored byte (0=unset/1=src/2=dst) seems simplest, no?
> >
> > It would avoid dragging KVM into the "track the migration state" business
> > at least for now.
> >
> > I guess the question in general is: What does KVM need to do with the
> > migration role beyond generic checks on the migration related calls?
>
> I haven't thought of any other uses for it. But if we do want those generic
> checks, I'd still lean toward the stored byte.
We can easily add the KVM role later on if real KVM generic need for
carrying the migration role comes up. We can just have the vendor code
do the checks for now.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
2026-09-24 5:34 ` Kishen Maloor
@ 2026-09-24 7:15 ` Tony Lindgren
0 siblings, 0 replies; 78+ messages in thread
From: Tony Lindgren @ 2026-09-24 7:15 UTC (permalink / raw)
To: Kishen Maloor
Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Artem Bityutskiy,
Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky,
Marc Zyngier, Oliver Upton, Steven Price, Anup Patel,
Samuel Ortiz, Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg,
Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm
On Wed, Sep 23, 2026 at 10:34:14PM -0700, Kishen Maloor wrote:
> On 9/22/26 11:50 PM, Tony Lindgren wrote:
> > Based on what we've discussed, my preference is the following:
> >
> > Keep the current buf naming. For the EXPORT/IMPORT type functions the use
> > should be obvious from the transfer type.
> >
> > Reserve enough space for a separate output buffer or results buffer or
> > whatever it might get called if such a use case ever pops up.
>
> Reserved space can be named later without changing sizeof, so either way
> works. I'd mildly prefer to declare the second kvm_transfer_buffer just
> because declaring both would settle the second buffer's semantics now. Not
> something I'd push hard on though if you prefer reserving.
Let's just use buf and reserved space then. There is no known usecase
needing separate in and out buffers for EXPORT or IMPORT. And the separate
in and out buffers would have to be needed the same time rather than first
in and then out..
> > Add the datasize to struct kvm_transfer_buffer like you suggested and
> > rename size to bufsize.
And to recap, with the bufsize and datasize in the buffer, the needs we
discussed for separate in and out buffers went away.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
2026-09-23 12:05 ` Artem Bityutskiy
@ 2026-09-24 21:19 ` Peter Xu
2026-09-28 14:15 ` Artem Bityutskiy
0 siblings, 1 reply; 78+ messages in thread
From: Peter Xu @ 2026-09-24 21:19 UTC (permalink / raw)
To: Artem Bityutskiy
Cc: Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas,
Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
On Wed, Sep 23, 2026 at 03:05:38PM +0300, Artem Bityutskiy wrote:
[...]
> I would like to make sure I understand correctly though. You refer to QEMU
> post-copy recovery, which is basically all about restoring network
> connectivity and finishing the migration. Is this right?
Correct.
>
> And to make sure we are on the same page, here is how I see things at a
> high level.
>
> 1. Post-copy recovery is about handling the state split between source and
> destination. I think this aspect is going to be similar, if not the same,
> for traditional and CoCo VM migration. Same problem, same recovery
> strategies, I'd guess.
>
> 2. The abort token stuff I discussed is about pre-copy. It is an artifact of
> the switchover: when src is paused and dst is allowed to start, an abort
> token is needed to reverse this process. And whether post-copy is used or
> not is orthogonal to the abort token stuff.
>
> Did I miss something?
I believe we're on the same page.
I mentioned the recovery feature because both of them (even if ABORT is
part of precopy rather than postcopy) describe such an use case where
network interruption can cause some form of split brain of the VM, causing
neither side be able to continue.
We used to not have such case with precopy, but then this ABORT / START
message can make it happen similarly like postcopy. Said that, the window
is much smaller than postcopy.
>
> >
> > > Specifically about TDX - the pause seamcall will return an error if TDIs
> > > are not unassigned.
> > >
> > > While this is something that is not implemented in Linux yet, I believe the
> > > model will be that there is some uAPI to unassign TDIs, and it is not
> > > related to migration. QEMU would just need to exercise this uAPI at the
> > > right time.
> >
> > OK, this sounds working, but then it means the migration will be visible to
> > the guest. I wonder whether there's any attempt to make it more
> > transparent, but we can also leave this question for later.
>
> Correct. With a disclaimer that I am not a TDX Connect expert, I had the
> impression that this is more of a compromise solution. The VMM is an
> untrusted entity in the CoCo model, so the trust between the TD and the
> PCIe/CXL TDX Connect device is built by the TD itself. It is the TD
> establishing the cryptographic trust with a specific physical PCIe/CXL
> device. Directly, not via VMM. This is part of the industry standard SPDM
> protocol, which stands for Security Protocol and Data Model.
>
> As I understand it, with the current SPDM protocol, this trust cannot be
> transparently moved from one host to another: the destination host has a
> different physical device, with different cryptographic keys, and the TD
> would need to re-establish trust explicitly. VMM cannot do it on behalf of
> the TD in a transparent manner. I would only speculate that this means
> preserving the device state is a hard problem.
>
> Maybe future TDX Connect and SPDM revisions will solve this problem, but for
> now, this is the compromise solution we have.
>
> But again, take it with a grain of salt, it is more of my intuition than
> based on concrete knowledge.
AFAIU, preserving device states were a hard problem even on non-CoCo
before, but then I guess people thought VFIO performs so good, after that
people managed to work the problem out.. and now more people start to rely
on VFIO precopy migrations working in the clusters, non-CoCo.
I had a gut feeling it will happen too for CoCo some day, that unplug
approach was exactly what happens before VFIO migration is implemented...
But yes, let's leave this for later, thanks for sharing.
> > > The reason I am asking is that my assumption was that it is not important.
> > > But if it is, I will come back to the TDX module architects with a
> > > request to revise the design to support independent dirty page tracking. Of
> > > course they may have some reasons for not doing it, but I would try at
> > > least.
> >
> > Thanks, I'll talk to our team and revisit this after I collect answers.
>
> Many thanks!
I got some feedback on this, I'll try to provide a summary.
So, first of all, calc_dirty_rate isn't seem to be widely used across our
customers.
However, we do have customer case using calc_dirty_rate to evaluate
migrations of a VM fleet for cases like from one data centre to another.
I think it makes sense because the normal "try to migrate and fallback
otherwise" idea applies well to one VM, but perhaps not that good on a
fleet.
When a fleet is involved, we don't want to migrate 400 VMs then found
there're 30 critical VMs too busy and can't migrate, then due to whatever
reason (inter-VM communication / service locality ?) one is forced to
migrate that 400 VMs backwards.
IOW, it seems helpful to provide high-level evalutions of migration
decisions over a full cluster, concurrently and efficiently.
> > Hmm, I was expecting PML is still superior in most cases. For "depending
> > on workloads", is that perhaps when (1) huge pages are used, and (2) the
> > workload writes only a small portion of guest memory?
>
> Sorry, I did not communicate it correctly.Sean did not talk about huge
> pages. I need to be very careful here. What I think was Sean's point is that
> VM exits forced by the WP-based dirty tracking cause "back-pressure" as he
> put it, meaning they work as a natural way to slow down vCPUs and improve
> migration convergence. PML-based tracking does not cause as much
> back-pressure, so, depending on workload, they may require artificial vCPU
> throttling. But disclaimer, this is not a cite, this is my interpretation.
No worires, thanks for sharing your thoughts. And I agree there is that
back pressure effect. Migration performance is one of the most weird
performance engineering topics I'm aware of for sure; sometimes, the better
a work done, the less likely it converges..
>
> > > Then he learned that TDX module's dirty scanning does not use PML, and
> > > was understandably surprised. Sean was concerned about dirty scanning
> > > performance.
> >
> > I'm definitely surprised too that PML isn't used. Could I ask if there's
> > any simple reason not to use it for TDX? Per my understanding, PML works
> > with all kinds of loads, and I was expecting PML to be efficient and most
> > ideal.
>
> Another point where I need to be careful to not miscommunicate. The honest
> answer is that I do not know for sure why PML specifically was not chosen.
> Please take my comments below with a grain of salt - I am a software person
> who tries to understand the design decisions made in the TDX module, but I
> am not a TDX module architect.
>
> Current Intel processors do not support PML for DMA - a TDI's DMA writes
> would go untracked by PML. But they do set the Secure EPT Dirty bit, so
> scanning works for catching both CPU and DMA writes.
>
> My speculation is that once non-blocking scanning had to be built to cover
> the TDX Connect case, it made sense to use it as the single mechanism for TD
> migration in general.
>
> AFAIU, the TDX guest migration implementation benchmarking results are
> satisfactory with the scanning approach, but I do not have hard numbers to
> share. I also feel that PML could offer better performance, at least for
> memory-intensive workloads. But this is intuition only.
My gut feeling is DMA shouldn't be a blocker for PML: AFAIU we don't track
DMA from KVM side. Assigned device should have its own dirty tracking for
DMAs, either via device's own tracking facilities, or the IOMMU on the
host. Feel free to refer to vfio_listener_log_sync() in QEMU. In all
cases, it'll be great you could share the reason if you have more solid
clues.
>
> > Another approach is if TDX can take over the bitmap buffer from the
> > relevant kvm memslots, update directly there alongside setting D bits in
> > EPT PTEs; after all IIUC we assumed dirty info not part of confidential
> > materials. But that sounds more complex than PML if it's already working
> > for years.
>
> Well, then the dirty bitmap specifics would become the ABI - the hard
> contract between the Intel platform and the OS.
>
> But this is effectively what TDX module dirty scanning does already today:
> on input you give it an array of up to 512 GPAs, on the output it marks
> which entries are migration candidates and also for what reason.
Are we talking about the memory export/import API or GET_DIRTY_LOG? IIUC,
GET_DIRTY_LOG always applies to a whole memslot,
struct kvm_dirty_log {
__u32 slot;
__u32 padding1;
union {
void *dirty_bitmap; /* one bit per page */
__u64 padding2;
};
};
>
> Keep in mind that dirty pages are the majority of migration candidates, but
> not all of them. Sometimes a migration candidate can be a page that was
> already exported, but then was, for example, converted from private to
> shared, or unaccepted by the TD (gone, in other words). In this case the TDX
> module flags it as a migration candidate too. The memory export seamcall
> treats it differently too - instead of exporting encrypted page data, it
> exports a small record indicating that the page has changed its status
> (gone).
Yes, it makes sense.
So can I inteprete this as GET_DIRTY_LOG works seamlessly for both private
and shared pages (or even, unaccepted pages)?
Then I assume it means MEMORY.EXPORT should also be able to read shared or
unaccepted pages too, am I right? Same to when apply with IMPORT. Another
counter example is MEMORY.EXPORT returns a flag saying "this page is
shared, go read it directly from HVA", but then QEMU reading it may race
with a concurrent shared->private conversion crashing VMM.
Looks to me MEMORY.EXPORT must support shared too, then.
>
> IOW, in the TDX migration case, it is not just dirty pages. In an abstract
> way, it is useful to think of it as the TDX module tracking both page data
> and metadata changes.
>
> We (me, Kishen, Tony) call them "dirty pages" for simplicity, but TDX specs
> use the term "migration candidates". But again, most of them are dirty
> pages.
Yes, I didn't notice it before, but now I see that marking converted or
unaccepted pages to be dirty makes sense. IIUC it's because that info
(shared, or private, or unaccepted) is part of page [meta]data that needs
to be migrated to reconstruct the whole VM on the other host.
>
> > The current scan approach sounds like unpredictable in terms of downtime,
> > in that even if with a scanner I don't see how TDX can guaratee the
> > downtime for the last dirty sync from QEMU, which will completely be part
> > of the blackout downtime.
>
> Could you help me understand exactly what you mean by "predictable" here?
> Let me walk through how I see it. Please correct me if I am wrong.
>
> For a traditional VM: at some point QEMU decides that pre-copy has
> converged. But the source VM keeps running until it is actually paused, and
> it can dirty more pages. QEMU has no way to know in advance how many more
> dirty pages there will be by the time the source VM is actually paused -
> could be a few, could be a lot. So the time it takes to find and copy them
> during the downtime is not fully predictable.
Correct.
>
> The same logic applies to a TDX guest using dirty scanning. Suppose the
> final dirty scan is slower than a theoretical TDX PML-based approach would
> have been. The prescan optimization I described in the previous e-mail
> should help with this in an average case, but let's assume it does not
> help for some special case - when the TD touches most of its memory, so
> all EPT sub-trees end up touched. This should be rare, but let's assume it
> happens.
>
> In this case, whether that scan slowdown will actually matter for the
> overall downtime also depends on the network. If the network is very fast,
> the scan itself can become the dominant part of the downtime. If the
> network is the slower part, the scan slowdown may barely matter.
>
> So my understanding is that downtime is never fully predictable, for
> either type of VM.
Right, but IMHO background scan of EPT pgtable dirty bits adds a completely
new reason to introduce downtime, and when I said "unpredictable", it is
about that part. Also, I worry in some worst case this can be pretty large.
So we have two overheads here at this stage, unpredictable:
(a) Scanning EPT pgtable, when very unlucky, can take a lot of time to
finally reports to a GET_DIRTY_LOG request,
(b) Migrating of dirty pages during blackout phase, which should be
roughly linear to how many dirty pages we just collected. (NOTE! I
think we may have way to fix this (b) or optimize it.. but this is
off-topic; let's focus on the difference of (a) and (b) first)
When with PML, IIUC (a) is predictable: we have the bitmap on hand, plus a
maximum of some (my memory is, 512?) PML entries to flush per vCPU.
That'll be flushed automatically when we do vm_stop(), likely also
concurrently, atomically updating the bitmaps. I never measured it, but it
is bounded, and sounds pretty fast.
When with scanning, (a) seems more unpredictable. That's the part I was
slightly concerned. But now after thinking a bit more, it seems fine.
Please read below.
>
> Intuitively, the TDX case does feel "less predictable". The open
> question for me is whether the degree of unpredictability is large enough
> to bother users. My attitude is to focus on getting something simple done
> first, learn from real-world behavior, and improve it later if needed,
> including exploring PML. The best is the enemy of the good sort of
> attitude.
Yes, I think it's always fine we start with whatever is most feasible.
I think it actually may not be that bad. The last sync is special at least
on how QEMU treats it, it should look like:
- GET_DIRTY_LOG, to do last math, decide to switchover, <------ [1]
- vm_stop()
- GET_DIRTY_LOG, this collects all rest dirty bits <------ [2]
- migrates the dirty pages, device states, etc.
So I expect there should be normally very small window between two
continuous GET_DIRTY_LOG across system. Only [2] will be part of downtime.
Since you explained to me on how the background rescan roughly works, by
relying on A bit in pgtable directory entries, I do feel like in this case
most of the memory regions shouldn't be accessed during small window of
[1]->[2], then the range to scan should be very much under control too. In
reality, it will likely be even smaller, [1]->vm_stop(), because after that
vCPUs are halted.
So it may not really be an issue in practise, but it still depends. In all
cases, some measurements after PoC ready would be nice on some large and
relatively busy VMs.
>
> > Whenever switchover decision made, QEMU stops the VM, do the last time
> > sync, move anything left. So that last one will matter a lot if we care
> > about downtime.
>
> Could you please help me understand: in your experience, how much do users
> care about the overall migration time?
>
> My current assumption is that users mostly care about downtime, and care
> little about the overall migration time. I guess nobody wants migration to
> take days, but if it takes, say, 10 minutes, I assume no one would put in a
> lot of effort to improve it to 9 minutes.
I agree. I think some use case may care about total migration time, but I
would say in most cases, "10min or 9min total migration time" difference is
less of a concern than downtime effects.
Thanks,
--
Peter Xu
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
2026-09-24 4:27 ` Kishen Maloor
@ 2026-09-25 14:18 ` Peter Xu
2026-09-29 1:28 ` Kishen Maloor
0 siblings, 1 reply; 78+ messages in thread
From: Peter Xu @ 2026-09-25 14:18 UTC (permalink / raw)
To: Kishen Maloor
Cc: Artem Bityutskiy, Tony Lindgren, Paolo Bonzini,
Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta,
Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price,
Anup Patel, Samuel Ortiz, Jakub Růžička,
Jörg Rödel, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On Wed, Sep 23, 2026 at 09:27:52PM -0700, Kishen Maloor wrote:
> Hi Peter,
Hi, Kishen,
[...]
> > IIUC this is the fundamental reason why you hit that PAM issue: at a
> > specific GPA, QEMU/KVM can map different things, hence GPA is not yet an
> > unified identifier for a physical page that guest uses. However, (slot_id,
> > slot_offset) will be.
> >
> > If the migration memory access API will be using that instead of GPA, I
> > think problems will be gone because then the PAM regions (likely, pc.ram
> > underneath) will be migrated in form of KVM memslots, destination apply
> > that to the PAM's (slot_id, slot_offset), then when switchover, that PAM
> > memory region will be enabled by QEMU and it will be visible to destination
> > QEMU when VM starts on destination.
>
> I think using slot+offset won't make the problem vanish, because the PAM window
> in the underlying pc.ram region is not addressable by KVM on the destination at
> that time. KVM's memslots come from QEMU's FlatView at any moment and QEMU
> reconfigures them on the fly. In the destination's boot-time setup, it seems
> like pc.ram's memslot is split around that window, and what covers the
> range instead is a separate, read-only memslot for the overlay, which is an
> allocation with its own backing, not the pc.ram bytes underneath.
> And since only the flattened view is ever registered, there's no slot+offset
> that refers to those bytes either. Only when the source's PAM configuration
> lands at the destination, which happens after memory migration concludes, do
> those pc.ram bytes become visible to KVM.
>
> In the TDX case, QEMU does not create that overlay at the start, so the
> destination has a memslot corresponding to the pc.ram region that stays put.
> KVM is able to do a GFN->PFN translation to hand to TDX's IMPORT ABI.
>
> So, it appears that a UAPI change will not address this corner case.
Yes, I think you're right, the ABI isn't the major issue. If from KVM's
perspective it's always convertable between GPA <-> slots, then it's the
same.
I believe my mindset when replying was pretty much in QEMU's perspective,
where in qemu we can have two ramblocks plugged into the same GPA range,
only one of them will be visible to KVM and guest (e.g. which one has
higher MemoryRegion priority, but it's not the only factor). What migration
module does right now is, it allows both ramblocks to be migrated with no
issue, even if one is not visible, but if it used to be touched, since that
ramblock will maintain its own dirty bitmap (GET_DIRTY_LOG on that kvm
memslot when it was visible to KVM, or maybe set within QEMU userspace
somehow).
IOW, what QEMU could do here is, after MEMORY.EXPORT, convert the GPA
address space into ramblock ranges in QEMU, migrate with that not GPA. In
case of PAM, it's part of pc.ram. Dest QEMU sees it is pc.ram, it should
"apply" those data to pc.ram ramblock only.
For non-CoCo, it's as easy as writting to some host HVA pointer.
Now, the question is, the new migration API only allows applying data in
GPA ranges. I don't think with TDX there's a way to "apply" the data..
dest QEMU just booted, this specific portion of pc.ram may not be mapped at
that GPA source fetched due to reset status of PAM registers.
That (rather than the ABI interface), might be the real thing I wanted to
point out.
Maybe it means TDX just can't work with it by definition? I think it'll be
fine, and now I wonder if it means PAM will be working for TDX only if PAM
boots too early so it was before TDX initializes, then after TDX enabled
anything like PAM will not work anymore?
So it seems TDX will "lock" the memory footprint in place when enabled,
allow accept/unaccept (or say, plug / unplug) memories, but anything like
"flipping this to that" will not work.
Then there's a very corner case question I want to double check, and I
apologize if this is stupid only due to my ignorance on TDX knowledge: can
someone migrate a VM too early so TDX is just hasn't been enabled at all?
[...]
> > I'm a bit surprised that TDX will also monitor how many times the same GPA
> > is updated per iteration. What if below happens:
>
> Yes, and maybe that is because it won't know the ordering if the same
> GPA were caught at different times in the same epoch and fed to different
> migration streams (which one is newer?)
> On regular VMs, I believe QEMU doesn't migrate the same GPA twice in one round.
> As I understand it, a migrated GPA is revisited only in the next round.
> I suppose TDX just makes that behavior architectural.
>
> > ...
> > ITERATION sync n
> > GET_DIRTY_LOG, see page P dirty
> > migrate page P
> > GET_DIRTY_LOG, see page P dirty again
> > migrate page P again <--------------------- [a]
>
> QEMU shouldn't transfer P again in the same round, right?
I believe yes with current QEMU, I can't think of anything otherwise. But
still, this is very specific impl detail. There's definitely no issue
migrating one page twice or more in non-CoCo.
I can give one example to illustrate what could happen.
In postcopy, we support preemption mode, which is simply a separate fast
path for requested / urgent pages. It's possible while background thread
transferring one page, the fast path saw a request on this same page. The
current algorithm is simple, it will wait for that in progress background
send to complete.
But logically, we could do it the other way too: send the page again on
fast path, in postcopy the page content is guaranteed to be identical and
unchnaged, it means the fast path can land this page earlier, reducing
fault latency. If so, a minimum cap we need is MEMORY.EXPORT be able to be
done twice, so the fast path can read the 2nd time. IMPORT is more
flexible, because QEMU can maintain what has been applied, so logically
background loader should be able to skip the 2nd IMPORT.
That is not a good example, at least because it's postcopy and doesn't
happen with precopy. So far, I also don't think a major risk, but I confess
I don't understand why TDX needs to add hard requirement on "only sample
one page once per iteration": even if VMM sampled a page twice, it's still
encrypted and confidential. I believe it has something to do with the
whole attestation logic. I think it'll be more flexible with less
restrictions, but no issue I see either, hence please only treat that a
verbose FYI.
>
> > ITERATION sync n+1
> > ...
> >
> > Would above crash on destination TDX at step [a] applying P 2nd time? Or
>
> If P were migrated again, then I believe it would fail to import on the
> destination, and also abort the import session, so migration would fail.
I wonder if we can just fail the 2nd IMPORT without abort the whole
process, then if userapp wants to detect it there's a way to (similar to an
-EEXIST). Not a request, more like a pure question.
> TDX records the current epoch against a GPA on every import, and I think
> that's how it catches a 2nd import attempt.
>
> > does it mean GET_DIRTY_LOG can only be invoked once per iteration?
>
> We don't touch QEMU's stock dirty logging/transfer flows so however it
> handles it currently stays.
I see, then it should be good. Just to say, qemu may sync dirty bitmaps in
the background nowadays, can refer to cpu_throttle_dirty_sync_timer_tick().
So it can be decoupled from iteration runs.
Thanks,
--
Peter Xu
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
2026-08-31 7:13 [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Tony Lindgren
` (5 preceding siblings ...)
2026-09-18 18:36 ` Ionut Mihalcea
@ 2026-09-25 16:03 ` Serge Hallyn (AMD)
2026-09-28 3:24 ` Kishen Maloor
6 siblings, 1 reply; 78+ messages in thread
From: Serge Hallyn (AMD) @ 2026-09-25 16:03 UTC (permalink / raw)
To: Tony Lindgren
Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Artem Bityutskiy,
Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky,
Marc Zyngier, Oliver Upton, Steven Price, Anup Patel,
Samuel Ortiz, Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
On Mon, Aug 31, 2026 at 10:13:00AM +0300, Tony Lindgren wrote:
> Hi all,
>
> As discussed in a recent PUCK call, Sean suggested we post what Intel is
> using for the KVM live migration API for TDX as an example to see if we
> can come up with APIs that are not vendor specific.
>
> The goal of this patch series is to start a discussion about live migration
> kernel APIs for CoCo guests.
>
> Tom, since you mentioned that AMD SEV-SNP and Intel TDX live migration
> sound similar, can you please take a look how the API might work for
> SEV-SNP?
>
> For CoCo VMs, the guest memory and vCPU states are not accessible to the
> userspace or KVM for live migration. The memory and vCPU states need to be
> extracted into encrypted blobs on the source, and decrypted on the
> destination. Before live migration, an encryption key needs to be
> negotiated between the source and destination.
>
> The layer handling the encryption for live migration is implementation
> specific. It can be the TDX module or Coconut-SVSM for example.
>
> For CoCo VMs, the KVM_MEMORY_ENCRYPT_OP ioctl() has been used with vendor
> specific sub-commands. Adding more vendor specific sub-commands is an
> option also for live migration. However, depending on how similar the KVM
> needs are, it may be possible to have a common API.
>
> For TDX, we're using a group of ioctl()s that might be possible to adapt
> also for other CoCo implementations.
>
> Artem has put together a brief description below of the example API and the
> migration flow:
Hi Tony,
is there any pubically available qemu git branch or patchset to show
how you currently are driving this?
> Example API
> ===========
>
> - KVM_CAP_LIVE_MIGRATION - if a VM supports live migration through this
> uAPI.
> - KVM_MIGRATE_CMD - the main ioctl that drives the migration phases. Each
> command takes vendor-specific flags and a buffer for the blob that travels
> between the hosts.
> - KVM_MIGRATE_SETUP - establish the migration session and transfer the
> immutable VM state.
> - KVM_MIGRATE_ITERATION - close a memory copy round.
> - KVM_MIGRATE_STOP_AND_COPY - pause the VM and transfer the remaining VM
> state.
> - KVM_MIGRATE_END - complete the migration, or abort it.
> - KVM_EXPORT_MEMORY - export memory pages on the source host.
> - KVM_IMPORT_MEMORY - import memory pages on the destination host.
> - KVM_EXPORT_VCPU - export vCPU state on the source host.
> - KVM_IMPORT_VCPU - import vCPU state on the destination host.
>
> Dirty page tracking does not add a new uAPI. Userspace keeps using
> KVM_GET_DIRTY_LOG and KVM_CLEAR_DIRTY_LOG.
>
> Migration flow
> ==============
>
> Source host Destination host
> =========== ================
>
> CMD(SETUP/SESSION) <--- setup msgs ---> CMD(SETUP/SESSION)
> | (repeated) |
> CMD(SETUP/IMMUTABLE_STATE) - immutable state -> CMD(SETUP/IMMUTABLE_STATE)
> | |
> KVM_GET_DIRTY_LOG |
> KVM_EXPORT_MEMORY --- memory data ---> KVM_IMPORT_MEMORY
> CMD(ITERATION) --- epoch token ---> CMD(ITERATION)
> | (repeat until convergence) |
> CMD(STOP_AND_COPY/PAUSE) |
> CMD(STOP_AND_COPY/TD_STATE) --- VM state ------> CMD(STOP_AND_COPY/TD_STATE)
> KVM_EXPORT_VCPU --- vCPU state ----> KVM_IMPORT_VCPU
> KVM_EXPORT_MEMORY -- final memory ---> KVM_IMPORT_MEMORY
> CMD(ITERATION/DONE) --- start token ---> CMD(ITERATION)
> | |
> CMD(END) CMD(END)
>
> For the TDX implementation, the above map to the TDX module SEAMCALLs.
>
> Regards,
>
> Tony
>
> Changes since v1 at [0] below:
>
> - Drop KVM_MIGRATE_CMD sub-command PREPARE, SETUP sub-command has been
> enough for TDX at least
>
> - Rename KVM_MIGRATE_CMD sub-command KVM_MIGRATE_TOKEN to
> KVM_MIGRATE_ITERATION
>
> - Rename KVM_MIGRATE_CMD sub-command KVM_MIGRATE_SOURCE_BLACKOUT to
> KVM_MIGRATE_STOP_AND_COPY
>
> - Add x86 ioctl handling
>
> [0] https://lore.kernel.org/kvm/20251006113524.1573116-1-tony.lindgren@linux.intel.com/
>
>
> Tony Lindgren (4):
> Documentation: KVM: Add live migration API for confidential guests
> KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD
> KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY
> KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU
>
> Documentation/virt/kvm/api.rst | 205 +++++++++++++++++++++++++++++
> arch/x86/include/asm/kvm-x86-ops.h | 6 +
> arch/x86/include/asm/kvm_host.h | 6 +
> arch/x86/kvm/x86.c | 102 ++++++++++++++
> include/uapi/linux/kvm.h | 43 ++++++
> 5 files changed, 362 insertions(+)
>
>
> base-commit: dc59e4fea9d83f03bad6bddf3fa2e52491777482
> --
> 2.43.0
>
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
2026-09-25 16:03 ` Serge Hallyn (AMD)
@ 2026-09-28 3:24 ` Kishen Maloor
0 siblings, 0 replies; 78+ messages in thread
From: Kishen Maloor @ 2026-09-28 3:24 UTC (permalink / raw)
To: Serge Hallyn (AMD), Tony Lindgren
Cc: Paolo Bonzini, Sean Christopherson, Peter Xu, Artem Bityutskiy,
Fabiano Rosas, Jon Grimm, Pankaj Gupta, Tom Lendacky,
Marc Zyngier, Oliver Upton, Steven Price, Anup Patel,
Samuel Ortiz, Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Mika Westerberg,
Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun, kvm
On 9/25/26 9:03 AM, Serge Hallyn (AMD) wrote:
> On Mon, Aug 31, 2026 at 10:13:00AM +0300, Tony Lindgren wrote:
> ...
>
> is there any pubically available qemu git branch or patchset to show
> how you currently are driving this?
>
Hi Serge,
Nothing public yet. Our QEMU code is throwaway scaffolding for exercising the
ioctls. We're working on something more suitable to post as an RFC.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
2026-09-24 21:19 ` Peter Xu
@ 2026-09-28 14:15 ` Artem Bityutskiy
2026-09-29 21:05 ` Peter Xu
0 siblings, 1 reply; 78+ messages in thread
From: Artem Bityutskiy @ 2026-09-28 14:15 UTC (permalink / raw)
To: Peter Xu
Cc: Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas,
Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
Hi Peter, thanks for reply again.
On Thu, 2026-09-24 at 17:19 -0400, Peter Xu wrote:
> > > > The reason I am asking is that my assumption was that it is not important.
> > > > But if it is, I will come back to the TDX module architects with a
> > > > request to revise the design to support independent dirty page tracking. Of
> > > > course they may have some reasons for not doing it, but I would try at
> > > > least.
> > >
> > > Thanks, I'll talk to our team and revisit this after I collect answers.
> >
> > Many thanks!
>
> I got some feedback on this, I'll try to provide a summary.
>
> So, first of all, calc_dirty_rate isn't seem to be widely used across our
> customers.
>
> However, we do have customer case using calc_dirty_rate to evaluate
> migrations of a VM fleet for cases like from one data centre to another.
>
> I think it makes sense because the normal "try to migrate and fallback
> otherwise" idea applies well to one VM, but perhaps not that good on a
> fleet.
Yes, that makes sense.
> When a fleet is involved, we don't want to migrate 400 VMs then found
> there're 30 critical VMs too busy and can't migrate, then due to whatever
> reason (inter-VM communication / service locality ?) one is forced to
> migrate that 400 VMs backwards.
>
> IOW, it seems helpful to provide high-level evalutions of migration
> decisions over a full cluster, concurrently and efficiently.
Thank you. I'll work on this internally. It will take time.
Just to give wider context: Sean and Paolo gave us feedback regarding the
entire SEPT scan approach - they believe it is too costly and won't scale,
and suggested using PML instead. For now, dirty scanning is the best we
have, but I continue to explore other options internally, and standalone
dirty tracking is one of them.
> > AFAIU, the TDX guest migration implementation benchmarking results are
> > satisfactory with the scanning approach, but I do not have hard numbers to
> > share. I also feel that PML could offer better performance, at least for
> > memory-intensive workloads. But this is intuition only.
>
> My gut feeling is DMA shouldn't be a blocker for PML: AFAIU we don't track
> DMA from KVM side. Assigned device should have its own dirty tracking for
> DMAs, either via device's own tracking facilities, or the IOMMU on the
> host. Feel free to refer to vfio_listener_log_sync() in QEMU. In all
> cases, it'll be great you could share the reason if you have more solid
> clues.
Yeah, this is also part of the internal exploration I mentioned above too.
> >
> Are we talking about the memory export/import API or GET_DIRTY_LOG? IIUC,
> GET_DIRTY_LOG always applies to a whole memslot,
>
> struct kvm_dirty_log {
> __u32 slot;
> __u32 padding1;
> union {
> void *dirty_bitmap; /* one bit per page */
> __u64 padding2;
> };
> };
I apologize, I worte something unrelated to the context. 512 GPAs at a time
is the limit for exporting the memory, not for dirty scanning.
For the dirty scanning seamcall (TDH.MEM.SCAN.RANGE), the limit is 512 * 512
GPAs at a time, which is 262,144 GPAs, or 1GiB. So if a memslot is larger
than 1GiB, multiple calls to the dirty scanning seamcall are needed.
> > Keep in mind that dirty pages are the majority of migration candidates, but
> > not all of them. Sometimes a migration candidate can be a page that was
> > already exported, but then was, for example, converted from private to
> > shared, or unaccepted by the TD (gone, in other words). In this case the TDX
> > module flags it as a migration candidate too. The memory export seamcall
> > treats it differently too - instead of exporting encrypted page data, it
> > exports a small record indicating that the page has changed its status
> > (gone).
>
> Yes, it makes sense.
>
> So can I inteprete this as GET_DIRTY_LOG works seamlessly for both private
> and shared pages (or even, unaccepted pages)?
>
> Then I assume it means MEMORY.EXPORT should also be able to read shared or
> unaccepted pages too, am I right? Same to when apply with IMPORT. Another
> counter example is MEMORY.EXPORT returns a flag saying "this page is
> shared, go read it directly from HVA", but then QEMU reading it may race
> with a concurrent shared->private conversion crashing VMM.
>
> Looks to me MEMORY.EXPORT must support shared too, then.
Hmm... First of all, it does sound like a possible approach. But it is not
the approach we took in our PoC today. I hope Kishen will chime in to
correct me.
Here is how I saw this, but I may be missing something (my excuse is that I
am still new to the team and still learning).
1. QEMU has a bitmap of shared pages in RAMBlockAttributes, so it can
distinguish shared pages.
2. In general, QEMU does not distinguish private vs unaccepted pages, so
unaccepted pages are treated as private pages.
Dirty tracking:
- QEMU uses the same KVM_DIRTY_LOG mechanism for tracking shared, private,
and unaccepted pages.
- For shared GFNs, KVM uses the normal VM dirty tracking mechanism. For
private and unaccepted GFNs, KVM goes to the TDX-specific code.
- But the final bitmap that QEMU sees covers all page types.
- KVM calls TDH.MEM.SCAN.RANGE on both private and unaccepted GFNs.
- For private GFNs, the TDX module reports it as a migration candidate if
its data changed or its status changed (e.g., converted to shared or
unaccepted).
- For unaccepted GFNs, the TDX module reports it as a migration candidate in
the first round (so it appears as dirty in KVM_DIRTY_LOG reply). Then it
reports it as clean, unless its status changes - it becomes accepted.
Page export:
- QEMU migrates shared pages the old way - it does not try to use the
proposed CoCo migration uAPI for that.
- For private and unaccepted pages, QEMU uses the CoCo migration uAPI.
- Our export uAPI PoC implementation does not try to check GFN type - it
just feeds them all to the TDX module TDH.EXPORT.MEM seamcall. Here
is what TDX module does depending on the page type:
- Shared pages: just skip, no errors.
- Private pages: export the data in encrypted form.
- Unaccepted pages: export a small record telling that the page is
unaccepted. This record should be delivered to the destination and
imported there, just like private pages.
But clearly this is part of the uAPI contract that must be discussed and
made explicit. What I describe above is obviously our PoC implementation,
plus TDX module behavior details.
> > The same logic applies to a TDX guest using dirty scanning. Suppose the
> > final dirty scan is slower than a theoretical TDX PML-based approach would
> > have been. The prescan optimization I described in the previous e-mail
> > should help with this in an average case, but let's assume it does not
> > help for some special case - when the TD touches most of its memory, so
> > all EPT sub-trees end up touched. This should be rare, but let's assume it
> > happens.
> >
> > In this case, whether that scan slowdown will actually matter for the
> > overall downtime also depends on the network. If the network is very fast,
> > the scan itself can become the dominant part of the downtime. If the
> > network is the slower part, the scan slowdown may barely matter.
> >
> > So my understanding is that downtime is never fully predictable, for
> > either type of VM.
>
> Right, but IMHO background scan of EPT pgtable dirty bits adds a completely
> new reason to introduce downtime, and when I said "unpredictable", it is
> about that part. Also, I worry in some worst case this can be pretty large.
Just to clarify on the "background" part. Yes, it is "background" relative
to the TD - some CPU is running it, in parallel with vCPUs running on other
CPUs.
But there is no background activity in the TDX module itself. All the
seamcalls run synchronously on the CPU that invokes them.
This means, for example, that to speed up memory export, one can run the
TDH.EXPORT.MEM seamcall for different GPAs in parallel on different CPUs.
Same for dirty scanning - if one could run the TDH.MEM.SCAN.RANGE seamcall
for different GPAs in parallel on different CPUs, that would make scanning
much faster.
Also a bit separately, one point to keep in mind is that in CoCo the page
export and import are heavy, compute-intensive crypto operations. So when we
talk about slower scanning, we need to keep in mind that it is not that slow
in relation to the export crypto. But I understand that this is not an
apples-to-apples comparison:
- Scanning is potentially about a large SEPT in a VM with terabytes of
memory.
- Exporting is only about the pages found to be dirty.
So this is only to remind that in the CoCo case, memory export also has a
high price tag, compared to a traditional VM.
> So we have two overheads here at this stage, unpredictable:
>
> (a) Scanning EPT pgtable, when very unlucky, can take a lot of time to
> finally reports to a GET_DIRTY_LOG request,
>
> (b) Migrating of dirty pages during blackout phase, which should be
> roughly linear to how many dirty pages we just collected. (NOTE! I
> think we may have way to fix this (b) or optimize it.. but this is
> off-topic; let's focus on the difference of (a) and (b) first)
>
> When with PML, IIUC (a) is predictable: we have the bitmap on hand, plus a
> maximum of some (my memory is, 512?) PML entries to flush per vCPU.
> That'll be flushed automatically when we do vm_stop(), likely also
> concurrently, atomically updating the bitmaps. I never measured it, but it
> is bounded, and sounds pretty fast.
Yes, I agree.
Now my secret desire is that TDX module can eventually plug PML under the
hood, consider it a "hardware accelerator" without changing the ABI. But I
do not know whether keeping the ABI unchanged is possible, or how soon it
could happen. This is something I am working on internally with Intel TDX
module team.
> When with scanning, (a) seems more unpredictable. That's the part I was
> slightly concerned. But now after thinking a bit more, it seems fine.
> Please read below.
Sure, thanks.
> > Intuitively, the TDX case does feel "less predictable". The open
> > question for me is whether the degree of unpredictability is large enough
> > to bother users. My attitude is to focus on getting something simple done
> > first, learn from real-world behavior, and improve it later if needed,
> > including exploring PML. The best is the enemy of the good sort of
> > attitude.
>
> Yes, I think it's always fine we start with whatever is most feasible.
>
> I think it actually may not be that bad. The last sync is special at least
> on how QEMU treats it, it should look like:
>
> - GET_DIRTY_LOG, to do last math, decide to switchover, <------ [1]
> - vm_stop()
> - GET_DIRTY_LOG, this collects all rest dirty bits <------ [2]
> - migrates the dirty pages, device states, etc.
>
> So I expect there should be normally very small window between two
> continuous GET_DIRTY_LOG across system. Only [2] will be part of downtime.
>
> Since you explained to me on how the background rescan roughly works, by
> relying on A bit in pgtable directory entries, I do feel like in this case
> most of the memory regions shouldn't be accessed during small window of
> [1]->[2], then the range to scan should be very much under control too. In
> reality, it will likely be even smaller, [1]->vm_stop(), because after that
> vCPUs are halted.
Yes, I agree. Just to flag the word "background" again, and to make sure we
are aligned - this "rescan" happens between [1] and [2]. The idea is that
[2] will be very fast after the "rescan". But it does increase the time
between [1] and [2], and the TD has time to dirty more pages.
> So it may not really be an issue in practise, but it still depends. In all
> cases, some measurements after PoC ready would be nice on some large and
> relatively busy VMs.
Yes, I agree.
>
Thanks, Artem.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
2026-09-25 14:18 ` Peter Xu
@ 2026-09-29 1:28 ` Kishen Maloor
2026-09-30 20:42 ` Peter Xu
0 siblings, 1 reply; 78+ messages in thread
From: Kishen Maloor @ 2026-09-29 1:28 UTC (permalink / raw)
To: Peter Xu
Cc: Artem Bityutskiy, Tony Lindgren, Paolo Bonzini,
Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta,
Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price,
Anup Patel, Samuel Ortiz, Jakub Růžička,
Jörg Rödel, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
Hi Peter,
Great comments!
On 9/25/26 7:18 AM, Peter Xu wrote:
> On Wed, Sep 23, 2026 at 09:27:52PM -0700, Kishen Maloor wrote:
> ...
> Yes, I think you're right, the ABI isn't the major issue. If from KVM's
> perspective it's always convertable between GPA <-> slots, then it's the
> same.
Correct. It's just that when the bytes aren't in any memslot there's no way
to reach pc.ram's HVA on the destination in KVM's view to write to it.
> I believe my mindset when replying was pretty much in QEMU's perspective,
> where in qemu we can have two ramblocks plugged into the same GPA range,
> only one of them will be visible to KVM and guest (e.g. which one has
> higher MemoryRegion priority, but it's not the only factor). What migration
I understand, and it matches what I found.
> module does right now is, it allows both ramblocks to be migrated with no
> issue, even if one is not visible, but if it used to be touched, since that
> ramblock will maintain its own dirty bitmap (GET_DIRTY_LOG on that kvm
> memslot when it was visible to KVM, or maybe set within QEMU userspace
> somehow).
Correct, and regular VM migration bypasses any notion of GPAs/memslots
entirely. It migrates RAMBlocks by block+offset. QEMU sets every RAMBlock's
migration bitmap to all ones before the first pass, so everything is
transferred in round 1 regardless of dirty state.
> IOW, what QEMU could do here is, after MEMORY.EXPORT, convert the GPA
> address space into ramblock ranges in QEMU, migrate with that not GPA. In
> case of PAM, it's part of pc.ram. Dest QEMU sees it is pc.ram, it should
> "apply" those data to pc.ram ramblock only.
>
> For non-CoCo, it's as easy as writting to some host HVA pointer.
I understand that this is how it works for regular VMs as the
destination is able to memcpy into the HVA obtained using the block+offset it
receives over the stream.
Also, when migration is kicked off after the source has fully booted, the
PAM window already maps to the source's pc.ram area, so the export transfers
the right bytes. The gap is purely on the destination because its layout is
still in the boot-time configuration until device state lands.
> Now, the question is, the new migration API only allows applying data in
> GPA ranges. I don't think with TDX there's a way to "apply" the data..
> dest QEMU just booted, this specific portion of pc.ram may not be mapped at
> that GPA source fetched due to reset status of PAM registers.
Correct. In TDX, the page is placed by the import SEAMCALL at the GPA it was
exported from. So QEMU has to turn the block+offset back into a GPA, and that
GPA has to be covered by a memslot for KVM to resolve it, which it is unable to
do at that point.
> That (rather than the ABI interface), might be the real thing I wanted to
> point out.
Understood. I did also wonder whether other VMMs organize their guest memory
hierarchy in similar ways.
> Maybe it means TDX just can't work with it by definition? I think it'll be
> fine, and now I wonder if it means PAM will be working for TDX only if PAM
> boots too early so it was before TDX initializes, then after TDX enabled
> anything like PAM will not work anymore?
Upstream QEMU gates the separate pc.rom allocation on !is_tdx_vm(), so there's
no second allocation to flip to. I don't see any other TDX gating around this,
so I presume that PAM operates as usual otherwise.
It doesn't appear to be a question of PAM coming up before TDX initializes.
> So it seems TDX will "lock" the memory footprint in place when enabled,
> allow accept/unaccept (or say, plug / unplug) memories, but anything like
> "flipping this to that" will not work.
Yes, I think so. The TDX module holds the GPA->page binding, so any change has
to pass through it. Adding and removing pages plumb through the module.
But a sort of content preserving remap of GPAs in the way QEMU does for PAM
is not possible AFAIU.
> Then there's a very corner case question I want to double check, and I
> apologize if this is stupid only due to my ignorance on TDX knowledge: can
> someone migrate a VM too early so TDX is just hasn't been enabled at all?
Not a stupid question at all. I've actually tried this. TDX VMs appear to
migrate fine beyond a certain point mid-boot of the source. Kicking off
migration any earlier causes the destination not to resume, even though the
migration itself reports success. I haven't yet confirmed what that point is
to explain it.
For now we're trying to settle on sound fundamentals for the UAPI and flow,
but this is definitely an area to dig further into.
>>> I'm a bit surprised that TDX will also monitor how many times the same GPA
>>> is updated per iteration. What if below happens:
>>
>> Yes, and maybe that is because it won't know the ordering if the same
>> GPA were caught at different times in the same epoch and fed to different
>> migration streams (which one is newer?)
>> On regular VMs, I believe QEMU doesn't migrate the same GPA twice in one round.
>> As I understand it, a migrated GPA is revisited only in the next round.
>> I suppose TDX just makes that behavior architectural.
>>
>>> ...
>>> ITERATION sync n
>>> GET_DIRTY_LOG, see page P dirty
>>> migrate page P
>>> GET_DIRTY_LOG, see page P dirty again
>>> migrate page P again <--------------------- [a]
>>
>> QEMU shouldn't transfer P again in the same round, right?
>
> I believe yes with current QEMU, I can't think of anything otherwise. But
> still, this is very specific impl detail. There's definitely no issue
> migrating one page twice or more in non-CoCo.
Thanks for confirming this. But do you think it would be a problem during
pre-copy with multifd on regular VM migration if we didn't have this
"transfer once per round" logic? If the same page got fed to different
multifd queues, couldn't the older copy land after the newer one at the
destination? [1]
With a single channel it shouldn't matter, but it could with multiple
channels. I think since the bitmap sweep is shared with multifd, QEMU ends up
transferring a page only once per round either way.
At least this was the rationale I conceived in trying to explain the TDX
policy of not allowing multiple imports of the same page in one epoch.
Though the underlying concern may not be TDX-specific.
> I can give one example to illustrate what could happen.
>
> In postcopy, we support preemption mode, which is simply a separate fast
> path for requested / urgent pages. It's possible while background thread
> transferring one page, the fast path saw a request on this same page. The
> current algorithm is simple, it will wait for that in progress background
> send to complete.
>
> But logically, we could do it the other way too: send the page again on
> fast path, in postcopy the page content is guaranteed to be identical and
> unchnaged, it means the fast path can land this page earlier, reducing
> fault latency. If so, a minimum cap we need is MEMORY.EXPORT be able to be
> done twice, so the fast path can read the 2nd time. IMPORT is more
We shall check about this. TDH.EXPORT.MEM (the export-side SEAMCALL on the
source) may already permit a 2nd export of the same page during post-copy.
If it doesn't, there's a good case for asking for it, since it would give
both the background and preemption threads a shot at fulfilling a request
ASAP. The import side might still reject it, so your suggestion below looks
like the right place to handle that.
> flexible, because QEMU can maintain what has been applied, so logically
> background loader should be able to skip the 2nd IMPORT.
This would make sense to do I suppose, because accepting a 2nd IMPORT would
clobber a page that the destination previously received and has itself
modified.
> That is not a good example, at least because it's postcopy and doesn't
> happen with precopy. So far, I also don't think a major risk, but I confess
> I don't understand why TDX needs to add hard requirement on "only sample
> one page once per iteration": even if VMM sampled a page twice, it's still
> encrypted and confidential. I believe it has something to do with the
> whole attestation logic. I think it'll be more flexible with less
> restrictions, but no issue I see either, hence please only treat that a
> verbose FYI.
I think the transfer once per iteration might have everything to do with
reason [1] I mentioned above. I've generally noticed that the TDX migration
architecture reflects established practices of VMMs. That being said, we
should look for instances where a restriction deviates and poses a problem.
>
>>
>>> ITERATION sync n+1
>>> ...
>>>
>>> Would above crash on destination TDX at step [a] applying P 2nd time? Or
>>
>> If P were migrated again, then I believe it would fail to import on the
>> destination, and also abort the import session, so migration would fail.
>
> I wonder if we can just fail the 2nd IMPORT without abort the whole
> process, then if userapp wants to detect it there's a way to (similar to an
> -EEXIST). Not a request, more like a pure question.
Understood. I think it's TDX's way of assuring correctness in case a VMM
transfers a page more than once in a round. QEMU doesn't. In theory it
could relax that when only one migration stream is registered as the
ordering would be unambiguous there, so a second import is provably newer
and could just be accepted.
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
2026-09-28 14:15 ` Artem Bityutskiy
@ 2026-09-29 21:05 ` Peter Xu
0 siblings, 0 replies; 78+ messages in thread
From: Peter Xu @ 2026-09-29 21:05 UTC (permalink / raw)
To: Artem Bityutskiy
Cc: Tony Lindgren, Paolo Bonzini, Sean Christopherson, Fabiano Rosas,
Jon Grimm, Pankaj Gupta, Tom Lendacky, Marc Zyngier, Oliver Upton,
Steven Price, Anup Patel, Samuel Ortiz,
Jakub Růžička, Jörg Rödel,
Vishal Annapurve, Elena Reshetova, Kai Huang, Kishen Maloor,
Mika Westerberg, Peter Fang, Rick Edgecombe, Xiaoyao Li, Xu Yilun,
kvm
On Mon, Sep 28, 2026 at 05:15:00PM +0300, Artem Bityutskiy wrote:
> Hi Peter, thanks for reply again.
>
> On Thu, 2026-09-24 at 17:19 -0400, Peter Xu wrote:
> > > > > The reason I am asking is that my assumption was that it is not important.
> > > > > But if it is, I will come back to the TDX module architects with a
> > > > > request to revise the design to support independent dirty page tracking. Of
> > > > > course they may have some reasons for not doing it, but I would try at
> > > > > least.
> > > >
> > > > Thanks, I'll talk to our team and revisit this after I collect answers.
> > >
> > > Many thanks!
> >
> > I got some feedback on this, I'll try to provide a summary.
> >
> > So, first of all, calc_dirty_rate isn't seem to be widely used across our
> > customers.
> >
> > However, we do have customer case using calc_dirty_rate to evaluate
> > migrations of a VM fleet for cases like from one data centre to another.
> >
> > I think it makes sense because the normal "try to migrate and fallback
> > otherwise" idea applies well to one VM, but perhaps not that good on a
> > fleet.
>
> Yes, that makes sense.
>
> > When a fleet is involved, we don't want to migrate 400 VMs then found
> > there're 30 critical VMs too busy and can't migrate, then due to whatever
> > reason (inter-VM communication / service locality ?) one is forced to
> > migrate that 400 VMs backwards.
> >
> > IOW, it seems helpful to provide high-level evalutions of migration
> > decisions over a full cluster, concurrently and efficiently.
>
> Thank you. I'll work on this internally. It will take time.
Yes, thanks. Nothing urgent, good to know it's on the map.
>
> Just to give wider context: Sean and Paolo gave us feedback regarding the
> entire SEPT scan approach - they believe it is too costly and won't scale,
> and suggested using PML instead. For now, dirty scanning is the best we
> have, but I continue to explore other options internally, and standalone
> dirty tracking is one of them.
Ok.
>
> > > AFAIU, the TDX guest migration implementation benchmarking results are
> > > satisfactory with the scanning approach, but I do not have hard numbers to
> > > share. I also feel that PML could offer better performance, at least for
> > > memory-intensive workloads. But this is intuition only.
> >
> > My gut feeling is DMA shouldn't be a blocker for PML: AFAIU we don't track
> > DMA from KVM side. Assigned device should have its own dirty tracking for
> > DMAs, either via device's own tracking facilities, or the IOMMU on the
> > host. Feel free to refer to vfio_listener_log_sync() in QEMU. In all
> > cases, it'll be great you could share the reason if you have more solid
> > clues.
>
> Yeah, this is also part of the internal exploration I mentioned above too.
>
> > >
> > Are we talking about the memory export/import API or GET_DIRTY_LOG? IIUC,
> > GET_DIRTY_LOG always applies to a whole memslot,
> >
> > struct kvm_dirty_log {
> > __u32 slot;
> > __u32 padding1;
> > union {
> > void *dirty_bitmap; /* one bit per page */
> > __u64 padding2;
> > };
> > };
>
> I apologize, I worte something unrelated to the context. 512 GPAs at a time
> is the limit for exporting the memory, not for dirty scanning.
>
> For the dirty scanning seamcall (TDH.MEM.SCAN.RANGE), the limit is 512 * 512
> GPAs at a time, which is 262,144 GPAs, or 1GiB. So if a memslot is larger
> than 1GiB, multiple calls to the dirty scanning seamcall are needed.
>
> > > Keep in mind that dirty pages are the majority of migration candidates, but
> > > not all of them. Sometimes a migration candidate can be a page that was
> > > already exported, but then was, for example, converted from private to
> > > shared, or unaccepted by the TD (gone, in other words). In this case the TDX
> > > module flags it as a migration candidate too. The memory export seamcall
> > > treats it differently too - instead of exporting encrypted page data, it
> > > exports a small record indicating that the page has changed its status
> > > (gone).
> >
> > Yes, it makes sense.
> >
> > So can I inteprete this as GET_DIRTY_LOG works seamlessly for both private
> > and shared pages (or even, unaccepted pages)?
> >
> > Then I assume it means MEMORY.EXPORT should also be able to read shared or
> > unaccepted pages too, am I right? Same to when apply with IMPORT. Another
> > counter example is MEMORY.EXPORT returns a flag saying "this page is
> > shared, go read it directly from HVA", but then QEMU reading it may race
> > with a concurrent shared->private conversion crashing VMM.
> >
> > Looks to me MEMORY.EXPORT must support shared too, then.
>
> Hmm... First of all, it does sound like a possible approach. But it is not
> the approach we took in our PoC today. I hope Kishen will chime in to
> correct me.
>
> Here is how I saw this, but I may be missing something (my excuse is that I
> am still new to the team and still learning).
>
> 1. QEMU has a bitmap of shared pages in RAMBlockAttributes, so it can
> distinguish shared pages.
> 2. In general, QEMU does not distinguish private vs unaccepted pages, so
> unaccepted pages are treated as private pages.
Yes, the latter seems uncontroversial.
The 1st one is true, and it just reminded me if the conversion is
synchronous and one step requires the hypercall to QEMU, then indeed
background conversion can be avoided by some form of userspace locking.
Perhaps, a rwlock suites, each vCPU takes it for write whenever page
conversion requested from the guest (private <-> shared; nothing about
"accepted" that matters). Then the migration threads, one or multiple,
take the read lock, lookup the bit, do MEM.EXPORT, unlock.
Then it seems fine in general, except that I donno if things can still go
wrong when there are multiple versions of "if this page is private or
shared". Say, minimum of three?
(a) QEMU maintains the bitmap in RAMBlockAttributes, each bit represents
if the page is shared or private
(b) KVM should maintain one, looks to me, kvm->mem_attr_array
(c) Hardware / Firmware may maintain its own, in case of TDX, is that one
bit on the SEPT pgtable?
They don't change together, AFAIU, they change in order, I believe
(c)->(a)->(b) if my above understanding is correct.
Then, what if they report different things?
Say, during migration the guest wants to convert a page from shared to
private. (c) can be already done saying one page "private" now for TDX,
(a) tries to mark it "private" too, but now assuming page being accessed
(read lock held), it may be trying to take a write lock and sleep, which
means (b) will be "shared" so far.
So what happens is, QEMU thinks this page "shared" because the conversion
hasn't take place waiting for the write lock, however at least TDX may
think it already "private" instead.
Then QEMU logically can access HVA of that page, with (a)=shared,
(b)=shared, (c)=private.
Would it cause trouble?
>
> Dirty tracking:
>
> - QEMU uses the same KVM_DIRTY_LOG mechanism for tracking shared, private,
> and unaccepted pages.
> - For shared GFNs, KVM uses the normal VM dirty tracking mechanism. For
> private and unaccepted GFNs, KVM goes to the TDX-specific code.
> - But the final bitmap that QEMU sees covers all page types.
> - KVM calls TDH.MEM.SCAN.RANGE on both private and unaccepted GFNs.
> - For private GFNs, the TDX module reports it as a migration candidate if
> its data changed or its status changed (e.g., converted to shared or
> unaccepted).
> - For unaccepted GFNs, the TDX module reports it as a migration candidate in
> the first round (so it appears as dirty in KVM_DIRTY_LOG reply). Then it
> reports it as clean, unless its status changes - it becomes accepted.
>
> Page export:
>
> - QEMU migrates shared pages the old way - it does not try to use the
> proposed CoCo migration uAPI for that.
> - For private and unaccepted pages, QEMU uses the CoCo migration uAPI.
> - Our export uAPI PoC implementation does not try to check GFN type - it
> just feeds them all to the TDX module TDH.EXPORT.MEM seamcall. Here
> is what TDX module does depending on the page type:
> - Shared pages: just skip, no errors.
> - Private pages: export the data in encrypted form.
> - Unaccepted pages: export a small record telling that the page is
> unaccepted. This record should be delivered to the destination and
> imported there, just like private pages.
>
> But clearly this is part of the uAPI contract that must be discussed and
> made explicit. What I describe above is obviously our PoC implementation,
> plus TDX module behavior details.
>
> > > The same logic applies to a TDX guest using dirty scanning. Suppose the
> > > final dirty scan is slower than a theoretical TDX PML-based approach would
> > > have been. The prescan optimization I described in the previous e-mail
> > > should help with this in an average case, but let's assume it does not
> > > help for some special case - when the TD touches most of its memory, so
> > > all EPT sub-trees end up touched. This should be rare, but let's assume it
> > > happens.
> > >
> > > In this case, whether that scan slowdown will actually matter for the
> > > overall downtime also depends on the network. If the network is very fast,
> > > the scan itself can become the dominant part of the downtime. If the
> > > network is the slower part, the scan slowdown may barely matter.
> > >
> > > So my understanding is that downtime is never fully predictable, for
> > > either type of VM.
> >
> > Right, but IMHO background scan of EPT pgtable dirty bits adds a completely
> > new reason to introduce downtime, and when I said "unpredictable", it is
> > about that part. Also, I worry in some worst case this can be pretty large.
>
> Just to clarify on the "background" part. Yes, it is "background" relative
> to the TD - some CPU is running it, in parallel with vCPUs running on other
> CPUs.
>
> But there is no background activity in the TDX module itself. All the
> seamcalls run synchronously on the CPU that invokes them.
>
> This means, for example, that to speed up memory export, one can run the
> TDH.EXPORT.MEM seamcall for different GPAs in parallel on different CPUs.
>
> Same for dirty scanning - if one could run the TDH.MEM.SCAN.RANGE seamcall
> for different GPAs in parallel on different CPUs, that would make scanning
> much faster.
>
> Also a bit separately, one point to keep in mind is that in CoCo the page
> export and import are heavy, compute-intensive crypto operations. So when we
> talk about slower scanning, we need to keep in mind that it is not that slow
> in relation to the export crypto. But I understand that this is not an
> apples-to-apples comparison:
>
> - Scanning is potentially about a large SEPT in a VM with terabytes of
> memory.
> - Exporting is only about the pages found to be dirty.
>
> So this is only to remind that in the CoCo case, memory export also has a
> high price tag, compared to a traditional VM.
Ok, I'll keep that in mind.
>
> > So we have two overheads here at this stage, unpredictable:
> >
> > (a) Scanning EPT pgtable, when very unlucky, can take a lot of time to
> > finally reports to a GET_DIRTY_LOG request,
> >
> > (b) Migrating of dirty pages during blackout phase, which should be
> > roughly linear to how many dirty pages we just collected. (NOTE! I
> > think we may have way to fix this (b) or optimize it.. but this is
> > off-topic; let's focus on the difference of (a) and (b) first)
> >
> > When with PML, IIUC (a) is predictable: we have the bitmap on hand, plus a
> > maximum of some (my memory is, 512?) PML entries to flush per vCPU.
> > That'll be flushed automatically when we do vm_stop(), likely also
> > concurrently, atomically updating the bitmaps. I never measured it, but it
> > is bounded, and sounds pretty fast.
>
> Yes, I agree.
>
> Now my secret desire is that TDX module can eventually plug PML under the
> hood, consider it a "hardware accelerator" without changing the ABI. But I
> do not know whether keeping the ABI unchanged is possible, or how soon it
> could happen. This is something I am working on internally with Intel TDX
> module team.
>
> > When with scanning, (a) seems more unpredictable. That's the part I was
> > slightly concerned. But now after thinking a bit more, it seems fine.
> > Please read below.
>
> Sure, thanks.
>
> > > Intuitively, the TDX case does feel "less predictable". The open
> > > question for me is whether the degree of unpredictability is large enough
> > > to bother users. My attitude is to focus on getting something simple done
> > > first, learn from real-world behavior, and improve it later if needed,
> > > including exploring PML. The best is the enemy of the good sort of
> > > attitude.
> >
> > Yes, I think it's always fine we start with whatever is most feasible.
> >
> > I think it actually may not be that bad. The last sync is special at least
> > on how QEMU treats it, it should look like:
> >
> > - GET_DIRTY_LOG, to do last math, decide to switchover, <------ [1]
> > - vm_stop()
> > - GET_DIRTY_LOG, this collects all rest dirty bits <------ [2]
> > - migrates the dirty pages, device states, etc.
> >
> > So I expect there should be normally very small window between two
> > continuous GET_DIRTY_LOG across system. Only [2] will be part of downtime.
> >
> > Since you explained to me on how the background rescan roughly works, by
> > relying on A bit in pgtable directory entries, I do feel like in this case
> > most of the memory regions shouldn't be accessed during small window of
> > [1]->[2], then the range to scan should be very much under control too. In
> > reality, it will likely be even smaller, [1]->vm_stop(), because after that
> > vCPUs are halted.
>
> Yes, I agree. Just to flag the word "background" again, and to make sure we
> are aligned - this "rescan" happens between [1] and [2]. The idea is that
> [2] will be very fast after the "rescan". But it does increase the time
> between [1] and [2], and the TD has time to dirty more pages.
In general, looks like either way we can still stick with GET_DIRTY_LOG.
PML would be better, otherwise from a review on the API level, that's still
ok.
Thanks,
--
Peter Xu
^ permalink raw reply [flat|nested] 78+ messages in thread
* Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
2026-09-29 1:28 ` Kishen Maloor
@ 2026-09-30 20:42 ` Peter Xu
0 siblings, 0 replies; 78+ messages in thread
From: Peter Xu @ 2026-09-30 20:42 UTC (permalink / raw)
To: Kishen Maloor
Cc: Artem Bityutskiy, Tony Lindgren, Paolo Bonzini,
Sean Christopherson, Fabiano Rosas, Jon Grimm, Pankaj Gupta,
Tom Lendacky, Marc Zyngier, Oliver Upton, Steven Price,
Anup Patel, Samuel Ortiz, Jakub Růžička,
Jörg Rödel, Vishal Annapurve, Elena Reshetova,
Kai Huang, Mika Westerberg, Peter Fang, Rick Edgecombe,
Xiaoyao Li, Xu Yilun, kvm
On Mon, Sep 28, 2026 at 06:28:39PM -0700, Kishen Maloor wrote:
> Hi Peter,
Hi, Kishen,
[...]
> > So it seems TDX will "lock" the memory footprint in place when enabled,
> > allow accept/unaccept (or say, plug / unplug) memories, but anything like
> > "flipping this to that" will not work.
>
> Yes, I think so. The TDX module holds the GPA->page binding, so any change has
> to pass through it. Adding and removing pages plumb through the module.
> But a sort of content preserving remap of GPAs in the way QEMU does for PAM
> is not possible AFAIU.
This still looks fine for TDX then, but I think we may want to make sure it
works properly for all the rest consumers of this API that we're aware of,
say, AMD, ARM CCA, etc.
Especially, MEM.IMPORT seems trickier to me: MEM.EXPORT even if described
in in GPA address space, can be converted to other address spaces at will,
as long as QEMU has full knowledge of the mappings. MEM.IMPORT, OTOH, may
not trivially work in GPA terms, if memory mappings can overlap.
>
> > Then there's a very corner case question I want to double check, and I
> > apologize if this is stupid only due to my ignorance on TDX knowledge: can
> > someone migrate a VM too early so TDX is just hasn't been enabled at all?
>
> Not a stupid question at all. I've actually tried this. TDX VMs appear to
> migrate fine beyond a certain point mid-boot of the source. Kicking off
> migration any earlier causes the destination not to resume, even though the
> migration itself reports success. I haven't yet confirmed what that point is
> to explain it.
>
> For now we're trying to settle on sound fundamentals for the UAPI and flow,
> but this is definitely an area to dig further into.
Ok. Before any fixing lands, perhaps one can add a temprorary migration
blocker in QEMU and remove it after that mid-boot phase.
>
> >>> I'm a bit surprised that TDX will also monitor how many times the same GPA
> >>> is updated per iteration. What if below happens:
> >>
> >> Yes, and maybe that is because it won't know the ordering if the same
> >> GPA were caught at different times in the same epoch and fed to different
> >> migration streams (which one is newer?)
> >> On regular VMs, I believe QEMU doesn't migrate the same GPA twice in one round.
> >> As I understand it, a migrated GPA is revisited only in the next round.
> >> I suppose TDX just makes that behavior architectural.
> >>
> >>> ...
> >>> ITERATION sync n
> >>> GET_DIRTY_LOG, see page P dirty
> >>> migrate page P
> >>> GET_DIRTY_LOG, see page P dirty again
> >>> migrate page P again <--------------------- [a]
> >>
> >> QEMU shouldn't transfer P again in the same round, right?
> >
> > I believe yes with current QEMU, I can't think of anything otherwise. But
> > still, this is very specific impl detail. There's definitely no issue
> > migrating one page twice or more in non-CoCo.
>
> Thanks for confirming this. But do you think it would be a problem during
> pre-copy with multifd on regular VM migration if we didn't have this
> "transfer once per round" logic? If the same page got fed to different
> multifd queues, couldn't the older copy land after the newer one at the
> destination? [1]
Multifd has a flush logic per-iteration, so no chance an old version page
lands after a new version. Can refer to multifd_ram_round_notify().
QEMU sometimes calls it "round" in the code or comment but it is the
ITERATION term we're talking about here.
NOTE, that function was merged very recently; you may need to pull latest
QEMU, but the logic existed for a long time, so even old multifd behaves
that way.
>
> With a single channel it shouldn't matter, but it could with multiple
> channels. I think since the bitmap sweep is shared with multifd, QEMU ends up
> transferring a page only once per round either way.
Normally, yes, the scanner submits pages to multifd threads, the scanner,
if scan one round of bitmap, can only submit one page once until the next
round / ITERATION.
But it may also be a matter of how that ioctl(ITERATION) will be done
inside QEMU; I'll need to reference to the QEMU PoC branch when ready to be
sure I have the correct understanding. E.g., IIUC that ioctl(ITERATION)
needs to be done similarly "per-round" here to make sure one page appears
once, and we should also always keep in mind QEMU can still have
GET_DIRTY_LOG run in the background concurrently of the scanner, as I
mentioned that too somewhere else (which means one page can be dirtied
right after sending it, but as long as the scanner moves on with the next
page it looks still OK).
Thanks,
--
Peter Xu
^ permalink raw reply [flat|nested] 78+ messages in thread
end of thread, other threads:[~2026-09-30 20:42 UTC | newest]
Thread overview: 78+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-31 7:13 [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Tony Lindgren
2026-08-31 7:13 ` [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests Tony Lindgren
2026-08-31 7:20 ` sashiko-bot
2026-09-18 11:35 ` Peter Xu
2026-09-21 4:20 ` Tony Lindgren
2026-09-24 1:50 ` Wei Wang
2026-09-24 4:51 ` Tony Lindgren
2026-08-31 7:13 ` [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD Tony Lindgren
2026-08-31 7:23 ` sashiko-bot
2026-09-01 6:03 ` Tony Lindgren
2026-09-07 11:53 ` Tony Lindgren
2026-09-07 13:15 ` Jörg Rödel
2026-09-07 13:32 ` Artem Bityutskiy
2026-09-08 4:15 ` Tony Lindgren
2026-09-08 4:43 ` Tony Lindgren
2026-09-09 0:22 ` Kishen Maloor
2026-09-09 6:57 ` Tony Lindgren
2026-09-10 1:11 ` Kishen Maloor
2026-09-10 6:33 ` Tony Lindgren
2026-09-11 1:40 ` Kishen Maloor
2026-09-11 4:23 ` Tony Lindgren
2026-09-15 0:14 ` Kishen Maloor
2026-09-15 4:44 ` Tony Lindgren
2026-09-15 15:53 ` Kishen Maloor
2026-09-16 5:09 ` Tony Lindgren
2026-09-17 3:31 ` Kishen Maloor
2026-09-17 6:42 ` Tony Lindgren
2026-09-18 4:32 ` Kishen Maloor
2026-09-18 5:58 ` Tony Lindgren
2026-09-21 0:13 ` Kishen Maloor
2026-09-21 6:52 ` Tony Lindgren
2026-09-21 9:24 ` Tony Lindgren
2026-09-21 10:58 ` Tony Lindgren
2026-09-22 3:57 ` Kishen Maloor
2026-09-22 5:25 ` Tony Lindgren
2026-09-23 0:38 ` Kishen Maloor
2026-09-23 6:04 ` Tony Lindgren
2026-09-24 5:53 ` Kishen Maloor
2026-09-24 6:59 ` Tony Lindgren
2026-09-18 4:33 ` Kishen Maloor
2026-09-21 5:58 ` Tony Lindgren
2026-09-21 6:56 ` Tony Lindgren
2026-09-22 3:56 ` Kishen Maloor
2026-09-22 6:27 ` Tony Lindgren
2026-09-23 0:37 ` Kishen Maloor
2026-09-23 6:50 ` Tony Lindgren
2026-09-24 5:34 ` Kishen Maloor
2026-09-24 7:15 ` Tony Lindgren
2026-08-31 7:13 ` [RFC PATCH v2 3/4] KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY Tony Lindgren
2026-08-31 7:23 ` sashiko-bot
2026-09-01 6:10 ` Tony Lindgren
2026-08-31 7:13 ` [RFC PATCH v2 4/4] KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU Tony Lindgren
2026-08-31 7:23 ` sashiko-bot
2026-09-01 6:12 ` Tony Lindgren
2026-09-04 18:24 ` [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Artem Bityutskiy
2026-09-17 21:27 ` Peter Xu
2026-09-18 12:46 ` Artem Bityutskiy
2026-09-18 15:53 ` Peter Xu
2026-09-22 8:09 ` Artem Bityutskiy
2026-09-22 9:42 ` Tony Lindgren
2026-09-22 11:54 ` Artem Bityutskiy
2026-09-23 4:20 ` Tony Lindgren
2026-09-22 21:18 ` Peter Xu
2026-09-23 12:05 ` Artem Bityutskiy
2026-09-24 21:19 ` Peter Xu
2026-09-28 14:15 ` Artem Bityutskiy
2026-09-29 21:05 ` Peter Xu
2026-09-23 15:28 ` Serge Hallyn (AMD)
2026-09-20 23:56 ` Kishen Maloor
2026-09-23 21:36 ` Peter Xu
2026-09-24 4:27 ` Kishen Maloor
2026-09-25 14:18 ` Peter Xu
2026-09-29 1:28 ` Kishen Maloor
2026-09-30 20:42 ` Peter Xu
2026-09-18 18:36 ` Ionut Mihalcea
2026-09-21 4:35 ` Tony Lindgren
2026-09-25 16:03 ` Serge Hallyn (AMD)
2026-09-28 3:24 ` Kishen Maloor
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.