From: Artem Bityutskiy <dedekind1@gmail.com>
To: Peter Xu <peterx@redhat.com>
Cc: "Tony Lindgren" <tony.lindgren@linux.intel.com>,
"Paolo Bonzini" <pbonzini@redhat.com>,
"Sean Christopherson" <seanjc@google.com>,
"Fabiano Rosas" <farosas@suse.de>,
"Jon Grimm" <Jon.Grimm@amd.com>,
"Pankaj Gupta" <pankaj.gupta@amd.com>,
"Tom Lendacky" <thomas.lendacky@amd.com>,
"Marc Zyngier" <maz@kernel.org>,
"Oliver Upton" <oliver.upton@linux.dev>,
"Steven Price" <steven.price@arm.com>,
"Anup Patel" <anup@brainfault.org>,
"Samuel Ortiz" <sameo@rivosinc.com>,
"Jakub Růžička" <jakub.ruzicka@matfyz.cz>,
"Jörg Rödel" <joro@8bytes.org>,
"Vishal Annapurve" <vannapurve@google.com>,
"Elena Reshetova" <elena.reshetova@intel.com>,
"Kai Huang" <kai.huang@intel.com>,
"Kishen Maloor" <kishen.maloor@intel.com>,
"Mika Westerberg" <mika.westerberg@linux.intel.com>,
"Peter Fang" <peter.fang@intel.com>,
"Rick Edgecombe" <rick.p.edgecombe@intel.com>,
"Xiaoyao Li" <xiaoyao.li@intel.com>,
"Xu Yilun" <yilun.xu@linux.intel.com>,
kvm@vger.kernel.org
Subject: Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
Date: Fri, 18 Sep 2026 15:46:32 +0300 [thread overview]
Message-ID: <84bf61e0e810859ed735dc92ab94167727c2e560.camel@gmail.com> (raw)
In-Reply-To: <aqxbMGLwvtT8zkEb@zhexu-thinkpadt14gen5.rmtcaon.csb>
Hi Peter,
thank you for your good comments and questions. A quick disclaimer before I
address them. In my answers I try to keep three distinct things separate:
1. The CoCo migration uAPI - the generic interface we ultimately want to
converge on. This is the end goal, not necessarily what these patches
propose. What is good for this uAPI is my priority at this point.
2. The TDX migration model - the mechanics offered by Intel TDX module for
TDX guest live migration today.
3. Our PoC - a concrete uAPI proposal and its example implementation for TDX
guests.
On Thu, 2026-09-17 at 17:27 -0400, Peter Xu wrote:
> > Some example high-level topics that would be nice to get feedback on:
> >
> > - Can we come up with a single set of generic migration uAPIs for different
> > CoCo models?
> > - Or should some uAPIs be generic while others are vendor-specific?
> > - Or should each CoCo model have its own vendor-specific set of migration
> > uAPIs?
>
> It's always good if we can put together as much function to be shared with
> generic ioctls as possible. At some point, IMHO we need to collect such
> information somehow, so when merging the generic API we know what vendor
> specific API will be needed. Hopefully this series is a good start.
Thanks, agreed.
On that note - does anyone know of a good doc describing the AMD, ARM, or
other CoCo migration models? My knowledge is limited to TDX, so it is hard
to tell what is common and what is TDX-specific.
FYI, I am working on a TDX migration model document. It describes what the
TDX module offers, but unlike the specs it is oriented towards software
engineers: much easier to read and it does not require deep TDX knowledge.
It is all based on public specs, just distilled into readable mental models.
I plan to publish it publicly. I am about 80% done.
> > - Should the same uAPIs also support traditional VMs? But the only use-case
> > I imagine here is "for testing purposes".
>
> This is an interesting idea, I think this could be useful. Especially, I
> wonder if you already have it done and PoC branches you can share, so that
> I can play with it.
We do not have code. But Kishen spent time playing with it, and I think he
concluded not to proceed with this. But he might have evaluated it from
the "unify all migration into a single generic API" perspective. But may
be as a "this is a test framework" perspective is different, at least I feel
it may be the case. I think Kishen can provide more insight if needed, he
is in CC.
> >
> > > Migration flow
> > > ==============
> > >
> > > Source host Destination host
> > > =========== ================
> > >
> > > CMD(SETUP/SESSION) <--- setup msgs ---> CMD(SETUP/SESSION)
> > > | (repeated) |
> > > CMD(SETUP/IMMUTABLE_STATE) - immutable state -> CMD(SETUP/IMMUTABLE_STATE)
> > > | |
> > > KVM_GET_DIRTY_LOG |
> > > KVM_EXPORT_MEMORY --- memory data ---> KVM_IMPORT_MEMORY
> > > CMD(ITERATION) --- epoch token ---> CMD(ITERATION)
>
> Could you elaborate this ITERATION operation? Is that something the
> userapp must do after full scan of a round of guest memory?
= Why no duplicate instances? =
First, let me answer the "why" question you asked further below: "why does
the TDX module require that only the source or only the destination runs,
never both?". Answering it first makes the tokens easier to understand. And
the tokens are why we proposed the "ITERATION" operation.
I believe this is not TDX-specific, it is a confidential computing
requirement. In short, cloning would give the VMM a very powerful primitive
to attack confidential VMs. Here are a couple of example attack approaches:
- When you attest a CoCo VM remotely, you get an assurance that you talk to
this one specific instance. If duplication were allowed, many instances
could exist, and that assurance is gone.
- With a clone you can security-upgrade the state of one copy, attest the
upgraded state, and then use the pre-upgraded copy: the user believes they
are working with an up-to-date CoCo VM, but in fact use an older, possibly
vulnerable version.
This is also why TDX migration requires that not only must the two copies
never run at the same time, but after migration the destination must be
exactly the same as the source. For example, the destination must not end up
using an older copy of a page.
= ITERATION Operation =
In the QEMU model, the pre-copy phase is a set of rounds:
1. Get the list of dirty pages.
2. Copy them to the destination.
3. Repeat until the convergence criteria are met.
ITERATION is the explicit uAPI that ends the current pre-copy round. There
is no equivalent uAPI for traditional VMs today.
In the TDX model, ending a round needs an extra step: the source generates
an epoch token, and the destination imports it. Two seamcalls do this:
- TDH.EXPORT.TRACK generates the epoch token on the source.
- TDH.IMPORT.TRACK imports it on the destination.
The token enforces integrity and ordering, for example:
- Every page exported on the source must be imported on the destination.
- Once a newer version of a page is imported, an older version can no longer
be imported.
The final round is special. It uses the "done" flag, which is passed to the
`TDH.EXPORT.TRACK` seamcall, and makes it export a special variant of the
epoch token that is called the start token. On top of the integrity and
ordering guarantees, the start token is what allows the destination to
start: until the destination imports it, the TDX module will not let the
destination TD run with partial state.
> > > | (repeat until convergence) |
> > > CMD(STOP_AND_COPY/PAUSE) |
> > > CMD(STOP_AND_COPY/TD_STATE) --- VM state ------> CMD(STOP_AND_COPY/TD_STATE)
> > > KVM_EXPORT_VCPU --- vCPU state ----> KVM_IMPORT_VCPU
> > > KVM_EXPORT_MEMORY -- final memory ---> KVM_IMPORT_MEMORY
>
> When read/write encrypted memories, two questions:
>
> - Is there an upper bound of the buffer size per-page?
So generally, the assumption is that exporting N pages requires M pages,
M > N, because there may be some metadata (e.g., MACs for integrity
checks). In case of TDX module, M is predictable and can be calculated in
advance.
I believe in our current PoC, the ioctl requires the buffer size to be large
enough to hold all the requested pages. But this is specific to our current
PoC implementation.
In general, I feel like if buffer size is not enough, the uAPI could fill it
with as much data as fits, and communicate back about what GPAs were
exported. The caller could export the rest separately.
Context: I am new in the Intel TDX live migration team, and did not
participate in TDX PoC, that's why I use "I believe". I am catching up. But
I assume others will (Tony, Kishen) will correct me if I am wrong.
> - Does this operation supports concurrency? If it supports, how well it
> scales per expectation (e.g. is there known big lock for that)?
From the TDX module perspective, parallel exports of different GPAs can run
on multiple CPUs, so I expect the QEMU multifd model to work and scale.
In our current PoC the ioctl does not take a VM-wide lock, and concurrency
is per-stream. I can expand on the stream concept if needed, but it is
exactly about parallel import/export of memory and vCPU state.
In our PoC we are focusing on the basics, but multifd support is definitely
a goal too, just later. Kishen was already prototyping it in QEMU.
From the uAPI point of view, I believe parallel export/import should be
allowed. If a specific CoCo VM has issues with that, it would need to
serialize the operations internally, I'd say.
> - Does this operation supports concurrency? If it supports, how well it
> scales per expectation (e.g. is there known big lock for that)?
>
> Similar question to the vCPU getter and setter. For now even without CoCo
> we serialize vCPU get/set, but I want to understand the potential of
> concurrent operations, and see if there's anything special for CoCo from
> that regard.
Similar to memory import/export: the TDX module explicitly allows vCPU state
to be exported and imported in parallel.
Our current PoC does not take advantage of this yet. vCPU export currently
grabs the KVM MMU write lock, so vCPU exports are serialized today. I think
Tony can comment more on the technical difficulties there.
From the uAPI point of view, I'd propose to allow concurrent vCPU
operations.
> >
> > It describes what the TDX module offers today and focuses on pre-copy
> > migration. This is our interpretation of the TDX specifications, not a
>
> IMHO we should really take postcopy into account when designing the API and
> state machine. We don't need to implement it in the first version, even
> until merging, but we need to make sure postcopy will be new ioctls on top
> of existing and it should have no major loopholes that it'll need a new set
> of APIs.
I totally agree. I have not yet dug into the TDX module implementation
details for post-copy, but I know it is supported and I know the basics. I
plan to study it in detail later. So far I have not noticed anything that
would prevent adding post-copy on top later.
>
> For example, I think we should consider KVM_EXPORT_MEMORY being usable
> after END on source, KVM_IMPORT_MEMORY while TD is in operation, etc. We
> should likely also need to still picture the rough process of postcopy,
> reserve those APIs since the start (but return -EINVAL or something).
Yes, agreed. I will spend more time looking at this. But at this point, I
just assumed that the proposed uAPIs can be used at the post-copy phase in
parallel with on-demand page delivery.
Just FYI, TDX module model allows for this, but we did not try it.
>
> AFAIU, postcopy is so far still the best solution for extremely large or
> extremely busy VMs regarding user experience, and it will happen to CoCo
> VMs one day or another.
Sure, thanks for sharing.
> > destination may run, but never both. In other words, cloning a TD is not
> > allowed.
> > - When migration completes, the destination must have the same memory and
> > vCPU state as the source. It must not end up with a partial or mixed
> > state.
>
> If such happens, it's definitely a bug, even without CoCo. Anything
> specific about CoCo? Like, whole-VM checksum?
Well, in CoCo VMs it is not just a bug, it is something the CoCo framework
needs to make impossible, because VMM is considered to be untrusted, it can
try to manipulate things and half-migrate, use it not as a bug but as attack
vector. In TDX case, the TDX module will not allow you to run the TD - the
TDH.VP.ENTER seamcall will fail.
Regarding checksums: there is no single whole-VM checksum in TDX migration
model. Instead integrity is enforced continuously - every exported blob
carries a MACs that the destination TDX module verifies on import, and the
epoch/start tokens guarantee that everything was imported, in order.
> I want to understand what is extra for a CoCo VM in terms of "pause", say,
> what's more than "stopping the vCPU threads".
TDX module guarantees the source won't run, even if VMM tries, the
TDH.VP.ENTER seamcall will fail. So the source TD state is effectively
frozen and cannot be modified by the VMM.
>
> I saw there's mention of PRE_COPY_STOP state. One example question is,
> when reaching this state, can the guest memory still change? What happens
> if some emulated device are still DMAing to the guest memory (assuming
> flipped from private to shared)? In case of future IO zone support, what
> happens if in case of VFIO-PCI assigned doing encrypted DMA?
>
> From that regard, VFIO has the P2P state where it quiesce initiation of any
> DMA from this specific device, then another round to fully stop all devices
> into STOP_COPY phase. I wonder if CoCo VMs need similar treatment.
Let me split this by device type, because TDX treats them very differently.
Emulated (virtio-net, virtio-blk, etc.) only use shared memory - they cannot
read or DMA into TD private memory. So full device state lives in shared
memory, and QEMU migrates them exactly the same way as for a traditional VM.
This is entirely outside the TDX module migration model and outside the
proposed uAPI - the uAPI is only for TD private memory.
Directly assigned devices are only possible with TDX Connect, where a
physical PCIe/CXL function (a "TDI") is assigned to the TD and can DMA into
private memory over a cryptographically protected link. This is not
implemented in Linux yet. For migration, the TDX module requires all TDIs to
be unassigned before the source TD is paused - the TDH.EXPORT.PAUSE seamcall
actually checks this. Unassigning a TDI tears down its whole TD-private
footprint (MMIO unmapped from the Secure EPT, trusted DMA mappings removed),
so no device-specific state is left to migrate. From the TD's point of view
it is a full hot-unplug on the source and a fresh hot-plug on the
destination.
Our current TDX guest migration PoC is built on this assumption.
> Could you elaborate what's the relations between TDH.MEM.SCAN.RANGE and the
> GET_DIRTY_LOG ioctl? I recall above mentioned GET_DIRTY_LOG will be
> available even for CoCo, which makes sense assuming dirty information isn't
> confidential. However then I don't understand what TDH.MEM.SCAN.RANGE
> plays the role here.
A few things here.
First, our plan is that the standard GET_DIRTY_LOG ioctl is backed by the
`TDH.MEM.SCAN.RANGE` seamcall - that is how dirty tracking is implemented
for a TD. So we do not propose any special uAPI for dirty tracking.
And I think you are right that the dirty information is not confidential -
the TDX module exposes dirty page information to the VMM.
Second, FYI, in current TDX migration model, dirty page scanning
(`TDH.MEM.SCAN.RANGE`) is only allowed during a migration session. The TDX
module returns an error if the seamcall is issued before the session is set
up (i.e. before the source and destination TDX modules have exchanged the
migration key and established trust - what the SETUP command does in the
proposed uAPI).
In other words, with our PoC, if someone tries to use GET_DIRTY_LOG without
going through the migration setup - the ioctl will return an error.
But I wish it were an independent feature instead. Then we could work on
upstreaming it on its own - Tony estimates it is about 20% of the current
TDX migration PoC code.
I already raised this with the Intel TDX module architects, and they asked
for use-cases. The only one we came up with is QEMU estimating the TD dirty
rate before starting migration (the calc_dirty_rate command). I understand
their position: without a use-case there is little reason to implement it,
and they would also need to study the security implications - can it help
an attacker in some way?
So if you or anyone else can educate me about use-cases for independent
dirty page scanning, I would really appreciate it - I could take them back
to the TDX module architects.
Thanks,
Artem.
next prev parent reply other threads:[~2026-09-18 12:46 UTC|newest]
Thread overview: 85+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-31 7:13 [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Tony Lindgren
2026-08-31 7:13 ` [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests Tony Lindgren
2026-08-31 7:20 ` sashiko-bot
2026-09-18 11:35 ` Peter Xu
2026-09-21 4:20 ` Tony Lindgren
2026-09-24 1:50 ` Wei Wang
2026-09-24 4:51 ` Tony Lindgren
2026-08-31 7:13 ` [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD Tony Lindgren
2026-08-31 7:23 ` sashiko-bot
2026-09-01 6:03 ` Tony Lindgren
2026-09-07 11:53 ` Tony Lindgren
2026-09-07 13:15 ` Jörg Rödel
2026-09-07 13:32 ` Artem Bityutskiy
2026-09-08 4:15 ` Tony Lindgren
2026-09-08 4:43 ` Tony Lindgren
2026-09-09 0:22 ` Kishen Maloor
2026-09-09 6:57 ` Tony Lindgren
2026-09-10 1:11 ` Kishen Maloor
2026-09-10 6:33 ` Tony Lindgren
2026-09-11 1:40 ` Kishen Maloor
2026-09-11 4:23 ` Tony Lindgren
2026-09-15 0:14 ` Kishen Maloor
2026-09-15 4:44 ` Tony Lindgren
2026-09-15 15:53 ` Kishen Maloor
2026-09-16 5:09 ` Tony Lindgren
2026-09-17 3:31 ` Kishen Maloor
2026-09-17 6:42 ` Tony Lindgren
2026-09-18 4:32 ` Kishen Maloor
2026-09-18 5:58 ` Tony Lindgren
2026-09-21 0:13 ` Kishen Maloor
2026-09-21 6:52 ` Tony Lindgren
2026-09-21 9:24 ` Tony Lindgren
2026-09-21 10:58 ` Tony Lindgren
2026-09-22 3:57 ` Kishen Maloor
2026-09-22 5:25 ` Tony Lindgren
2026-09-23 0:38 ` Kishen Maloor
2026-09-23 6:04 ` Tony Lindgren
2026-09-24 5:53 ` Kishen Maloor
2026-09-24 6:59 ` Tony Lindgren
2026-09-18 4:33 ` Kishen Maloor
2026-09-21 5:58 ` Tony Lindgren
2026-09-21 6:56 ` Tony Lindgren
2026-09-22 3:56 ` Kishen Maloor
2026-09-22 6:27 ` Tony Lindgren
2026-09-23 0:37 ` Kishen Maloor
2026-09-23 6:50 ` Tony Lindgren
2026-09-24 5:34 ` Kishen Maloor
2026-09-24 7:15 ` Tony Lindgren
2026-10-08 9:22 ` Tony Lindgren
2026-08-31 7:13 ` [RFC PATCH v2 3/4] KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY Tony Lindgren
2026-08-31 7:23 ` sashiko-bot
2026-09-01 6:10 ` Tony Lindgren
2026-08-31 7:13 ` [RFC PATCH v2 4/4] KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU Tony Lindgren
2026-08-31 7:23 ` sashiko-bot
2026-09-01 6:12 ` Tony Lindgren
2026-09-04 18:24 ` [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Artem Bityutskiy
2026-09-17 21:27 ` Peter Xu
2026-09-18 12:46 ` Artem Bityutskiy [this message]
2026-09-18 15:53 ` Peter Xu
2026-09-22 8:09 ` Artem Bityutskiy
2026-09-22 9:42 ` Tony Lindgren
2026-09-22 11:54 ` Artem Bityutskiy
2026-09-23 4:20 ` Tony Lindgren
2026-09-22 21:18 ` Peter Xu
2026-09-23 12:05 ` Artem Bityutskiy
2026-09-24 21:19 ` Peter Xu
2026-09-28 14:15 ` Artem Bityutskiy
2026-09-29 21:05 ` Peter Xu
2026-10-02 19:57 ` Artem Bityutskiy
2026-10-07 20:00 ` Peter Xu
2026-09-23 15:28 ` Serge Hallyn (AMD)
2026-09-20 23:56 ` Kishen Maloor
2026-09-23 21:36 ` Peter Xu
2026-09-24 4:27 ` Kishen Maloor
2026-09-25 14:18 ` Peter Xu
2026-09-29 1:28 ` Kishen Maloor
2026-09-30 20:42 ` Peter Xu
2026-10-07 4:27 ` Kishen Maloor
2026-10-07 20:13 ` Peter Xu
2026-10-08 6:23 ` Tony Lindgren
2026-10-08 14:31 ` Peter Xu
2026-09-18 18:36 ` Ionut Mihalcea
2026-09-21 4:35 ` Tony Lindgren
2026-09-25 16:03 ` Serge Hallyn (AMD)
2026-09-28 3:24 ` Kishen Maloor
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=84bf61e0e810859ed735dc92ab94167727c2e560.camel@gmail.com \
--to=dedekind1@gmail.com \
--cc=Jon.Grimm@amd.com \
--cc=anup@brainfault.org \
--cc=elena.reshetova@intel.com \
--cc=farosas@suse.de \
--cc=jakub.ruzicka@matfyz.cz \
--cc=joro@8bytes.org \
--cc=kai.huang@intel.com \
--cc=kishen.maloor@intel.com \
--cc=kvm@vger.kernel.org \
--cc=maz@kernel.org \
--cc=mika.westerberg@linux.intel.com \
--cc=oliver.upton@linux.dev \
--cc=pankaj.gupta@amd.com \
--cc=pbonzini@redhat.com \
--cc=peter.fang@intel.com \
--cc=peterx@redhat.com \
--cc=rick.p.edgecombe@intel.com \
--cc=sameo@rivosinc.com \
--cc=seanjc@google.com \
--cc=steven.price@arm.com \
--cc=thomas.lendacky@amd.com \
--cc=tony.lindgren@linux.intel.com \
--cc=vannapurve@google.com \
--cc=xiaoyao.li@intel.com \
--cc=yilun.xu@linux.intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox