From: Kishen Maloor <kishen.maloor@intel.com>
To: Peter Xu <peterx@redhat.com>
Cc: "Artem Bityutskiy" <dedekind1@gmail.com>,
"Tony Lindgren" <tony.lindgren@linux.intel.com>,
"Paolo Bonzini" <pbonzini@redhat.com>,
"Sean Christopherson" <seanjc@google.com>,
"Fabiano Rosas" <farosas@suse.de>,
"Jon Grimm" <Jon.Grimm@amd.com>,
"Pankaj Gupta" <pankaj.gupta@amd.com>,
"Tom Lendacky" <thomas.lendacky@amd.com>,
"Marc Zyngier" <maz@kernel.org>,
"Oliver Upton" <oliver.upton@linux.dev>,
"Steven Price" <steven.price@arm.com>,
"Anup Patel" <anup@brainfault.org>,
"Samuel Ortiz" <sameo@rivosinc.com>,
"Jakub Růžička" <jakub.ruzicka@matfyz.cz>,
"Jörg Rödel" <joro@8bytes.org>,
"Vishal Annapurve" <vannapurve@google.com>,
"Elena Reshetova" <elena.reshetova@intel.com>,
"Kai Huang" <kai.huang@intel.com>,
"Mika Westerberg" <mika.westerberg@linux.intel.com>,
"Peter Fang" <peter.fang@intel.com>,
"Rick Edgecombe" <rick.p.edgecombe@intel.com>,
"Xiaoyao Li" <xiaoyao.li@intel.com>,
"Xu Yilun" <yilun.xu@linux.intel.com>,
kvm@vger.kernel.org
Subject: Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
Date: Mon, 28 Sep 2026 18:28:39 -0700 [thread overview]
Message-ID: <784c4ecc-da9d-44b8-9adc-1d1e92d20ae3@intel.com> (raw)
In-Reply-To: <araC0YkwLVPhGfy0@zhexu-thinkpadt14gen5.rmtcaon.csb>
Hi Peter,
Great comments!
On 9/25/26 7:18 AM, Peter Xu wrote:
> On Wed, Sep 23, 2026 at 09:27:52PM -0700, Kishen Maloor wrote:
> ...
> Yes, I think you're right, the ABI isn't the major issue. If from KVM's
> perspective it's always convertable between GPA <-> slots, then it's the
> same.
Correct. It's just that when the bytes aren't in any memslot there's no way
to reach pc.ram's HVA on the destination in KVM's view to write to it.
> I believe my mindset when replying was pretty much in QEMU's perspective,
> where in qemu we can have two ramblocks plugged into the same GPA range,
> only one of them will be visible to KVM and guest (e.g. which one has
> higher MemoryRegion priority, but it's not the only factor). What migration
I understand, and it matches what I found.
> module does right now is, it allows both ramblocks to be migrated with no
> issue, even if one is not visible, but if it used to be touched, since that
> ramblock will maintain its own dirty bitmap (GET_DIRTY_LOG on that kvm
> memslot when it was visible to KVM, or maybe set within QEMU userspace
> somehow).
Correct, and regular VM migration bypasses any notion of GPAs/memslots
entirely. It migrates RAMBlocks by block+offset. QEMU sets every RAMBlock's
migration bitmap to all ones before the first pass, so everything is
transferred in round 1 regardless of dirty state.
> IOW, what QEMU could do here is, after MEMORY.EXPORT, convert the GPA
> address space into ramblock ranges in QEMU, migrate with that not GPA. In
> case of PAM, it's part of pc.ram. Dest QEMU sees it is pc.ram, it should
> "apply" those data to pc.ram ramblock only.
>
> For non-CoCo, it's as easy as writting to some host HVA pointer.
I understand that this is how it works for regular VMs as the
destination is able to memcpy into the HVA obtained using the block+offset it
receives over the stream.
Also, when migration is kicked off after the source has fully booted, the
PAM window already maps to the source's pc.ram area, so the export transfers
the right bytes. The gap is purely on the destination because its layout is
still in the boot-time configuration until device state lands.
> Now, the question is, the new migration API only allows applying data in
> GPA ranges. I don't think with TDX there's a way to "apply" the data..
> dest QEMU just booted, this specific portion of pc.ram may not be mapped at
> that GPA source fetched due to reset status of PAM registers.
Correct. In TDX, the page is placed by the import SEAMCALL at the GPA it was
exported from. So QEMU has to turn the block+offset back into a GPA, and that
GPA has to be covered by a memslot for KVM to resolve it, which it is unable to
do at that point.
> That (rather than the ABI interface), might be the real thing I wanted to
> point out.
Understood. I did also wonder whether other VMMs organize their guest memory
hierarchy in similar ways.
> Maybe it means TDX just can't work with it by definition? I think it'll be
> fine, and now I wonder if it means PAM will be working for TDX only if PAM
> boots too early so it was before TDX initializes, then after TDX enabled
> anything like PAM will not work anymore?
Upstream QEMU gates the separate pc.rom allocation on !is_tdx_vm(), so there's
no second allocation to flip to. I don't see any other TDX gating around this,
so I presume that PAM operates as usual otherwise.
It doesn't appear to be a question of PAM coming up before TDX initializes.
> So it seems TDX will "lock" the memory footprint in place when enabled,
> allow accept/unaccept (or say, plug / unplug) memories, but anything like
> "flipping this to that" will not work.
Yes, I think so. The TDX module holds the GPA->page binding, so any change has
to pass through it. Adding and removing pages plumb through the module.
But a sort of content preserving remap of GPAs in the way QEMU does for PAM
is not possible AFAIU.
> Then there's a very corner case question I want to double check, and I
> apologize if this is stupid only due to my ignorance on TDX knowledge: can
> someone migrate a VM too early so TDX is just hasn't been enabled at all?
Not a stupid question at all. I've actually tried this. TDX VMs appear to
migrate fine beyond a certain point mid-boot of the source. Kicking off
migration any earlier causes the destination not to resume, even though the
migration itself reports success. I haven't yet confirmed what that point is
to explain it.
For now we're trying to settle on sound fundamentals for the UAPI and flow,
but this is definitely an area to dig further into.
>>> I'm a bit surprised that TDX will also monitor how many times the same GPA
>>> is updated per iteration. What if below happens:
>>
>> Yes, and maybe that is because it won't know the ordering if the same
>> GPA were caught at different times in the same epoch and fed to different
>> migration streams (which one is newer?)
>> On regular VMs, I believe QEMU doesn't migrate the same GPA twice in one round.
>> As I understand it, a migrated GPA is revisited only in the next round.
>> I suppose TDX just makes that behavior architectural.
>>
>>> ...
>>> ITERATION sync n
>>> GET_DIRTY_LOG, see page P dirty
>>> migrate page P
>>> GET_DIRTY_LOG, see page P dirty again
>>> migrate page P again <--------------------- [a]
>>
>> QEMU shouldn't transfer P again in the same round, right?
>
> I believe yes with current QEMU, I can't think of anything otherwise. But
> still, this is very specific impl detail. There's definitely no issue
> migrating one page twice or more in non-CoCo.
Thanks for confirming this. But do you think it would be a problem during
pre-copy with multifd on regular VM migration if we didn't have this
"transfer once per round" logic? If the same page got fed to different
multifd queues, couldn't the older copy land after the newer one at the
destination? [1]
With a single channel it shouldn't matter, but it could with multiple
channels. I think since the bitmap sweep is shared with multifd, QEMU ends up
transferring a page only once per round either way.
At least this was the rationale I conceived in trying to explain the TDX
policy of not allowing multiple imports of the same page in one epoch.
Though the underlying concern may not be TDX-specific.
> I can give one example to illustrate what could happen.
>
> In postcopy, we support preemption mode, which is simply a separate fast
> path for requested / urgent pages. It's possible while background thread
> transferring one page, the fast path saw a request on this same page. The
> current algorithm is simple, it will wait for that in progress background
> send to complete.
>
> But logically, we could do it the other way too: send the page again on
> fast path, in postcopy the page content is guaranteed to be identical and
> unchnaged, it means the fast path can land this page earlier, reducing
> fault latency. If so, a minimum cap we need is MEMORY.EXPORT be able to be
> done twice, so the fast path can read the 2nd time. IMPORT is more
We shall check about this. TDH.EXPORT.MEM (the export-side SEAMCALL on the
source) may already permit a 2nd export of the same page during post-copy.
If it doesn't, there's a good case for asking for it, since it would give
both the background and preemption threads a shot at fulfilling a request
ASAP. The import side might still reject it, so your suggestion below looks
like the right place to handle that.
> flexible, because QEMU can maintain what has been applied, so logically
> background loader should be able to skip the 2nd IMPORT.
This would make sense to do I suppose, because accepting a 2nd IMPORT would
clobber a page that the destination previously received and has itself
modified.
> That is not a good example, at least because it's postcopy and doesn't
> happen with precopy. So far, I also don't think a major risk, but I confess
> I don't understand why TDX needs to add hard requirement on "only sample
> one page once per iteration": even if VMM sampled a page twice, it's still
> encrypted and confidential. I believe it has something to do with the
> whole attestation logic. I think it'll be more flexible with less
> restrictions, but no issue I see either, hence please only treat that a
> verbose FYI.
I think the transfer once per iteration might have everything to do with
reason [1] I mentioned above. I've generally noticed that the TDX migration
architecture reflects established practices of VMMs. That being said, we
should look for instances where a restriction deviates and poses a problem.
>
>>
>>> ITERATION sync n+1
>>> ...
>>>
>>> Would above crash on destination TDX at step [a] applying P 2nd time? Or
>>
>> If P were migrated again, then I believe it would fail to import on the
>> destination, and also abort the import session, so migration would fail.
>
> I wonder if we can just fail the 2nd IMPORT without abort the whole
> process, then if userapp wants to detect it there's a way to (similar to an
> -EEXIST). Not a request, more like a pure question.
Understood. I think it's TDX's way of assuring correctness in case a VMM
transfers a page more than once in a round. QEMU doesn't. In theory it
could relax that when only one migration stream is registered as the
ordering would be unambiguous there, so a second import is provably newer
and could just be accepted.
next prev parent reply other threads:[~2026-09-29 1:28 UTC|newest]
Thread overview: 85+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-31 7:13 [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Tony Lindgren
2026-08-31 7:13 ` [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests Tony Lindgren
2026-08-31 7:20 ` sashiko-bot
2026-09-18 11:35 ` Peter Xu
2026-09-21 4:20 ` Tony Lindgren
2026-09-24 1:50 ` Wei Wang
2026-09-24 4:51 ` Tony Lindgren
2026-08-31 7:13 ` [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD Tony Lindgren
2026-08-31 7:23 ` sashiko-bot
2026-09-01 6:03 ` Tony Lindgren
2026-09-07 11:53 ` Tony Lindgren
2026-09-07 13:15 ` Jörg Rödel
2026-09-07 13:32 ` Artem Bityutskiy
2026-09-08 4:15 ` Tony Lindgren
2026-09-08 4:43 ` Tony Lindgren
2026-09-09 0:22 ` Kishen Maloor
2026-09-09 6:57 ` Tony Lindgren
2026-09-10 1:11 ` Kishen Maloor
2026-09-10 6:33 ` Tony Lindgren
2026-09-11 1:40 ` Kishen Maloor
2026-09-11 4:23 ` Tony Lindgren
2026-09-15 0:14 ` Kishen Maloor
2026-09-15 4:44 ` Tony Lindgren
2026-09-15 15:53 ` Kishen Maloor
2026-09-16 5:09 ` Tony Lindgren
2026-09-17 3:31 ` Kishen Maloor
2026-09-17 6:42 ` Tony Lindgren
2026-09-18 4:32 ` Kishen Maloor
2026-09-18 5:58 ` Tony Lindgren
2026-09-21 0:13 ` Kishen Maloor
2026-09-21 6:52 ` Tony Lindgren
2026-09-21 9:24 ` Tony Lindgren
2026-09-21 10:58 ` Tony Lindgren
2026-09-22 3:57 ` Kishen Maloor
2026-09-22 5:25 ` Tony Lindgren
2026-09-23 0:38 ` Kishen Maloor
2026-09-23 6:04 ` Tony Lindgren
2026-09-24 5:53 ` Kishen Maloor
2026-09-24 6:59 ` Tony Lindgren
2026-09-18 4:33 ` Kishen Maloor
2026-09-21 5:58 ` Tony Lindgren
2026-09-21 6:56 ` Tony Lindgren
2026-09-22 3:56 ` Kishen Maloor
2026-09-22 6:27 ` Tony Lindgren
2026-09-23 0:37 ` Kishen Maloor
2026-09-23 6:50 ` Tony Lindgren
2026-09-24 5:34 ` Kishen Maloor
2026-09-24 7:15 ` Tony Lindgren
2026-10-08 9:22 ` Tony Lindgren
2026-08-31 7:13 ` [RFC PATCH v2 3/4] KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY Tony Lindgren
2026-08-31 7:23 ` sashiko-bot
2026-09-01 6:10 ` Tony Lindgren
2026-08-31 7:13 ` [RFC PATCH v2 4/4] KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU Tony Lindgren
2026-08-31 7:23 ` sashiko-bot
2026-09-01 6:12 ` Tony Lindgren
2026-09-04 18:24 ` [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Artem Bityutskiy
2026-09-17 21:27 ` Peter Xu
2026-09-18 12:46 ` Artem Bityutskiy
2026-09-18 15:53 ` Peter Xu
2026-09-22 8:09 ` Artem Bityutskiy
2026-09-22 9:42 ` Tony Lindgren
2026-09-22 11:54 ` Artem Bityutskiy
2026-09-23 4:20 ` Tony Lindgren
2026-09-22 21:18 ` Peter Xu
2026-09-23 12:05 ` Artem Bityutskiy
2026-09-24 21:19 ` Peter Xu
2026-09-28 14:15 ` Artem Bityutskiy
2026-09-29 21:05 ` Peter Xu
2026-10-02 19:57 ` Artem Bityutskiy
2026-10-07 20:00 ` Peter Xu
2026-09-23 15:28 ` Serge Hallyn (AMD)
2026-09-20 23:56 ` Kishen Maloor
2026-09-23 21:36 ` Peter Xu
2026-09-24 4:27 ` Kishen Maloor
2026-09-25 14:18 ` Peter Xu
2026-09-29 1:28 ` Kishen Maloor [this message]
2026-09-30 20:42 ` Peter Xu
2026-10-07 4:27 ` Kishen Maloor
2026-10-07 20:13 ` Peter Xu
2026-10-08 6:23 ` Tony Lindgren
2026-10-08 14:31 ` Peter Xu
2026-09-18 18:36 ` Ionut Mihalcea
2026-09-21 4:35 ` Tony Lindgren
2026-09-25 16:03 ` Serge Hallyn (AMD)
2026-09-28 3:24 ` Kishen Maloor
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=784c4ecc-da9d-44b8-9adc-1d1e92d20ae3@intel.com \
--to=kishen.maloor@intel.com \
--cc=Jon.Grimm@amd.com \
--cc=anup@brainfault.org \
--cc=dedekind1@gmail.com \
--cc=elena.reshetova@intel.com \
--cc=farosas@suse.de \
--cc=jakub.ruzicka@matfyz.cz \
--cc=joro@8bytes.org \
--cc=kai.huang@intel.com \
--cc=kvm@vger.kernel.org \
--cc=maz@kernel.org \
--cc=mika.westerberg@linux.intel.com \
--cc=oliver.upton@linux.dev \
--cc=pankaj.gupta@amd.com \
--cc=pbonzini@redhat.com \
--cc=peter.fang@intel.com \
--cc=peterx@redhat.com \
--cc=rick.p.edgecombe@intel.com \
--cc=sameo@rivosinc.com \
--cc=seanjc@google.com \
--cc=steven.price@arm.com \
--cc=thomas.lendacky@amd.com \
--cc=tony.lindgren@linux.intel.com \
--cc=vannapurve@google.com \
--cc=xiaoyao.li@intel.com \
--cc=yilun.xu@linux.intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox