From: Kishen Maloor <kishen.maloor@intel.com>
To: Peter Xu <peterx@redhat.com>
Cc: "Artem Bityutskiy" <dedekind1@gmail.com>,
"Tony Lindgren" <tony.lindgren@linux.intel.com>,
"Paolo Bonzini" <pbonzini@redhat.com>,
"Sean Christopherson" <seanjc@google.com>,
"Fabiano Rosas" <farosas@suse.de>,
"Jon Grimm" <Jon.Grimm@amd.com>,
"Pankaj Gupta" <pankaj.gupta@amd.com>,
"Tom Lendacky" <thomas.lendacky@amd.com>,
"Marc Zyngier" <maz@kernel.org>,
"Oliver Upton" <oliver.upton@linux.dev>,
"Steven Price" <steven.price@arm.com>,
"Anup Patel" <anup@brainfault.org>,
"Samuel Ortiz" <sameo@rivosinc.com>,
"Jakub Růžička" <jakub.ruzicka@matfyz.cz>,
"Jörg Rödel" <joro@8bytes.org>,
"Vishal Annapurve" <vannapurve@google.com>,
"Elena Reshetova" <elena.reshetova@intel.com>,
"Kai Huang" <kai.huang@intel.com>,
"Mika Westerberg" <mika.westerberg@linux.intel.com>,
"Peter Fang" <peter.fang@intel.com>,
"Rick Edgecombe" <rick.p.edgecombe@intel.com>,
"Xiaoyao Li" <xiaoyao.li@intel.com>,
"Xu Yilun" <yilun.xu@linux.intel.com>,
kvm@vger.kernel.org
Subject: Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
Date: Wed, 23 Sep 2026 21:27:52 -0700 [thread overview]
Message-ID: <9fa229ea-818c-4ba9-81c6-be7ead5cb60a@intel.com> (raw)
In-Reply-To: <arRGYbvXgq8LLYQH@zhexu-thinkpadt14gen5.rmtcaon.csb>
Hi Peter,
Thanks! It really helps to go over details.
On 9/23/26 2:36 PM, Peter Xu wrote:
> ...
>
> I see this one slightly differently, and I have a major question on the
> choice of GPA for KVM's memory access API.
>
> If reusing GET_DIRTY_LOG, it means slot_id works for CoCo like before
> because that's the old interface there.
It works because we map backwards from GPAs in TDX's dirty scan results to
a KVM memslot on the source, and a bit in its dirty bitmap, in keeping with
the GET_DIRTY_LOG contract.
> But the new KVM_EXPORT_MEM (vice versa) used GPA arrays. Could I ask why
> the change?
The migration ABI (at least in TDX, possibly others) is GPA shaped. Also,
everything downstream like SEPT entries key off GPAs. The EXPORT and IMPORT
ABIs take GPAs alongside the encrypted blob, so both the source and the
destination need it at the call site that implements the UAPI. Passing GFNs
in the UAPI is merely a convenience in that respect. If we passed in
slot+offset instead, it would just be an indirection that needs to resolve
back to the GPA anyway to issue the vendor migration call. But this is not
the blocking issue for the PAM case, as I'll explain below.
> IIUC this is the fundamental reason why you hit that PAM issue: at a
> specific GPA, QEMU/KVM can map different things, hence GPA is not yet an
> unified identifier for a physical page that guest uses. However, (slot_id,
> slot_offset) will be.
>
> If the migration memory access API will be using that instead of GPA, I
> think problems will be gone because then the PAM regions (likely, pc.ram
> underneath) will be migrated in form of KVM memslots, destination apply
> that to the PAM's (slot_id, slot_offset), then when switchover, that PAM
> memory region will be enabled by QEMU and it will be visible to destination
> QEMU when VM starts on destination.
I think using slot+offset won't make the problem vanish, because the PAM window
in the underlying pc.ram region is not addressable by KVM on the destination at
that time. KVM's memslots come from QEMU's FlatView at any moment and QEMU
reconfigures them on the fly. In the destination's boot-time setup, it seems
like pc.ram's memslot is split around that window, and what covers the
range instead is a separate, read-only memslot for the overlay, which is an
allocation with its own backing, not the pc.ram bytes underneath.
And since only the flattened view is ever registered, there's no slot+offset
that refers to those bytes either. Only when the source's PAM configuration
lands at the destination, which happens after memory migration concludes, do
those pc.ram bytes become visible to KVM.
In the TDX case, QEMU does not create that overlay at the start, so the
destination has a memslot corresponding to the pc.ram region that stays put.
KVM is able to do a GFN->PFN translation to hand to TDX's IMPORT ABI.
So, it appears that a UAPI change will not address this corner case.
> I'm actually not sure if such (e.g. PAM) will ever be supported at all in
> CoCo context, but it seems to me using (slot_id, slot_offset) (for each
> entry; it can also be an array) is more flexible than GPA arrays in that it
> allows GPA mapping to change on the fly if needed even during migration.
>
> I'm not sure if that's explicitly forbidden for CoCo, though.
A TD private page lives at a fixed GPA in the SEPT, and I believe the
EXPORT/IMPORT bundle binds the GPA into the page's integrity check, so a page
has to be imported at the GPA it was exported from.
>>> Could you elaborate this ITERATION operation? Is that something the
>>> userapp must do after full scan of a round of guest memory?
>>
>> Essentially, yes. TDX migration architecture delimits such pre-copy rounds as
>> "migration epochs" and emits an epoch token at each round boundary that needs to be
>> consumed on the destination. It allows the TDX module to verify that everything from
>> the prior round has been received at the destination and that two versions of the
>> same GPA aren't sent in the same epoch. It is userspace that decides where a round
>> ends, but any deviation from this model and migration would fail.
>
> I'm a bit surprised that TDX will also monitor how many times the same GPA
> is updated per iteration. What if below happens:
Yes, and maybe that is because it won't know the ordering if the same
GPA were caught at different times in the same epoch and fed to different
migration streams (which one is newer?)
On regular VMs, I believe QEMU doesn't migrate the same GPA twice in one round.
As I understand it, a migrated GPA is revisited only in the next round.
I suppose TDX just makes that behavior architectural.
> ...
> ITERATION sync n
> GET_DIRTY_LOG, see page P dirty
> migrate page P
> GET_DIRTY_LOG, see page P dirty again
> migrate page P again <--------------------- [a]
QEMU shouldn't transfer P again in the same round, right?
> ITERATION sync n+1
> ...
>
> Would above crash on destination TDX at step [a] applying P 2nd time? Or
If P were migrated again, then I believe it would fail to import on the
destination, and also abort the import session, so migration would fail.
TDX records the current epoch against a GPA on every import, and I think
that's how it catches a 2nd import attempt.
> does it mean GET_DIRTY_LOG can only be invoked once per iteration?
We don't touch QEMU's stock dirty logging/transfer flows so however it
handles it currently stays.
next prev parent reply other threads:[~2026-09-24 4:28 UTC|newest]
Thread overview: 85+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-31 7:13 [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Tony Lindgren
2026-08-31 7:13 ` [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests Tony Lindgren
2026-08-31 7:20 ` sashiko-bot
2026-09-18 11:35 ` Peter Xu
2026-09-21 4:20 ` Tony Lindgren
2026-09-24 1:50 ` Wei Wang
2026-09-24 4:51 ` Tony Lindgren
2026-08-31 7:13 ` [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD Tony Lindgren
2026-08-31 7:23 ` sashiko-bot
2026-09-01 6:03 ` Tony Lindgren
2026-09-07 11:53 ` Tony Lindgren
2026-09-07 13:15 ` Jörg Rödel
2026-09-07 13:32 ` Artem Bityutskiy
2026-09-08 4:15 ` Tony Lindgren
2026-09-08 4:43 ` Tony Lindgren
2026-09-09 0:22 ` Kishen Maloor
2026-09-09 6:57 ` Tony Lindgren
2026-09-10 1:11 ` Kishen Maloor
2026-09-10 6:33 ` Tony Lindgren
2026-09-11 1:40 ` Kishen Maloor
2026-09-11 4:23 ` Tony Lindgren
2026-09-15 0:14 ` Kishen Maloor
2026-09-15 4:44 ` Tony Lindgren
2026-09-15 15:53 ` Kishen Maloor
2026-09-16 5:09 ` Tony Lindgren
2026-09-17 3:31 ` Kishen Maloor
2026-09-17 6:42 ` Tony Lindgren
2026-09-18 4:32 ` Kishen Maloor
2026-09-18 5:58 ` Tony Lindgren
2026-09-21 0:13 ` Kishen Maloor
2026-09-21 6:52 ` Tony Lindgren
2026-09-21 9:24 ` Tony Lindgren
2026-09-21 10:58 ` Tony Lindgren
2026-09-22 3:57 ` Kishen Maloor
2026-09-22 5:25 ` Tony Lindgren
2026-09-23 0:38 ` Kishen Maloor
2026-09-23 6:04 ` Tony Lindgren
2026-09-24 5:53 ` Kishen Maloor
2026-09-24 6:59 ` Tony Lindgren
2026-09-18 4:33 ` Kishen Maloor
2026-09-21 5:58 ` Tony Lindgren
2026-09-21 6:56 ` Tony Lindgren
2026-09-22 3:56 ` Kishen Maloor
2026-09-22 6:27 ` Tony Lindgren
2026-09-23 0:37 ` Kishen Maloor
2026-09-23 6:50 ` Tony Lindgren
2026-09-24 5:34 ` Kishen Maloor
2026-09-24 7:15 ` Tony Lindgren
2026-10-08 9:22 ` Tony Lindgren
2026-08-31 7:13 ` [RFC PATCH v2 3/4] KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY Tony Lindgren
2026-08-31 7:23 ` sashiko-bot
2026-09-01 6:10 ` Tony Lindgren
2026-08-31 7:13 ` [RFC PATCH v2 4/4] KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU Tony Lindgren
2026-08-31 7:23 ` sashiko-bot
2026-09-01 6:12 ` Tony Lindgren
2026-09-04 18:24 ` [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Artem Bityutskiy
2026-09-17 21:27 ` Peter Xu
2026-09-18 12:46 ` Artem Bityutskiy
2026-09-18 15:53 ` Peter Xu
2026-09-22 8:09 ` Artem Bityutskiy
2026-09-22 9:42 ` Tony Lindgren
2026-09-22 11:54 ` Artem Bityutskiy
2026-09-23 4:20 ` Tony Lindgren
2026-09-22 21:18 ` Peter Xu
2026-09-23 12:05 ` Artem Bityutskiy
2026-09-24 21:19 ` Peter Xu
2026-09-28 14:15 ` Artem Bityutskiy
2026-09-29 21:05 ` Peter Xu
2026-10-02 19:57 ` Artem Bityutskiy
2026-10-07 20:00 ` Peter Xu
2026-09-23 15:28 ` Serge Hallyn (AMD)
2026-09-20 23:56 ` Kishen Maloor
2026-09-23 21:36 ` Peter Xu
2026-09-24 4:27 ` Kishen Maloor [this message]
2026-09-25 14:18 ` Peter Xu
2026-09-29 1:28 ` Kishen Maloor
2026-09-30 20:42 ` Peter Xu
2026-10-07 4:27 ` Kishen Maloor
2026-10-07 20:13 ` Peter Xu
2026-10-08 6:23 ` Tony Lindgren
2026-10-08 14:31 ` Peter Xu
2026-09-18 18:36 ` Ionut Mihalcea
2026-09-21 4:35 ` Tony Lindgren
2026-09-25 16:03 ` Serge Hallyn (AMD)
2026-09-28 3:24 ` Kishen Maloor
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=9fa229ea-818c-4ba9-81c6-be7ead5cb60a@intel.com \
--to=kishen.maloor@intel.com \
--cc=Jon.Grimm@amd.com \
--cc=anup@brainfault.org \
--cc=dedekind1@gmail.com \
--cc=elena.reshetova@intel.com \
--cc=farosas@suse.de \
--cc=jakub.ruzicka@matfyz.cz \
--cc=joro@8bytes.org \
--cc=kai.huang@intel.com \
--cc=kvm@vger.kernel.org \
--cc=maz@kernel.org \
--cc=mika.westerberg@linux.intel.com \
--cc=oliver.upton@linux.dev \
--cc=pankaj.gupta@amd.com \
--cc=pbonzini@redhat.com \
--cc=peter.fang@intel.com \
--cc=peterx@redhat.com \
--cc=rick.p.edgecombe@intel.com \
--cc=sameo@rivosinc.com \
--cc=seanjc@google.com \
--cc=steven.price@arm.com \
--cc=thomas.lendacky@amd.com \
--cc=tony.lindgren@linux.intel.com \
--cc=vannapurve@google.com \
--cc=xiaoyao.li@intel.com \
--cc=yilun.xu@linux.intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox