From: Peter Xu <peterx@redhat.com>
To: Artem Bityutskiy <dedekind1@gmail.com>
Cc: "Tony Lindgren" <tony.lindgren@linux.intel.com>,
"Paolo Bonzini" <pbonzini@redhat.com>,
"Sean Christopherson" <seanjc@google.com>,
"Fabiano Rosas" <farosas@suse.de>,
"Jon Grimm" <Jon.Grimm@amd.com>,
"Pankaj Gupta" <pankaj.gupta@amd.com>,
"Tom Lendacky" <thomas.lendacky@amd.com>,
"Marc Zyngier" <maz@kernel.org>,
"Oliver Upton" <oliver.upton@linux.dev>,
"Steven Price" <steven.price@arm.com>,
"Anup Patel" <anup@brainfault.org>,
"Samuel Ortiz" <sameo@rivosinc.com>,
"Jakub Růžička" <jakub.ruzicka@matfyz.cz>,
"Jörg Rödel" <joro@8bytes.org>,
"Vishal Annapurve" <vannapurve@google.com>,
"Elena Reshetova" <elena.reshetova@intel.com>,
"Kai Huang" <kai.huang@intel.com>,
"Kishen Maloor" <kishen.maloor@intel.com>,
"Mika Westerberg" <mika.westerberg@linux.intel.com>,
"Peter Fang" <peter.fang@intel.com>,
"Rick Edgecombe" <rick.p.edgecombe@intel.com>,
"Xiaoyao Li" <xiaoyao.li@intel.com>,
"Xu Yilun" <yilun.xu@linux.intel.com>,
kvm@vger.kernel.org
Subject: Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
Date: Wed, 7 Oct 2026 16:00:08 -0400 [thread overview]
Message-ID: <asakyIs1beS_I-o_@zhexu-thinkpadt14gen5.rmtcaon.csb> (raw)
In-Reply-To: <c7d5adf38d6338857f73bb4b3f014a500b4c2fe1.camel@gmail.com>
On Fri, Oct 02, 2026 at 10:57:46PM +0300, Artem Bityutskiy wrote:
> On Tue, 2026-09-29 at 17:05 -0400, Peter Xu wrote:
> > > Here is how I saw this, but I may be missing something (my excuse is that I
> > > am still new to the team and still learning).
> > >
> > > 1. QEMU has a bitmap of shared pages in RAMBlockAttributes, so it can
> > > distinguish shared pages.
> > > 2. In general, QEMU does not distinguish private vs unaccepted pages, so
> > > unaccepted pages are treated as private pages.
> >
> > Yes, the latter seems uncontroversial.
> >
> > The 1st one is true, and it just reminded me if the conversion is
> > synchronous and one step requires the hypercall to QEMU, then indeed
> > background conversion can be avoided by some form of userspace locking.
> >
> > Perhaps, a rwlock suites, each vCPU takes it for write whenever page
> > conversion requested from the guest (private <-> shared; nothing about
> > "accepted" that matters). Then the migration threads, one or multiple,
> > take the read lock, lookup the bit, do MEM.EXPORT, unlock.
> >
> > Then it seems fine in general, except that I donno if things can still go
> > wrong when there are multiple versions of "if this page is private or
> > shared". Say, minimum of three?
> >
> > (a) QEMU maintains the bitmap in RAMBlockAttributes, each bit represents
> > if the page is shared or private
> >
> > (b) KVM should maintain one, looks to me, kvm->mem_attr_array
> >
> > (c) Hardware / Firmware may maintain its own, in case of TDX, is that one
> > bit on the SEPT pgtable?
> >
> > They don't change together, AFAIU, they change in order, I believe
> > (c)->(a)->(b) if my above understanding is correct.
>
> It looks like in terms of which layer saves the page type change first:
>
> - Private->Shared: TDX -> KVM -> QEMU
> - Shared->Private: KVM -> QEMU -> TDX
>
> In both cases the TD initiates the change. This ends up with a TD exit,
> followed by KVM exiting to QEMU. QEMU calls kvm_convert_memory(), which
> calls back into KVM (KVM_SET_MEMORY_ATTRIBUTES). At this point SEPT did not
> change yet.
>
> Private->Shared:
> - KVM first removes the page from SEPT, so TDX sees the change first.
> - KVM updates own data (kvm->mem_attr_array). So KVM "gets" the change
> second.
> - QEMU updates RAMBlockAttributes, so QEMU "gets" the change last.
>
> Shared->Private:
> - KVM removes the page from the shared EPT, updates own data
> (kvm->mem_attr_array).
> - QEMU updates RAMBlockAttributes, discards backing storage.
> - Back to TD, which accepts the page. This causes an EPT violation, and
> KVM adds the page to SEPT (TDH.MEM.PAGE.AUG).
Oh, this reminded me that ram_block_attributes_state_change() is done after
the ioctl(KVM_SET_MEMORY_ATTRIBUTES); I believe I didn't notice this detail
and assumed the other way round. In that case, yes, KVM's page status will
always change before QEMU's.
>
> > Then, what if they report different things?
> >
> > Say, during migration the guest wants to convert a page from shared to
> > private. (c) can be already done saying one page "private" now for TDX,
> > (a) tries to mark it "private" too, but now assuming page being accessed
> > (read lock held), it may be trying to take a write lock and sleep, which
> > means (b) will be "shared" so far.
> >
> > So what happens is, QEMU thinks this page "shared" because the conversion
> > hasn't take place waiting for the write lock, however at least TDX may
> > think it already "private" instead.
> >
> > Then QEMU logically can access HVA of that page, with (a)=shared,
> > (b)=shared, (c)=private.
> >
> > Would it cause trouble?
>
> (c) should see it as private only after KVM and QEMU do.
>
> But I am not sure about the entire idea. Holding the read lock around the
> export ioctl means that a vCPU requesting a conversion waits until the
> export finishes. One export call may cover many MiB of crypto work, so the
> vCPU stalling may be significant, right?
>
> Let's check the 2 cases.
>
> Private -> Shared
>
> QEMU calls the export ioctl for a page that it thinks is private, but
> meanwhile it became shared. In this case, if the semantics of the export
> ioctl is that such pages are skipped, we should be fine, right? QEMU will
> just handle this page during the next round.
Yes, this should be benign as long as conversion of page status set the
dirty bit; QEMU guarantees to clear the D-bit before reading the page, then
it must read it again later.
>
> Shared -> Private
>
> QEMU tries to migrate a shared page, which meanwhile became private. Reading
> it would result in zeros or some stale data, right? Would it help if QEMU
> used a lock-check_if_still_shared-copy-release, and the same lock around
> kvm_convert_memory()?
AFAIU, this should SIGBUS QEMU, which is the major issue, that should be
what happens in current linux tree, where now only INIT_SHARED is available
for fault() processing.
I also think that's the plan for afterwards, but still worth check kernel
tree with TDX migration integrated: kvm_gmem_fault_user_mapping() should
decide what happens..
--
Peter Xu
next prev parent reply other threads:[~2026-10-07 20:00 UTC|newest]
Thread overview: 87+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-31 7:13 [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Tony Lindgren
2026-08-31 7:13 ` [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests Tony Lindgren
2026-08-31 7:20 ` sashiko-bot
2026-09-18 11:35 ` Peter Xu
2026-09-21 4:20 ` Tony Lindgren
2026-09-24 1:50 ` Wei Wang
2026-09-24 4:51 ` Tony Lindgren
2026-08-31 7:13 ` [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD Tony Lindgren
2026-08-31 7:23 ` sashiko-bot
2026-09-01 6:03 ` Tony Lindgren
2026-09-07 11:53 ` Tony Lindgren
2026-09-07 13:15 ` Jörg Rödel
2026-09-07 13:32 ` Artem Bityutskiy
2026-09-08 4:15 ` Tony Lindgren
2026-09-08 4:43 ` Tony Lindgren
2026-09-09 0:22 ` Kishen Maloor
2026-09-09 6:57 ` Tony Lindgren
2026-09-10 1:11 ` Kishen Maloor
2026-09-10 6:33 ` Tony Lindgren
2026-09-11 1:40 ` Kishen Maloor
2026-09-11 4:23 ` Tony Lindgren
2026-09-15 0:14 ` Kishen Maloor
2026-09-15 4:44 ` Tony Lindgren
2026-09-15 15:53 ` Kishen Maloor
2026-09-16 5:09 ` Tony Lindgren
2026-09-17 3:31 ` Kishen Maloor
2026-09-17 6:42 ` Tony Lindgren
2026-09-18 4:32 ` Kishen Maloor
2026-09-18 5:58 ` Tony Lindgren
2026-09-21 0:13 ` Kishen Maloor
2026-09-21 6:52 ` Tony Lindgren
2026-09-21 9:24 ` Tony Lindgren
2026-09-21 10:58 ` Tony Lindgren
2026-09-22 3:57 ` Kishen Maloor
2026-09-22 5:25 ` Tony Lindgren
2026-09-23 0:38 ` Kishen Maloor
2026-09-23 6:04 ` Tony Lindgren
2026-09-24 5:53 ` Kishen Maloor
2026-09-24 6:59 ` Tony Lindgren
2026-09-18 4:33 ` Kishen Maloor
2026-09-21 5:58 ` Tony Lindgren
2026-09-21 6:56 ` Tony Lindgren
2026-09-22 3:56 ` Kishen Maloor
2026-09-22 6:27 ` Tony Lindgren
2026-09-23 0:37 ` Kishen Maloor
2026-09-23 6:50 ` Tony Lindgren
2026-09-24 5:34 ` Kishen Maloor
2026-09-24 7:15 ` Tony Lindgren
2026-10-08 9:22 ` Tony Lindgren
2026-08-31 7:13 ` [RFC PATCH v2 3/4] KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY Tony Lindgren
2026-08-31 7:23 ` sashiko-bot
2026-09-01 6:10 ` Tony Lindgren
2026-08-31 7:13 ` [RFC PATCH v2 4/4] KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU Tony Lindgren
2026-08-31 7:23 ` sashiko-bot
2026-09-01 6:12 ` Tony Lindgren
2026-09-04 18:24 ` [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Artem Bityutskiy
2026-09-17 21:27 ` Peter Xu
2026-09-18 12:46 ` Artem Bityutskiy
2026-09-18 15:53 ` Peter Xu
2026-09-22 8:09 ` Artem Bityutskiy
2026-09-22 9:42 ` Tony Lindgren
2026-09-22 11:54 ` Artem Bityutskiy
2026-09-23 4:20 ` Tony Lindgren
2026-09-22 21:18 ` Peter Xu
2026-09-23 12:05 ` Artem Bityutskiy
2026-09-24 21:19 ` Peter Xu
2026-09-28 14:15 ` Artem Bityutskiy
2026-09-29 21:05 ` Peter Xu
2026-10-02 19:57 ` Artem Bityutskiy
2026-10-07 20:00 ` Peter Xu [this message]
2026-09-23 15:28 ` Serge Hallyn (AMD)
2026-10-09 6:11 ` Tony Lindgren
2026-09-20 23:56 ` Kishen Maloor
2026-09-23 21:36 ` Peter Xu
2026-09-24 4:27 ` Kishen Maloor
2026-09-25 14:18 ` Peter Xu
2026-09-29 1:28 ` Kishen Maloor
2026-09-30 20:42 ` Peter Xu
2026-10-07 4:27 ` Kishen Maloor
2026-10-07 20:13 ` Peter Xu
2026-10-08 6:23 ` Tony Lindgren
2026-10-08 14:31 ` Peter Xu
2026-10-09 4:29 ` Tony Lindgren
2026-09-18 18:36 ` Ionut Mihalcea
2026-09-21 4:35 ` Tony Lindgren
2026-09-25 16:03 ` Serge Hallyn (AMD)
2026-09-28 3:24 ` Kishen Maloor
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=asakyIs1beS_I-o_@zhexu-thinkpadt14gen5.rmtcaon.csb \
--to=peterx@redhat.com \
--cc=Jon.Grimm@amd.com \
--cc=anup@brainfault.org \
--cc=dedekind1@gmail.com \
--cc=elena.reshetova@intel.com \
--cc=farosas@suse.de \
--cc=jakub.ruzicka@matfyz.cz \
--cc=joro@8bytes.org \
--cc=kai.huang@intel.com \
--cc=kishen.maloor@intel.com \
--cc=kvm@vger.kernel.org \
--cc=maz@kernel.org \
--cc=mika.westerberg@linux.intel.com \
--cc=oliver.upton@linux.dev \
--cc=pankaj.gupta@amd.com \
--cc=pbonzini@redhat.com \
--cc=peter.fang@intel.com \
--cc=rick.p.edgecombe@intel.com \
--cc=sameo@rivosinc.com \
--cc=seanjc@google.com \
--cc=steven.price@arm.com \
--cc=thomas.lendacky@amd.com \
--cc=tony.lindgren@linux.intel.com \
--cc=vannapurve@google.com \
--cc=xiaoyao.li@intel.com \
--cc=yilun.xu@linux.intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox