Kernel KVM virtualization development
 help / color / mirror / Atom feed
From: Peter Xu <peterx@redhat.com>
To: Artem Bityutskiy <dedekind1@gmail.com>
Cc: "Tony Lindgren" <tony.lindgren@linux.intel.com>,
	"Paolo Bonzini" <pbonzini@redhat.com>,
	"Sean Christopherson" <seanjc@google.com>,
	"Fabiano Rosas" <farosas@suse.de>,
	"Jon Grimm" <Jon.Grimm@amd.com>,
	"Pankaj Gupta" <pankaj.gupta@amd.com>,
	"Tom Lendacky" <thomas.lendacky@amd.com>,
	"Marc Zyngier" <maz@kernel.org>,
	"Oliver Upton" <oliver.upton@linux.dev>,
	"Steven Price" <steven.price@arm.com>,
	"Anup Patel" <anup@brainfault.org>,
	"Samuel Ortiz" <sameo@rivosinc.com>,
	"Jakub Růžička" <jakub.ruzicka@matfyz.cz>,
	"Jörg Rödel" <joro@8bytes.org>,
	"Vishal Annapurve" <vannapurve@google.com>,
	"Elena Reshetova" <elena.reshetova@intel.com>,
	"Kai Huang" <kai.huang@intel.com>,
	"Kishen Maloor" <kishen.maloor@intel.com>,
	"Mika Westerberg" <mika.westerberg@linux.intel.com>,
	"Peter Fang" <peter.fang@intel.com>,
	"Rick Edgecombe" <rick.p.edgecombe@intel.com>,
	"Xiaoyao Li" <xiaoyao.li@intel.com>,
	"Xu Yilun" <yilun.xu@linux.intel.com>,
	kvm@vger.kernel.org
Subject: Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
Date: Wed, 7 Oct 2026 16:00:08 -0400	[thread overview]
Message-ID: <asakyIs1beS_I-o_@zhexu-thinkpadt14gen5.rmtcaon.csb> (raw)
In-Reply-To: <c7d5adf38d6338857f73bb4b3f014a500b4c2fe1.camel@gmail.com>

On Fri, Oct 02, 2026 at 10:57:46PM +0300, Artem Bityutskiy wrote:
> On Tue, 2026-09-29 at 17:05 -0400, Peter Xu wrote:
> > > Here is how I saw this, but I may be missing something (my excuse is that I
> > > am still new to the team and still learning).
> > > 
> > > 1. QEMU has a bitmap of shared pages in RAMBlockAttributes, so it can
> > >    distinguish shared pages.
> > > 2. In general, QEMU does not distinguish private vs unaccepted pages, so
> > >    unaccepted pages are treated as private pages.
> > 
> > Yes, the latter seems uncontroversial.
> > 
> > The 1st one is true, and it just reminded me if the conversion is
> > synchronous and one step requires the hypercall to QEMU, then indeed
> > background conversion can be avoided by some form of userspace locking.
> > 
> > Perhaps, a rwlock suites, each vCPU takes it for write whenever page
> > conversion requested from the guest (private <-> shared; nothing about
> > "accepted" that matters).  Then the migration threads, one or multiple,
> > take the read lock, lookup the bit, do MEM.EXPORT, unlock.
> > 
> > Then it seems fine in general, except that I donno if things can still go
> > wrong when there are multiple versions of "if this page is private or
> > shared".  Say, minimum of three?
> > 
> >   (a) QEMU maintains the bitmap in RAMBlockAttributes, each bit represents
> >       if the page is shared or private
> > 
> >   (b) KVM should maintain one, looks to me, kvm->mem_attr_array
> > 
> >   (c) Hardware / Firmware may maintain its own, in case of TDX, is that one
> >       bit on the SEPT pgtable?
> > 
> > They don't change together, AFAIU, they change in order, I believe
> > (c)->(a)->(b) if my above understanding is correct.
> 
> It looks like in terms of which layer saves the page type change first:
> 
> - Private->Shared: TDX -> KVM  -> QEMU
> - Shared->Private: KVM -> QEMU -> TDX
> 
> In both cases the TD initiates the change. This ends up with a TD exit,
> followed by KVM exiting to QEMU. QEMU calls kvm_convert_memory(), which
> calls back into KVM (KVM_SET_MEMORY_ATTRIBUTES). At this point SEPT did not
> change yet.
> 
> Private->Shared:
> - KVM first removes the page from SEPT, so TDX sees the change first.
> - KVM updates own data (kvm->mem_attr_array). So KVM "gets" the change
>   second.
> - QEMU updates RAMBlockAttributes, so QEMU "gets" the change last.
> 
> Shared->Private:
> - KVM removes the page from the shared EPT, updates own data
>   (kvm->mem_attr_array).
> - QEMU updates RAMBlockAttributes, discards backing storage.
> - Back to TD, which accepts the page. This causes an EPT violation, and
>   KVM adds the page to SEPT (TDH.MEM.PAGE.AUG).

Oh, this reminded me that ram_block_attributes_state_change() is done after
the ioctl(KVM_SET_MEMORY_ATTRIBUTES); I believe I didn't notice this detail
and assumed the other way round.  In that case, yes, KVM's page status will
always change before QEMU's.

> 
> > Then, what if they report different things?
> > 
> > Say, during migration the guest wants to convert a page from shared to
> > private.  (c) can be already done saying one page "private" now for TDX,
> > (a) tries to mark it "private" too, but now assuming page being accessed
> > (read lock held), it may be trying to take a write lock and sleep, which
> > means (b) will be "shared" so far.
> > 
> > So what happens is, QEMU thinks this page "shared" because the conversion
> > hasn't take place waiting for the write lock, however at least TDX may
> > think it already "private" instead.
> > 
> > Then QEMU logically can access HVA of that page, with (a)=shared,
> > (b)=shared, (c)=private.
> > 
> > Would it cause trouble?
> 
> (c) should see it as private only after KVM and QEMU do.
> 
> But I am not sure about the entire idea. Holding the read lock around the
> export ioctl means that a vCPU requesting a conversion waits until the
> export finishes. One export call may cover many MiB of crypto work, so the
> vCPU stalling may be significant, right?
> 
> Let's check the 2 cases.
> 
> Private -> Shared
> 
> QEMU calls the export ioctl for a page that it thinks is private, but
> meanwhile it became shared. In this case, if the semantics of the export
> ioctl is that such pages are skipped, we should be fine, right? QEMU will
> just handle this page during the next round.

Yes, this should be benign as long as conversion of page status set the
dirty bit; QEMU guarantees to clear the D-bit before reading the page, then
it must read it again later.

> 
> Shared -> Private
> 
> QEMU tries to migrate a shared page, which meanwhile became private. Reading
> it would result in zeros or some stale data, right? Would it help if QEMU
> used a lock-check_if_still_shared-copy-release, and the same lock around
> kvm_convert_memory()?

AFAIU, this should SIGBUS QEMU, which is the major issue, that should be
what happens in current linux tree, where now only INIT_SHARED is available
for fault() processing.

I also think that's the plan for afterwards, but still worth check kernel
tree with TDX migration integrated: kvm_gmem_fault_user_mapping() should
decide what happens..

-- 
Peter Xu


  reply	other threads:[~2026-10-07 20:00 UTC|newest]

Thread overview: 87+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-31  7:13 [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Tony Lindgren
2026-08-31  7:13 ` [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests Tony Lindgren
2026-08-31  7:20   ` sashiko-bot
2026-09-18 11:35   ` Peter Xu
2026-09-21  4:20     ` Tony Lindgren
2026-09-24  1:50     ` Wei Wang
2026-09-24  4:51       ` Tony Lindgren
2026-08-31  7:13 ` [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD Tony Lindgren
2026-08-31  7:23   ` sashiko-bot
2026-09-01  6:03     ` Tony Lindgren
2026-09-07 11:53   ` Tony Lindgren
2026-09-07 13:15     ` Jörg Rödel
2026-09-07 13:32       ` Artem Bityutskiy
2026-09-08  4:15         ` Tony Lindgren
2026-09-08  4:43         ` Tony Lindgren
2026-09-09  0:22           ` Kishen Maloor
2026-09-09  6:57             ` Tony Lindgren
2026-09-10  1:11               ` Kishen Maloor
2026-09-10  6:33                 ` Tony Lindgren
2026-09-11  1:40                   ` Kishen Maloor
2026-09-11  4:23                     ` Tony Lindgren
2026-09-15  0:14                       ` Kishen Maloor
2026-09-15  4:44                         ` Tony Lindgren
2026-09-15 15:53                           ` Kishen Maloor
2026-09-16  5:09                             ` Tony Lindgren
2026-09-17  3:31                               ` Kishen Maloor
2026-09-17  6:42                                 ` Tony Lindgren
2026-09-18  4:32                                   ` Kishen Maloor
2026-09-18  5:58                                     ` Tony Lindgren
2026-09-21  0:13                                       ` Kishen Maloor
2026-09-21  6:52                                         ` Tony Lindgren
2026-09-21  9:24                                           ` Tony Lindgren
2026-09-21 10:58                                             ` Tony Lindgren
2026-09-22  3:57                                           ` Kishen Maloor
2026-09-22  5:25                                             ` Tony Lindgren
2026-09-23  0:38                                               ` Kishen Maloor
2026-09-23  6:04                                                 ` Tony Lindgren
2026-09-24  5:53                                                   ` Kishen Maloor
2026-09-24  6:59                                                     ` Tony Lindgren
2026-09-18  4:33   ` Kishen Maloor
2026-09-21  5:58     ` Tony Lindgren
2026-09-21  6:56       ` Tony Lindgren
2026-09-22  3:56         ` Kishen Maloor
2026-09-22  6:27           ` Tony Lindgren
2026-09-23  0:37             ` Kishen Maloor
2026-09-23  6:50               ` Tony Lindgren
2026-09-24  5:34                 ` Kishen Maloor
2026-09-24  7:15                   ` Tony Lindgren
2026-10-08  9:22                     ` Tony Lindgren
2026-08-31  7:13 ` [RFC PATCH v2 3/4] KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY Tony Lindgren
2026-08-31  7:23   ` sashiko-bot
2026-09-01  6:10     ` Tony Lindgren
2026-08-31  7:13 ` [RFC PATCH v2 4/4] KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU Tony Lindgren
2026-08-31  7:23   ` sashiko-bot
2026-09-01  6:12     ` Tony Lindgren
2026-09-04 18:24 ` [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Artem Bityutskiy
2026-09-17 21:27   ` Peter Xu
2026-09-18 12:46     ` Artem Bityutskiy
2026-09-18 15:53       ` Peter Xu
2026-09-22  8:09         ` Artem Bityutskiy
2026-09-22  9:42           ` Tony Lindgren
2026-09-22 11:54             ` Artem Bityutskiy
2026-09-23  4:20               ` Tony Lindgren
2026-09-22 21:18           ` Peter Xu
2026-09-23 12:05             ` Artem Bityutskiy
2026-09-24 21:19               ` Peter Xu
2026-09-28 14:15                 ` Artem Bityutskiy
2026-09-29 21:05                   ` Peter Xu
2026-10-02 19:57                     ` Artem Bityutskiy
2026-10-07 20:00                       ` Peter Xu [this message]
2026-09-23 15:28           ` Serge Hallyn (AMD)
2026-10-09  6:11       ` Tony Lindgren
2026-09-20 23:56     ` Kishen Maloor
2026-09-23 21:36       ` Peter Xu
2026-09-24  4:27         ` Kishen Maloor
2026-09-25 14:18           ` Peter Xu
2026-09-29  1:28             ` Kishen Maloor
2026-09-30 20:42               ` Peter Xu
2026-10-07  4:27                 ` Kishen Maloor
2026-10-07 20:13                   ` Peter Xu
2026-10-08  6:23                     ` Tony Lindgren
2026-10-08 14:31                       ` Peter Xu
2026-10-09  4:29                         ` Tony Lindgren
2026-09-18 18:36 ` Ionut Mihalcea
2026-09-21  4:35   ` Tony Lindgren
2026-09-25 16:03 ` Serge Hallyn (AMD)
2026-09-28  3:24   ` Kishen Maloor

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=asakyIs1beS_I-o_@zhexu-thinkpadt14gen5.rmtcaon.csb \
    --to=peterx@redhat.com \
    --cc=Jon.Grimm@amd.com \
    --cc=anup@brainfault.org \
    --cc=dedekind1@gmail.com \
    --cc=elena.reshetova@intel.com \
    --cc=farosas@suse.de \
    --cc=jakub.ruzicka@matfyz.cz \
    --cc=joro@8bytes.org \
    --cc=kai.huang@intel.com \
    --cc=kishen.maloor@intel.com \
    --cc=kvm@vger.kernel.org \
    --cc=maz@kernel.org \
    --cc=mika.westerberg@linux.intel.com \
    --cc=oliver.upton@linux.dev \
    --cc=pankaj.gupta@amd.com \
    --cc=pbonzini@redhat.com \
    --cc=peter.fang@intel.com \
    --cc=rick.p.edgecombe@intel.com \
    --cc=sameo@rivosinc.com \
    --cc=seanjc@google.com \
    --cc=steven.price@arm.com \
    --cc=thomas.lendacky@amd.com \
    --cc=tony.lindgren@linux.intel.com \
    --cc=vannapurve@google.com \
    --cc=xiaoyao.li@intel.com \
    --cc=yilun.xu@linux.intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox