Kernel KVM virtualization development
 help / color / mirror / Atom feed
From: Artem Bityutskiy <dedekind1@gmail.com>
To: Peter Xu <peterx@redhat.com>
Cc: "Tony Lindgren" <tony.lindgren@linux.intel.com>,
	"Paolo Bonzini" <pbonzini@redhat.com>,
	"Sean Christopherson" <seanjc@google.com>,
	"Fabiano Rosas" <farosas@suse.de>,
	"Jon Grimm" <Jon.Grimm@amd.com>,
	"Pankaj Gupta" <pankaj.gupta@amd.com>,
	"Tom Lendacky" <thomas.lendacky@amd.com>,
	"Marc Zyngier" <maz@kernel.org>,
	"Oliver Upton" <oliver.upton@linux.dev>,
	"Steven Price" <steven.price@arm.com>,
	"Anup Patel" <anup@brainfault.org>,
	"Samuel Ortiz" <sameo@rivosinc.com>,
	"Jakub Růžička" <jakub.ruzicka@matfyz.cz>,
	"Jörg Rödel" <joro@8bytes.org>,
	"Vishal Annapurve" <vannapurve@google.com>,
	"Elena Reshetova" <elena.reshetova@intel.com>,
	"Kai Huang" <kai.huang@intel.com>,
	"Kishen Maloor" <kishen.maloor@intel.com>,
	"Mika Westerberg" <mika.westerberg@linux.intel.com>,
	"Peter Fang" <peter.fang@intel.com>,
	"Rick Edgecombe" <rick.p.edgecombe@intel.com>,
	"Xiaoyao Li" <xiaoyao.li@intel.com>,
	"Xu Yilun" <yilun.xu@linux.intel.com>,
	kvm@vger.kernel.org
Subject: Re: [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration
Date: Mon, 28 Sep 2026 17:15:00 +0300	[thread overview]
Message-ID: <59384511c6070abfd048b37f5ec2831e5f8bb715.camel@gmail.com> (raw)
In-Reply-To: <arWT9fwpCkFneDKn@zhexu-thinkpadt14gen5.rmtcaon.csb>

Hi Peter, thanks for reply again.

On Thu, 2026-09-24 at 17:19 -0400, Peter Xu wrote:
> > > > The reason I am asking is that my assumption was that it is not important.
> > > > But if it is, I will come back to the TDX module architects with a
> > > > request to revise the design to support independent dirty page tracking. Of
> > > > course they may have some reasons for not doing it, but I would try at
> > > > least.
> > > 
> > > Thanks, I'll talk to our team and revisit this after I collect answers.
> > 
> > Many thanks!
> 
> I got some feedback on this, I'll try to provide a summary.
> 
> So, first of all, calc_dirty_rate isn't seem to be widely used across our
> customers.
> 
> However, we do have customer case using calc_dirty_rate to evaluate
> migrations of a VM fleet for cases like from one data centre to another.
> 
> I think it makes sense because the normal "try to migrate and fallback
> otherwise" idea applies well to one VM, but perhaps not that good on a
> fleet.

Yes, that makes sense.

> When a fleet is involved, we don't want to migrate 400 VMs then found
> there're 30 critical VMs too busy and can't migrate, then due to whatever
> reason (inter-VM communication / service locality ?) one is forced to
> migrate that 400 VMs backwards.
> 
> IOW, it seems helpful to provide high-level evalutions of migration
> decisions over a full cluster, concurrently and efficiently.

Thank you. I'll work on this internally. It will take time.

Just to give wider context: Sean and Paolo gave us feedback regarding the
entire SEPT scan approach - they believe it is too costly and won't scale,
and suggested using PML instead. For now, dirty scanning is the best we
have, but I continue to explore other options internally, and standalone
dirty tracking is one of them.

> > AFAIU, the TDX guest migration implementation benchmarking results are
> > satisfactory with the scanning approach, but I do not have hard numbers to
> > share. I also feel that PML could offer better performance, at least for
> > memory-intensive workloads. But this is intuition only.	
> 
> My gut feeling is DMA shouldn't be a blocker for PML: AFAIU we don't track
> DMA from KVM side. Assigned device should have its own dirty tracking for
> DMAs, either via device's own tracking facilities, or the IOMMU on the
> host.  Feel free to refer to vfio_listener_log_sync() in QEMU.  In all
> cases, it'll be great you could share the reason if you have more solid
> clues.

Yeah, this is also part of the internal exploration I mentioned above too.

> > 
> Are we talking about the memory export/import API or GET_DIRTY_LOG?  IIUC,
> GET_DIRTY_LOG always applies to a whole memslot,
> 
> struct kvm_dirty_log {
> 	__u32 slot;
> 	__u32 padding1;
> 	union {
> 		void *dirty_bitmap; /* one bit per page */
> 		__u64 padding2;
> 	};
> };

I apologize, I worte something unrelated to the context. 512 GPAs at a time
is the limit for exporting the memory, not for dirty scanning.

For the dirty scanning seamcall (TDH.MEM.SCAN.RANGE), the limit is 512 * 512
GPAs at a time, which is 262,144 GPAs, or 1GiB. So if a memslot is larger
than 1GiB, multiple calls to the dirty scanning seamcall are needed.

> > Keep in mind that dirty pages are the majority of migration candidates, but
> > not all of them. Sometimes a migration candidate can be a page that was
> > already exported, but then was, for example, converted from private to
> > shared, or unaccepted by the TD (gone, in other words). In this case the TDX
> > module flags it as a migration candidate too. The memory export seamcall
> > treats it differently too - instead of exporting encrypted page data, it
> > exports a small record indicating that the page has changed its status
> > (gone).
> 
> Yes, it makes sense.
> 
> So can I inteprete this as GET_DIRTY_LOG works seamlessly for both private
> and shared pages (or even, unaccepted pages)?
> 
> Then I assume it means MEMORY.EXPORT should also be able to read shared or
> unaccepted pages too, am I right?  Same to when apply with IMPORT.  Another
> counter example is MEMORY.EXPORT returns a flag saying "this page is
> shared, go read it directly from HVA", but then QEMU reading it may race
> with a concurrent shared->private conversion crashing VMM.
> 
> Looks to me MEMORY.EXPORT must support shared too, then.

Hmm... First of all, it does sound like a possible approach. But it is not
the approach we took in our PoC today. I hope Kishen will chime in to
correct me.

Here is how I saw this, but I may be missing something (my excuse is that I
am still new to the team and still learning).

1. QEMU has a bitmap of shared pages in RAMBlockAttributes, so it can
   distinguish shared pages.
2. In general, QEMU does not distinguish private vs unaccepted pages, so
   unaccepted pages are treated as private pages.

Dirty tracking:

- QEMU uses the same KVM_DIRTY_LOG mechanism for tracking shared, private,
  and unaccepted pages.
- For shared GFNs, KVM uses the normal VM dirty tracking mechanism. For
  private and unaccepted GFNs, KVM goes to the TDX-specific code.
- But the final bitmap that QEMU sees covers all page types.
- KVM calls TDH.MEM.SCAN.RANGE on both private and unaccepted GFNs.
- For private GFNs, the TDX module reports it as a migration candidate if
  its data changed or its status changed (e.g., converted to shared or
  unaccepted).
- For unaccepted GFNs, the TDX module reports it as a migration candidate in
  the first round (so it appears as dirty in KVM_DIRTY_LOG reply). Then it
  reports it as clean, unless its status changes - it becomes accepted.

Page export:

- QEMU migrates shared pages the old way - it does not try to use the
  proposed CoCo migration uAPI for that.
- For private and unaccepted pages, QEMU uses the CoCo migration uAPI.
- Our export uAPI PoC implementation does not try to check GFN type - it
  just   feeds them all to the TDX module TDH.EXPORT.MEM seamcall. Here
  is what TDX module does depending on the page type:
  - Shared pages: just skip, no errors.
  - Private pages: export the data in encrypted form.
  - Unaccepted pages: export a small record telling that the page is
    unaccepted. This record should be delivered to the destination and
    imported there, just like private pages.

But clearly this is part of the uAPI contract that must be discussed and
made explicit. What I describe above is obviously our PoC implementation,
plus TDX module behavior details.

> > The same logic applies to a TDX guest using dirty scanning. Suppose the
> > final dirty scan is slower than a theoretical TDX PML-based approach would
> > have been. The prescan optimization I described in the previous e-mail
> > should help with this in an average case, but let's assume it does not
> > help for some special case - when the TD touches most of its memory, so
> > all EPT sub-trees end up touched. This should be rare, but let's assume it
> > happens.
> > 
> > In this case, whether that scan slowdown will actually matter for the
> > overall downtime also depends on the network. If the network is very fast,
> > the scan itself can become the dominant part of the downtime. If the
> > network is the slower part, the scan slowdown may barely matter.
> > 
> > So my understanding is that downtime is never fully predictable, for
> > either type of VM.
> 
> Right, but IMHO background scan of EPT pgtable dirty bits adds a completely
> new reason to introduce downtime, and when I said "unpredictable", it is
> about that part.  Also, I worry in some worst case this can be pretty large.

Just to clarify on the "background" part. Yes, it is "background" relative
to the TD - some CPU is running it, in parallel with vCPUs running on other
CPUs.

But there is no background activity in the TDX module itself. All the
seamcalls run synchronously on the CPU that invokes them.

This means, for example, that to speed up memory export, one can run the
TDH.EXPORT.MEM seamcall for different GPAs in parallel on different CPUs.

Same for dirty scanning - if one could run the TDH.MEM.SCAN.RANGE seamcall
for different GPAs in parallel on different CPUs, that would make scanning
much faster.

Also a bit separately, one point to keep in mind is that in CoCo the page
export and import are heavy, compute-intensive crypto operations. So when we
talk about slower scanning, we need to keep in mind that it is not that slow
in relation to the export crypto. But I understand that this is not an
apples-to-apples comparison:

- Scanning is potentially about a large SEPT in a VM with terabytes of
  memory.
- Exporting is only about the pages found to be dirty.

So this is only to remind that in the CoCo case, memory export also has a
high price tag, compared to a traditional VM.

> So we have two overheads here at this stage, unpredictable:
> 
>   (a) Scanning EPT pgtable, when very unlucky, can take a lot of time to
>       finally reports to a GET_DIRTY_LOG request,
> 
>   (b) Migrating of dirty pages during blackout phase, which should be
>       roughly linear to how many dirty pages we just collected.  (NOTE!  I
>       think we may have way to fix this (b) or optimize it.. but this is
>       off-topic; let's focus on the difference of (a) and (b) first)
> 
> When with PML, IIUC (a) is predictable: we have the bitmap on hand, plus a
> maximum of some (my memory is, 512?) PML entries to flush per vCPU.
> That'll be flushed automatically when we do vm_stop(), likely also
> concurrently, atomically updating the bitmaps.  I never measured it, but it
> is bounded, and sounds pretty fast.

Yes, I agree.

Now my secret desire is that TDX module can eventually plug PML under the
hood, consider it a "hardware accelerator" without changing the ABI. But I
do not know whether keeping the ABI unchanged is possible, or how soon it
could happen. This is something I am working on internally with Intel TDX
module team.

> When with scanning, (a) seems more unpredictable.  That's the part I was
> slightly concerned.  But now after thinking a bit more, it seems fine.
> Please read below.

Sure, thanks.

> > Intuitively, the TDX case does feel "less predictable". The open
> > question for me is whether the degree of unpredictability is large enough
> > to bother users. My attitude is to focus on getting something simple done
> > first, learn from real-world behavior, and improve it later if needed,
> > including exploring PML. The best is the enemy of the good sort of
> > attitude.
> 
> Yes, I think it's always fine we start with whatever is most feasible.
> 
> I think it actually may not be that bad. The last sync is special at least
> on how QEMU treats it, it should look like:
> 
>   - GET_DIRTY_LOG, to do last math, decide to switchover, <------   [1]
>   - vm_stop()
>   - GET_DIRTY_LOG, this collects all rest dirty bits      <------   [2]
>   - migrates the dirty pages, device states, etc.
> 
> So I expect there should be normally very small window between two
> continuous GET_DIRTY_LOG across system.  Only [2] will be part of downtime.
> 
> Since you explained to me on how the background rescan roughly works, by
> relying on A bit in pgtable directory entries, I do feel like in this case
> most of the memory regions shouldn't be accessed during small window of
> [1]->[2], then the range to scan should be very much under control too.  In
> reality, it will likely be even smaller, [1]->vm_stop(), because after that
> vCPUs are halted.

Yes, I agree. Just to flag the word "background" again, and to make sure we
are aligned - this "rescan" happens between [1] and [2]. The idea is that
[2] will be very fast after the "rescan". But it does increase the time
between [1] and [2], and the TD has time to dirty more pages.

> So it may not really be an issue in practise, but it still depends. In all
> cases, some measurements after PoC ready would be nice on some large and
> relatively busy VMs.

Yes, I agree.
> 

Thanks, Artem.

  reply	other threads:[~2026-09-28 14:15 UTC|newest]

Thread overview: 85+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-31  7:13 [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Tony Lindgren
2026-08-31  7:13 ` [RFC PATCH v2 1/4] Documentation: KVM: Add live migration API for confidential guests Tony Lindgren
2026-08-31  7:20   ` sashiko-bot
2026-09-18 11:35   ` Peter Xu
2026-09-21  4:20     ` Tony Lindgren
2026-09-24  1:50     ` Wei Wang
2026-09-24  4:51       ` Tony Lindgren
2026-08-31  7:13 ` [RFC PATCH v2 2/4] KVM: x86: Add optional KVM_CAP_LIVE_MIGRATION and KVM_MIGRATE_CMD Tony Lindgren
2026-08-31  7:23   ` sashiko-bot
2026-09-01  6:03     ` Tony Lindgren
2026-09-07 11:53   ` Tony Lindgren
2026-09-07 13:15     ` Jörg Rödel
2026-09-07 13:32       ` Artem Bityutskiy
2026-09-08  4:15         ` Tony Lindgren
2026-09-08  4:43         ` Tony Lindgren
2026-09-09  0:22           ` Kishen Maloor
2026-09-09  6:57             ` Tony Lindgren
2026-09-10  1:11               ` Kishen Maloor
2026-09-10  6:33                 ` Tony Lindgren
2026-09-11  1:40                   ` Kishen Maloor
2026-09-11  4:23                     ` Tony Lindgren
2026-09-15  0:14                       ` Kishen Maloor
2026-09-15  4:44                         ` Tony Lindgren
2026-09-15 15:53                           ` Kishen Maloor
2026-09-16  5:09                             ` Tony Lindgren
2026-09-17  3:31                               ` Kishen Maloor
2026-09-17  6:42                                 ` Tony Lindgren
2026-09-18  4:32                                   ` Kishen Maloor
2026-09-18  5:58                                     ` Tony Lindgren
2026-09-21  0:13                                       ` Kishen Maloor
2026-09-21  6:52                                         ` Tony Lindgren
2026-09-21  9:24                                           ` Tony Lindgren
2026-09-21 10:58                                             ` Tony Lindgren
2026-09-22  3:57                                           ` Kishen Maloor
2026-09-22  5:25                                             ` Tony Lindgren
2026-09-23  0:38                                               ` Kishen Maloor
2026-09-23  6:04                                                 ` Tony Lindgren
2026-09-24  5:53                                                   ` Kishen Maloor
2026-09-24  6:59                                                     ` Tony Lindgren
2026-09-18  4:33   ` Kishen Maloor
2026-09-21  5:58     ` Tony Lindgren
2026-09-21  6:56       ` Tony Lindgren
2026-09-22  3:56         ` Kishen Maloor
2026-09-22  6:27           ` Tony Lindgren
2026-09-23  0:37             ` Kishen Maloor
2026-09-23  6:50               ` Tony Lindgren
2026-09-24  5:34                 ` Kishen Maloor
2026-09-24  7:15                   ` Tony Lindgren
2026-10-08  9:22                     ` Tony Lindgren
2026-08-31  7:13 ` [RFC PATCH v2 3/4] KVM: x86: Add optional KVM_EXPORT_MEMORY and KVM_IMPORT_MEMORY Tony Lindgren
2026-08-31  7:23   ` sashiko-bot
2026-09-01  6:10     ` Tony Lindgren
2026-08-31  7:13 ` [RFC PATCH v2 4/4] KVM: x86: Add optional KVM_EXPORT_VCPU and KVM_IMPORT_VCPU Tony Lindgren
2026-08-31  7:23   ` sashiko-bot
2026-09-01  6:12     ` Tony Lindgren
2026-09-04 18:24 ` [RFC PATCH v2 0/4] Add KVM API for confidential guest live migration Artem Bityutskiy
2026-09-17 21:27   ` Peter Xu
2026-09-18 12:46     ` Artem Bityutskiy
2026-09-18 15:53       ` Peter Xu
2026-09-22  8:09         ` Artem Bityutskiy
2026-09-22  9:42           ` Tony Lindgren
2026-09-22 11:54             ` Artem Bityutskiy
2026-09-23  4:20               ` Tony Lindgren
2026-09-22 21:18           ` Peter Xu
2026-09-23 12:05             ` Artem Bityutskiy
2026-09-24 21:19               ` Peter Xu
2026-09-28 14:15                 ` Artem Bityutskiy [this message]
2026-09-29 21:05                   ` Peter Xu
2026-10-02 19:57                     ` Artem Bityutskiy
2026-10-07 20:00                       ` Peter Xu
2026-09-23 15:28           ` Serge Hallyn (AMD)
2026-09-20 23:56     ` Kishen Maloor
2026-09-23 21:36       ` Peter Xu
2026-09-24  4:27         ` Kishen Maloor
2026-09-25 14:18           ` Peter Xu
2026-09-29  1:28             ` Kishen Maloor
2026-09-30 20:42               ` Peter Xu
2026-10-07  4:27                 ` Kishen Maloor
2026-10-07 20:13                   ` Peter Xu
2026-10-08  6:23                     ` Tony Lindgren
2026-10-08 14:31                       ` Peter Xu
2026-09-18 18:36 ` Ionut Mihalcea
2026-09-21  4:35   ` Tony Lindgren
2026-09-25 16:03 ` Serge Hallyn (AMD)
2026-09-28  3:24   ` Kishen Maloor

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=59384511c6070abfd048b37f5ec2831e5f8bb715.camel@gmail.com \
    --to=dedekind1@gmail.com \
    --cc=Jon.Grimm@amd.com \
    --cc=anup@brainfault.org \
    --cc=elena.reshetova@intel.com \
    --cc=farosas@suse.de \
    --cc=jakub.ruzicka@matfyz.cz \
    --cc=joro@8bytes.org \
    --cc=kai.huang@intel.com \
    --cc=kishen.maloor@intel.com \
    --cc=kvm@vger.kernel.org \
    --cc=maz@kernel.org \
    --cc=mika.westerberg@linux.intel.com \
    --cc=oliver.upton@linux.dev \
    --cc=pankaj.gupta@amd.com \
    --cc=pbonzini@redhat.com \
    --cc=peter.fang@intel.com \
    --cc=peterx@redhat.com \
    --cc=rick.p.edgecombe@intel.com \
    --cc=sameo@rivosinc.com \
    --cc=seanjc@google.com \
    --cc=steven.price@arm.com \
    --cc=thomas.lendacky@amd.com \
    --cc=tony.lindgren@linux.intel.com \
    --cc=vannapurve@google.com \
    --cc=xiaoyao.li@intel.com \
    --cc=yilun.xu@linux.intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox