All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH 0/2] migration: defer a post_load which only rearranges memory
@ 2026-09-11 14:23 Denis V. Lunev
  2026-09-11 14:23 ` [PATCH 1/2] migration: let a vmstate defer its post_load to end of stream Denis V. Lunev
                   ` (3 more replies)
  0 siblings, 4 replies; 8+ messages in thread
From: Denis V. Lunev @ 2026-09-11 14:23 UTC (permalink / raw)
  To: qemu-devel
  Cc: den, Peter Xu, Fabiano Rosas, Paolo Bonzini, Zhao Liu,
	Denis V. Lunev

Restoring a big Windows guest spends most of its destination-side time in
post_load hooks which do nothing but move memory regions around. Each one
ends a memory transaction, and a transaction commit re-renders every
flatview it touches at a cost which grows with the number of regions in
the machine. A hook which runs once per vCPU therefore pays that render
once per vCPU, and the machine gets slower to migrate the bigger it is.

The Hyper-V SynIC is the case that hurts: restoring the synthetic
interrupt controller maps a message page and an event page per vCPU, so a
64-vCPU guest forces 128 remaps, each with its own rebuild, while the rest
of the stream is still being read.

Patch 1 adds post_load_deferrable. A vmsd which sets it has its hook
queued during the load and run once the stream has been consumed, in the
order the hooks would have fired, with the whole drain sharing one memory
transaction. Patch 2 sets it on the SynIC subsection.

It is opt-in rather than automatic, and the three preconditions are
spelled out on the field: the hook must not fail, nothing later in the
load may depend on what it does, and it must not read guest memory or
resolve an address space. A hook which breaks the first is fatal rather
than silently reported, because by drain time the source may already have
been told the migration succeeded.

Deferring is not free in general, which is the other reason it is opt-in.
Deferring the APIC post_load, whose cost is a synchronous run_on_cpu per
vCPU rather than a memory remap, moves 3 ms out of the section walk and
pays about 9 ms of drain for it. Deferral helps a hook which repeats
topology work; it makes a hook which does cross-thread work worse.

Measurements
------------

Destination-side non-iterable load, ie. the sum of vmstate_downtime_load
over non-iterable sections, on a guest which has actually programmed its
Hyper-V state. Five interleaved rounds per point on an otherwise idle
host, twice; medians, with the spread across all ten rounds.

  upstream                          377 ms   (363-400)
  + pci mapping transactions        247 ms   (242-255)
  + this series                      96 ms   (86-97)

Two postings against one problem, so the whole ladder is shown. The first
step is a pci pair which batches a device's BAR and bridge window updates
into a single transaction, posted separately and now queued in Michael's
tree:

  https://lore.kernel.org/qemu-devel/20260903184542.2629976-1-den@openvz.org/

Those two are listed because they change what a rebuild costs, and so
change what this series is worth. Together the postings take the load
from 377 ms to 96 ms; this series is the 247 ms to 96 ms step.

Where it goes: the cpu sections fall from 154 ms to 1.3 ms. The deferred
hooks themselves cost 77 us at the drain, so the work is removed rather
than moved somewhere the per-section metric cannot see.

The saving scales with vCPU count, since that is how many times the remap
repeats, and with the number of memory regions in the machine, since that
is what a rebuild costs.

Guest under test
----------------

Windows Server 2022, installed unattended, idle at the console:

  -machine q35,accel=kvm
  -cpu host,hv-synic,hv-stimer,hv-stimer-direct,hv-vapic,hv-runtime,
       hv-time,hv-ipi,hv-crash,hv-reset,hv-frequencies,hv-vpindex,
       hv-spinlocks=0x1fff
  -smp 64,sockets=2,cores=32,threads=1
  -m 4G
  65 pcie-root-ports, 9 virtio devices behind them, qxl

Host: AMD EPYC 7443P, 24 cores / 48 threads.

hv-synic is the flag that matters. Without it the guest never programs
the SynIC pages and the effect under test does not exist.

Measured with a save/restore harness rather than a live migration: a
restore from a captured stream walks the same qemu_loadvm_state_main()
path a destination does, which removes libvirt, the network and the
second host from the measurement.

CC: Peter Xu <peterx@redhat.com>
CC: Fabiano Rosas <farosas@suse.de>
CC: Paolo Bonzini <pbonzini@redhat.com>
CC: Zhao Liu <zhao1.liu@intel.com>
Signed-off-by: Denis V. Lunev <den@virtuozzo.com>

Denis V. Lunev (2):
  migration: let a vmstate defer its post_load to end of stream
  target/i386: defer the Hyper-V SynIC post_load

 include/migration/vmstate.h |  40 +++++++++++++
 migration/savevm.c          |  23 ++++++++
 migration/vmstate.c         |  67 +++++++++++++++++++++-
 target/i386/machine.c       |   1 +
 tests/unit/test-vmstate.c   | 111 ++++++++++++++++++++++++++++++++++++
 5 files changed, 240 insertions(+), 2 deletions(-)

-- 
2.53.0



^ permalink raw reply	[flat|nested] 8+ messages in thread

* [PATCH 1/2] migration: let a vmstate defer its post_load to end of stream
  2026-09-11 14:23 [PATCH 0/2] migration: defer a post_load which only rearranges memory Denis V. Lunev
@ 2026-09-11 14:23 ` Denis V. Lunev
  2026-10-01 19:52   ` Peter Xu
  2026-09-11 14:23 ` [PATCH 2/2] target/i386: defer the Hyper-V SynIC post_load Denis V. Lunev
                   ` (2 subsequent siblings)
  3 siblings, 1 reply; 8+ messages in thread
From: Denis V. Lunev @ 2026-09-11 14:23 UTC (permalink / raw)
  To: qemu-devel
  Cc: den, Peter Xu, Fabiano Rosas, Paolo Bonzini, Zhao Liu,
	Denis V. Lunev

From: Denis V. Lunev <den@openvz.org>

A post_load hook which only rearranges the memory topology forces a
flatview rebuild as its section is read, and the cost of a rebuild grows
with the number of regions in the machine. A device which does this once
per vCPU therefore scales badly on the destination.

Add post_load_deferrable. A vmsd which sets it has its hook queued during
the load and run once the stream has been consumed, in the order the
hooks would have fired, with the whole drain sharing one memory
transaction.

Opt-in, because deferral is not free in general. A hook which can fail
must not be deferred: failing after the stream is consumed means the
source has already been told the migration succeeded, and may release a
guest the destination never started. A hook which reads guest memory must
not be deferred either, since the drain runs with the topology in flux.
Postcopy is excluded because its listen thread walks the same stream
concurrently.

CC: Peter Xu <peterx@redhat.com>
CC: Fabiano Rosas <farosas@suse.de>
CC: Paolo Bonzini <pbonzini@redhat.com>
CC: Zhao Liu <zhao1.liu@intel.com>
Signed-off-by: Denis V. Lunev <den@virtuozzo.com>
---
 include/migration/vmstate.h |  40 +++++++++++++
 migration/savevm.c          |  23 ++++++++
 migration/vmstate.c         |  67 +++++++++++++++++++++-
 tests/unit/test-vmstate.c   | 111 ++++++++++++++++++++++++++++++++++++
 4 files changed, 239 insertions(+), 2 deletions(-)

diff --git a/include/migration/vmstate.h b/include/migration/vmstate.h
index e72c3fae9a..16045319d5 100644
--- a/include/migration/vmstate.h
+++ b/include/migration/vmstate.h
@@ -303,6 +303,29 @@ struct VMStateDescription {
     bool (*pre_load_errp)(void *opaque, Error **errp);
     int (*post_load)(void *opaque, int version_id);
     bool (*post_load_errp)(void *opaque, int version_id, Error **errp);
+
+    /*
+     * Run .post_load() once the whole stream has been loaded rather than
+     * as this section is read, so that it sees a machine whose devices
+     * have all been restored, and so that several of them can share the
+     * work they would each repeat.
+     *
+     * Three things must hold of a hook before it may be deferred.
+     *
+     * It must not fail. By the time the queue is drained the stream has
+     * been consumed, so a source may already have been told the migration
+     * succeeded and may have released the guest; there is nothing left to
+     * report a failure to. A hook which fails here is a bug in its vmsd
+     * and is fatal.
+     *
+     * Nothing else in the load may depend on what it does. The hooks run
+     * after every section has been read, so anything a later section needs
+     * to observe must not be produced here.
+     *
+     * It must not read guest memory or resolve an address space, because
+     * the drain runs as one batch with the memory topology in flux.
+     */
+    bool post_load_deferrable;
     int (*pre_save)(void *opaque);
     bool (*pre_save_errp)(void *opaque, Error **errp);
 
@@ -1299,6 +1322,23 @@ bool vmstate_save_vmsd(QEMUFile *f, const VMStateDescription *vmsd,
 
 bool vmstate_section_needed(const VMStateDescription *vmsd, void *opaque);
 
+/**
+ * vmstate_post_load_defer_begin: Queue deferrable post_load hooks
+ *
+ * Between this and vmstate_post_load_defer_finish(), a post_load hook whose
+ * vmsd sets post_load_deferrable is recorded rather than called. Hooks are
+ * queued in the order they would have run.
+ */
+void vmstate_post_load_defer_begin(void);
+
+/**
+ * vmstate_post_load_defer_finish: Stop deferring and drain the queue
+ * @run: run the queued hooks, in order; when false, discard them
+ *
+ * Returns false if a hook failed, in which case the rest are discarded.
+ */
+bool vmstate_post_load_defer_finish(bool run, Error **errp);
+
 #define  VMSTATE_INSTANCE_ID_ANY  -1
 
 /* Returns: 0 on success, -1 on failure */
diff --git a/migration/savevm.c b/migration/savevm.c
index 4b590ea672..2352dcf684 100644
--- a/migration/savevm.c
+++ b/migration/savevm.c
@@ -3128,6 +3128,7 @@ int qemu_loadvm_state(QEMUFile *f, Error **errp)
 {
     MigrationState *s = migrate_get_current();
     MigrationIncomingState *mis = migration_incoming_get_current();
+    bool defer_post_load;
     int ret;
 
     if (qemu_savevm_state_blocked(errp)) {
@@ -3147,7 +3148,29 @@ int qemu_loadvm_state(QEMUFile *f, Error **errp)
 
     cpu_synchronize_all_pre_loadvm();
 
+    /*
+     * The postcopy listen thread walks the same stream concurrently, so the
+     * queue would need locking and a defined owner for the drain.
+     */
+    defer_post_load = !migrate_postcopy_ram();
+    if (defer_post_load) {
+        vmstate_post_load_defer_begin();
+    }
+
     ret = qemu_loadvm_state_main(f, mis, errp);
+
+    if (defer_post_load) {
+        /*
+         * A deferrable hook only rearranges the memory topology, so the
+         * whole drain can share one flatview rebuild.
+         */
+        memory_region_transaction_begin();
+        if (!vmstate_post_load_defer_finish(ret == 0, errp)) {
+            ret = -EINVAL;
+        }
+        memory_region_transaction_commit();
+    }
+
     qemu_event_set(&mis->main_thread_load_event);
 
     trace_qemu_loadvm_state_post_main(ret);
diff --git a/migration/vmstate.c b/migration/vmstate.c
index 24004565bf..3745f80dc8 100644
--- a/migration/vmstate.c
+++ b/migration/vmstate.c
@@ -250,8 +250,19 @@ static bool vmstate_load_field(QEMUFile *f, void *pv, size_t size,
     return true;
 }
 
-static bool vmstate_post_load(const VMStateDescription *vmsd,
-                              void *opaque, int version_id, Error **errp)
+typedef struct VMStateDeferredPostLoad {
+    const VMStateDescription *vmsd;
+    void *opaque;
+    int version_id;
+    QSIMPLEQ_ENTRY(VMStateDeferredPostLoad) entry;
+} VMStateDeferredPostLoad;
+
+static QSIMPLEQ_HEAD(, VMStateDeferredPostLoad) vmstate_deferred_post_loads =
+    QSIMPLEQ_HEAD_INITIALIZER(vmstate_deferred_post_loads);
+static bool vmstate_defer_post_load;
+
+static bool vmstate_do_post_load(const VMStateDescription *vmsd,
+                                 void *opaque, int version_id, Error **errp)
 {
     ERRP_GUARD();
 
@@ -277,6 +288,58 @@ static bool vmstate_post_load(const VMStateDescription *vmsd,
     return true;
 }
 
+static bool vmstate_post_load(const VMStateDescription *vmsd,
+                              void *opaque, int version_id, Error **errp)
+{
+    VMStateDeferredPostLoad *d;
+
+    if (!vmstate_defer_post_load || !vmsd->post_load_deferrable) {
+        return vmstate_do_post_load(vmsd, opaque, version_id, errp);
+    }
+
+    d = g_new(VMStateDeferredPostLoad, 1);
+    d->vmsd = vmsd;
+    d->opaque = opaque;
+    d->version_id = version_id;
+    QSIMPLEQ_INSERT_TAIL(&vmstate_deferred_post_loads, d, entry);
+
+    return true;
+}
+
+void vmstate_post_load_defer_begin(void)
+{
+    assert(!vmstate_defer_post_load);
+    assert(QSIMPLEQ_EMPTY(&vmstate_deferred_post_loads));
+    vmstate_defer_post_load = true;
+}
+
+bool vmstate_post_load_defer_finish(bool run, Error **errp)
+{
+    VMStateDeferredPostLoad *d;
+    bool ok = true;
+
+    vmstate_defer_post_load = false;
+
+    while ((d = QSIMPLEQ_FIRST(&vmstate_deferred_post_loads))) {
+        QSIMPLEQ_REMOVE_HEAD(&vmstate_deferred_post_loads, entry);
+        if (run) {
+            ERRP_GUARD();
+
+            if (!vmstate_do_post_load(d->vmsd, d->opaque, d->version_id,
+                                      errp)) {
+                error_prepend(errp, "deferrable post load hook failed, which "
+                              "its vmsd promised could not happen: ");
+                error_report_err(*errp);
+                *errp = NULL;
+                abort();
+            }
+        }
+        g_free(d);
+    }
+
+    return ok;
+}
+
 /*
  * Try to prepare loading the next element, the object pointer to be put
  * into @next_elem.  When @next_elem is NULL, it means we should skip
diff --git a/tests/unit/test-vmstate.c b/tests/unit/test-vmstate.c
index df1fb4c778..e1e64b23c2 100644
--- a/tests/unit/test-vmstate.c
+++ b/tests/unit/test-vmstate.c
@@ -1619,6 +1619,116 @@ static void test_tmp_struct(void)
     g_assert_cmpint(obj.f, ==, 8); /* From the child->parent */
 }
 
+/* Deferred post_load */
+
+static int defer_order;
+static int defer_a_ran;
+static int defer_b_ran;
+static int defer_inline_ran;
+
+static int defer_a_post_load(void *opaque, int version_id)
+{
+    defer_a_ran = ++defer_order;
+    return 0;
+}
+
+static int defer_b_post_load(void *opaque, int version_id)
+{
+    defer_b_ran = ++defer_order;
+    return 0;
+}
+
+static int defer_inline_post_load(void *opaque, int version_id)
+{
+    defer_inline_ran = ++defer_order;
+    return 0;
+}
+
+static const VMStateDescription vmstate_defer_a = {
+    .name = "test/defer_a",
+    .version_id = 1,
+    .post_load = defer_a_post_load,
+    .post_load_deferrable = true,
+    .fields = (const VMStateField[]) {
+        VMSTATE_UINT32(a, TestStruct),
+        VMSTATE_END_OF_LIST()
+    }
+};
+
+static const VMStateDescription vmstate_defer_b = {
+    .name = "test/defer_b",
+    .version_id = 1,
+    .post_load = defer_b_post_load,
+    .post_load_deferrable = true,
+    .fields = (const VMStateField[]) {
+        VMSTATE_UINT32(a, TestStruct),
+        VMSTATE_END_OF_LIST()
+    }
+};
+
+static const VMStateDescription vmstate_defer_inline = {
+    .name = "test/defer_inline",
+    .version_id = 1,
+    .post_load = defer_inline_post_load,
+    .fields = (const VMStateField[]) {
+        VMSTATE_UINT32(a, TestStruct),
+        VMSTATE_END_OF_LIST()
+    }
+};
+
+static void defer_reset(void)
+{
+    defer_order = 0;
+    defer_a_ran = 0;
+    defer_b_ran = 0;
+    defer_inline_ran = 0;
+}
+
+static void test_post_load_defer(void)
+{
+    uint8_t const wire[] = {
+        /* uint32 a */ 0x00, 0x00, 0x00, 0x01,
+        QEMU_VM_EOF,
+    };
+    TestStruct obj;
+    Error *err = NULL;
+
+    /* Not deferring: a deferrable hook still runs as the section is read */
+    defer_reset();
+    memset(&obj, 0, sizeof(obj));
+    SUCCESS(load_vmstate_one(&vmstate_defer_a, &obj, 1, wire, sizeof(wire)));
+    g_assert_cmpint(defer_a_ran, ==, 1);
+
+    /*
+     * Deferring holds back the hooks which opted in and replays them in the
+     * order they would have run. A hook which did not opt in is unaffected.
+     */
+    defer_reset();
+    memset(&obj, 0, sizeof(obj));
+    vmstate_post_load_defer_begin();
+    SUCCESS(load_vmstate_one(&vmstate_defer_a, &obj, 1, wire, sizeof(wire)));
+    SUCCESS(load_vmstate_one(&vmstate_defer_inline, &obj, 1, wire,
+                             sizeof(wire)));
+    SUCCESS(load_vmstate_one(&vmstate_defer_b, &obj, 1, wire, sizeof(wire)));
+    g_assert_cmpint(defer_a_ran, ==, 0);
+    g_assert_cmpint(defer_b_ran, ==, 0);
+    g_assert_cmpint(defer_inline_ran, ==, 1);
+
+    g_assert(vmstate_post_load_defer_finish(true, &err));
+    g_assert(!err);
+    g_assert_cmpint(defer_a_ran, ==, 2);
+    g_assert_cmpint(defer_b_ran, ==, 3);
+
+    /* A discarded queue runs nothing */
+    defer_reset();
+    memset(&obj, 0, sizeof(obj));
+    vmstate_post_load_defer_begin();
+    SUCCESS(load_vmstate_one(&vmstate_defer_a, &obj, 1, wire, sizeof(wire)));
+    g_assert(vmstate_post_load_defer_finish(false, &err));
+    g_assert(!err);
+    g_assert_cmpint(defer_a_ran, ==, 0);
+}
+
 int main(int argc, char **argv)
 {
     g_autofree char *temp_file = g_strdup_printf("%s/vmst.test.XXXXXX",
@@ -1663,6 +1773,7 @@ int main(int argc, char **argv)
     g_test_add_func("/vmstate/qlist/save/saveqlist", test_save_qlist);
     g_test_add_func("/vmstate/qlist/load/loadqlist", test_load_qlist);
     g_test_add_func("/vmstate/tmp_struct", test_tmp_struct);
+    g_test_add_func("/vmstate/post_load/defer", test_post_load_defer);
     g_test_run();
 
     close(temp_fd);
-- 
2.53.0



^ permalink raw reply related	[flat|nested] 8+ messages in thread

* [PATCH 2/2] target/i386: defer the Hyper-V SynIC post_load
  2026-09-11 14:23 [PATCH 0/2] migration: defer a post_load which only rearranges memory Denis V. Lunev
  2026-09-11 14:23 ` [PATCH 1/2] migration: let a vmstate defer its post_load to end of stream Denis V. Lunev
@ 2026-09-11 14:23 ` Denis V. Lunev
  2026-09-16 15:21 ` [PATCH 0/2] migration: defer a post_load which only rearranges memory Denis V. Lunev
  2026-10-01 19:46 ` Peter Xu
  3 siblings, 0 replies; 8+ messages in thread
From: Denis V. Lunev @ 2026-09-11 14:23 UTC (permalink / raw)
  To: qemu-devel
  Cc: den, Peter Xu, Fabiano Rosas, Paolo Bonzini, Zhao Liu,
	Denis V. Lunev

From: Denis V. Lunev <den@openvz.org>

hyperv_synic_post_load() maps the SynIC message and event pages, and a
guest which has programmed the SynIC carries the subsection on every
vCPU, so restoring one rebuilds the flatview per vCPU rather than once.

The hook cannot fail and does not read guest memory, so it can run with
the rest of the deferred hooks once the stream has been consumed. The
saving scales with vCPU count.

CC: Peter Xu <peterx@redhat.com>
CC: Fabiano Rosas <farosas@suse.de>
CC: Paolo Bonzini <pbonzini@redhat.com>
CC: Zhao Liu <zhao1.liu@intel.com>
Signed-off-by: Denis V. Lunev <den@virtuozzo.com>
---
 target/i386/machine.c | 1 +
 1 file changed, 1 insertion(+)

diff --git a/target/i386/machine.c b/target/i386/machine.c
index 8d69d7e25e..eae74e33de 100644
--- a/target/i386/machine.c
+++ b/target/i386/machine.c
@@ -873,6 +873,7 @@ static const VMStateDescription vmstate_msr_hyperv_synic = {
     .minimum_version_id = 1,
     .needed = hyperv_synic_enable_needed,
     .post_load = hyperv_synic_post_load,
+    .post_load_deferrable = true,
     .fields = (const VMStateField[]) {
         VMSTATE_UINT64(env.msr_hv_synic_control, X86CPU),
         VMSTATE_UINT64(env.msr_hv_synic_evt_page, X86CPU),
-- 
2.53.0



^ permalink raw reply related	[flat|nested] 8+ messages in thread

* Re: [PATCH 0/2] migration: defer a post_load which only rearranges memory
  2026-09-11 14:23 [PATCH 0/2] migration: defer a post_load which only rearranges memory Denis V. Lunev
  2026-09-11 14:23 ` [PATCH 1/2] migration: let a vmstate defer its post_load to end of stream Denis V. Lunev
  2026-09-11 14:23 ` [PATCH 2/2] target/i386: defer the Hyper-V SynIC post_load Denis V. Lunev
@ 2026-09-16 15:21 ` Denis V. Lunev
  2026-10-01 19:46 ` Peter Xu
  3 siblings, 0 replies; 8+ messages in thread
From: Denis V. Lunev @ 2026-09-16 15:21 UTC (permalink / raw)
  To: Denis V. Lunev, qemu-devel
  Cc: Peter Xu, Fabiano Rosas, Paolo Bonzini, Zhao Liu

On 9/11/26 16:23, Denis V. Lunev wrote:
> Restoring a big Windows guest spends most of its destination-side time in
> post_load hooks which do nothing but move memory regions around. Each one
> ends a memory transaction, and a transaction commit re-renders every
> flatview it touches at a cost which grows with the number of regions in
> the machine. A hook which runs once per vCPU therefore pays that render
> once per vCPU, and the machine gets slower to migrate the bigger it is.
>
> The Hyper-V SynIC is the case that hurts: restoring the synthetic
> interrupt controller maps a message page and an event page per vCPU, so a
> 64-vCPU guest forces 128 remaps, each with its own rebuild, while the rest
> of the stream is still being read.
>
> Patch 1 adds post_load_deferrable. A vmsd which sets it has its hook
> queued during the load and run once the stream has been consumed, in the
> order the hooks would have fired, with the whole drain sharing one memory
> transaction. Patch 2 sets it on the SynIC subsection.
>
> It is opt-in rather than automatic, and the three preconditions are
> spelled out on the field: the hook must not fail, nothing later in the
> load may depend on what it does, and it must not read guest memory or
> resolve an address space. A hook which breaks the first is fatal rather
> than silently reported, because by drain time the source may already have
> been told the migration succeeded.
>
> Deferring is not free in general, which is the other reason it is opt-in.
> Deferring the APIC post_load, whose cost is a synchronous run_on_cpu per
> vCPU rather than a memory remap, moves 3 ms out of the section walk and
> pays about 9 ms of drain for it. Deferral helps a hook which repeats
> topology work; it makes a hook which does cross-thread work worse.
>
> Measurements
> ------------
>
> Destination-side non-iterable load, ie. the sum of vmstate_downtime_load
> over non-iterable sections, on a guest which has actually programmed its
> Hyper-V state. Five interleaved rounds per point on an otherwise idle
> host, twice; medians, with the spread across all ten rounds.
>
>   upstream                          377 ms   (363-400)
>   + pci mapping transactions        247 ms   (242-255)
>   + this series                      96 ms   (86-97)
>
> Two postings against one problem, so the whole ladder is shown. The first
> step is a pci pair which batches a device's BAR and bridge window updates
> into a single transaction, posted separately and now queued in Michael's
> tree:
>
>   https://lore.kernel.org/qemu-devel/20260903184542.2629976-1-den@openvz.org/
>
> Those two are listed because they change what a rebuild costs, and so
> change what this series is worth. Together the postings take the load
> from 377 ms to 96 ms; this series is the 247 ms to 96 ms step.
>
> Where it goes: the cpu sections fall from 154 ms to 1.3 ms. The deferred
> hooks themselves cost 77 us at the drain, so the work is removed rather
> than moved somewhere the per-section metric cannot see.
>
> The saving scales with vCPU count, since that is how many times the remap
> repeats, and with the number of memory regions in the machine, since that
> is what a rebuild costs.
>
> Guest under test
> ----------------
>
> Windows Server 2022, installed unattended, idle at the console:
>
>   -machine q35,accel=kvm
>   -cpu host,hv-synic,hv-stimer,hv-stimer-direct,hv-vapic,hv-runtime,
>        hv-time,hv-ipi,hv-crash,hv-reset,hv-frequencies,hv-vpindex,
>        hv-spinlocks=0x1fff
>   -smp 64,sockets=2,cores=32,threads=1
>   -m 4G
>   65 pcie-root-ports, 9 virtio devices behind them, qxl
>
> Host: AMD EPYC 7443P, 24 cores / 48 threads.
>
> hv-synic is the flag that matters. Without it the guest never programs
> the SynIC pages and the effect under test does not exist.
>
> Measured with a save/restore harness rather than a live migration: a
> restore from a captured stream walks the same qemu_loadvm_state_main()
> path a destination does, which removes libvirt, the network and the
> second host from the measurement.
>
> CC: Peter Xu <peterx@redhat.com>
> CC: Fabiano Rosas <farosas@suse.de>
> CC: Paolo Bonzini <pbonzini@redhat.com>
> CC: Zhao Liu <zhao1.liu@intel.com>
> Signed-off-by: Denis V. Lunev <den@virtuozzo.com>
>
> Denis V. Lunev (2):
>   migration: let a vmstate defer its post_load to end of stream
>   target/i386: defer the Hyper-V SynIC post_load
>
>  include/migration/vmstate.h |  40 +++++++++++++
>  migration/savevm.c          |  23 ++++++++
>  migration/vmstate.c         |  67 +++++++++++++++++++++-
>  target/i386/machine.c       |   1 +
>  tests/unit/test-vmstate.c   | 111 ++++++++++++++++++++++++++++++++++++
>  5 files changed, 240 insertions(+), 2 deletions(-)
>
please disregard this for now. I'll come back with v2 which
will put as around 30 ms.

Flatview rebuild should be optimized.

Sorry for noise,
    Den


^ permalink raw reply	[flat|nested] 8+ messages in thread

* Re: [PATCH 0/2] migration: defer a post_load which only rearranges memory
  2026-09-11 14:23 [PATCH 0/2] migration: defer a post_load which only rearranges memory Denis V. Lunev
                   ` (2 preceding siblings ...)
  2026-09-16 15:21 ` [PATCH 0/2] migration: defer a post_load which only rearranges memory Denis V. Lunev
@ 2026-10-01 19:46 ` Peter Xu
  2026-10-01 20:14   ` Denis V. Lunev
  3 siblings, 1 reply; 8+ messages in thread
From: Peter Xu @ 2026-10-01 19:46 UTC (permalink / raw)
  To: Denis V. Lunev
  Cc: qemu-devel, Fabiano Rosas, Paolo Bonzini, Zhao Liu,
	Denis V. Lunev

On Fri, Sep 11, 2026 at 04:23:43PM +0200, Denis V. Lunev wrote:
> Restoring a big Windows guest spends most of its destination-side time in
> post_load hooks which do nothing but move memory regions around. Each one
> ends a memory transaction, and a transaction commit re-renders every
> flatview it touches at a cost which grows with the number of regions in
> the machine. A hook which runs once per vCPU therefore pays that render
> once per vCPU, and the machine gets slower to migrate the bigger it is.
> 
> The Hyper-V SynIC is the case that hurts: restoring the synthetic
> interrupt controller maps a message page and an event page per vCPU, so a
> 64-vCPU guest forces 128 remaps, each with its own rebuild, while the rest
> of the stream is still being read.

This is partly a known issue, not from Hyper-V, but from virtio mmio
regions.. please see:

https://wiki.qemu.org/ToDo/LiveMigration#Optimize_memory_updates_for_non-iterative_vmstates
https://lore.kernel.org/r/20230317081904.24389-1-xuchuangxclwt@bytedance.com

I believe we also thought about do MR update per-device, so batching but
smaller scale, easier to make sure no illegal access to a stale flatview.

So in general, I agree this approach might be the right way to do, which is
to shrink the transaction to be smaller than "batch everything".. as what
Chuang used to do.  I still have some pure questions inline.

> 
> Patch 1 adds post_load_deferrable. A vmsd which sets it has its hook

Nit, IMHO if so it needs to be called "deferred", as "deferrable" implies
the defer is optional.

> queued during the load and run once the stream has been consumed, in the
> order the hooks would have fired, with the whole drain sharing one memory
> transaction. Patch 2 sets it on the SynIC subsection.
> 
> It is opt-in rather than automatic, and the three preconditions are
> spelled out on the field: the hook must not fail, nothing later in the

Actually, I _think_ maybe it can still fail.. IIUC source QEMU only dies if
it receives shut from migrate_send_rp_shut(), which is after the deferred
loads at least with the current change.  Worth check..

> load may depend on what it does, and it must not read guest memory or
> resolve an address space.

Yes, but I think this is partial of the whole picture: IIUC if this
deferred hook may inject some MRs that may be accessed by other VMSD
loaders, or anything (including hard-coded loading process), I think it's
an issue too.

So personally I don't like this API very much yet on how it was defined;
it'll be very hard to be used right unless we fully understand what will
happen..

I wonder if there's better way to define the API to be clearer.

Since all the known issues about this is about MR updates: virtio MMIO
regions, hyper-v, pci bar/bridge (mentioned below), I wonder if this can be
something dedicated to MR updates, and maybe it doesn't need to be
"deferred", just grouped together properly into one transaction, which can
happen in the middle too or maybe it doesn't matter much. Then it applies
some form of limitation to what can be split out from normal VMSD flow. The
current API relies on allowing to defer anything, which is fine but very
hard to control, and we may face tricky bugs if users grows but when
they're not used right..

> A hook which breaks the first is fatal rather
> than silently reported, because by drain time the source may already have
> been told the migration succeeded.
> 
> Deferring is not free in general, which is the other reason it is opt-in.
> Deferring the APIC post_load, whose cost is a synchronous run_on_cpu per
> vCPU rather than a memory remap, moves 3 ms out of the section walk and
> pays about 9 ms of drain for it. Deferral helps a hook which repeats

I'm just curious: why something will take 9ms if deferred, even if it used
to take 3ms?  I think I misread something, but I can't tell myself.

> topology work; it makes a hook which does cross-thread work worse.
> 
> Measurements
> ------------
> 
> Destination-side non-iterable load, ie. the sum of vmstate_downtime_load
> over non-iterable sections, on a guest which has actually programmed its
> Hyper-V state. Five interleaved rounds per point on an otherwise idle
> host, twice; medians, with the spread across all ten rounds.
> 
>   upstream                          377 ms   (363-400)
>   + pci mapping transactions        247 ms   (242-255)
>   + this series                      96 ms   (86-97)

Definitely a great improvement.  I think we need this, just one way or
another. I may have some other trivial comments later in the patch.  Thanks
for working on it.

> 
> Two postings against one problem, so the whole ladder is shown. The first
> step is a pci pair which batches a device's BAR and bridge window updates
> into a single transaction, posted separately and now queued in Michael's
> tree:
> 
>   https://lore.kernel.org/qemu-devel/20260903184542.2629976-1-den@openvz.org/
> 
> Those two are listed because they change what a rebuild costs, and so
> change what this series is worth. Together the postings take the load
> from 377 ms to 96 ms; this series is the 247 ms to 96 ms step.
> 
> Where it goes: the cpu sections fall from 154 ms to 1.3 ms. The deferred
> hooks themselves cost 77 us at the drain, so the work is removed rather
> than moved somewhere the per-section metric cannot see.
> 
> The saving scales with vCPU count, since that is how many times the remap
> repeats, and with the number of memory regions in the machine, since that
> is what a rebuild costs.
> 
> Guest under test
> ----------------
> 
> Windows Server 2022, installed unattended, idle at the console:
> 
>   -machine q35,accel=kvm
>   -cpu host,hv-synic,hv-stimer,hv-stimer-direct,hv-vapic,hv-runtime,
>        hv-time,hv-ipi,hv-crash,hv-reset,hv-frequencies,hv-vpindex,
>        hv-spinlocks=0x1fff
>   -smp 64,sockets=2,cores=32,threads=1
>   -m 4G
>   65 pcie-root-ports, 9 virtio devices behind them, qxl
> 
> Host: AMD EPYC 7443P, 24 cores / 48 threads.
> 
> hv-synic is the flag that matters. Without it the guest never programs
> the SynIC pages and the effect under test does not exist.
> 
> Measured with a save/restore harness rather than a live migration: a
> restore from a captured stream walks the same qemu_loadvm_state_main()
> path a destination does, which removes libvirt, the network and the
> second host from the measurement.
> 
> CC: Peter Xu <peterx@redhat.com>
> CC: Fabiano Rosas <farosas@suse.de>
> CC: Paolo Bonzini <pbonzini@redhat.com>
> CC: Zhao Liu <zhao1.liu@intel.com>
> Signed-off-by: Denis V. Lunev <den@virtuozzo.com>
> 
> Denis V. Lunev (2):
>   migration: let a vmstate defer its post_load to end of stream
>   target/i386: defer the Hyper-V SynIC post_load
> 
>  include/migration/vmstate.h |  40 +++++++++++++
>  migration/savevm.c          |  23 ++++++++
>  migration/vmstate.c         |  67 +++++++++++++++++++++-
>  target/i386/machine.c       |   1 +
>  tests/unit/test-vmstate.c   | 111 ++++++++++++++++++++++++++++++++++++
>  5 files changed, 240 insertions(+), 2 deletions(-)
> 
> -- 
> 2.53.0
> 

-- 
Peter Xu



^ permalink raw reply	[flat|nested] 8+ messages in thread

* Re: [PATCH 1/2] migration: let a vmstate defer its post_load to end of stream
  2026-09-11 14:23 ` [PATCH 1/2] migration: let a vmstate defer its post_load to end of stream Denis V. Lunev
@ 2026-10-01 19:52   ` Peter Xu
  0 siblings, 0 replies; 8+ messages in thread
From: Peter Xu @ 2026-10-01 19:52 UTC (permalink / raw)
  To: Denis V. Lunev
  Cc: qemu-devel, Fabiano Rosas, Paolo Bonzini, Zhao Liu,
	Denis V. Lunev

On Fri, Sep 11, 2026 at 04:23:44PM +0200, Denis V. Lunev wrote:
> From: Denis V. Lunev <den@openvz.org>
> 
> A post_load hook which only rearranges the memory topology forces a
> flatview rebuild as its section is read, and the cost of a rebuild grows
> with the number of regions in the machine. A device which does this once
> per vCPU therefore scales badly on the destination.
> 
> Add post_load_deferrable. A vmsd which sets it has its hook queued during
> the load and run once the stream has been consumed, in the order the
> hooks would have fired, with the whole drain sharing one memory
> transaction.
> 
> Opt-in, because deferral is not free in general. A hook which can fail
> must not be deferred: failing after the stream is consumed means the
> source has already been told the migration succeeded, and may release a
> guest the destination never started. A hook which reads guest memory must
> not be deferred either, since the drain runs with the topology in flux.
> Postcopy is excluded because its listen thread walks the same stream
> concurrently.
> 
> CC: Peter Xu <peterx@redhat.com>
> CC: Fabiano Rosas <farosas@suse.de>
> CC: Paolo Bonzini <pbonzini@redhat.com>
> CC: Zhao Liu <zhao1.liu@intel.com>
> Signed-off-by: Denis V. Lunev <den@virtuozzo.com>
> ---
>  include/migration/vmstate.h |  40 +++++++++++++
>  migration/savevm.c          |  23 ++++++++
>  migration/vmstate.c         |  67 +++++++++++++++++++++-
>  tests/unit/test-vmstate.c   | 111 ++++++++++++++++++++++++++++++++++++
>  4 files changed, 239 insertions(+), 2 deletions(-)
> 
> diff --git a/include/migration/vmstate.h b/include/migration/vmstate.h
> index e72c3fae9a..16045319d5 100644
> --- a/include/migration/vmstate.h
> +++ b/include/migration/vmstate.h
> @@ -303,6 +303,29 @@ struct VMStateDescription {
>      bool (*pre_load_errp)(void *opaque, Error **errp);
>      int (*post_load)(void *opaque, int version_id);
>      bool (*post_load_errp)(void *opaque, int version_id, Error **errp);
> +
> +    /*
> +     * Run .post_load() once the whole stream has been loaded rather than
> +     * as this section is read, so that it sees a machine whose devices
> +     * have all been restored, and so that several of them can share the
> +     * work they would each repeat.
> +     *
> +     * Three things must hold of a hook before it may be deferred.
> +     *
> +     * It must not fail. By the time the queue is drained the stream has
> +     * been consumed, so a source may already have been told the migration
> +     * succeeded and may have released the guest; there is nothing left to
> +     * report a failure to. A hook which fails here is a bug in its vmsd
> +     * and is fatal.
> +     *
> +     * Nothing else in the load may depend on what it does. The hooks run
> +     * after every section has been read, so anything a later section needs
> +     * to observe must not be produced here.
> +     *
> +     * It must not read guest memory or resolve an address space, because
> +     * the drain runs as one batch with the memory topology in flux.
> +     */
> +    bool post_load_deferrable;
>      int (*pre_save)(void *opaque);
>      bool (*pre_save_errp)(void *opaque, Error **errp);
>  
> @@ -1299,6 +1322,23 @@ bool vmstate_save_vmsd(QEMUFile *f, const VMStateDescription *vmsd,
>  
>  bool vmstate_section_needed(const VMStateDescription *vmsd, void *opaque);
>  
> +/**
> + * vmstate_post_load_defer_begin: Queue deferrable post_load hooks
> + *
> + * Between this and vmstate_post_load_defer_finish(), a post_load hook whose
> + * vmsd sets post_load_deferrable is recorded rather than called. Hooks are
> + * queued in the order they would have run.
> + */
> +void vmstate_post_load_defer_begin(void);
> +
> +/**
> + * vmstate_post_load_defer_finish: Stop deferring and drain the queue
> + * @run: run the queued hooks, in order; when false, discard them
> + *
> + * Returns false if a hook failed, in which case the rest are discarded.
> + */
> +bool vmstate_post_load_defer_finish(bool run, Error **errp);
> +
>  #define  VMSTATE_INSTANCE_ID_ANY  -1
>  
>  /* Returns: 0 on success, -1 on failure */
> diff --git a/migration/savevm.c b/migration/savevm.c
> index 4b590ea672..2352dcf684 100644
> --- a/migration/savevm.c
> +++ b/migration/savevm.c
> @@ -3128,6 +3128,7 @@ int qemu_loadvm_state(QEMUFile *f, Error **errp)
>  {
>      MigrationState *s = migrate_get_current();
>      MigrationIncomingState *mis = migration_incoming_get_current();
> +    bool defer_post_load;
>      int ret;
>  
>      if (qemu_savevm_state_blocked(errp)) {
> @@ -3147,7 +3148,29 @@ int qemu_loadvm_state(QEMUFile *f, Error **errp)
>  
>      cpu_synchronize_all_pre_loadvm();
>  
> +    /*
> +     * The postcopy listen thread walks the same stream concurrently, so the
> +     * queue would need locking and a defined owner for the drain.
> +     */
> +    defer_post_load = !migrate_postcopy_ram();

Just to mention we have two other paths that will enable this too:

qmp_xen_load_devices_state
load_snapshot

It looks all fine at least considering the hyper-v scope, but still raise
this in case it's not expected.

The other thing is, IMHO we should enable this feature for both precopy and
postcopy.  Postcopy loads that in the package, please kindly share more on
the above comment on disabling it; I didn't directly get that part.

> +    if (defer_post_load) {
> +        vmstate_post_load_defer_begin();
> +    }
> +
>      ret = qemu_loadvm_state_main(f, mis, errp);
> +
> +    if (defer_post_load) {
> +        /*
> +         * A deferrable hook only rearranges the memory topology, so the
> +         * whole drain can share one flatview rebuild.
> +         */
> +        memory_region_transaction_begin();

IMHO we should document this behavior in the "deferrable" interface above,
this is important knowledge to know that the whole deferred vmsd loads are
wrapped with boosted transaction depths and all memory access is illegal.

AFAIU, that partly supplement my other reply, that the whole thing is not
about "defer" or not, but about "batching MR update".  It can still be at
the end of course, but it's not the major goal, the major goal is about
memory updates.

Thanks,

> +        if (!vmstate_post_load_defer_finish(ret == 0, errp)) {
> +            ret = -EINVAL;
> +        }
> +        memory_region_transaction_commit();
> +    }

-- 
Peter Xu



^ permalink raw reply	[flat|nested] 8+ messages in thread

* Re: [PATCH 0/2] migration: defer a post_load which only rearranges memory
  2026-10-01 19:46 ` Peter Xu
@ 2026-10-01 20:14   ` Denis V. Lunev
  2026-10-01 21:02     ` Peter Xu
  0 siblings, 1 reply; 8+ messages in thread
From: Denis V. Lunev @ 2026-10-01 20:14 UTC (permalink / raw)
  To: Peter Xu, Denis V. Lunev
  Cc: qemu-devel, Fabiano Rosas, Paolo Bonzini, Zhao Liu

On 10/1/26 21:46, Peter Xu wrote:
> On Fri, Sep 11, 2026 at 04:23:43PM +0200, Denis V. Lunev wrote:
>> Restoring a big Windows guest spends most of its destination-side time in
>> post_load hooks which do nothing but move memory regions around. Each one
>> ends a memory transaction, and a transaction commit re-renders every
>> flatview it touches at a cost which grows with the number of regions in
>> the machine. A hook which runs once per vCPU therefore pays that render
>> once per vCPU, and the machine gets slower to migrate the bigger it is.
>>
>> The Hyper-V SynIC is the case that hurts: restoring the synthetic
>> interrupt controller maps a message page and an event page per vCPU, so a
>> 64-vCPU guest forces 128 remaps, each with its own rebuild, while the rest
>> of the stream is still being read.
> This is partly a known issue, not from Hyper-V, but from virtio mmio
> regions.. please see:
>
> https://wiki.qemu.org/ToDo/LiveMigration#Optimize_memory_updates_for_non-iterative_vmstates
> https://lore.kernel.org/r/20230317081904.24389-1-xuchuangxclwt@bytedance.com
>
> I believe we also thought about do MR update per-device, so batching but
> smaller scale, easier to make sure no illegal access to a stale flatview.
>
> So in general, I agree this approach might be the right way to do, which is
> to shrink the transaction to be smaller than "batch everything".. as what
> Chuang used to do.  I still have some pure questions inline.
>
>> Patch 1 adds post_load_deferrable. A vmsd which sets it has its hook
> Nit, IMHO if so it needs to be called "deferred", as "deferrable" implies
> the defer is optional.
>
>> queued during the load and run once the stream has been consumed, in the
>> order the hooks would have fired, with the whole drain sharing one memory
>> transaction. Patch 2 sets it on the SynIC subsection.
>>
>> It is opt-in rather than automatic, and the three preconditions are
>> spelled out on the field: the hook must not fail, nothing later in the
> Actually, I _think_ maybe it can still fail.. IIUC source QEMU only dies if
> it receives shut from migrate_send_rp_shut(), which is after the deferred
> loads at least with the current change.  Worth check..
>
>> load may depend on what it does, and it must not read guest memory or
>> resolve an address space.
> Yes, but I think this is partial of the whole picture: IIUC if this
> deferred hook may inject some MRs that may be accessed by other VMSD
> loaders, or anything (including hard-coded loading process), I think it's
> an issue too.
>
> So personally I don't like this API very much yet on how it was defined;
> it'll be very hard to be used right unless we fully understand what will
> happen..
>
> I wonder if there's better way to define the API to be clearer.
>
> Since all the known issues about this is about MR updates: virtio MMIO
> regions, hyper-v, pci bar/bridge (mentioned below), I wonder if this can be
> something dedicated to MR updates, and maybe it doesn't need to be
> "deferred", just grouped together properly into one transaction, which can
> happen in the middle too or maybe it doesn't matter much. Then it applies
> some form of limitation to what can be split out from normal VMSD flow. The
> current API relies on allowing to defer anything, which is fine but very
> hard to control, and we may face tricky bugs if users grows but when
> they're not used right..
>
>> A hook which breaks the first is fatal rather
>> than silently reported, because by drain time the source may already have
>> been told the migration succeeded.
>>
>> Deferring is not free in general, which is the other reason it is opt-in.
>> Deferring the APIC post_load, whose cost is a synchronous run_on_cpu per
>> vCPU rather than a memory remap, moves 3 ms out of the section walk and
>> pays about 9 ms of drain for it. Deferral helps a hook which repeats
> I'm just curious: why something will take 9ms if deferred, even if it used
> to take 3ms?  I think I misread something, but I can't tell myself.
>
>> topology work; it makes a hook which does cross-thread work worse.
>>
>> Measurements
>> ------------
>>
>> Destination-side non-iterable load, ie. the sum of vmstate_downtime_load
>> over non-iterable sections, on a guest which has actually programmed its
>> Hyper-V state. Five interleaved rounds per point on an otherwise idle
>> host, twice; medians, with the spread across all ten rounds.
>>
>>   upstream                          377 ms   (363-400)
>>   + pci mapping transactions        247 ms   (242-255)
>>   + this series                      96 ms   (86-97)
> Definitely a great improvement.  I think we need this, just one way or
> another. I may have some other trivial comments later in the patch.  Thanks
> for working on it.

Right now I have series staying at 32 ms with boost to ordinary
QEMU startup, that is why I have said that I'll send follow up.
This patch will be included along as flatview rebuild on startup.

Anyway, right now I need to understand and eat your notes to
send v2. This just intersects with upcoming release of downstream,
but this change is in my priority list and I hope to finish
v2 early next week.

Thank you for you time,
    Den


^ permalink raw reply	[flat|nested] 8+ messages in thread

* Re: [PATCH 0/2] migration: defer a post_load which only rearranges memory
  2026-10-01 20:14   ` Denis V. Lunev
@ 2026-10-01 21:02     ` Peter Xu
  0 siblings, 0 replies; 8+ messages in thread
From: Peter Xu @ 2026-10-01 21:02 UTC (permalink / raw)
  To: Denis V. Lunev
  Cc: Denis V. Lunev, qemu-devel, Fabiano Rosas, Paolo Bonzini,
	Zhao Liu

On Thu, Oct 01, 2026 at 10:14:20PM +0200, Denis V. Lunev wrote:
> Right now I have series staying at 32 ms with boost to ordinary
> QEMU startup, that is why I have said that I'll send follow up.
> This patch will be included along as flatview rebuild on startup.
> 
> Anyway, right now I need to understand and eat your notes to
> send v2. This just intersects with upcoming release of downstream,
> but this change is in my priority list and I hope to finish
> v2 early next week.

Ah I didn't notice the other email when replying...  Take your time, feel
free to take whatever comments still makes sense to you, thanks.

-- 
Peter Xu



^ permalink raw reply	[flat|nested] 8+ messages in thread

end of thread, other threads:[~2026-10-01 21:03 UTC | newest]

Thread overview: 8+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-11 14:23 [PATCH 0/2] migration: defer a post_load which only rearranges memory Denis V. Lunev
2026-09-11 14:23 ` [PATCH 1/2] migration: let a vmstate defer its post_load to end of stream Denis V. Lunev
2026-10-01 19:52   ` Peter Xu
2026-09-11 14:23 ` [PATCH 2/2] target/i386: defer the Hyper-V SynIC post_load Denis V. Lunev
2026-09-16 15:21 ` [PATCH 0/2] migration: defer a post_load which only rearranges memory Denis V. Lunev
2026-10-01 19:46 ` Peter Xu
2026-10-01 20:14   ` Denis V. Lunev
2026-10-01 21:02     ` Peter Xu

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.