All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH 0/2] migration: defer a post_load which only rearranges memory
@ 2026-09-11 14:23 Denis V. Lunev
  2026-09-11 14:23 ` [PATCH 1/2] migration: let a vmstate defer its post_load to end of stream Denis V. Lunev
  2026-09-11 14:23 ` [PATCH 2/2] target/i386: defer the Hyper-V SynIC post_load Denis V. Lunev
  0 siblings, 2 replies; 3+ messages in thread
From: Denis V. Lunev @ 2026-09-11 14:23 UTC (permalink / raw)
  To: qemu-devel
  Cc: den, Peter Xu, Fabiano Rosas, Paolo Bonzini, Zhao Liu,
	Denis V. Lunev

Restoring a big Windows guest spends most of its destination-side time in
post_load hooks which do nothing but move memory regions around. Each one
ends a memory transaction, and a transaction commit re-renders every
flatview it touches at a cost which grows with the number of regions in
the machine. A hook which runs once per vCPU therefore pays that render
once per vCPU, and the machine gets slower to migrate the bigger it is.

The Hyper-V SynIC is the case that hurts: restoring the synthetic
interrupt controller maps a message page and an event page per vCPU, so a
64-vCPU guest forces 128 remaps, each with its own rebuild, while the rest
of the stream is still being read.

Patch 1 adds post_load_deferrable. A vmsd which sets it has its hook
queued during the load and run once the stream has been consumed, in the
order the hooks would have fired, with the whole drain sharing one memory
transaction. Patch 2 sets it on the SynIC subsection.

It is opt-in rather than automatic, and the three preconditions are
spelled out on the field: the hook must not fail, nothing later in the
load may depend on what it does, and it must not read guest memory or
resolve an address space. A hook which breaks the first is fatal rather
than silently reported, because by drain time the source may already have
been told the migration succeeded.

Deferring is not free in general, which is the other reason it is opt-in.
Deferring the APIC post_load, whose cost is a synchronous run_on_cpu per
vCPU rather than a memory remap, moves 3 ms out of the section walk and
pays about 9 ms of drain for it. Deferral helps a hook which repeats
topology work; it makes a hook which does cross-thread work worse.

Measurements
------------

Destination-side non-iterable load, ie. the sum of vmstate_downtime_load
over non-iterable sections, on a guest which has actually programmed its
Hyper-V state. Five interleaved rounds per point on an otherwise idle
host, twice; medians, with the spread across all ten rounds.

  upstream                          377 ms   (363-400)
  + pci mapping transactions        247 ms   (242-255)
  + this series                      96 ms   (86-97)

Two postings against one problem, so the whole ladder is shown. The first
step is a pci pair which batches a device's BAR and bridge window updates
into a single transaction, posted separately and now queued in Michael's
tree:

  https://lore.kernel.org/qemu-devel/20260903184542.2629976-1-den@openvz.org/

Those two are listed because they change what a rebuild costs, and so
change what this series is worth. Together the postings take the load
from 377 ms to 96 ms; this series is the 247 ms to 96 ms step.

Where it goes: the cpu sections fall from 154 ms to 1.3 ms. The deferred
hooks themselves cost 77 us at the drain, so the work is removed rather
than moved somewhere the per-section metric cannot see.

The saving scales with vCPU count, since that is how many times the remap
repeats, and with the number of memory regions in the machine, since that
is what a rebuild costs.

Guest under test
----------------

Windows Server 2022, installed unattended, idle at the console:

  -machine q35,accel=kvm
  -cpu host,hv-synic,hv-stimer,hv-stimer-direct,hv-vapic,hv-runtime,
       hv-time,hv-ipi,hv-crash,hv-reset,hv-frequencies,hv-vpindex,
       hv-spinlocks=0x1fff
  -smp 64,sockets=2,cores=32,threads=1
  -m 4G
  65 pcie-root-ports, 9 virtio devices behind them, qxl

Host: AMD EPYC 7443P, 24 cores / 48 threads.

hv-synic is the flag that matters. Without it the guest never programs
the SynIC pages and the effect under test does not exist.

Measured with a save/restore harness rather than a live migration: a
restore from a captured stream walks the same qemu_loadvm_state_main()
path a destination does, which removes libvirt, the network and the
second host from the measurement.

CC: Peter Xu <peterx@redhat.com>
CC: Fabiano Rosas <farosas@suse.de>
CC: Paolo Bonzini <pbonzini@redhat.com>
CC: Zhao Liu <zhao1.liu@intel.com>
Signed-off-by: Denis V. Lunev <den@virtuozzo.com>

Denis V. Lunev (2):
  migration: let a vmstate defer its post_load to end of stream
  target/i386: defer the Hyper-V SynIC post_load

 include/migration/vmstate.h |  40 +++++++++++++
 migration/savevm.c          |  23 ++++++++
 migration/vmstate.c         |  67 +++++++++++++++++++++-
 target/i386/machine.c       |   1 +
 tests/unit/test-vmstate.c   | 111 ++++++++++++++++++++++++++++++++++++
 5 files changed, 240 insertions(+), 2 deletions(-)

-- 
2.53.0



^ permalink raw reply	[flat|nested] 3+ messages in thread

* [PATCH 1/2] migration: let a vmstate defer its post_load to end of stream
  2026-09-11 14:23 [PATCH 0/2] migration: defer a post_load which only rearranges memory Denis V. Lunev
@ 2026-09-11 14:23 ` Denis V. Lunev
  2026-09-11 14:23 ` [PATCH 2/2] target/i386: defer the Hyper-V SynIC post_load Denis V. Lunev
  1 sibling, 0 replies; 3+ messages in thread
From: Denis V. Lunev @ 2026-09-11 14:23 UTC (permalink / raw)
  To: qemu-devel
  Cc: den, Peter Xu, Fabiano Rosas, Paolo Bonzini, Zhao Liu,
	Denis V. Lunev

From: Denis V. Lunev <den@openvz.org>

A post_load hook which only rearranges the memory topology forces a
flatview rebuild as its section is read, and the cost of a rebuild grows
with the number of regions in the machine. A device which does this once
per vCPU therefore scales badly on the destination.

Add post_load_deferrable. A vmsd which sets it has its hook queued during
the load and run once the stream has been consumed, in the order the
hooks would have fired, with the whole drain sharing one memory
transaction.

Opt-in, because deferral is not free in general. A hook which can fail
must not be deferred: failing after the stream is consumed means the
source has already been told the migration succeeded, and may release a
guest the destination never started. A hook which reads guest memory must
not be deferred either, since the drain runs with the topology in flux.
Postcopy is excluded because its listen thread walks the same stream
concurrently.

CC: Peter Xu <peterx@redhat.com>
CC: Fabiano Rosas <farosas@suse.de>
CC: Paolo Bonzini <pbonzini@redhat.com>
CC: Zhao Liu <zhao1.liu@intel.com>
Signed-off-by: Denis V. Lunev <den@virtuozzo.com>
---
 include/migration/vmstate.h |  40 +++++++++++++
 migration/savevm.c          |  23 ++++++++
 migration/vmstate.c         |  67 +++++++++++++++++++++-
 tests/unit/test-vmstate.c   | 111 ++++++++++++++++++++++++++++++++++++
 4 files changed, 239 insertions(+), 2 deletions(-)

diff --git a/include/migration/vmstate.h b/include/migration/vmstate.h
index e72c3fae9a..16045319d5 100644
--- a/include/migration/vmstate.h
+++ b/include/migration/vmstate.h
@@ -303,6 +303,29 @@ struct VMStateDescription {
     bool (*pre_load_errp)(void *opaque, Error **errp);
     int (*post_load)(void *opaque, int version_id);
     bool (*post_load_errp)(void *opaque, int version_id, Error **errp);
+
+    /*
+     * Run .post_load() once the whole stream has been loaded rather than
+     * as this section is read, so that it sees a machine whose devices
+     * have all been restored, and so that several of them can share the
+     * work they would each repeat.
+     *
+     * Three things must hold of a hook before it may be deferred.
+     *
+     * It must not fail. By the time the queue is drained the stream has
+     * been consumed, so a source may already have been told the migration
+     * succeeded and may have released the guest; there is nothing left to
+     * report a failure to. A hook which fails here is a bug in its vmsd
+     * and is fatal.
+     *
+     * Nothing else in the load may depend on what it does. The hooks run
+     * after every section has been read, so anything a later section needs
+     * to observe must not be produced here.
+     *
+     * It must not read guest memory or resolve an address space, because
+     * the drain runs as one batch with the memory topology in flux.
+     */
+    bool post_load_deferrable;
     int (*pre_save)(void *opaque);
     bool (*pre_save_errp)(void *opaque, Error **errp);
 
@@ -1299,6 +1322,23 @@ bool vmstate_save_vmsd(QEMUFile *f, const VMStateDescription *vmsd,
 
 bool vmstate_section_needed(const VMStateDescription *vmsd, void *opaque);
 
+/**
+ * vmstate_post_load_defer_begin: Queue deferrable post_load hooks
+ *
+ * Between this and vmstate_post_load_defer_finish(), a post_load hook whose
+ * vmsd sets post_load_deferrable is recorded rather than called. Hooks are
+ * queued in the order they would have run.
+ */
+void vmstate_post_load_defer_begin(void);
+
+/**
+ * vmstate_post_load_defer_finish: Stop deferring and drain the queue
+ * @run: run the queued hooks, in order; when false, discard them
+ *
+ * Returns false if a hook failed, in which case the rest are discarded.
+ */
+bool vmstate_post_load_defer_finish(bool run, Error **errp);
+
 #define  VMSTATE_INSTANCE_ID_ANY  -1
 
 /* Returns: 0 on success, -1 on failure */
diff --git a/migration/savevm.c b/migration/savevm.c
index 4b590ea672..2352dcf684 100644
--- a/migration/savevm.c
+++ b/migration/savevm.c
@@ -3128,6 +3128,7 @@ int qemu_loadvm_state(QEMUFile *f, Error **errp)
 {
     MigrationState *s = migrate_get_current();
     MigrationIncomingState *mis = migration_incoming_get_current();
+    bool defer_post_load;
     int ret;
 
     if (qemu_savevm_state_blocked(errp)) {
@@ -3147,7 +3148,29 @@ int qemu_loadvm_state(QEMUFile *f, Error **errp)
 
     cpu_synchronize_all_pre_loadvm();
 
+    /*
+     * The postcopy listen thread walks the same stream concurrently, so the
+     * queue would need locking and a defined owner for the drain.
+     */
+    defer_post_load = !migrate_postcopy_ram();
+    if (defer_post_load) {
+        vmstate_post_load_defer_begin();
+    }
+
     ret = qemu_loadvm_state_main(f, mis, errp);
+
+    if (defer_post_load) {
+        /*
+         * A deferrable hook only rearranges the memory topology, so the
+         * whole drain can share one flatview rebuild.
+         */
+        memory_region_transaction_begin();
+        if (!vmstate_post_load_defer_finish(ret == 0, errp)) {
+            ret = -EINVAL;
+        }
+        memory_region_transaction_commit();
+    }
+
     qemu_event_set(&mis->main_thread_load_event);
 
     trace_qemu_loadvm_state_post_main(ret);
diff --git a/migration/vmstate.c b/migration/vmstate.c
index 24004565bf..3745f80dc8 100644
--- a/migration/vmstate.c
+++ b/migration/vmstate.c
@@ -250,8 +250,19 @@ static bool vmstate_load_field(QEMUFile *f, void *pv, size_t size,
     return true;
 }
 
-static bool vmstate_post_load(const VMStateDescription *vmsd,
-                              void *opaque, int version_id, Error **errp)
+typedef struct VMStateDeferredPostLoad {
+    const VMStateDescription *vmsd;
+    void *opaque;
+    int version_id;
+    QSIMPLEQ_ENTRY(VMStateDeferredPostLoad) entry;
+} VMStateDeferredPostLoad;
+
+static QSIMPLEQ_HEAD(, VMStateDeferredPostLoad) vmstate_deferred_post_loads =
+    QSIMPLEQ_HEAD_INITIALIZER(vmstate_deferred_post_loads);
+static bool vmstate_defer_post_load;
+
+static bool vmstate_do_post_load(const VMStateDescription *vmsd,
+                                 void *opaque, int version_id, Error **errp)
 {
     ERRP_GUARD();
 
@@ -277,6 +288,58 @@ static bool vmstate_post_load(const VMStateDescription *vmsd,
     return true;
 }
 
+static bool vmstate_post_load(const VMStateDescription *vmsd,
+                              void *opaque, int version_id, Error **errp)
+{
+    VMStateDeferredPostLoad *d;
+
+    if (!vmstate_defer_post_load || !vmsd->post_load_deferrable) {
+        return vmstate_do_post_load(vmsd, opaque, version_id, errp);
+    }
+
+    d = g_new(VMStateDeferredPostLoad, 1);
+    d->vmsd = vmsd;
+    d->opaque = opaque;
+    d->version_id = version_id;
+    QSIMPLEQ_INSERT_TAIL(&vmstate_deferred_post_loads, d, entry);
+
+    return true;
+}
+
+void vmstate_post_load_defer_begin(void)
+{
+    assert(!vmstate_defer_post_load);
+    assert(QSIMPLEQ_EMPTY(&vmstate_deferred_post_loads));
+    vmstate_defer_post_load = true;
+}
+
+bool vmstate_post_load_defer_finish(bool run, Error **errp)
+{
+    VMStateDeferredPostLoad *d;
+    bool ok = true;
+
+    vmstate_defer_post_load = false;
+
+    while ((d = QSIMPLEQ_FIRST(&vmstate_deferred_post_loads))) {
+        QSIMPLEQ_REMOVE_HEAD(&vmstate_deferred_post_loads, entry);
+        if (run) {
+            ERRP_GUARD();
+
+            if (!vmstate_do_post_load(d->vmsd, d->opaque, d->version_id,
+                                      errp)) {
+                error_prepend(errp, "deferrable post load hook failed, which "
+                              "its vmsd promised could not happen: ");
+                error_report_err(*errp);
+                *errp = NULL;
+                abort();
+            }
+        }
+        g_free(d);
+    }
+
+    return ok;
+}
+
 /*
  * Try to prepare loading the next element, the object pointer to be put
  * into @next_elem.  When @next_elem is NULL, it means we should skip
diff --git a/tests/unit/test-vmstate.c b/tests/unit/test-vmstate.c
index df1fb4c778..e1e64b23c2 100644
--- a/tests/unit/test-vmstate.c
+++ b/tests/unit/test-vmstate.c
@@ -1619,6 +1619,116 @@ static void test_tmp_struct(void)
     g_assert_cmpint(obj.f, ==, 8); /* From the child->parent */
 }
 
+/* Deferred post_load */
+
+static int defer_order;
+static int defer_a_ran;
+static int defer_b_ran;
+static int defer_inline_ran;
+
+static int defer_a_post_load(void *opaque, int version_id)
+{
+    defer_a_ran = ++defer_order;
+    return 0;
+}
+
+static int defer_b_post_load(void *opaque, int version_id)
+{
+    defer_b_ran = ++defer_order;
+    return 0;
+}
+
+static int defer_inline_post_load(void *opaque, int version_id)
+{
+    defer_inline_ran = ++defer_order;
+    return 0;
+}
+
+static const VMStateDescription vmstate_defer_a = {
+    .name = "test/defer_a",
+    .version_id = 1,
+    .post_load = defer_a_post_load,
+    .post_load_deferrable = true,
+    .fields = (const VMStateField[]) {
+        VMSTATE_UINT32(a, TestStruct),
+        VMSTATE_END_OF_LIST()
+    }
+};
+
+static const VMStateDescription vmstate_defer_b = {
+    .name = "test/defer_b",
+    .version_id = 1,
+    .post_load = defer_b_post_load,
+    .post_load_deferrable = true,
+    .fields = (const VMStateField[]) {
+        VMSTATE_UINT32(a, TestStruct),
+        VMSTATE_END_OF_LIST()
+    }
+};
+
+static const VMStateDescription vmstate_defer_inline = {
+    .name = "test/defer_inline",
+    .version_id = 1,
+    .post_load = defer_inline_post_load,
+    .fields = (const VMStateField[]) {
+        VMSTATE_UINT32(a, TestStruct),
+        VMSTATE_END_OF_LIST()
+    }
+};
+
+static void defer_reset(void)
+{
+    defer_order = 0;
+    defer_a_ran = 0;
+    defer_b_ran = 0;
+    defer_inline_ran = 0;
+}
+
+static void test_post_load_defer(void)
+{
+    uint8_t const wire[] = {
+        /* uint32 a */ 0x00, 0x00, 0x00, 0x01,
+        QEMU_VM_EOF,
+    };
+    TestStruct obj;
+    Error *err = NULL;
+
+    /* Not deferring: a deferrable hook still runs as the section is read */
+    defer_reset();
+    memset(&obj, 0, sizeof(obj));
+    SUCCESS(load_vmstate_one(&vmstate_defer_a, &obj, 1, wire, sizeof(wire)));
+    g_assert_cmpint(defer_a_ran, ==, 1);
+
+    /*
+     * Deferring holds back the hooks which opted in and replays them in the
+     * order they would have run. A hook which did not opt in is unaffected.
+     */
+    defer_reset();
+    memset(&obj, 0, sizeof(obj));
+    vmstate_post_load_defer_begin();
+    SUCCESS(load_vmstate_one(&vmstate_defer_a, &obj, 1, wire, sizeof(wire)));
+    SUCCESS(load_vmstate_one(&vmstate_defer_inline, &obj, 1, wire,
+                             sizeof(wire)));
+    SUCCESS(load_vmstate_one(&vmstate_defer_b, &obj, 1, wire, sizeof(wire)));
+    g_assert_cmpint(defer_a_ran, ==, 0);
+    g_assert_cmpint(defer_b_ran, ==, 0);
+    g_assert_cmpint(defer_inline_ran, ==, 1);
+
+    g_assert(vmstate_post_load_defer_finish(true, &err));
+    g_assert(!err);
+    g_assert_cmpint(defer_a_ran, ==, 2);
+    g_assert_cmpint(defer_b_ran, ==, 3);
+
+    /* A discarded queue runs nothing */
+    defer_reset();
+    memset(&obj, 0, sizeof(obj));
+    vmstate_post_load_defer_begin();
+    SUCCESS(load_vmstate_one(&vmstate_defer_a, &obj, 1, wire, sizeof(wire)));
+    g_assert(vmstate_post_load_defer_finish(false, &err));
+    g_assert(!err);
+    g_assert_cmpint(defer_a_ran, ==, 0);
+}
+
 int main(int argc, char **argv)
 {
     g_autofree char *temp_file = g_strdup_printf("%s/vmst.test.XXXXXX",
@@ -1663,6 +1773,7 @@ int main(int argc, char **argv)
     g_test_add_func("/vmstate/qlist/save/saveqlist", test_save_qlist);
     g_test_add_func("/vmstate/qlist/load/loadqlist", test_load_qlist);
     g_test_add_func("/vmstate/tmp_struct", test_tmp_struct);
+    g_test_add_func("/vmstate/post_load/defer", test_post_load_defer);
     g_test_run();
 
     close(temp_fd);
-- 
2.53.0



^ permalink raw reply related	[flat|nested] 3+ messages in thread

* [PATCH 2/2] target/i386: defer the Hyper-V SynIC post_load
  2026-09-11 14:23 [PATCH 0/2] migration: defer a post_load which only rearranges memory Denis V. Lunev
  2026-09-11 14:23 ` [PATCH 1/2] migration: let a vmstate defer its post_load to end of stream Denis V. Lunev
@ 2026-09-11 14:23 ` Denis V. Lunev
  1 sibling, 0 replies; 3+ messages in thread
From: Denis V. Lunev @ 2026-09-11 14:23 UTC (permalink / raw)
  To: qemu-devel
  Cc: den, Peter Xu, Fabiano Rosas, Paolo Bonzini, Zhao Liu,
	Denis V. Lunev

From: Denis V. Lunev <den@openvz.org>

hyperv_synic_post_load() maps the SynIC message and event pages, and a
guest which has programmed the SynIC carries the subsection on every
vCPU, so restoring one rebuilds the flatview per vCPU rather than once.

The hook cannot fail and does not read guest memory, so it can run with
the rest of the deferred hooks once the stream has been consumed. The
saving scales with vCPU count.

CC: Peter Xu <peterx@redhat.com>
CC: Fabiano Rosas <farosas@suse.de>
CC: Paolo Bonzini <pbonzini@redhat.com>
CC: Zhao Liu <zhao1.liu@intel.com>
Signed-off-by: Denis V. Lunev <den@virtuozzo.com>
---
 target/i386/machine.c | 1 +
 1 file changed, 1 insertion(+)

diff --git a/target/i386/machine.c b/target/i386/machine.c
index 8d69d7e25e..eae74e33de 100644
--- a/target/i386/machine.c
+++ b/target/i386/machine.c
@@ -873,6 +873,7 @@ static const VMStateDescription vmstate_msr_hyperv_synic = {
     .minimum_version_id = 1,
     .needed = hyperv_synic_enable_needed,
     .post_load = hyperv_synic_post_load,
+    .post_load_deferrable = true,
     .fields = (const VMStateField[]) {
         VMSTATE_UINT64(env.msr_hv_synic_control, X86CPU),
         VMSTATE_UINT64(env.msr_hv_synic_evt_page, X86CPU),
-- 
2.53.0



^ permalink raw reply related	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-09-11 14:24 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-11 14:23 [PATCH 0/2] migration: defer a post_load which only rearranges memory Denis V. Lunev
2026-09-11 14:23 ` [PATCH 1/2] migration: let a vmstate defer its post_load to end of stream Denis V. Lunev
2026-09-11 14:23 ` [PATCH 2/2] target/i386: defer the Hyper-V SynIC post_load Denis V. Lunev

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.