All of lore.kernel.org
 help / color / mirror / Atom feed
* [RFC PATCH v2 0/9] igb: Add experimental VF live migration support
@ 2026-09-02 19:20 Cédric Le Goater
  2026-09-02 19:20 ` [RFC PATCH v2 1/9] igb: Add x-vf-migration property and DVSEC extended capability Cédric Le Goater
                   ` (9 more replies)
  0 siblings, 10 replies; 16+ messages in thread
From: Cédric Le Goater @ 2026-09-02 19:20 UTC (permalink / raw)
  To: qemu-devel
  Cc: Akihiko Odaki, Sriram Yagnaraman, Jason Wang, Alex Williamson,
	Peter Xu, Cédric Le Goater

Hello,

Live migration of VFIO-passthrough devices - SR-IOV VFs, vGPUs - is a
growing requirement, but real hardware with migration support is
scarce and hard to debug. An emulated device provides a fully
controlled testbed for developing and validating the entire software
stack - vfio-pci variant drivers, VFIO core migration v2 framework,
QEMU, libvirt - and for tuning complex migration policies such as
downtime convergence. It also serves as an educational reference for
understanding VFIO migration end-to-end, from device state
serialization to dirty page tracking.

This series adds an experimental VF live migration interface to the
emulated igb (82576) device. It enables a vfio-pci variant driver
(igb-vfio-pci) to migrate VFs using the standard VFIO migration v2
protocol with stop-copy and pre-copy support.

The target scenario is nested virtualization:

  L0 QEMU (these patches)
    igb PF with x-vf-migration=on
    └── VFs with migration DVSEC

  L1 kernel
    igb-vfio-pci variant driver [1]
    translates VFIO migration v2 ioctls → DVSEC config writes

  L1 QEMU (stock, unmodified)
    vfio-pci device model, standard migration fd

  L2 guest
    standard igbvf driver, unaware of migration

The L1 QEMU is completely unmodified -- it sees a standard VFIO
migratable device and uses the normal migration fd path.

* Design

The migration interface is exposed through a DVSEC (Designated
Vendor-Specific Extended Capability, PCIe cap id 0x23) at offset
0x160 in VF extended config space. The DVSEC uses a command doorbell
model - all commands are synchronous via PCI config space writes.

Device state is serialized as a versioned blob of per-VF register
(offset, value) pairs covering control, interrupt, RX/TX queue,
receive address (RA/RA2), etc. plus TX context descriptors and
VFRE/VFTE enable bits. The buffer address is a guest physical
address (GPA) written by the driver via virt_to_phys; the device
accesses guest RAM directly through the system address space.

Dirty page tracking is implemented with per-range bitmaps maintained
in IGBCore. All VF DMA paths in igb_core.c (TX data, RX data,
descriptor writeback) are instrumented to record touched pages. The
variant driver registers tracked IOVA ranges and queries dirty bitmaps
through a shared buffer. Buffer structures include len, flags, and
reserved fields for future extensibility.

* Caveats

The x-vf-migration property is experimental (x- prefix, default off).

The dirty bitmaps are maintained inside the device, which is not
realistic for discrete NICs without on-chip DRAM.

* Testing

The target scenario is nested virtualization: L0 runs QEMU with an
igb PF (x-vf-migration=on), L1 runs the igb-vfio-pci variant driver
and an unmodified QEMU, and L2 runs a standard igbvf driver.

Migration under iperf3 load works correctly: dirty page tracking
converges (from ~2000 pages per PRE_COPY iteration down to ~280 at
STOP_COPY), and STOP_COPY stays under 250ms.

* Todo

  1. Add migration blocker when x-vf-migration=on (no VMState yet) or
     add VMState support for L0 migration (dirty bitmaps, tracking
     engines, DVSEC registers, stats)
  2. Add PRE_COPY state transfer to validate device INIT data (magic,
     version, etc.)
  3. Add qtests for migration state machine transitions, dirty page
     tracking ?

* Ideas

  1. RX bandwidth throttle (x-mig-rx-limit, uint32, default 0)

     Return false from can_receive when the per-VF packet count in the
     current tracking interval exceeds the limit. Reduces DMA writes
     and dirty pages realistically.

  2. Migration phase timing (GET_STATS extension)

     Add per-VF timestamps: precopy_start_ns, stopcopy_start_ns,
     precopy_duration_ns, stopcopy_duration_ns,
     state_transition_count. Expose via GET_STATS.

  3. Hot page simulation (x-mig-hot-pages, uint32, default 0)

     Re-set the first N bitmap bits after each DIRTY_QUERY, simulating
     workloads with hot pages that prevent convergence.

  4. Error injection (x-mig-inject-error, uint32, default 0)

     One-shot error code injection before command dispatch. A separate
     x-mig-inject-dma-fail (bool) for persistent DMA failure testing.

* Credits

Alex Williamson suggested the overall approach of a variant driver
with the "x-vf-migration" device property to gate the feature. Thanks
for the ever ongoing support and valuable discussions throughout these
years.

* AI disclaimer

The lack of a migration-capable device has been a recurring pain point
for VFIO development over the years, and we hope this proposal
demonstrates the value of having one.

Claude was used to analyze the IGB PF and VF internal state and
identify the pain points of a working live migration of such devices.
The generated code served as a starting point but *significant* time
was then spent cleaning up, reworking, and shaping it into a clear,
reviewable IGB model extension.

As QEMU does not yet accept AI-assisted contributions, this series is
submitted as an RFC.

Thanks,

C.

[1] https://github.com/legoater/vfio-pci-extras

* Changes since rfc-v1

  - Migration BAR replaced with DVSEC at offset 0x160 (no BAR needed)
  - Wire-format structs (IgbMigBlob, IgbMigRegPair, IgbMigTxCtx)
    replace raw pointer arithmetic and memcpy
  - RA entries separated from fixed regs, scanned by pool bit
  - VFN relocation support (offset remapping + RA pool bit swap)
  - GPA buffer (address_space_read/write) replaces PCI DMA through PF
  - State blob validation on load (magic, version, error codes)
  - NEED_WORDS macro and igb_vf_offset_valid removed
  - Error codes renumbered: removed BAD_VFN, added UNK_CMD (1-11)

  Reported by Akihiko Odaki:

  - propagate_irqs: clear VF bits before OR (EIMS/EIAC/EIAM)
  - propagate_ivar: clear IVAR entry when source VTIVAR is invalid
  - rearm_irqs: restore actual PVTEICR causes, not all three
  - Dirty bitmap allocation uses BITS_TO_LONGS (heap corruption fix)
  - Dirty query validates range before g_malloc0 (memory exhaustion)
  - Dirty bits cleared only after successful bitmap DMA write
  - Dirty query buffer uses struct offsets (layout mismatch fix)
  - Load path: register offsets validated against VF whitelist
  - VMBMEM (mailbox payload) documented as transient, not serialized
  - dma_writes counter: consistently uint64_t
  - Dirty range_size: consistently uint64_t (was truncated to 32 bits)
  - ERROR->STOP: quiesce VF (clear VFRE/VFTE) on transition
  - rearm_irqs: runs on re || te, not just re (TX-only VF fix)
  - Stats DMA-written atomically via GET_STATS (no split MMIO tear)
  - Bisectability: DVSEC + state machine introduced together
  - Removed NAPI reference and "===" comment decoration

Cédric Le Goater (9):
  igb: Add x-vf-migration property and DVSEC extended capability
  igb: Add migration state machine via extended config space
  igb: Add VF state serialization for live migration
  igb: Add VF post-load fixups for live migration
  igb: Add dirty page tracking for IGBVF migration
  igb: Quiesce VFs on STOP and include PF enable state in migration
  igb: Fix post-migration RX ring deadlock
  igb: Add dirty page tracking statistics
  docs: Add igb VF migration testing setup guide

 MAINTAINERS                           |    6 +
 docs/system/device-emulation.rst      |    1 +
 docs/system/devices/igb-migration.rst |  417 +++++++++
 docs/system/devices/igb.rst           |    6 +
 hw/net/igb_common.h                   |   16 +
 hw/net/igb_core.h                     |   11 +
 hw/net/igb_migration.h                |  186 ++++
 hw/net/igb.c                          |    7 +
 hw/net/igb_core.c                     |  124 ++-
 hw/net/igb_migration.c                | 1128 +++++++++++++++++++++++++
 hw/net/igbvf.c                        |   52 +-
 hw/net/meson.build                    |    2 +-
 hw/net/trace-events                   |   13 +
 13 files changed, 1945 insertions(+), 24 deletions(-)
 create mode 100644 docs/system/devices/igb-migration.rst
 create mode 100644 hw/net/igb_migration.h
 create mode 100644 hw/net/igb_migration.c

-- 
2.55.0



^ permalink raw reply	[flat|nested] 16+ messages in thread

* [RFC PATCH v2 1/9] igb: Add x-vf-migration property and DVSEC extended capability
  2026-09-02 19:20 [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Cédric Le Goater
@ 2026-09-02 19:20 ` Cédric Le Goater
  2026-09-03 19:57   ` Alex Williamson
  2026-09-02 19:20 ` [RFC PATCH v2 2/9] igb: Add migration state machine via extended config space Cédric Le Goater
                   ` (8 subsequent siblings)
  9 siblings, 1 reply; 16+ messages in thread
From: Cédric Le Goater @ 2026-09-02 19:20 UTC (permalink / raw)
  To: qemu-devel
  Cc: Akihiko Odaki, Sriram Yagnaraman, Jason Wang, Alex Williamson,
	Peter Xu, Cédric Le Goater

Add an "x-vf-migration" property to the IGB PF device and expose a
DVSEC (Designated Vendor-Specific Extended Capability) at offset 0x160
in VF extended config space when migration is enabled.

The DVSEC provides the register interface for VF live migration:
 - CAPS:       supported features (state migration)
 - CTRL:       command doorbell
 - STATUS:     state and error reporting
 - BUF_ADDR:   shared DMA buffer address (GPA)

Move IgbVfState from igbvf.c to igb_common.h so it can be shared with
the migration module, and add the migration state field.

AI-used-for: code (prototype)
Signed-off-by: Cédric Le Goater <clg@redhat.com>
---
 MAINTAINERS            |   5 ++
 hw/net/igb_common.h    |  16 +++++++
 hw/net/igb_migration.h |  72 ++++++++++++++++++++++++++++
 hw/net/igb.c           |   2 +
 hw/net/igb_migration.c | 106 +++++++++++++++++++++++++++++++++++++++++
 hw/net/igbvf.c         |  52 ++++++++++++++++----
 hw/net/meson.build     |   2 +-
 7 files changed, 245 insertions(+), 10 deletions(-)
 create mode 100644 hw/net/igb_migration.h
 create mode 100644 hw/net/igb_migration.c

diff --git a/MAINTAINERS b/MAINTAINERS
index 4a49a40294eb..f88b526be238 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -2816,6 +2816,11 @@ F: tests/functional/x86_64/test_netdev_ethtool.py
 F: tests/qtest/igb-test.c
 F: tests/qtest/libqos/igb.c
 
+igb VF migration
+M: Cédric Le Goater <clg@redhat.com>
+S: Maintained
+F: hw/net/igb_migration.*
+
 eepro100
 M: Stefan Weil <sw@weilnetz.de>
 S: Maintained
diff --git a/hw/net/igb_common.h b/hw/net/igb_common.h
index b316a5bcfa5c..f0e2529e757b 100644
--- a/hw/net/igb_common.h
+++ b/hw/net/igb_common.h
@@ -26,7 +26,9 @@
 #ifndef HW_NET_IGB_COMMON_H
 #define HW_NET_IGB_COMMON_H
 
+#include "hw/pci/pci_device.h"
 #include "igb_regs.h"
+#include "igb_migration.h"
 
 #define TYPE_IGBVF "igbvf"
 
@@ -154,4 +156,18 @@ uint64_t igb_mmio_read(void *opaque, hwaddr addr, unsigned size);
 void igb_mmio_write(void *opaque, hwaddr addr, uint64_t val, unsigned size);
 void igb_vf_reset(void *opaque, uint16_t vfn);
 
+OBJECT_DECLARE_SIMPLE_TYPE(IgbVfState, IGBVF)
+
+struct IgbVfState {
+    PCIDevice parent_obj;
+
+    uint16_t vfn;
+    bool migration_enabled;
+
+    MemoryRegion mmio;
+    MemoryRegion msix;
+
+    IgbVfMigState mig;
+};
+
 #endif
diff --git a/hw/net/igb_migration.h b/hw/net/igb_migration.h
new file mode 100644
index 000000000000..3da28e11e49e
--- /dev/null
+++ b/hw/net/igb_migration.h
@@ -0,0 +1,72 @@
+/*
+ * QEMU Intel 82576 SR/IOV VF Migration Support
+ *
+ * Copyright (c) 2026 Red Hat, Inc.
+ *
+ * SPDX-License-Identifier: GPL-2.0-or-later
+ */
+
+#ifndef HW_NET_IGB_MIGRATION_H
+#define HW_NET_IGB_MIGRATION_H
+
+#include "hw/pci/pci_device.h"
+
+/*
+ * Migration interface exposed as a DVSEC (Designated Vendor-Specific
+ * Extended Capability) in VF extended config space.
+ *
+ * DVSEC layout at IGB_MIG_DVSEC_OFFSET (0x160):
+ *
+ *   +0x00  PCIe extended cap header   (cap_id=0x23, ver=1, next)
+ *   +0x04  DVSEC header 1             (len | rev | vendor_id)
+ *   +0x08  DVSEC header 2             (DVSEC ID)
+ *   +0x0A  Reserved                   (padding for DWORD alignment)
+ *   +0x0C  CAPS                       (RO: F_STATE[0])
+ *   +0x10  CTRL                       (WO: doorbell command)
+ *   +0x14  STATUS                     (RO: state[7:0], error_code[15:8])
+ *   +0x18  BUF_ADDR_LO                (RW: shared buffer GPA low)
+ *   +0x1C  BUF_ADDR_HI                (RW: shared buffer GPA high)
+ */
+
+#define IGB_MIG_DVSEC_OFFSET    0x160
+#define IGB_MIG_DVSEC_SIZE      0x20
+#define IGB_MIG_DVSEC_VER       1
+#define IGB_MIG_DVSEC_ID        1
+
+/* Register offsets relative to DVSEC base */
+#define IGB_MIG_CAPS            0x0C
+#define IGB_MIG_CTRL            0x10
+#define IGB_MIG_STATUS          0x14
+#define IGB_MIG_BUF_ADDR_LO     0x18
+#define IGB_MIG_BUF_ADDR_HI     0x1C
+
+/* CAPS register layout */
+#define IGB_MIG_CAP_F_STATE             (1u << 0)
+
+/* STATUS register: state in [7:0], error code [15:8] */
+#define IGB_MIG_STATUS_STATE_MASK       0xFF
+#define IGB_MIG_STATUS_ERROR_CODE_SHIFT 8
+#define IGB_MIG_STATUS_ERR(code) \
+    ((uint32_t)(code) << IGB_MIG_STATUS_ERROR_CODE_SHIFT)
+
+/* Device states (based on VFIO migration v2) */
+#define IGB_MIG_STATE_ERROR             0
+#define IGB_MIG_STATE_STOP              1
+#define IGB_MIG_STATE_RUNNING           2
+#define IGB_MIG_STATE_STOP_COPY         3
+#define IGB_MIG_STATE_RESUMING          4
+
+typedef struct IgbVfMigState {
+    uint32_t mig_state;
+    uint64_t mig_data_buf_addr;
+} IgbVfMigState;
+
+typedef struct IgbVfState IgbVfState;
+
+bool igbvf_add_migration_dvsec(PCIDevice *dev, Error **errp);
+void igbvf_mig_state_reset(IgbVfState *s);
+uint32_t igbvf_mig_config_read(IgbVfState *s, uint32_t addr, int size);
+bool igbvf_mig_config_write(IgbVfState *s, uint32_t addr, uint32_t val,
+                            int size);
+
+#endif
diff --git a/hw/net/igb.c b/hw/net/igb.c
index c076807e7110..7268e5473fc3 100644
--- a/hw/net/igb.c
+++ b/hw/net/igb.c
@@ -79,6 +79,7 @@ struct IGBState {
 
     IGBCore core;
     bool has_flr;
+    bool vf_migration;
 };
 
 #define IGB_CAP_SRIOV_OFFSET    (0x160)
@@ -597,6 +598,7 @@ static const VMStateDescription igb_vmstate = {
 static const Property igb_properties[] = {
     DEFINE_NIC_PROPERTIES(IGBState, conf),
     DEFINE_PROP_BOOL("x-pcie-flr-init", IGBState, has_flr, true),
+    DEFINE_PROP_BOOL("x-vf-migration", IGBState, vf_migration, false),
 };
 
 static void igb_class_init(ObjectClass *class, const void *data)
diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
new file mode 100644
index 000000000000..4dfebd82344c
--- /dev/null
+++ b/hw/net/igb_migration.c
@@ -0,0 +1,106 @@
+/*
+ * QEMU Intel 82576 SR/IOV VF Migration Support
+ *
+ * Copyright (c) 2026 Red Hat, Inc.
+ *
+ * SPDX-License-Identifier: GPL-2.0-or-later
+ */
+
+#include "qemu/osdep.h"
+#include "hw/pci/pci_device.h"
+#include "hw/pci/pcie.h"
+#include "igb_common.h"
+#include "igb_migration.h"
+
+static void igbvf_mig_update_status(IgbVfState *s, uint8_t err)
+{
+    IgbVfMigState *ms = &s->mig;
+    PCIDevice *dev = PCI_DEVICE(s);
+    uint32_t status;
+
+    status = ms->mig_state & IGB_MIG_STATUS_STATE_MASK;
+
+    if (err) {
+        status = IGB_MIG_STATE_ERROR | IGB_MIG_STATUS_ERR(err);
+    }
+
+    pci_set_long(dev->config + IGB_MIG_DVSEC_OFFSET + IGB_MIG_STATUS, status);
+}
+
+
+bool igbvf_add_migration_dvsec(PCIDevice *dev, Error **errp)
+{
+    uint16_t offset = IGB_MIG_DVSEC_OFFSET;
+    uint32_t caps;
+
+    pcie_add_capability(dev, PCI_EXT_CAP_ID_DVSEC, 1, offset,
+                        IGB_MIG_DVSEC_SIZE);
+
+    /* DVSEC header 1: length[31:20] | rev[19:16] | vendor_id[15:0] */
+    pci_set_long(dev->config + offset + 0x4,
+                 (IGB_MIG_DVSEC_SIZE << 20) |
+                 (IGB_MIG_DVSEC_VER << 16) |
+                 PCI_VENDOR_ID_INTEL);
+
+    /* DVSEC header 2: DVSEC ID */
+    pci_set_word(dev->config + offset + 0x8, IGB_MIG_DVSEC_ID);
+
+    /* CAPS: features (state migration only) */
+    caps = IGB_MIG_CAP_F_STATE;
+    pci_set_long(dev->config + offset + IGB_MIG_CAPS, caps);
+
+    /* STATUS: initial state is RUNNING */
+    pci_set_long(dev->config + offset + IGB_MIG_STATUS,
+                 IGB_MIG_STATE_RUNNING);
+
+    /* BUF_ADDR_LO and BUF_ADDR_HI are writable */
+    memset(dev->wmask + offset + IGB_MIG_BUF_ADDR_LO, 0xff, 4);
+    memset(dev->wmask + offset + IGB_MIG_BUF_ADDR_HI, 0xff, 4);
+
+    return true;
+}
+
+uint32_t igbvf_mig_config_read(IgbVfState *s, uint32_t addr, int size)
+{
+    PCIDevice *dev = PCI_DEVICE(s);
+
+    return pci_default_read_config(dev, addr, size);
+}
+
+bool igbvf_mig_config_write(IgbVfState *s, uint32_t addr, uint32_t val,
+                            int size)
+{
+    PCIDevice *dev = PCI_DEVICE(s);
+    uint32_t offset = addr - IGB_MIG_DVSEC_OFFSET;
+
+    switch (offset) {
+    case IGB_MIG_CTRL:
+        /* Command handling will be added in a later commit */
+        break;
+
+    case IGB_MIG_BUF_ADDR_LO:
+    case IGB_MIG_BUF_ADDR_HI:
+        pci_default_write_config(dev, addr, val, size);
+        break;
+
+    default:
+        break;
+    }
+
+    return true;
+}
+
+void igbvf_mig_state_reset(IgbVfState *s)
+{
+    IgbVfMigState *ms = &s->mig;
+
+    ms->mig_state = IGB_MIG_STATE_RUNNING;
+    ms->mig_data_buf_addr = 0;
+
+    pci_set_long(PCI_DEVICE(s)->config +
+                 IGB_MIG_DVSEC_OFFSET + IGB_MIG_BUF_ADDR_LO, 0);
+    pci_set_long(PCI_DEVICE(s)->config +
+                 IGB_MIG_DVSEC_OFFSET + IGB_MIG_BUF_ADDR_HI, 0);
+
+    igbvf_mig_update_status(s, 0);
+}
diff --git a/hw/net/igbvf.c b/hw/net/igbvf.c
index 9a165c7063ee..30dfdb574ac7 100644
--- a/hw/net/igbvf.c
+++ b/hw/net/igbvf.c
@@ -38,27 +38,21 @@
  */
 
 #include "qemu/osdep.h"
+#include "qemu/range.h"
 #include "hw/core/hw-error.h"
 #include "hw/net/mii.h"
 #include "hw/pci/pci_device.h"
 #include "hw/pci/pcie.h"
+#include "hw/pci/pcie_sriov.h"
 #include "hw/pci/msix.h"
 #include "net/eth.h"
 #include "net/net.h"
 #include "igb_common.h"
 #include "igb_core.h"
+#include "igb_migration.h"
 #include "trace.h"
 #include "qapi/error.h"
 
-OBJECT_DECLARE_SIMPLE_TYPE(IgbVfState, IGBVF)
-
-struct IgbVfState {
-    PCIDevice parent_obj;
-
-    MemoryRegion mmio;
-    MemoryRegion msix;
-};
-
 static hwaddr vf_to_pf_addr(hwaddr addr, uint16_t vfn, bool write)
 {
     switch (addr) {
@@ -199,10 +193,35 @@ static hwaddr vf_to_pf_addr(hwaddr addr, uint16_t vfn, bool write)
     return HWADDR_MAX;
 }
 
+static bool igbvf_addr_in_dvsec(uint32_t addr, int len)
+{
+    return ranges_overlap(addr, len,
+                          IGB_MIG_DVSEC_OFFSET, IGB_MIG_DVSEC_SIZE);
+}
+
+static uint32_t igbvf_read_config(PCIDevice *dev, uint32_t addr, int size)
+{
+    IgbVfState *s = IGBVF(dev);
+
+    if (s->migration_enabled && igbvf_addr_in_dvsec(addr, size)) {
+        return igbvf_mig_config_read(s, addr, size);
+    }
+
+    return pci_default_read_config(dev, addr, size);
+}
+
 static void igbvf_write_config(PCIDevice *dev, uint32_t addr, uint32_t val,
     int len)
 {
+    IgbVfState *s = IGBVF(dev);
+
     trace_igbvf_write_config(addr, val, len);
+
+    if (s->migration_enabled && igbvf_addr_in_dvsec(addr, len)) {
+        igbvf_mig_config_write(s, addr, val, len);
+        return;
+    }
+
     pci_default_write_config(dev, addr, val, len);
     if (object_property_get_bool(OBJECT(pcie_sriov_get_pf(dev)),
                                  "x-pcie-flr-init", &error_abort)) {
@@ -282,13 +301,27 @@ static void igbvf_pci_realize(PCIDevice *dev, Error **errp)
     }
 
     pcie_ari_init(dev, 0x150);
+
+    if (object_property_get_bool(OBJECT(pcie_sriov_get_pf(dev)),
+                                 "x-vf-migration", &error_abort)) {
+        s->vfn = pcie_sriov_vf_number(dev);
+        s->migration_enabled = true;
+        if (!igbvf_add_migration_dvsec(dev, errp)) {
+            return;
+        }
+    }
 }
 
 static void igbvf_qdev_reset_hold(Object *obj, ResetType type)
 {
     PCIDevice *vf = PCI_DEVICE(obj);
+    IgbVfState *s = IGBVF(vf);
 
     igb_vf_reset(pcie_sriov_get_pf(vf), pcie_sriov_vf_number(vf));
+
+    if (s->migration_enabled) {
+        igbvf_mig_state_reset(s);
+    }
 }
 
 static void igbvf_pci_uninit(PCIDevice *dev)
@@ -309,6 +342,7 @@ static void igbvf_class_init(ObjectClass *class, const void *data)
 
     c->realize = igbvf_pci_realize;
     c->exit = igbvf_pci_uninit;
+    c->config_read = igbvf_read_config;
     c->vendor_id = PCI_VENDOR_ID_INTEL;
     c->device_id = E1000_DEV_ID_82576_VF;
     c->revision = 1;
diff --git a/hw/net/meson.build b/hw/net/meson.build
index 84f142df222a..bb4b449b25ba 100644
--- a/hw/net/meson.build
+++ b/hw/net/meson.build
@@ -11,7 +11,7 @@ system_ss.add(when: 'CONFIG_E1000_PCI', if_true: files('e1000.c', 'e1000x_common
 system_ss.add(when: 'CONFIG_E1000E_PCI_EXPRESS', if_true: files('net_tx_pkt.c', 'net_rx_pkt.c'))
 system_ss.add(when: 'CONFIG_E1000E_PCI_EXPRESS', if_true: files('e1000e.c', 'e1000e_core.c', 'e1000x_common.c'))
 system_ss.add(when: 'CONFIG_IGB_PCI_EXPRESS', if_true: files('net_tx_pkt.c', 'net_rx_pkt.c'))
-system_ss.add(when: 'CONFIG_IGB_PCI_EXPRESS', if_true: files('igb.c', 'igbvf.c', 'igb_core.c'))
+system_ss.add(when: 'CONFIG_IGB_PCI_EXPRESS', if_true: files('igb.c', 'igbvf.c', 'igb_core.c', 'igb_migration.c'))
 system_ss.add(when: 'CONFIG_RTL8139_PCI', if_true: files('rtl8139.c'))
 system_ss.add(when: 'CONFIG_TULIP', if_true: files('tulip.c'))
 system_ss.add(when: 'CONFIG_VMXNET3_PCI', if_true: files('net_tx_pkt.c', 'net_rx_pkt.c'))
-- 
2.55.0



^ permalink raw reply related	[flat|nested] 16+ messages in thread

* [RFC PATCH v2 2/9] igb: Add migration state machine via extended config space
  2026-09-02 19:20 [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Cédric Le Goater
  2026-09-02 19:20 ` [RFC PATCH v2 1/9] igb: Add x-vf-migration property and DVSEC extended capability Cédric Le Goater
@ 2026-09-02 19:20 ` Cédric Le Goater
  2026-09-02 19:20 ` [RFC PATCH v2 3/9] igb: Add VF state serialization for live migration Cédric Le Goater
                   ` (7 subsequent siblings)
  9 siblings, 0 replies; 16+ messages in thread
From: Cédric Le Goater @ 2026-09-02 19:20 UTC (permalink / raw)
  To: qemu-devel
  Cc: Akihiko Odaki, Sriram Yagnaraman, Jason Wang, Alex Williamson,
	Peter Xu, Cédric Le Goater

Implement the VFIO migration v2 state machine and command dispatch
through the CTRL doorbell register in the DVSEC capability.

Add three CTRL commands:
 - SET_STATE: transitions follow the VFIO migration protocol
     RUNNING <-> STOP
     STOP <-> STOP_COPY
     STOP <-> RESUMING
     ERROR -> STOP
 - SAVE: DMA-writes the device state to a guest buffer
 - LOAD: DMA-reads it back

Add a DATA_SIZE register at +0x20 to report the state blob size. The
save/load stubs are filled in by the next patch.

AI-used-for: code (prototype)
Signed-off-by: Cédric Le Goater <clg@redhat.com>
---
 hw/net/igb_migration.h |  25 ++++-
 hw/net/igb_migration.c | 217 ++++++++++++++++++++++++++++++++++++++++-
 hw/net/trace-events    |   6 ++
 3 files changed, 246 insertions(+), 2 deletions(-)

diff --git a/hw/net/igb_migration.h b/hw/net/igb_migration.h
index 3da28e11e49e..ea40ac65c54b 100644
--- a/hw/net/igb_migration.h
+++ b/hw/net/igb_migration.h
@@ -26,10 +26,11 @@
  *   +0x14  STATUS                     (RO: state[7:0], error_code[15:8])
  *   +0x18  BUF_ADDR_LO                (RW: shared buffer GPA low)
  *   +0x1C  BUF_ADDR_HI                (RW: shared buffer GPA high)
+ *   +0x20  DATA_SIZE                  (RO: max state blob size in bytes)
  */
 
 #define IGB_MIG_DVSEC_OFFSET    0x160
-#define IGB_MIG_DVSEC_SIZE      0x20
+#define IGB_MIG_DVSEC_SIZE      0x24
 #define IGB_MIG_DVSEC_VER       1
 #define IGB_MIG_DVSEC_ID        1
 
@@ -39,10 +40,20 @@
 #define IGB_MIG_STATUS          0x14
 #define IGB_MIG_BUF_ADDR_LO     0x18
 #define IGB_MIG_BUF_ADDR_HI     0x1C
+#define IGB_MIG_DATA_SIZE       0x20
 
 /* CAPS register layout */
 #define IGB_MIG_CAP_F_STATE             (1u << 0)
 
+/* CTRL register: command in [7:0] */
+#define IGB_MIG_CTRL_CMD_MASK           0xFF
+#define IGB_MIG_CTRL_ARG_SHIFT          8
+
+/* CTRL commands */
+#define IGB_MIG_CMD_SET_STATE           1
+#define IGB_MIG_CMD_SAVE                2
+#define IGB_MIG_CMD_LOAD                3
+
 /* STATUS register: state in [7:0], error code [15:8] */
 #define IGB_MIG_STATUS_STATE_MASK       0xFF
 #define IGB_MIG_STATUS_ERROR_CODE_SHIFT 8
@@ -56,8 +67,20 @@
 #define IGB_MIG_STATE_STOP_COPY         3
 #define IGB_MIG_STATE_RESUMING          4
 
+/* Error codes */
+#define IGB_MIG_ERR_UNK_CMD             1
+#define IGB_MIG_ERR_BAD_STATE           2
+#define IGB_MIG_ERR_NO_BUFFER           3
+#define IGB_MIG_ERR_DMA_FAILED          4
+#define IGB_MIG_ERR_BAD_SIZE            5
+
+/* Shared buffer constants */
+#define IGB_VF_STATE_MAX_SIZE           4096
+
 typedef struct IgbVfMigState {
     uint32_t mig_state;
+    uint32_t mig_data[IGB_VF_STATE_MAX_SIZE / sizeof(uint32_t)];
+    uint32_t mig_data_size;
     uint64_t mig_data_buf_addr;
 } IgbVfMigState;
 
diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
index 4dfebd82344c..8e7e6fac9b9b 100644
--- a/hw/net/igb_migration.c
+++ b/hw/net/igb_migration.c
@@ -7,10 +7,177 @@
  */
 
 #include "qemu/osdep.h"
+#include "qemu/log.h"
 #include "hw/pci/pci_device.h"
 #include "hw/pci/pcie.h"
 #include "igb_common.h"
 #include "igb_migration.h"
+#include "system/address-spaces.h"
+#include "trace.h"
+
+/*
+ * Per-VF state serialization / deserialization
+ */
+
+static int igb_core_vf_save_state(IgbVfState *s, void *buf, size_t buf_size)
+{
+    int size = 0;
+
+    trace_igbvf_mig_save_state(s->vfn, size);
+    return size;
+}
+
+static int igb_core_vf_max_data_size(IgbVfState *s)
+{
+    return sizeof(s->mig.mig_data);
+}
+
+static int igb_core_vf_load_state(IgbVfState *s, const void *buf, size_t size)
+{
+    trace_igbvf_mig_load_state(s->vfn, (uint32_t)size);
+    return 0;
+}
+
+static int igbvf_mig_load(IgbVfState *s, const void *buf, size_t size)
+{
+    int ret;
+
+    ret = igb_core_vf_load_state(s, buf, size);
+    if (ret < 0) {
+        return ret;
+    }
+
+    return 0;
+}
+
+/*
+ * Migration command handlers
+ */
+
+static void igbvf_mig_update_data_size(IgbVfState *s, uint32_t size)
+{
+    IgbVfMigState *ms = &s->mig;
+
+    ms->mig_data_size = size;
+    pci_set_long(PCI_DEVICE(s)->config +
+                 IGB_MIG_DVSEC_OFFSET + IGB_MIG_DATA_SIZE, size);
+}
+
+static uint8_t igbvf_mig_cmd_save(IgbVfState *s)
+{
+    IgbVfMigState *ms = &s->mig;
+    MemTxResult r;
+    int ret;
+
+    if (ms->mig_state != IGB_MIG_STATE_STOP_COPY) {
+        return IGB_MIG_ERR_BAD_STATE;
+    }
+
+    if (!ms->mig_data_buf_addr) {
+        return IGB_MIG_ERR_NO_BUFFER;
+    }
+
+    ret = igb_core_vf_save_state(s, ms->mig_data, sizeof(ms->mig_data));
+    if (ret < 0) {
+        return -ret;
+    }
+    igbvf_mig_update_data_size(s, ret);
+
+    r = address_space_write(&address_space_memory, ms->mig_data_buf_addr,
+                            MEMTXATTRS_UNSPECIFIED,
+                            ms->mig_data, ms->mig_data_size);
+    if (r != MEMTX_OK) {
+        return IGB_MIG_ERR_DMA_FAILED;
+    }
+
+    return 0;
+}
+
+static uint8_t igbvf_mig_cmd_load(IgbVfState *s)
+{
+    IgbVfMigState *ms = &s->mig;
+    MemTxResult r;
+    int ret;
+
+    if (ms->mig_state != IGB_MIG_STATE_RESUMING) {
+        return IGB_MIG_ERR_BAD_STATE;
+    }
+
+    if (!ms->mig_data_buf_addr) {
+        return IGB_MIG_ERR_NO_BUFFER;
+    }
+
+    if (ms->mig_data_size == 0 ||
+        ms->mig_data_size > sizeof(ms->mig_data)) {
+        return IGB_MIG_ERR_BAD_SIZE;
+    }
+
+    r = address_space_read(&address_space_memory, ms->mig_data_buf_addr,
+                           MEMTXATTRS_UNSPECIFIED,
+                           ms->mig_data, ms->mig_data_size);
+    if (r != MEMTX_OK) {
+        return IGB_MIG_ERR_DMA_FAILED;
+    }
+
+    ret = igbvf_mig_load(s, ms->mig_data, ms->mig_data_size);
+    if (ret < 0) {
+        return -ret;
+    }
+
+    return 0;
+}
+
+static uint8_t igbvf_mig_set_state(IgbVfState *s, uint32_t new_state)
+{
+    IgbVfMigState *ms = &s->mig;
+    uint32_t old = ms->mig_state;
+    int ret;
+
+    switch (new_state) {
+    case IGB_MIG_STATE_STOP:
+        if (old != IGB_MIG_STATE_RUNNING &&
+            old != IGB_MIG_STATE_STOP_COPY &&
+            old != IGB_MIG_STATE_RESUMING &&
+            old != IGB_MIG_STATE_ERROR) {
+            return IGB_MIG_ERR_BAD_STATE;
+        }
+        /* Restore DATA_SIZE to max, same as at reset */
+        igbvf_mig_update_data_size(s, igb_core_vf_max_data_size(s));
+        break;
+
+    case IGB_MIG_STATE_RUNNING:
+        if (old != IGB_MIG_STATE_STOP) {
+            return IGB_MIG_ERR_BAD_STATE;
+        }
+        break;
+
+    case IGB_MIG_STATE_STOP_COPY:
+        if (old != IGB_MIG_STATE_STOP) {
+            return IGB_MIG_ERR_BAD_STATE;
+        }
+        ret = igb_core_vf_save_state(s, ms->mig_data, sizeof(ms->mig_data));
+        if (ret < 0) {
+            return -ret;
+        }
+        igbvf_mig_update_data_size(s, ret);
+        break;
+
+    case IGB_MIG_STATE_RESUMING:
+        if (old != IGB_MIG_STATE_STOP) {
+            return IGB_MIG_ERR_BAD_STATE;
+        }
+        memset(ms->mig_data, 0, sizeof(ms->mig_data));
+        igbvf_mig_update_data_size(s, 0);
+        break;
+
+    default:
+        return IGB_MIG_ERR_BAD_STATE;
+    }
+
+    ms->mig_state = new_state;
+    trace_igbvf_mig_set_state(s->vfn, old, new_state);
+    return 0;
+}
 
 static void igbvf_mig_update_status(IgbVfState *s, uint8_t err)
 {
@@ -27,6 +194,38 @@ static void igbvf_mig_update_status(IgbVfState *s, uint8_t err)
     pci_set_long(dev->config + IGB_MIG_DVSEC_OFFSET + IGB_MIG_STATUS, status);
 }
 
+static void igbvf_mig_cmd_ctrl(IgbVfState *s, uint32_t val)
+{
+    uint32_t cmd = val & IGB_MIG_CTRL_CMD_MASK;
+    uint32_t arg = val >> IGB_MIG_CTRL_ARG_SHIFT;
+    uint8_t err = 0;
+
+    switch (cmd) {
+    case IGB_MIG_CMD_SET_STATE:
+        err = igbvf_mig_set_state(s, arg);
+        break;
+
+    case IGB_MIG_CMD_SAVE:
+        err = igbvf_mig_cmd_save(s);
+        break;
+
+    case IGB_MIG_CMD_LOAD:
+        igbvf_mig_update_data_size(s, arg);
+        err = igbvf_mig_cmd_load(s);
+        break;
+
+    default:
+        err = IGB_MIG_ERR_UNK_CMD;
+        break;
+    }
+
+    if (err) {
+        qemu_log_mask(LOG_GUEST_ERROR,
+                      "igbvf: VF%u CTRL cmd %u failed (error %u)\n",
+                      s->vfn, cmd, err);
+    }
+    igbvf_mig_update_status(s, err);
+}
 
 bool igbvf_add_migration_dvsec(PCIDevice *dev, Error **errp)
 {
@@ -57,6 +256,8 @@ bool igbvf_add_migration_dvsec(PCIDevice *dev, Error **errp)
     memset(dev->wmask + offset + IGB_MIG_BUF_ADDR_LO, 0xff, 4);
     memset(dev->wmask + offset + IGB_MIG_BUF_ADDR_HI, 0xff, 4);
 
+    /* DATA_SIZE is set by igbvf_mig_state_reset() */
+
     return true;
 }
 
@@ -67,6 +268,16 @@ uint32_t igbvf_mig_config_read(IgbVfState *s, uint32_t addr, int size)
     return pci_default_read_config(dev, addr, size);
 }
 
+static uint64_t igbvf_mig_get_buf_addr(IgbVfState *s)
+{
+    PCIDevice *dev = PCI_DEVICE(s);
+    uint32_t lo, hi;
+
+    lo = pci_get_long(dev->config + IGB_MIG_DVSEC_OFFSET + IGB_MIG_BUF_ADDR_LO);
+    hi = pci_get_long(dev->config + IGB_MIG_DVSEC_OFFSET + IGB_MIG_BUF_ADDR_HI);
+    return ((uint64_t)hi << 32) | lo;
+}
+
 bool igbvf_mig_config_write(IgbVfState *s, uint32_t addr, uint32_t val,
                             int size)
 {
@@ -75,7 +286,8 @@ bool igbvf_mig_config_write(IgbVfState *s, uint32_t addr, uint32_t val,
 
     switch (offset) {
     case IGB_MIG_CTRL:
-        /* Command handling will be added in a later commit */
+        s->mig.mig_data_buf_addr = igbvf_mig_get_buf_addr(s);
+        igbvf_mig_cmd_ctrl(s, val);
         break;
 
     case IGB_MIG_BUF_ADDR_LO:
@@ -94,8 +306,11 @@ void igbvf_mig_state_reset(IgbVfState *s)
 {
     IgbVfMigState *ms = &s->mig;
 
+    trace_igbvf_mig_reset(s->vfn);
     ms->mig_state = IGB_MIG_STATE_RUNNING;
     ms->mig_data_buf_addr = 0;
+    igbvf_mig_update_data_size(s, igb_core_vf_max_data_size(s));
+    memset(ms->mig_data, 0, sizeof(ms->mig_data));
 
     pci_set_long(PCI_DEVICE(s)->config +
                  IGB_MIG_DVSEC_OFFSET + IGB_MIG_BUF_ADDR_LO, 0);
diff --git a/hw/net/trace-events b/hw/net/trace-events
index 001a20b0e2ac..7057fbe5f16a 100644
--- a/hw/net/trace-events
+++ b/hw/net/trace-events
@@ -295,6 +295,12 @@ igb_wrn_rx_desc_modes_not_supp(int desc_type) "Not supported descriptor type: %d
 # igbvf.c
 igbvf_wrn_io_addr_unknown(uint64_t addr) "IO unknown register 0x%"PRIx64
 
+# igb_migration.c
+igbvf_mig_set_state(uint16_t vfn, uint32_t old_state, uint32_t new_state) "VF%u: state %u -> %u"
+igbvf_mig_save_state(uint16_t vfn, uint32_t size) "VF%u: saved %u bytes of device state"
+igbvf_mig_load_state(uint16_t vfn, uint32_t size) "VF%u: loaded %u bytes of device state"
+igbvf_mig_reset(uint16_t vfn) "VF%u: migration state reset"
+
 # spapr_llan.c
 spapr_vlan_get_rx_bd_from_pool_found(int pool, int32_t count, uint32_t rx_bufs) "pool=%d count=%"PRId32" rxbufs=%"PRIu32
 spapr_vlan_get_rx_bd_from_page(int buf_ptr, uint64_t bd) "use_buf_ptr=%d bd=0x%016"PRIx64
-- 
2.55.0



^ permalink raw reply related	[flat|nested] 16+ messages in thread

* [RFC PATCH v2 3/9] igb: Add VF state serialization for live migration
  2026-09-02 19:20 [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Cédric Le Goater
  2026-09-02 19:20 ` [RFC PATCH v2 1/9] igb: Add x-vf-migration property and DVSEC extended capability Cédric Le Goater
  2026-09-02 19:20 ` [RFC PATCH v2 2/9] igb: Add migration state machine via extended config space Cédric Le Goater
@ 2026-09-02 19:20 ` Cédric Le Goater
  2026-09-08  8:01   ` Akihiko Odaki
  2026-09-02 19:20 ` [RFC PATCH v2 4/9] igb: Add VF post-load fixups " Cédric Le Goater
                   ` (6 subsequent siblings)
  9 siblings, 1 reply; 16+ messages in thread
From: Cédric Le Goater @ 2026-09-02 19:20 UTC (permalink / raw)
  To: qemu-devel
  Cc: Akihiko Odaki, Sriram Yagnaraman, Jason Wang, Alex Williamson,
	Peter Xu, Cédric Le Goater

Implement per-VF state serialization and deserialization for the SAVE
and LOAD commands. The wire format consists of a header (magic,
version, VF number, register count), per-VF register offset/value
pairs from a whitelist, RA table entries owned by the VF, and TX
queue contexts.

Register offsets are relocated on load so a VF can migrate to a
different VF number on the destination. RA pool ownership bits are
swapped accordingly.

AI-used-for: analysis, code (prototype)
Signed-off-by: Cédric Le Goater <clg@redhat.com>
---
 hw/net/igb_core.h      |   2 +
 hw/net/igb_migration.h |   2 +
 hw/net/igb.c           |   5 +
 hw/net/igb_migration.c | 347 ++++++++++++++++++++++++++++++++++++++++-
 4 files changed, 354 insertions(+), 2 deletions(-)

diff --git a/hw/net/igb_core.h b/hw/net/igb_core.h
index d70b54e318f1..60724e2824ab 100644
--- a/hw/net/igb_core.h
+++ b/hw/net/igb_core.h
@@ -143,4 +143,6 @@ igb_receive_iov(IGBCore *core, const struct iovec *iov, int iovcnt);
 void
 igb_start_recv(IGBCore *core);
 
+IGBCore *igb_pf_get_core(void *pf);
+
 #endif
diff --git a/hw/net/igb_migration.h b/hw/net/igb_migration.h
index ea40ac65c54b..b2f601e74346 100644
--- a/hw/net/igb_migration.h
+++ b/hw/net/igb_migration.h
@@ -73,6 +73,8 @@
 #define IGB_MIG_ERR_NO_BUFFER           3
 #define IGB_MIG_ERR_DMA_FAILED          4
 #define IGB_MIG_ERR_BAD_SIZE            5
+#define IGB_MIG_ERR_BAD_MAGIC           6
+#define IGB_MIG_ERR_BAD_VERSION         7
 
 /* Shared buffer constants */
 #define IGB_VF_STATE_MAX_SIZE           4096
diff --git a/hw/net/igb.c b/hw/net/igb.c
index 7268e5473fc3..f39f2bc3a04e 100644
--- a/hw/net/igb.c
+++ b/hw/net/igb.c
@@ -133,6 +133,11 @@ void igb_vf_reset(void *opaque, uint16_t vfn)
     igb_core_vf_reset(&s->core, vfn);
 }
 
+IGBCore *igb_pf_get_core(void *pf)
+{
+    return &IGB(pf)->core;
+}
+
 static bool
 igb_io_get_reg_index(IGBState *s, uint32_t *idx)
 {
diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
index 8e7e6fac9b9b..c34035974620 100644
--- a/hw/net/igb_migration.c
+++ b/hw/net/igb_migration.c
@@ -10,18 +10,241 @@
 #include "qemu/log.h"
 #include "hw/pci/pci_device.h"
 #include "hw/pci/pcie.h"
+#include "net/eth.h"
+#include "net/net.h"
 #include "igb_common.h"
+#include "igb_core.h"
 #include "igb_migration.h"
 #include "system/address-spaces.h"
 #include "trace.h"
 
+static IGBCore *igbvf_get_core(IgbVfState *s)
+{
+    return igb_pf_get_core(pcie_sriov_get_pf(PCI_DEVICE(s)));
+}
+
 /*
  * Per-VF state serialization / deserialization
  */
 
+#define IGB_MIG_BLOB_MAGIC        0x4D494742  /* "MIGB" */
+#define IGB_MIG_BLOB_VERSION      1
+
+typedef struct IgbMigRegPair {
+    uint32_t offset;
+    uint32_t value;
+} IgbMigRegPair;
+
+typedef struct IgbMigTxCtx {
+    uint32_t ctx_desc[8];         /* 2 × adv_tx_context_desc (4 dwords each) */
+    uint32_t first_cmd_type_len;
+    uint32_t first_olinfo_status;
+    uint32_t first;
+    uint32_t skip_cp;
+} IgbMigTxCtx;
+
+#define IGB_VF_MAX_FIXED_REGS     64
+#define IGB_VF_MAX_RA_REGS        48  /* (16 + 8) RA entries × 2 (RAL+RAH) */
+
+typedef struct IgbMigBlob {
+    uint32_t magic;
+    uint32_t version;
+    uint32_t vfn;
+    uint32_t num_regs;
+    IgbMigRegPair regs[IGB_VF_MAX_FIXED_REGS];
+    uint32_t num_ra;
+    IgbMigRegPair ra[IGB_VF_MAX_RA_REGS];
+    uint32_t num_tx_ctx;
+    IgbMigTxCtx tx_ctx[2];
+} IgbMigBlob;
+
+#define IGB_MIG_BLOB_SIZE            sizeof(IgbMigBlob)
+
+QEMU_BUILD_BUG_ON(IGB_MIG_BLOB_SIZE > IGB_VF_STATE_MAX_SIZE);
+
+/* Register offsets that constitute a VF's state slice */
+static int igb_vf_reg_list(uint16_t vfn, uint32_t *offsets)
+{
+    int n = 0;
+    int q0 = vfn;
+    int q1 = vfn + IGB_NUM_VM_POOLS;
+
+    /* Per-VF control and interrupt registers */
+    offsets[n++] = E1000_PVTCTRL(vfn) >> 2;
+    offsets[n++] = E1000_PVTEICS(vfn) >> 2;
+    offsets[n++] = E1000_PVTEIMS(vfn) >> 2;
+    offsets[n++] = E1000_PVTEIMC(vfn) >> 2;
+    offsets[n++] = E1000_PVTEIAC(vfn) >> 2;
+    offsets[n++] = E1000_PVTEIAM(vfn) >> 2;
+    offsets[n++] = E1000_PVTEICR(vfn) >> 2;
+
+    /* Per-VF statistics */
+    offsets[n++] = E1000_PVFGPRC(vfn) >> 2;
+    offsets[n++] = E1000_PVFGPTC(vfn) >> 2;
+    offsets[n++] = E1000_PVFGORC(vfn) >> 2;
+    offsets[n++] = E1000_PVFGOTC(vfn) >> 2;
+    offsets[n++] = E1000_PVFMPRC(vfn) >> 2;
+    offsets[n++] = E1000_PVFGPRLBC(vfn) >> 2;
+    offsets[n++] = E1000_PVFGPTLBC(vfn) >> 2;
+    offsets[n++] = E1000_PVFGORLBC(vfn) >> 2;
+    offsets[n++] = E1000_PVFGOTLBC(vfn) >> 2;
+
+    /*
+     * Mailbox control registers only - the 16-dword payload buffer
+     * (VMBMEM) is transient and drained on quiesce.
+     */
+    offsets[n++] = E1000_V2PMAILBOX(vfn) >> 2;
+    offsets[n++] = E1000_P2VMAILBOX(vfn) >> 2;
+
+    /* Per-VF config */
+    offsets[n++] = E1000_VMOLR(vfn) >> 2;
+    offsets[n++] = E1000_VMVIR(vfn) >> 2;
+    offsets[n++] = E1000_PSRTYPE(vfn) >> 2;
+
+    /*
+     * VF receive addresses (RA/RA2) are saved dynamically in
+     * igb_core_vf_save_state by scanning for entries whose pool
+     * bits match this VF - the PF driver chooses the RA slot.
+     */
+
+    /* Interrupt routing */
+    offsets[n++] = (E1000_VTIVAR + vfn * 4) >> 2;
+    offsets[n++] = (E1000_VTIVAR_MISC + vfn * 4) >> 2;
+
+    /*
+     * EITR (Extended Interrupt Throttle Register) - 3 vectors per VF.
+     * Each VF has 3 MSI-X vectors, each with its own EITR controlling
+     * interrupt coalescing. Without saving these, interrupt
+     * throttling resets to zero after migration which can cause
+     * interrupt storms or latency changes. VF N uses PF EITR indices
+     * (22 - N*3) .. (24 - N*3).
+     */
+    {
+        int eitr_base = 22 - vfn * 3;
+        offsets[n++] = E1000_EITR(eitr_base) >> 2;
+        offsets[n++] = E1000_EITR(eitr_base + 1) >> 2;
+        offsets[n++] = E1000_EITR(eitr_base + 2) >> 2;
+    }
+
+    /* RX and TX queue registers for queues q0 and q1 */
+#define ADD_QUEUE_REGS(q) do { \
+    offsets[n++] = E1000_RDBAL(q) >> 2; \
+    offsets[n++] = E1000_RDBAH(q) >> 2; \
+    offsets[n++] = E1000_RDLEN(q) >> 2; \
+    offsets[n++] = E1000_SRRCTL(q) >> 2; \
+    offsets[n++] = E1000_RDH(q) >> 2; \
+    offsets[n++] = E1000_RDT(q) >> 2; \
+    offsets[n++] = E1000_RXDCTL(q) >> 2; \
+    offsets[n++] = E1000_RXCTL(q) >> 2; \
+    offsets[n++] = E1000_RQDPC(q) >> 2; \
+    offsets[n++] = E1000_TDBAL(q) >> 2; \
+    offsets[n++] = E1000_TDBAH(q) >> 2; \
+    offsets[n++] = E1000_TDLEN(q) >> 2; \
+    offsets[n++] = E1000_TDH(q) >> 2; \
+    offsets[n++] = E1000_TDT(q) >> 2; \
+    offsets[n++] = E1000_TXDCTL(q) >> 2; \
+    offsets[n++] = E1000_TXCTL(q) >> 2; \
+    offsets[n++] = E1000_TDWBAL(q) >> 2; \
+    offsets[n++] = E1000_TDWBAH(q) >> 2; \
+} while (0)
+
+    ADD_QUEUE_REGS(q0);
+    ADD_QUEUE_REGS(q1);
+#undef ADD_QUEUE_REGS
+
+    g_assert(n <= IGB_VF_MAX_FIXED_REGS);
+    return n;
+}
+
+/*
+ * Scan RA and RA2 arrays for receive address entries assigned to
+ * this VF. The PF driver picks the RA slot, so we cannot use a
+ * fixed index - instead check each entry's pool bits.
+ */
+static int igb_core_vf_save_ra(IGBCore *core, uint16_t vfn,
+                               IgbMigRegPair *regs)
+{
+    uint32_t vf_pool_bit = E1000_RAH_POOL_1 << vfn;
+    int n = 0;
+    static const struct {
+        uint32_t base;
+        int count;
+    } ra_banks[] = {
+        { RA,  16 },
+        { RA2,  8 },
+    };
+
+    for (int i = 0; i < ARRAY_SIZE(ra_banks); i++) {
+        for (int j = 0; j < ra_banks[i].count; j++) {
+            uint32_t ral_off = ra_banks[i].base + j * 2;
+            uint32_t rah_off = ra_banks[i].base + j * 2 + 1;
+            uint32_t rah_val = core->mac[rah_off];
+
+            if ((rah_val & E1000_RAH_AV) && (rah_val & vf_pool_bit)) {
+                regs[n].offset = cpu_to_le32(ral_off);
+                regs[n].value = cpu_to_le32(core->mac[ral_off]);
+                n++;
+                regs[n].offset = cpu_to_le32(rah_off);
+                regs[n].value = cpu_to_le32(rah_val);
+                n++;
+            }
+        }
+    }
+    return n;
+}
+
+static void igb_core_vf_save_tx_ctx(IGBCore *core, int queue,
+                                    IgbMigTxCtx *tx)
+{
+    struct igb_tx *src = &core->tx[queue];
+
+    memcpy(tx->ctx_desc, src->ctx, sizeof(tx->ctx_desc));
+    tx->first_cmd_type_len = cpu_to_le32(src->first_cmd_type_len);
+    tx->first_olinfo_status = cpu_to_le32(src->first_olinfo_status);
+    tx->first = cpu_to_le32(src->first);
+    tx->skip_cp = cpu_to_le32(src->skip_cp);
+}
+
 static int igb_core_vf_save_state(IgbVfState *s, void *buf, size_t buf_size)
 {
-    int size = 0;
+    int size = IGB_MIG_BLOB_SIZE;
+    IGBCore *core = igbvf_get_core(s);
+    IgbMigBlob *blob = buf;
+    uint32_t offsets[IGB_VF_MAX_FIXED_REGS];
+    int num_regs;
+    int q0 = s->vfn;
+    int q1 = s->vfn + IGB_NUM_VM_POOLS;
+
+    /*
+     * Save PVT shadow registers (PVTEIMS/PVTEIAC/PVTEIAM) instead of
+     * extracting from PF aggregates - the L1 PF driver may have
+     * transiently cleared EIMS via EIMC. The load path ORs them back.
+     */
+    num_regs = igb_vf_reg_list(s->vfn, offsets);
+
+    if (!buf) {
+        return size;
+    }
+
+    if (size > buf_size) {
+        return -IGB_MIG_ERR_BAD_SIZE;
+    }
+
+    blob->magic = cpu_to_le32(IGB_MIG_BLOB_MAGIC);
+    blob->version = cpu_to_le32(IGB_MIG_BLOB_VERSION);
+    blob->vfn = cpu_to_le32(s->vfn);
+
+    blob->num_regs = cpu_to_le32(num_regs);
+    for (int i = 0; i < num_regs; i++) {
+        blob->regs[i].offset = cpu_to_le32(offsets[i]);
+        blob->regs[i].value = cpu_to_le32(core->mac[offsets[i]]);
+    }
+
+    blob->num_ra = cpu_to_le32(igb_core_vf_save_ra(core, s->vfn, blob->ra));
+
+    blob->num_tx_ctx = cpu_to_le32(2);
+    igb_core_vf_save_tx_ctx(core, q0, &blob->tx_ctx[0]);
+    igb_core_vf_save_tx_ctx(core, q1, &blob->tx_ctx[1]);
 
     trace_igbvf_mig_save_state(s->vfn, size);
     return size;
@@ -29,11 +252,131 @@ static int igb_core_vf_save_state(IgbVfState *s, void *buf, size_t buf_size)
 
 static int igb_core_vf_max_data_size(IgbVfState *s)
 {
-    return sizeof(s->mig.mig_data);
+    int size = igb_core_vf_save_state(s, NULL, 0);
+
+    g_assert(size > 0 && size <= IGB_VF_STATE_MAX_SIZE);
+    return size;
+}
+
+static void igb_core_vf_load_tx_ctx(IGBCore *core, int queue,
+                                    const IgbMigTxCtx *tx)
+{
+    struct igb_tx *dst = &core->tx[queue];
+
+    /*
+     * Preserve the destination's tx_pkt - it's a host-side object,
+     * not guest state
+     */
+    memcpy(dst->ctx, tx->ctx_desc, sizeof(dst->ctx));
+    dst->first_cmd_type_len = le32_to_cpu(tx->first_cmd_type_len);
+    dst->first_olinfo_status = le32_to_cpu(tx->first_olinfo_status);
+    dst->first = le32_to_cpu(tx->first);
+    dst->skip_cp = le32_to_cpu(tx->skip_cp);
+}
+
+static uint32_t igb_vf_relocate_offset(uint32_t offset,
+                                       const uint32_t *src_offsets,
+                                       const uint32_t *dst_offsets,
+                                       int num_offsets)
+{
+    for (int i = 0; i < num_offsets; i++) {
+        if (src_offsets[i] == offset) {
+            return dst_offsets[i];
+        }
+    }
+    return 0;
 }
 
 static int igb_core_vf_load_state(IgbVfState *s, const void *buf, size_t size)
 {
+    IGBCore *core = igbvf_get_core(s);
+    uint32_t src_offsets[IGB_VF_MAX_FIXED_REGS];
+    uint32_t dst_offsets[IGB_VF_MAX_FIXED_REGS];
+    int q0 = s->vfn;
+    int q1 = s->vfn + IGB_NUM_VM_POOLS;
+
+    if (size < IGB_MIG_BLOB_SIZE) {
+        return -IGB_MIG_ERR_BAD_SIZE;
+    }
+
+    const IgbMigBlob *blob = buf;
+
+    uint32_t magic = le32_to_cpu(blob->magic);
+    uint32_t version = le32_to_cpu(blob->version);
+    uint32_t saved_vfn = le32_to_cpu(blob->vfn);
+    uint32_t num_regs = le32_to_cpu(blob->num_regs);
+
+    if (magic != IGB_MIG_BLOB_MAGIC) {
+        return -IGB_MIG_ERR_BAD_MAGIC;
+    }
+    if (version != IGB_MIG_BLOB_VERSION) {
+        return -IGB_MIG_ERR_BAD_VERSION;
+    }
+    if (num_regs > IGB_VF_MAX_FIXED_REGS) {
+        return -IGB_MIG_ERR_BAD_SIZE;
+    }
+
+    uint32_t num_ra = le32_to_cpu(blob->num_ra);
+    if (num_ra > IGB_VF_MAX_RA_REGS) {
+        return -IGB_MIG_ERR_BAD_SIZE;
+    }
+
+    int num_offsets = igb_vf_reg_list(saved_vfn, src_offsets);
+    igb_vf_reg_list(s->vfn, dst_offsets);
+
+    for (uint32_t i = 0; i < num_regs; i++) {
+        uint32_t src_off = le32_to_cpu(blob->regs[i].offset);
+        uint32_t value = le32_to_cpu(blob->regs[i].value);
+        uint32_t offset = igb_vf_relocate_offset(src_off,
+            src_offsets, dst_offsets, num_offsets);
+        if (!offset) {
+            return -IGB_MIG_ERR_BAD_SIZE;
+        }
+
+        core->mac[offset] = value;
+
+        /*
+         * Sync EITR to eitr_guest_value[] shadow array, stripping
+         * E1000_EITR_CNT_IGNR so guest register readback returns the
+         * correct value.
+         */
+        if (offset >= EITR0 && offset < EITR0 + IGB_INTR_NUM) {
+            core->eitr_guest_value[offset - EITR0] =
+                value & ~E1000_EITR_CNT_IGNR;
+        }
+    }
+
+    /*
+     * MSI-X table/PBA is not saved - L1's VFIO reprograms it with
+     * destination-specific IRTE references after migration.
+     */
+
+    uint32_t src_pool = E1000_RAH_POOL_1 << saved_vfn;
+    uint32_t dst_pool = E1000_RAH_POOL_1 << s->vfn;
+
+    for (uint32_t i = 0; i < num_ra; i++) {
+        uint32_t offset = le32_to_cpu(blob->ra[i].offset);
+        uint32_t value = le32_to_cpu(blob->ra[i].value);
+
+        /* RAH entries: swap pool ownership bits */
+        if (offset >= RA && offset < RA + 32 && (offset - RA) % 2 == 1) {
+            value = (value & ~src_pool) | dst_pool;
+        }
+        if (offset >= RA2 && offset < RA2 + 16 && (offset - RA2) % 2 == 1) {
+            value = (value & ~src_pool) | dst_pool;
+        }
+
+        core->mac[offset] = value;
+    }
+
+    uint32_t num_tx = le32_to_cpu(blob->num_tx_ctx);
+    if (num_tx != 2) {
+        return -IGB_MIG_ERR_BAD_SIZE;
+    }
+
+    igb_core_vf_load_tx_ctx(core, q0, &blob->tx_ctx[0]);
+    igb_core_vf_load_tx_ctx(core, q1, &blob->tx_ctx[1]);
+
     trace_igbvf_mig_load_state(s->vfn, (uint32_t)size);
     return 0;
 }
-- 
2.55.0



^ permalink raw reply related	[flat|nested] 16+ messages in thread

* [RFC PATCH v2 4/9] igb: Add VF post-load fixups for live migration
  2026-09-02 19:20 [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Cédric Le Goater
                   ` (2 preceding siblings ...)
  2026-09-02 19:20 ` [RFC PATCH v2 3/9] igb: Add VF state serialization for live migration Cédric Le Goater
@ 2026-09-02 19:20 ` Cédric Le Goater
  2026-09-08  8:10   ` Akihiko Odaki
  2026-09-02 19:20 ` [RFC PATCH v2 5/9] igb: Add dirty page tracking for IGBVF migration Cédric Le Goater
                   ` (5 subsequent siblings)
  9 siblings, 1 reply; 16+ messages in thread
From: Cédric Le Goater @ 2026-09-02 19:20 UTC (permalink / raw)
  To: qemu-devel
  Cc: Akihiko Odaki, Sriram Yagnaraman, Jason Wang, Alex Williamson,
	Peter Xu, Cédric Le Goater

After restoring per-VF register state, propagate the VF's PVT shadow
values back into the PF's aggregate EIMS/EIAC/EIAM registers and
re-apply the VTIVAR interrupt vector routing to the shared IVAR0.

AI-used-for: analysis, code (prototype)
Signed-off-by: Cédric Le Goater <clg@redhat.com>
---
 hw/net/igb_core.h      |  3 ++
 hw/net/igb_core.c      | 66 ++++++++++++++++++++++++++++++++++++++++++
 hw/net/igb_migration.c |  7 +++++
 3 files changed, 76 insertions(+)

diff --git a/hw/net/igb_core.h b/hw/net/igb_core.h
index 60724e2824ab..22e10e4e6d0b 100644
--- a/hw/net/igb_core.h
+++ b/hw/net/igb_core.h
@@ -145,4 +145,7 @@ igb_start_recv(IGBCore *core);
 
 IGBCore *igb_pf_get_core(void *pf);
 
+void igb_core_vf_propagate_irqs(IGBCore *core, uint16_t vfn);
+void igb_core_vf_propagate_ivar(IGBCore *core, uint16_t vfn);
+
 #endif
diff --git a/hw/net/igb_core.c b/hw/net/igb_core.c
index 2a4883907353..01745fe756d0 100644
--- a/hw/net/igb_core.c
+++ b/hw/net/igb_core.c
@@ -4552,3 +4552,69 @@ igb_core_post_load(IGBCore *core)
 
     return 0;
 }
+
+/*
+ * Propagate VF interrupt state to PF aggregates after loading VF
+ * registers. The load path writes directly to mac[] bypassing the
+ * register handlers that OR VF bits into EIMS/EIAC/EIAM. Also clear
+ * stale VF bits in EICR that may have been set by packets arriving
+ * between PF vmstate restore and VF state load.
+ */
+void igb_core_vf_propagate_irqs(IGBCore *core, uint16_t vfn)
+{
+    uint32_t shift = 22 - vfn * IGBVF_MSIX_VEC_NUM;
+    uint32_t vf_mask = 0x7 << shift;
+    uint32_t pvt_idx;
+
+    core->mac[EIMS] &= ~vf_mask;
+    pvt_idx = PVTEIMS0 + vfn * 0x40;
+    core->mac[EIMS] |= (core->mac[pvt_idx] & 0x7) << shift;
+
+    core->mac[EIAC] &= ~vf_mask;
+    pvt_idx = PVTEIAC0 + vfn * 0x40;
+    core->mac[EIAC] |= (core->mac[pvt_idx] & 0x7) << shift;
+
+    core->mac[EIAM] &= ~vf_mask;
+    pvt_idx = PVTEIAM0 + vfn * 0x40;
+    core->mac[EIAM] |= (core->mac[pvt_idx] & 0x7) << shift;
+
+    core->mac[EICR] &= ~vf_mask;
+}
+
+/*
+ * Re-apply VTIVAR -> IVAR0 interrupt routing. The L1 PF driver
+ * may have overwritten the shared IVAR0 entries with its own
+ * queue routing after L0 vmstate restore.
+ */
+void igb_core_vf_propagate_ivar(IGBCore *core, uint16_t vfn)
+{
+    uint32_t vtivar = core->mac[VTIVAR + vfn];
+    int n;
+    uint8_t ent;
+    uint32_t mask;
+
+    n = igb_ivar_entry_rx(vfn);
+    mask = 0xffU << (8 * (n % 4));
+    if (vtivar & E1000_IVAR_VALID) {
+        ent = E1000_IVAR_VALID |
+              (24 - vfn * IGBVF_MSIX_VEC_NUM - (2 - (vtivar & 0x7)));
+        core->mac[IVAR0 + n / 4] =
+            (core->mac[IVAR0 + n / 4] & ~mask) |
+            ((uint32_t)ent << (8 * (n % 4)));
+    } else {
+        core->mac[IVAR0 + n / 4] &= ~mask;
+    }
+
+    n = igb_ivar_entry_tx(vfn);
+    mask = 0xffU << (8 * (n % 4));
+    ent = vtivar >> 8;
+    if (ent & E1000_IVAR_VALID) {
+        ent = E1000_IVAR_VALID |
+              (24 - vfn * IGBVF_MSIX_VEC_NUM - (2 - (ent & 0x7)));
+        core->mac[IVAR0 + n / 4] =
+            (core->mac[IVAR0 + n / 4] & ~mask) |
+            ((uint32_t)ent << (8 * (n % 4)));
+    } else {
+        core->mac[IVAR0 + n / 4] &= ~mask;
+    }
+}
diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
index c34035974620..af7257ccc4fd 100644
--- a/hw/net/igb_migration.c
+++ b/hw/net/igb_migration.c
@@ -384,12 +384,19 @@ static int igb_core_vf_load_state(IgbVfState *s, const void *buf, size_t size)
 static int igbvf_mig_load(IgbVfState *s, const void *buf, size_t size)
 {
     int ret;
+    IGBCore *core = igbvf_get_core(s);
 
     ret = igb_core_vf_load_state(s, buf, size);
     if (ret < 0) {
         return ret;
     }
 
+    /*
+     * Post-load: sync VF interrupt and routing state to PF aggregates
+     */
+    igb_core_vf_propagate_irqs(core, s->vfn);
+    igb_core_vf_propagate_ivar(core, s->vfn);
+
     return 0;
 }
 
-- 
2.55.0



^ permalink raw reply related	[flat|nested] 16+ messages in thread

* [RFC PATCH v2 5/9] igb: Add dirty page tracking for IGBVF migration
  2026-09-02 19:20 [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Cédric Le Goater
                   ` (3 preceding siblings ...)
  2026-09-02 19:20 ` [RFC PATCH v2 4/9] igb: Add VF post-load fixups " Cédric Le Goater
@ 2026-09-02 19:20 ` Cédric Le Goater
  2026-09-02 19:20 ` [RFC PATCH v2 6/9] igb: Quiesce VFs on STOP and include PF enable state in migration Cédric Le Goater
                   ` (4 subsequent siblings)
  9 siblings, 0 replies; 16+ messages in thread
From: Cédric Le Goater @ 2026-09-02 19:20 UTC (permalink / raw)
  To: qemu-devel
  Cc: Akihiko Odaki, Sriram Yagnaraman, Jason Wang, Alex Williamson,
	Peter Xu, Cédric Le Goater

Add per-VF dirty page tracking using bitmaps allocated per IOVA range.
The DMA data path is instrumented via igb_core_dirty_track_dma() calls at all
five DMA write sites: TX descriptor writeback (head pointer and status),
RX descriptor writeback, RX header fragment, and RX payload fragment.

The migration DVSEC exposes DIRTY_ENABLE, DIRTY_DISABLE and
DIRTY_QUERY commands. Range parameters and query results are exchanged
through the shared DMA buffer.

Add the PRE_COPY device state to support pre-copy live migration with
concurrent dirty tracking. PRE_COPY transitions:
 - RUNNING -> PRE_COPY
 - PRE_COPY -> STOP, RUNNING, STOP_COPY

Dirty tracking is automatically disabled when leaving PRE_COPY or
STOP_COPY.

AI-used-for: code (prototype)
Signed-off-by: Cédric Le Goater <clg@redhat.com>
---
 hw/net/igb_core.h      |   4 +
 hw/net/igb_migration.h |  65 +++++++-
 hw/net/igb_core.c      |  42 ++++--
 hw/net/igb_migration.c | 326 ++++++++++++++++++++++++++++++++++++++++-
 hw/net/trace-events    |   5 +
 5 files changed, 422 insertions(+), 20 deletions(-)

diff --git a/hw/net/igb_core.h b/hw/net/igb_core.h
index 22e10e4e6d0b..7fdc41690b36 100644
--- a/hw/net/igb_core.h
+++ b/hw/net/igb_core.h
@@ -40,6 +40,8 @@
 #ifndef HW_NET_IGB_CORE_H
 #define HW_NET_IGB_CORE_H
 
+#include "igb_migration.h"
+
 #define E1000E_MAC_SIZE         (0x8000)
 #define IGB_EEPROM_SIZE         (1024)
 
@@ -99,6 +101,8 @@ struct IGBCore {
     void (*owner_start_recv)(PCIDevice *d);
 
     int64_t timadj;
+
+    IGBVfDirtyState vf_dirty[IGB_MAX_VF_FUNCTIONS];
 };
 
 void
diff --git a/hw/net/igb_migration.h b/hw/net/igb_migration.h
index b2f601e74346..996de2d2f50b 100644
--- a/hw/net/igb_migration.h
+++ b/hw/net/igb_migration.h
@@ -21,7 +21,8 @@
  *   +0x04  DVSEC header 1             (len | rev | vendor_id)
  *   +0x08  DVSEC header 2             (DVSEC ID)
  *   +0x0A  Reserved                   (padding for DWORD alignment)
- *   +0x0C  CAPS                       (RO: F_STATE[0])
+ *   +0x0C  CAPS                       (RO: F_STATE[0], F_DIRTY[1],
+ *                                      max_ranges[11:8], pgsize[16:12])
  *   +0x10  CTRL                       (WO: doorbell command)
  *   +0x14  STATUS                     (RO: state[7:0], error_code[15:8])
  *   +0x18  BUF_ADDR_LO                (RW: shared buffer GPA low)
@@ -44,6 +45,12 @@
 
 /* CAPS register layout */
 #define IGB_MIG_CAP_F_STATE             (1u << 0)
+#define IGB_MIG_CAP_F_DIRTY             (1u << 1)
+#define IGB_MIG_CAPS_MAX_RANGES_SHIFT   8
+#define IGB_MIG_CAPS_MAX_RANGES         4
+#define IGB_MIG_CAPS_PGSIZE_SHIFT       12
+#define IGB_MIG_CAPS_PGSIZE_4K          (1u << 12)
+#define IGB_MIG_CAPS_PGSIZE_64K         (1u << 16)
 
 /* CTRL register: command in [7:0] */
 #define IGB_MIG_CTRL_CMD_MASK           0xFF
@@ -53,6 +60,9 @@
 #define IGB_MIG_CMD_SET_STATE           1
 #define IGB_MIG_CMD_SAVE                2
 #define IGB_MIG_CMD_LOAD                3
+#define IGB_MIG_CMD_DIRTY_ENABLE        4
+#define IGB_MIG_CMD_DIRTY_DISABLE       5
+#define IGB_MIG_CMD_DIRTY_QUERY         6
 
 /* STATUS register: state in [7:0], error code [15:8] */
 #define IGB_MIG_STATUS_STATE_MASK       0xFF
@@ -66,6 +76,7 @@
 #define IGB_MIG_STATE_RUNNING           2
 #define IGB_MIG_STATE_STOP_COPY         3
 #define IGB_MIG_STATE_RESUMING          4
+#define IGB_MIG_STATE_PRE_COPY          5
 
 /* Error codes */
 #define IGB_MIG_ERR_UNK_CMD             1
@@ -75,10 +86,29 @@
 #define IGB_MIG_ERR_BAD_SIZE            5
 #define IGB_MIG_ERR_BAD_MAGIC           6
 #define IGB_MIG_ERR_BAD_VERSION         7
+#define IGB_MIG_ERR_TOO_MANY_RANGES     8
+#define IGB_MIG_ERR_BAD_RANGE           9
+#define IGB_MIG_ERR_BAD_PGSIZE          10
+#define IGB_MIG_ERR_NOT_ENABLED         11
 
 /* Shared buffer constants */
 #define IGB_VF_STATE_MAX_SIZE           4096
 
+#define IGB_MIG_DIRTY_DEFAULT_PGSIZE        4096
+
+typedef struct IGBVfDirtyRange {
+    uint64_t iova;
+    uint64_t size;
+    uint64_t page_size;
+    unsigned long *bitmap;
+    uint64_t nbits;
+} IGBVfDirtyRange;
+
+typedef struct IGBVfDirtyState {
+    IGBVfDirtyRange ranges[IGB_MIG_CAPS_MAX_RANGES];
+    uint32_t num_ranges;
+} IGBVfDirtyState;
+
 typedef struct IgbVfMigState {
     uint32_t mig_state;
     uint32_t mig_data[IGB_VF_STATE_MAX_SIZE / sizeof(uint32_t)];
@@ -86,6 +116,36 @@ typedef struct IgbVfMigState {
     uint64_t mig_data_buf_addr;
 } IgbVfMigState;
 
+/*
+ * DMA buffer layouts for dirty tracking commands.
+ *
+ * DIRTY_ENABLE: driver writes igb_mig_dirty_enable_req to buffer
+ *               before cmd.
+ * DIRTY_QUERY: driver writes iova/size fields, device writes
+ *              response + bitmap.
+ */
+struct igb_mig_dirty_enable_req {
+    uint32_t len;
+    uint32_t flags;
+    uint64_t pgsize;
+    uint64_t range_iova;
+    uint64_t range_size;
+    uint32_t reserved[4];
+};
+
+struct igb_mig_dirty_query {
+    uint32_t len;
+    uint32_t flags;
+    uint64_t iova;
+    uint64_t size;
+    uint32_t bitmap_size;
+    uint32_t dirty_page_count;
+    uint64_t dma_writes;
+    uint32_t reserved[6];
+    uint8_t bitmap[];
+};
+
+typedef struct IGBCore IGBCore;
 typedef struct IgbVfState IgbVfState;
 
 bool igbvf_add_migration_dvsec(PCIDevice *dev, Error **errp);
@@ -94,4 +154,7 @@ uint32_t igbvf_mig_config_read(IgbVfState *s, uint32_t addr, int size);
 bool igbvf_mig_config_write(IgbVfState *s, uint32_t addr, uint32_t val,
                             int size);
 
+void igb_core_dirty_track_dma(IGBCore *core, int vfn,
+                              dma_addr_t addr, dma_addr_t len);
+
 #endif
diff --git a/hw/net/igb_core.c b/hw/net/igb_core.c
index 01745fe756d0..12329a1aff58 100644
--- a/hw/net/igb_core.c
+++ b/hw/net/igb_core.c
@@ -824,6 +824,16 @@ igb_rx_ring_init(IGBCore *core, E1000E_RxRing *rxr, int idx)
     rxr->i      = &i[idx];
 }
 
+static inline void
+igb_pci_dma_write(IGBCore *core, PCIDevice *dev,
+                  dma_addr_t addr, const void *buf, dma_addr_t len)
+{
+    pci_dma_write(dev, addr, buf, len);
+    if (pci_is_vf(dev)) {
+        igb_core_dirty_track_dma(core, pcie_sriov_vf_number(dev), addr, len);
+    }
+}
+
 static uint32_t
 igb_txdesc_writeback(IGBCore *core, dma_addr_t base,
                      union e1000_adv_tx_desc *tx_desc,
@@ -847,13 +857,15 @@ igb_txdesc_writeback(IGBCore *core, dma_addr_t base,
 
     if (tdwba & 1) {
         uint32_t buffer = cpu_to_le32(core->mac[txi->dh]);
-        pci_dma_write(d, tdwba & ~3, &buffer, sizeof(buffer));
+        igb_pci_dma_write(core, d,
+                          tdwba & ~3, &buffer, sizeof(buffer));
     } else {
         uint32_t status = le32_to_cpu(tx_desc->wb.status) | E1000_TXD_STAT_DD;
 
         tx_desc->wb.status = cpu_to_le32(status);
-        pci_dma_write(d, base + offsetof(union e1000_adv_tx_desc, wb),
-            &tx_desc->wb, sizeof(tx_desc->wb));
+        igb_pci_dma_write(core, d,
+                          base + offsetof(union e1000_adv_tx_desc, wb),
+                          &tx_desc->wb, sizeof(tx_desc->wb));
     }
 
     return igb_tx_wb_eic(core, txi->idx);
@@ -1598,11 +1610,12 @@ igb_pci_dma_write_rx_desc(IGBCore *core, PCIDevice *dev, dma_addr_t addr,
         uint8_t status = d->status;
 
         d->status &= ~E1000_RXD_STAT_DD;
-        pci_dma_write(dev, addr, desc, len);
+        igb_pci_dma_write(core, dev, addr, desc, len);
 
         if (status & E1000_RXD_STAT_DD) {
             d->status = status;
-            pci_dma_write(dev, addr + offset, &status, sizeof(status));
+            igb_pci_dma_write(core, dev,
+                              addr + offset, &status, sizeof(status));
         }
     } else {
         union e1000_adv_rx_desc *d = &desc->adv;
@@ -1611,11 +1624,12 @@ igb_pci_dma_write_rx_desc(IGBCore *core, PCIDevice *dev, dma_addr_t addr,
         uint32_t status = d->wb.upper.status_error;
 
         d->wb.upper.status_error &= ~E1000_RXD_STAT_DD;
-        pci_dma_write(dev, addr, desc, len);
+        igb_pci_dma_write(core, dev, addr, desc, len);
 
         if (status & E1000_RXD_STAT_DD) {
             d->wb.upper.status_error = status;
-            pci_dma_write(dev, addr + offset, &status, sizeof(status));
+            igb_pci_dma_write(core, dev,
+                              addr + offset, &status, sizeof(status));
         }
     }
 }
@@ -1737,9 +1751,9 @@ igb_write_hdr_frag_to_rx_buffers(IGBCore *core,
 {
     assert(data_len <= pdma_st->rx_desc_header_buf_size -
                        pdma_st->bastate.written[0]);
-    pci_dma_write(d,
-                  pdma_st->ba[0] + pdma_st->bastate.written[0],
-                  data, data_len);
+    igb_pci_dma_write(core, d,
+                      pdma_st->ba[0] + pdma_st->bastate.written[0],
+                      data, data_len);
     pdma_st->bastate.written[0] += data_len;
     pdma_st->bastate.cur_idx = 1;
 }
@@ -1804,10 +1818,10 @@ igb_write_payload_frag_to_rx_buffers(IGBCore *core,
             data,
             bytes_to_write);
 
-        pci_dma_write(d,
-                      pdma_st->ba[pdma_st->bastate.cur_idx] +
-                      pdma_st->bastate.written[pdma_st->bastate.cur_idx],
-                      data, bytes_to_write);
+        igb_pci_dma_write(core, d,
+                          pdma_st->ba[pdma_st->bastate.cur_idx] +
+                          pdma_st->bastate.written[pdma_st->bastate.cur_idx],
+                          data, bytes_to_write);
 
         pdma_st->bastate.written[pdma_st->bastate.cur_idx] += bytes_to_write;
         data += bytes_to_write;
diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
index af7257ccc4fd..cfa73673e39a 100644
--- a/hw/net/igb_migration.c
+++ b/hw/net/igb_migration.c
@@ -8,6 +8,8 @@
 
 #include "qemu/osdep.h"
 #include "qemu/log.h"
+#include "qemu/bitmap.h"
+#include "qemu/units.h"
 #include "hw/pci/pci_device.h"
 #include "hw/pci/pcie.h"
 #include "net/eth.h"
@@ -400,6 +402,285 @@ static int igbvf_mig_load(IgbVfState *s, const void *buf, size_t size)
     return 0;
 }
 
+/*
+ * Per-VF dirty page tracking
+ *
+ * All VF DMA writes in igb_core.c go through igb_pci_dma_write(),
+ * which calls igb_core_dirty_track_dma() to mark the target page in a
+ * per-range bitmap before performing the actual DMA.
+ *
+ * The IGBCore::vf_dirty[] bitmaps live in IGBCore so they are easily
+ * accessible from the core TX and RX paths without reaching back into
+ * VF state.
+ */
+
+void igb_core_dirty_track_dma(IGBCore *core, int vfn,
+                              dma_addr_t addr, dma_addr_t len)
+{
+    IGBVfDirtyState *ds = &core->vf_dirty[vfn];
+    bool matched = false;
+    uint32_t i;
+
+    if (!ds->num_ranges) {
+        return;
+    }
+
+    trace_igb_core_dirty_track_dma(vfn, addr, len);
+
+    for (i = 0; i < ds->num_ranges; i++) {
+        IGBVfDirtyRange *r = &ds->ranges[i];
+        uint64_t r_end = r->iova + r->size;
+        uint64_t dma_end = addr + len;
+        uint64_t start, end, start_page, end_page, page;
+
+        if (addr >= r_end || dma_end <= r->iova) {
+            continue;
+        }
+
+        matched = true;
+        start = MAX(addr, r->iova);
+        end = MIN(dma_end, r_end);
+
+        start_page = (start - r->iova) / r->page_size;
+        end_page = (end - 1 - r->iova) / r->page_size;
+
+        for (page = start_page; page <= end_page; page++) {
+            if (page < r->nbits) {
+                set_bit(page, r->bitmap);
+            }
+        }
+    }
+
+    if (!matched) {
+        trace_igb_core_dirty_track_dma_drop(vfn, addr, len);
+    }
+}
+
+static IGBVfDirtyState *igb_core_vf_dirty_state(IgbVfState *s)
+{
+    IGBCore *core = igbvf_get_core(s);
+    return &core->vf_dirty[s->vfn];
+}
+
+#define IGB_MIG_DIRTY_MAX_PAGES      ((256ULL * GiB) / (4 * KiB))
+
+static uint32_t igb_core_vf_dirty_enable(IgbVfState *s, uint64_t pgsize,
+                                         uint64_t range_iova,
+                                         uint64_t range_size)
+{
+    uint32_t caps = pci_get_long(PCI_DEVICE(s)->config +
+                                 IGB_MIG_DVSEC_OFFSET + IGB_MIG_CAPS);
+    IGBVfDirtyState *ds = igb_core_vf_dirty_state(s);
+    IGBVfDirtyRange *r;
+
+    if (ds->num_ranges >= IGB_MIG_CAPS_MAX_RANGES) {
+        return IGB_MIG_ERR_TOO_MANY_RANGES;
+    }
+
+    if (!range_size) {
+        return IGB_MIG_ERR_BAD_RANGE;
+    }
+
+    /* Validate page size against CAPS supported page size bitmask */
+    if (!is_power_of_2(pgsize) || !(pgsize & caps)) {
+        return IGB_MIG_ERR_BAD_PGSIZE;
+    }
+
+    if ((range_iova % pgsize) || (range_size % pgsize)) {
+        return IGB_MIG_ERR_BAD_PGSIZE;
+    }
+
+    if (range_size / pgsize > IGB_MIG_DIRTY_MAX_PAGES) {
+        return IGB_MIG_ERR_BAD_RANGE;
+    }
+
+    r = &ds->ranges[ds->num_ranges];
+    r->iova = range_iova;
+    r->size = range_size;
+    r->page_size = pgsize;
+    r->nbits = range_size / pgsize;
+    r->bitmap = bitmap_new(r->nbits);
+    ds->num_ranges++;
+    return 0;
+}
+
+static void igb_core_vf_dirty_disable(IgbVfState *s)
+{
+    IGBVfDirtyState *ds = igb_core_vf_dirty_state(s);
+    uint32_t i;
+
+    for (i = 0; i < ds->num_ranges; i++) {
+        IGBVfDirtyRange *r = &ds->ranges[i];
+
+        g_free(r->bitmap);
+        r->bitmap = NULL;
+        r->nbits = 0;
+    }
+    ds->num_ranges = 0;
+    trace_igbvf_mig_dirty_disable(s->vfn);
+}
+
+static bool igb_core_vf_dirty_enabled(IgbVfState *s)
+{
+    return igb_core_vf_dirty_state(s)->num_ranges > 0;
+}
+
+static void igb_core_vf_dirty_query(IGBVfDirtyRange *r,
+                                    uint64_t range_iova, uint64_t range_size,
+                                    void *buf, size_t buf_size,
+                                    size_t *out_size)
+{
+    uint64_t start_page = (range_iova - r->iova) / r->page_size;
+    uint64_t range_pages = range_size / r->page_size;
+    uint64_t count = MIN(range_pages, (uint64_t)buf_size * 8);
+
+    memset(buf, 0, buf_size);
+
+    if (start_page < r->nbits) {
+        uint64_t avail = r->nbits - start_page;
+        uint64_t n = MIN(count, avail);
+
+        bitmap_copy_with_src_offset(buf, r->bitmap, start_page, n);
+    }
+    *out_size = bitmap_empty(buf, count) ? 0 : DIV_ROUND_UP(count, 8);
+}
+
+static void igb_core_vf_dirty_query_commit(IGBVfDirtyRange *r,
+                                           uint64_t range_iova,
+                                           uint64_t range_size)
+{
+    uint64_t start_page = (range_iova - r->iova) / r->page_size;
+
+    if (start_page < r->nbits) {
+        uint64_t avail = r->nbits - start_page;
+        uint64_t range_pages = range_size / r->page_size;
+        uint64_t n = MIN(range_pages, avail);
+
+        bitmap_clear(r->bitmap, start_page, n);
+    }
+}
+
+static uint8_t igbvf_mig_cmd_dirty_enable(IgbVfState *s)
+{
+    IgbVfMigState *ms = &s->mig;
+    struct igb_mig_dirty_enable_req req;
+    MemTxResult r;
+    uint32_t status;
+
+    if (!ms->mig_data_buf_addr) {
+        return IGB_MIG_ERR_NO_BUFFER;
+    }
+
+    r = address_space_read(&address_space_memory, ms->mig_data_buf_addr,
+                           MEMTXATTRS_UNSPECIFIED, &req, sizeof(req));
+    if (r != MEMTX_OK) {
+        return IGB_MIG_ERR_DMA_FAILED;
+    }
+
+    /* TODO: validate req.len and req.flags */
+
+    status = igb_core_vf_dirty_enable(s, le64_to_cpu(req.pgsize),
+                                      le64_to_cpu(req.range_iova),
+                                      le64_to_cpu(req.range_size));
+    if (status != 0) {
+        return status;
+    }
+    trace_igbvf_mig_dirty_enable(s->vfn, le64_to_cpu(req.pgsize),
+                                 le64_to_cpu(req.range_size) /
+                                 le64_to_cpu(req.pgsize));
+    return 0;
+}
+
+static IGBVfDirtyRange *igb_core_vf_dirty_range_valid(IGBVfDirtyState *ds,
+                                                      uint64_t range_iova,
+                                                      uint64_t range_size)
+{
+    if (!range_size) {
+        return NULL;
+    }
+
+    for (uint32_t i = 0; i < ds->num_ranges; i++) {
+        IGBVfDirtyRange *r = &ds->ranges[i];
+
+        if (range_iova >= r->iova &&
+            range_iova + range_size <= r->iova + r->size &&
+            QEMU_IS_ALIGNED(range_iova, r->page_size) &&
+            QEMU_IS_ALIGNED(range_size, r->page_size)) {
+            return r;
+        }
+    }
+    return NULL;
+}
+
+static uint8_t igbvf_mig_cmd_dirty_query(IgbVfState *s)
+{
+    IgbVfMigState *ms = &s->mig;
+    IGBVfDirtyState *ds = &igbvf_get_core(s)->vf_dirty[s->vfn];
+    uint64_t buf_addr = ms->mig_data_buf_addr;
+    uint64_t range_iova = 0, range_size = 0;
+    IGBVfDirtyRange *range;
+    uint32_t bmp_bytes, dirty_pages;
+    uint32_t val32;
+    size_t out_size;
+    g_autofree void *bitmap = NULL;
+
+    if (!buf_addr) {
+        return IGB_MIG_ERR_NO_BUFFER;
+    }
+
+    if (!igb_core_vf_dirty_enabled(s)) {
+        return IGB_MIG_ERR_NOT_ENABLED;
+    }
+
+    address_space_read(&address_space_memory,
+                       buf_addr + offsetof(struct igb_mig_dirty_query, iova),
+                       MEMTXATTRS_UNSPECIFIED, &range_iova, sizeof(range_iova));
+    range_iova = le64_to_cpu(range_iova);
+    address_space_read(&address_space_memory,
+                       buf_addr + offsetof(struct igb_mig_dirty_query, size),
+                       MEMTXATTRS_UNSPECIFIED, &range_size, sizeof(range_size));
+    range_size = le64_to_cpu(range_size);
+
+    range = igb_core_vf_dirty_range_valid(ds, range_iova, range_size);
+    if (!range) {
+        return IGB_MIG_ERR_BAD_RANGE;
+    }
+
+    bmp_bytes = BITS_TO_LONGS(range_size / range->page_size) *
+        sizeof(unsigned long);
+    bitmap = g_malloc0(bmp_bytes);
+
+    igb_core_vf_dirty_query(range, range_iova, range_size,
+                            bitmap, bmp_bytes, &out_size);
+
+    if (out_size) {
+        if (address_space_write(&address_space_memory,
+                                buf_addr +
+                                offsetof(struct igb_mig_dirty_query, bitmap),
+                                MEMTXATTRS_UNSPECIFIED, bitmap, out_size)) {
+            return IGB_MIG_ERR_DMA_FAILED;
+        }
+    }
+
+    igb_core_vf_dirty_query_commit(range, range_iova, range_size);
+
+    dirty_pages = bitmap_count_one(bitmap, range_size / range->page_size);
+
+    val32 = cpu_to_le32(out_size);
+    address_space_write(&address_space_memory,
+                        buf_addr + offsetof(struct igb_mig_dirty_query,
+                                            bitmap_size),
+                        MEMTXATTRS_UNSPECIFIED, &val32, sizeof(val32));
+    val32 = cpu_to_le32(dirty_pages);
+    address_space_write(&address_space_memory,
+                        buf_addr + offsetof(struct igb_mig_dirty_query,
+                                            dirty_page_count),
+                        MEMTXATTRS_UNSPECIFIED, &val32, sizeof(val32));
+
+    trace_igbvf_mig_dirty_query(s->vfn, (uint64_t)out_size, dirty_pages);
+    return 0;
+}
+
 /*
  * Migration command handlers
  */
@@ -419,7 +700,8 @@ static uint8_t igbvf_mig_cmd_save(IgbVfState *s)
     MemTxResult r;
     int ret;
 
-    if (ms->mig_state != IGB_MIG_STATE_STOP_COPY) {
+    if (ms->mig_state != IGB_MIG_STATE_STOP_COPY &&
+        ms->mig_state != IGB_MIG_STATE_PRE_COPY) {
         return IGB_MIG_ERR_BAD_STATE;
     }
 
@@ -487,22 +769,33 @@ static uint8_t igbvf_mig_set_state(IgbVfState *s, uint32_t new_state)
     case IGB_MIG_STATE_STOP:
         if (old != IGB_MIG_STATE_RUNNING &&
             old != IGB_MIG_STATE_STOP_COPY &&
+            old != IGB_MIG_STATE_PRE_COPY &&
             old != IGB_MIG_STATE_RESUMING &&
             old != IGB_MIG_STATE_ERROR) {
             return IGB_MIG_ERR_BAD_STATE;
         }
+        if (old == IGB_MIG_STATE_PRE_COPY ||
+            old == IGB_MIG_STATE_STOP_COPY ||
+            old == IGB_MIG_STATE_ERROR) {
+            igb_core_vf_dirty_disable(s);
+        }
         /* Restore DATA_SIZE to max, same as at reset */
         igbvf_mig_update_data_size(s, igb_core_vf_max_data_size(s));
         break;
 
     case IGB_MIG_STATE_RUNNING:
-        if (old != IGB_MIG_STATE_STOP) {
+        if (old != IGB_MIG_STATE_STOP &&
+            old != IGB_MIG_STATE_PRE_COPY) {
             return IGB_MIG_ERR_BAD_STATE;
         }
+        if (old == IGB_MIG_STATE_PRE_COPY) {
+            igb_core_vf_dirty_disable(s);
+        }
         break;
 
     case IGB_MIG_STATE_STOP_COPY:
-        if (old != IGB_MIG_STATE_STOP) {
+        if (old != IGB_MIG_STATE_STOP &&
+            old != IGB_MIG_STATE_PRE_COPY) {
             return IGB_MIG_ERR_BAD_STATE;
         }
         ret = igb_core_vf_save_state(s, ms->mig_data, sizeof(ms->mig_data));
@@ -520,6 +813,12 @@ static uint8_t igbvf_mig_set_state(IgbVfState *s, uint32_t new_state)
         igbvf_mig_update_data_size(s, 0);
         break;
 
+    case IGB_MIG_STATE_PRE_COPY:
+        if (old != IGB_MIG_STATE_RUNNING) {
+            return IGB_MIG_ERR_BAD_STATE;
+        }
+        break;
+
     default:
         return IGB_MIG_ERR_BAD_STATE;
     }
@@ -564,6 +863,18 @@ static void igbvf_mig_cmd_ctrl(IgbVfState *s, uint32_t val)
         err = igbvf_mig_cmd_load(s);
         break;
 
+    case IGB_MIG_CMD_DIRTY_ENABLE:
+        err = igbvf_mig_cmd_dirty_enable(s);
+        break;
+
+    case IGB_MIG_CMD_DIRTY_DISABLE:
+        igb_core_vf_dirty_disable(s);
+        break;
+
+    case IGB_MIG_CMD_DIRTY_QUERY:
+        err = igbvf_mig_cmd_dirty_query(s);
+        break;
+
     default:
         err = IGB_MIG_ERR_UNK_CMD;
         break;
@@ -594,8 +905,10 @@ bool igbvf_add_migration_dvsec(PCIDevice *dev, Error **errp)
     /* DVSEC header 2: DVSEC ID */
     pci_set_word(dev->config + offset + 0x8, IGB_MIG_DVSEC_ID);
 
-    /* CAPS: features (state migration only) */
-    caps = IGB_MIG_CAP_F_STATE;
+    /* CAPS: features | max_ranges | supported page sizes (4K) */
+    caps = IGB_MIG_CAP_F_STATE | IGB_MIG_CAP_F_DIRTY |
+           (IGB_MIG_CAPS_MAX_RANGES << IGB_MIG_CAPS_MAX_RANGES_SHIFT) |
+           IGB_MIG_CAPS_PGSIZE_4K;
     pci_set_long(dev->config + offset + IGB_MIG_CAPS, caps);
 
     /* STATUS: initial state is RUNNING */
@@ -657,6 +970,9 @@ void igbvf_mig_state_reset(IgbVfState *s)
     IgbVfMigState *ms = &s->mig;
 
     trace_igbvf_mig_reset(s->vfn);
+
+    igb_core_vf_dirty_disable(s);
+
     ms->mig_state = IGB_MIG_STATE_RUNNING;
     ms->mig_data_buf_addr = 0;
     igbvf_mig_update_data_size(s, igb_core_vf_max_data_size(s));
diff --git a/hw/net/trace-events b/hw/net/trace-events
index 7057fbe5f16a..c606191df2ef 100644
--- a/hw/net/trace-events
+++ b/hw/net/trace-events
@@ -300,6 +300,11 @@ igbvf_mig_set_state(uint16_t vfn, uint32_t old_state, uint32_t new_state) "VF%u:
 igbvf_mig_save_state(uint16_t vfn, uint32_t size) "VF%u: saved %u bytes of device state"
 igbvf_mig_load_state(uint16_t vfn, uint32_t size) "VF%u: loaded %u bytes of device state"
 igbvf_mig_reset(uint16_t vfn) "VF%u: migration state reset"
+igbvf_mig_dirty_enable(uint16_t vfn, uint64_t pgsize, uint64_t nbits) "VF%u: dirty tracking enabled pgsize=%"PRIu64" nbits=%"PRIu64
+igbvf_mig_dirty_disable(uint16_t vfn) "VF%u: dirty tracking disabled"
+igbvf_mig_dirty_query(uint16_t vfn, uint64_t size, uint32_t dirty_pages) "VF%u: dirty query returned %"PRIu64" bytes, %u dirty pages"
+igb_core_dirty_track_dma(int vfn, uint64_t addr, uint64_t len) "VF%d: dirty DMA addr=0x%"PRIx64" len=%"PRIu64
+igb_core_dirty_track_dma_drop(int vfn, uint64_t addr, uint64_t len) "VF%d: dirty DMA dropped addr=0x%"PRIx64" len=%"PRIu64" no matching range"
 
 # spapr_llan.c
 spapr_vlan_get_rx_bd_from_pool_found(int pool, int32_t count, uint32_t rx_bufs) "pool=%d count=%"PRId32" rxbufs=%"PRIu32
-- 
2.55.0



^ permalink raw reply related	[flat|nested] 16+ messages in thread

* [RFC PATCH v2 6/9] igb: Quiesce VFs on STOP and include PF enable state in migration
  2026-09-02 19:20 [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Cédric Le Goater
                   ` (4 preceding siblings ...)
  2026-09-02 19:20 ` [RFC PATCH v2 5/9] igb: Add dirty page tracking for IGBVF migration Cédric Le Goater
@ 2026-09-02 19:20 ` Cédric Le Goater
  2026-09-02 19:20 ` [RFC PATCH v2 7/9] igb: Fix post-migration RX ring deadlock Cédric Le Goater
                   ` (3 subsequent siblings)
  9 siblings, 0 replies; 16+ messages in thread
From: Cédric Le Goater @ 2026-09-02 19:20 UTC (permalink / raw)
  To: qemu-devel
  Cc: Akihiko Odaki, Sriram Yagnaraman, Jason Wang, Alex Williamson,
	Peter Xu, Cédric Le Goater

Quiesce VFs by clearing their VFRE/VFTE bits on transitions to STOP
(from RUNNING, PRE_COPY, or ERROR) and to STOP_COPY (from PRE_COPY).
This prevents further DMA while device state is being serialized.
Restore the saved VFRE/VFTE state on STOP->RUNNING so the VF can
resume normal operation.

Include the per-VF VFRE/VFTE enable bits in the migration blob so the
destination knows whether receive and transmit were active before
quiesce. On reset, both default to true.

Add a QUIESCED bit to the STATUS register so the driver can confirm
DMA has been drained before reading device state.

Clear the VFLRE (VF Level Reset Event) bit before restoring state on
the destination to prevent the PF watchdog from seeing a stale reset
indication and overwriting the loaded registers.

AI-used-for: analysis, code (prototype)
Signed-off-by: Cédric Le Goater <clg@redhat.com>
---
 hw/net/igb_migration.h |  6 ++-
 hw/net/igb_migration.c | 86 +++++++++++++++++++++++++++++++++++++++++-
 hw/net/trace-events    |  6 ++-
 3 files changed, 93 insertions(+), 5 deletions(-)

diff --git a/hw/net/igb_migration.h b/hw/net/igb_migration.h
index 996de2d2f50b..2bb9a0ed36ce 100644
--- a/hw/net/igb_migration.h
+++ b/hw/net/igb_migration.h
@@ -24,7 +24,8 @@
  *   +0x0C  CAPS                       (RO: F_STATE[0], F_DIRTY[1],
  *                                      max_ranges[11:8], pgsize[16:12])
  *   +0x10  CTRL                       (WO: doorbell command)
- *   +0x14  STATUS                     (RO: state[7:0], error_code[15:8])
+ *   +0x14  STATUS                     (RO: state[7:0], error_code[15:8],
+ *                                      QUIESCED[16])
  *   +0x18  BUF_ADDR_LO                (RW: shared buffer GPA low)
  *   +0x1C  BUF_ADDR_HI                (RW: shared buffer GPA high)
  *   +0x20  DATA_SIZE                  (RO: max state blob size in bytes)
@@ -69,6 +70,7 @@
 #define IGB_MIG_STATUS_ERROR_CODE_SHIFT 8
 #define IGB_MIG_STATUS_ERR(code) \
     ((uint32_t)(code) << IGB_MIG_STATUS_ERROR_CODE_SHIFT)
+#define IGB_MIG_STATUS_QUIESCED         (1u << 16)
 
 /* Device states (based on VFIO migration v2) */
 #define IGB_MIG_STATE_ERROR             0
@@ -114,6 +116,8 @@ typedef struct IgbVfMigState {
     uint32_t mig_data[IGB_VF_STATE_MAX_SIZE / sizeof(uint32_t)];
     uint32_t mig_data_size;
     uint64_t mig_data_buf_addr;
+    bool mig_saved_vfre;
+    bool mig_saved_vfte;
 } IgbVfMigState;
 
 /*
diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
index cfa73673e39a..435bd0d1623b 100644
--- a/hw/net/igb_migration.c
+++ b/hw/net/igb_migration.c
@@ -58,6 +58,8 @@ typedef struct IgbMigBlob {
     IgbMigRegPair ra[IGB_VF_MAX_RA_REGS];
     uint32_t num_tx_ctx;
     IgbMigTxCtx tx_ctx[2];
+    uint32_t vfre;
+    uint32_t vfte;
 } IgbMigBlob;
 
 #define IGB_MIG_BLOB_SIZE            sizeof(IgbMigBlob)
@@ -210,6 +212,7 @@ static void igb_core_vf_save_tx_ctx(IGBCore *core, int queue,
 static int igb_core_vf_save_state(IgbVfState *s, void *buf, size_t buf_size)
 {
     int size = IGB_MIG_BLOB_SIZE;
+    IgbVfMigState *ms = &s->mig;
     IGBCore *core = igbvf_get_core(s);
     IgbMigBlob *blob = buf;
     uint32_t offsets[IGB_VF_MAX_FIXED_REGS];
@@ -248,7 +251,12 @@ static int igb_core_vf_save_state(IgbVfState *s, void *buf, size_t buf_size)
     igb_core_vf_save_tx_ctx(core, q0, &blob->tx_ctx[0]);
     igb_core_vf_save_tx_ctx(core, q1, &blob->tx_ctx[1]);
 
-    trace_igbvf_mig_save_state(s->vfn, size);
+    blob->vfre = cpu_to_le32(ms->mig_saved_vfre);
+    blob->vfte = cpu_to_le32(ms->mig_saved_vfte);
+
+    trace_igbvf_mig_save_state(s->vfn, size, ms->mig_saved_vfre,
+                               ms->mig_saved_vfte,
+                               core->mac[VFRE]);
     return size;
 }
 
@@ -291,6 +299,7 @@ static uint32_t igb_vf_relocate_offset(uint32_t offset,
 
 static int igb_core_vf_load_state(IgbVfState *s, const void *buf, size_t size)
 {
+    IgbVfMigState *ms = &s->mig;
     IGBCore *core = igbvf_get_core(s);
     uint32_t src_offsets[IGB_VF_MAX_FIXED_REGS];
     uint32_t dst_offsets[IGB_VF_MAX_FIXED_REGS];
@@ -379,7 +388,12 @@ static int igb_core_vf_load_state(IgbVfState *s, const void *buf, size_t size)
     igb_core_vf_load_tx_ctx(core, q0, &blob->tx_ctx[0]);
     igb_core_vf_load_tx_ctx(core, q1, &blob->tx_ctx[1]);
 
-    trace_igbvf_mig_load_state(s->vfn, (uint32_t)size);
+    ms->mig_saved_vfre = !!le32_to_cpu(blob->vfre);
+    ms->mig_saved_vfte = !!le32_to_cpu(blob->vfte);
+
+    trace_igbvf_mig_load_state(s->vfn, (uint32_t)size,
+                               ms->mig_saved_vfre,
+                               ms->mig_saved_vfte);
     return 0;
 }
 
@@ -388,6 +402,12 @@ static int igbvf_mig_load(IgbVfState *s, const void *buf, size_t size)
     int ret;
     IGBCore *core = igbvf_get_core(s);
 
+    /*
+     * Pre-load: Clear the VFLRE bit before restoring state so the PF
+     * watchdog does not overwrite what we are about to load.
+     */
+    core->mac[VFLRE] &= ~BIT(s->vfn);
+
     ret = igb_core_vf_load_state(s, buf, size);
     if (ret < 0) {
         return ret;
@@ -759,6 +779,45 @@ static uint8_t igbvf_mig_cmd_load(IgbVfState *s)
     return 0;
 }
 
+/* Quiesce a VF by disabling its RX and TX at the PF level. */
+static void igb_core_vf_quiesce(IgbVfState *s)
+{
+    IgbVfMigState *ms = &s->mig;
+    IGBCore *core = igbvf_get_core(s);
+
+    ms->mig_saved_vfre = !!(core->mac[VFRE] & BIT(s->vfn));
+    ms->mig_saved_vfte = !!(core->mac[VFTE] & BIT(s->vfn));
+
+    core->mac[VFRE] &= ~BIT(s->vfn);
+    core->mac[VFTE] &= ~BIT(s->vfn);
+    trace_igbvf_mig_quiesce(s->vfn, core->mac[VFRE], core->mac[VFTE]);
+}
+
+static void igb_core_vf_unquiesce(IgbVfState *s)
+{
+    IgbVfMigState *ms = &s->mig;
+    IGBCore *core = igbvf_get_core(s);
+    bool re = ms->mig_saved_vfre;
+    bool te = ms->mig_saved_vfte;
+
+    if (re) {
+        core->mac[VFRE] |= BIT(s->vfn);
+    } else {
+        core->mac[VFRE] &= ~BIT(s->vfn);
+    }
+    if (te) {
+        core->mac[VFTE] |= BIT(s->vfn);
+    } else {
+        core->mac[VFTE] &= ~BIT(s->vfn);
+    }
+
+    trace_igbvf_mig_unquiesce(s->vfn, core->mac[VFRE], core->mac[VFTE]);
+
+    if (re) {
+        igb_start_recv(core);
+    }
+}
+
 static uint8_t igbvf_mig_set_state(IgbVfState *s, uint32_t new_state)
 {
     IgbVfMigState *ms = &s->mig;
@@ -779,6 +838,11 @@ static uint8_t igbvf_mig_set_state(IgbVfState *s, uint32_t new_state)
             old == IGB_MIG_STATE_ERROR) {
             igb_core_vf_dirty_disable(s);
         }
+        if (old == IGB_MIG_STATE_RUNNING ||
+            old == IGB_MIG_STATE_PRE_COPY ||
+            old == IGB_MIG_STATE_ERROR) {
+            igb_core_vf_quiesce(s);
+        }
         /* Restore DATA_SIZE to max, same as at reset */
         igbvf_mig_update_data_size(s, igb_core_vf_max_data_size(s));
         break;
@@ -791,6 +855,9 @@ static uint8_t igbvf_mig_set_state(IgbVfState *s, uint32_t new_state)
         if (old == IGB_MIG_STATE_PRE_COPY) {
             igb_core_vf_dirty_disable(s);
         }
+        if (old == IGB_MIG_STATE_STOP) {
+            igb_core_vf_unquiesce(s);
+        }
         break;
 
     case IGB_MIG_STATE_STOP_COPY:
@@ -798,6 +865,9 @@ static uint8_t igbvf_mig_set_state(IgbVfState *s, uint32_t new_state)
             old != IGB_MIG_STATE_PRE_COPY) {
             return IGB_MIG_ERR_BAD_STATE;
         }
+        if (old == IGB_MIG_STATE_PRE_COPY) {
+            igb_core_vf_quiesce(s);
+        }
         ret = igb_core_vf_save_state(s, ms->mig_data, sizeof(ms->mig_data));
         if (ret < 0) {
             return -ret;
@@ -840,6 +910,16 @@ static void igbvf_mig_update_status(IgbVfState *s, uint8_t err)
         status = IGB_MIG_STATE_ERROR | IGB_MIG_STATUS_ERR(err);
     }
 
+    /*
+     * QUIESCED tells the driver it is safe to read device state.
+     * In STOP and STOP_COPY, igb_core_vf_quiesce() has already
+     * cleared VFRE/VFTE so no further VF DMA can occur.
+     */
+    if (ms->mig_state == IGB_MIG_STATE_STOP ||
+        ms->mig_state == IGB_MIG_STATE_STOP_COPY) {
+        status |= IGB_MIG_STATUS_QUIESCED;
+    }
+
     pci_set_long(dev->config + IGB_MIG_DVSEC_OFFSET + IGB_MIG_STATUS, status);
 }
 
@@ -977,6 +1057,8 @@ void igbvf_mig_state_reset(IgbVfState *s)
     ms->mig_data_buf_addr = 0;
     igbvf_mig_update_data_size(s, igb_core_vf_max_data_size(s));
     memset(ms->mig_data, 0, sizeof(ms->mig_data));
+    ms->mig_saved_vfre = true;
+    ms->mig_saved_vfte = true;
 
     pci_set_long(PCI_DEVICE(s)->config +
                  IGB_MIG_DVSEC_OFFSET + IGB_MIG_BUF_ADDR_LO, 0);
diff --git a/hw/net/trace-events b/hw/net/trace-events
index c606191df2ef..eb5c61a8f6fb 100644
--- a/hw/net/trace-events
+++ b/hw/net/trace-events
@@ -297,14 +297,16 @@ igbvf_wrn_io_addr_unknown(uint64_t addr) "IO unknown register 0x%"PRIx64
 
 # igb_migration.c
 igbvf_mig_set_state(uint16_t vfn, uint32_t old_state, uint32_t new_state) "VF%u: state %u -> %u"
-igbvf_mig_save_state(uint16_t vfn, uint32_t size) "VF%u: saved %u bytes of device state"
-igbvf_mig_load_state(uint16_t vfn, uint32_t size) "VF%u: loaded %u bytes of device state"
+igbvf_mig_save_state(uint16_t vfn, int size, bool vfre, bool vfte, uint32_t reg_vfre) "VF%u: saved %d bytes vfre=%d vfte=%d VFRE=0x%x"
+igbvf_mig_load_state(uint16_t vfn, uint32_t size, bool vfre, bool vfte) "VF%u: loaded %u bytes vfre=%d vfte=%d"
 igbvf_mig_reset(uint16_t vfn) "VF%u: migration state reset"
 igbvf_mig_dirty_enable(uint16_t vfn, uint64_t pgsize, uint64_t nbits) "VF%u: dirty tracking enabled pgsize=%"PRIu64" nbits=%"PRIu64
 igbvf_mig_dirty_disable(uint16_t vfn) "VF%u: dirty tracking disabled"
 igbvf_mig_dirty_query(uint16_t vfn, uint64_t size, uint32_t dirty_pages) "VF%u: dirty query returned %"PRIu64" bytes, %u dirty pages"
 igb_core_dirty_track_dma(int vfn, uint64_t addr, uint64_t len) "VF%d: dirty DMA addr=0x%"PRIx64" len=%"PRIu64
 igb_core_dirty_track_dma_drop(int vfn, uint64_t addr, uint64_t len) "VF%d: dirty DMA dropped addr=0x%"PRIx64" len=%"PRIu64" no matching range"
+igbvf_mig_quiesce(uint16_t vfn, uint32_t vfre, uint32_t vfte) "VF%u: quiesce VFRE=0x%x VFTE=0x%x"
+igbvf_mig_unquiesce(uint16_t vfn, uint32_t vfre, uint32_t vfte) "VF%u: unquiesce VFRE=0x%x VFTE=0x%x"
 
 # spapr_llan.c
 spapr_vlan_get_rx_bd_from_pool_found(int pool, int32_t count, uint32_t rx_bufs) "pool=%d count=%"PRId32" rxbufs=%"PRIu32
-- 
2.55.0



^ permalink raw reply related	[flat|nested] 16+ messages in thread

* [RFC PATCH v2 7/9] igb: Fix post-migration RX ring deadlock
  2026-09-02 19:20 [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Cédric Le Goater
                   ` (5 preceding siblings ...)
  2026-09-02 19:20 ` [RFC PATCH v2 6/9] igb: Quiesce VFs on STOP and include PF enable state in migration Cédric Le Goater
@ 2026-09-02 19:20 ` Cédric Le Goater
  2026-09-02 19:20 ` [RFC PATCH v2 8/9] igb: Add dirty page tracking statistics Cédric Le Goater
                   ` (2 subsequent siblings)
  9 siblings, 0 replies; 16+ messages in thread
From: Cédric Le Goater @ 2026-09-02 19:20 UTC (permalink / raw)
  To: qemu-devel
  Cc: Akihiko Odaki, Sriram Yagnaraman, Jason Wang, Alex Williamson,
	Peter Xu, Cédric Le Goater

After restoring VFRE/VFTE on the destination, the VF's RX rings may
be stalled because the guest driver is waiting for an interrupt that
was pending at the time of migration. Re-raise any pending EICR
causes for the VF (both RX and TX) via igb_set_eics() to break the
deadlock.

AI-used-for: analysis, code (prototype)
Signed-off-by: Cédric Le Goater <clg@redhat.com>
---
 hw/net/igb_core.h      |  1 +
 hw/net/igb_core.c      | 16 ++++++++++++++++
 hw/net/igb_migration.c |  4 ++--
 3 files changed, 19 insertions(+), 2 deletions(-)

diff --git a/hw/net/igb_core.h b/hw/net/igb_core.h
index 7fdc41690b36..7d6f44b7ffe9 100644
--- a/hw/net/igb_core.h
+++ b/hw/net/igb_core.h
@@ -150,6 +150,7 @@ igb_start_recv(IGBCore *core);
 IGBCore *igb_pf_get_core(void *pf);
 
 void igb_core_vf_propagate_irqs(IGBCore *core, uint16_t vfn);
+void igb_core_vf_rearm_irqs(IGBCore *core, uint16_t vfn);
 void igb_core_vf_propagate_ivar(IGBCore *core, uint16_t vfn);
 
 #endif
diff --git a/hw/net/igb_core.c b/hw/net/igb_core.c
index 12329a1aff58..1de8d95102e4 100644
--- a/hw/net/igb_core.c
+++ b/hw/net/igb_core.c
@@ -4632,3 +4632,19 @@ void igb_core_vf_propagate_ivar(IGBCore *core, uint16_t vfn)
         core->mac[IVAR0 + n / 4] &= ~mask;
     }
 }
+
+/*
+ * Re-apply VF interrupt enables to PF aggregates and raise the
+ * pending causes so the guest driver resumes polling after migration.
+ */
+void igb_core_vf_rearm_irqs(IGBCore *core, uint16_t vfn)
+{
+    uint32_t shift = 22 - vfn * IGBVF_MSIX_VEC_NUM;
+    uint32_t pvt_idx = PVTEICR0 + vfn * 0x40;
+    uint32_t causes = (core->mac[pvt_idx] & 0x7) << shift;
+
+    igb_core_vf_propagate_irqs(core, vfn);
+    if (causes) {
+        igb_set_eics(core, EICS, causes);
+    }
+}
diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
index 435bd0d1623b..46299ff77ab9 100644
--- a/hw/net/igb_migration.c
+++ b/hw/net/igb_migration.c
@@ -813,8 +813,8 @@ static void igb_core_vf_unquiesce(IgbVfState *s)
 
     trace_igbvf_mig_unquiesce(s->vfn, core->mac[VFRE], core->mac[VFTE]);
 
-    if (re) {
-        igb_start_recv(core);
+    if (re || te) {
+        igb_core_vf_rearm_irqs(core, s->vfn);
     }
 }
 
-- 
2.55.0



^ permalink raw reply related	[flat|nested] 16+ messages in thread

* [RFC PATCH v2 8/9] igb: Add dirty page tracking statistics
  2026-09-02 19:20 [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Cédric Le Goater
                   ` (6 preceding siblings ...)
  2026-09-02 19:20 ` [RFC PATCH v2 7/9] igb: Fix post-migration RX ring deadlock Cédric Le Goater
@ 2026-09-02 19:20 ` Cédric Le Goater
  2026-09-02 19:20 ` [RFC PATCH v2 9/9] docs: Add igb VF migration testing setup guide Cédric Le Goater
  2026-09-09  7:19 ` [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Akihiko Odaki
  9 siblings, 0 replies; 16+ messages in thread
From: Cédric Le Goater @ 2026-09-02 19:20 UTC (permalink / raw)
  To: qemu-devel
  Cc: Akihiko Odaki, Sriram Yagnaraman, Jason Wang, Alex Williamson,
	Peter Xu, Cédric Le Goater

Add a GET_STATS command (cmd 7) to the migration DVSEC for monitoring
dirty page tracking and DMA activity per VF. The device DMA-writes an
igb_mig_stats_resp struct to the shared buffer.

The dma_writes counter is also reported in the dirty query buffer so
the driver gets it alongside the bitmap without an extra config read.

Counters are reset on first DIRTY_ENABLE or device reset, so the
driver can read final values after DIRTY_DISABLE.

AI-used-for: code (prototype)
Signed-off-by: Cédric Le Goater <clg@redhat.com>
---
 hw/net/igb_core.h      |  1 +
 hw/net/igb_migration.h | 22 ++++++++++++++
 hw/net/igb_migration.c | 65 ++++++++++++++++++++++++++++++++++++++++--
 hw/net/trace-events    |  2 +-
 4 files changed, 86 insertions(+), 4 deletions(-)

diff --git a/hw/net/igb_core.h b/hw/net/igb_core.h
index 7d6f44b7ffe9..380cc88b7959 100644
--- a/hw/net/igb_core.h
+++ b/hw/net/igb_core.h
@@ -103,6 +103,7 @@ struct IGBCore {
     int64_t timadj;
 
     IGBVfDirtyState vf_dirty[IGB_MAX_VF_FUNCTIONS];
+    IgbVfMigStats vf_mig_stats[IGB_MAX_VF_FUNCTIONS];
 };
 
 void
diff --git a/hw/net/igb_migration.h b/hw/net/igb_migration.h
index 2bb9a0ed36ce..b334bb8f78ba 100644
--- a/hw/net/igb_migration.h
+++ b/hw/net/igb_migration.h
@@ -64,6 +64,7 @@
 #define IGB_MIG_CMD_DIRTY_ENABLE        4
 #define IGB_MIG_CMD_DIRTY_DISABLE       5
 #define IGB_MIG_CMD_DIRTY_QUERY         6
+#define IGB_MIG_CMD_GET_STATS           7
 
 /* STATUS register: state in [7:0], error code [15:8] */
 #define IGB_MIG_STATUS_STATE_MASK       0xFF
@@ -111,6 +112,15 @@ typedef struct IGBVfDirtyState {
     uint32_t num_ranges;
 } IGBVfDirtyState;
 
+typedef struct IgbVfMigStats {
+    uint64_t dma_writes;
+    uint64_t dma_bytes;
+    uint32_t dirty_pages_set;
+    uint32_t dirty_pages_cleared;
+    uint32_t dirty_page_count;
+    uint32_t dirty_query_count;
+} IgbVfMigStats;
+
 typedef struct IgbVfMigState {
     uint32_t mig_state;
     uint32_t mig_data[IGB_VF_STATE_MAX_SIZE / sizeof(uint32_t)];
@@ -149,6 +159,18 @@ struct igb_mig_dirty_query {
     uint8_t bitmap[];
 };
 
+/*
+ * GET_STATS:    device writes igb_mig_stats_resp to buffer.
+ */
+struct igb_mig_stats_resp {
+    uint64_t dma_writes;
+    uint64_t dma_bytes;
+    uint32_t dirty_pages_set;
+    uint32_t dirty_pages_cleared;
+    uint32_t dirty_page_count;
+    uint32_t dirty_query_count;
+};
+
 typedef struct IGBCore IGBCore;
 typedef struct IgbVfState IgbVfState;
 
diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
index 46299ff77ab9..2504d5bd308b 100644
--- a/hw/net/igb_migration.c
+++ b/hw/net/igb_migration.c
@@ -438,6 +438,7 @@ void igb_core_dirty_track_dma(IGBCore *core, int vfn,
                               dma_addr_t addr, dma_addr_t len)
 {
     IGBVfDirtyState *ds = &core->vf_dirty[vfn];
+    IgbVfMigStats *stats = &core->vf_mig_stats[vfn];
     bool matched = false;
     uint32_t i;
 
@@ -447,6 +448,9 @@ void igb_core_dirty_track_dma(IGBCore *core, int vfn,
 
     trace_igb_core_dirty_track_dma(vfn, addr, len);
 
+    stats->dma_writes++;
+    stats->dma_bytes += len;
+
     for (i = 0; i < ds->num_ranges; i++) {
         IGBVfDirtyRange *r = &ds->ranges[i];
         uint64_t r_end = r->iova + r->size;
@@ -466,7 +470,10 @@ void igb_core_dirty_track_dma(IGBCore *core, int vfn,
 
         for (page = start_page; page <= end_page; page++) {
             if (page < r->nbits) {
-                set_bit(page, r->bitmap);
+                if (!test_and_set_bit(page, r->bitmap)) {
+                    stats->dirty_pages_set++;
+                    stats->dirty_page_count++;
+                }
             }
         }
     }
@@ -493,6 +500,15 @@ static uint32_t igb_core_vf_dirty_enable(IgbVfState *s, uint64_t pgsize,
     IGBVfDirtyState *ds = igb_core_vf_dirty_state(s);
     IGBVfDirtyRange *r;
 
+    /*
+     * Reset stats on first enable so the driver can read them after
+     * disable
+     */
+    if (ds->num_ranges == 0) {
+        memset(&igbvf_get_core(s)->vf_mig_stats[s->vfn], 0,
+               sizeof(IgbVfMigStats));
+    }
+
     if (ds->num_ranges >= IGB_MIG_CAPS_MAX_RANGES) {
         return IGB_MIG_ERR_TOO_MANY_RANGES;
     }
@@ -635,12 +651,14 @@ static IGBVfDirtyRange *igb_core_vf_dirty_range_valid(IGBVfDirtyState *ds,
 static uint8_t igbvf_mig_cmd_dirty_query(IgbVfState *s)
 {
     IgbVfMigState *ms = &s->mig;
+    IgbVfMigStats *stats = &igbvf_get_core(s)->vf_mig_stats[s->vfn];
     IGBVfDirtyState *ds = &igbvf_get_core(s)->vf_dirty[s->vfn];
     uint64_t buf_addr = ms->mig_data_buf_addr;
     uint64_t range_iova = 0, range_size = 0;
     IGBVfDirtyRange *range;
     uint32_t bmp_bytes, dirty_pages;
     uint32_t val32;
+    uint64_t val64;
     size_t out_size;
     g_autofree void *bitmap = NULL;
 
@@ -686,6 +704,10 @@ static uint8_t igbvf_mig_cmd_dirty_query(IgbVfState *s)
 
     dirty_pages = bitmap_count_one(bitmap, range_size / range->page_size);
 
+    stats->dirty_pages_cleared += dirty_pages;
+    stats->dirty_page_count -= MIN(stats->dirty_page_count, dirty_pages);
+    stats->dirty_query_count++;
+
     val32 = cpu_to_le32(out_size);
     address_space_write(&address_space_memory,
                         buf_addr + offsetof(struct igb_mig_dirty_query,
@@ -696,8 +718,13 @@ static uint8_t igbvf_mig_cmd_dirty_query(IgbVfState *s)
                         buf_addr + offsetof(struct igb_mig_dirty_query,
                                             dirty_page_count),
                         MEMTXATTRS_UNSPECIFIED, &val32, sizeof(val32));
-
-    trace_igbvf_mig_dirty_query(s->vfn, (uint64_t)out_size, dirty_pages);
+    val64 = cpu_to_le64(stats->dma_writes);
+    address_space_write(&address_space_memory,
+                        buf_addr + offsetof(struct igb_mig_dirty_query,
+                                            dma_writes),
+                        MEMTXATTRS_UNSPECIFIED, &val64, sizeof(val64));
+    trace_igbvf_mig_dirty_query(s->vfn, (uint64_t)out_size, dirty_pages,
+                                stats->dma_writes);
     return 0;
 }
 
@@ -898,6 +925,32 @@ static uint8_t igbvf_mig_set_state(IgbVfState *s, uint32_t new_state)
     return 0;
 }
 
+static uint8_t igbvf_mig_cmd_get_stats(IgbVfState *s)
+{
+    IgbVfMigState *ms = &s->mig;
+    IgbVfMigStats *stats = &igbvf_get_core(s)->vf_mig_stats[s->vfn];
+    struct igb_mig_stats_resp resp;
+    MemTxResult r;
+
+    if (!ms->mig_data_buf_addr) {
+        return IGB_MIG_ERR_NO_BUFFER;
+    }
+
+    resp.dma_writes = cpu_to_le64(stats->dma_writes);
+    resp.dma_bytes = cpu_to_le64(stats->dma_bytes);
+    resp.dirty_pages_set = cpu_to_le32(stats->dirty_pages_set);
+    resp.dirty_pages_cleared = cpu_to_le32(stats->dirty_pages_cleared);
+    resp.dirty_page_count = cpu_to_le32(stats->dirty_page_count);
+    resp.dirty_query_count = cpu_to_le32(stats->dirty_query_count);
+
+    r = address_space_write(&address_space_memory, ms->mig_data_buf_addr,
+                            MEMTXATTRS_UNSPECIFIED, &resp, sizeof(resp));
+    if (r != MEMTX_OK) {
+        return IGB_MIG_ERR_DMA_FAILED;
+    }
+    return 0;
+}
+
 static void igbvf_mig_update_status(IgbVfState *s, uint8_t err)
 {
     IgbVfMigState *ms = &s->mig;
@@ -955,6 +1008,10 @@ static void igbvf_mig_cmd_ctrl(IgbVfState *s, uint32_t val)
         err = igbvf_mig_cmd_dirty_query(s);
         break;
 
+    case IGB_MIG_CMD_GET_STATS:
+        err = igbvf_mig_cmd_get_stats(s);
+        break;
+
     default:
         err = IGB_MIG_ERR_UNK_CMD;
         break;
@@ -1059,6 +1116,8 @@ void igbvf_mig_state_reset(IgbVfState *s)
     memset(ms->mig_data, 0, sizeof(ms->mig_data));
     ms->mig_saved_vfre = true;
     ms->mig_saved_vfte = true;
+    memset(&igbvf_get_core(s)->vf_mig_stats[s->vfn], 0,
+           sizeof(IgbVfMigStats));
 
     pci_set_long(PCI_DEVICE(s)->config +
                  IGB_MIG_DVSEC_OFFSET + IGB_MIG_BUF_ADDR_LO, 0);
diff --git a/hw/net/trace-events b/hw/net/trace-events
index eb5c61a8f6fb..6c2c5493094a 100644
--- a/hw/net/trace-events
+++ b/hw/net/trace-events
@@ -302,7 +302,7 @@ igbvf_mig_load_state(uint16_t vfn, uint32_t size, bool vfre, bool vfte) "VF%u: l
 igbvf_mig_reset(uint16_t vfn) "VF%u: migration state reset"
 igbvf_mig_dirty_enable(uint16_t vfn, uint64_t pgsize, uint64_t nbits) "VF%u: dirty tracking enabled pgsize=%"PRIu64" nbits=%"PRIu64
 igbvf_mig_dirty_disable(uint16_t vfn) "VF%u: dirty tracking disabled"
-igbvf_mig_dirty_query(uint16_t vfn, uint64_t size, uint32_t dirty_pages) "VF%u: dirty query returned %"PRIu64" bytes, %u dirty pages"
+igbvf_mig_dirty_query(uint16_t vfn, uint64_t size, uint32_t dirty_pages, uint64_t dma_writes) "VF%u: dirty query returned %"PRIu64" bytes, %u dirty pages (dma_writes=%"PRIu64")"
 igb_core_dirty_track_dma(int vfn, uint64_t addr, uint64_t len) "VF%d: dirty DMA addr=0x%"PRIx64" len=%"PRIu64
 igb_core_dirty_track_dma_drop(int vfn, uint64_t addr, uint64_t len) "VF%d: dirty DMA dropped addr=0x%"PRIx64" len=%"PRIu64" no matching range"
 igbvf_mig_quiesce(uint16_t vfn, uint32_t vfre, uint32_t vfte) "VF%u: quiesce VFRE=0x%x VFTE=0x%x"
-- 
2.55.0



^ permalink raw reply related	[flat|nested] 16+ messages in thread

* [RFC PATCH v2 9/9] docs: Add igb VF migration testing setup guide
  2026-09-02 19:20 [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Cédric Le Goater
                   ` (7 preceding siblings ...)
  2026-09-02 19:20 ` [RFC PATCH v2 8/9] igb: Add dirty page tracking statistics Cédric Le Goater
@ 2026-09-02 19:20 ` Cédric Le Goater
  2026-09-08  8:39   ` Akihiko Odaki
  2026-09-09  7:19 ` [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Akihiko Odaki
  9 siblings, 1 reply; 16+ messages in thread
From: Cédric Le Goater @ 2026-09-02 19:20 UTC (permalink / raw)
  To: qemu-devel
  Cc: Akihiko Odaki, Sriram Yagnaraman, Jason Wang, Alex Williamson,
	Peter Xu, Cédric Le Goater

Document the igb VF migration interface: DVSEC register layout, dirty
page tracking, testing setup with nested virtualization.

AI-used-for: docs
Signed-off-by: Cédric Le Goater <clg@redhat.com>
---
 MAINTAINERS                           |   1 +
 docs/system/device-emulation.rst      |   1 +
 docs/system/devices/igb-migration.rst | 417 ++++++++++++++++++++++++++
 docs/system/devices/igb.rst           |   6 +
 4 files changed, 425 insertions(+)
 create mode 100644 docs/system/devices/igb-migration.rst

diff --git a/MAINTAINERS b/MAINTAINERS
index f88b526be238..4a2f1357e936 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -2820,6 +2820,7 @@ igb VF migration
 M: Cédric Le Goater <clg@redhat.com>
 S: Maintained
 F: hw/net/igb_migration.*
+F: docs/system/devices/igb-migration.rst
 
 eepro100
 M: Stefan Weil <sw@weilnetz.de>
diff --git a/docs/system/device-emulation.rst b/docs/system/device-emulation.rst
index 40054bb7dfcc..75f423b795cd 100644
--- a/docs/system/device-emulation.rst
+++ b/docs/system/device-emulation.rst
@@ -90,6 +90,7 @@ Emulated Devices
    devices/cxl.rst
    devices/emmc.rst
    devices/igb.rst
+   devices/igb-migration.rst
    devices/ivshmem-flat.rst
    devices/ivshmem.rst
    devices/keyboard.rst
diff --git a/docs/system/devices/igb-migration.rst b/docs/system/devices/igb-migration.rst
new file mode 100644
index 000000000000..d72f14f5fe63
--- /dev/null
+++ b/docs/system/devices/igb-migration.rst
@@ -0,0 +1,417 @@
+.. SPDX-License-Identifier: GPL-2.0-or-later
+.. _igb-migration:
+
+igb VF Migration
+----------------
+
+Live migration of VFIO-passthrough devices (SR-IOV VFs, vGPUs) is a
+growing requirement, but real hardware with migration support is scarce
+and hard to debug. An emulated device provides a fully controlled
+testbed for developing and validating the entire software stack --
+vfio-pci variant drivers, VFIO core migration v2 framework, QEMU,
+libvirt -- and for tuning complex migration policies such as downtime
+convergence. It also serves as an educational reference for
+understanding VFIO migration end-to-end, from device state
+serialization to dirty page tracking.
+
+The igb device supports an experimental VF migration interface that allows
+the `igb-vfio-pci`_ variant driver to migrate VF state during live
+migration using the standard VFIO migration v2 protocol with stop-copy
+and pre-copy support.
+
+This is enabled with the ``x-vf-migration`` property::
+
+  -device igb,x-vf-migration=on,...
+
+Each emulated VF then advertises a DVSEC discovered by the
+`igb-vfio-pci`_ variant driver at bind time. This feature is
+experimental (``x-`` prefix, default off).
+
+Architecture
+~~~~~~~~~~~~
+
+The target scenario is nested virtualization::
+
+  L0 QEMU
+    igb PF with x-vf-migration=on
+    └── VFs with migration DVSEC
+
+  L1 kernel
+    igb-vfio-pci variant driver
+    translates VFIO migration v2 ioctls → DVSEC config writes
+
+  L1 QEMU (stock, unmodified)
+    vfio-pci device model, standard migration fd
+
+  L2 guest
+    standard igbvf driver, unaware of migration
+
+The L1 QEMU is completely unmodified -- it sees a standard VFIO
+migratable device and uses the normal migration fd path. The
+`igb-vfio-pci`_ variant driver handles the translation between
+VFIO migration v2 ioctls and DVSEC config writes.
+
+Design
+~~~~~~
+
+The migration interface is exposed through a DVSEC at offset 0x160
+in VF extended config space (see `DVSEC register layout`_ below for
+the full register map).
+
+Device state is serialized as a versioned blob of per-VF register
+(offset, value) pairs covering control, interrupt, RX/TX queue,
+receive address (RA/RA2), etc. plus TX context descriptors and
+VFRE/VFTE enable bits. The buffer address is a guest physical address
+(GPA) written by the driver via ``virt_to_phys``; the device accesses
+guest RAM directly through the system address space.
+
+Dirty page tracking is implemented with per-range bitmaps maintained
+in IGBCore. All VF DMA paths in ``igb_core.c`` (TX data, RX data,
+descriptor writeback) are instrumented to record touched pages. The
+`igb-vfio-pci`_ variant driver registers tracked IOVA ranges and
+queries dirty bitmaps through a shared buffer. Buffer structures
+include len, flags, and reserved fields for future extensibility.
+
+The dirty bitmaps are maintained inside the device, which is not
+realistic for discrete NICs without on-chip DRAM.
+
+DVSEC register layout
+~~~~~~~~~~~~~~~~~~~~~
+
+The migration DVSEC (36 bytes at offset ``0x160``) uses a command
+doorbell model. All commands are synchronous -- the device completes
+the operation before the config write returns::
+
+  Offset  Name          Access  Description
+  +0x00   ExtCap Hdr    RO      PCIe extended cap (id=0x23, ver=1)
+  +0x04   DVSEC Hdr 1   RO      length[31:20] | rev[19:16] | vendor_id[15:0]
+  +0x08   DVSEC Hdr 2   RO      DVSEC ID (1)
+  +0x0A   Reserved      -       Padding for DWORD alignment
+  +0x0C   CAPS          RO      F_STATE[0], F_DIRTY[1], max_ranges[11:8], pgsize[16:12]
+  +0x10   CTRL          WO      Doorbell: cmd[7:0], arg[31:8]
+  +0x14   STATUS        RO      state[7:0], error_code[15:8], quiesced[16]
+  +0x18   BUF_ADDR_LO   RW      Shared DMA buffer GPA (low 32 bits)
+  +0x1C   BUF_ADDR_HI   RW      Shared DMA buffer GPA (high 32 bits)
+  +0x20   DATA_SIZE     RO      State blob size in bytes
+
+CTRL commands::
+
+  Cmd  Name            Arg             Description
+  1    SET_STATE       state[31:8]     Set migration state
+  2    SAVE            -               DMA-write state to buffer
+  3    LOAD            size[31:8]      DMA-read state from buffer
+  4    DIRTY_ENABLE    -               Enable dirty tracking (params in DMA buffer)
+  5    DIRTY_DISABLE   -               Disable dirty tracking
+  6    DIRTY_QUERY     -               Query dirty bitmap (via DMA buffer)
+  7    GET_STATS       -               Query statistics (via DMA buffer)
+
+The driver sets ``BUF_ADDR_LO/HI`` before issuing commands that use a
+DMA buffer (SAVE, LOAD, DIRTY_ENABLE, DIRTY_QUERY, GET_STATS). The
+buffer address is latched per CTRL write, so the driver can use
+different buffers for different commands. The buffer address is a
+guest physical address (GPA).
+
+State transitions follow the VFIO migration v2 state machine. The
+driver issues ``SET_STATE`` with the target state in the arg field and
+reads ``STATUS`` to confirm the transition. Device states::
+
+  0  ERROR       1  STOP       2  RUNNING
+  3  STOP_COPY   4  RESUMING   5  PRE_COPY
+
+``DATA_SIZE`` reflects the state blob size. At reset and in ``STOP``
+state it holds the maximum size the driver should allocate. After
+``SET_STATE(STOP_COPY)`` or ``SAVE`` it holds the actual serialized
+size. The driver reads it after entering ``STOP_COPY`` to allocate an
+exact-sized DMA buffer before issuing ``SAVE``.
+
+The state blob is a versioned sequence of register (offset, value)
+pairs with magic ``0x4D494742`` ("MIGB").
+
+When ``STATUS`` state is ``ERROR`` (0), bits [15:8] contain an error
+code identifying the failure::
+
+  1   UNK_CMD           Unknown CTRL command
+  2   BAD_STATE         Command issued in wrong migration state
+  3   NO_BUFFER         Command requires buffer but BUF_ADDR not set
+  4   DMA_FAILED        DMA transfer to/from buffer failed
+  5   BAD_SIZE          State blob too large or empty
+  6   BAD_MAGIC         State blob magic mismatch
+  7   BAD_VERSION       State blob version mismatch
+  8   TOO_MANY_RANGES   Exceeds max_ranges from CAPS
+  9   BAD_RANGE         Invalid range (zero size, misaligned, not contained)
+  10  BAD_PGSIZE        Invalid or misaligned page size
+  11  NOT_ENABLED       Dirty query without prior enable
+
+Dirty page tracking
+~~~~~~~~~~~~~~~~~~~
+
+The migration interface supports per-VF dirty page tracking, advertised
+by the ``F_DIRTY`` flag (bit 1) in ``CAPS``. This allows the variant
+driver to enter ``PRE_COPY`` state while the VM continues to run,
+iterating on dirty pages to reduce the final stop-and-copy window.
+
+The device maintains one dirty tracking engine per range, each with its
+own bitmap scoped to the range boundaries. The ``CAPS`` register
+advertises the maximum number of ranges in bits [11:8].
+
+Dirty tracking is controlled through CTRL commands:
+
+- **DIRTY_ENABLE** (4): the driver fills an ``igb_mig_dirty_enable_req``
+  struct in the DMA buffer with the page size, range IOVA, and range
+  size, then issues the command. The device allocates a bitmap for the
+  range and begins recording pages touched by DMA. Supported page sizes
+  are advertised in ``CAPS`` bits [16:12] (bit N = 2^N bytes). The driver
+  checks ``STATUS`` for errors after the command completes.
+- **DIRTY_DISABLE** (5): tears down all ranges and stops tracking.
+- **DIRTY_QUERY** (6): the driver writes (iova, size) into the
+  ``igb_mig_dirty_query`` DMA buffer, then issues the command. The
+  device validates the range, copies the dirty bitmap into the buffer,
+  and clears the tracked bits after successful DMA. The driver checks
+  ``STATUS`` for errors after the command.
+
+Dirty enable DMA buffer
+~~~~~~~~~~~~~~~~~~~~~~~
+
+The ``DIRTY_ENABLE`` command reads its parameters from the DMA buffer.
+The ``len`` field holds the total structure size (including reserved
+bytes) so the device can detect newer formats. ``flags`` and
+``reserved`` must be zero::
+
+  Offset  Field        Type      Description
+  0x00    len          uint32    Structure size in bytes
+  0x04    flags        uint32    Reserved, must be 0
+  0x08    pgsize       uint64    Page granularity (must match a CAPS pgsize bit)
+  0x10    range_iova   uint64    Tracked range start address
+  0x18    range_size   uint64    Tracked range size in bytes
+  0x20    reserved[4]  uint32    Reserved, must be 0
+
+Dirty query DMA buffer
+~~~~~~~~~~~~~~~~~~~~~~
+
+The ``DIRTY_QUERY`` command uses a shared DMA buffer for both request
+and response. The ``len`` field holds the total buffer size (header +
+bitmap). ``flags`` and ``reserved`` must be zero::
+
+  Offset  Field              Written by  Description
+  0x00    len                driver      Total buffer size in bytes
+  0x04    flags              driver      Reserved, must be 0
+  0x08    iova               driver      Query range start
+  0x10    size               driver      Query range size
+  0x18    bitmap_size        device      Bytes written to bitmap
+  0x1C    dirty_page_count   device      Number of set bits
+  0x20    dma_writes         device      DMA write count (diagnostic)
+  0x28    reserved[6]        -           Reserved, must be 0
+  0x40    bitmap[]           device      Dirty page bitmap
+
+Migration statistics
+~~~~~~~~~~~~~~~~~~~~
+
+The ``GET_STATS`` command DMA-writes a statistics response into the
+driver-provided buffer. The driver sets ``BUF_ADDR_LO/HI`` and issues
+the command; the device writes the response and returns::
+
+  Offset  Field                Type      Description
+  0x00    dma_writes           uint64    DMA write operations tracked
+  0x08    dma_bytes            uint64    DMA bytes written
+  0x10    dirty_pages_set      uint32    Dirty pages marked since enable
+  0x14    dirty_pages_cleared  uint32    Dirty pages cleared by queries
+  0x18    dirty_page_count     uint32    Current dirty pages (set - cleared)
+  0x1C    dirty_query_count    uint32    Number of QUERY operations
+
+The variant driver exposes these via debugfs at
+``/sys/kernel/debug/vfio/<device>/migration/dirty/stats``.
+
+Testing setup
+~~~~~~~~~~~~~
+
+The target scenario is nested virtualization: L0 runs QEMU with an
+igb PF (``x-vf-migration=on``), L1 runs the `igb-vfio-pci`_ variant
+driver and an unmodified QEMU, and L2 runs a standard igbvf driver.
+See `Architecture`_ above for the full stack diagram.
+
+NetworkManager configuration
+^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
+
+In a nested setup, the L1 VMs (source and destination) have emulated
+igb PFs connected to the L0 bridge. By default, NetworkManager
+acquires DHCP leases on those PF interfaces and on any igb VFs created
+later. This causes the VF MAC address to be learned on the L0 bridge,
+which can misdirect iperf3 traffic after migration.
+
+To prevent this, configure NetworkManager on **both L1 VMs** and on the
+**L2 guest disk image**.
+
+L1 VMs (source and destination)
+...............................
+
+1. Prevent NetworkManager from managing igbvf interfaces:
+
+.. code-block:: bash
+
+   cat > /etc/NetworkManager/conf.d/99-no-igbvf.conf <<EOF
+   [keyfile]
+   unmanaged-devices=driver:igbvf
+   EOF
+
+2. Disable IP on the igb PF connections (keep the interfaces UP for
+   bridging, but with no DHCP lease):
+
+.. code-block:: bash
+
+   # Identify the NM connections for the igb PFs (NOT the virtio management NIC)
+   nmcli -t -f NAME,DEVICE connection show
+
+   # For each igb PF connection:
+   nmcli connection modify "<igb-pf-connection>" ipv4.method disabled ipv6.method disabled
+
+3. Reload NetworkManager:
+
+.. code-block:: bash
+
+   nmcli general reload
+
+L2 guest disk image
+...................
+
+Use ``virt-customize`` to add the igbvf unmanaged config to the guest
+image (offline, before any test run):
+
+.. code-block:: bash
+
+   virt-customize -a /srv/migration/rhel10.qcow2 \
+     --write /etc/NetworkManager/conf.d/99-no-igbvf.conf:'[keyfile]
+   unmanaged-devices=driver:igbvf'
+
+Network diagram
+^^^^^^^^^^^^^^^
+
+The diagram below shows the nested setup where the source and
+destination hosts are themselves VMs (L1) running on a physical
+host (L0) that emulates the igb NIC::
+
+  ┌────────────────────────────────────────────────────────────────────────────┐
+  │  L0: physical host                                                         │
+  │                                                                            │
+  │  virbr0  192.168.199.1/24                                                  │
+  │  ├── NFS server: /srv/migration                                            │
+  │  └── iperf3 client: iperf3 -c 192.168.199.200 -t 60 -i 1                   │
+  │      │                                                                     │
+  │      │  L0 virbr0 bridge (192.168.199.0/24)                                │
+  │  ────┼──────────┬──────────────────┬──────────────────────────────         │
+  │      │          │                  │                                       │
+  │      │     ┌────┴────┐        ┌────┴────┐                                  │
+  │      │     │ virtio  │        │ emulated│                                  │
+  │      │     │c0:ff:ee:│        │  igb PF │    L0 QEMU (vm6)                 │
+  │      │     │ :00:06  │        │ + igbvf │    tracks DMA dirty pages        │
+  │      │     └────┬────┘        └────┬────┘                                  │
+  │      │          │                  │                                       │
+  │  ┌───┼──────────┼──────────────────┼──────────────────────────────────┐    │
+  │  │   │  L1: vm6 (source)           │                                  │    │
+  │  │   │  enp1s0: 192.168.199.6      │                                  │    │
+  │  │   │  (management)               │                                  │    │
+  │  │   │                        enp8s0 (igb PF, no IP)                  │    │
+  │  │   │                             │                                  │    │
+  │  │   │                        igb VF0 ──► igb-vfio-pci (VFIO)         │    │
+  │  │   │                             │      dirty_sync → L0 igbvf       │    │
+  │  │   │                             │                                  │    │
+  │  │   │   virbr0                    │ VFIO passthrough                 │    │
+  │  │   │   192.168.200.1/24          │                                  │    │
+  │  │   │       │                     │                                  │    │
+  │  │   │  ┌────┼─────────────────────┼───────────────────────────┐      │    │
+  │  │   │  │    │  L2: rhel10 guest   │                           │      │    │
+  │  │   │  │    │                     │                           │      │    │
+  │  │   │  │  virtio NIC           igb VF (enp7s0)                │      │    │
+  │  │   │  │  192.168.200.130/24   192.168.199.200/24             │      │    │
+  │  │   │  │  (SSH login)          (iperf3 data path)             │      │    │
+  │  │   │  │                          │                           │      │    │
+  │  │   │  │            iperf3 -s -D  │ (listens on 0.0.0.0)      │      │    │
+  │  │   │  └──────────────────────────┼───────────────────────────┘      │    │
+  │  │   │                             │                                  │    │
+  │  │   │  virsh migrate --live ──────┼──────────────────► vm7           │    │
+  │  │   │                             │                                  │    │
+  │  └───┼─────────────────────────────┼──────────────────────────────────┘    │
+  │      │                             │                                       │
+  │      │          iperf3 traffic     │                                       │
+  │      └─────────────────────────────┘                                       │
+  │                                                                            │
+  │  ────────────────────────────────────────────────────────────────          │
+  │      │                  │                                                  │
+  │      │     ┌────────┐   │   ┌─────────┐                                    │
+  │      │     │ virtio │   │   │emulated │    L0 QEMU (vm7)                   │
+  │      │     │c0:ff:ee│   │   │ igb PF  │                                    │
+  │      │     │ :00:07 │   │   │ + igbvf │                                    │
+  │      │     └────┬───┘   │   └────┬────┘                                    │
+  │  ┌──────────────┼───────┼────────┼────────────────────────────────────┐    │
+  │  │   L1: vm7 (destination)       │                                    │    │
+  │  │   enp1s0: 192.168.199.7       │                                    │    │
+  │  │   (management)           enp8s0 (igb PF, no IP)                    │    │
+  │  │                               │                                    │    │
+  │  │                          igb VF0 ──► igb-vfio-pci (VFIO)           │    │
+  │  │                               │                                    │    │
+  │  │   virbr0                      │ VFIO passthrough                   │    │
+  │  │   192.168.200.1/24            │                                    │    │
+  │  │       │                       │                                    │    │
+  │  │  ┌────┼───────────────────────┼────────────────────────────┐       │    │
+  │  │  │    │  L2: rhel10 (after migration)                      │       │    │
+  │  │  │    │                       │                            │       │    │
+  │  │  │  virtio NIC             igb VF (enp7s0)                 │       │    │
+  │  │  │  192.168.200.130/24     192.168.199.200/24              │       │    │
+  │  │  │                            │                            │       │    │
+  │  │  │              iperf3 -s -D  │ (connection survives)      │       │    │
+  │  │  └────────────────────────────┼────────────────────────────┘       │    │
+  │  └───────────────────────────────┼────────────────────────────────────┘    │
+  │                                  │                                         │
+  │      iperf3 traffic resumes ─────┘                                         │
+  │      (same IP, same MAC, same L2 segment → transparent to client)          │
+  └────────────────────────────────────────────────────────────────────────────┘
+
+Migration under iperf3 load works correctly: dirty page tracking
+converges (from ~2000 pages per PRE_COPY iteration down to ~280 at
+STOP_COPY), and STOP_COPY stays under 250ms.
+
+Todo
+~~~~
+
+1. Add migration blocker when ``x-vf-migration=on`` (no VMState yet) or
+   add VMState support for L0 migration (dirty bitmaps, tracking
+   engines, DVSEC registers, stats)
+2. Add PRE_COPY state transfer to validate device INIT data (magic,
+   version, etc.)
+3. Add qtests for migration state machine transitions, dirty page
+   tracking
+
+Ideas
+~~~~~
+
+1. **RX bandwidth throttle** (``x-mig-rx-limit``, uint32, default 0)
+
+   Return false from ``can_receive`` when the per-VF packet count in the
+   current tracking interval exceeds the limit. Reduces DMA writes and
+   dirty pages realistically.
+
+2. **Migration phase timing** (GET_STATS extension)
+
+   Add per-VF timestamps: ``precopy_start_ns``, ``stopcopy_start_ns``,
+   ``precopy_duration_ns``, ``stopcopy_duration_ns``,
+   ``state_transition_count``. Expose via GET_STATS.
+
+3. **Hot page simulation** (``x-mig-hot-pages``, uint32, default 0)
+
+   Re-set the first N bitmap bits after each DIRTY_QUERY, simulating
+   workloads with hot pages that prevent convergence.
+
+4. **Error injection** (``x-mig-inject-error``, uint32, default 0)
+
+   One-shot error code injection before command dispatch. A separate
+   ``x-mig-inject-dma-fail`` (bool) for persistent DMA failure testing.
+
+AI disclaimer
+~~~~~~~~~~~~~
+
+Claude was used to analyze the IGB PF and VF internal state and
+identify the pain points of a working live migration of such devices.
+The generated code served as a starting point but *significant* time
+was then spent cleaning up, reworking, and shaping it into a clear,
+reviewable IGB model extension.
+
+.. _igb-vfio-pci: https://github.com/legoater/vfio-pci-extras
diff --git a/docs/system/devices/igb.rst b/docs/system/devices/igb.rst
index 50f625fd77e4..00271dbc92c3 100644
--- a/docs/system/devices/igb.rst
+++ b/docs/system/devices/igb.rst
@@ -64,6 +64,12 @@ command:
 
   pyvenv/bin/meson test --suite thorough func-x86_64-netdev_ethtool
 
+VF Migration (experimental)
+===========================
+
+See :ref:`igb-migration` for details on the experimental VF live migration
+interface.
+
 References
 ==========
 
-- 
2.55.0



^ permalink raw reply related	[flat|nested] 16+ messages in thread

* Re: [RFC PATCH v2 1/9] igb: Add x-vf-migration property and DVSEC extended capability
  2026-09-02 19:20 ` [RFC PATCH v2 1/9] igb: Add x-vf-migration property and DVSEC extended capability Cédric Le Goater
@ 2026-09-03 19:57   ` Alex Williamson
  2026-09-07 20:58     ` Cédric Le Goater
  0 siblings, 1 reply; 16+ messages in thread
From: Alex Williamson @ 2026-09-03 19:57 UTC (permalink / raw)
  To: Cédric Le Goater
  Cc: qemu-devel, Akihiko Odaki, Sriram Yagnaraman, Jason Wang,
	Peter Xu, alex

On Wed,  2 Sep 2026 21:20:46 +0200
Cédric Le Goater <clg@redhat.com> wrote:

> diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
> new file mode 100644
> index 000000000000..4dfebd82344c
> --- /dev/null
> +++ b/hw/net/igb_migration.c
> @@ -0,0 +1,106 @@
> +/*
> + * QEMU Intel 82576 SR/IOV VF Migration Support
> + *
> + * Copyright (c) 2026 Red Hat, Inc.
> + *
> + * SPDX-License-Identifier: GPL-2.0-or-later
> + */
> +
> +#include "qemu/osdep.h"
> +#include "hw/pci/pci_device.h"
> +#include "hw/pci/pcie.h"
> +#include "igb_common.h"
> +#include "igb_migration.h"
> +
> +static void igbvf_mig_update_status(IgbVfState *s, uint8_t err)
> +{
> +    IgbVfMigState *ms = &s->mig;
> +    PCIDevice *dev = PCI_DEVICE(s);
> +    uint32_t status;
> +
> +    status = ms->mig_state & IGB_MIG_STATUS_STATE_MASK;
> +
> +    if (err) {
> +        status = IGB_MIG_STATE_ERROR | IGB_MIG_STATUS_ERR(err);
> +    }
> +
> +    pci_set_long(dev->config + IGB_MIG_DVSEC_OFFSET + IGB_MIG_STATUS, status);
> +}
> +
> +
> +bool igbvf_add_migration_dvsec(PCIDevice *dev, Error **errp)
> +{
> +    uint16_t offset = IGB_MIG_DVSEC_OFFSET;
> +    uint32_t caps;
> +
> +    pcie_add_capability(dev, PCI_EXT_CAP_ID_DVSEC, 1, offset,
> +                        IGB_MIG_DVSEC_SIZE);
> +
> +    /* DVSEC header 1: length[31:20] | rev[19:16] | vendor_id[15:0] */
> +    pci_set_long(dev->config + offset + 0x4,
> +                 (IGB_MIG_DVSEC_SIZE << 20) |
> +                 (IGB_MIG_DVSEC_VER << 16) |
> +                 PCI_VENDOR_ID_INTEL);

Let's not co-opt an Intel DVSEC ID, we should use a vendor ID that we
have a claim to, the RedHat/Qumranet one, I'd guess.  We may need to
add DVSEC IDs to the spreadsheet for whoever is tracking Device IDs.
Gerd?

> +
> +    /* DVSEC header 2: DVSEC ID */
> +    pci_set_word(dev->config + offset + 0x8, IGB_MIG_DVSEC_ID);
> +
> +    /* CAPS: features (state migration only) */
> +    caps = IGB_MIG_CAP_F_STATE;
> +    pci_set_long(dev->config + offset + IGB_MIG_CAPS, caps);
> +
> +    /* STATUS: initial state is RUNNING */
> +    pci_set_long(dev->config + offset + IGB_MIG_STATUS,
> +                 IGB_MIG_STATE_RUNNING);
> +
> +    /* BUF_ADDR_LO and BUF_ADDR_HI are writable */
> +    memset(dev->wmask + offset + IGB_MIG_BUF_ADDR_LO, 0xff, 4);
> +    memset(dev->wmask + offset + IGB_MIG_BUF_ADDR_HI, 0xff, 4);
> +
> +    return true;
> +}
...
> diff --git a/hw/net/igbvf.c b/hw/net/igbvf.c
> index 9a165c7063ee..30dfdb574ac7 100644
> --- a/hw/net/igbvf.c
> +++ b/hw/net/igbvf.c
> @@ -38,27 +38,21 @@
>   */
>  
>  #include "qemu/osdep.h"
> +#include "qemu/range.h"
>  #include "hw/core/hw-error.h"
>  #include "hw/net/mii.h"
>  #include "hw/pci/pci_device.h"
>  #include "hw/pci/pcie.h"
> +#include "hw/pci/pcie_sriov.h"
>  #include "hw/pci/msix.h"
>  #include "net/eth.h"
>  #include "net/net.h"
>  #include "igb_common.h"
>  #include "igb_core.h"
> +#include "igb_migration.h"
>  #include "trace.h"
>  #include "qapi/error.h"
>  
> -OBJECT_DECLARE_SIMPLE_TYPE(IgbVfState, IGBVF)
> -
> -struct IgbVfState {
> -    PCIDevice parent_obj;
> -
> -    MemoryRegion mmio;
> -    MemoryRegion msix;
> -};
> -
>  static hwaddr vf_to_pf_addr(hwaddr addr, uint16_t vfn, bool write)
>  {
>      switch (addr) {
> @@ -199,10 +193,35 @@ static hwaddr vf_to_pf_addr(hwaddr addr, uint16_t vfn, bool write)
>      return HWADDR_MAX;
>  }
>  
> +static bool igbvf_addr_in_dvsec(uint32_t addr, int len)
> +{
> +    return ranges_overlap(addr, len,
> +                          IGB_MIG_DVSEC_OFFSET, IGB_MIG_DVSEC_SIZE);
> +}
> +
> +static uint32_t igbvf_read_config(PCIDevice *dev, uint32_t addr, int size)
> +{
> +    IgbVfState *s = IGBVF(dev);
> +
> +    if (s->migration_enabled && igbvf_addr_in_dvsec(addr, size)) {
> +        return igbvf_mig_config_read(s, addr, size);
> +    }
> +
> +    return pci_default_read_config(dev, addr, size);
> +}
> +
>  static void igbvf_write_config(PCIDevice *dev, uint32_t addr, uint32_t val,
>      int len)
>  {
> +    IgbVfState *s = IGBVF(dev);
> +
>      trace_igbvf_write_config(addr, val, len);
> +
> +    if (s->migration_enabled && igbvf_addr_in_dvsec(addr, len)) {
> +        igbvf_mig_config_write(s, addr, val, len);
> +        return;
> +    }
> +
>      pci_default_write_config(dev, addr, val, len);
>      if (object_property_get_bool(OBJECT(pcie_sriov_get_pf(dev)),
>                                   "x-pcie-flr-init", &error_abort)) {
> @@ -282,13 +301,27 @@ static void igbvf_pci_realize(PCIDevice *dev, Error **errp)
>      }
>  
>      pcie_ari_init(dev, 0x150);
> +
> +    if (object_property_get_bool(OBJECT(pcie_sriov_get_pf(dev)),
> +                                 "x-vf-migration", &error_abort)) {
> +        s->vfn = pcie_sriov_vf_number(dev);
> +        s->migration_enabled = true;
> +        if (!igbvf_add_migration_dvsec(dev, errp)) {
> +            return;
> +        }

The migration blocker noted later in the docs should be added here.
Trivial to add, avoids the internal migration state being reset by the
L0 VM being migrated, doesn't seem worth extending this driver's VMState
while the feature is experimental.

> +    }
>  }
>  
>  static void igbvf_qdev_reset_hold(Object *obj, ResetType type)
>  {
>      PCIDevice *vf = PCI_DEVICE(obj);
> +    IgbVfState *s = IGBVF(vf);
>  
>      igb_vf_reset(pcie_sriov_get_pf(vf), pcie_sriov_vf_number(vf));
> +
> +    if (s->migration_enabled) {
> +        igbvf_mig_state_reset(s);
> +    }

Hmm, I think reset is more complicated that this.  This seems to define
that any device reset will reset the migration state.  The vfio
migration protocol only defines that a VFIO_DEVICE_RESET returns the
device to running.

Consider the case of an FLR triggered by the L2 guest.  L1 QEMU passes
through the config space write, that lands in vfio-pci core config
space handling in the L1 variant driver, which turns into a
pci_reset_function() in the L1 kernel and I think lands here in the L0
QEMU.  Therefore, it seems like the L2 guest can corrupt the migration
state.

As above, the migration state is defined to be reset via the RESET
ioctl, which also turns into a pci_reset_function() in the L1 kernel.
So L0 QEMU can't tell the difference here.

I think that means that the variant driver itself needs to own this
part of the protocol, performing the migration state housekeeping on
RESET ioctl, while both allow the state to persist on other resets.  In
that sense QEMU cannot emulate this DVSEC as normal config space, it's
a persistent control plane that lives in the VMM and happens to be
accessed through config space.  Thanks,

Alex

>  }
>  
>  static void igbvf_pci_uninit(PCIDevice *dev)
> @@ -309,6 +342,7 @@ static void igbvf_class_init(ObjectClass *class, const void *data)
>  
>      c->realize = igbvf_pci_realize;
>      c->exit = igbvf_pci_uninit;
> +    c->config_read = igbvf_read_config;
>      c->vendor_id = PCI_VENDOR_ID_INTEL;
>      c->device_id = E1000_DEV_ID_82576_VF;
>      c->revision = 1;
> diff --git a/hw/net/meson.build b/hw/net/meson.build
> index 84f142df222a..bb4b449b25ba 100644
> --- a/hw/net/meson.build
> +++ b/hw/net/meson.build
> @@ -11,7 +11,7 @@ system_ss.add(when: 'CONFIG_E1000_PCI', if_true: files('e1000.c', 'e1000x_common
>  system_ss.add(when: 'CONFIG_E1000E_PCI_EXPRESS', if_true: files('net_tx_pkt.c', 'net_rx_pkt.c'))
>  system_ss.add(when: 'CONFIG_E1000E_PCI_EXPRESS', if_true: files('e1000e.c', 'e1000e_core.c', 'e1000x_common.c'))
>  system_ss.add(when: 'CONFIG_IGB_PCI_EXPRESS', if_true: files('net_tx_pkt.c', 'net_rx_pkt.c'))
> -system_ss.add(when: 'CONFIG_IGB_PCI_EXPRESS', if_true: files('igb.c', 'igbvf.c', 'igb_core.c'))
> +system_ss.add(when: 'CONFIG_IGB_PCI_EXPRESS', if_true: files('igb.c', 'igbvf.c', 'igb_core.c', 'igb_migration.c'))
>  system_ss.add(when: 'CONFIG_RTL8139_PCI', if_true: files('rtl8139.c'))
>  system_ss.add(when: 'CONFIG_TULIP', if_true: files('tulip.c'))
>  system_ss.add(when: 'CONFIG_VMXNET3_PCI', if_true: files('net_tx_pkt.c', 'net_rx_pkt.c'))



^ permalink raw reply	[flat|nested] 16+ messages in thread

* Re: [RFC PATCH v2 1/9] igb: Add x-vf-migration property and DVSEC extended capability
  2026-09-03 19:57   ` Alex Williamson
@ 2026-09-07 20:58     ` Cédric Le Goater
  0 siblings, 0 replies; 16+ messages in thread
From: Cédric Le Goater @ 2026-09-07 20:58 UTC (permalink / raw)
  To: Alex Williamson
  Cc: qemu-devel, Akihiko Odaki, Sriram Yagnaraman, Jason Wang,
	Peter Xu

On 9/3/26 21:57, Alex Williamson wrote:
> On Wed,  2 Sep 2026 21:20:46 +0200
> Cédric Le Goater <clg@redhat.com> wrote:
> 
>> diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
>> new file mode 100644
>> index 000000000000..4dfebd82344c
>> --- /dev/null
>> +++ b/hw/net/igb_migration.c
>> @@ -0,0 +1,106 @@
>> +/*
>> + * QEMU Intel 82576 SR/IOV VF Migration Support
>> + *
>> + * Copyright (c) 2026 Red Hat, Inc.
>> + *
>> + * SPDX-License-Identifier: GPL-2.0-or-later
>> + */
>> +
>> +#include "qemu/osdep.h"
>> +#include "hw/pci/pci_device.h"
>> +#include "hw/pci/pcie.h"
>> +#include "igb_common.h"
>> +#include "igb_migration.h"
>> +
>> +static void igbvf_mig_update_status(IgbVfState *s, uint8_t err)
>> +{
>> +    IgbVfMigState *ms = &s->mig;
>> +    PCIDevice *dev = PCI_DEVICE(s);
>> +    uint32_t status;
>> +
>> +    status = ms->mig_state & IGB_MIG_STATUS_STATE_MASK;
>> +
>> +    if (err) {
>> +        status = IGB_MIG_STATE_ERROR | IGB_MIG_STATUS_ERR(err);
>> +    }
>> +
>> +    pci_set_long(dev->config + IGB_MIG_DVSEC_OFFSET + IGB_MIG_STATUS, status);
>> +}
>> +
>> +
>> +bool igbvf_add_migration_dvsec(PCIDevice *dev, Error **errp)
>> +{
>> +    uint16_t offset = IGB_MIG_DVSEC_OFFSET;
>> +    uint32_t caps;
>> +
>> +    pcie_add_capability(dev, PCI_EXT_CAP_ID_DVSEC, 1, offset,
>> +                        IGB_MIG_DVSEC_SIZE);
>> +
>> +    /* DVSEC header 1: length[31:20] | rev[19:16] | vendor_id[15:0] */
>> +    pci_set_long(dev->config + offset + 0x4,
>> +                 (IGB_MIG_DVSEC_SIZE << 20) |
>> +                 (IGB_MIG_DVSEC_VER << 16) |
>> +                 PCI_VENDOR_ID_INTEL);
> 
> Let's not co-opt an Intel DVSEC ID, we should use a vendor ID that we
> have a claim to, the RedHat/Qumranet one, I'd guess.  We may need to
> add DVSEC IDs to the spreadsheet for whoever is tracking Device IDs.
> Gerd?
> 
>> +
>> +    /* DVSEC header 2: DVSEC ID */
>> +    pci_set_word(dev->config + offset + 0x8, IGB_MIG_DVSEC_ID);
>> +
>> +    /* CAPS: features (state migration only) */
>> +    caps = IGB_MIG_CAP_F_STATE;
>> +    pci_set_long(dev->config + offset + IGB_MIG_CAPS, caps);
>> +
>> +    /* STATUS: initial state is RUNNING */
>> +    pci_set_long(dev->config + offset + IGB_MIG_STATUS,
>> +                 IGB_MIG_STATE_RUNNING);
>> +
>> +    /* BUF_ADDR_LO and BUF_ADDR_HI are writable */
>> +    memset(dev->wmask + offset + IGB_MIG_BUF_ADDR_LO, 0xff, 4);
>> +    memset(dev->wmask + offset + IGB_MIG_BUF_ADDR_HI, 0xff, 4);
>> +
>> +    return true;
>> +}
> ...
>> diff --git a/hw/net/igbvf.c b/hw/net/igbvf.c
>> index 9a165c7063ee..30dfdb574ac7 100644
>> --- a/hw/net/igbvf.c
>> +++ b/hw/net/igbvf.c
>> @@ -38,27 +38,21 @@
>>    */
>>   
>>   #include "qemu/osdep.h"
>> +#include "qemu/range.h"
>>   #include "hw/core/hw-error.h"
>>   #include "hw/net/mii.h"
>>   #include "hw/pci/pci_device.h"
>>   #include "hw/pci/pcie.h"
>> +#include "hw/pci/pcie_sriov.h"
>>   #include "hw/pci/msix.h"
>>   #include "net/eth.h"
>>   #include "net/net.h"
>>   #include "igb_common.h"
>>   #include "igb_core.h"
>> +#include "igb_migration.h"
>>   #include "trace.h"
>>   #include "qapi/error.h"
>>   
>> -OBJECT_DECLARE_SIMPLE_TYPE(IgbVfState, IGBVF)
>> -
>> -struct IgbVfState {
>> -    PCIDevice parent_obj;
>> -
>> -    MemoryRegion mmio;
>> -    MemoryRegion msix;
>> -};
>> -
>>   static hwaddr vf_to_pf_addr(hwaddr addr, uint16_t vfn, bool write)
>>   {
>>       switch (addr) {
>> @@ -199,10 +193,35 @@ static hwaddr vf_to_pf_addr(hwaddr addr, uint16_t vfn, bool write)
>>       return HWADDR_MAX;
>>   }
>>   
>> +static bool igbvf_addr_in_dvsec(uint32_t addr, int len)
>> +{
>> +    return ranges_overlap(addr, len,
>> +                          IGB_MIG_DVSEC_OFFSET, IGB_MIG_DVSEC_SIZE);
>> +}
>> +
>> +static uint32_t igbvf_read_config(PCIDevice *dev, uint32_t addr, int size)
>> +{
>> +    IgbVfState *s = IGBVF(dev);
>> +
>> +    if (s->migration_enabled && igbvf_addr_in_dvsec(addr, size)) {
>> +        return igbvf_mig_config_read(s, addr, size);
>> +    }
>> +
>> +    return pci_default_read_config(dev, addr, size);
>> +}
>> +
>>   static void igbvf_write_config(PCIDevice *dev, uint32_t addr, uint32_t val,
>>       int len)
>>   {
>> +    IgbVfState *s = IGBVF(dev);
>> +
>>       trace_igbvf_write_config(addr, val, len);
>> +
>> +    if (s->migration_enabled && igbvf_addr_in_dvsec(addr, len)) {
>> +        igbvf_mig_config_write(s, addr, val, len);
>> +        return;
>> +    }
>> +
>>       pci_default_write_config(dev, addr, val, len);
>>       if (object_property_get_bool(OBJECT(pcie_sriov_get_pf(dev)),
>>                                    "x-pcie-flr-init", &error_abort)) {
>> @@ -282,13 +301,27 @@ static void igbvf_pci_realize(PCIDevice *dev, Error **errp)
>>       }
>>   
>>       pcie_ari_init(dev, 0x150);
>> +
>> +    if (object_property_get_bool(OBJECT(pcie_sriov_get_pf(dev)),
>> +                                 "x-vf-migration", &error_abort)) {
>> +        s->vfn = pcie_sriov_vf_number(dev);
>> +        s->migration_enabled = true;
>> +        if (!igbvf_add_migration_dvsec(dev, errp)) {
>> +            return;
>> +        }
> 
> The migration blocker noted later in the docs should be added here.
> Trivial to add, avoids the internal migration state being reset by the
> L0 VM being migrated, doesn't seem worth extending this driver's VMState
> while the feature is experimental.
> 
>> +    }
>>   }
>>   
>>   static void igbvf_qdev_reset_hold(Object *obj, ResetType type)
>>   {
>>       PCIDevice *vf = PCI_DEVICE(obj);
>> +    IgbVfState *s = IGBVF(vf);
>>   
>>       igb_vf_reset(pcie_sriov_get_pf(vf), pcie_sriov_vf_number(vf));
>> +
>> +    if (s->migration_enabled) {
>> +        igbvf_mig_state_reset(s);
>> +    }
> 
> Hmm, I think reset is more complicated that this.  This seems to define
> that any device reset will reset the migration state.  The vfio
> migration protocol only defines that a VFIO_DEVICE_RESET returns the
> device to running.
> 
> Consider the case of an FLR triggered by the L2 guest.  L1 QEMU passes
> through the config space write, that lands in vfio-pci core config
> space handling in the L1 variant driver, which turns into a
> pci_reset_function() in the L1 kernel and I think lands here in the L0
> QEMU.  Therefore, it seems like the L2 guest can corrupt the migration
> state.
> 
> As above, the migration state is defined to be reset via the RESET
> ioctl, which also turns into a pci_reset_function() in the L1 kernel.
> So L0 QEMU can't tell the difference here.
> 
> I think that means that the variant driver itself needs to own this
> part of the protocol, performing the migration state housekeeping on
> RESET ioctl, while both allow the state to persist on other resets.  In
> that sense QEMU cannot emulate this DVSEC as normal config space, it's
> a persistent control plane that lives in the VMM and happens to be
> accessed through config space.  Thanks,


Hi Alex,
  
You are right. The key principle: The DVSEC is a persistent control
plane living in the VMM, not device state. Device reset (VFIO ioctl
and L2 FLR) resets the device, not the mailbox. The variant driver
owns the reset, not the emulated device in L0 QEMU.

I missed two things : 1. transferring the DVSEC hiding (as it was done
for the mig BAR) and 2. reset of migration state.

Looking closer at it, the changes are small.

In QEMU:

The cold boot initialization sets all fields to defaults (zero). The
only change from today is that the initial state is ERROR, which means
migration isn't operational until the driver binds. The variant driver
activates the control plane by cycling the machine state to RUNNING.

The VF functional registers (queues, DMA rings, interrupts ... ) are
reset normally by igb_core_vf_reset(), and the migration control plane
(IgbVfMigState) persists in memory. It survives pci_do_device_reset()
and the variant driver sets the fields as needed through the normal
DVSEC command flow.

BUF_ADDR could be restored from the IgbVfMigState cache values since
it's a RW reg, but even that isn't a strong requirement.

So igbvf_mig_state_reset() is no longer needed in the reset path.
That's all for QEMU.

Driver :

Variant driver owns all the cleanup after reset: closes migration fds,
disables dirty tracking, reads and cycles the DVSEC state back to
RUNNING.

The first path, VFIO_DEVICE_RESET, needs a custom ioctl handler to run
the cleanup.

Second path, L2 FLR. To hide the DVSEC range from the L2 guest, the
driver implements custom VFIO config space reads and writes. The
config write handler can detect the FLR and implement the same cleanup
logic.

That's for v3.

Thanks,

C.




^ permalink raw reply	[flat|nested] 16+ messages in thread

* Re: [RFC PATCH v2 3/9] igb: Add VF state serialization for live migration
  2026-09-02 19:20 ` [RFC PATCH v2 3/9] igb: Add VF state serialization for live migration Cédric Le Goater
@ 2026-09-08  8:01   ` Akihiko Odaki
  0 siblings, 0 replies; 16+ messages in thread
From: Akihiko Odaki @ 2026-09-08  8:01 UTC (permalink / raw)
  To: Cédric Le Goater, qemu-devel
  Cc: Sriram Yagnaraman, Jason Wang, Alex Williamson, Peter Xu



On 2026/09/03 4:20, Cédric Le Goater wrote:
> Implement per-VF state serialization and deserialization for the SAVE
> and LOAD commands. The wire format consists of a header (magic,
> version, VF number, register count), per-VF register offset/value
> pairs from a whitelist, RA table entries owned by the VF, and TX
> queue contexts.
> 
> Register offsets are relocated on load so a VF can migrate to a
> different VF number on the destination. RA pool ownership bits are
> swapped accordingly.
> 
> AI-used-for: analysis, code (prototype)
> Signed-off-by: Cédric Le Goater <clg@redhat.com>
> ---
>   hw/net/igb_core.h      |   2 +
>   hw/net/igb_migration.h |   2 +
>   hw/net/igb.c           |   5 +
>   hw/net/igb_migration.c | 347 ++++++++++++++++++++++++++++++++++++++++-
>   4 files changed, 354 insertions(+), 2 deletions(-)
> 
> diff --git a/hw/net/igb_core.h b/hw/net/igb_core.h
> index d70b54e318f1..60724e2824ab 100644
> --- a/hw/net/igb_core.h
> +++ b/hw/net/igb_core.h
> @@ -143,4 +143,6 @@ igb_receive_iov(IGBCore *core, const struct iovec *iov, int iovcnt);
>   void
>   igb_start_recv(IGBCore *core);
>   
> +IGBCore *igb_pf_get_core(void *pf);
> +
>   #endif
> diff --git a/hw/net/igb_migration.h b/hw/net/igb_migration.h
> index ea40ac65c54b..b2f601e74346 100644
> --- a/hw/net/igb_migration.h
> +++ b/hw/net/igb_migration.h
> @@ -73,6 +73,8 @@
>   #define IGB_MIG_ERR_NO_BUFFER           3
>   #define IGB_MIG_ERR_DMA_FAILED          4
>   #define IGB_MIG_ERR_BAD_SIZE            5
> +#define IGB_MIG_ERR_BAD_MAGIC           6
> +#define IGB_MIG_ERR_BAD_VERSION         7
>   
>   /* Shared buffer constants */
>   #define IGB_VF_STATE_MAX_SIZE           4096
> diff --git a/hw/net/igb.c b/hw/net/igb.c
> index 7268e5473fc3..f39f2bc3a04e 100644
> --- a/hw/net/igb.c
> +++ b/hw/net/igb.c
> @@ -133,6 +133,11 @@ void igb_vf_reset(void *opaque, uint16_t vfn)
>       igb_core_vf_reset(&s->core, vfn);
>   }
>   
> +IGBCore *igb_pf_get_core(void *pf)
> +{
> +    return &IGB(pf)->core;
> +}
> +
>   static bool
>   igb_io_get_reg_index(IGBState *s, uint32_t *idx)
>   {
> diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
> index 8e7e6fac9b9b..c34035974620 100644
> --- a/hw/net/igb_migration.c
> +++ b/hw/net/igb_migration.c
> @@ -10,18 +10,241 @@
>   #include "qemu/log.h"
>   #include "hw/pci/pci_device.h"
>   #include "hw/pci/pcie.h"
> +#include "net/eth.h"
> +#include "net/net.h"
>   #include "igb_common.h"
> +#include "igb_core.h"
>   #include "igb_migration.h"
>   #include "system/address-spaces.h"
>   #include "trace.h"
>   
> +static IGBCore *igbvf_get_core(IgbVfState *s)
> +{
> +    return igb_pf_get_core(pcie_sriov_get_pf(PCI_DEVICE(s)));
> +}
> +
>   /*
>    * Per-VF state serialization / deserialization
>    */
>   
> +#define IGB_MIG_BLOB_MAGIC        0x4D494742  /* "MIGB" */
> +#define IGB_MIG_BLOB_VERSION      1
> +
> +typedef struct IgbMigRegPair {
> +    uint32_t offset;
> +    uint32_t value;
> +} IgbMigRegPair;
> +
> +typedef struct IgbMigTxCtx {
> +    uint32_t ctx_desc[8];         /* 2 × adv_tx_context_desc (4 dwords each) */
> +    uint32_t first_cmd_type_len;
> +    uint32_t first_olinfo_status;
> +    uint32_t first;
> +    uint32_t skip_cp;
> +} IgbMigTxCtx;
> +
> +#define IGB_VF_MAX_FIXED_REGS     64
> +#define IGB_VF_MAX_RA_REGS        48  /* (16 + 8) RA entries × 2 (RAL+RAH) */
> +
> +typedef struct IgbMigBlob {
> +    uint32_t magic;
> +    uint32_t version;
> +    uint32_t vfn;
> +    uint32_t num_regs;
> +    IgbMigRegPair regs[IGB_VF_MAX_FIXED_REGS];
> +    uint32_t num_ra;
> +    IgbMigRegPair ra[IGB_VF_MAX_RA_REGS];
> +    uint32_t num_tx_ctx;
> +    IgbMigTxCtx tx_ctx[2];
> +} IgbMigBlob;
> +
> +#define IGB_MIG_BLOB_SIZE            sizeof(IgbMigBlob)
> +
> +QEMU_BUILD_BUG_ON(IGB_MIG_BLOB_SIZE > IGB_VF_STATE_MAX_SIZE);
> +
> +/* Register offsets that constitute a VF's state slice */
> +static int igb_vf_reg_list(uint16_t vfn, uint32_t *offsets)
> +{
> +    int n = 0;
> +    int q0 = vfn;
> +    int q1 = vfn + IGB_NUM_VM_POOLS;
> +
> +    /* Per-VF control and interrupt registers */
> +    offsets[n++] = E1000_PVTCTRL(vfn) >> 2;
> +    offsets[n++] = E1000_PVTEICS(vfn) >> 2;
> +    offsets[n++] = E1000_PVTEIMS(vfn) >> 2;
> +    offsets[n++] = E1000_PVTEIMC(vfn) >> 2;
> +    offsets[n++] = E1000_PVTEIAC(vfn) >> 2;
> +    offsets[n++] = E1000_PVTEIAM(vfn) >> 2;
> +    offsets[n++] = E1000_PVTEICR(vfn) >> 2;
> +
> +    /* Per-VF statistics */
> +    offsets[n++] = E1000_PVFGPRC(vfn) >> 2;
> +    offsets[n++] = E1000_PVFGPTC(vfn) >> 2;
> +    offsets[n++] = E1000_PVFGORC(vfn) >> 2;
> +    offsets[n++] = E1000_PVFGOTC(vfn) >> 2;
> +    offsets[n++] = E1000_PVFMPRC(vfn) >> 2;
> +    offsets[n++] = E1000_PVFGPRLBC(vfn) >> 2;
> +    offsets[n++] = E1000_PVFGPTLBC(vfn) >> 2;
> +    offsets[n++] = E1000_PVFGORLBC(vfn) >> 2;
> +    offsets[n++] = E1000_PVFGOTLBC(vfn) >> 2;
> +
> +    /*
> +     * Mailbox control registers only - the 16-dword payload buffer
> +     * (VMBMEM) is transient and drained on quiesce.
> +     */
> +    offsets[n++] = E1000_V2PMAILBOX(vfn) >> 2;
> +    offsets[n++] = E1000_P2VMAILBOX(vfn) >> 2;
> +
> +    /* Per-VF config */
> +    offsets[n++] = E1000_VMOLR(vfn) >> 2;
> +    offsets[n++] = E1000_VMVIR(vfn) >> 2;
> +    offsets[n++] = E1000_PSRTYPE(vfn) >> 2;
> +
> +    /*
> +     * VF receive addresses (RA/RA2) are saved dynamically in
> +     * igb_core_vf_save_state by scanning for entries whose pool
> +     * bits match this VF - the PF driver chooses the RA slot.
> +     */
> +
> +    /* Interrupt routing */
> +    offsets[n++] = (E1000_VTIVAR + vfn * 4) >> 2;
> +    offsets[n++] = (E1000_VTIVAR_MISC + vfn * 4) >> 2;
> +
> +    /*
> +     * EITR (Extended Interrupt Throttle Register) - 3 vectors per VF.
> +     * Each VF has 3 MSI-X vectors, each with its own EITR controlling
> +     * interrupt coalescing. Without saving these, interrupt
> +     * throttling resets to zero after migration which can cause
> +     * interrupt storms or latency changes. VF N uses PF EITR indices
> +     * (22 - N*3) .. (24 - N*3).
> +     */
> +    {
> +        int eitr_base = 22 - vfn * 3;
> +        offsets[n++] = E1000_EITR(eitr_base) >> 2;
> +        offsets[n++] = E1000_EITR(eitr_base + 1) >> 2;
> +        offsets[n++] = E1000_EITR(eitr_base + 2) >> 2;
> +    }
> +
> +    /* RX and TX queue registers for queues q0 and q1 */
> +#define ADD_QUEUE_REGS(q) do { \
> +    offsets[n++] = E1000_RDBAL(q) >> 2; \
> +    offsets[n++] = E1000_RDBAH(q) >> 2; \
> +    offsets[n++] = E1000_RDLEN(q) >> 2; \
> +    offsets[n++] = E1000_SRRCTL(q) >> 2; \
> +    offsets[n++] = E1000_RDH(q) >> 2; \
> +    offsets[n++] = E1000_RDT(q) >> 2; \
> +    offsets[n++] = E1000_RXDCTL(q) >> 2; \
> +    offsets[n++] = E1000_RXCTL(q) >> 2; \
> +    offsets[n++] = E1000_RQDPC(q) >> 2; \
> +    offsets[n++] = E1000_TDBAL(q) >> 2; \
> +    offsets[n++] = E1000_TDBAH(q) >> 2; \
> +    offsets[n++] = E1000_TDLEN(q) >> 2; \
> +    offsets[n++] = E1000_TDH(q) >> 2; \
> +    offsets[n++] = E1000_TDT(q) >> 2; \
> +    offsets[n++] = E1000_TXDCTL(q) >> 2; \
> +    offsets[n++] = E1000_TXCTL(q) >> 2; \
> +    offsets[n++] = E1000_TDWBAL(q) >> 2; \
> +    offsets[n++] = E1000_TDWBAH(q) >> 2; \
> +} while (0)
> +
> +    ADD_QUEUE_REGS(q0);
> +    ADD_QUEUE_REGS(q1);
> +#undef ADD_QUEUE_REGS
> +
> +    g_assert(n <= IGB_VF_MAX_FIXED_REGS);
> +    return n;
> +}
> +
> +/*
> + * Scan RA and RA2 arrays for receive address entries assigned to
> + * this VF. The PF driver picks the RA slot, so we cannot use a
> + * fixed index - instead check each entry's pool bits.
> + */
> +static int igb_core_vf_save_ra(IGBCore *core, uint16_t vfn,
> +                               IgbMigRegPair *regs)
> +{
> +    uint32_t vf_pool_bit = E1000_RAH_POOL_1 << vfn;
> +    int n = 0;
> +    static const struct {
> +        uint32_t base;
> +        int count;
> +    } ra_banks[] = {
> +        { RA,  16 },
> +        { RA2,  8 },
> +    };
> +
> +    for (int i = 0; i < ARRAY_SIZE(ra_banks); i++) {
> +        for (int j = 0; j < ra_banks[i].count; j++) {
> +            uint32_t ral_off = ra_banks[i].base + j * 2;
> +            uint32_t rah_off = ra_banks[i].base + j * 2 + 1;
> +            uint32_t rah_val = core->mac[rah_off];
> +
> +            if ((rah_val & E1000_RAH_AV) && (rah_val & vf_pool_bit)) {
> +                regs[n].offset = cpu_to_le32(ral_off);
> +                regs[n].value = cpu_to_le32(core->mac[ral_off]);
> +                n++;
> +                regs[n].offset = cpu_to_le32(rah_off);
> +                regs[n].value = cpu_to_le32(rah_val);
> +                n++;
> +            }
> +        }
> +    }
> +    return n;
> +}
> +
> +static void igb_core_vf_save_tx_ctx(IGBCore *core, int queue,
> +                                    IgbMigTxCtx *tx)
> +{
> +    struct igb_tx *src = &core->tx[queue];
> +
> +    memcpy(tx->ctx_desc, src->ctx, sizeof(tx->ctx_desc));
> +    tx->first_cmd_type_len = cpu_to_le32(src->first_cmd_type_len);
> +    tx->first_olinfo_status = cpu_to_le32(src->first_olinfo_status);
> +    tx->first = cpu_to_le32(src->first);
> +    tx->skip_cp = cpu_to_le32(src->skip_cp);
> +}
> +
>   static int igb_core_vf_save_state(IgbVfState *s, void *buf, size_t buf_size)
>   {
> -    int size = 0;
> +    int size = IGB_MIG_BLOB_SIZE;
> +    IGBCore *core = igbvf_get_core(s);
> +    IgbMigBlob *blob = buf;
> +    uint32_t offsets[IGB_VF_MAX_FIXED_REGS];
> +    int num_regs;
> +    int q0 = s->vfn;
> +    int q1 = s->vfn + IGB_NUM_VM_POOLS;
> +
> +    /*
> +     * Save PVT shadow registers (PVTEIMS/PVTEIAC/PVTEIAM) instead of
> +     * extracting from PF aggregates - the L1 PF driver may have
> +     * transiently cleared EIMS via EIMC. The load path ORs them back.
> +     */
> +    num_regs = igb_vf_reg_list(s->vfn, offsets);
> +
> +    if (!buf) {
> +        return size;
> +    }
> +
> +    if (size > buf_size) {
> +        return -IGB_MIG_ERR_BAD_SIZE;
> +    }
> +
> +    blob->magic = cpu_to_le32(IGB_MIG_BLOB_MAGIC);
> +    blob->version = cpu_to_le32(IGB_MIG_BLOB_VERSION);
> +    blob->vfn = cpu_to_le32(s->vfn);
> +
> +    blob->num_regs = cpu_to_le32(num_regs);
> +    for (int i = 0; i < num_regs; i++) {
> +        blob->regs[i].offset = cpu_to_le32(offsets[i]);
> +        blob->regs[i].value = cpu_to_le32(core->mac[offsets[i]]);
> +    }
> +
> +    blob->num_ra = cpu_to_le32(igb_core_vf_save_ra(core, s->vfn, blob->ra));

The serialized state contains RA entries but excludes VFTA/VLVF and 
MTA/UTA state. A guest-configured VLAN therefore resumes on a 
destination lacking its filter membership: packets are rejected by the 
global VLAN check or have their VF queue removed  in 
igb_receive_assign(),. Likewise, restored VMOLR.ROMPE cannot receive 
subscribed multicast without its MTA hash bits. These are 
guest-requested settings communicated through the PF mailbox, and the 
resumed guest does not recreate them automatically.

> +
> +    blob->num_tx_ctx = cpu_to_le32(2);
> +    igb_core_vf_save_tx_ctx(core, q0, &blob->tx_ctx[0]);
> +    igb_core_vf_save_tx_ctx(core, q1, &blob->tx_ctx[1]);
>   
>       trace_igbvf_mig_save_state(s->vfn, size);
>       return size;
> @@ -29,11 +252,131 @@ static int igb_core_vf_save_state(IgbVfState *s, void *buf, size_t buf_size)
>   
>   static int igb_core_vf_max_data_size(IgbVfState *s)
>   {
> -    return sizeof(s->mig.mig_data);
> +    int size = igb_core_vf_save_state(s, NULL, 0);
> +
> +    g_assert(size > 0 && size <= IGB_VF_STATE_MAX_SIZE);
> +    return size;
> +}
> +
> +static void igb_core_vf_load_tx_ctx(IGBCore *core, int queue,
> +                                    const IgbMigTxCtx *tx)
> +{
> +    struct igb_tx *dst = &core->tx[queue];
> +
> +    /*
> +     * Preserve the destination's tx_pkt - it's a host-side object,
> +     * not guest state
> +     */
> +    memcpy(dst->ctx, tx->ctx_desc, sizeof(dst->ctx));
> +    dst->first_cmd_type_len = le32_to_cpu(tx->first_cmd_type_len);
> +    dst->first_olinfo_status = le32_to_cpu(tx->first_olinfo_status);
> +    dst->first = le32_to_cpu(tx->first);
> +    dst->skip_cp = le32_to_cpu(tx->skip_cp);
> +}
> +
> +static uint32_t igb_vf_relocate_offset(uint32_t offset,
> +                                       const uint32_t *src_offsets,
> +                                       const uint32_t *dst_offsets,
> +                                       int num_offsets)
> +{
> +    for (int i = 0; i < num_offsets; i++) {
> +        if (src_offsets[i] == offset) {
> +            return dst_offsets[i];
> +        }
> +    }
> +    return 0;
>   }
>   
>   static int igb_core_vf_load_state(IgbVfState *s, const void *buf, size_t size)
>   {
> +    IGBCore *core = igbvf_get_core(s);
> +    uint32_t src_offsets[IGB_VF_MAX_FIXED_REGS];
> +    uint32_t dst_offsets[IGB_VF_MAX_FIXED_REGS];
> +    int q0 = s->vfn;
> +    int q1 = s->vfn + IGB_NUM_VM_POOLS;
> +
> +    if (size < IGB_MIG_BLOB_SIZE) {
> +        return -IGB_MIG_ERR_BAD_SIZE;
> +    }
> +
> +    const IgbMigBlob *blob = buf;
> +
> +    uint32_t magic = le32_to_cpu(blob->magic);
> +    uint32_t version = le32_to_cpu(blob->version);
> +    uint32_t saved_vfn = le32_to_cpu(blob->vfn);
> +    uint32_t num_regs = le32_to_cpu(blob->num_regs);
> +
> +    if (magic != IGB_MIG_BLOB_MAGIC) {
> +        return -IGB_MIG_ERR_BAD_MAGIC;
> +    }
> +    if (version != IGB_MIG_BLOB_VERSION) {
> +        return -IGB_MIG_ERR_BAD_VERSION;
> +    }
> +    if (num_regs > IGB_VF_MAX_FIXED_REGS) {
> +        return -IGB_MIG_ERR_BAD_SIZE;
> +    }
> +
> +    uint32_t num_ra = le32_to_cpu(blob->num_ra);
> +    if (num_ra > IGB_VF_MAX_RA_REGS) {
> +        return -IGB_MIG_ERR_BAD_SIZE;
> +    }
> +
> +    int num_offsets = igb_vf_reg_list(saved_vfn, src_offsets);
> +    igb_vf_reg_list(s->vfn, dst_offsets);
> +
> +    for (uint32_t i = 0; i < num_regs; i++) {
> +        uint32_t src_off = le32_to_cpu(blob->regs[i].offset);
> +        uint32_t value = le32_to_cpu(blob->regs[i].value);
> +        uint32_t offset = igb_vf_relocate_offset(src_off,
> +            src_offsets, dst_offsets, num_offsets);
> +        if (!offset) {
> +            return -IGB_MIG_ERR_BAD_SIZE;
> +        }
> +
> +        core->mac[offset] = value;
> +
> +        /*
> +         * Sync EITR to eitr_guest_value[] shadow array, stripping
> +         * E1000_EITR_CNT_IGNR so guest register readback returns the
> +         * correct value.
> +         */
> +        if (offset >= EITR0 && offset < EITR0 + IGB_INTR_NUM) {
> +            core->eitr_guest_value[offset - EITR0] =
> +                value & ~E1000_EITR_CNT_IGNR;
> +        }
> +    }
> +
> +    /*
> +     * MSI-X table/PBA is not saved - L1's VFIO reprograms it with
> +     * destination-specific IRTE references after migration.
> +     */
> +
> +    uint32_t src_pool = E1000_RAH_POOL_1 << saved_vfn;
> +    uint32_t dst_pool = E1000_RAH_POOL_1 << s->vfn;
> +
> +    for (uint32_t i = 0; i < num_ra; i++) {
> +        uint32_t offset = le32_to_cpu(blob->ra[i].offset);
> +        uint32_t value = le32_to_cpu(blob->ra[i].value);
> +
> +        /* RAH entries: swap pool ownership bits */
> +        if (offset >= RA && offset < RA + 32 && (offset - RA) % 2 == 1) {
> +            value = (value & ~src_pool) | dst_pool;
> +        }
> +        if (offset >= RA2 && offset < RA2 + 16 && (offset - RA2) % 2 == 1) {
> +            value = (value & ~src_pool) | dst_pool;
> +        }
> +
> +        core->mac[offset] = value;

Please validate RA offsets before writing core->mac. A valid-sized LOAD 
blob with one RA entry and an out-of-range offset produces an 
out-of-bounds 32-bit host write. LOAD reads this blob directly from 
guest RAM. Please validate all offsets and saved_vfn before modifying state.

This loop also changes the VF pool bit but preserves the source RA slot. 
A valid VF0→VF1 migration onto a destination already using VF0 
overwrites destination VF0’s primary MAC entry. Linux assigns the 
primary entry as rar_entry_count - (vf + 1); receive filtering consumes 
these slots directly in igb_receive_assign(). Thus the unrelated 
destination VF loses unicast reception. Existing destination entries 
owned by the migrating VF are also left behind.

Regards,
Akihiko Odaki

> +    }
> +
> +    uint32_t num_tx = le32_to_cpu(blob->num_tx_ctx);
> +    if (num_tx != 2) {
> +        return -IGB_MIG_ERR_BAD_SIZE;
> +    }
> +
> +    igb_core_vf_load_tx_ctx(core, q0, &blob->tx_ctx[0]);
> +    igb_core_vf_load_tx_ctx(core, q1, &blob->tx_ctx[1]);
> +
>       trace_igbvf_mig_load_state(s->vfn, (uint32_t)size);
>       return 0;
>   }



^ permalink raw reply	[flat|nested] 16+ messages in thread

* Re: [RFC PATCH v2 4/9] igb: Add VF post-load fixups for live migration
  2026-09-02 19:20 ` [RFC PATCH v2 4/9] igb: Add VF post-load fixups " Cédric Le Goater
@ 2026-09-08  8:10   ` Akihiko Odaki
  0 siblings, 0 replies; 16+ messages in thread
From: Akihiko Odaki @ 2026-09-08  8:10 UTC (permalink / raw)
  To: Cédric Le Goater, qemu-devel
  Cc: Sriram Yagnaraman, Jason Wang, Alex Williamson, Peter Xu

On 2026/09/03 4:20, Cédric Le Goater wrote:
> After restoring per-VF register state, propagate the VF's PVT shadow
> values back into the PF's aggregate EIMS/EIAC/EIAM registers and
> re-apply the VTIVAR interrupt vector routing to the shared IVAR0.
> 
> AI-used-for: analysis, code (prototype)
> Signed-off-by: Cédric Le Goater <clg@redhat.com>
> ---
>   hw/net/igb_core.h      |  3 ++
>   hw/net/igb_core.c      | 66 ++++++++++++++++++++++++++++++++++++++++++
>   hw/net/igb_migration.c |  7 +++++
>   3 files changed, 76 insertions(+)
> 
> diff --git a/hw/net/igb_core.h b/hw/net/igb_core.h
> index 60724e2824ab..22e10e4e6d0b 100644
> --- a/hw/net/igb_core.h
> +++ b/hw/net/igb_core.h
> @@ -145,4 +145,7 @@ igb_start_recv(IGBCore *core);
>   
>   IGBCore *igb_pf_get_core(void *pf);
>   
> +void igb_core_vf_propagate_irqs(IGBCore *core, uint16_t vfn);
> +void igb_core_vf_propagate_ivar(IGBCore *core, uint16_t vfn);
> +
>   #endif
> diff --git a/hw/net/igb_core.c b/hw/net/igb_core.c
> index 2a4883907353..01745fe756d0 100644
> --- a/hw/net/igb_core.c
> +++ b/hw/net/igb_core.c
> @@ -4552,3 +4552,69 @@ igb_core_post_load(IGBCore *core)
>   
>       return 0;
>   }
> +
> +/*
> + * Propagate VF interrupt state to PF aggregates after loading VF
> + * registers. The load path writes directly to mac[] bypassing the
> + * register handlers that OR VF bits into EIMS/EIAC/EIAM. Also clear
> + * stale VF bits in EICR that may have been set by packets arriving
> + * between PF vmstate restore and VF state load.
> + */
> +void igb_core_vf_propagate_irqs(IGBCore *core, uint16_t vfn)
> +{
> +    uint32_t shift = 22 - vfn * IGBVF_MSIX_VEC_NUM;
> +    uint32_t vf_mask = 0x7 << shift;
> +    uint32_t pvt_idx;
> +
> +    core->mac[EIMS] &= ~vf_mask;

Restore the effective interrupt mask rather than the last EIMS write.

This clears the destination VF mask and copies PVTEIMS, but that shadow 
records only the last write: igb_set_vteims() assigns it before applying 
set-bit semantics to aggregate EIMS. For example, source writes EIMS=1, 
then EIMS=2, leave effective mask 3 and shadow 2; migration restores 2, 
suppressing vector 0 until explicitly enabled again. Conversely, EIMC 
updates aggregate EIMS without updating PVTEIMS, so migration reenables 
explicitly masked vectors. Patch 7 repeats this fixup on unquiesce 
without repairing the bookkeeping.

Regards,
Akihiko Odaki

> +    pvt_idx = PVTEIMS0 + vfn * 0x40;
> +    core->mac[EIMS] |= (core->mac[pvt_idx] & 0x7) << shift;
> +
> +    core->mac[EIAC] &= ~vf_mask;
> +    pvt_idx = PVTEIAC0 + vfn * 0x40;
> +    core->mac[EIAC] |= (core->mac[pvt_idx] & 0x7) << shift;
> +
> +    core->mac[EIAM] &= ~vf_mask;
> +    pvt_idx = PVTEIAM0 + vfn * 0x40;
> +    core->mac[EIAM] |= (core->mac[pvt_idx] & 0x7) << shift;
> +
> +    core->mac[EICR] &= ~vf_mask;
> +}
> +
> +/*
> + * Re-apply VTIVAR -> IVAR0 interrupt routing. The L1 PF driver
> + * may have overwritten the shared IVAR0 entries with its own
> + * queue routing after L0 vmstate restore.
> + */
> +void igb_core_vf_propagate_ivar(IGBCore *core, uint16_t vfn)
> +{
> +    uint32_t vtivar = core->mac[VTIVAR + vfn];
> +    int n;
> +    uint8_t ent;
> +    uint32_t mask;
> +
> +    n = igb_ivar_entry_rx(vfn);
> +    mask = 0xffU << (8 * (n % 4));
> +    if (vtivar & E1000_IVAR_VALID) {
> +        ent = E1000_IVAR_VALID |
> +              (24 - vfn * IGBVF_MSIX_VEC_NUM - (2 - (vtivar & 0x7)));
> +        core->mac[IVAR0 + n / 4] =
> +            (core->mac[IVAR0 + n / 4] & ~mask) |
> +            ((uint32_t)ent << (8 * (n % 4)));
> +    } else {
> +        core->mac[IVAR0 + n / 4] &= ~mask;
> +    }
> +
> +    n = igb_ivar_entry_tx(vfn);
> +    mask = 0xffU << (8 * (n % 4));
> +    ent = vtivar >> 8;
> +    if (ent & E1000_IVAR_VALID) {
> +        ent = E1000_IVAR_VALID |
> +              (24 - vfn * IGBVF_MSIX_VEC_NUM - (2 - (ent & 0x7)));
> +        core->mac[IVAR0 + n / 4] =
> +            (core->mac[IVAR0 + n / 4] & ~mask) |
> +            ((uint32_t)ent << (8 * (n % 4)));
> +    } else {
> +        core->mac[IVAR0 + n / 4] &= ~mask;
> +    }
> +}
> diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
> index c34035974620..af7257ccc4fd 100644
> --- a/hw/net/igb_migration.c
> +++ b/hw/net/igb_migration.c
> @@ -384,12 +384,19 @@ static int igb_core_vf_load_state(IgbVfState *s, const void *buf, size_t size)
>   static int igbvf_mig_load(IgbVfState *s, const void *buf, size_t size)
>   {
>       int ret;
> +    IGBCore *core = igbvf_get_core(s);
>   
>       ret = igb_core_vf_load_state(s, buf, size);
>       if (ret < 0) {
>           return ret;
>       }
>   
> +    /*
> +     * Post-load: sync VF interrupt and routing state to PF aggregates
> +     */
> +    igb_core_vf_propagate_irqs(core, s->vfn);
> +    igb_core_vf_propagate_ivar(core, s->vfn);
> +
>       return 0;
>   }
>   



^ permalink raw reply	[flat|nested] 16+ messages in thread

* Re: [RFC PATCH v2 9/9] docs: Add igb VF migration testing setup guide
  2026-09-02 19:20 ` [RFC PATCH v2 9/9] docs: Add igb VF migration testing setup guide Cédric Le Goater
@ 2026-09-08  8:39   ` Akihiko Odaki
  0 siblings, 0 replies; 16+ messages in thread
From: Akihiko Odaki @ 2026-09-08  8:39 UTC (permalink / raw)
  To: Cédric Le Goater, qemu-devel
  Cc: Sriram Yagnaraman, Jason Wang, Alex Williamson, Peter Xu

On 2026/09/03 4:20, Cédric Le Goater wrote:
> Document the igb VF migration interface: DVSEC register layout, dirty
> page tracking, testing setup with nested virtualization.
> 
> AI-used-for: docs
> Signed-off-by: Cédric Le Goater <clg@redhat.com>
> ---
>   MAINTAINERS                           |   1 +
>   docs/system/device-emulation.rst      |   1 +
>   docs/system/devices/igb-migration.rst | 417 ++++++++++++++++++++++++++
>   docs/system/devices/igb.rst           |   6 +
>   4 files changed, 425 insertions(+)
>   create mode 100644 docs/system/devices/igb-migration.rst
> 
> diff --git a/MAINTAINERS b/MAINTAINERS
> index f88b526be238..4a2f1357e936 100644
> --- a/MAINTAINERS
> +++ b/MAINTAINERS
> @@ -2820,6 +2820,7 @@ igb VF migration
>   M: Cédric Le Goater <clg@redhat.com>
>   S: Maintained
>   F: hw/net/igb_migration.*
> +F: docs/system/devices/igb-migration.rst
>   
>   eepro100
>   M: Stefan Weil <sw@weilnetz.de>
> diff --git a/docs/system/device-emulation.rst b/docs/system/device-emulation.rst
> index 40054bb7dfcc..75f423b795cd 100644
> --- a/docs/system/device-emulation.rst
> +++ b/docs/system/device-emulation.rst
> @@ -90,6 +90,7 @@ Emulated Devices
>      devices/cxl.rst
>      devices/emmc.rst
>      devices/igb.rst
> +   devices/igb-migration.rst
>      devices/ivshmem-flat.rst
>      devices/ivshmem.rst
>      devices/keyboard.rst
> diff --git a/docs/system/devices/igb-migration.rst b/docs/system/devices/igb-migration.rst
> new file mode 100644
> index 000000000000..d72f14f5fe63
> --- /dev/null
> +++ b/docs/system/devices/igb-migration.rst
> @@ -0,0 +1,417 @@
> +.. SPDX-License-Identifier: GPL-2.0-or-later
> +.. _igb-migration:
> +
> +igb VF Migration
> +----------------
> +
> +Live migration of VFIO-passthrough devices (SR-IOV VFs, vGPUs) is a
> +growing requirement, but real hardware with migration support is scarce
> +and hard to debug. An emulated device provides a fully controlled
> +testbed for developing and validating the entire software stack --
> +vfio-pci variant drivers, VFIO core migration v2 framework, QEMU,
> +libvirt -- and for tuning complex migration policies such as downtime
> +convergence. It also serves as an educational reference for
> +understanding VFIO migration end-to-end, from device state
> +serialization to dirty page tracking.
> +
> +The igb device supports an experimental VF migration interface that allows
> +the `igb-vfio-pci`_ variant driver to migrate VF state during live
> +migration using the standard VFIO migration v2 protocol with stop-copy
> +and pre-copy support.
> +
> +This is enabled with the ``x-vf-migration`` property::
> +
> +  -device igb,x-vf-migration=on,...
> +
> +Each emulated VF then advertises a DVSEC discovered by the
> +`igb-vfio-pci`_ variant driver at bind time. This feature is
> +experimental (``x-`` prefix, default off).
> +
> +Architecture
> +~~~~~~~~~~~~
> +
> +The target scenario is nested virtualization::
> +
> +  L0 QEMU
> +    igb PF with x-vf-migration=on
> +    └── VFs with migration DVSEC
> +
> +  L1 kernel
> +    igb-vfio-pci variant driver
> +    translates VFIO migration v2 ioctls → DVSEC config writes
> +
> +  L1 QEMU (stock, unmodified)
> +    vfio-pci device model, standard migration fd
> +
> +  L2 guest
> +    standard igbvf driver, unaware of migration
> +
> +The L1 QEMU is completely unmodified -- it sees a standard VFIO
> +migratable device and uses the normal migration fd path. The
> +`igb-vfio-pci`_ variant driver handles the translation between
> +VFIO migration v2 ioctls and DVSEC config writes.
> +
> +Design
> +~~~~~~
> +
> +The migration interface is exposed through a DVSEC at offset 0x160
> +in VF extended config space (see `DVSEC register layout`_ below for
> +the full register map).
> +
> +Device state is serialized as a versioned blob of per-VF register
> +(offset, value) pairs covering control, interrupt, RX/TX queue,
> +receive address (RA/RA2), etc. plus TX context descriptors and
> +VFRE/VFTE enable bits. The buffer address is a guest physical address
> +(GPA) written by the driver via ``virt_to_phys``; the device accesses
> +guest RAM directly through the system address space.
> +
> +Dirty page tracking is implemented with per-range bitmaps maintained
> +in IGBCore. All VF DMA paths in ``igb_core.c`` (TX data, RX data,
> +descriptor writeback) are instrumented to record touched pages. The
> +`igb-vfio-pci`_ variant driver registers tracked IOVA ranges and
> +queries dirty bitmaps through a shared buffer. Buffer structures
> +include len, flags, and reserved fields for future extensibility.
> +
> +The dirty bitmaps are maintained inside the device, which is not
> +realistic for discrete NICs without on-chip DRAM.
> +
> +DVSEC register layout
> +~~~~~~~~~~~~~~~~~~~~~
> +
> +The migration DVSEC (36 bytes at offset ``0x160``) uses a command
> +doorbell model. All commands are synchronous -- the device completes
> +the operation before the config write returns::
> +
> +  Offset  Name          Access  Description
> +  +0x00   ExtCap Hdr    RO      PCIe extended cap (id=0x23, ver=1)
> +  +0x04   DVSEC Hdr 1   RO      length[31:20] | rev[19:16] | vendor_id[15:0]
> +  +0x08   DVSEC Hdr 2   RO      DVSEC ID (1)
> +  +0x0A   Reserved      -       Padding for DWORD alignment
> +  +0x0C   CAPS          RO      F_STATE[0], F_DIRTY[1], max_ranges[11:8], pgsize[16:12]
> +  +0x10   CTRL          WO      Doorbell: cmd[7:0], arg[31:8]
> +  +0x14   STATUS        RO      state[7:0], error_code[15:8], quiesced[16]
> +  +0x18   BUF_ADDR_LO   RW      Shared DMA buffer GPA (low 32 bits)
> +  +0x1C   BUF_ADDR_HI   RW      Shared DMA buffer GPA (high 32 bits)
> +  +0x20   DATA_SIZE     RO      State blob size in bytes
> +
> +CTRL commands::
> +
> +  Cmd  Name            Arg             Description
> +  1    SET_STATE       state[31:8]     Set migration state
> +  2    SAVE            -               DMA-write state to buffer
> +  3    LOAD            size[31:8]      DMA-read state from buffer
> +  4    DIRTY_ENABLE    -               Enable dirty tracking (params in DMA buffer)
> +  5    DIRTY_DISABLE   -               Disable dirty tracking
> +  6    DIRTY_QUERY     -               Query dirty bitmap (via DMA buffer)
> +  7    GET_STATS       -               Query statistics (via DMA buffer)
> +
> +The driver sets ``BUF_ADDR_LO/HI`` before issuing commands that use a
> +DMA buffer (SAVE, LOAD, DIRTY_ENABLE, DIRTY_QUERY, GET_STATS). The
> +buffer address is latched per CTRL write, so the driver can use
> +different buffers for different commands. The buffer address is a
> +guest physical address (GPA).
> +
> +State transitions follow the VFIO migration v2 state machine. The
> +driver issues ``SET_STATE`` with the target state in the arg field and
> +reads ``STATUS`` to confirm the transition. Device states::
> +
> +  0  ERROR       1  STOP       2  RUNNING
> +  3  STOP_COPY   4  RESUMING   5  PRE_COPY
> +
> +``DATA_SIZE`` reflects the state blob size. At reset and in ``STOP``
> +state it holds the maximum size the driver should allocate. After
> +``SET_STATE(STOP_COPY)`` or ``SAVE`` it holds the actual serialized
> +size. The driver reads it after entering ``STOP_COPY`` to allocate an
> +exact-sized DMA buffer before issuing ``SAVE``.
> +
> +The state blob is a versioned sequence of register (offset, value)
> +pairs with magic ``0x4D494742`` ("MIGB").
> +
> +When ``STATUS`` state is ``ERROR`` (0), bits [15:8] contain an error
> +code identifying the failure::
> +
> +  1   UNK_CMD           Unknown CTRL command
> +  2   BAD_STATE         Command issued in wrong migration state
> +  3   NO_BUFFER         Command requires buffer but BUF_ADDR not set
> +  4   DMA_FAILED        DMA transfer to/from buffer failed
> +  5   BAD_SIZE          State blob too large or empty
> +  6   BAD_MAGIC         State blob magic mismatch
> +  7   BAD_VERSION       State blob version mismatch
> +  8   TOO_MANY_RANGES   Exceeds max_ranges from CAPS
> +  9   BAD_RANGE         Invalid range (zero size, misaligned, not contained)
> +  10  BAD_PGSIZE        Invalid or misaligned page size
> +  11  NOT_ENABLED       Dirty query without prior enable
> +
> +Dirty page tracking
> +~~~~~~~~~~~~~~~~~~~
> +
> +The migration interface supports per-VF dirty page tracking, advertised
> +by the ``F_DIRTY`` flag (bit 1) in ``CAPS``. This allows the variant
> +driver to enter ``PRE_COPY`` state while the VM continues to run,
> +iterating on dirty pages to reduce the final stop-and-copy window.
> +
> +The device maintains one dirty tracking engine per range, each with its
> +own bitmap scoped to the range boundaries. The ``CAPS`` register
> +advertises the maximum number of ranges in bits [11:8].
> +
> +Dirty tracking is controlled through CTRL commands:
> +
> +- **DIRTY_ENABLE** (4): the driver fills an ``igb_mig_dirty_enable_req``
> +  struct in the DMA buffer with the page size, range IOVA, and range
> +  size, then issues the command. The device allocates a bitmap for the
> +  range and begins recording pages touched by DMA. Supported page sizes
> +  are advertised in ``CAPS`` bits [16:12] (bit N = 2^N bytes). The driver
> +  checks ``STATUS`` for errors after the command completes.
> +- **DIRTY_DISABLE** (5): tears down all ranges and stops tracking.
> +- **DIRTY_QUERY** (6): the driver writes (iova, size) into the
> +  ``igb_mig_dirty_query`` DMA buffer, then issues the command. The
> +  device validates the range, copies the dirty bitmap into the buffer,
> +  and clears the tracked bits after successful DMA. The driver checks
> +  ``STATUS`` for errors after the command.
> +
> +Dirty enable DMA buffer
> +~~~~~~~~~~~~~~~~~~~~~~~
> +
> +The ``DIRTY_ENABLE`` command reads its parameters from the DMA buffer.
> +The ``len`` field holds the total structure size (including reserved
> +bytes) so the device can detect newer formats. ``flags`` and
> +``reserved`` must be zero::
> +
> +  Offset  Field        Type      Description
> +  0x00    len          uint32    Structure size in bytes
> +  0x04    flags        uint32    Reserved, must be 0
> +  0x08    pgsize       uint64    Page granularity (must match a CAPS pgsize bit)
> +  0x10    range_iova   uint64    Tracked range start address
> +  0x18    range_size   uint64    Tracked range size in bytes
> +  0x20    reserved[4]  uint32    Reserved, must be 0
> +
> +Dirty query DMA buffer
> +~~~~~~~~~~~~~~~~~~~~~~
> +
> +The ``DIRTY_QUERY`` command uses a shared DMA buffer for both request
> +and response. The ``len`` field holds the total buffer size (header +
> +bitmap). ``flags`` and ``reserved`` must be zero::
> +
> +  Offset  Field              Written by  Description
> +  0x00    len                driver      Total buffer size in bytes
> +  0x04    flags              driver      Reserved, must be 0
> +  0x08    iova               driver      Query range start
> +  0x10    size               driver      Query range size
> +  0x18    bitmap_size        device      Bytes written to bitmap
> +  0x1C    dirty_page_count   device      Number of set bits
> +  0x20    dma_writes         device      DMA write count (diagnostic)
> +  0x28    reserved[6]        -           Reserved, must be 0
> +  0x40    bitmap[]           device      Dirty page bitmap
> +
> +Migration statistics
> +~~~~~~~~~~~~~~~~~~~~
> +
> +The ``GET_STATS`` command DMA-writes a statistics response into the
> +driver-provided buffer. The driver sets ``BUF_ADDR_LO/HI`` and issues
> +the command; the device writes the response and returns::
> +
> +  Offset  Field                Type      Description
> +  0x00    dma_writes           uint64    DMA write operations tracked
> +  0x08    dma_bytes            uint64    DMA bytes written
> +  0x10    dirty_pages_set      uint32    Dirty pages marked since enable
> +  0x14    dirty_pages_cleared  uint32    Dirty pages cleared by queries
> +  0x18    dirty_page_count     uint32    Current dirty pages (set - cleared)
> +  0x1C    dirty_query_count    uint32    Number of QUERY operations
> +
> +The variant driver exposes these via debugfs at
> +``/sys/kernel/debug/vfio/<device>/migration/dirty/stats``.
> +
> +Testing setup
> +~~~~~~~~~~~~~
> +
> +The target scenario is nested virtualization: L0 runs QEMU with an
> +igb PF (``x-vf-migration=on``), L1 runs the `igb-vfio-pci`_ variant
> +driver and an unmodified QEMU, and L2 runs a standard igbvf driver.
> +See `Architecture`_ above for the full stack diagram.
> +
> +NetworkManager configuration
> +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
> +
> +In a nested setup, the L1 VMs (source and destination) have emulated
> +igb PFs connected to the L0 bridge. By default, NetworkManager
> +acquires DHCP leases on those PF interfaces and on any igb VFs created
> +later. This causes the VF MAC address to be learned on the L0 bridge,
> +which can misdirect iperf3 traffic after migration.
> +
> +To prevent this, configure NetworkManager on **both L1 VMs** and on the
> +**L2 guest disk image**.
> +
> +L1 VMs (source and destination)
> +...............................
> +
> +1. Prevent NetworkManager from managing igbvf interfaces:
> +
> +.. code-block:: bash
> +
> +   cat > /etc/NetworkManager/conf.d/99-no-igbvf.conf <<EOF
> +   [keyfile]
> +   unmanaged-devices=driver:igbvf
> +   EOF
> +
> +2. Disable IP on the igb PF connections (keep the interfaces UP for
> +   bridging, but with no DHCP lease):
> +
> +.. code-block:: bash
> +
> +   # Identify the NM connections for the igb PFs (NOT the virtio management NIC)
> +   nmcli -t -f NAME,DEVICE connection show
> +
> +   # For each igb PF connection:
> +   nmcli connection modify "<igb-pf-connection>" ipv4.method disabled ipv6.method disabled
> +
> +3. Reload NetworkManager:
> +
> +.. code-block:: bash
> +
> +   nmcli general reload
> +
> +L2 guest disk image
> +...................
> +
> +Use ``virt-customize`` to add the igbvf unmanaged config to the guest
> +image (offline, before any test run):
> +
> +.. code-block:: bash
> +
> +   virt-customize -a /srv/migration/rhel10.qcow2 \
> +     --write /etc/NetworkManager/conf.d/99-no-igbvf.conf:'[keyfile]
> +   unmanaged-devices=driver:igbvf'
> +
> +Network diagram
> +^^^^^^^^^^^^^^^
> +
> +The diagram below shows the nested setup where the source and
> +destination hosts are themselves VMs (L1) running on a physical
> +host (L0) that emulates the igb NIC::
> +
> +  ┌────────────────────────────────────────────────────────────────────────────┐
> +  │  L0: physical host                                                         │
> +  │                                                                            │
> +  │  virbr0  192.168.199.1/24                                                  │
> +  │  ├── NFS server: /srv/migration                                            │
> +  │  └── iperf3 client: iperf3 -c 192.168.199.200 -t 60 -i 1                   │
> +  │      │                                                                     │
> +  │      │  L0 virbr0 bridge (192.168.199.0/24)                                │
> +  │  ────┼──────────┬──────────────────┬──────────────────────────────         │
> +  │      │          │                  │                                       │
> +  │      │     ┌────┴────┐        ┌────┴────┐                                  │
> +  │      │     │ virtio  │        │ emulated│                                  │
> +  │      │     │c0:ff:ee:│        │  igb PF │    L0 QEMU (vm6)                 │
> +  │      │     │ :00:06  │        │ + igbvf │    tracks DMA dirty pages        │
> +  │      │     └────┬────┘        └────┬────┘                                  │
> +  │      │          │                  │                                       │
> +  │  ┌───┼──────────┼──────────────────┼──────────────────────────────────┐    │
> +  │  │   │  L1: vm6 (source)           │                                  │    │
> +  │  │   │  enp1s0: 192.168.199.6      │                                  │    │
> +  │  │   │  (management)               │                                  │    │
> +  │  │   │                        enp8s0 (igb PF, no IP)                  │    │
> +  │  │   │                             │                                  │    │
> +  │  │   │                        igb VF0 ──► igb-vfio-pci (VFIO)         │    │
> +  │  │   │                             │      dirty_sync → L0 igbvf       │    │
> +  │  │   │                             │                                  │    │
> +  │  │   │   virbr0                    │ VFIO passthrough                 │    │
> +  │  │   │   192.168.200.1/24          │                                  │    │
> +  │  │   │       │                     │                                  │    │
> +  │  │   │  ┌────┼─────────────────────┼───────────────────────────┐      │    │
> +  │  │   │  │    │  L2: rhel10 guest   │                           │      │    │
> +  │  │   │  │    │                     │                           │      │    │
> +  │  │   │  │  virtio NIC           igb VF (enp7s0)                │      │    │
> +  │  │   │  │  192.168.200.130/24   192.168.199.200/24             │      │    │
> +  │  │   │  │  (SSH login)          (iperf3 data path)             │      │    │
> +  │  │   │  │                          │                           │      │    │
> +  │  │   │  │            iperf3 -s -D  │ (listens on 0.0.0.0)      │      │    │
> +  │  │   │  └──────────────────────────┼───────────────────────────┘      │    │
> +  │  │   │                             │                                  │    │
> +  │  │   │  virsh migrate --live ──────┼──────────────────► vm7           │    │
> +  │  │   │                             │                                  │    │
> +  │  └───┼─────────────────────────────┼──────────────────────────────────┘    │
> +  │      │                             │                                       │
> +  │      │          iperf3 traffic     │                                       │
> +  │      └─────────────────────────────┘                                       │
> +  │                                                                            │
> +  │  ────────────────────────────────────────────────────────────────          │
> +  │      │                  │                                                  │
> +  │      │     ┌────────┐   │   ┌─────────┐                                    │
> +  │      │     │ virtio │   │   │emulated │    L0 QEMU (vm7)                   │
> +  │      │     │c0:ff:ee│   │   │ igb PF  │                                    │
> +  │      │     │ :00:07 │   │   │ + igbvf │                                    │
> +  │      │     └────┬───┘   │   └────┬────┘                                    │
> +  │  ┌──────────────┼───────┼────────┼────────────────────────────────────┐    │
> +  │  │   L1: vm7 (destination)       │                                    │    │
> +  │  │   enp1s0: 192.168.199.7       │                                    │    │
> +  │  │   (management)           enp8s0 (igb PF, no IP)                    │    │
> +  │  │                               │                                    │    │
> +  │  │                          igb VF0 ──► igb-vfio-pci (VFIO)           │    │
> +  │  │                               │                                    │    │
> +  │  │   virbr0                      │ VFIO passthrough                   │    │
> +  │  │   192.168.200.1/24            │                                    │    │
> +  │  │       │                       │                                    │    │
> +  │  │  ┌────┼───────────────────────┼────────────────────────────┐       │    │
> +  │  │  │    │  L2: rhel10 (after migration)                      │       │    │
> +  │  │  │    │                       │                            │       │    │
> +  │  │  │  virtio NIC             igb VF (enp7s0)                 │       │    │
> +  │  │  │  192.168.200.130/24     192.168.199.200/24              │       │    │
> +  │  │  │                            │                            │       │    │
> +  │  │  │              iperf3 -s -D  │ (connection survives)      │       │    │
> +  │  │  └────────────────────────────┼────────────────────────────┘       │    │
> +  │  └───────────────────────────────┼────────────────────────────────────┘    │
> +  │                                  │                                         │
> +  │      iperf3 traffic resumes ─────┘                                         │
> +  │      (same IP, same MAC, same L2 segment → transparent to client)          │
> +  └────────────────────────────────────────────────────────────────────────────┘
> +
> +Migration under iperf3 load works correctly: dirty page tracking
> +converges (from ~2000 pages per PRE_COPY iteration down to ~280 at
> +STOP_COPY), and STOP_COPY stays under 250ms.
> +
> +Todo
> +~~~~
> +
> +1. Add migration blocker when ``x-vf-migration=on`` (no VMState yet) or
> +   add VMState support for L0 migration (dirty bitmaps, tracking
> +   engines, DVSEC registers, stats)
> +2. Add PRE_COPY state transfer to validate device INIT data (magic,
> +   version, etc.)
> +3. Add qtests for migration state machine transitions, dirty page
> +   tracking
> +
> +Ideas
> +~~~~~
> +
> +1. **RX bandwidth throttle** (``x-mig-rx-limit``, uint32, default 0)
> +
> +   Return false from ``can_receive`` when the per-VF packet count in the
> +   current tracking interval exceeds the limit. Reduces DMA writes and
> +   dirty pages realistically.
> +
> +2. **Migration phase timing** (GET_STATS extension)
> +
> +   Add per-VF timestamps: ``precopy_start_ns``, ``stopcopy_start_ns``,
> +   ``precopy_duration_ns``, ``stopcopy_duration_ns``,
> +   ``state_transition_count``. Expose via GET_STATS.
> +
> +3. **Hot page simulation** (``x-mig-hot-pages``, uint32, default 0)
> +
> +   Re-set the first N bitmap bits after each DIRTY_QUERY, simulating
> +   workloads with hot pages that prevent convergence.
> +
> +4. **Error injection** (``x-mig-inject-error``, uint32, default 0)
> +
> +   One-shot error code injection before command dispatch. A separate
> +   ``x-mig-inject-dma-fail`` (bool) for persistent DMA failure testing.
> +
> +AI disclaimer
> +~~~~~~~~~~~~~
> +
> +Claude was used to analyze the IGB PF and VF internal state and
> +identify the pain points of a working live migration of such devices.
> +The generated code served as a starting point but *significant* time
> +was then spent cleaning up, reworking, and shaping it into a clear,
> +reviewable IGB model extension.

I don't think it's really a good idea to keep this "AI disclaimer" in 
the documentation.

> +
> +.. _igb-vfio-pci: https://github.com/legoater/vfio-pci-extras
> diff --git a/docs/system/devices/igb.rst b/docs/system/devices/igb.rst
> index 50f625fd77e4..00271dbc92c3 100644
> --- a/docs/system/devices/igb.rst
> +++ b/docs/system/devices/igb.rst
> @@ -64,6 +64,12 @@ command:
>   
>     pyvenv/bin/meson test --suite thorough func-x86_64-netdev_ethtool
>   
> +VF Migration (experimental)
> +===========================
> +
> +See :ref:`igb-migration` for details on the experimental VF live migration
> +interface.
> +
>   References
>   ==========
>   



^ permalink raw reply	[flat|nested] 16+ messages in thread

* Re: [RFC PATCH v2 0/9] igb: Add experimental VF live migration support
  2026-09-02 19:20 [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Cédric Le Goater
                   ` (8 preceding siblings ...)
  2026-09-02 19:20 ` [RFC PATCH v2 9/9] docs: Add igb VF migration testing setup guide Cédric Le Goater
@ 2026-09-09  7:19 ` Akihiko Odaki
  9 siblings, 0 replies; 16+ messages in thread
From: Akihiko Odaki @ 2026-09-09  7:19 UTC (permalink / raw)
  To: Cédric Le Goater, qemu-devel
  Cc: Sriram Yagnaraman, Jason Wang, Alex Williamson, Peter Xu

On 2026/09/03 4:20, Cédric Le Goater wrote:
> Hello,
> 
> Live migration of VFIO-passthrough devices - SR-IOV VFs, vGPUs - is a
> growing requirement, but real hardware with migration support is
> scarce and hard to debug. An emulated device provides a fully
> controlled testbed for developing and validating the entire software
> stack - vfio-pci variant drivers, VFIO core migration v2 framework,
> QEMU, libvirt - and for tuning complex migration policies such as
> downtime convergence. It also serves as an educational reference for
> understanding VFIO migration end-to-end, from device state
> serialization to dirty page tracking.
> 
> This series adds an experimental VF live migration interface to the
> emulated igb (82576) device. It enables a vfio-pci variant driver
> (igb-vfio-pci) to migrate VFs using the standard VFIO migration v2
> protocol with stop-copy and pre-copy support.
> 
> The target scenario is nested virtualization:
> 
>    L0 QEMU (these patches)
>      igb PF with x-vf-migration=on
>      └── VFs with migration DVSEC
> 
>    L1 kernel
>      igb-vfio-pci variant driver [1]
>      translates VFIO migration v2 ioctls → DVSEC config writes
> 
>    L1 QEMU (stock, unmodified)
>      vfio-pci device model, standard migration fd
> 
>    L2 guest
>      standard igbvf driver, unaware of migration
> 
> The L1 QEMU is completely unmodified -- it sees a standard VFIO
> migratable device and uses the normal migration fd path.
> 
> * Design
> 
> The migration interface is exposed through a DVSEC (Designated
> Vendor-Specific Extended Capability, PCIe cap id 0x23) at offset
> 0x160 in VF extended config space. The DVSEC uses a command doorbell
> model - all commands are synchronous via PCI config space writes.
> 
> Device state is serialized as a versioned blob of per-VF register
> (offset, value) pairs covering control, interrupt, RX/TX queue,
> receive address (RA/RA2), etc. plus TX context descriptors and
> VFRE/VFTE enable bits. The buffer address is a guest physical
> address (GPA) written by the driver via virt_to_phys; the device
> accesses guest RAM directly through the system address space.
> 
> Dirty page tracking is implemented with per-range bitmaps maintained
> in IGBCore. All VF DMA paths in igb_core.c (TX data, RX data,
> descriptor writeback) are instrumented to record touched pages. The
> variant driver registers tracked IOVA ranges and queries dirty bitmaps
> through a shared buffer. Buffer structures include len, flags, and
> reserved fields for future extensibility.
> 
> * Caveats
> 
> The x-vf-migration property is experimental (x- prefix, default off).
> 
> The dirty bitmaps are maintained inside the device, which is not
> realistic for discrete NICs without on-chip DRAM.
> 
> * Testing
> 
> The target scenario is nested virtualization: L0 runs QEMU with an
> igb PF (x-vf-migration=on), L1 runs the igb-vfio-pci variant driver
> and an unmodified QEMU, and L2 runs a standard igbvf driver.
> 
> Migration under iperf3 load works correctly: dirty page tracking
> converges (from ~2000 pages per PRE_COPY iteration down to ~280 at
> STOP_COPY), and STOP_COPY stays under 250ms.
> 
> * Todo
> 
>    1. Add migration blocker when x-vf-migration=on (no VMState yet) or
>       add VMState support for L0 migration (dirty bitmaps, tracking
>       engines, DVSEC registers, stats)
>    2. Add PRE_COPY state transfer to validate device INIT data (magic,
>       version, etc.)
>    3. Add qtests for migration state machine transitions, dirty page
>       tracking ?
> 
> * Ideas
> 
>    1. RX bandwidth throttle (x-mig-rx-limit, uint32, default 0)
> 
>       Return false from can_receive when the per-VF packet count in the
>       current tracking interval exceeds the limit. Reduces DMA writes
>       and dirty pages realistically.
> 
>    2. Migration phase timing (GET_STATS extension)
> 
>       Add per-VF timestamps: precopy_start_ns, stopcopy_start_ns,
>       precopy_duration_ns, stopcopy_duration_ns,
>       state_transition_count. Expose via GET_STATS.
> 
>    3. Hot page simulation (x-mig-hot-pages, uint32, default 0)
> 
>       Re-set the first N bitmap bits after each DIRTY_QUERY, simulating
>       workloads with hot pages that prevent convergence.
> 
>    4. Error injection (x-mig-inject-error, uint32, default 0)
> 
>       One-shot error code injection before command dispatch. A separate
>       x-mig-inject-dma-fail (bool) for persistent DMA failure testing.
> 
> * Credits
> 
> Alex Williamson suggested the overall approach of a variant driver
> with the "x-vf-migration" device property to gate the feature. Thanks
> for the ever ongoing support and valuable discussions throughout these
> years.
> 
> * AI disclaimer
> 
> The lack of a migration-capable device has been a recurring pain point
> for VFIO development over the years, and we hope this proposal
> demonstrates the value of having one.
> 
> Claude was used to analyze the IGB PF and VF internal state and
> identify the pain points of a working live migration of such devices.
> The generated code served as a starting point but *significant* time
> was then spent cleaning up, reworking, and shaping it into a clear,
> reviewable IGB model extension.
> 
> As QEMU does not yet accept AI-assisted contributions, this series is
> submitted as an RFC.

This largely repeats the concerns I raised previously [1]. Your 
description of the *significant* time spent cleaning up and reworking 
the generated code suggests that this series is intended for upstream 
inclusion once the implementation direction is agreed. My concerns as a 
maintainer are therefore correctness and maintainability.

The missing migration blocker is a concrete example: it is a basic 
correctness issue that should be straightforward to address, yet it 
remains a TODO. My comments on individual patches also raise the same 
kinds of issues I pointed out in the previous version.

Many of the correctness problems found in v1 and v2 stem from the choice 
to build on igb, which is more complex than virtio-net. That is why I 
suggested using virtio-net as a simpler basis, unless there is a problem 
that rules it out.

I would be willing to support allowing AI use in the project for work 
like this. I would still need to see these correctness and 
maintainability concerns addressed before I could support merging
the series.

Regards,
Akihiko Odaki

[1] 
https://lore.kernel.org/qemu-devel/20260727053935.1392269-1-clg@redhat.com/


^ permalink raw reply	[flat|nested] 16+ messages in thread

end of thread, other threads:[~2026-09-09  7:20 UTC | newest]

Thread overview: 16+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-02 19:20 [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Cédric Le Goater
2026-09-02 19:20 ` [RFC PATCH v2 1/9] igb: Add x-vf-migration property and DVSEC extended capability Cédric Le Goater
2026-09-03 19:57   ` Alex Williamson
2026-09-07 20:58     ` Cédric Le Goater
2026-09-02 19:20 ` [RFC PATCH v2 2/9] igb: Add migration state machine via extended config space Cédric Le Goater
2026-09-02 19:20 ` [RFC PATCH v2 3/9] igb: Add VF state serialization for live migration Cédric Le Goater
2026-09-08  8:01   ` Akihiko Odaki
2026-09-02 19:20 ` [RFC PATCH v2 4/9] igb: Add VF post-load fixups " Cédric Le Goater
2026-09-08  8:10   ` Akihiko Odaki
2026-09-02 19:20 ` [RFC PATCH v2 5/9] igb: Add dirty page tracking for IGBVF migration Cédric Le Goater
2026-09-02 19:20 ` [RFC PATCH v2 6/9] igb: Quiesce VFs on STOP and include PF enable state in migration Cédric Le Goater
2026-09-02 19:20 ` [RFC PATCH v2 7/9] igb: Fix post-migration RX ring deadlock Cédric Le Goater
2026-09-02 19:20 ` [RFC PATCH v2 8/9] igb: Add dirty page tracking statistics Cédric Le Goater
2026-09-02 19:20 ` [RFC PATCH v2 9/9] docs: Add igb VF migration testing setup guide Cédric Le Goater
2026-09-08  8:39   ` Akihiko Odaki
2026-09-09  7:19 ` [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Akihiko Odaki

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.