* [RFC PATCH v2 0/9] igb: Add experimental VF live migration support
@ 2026-09-02 19:20 Cédric Le Goater
2026-09-02 19:20 ` [RFC PATCH v2 1/9] igb: Add x-vf-migration property and DVSEC extended capability Cédric Le Goater
` (9 more replies)
0 siblings, 10 replies; 23+ messages in thread
From: Cédric Le Goater @ 2026-09-02 19:20 UTC (permalink / raw)
To: qemu-devel
Cc: Akihiko Odaki, Sriram Yagnaraman, Jason Wang, Alex Williamson,
Peter Xu, Cédric Le Goater
Hello,
Live migration of VFIO-passthrough devices - SR-IOV VFs, vGPUs - is a
growing requirement, but real hardware with migration support is
scarce and hard to debug. An emulated device provides a fully
controlled testbed for developing and validating the entire software
stack - vfio-pci variant drivers, VFIO core migration v2 framework,
QEMU, libvirt - and for tuning complex migration policies such as
downtime convergence. It also serves as an educational reference for
understanding VFIO migration end-to-end, from device state
serialization to dirty page tracking.
This series adds an experimental VF live migration interface to the
emulated igb (82576) device. It enables a vfio-pci variant driver
(igb-vfio-pci) to migrate VFs using the standard VFIO migration v2
protocol with stop-copy and pre-copy support.
The target scenario is nested virtualization:
L0 QEMU (these patches)
igb PF with x-vf-migration=on
└── VFs with migration DVSEC
L1 kernel
igb-vfio-pci variant driver [1]
translates VFIO migration v2 ioctls → DVSEC config writes
L1 QEMU (stock, unmodified)
vfio-pci device model, standard migration fd
L2 guest
standard igbvf driver, unaware of migration
The L1 QEMU is completely unmodified -- it sees a standard VFIO
migratable device and uses the normal migration fd path.
* Design
The migration interface is exposed through a DVSEC (Designated
Vendor-Specific Extended Capability, PCIe cap id 0x23) at offset
0x160 in VF extended config space. The DVSEC uses a command doorbell
model - all commands are synchronous via PCI config space writes.
Device state is serialized as a versioned blob of per-VF register
(offset, value) pairs covering control, interrupt, RX/TX queue,
receive address (RA/RA2), etc. plus TX context descriptors and
VFRE/VFTE enable bits. The buffer address is a guest physical
address (GPA) written by the driver via virt_to_phys; the device
accesses guest RAM directly through the system address space.
Dirty page tracking is implemented with per-range bitmaps maintained
in IGBCore. All VF DMA paths in igb_core.c (TX data, RX data,
descriptor writeback) are instrumented to record touched pages. The
variant driver registers tracked IOVA ranges and queries dirty bitmaps
through a shared buffer. Buffer structures include len, flags, and
reserved fields for future extensibility.
* Caveats
The x-vf-migration property is experimental (x- prefix, default off).
The dirty bitmaps are maintained inside the device, which is not
realistic for discrete NICs without on-chip DRAM.
* Testing
The target scenario is nested virtualization: L0 runs QEMU with an
igb PF (x-vf-migration=on), L1 runs the igb-vfio-pci variant driver
and an unmodified QEMU, and L2 runs a standard igbvf driver.
Migration under iperf3 load works correctly: dirty page tracking
converges (from ~2000 pages per PRE_COPY iteration down to ~280 at
STOP_COPY), and STOP_COPY stays under 250ms.
* Todo
1. Add migration blocker when x-vf-migration=on (no VMState yet) or
add VMState support for L0 migration (dirty bitmaps, tracking
engines, DVSEC registers, stats)
2. Add PRE_COPY state transfer to validate device INIT data (magic,
version, etc.)
3. Add qtests for migration state machine transitions, dirty page
tracking ?
* Ideas
1. RX bandwidth throttle (x-mig-rx-limit, uint32, default 0)
Return false from can_receive when the per-VF packet count in the
current tracking interval exceeds the limit. Reduces DMA writes
and dirty pages realistically.
2. Migration phase timing (GET_STATS extension)
Add per-VF timestamps: precopy_start_ns, stopcopy_start_ns,
precopy_duration_ns, stopcopy_duration_ns,
state_transition_count. Expose via GET_STATS.
3. Hot page simulation (x-mig-hot-pages, uint32, default 0)
Re-set the first N bitmap bits after each DIRTY_QUERY, simulating
workloads with hot pages that prevent convergence.
4. Error injection (x-mig-inject-error, uint32, default 0)
One-shot error code injection before command dispatch. A separate
x-mig-inject-dma-fail (bool) for persistent DMA failure testing.
* Credits
Alex Williamson suggested the overall approach of a variant driver
with the "x-vf-migration" device property to gate the feature. Thanks
for the ever ongoing support and valuable discussions throughout these
years.
* AI disclaimer
The lack of a migration-capable device has been a recurring pain point
for VFIO development over the years, and we hope this proposal
demonstrates the value of having one.
Claude was used to analyze the IGB PF and VF internal state and
identify the pain points of a working live migration of such devices.
The generated code served as a starting point but *significant* time
was then spent cleaning up, reworking, and shaping it into a clear,
reviewable IGB model extension.
As QEMU does not yet accept AI-assisted contributions, this series is
submitted as an RFC.
Thanks,
C.
[1] https://github.com/legoater/vfio-pci-extras
* Changes since rfc-v1
- Migration BAR replaced with DVSEC at offset 0x160 (no BAR needed)
- Wire-format structs (IgbMigBlob, IgbMigRegPair, IgbMigTxCtx)
replace raw pointer arithmetic and memcpy
- RA entries separated from fixed regs, scanned by pool bit
- VFN relocation support (offset remapping + RA pool bit swap)
- GPA buffer (address_space_read/write) replaces PCI DMA through PF
- State blob validation on load (magic, version, error codes)
- NEED_WORDS macro and igb_vf_offset_valid removed
- Error codes renumbered: removed BAD_VFN, added UNK_CMD (1-11)
Reported by Akihiko Odaki:
- propagate_irqs: clear VF bits before OR (EIMS/EIAC/EIAM)
- propagate_ivar: clear IVAR entry when source VTIVAR is invalid
- rearm_irqs: restore actual PVTEICR causes, not all three
- Dirty bitmap allocation uses BITS_TO_LONGS (heap corruption fix)
- Dirty query validates range before g_malloc0 (memory exhaustion)
- Dirty bits cleared only after successful bitmap DMA write
- Dirty query buffer uses struct offsets (layout mismatch fix)
- Load path: register offsets validated against VF whitelist
- VMBMEM (mailbox payload) documented as transient, not serialized
- dma_writes counter: consistently uint64_t
- Dirty range_size: consistently uint64_t (was truncated to 32 bits)
- ERROR->STOP: quiesce VF (clear VFRE/VFTE) on transition
- rearm_irqs: runs on re || te, not just re (TX-only VF fix)
- Stats DMA-written atomically via GET_STATS (no split MMIO tear)
- Bisectability: DVSEC + state machine introduced together
- Removed NAPI reference and "===" comment decoration
Cédric Le Goater (9):
igb: Add x-vf-migration property and DVSEC extended capability
igb: Add migration state machine via extended config space
igb: Add VF state serialization for live migration
igb: Add VF post-load fixups for live migration
igb: Add dirty page tracking for IGBVF migration
igb: Quiesce VFs on STOP and include PF enable state in migration
igb: Fix post-migration RX ring deadlock
igb: Add dirty page tracking statistics
docs: Add igb VF migration testing setup guide
MAINTAINERS | 6 +
docs/system/device-emulation.rst | 1 +
docs/system/devices/igb-migration.rst | 417 +++++++++
docs/system/devices/igb.rst | 6 +
hw/net/igb_common.h | 16 +
hw/net/igb_core.h | 11 +
hw/net/igb_migration.h | 186 ++++
hw/net/igb.c | 7 +
hw/net/igb_core.c | 124 ++-
hw/net/igb_migration.c | 1128 +++++++++++++++++++++++++
hw/net/igbvf.c | 52 +-
hw/net/meson.build | 2 +-
hw/net/trace-events | 13 +
13 files changed, 1945 insertions(+), 24 deletions(-)
create mode 100644 docs/system/devices/igb-migration.rst
create mode 100644 hw/net/igb_migration.h
create mode 100644 hw/net/igb_migration.c
--
2.55.0
^ permalink raw reply [flat|nested] 23+ messages in thread
* [RFC PATCH v2 1/9] igb: Add x-vf-migration property and DVSEC extended capability
2026-09-02 19:20 [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Cédric Le Goater
@ 2026-09-02 19:20 ` Cédric Le Goater
2026-09-03 19:57 ` Alex Williamson
2026-09-02 19:20 ` [RFC PATCH v2 2/9] igb: Add migration state machine via extended config space Cédric Le Goater
` (8 subsequent siblings)
9 siblings, 1 reply; 23+ messages in thread
From: Cédric Le Goater @ 2026-09-02 19:20 UTC (permalink / raw)
To: qemu-devel
Cc: Akihiko Odaki, Sriram Yagnaraman, Jason Wang, Alex Williamson,
Peter Xu, Cédric Le Goater
Add an "x-vf-migration" property to the IGB PF device and expose a
DVSEC (Designated Vendor-Specific Extended Capability) at offset 0x160
in VF extended config space when migration is enabled.
The DVSEC provides the register interface for VF live migration:
- CAPS: supported features (state migration)
- CTRL: command doorbell
- STATUS: state and error reporting
- BUF_ADDR: shared DMA buffer address (GPA)
Move IgbVfState from igbvf.c to igb_common.h so it can be shared with
the migration module, and add the migration state field.
AI-used-for: code (prototype)
Signed-off-by: Cédric Le Goater <clg@redhat.com>
---
MAINTAINERS | 5 ++
hw/net/igb_common.h | 16 +++++++
hw/net/igb_migration.h | 72 ++++++++++++++++++++++++++++
hw/net/igb.c | 2 +
hw/net/igb_migration.c | 106 +++++++++++++++++++++++++++++++++++++++++
hw/net/igbvf.c | 52 ++++++++++++++++----
hw/net/meson.build | 2 +-
7 files changed, 245 insertions(+), 10 deletions(-)
create mode 100644 hw/net/igb_migration.h
create mode 100644 hw/net/igb_migration.c
diff --git a/MAINTAINERS b/MAINTAINERS
index 4a49a40294eb..f88b526be238 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -2816,6 +2816,11 @@ F: tests/functional/x86_64/test_netdev_ethtool.py
F: tests/qtest/igb-test.c
F: tests/qtest/libqos/igb.c
+igb VF migration
+M: Cédric Le Goater <clg@redhat.com>
+S: Maintained
+F: hw/net/igb_migration.*
+
eepro100
M: Stefan Weil <sw@weilnetz.de>
S: Maintained
diff --git a/hw/net/igb_common.h b/hw/net/igb_common.h
index b316a5bcfa5c..f0e2529e757b 100644
--- a/hw/net/igb_common.h
+++ b/hw/net/igb_common.h
@@ -26,7 +26,9 @@
#ifndef HW_NET_IGB_COMMON_H
#define HW_NET_IGB_COMMON_H
+#include "hw/pci/pci_device.h"
#include "igb_regs.h"
+#include "igb_migration.h"
#define TYPE_IGBVF "igbvf"
@@ -154,4 +156,18 @@ uint64_t igb_mmio_read(void *opaque, hwaddr addr, unsigned size);
void igb_mmio_write(void *opaque, hwaddr addr, uint64_t val, unsigned size);
void igb_vf_reset(void *opaque, uint16_t vfn);
+OBJECT_DECLARE_SIMPLE_TYPE(IgbVfState, IGBVF)
+
+struct IgbVfState {
+ PCIDevice parent_obj;
+
+ uint16_t vfn;
+ bool migration_enabled;
+
+ MemoryRegion mmio;
+ MemoryRegion msix;
+
+ IgbVfMigState mig;
+};
+
#endif
diff --git a/hw/net/igb_migration.h b/hw/net/igb_migration.h
new file mode 100644
index 000000000000..3da28e11e49e
--- /dev/null
+++ b/hw/net/igb_migration.h
@@ -0,0 +1,72 @@
+/*
+ * QEMU Intel 82576 SR/IOV VF Migration Support
+ *
+ * Copyright (c) 2026 Red Hat, Inc.
+ *
+ * SPDX-License-Identifier: GPL-2.0-or-later
+ */
+
+#ifndef HW_NET_IGB_MIGRATION_H
+#define HW_NET_IGB_MIGRATION_H
+
+#include "hw/pci/pci_device.h"
+
+/*
+ * Migration interface exposed as a DVSEC (Designated Vendor-Specific
+ * Extended Capability) in VF extended config space.
+ *
+ * DVSEC layout at IGB_MIG_DVSEC_OFFSET (0x160):
+ *
+ * +0x00 PCIe extended cap header (cap_id=0x23, ver=1, next)
+ * +0x04 DVSEC header 1 (len | rev | vendor_id)
+ * +0x08 DVSEC header 2 (DVSEC ID)
+ * +0x0A Reserved (padding for DWORD alignment)
+ * +0x0C CAPS (RO: F_STATE[0])
+ * +0x10 CTRL (WO: doorbell command)
+ * +0x14 STATUS (RO: state[7:0], error_code[15:8])
+ * +0x18 BUF_ADDR_LO (RW: shared buffer GPA low)
+ * +0x1C BUF_ADDR_HI (RW: shared buffer GPA high)
+ */
+
+#define IGB_MIG_DVSEC_OFFSET 0x160
+#define IGB_MIG_DVSEC_SIZE 0x20
+#define IGB_MIG_DVSEC_VER 1
+#define IGB_MIG_DVSEC_ID 1
+
+/* Register offsets relative to DVSEC base */
+#define IGB_MIG_CAPS 0x0C
+#define IGB_MIG_CTRL 0x10
+#define IGB_MIG_STATUS 0x14
+#define IGB_MIG_BUF_ADDR_LO 0x18
+#define IGB_MIG_BUF_ADDR_HI 0x1C
+
+/* CAPS register layout */
+#define IGB_MIG_CAP_F_STATE (1u << 0)
+
+/* STATUS register: state in [7:0], error code [15:8] */
+#define IGB_MIG_STATUS_STATE_MASK 0xFF
+#define IGB_MIG_STATUS_ERROR_CODE_SHIFT 8
+#define IGB_MIG_STATUS_ERR(code) \
+ ((uint32_t)(code) << IGB_MIG_STATUS_ERROR_CODE_SHIFT)
+
+/* Device states (based on VFIO migration v2) */
+#define IGB_MIG_STATE_ERROR 0
+#define IGB_MIG_STATE_STOP 1
+#define IGB_MIG_STATE_RUNNING 2
+#define IGB_MIG_STATE_STOP_COPY 3
+#define IGB_MIG_STATE_RESUMING 4
+
+typedef struct IgbVfMigState {
+ uint32_t mig_state;
+ uint64_t mig_data_buf_addr;
+} IgbVfMigState;
+
+typedef struct IgbVfState IgbVfState;
+
+bool igbvf_add_migration_dvsec(PCIDevice *dev, Error **errp);
+void igbvf_mig_state_reset(IgbVfState *s);
+uint32_t igbvf_mig_config_read(IgbVfState *s, uint32_t addr, int size);
+bool igbvf_mig_config_write(IgbVfState *s, uint32_t addr, uint32_t val,
+ int size);
+
+#endif
diff --git a/hw/net/igb.c b/hw/net/igb.c
index c076807e7110..7268e5473fc3 100644
--- a/hw/net/igb.c
+++ b/hw/net/igb.c
@@ -79,6 +79,7 @@ struct IGBState {
IGBCore core;
bool has_flr;
+ bool vf_migration;
};
#define IGB_CAP_SRIOV_OFFSET (0x160)
@@ -597,6 +598,7 @@ static const VMStateDescription igb_vmstate = {
static const Property igb_properties[] = {
DEFINE_NIC_PROPERTIES(IGBState, conf),
DEFINE_PROP_BOOL("x-pcie-flr-init", IGBState, has_flr, true),
+ DEFINE_PROP_BOOL("x-vf-migration", IGBState, vf_migration, false),
};
static void igb_class_init(ObjectClass *class, const void *data)
diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
new file mode 100644
index 000000000000..4dfebd82344c
--- /dev/null
+++ b/hw/net/igb_migration.c
@@ -0,0 +1,106 @@
+/*
+ * QEMU Intel 82576 SR/IOV VF Migration Support
+ *
+ * Copyright (c) 2026 Red Hat, Inc.
+ *
+ * SPDX-License-Identifier: GPL-2.0-or-later
+ */
+
+#include "qemu/osdep.h"
+#include "hw/pci/pci_device.h"
+#include "hw/pci/pcie.h"
+#include "igb_common.h"
+#include "igb_migration.h"
+
+static void igbvf_mig_update_status(IgbVfState *s, uint8_t err)
+{
+ IgbVfMigState *ms = &s->mig;
+ PCIDevice *dev = PCI_DEVICE(s);
+ uint32_t status;
+
+ status = ms->mig_state & IGB_MIG_STATUS_STATE_MASK;
+
+ if (err) {
+ status = IGB_MIG_STATE_ERROR | IGB_MIG_STATUS_ERR(err);
+ }
+
+ pci_set_long(dev->config + IGB_MIG_DVSEC_OFFSET + IGB_MIG_STATUS, status);
+}
+
+
+bool igbvf_add_migration_dvsec(PCIDevice *dev, Error **errp)
+{
+ uint16_t offset = IGB_MIG_DVSEC_OFFSET;
+ uint32_t caps;
+
+ pcie_add_capability(dev, PCI_EXT_CAP_ID_DVSEC, 1, offset,
+ IGB_MIG_DVSEC_SIZE);
+
+ /* DVSEC header 1: length[31:20] | rev[19:16] | vendor_id[15:0] */
+ pci_set_long(dev->config + offset + 0x4,
+ (IGB_MIG_DVSEC_SIZE << 20) |
+ (IGB_MIG_DVSEC_VER << 16) |
+ PCI_VENDOR_ID_INTEL);
+
+ /* DVSEC header 2: DVSEC ID */
+ pci_set_word(dev->config + offset + 0x8, IGB_MIG_DVSEC_ID);
+
+ /* CAPS: features (state migration only) */
+ caps = IGB_MIG_CAP_F_STATE;
+ pci_set_long(dev->config + offset + IGB_MIG_CAPS, caps);
+
+ /* STATUS: initial state is RUNNING */
+ pci_set_long(dev->config + offset + IGB_MIG_STATUS,
+ IGB_MIG_STATE_RUNNING);
+
+ /* BUF_ADDR_LO and BUF_ADDR_HI are writable */
+ memset(dev->wmask + offset + IGB_MIG_BUF_ADDR_LO, 0xff, 4);
+ memset(dev->wmask + offset + IGB_MIG_BUF_ADDR_HI, 0xff, 4);
+
+ return true;
+}
+
+uint32_t igbvf_mig_config_read(IgbVfState *s, uint32_t addr, int size)
+{
+ PCIDevice *dev = PCI_DEVICE(s);
+
+ return pci_default_read_config(dev, addr, size);
+}
+
+bool igbvf_mig_config_write(IgbVfState *s, uint32_t addr, uint32_t val,
+ int size)
+{
+ PCIDevice *dev = PCI_DEVICE(s);
+ uint32_t offset = addr - IGB_MIG_DVSEC_OFFSET;
+
+ switch (offset) {
+ case IGB_MIG_CTRL:
+ /* Command handling will be added in a later commit */
+ break;
+
+ case IGB_MIG_BUF_ADDR_LO:
+ case IGB_MIG_BUF_ADDR_HI:
+ pci_default_write_config(dev, addr, val, size);
+ break;
+
+ default:
+ break;
+ }
+
+ return true;
+}
+
+void igbvf_mig_state_reset(IgbVfState *s)
+{
+ IgbVfMigState *ms = &s->mig;
+
+ ms->mig_state = IGB_MIG_STATE_RUNNING;
+ ms->mig_data_buf_addr = 0;
+
+ pci_set_long(PCI_DEVICE(s)->config +
+ IGB_MIG_DVSEC_OFFSET + IGB_MIG_BUF_ADDR_LO, 0);
+ pci_set_long(PCI_DEVICE(s)->config +
+ IGB_MIG_DVSEC_OFFSET + IGB_MIG_BUF_ADDR_HI, 0);
+
+ igbvf_mig_update_status(s, 0);
+}
diff --git a/hw/net/igbvf.c b/hw/net/igbvf.c
index 9a165c7063ee..30dfdb574ac7 100644
--- a/hw/net/igbvf.c
+++ b/hw/net/igbvf.c
@@ -38,27 +38,21 @@
*/
#include "qemu/osdep.h"
+#include "qemu/range.h"
#include "hw/core/hw-error.h"
#include "hw/net/mii.h"
#include "hw/pci/pci_device.h"
#include "hw/pci/pcie.h"
+#include "hw/pci/pcie_sriov.h"
#include "hw/pci/msix.h"
#include "net/eth.h"
#include "net/net.h"
#include "igb_common.h"
#include "igb_core.h"
+#include "igb_migration.h"
#include "trace.h"
#include "qapi/error.h"
-OBJECT_DECLARE_SIMPLE_TYPE(IgbVfState, IGBVF)
-
-struct IgbVfState {
- PCIDevice parent_obj;
-
- MemoryRegion mmio;
- MemoryRegion msix;
-};
-
static hwaddr vf_to_pf_addr(hwaddr addr, uint16_t vfn, bool write)
{
switch (addr) {
@@ -199,10 +193,35 @@ static hwaddr vf_to_pf_addr(hwaddr addr, uint16_t vfn, bool write)
return HWADDR_MAX;
}
+static bool igbvf_addr_in_dvsec(uint32_t addr, int len)
+{
+ return ranges_overlap(addr, len,
+ IGB_MIG_DVSEC_OFFSET, IGB_MIG_DVSEC_SIZE);
+}
+
+static uint32_t igbvf_read_config(PCIDevice *dev, uint32_t addr, int size)
+{
+ IgbVfState *s = IGBVF(dev);
+
+ if (s->migration_enabled && igbvf_addr_in_dvsec(addr, size)) {
+ return igbvf_mig_config_read(s, addr, size);
+ }
+
+ return pci_default_read_config(dev, addr, size);
+}
+
static void igbvf_write_config(PCIDevice *dev, uint32_t addr, uint32_t val,
int len)
{
+ IgbVfState *s = IGBVF(dev);
+
trace_igbvf_write_config(addr, val, len);
+
+ if (s->migration_enabled && igbvf_addr_in_dvsec(addr, len)) {
+ igbvf_mig_config_write(s, addr, val, len);
+ return;
+ }
+
pci_default_write_config(dev, addr, val, len);
if (object_property_get_bool(OBJECT(pcie_sriov_get_pf(dev)),
"x-pcie-flr-init", &error_abort)) {
@@ -282,13 +301,27 @@ static void igbvf_pci_realize(PCIDevice *dev, Error **errp)
}
pcie_ari_init(dev, 0x150);
+
+ if (object_property_get_bool(OBJECT(pcie_sriov_get_pf(dev)),
+ "x-vf-migration", &error_abort)) {
+ s->vfn = pcie_sriov_vf_number(dev);
+ s->migration_enabled = true;
+ if (!igbvf_add_migration_dvsec(dev, errp)) {
+ return;
+ }
+ }
}
static void igbvf_qdev_reset_hold(Object *obj, ResetType type)
{
PCIDevice *vf = PCI_DEVICE(obj);
+ IgbVfState *s = IGBVF(vf);
igb_vf_reset(pcie_sriov_get_pf(vf), pcie_sriov_vf_number(vf));
+
+ if (s->migration_enabled) {
+ igbvf_mig_state_reset(s);
+ }
}
static void igbvf_pci_uninit(PCIDevice *dev)
@@ -309,6 +342,7 @@ static void igbvf_class_init(ObjectClass *class, const void *data)
c->realize = igbvf_pci_realize;
c->exit = igbvf_pci_uninit;
+ c->config_read = igbvf_read_config;
c->vendor_id = PCI_VENDOR_ID_INTEL;
c->device_id = E1000_DEV_ID_82576_VF;
c->revision = 1;
diff --git a/hw/net/meson.build b/hw/net/meson.build
index 84f142df222a..bb4b449b25ba 100644
--- a/hw/net/meson.build
+++ b/hw/net/meson.build
@@ -11,7 +11,7 @@ system_ss.add(when: 'CONFIG_E1000_PCI', if_true: files('e1000.c', 'e1000x_common
system_ss.add(when: 'CONFIG_E1000E_PCI_EXPRESS', if_true: files('net_tx_pkt.c', 'net_rx_pkt.c'))
system_ss.add(when: 'CONFIG_E1000E_PCI_EXPRESS', if_true: files('e1000e.c', 'e1000e_core.c', 'e1000x_common.c'))
system_ss.add(when: 'CONFIG_IGB_PCI_EXPRESS', if_true: files('net_tx_pkt.c', 'net_rx_pkt.c'))
-system_ss.add(when: 'CONFIG_IGB_PCI_EXPRESS', if_true: files('igb.c', 'igbvf.c', 'igb_core.c'))
+system_ss.add(when: 'CONFIG_IGB_PCI_EXPRESS', if_true: files('igb.c', 'igbvf.c', 'igb_core.c', 'igb_migration.c'))
system_ss.add(when: 'CONFIG_RTL8139_PCI', if_true: files('rtl8139.c'))
system_ss.add(when: 'CONFIG_TULIP', if_true: files('tulip.c'))
system_ss.add(when: 'CONFIG_VMXNET3_PCI', if_true: files('net_tx_pkt.c', 'net_rx_pkt.c'))
--
2.55.0
^ permalink raw reply related [flat|nested] 23+ messages in thread
* [RFC PATCH v2 2/9] igb: Add migration state machine via extended config space
2026-09-02 19:20 [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Cédric Le Goater
2026-09-02 19:20 ` [RFC PATCH v2 1/9] igb: Add x-vf-migration property and DVSEC extended capability Cédric Le Goater
@ 2026-09-02 19:20 ` Cédric Le Goater
2026-09-02 19:20 ` [RFC PATCH v2 3/9] igb: Add VF state serialization for live migration Cédric Le Goater
` (7 subsequent siblings)
9 siblings, 0 replies; 23+ messages in thread
From: Cédric Le Goater @ 2026-09-02 19:20 UTC (permalink / raw)
To: qemu-devel
Cc: Akihiko Odaki, Sriram Yagnaraman, Jason Wang, Alex Williamson,
Peter Xu, Cédric Le Goater
Implement the VFIO migration v2 state machine and command dispatch
through the CTRL doorbell register in the DVSEC capability.
Add three CTRL commands:
- SET_STATE: transitions follow the VFIO migration protocol
RUNNING <-> STOP
STOP <-> STOP_COPY
STOP <-> RESUMING
ERROR -> STOP
- SAVE: DMA-writes the device state to a guest buffer
- LOAD: DMA-reads it back
Add a DATA_SIZE register at +0x20 to report the state blob size. The
save/load stubs are filled in by the next patch.
AI-used-for: code (prototype)
Signed-off-by: Cédric Le Goater <clg@redhat.com>
---
hw/net/igb_migration.h | 25 ++++-
hw/net/igb_migration.c | 217 ++++++++++++++++++++++++++++++++++++++++-
hw/net/trace-events | 6 ++
3 files changed, 246 insertions(+), 2 deletions(-)
diff --git a/hw/net/igb_migration.h b/hw/net/igb_migration.h
index 3da28e11e49e..ea40ac65c54b 100644
--- a/hw/net/igb_migration.h
+++ b/hw/net/igb_migration.h
@@ -26,10 +26,11 @@
* +0x14 STATUS (RO: state[7:0], error_code[15:8])
* +0x18 BUF_ADDR_LO (RW: shared buffer GPA low)
* +0x1C BUF_ADDR_HI (RW: shared buffer GPA high)
+ * +0x20 DATA_SIZE (RO: max state blob size in bytes)
*/
#define IGB_MIG_DVSEC_OFFSET 0x160
-#define IGB_MIG_DVSEC_SIZE 0x20
+#define IGB_MIG_DVSEC_SIZE 0x24
#define IGB_MIG_DVSEC_VER 1
#define IGB_MIG_DVSEC_ID 1
@@ -39,10 +40,20 @@
#define IGB_MIG_STATUS 0x14
#define IGB_MIG_BUF_ADDR_LO 0x18
#define IGB_MIG_BUF_ADDR_HI 0x1C
+#define IGB_MIG_DATA_SIZE 0x20
/* CAPS register layout */
#define IGB_MIG_CAP_F_STATE (1u << 0)
+/* CTRL register: command in [7:0] */
+#define IGB_MIG_CTRL_CMD_MASK 0xFF
+#define IGB_MIG_CTRL_ARG_SHIFT 8
+
+/* CTRL commands */
+#define IGB_MIG_CMD_SET_STATE 1
+#define IGB_MIG_CMD_SAVE 2
+#define IGB_MIG_CMD_LOAD 3
+
/* STATUS register: state in [7:0], error code [15:8] */
#define IGB_MIG_STATUS_STATE_MASK 0xFF
#define IGB_MIG_STATUS_ERROR_CODE_SHIFT 8
@@ -56,8 +67,20 @@
#define IGB_MIG_STATE_STOP_COPY 3
#define IGB_MIG_STATE_RESUMING 4
+/* Error codes */
+#define IGB_MIG_ERR_UNK_CMD 1
+#define IGB_MIG_ERR_BAD_STATE 2
+#define IGB_MIG_ERR_NO_BUFFER 3
+#define IGB_MIG_ERR_DMA_FAILED 4
+#define IGB_MIG_ERR_BAD_SIZE 5
+
+/* Shared buffer constants */
+#define IGB_VF_STATE_MAX_SIZE 4096
+
typedef struct IgbVfMigState {
uint32_t mig_state;
+ uint32_t mig_data[IGB_VF_STATE_MAX_SIZE / sizeof(uint32_t)];
+ uint32_t mig_data_size;
uint64_t mig_data_buf_addr;
} IgbVfMigState;
diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
index 4dfebd82344c..8e7e6fac9b9b 100644
--- a/hw/net/igb_migration.c
+++ b/hw/net/igb_migration.c
@@ -7,10 +7,177 @@
*/
#include "qemu/osdep.h"
+#include "qemu/log.h"
#include "hw/pci/pci_device.h"
#include "hw/pci/pcie.h"
#include "igb_common.h"
#include "igb_migration.h"
+#include "system/address-spaces.h"
+#include "trace.h"
+
+/*
+ * Per-VF state serialization / deserialization
+ */
+
+static int igb_core_vf_save_state(IgbVfState *s, void *buf, size_t buf_size)
+{
+ int size = 0;
+
+ trace_igbvf_mig_save_state(s->vfn, size);
+ return size;
+}
+
+static int igb_core_vf_max_data_size(IgbVfState *s)
+{
+ return sizeof(s->mig.mig_data);
+}
+
+static int igb_core_vf_load_state(IgbVfState *s, const void *buf, size_t size)
+{
+ trace_igbvf_mig_load_state(s->vfn, (uint32_t)size);
+ return 0;
+}
+
+static int igbvf_mig_load(IgbVfState *s, const void *buf, size_t size)
+{
+ int ret;
+
+ ret = igb_core_vf_load_state(s, buf, size);
+ if (ret < 0) {
+ return ret;
+ }
+
+ return 0;
+}
+
+/*
+ * Migration command handlers
+ */
+
+static void igbvf_mig_update_data_size(IgbVfState *s, uint32_t size)
+{
+ IgbVfMigState *ms = &s->mig;
+
+ ms->mig_data_size = size;
+ pci_set_long(PCI_DEVICE(s)->config +
+ IGB_MIG_DVSEC_OFFSET + IGB_MIG_DATA_SIZE, size);
+}
+
+static uint8_t igbvf_mig_cmd_save(IgbVfState *s)
+{
+ IgbVfMigState *ms = &s->mig;
+ MemTxResult r;
+ int ret;
+
+ if (ms->mig_state != IGB_MIG_STATE_STOP_COPY) {
+ return IGB_MIG_ERR_BAD_STATE;
+ }
+
+ if (!ms->mig_data_buf_addr) {
+ return IGB_MIG_ERR_NO_BUFFER;
+ }
+
+ ret = igb_core_vf_save_state(s, ms->mig_data, sizeof(ms->mig_data));
+ if (ret < 0) {
+ return -ret;
+ }
+ igbvf_mig_update_data_size(s, ret);
+
+ r = address_space_write(&address_space_memory, ms->mig_data_buf_addr,
+ MEMTXATTRS_UNSPECIFIED,
+ ms->mig_data, ms->mig_data_size);
+ if (r != MEMTX_OK) {
+ return IGB_MIG_ERR_DMA_FAILED;
+ }
+
+ return 0;
+}
+
+static uint8_t igbvf_mig_cmd_load(IgbVfState *s)
+{
+ IgbVfMigState *ms = &s->mig;
+ MemTxResult r;
+ int ret;
+
+ if (ms->mig_state != IGB_MIG_STATE_RESUMING) {
+ return IGB_MIG_ERR_BAD_STATE;
+ }
+
+ if (!ms->mig_data_buf_addr) {
+ return IGB_MIG_ERR_NO_BUFFER;
+ }
+
+ if (ms->mig_data_size == 0 ||
+ ms->mig_data_size > sizeof(ms->mig_data)) {
+ return IGB_MIG_ERR_BAD_SIZE;
+ }
+
+ r = address_space_read(&address_space_memory, ms->mig_data_buf_addr,
+ MEMTXATTRS_UNSPECIFIED,
+ ms->mig_data, ms->mig_data_size);
+ if (r != MEMTX_OK) {
+ return IGB_MIG_ERR_DMA_FAILED;
+ }
+
+ ret = igbvf_mig_load(s, ms->mig_data, ms->mig_data_size);
+ if (ret < 0) {
+ return -ret;
+ }
+
+ return 0;
+}
+
+static uint8_t igbvf_mig_set_state(IgbVfState *s, uint32_t new_state)
+{
+ IgbVfMigState *ms = &s->mig;
+ uint32_t old = ms->mig_state;
+ int ret;
+
+ switch (new_state) {
+ case IGB_MIG_STATE_STOP:
+ if (old != IGB_MIG_STATE_RUNNING &&
+ old != IGB_MIG_STATE_STOP_COPY &&
+ old != IGB_MIG_STATE_RESUMING &&
+ old != IGB_MIG_STATE_ERROR) {
+ return IGB_MIG_ERR_BAD_STATE;
+ }
+ /* Restore DATA_SIZE to max, same as at reset */
+ igbvf_mig_update_data_size(s, igb_core_vf_max_data_size(s));
+ break;
+
+ case IGB_MIG_STATE_RUNNING:
+ if (old != IGB_MIG_STATE_STOP) {
+ return IGB_MIG_ERR_BAD_STATE;
+ }
+ break;
+
+ case IGB_MIG_STATE_STOP_COPY:
+ if (old != IGB_MIG_STATE_STOP) {
+ return IGB_MIG_ERR_BAD_STATE;
+ }
+ ret = igb_core_vf_save_state(s, ms->mig_data, sizeof(ms->mig_data));
+ if (ret < 0) {
+ return -ret;
+ }
+ igbvf_mig_update_data_size(s, ret);
+ break;
+
+ case IGB_MIG_STATE_RESUMING:
+ if (old != IGB_MIG_STATE_STOP) {
+ return IGB_MIG_ERR_BAD_STATE;
+ }
+ memset(ms->mig_data, 0, sizeof(ms->mig_data));
+ igbvf_mig_update_data_size(s, 0);
+ break;
+
+ default:
+ return IGB_MIG_ERR_BAD_STATE;
+ }
+
+ ms->mig_state = new_state;
+ trace_igbvf_mig_set_state(s->vfn, old, new_state);
+ return 0;
+}
static void igbvf_mig_update_status(IgbVfState *s, uint8_t err)
{
@@ -27,6 +194,38 @@ static void igbvf_mig_update_status(IgbVfState *s, uint8_t err)
pci_set_long(dev->config + IGB_MIG_DVSEC_OFFSET + IGB_MIG_STATUS, status);
}
+static void igbvf_mig_cmd_ctrl(IgbVfState *s, uint32_t val)
+{
+ uint32_t cmd = val & IGB_MIG_CTRL_CMD_MASK;
+ uint32_t arg = val >> IGB_MIG_CTRL_ARG_SHIFT;
+ uint8_t err = 0;
+
+ switch (cmd) {
+ case IGB_MIG_CMD_SET_STATE:
+ err = igbvf_mig_set_state(s, arg);
+ break;
+
+ case IGB_MIG_CMD_SAVE:
+ err = igbvf_mig_cmd_save(s);
+ break;
+
+ case IGB_MIG_CMD_LOAD:
+ igbvf_mig_update_data_size(s, arg);
+ err = igbvf_mig_cmd_load(s);
+ break;
+
+ default:
+ err = IGB_MIG_ERR_UNK_CMD;
+ break;
+ }
+
+ if (err) {
+ qemu_log_mask(LOG_GUEST_ERROR,
+ "igbvf: VF%u CTRL cmd %u failed (error %u)\n",
+ s->vfn, cmd, err);
+ }
+ igbvf_mig_update_status(s, err);
+}
bool igbvf_add_migration_dvsec(PCIDevice *dev, Error **errp)
{
@@ -57,6 +256,8 @@ bool igbvf_add_migration_dvsec(PCIDevice *dev, Error **errp)
memset(dev->wmask + offset + IGB_MIG_BUF_ADDR_LO, 0xff, 4);
memset(dev->wmask + offset + IGB_MIG_BUF_ADDR_HI, 0xff, 4);
+ /* DATA_SIZE is set by igbvf_mig_state_reset() */
+
return true;
}
@@ -67,6 +268,16 @@ uint32_t igbvf_mig_config_read(IgbVfState *s, uint32_t addr, int size)
return pci_default_read_config(dev, addr, size);
}
+static uint64_t igbvf_mig_get_buf_addr(IgbVfState *s)
+{
+ PCIDevice *dev = PCI_DEVICE(s);
+ uint32_t lo, hi;
+
+ lo = pci_get_long(dev->config + IGB_MIG_DVSEC_OFFSET + IGB_MIG_BUF_ADDR_LO);
+ hi = pci_get_long(dev->config + IGB_MIG_DVSEC_OFFSET + IGB_MIG_BUF_ADDR_HI);
+ return ((uint64_t)hi << 32) | lo;
+}
+
bool igbvf_mig_config_write(IgbVfState *s, uint32_t addr, uint32_t val,
int size)
{
@@ -75,7 +286,8 @@ bool igbvf_mig_config_write(IgbVfState *s, uint32_t addr, uint32_t val,
switch (offset) {
case IGB_MIG_CTRL:
- /* Command handling will be added in a later commit */
+ s->mig.mig_data_buf_addr = igbvf_mig_get_buf_addr(s);
+ igbvf_mig_cmd_ctrl(s, val);
break;
case IGB_MIG_BUF_ADDR_LO:
@@ -94,8 +306,11 @@ void igbvf_mig_state_reset(IgbVfState *s)
{
IgbVfMigState *ms = &s->mig;
+ trace_igbvf_mig_reset(s->vfn);
ms->mig_state = IGB_MIG_STATE_RUNNING;
ms->mig_data_buf_addr = 0;
+ igbvf_mig_update_data_size(s, igb_core_vf_max_data_size(s));
+ memset(ms->mig_data, 0, sizeof(ms->mig_data));
pci_set_long(PCI_DEVICE(s)->config +
IGB_MIG_DVSEC_OFFSET + IGB_MIG_BUF_ADDR_LO, 0);
diff --git a/hw/net/trace-events b/hw/net/trace-events
index 001a20b0e2ac..7057fbe5f16a 100644
--- a/hw/net/trace-events
+++ b/hw/net/trace-events
@@ -295,6 +295,12 @@ igb_wrn_rx_desc_modes_not_supp(int desc_type) "Not supported descriptor type: %d
# igbvf.c
igbvf_wrn_io_addr_unknown(uint64_t addr) "IO unknown register 0x%"PRIx64
+# igb_migration.c
+igbvf_mig_set_state(uint16_t vfn, uint32_t old_state, uint32_t new_state) "VF%u: state %u -> %u"
+igbvf_mig_save_state(uint16_t vfn, uint32_t size) "VF%u: saved %u bytes of device state"
+igbvf_mig_load_state(uint16_t vfn, uint32_t size) "VF%u: loaded %u bytes of device state"
+igbvf_mig_reset(uint16_t vfn) "VF%u: migration state reset"
+
# spapr_llan.c
spapr_vlan_get_rx_bd_from_pool_found(int pool, int32_t count, uint32_t rx_bufs) "pool=%d count=%"PRId32" rxbufs=%"PRIu32
spapr_vlan_get_rx_bd_from_page(int buf_ptr, uint64_t bd) "use_buf_ptr=%d bd=0x%016"PRIx64
--
2.55.0
^ permalink raw reply related [flat|nested] 23+ messages in thread
* [RFC PATCH v2 3/9] igb: Add VF state serialization for live migration
2026-09-02 19:20 [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Cédric Le Goater
2026-09-02 19:20 ` [RFC PATCH v2 1/9] igb: Add x-vf-migration property and DVSEC extended capability Cédric Le Goater
2026-09-02 19:20 ` [RFC PATCH v2 2/9] igb: Add migration state machine via extended config space Cédric Le Goater
@ 2026-09-02 19:20 ` Cédric Le Goater
2026-09-08 8:01 ` Akihiko Odaki
2026-09-02 19:20 ` [RFC PATCH v2 4/9] igb: Add VF post-load fixups " Cédric Le Goater
` (6 subsequent siblings)
9 siblings, 1 reply; 23+ messages in thread
From: Cédric Le Goater @ 2026-09-02 19:20 UTC (permalink / raw)
To: qemu-devel
Cc: Akihiko Odaki, Sriram Yagnaraman, Jason Wang, Alex Williamson,
Peter Xu, Cédric Le Goater
Implement per-VF state serialization and deserialization for the SAVE
and LOAD commands. The wire format consists of a header (magic,
version, VF number, register count), per-VF register offset/value
pairs from a whitelist, RA table entries owned by the VF, and TX
queue contexts.
Register offsets are relocated on load so a VF can migrate to a
different VF number on the destination. RA pool ownership bits are
swapped accordingly.
AI-used-for: analysis, code (prototype)
Signed-off-by: Cédric Le Goater <clg@redhat.com>
---
hw/net/igb_core.h | 2 +
hw/net/igb_migration.h | 2 +
hw/net/igb.c | 5 +
hw/net/igb_migration.c | 347 ++++++++++++++++++++++++++++++++++++++++-
4 files changed, 354 insertions(+), 2 deletions(-)
diff --git a/hw/net/igb_core.h b/hw/net/igb_core.h
index d70b54e318f1..60724e2824ab 100644
--- a/hw/net/igb_core.h
+++ b/hw/net/igb_core.h
@@ -143,4 +143,6 @@ igb_receive_iov(IGBCore *core, const struct iovec *iov, int iovcnt);
void
igb_start_recv(IGBCore *core);
+IGBCore *igb_pf_get_core(void *pf);
+
#endif
diff --git a/hw/net/igb_migration.h b/hw/net/igb_migration.h
index ea40ac65c54b..b2f601e74346 100644
--- a/hw/net/igb_migration.h
+++ b/hw/net/igb_migration.h
@@ -73,6 +73,8 @@
#define IGB_MIG_ERR_NO_BUFFER 3
#define IGB_MIG_ERR_DMA_FAILED 4
#define IGB_MIG_ERR_BAD_SIZE 5
+#define IGB_MIG_ERR_BAD_MAGIC 6
+#define IGB_MIG_ERR_BAD_VERSION 7
/* Shared buffer constants */
#define IGB_VF_STATE_MAX_SIZE 4096
diff --git a/hw/net/igb.c b/hw/net/igb.c
index 7268e5473fc3..f39f2bc3a04e 100644
--- a/hw/net/igb.c
+++ b/hw/net/igb.c
@@ -133,6 +133,11 @@ void igb_vf_reset(void *opaque, uint16_t vfn)
igb_core_vf_reset(&s->core, vfn);
}
+IGBCore *igb_pf_get_core(void *pf)
+{
+ return &IGB(pf)->core;
+}
+
static bool
igb_io_get_reg_index(IGBState *s, uint32_t *idx)
{
diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
index 8e7e6fac9b9b..c34035974620 100644
--- a/hw/net/igb_migration.c
+++ b/hw/net/igb_migration.c
@@ -10,18 +10,241 @@
#include "qemu/log.h"
#include "hw/pci/pci_device.h"
#include "hw/pci/pcie.h"
+#include "net/eth.h"
+#include "net/net.h"
#include "igb_common.h"
+#include "igb_core.h"
#include "igb_migration.h"
#include "system/address-spaces.h"
#include "trace.h"
+static IGBCore *igbvf_get_core(IgbVfState *s)
+{
+ return igb_pf_get_core(pcie_sriov_get_pf(PCI_DEVICE(s)));
+}
+
/*
* Per-VF state serialization / deserialization
*/
+#define IGB_MIG_BLOB_MAGIC 0x4D494742 /* "MIGB" */
+#define IGB_MIG_BLOB_VERSION 1
+
+typedef struct IgbMigRegPair {
+ uint32_t offset;
+ uint32_t value;
+} IgbMigRegPair;
+
+typedef struct IgbMigTxCtx {
+ uint32_t ctx_desc[8]; /* 2 × adv_tx_context_desc (4 dwords each) */
+ uint32_t first_cmd_type_len;
+ uint32_t first_olinfo_status;
+ uint32_t first;
+ uint32_t skip_cp;
+} IgbMigTxCtx;
+
+#define IGB_VF_MAX_FIXED_REGS 64
+#define IGB_VF_MAX_RA_REGS 48 /* (16 + 8) RA entries × 2 (RAL+RAH) */
+
+typedef struct IgbMigBlob {
+ uint32_t magic;
+ uint32_t version;
+ uint32_t vfn;
+ uint32_t num_regs;
+ IgbMigRegPair regs[IGB_VF_MAX_FIXED_REGS];
+ uint32_t num_ra;
+ IgbMigRegPair ra[IGB_VF_MAX_RA_REGS];
+ uint32_t num_tx_ctx;
+ IgbMigTxCtx tx_ctx[2];
+} IgbMigBlob;
+
+#define IGB_MIG_BLOB_SIZE sizeof(IgbMigBlob)
+
+QEMU_BUILD_BUG_ON(IGB_MIG_BLOB_SIZE > IGB_VF_STATE_MAX_SIZE);
+
+/* Register offsets that constitute a VF's state slice */
+static int igb_vf_reg_list(uint16_t vfn, uint32_t *offsets)
+{
+ int n = 0;
+ int q0 = vfn;
+ int q1 = vfn + IGB_NUM_VM_POOLS;
+
+ /* Per-VF control and interrupt registers */
+ offsets[n++] = E1000_PVTCTRL(vfn) >> 2;
+ offsets[n++] = E1000_PVTEICS(vfn) >> 2;
+ offsets[n++] = E1000_PVTEIMS(vfn) >> 2;
+ offsets[n++] = E1000_PVTEIMC(vfn) >> 2;
+ offsets[n++] = E1000_PVTEIAC(vfn) >> 2;
+ offsets[n++] = E1000_PVTEIAM(vfn) >> 2;
+ offsets[n++] = E1000_PVTEICR(vfn) >> 2;
+
+ /* Per-VF statistics */
+ offsets[n++] = E1000_PVFGPRC(vfn) >> 2;
+ offsets[n++] = E1000_PVFGPTC(vfn) >> 2;
+ offsets[n++] = E1000_PVFGORC(vfn) >> 2;
+ offsets[n++] = E1000_PVFGOTC(vfn) >> 2;
+ offsets[n++] = E1000_PVFMPRC(vfn) >> 2;
+ offsets[n++] = E1000_PVFGPRLBC(vfn) >> 2;
+ offsets[n++] = E1000_PVFGPTLBC(vfn) >> 2;
+ offsets[n++] = E1000_PVFGORLBC(vfn) >> 2;
+ offsets[n++] = E1000_PVFGOTLBC(vfn) >> 2;
+
+ /*
+ * Mailbox control registers only - the 16-dword payload buffer
+ * (VMBMEM) is transient and drained on quiesce.
+ */
+ offsets[n++] = E1000_V2PMAILBOX(vfn) >> 2;
+ offsets[n++] = E1000_P2VMAILBOX(vfn) >> 2;
+
+ /* Per-VF config */
+ offsets[n++] = E1000_VMOLR(vfn) >> 2;
+ offsets[n++] = E1000_VMVIR(vfn) >> 2;
+ offsets[n++] = E1000_PSRTYPE(vfn) >> 2;
+
+ /*
+ * VF receive addresses (RA/RA2) are saved dynamically in
+ * igb_core_vf_save_state by scanning for entries whose pool
+ * bits match this VF - the PF driver chooses the RA slot.
+ */
+
+ /* Interrupt routing */
+ offsets[n++] = (E1000_VTIVAR + vfn * 4) >> 2;
+ offsets[n++] = (E1000_VTIVAR_MISC + vfn * 4) >> 2;
+
+ /*
+ * EITR (Extended Interrupt Throttle Register) - 3 vectors per VF.
+ * Each VF has 3 MSI-X vectors, each with its own EITR controlling
+ * interrupt coalescing. Without saving these, interrupt
+ * throttling resets to zero after migration which can cause
+ * interrupt storms or latency changes. VF N uses PF EITR indices
+ * (22 - N*3) .. (24 - N*3).
+ */
+ {
+ int eitr_base = 22 - vfn * 3;
+ offsets[n++] = E1000_EITR(eitr_base) >> 2;
+ offsets[n++] = E1000_EITR(eitr_base + 1) >> 2;
+ offsets[n++] = E1000_EITR(eitr_base + 2) >> 2;
+ }
+
+ /* RX and TX queue registers for queues q0 and q1 */
+#define ADD_QUEUE_REGS(q) do { \
+ offsets[n++] = E1000_RDBAL(q) >> 2; \
+ offsets[n++] = E1000_RDBAH(q) >> 2; \
+ offsets[n++] = E1000_RDLEN(q) >> 2; \
+ offsets[n++] = E1000_SRRCTL(q) >> 2; \
+ offsets[n++] = E1000_RDH(q) >> 2; \
+ offsets[n++] = E1000_RDT(q) >> 2; \
+ offsets[n++] = E1000_RXDCTL(q) >> 2; \
+ offsets[n++] = E1000_RXCTL(q) >> 2; \
+ offsets[n++] = E1000_RQDPC(q) >> 2; \
+ offsets[n++] = E1000_TDBAL(q) >> 2; \
+ offsets[n++] = E1000_TDBAH(q) >> 2; \
+ offsets[n++] = E1000_TDLEN(q) >> 2; \
+ offsets[n++] = E1000_TDH(q) >> 2; \
+ offsets[n++] = E1000_TDT(q) >> 2; \
+ offsets[n++] = E1000_TXDCTL(q) >> 2; \
+ offsets[n++] = E1000_TXCTL(q) >> 2; \
+ offsets[n++] = E1000_TDWBAL(q) >> 2; \
+ offsets[n++] = E1000_TDWBAH(q) >> 2; \
+} while (0)
+
+ ADD_QUEUE_REGS(q0);
+ ADD_QUEUE_REGS(q1);
+#undef ADD_QUEUE_REGS
+
+ g_assert(n <= IGB_VF_MAX_FIXED_REGS);
+ return n;
+}
+
+/*
+ * Scan RA and RA2 arrays for receive address entries assigned to
+ * this VF. The PF driver picks the RA slot, so we cannot use a
+ * fixed index - instead check each entry's pool bits.
+ */
+static int igb_core_vf_save_ra(IGBCore *core, uint16_t vfn,
+ IgbMigRegPair *regs)
+{
+ uint32_t vf_pool_bit = E1000_RAH_POOL_1 << vfn;
+ int n = 0;
+ static const struct {
+ uint32_t base;
+ int count;
+ } ra_banks[] = {
+ { RA, 16 },
+ { RA2, 8 },
+ };
+
+ for (int i = 0; i < ARRAY_SIZE(ra_banks); i++) {
+ for (int j = 0; j < ra_banks[i].count; j++) {
+ uint32_t ral_off = ra_banks[i].base + j * 2;
+ uint32_t rah_off = ra_banks[i].base + j * 2 + 1;
+ uint32_t rah_val = core->mac[rah_off];
+
+ if ((rah_val & E1000_RAH_AV) && (rah_val & vf_pool_bit)) {
+ regs[n].offset = cpu_to_le32(ral_off);
+ regs[n].value = cpu_to_le32(core->mac[ral_off]);
+ n++;
+ regs[n].offset = cpu_to_le32(rah_off);
+ regs[n].value = cpu_to_le32(rah_val);
+ n++;
+ }
+ }
+ }
+ return n;
+}
+
+static void igb_core_vf_save_tx_ctx(IGBCore *core, int queue,
+ IgbMigTxCtx *tx)
+{
+ struct igb_tx *src = &core->tx[queue];
+
+ memcpy(tx->ctx_desc, src->ctx, sizeof(tx->ctx_desc));
+ tx->first_cmd_type_len = cpu_to_le32(src->first_cmd_type_len);
+ tx->first_olinfo_status = cpu_to_le32(src->first_olinfo_status);
+ tx->first = cpu_to_le32(src->first);
+ tx->skip_cp = cpu_to_le32(src->skip_cp);
+}
+
static int igb_core_vf_save_state(IgbVfState *s, void *buf, size_t buf_size)
{
- int size = 0;
+ int size = IGB_MIG_BLOB_SIZE;
+ IGBCore *core = igbvf_get_core(s);
+ IgbMigBlob *blob = buf;
+ uint32_t offsets[IGB_VF_MAX_FIXED_REGS];
+ int num_regs;
+ int q0 = s->vfn;
+ int q1 = s->vfn + IGB_NUM_VM_POOLS;
+
+ /*
+ * Save PVT shadow registers (PVTEIMS/PVTEIAC/PVTEIAM) instead of
+ * extracting from PF aggregates - the L1 PF driver may have
+ * transiently cleared EIMS via EIMC. The load path ORs them back.
+ */
+ num_regs = igb_vf_reg_list(s->vfn, offsets);
+
+ if (!buf) {
+ return size;
+ }
+
+ if (size > buf_size) {
+ return -IGB_MIG_ERR_BAD_SIZE;
+ }
+
+ blob->magic = cpu_to_le32(IGB_MIG_BLOB_MAGIC);
+ blob->version = cpu_to_le32(IGB_MIG_BLOB_VERSION);
+ blob->vfn = cpu_to_le32(s->vfn);
+
+ blob->num_regs = cpu_to_le32(num_regs);
+ for (int i = 0; i < num_regs; i++) {
+ blob->regs[i].offset = cpu_to_le32(offsets[i]);
+ blob->regs[i].value = cpu_to_le32(core->mac[offsets[i]]);
+ }
+
+ blob->num_ra = cpu_to_le32(igb_core_vf_save_ra(core, s->vfn, blob->ra));
+
+ blob->num_tx_ctx = cpu_to_le32(2);
+ igb_core_vf_save_tx_ctx(core, q0, &blob->tx_ctx[0]);
+ igb_core_vf_save_tx_ctx(core, q1, &blob->tx_ctx[1]);
trace_igbvf_mig_save_state(s->vfn, size);
return size;
@@ -29,11 +252,131 @@ static int igb_core_vf_save_state(IgbVfState *s, void *buf, size_t buf_size)
static int igb_core_vf_max_data_size(IgbVfState *s)
{
- return sizeof(s->mig.mig_data);
+ int size = igb_core_vf_save_state(s, NULL, 0);
+
+ g_assert(size > 0 && size <= IGB_VF_STATE_MAX_SIZE);
+ return size;
+}
+
+static void igb_core_vf_load_tx_ctx(IGBCore *core, int queue,
+ const IgbMigTxCtx *tx)
+{
+ struct igb_tx *dst = &core->tx[queue];
+
+ /*
+ * Preserve the destination's tx_pkt - it's a host-side object,
+ * not guest state
+ */
+ memcpy(dst->ctx, tx->ctx_desc, sizeof(dst->ctx));
+ dst->first_cmd_type_len = le32_to_cpu(tx->first_cmd_type_len);
+ dst->first_olinfo_status = le32_to_cpu(tx->first_olinfo_status);
+ dst->first = le32_to_cpu(tx->first);
+ dst->skip_cp = le32_to_cpu(tx->skip_cp);
+}
+
+static uint32_t igb_vf_relocate_offset(uint32_t offset,
+ const uint32_t *src_offsets,
+ const uint32_t *dst_offsets,
+ int num_offsets)
+{
+ for (int i = 0; i < num_offsets; i++) {
+ if (src_offsets[i] == offset) {
+ return dst_offsets[i];
+ }
+ }
+ return 0;
}
static int igb_core_vf_load_state(IgbVfState *s, const void *buf, size_t size)
{
+ IGBCore *core = igbvf_get_core(s);
+ uint32_t src_offsets[IGB_VF_MAX_FIXED_REGS];
+ uint32_t dst_offsets[IGB_VF_MAX_FIXED_REGS];
+ int q0 = s->vfn;
+ int q1 = s->vfn + IGB_NUM_VM_POOLS;
+
+ if (size < IGB_MIG_BLOB_SIZE) {
+ return -IGB_MIG_ERR_BAD_SIZE;
+ }
+
+ const IgbMigBlob *blob = buf;
+
+ uint32_t magic = le32_to_cpu(blob->magic);
+ uint32_t version = le32_to_cpu(blob->version);
+ uint32_t saved_vfn = le32_to_cpu(blob->vfn);
+ uint32_t num_regs = le32_to_cpu(blob->num_regs);
+
+ if (magic != IGB_MIG_BLOB_MAGIC) {
+ return -IGB_MIG_ERR_BAD_MAGIC;
+ }
+ if (version != IGB_MIG_BLOB_VERSION) {
+ return -IGB_MIG_ERR_BAD_VERSION;
+ }
+ if (num_regs > IGB_VF_MAX_FIXED_REGS) {
+ return -IGB_MIG_ERR_BAD_SIZE;
+ }
+
+ uint32_t num_ra = le32_to_cpu(blob->num_ra);
+ if (num_ra > IGB_VF_MAX_RA_REGS) {
+ return -IGB_MIG_ERR_BAD_SIZE;
+ }
+
+ int num_offsets = igb_vf_reg_list(saved_vfn, src_offsets);
+ igb_vf_reg_list(s->vfn, dst_offsets);
+
+ for (uint32_t i = 0; i < num_regs; i++) {
+ uint32_t src_off = le32_to_cpu(blob->regs[i].offset);
+ uint32_t value = le32_to_cpu(blob->regs[i].value);
+ uint32_t offset = igb_vf_relocate_offset(src_off,
+ src_offsets, dst_offsets, num_offsets);
+ if (!offset) {
+ return -IGB_MIG_ERR_BAD_SIZE;
+ }
+
+ core->mac[offset] = value;
+
+ /*
+ * Sync EITR to eitr_guest_value[] shadow array, stripping
+ * E1000_EITR_CNT_IGNR so guest register readback returns the
+ * correct value.
+ */
+ if (offset >= EITR0 && offset < EITR0 + IGB_INTR_NUM) {
+ core->eitr_guest_value[offset - EITR0] =
+ value & ~E1000_EITR_CNT_IGNR;
+ }
+ }
+
+ /*
+ * MSI-X table/PBA is not saved - L1's VFIO reprograms it with
+ * destination-specific IRTE references after migration.
+ */
+
+ uint32_t src_pool = E1000_RAH_POOL_1 << saved_vfn;
+ uint32_t dst_pool = E1000_RAH_POOL_1 << s->vfn;
+
+ for (uint32_t i = 0; i < num_ra; i++) {
+ uint32_t offset = le32_to_cpu(blob->ra[i].offset);
+ uint32_t value = le32_to_cpu(blob->ra[i].value);
+
+ /* RAH entries: swap pool ownership bits */
+ if (offset >= RA && offset < RA + 32 && (offset - RA) % 2 == 1) {
+ value = (value & ~src_pool) | dst_pool;
+ }
+ if (offset >= RA2 && offset < RA2 + 16 && (offset - RA2) % 2 == 1) {
+ value = (value & ~src_pool) | dst_pool;
+ }
+
+ core->mac[offset] = value;
+ }
+
+ uint32_t num_tx = le32_to_cpu(blob->num_tx_ctx);
+ if (num_tx != 2) {
+ return -IGB_MIG_ERR_BAD_SIZE;
+ }
+
+ igb_core_vf_load_tx_ctx(core, q0, &blob->tx_ctx[0]);
+ igb_core_vf_load_tx_ctx(core, q1, &blob->tx_ctx[1]);
+
trace_igbvf_mig_load_state(s->vfn, (uint32_t)size);
return 0;
}
--
2.55.0
^ permalink raw reply related [flat|nested] 23+ messages in thread
* [RFC PATCH v2 4/9] igb: Add VF post-load fixups for live migration
2026-09-02 19:20 [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Cédric Le Goater
` (2 preceding siblings ...)
2026-09-02 19:20 ` [RFC PATCH v2 3/9] igb: Add VF state serialization for live migration Cédric Le Goater
@ 2026-09-02 19:20 ` Cédric Le Goater
2026-09-08 8:10 ` Akihiko Odaki
2026-09-02 19:20 ` [RFC PATCH v2 5/9] igb: Add dirty page tracking for IGBVF migration Cédric Le Goater
` (5 subsequent siblings)
9 siblings, 1 reply; 23+ messages in thread
From: Cédric Le Goater @ 2026-09-02 19:20 UTC (permalink / raw)
To: qemu-devel
Cc: Akihiko Odaki, Sriram Yagnaraman, Jason Wang, Alex Williamson,
Peter Xu, Cédric Le Goater
After restoring per-VF register state, propagate the VF's PVT shadow
values back into the PF's aggregate EIMS/EIAC/EIAM registers and
re-apply the VTIVAR interrupt vector routing to the shared IVAR0.
AI-used-for: analysis, code (prototype)
Signed-off-by: Cédric Le Goater <clg@redhat.com>
---
hw/net/igb_core.h | 3 ++
hw/net/igb_core.c | 66 ++++++++++++++++++++++++++++++++++++++++++
hw/net/igb_migration.c | 7 +++++
3 files changed, 76 insertions(+)
diff --git a/hw/net/igb_core.h b/hw/net/igb_core.h
index 60724e2824ab..22e10e4e6d0b 100644
--- a/hw/net/igb_core.h
+++ b/hw/net/igb_core.h
@@ -145,4 +145,7 @@ igb_start_recv(IGBCore *core);
IGBCore *igb_pf_get_core(void *pf);
+void igb_core_vf_propagate_irqs(IGBCore *core, uint16_t vfn);
+void igb_core_vf_propagate_ivar(IGBCore *core, uint16_t vfn);
+
#endif
diff --git a/hw/net/igb_core.c b/hw/net/igb_core.c
index 2a4883907353..01745fe756d0 100644
--- a/hw/net/igb_core.c
+++ b/hw/net/igb_core.c
@@ -4552,3 +4552,69 @@ igb_core_post_load(IGBCore *core)
return 0;
}
+
+/*
+ * Propagate VF interrupt state to PF aggregates after loading VF
+ * registers. The load path writes directly to mac[] bypassing the
+ * register handlers that OR VF bits into EIMS/EIAC/EIAM. Also clear
+ * stale VF bits in EICR that may have been set by packets arriving
+ * between PF vmstate restore and VF state load.
+ */
+void igb_core_vf_propagate_irqs(IGBCore *core, uint16_t vfn)
+{
+ uint32_t shift = 22 - vfn * IGBVF_MSIX_VEC_NUM;
+ uint32_t vf_mask = 0x7 << shift;
+ uint32_t pvt_idx;
+
+ core->mac[EIMS] &= ~vf_mask;
+ pvt_idx = PVTEIMS0 + vfn * 0x40;
+ core->mac[EIMS] |= (core->mac[pvt_idx] & 0x7) << shift;
+
+ core->mac[EIAC] &= ~vf_mask;
+ pvt_idx = PVTEIAC0 + vfn * 0x40;
+ core->mac[EIAC] |= (core->mac[pvt_idx] & 0x7) << shift;
+
+ core->mac[EIAM] &= ~vf_mask;
+ pvt_idx = PVTEIAM0 + vfn * 0x40;
+ core->mac[EIAM] |= (core->mac[pvt_idx] & 0x7) << shift;
+
+ core->mac[EICR] &= ~vf_mask;
+}
+
+/*
+ * Re-apply VTIVAR -> IVAR0 interrupt routing. The L1 PF driver
+ * may have overwritten the shared IVAR0 entries with its own
+ * queue routing after L0 vmstate restore.
+ */
+void igb_core_vf_propagate_ivar(IGBCore *core, uint16_t vfn)
+{
+ uint32_t vtivar = core->mac[VTIVAR + vfn];
+ int n;
+ uint8_t ent;
+ uint32_t mask;
+
+ n = igb_ivar_entry_rx(vfn);
+ mask = 0xffU << (8 * (n % 4));
+ if (vtivar & E1000_IVAR_VALID) {
+ ent = E1000_IVAR_VALID |
+ (24 - vfn * IGBVF_MSIX_VEC_NUM - (2 - (vtivar & 0x7)));
+ core->mac[IVAR0 + n / 4] =
+ (core->mac[IVAR0 + n / 4] & ~mask) |
+ ((uint32_t)ent << (8 * (n % 4)));
+ } else {
+ core->mac[IVAR0 + n / 4] &= ~mask;
+ }
+
+ n = igb_ivar_entry_tx(vfn);
+ mask = 0xffU << (8 * (n % 4));
+ ent = vtivar >> 8;
+ if (ent & E1000_IVAR_VALID) {
+ ent = E1000_IVAR_VALID |
+ (24 - vfn * IGBVF_MSIX_VEC_NUM - (2 - (ent & 0x7)));
+ core->mac[IVAR0 + n / 4] =
+ (core->mac[IVAR0 + n / 4] & ~mask) |
+ ((uint32_t)ent << (8 * (n % 4)));
+ } else {
+ core->mac[IVAR0 + n / 4] &= ~mask;
+ }
+}
diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
index c34035974620..af7257ccc4fd 100644
--- a/hw/net/igb_migration.c
+++ b/hw/net/igb_migration.c
@@ -384,12 +384,19 @@ static int igb_core_vf_load_state(IgbVfState *s, const void *buf, size_t size)
static int igbvf_mig_load(IgbVfState *s, const void *buf, size_t size)
{
int ret;
+ IGBCore *core = igbvf_get_core(s);
ret = igb_core_vf_load_state(s, buf, size);
if (ret < 0) {
return ret;
}
+ /*
+ * Post-load: sync VF interrupt and routing state to PF aggregates
+ */
+ igb_core_vf_propagate_irqs(core, s->vfn);
+ igb_core_vf_propagate_ivar(core, s->vfn);
+
return 0;
}
--
2.55.0
^ permalink raw reply related [flat|nested] 23+ messages in thread
* [RFC PATCH v2 5/9] igb: Add dirty page tracking for IGBVF migration
2026-09-02 19:20 [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Cédric Le Goater
` (3 preceding siblings ...)
2026-09-02 19:20 ` [RFC PATCH v2 4/9] igb: Add VF post-load fixups " Cédric Le Goater
@ 2026-09-02 19:20 ` Cédric Le Goater
2026-09-02 19:20 ` [RFC PATCH v2 6/9] igb: Quiesce VFs on STOP and include PF enable state in migration Cédric Le Goater
` (4 subsequent siblings)
9 siblings, 0 replies; 23+ messages in thread
From: Cédric Le Goater @ 2026-09-02 19:20 UTC (permalink / raw)
To: qemu-devel
Cc: Akihiko Odaki, Sriram Yagnaraman, Jason Wang, Alex Williamson,
Peter Xu, Cédric Le Goater
Add per-VF dirty page tracking using bitmaps allocated per IOVA range.
The DMA data path is instrumented via igb_core_dirty_track_dma() calls at all
five DMA write sites: TX descriptor writeback (head pointer and status),
RX descriptor writeback, RX header fragment, and RX payload fragment.
The migration DVSEC exposes DIRTY_ENABLE, DIRTY_DISABLE and
DIRTY_QUERY commands. Range parameters and query results are exchanged
through the shared DMA buffer.
Add the PRE_COPY device state to support pre-copy live migration with
concurrent dirty tracking. PRE_COPY transitions:
- RUNNING -> PRE_COPY
- PRE_COPY -> STOP, RUNNING, STOP_COPY
Dirty tracking is automatically disabled when leaving PRE_COPY or
STOP_COPY.
AI-used-for: code (prototype)
Signed-off-by: Cédric Le Goater <clg@redhat.com>
---
hw/net/igb_core.h | 4 +
hw/net/igb_migration.h | 65 +++++++-
hw/net/igb_core.c | 42 ++++--
hw/net/igb_migration.c | 326 ++++++++++++++++++++++++++++++++++++++++-
hw/net/trace-events | 5 +
5 files changed, 422 insertions(+), 20 deletions(-)
diff --git a/hw/net/igb_core.h b/hw/net/igb_core.h
index 22e10e4e6d0b..7fdc41690b36 100644
--- a/hw/net/igb_core.h
+++ b/hw/net/igb_core.h
@@ -40,6 +40,8 @@
#ifndef HW_NET_IGB_CORE_H
#define HW_NET_IGB_CORE_H
+#include "igb_migration.h"
+
#define E1000E_MAC_SIZE (0x8000)
#define IGB_EEPROM_SIZE (1024)
@@ -99,6 +101,8 @@ struct IGBCore {
void (*owner_start_recv)(PCIDevice *d);
int64_t timadj;
+
+ IGBVfDirtyState vf_dirty[IGB_MAX_VF_FUNCTIONS];
};
void
diff --git a/hw/net/igb_migration.h b/hw/net/igb_migration.h
index b2f601e74346..996de2d2f50b 100644
--- a/hw/net/igb_migration.h
+++ b/hw/net/igb_migration.h
@@ -21,7 +21,8 @@
* +0x04 DVSEC header 1 (len | rev | vendor_id)
* +0x08 DVSEC header 2 (DVSEC ID)
* +0x0A Reserved (padding for DWORD alignment)
- * +0x0C CAPS (RO: F_STATE[0])
+ * +0x0C CAPS (RO: F_STATE[0], F_DIRTY[1],
+ * max_ranges[11:8], pgsize[16:12])
* +0x10 CTRL (WO: doorbell command)
* +0x14 STATUS (RO: state[7:0], error_code[15:8])
* +0x18 BUF_ADDR_LO (RW: shared buffer GPA low)
@@ -44,6 +45,12 @@
/* CAPS register layout */
#define IGB_MIG_CAP_F_STATE (1u << 0)
+#define IGB_MIG_CAP_F_DIRTY (1u << 1)
+#define IGB_MIG_CAPS_MAX_RANGES_SHIFT 8
+#define IGB_MIG_CAPS_MAX_RANGES 4
+#define IGB_MIG_CAPS_PGSIZE_SHIFT 12
+#define IGB_MIG_CAPS_PGSIZE_4K (1u << 12)
+#define IGB_MIG_CAPS_PGSIZE_64K (1u << 16)
/* CTRL register: command in [7:0] */
#define IGB_MIG_CTRL_CMD_MASK 0xFF
@@ -53,6 +60,9 @@
#define IGB_MIG_CMD_SET_STATE 1
#define IGB_MIG_CMD_SAVE 2
#define IGB_MIG_CMD_LOAD 3
+#define IGB_MIG_CMD_DIRTY_ENABLE 4
+#define IGB_MIG_CMD_DIRTY_DISABLE 5
+#define IGB_MIG_CMD_DIRTY_QUERY 6
/* STATUS register: state in [7:0], error code [15:8] */
#define IGB_MIG_STATUS_STATE_MASK 0xFF
@@ -66,6 +76,7 @@
#define IGB_MIG_STATE_RUNNING 2
#define IGB_MIG_STATE_STOP_COPY 3
#define IGB_MIG_STATE_RESUMING 4
+#define IGB_MIG_STATE_PRE_COPY 5
/* Error codes */
#define IGB_MIG_ERR_UNK_CMD 1
@@ -75,10 +86,29 @@
#define IGB_MIG_ERR_BAD_SIZE 5
#define IGB_MIG_ERR_BAD_MAGIC 6
#define IGB_MIG_ERR_BAD_VERSION 7
+#define IGB_MIG_ERR_TOO_MANY_RANGES 8
+#define IGB_MIG_ERR_BAD_RANGE 9
+#define IGB_MIG_ERR_BAD_PGSIZE 10
+#define IGB_MIG_ERR_NOT_ENABLED 11
/* Shared buffer constants */
#define IGB_VF_STATE_MAX_SIZE 4096
+#define IGB_MIG_DIRTY_DEFAULT_PGSIZE 4096
+
+typedef struct IGBVfDirtyRange {
+ uint64_t iova;
+ uint64_t size;
+ uint64_t page_size;
+ unsigned long *bitmap;
+ uint64_t nbits;
+} IGBVfDirtyRange;
+
+typedef struct IGBVfDirtyState {
+ IGBVfDirtyRange ranges[IGB_MIG_CAPS_MAX_RANGES];
+ uint32_t num_ranges;
+} IGBVfDirtyState;
+
typedef struct IgbVfMigState {
uint32_t mig_state;
uint32_t mig_data[IGB_VF_STATE_MAX_SIZE / sizeof(uint32_t)];
@@ -86,6 +116,36 @@ typedef struct IgbVfMigState {
uint64_t mig_data_buf_addr;
} IgbVfMigState;
+/*
+ * DMA buffer layouts for dirty tracking commands.
+ *
+ * DIRTY_ENABLE: driver writes igb_mig_dirty_enable_req to buffer
+ * before cmd.
+ * DIRTY_QUERY: driver writes iova/size fields, device writes
+ * response + bitmap.
+ */
+struct igb_mig_dirty_enable_req {
+ uint32_t len;
+ uint32_t flags;
+ uint64_t pgsize;
+ uint64_t range_iova;
+ uint64_t range_size;
+ uint32_t reserved[4];
+};
+
+struct igb_mig_dirty_query {
+ uint32_t len;
+ uint32_t flags;
+ uint64_t iova;
+ uint64_t size;
+ uint32_t bitmap_size;
+ uint32_t dirty_page_count;
+ uint64_t dma_writes;
+ uint32_t reserved[6];
+ uint8_t bitmap[];
+};
+
+typedef struct IGBCore IGBCore;
typedef struct IgbVfState IgbVfState;
bool igbvf_add_migration_dvsec(PCIDevice *dev, Error **errp);
@@ -94,4 +154,7 @@ uint32_t igbvf_mig_config_read(IgbVfState *s, uint32_t addr, int size);
bool igbvf_mig_config_write(IgbVfState *s, uint32_t addr, uint32_t val,
int size);
+void igb_core_dirty_track_dma(IGBCore *core, int vfn,
+ dma_addr_t addr, dma_addr_t len);
+
#endif
diff --git a/hw/net/igb_core.c b/hw/net/igb_core.c
index 01745fe756d0..12329a1aff58 100644
--- a/hw/net/igb_core.c
+++ b/hw/net/igb_core.c
@@ -824,6 +824,16 @@ igb_rx_ring_init(IGBCore *core, E1000E_RxRing *rxr, int idx)
rxr->i = &i[idx];
}
+static inline void
+igb_pci_dma_write(IGBCore *core, PCIDevice *dev,
+ dma_addr_t addr, const void *buf, dma_addr_t len)
+{
+ pci_dma_write(dev, addr, buf, len);
+ if (pci_is_vf(dev)) {
+ igb_core_dirty_track_dma(core, pcie_sriov_vf_number(dev), addr, len);
+ }
+}
+
static uint32_t
igb_txdesc_writeback(IGBCore *core, dma_addr_t base,
union e1000_adv_tx_desc *tx_desc,
@@ -847,13 +857,15 @@ igb_txdesc_writeback(IGBCore *core, dma_addr_t base,
if (tdwba & 1) {
uint32_t buffer = cpu_to_le32(core->mac[txi->dh]);
- pci_dma_write(d, tdwba & ~3, &buffer, sizeof(buffer));
+ igb_pci_dma_write(core, d,
+ tdwba & ~3, &buffer, sizeof(buffer));
} else {
uint32_t status = le32_to_cpu(tx_desc->wb.status) | E1000_TXD_STAT_DD;
tx_desc->wb.status = cpu_to_le32(status);
- pci_dma_write(d, base + offsetof(union e1000_adv_tx_desc, wb),
- &tx_desc->wb, sizeof(tx_desc->wb));
+ igb_pci_dma_write(core, d,
+ base + offsetof(union e1000_adv_tx_desc, wb),
+ &tx_desc->wb, sizeof(tx_desc->wb));
}
return igb_tx_wb_eic(core, txi->idx);
@@ -1598,11 +1610,12 @@ igb_pci_dma_write_rx_desc(IGBCore *core, PCIDevice *dev, dma_addr_t addr,
uint8_t status = d->status;
d->status &= ~E1000_RXD_STAT_DD;
- pci_dma_write(dev, addr, desc, len);
+ igb_pci_dma_write(core, dev, addr, desc, len);
if (status & E1000_RXD_STAT_DD) {
d->status = status;
- pci_dma_write(dev, addr + offset, &status, sizeof(status));
+ igb_pci_dma_write(core, dev,
+ addr + offset, &status, sizeof(status));
}
} else {
union e1000_adv_rx_desc *d = &desc->adv;
@@ -1611,11 +1624,12 @@ igb_pci_dma_write_rx_desc(IGBCore *core, PCIDevice *dev, dma_addr_t addr,
uint32_t status = d->wb.upper.status_error;
d->wb.upper.status_error &= ~E1000_RXD_STAT_DD;
- pci_dma_write(dev, addr, desc, len);
+ igb_pci_dma_write(core, dev, addr, desc, len);
if (status & E1000_RXD_STAT_DD) {
d->wb.upper.status_error = status;
- pci_dma_write(dev, addr + offset, &status, sizeof(status));
+ igb_pci_dma_write(core, dev,
+ addr + offset, &status, sizeof(status));
}
}
}
@@ -1737,9 +1751,9 @@ igb_write_hdr_frag_to_rx_buffers(IGBCore *core,
{
assert(data_len <= pdma_st->rx_desc_header_buf_size -
pdma_st->bastate.written[0]);
- pci_dma_write(d,
- pdma_st->ba[0] + pdma_st->bastate.written[0],
- data, data_len);
+ igb_pci_dma_write(core, d,
+ pdma_st->ba[0] + pdma_st->bastate.written[0],
+ data, data_len);
pdma_st->bastate.written[0] += data_len;
pdma_st->bastate.cur_idx = 1;
}
@@ -1804,10 +1818,10 @@ igb_write_payload_frag_to_rx_buffers(IGBCore *core,
data,
bytes_to_write);
- pci_dma_write(d,
- pdma_st->ba[pdma_st->bastate.cur_idx] +
- pdma_st->bastate.written[pdma_st->bastate.cur_idx],
- data, bytes_to_write);
+ igb_pci_dma_write(core, d,
+ pdma_st->ba[pdma_st->bastate.cur_idx] +
+ pdma_st->bastate.written[pdma_st->bastate.cur_idx],
+ data, bytes_to_write);
pdma_st->bastate.written[pdma_st->bastate.cur_idx] += bytes_to_write;
data += bytes_to_write;
diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
index af7257ccc4fd..cfa73673e39a 100644
--- a/hw/net/igb_migration.c
+++ b/hw/net/igb_migration.c
@@ -8,6 +8,8 @@
#include "qemu/osdep.h"
#include "qemu/log.h"
+#include "qemu/bitmap.h"
+#include "qemu/units.h"
#include "hw/pci/pci_device.h"
#include "hw/pci/pcie.h"
#include "net/eth.h"
@@ -400,6 +402,285 @@ static int igbvf_mig_load(IgbVfState *s, const void *buf, size_t size)
return 0;
}
+/*
+ * Per-VF dirty page tracking
+ *
+ * All VF DMA writes in igb_core.c go through igb_pci_dma_write(),
+ * which calls igb_core_dirty_track_dma() to mark the target page in a
+ * per-range bitmap before performing the actual DMA.
+ *
+ * The IGBCore::vf_dirty[] bitmaps live in IGBCore so they are easily
+ * accessible from the core TX and RX paths without reaching back into
+ * VF state.
+ */
+
+void igb_core_dirty_track_dma(IGBCore *core, int vfn,
+ dma_addr_t addr, dma_addr_t len)
+{
+ IGBVfDirtyState *ds = &core->vf_dirty[vfn];
+ bool matched = false;
+ uint32_t i;
+
+ if (!ds->num_ranges) {
+ return;
+ }
+
+ trace_igb_core_dirty_track_dma(vfn, addr, len);
+
+ for (i = 0; i < ds->num_ranges; i++) {
+ IGBVfDirtyRange *r = &ds->ranges[i];
+ uint64_t r_end = r->iova + r->size;
+ uint64_t dma_end = addr + len;
+ uint64_t start, end, start_page, end_page, page;
+
+ if (addr >= r_end || dma_end <= r->iova) {
+ continue;
+ }
+
+ matched = true;
+ start = MAX(addr, r->iova);
+ end = MIN(dma_end, r_end);
+
+ start_page = (start - r->iova) / r->page_size;
+ end_page = (end - 1 - r->iova) / r->page_size;
+
+ for (page = start_page; page <= end_page; page++) {
+ if (page < r->nbits) {
+ set_bit(page, r->bitmap);
+ }
+ }
+ }
+
+ if (!matched) {
+ trace_igb_core_dirty_track_dma_drop(vfn, addr, len);
+ }
+}
+
+static IGBVfDirtyState *igb_core_vf_dirty_state(IgbVfState *s)
+{
+ IGBCore *core = igbvf_get_core(s);
+ return &core->vf_dirty[s->vfn];
+}
+
+#define IGB_MIG_DIRTY_MAX_PAGES ((256ULL * GiB) / (4 * KiB))
+
+static uint32_t igb_core_vf_dirty_enable(IgbVfState *s, uint64_t pgsize,
+ uint64_t range_iova,
+ uint64_t range_size)
+{
+ uint32_t caps = pci_get_long(PCI_DEVICE(s)->config +
+ IGB_MIG_DVSEC_OFFSET + IGB_MIG_CAPS);
+ IGBVfDirtyState *ds = igb_core_vf_dirty_state(s);
+ IGBVfDirtyRange *r;
+
+ if (ds->num_ranges >= IGB_MIG_CAPS_MAX_RANGES) {
+ return IGB_MIG_ERR_TOO_MANY_RANGES;
+ }
+
+ if (!range_size) {
+ return IGB_MIG_ERR_BAD_RANGE;
+ }
+
+ /* Validate page size against CAPS supported page size bitmask */
+ if (!is_power_of_2(pgsize) || !(pgsize & caps)) {
+ return IGB_MIG_ERR_BAD_PGSIZE;
+ }
+
+ if ((range_iova % pgsize) || (range_size % pgsize)) {
+ return IGB_MIG_ERR_BAD_PGSIZE;
+ }
+
+ if (range_size / pgsize > IGB_MIG_DIRTY_MAX_PAGES) {
+ return IGB_MIG_ERR_BAD_RANGE;
+ }
+
+ r = &ds->ranges[ds->num_ranges];
+ r->iova = range_iova;
+ r->size = range_size;
+ r->page_size = pgsize;
+ r->nbits = range_size / pgsize;
+ r->bitmap = bitmap_new(r->nbits);
+ ds->num_ranges++;
+ return 0;
+}
+
+static void igb_core_vf_dirty_disable(IgbVfState *s)
+{
+ IGBVfDirtyState *ds = igb_core_vf_dirty_state(s);
+ uint32_t i;
+
+ for (i = 0; i < ds->num_ranges; i++) {
+ IGBVfDirtyRange *r = &ds->ranges[i];
+
+ g_free(r->bitmap);
+ r->bitmap = NULL;
+ r->nbits = 0;
+ }
+ ds->num_ranges = 0;
+ trace_igbvf_mig_dirty_disable(s->vfn);
+}
+
+static bool igb_core_vf_dirty_enabled(IgbVfState *s)
+{
+ return igb_core_vf_dirty_state(s)->num_ranges > 0;
+}
+
+static void igb_core_vf_dirty_query(IGBVfDirtyRange *r,
+ uint64_t range_iova, uint64_t range_size,
+ void *buf, size_t buf_size,
+ size_t *out_size)
+{
+ uint64_t start_page = (range_iova - r->iova) / r->page_size;
+ uint64_t range_pages = range_size / r->page_size;
+ uint64_t count = MIN(range_pages, (uint64_t)buf_size * 8);
+
+ memset(buf, 0, buf_size);
+
+ if (start_page < r->nbits) {
+ uint64_t avail = r->nbits - start_page;
+ uint64_t n = MIN(count, avail);
+
+ bitmap_copy_with_src_offset(buf, r->bitmap, start_page, n);
+ }
+ *out_size = bitmap_empty(buf, count) ? 0 : DIV_ROUND_UP(count, 8);
+}
+
+static void igb_core_vf_dirty_query_commit(IGBVfDirtyRange *r,
+ uint64_t range_iova,
+ uint64_t range_size)
+{
+ uint64_t start_page = (range_iova - r->iova) / r->page_size;
+
+ if (start_page < r->nbits) {
+ uint64_t avail = r->nbits - start_page;
+ uint64_t range_pages = range_size / r->page_size;
+ uint64_t n = MIN(range_pages, avail);
+
+ bitmap_clear(r->bitmap, start_page, n);
+ }
+}
+
+static uint8_t igbvf_mig_cmd_dirty_enable(IgbVfState *s)
+{
+ IgbVfMigState *ms = &s->mig;
+ struct igb_mig_dirty_enable_req req;
+ MemTxResult r;
+ uint32_t status;
+
+ if (!ms->mig_data_buf_addr) {
+ return IGB_MIG_ERR_NO_BUFFER;
+ }
+
+ r = address_space_read(&address_space_memory, ms->mig_data_buf_addr,
+ MEMTXATTRS_UNSPECIFIED, &req, sizeof(req));
+ if (r != MEMTX_OK) {
+ return IGB_MIG_ERR_DMA_FAILED;
+ }
+
+ /* TODO: validate req.len and req.flags */
+
+ status = igb_core_vf_dirty_enable(s, le64_to_cpu(req.pgsize),
+ le64_to_cpu(req.range_iova),
+ le64_to_cpu(req.range_size));
+ if (status != 0) {
+ return status;
+ }
+ trace_igbvf_mig_dirty_enable(s->vfn, le64_to_cpu(req.pgsize),
+ le64_to_cpu(req.range_size) /
+ le64_to_cpu(req.pgsize));
+ return 0;
+}
+
+static IGBVfDirtyRange *igb_core_vf_dirty_range_valid(IGBVfDirtyState *ds,
+ uint64_t range_iova,
+ uint64_t range_size)
+{
+ if (!range_size) {
+ return NULL;
+ }
+
+ for (uint32_t i = 0; i < ds->num_ranges; i++) {
+ IGBVfDirtyRange *r = &ds->ranges[i];
+
+ if (range_iova >= r->iova &&
+ range_iova + range_size <= r->iova + r->size &&
+ QEMU_IS_ALIGNED(range_iova, r->page_size) &&
+ QEMU_IS_ALIGNED(range_size, r->page_size)) {
+ return r;
+ }
+ }
+ return NULL;
+}
+
+static uint8_t igbvf_mig_cmd_dirty_query(IgbVfState *s)
+{
+ IgbVfMigState *ms = &s->mig;
+ IGBVfDirtyState *ds = &igbvf_get_core(s)->vf_dirty[s->vfn];
+ uint64_t buf_addr = ms->mig_data_buf_addr;
+ uint64_t range_iova = 0, range_size = 0;
+ IGBVfDirtyRange *range;
+ uint32_t bmp_bytes, dirty_pages;
+ uint32_t val32;
+ size_t out_size;
+ g_autofree void *bitmap = NULL;
+
+ if (!buf_addr) {
+ return IGB_MIG_ERR_NO_BUFFER;
+ }
+
+ if (!igb_core_vf_dirty_enabled(s)) {
+ return IGB_MIG_ERR_NOT_ENABLED;
+ }
+
+ address_space_read(&address_space_memory,
+ buf_addr + offsetof(struct igb_mig_dirty_query, iova),
+ MEMTXATTRS_UNSPECIFIED, &range_iova, sizeof(range_iova));
+ range_iova = le64_to_cpu(range_iova);
+ address_space_read(&address_space_memory,
+ buf_addr + offsetof(struct igb_mig_dirty_query, size),
+ MEMTXATTRS_UNSPECIFIED, &range_size, sizeof(range_size));
+ range_size = le64_to_cpu(range_size);
+
+ range = igb_core_vf_dirty_range_valid(ds, range_iova, range_size);
+ if (!range) {
+ return IGB_MIG_ERR_BAD_RANGE;
+ }
+
+ bmp_bytes = BITS_TO_LONGS(range_size / range->page_size) *
+ sizeof(unsigned long);
+ bitmap = g_malloc0(bmp_bytes);
+
+ igb_core_vf_dirty_query(range, range_iova, range_size,
+ bitmap, bmp_bytes, &out_size);
+
+ if (out_size) {
+ if (address_space_write(&address_space_memory,
+ buf_addr +
+ offsetof(struct igb_mig_dirty_query, bitmap),
+ MEMTXATTRS_UNSPECIFIED, bitmap, out_size)) {
+ return IGB_MIG_ERR_DMA_FAILED;
+ }
+ }
+
+ igb_core_vf_dirty_query_commit(range, range_iova, range_size);
+
+ dirty_pages = bitmap_count_one(bitmap, range_size / range->page_size);
+
+ val32 = cpu_to_le32(out_size);
+ address_space_write(&address_space_memory,
+ buf_addr + offsetof(struct igb_mig_dirty_query,
+ bitmap_size),
+ MEMTXATTRS_UNSPECIFIED, &val32, sizeof(val32));
+ val32 = cpu_to_le32(dirty_pages);
+ address_space_write(&address_space_memory,
+ buf_addr + offsetof(struct igb_mig_dirty_query,
+ dirty_page_count),
+ MEMTXATTRS_UNSPECIFIED, &val32, sizeof(val32));
+
+ trace_igbvf_mig_dirty_query(s->vfn, (uint64_t)out_size, dirty_pages);
+ return 0;
+}
+
/*
* Migration command handlers
*/
@@ -419,7 +700,8 @@ static uint8_t igbvf_mig_cmd_save(IgbVfState *s)
MemTxResult r;
int ret;
- if (ms->mig_state != IGB_MIG_STATE_STOP_COPY) {
+ if (ms->mig_state != IGB_MIG_STATE_STOP_COPY &&
+ ms->mig_state != IGB_MIG_STATE_PRE_COPY) {
return IGB_MIG_ERR_BAD_STATE;
}
@@ -487,22 +769,33 @@ static uint8_t igbvf_mig_set_state(IgbVfState *s, uint32_t new_state)
case IGB_MIG_STATE_STOP:
if (old != IGB_MIG_STATE_RUNNING &&
old != IGB_MIG_STATE_STOP_COPY &&
+ old != IGB_MIG_STATE_PRE_COPY &&
old != IGB_MIG_STATE_RESUMING &&
old != IGB_MIG_STATE_ERROR) {
return IGB_MIG_ERR_BAD_STATE;
}
+ if (old == IGB_MIG_STATE_PRE_COPY ||
+ old == IGB_MIG_STATE_STOP_COPY ||
+ old == IGB_MIG_STATE_ERROR) {
+ igb_core_vf_dirty_disable(s);
+ }
/* Restore DATA_SIZE to max, same as at reset */
igbvf_mig_update_data_size(s, igb_core_vf_max_data_size(s));
break;
case IGB_MIG_STATE_RUNNING:
- if (old != IGB_MIG_STATE_STOP) {
+ if (old != IGB_MIG_STATE_STOP &&
+ old != IGB_MIG_STATE_PRE_COPY) {
return IGB_MIG_ERR_BAD_STATE;
}
+ if (old == IGB_MIG_STATE_PRE_COPY) {
+ igb_core_vf_dirty_disable(s);
+ }
break;
case IGB_MIG_STATE_STOP_COPY:
- if (old != IGB_MIG_STATE_STOP) {
+ if (old != IGB_MIG_STATE_STOP &&
+ old != IGB_MIG_STATE_PRE_COPY) {
return IGB_MIG_ERR_BAD_STATE;
}
ret = igb_core_vf_save_state(s, ms->mig_data, sizeof(ms->mig_data));
@@ -520,6 +813,12 @@ static uint8_t igbvf_mig_set_state(IgbVfState *s, uint32_t new_state)
igbvf_mig_update_data_size(s, 0);
break;
+ case IGB_MIG_STATE_PRE_COPY:
+ if (old != IGB_MIG_STATE_RUNNING) {
+ return IGB_MIG_ERR_BAD_STATE;
+ }
+ break;
+
default:
return IGB_MIG_ERR_BAD_STATE;
}
@@ -564,6 +863,18 @@ static void igbvf_mig_cmd_ctrl(IgbVfState *s, uint32_t val)
err = igbvf_mig_cmd_load(s);
break;
+ case IGB_MIG_CMD_DIRTY_ENABLE:
+ err = igbvf_mig_cmd_dirty_enable(s);
+ break;
+
+ case IGB_MIG_CMD_DIRTY_DISABLE:
+ igb_core_vf_dirty_disable(s);
+ break;
+
+ case IGB_MIG_CMD_DIRTY_QUERY:
+ err = igbvf_mig_cmd_dirty_query(s);
+ break;
+
default:
err = IGB_MIG_ERR_UNK_CMD;
break;
@@ -594,8 +905,10 @@ bool igbvf_add_migration_dvsec(PCIDevice *dev, Error **errp)
/* DVSEC header 2: DVSEC ID */
pci_set_word(dev->config + offset + 0x8, IGB_MIG_DVSEC_ID);
- /* CAPS: features (state migration only) */
- caps = IGB_MIG_CAP_F_STATE;
+ /* CAPS: features | max_ranges | supported page sizes (4K) */
+ caps = IGB_MIG_CAP_F_STATE | IGB_MIG_CAP_F_DIRTY |
+ (IGB_MIG_CAPS_MAX_RANGES << IGB_MIG_CAPS_MAX_RANGES_SHIFT) |
+ IGB_MIG_CAPS_PGSIZE_4K;
pci_set_long(dev->config + offset + IGB_MIG_CAPS, caps);
/* STATUS: initial state is RUNNING */
@@ -657,6 +970,9 @@ void igbvf_mig_state_reset(IgbVfState *s)
IgbVfMigState *ms = &s->mig;
trace_igbvf_mig_reset(s->vfn);
+
+ igb_core_vf_dirty_disable(s);
+
ms->mig_state = IGB_MIG_STATE_RUNNING;
ms->mig_data_buf_addr = 0;
igbvf_mig_update_data_size(s, igb_core_vf_max_data_size(s));
diff --git a/hw/net/trace-events b/hw/net/trace-events
index 7057fbe5f16a..c606191df2ef 100644
--- a/hw/net/trace-events
+++ b/hw/net/trace-events
@@ -300,6 +300,11 @@ igbvf_mig_set_state(uint16_t vfn, uint32_t old_state, uint32_t new_state) "VF%u:
igbvf_mig_save_state(uint16_t vfn, uint32_t size) "VF%u: saved %u bytes of device state"
igbvf_mig_load_state(uint16_t vfn, uint32_t size) "VF%u: loaded %u bytes of device state"
igbvf_mig_reset(uint16_t vfn) "VF%u: migration state reset"
+igbvf_mig_dirty_enable(uint16_t vfn, uint64_t pgsize, uint64_t nbits) "VF%u: dirty tracking enabled pgsize=%"PRIu64" nbits=%"PRIu64
+igbvf_mig_dirty_disable(uint16_t vfn) "VF%u: dirty tracking disabled"
+igbvf_mig_dirty_query(uint16_t vfn, uint64_t size, uint32_t dirty_pages) "VF%u: dirty query returned %"PRIu64" bytes, %u dirty pages"
+igb_core_dirty_track_dma(int vfn, uint64_t addr, uint64_t len) "VF%d: dirty DMA addr=0x%"PRIx64" len=%"PRIu64
+igb_core_dirty_track_dma_drop(int vfn, uint64_t addr, uint64_t len) "VF%d: dirty DMA dropped addr=0x%"PRIx64" len=%"PRIu64" no matching range"
# spapr_llan.c
spapr_vlan_get_rx_bd_from_pool_found(int pool, int32_t count, uint32_t rx_bufs) "pool=%d count=%"PRId32" rxbufs=%"PRIu32
--
2.55.0
^ permalink raw reply related [flat|nested] 23+ messages in thread
* [RFC PATCH v2 6/9] igb: Quiesce VFs on STOP and include PF enable state in migration
2026-09-02 19:20 [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Cédric Le Goater
` (4 preceding siblings ...)
2026-09-02 19:20 ` [RFC PATCH v2 5/9] igb: Add dirty page tracking for IGBVF migration Cédric Le Goater
@ 2026-09-02 19:20 ` Cédric Le Goater
2026-09-02 19:20 ` [RFC PATCH v2 7/9] igb: Fix post-migration RX ring deadlock Cédric Le Goater
` (3 subsequent siblings)
9 siblings, 0 replies; 23+ messages in thread
From: Cédric Le Goater @ 2026-09-02 19:20 UTC (permalink / raw)
To: qemu-devel
Cc: Akihiko Odaki, Sriram Yagnaraman, Jason Wang, Alex Williamson,
Peter Xu, Cédric Le Goater
Quiesce VFs by clearing their VFRE/VFTE bits on transitions to STOP
(from RUNNING, PRE_COPY, or ERROR) and to STOP_COPY (from PRE_COPY).
This prevents further DMA while device state is being serialized.
Restore the saved VFRE/VFTE state on STOP->RUNNING so the VF can
resume normal operation.
Include the per-VF VFRE/VFTE enable bits in the migration blob so the
destination knows whether receive and transmit were active before
quiesce. On reset, both default to true.
Add a QUIESCED bit to the STATUS register so the driver can confirm
DMA has been drained before reading device state.
Clear the VFLRE (VF Level Reset Event) bit before restoring state on
the destination to prevent the PF watchdog from seeing a stale reset
indication and overwriting the loaded registers.
AI-used-for: analysis, code (prototype)
Signed-off-by: Cédric Le Goater <clg@redhat.com>
---
hw/net/igb_migration.h | 6 ++-
hw/net/igb_migration.c | 86 +++++++++++++++++++++++++++++++++++++++++-
hw/net/trace-events | 6 ++-
3 files changed, 93 insertions(+), 5 deletions(-)
diff --git a/hw/net/igb_migration.h b/hw/net/igb_migration.h
index 996de2d2f50b..2bb9a0ed36ce 100644
--- a/hw/net/igb_migration.h
+++ b/hw/net/igb_migration.h
@@ -24,7 +24,8 @@
* +0x0C CAPS (RO: F_STATE[0], F_DIRTY[1],
* max_ranges[11:8], pgsize[16:12])
* +0x10 CTRL (WO: doorbell command)
- * +0x14 STATUS (RO: state[7:0], error_code[15:8])
+ * +0x14 STATUS (RO: state[7:0], error_code[15:8],
+ * QUIESCED[16])
* +0x18 BUF_ADDR_LO (RW: shared buffer GPA low)
* +0x1C BUF_ADDR_HI (RW: shared buffer GPA high)
* +0x20 DATA_SIZE (RO: max state blob size in bytes)
@@ -69,6 +70,7 @@
#define IGB_MIG_STATUS_ERROR_CODE_SHIFT 8
#define IGB_MIG_STATUS_ERR(code) \
((uint32_t)(code) << IGB_MIG_STATUS_ERROR_CODE_SHIFT)
+#define IGB_MIG_STATUS_QUIESCED (1u << 16)
/* Device states (based on VFIO migration v2) */
#define IGB_MIG_STATE_ERROR 0
@@ -114,6 +116,8 @@ typedef struct IgbVfMigState {
uint32_t mig_data[IGB_VF_STATE_MAX_SIZE / sizeof(uint32_t)];
uint32_t mig_data_size;
uint64_t mig_data_buf_addr;
+ bool mig_saved_vfre;
+ bool mig_saved_vfte;
} IgbVfMigState;
/*
diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
index cfa73673e39a..435bd0d1623b 100644
--- a/hw/net/igb_migration.c
+++ b/hw/net/igb_migration.c
@@ -58,6 +58,8 @@ typedef struct IgbMigBlob {
IgbMigRegPair ra[IGB_VF_MAX_RA_REGS];
uint32_t num_tx_ctx;
IgbMigTxCtx tx_ctx[2];
+ uint32_t vfre;
+ uint32_t vfte;
} IgbMigBlob;
#define IGB_MIG_BLOB_SIZE sizeof(IgbMigBlob)
@@ -210,6 +212,7 @@ static void igb_core_vf_save_tx_ctx(IGBCore *core, int queue,
static int igb_core_vf_save_state(IgbVfState *s, void *buf, size_t buf_size)
{
int size = IGB_MIG_BLOB_SIZE;
+ IgbVfMigState *ms = &s->mig;
IGBCore *core = igbvf_get_core(s);
IgbMigBlob *blob = buf;
uint32_t offsets[IGB_VF_MAX_FIXED_REGS];
@@ -248,7 +251,12 @@ static int igb_core_vf_save_state(IgbVfState *s, void *buf, size_t buf_size)
igb_core_vf_save_tx_ctx(core, q0, &blob->tx_ctx[0]);
igb_core_vf_save_tx_ctx(core, q1, &blob->tx_ctx[1]);
- trace_igbvf_mig_save_state(s->vfn, size);
+ blob->vfre = cpu_to_le32(ms->mig_saved_vfre);
+ blob->vfte = cpu_to_le32(ms->mig_saved_vfte);
+
+ trace_igbvf_mig_save_state(s->vfn, size, ms->mig_saved_vfre,
+ ms->mig_saved_vfte,
+ core->mac[VFRE]);
return size;
}
@@ -291,6 +299,7 @@ static uint32_t igb_vf_relocate_offset(uint32_t offset,
static int igb_core_vf_load_state(IgbVfState *s, const void *buf, size_t size)
{
+ IgbVfMigState *ms = &s->mig;
IGBCore *core = igbvf_get_core(s);
uint32_t src_offsets[IGB_VF_MAX_FIXED_REGS];
uint32_t dst_offsets[IGB_VF_MAX_FIXED_REGS];
@@ -379,7 +388,12 @@ static int igb_core_vf_load_state(IgbVfState *s, const void *buf, size_t size)
igb_core_vf_load_tx_ctx(core, q0, &blob->tx_ctx[0]);
igb_core_vf_load_tx_ctx(core, q1, &blob->tx_ctx[1]);
- trace_igbvf_mig_load_state(s->vfn, (uint32_t)size);
+ ms->mig_saved_vfre = !!le32_to_cpu(blob->vfre);
+ ms->mig_saved_vfte = !!le32_to_cpu(blob->vfte);
+
+ trace_igbvf_mig_load_state(s->vfn, (uint32_t)size,
+ ms->mig_saved_vfre,
+ ms->mig_saved_vfte);
return 0;
}
@@ -388,6 +402,12 @@ static int igbvf_mig_load(IgbVfState *s, const void *buf, size_t size)
int ret;
IGBCore *core = igbvf_get_core(s);
+ /*
+ * Pre-load: Clear the VFLRE bit before restoring state so the PF
+ * watchdog does not overwrite what we are about to load.
+ */
+ core->mac[VFLRE] &= ~BIT(s->vfn);
+
ret = igb_core_vf_load_state(s, buf, size);
if (ret < 0) {
return ret;
@@ -759,6 +779,45 @@ static uint8_t igbvf_mig_cmd_load(IgbVfState *s)
return 0;
}
+/* Quiesce a VF by disabling its RX and TX at the PF level. */
+static void igb_core_vf_quiesce(IgbVfState *s)
+{
+ IgbVfMigState *ms = &s->mig;
+ IGBCore *core = igbvf_get_core(s);
+
+ ms->mig_saved_vfre = !!(core->mac[VFRE] & BIT(s->vfn));
+ ms->mig_saved_vfte = !!(core->mac[VFTE] & BIT(s->vfn));
+
+ core->mac[VFRE] &= ~BIT(s->vfn);
+ core->mac[VFTE] &= ~BIT(s->vfn);
+ trace_igbvf_mig_quiesce(s->vfn, core->mac[VFRE], core->mac[VFTE]);
+}
+
+static void igb_core_vf_unquiesce(IgbVfState *s)
+{
+ IgbVfMigState *ms = &s->mig;
+ IGBCore *core = igbvf_get_core(s);
+ bool re = ms->mig_saved_vfre;
+ bool te = ms->mig_saved_vfte;
+
+ if (re) {
+ core->mac[VFRE] |= BIT(s->vfn);
+ } else {
+ core->mac[VFRE] &= ~BIT(s->vfn);
+ }
+ if (te) {
+ core->mac[VFTE] |= BIT(s->vfn);
+ } else {
+ core->mac[VFTE] &= ~BIT(s->vfn);
+ }
+
+ trace_igbvf_mig_unquiesce(s->vfn, core->mac[VFRE], core->mac[VFTE]);
+
+ if (re) {
+ igb_start_recv(core);
+ }
+}
+
static uint8_t igbvf_mig_set_state(IgbVfState *s, uint32_t new_state)
{
IgbVfMigState *ms = &s->mig;
@@ -779,6 +838,11 @@ static uint8_t igbvf_mig_set_state(IgbVfState *s, uint32_t new_state)
old == IGB_MIG_STATE_ERROR) {
igb_core_vf_dirty_disable(s);
}
+ if (old == IGB_MIG_STATE_RUNNING ||
+ old == IGB_MIG_STATE_PRE_COPY ||
+ old == IGB_MIG_STATE_ERROR) {
+ igb_core_vf_quiesce(s);
+ }
/* Restore DATA_SIZE to max, same as at reset */
igbvf_mig_update_data_size(s, igb_core_vf_max_data_size(s));
break;
@@ -791,6 +855,9 @@ static uint8_t igbvf_mig_set_state(IgbVfState *s, uint32_t new_state)
if (old == IGB_MIG_STATE_PRE_COPY) {
igb_core_vf_dirty_disable(s);
}
+ if (old == IGB_MIG_STATE_STOP) {
+ igb_core_vf_unquiesce(s);
+ }
break;
case IGB_MIG_STATE_STOP_COPY:
@@ -798,6 +865,9 @@ static uint8_t igbvf_mig_set_state(IgbVfState *s, uint32_t new_state)
old != IGB_MIG_STATE_PRE_COPY) {
return IGB_MIG_ERR_BAD_STATE;
}
+ if (old == IGB_MIG_STATE_PRE_COPY) {
+ igb_core_vf_quiesce(s);
+ }
ret = igb_core_vf_save_state(s, ms->mig_data, sizeof(ms->mig_data));
if (ret < 0) {
return -ret;
@@ -840,6 +910,16 @@ static void igbvf_mig_update_status(IgbVfState *s, uint8_t err)
status = IGB_MIG_STATE_ERROR | IGB_MIG_STATUS_ERR(err);
}
+ /*
+ * QUIESCED tells the driver it is safe to read device state.
+ * In STOP and STOP_COPY, igb_core_vf_quiesce() has already
+ * cleared VFRE/VFTE so no further VF DMA can occur.
+ */
+ if (ms->mig_state == IGB_MIG_STATE_STOP ||
+ ms->mig_state == IGB_MIG_STATE_STOP_COPY) {
+ status |= IGB_MIG_STATUS_QUIESCED;
+ }
+
pci_set_long(dev->config + IGB_MIG_DVSEC_OFFSET + IGB_MIG_STATUS, status);
}
@@ -977,6 +1057,8 @@ void igbvf_mig_state_reset(IgbVfState *s)
ms->mig_data_buf_addr = 0;
igbvf_mig_update_data_size(s, igb_core_vf_max_data_size(s));
memset(ms->mig_data, 0, sizeof(ms->mig_data));
+ ms->mig_saved_vfre = true;
+ ms->mig_saved_vfte = true;
pci_set_long(PCI_DEVICE(s)->config +
IGB_MIG_DVSEC_OFFSET + IGB_MIG_BUF_ADDR_LO, 0);
diff --git a/hw/net/trace-events b/hw/net/trace-events
index c606191df2ef..eb5c61a8f6fb 100644
--- a/hw/net/trace-events
+++ b/hw/net/trace-events
@@ -297,14 +297,16 @@ igbvf_wrn_io_addr_unknown(uint64_t addr) "IO unknown register 0x%"PRIx64
# igb_migration.c
igbvf_mig_set_state(uint16_t vfn, uint32_t old_state, uint32_t new_state) "VF%u: state %u -> %u"
-igbvf_mig_save_state(uint16_t vfn, uint32_t size) "VF%u: saved %u bytes of device state"
-igbvf_mig_load_state(uint16_t vfn, uint32_t size) "VF%u: loaded %u bytes of device state"
+igbvf_mig_save_state(uint16_t vfn, int size, bool vfre, bool vfte, uint32_t reg_vfre) "VF%u: saved %d bytes vfre=%d vfte=%d VFRE=0x%x"
+igbvf_mig_load_state(uint16_t vfn, uint32_t size, bool vfre, bool vfte) "VF%u: loaded %u bytes vfre=%d vfte=%d"
igbvf_mig_reset(uint16_t vfn) "VF%u: migration state reset"
igbvf_mig_dirty_enable(uint16_t vfn, uint64_t pgsize, uint64_t nbits) "VF%u: dirty tracking enabled pgsize=%"PRIu64" nbits=%"PRIu64
igbvf_mig_dirty_disable(uint16_t vfn) "VF%u: dirty tracking disabled"
igbvf_mig_dirty_query(uint16_t vfn, uint64_t size, uint32_t dirty_pages) "VF%u: dirty query returned %"PRIu64" bytes, %u dirty pages"
igb_core_dirty_track_dma(int vfn, uint64_t addr, uint64_t len) "VF%d: dirty DMA addr=0x%"PRIx64" len=%"PRIu64
igb_core_dirty_track_dma_drop(int vfn, uint64_t addr, uint64_t len) "VF%d: dirty DMA dropped addr=0x%"PRIx64" len=%"PRIu64" no matching range"
+igbvf_mig_quiesce(uint16_t vfn, uint32_t vfre, uint32_t vfte) "VF%u: quiesce VFRE=0x%x VFTE=0x%x"
+igbvf_mig_unquiesce(uint16_t vfn, uint32_t vfre, uint32_t vfte) "VF%u: unquiesce VFRE=0x%x VFTE=0x%x"
# spapr_llan.c
spapr_vlan_get_rx_bd_from_pool_found(int pool, int32_t count, uint32_t rx_bufs) "pool=%d count=%"PRId32" rxbufs=%"PRIu32
--
2.55.0
^ permalink raw reply related [flat|nested] 23+ messages in thread
* [RFC PATCH v2 7/9] igb: Fix post-migration RX ring deadlock
2026-09-02 19:20 [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Cédric Le Goater
` (5 preceding siblings ...)
2026-09-02 19:20 ` [RFC PATCH v2 6/9] igb: Quiesce VFs on STOP and include PF enable state in migration Cédric Le Goater
@ 2026-09-02 19:20 ` Cédric Le Goater
2026-09-02 19:20 ` [RFC PATCH v2 8/9] igb: Add dirty page tracking statistics Cédric Le Goater
` (2 subsequent siblings)
9 siblings, 0 replies; 23+ messages in thread
From: Cédric Le Goater @ 2026-09-02 19:20 UTC (permalink / raw)
To: qemu-devel
Cc: Akihiko Odaki, Sriram Yagnaraman, Jason Wang, Alex Williamson,
Peter Xu, Cédric Le Goater
After restoring VFRE/VFTE on the destination, the VF's RX rings may
be stalled because the guest driver is waiting for an interrupt that
was pending at the time of migration. Re-raise any pending EICR
causes for the VF (both RX and TX) via igb_set_eics() to break the
deadlock.
AI-used-for: analysis, code (prototype)
Signed-off-by: Cédric Le Goater <clg@redhat.com>
---
hw/net/igb_core.h | 1 +
hw/net/igb_core.c | 16 ++++++++++++++++
hw/net/igb_migration.c | 4 ++--
3 files changed, 19 insertions(+), 2 deletions(-)
diff --git a/hw/net/igb_core.h b/hw/net/igb_core.h
index 7fdc41690b36..7d6f44b7ffe9 100644
--- a/hw/net/igb_core.h
+++ b/hw/net/igb_core.h
@@ -150,6 +150,7 @@ igb_start_recv(IGBCore *core);
IGBCore *igb_pf_get_core(void *pf);
void igb_core_vf_propagate_irqs(IGBCore *core, uint16_t vfn);
+void igb_core_vf_rearm_irqs(IGBCore *core, uint16_t vfn);
void igb_core_vf_propagate_ivar(IGBCore *core, uint16_t vfn);
#endif
diff --git a/hw/net/igb_core.c b/hw/net/igb_core.c
index 12329a1aff58..1de8d95102e4 100644
--- a/hw/net/igb_core.c
+++ b/hw/net/igb_core.c
@@ -4632,3 +4632,19 @@ void igb_core_vf_propagate_ivar(IGBCore *core, uint16_t vfn)
core->mac[IVAR0 + n / 4] &= ~mask;
}
}
+
+/*
+ * Re-apply VF interrupt enables to PF aggregates and raise the
+ * pending causes so the guest driver resumes polling after migration.
+ */
+void igb_core_vf_rearm_irqs(IGBCore *core, uint16_t vfn)
+{
+ uint32_t shift = 22 - vfn * IGBVF_MSIX_VEC_NUM;
+ uint32_t pvt_idx = PVTEICR0 + vfn * 0x40;
+ uint32_t causes = (core->mac[pvt_idx] & 0x7) << shift;
+
+ igb_core_vf_propagate_irqs(core, vfn);
+ if (causes) {
+ igb_set_eics(core, EICS, causes);
+ }
+}
diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
index 435bd0d1623b..46299ff77ab9 100644
--- a/hw/net/igb_migration.c
+++ b/hw/net/igb_migration.c
@@ -813,8 +813,8 @@ static void igb_core_vf_unquiesce(IgbVfState *s)
trace_igbvf_mig_unquiesce(s->vfn, core->mac[VFRE], core->mac[VFTE]);
- if (re) {
- igb_start_recv(core);
+ if (re || te) {
+ igb_core_vf_rearm_irqs(core, s->vfn);
}
}
--
2.55.0
^ permalink raw reply related [flat|nested] 23+ messages in thread
* [RFC PATCH v2 8/9] igb: Add dirty page tracking statistics
2026-09-02 19:20 [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Cédric Le Goater
` (6 preceding siblings ...)
2026-09-02 19:20 ` [RFC PATCH v2 7/9] igb: Fix post-migration RX ring deadlock Cédric Le Goater
@ 2026-09-02 19:20 ` Cédric Le Goater
2026-09-02 19:20 ` [RFC PATCH v2 9/9] docs: Add igb VF migration testing setup guide Cédric Le Goater
2026-09-09 7:19 ` [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Akihiko Odaki
9 siblings, 0 replies; 23+ messages in thread
From: Cédric Le Goater @ 2026-09-02 19:20 UTC (permalink / raw)
To: qemu-devel
Cc: Akihiko Odaki, Sriram Yagnaraman, Jason Wang, Alex Williamson,
Peter Xu, Cédric Le Goater
Add a GET_STATS command (cmd 7) to the migration DVSEC for monitoring
dirty page tracking and DMA activity per VF. The device DMA-writes an
igb_mig_stats_resp struct to the shared buffer.
The dma_writes counter is also reported in the dirty query buffer so
the driver gets it alongside the bitmap without an extra config read.
Counters are reset on first DIRTY_ENABLE or device reset, so the
driver can read final values after DIRTY_DISABLE.
AI-used-for: code (prototype)
Signed-off-by: Cédric Le Goater <clg@redhat.com>
---
hw/net/igb_core.h | 1 +
hw/net/igb_migration.h | 22 ++++++++++++++
hw/net/igb_migration.c | 65 ++++++++++++++++++++++++++++++++++++++++--
hw/net/trace-events | 2 +-
4 files changed, 86 insertions(+), 4 deletions(-)
diff --git a/hw/net/igb_core.h b/hw/net/igb_core.h
index 7d6f44b7ffe9..380cc88b7959 100644
--- a/hw/net/igb_core.h
+++ b/hw/net/igb_core.h
@@ -103,6 +103,7 @@ struct IGBCore {
int64_t timadj;
IGBVfDirtyState vf_dirty[IGB_MAX_VF_FUNCTIONS];
+ IgbVfMigStats vf_mig_stats[IGB_MAX_VF_FUNCTIONS];
};
void
diff --git a/hw/net/igb_migration.h b/hw/net/igb_migration.h
index 2bb9a0ed36ce..b334bb8f78ba 100644
--- a/hw/net/igb_migration.h
+++ b/hw/net/igb_migration.h
@@ -64,6 +64,7 @@
#define IGB_MIG_CMD_DIRTY_ENABLE 4
#define IGB_MIG_CMD_DIRTY_DISABLE 5
#define IGB_MIG_CMD_DIRTY_QUERY 6
+#define IGB_MIG_CMD_GET_STATS 7
/* STATUS register: state in [7:0], error code [15:8] */
#define IGB_MIG_STATUS_STATE_MASK 0xFF
@@ -111,6 +112,15 @@ typedef struct IGBVfDirtyState {
uint32_t num_ranges;
} IGBVfDirtyState;
+typedef struct IgbVfMigStats {
+ uint64_t dma_writes;
+ uint64_t dma_bytes;
+ uint32_t dirty_pages_set;
+ uint32_t dirty_pages_cleared;
+ uint32_t dirty_page_count;
+ uint32_t dirty_query_count;
+} IgbVfMigStats;
+
typedef struct IgbVfMigState {
uint32_t mig_state;
uint32_t mig_data[IGB_VF_STATE_MAX_SIZE / sizeof(uint32_t)];
@@ -149,6 +159,18 @@ struct igb_mig_dirty_query {
uint8_t bitmap[];
};
+/*
+ * GET_STATS: device writes igb_mig_stats_resp to buffer.
+ */
+struct igb_mig_stats_resp {
+ uint64_t dma_writes;
+ uint64_t dma_bytes;
+ uint32_t dirty_pages_set;
+ uint32_t dirty_pages_cleared;
+ uint32_t dirty_page_count;
+ uint32_t dirty_query_count;
+};
+
typedef struct IGBCore IGBCore;
typedef struct IgbVfState IgbVfState;
diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
index 46299ff77ab9..2504d5bd308b 100644
--- a/hw/net/igb_migration.c
+++ b/hw/net/igb_migration.c
@@ -438,6 +438,7 @@ void igb_core_dirty_track_dma(IGBCore *core, int vfn,
dma_addr_t addr, dma_addr_t len)
{
IGBVfDirtyState *ds = &core->vf_dirty[vfn];
+ IgbVfMigStats *stats = &core->vf_mig_stats[vfn];
bool matched = false;
uint32_t i;
@@ -447,6 +448,9 @@ void igb_core_dirty_track_dma(IGBCore *core, int vfn,
trace_igb_core_dirty_track_dma(vfn, addr, len);
+ stats->dma_writes++;
+ stats->dma_bytes += len;
+
for (i = 0; i < ds->num_ranges; i++) {
IGBVfDirtyRange *r = &ds->ranges[i];
uint64_t r_end = r->iova + r->size;
@@ -466,7 +470,10 @@ void igb_core_dirty_track_dma(IGBCore *core, int vfn,
for (page = start_page; page <= end_page; page++) {
if (page < r->nbits) {
- set_bit(page, r->bitmap);
+ if (!test_and_set_bit(page, r->bitmap)) {
+ stats->dirty_pages_set++;
+ stats->dirty_page_count++;
+ }
}
}
}
@@ -493,6 +500,15 @@ static uint32_t igb_core_vf_dirty_enable(IgbVfState *s, uint64_t pgsize,
IGBVfDirtyState *ds = igb_core_vf_dirty_state(s);
IGBVfDirtyRange *r;
+ /*
+ * Reset stats on first enable so the driver can read them after
+ * disable
+ */
+ if (ds->num_ranges == 0) {
+ memset(&igbvf_get_core(s)->vf_mig_stats[s->vfn], 0,
+ sizeof(IgbVfMigStats));
+ }
+
if (ds->num_ranges >= IGB_MIG_CAPS_MAX_RANGES) {
return IGB_MIG_ERR_TOO_MANY_RANGES;
}
@@ -635,12 +651,14 @@ static IGBVfDirtyRange *igb_core_vf_dirty_range_valid(IGBVfDirtyState *ds,
static uint8_t igbvf_mig_cmd_dirty_query(IgbVfState *s)
{
IgbVfMigState *ms = &s->mig;
+ IgbVfMigStats *stats = &igbvf_get_core(s)->vf_mig_stats[s->vfn];
IGBVfDirtyState *ds = &igbvf_get_core(s)->vf_dirty[s->vfn];
uint64_t buf_addr = ms->mig_data_buf_addr;
uint64_t range_iova = 0, range_size = 0;
IGBVfDirtyRange *range;
uint32_t bmp_bytes, dirty_pages;
uint32_t val32;
+ uint64_t val64;
size_t out_size;
g_autofree void *bitmap = NULL;
@@ -686,6 +704,10 @@ static uint8_t igbvf_mig_cmd_dirty_query(IgbVfState *s)
dirty_pages = bitmap_count_one(bitmap, range_size / range->page_size);
+ stats->dirty_pages_cleared += dirty_pages;
+ stats->dirty_page_count -= MIN(stats->dirty_page_count, dirty_pages);
+ stats->dirty_query_count++;
+
val32 = cpu_to_le32(out_size);
address_space_write(&address_space_memory,
buf_addr + offsetof(struct igb_mig_dirty_query,
@@ -696,8 +718,13 @@ static uint8_t igbvf_mig_cmd_dirty_query(IgbVfState *s)
buf_addr + offsetof(struct igb_mig_dirty_query,
dirty_page_count),
MEMTXATTRS_UNSPECIFIED, &val32, sizeof(val32));
-
- trace_igbvf_mig_dirty_query(s->vfn, (uint64_t)out_size, dirty_pages);
+ val64 = cpu_to_le64(stats->dma_writes);
+ address_space_write(&address_space_memory,
+ buf_addr + offsetof(struct igb_mig_dirty_query,
+ dma_writes),
+ MEMTXATTRS_UNSPECIFIED, &val64, sizeof(val64));
+ trace_igbvf_mig_dirty_query(s->vfn, (uint64_t)out_size, dirty_pages,
+ stats->dma_writes);
return 0;
}
@@ -898,6 +925,32 @@ static uint8_t igbvf_mig_set_state(IgbVfState *s, uint32_t new_state)
return 0;
}
+static uint8_t igbvf_mig_cmd_get_stats(IgbVfState *s)
+{
+ IgbVfMigState *ms = &s->mig;
+ IgbVfMigStats *stats = &igbvf_get_core(s)->vf_mig_stats[s->vfn];
+ struct igb_mig_stats_resp resp;
+ MemTxResult r;
+
+ if (!ms->mig_data_buf_addr) {
+ return IGB_MIG_ERR_NO_BUFFER;
+ }
+
+ resp.dma_writes = cpu_to_le64(stats->dma_writes);
+ resp.dma_bytes = cpu_to_le64(stats->dma_bytes);
+ resp.dirty_pages_set = cpu_to_le32(stats->dirty_pages_set);
+ resp.dirty_pages_cleared = cpu_to_le32(stats->dirty_pages_cleared);
+ resp.dirty_page_count = cpu_to_le32(stats->dirty_page_count);
+ resp.dirty_query_count = cpu_to_le32(stats->dirty_query_count);
+
+ r = address_space_write(&address_space_memory, ms->mig_data_buf_addr,
+ MEMTXATTRS_UNSPECIFIED, &resp, sizeof(resp));
+ if (r != MEMTX_OK) {
+ return IGB_MIG_ERR_DMA_FAILED;
+ }
+ return 0;
+}
+
static void igbvf_mig_update_status(IgbVfState *s, uint8_t err)
{
IgbVfMigState *ms = &s->mig;
@@ -955,6 +1008,10 @@ static void igbvf_mig_cmd_ctrl(IgbVfState *s, uint32_t val)
err = igbvf_mig_cmd_dirty_query(s);
break;
+ case IGB_MIG_CMD_GET_STATS:
+ err = igbvf_mig_cmd_get_stats(s);
+ break;
+
default:
err = IGB_MIG_ERR_UNK_CMD;
break;
@@ -1059,6 +1116,8 @@ void igbvf_mig_state_reset(IgbVfState *s)
memset(ms->mig_data, 0, sizeof(ms->mig_data));
ms->mig_saved_vfre = true;
ms->mig_saved_vfte = true;
+ memset(&igbvf_get_core(s)->vf_mig_stats[s->vfn], 0,
+ sizeof(IgbVfMigStats));
pci_set_long(PCI_DEVICE(s)->config +
IGB_MIG_DVSEC_OFFSET + IGB_MIG_BUF_ADDR_LO, 0);
diff --git a/hw/net/trace-events b/hw/net/trace-events
index eb5c61a8f6fb..6c2c5493094a 100644
--- a/hw/net/trace-events
+++ b/hw/net/trace-events
@@ -302,7 +302,7 @@ igbvf_mig_load_state(uint16_t vfn, uint32_t size, bool vfre, bool vfte) "VF%u: l
igbvf_mig_reset(uint16_t vfn) "VF%u: migration state reset"
igbvf_mig_dirty_enable(uint16_t vfn, uint64_t pgsize, uint64_t nbits) "VF%u: dirty tracking enabled pgsize=%"PRIu64" nbits=%"PRIu64
igbvf_mig_dirty_disable(uint16_t vfn) "VF%u: dirty tracking disabled"
-igbvf_mig_dirty_query(uint16_t vfn, uint64_t size, uint32_t dirty_pages) "VF%u: dirty query returned %"PRIu64" bytes, %u dirty pages"
+igbvf_mig_dirty_query(uint16_t vfn, uint64_t size, uint32_t dirty_pages, uint64_t dma_writes) "VF%u: dirty query returned %"PRIu64" bytes, %u dirty pages (dma_writes=%"PRIu64")"
igb_core_dirty_track_dma(int vfn, uint64_t addr, uint64_t len) "VF%d: dirty DMA addr=0x%"PRIx64" len=%"PRIu64
igb_core_dirty_track_dma_drop(int vfn, uint64_t addr, uint64_t len) "VF%d: dirty DMA dropped addr=0x%"PRIx64" len=%"PRIu64" no matching range"
igbvf_mig_quiesce(uint16_t vfn, uint32_t vfre, uint32_t vfte) "VF%u: quiesce VFRE=0x%x VFTE=0x%x"
--
2.55.0
^ permalink raw reply related [flat|nested] 23+ messages in thread
* [RFC PATCH v2 9/9] docs: Add igb VF migration testing setup guide
2026-09-02 19:20 [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Cédric Le Goater
` (7 preceding siblings ...)
2026-09-02 19:20 ` [RFC PATCH v2 8/9] igb: Add dirty page tracking statistics Cédric Le Goater
@ 2026-09-02 19:20 ` Cédric Le Goater
2026-09-08 8:39 ` Akihiko Odaki
2026-09-09 7:19 ` [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Akihiko Odaki
9 siblings, 1 reply; 23+ messages in thread
From: Cédric Le Goater @ 2026-09-02 19:20 UTC (permalink / raw)
To: qemu-devel
Cc: Akihiko Odaki, Sriram Yagnaraman, Jason Wang, Alex Williamson,
Peter Xu, Cédric Le Goater
Document the igb VF migration interface: DVSEC register layout, dirty
page tracking, testing setup with nested virtualization.
AI-used-for: docs
Signed-off-by: Cédric Le Goater <clg@redhat.com>
---
MAINTAINERS | 1 +
docs/system/device-emulation.rst | 1 +
docs/system/devices/igb-migration.rst | 417 ++++++++++++++++++++++++++
docs/system/devices/igb.rst | 6 +
4 files changed, 425 insertions(+)
create mode 100644 docs/system/devices/igb-migration.rst
diff --git a/MAINTAINERS b/MAINTAINERS
index f88b526be238..4a2f1357e936 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -2820,6 +2820,7 @@ igb VF migration
M: Cédric Le Goater <clg@redhat.com>
S: Maintained
F: hw/net/igb_migration.*
+F: docs/system/devices/igb-migration.rst
eepro100
M: Stefan Weil <sw@weilnetz.de>
diff --git a/docs/system/device-emulation.rst b/docs/system/device-emulation.rst
index 40054bb7dfcc..75f423b795cd 100644
--- a/docs/system/device-emulation.rst
+++ b/docs/system/device-emulation.rst
@@ -90,6 +90,7 @@ Emulated Devices
devices/cxl.rst
devices/emmc.rst
devices/igb.rst
+ devices/igb-migration.rst
devices/ivshmem-flat.rst
devices/ivshmem.rst
devices/keyboard.rst
diff --git a/docs/system/devices/igb-migration.rst b/docs/system/devices/igb-migration.rst
new file mode 100644
index 000000000000..d72f14f5fe63
--- /dev/null
+++ b/docs/system/devices/igb-migration.rst
@@ -0,0 +1,417 @@
+.. SPDX-License-Identifier: GPL-2.0-or-later
+.. _igb-migration:
+
+igb VF Migration
+----------------
+
+Live migration of VFIO-passthrough devices (SR-IOV VFs, vGPUs) is a
+growing requirement, but real hardware with migration support is scarce
+and hard to debug. An emulated device provides a fully controlled
+testbed for developing and validating the entire software stack --
+vfio-pci variant drivers, VFIO core migration v2 framework, QEMU,
+libvirt -- and for tuning complex migration policies such as downtime
+convergence. It also serves as an educational reference for
+understanding VFIO migration end-to-end, from device state
+serialization to dirty page tracking.
+
+The igb device supports an experimental VF migration interface that allows
+the `igb-vfio-pci`_ variant driver to migrate VF state during live
+migration using the standard VFIO migration v2 protocol with stop-copy
+and pre-copy support.
+
+This is enabled with the ``x-vf-migration`` property::
+
+ -device igb,x-vf-migration=on,...
+
+Each emulated VF then advertises a DVSEC discovered by the
+`igb-vfio-pci`_ variant driver at bind time. This feature is
+experimental (``x-`` prefix, default off).
+
+Architecture
+~~~~~~~~~~~~
+
+The target scenario is nested virtualization::
+
+ L0 QEMU
+ igb PF with x-vf-migration=on
+ └── VFs with migration DVSEC
+
+ L1 kernel
+ igb-vfio-pci variant driver
+ translates VFIO migration v2 ioctls → DVSEC config writes
+
+ L1 QEMU (stock, unmodified)
+ vfio-pci device model, standard migration fd
+
+ L2 guest
+ standard igbvf driver, unaware of migration
+
+The L1 QEMU is completely unmodified -- it sees a standard VFIO
+migratable device and uses the normal migration fd path. The
+`igb-vfio-pci`_ variant driver handles the translation between
+VFIO migration v2 ioctls and DVSEC config writes.
+
+Design
+~~~~~~
+
+The migration interface is exposed through a DVSEC at offset 0x160
+in VF extended config space (see `DVSEC register layout`_ below for
+the full register map).
+
+Device state is serialized as a versioned blob of per-VF register
+(offset, value) pairs covering control, interrupt, RX/TX queue,
+receive address (RA/RA2), etc. plus TX context descriptors and
+VFRE/VFTE enable bits. The buffer address is a guest physical address
+(GPA) written by the driver via ``virt_to_phys``; the device accesses
+guest RAM directly through the system address space.
+
+Dirty page tracking is implemented with per-range bitmaps maintained
+in IGBCore. All VF DMA paths in ``igb_core.c`` (TX data, RX data,
+descriptor writeback) are instrumented to record touched pages. The
+`igb-vfio-pci`_ variant driver registers tracked IOVA ranges and
+queries dirty bitmaps through a shared buffer. Buffer structures
+include len, flags, and reserved fields for future extensibility.
+
+The dirty bitmaps are maintained inside the device, which is not
+realistic for discrete NICs without on-chip DRAM.
+
+DVSEC register layout
+~~~~~~~~~~~~~~~~~~~~~
+
+The migration DVSEC (36 bytes at offset ``0x160``) uses a command
+doorbell model. All commands are synchronous -- the device completes
+the operation before the config write returns::
+
+ Offset Name Access Description
+ +0x00 ExtCap Hdr RO PCIe extended cap (id=0x23, ver=1)
+ +0x04 DVSEC Hdr 1 RO length[31:20] | rev[19:16] | vendor_id[15:0]
+ +0x08 DVSEC Hdr 2 RO DVSEC ID (1)
+ +0x0A Reserved - Padding for DWORD alignment
+ +0x0C CAPS RO F_STATE[0], F_DIRTY[1], max_ranges[11:8], pgsize[16:12]
+ +0x10 CTRL WO Doorbell: cmd[7:0], arg[31:8]
+ +0x14 STATUS RO state[7:0], error_code[15:8], quiesced[16]
+ +0x18 BUF_ADDR_LO RW Shared DMA buffer GPA (low 32 bits)
+ +0x1C BUF_ADDR_HI RW Shared DMA buffer GPA (high 32 bits)
+ +0x20 DATA_SIZE RO State blob size in bytes
+
+CTRL commands::
+
+ Cmd Name Arg Description
+ 1 SET_STATE state[31:8] Set migration state
+ 2 SAVE - DMA-write state to buffer
+ 3 LOAD size[31:8] DMA-read state from buffer
+ 4 DIRTY_ENABLE - Enable dirty tracking (params in DMA buffer)
+ 5 DIRTY_DISABLE - Disable dirty tracking
+ 6 DIRTY_QUERY - Query dirty bitmap (via DMA buffer)
+ 7 GET_STATS - Query statistics (via DMA buffer)
+
+The driver sets ``BUF_ADDR_LO/HI`` before issuing commands that use a
+DMA buffer (SAVE, LOAD, DIRTY_ENABLE, DIRTY_QUERY, GET_STATS). The
+buffer address is latched per CTRL write, so the driver can use
+different buffers for different commands. The buffer address is a
+guest physical address (GPA).
+
+State transitions follow the VFIO migration v2 state machine. The
+driver issues ``SET_STATE`` with the target state in the arg field and
+reads ``STATUS`` to confirm the transition. Device states::
+
+ 0 ERROR 1 STOP 2 RUNNING
+ 3 STOP_COPY 4 RESUMING 5 PRE_COPY
+
+``DATA_SIZE`` reflects the state blob size. At reset and in ``STOP``
+state it holds the maximum size the driver should allocate. After
+``SET_STATE(STOP_COPY)`` or ``SAVE`` it holds the actual serialized
+size. The driver reads it after entering ``STOP_COPY`` to allocate an
+exact-sized DMA buffer before issuing ``SAVE``.
+
+The state blob is a versioned sequence of register (offset, value)
+pairs with magic ``0x4D494742`` ("MIGB").
+
+When ``STATUS`` state is ``ERROR`` (0), bits [15:8] contain an error
+code identifying the failure::
+
+ 1 UNK_CMD Unknown CTRL command
+ 2 BAD_STATE Command issued in wrong migration state
+ 3 NO_BUFFER Command requires buffer but BUF_ADDR not set
+ 4 DMA_FAILED DMA transfer to/from buffer failed
+ 5 BAD_SIZE State blob too large or empty
+ 6 BAD_MAGIC State blob magic mismatch
+ 7 BAD_VERSION State blob version mismatch
+ 8 TOO_MANY_RANGES Exceeds max_ranges from CAPS
+ 9 BAD_RANGE Invalid range (zero size, misaligned, not contained)
+ 10 BAD_PGSIZE Invalid or misaligned page size
+ 11 NOT_ENABLED Dirty query without prior enable
+
+Dirty page tracking
+~~~~~~~~~~~~~~~~~~~
+
+The migration interface supports per-VF dirty page tracking, advertised
+by the ``F_DIRTY`` flag (bit 1) in ``CAPS``. This allows the variant
+driver to enter ``PRE_COPY`` state while the VM continues to run,
+iterating on dirty pages to reduce the final stop-and-copy window.
+
+The device maintains one dirty tracking engine per range, each with its
+own bitmap scoped to the range boundaries. The ``CAPS`` register
+advertises the maximum number of ranges in bits [11:8].
+
+Dirty tracking is controlled through CTRL commands:
+
+- **DIRTY_ENABLE** (4): the driver fills an ``igb_mig_dirty_enable_req``
+ struct in the DMA buffer with the page size, range IOVA, and range
+ size, then issues the command. The device allocates a bitmap for the
+ range and begins recording pages touched by DMA. Supported page sizes
+ are advertised in ``CAPS`` bits [16:12] (bit N = 2^N bytes). The driver
+ checks ``STATUS`` for errors after the command completes.
+- **DIRTY_DISABLE** (5): tears down all ranges and stops tracking.
+- **DIRTY_QUERY** (6): the driver writes (iova, size) into the
+ ``igb_mig_dirty_query`` DMA buffer, then issues the command. The
+ device validates the range, copies the dirty bitmap into the buffer,
+ and clears the tracked bits after successful DMA. The driver checks
+ ``STATUS`` for errors after the command.
+
+Dirty enable DMA buffer
+~~~~~~~~~~~~~~~~~~~~~~~
+
+The ``DIRTY_ENABLE`` command reads its parameters from the DMA buffer.
+The ``len`` field holds the total structure size (including reserved
+bytes) so the device can detect newer formats. ``flags`` and
+``reserved`` must be zero::
+
+ Offset Field Type Description
+ 0x00 len uint32 Structure size in bytes
+ 0x04 flags uint32 Reserved, must be 0
+ 0x08 pgsize uint64 Page granularity (must match a CAPS pgsize bit)
+ 0x10 range_iova uint64 Tracked range start address
+ 0x18 range_size uint64 Tracked range size in bytes
+ 0x20 reserved[4] uint32 Reserved, must be 0
+
+Dirty query DMA buffer
+~~~~~~~~~~~~~~~~~~~~~~
+
+The ``DIRTY_QUERY`` command uses a shared DMA buffer for both request
+and response. The ``len`` field holds the total buffer size (header +
+bitmap). ``flags`` and ``reserved`` must be zero::
+
+ Offset Field Written by Description
+ 0x00 len driver Total buffer size in bytes
+ 0x04 flags driver Reserved, must be 0
+ 0x08 iova driver Query range start
+ 0x10 size driver Query range size
+ 0x18 bitmap_size device Bytes written to bitmap
+ 0x1C dirty_page_count device Number of set bits
+ 0x20 dma_writes device DMA write count (diagnostic)
+ 0x28 reserved[6] - Reserved, must be 0
+ 0x40 bitmap[] device Dirty page bitmap
+
+Migration statistics
+~~~~~~~~~~~~~~~~~~~~
+
+The ``GET_STATS`` command DMA-writes a statistics response into the
+driver-provided buffer. The driver sets ``BUF_ADDR_LO/HI`` and issues
+the command; the device writes the response and returns::
+
+ Offset Field Type Description
+ 0x00 dma_writes uint64 DMA write operations tracked
+ 0x08 dma_bytes uint64 DMA bytes written
+ 0x10 dirty_pages_set uint32 Dirty pages marked since enable
+ 0x14 dirty_pages_cleared uint32 Dirty pages cleared by queries
+ 0x18 dirty_page_count uint32 Current dirty pages (set - cleared)
+ 0x1C dirty_query_count uint32 Number of QUERY operations
+
+The variant driver exposes these via debugfs at
+``/sys/kernel/debug/vfio/<device>/migration/dirty/stats``.
+
+Testing setup
+~~~~~~~~~~~~~
+
+The target scenario is nested virtualization: L0 runs QEMU with an
+igb PF (``x-vf-migration=on``), L1 runs the `igb-vfio-pci`_ variant
+driver and an unmodified QEMU, and L2 runs a standard igbvf driver.
+See `Architecture`_ above for the full stack diagram.
+
+NetworkManager configuration
+^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
+
+In a nested setup, the L1 VMs (source and destination) have emulated
+igb PFs connected to the L0 bridge. By default, NetworkManager
+acquires DHCP leases on those PF interfaces and on any igb VFs created
+later. This causes the VF MAC address to be learned on the L0 bridge,
+which can misdirect iperf3 traffic after migration.
+
+To prevent this, configure NetworkManager on **both L1 VMs** and on the
+**L2 guest disk image**.
+
+L1 VMs (source and destination)
+...............................
+
+1. Prevent NetworkManager from managing igbvf interfaces:
+
+.. code-block:: bash
+
+ cat > /etc/NetworkManager/conf.d/99-no-igbvf.conf <<EOF
+ [keyfile]
+ unmanaged-devices=driver:igbvf
+ EOF
+
+2. Disable IP on the igb PF connections (keep the interfaces UP for
+ bridging, but with no DHCP lease):
+
+.. code-block:: bash
+
+ # Identify the NM connections for the igb PFs (NOT the virtio management NIC)
+ nmcli -t -f NAME,DEVICE connection show
+
+ # For each igb PF connection:
+ nmcli connection modify "<igb-pf-connection>" ipv4.method disabled ipv6.method disabled
+
+3. Reload NetworkManager:
+
+.. code-block:: bash
+
+ nmcli general reload
+
+L2 guest disk image
+...................
+
+Use ``virt-customize`` to add the igbvf unmanaged config to the guest
+image (offline, before any test run):
+
+.. code-block:: bash
+
+ virt-customize -a /srv/migration/rhel10.qcow2 \
+ --write /etc/NetworkManager/conf.d/99-no-igbvf.conf:'[keyfile]
+ unmanaged-devices=driver:igbvf'
+
+Network diagram
+^^^^^^^^^^^^^^^
+
+The diagram below shows the nested setup where the source and
+destination hosts are themselves VMs (L1) running on a physical
+host (L0) that emulates the igb NIC::
+
+ ┌────────────────────────────────────────────────────────────────────────────┐
+ │ L0: physical host │
+ │ │
+ │ virbr0 192.168.199.1/24 │
+ │ ├── NFS server: /srv/migration │
+ │ └── iperf3 client: iperf3 -c 192.168.199.200 -t 60 -i 1 │
+ │ │ │
+ │ │ L0 virbr0 bridge (192.168.199.0/24) │
+ │ ────┼──────────┬──────────────────┬────────────────────────────── │
+ │ │ │ │ │
+ │ │ ┌────┴────┐ ┌────┴────┐ │
+ │ │ │ virtio │ │ emulated│ │
+ │ │ │c0:ff:ee:│ │ igb PF │ L0 QEMU (vm6) │
+ │ │ │ :00:06 │ │ + igbvf │ tracks DMA dirty pages │
+ │ │ └────┬────┘ └────┬────┘ │
+ │ │ │ │ │
+ │ ┌───┼──────────┼──────────────────┼──────────────────────────────────┐ │
+ │ │ │ L1: vm6 (source) │ │ │
+ │ │ │ enp1s0: 192.168.199.6 │ │ │
+ │ │ │ (management) │ │ │
+ │ │ │ enp8s0 (igb PF, no IP) │ │
+ │ │ │ │ │ │
+ │ │ │ igb VF0 ──► igb-vfio-pci (VFIO) │ │
+ │ │ │ │ dirty_sync → L0 igbvf │ │
+ │ │ │ │ │ │
+ │ │ │ virbr0 │ VFIO passthrough │ │
+ │ │ │ 192.168.200.1/24 │ │ │
+ │ │ │ │ │ │ │
+ │ │ │ ┌────┼─────────────────────┼───────────────────────────┐ │ │
+ │ │ │ │ │ L2: rhel10 guest │ │ │ │
+ │ │ │ │ │ │ │ │ │
+ │ │ │ │ virtio NIC igb VF (enp7s0) │ │ │
+ │ │ │ │ 192.168.200.130/24 192.168.199.200/24 │ │ │
+ │ │ │ │ (SSH login) (iperf3 data path) │ │ │
+ │ │ │ │ │ │ │ │
+ │ │ │ │ iperf3 -s -D │ (listens on 0.0.0.0) │ │ │
+ │ │ │ └──────────────────────────┼───────────────────────────┘ │ │
+ │ │ │ │ │ │
+ │ │ │ virsh migrate --live ──────┼──────────────────► vm7 │ │
+ │ │ │ │ │ │
+ │ └───┼─────────────────────────────┼──────────────────────────────────┘ │
+ │ │ │ │
+ │ │ iperf3 traffic │ │
+ │ └─────────────────────────────┘ │
+ │ │
+ │ ──────────────────────────────────────────────────────────────── │
+ │ │ │ │
+ │ │ ┌────────┐ │ ┌─────────┐ │
+ │ │ │ virtio │ │ │emulated │ L0 QEMU (vm7) │
+ │ │ │c0:ff:ee│ │ │ igb PF │ │
+ │ │ │ :00:07 │ │ │ + igbvf │ │
+ │ │ └────┬───┘ │ └────┬────┘ │
+ │ ┌──────────────┼───────┼────────┼────────────────────────────────────┐ │
+ │ │ L1: vm7 (destination) │ │ │
+ │ │ enp1s0: 192.168.199.7 │ │ │
+ │ │ (management) enp8s0 (igb PF, no IP) │ │
+ │ │ │ │ │
+ │ │ igb VF0 ──► igb-vfio-pci (VFIO) │ │
+ │ │ │ │ │
+ │ │ virbr0 │ VFIO passthrough │ │
+ │ │ 192.168.200.1/24 │ │ │
+ │ │ │ │ │ │
+ │ │ ┌────┼───────────────────────┼────────────────────────────┐ │ │
+ │ │ │ │ L2: rhel10 (after migration) │ │ │
+ │ │ │ │ │ │ │ │
+ │ │ │ virtio NIC igb VF (enp7s0) │ │ │
+ │ │ │ 192.168.200.130/24 192.168.199.200/24 │ │ │
+ │ │ │ │ │ │ │
+ │ │ │ iperf3 -s -D │ (connection survives) │ │ │
+ │ │ └────────────────────────────┼────────────────────────────┘ │ │
+ │ └───────────────────────────────┼────────────────────────────────────┘ │
+ │ │ │
+ │ iperf3 traffic resumes ─────┘ │
+ │ (same IP, same MAC, same L2 segment → transparent to client) │
+ └────────────────────────────────────────────────────────────────────────────┘
+
+Migration under iperf3 load works correctly: dirty page tracking
+converges (from ~2000 pages per PRE_COPY iteration down to ~280 at
+STOP_COPY), and STOP_COPY stays under 250ms.
+
+Todo
+~~~~
+
+1. Add migration blocker when ``x-vf-migration=on`` (no VMState yet) or
+ add VMState support for L0 migration (dirty bitmaps, tracking
+ engines, DVSEC registers, stats)
+2. Add PRE_COPY state transfer to validate device INIT data (magic,
+ version, etc.)
+3. Add qtests for migration state machine transitions, dirty page
+ tracking
+
+Ideas
+~~~~~
+
+1. **RX bandwidth throttle** (``x-mig-rx-limit``, uint32, default 0)
+
+ Return false from ``can_receive`` when the per-VF packet count in the
+ current tracking interval exceeds the limit. Reduces DMA writes and
+ dirty pages realistically.
+
+2. **Migration phase timing** (GET_STATS extension)
+
+ Add per-VF timestamps: ``precopy_start_ns``, ``stopcopy_start_ns``,
+ ``precopy_duration_ns``, ``stopcopy_duration_ns``,
+ ``state_transition_count``. Expose via GET_STATS.
+
+3. **Hot page simulation** (``x-mig-hot-pages``, uint32, default 0)
+
+ Re-set the first N bitmap bits after each DIRTY_QUERY, simulating
+ workloads with hot pages that prevent convergence.
+
+4. **Error injection** (``x-mig-inject-error``, uint32, default 0)
+
+ One-shot error code injection before command dispatch. A separate
+ ``x-mig-inject-dma-fail`` (bool) for persistent DMA failure testing.
+
+AI disclaimer
+~~~~~~~~~~~~~
+
+Claude was used to analyze the IGB PF and VF internal state and
+identify the pain points of a working live migration of such devices.
+The generated code served as a starting point but *significant* time
+was then spent cleaning up, reworking, and shaping it into a clear,
+reviewable IGB model extension.
+
+.. _igb-vfio-pci: https://github.com/legoater/vfio-pci-extras
diff --git a/docs/system/devices/igb.rst b/docs/system/devices/igb.rst
index 50f625fd77e4..00271dbc92c3 100644
--- a/docs/system/devices/igb.rst
+++ b/docs/system/devices/igb.rst
@@ -64,6 +64,12 @@ command:
pyvenv/bin/meson test --suite thorough func-x86_64-netdev_ethtool
+VF Migration (experimental)
+===========================
+
+See :ref:`igb-migration` for details on the experimental VF live migration
+interface.
+
References
==========
--
2.55.0
^ permalink raw reply related [flat|nested] 23+ messages in thread
* Re: [RFC PATCH v2 1/9] igb: Add x-vf-migration property and DVSEC extended capability
2026-09-02 19:20 ` [RFC PATCH v2 1/9] igb: Add x-vf-migration property and DVSEC extended capability Cédric Le Goater
@ 2026-09-03 19:57 ` Alex Williamson
2026-09-07 20:58 ` Cédric Le Goater
0 siblings, 1 reply; 23+ messages in thread
From: Alex Williamson @ 2026-09-03 19:57 UTC (permalink / raw)
To: Cédric Le Goater
Cc: qemu-devel, Akihiko Odaki, Sriram Yagnaraman, Jason Wang,
Peter Xu, alex
On Wed, 2 Sep 2026 21:20:46 +0200
Cédric Le Goater <clg@redhat.com> wrote:
> diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
> new file mode 100644
> index 000000000000..4dfebd82344c
> --- /dev/null
> +++ b/hw/net/igb_migration.c
> @@ -0,0 +1,106 @@
> +/*
> + * QEMU Intel 82576 SR/IOV VF Migration Support
> + *
> + * Copyright (c) 2026 Red Hat, Inc.
> + *
> + * SPDX-License-Identifier: GPL-2.0-or-later
> + */
> +
> +#include "qemu/osdep.h"
> +#include "hw/pci/pci_device.h"
> +#include "hw/pci/pcie.h"
> +#include "igb_common.h"
> +#include "igb_migration.h"
> +
> +static void igbvf_mig_update_status(IgbVfState *s, uint8_t err)
> +{
> + IgbVfMigState *ms = &s->mig;
> + PCIDevice *dev = PCI_DEVICE(s);
> + uint32_t status;
> +
> + status = ms->mig_state & IGB_MIG_STATUS_STATE_MASK;
> +
> + if (err) {
> + status = IGB_MIG_STATE_ERROR | IGB_MIG_STATUS_ERR(err);
> + }
> +
> + pci_set_long(dev->config + IGB_MIG_DVSEC_OFFSET + IGB_MIG_STATUS, status);
> +}
> +
> +
> +bool igbvf_add_migration_dvsec(PCIDevice *dev, Error **errp)
> +{
> + uint16_t offset = IGB_MIG_DVSEC_OFFSET;
> + uint32_t caps;
> +
> + pcie_add_capability(dev, PCI_EXT_CAP_ID_DVSEC, 1, offset,
> + IGB_MIG_DVSEC_SIZE);
> +
> + /* DVSEC header 1: length[31:20] | rev[19:16] | vendor_id[15:0] */
> + pci_set_long(dev->config + offset + 0x4,
> + (IGB_MIG_DVSEC_SIZE << 20) |
> + (IGB_MIG_DVSEC_VER << 16) |
> + PCI_VENDOR_ID_INTEL);
Let's not co-opt an Intel DVSEC ID, we should use a vendor ID that we
have a claim to, the RedHat/Qumranet one, I'd guess. We may need to
add DVSEC IDs to the spreadsheet for whoever is tracking Device IDs.
Gerd?
> +
> + /* DVSEC header 2: DVSEC ID */
> + pci_set_word(dev->config + offset + 0x8, IGB_MIG_DVSEC_ID);
> +
> + /* CAPS: features (state migration only) */
> + caps = IGB_MIG_CAP_F_STATE;
> + pci_set_long(dev->config + offset + IGB_MIG_CAPS, caps);
> +
> + /* STATUS: initial state is RUNNING */
> + pci_set_long(dev->config + offset + IGB_MIG_STATUS,
> + IGB_MIG_STATE_RUNNING);
> +
> + /* BUF_ADDR_LO and BUF_ADDR_HI are writable */
> + memset(dev->wmask + offset + IGB_MIG_BUF_ADDR_LO, 0xff, 4);
> + memset(dev->wmask + offset + IGB_MIG_BUF_ADDR_HI, 0xff, 4);
> +
> + return true;
> +}
...
> diff --git a/hw/net/igbvf.c b/hw/net/igbvf.c
> index 9a165c7063ee..30dfdb574ac7 100644
> --- a/hw/net/igbvf.c
> +++ b/hw/net/igbvf.c
> @@ -38,27 +38,21 @@
> */
>
> #include "qemu/osdep.h"
> +#include "qemu/range.h"
> #include "hw/core/hw-error.h"
> #include "hw/net/mii.h"
> #include "hw/pci/pci_device.h"
> #include "hw/pci/pcie.h"
> +#include "hw/pci/pcie_sriov.h"
> #include "hw/pci/msix.h"
> #include "net/eth.h"
> #include "net/net.h"
> #include "igb_common.h"
> #include "igb_core.h"
> +#include "igb_migration.h"
> #include "trace.h"
> #include "qapi/error.h"
>
> -OBJECT_DECLARE_SIMPLE_TYPE(IgbVfState, IGBVF)
> -
> -struct IgbVfState {
> - PCIDevice parent_obj;
> -
> - MemoryRegion mmio;
> - MemoryRegion msix;
> -};
> -
> static hwaddr vf_to_pf_addr(hwaddr addr, uint16_t vfn, bool write)
> {
> switch (addr) {
> @@ -199,10 +193,35 @@ static hwaddr vf_to_pf_addr(hwaddr addr, uint16_t vfn, bool write)
> return HWADDR_MAX;
> }
>
> +static bool igbvf_addr_in_dvsec(uint32_t addr, int len)
> +{
> + return ranges_overlap(addr, len,
> + IGB_MIG_DVSEC_OFFSET, IGB_MIG_DVSEC_SIZE);
> +}
> +
> +static uint32_t igbvf_read_config(PCIDevice *dev, uint32_t addr, int size)
> +{
> + IgbVfState *s = IGBVF(dev);
> +
> + if (s->migration_enabled && igbvf_addr_in_dvsec(addr, size)) {
> + return igbvf_mig_config_read(s, addr, size);
> + }
> +
> + return pci_default_read_config(dev, addr, size);
> +}
> +
> static void igbvf_write_config(PCIDevice *dev, uint32_t addr, uint32_t val,
> int len)
> {
> + IgbVfState *s = IGBVF(dev);
> +
> trace_igbvf_write_config(addr, val, len);
> +
> + if (s->migration_enabled && igbvf_addr_in_dvsec(addr, len)) {
> + igbvf_mig_config_write(s, addr, val, len);
> + return;
> + }
> +
> pci_default_write_config(dev, addr, val, len);
> if (object_property_get_bool(OBJECT(pcie_sriov_get_pf(dev)),
> "x-pcie-flr-init", &error_abort)) {
> @@ -282,13 +301,27 @@ static void igbvf_pci_realize(PCIDevice *dev, Error **errp)
> }
>
> pcie_ari_init(dev, 0x150);
> +
> + if (object_property_get_bool(OBJECT(pcie_sriov_get_pf(dev)),
> + "x-vf-migration", &error_abort)) {
> + s->vfn = pcie_sriov_vf_number(dev);
> + s->migration_enabled = true;
> + if (!igbvf_add_migration_dvsec(dev, errp)) {
> + return;
> + }
The migration blocker noted later in the docs should be added here.
Trivial to add, avoids the internal migration state being reset by the
L0 VM being migrated, doesn't seem worth extending this driver's VMState
while the feature is experimental.
> + }
> }
>
> static void igbvf_qdev_reset_hold(Object *obj, ResetType type)
> {
> PCIDevice *vf = PCI_DEVICE(obj);
> + IgbVfState *s = IGBVF(vf);
>
> igb_vf_reset(pcie_sriov_get_pf(vf), pcie_sriov_vf_number(vf));
> +
> + if (s->migration_enabled) {
> + igbvf_mig_state_reset(s);
> + }
Hmm, I think reset is more complicated that this. This seems to define
that any device reset will reset the migration state. The vfio
migration protocol only defines that a VFIO_DEVICE_RESET returns the
device to running.
Consider the case of an FLR triggered by the L2 guest. L1 QEMU passes
through the config space write, that lands in vfio-pci core config
space handling in the L1 variant driver, which turns into a
pci_reset_function() in the L1 kernel and I think lands here in the L0
QEMU. Therefore, it seems like the L2 guest can corrupt the migration
state.
As above, the migration state is defined to be reset via the RESET
ioctl, which also turns into a pci_reset_function() in the L1 kernel.
So L0 QEMU can't tell the difference here.
I think that means that the variant driver itself needs to own this
part of the protocol, performing the migration state housekeeping on
RESET ioctl, while both allow the state to persist on other resets. In
that sense QEMU cannot emulate this DVSEC as normal config space, it's
a persistent control plane that lives in the VMM and happens to be
accessed through config space. Thanks,
Alex
> }
>
> static void igbvf_pci_uninit(PCIDevice *dev)
> @@ -309,6 +342,7 @@ static void igbvf_class_init(ObjectClass *class, const void *data)
>
> c->realize = igbvf_pci_realize;
> c->exit = igbvf_pci_uninit;
> + c->config_read = igbvf_read_config;
> c->vendor_id = PCI_VENDOR_ID_INTEL;
> c->device_id = E1000_DEV_ID_82576_VF;
> c->revision = 1;
> diff --git a/hw/net/meson.build b/hw/net/meson.build
> index 84f142df222a..bb4b449b25ba 100644
> --- a/hw/net/meson.build
> +++ b/hw/net/meson.build
> @@ -11,7 +11,7 @@ system_ss.add(when: 'CONFIG_E1000_PCI', if_true: files('e1000.c', 'e1000x_common
> system_ss.add(when: 'CONFIG_E1000E_PCI_EXPRESS', if_true: files('net_tx_pkt.c', 'net_rx_pkt.c'))
> system_ss.add(when: 'CONFIG_E1000E_PCI_EXPRESS', if_true: files('e1000e.c', 'e1000e_core.c', 'e1000x_common.c'))
> system_ss.add(when: 'CONFIG_IGB_PCI_EXPRESS', if_true: files('net_tx_pkt.c', 'net_rx_pkt.c'))
> -system_ss.add(when: 'CONFIG_IGB_PCI_EXPRESS', if_true: files('igb.c', 'igbvf.c', 'igb_core.c'))
> +system_ss.add(when: 'CONFIG_IGB_PCI_EXPRESS', if_true: files('igb.c', 'igbvf.c', 'igb_core.c', 'igb_migration.c'))
> system_ss.add(when: 'CONFIG_RTL8139_PCI', if_true: files('rtl8139.c'))
> system_ss.add(when: 'CONFIG_TULIP', if_true: files('tulip.c'))
> system_ss.add(when: 'CONFIG_VMXNET3_PCI', if_true: files('net_tx_pkt.c', 'net_rx_pkt.c'))
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [RFC PATCH v2 1/9] igb: Add x-vf-migration property and DVSEC extended capability
2026-09-03 19:57 ` Alex Williamson
@ 2026-09-07 20:58 ` Cédric Le Goater
0 siblings, 0 replies; 23+ messages in thread
From: Cédric Le Goater @ 2026-09-07 20:58 UTC (permalink / raw)
To: Alex Williamson
Cc: qemu-devel, Akihiko Odaki, Sriram Yagnaraman, Jason Wang,
Peter Xu
On 9/3/26 21:57, Alex Williamson wrote:
> On Wed, 2 Sep 2026 21:20:46 +0200
> Cédric Le Goater <clg@redhat.com> wrote:
>
>> diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
>> new file mode 100644
>> index 000000000000..4dfebd82344c
>> --- /dev/null
>> +++ b/hw/net/igb_migration.c
>> @@ -0,0 +1,106 @@
>> +/*
>> + * QEMU Intel 82576 SR/IOV VF Migration Support
>> + *
>> + * Copyright (c) 2026 Red Hat, Inc.
>> + *
>> + * SPDX-License-Identifier: GPL-2.0-or-later
>> + */
>> +
>> +#include "qemu/osdep.h"
>> +#include "hw/pci/pci_device.h"
>> +#include "hw/pci/pcie.h"
>> +#include "igb_common.h"
>> +#include "igb_migration.h"
>> +
>> +static void igbvf_mig_update_status(IgbVfState *s, uint8_t err)
>> +{
>> + IgbVfMigState *ms = &s->mig;
>> + PCIDevice *dev = PCI_DEVICE(s);
>> + uint32_t status;
>> +
>> + status = ms->mig_state & IGB_MIG_STATUS_STATE_MASK;
>> +
>> + if (err) {
>> + status = IGB_MIG_STATE_ERROR | IGB_MIG_STATUS_ERR(err);
>> + }
>> +
>> + pci_set_long(dev->config + IGB_MIG_DVSEC_OFFSET + IGB_MIG_STATUS, status);
>> +}
>> +
>> +
>> +bool igbvf_add_migration_dvsec(PCIDevice *dev, Error **errp)
>> +{
>> + uint16_t offset = IGB_MIG_DVSEC_OFFSET;
>> + uint32_t caps;
>> +
>> + pcie_add_capability(dev, PCI_EXT_CAP_ID_DVSEC, 1, offset,
>> + IGB_MIG_DVSEC_SIZE);
>> +
>> + /* DVSEC header 1: length[31:20] | rev[19:16] | vendor_id[15:0] */
>> + pci_set_long(dev->config + offset + 0x4,
>> + (IGB_MIG_DVSEC_SIZE << 20) |
>> + (IGB_MIG_DVSEC_VER << 16) |
>> + PCI_VENDOR_ID_INTEL);
>
> Let's not co-opt an Intel DVSEC ID, we should use a vendor ID that we
> have a claim to, the RedHat/Qumranet one, I'd guess. We may need to
> add DVSEC IDs to the spreadsheet for whoever is tracking Device IDs.
> Gerd?
>
>> +
>> + /* DVSEC header 2: DVSEC ID */
>> + pci_set_word(dev->config + offset + 0x8, IGB_MIG_DVSEC_ID);
>> +
>> + /* CAPS: features (state migration only) */
>> + caps = IGB_MIG_CAP_F_STATE;
>> + pci_set_long(dev->config + offset + IGB_MIG_CAPS, caps);
>> +
>> + /* STATUS: initial state is RUNNING */
>> + pci_set_long(dev->config + offset + IGB_MIG_STATUS,
>> + IGB_MIG_STATE_RUNNING);
>> +
>> + /* BUF_ADDR_LO and BUF_ADDR_HI are writable */
>> + memset(dev->wmask + offset + IGB_MIG_BUF_ADDR_LO, 0xff, 4);
>> + memset(dev->wmask + offset + IGB_MIG_BUF_ADDR_HI, 0xff, 4);
>> +
>> + return true;
>> +}
> ...
>> diff --git a/hw/net/igbvf.c b/hw/net/igbvf.c
>> index 9a165c7063ee..30dfdb574ac7 100644
>> --- a/hw/net/igbvf.c
>> +++ b/hw/net/igbvf.c
>> @@ -38,27 +38,21 @@
>> */
>>
>> #include "qemu/osdep.h"
>> +#include "qemu/range.h"
>> #include "hw/core/hw-error.h"
>> #include "hw/net/mii.h"
>> #include "hw/pci/pci_device.h"
>> #include "hw/pci/pcie.h"
>> +#include "hw/pci/pcie_sriov.h"
>> #include "hw/pci/msix.h"
>> #include "net/eth.h"
>> #include "net/net.h"
>> #include "igb_common.h"
>> #include "igb_core.h"
>> +#include "igb_migration.h"
>> #include "trace.h"
>> #include "qapi/error.h"
>>
>> -OBJECT_DECLARE_SIMPLE_TYPE(IgbVfState, IGBVF)
>> -
>> -struct IgbVfState {
>> - PCIDevice parent_obj;
>> -
>> - MemoryRegion mmio;
>> - MemoryRegion msix;
>> -};
>> -
>> static hwaddr vf_to_pf_addr(hwaddr addr, uint16_t vfn, bool write)
>> {
>> switch (addr) {
>> @@ -199,10 +193,35 @@ static hwaddr vf_to_pf_addr(hwaddr addr, uint16_t vfn, bool write)
>> return HWADDR_MAX;
>> }
>>
>> +static bool igbvf_addr_in_dvsec(uint32_t addr, int len)
>> +{
>> + return ranges_overlap(addr, len,
>> + IGB_MIG_DVSEC_OFFSET, IGB_MIG_DVSEC_SIZE);
>> +}
>> +
>> +static uint32_t igbvf_read_config(PCIDevice *dev, uint32_t addr, int size)
>> +{
>> + IgbVfState *s = IGBVF(dev);
>> +
>> + if (s->migration_enabled && igbvf_addr_in_dvsec(addr, size)) {
>> + return igbvf_mig_config_read(s, addr, size);
>> + }
>> +
>> + return pci_default_read_config(dev, addr, size);
>> +}
>> +
>> static void igbvf_write_config(PCIDevice *dev, uint32_t addr, uint32_t val,
>> int len)
>> {
>> + IgbVfState *s = IGBVF(dev);
>> +
>> trace_igbvf_write_config(addr, val, len);
>> +
>> + if (s->migration_enabled && igbvf_addr_in_dvsec(addr, len)) {
>> + igbvf_mig_config_write(s, addr, val, len);
>> + return;
>> + }
>> +
>> pci_default_write_config(dev, addr, val, len);
>> if (object_property_get_bool(OBJECT(pcie_sriov_get_pf(dev)),
>> "x-pcie-flr-init", &error_abort)) {
>> @@ -282,13 +301,27 @@ static void igbvf_pci_realize(PCIDevice *dev, Error **errp)
>> }
>>
>> pcie_ari_init(dev, 0x150);
>> +
>> + if (object_property_get_bool(OBJECT(pcie_sriov_get_pf(dev)),
>> + "x-vf-migration", &error_abort)) {
>> + s->vfn = pcie_sriov_vf_number(dev);
>> + s->migration_enabled = true;
>> + if (!igbvf_add_migration_dvsec(dev, errp)) {
>> + return;
>> + }
>
> The migration blocker noted later in the docs should be added here.
> Trivial to add, avoids the internal migration state being reset by the
> L0 VM being migrated, doesn't seem worth extending this driver's VMState
> while the feature is experimental.
>
>> + }
>> }
>>
>> static void igbvf_qdev_reset_hold(Object *obj, ResetType type)
>> {
>> PCIDevice *vf = PCI_DEVICE(obj);
>> + IgbVfState *s = IGBVF(vf);
>>
>> igb_vf_reset(pcie_sriov_get_pf(vf), pcie_sriov_vf_number(vf));
>> +
>> + if (s->migration_enabled) {
>> + igbvf_mig_state_reset(s);
>> + }
>
> Hmm, I think reset is more complicated that this. This seems to define
> that any device reset will reset the migration state. The vfio
> migration protocol only defines that a VFIO_DEVICE_RESET returns the
> device to running.
>
> Consider the case of an FLR triggered by the L2 guest. L1 QEMU passes
> through the config space write, that lands in vfio-pci core config
> space handling in the L1 variant driver, which turns into a
> pci_reset_function() in the L1 kernel and I think lands here in the L0
> QEMU. Therefore, it seems like the L2 guest can corrupt the migration
> state.
>
> As above, the migration state is defined to be reset via the RESET
> ioctl, which also turns into a pci_reset_function() in the L1 kernel.
> So L0 QEMU can't tell the difference here.
>
> I think that means that the variant driver itself needs to own this
> part of the protocol, performing the migration state housekeeping on
> RESET ioctl, while both allow the state to persist on other resets. In
> that sense QEMU cannot emulate this DVSEC as normal config space, it's
> a persistent control plane that lives in the VMM and happens to be
> accessed through config space. Thanks,
Hi Alex,
You are right. The key principle: The DVSEC is a persistent control
plane living in the VMM, not device state. Device reset (VFIO ioctl
and L2 FLR) resets the device, not the mailbox. The variant driver
owns the reset, not the emulated device in L0 QEMU.
I missed two things : 1. transferring the DVSEC hiding (as it was done
for the mig BAR) and 2. reset of migration state.
Looking closer at it, the changes are small.
In QEMU:
The cold boot initialization sets all fields to defaults (zero). The
only change from today is that the initial state is ERROR, which means
migration isn't operational until the driver binds. The variant driver
activates the control plane by cycling the machine state to RUNNING.
The VF functional registers (queues, DMA rings, interrupts ... ) are
reset normally by igb_core_vf_reset(), and the migration control plane
(IgbVfMigState) persists in memory. It survives pci_do_device_reset()
and the variant driver sets the fields as needed through the normal
DVSEC command flow.
BUF_ADDR could be restored from the IgbVfMigState cache values since
it's a RW reg, but even that isn't a strong requirement.
So igbvf_mig_state_reset() is no longer needed in the reset path.
That's all for QEMU.
Driver :
Variant driver owns all the cleanup after reset: closes migration fds,
disables dirty tracking, reads and cycles the DVSEC state back to
RUNNING.
The first path, VFIO_DEVICE_RESET, needs a custom ioctl handler to run
the cleanup.
Second path, L2 FLR. To hide the DVSEC range from the L2 guest, the
driver implements custom VFIO config space reads and writes. The
config write handler can detect the FLR and implement the same cleanup
logic.
That's for v3.
Thanks,
C.
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [RFC PATCH v2 3/9] igb: Add VF state serialization for live migration
2026-09-02 19:20 ` [RFC PATCH v2 3/9] igb: Add VF state serialization for live migration Cédric Le Goater
@ 2026-09-08 8:01 ` Akihiko Odaki
2026-09-15 7:47 ` Cédric Le Goater
0 siblings, 1 reply; 23+ messages in thread
From: Akihiko Odaki @ 2026-09-08 8:01 UTC (permalink / raw)
To: Cédric Le Goater, qemu-devel
Cc: Sriram Yagnaraman, Jason Wang, Alex Williamson, Peter Xu
On 2026/09/03 4:20, Cédric Le Goater wrote:
> Implement per-VF state serialization and deserialization for the SAVE
> and LOAD commands. The wire format consists of a header (magic,
> version, VF number, register count), per-VF register offset/value
> pairs from a whitelist, RA table entries owned by the VF, and TX
> queue contexts.
>
> Register offsets are relocated on load so a VF can migrate to a
> different VF number on the destination. RA pool ownership bits are
> swapped accordingly.
>
> AI-used-for: analysis, code (prototype)
> Signed-off-by: Cédric Le Goater <clg@redhat.com>
> ---
> hw/net/igb_core.h | 2 +
> hw/net/igb_migration.h | 2 +
> hw/net/igb.c | 5 +
> hw/net/igb_migration.c | 347 ++++++++++++++++++++++++++++++++++++++++-
> 4 files changed, 354 insertions(+), 2 deletions(-)
>
> diff --git a/hw/net/igb_core.h b/hw/net/igb_core.h
> index d70b54e318f1..60724e2824ab 100644
> --- a/hw/net/igb_core.h
> +++ b/hw/net/igb_core.h
> @@ -143,4 +143,6 @@ igb_receive_iov(IGBCore *core, const struct iovec *iov, int iovcnt);
> void
> igb_start_recv(IGBCore *core);
>
> +IGBCore *igb_pf_get_core(void *pf);
> +
> #endif
> diff --git a/hw/net/igb_migration.h b/hw/net/igb_migration.h
> index ea40ac65c54b..b2f601e74346 100644
> --- a/hw/net/igb_migration.h
> +++ b/hw/net/igb_migration.h
> @@ -73,6 +73,8 @@
> #define IGB_MIG_ERR_NO_BUFFER 3
> #define IGB_MIG_ERR_DMA_FAILED 4
> #define IGB_MIG_ERR_BAD_SIZE 5
> +#define IGB_MIG_ERR_BAD_MAGIC 6
> +#define IGB_MIG_ERR_BAD_VERSION 7
>
> /* Shared buffer constants */
> #define IGB_VF_STATE_MAX_SIZE 4096
> diff --git a/hw/net/igb.c b/hw/net/igb.c
> index 7268e5473fc3..f39f2bc3a04e 100644
> --- a/hw/net/igb.c
> +++ b/hw/net/igb.c
> @@ -133,6 +133,11 @@ void igb_vf_reset(void *opaque, uint16_t vfn)
> igb_core_vf_reset(&s->core, vfn);
> }
>
> +IGBCore *igb_pf_get_core(void *pf)
> +{
> + return &IGB(pf)->core;
> +}
> +
> static bool
> igb_io_get_reg_index(IGBState *s, uint32_t *idx)
> {
> diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
> index 8e7e6fac9b9b..c34035974620 100644
> --- a/hw/net/igb_migration.c
> +++ b/hw/net/igb_migration.c
> @@ -10,18 +10,241 @@
> #include "qemu/log.h"
> #include "hw/pci/pci_device.h"
> #include "hw/pci/pcie.h"
> +#include "net/eth.h"
> +#include "net/net.h"
> #include "igb_common.h"
> +#include "igb_core.h"
> #include "igb_migration.h"
> #include "system/address-spaces.h"
> #include "trace.h"
>
> +static IGBCore *igbvf_get_core(IgbVfState *s)
> +{
> + return igb_pf_get_core(pcie_sriov_get_pf(PCI_DEVICE(s)));
> +}
> +
> /*
> * Per-VF state serialization / deserialization
> */
>
> +#define IGB_MIG_BLOB_MAGIC 0x4D494742 /* "MIGB" */
> +#define IGB_MIG_BLOB_VERSION 1
> +
> +typedef struct IgbMigRegPair {
> + uint32_t offset;
> + uint32_t value;
> +} IgbMigRegPair;
> +
> +typedef struct IgbMigTxCtx {
> + uint32_t ctx_desc[8]; /* 2 × adv_tx_context_desc (4 dwords each) */
> + uint32_t first_cmd_type_len;
> + uint32_t first_olinfo_status;
> + uint32_t first;
> + uint32_t skip_cp;
> +} IgbMigTxCtx;
> +
> +#define IGB_VF_MAX_FIXED_REGS 64
> +#define IGB_VF_MAX_RA_REGS 48 /* (16 + 8) RA entries × 2 (RAL+RAH) */
> +
> +typedef struct IgbMigBlob {
> + uint32_t magic;
> + uint32_t version;
> + uint32_t vfn;
> + uint32_t num_regs;
> + IgbMigRegPair regs[IGB_VF_MAX_FIXED_REGS];
> + uint32_t num_ra;
> + IgbMigRegPair ra[IGB_VF_MAX_RA_REGS];
> + uint32_t num_tx_ctx;
> + IgbMigTxCtx tx_ctx[2];
> +} IgbMigBlob;
> +
> +#define IGB_MIG_BLOB_SIZE sizeof(IgbMigBlob)
> +
> +QEMU_BUILD_BUG_ON(IGB_MIG_BLOB_SIZE > IGB_VF_STATE_MAX_SIZE);
> +
> +/* Register offsets that constitute a VF's state slice */
> +static int igb_vf_reg_list(uint16_t vfn, uint32_t *offsets)
> +{
> + int n = 0;
> + int q0 = vfn;
> + int q1 = vfn + IGB_NUM_VM_POOLS;
> +
> + /* Per-VF control and interrupt registers */
> + offsets[n++] = E1000_PVTCTRL(vfn) >> 2;
> + offsets[n++] = E1000_PVTEICS(vfn) >> 2;
> + offsets[n++] = E1000_PVTEIMS(vfn) >> 2;
> + offsets[n++] = E1000_PVTEIMC(vfn) >> 2;
> + offsets[n++] = E1000_PVTEIAC(vfn) >> 2;
> + offsets[n++] = E1000_PVTEIAM(vfn) >> 2;
> + offsets[n++] = E1000_PVTEICR(vfn) >> 2;
> +
> + /* Per-VF statistics */
> + offsets[n++] = E1000_PVFGPRC(vfn) >> 2;
> + offsets[n++] = E1000_PVFGPTC(vfn) >> 2;
> + offsets[n++] = E1000_PVFGORC(vfn) >> 2;
> + offsets[n++] = E1000_PVFGOTC(vfn) >> 2;
> + offsets[n++] = E1000_PVFMPRC(vfn) >> 2;
> + offsets[n++] = E1000_PVFGPRLBC(vfn) >> 2;
> + offsets[n++] = E1000_PVFGPTLBC(vfn) >> 2;
> + offsets[n++] = E1000_PVFGORLBC(vfn) >> 2;
> + offsets[n++] = E1000_PVFGOTLBC(vfn) >> 2;
> +
> + /*
> + * Mailbox control registers only - the 16-dword payload buffer
> + * (VMBMEM) is transient and drained on quiesce.
> + */
> + offsets[n++] = E1000_V2PMAILBOX(vfn) >> 2;
> + offsets[n++] = E1000_P2VMAILBOX(vfn) >> 2;
> +
> + /* Per-VF config */
> + offsets[n++] = E1000_VMOLR(vfn) >> 2;
> + offsets[n++] = E1000_VMVIR(vfn) >> 2;
> + offsets[n++] = E1000_PSRTYPE(vfn) >> 2;
> +
> + /*
> + * VF receive addresses (RA/RA2) are saved dynamically in
> + * igb_core_vf_save_state by scanning for entries whose pool
> + * bits match this VF - the PF driver chooses the RA slot.
> + */
> +
> + /* Interrupt routing */
> + offsets[n++] = (E1000_VTIVAR + vfn * 4) >> 2;
> + offsets[n++] = (E1000_VTIVAR_MISC + vfn * 4) >> 2;
> +
> + /*
> + * EITR (Extended Interrupt Throttle Register) - 3 vectors per VF.
> + * Each VF has 3 MSI-X vectors, each with its own EITR controlling
> + * interrupt coalescing. Without saving these, interrupt
> + * throttling resets to zero after migration which can cause
> + * interrupt storms or latency changes. VF N uses PF EITR indices
> + * (22 - N*3) .. (24 - N*3).
> + */
> + {
> + int eitr_base = 22 - vfn * 3;
> + offsets[n++] = E1000_EITR(eitr_base) >> 2;
> + offsets[n++] = E1000_EITR(eitr_base + 1) >> 2;
> + offsets[n++] = E1000_EITR(eitr_base + 2) >> 2;
> + }
> +
> + /* RX and TX queue registers for queues q0 and q1 */
> +#define ADD_QUEUE_REGS(q) do { \
> + offsets[n++] = E1000_RDBAL(q) >> 2; \
> + offsets[n++] = E1000_RDBAH(q) >> 2; \
> + offsets[n++] = E1000_RDLEN(q) >> 2; \
> + offsets[n++] = E1000_SRRCTL(q) >> 2; \
> + offsets[n++] = E1000_RDH(q) >> 2; \
> + offsets[n++] = E1000_RDT(q) >> 2; \
> + offsets[n++] = E1000_RXDCTL(q) >> 2; \
> + offsets[n++] = E1000_RXCTL(q) >> 2; \
> + offsets[n++] = E1000_RQDPC(q) >> 2; \
> + offsets[n++] = E1000_TDBAL(q) >> 2; \
> + offsets[n++] = E1000_TDBAH(q) >> 2; \
> + offsets[n++] = E1000_TDLEN(q) >> 2; \
> + offsets[n++] = E1000_TDH(q) >> 2; \
> + offsets[n++] = E1000_TDT(q) >> 2; \
> + offsets[n++] = E1000_TXDCTL(q) >> 2; \
> + offsets[n++] = E1000_TXCTL(q) >> 2; \
> + offsets[n++] = E1000_TDWBAL(q) >> 2; \
> + offsets[n++] = E1000_TDWBAH(q) >> 2; \
> +} while (0)
> +
> + ADD_QUEUE_REGS(q0);
> + ADD_QUEUE_REGS(q1);
> +#undef ADD_QUEUE_REGS
> +
> + g_assert(n <= IGB_VF_MAX_FIXED_REGS);
> + return n;
> +}
> +
> +/*
> + * Scan RA and RA2 arrays for receive address entries assigned to
> + * this VF. The PF driver picks the RA slot, so we cannot use a
> + * fixed index - instead check each entry's pool bits.
> + */
> +static int igb_core_vf_save_ra(IGBCore *core, uint16_t vfn,
> + IgbMigRegPair *regs)
> +{
> + uint32_t vf_pool_bit = E1000_RAH_POOL_1 << vfn;
> + int n = 0;
> + static const struct {
> + uint32_t base;
> + int count;
> + } ra_banks[] = {
> + { RA, 16 },
> + { RA2, 8 },
> + };
> +
> + for (int i = 0; i < ARRAY_SIZE(ra_banks); i++) {
> + for (int j = 0; j < ra_banks[i].count; j++) {
> + uint32_t ral_off = ra_banks[i].base + j * 2;
> + uint32_t rah_off = ra_banks[i].base + j * 2 + 1;
> + uint32_t rah_val = core->mac[rah_off];
> +
> + if ((rah_val & E1000_RAH_AV) && (rah_val & vf_pool_bit)) {
> + regs[n].offset = cpu_to_le32(ral_off);
> + regs[n].value = cpu_to_le32(core->mac[ral_off]);
> + n++;
> + regs[n].offset = cpu_to_le32(rah_off);
> + regs[n].value = cpu_to_le32(rah_val);
> + n++;
> + }
> + }
> + }
> + return n;
> +}
> +
> +static void igb_core_vf_save_tx_ctx(IGBCore *core, int queue,
> + IgbMigTxCtx *tx)
> +{
> + struct igb_tx *src = &core->tx[queue];
> +
> + memcpy(tx->ctx_desc, src->ctx, sizeof(tx->ctx_desc));
> + tx->first_cmd_type_len = cpu_to_le32(src->first_cmd_type_len);
> + tx->first_olinfo_status = cpu_to_le32(src->first_olinfo_status);
> + tx->first = cpu_to_le32(src->first);
> + tx->skip_cp = cpu_to_le32(src->skip_cp);
> +}
> +
> static int igb_core_vf_save_state(IgbVfState *s, void *buf, size_t buf_size)
> {
> - int size = 0;
> + int size = IGB_MIG_BLOB_SIZE;
> + IGBCore *core = igbvf_get_core(s);
> + IgbMigBlob *blob = buf;
> + uint32_t offsets[IGB_VF_MAX_FIXED_REGS];
> + int num_regs;
> + int q0 = s->vfn;
> + int q1 = s->vfn + IGB_NUM_VM_POOLS;
> +
> + /*
> + * Save PVT shadow registers (PVTEIMS/PVTEIAC/PVTEIAM) instead of
> + * extracting from PF aggregates - the L1 PF driver may have
> + * transiently cleared EIMS via EIMC. The load path ORs them back.
> + */
> + num_regs = igb_vf_reg_list(s->vfn, offsets);
> +
> + if (!buf) {
> + return size;
> + }
> +
> + if (size > buf_size) {
> + return -IGB_MIG_ERR_BAD_SIZE;
> + }
> +
> + blob->magic = cpu_to_le32(IGB_MIG_BLOB_MAGIC);
> + blob->version = cpu_to_le32(IGB_MIG_BLOB_VERSION);
> + blob->vfn = cpu_to_le32(s->vfn);
> +
> + blob->num_regs = cpu_to_le32(num_regs);
> + for (int i = 0; i < num_regs; i++) {
> + blob->regs[i].offset = cpu_to_le32(offsets[i]);
> + blob->regs[i].value = cpu_to_le32(core->mac[offsets[i]]);
> + }
> +
> + blob->num_ra = cpu_to_le32(igb_core_vf_save_ra(core, s->vfn, blob->ra));
The serialized state contains RA entries but excludes VFTA/VLVF and
MTA/UTA state. A guest-configured VLAN therefore resumes on a
destination lacking its filter membership: packets are rejected by the
global VLAN check or have their VF queue removed in
igb_receive_assign(),. Likewise, restored VMOLR.ROMPE cannot receive
subscribed multicast without its MTA hash bits. These are
guest-requested settings communicated through the PF mailbox, and the
resumed guest does not recreate them automatically.
> +
> + blob->num_tx_ctx = cpu_to_le32(2);
> + igb_core_vf_save_tx_ctx(core, q0, &blob->tx_ctx[0]);
> + igb_core_vf_save_tx_ctx(core, q1, &blob->tx_ctx[1]);
>
> trace_igbvf_mig_save_state(s->vfn, size);
> return size;
> @@ -29,11 +252,131 @@ static int igb_core_vf_save_state(IgbVfState *s, void *buf, size_t buf_size)
>
> static int igb_core_vf_max_data_size(IgbVfState *s)
> {
> - return sizeof(s->mig.mig_data);
> + int size = igb_core_vf_save_state(s, NULL, 0);
> +
> + g_assert(size > 0 && size <= IGB_VF_STATE_MAX_SIZE);
> + return size;
> +}
> +
> +static void igb_core_vf_load_tx_ctx(IGBCore *core, int queue,
> + const IgbMigTxCtx *tx)
> +{
> + struct igb_tx *dst = &core->tx[queue];
> +
> + /*
> + * Preserve the destination's tx_pkt - it's a host-side object,
> + * not guest state
> + */
> + memcpy(dst->ctx, tx->ctx_desc, sizeof(dst->ctx));
> + dst->first_cmd_type_len = le32_to_cpu(tx->first_cmd_type_len);
> + dst->first_olinfo_status = le32_to_cpu(tx->first_olinfo_status);
> + dst->first = le32_to_cpu(tx->first);
> + dst->skip_cp = le32_to_cpu(tx->skip_cp);
> +}
> +
> +static uint32_t igb_vf_relocate_offset(uint32_t offset,
> + const uint32_t *src_offsets,
> + const uint32_t *dst_offsets,
> + int num_offsets)
> +{
> + for (int i = 0; i < num_offsets; i++) {
> + if (src_offsets[i] == offset) {
> + return dst_offsets[i];
> + }
> + }
> + return 0;
> }
>
> static int igb_core_vf_load_state(IgbVfState *s, const void *buf, size_t size)
> {
> + IGBCore *core = igbvf_get_core(s);
> + uint32_t src_offsets[IGB_VF_MAX_FIXED_REGS];
> + uint32_t dst_offsets[IGB_VF_MAX_FIXED_REGS];
> + int q0 = s->vfn;
> + int q1 = s->vfn + IGB_NUM_VM_POOLS;
> +
> + if (size < IGB_MIG_BLOB_SIZE) {
> + return -IGB_MIG_ERR_BAD_SIZE;
> + }
> +
> + const IgbMigBlob *blob = buf;
> +
> + uint32_t magic = le32_to_cpu(blob->magic);
> + uint32_t version = le32_to_cpu(blob->version);
> + uint32_t saved_vfn = le32_to_cpu(blob->vfn);
> + uint32_t num_regs = le32_to_cpu(blob->num_regs);
> +
> + if (magic != IGB_MIG_BLOB_MAGIC) {
> + return -IGB_MIG_ERR_BAD_MAGIC;
> + }
> + if (version != IGB_MIG_BLOB_VERSION) {
> + return -IGB_MIG_ERR_BAD_VERSION;
> + }
> + if (num_regs > IGB_VF_MAX_FIXED_REGS) {
> + return -IGB_MIG_ERR_BAD_SIZE;
> + }
> +
> + uint32_t num_ra = le32_to_cpu(blob->num_ra);
> + if (num_ra > IGB_VF_MAX_RA_REGS) {
> + return -IGB_MIG_ERR_BAD_SIZE;
> + }
> +
> + int num_offsets = igb_vf_reg_list(saved_vfn, src_offsets);
> + igb_vf_reg_list(s->vfn, dst_offsets);
> +
> + for (uint32_t i = 0; i < num_regs; i++) {
> + uint32_t src_off = le32_to_cpu(blob->regs[i].offset);
> + uint32_t value = le32_to_cpu(blob->regs[i].value);
> + uint32_t offset = igb_vf_relocate_offset(src_off,
> + src_offsets, dst_offsets, num_offsets);
> + if (!offset) {
> + return -IGB_MIG_ERR_BAD_SIZE;
> + }
> +
> + core->mac[offset] = value;
> +
> + /*
> + * Sync EITR to eitr_guest_value[] shadow array, stripping
> + * E1000_EITR_CNT_IGNR so guest register readback returns the
> + * correct value.
> + */
> + if (offset >= EITR0 && offset < EITR0 + IGB_INTR_NUM) {
> + core->eitr_guest_value[offset - EITR0] =
> + value & ~E1000_EITR_CNT_IGNR;
> + }
> + }
> +
> + /*
> + * MSI-X table/PBA is not saved - L1's VFIO reprograms it with
> + * destination-specific IRTE references after migration.
> + */
> +
> + uint32_t src_pool = E1000_RAH_POOL_1 << saved_vfn;
> + uint32_t dst_pool = E1000_RAH_POOL_1 << s->vfn;
> +
> + for (uint32_t i = 0; i < num_ra; i++) {
> + uint32_t offset = le32_to_cpu(blob->ra[i].offset);
> + uint32_t value = le32_to_cpu(blob->ra[i].value);
> +
> + /* RAH entries: swap pool ownership bits */
> + if (offset >= RA && offset < RA + 32 && (offset - RA) % 2 == 1) {
> + value = (value & ~src_pool) | dst_pool;
> + }
> + if (offset >= RA2 && offset < RA2 + 16 && (offset - RA2) % 2 == 1) {
> + value = (value & ~src_pool) | dst_pool;
> + }
> +
> + core->mac[offset] = value;
Please validate RA offsets before writing core->mac. A valid-sized LOAD
blob with one RA entry and an out-of-range offset produces an
out-of-bounds 32-bit host write. LOAD reads this blob directly from
guest RAM. Please validate all offsets and saved_vfn before modifying state.
This loop also changes the VF pool bit but preserves the source RA slot.
A valid VF0→VF1 migration onto a destination already using VF0
overwrites destination VF0’s primary MAC entry. Linux assigns the
primary entry as rar_entry_count - (vf + 1); receive filtering consumes
these slots directly in igb_receive_assign(). Thus the unrelated
destination VF loses unicast reception. Existing destination entries
owned by the migrating VF are also left behind.
Regards,
Akihiko Odaki
> + }
> +
> + uint32_t num_tx = le32_to_cpu(blob->num_tx_ctx);
> + if (num_tx != 2) {
> + return -IGB_MIG_ERR_BAD_SIZE;
> + }
> +
> + igb_core_vf_load_tx_ctx(core, q0, &blob->tx_ctx[0]);
> + igb_core_vf_load_tx_ctx(core, q1, &blob->tx_ctx[1]);
> +
> trace_igbvf_mig_load_state(s->vfn, (uint32_t)size);
> return 0;
> }
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [RFC PATCH v2 4/9] igb: Add VF post-load fixups for live migration
2026-09-02 19:20 ` [RFC PATCH v2 4/9] igb: Add VF post-load fixups " Cédric Le Goater
@ 2026-09-08 8:10 ` Akihiko Odaki
2026-09-15 7:58 ` Cédric Le Goater
0 siblings, 1 reply; 23+ messages in thread
From: Akihiko Odaki @ 2026-09-08 8:10 UTC (permalink / raw)
To: Cédric Le Goater, qemu-devel
Cc: Sriram Yagnaraman, Jason Wang, Alex Williamson, Peter Xu
On 2026/09/03 4:20, Cédric Le Goater wrote:
> After restoring per-VF register state, propagate the VF's PVT shadow
> values back into the PF's aggregate EIMS/EIAC/EIAM registers and
> re-apply the VTIVAR interrupt vector routing to the shared IVAR0.
>
> AI-used-for: analysis, code (prototype)
> Signed-off-by: Cédric Le Goater <clg@redhat.com>
> ---
> hw/net/igb_core.h | 3 ++
> hw/net/igb_core.c | 66 ++++++++++++++++++++++++++++++++++++++++++
> hw/net/igb_migration.c | 7 +++++
> 3 files changed, 76 insertions(+)
>
> diff --git a/hw/net/igb_core.h b/hw/net/igb_core.h
> index 60724e2824ab..22e10e4e6d0b 100644
> --- a/hw/net/igb_core.h
> +++ b/hw/net/igb_core.h
> @@ -145,4 +145,7 @@ igb_start_recv(IGBCore *core);
>
> IGBCore *igb_pf_get_core(void *pf);
>
> +void igb_core_vf_propagate_irqs(IGBCore *core, uint16_t vfn);
> +void igb_core_vf_propagate_ivar(IGBCore *core, uint16_t vfn);
> +
> #endif
> diff --git a/hw/net/igb_core.c b/hw/net/igb_core.c
> index 2a4883907353..01745fe756d0 100644
> --- a/hw/net/igb_core.c
> +++ b/hw/net/igb_core.c
> @@ -4552,3 +4552,69 @@ igb_core_post_load(IGBCore *core)
>
> return 0;
> }
> +
> +/*
> + * Propagate VF interrupt state to PF aggregates after loading VF
> + * registers. The load path writes directly to mac[] bypassing the
> + * register handlers that OR VF bits into EIMS/EIAC/EIAM. Also clear
> + * stale VF bits in EICR that may have been set by packets arriving
> + * between PF vmstate restore and VF state load.
> + */
> +void igb_core_vf_propagate_irqs(IGBCore *core, uint16_t vfn)
> +{
> + uint32_t shift = 22 - vfn * IGBVF_MSIX_VEC_NUM;
> + uint32_t vf_mask = 0x7 << shift;
> + uint32_t pvt_idx;
> +
> + core->mac[EIMS] &= ~vf_mask;
Restore the effective interrupt mask rather than the last EIMS write.
This clears the destination VF mask and copies PVTEIMS, but that shadow
records only the last write: igb_set_vteims() assigns it before applying
set-bit semantics to aggregate EIMS. For example, source writes EIMS=1,
then EIMS=2, leave effective mask 3 and shadow 2; migration restores 2,
suppressing vector 0 until explicitly enabled again. Conversely, EIMC
updates aggregate EIMS without updating PVTEIMS, so migration reenables
explicitly masked vectors. Patch 7 repeats this fixup on unquiesce
without repairing the bookkeeping.
Regards,
Akihiko Odaki
> + pvt_idx = PVTEIMS0 + vfn * 0x40;
> + core->mac[EIMS] |= (core->mac[pvt_idx] & 0x7) << shift;
> +
> + core->mac[EIAC] &= ~vf_mask;
> + pvt_idx = PVTEIAC0 + vfn * 0x40;
> + core->mac[EIAC] |= (core->mac[pvt_idx] & 0x7) << shift;
> +
> + core->mac[EIAM] &= ~vf_mask;
> + pvt_idx = PVTEIAM0 + vfn * 0x40;
> + core->mac[EIAM] |= (core->mac[pvt_idx] & 0x7) << shift;
> +
> + core->mac[EICR] &= ~vf_mask;
> +}
> +
> +/*
> + * Re-apply VTIVAR -> IVAR0 interrupt routing. The L1 PF driver
> + * may have overwritten the shared IVAR0 entries with its own
> + * queue routing after L0 vmstate restore.
> + */
> +void igb_core_vf_propagate_ivar(IGBCore *core, uint16_t vfn)
> +{
> + uint32_t vtivar = core->mac[VTIVAR + vfn];
> + int n;
> + uint8_t ent;
> + uint32_t mask;
> +
> + n = igb_ivar_entry_rx(vfn);
> + mask = 0xffU << (8 * (n % 4));
> + if (vtivar & E1000_IVAR_VALID) {
> + ent = E1000_IVAR_VALID |
> + (24 - vfn * IGBVF_MSIX_VEC_NUM - (2 - (vtivar & 0x7)));
> + core->mac[IVAR0 + n / 4] =
> + (core->mac[IVAR0 + n / 4] & ~mask) |
> + ((uint32_t)ent << (8 * (n % 4)));
> + } else {
> + core->mac[IVAR0 + n / 4] &= ~mask;
> + }
> +
> + n = igb_ivar_entry_tx(vfn);
> + mask = 0xffU << (8 * (n % 4));
> + ent = vtivar >> 8;
> + if (ent & E1000_IVAR_VALID) {
> + ent = E1000_IVAR_VALID |
> + (24 - vfn * IGBVF_MSIX_VEC_NUM - (2 - (ent & 0x7)));
> + core->mac[IVAR0 + n / 4] =
> + (core->mac[IVAR0 + n / 4] & ~mask) |
> + ((uint32_t)ent << (8 * (n % 4)));
> + } else {
> + core->mac[IVAR0 + n / 4] &= ~mask;
> + }
> +}
> diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
> index c34035974620..af7257ccc4fd 100644
> --- a/hw/net/igb_migration.c
> +++ b/hw/net/igb_migration.c
> @@ -384,12 +384,19 @@ static int igb_core_vf_load_state(IgbVfState *s, const void *buf, size_t size)
> static int igbvf_mig_load(IgbVfState *s, const void *buf, size_t size)
> {
> int ret;
> + IGBCore *core = igbvf_get_core(s);
>
> ret = igb_core_vf_load_state(s, buf, size);
> if (ret < 0) {
> return ret;
> }
>
> + /*
> + * Post-load: sync VF interrupt and routing state to PF aggregates
> + */
> + igb_core_vf_propagate_irqs(core, s->vfn);
> + igb_core_vf_propagate_ivar(core, s->vfn);
> +
> return 0;
> }
>
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [RFC PATCH v2 9/9] docs: Add igb VF migration testing setup guide
2026-09-02 19:20 ` [RFC PATCH v2 9/9] docs: Add igb VF migration testing setup guide Cédric Le Goater
@ 2026-09-08 8:39 ` Akihiko Odaki
2026-09-15 7:59 ` Cédric Le Goater
0 siblings, 1 reply; 23+ messages in thread
From: Akihiko Odaki @ 2026-09-08 8:39 UTC (permalink / raw)
To: Cédric Le Goater, qemu-devel
Cc: Sriram Yagnaraman, Jason Wang, Alex Williamson, Peter Xu
On 2026/09/03 4:20, Cédric Le Goater wrote:
> Document the igb VF migration interface: DVSEC register layout, dirty
> page tracking, testing setup with nested virtualization.
>
> AI-used-for: docs
> Signed-off-by: Cédric Le Goater <clg@redhat.com>
> ---
> MAINTAINERS | 1 +
> docs/system/device-emulation.rst | 1 +
> docs/system/devices/igb-migration.rst | 417 ++++++++++++++++++++++++++
> docs/system/devices/igb.rst | 6 +
> 4 files changed, 425 insertions(+)
> create mode 100644 docs/system/devices/igb-migration.rst
>
> diff --git a/MAINTAINERS b/MAINTAINERS
> index f88b526be238..4a2f1357e936 100644
> --- a/MAINTAINERS
> +++ b/MAINTAINERS
> @@ -2820,6 +2820,7 @@ igb VF migration
> M: Cédric Le Goater <clg@redhat.com>
> S: Maintained
> F: hw/net/igb_migration.*
> +F: docs/system/devices/igb-migration.rst
>
> eepro100
> M: Stefan Weil <sw@weilnetz.de>
> diff --git a/docs/system/device-emulation.rst b/docs/system/device-emulation.rst
> index 40054bb7dfcc..75f423b795cd 100644
> --- a/docs/system/device-emulation.rst
> +++ b/docs/system/device-emulation.rst
> @@ -90,6 +90,7 @@ Emulated Devices
> devices/cxl.rst
> devices/emmc.rst
> devices/igb.rst
> + devices/igb-migration.rst
> devices/ivshmem-flat.rst
> devices/ivshmem.rst
> devices/keyboard.rst
> diff --git a/docs/system/devices/igb-migration.rst b/docs/system/devices/igb-migration.rst
> new file mode 100644
> index 000000000000..d72f14f5fe63
> --- /dev/null
> +++ b/docs/system/devices/igb-migration.rst
> @@ -0,0 +1,417 @@
> +.. SPDX-License-Identifier: GPL-2.0-or-later
> +.. _igb-migration:
> +
> +igb VF Migration
> +----------------
> +
> +Live migration of VFIO-passthrough devices (SR-IOV VFs, vGPUs) is a
> +growing requirement, but real hardware with migration support is scarce
> +and hard to debug. An emulated device provides a fully controlled
> +testbed for developing and validating the entire software stack --
> +vfio-pci variant drivers, VFIO core migration v2 framework, QEMU,
> +libvirt -- and for tuning complex migration policies such as downtime
> +convergence. It also serves as an educational reference for
> +understanding VFIO migration end-to-end, from device state
> +serialization to dirty page tracking.
> +
> +The igb device supports an experimental VF migration interface that allows
> +the `igb-vfio-pci`_ variant driver to migrate VF state during live
> +migration using the standard VFIO migration v2 protocol with stop-copy
> +and pre-copy support.
> +
> +This is enabled with the ``x-vf-migration`` property::
> +
> + -device igb,x-vf-migration=on,...
> +
> +Each emulated VF then advertises a DVSEC discovered by the
> +`igb-vfio-pci`_ variant driver at bind time. This feature is
> +experimental (``x-`` prefix, default off).
> +
> +Architecture
> +~~~~~~~~~~~~
> +
> +The target scenario is nested virtualization::
> +
> + L0 QEMU
> + igb PF with x-vf-migration=on
> + └── VFs with migration DVSEC
> +
> + L1 kernel
> + igb-vfio-pci variant driver
> + translates VFIO migration v2 ioctls → DVSEC config writes
> +
> + L1 QEMU (stock, unmodified)
> + vfio-pci device model, standard migration fd
> +
> + L2 guest
> + standard igbvf driver, unaware of migration
> +
> +The L1 QEMU is completely unmodified -- it sees a standard VFIO
> +migratable device and uses the normal migration fd path. The
> +`igb-vfio-pci`_ variant driver handles the translation between
> +VFIO migration v2 ioctls and DVSEC config writes.
> +
> +Design
> +~~~~~~
> +
> +The migration interface is exposed through a DVSEC at offset 0x160
> +in VF extended config space (see `DVSEC register layout`_ below for
> +the full register map).
> +
> +Device state is serialized as a versioned blob of per-VF register
> +(offset, value) pairs covering control, interrupt, RX/TX queue,
> +receive address (RA/RA2), etc. plus TX context descriptors and
> +VFRE/VFTE enable bits. The buffer address is a guest physical address
> +(GPA) written by the driver via ``virt_to_phys``; the device accesses
> +guest RAM directly through the system address space.
> +
> +Dirty page tracking is implemented with per-range bitmaps maintained
> +in IGBCore. All VF DMA paths in ``igb_core.c`` (TX data, RX data,
> +descriptor writeback) are instrumented to record touched pages. The
> +`igb-vfio-pci`_ variant driver registers tracked IOVA ranges and
> +queries dirty bitmaps through a shared buffer. Buffer structures
> +include len, flags, and reserved fields for future extensibility.
> +
> +The dirty bitmaps are maintained inside the device, which is not
> +realistic for discrete NICs without on-chip DRAM.
> +
> +DVSEC register layout
> +~~~~~~~~~~~~~~~~~~~~~
> +
> +The migration DVSEC (36 bytes at offset ``0x160``) uses a command
> +doorbell model. All commands are synchronous -- the device completes
> +the operation before the config write returns::
> +
> + Offset Name Access Description
> + +0x00 ExtCap Hdr RO PCIe extended cap (id=0x23, ver=1)
> + +0x04 DVSEC Hdr 1 RO length[31:20] | rev[19:16] | vendor_id[15:0]
> + +0x08 DVSEC Hdr 2 RO DVSEC ID (1)
> + +0x0A Reserved - Padding for DWORD alignment
> + +0x0C CAPS RO F_STATE[0], F_DIRTY[1], max_ranges[11:8], pgsize[16:12]
> + +0x10 CTRL WO Doorbell: cmd[7:0], arg[31:8]
> + +0x14 STATUS RO state[7:0], error_code[15:8], quiesced[16]
> + +0x18 BUF_ADDR_LO RW Shared DMA buffer GPA (low 32 bits)
> + +0x1C BUF_ADDR_HI RW Shared DMA buffer GPA (high 32 bits)
> + +0x20 DATA_SIZE RO State blob size in bytes
> +
> +CTRL commands::
> +
> + Cmd Name Arg Description
> + 1 SET_STATE state[31:8] Set migration state
> + 2 SAVE - DMA-write state to buffer
> + 3 LOAD size[31:8] DMA-read state from buffer
> + 4 DIRTY_ENABLE - Enable dirty tracking (params in DMA buffer)
> + 5 DIRTY_DISABLE - Disable dirty tracking
> + 6 DIRTY_QUERY - Query dirty bitmap (via DMA buffer)
> + 7 GET_STATS - Query statistics (via DMA buffer)
> +
> +The driver sets ``BUF_ADDR_LO/HI`` before issuing commands that use a
> +DMA buffer (SAVE, LOAD, DIRTY_ENABLE, DIRTY_QUERY, GET_STATS). The
> +buffer address is latched per CTRL write, so the driver can use
> +different buffers for different commands. The buffer address is a
> +guest physical address (GPA).
> +
> +State transitions follow the VFIO migration v2 state machine. The
> +driver issues ``SET_STATE`` with the target state in the arg field and
> +reads ``STATUS`` to confirm the transition. Device states::
> +
> + 0 ERROR 1 STOP 2 RUNNING
> + 3 STOP_COPY 4 RESUMING 5 PRE_COPY
> +
> +``DATA_SIZE`` reflects the state blob size. At reset and in ``STOP``
> +state it holds the maximum size the driver should allocate. After
> +``SET_STATE(STOP_COPY)`` or ``SAVE`` it holds the actual serialized
> +size. The driver reads it after entering ``STOP_COPY`` to allocate an
> +exact-sized DMA buffer before issuing ``SAVE``.
> +
> +The state blob is a versioned sequence of register (offset, value)
> +pairs with magic ``0x4D494742`` ("MIGB").
> +
> +When ``STATUS`` state is ``ERROR`` (0), bits [15:8] contain an error
> +code identifying the failure::
> +
> + 1 UNK_CMD Unknown CTRL command
> + 2 BAD_STATE Command issued in wrong migration state
> + 3 NO_BUFFER Command requires buffer but BUF_ADDR not set
> + 4 DMA_FAILED DMA transfer to/from buffer failed
> + 5 BAD_SIZE State blob too large or empty
> + 6 BAD_MAGIC State blob magic mismatch
> + 7 BAD_VERSION State blob version mismatch
> + 8 TOO_MANY_RANGES Exceeds max_ranges from CAPS
> + 9 BAD_RANGE Invalid range (zero size, misaligned, not contained)
> + 10 BAD_PGSIZE Invalid or misaligned page size
> + 11 NOT_ENABLED Dirty query without prior enable
> +
> +Dirty page tracking
> +~~~~~~~~~~~~~~~~~~~
> +
> +The migration interface supports per-VF dirty page tracking, advertised
> +by the ``F_DIRTY`` flag (bit 1) in ``CAPS``. This allows the variant
> +driver to enter ``PRE_COPY`` state while the VM continues to run,
> +iterating on dirty pages to reduce the final stop-and-copy window.
> +
> +The device maintains one dirty tracking engine per range, each with its
> +own bitmap scoped to the range boundaries. The ``CAPS`` register
> +advertises the maximum number of ranges in bits [11:8].
> +
> +Dirty tracking is controlled through CTRL commands:
> +
> +- **DIRTY_ENABLE** (4): the driver fills an ``igb_mig_dirty_enable_req``
> + struct in the DMA buffer with the page size, range IOVA, and range
> + size, then issues the command. The device allocates a bitmap for the
> + range and begins recording pages touched by DMA. Supported page sizes
> + are advertised in ``CAPS`` bits [16:12] (bit N = 2^N bytes). The driver
> + checks ``STATUS`` for errors after the command completes.
> +- **DIRTY_DISABLE** (5): tears down all ranges and stops tracking.
> +- **DIRTY_QUERY** (6): the driver writes (iova, size) into the
> + ``igb_mig_dirty_query`` DMA buffer, then issues the command. The
> + device validates the range, copies the dirty bitmap into the buffer,
> + and clears the tracked bits after successful DMA. The driver checks
> + ``STATUS`` for errors after the command.
> +
> +Dirty enable DMA buffer
> +~~~~~~~~~~~~~~~~~~~~~~~
> +
> +The ``DIRTY_ENABLE`` command reads its parameters from the DMA buffer.
> +The ``len`` field holds the total structure size (including reserved
> +bytes) so the device can detect newer formats. ``flags`` and
> +``reserved`` must be zero::
> +
> + Offset Field Type Description
> + 0x00 len uint32 Structure size in bytes
> + 0x04 flags uint32 Reserved, must be 0
> + 0x08 pgsize uint64 Page granularity (must match a CAPS pgsize bit)
> + 0x10 range_iova uint64 Tracked range start address
> + 0x18 range_size uint64 Tracked range size in bytes
> + 0x20 reserved[4] uint32 Reserved, must be 0
> +
> +Dirty query DMA buffer
> +~~~~~~~~~~~~~~~~~~~~~~
> +
> +The ``DIRTY_QUERY`` command uses a shared DMA buffer for both request
> +and response. The ``len`` field holds the total buffer size (header +
> +bitmap). ``flags`` and ``reserved`` must be zero::
> +
> + Offset Field Written by Description
> + 0x00 len driver Total buffer size in bytes
> + 0x04 flags driver Reserved, must be 0
> + 0x08 iova driver Query range start
> + 0x10 size driver Query range size
> + 0x18 bitmap_size device Bytes written to bitmap
> + 0x1C dirty_page_count device Number of set bits
> + 0x20 dma_writes device DMA write count (diagnostic)
> + 0x28 reserved[6] - Reserved, must be 0
> + 0x40 bitmap[] device Dirty page bitmap
> +
> +Migration statistics
> +~~~~~~~~~~~~~~~~~~~~
> +
> +The ``GET_STATS`` command DMA-writes a statistics response into the
> +driver-provided buffer. The driver sets ``BUF_ADDR_LO/HI`` and issues
> +the command; the device writes the response and returns::
> +
> + Offset Field Type Description
> + 0x00 dma_writes uint64 DMA write operations tracked
> + 0x08 dma_bytes uint64 DMA bytes written
> + 0x10 dirty_pages_set uint32 Dirty pages marked since enable
> + 0x14 dirty_pages_cleared uint32 Dirty pages cleared by queries
> + 0x18 dirty_page_count uint32 Current dirty pages (set - cleared)
> + 0x1C dirty_query_count uint32 Number of QUERY operations
> +
> +The variant driver exposes these via debugfs at
> +``/sys/kernel/debug/vfio/<device>/migration/dirty/stats``.
> +
> +Testing setup
> +~~~~~~~~~~~~~
> +
> +The target scenario is nested virtualization: L0 runs QEMU with an
> +igb PF (``x-vf-migration=on``), L1 runs the `igb-vfio-pci`_ variant
> +driver and an unmodified QEMU, and L2 runs a standard igbvf driver.
> +See `Architecture`_ above for the full stack diagram.
> +
> +NetworkManager configuration
> +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
> +
> +In a nested setup, the L1 VMs (source and destination) have emulated
> +igb PFs connected to the L0 bridge. By default, NetworkManager
> +acquires DHCP leases on those PF interfaces and on any igb VFs created
> +later. This causes the VF MAC address to be learned on the L0 bridge,
> +which can misdirect iperf3 traffic after migration.
> +
> +To prevent this, configure NetworkManager on **both L1 VMs** and on the
> +**L2 guest disk image**.
> +
> +L1 VMs (source and destination)
> +...............................
> +
> +1. Prevent NetworkManager from managing igbvf interfaces:
> +
> +.. code-block:: bash
> +
> + cat > /etc/NetworkManager/conf.d/99-no-igbvf.conf <<EOF
> + [keyfile]
> + unmanaged-devices=driver:igbvf
> + EOF
> +
> +2. Disable IP on the igb PF connections (keep the interfaces UP for
> + bridging, but with no DHCP lease):
> +
> +.. code-block:: bash
> +
> + # Identify the NM connections for the igb PFs (NOT the virtio management NIC)
> + nmcli -t -f NAME,DEVICE connection show
> +
> + # For each igb PF connection:
> + nmcli connection modify "<igb-pf-connection>" ipv4.method disabled ipv6.method disabled
> +
> +3. Reload NetworkManager:
> +
> +.. code-block:: bash
> +
> + nmcli general reload
> +
> +L2 guest disk image
> +...................
> +
> +Use ``virt-customize`` to add the igbvf unmanaged config to the guest
> +image (offline, before any test run):
> +
> +.. code-block:: bash
> +
> + virt-customize -a /srv/migration/rhel10.qcow2 \
> + --write /etc/NetworkManager/conf.d/99-no-igbvf.conf:'[keyfile]
> + unmanaged-devices=driver:igbvf'
> +
> +Network diagram
> +^^^^^^^^^^^^^^^
> +
> +The diagram below shows the nested setup where the source and
> +destination hosts are themselves VMs (L1) running on a physical
> +host (L0) that emulates the igb NIC::
> +
> + ┌────────────────────────────────────────────────────────────────────────────┐
> + │ L0: physical host │
> + │ │
> + │ virbr0 192.168.199.1/24 │
> + │ ├── NFS server: /srv/migration │
> + │ └── iperf3 client: iperf3 -c 192.168.199.200 -t 60 -i 1 │
> + │ │ │
> + │ │ L0 virbr0 bridge (192.168.199.0/24) │
> + │ ────┼──────────┬──────────────────┬────────────────────────────── │
> + │ │ │ │ │
> + │ │ ┌────┴────┐ ┌────┴────┐ │
> + │ │ │ virtio │ │ emulated│ │
> + │ │ │c0:ff:ee:│ │ igb PF │ L0 QEMU (vm6) │
> + │ │ │ :00:06 │ │ + igbvf │ tracks DMA dirty pages │
> + │ │ └────┬────┘ └────┬────┘ │
> + │ │ │ │ │
> + │ ┌───┼──────────┼──────────────────┼──────────────────────────────────┐ │
> + │ │ │ L1: vm6 (source) │ │ │
> + │ │ │ enp1s0: 192.168.199.6 │ │ │
> + │ │ │ (management) │ │ │
> + │ │ │ enp8s0 (igb PF, no IP) │ │
> + │ │ │ │ │ │
> + │ │ │ igb VF0 ──► igb-vfio-pci (VFIO) │ │
> + │ │ │ │ dirty_sync → L0 igbvf │ │
> + │ │ │ │ │ │
> + │ │ │ virbr0 │ VFIO passthrough │ │
> + │ │ │ 192.168.200.1/24 │ │ │
> + │ │ │ │ │ │ │
> + │ │ │ ┌────┼─────────────────────┼───────────────────────────┐ │ │
> + │ │ │ │ │ L2: rhel10 guest │ │ │ │
> + │ │ │ │ │ │ │ │ │
> + │ │ │ │ virtio NIC igb VF (enp7s0) │ │ │
> + │ │ │ │ 192.168.200.130/24 192.168.199.200/24 │ │ │
> + │ │ │ │ (SSH login) (iperf3 data path) │ │ │
> + │ │ │ │ │ │ │ │
> + │ │ │ │ iperf3 -s -D │ (listens on 0.0.0.0) │ │ │
> + │ │ │ └──────────────────────────┼───────────────────────────┘ │ │
> + │ │ │ │ │ │
> + │ │ │ virsh migrate --live ──────┼──────────────────► vm7 │ │
> + │ │ │ │ │ │
> + │ └───┼─────────────────────────────┼──────────────────────────────────┘ │
> + │ │ │ │
> + │ │ iperf3 traffic │ │
> + │ └─────────────────────────────┘ │
> + │ │
> + │ ──────────────────────────────────────────────────────────────── │
> + │ │ │ │
> + │ │ ┌────────┐ │ ┌─────────┐ │
> + │ │ │ virtio │ │ │emulated │ L0 QEMU (vm7) │
> + │ │ │c0:ff:ee│ │ │ igb PF │ │
> + │ │ │ :00:07 │ │ │ + igbvf │ │
> + │ │ └────┬───┘ │ └────┬────┘ │
> + │ ┌──────────────┼───────┼────────┼────────────────────────────────────┐ │
> + │ │ L1: vm7 (destination) │ │ │
> + │ │ enp1s0: 192.168.199.7 │ │ │
> + │ │ (management) enp8s0 (igb PF, no IP) │ │
> + │ │ │ │ │
> + │ │ igb VF0 ──► igb-vfio-pci (VFIO) │ │
> + │ │ │ │ │
> + │ │ virbr0 │ VFIO passthrough │ │
> + │ │ 192.168.200.1/24 │ │ │
> + │ │ │ │ │ │
> + │ │ ┌────┼───────────────────────┼────────────────────────────┐ │ │
> + │ │ │ │ L2: rhel10 (after migration) │ │ │
> + │ │ │ │ │ │ │ │
> + │ │ │ virtio NIC igb VF (enp7s0) │ │ │
> + │ │ │ 192.168.200.130/24 192.168.199.200/24 │ │ │
> + │ │ │ │ │ │ │
> + │ │ │ iperf3 -s -D │ (connection survives) │ │ │
> + │ │ └────────────────────────────┼────────────────────────────┘ │ │
> + │ └───────────────────────────────┼────────────────────────────────────┘ │
> + │ │ │
> + │ iperf3 traffic resumes ─────┘ │
> + │ (same IP, same MAC, same L2 segment → transparent to client) │
> + └────────────────────────────────────────────────────────────────────────────┘
> +
> +Migration under iperf3 load works correctly: dirty page tracking
> +converges (from ~2000 pages per PRE_COPY iteration down to ~280 at
> +STOP_COPY), and STOP_COPY stays under 250ms.
> +
> +Todo
> +~~~~
> +
> +1. Add migration blocker when ``x-vf-migration=on`` (no VMState yet) or
> + add VMState support for L0 migration (dirty bitmaps, tracking
> + engines, DVSEC registers, stats)
> +2. Add PRE_COPY state transfer to validate device INIT data (magic,
> + version, etc.)
> +3. Add qtests for migration state machine transitions, dirty page
> + tracking
> +
> +Ideas
> +~~~~~
> +
> +1. **RX bandwidth throttle** (``x-mig-rx-limit``, uint32, default 0)
> +
> + Return false from ``can_receive`` when the per-VF packet count in the
> + current tracking interval exceeds the limit. Reduces DMA writes and
> + dirty pages realistically.
> +
> +2. **Migration phase timing** (GET_STATS extension)
> +
> + Add per-VF timestamps: ``precopy_start_ns``, ``stopcopy_start_ns``,
> + ``precopy_duration_ns``, ``stopcopy_duration_ns``,
> + ``state_transition_count``. Expose via GET_STATS.
> +
> +3. **Hot page simulation** (``x-mig-hot-pages``, uint32, default 0)
> +
> + Re-set the first N bitmap bits after each DIRTY_QUERY, simulating
> + workloads with hot pages that prevent convergence.
> +
> +4. **Error injection** (``x-mig-inject-error``, uint32, default 0)
> +
> + One-shot error code injection before command dispatch. A separate
> + ``x-mig-inject-dma-fail`` (bool) for persistent DMA failure testing.
> +
> +AI disclaimer
> +~~~~~~~~~~~~~
> +
> +Claude was used to analyze the IGB PF and VF internal state and
> +identify the pain points of a working live migration of such devices.
> +The generated code served as a starting point but *significant* time
> +was then spent cleaning up, reworking, and shaping it into a clear,
> +reviewable IGB model extension.
I don't think it's really a good idea to keep this "AI disclaimer" in
the documentation.
> +
> +.. _igb-vfio-pci: https://github.com/legoater/vfio-pci-extras
> diff --git a/docs/system/devices/igb.rst b/docs/system/devices/igb.rst
> index 50f625fd77e4..00271dbc92c3 100644
> --- a/docs/system/devices/igb.rst
> +++ b/docs/system/devices/igb.rst
> @@ -64,6 +64,12 @@ command:
>
> pyvenv/bin/meson test --suite thorough func-x86_64-netdev_ethtool
>
> +VF Migration (experimental)
> +===========================
> +
> +See :ref:`igb-migration` for details on the experimental VF live migration
> +interface.
> +
> References
> ==========
>
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [RFC PATCH v2 0/9] igb: Add experimental VF live migration support
2026-09-02 19:20 [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Cédric Le Goater
` (8 preceding siblings ...)
2026-09-02 19:20 ` [RFC PATCH v2 9/9] docs: Add igb VF migration testing setup guide Cédric Le Goater
@ 2026-09-09 7:19 ` Akihiko Odaki
2026-09-15 8:20 ` Cédric Le Goater
9 siblings, 1 reply; 23+ messages in thread
From: Akihiko Odaki @ 2026-09-09 7:19 UTC (permalink / raw)
To: Cédric Le Goater, qemu-devel
Cc: Sriram Yagnaraman, Jason Wang, Alex Williamson, Peter Xu
On 2026/09/03 4:20, Cédric Le Goater wrote:
> Hello,
>
> Live migration of VFIO-passthrough devices - SR-IOV VFs, vGPUs - is a
> growing requirement, but real hardware with migration support is
> scarce and hard to debug. An emulated device provides a fully
> controlled testbed for developing and validating the entire software
> stack - vfio-pci variant drivers, VFIO core migration v2 framework,
> QEMU, libvirt - and for tuning complex migration policies such as
> downtime convergence. It also serves as an educational reference for
> understanding VFIO migration end-to-end, from device state
> serialization to dirty page tracking.
>
> This series adds an experimental VF live migration interface to the
> emulated igb (82576) device. It enables a vfio-pci variant driver
> (igb-vfio-pci) to migrate VFs using the standard VFIO migration v2
> protocol with stop-copy and pre-copy support.
>
> The target scenario is nested virtualization:
>
> L0 QEMU (these patches)
> igb PF with x-vf-migration=on
> └── VFs with migration DVSEC
>
> L1 kernel
> igb-vfio-pci variant driver [1]
> translates VFIO migration v2 ioctls → DVSEC config writes
>
> L1 QEMU (stock, unmodified)
> vfio-pci device model, standard migration fd
>
> L2 guest
> standard igbvf driver, unaware of migration
>
> The L1 QEMU is completely unmodified -- it sees a standard VFIO
> migratable device and uses the normal migration fd path.
>
> * Design
>
> The migration interface is exposed through a DVSEC (Designated
> Vendor-Specific Extended Capability, PCIe cap id 0x23) at offset
> 0x160 in VF extended config space. The DVSEC uses a command doorbell
> model - all commands are synchronous via PCI config space writes.
>
> Device state is serialized as a versioned blob of per-VF register
> (offset, value) pairs covering control, interrupt, RX/TX queue,
> receive address (RA/RA2), etc. plus TX context descriptors and
> VFRE/VFTE enable bits. The buffer address is a guest physical
> address (GPA) written by the driver via virt_to_phys; the device
> accesses guest RAM directly through the system address space.
>
> Dirty page tracking is implemented with per-range bitmaps maintained
> in IGBCore. All VF DMA paths in igb_core.c (TX data, RX data,
> descriptor writeback) are instrumented to record touched pages. The
> variant driver registers tracked IOVA ranges and queries dirty bitmaps
> through a shared buffer. Buffer structures include len, flags, and
> reserved fields for future extensibility.
>
> * Caveats
>
> The x-vf-migration property is experimental (x- prefix, default off).
>
> The dirty bitmaps are maintained inside the device, which is not
> realistic for discrete NICs without on-chip DRAM.
>
> * Testing
>
> The target scenario is nested virtualization: L0 runs QEMU with an
> igb PF (x-vf-migration=on), L1 runs the igb-vfio-pci variant driver
> and an unmodified QEMU, and L2 runs a standard igbvf driver.
>
> Migration under iperf3 load works correctly: dirty page tracking
> converges (from ~2000 pages per PRE_COPY iteration down to ~280 at
> STOP_COPY), and STOP_COPY stays under 250ms.
>
> * Todo
>
> 1. Add migration blocker when x-vf-migration=on (no VMState yet) or
> add VMState support for L0 migration (dirty bitmaps, tracking
> engines, DVSEC registers, stats)
> 2. Add PRE_COPY state transfer to validate device INIT data (magic,
> version, etc.)
> 3. Add qtests for migration state machine transitions, dirty page
> tracking ?
>
> * Ideas
>
> 1. RX bandwidth throttle (x-mig-rx-limit, uint32, default 0)
>
> Return false from can_receive when the per-VF packet count in the
> current tracking interval exceeds the limit. Reduces DMA writes
> and dirty pages realistically.
>
> 2. Migration phase timing (GET_STATS extension)
>
> Add per-VF timestamps: precopy_start_ns, stopcopy_start_ns,
> precopy_duration_ns, stopcopy_duration_ns,
> state_transition_count. Expose via GET_STATS.
>
> 3. Hot page simulation (x-mig-hot-pages, uint32, default 0)
>
> Re-set the first N bitmap bits after each DIRTY_QUERY, simulating
> workloads with hot pages that prevent convergence.
>
> 4. Error injection (x-mig-inject-error, uint32, default 0)
>
> One-shot error code injection before command dispatch. A separate
> x-mig-inject-dma-fail (bool) for persistent DMA failure testing.
>
> * Credits
>
> Alex Williamson suggested the overall approach of a variant driver
> with the "x-vf-migration" device property to gate the feature. Thanks
> for the ever ongoing support and valuable discussions throughout these
> years.
>
> * AI disclaimer
>
> The lack of a migration-capable device has been a recurring pain point
> for VFIO development over the years, and we hope this proposal
> demonstrates the value of having one.
>
> Claude was used to analyze the IGB PF and VF internal state and
> identify the pain points of a working live migration of such devices.
> The generated code served as a starting point but *significant* time
> was then spent cleaning up, reworking, and shaping it into a clear,
> reviewable IGB model extension.
>
> As QEMU does not yet accept AI-assisted contributions, this series is
> submitted as an RFC.
This largely repeats the concerns I raised previously [1]. Your
description of the *significant* time spent cleaning up and reworking
the generated code suggests that this series is intended for upstream
inclusion once the implementation direction is agreed. My concerns as a
maintainer are therefore correctness and maintainability.
The missing migration blocker is a concrete example: it is a basic
correctness issue that should be straightforward to address, yet it
remains a TODO. My comments on individual patches also raise the same
kinds of issues I pointed out in the previous version.
Many of the correctness problems found in v1 and v2 stem from the choice
to build on igb, which is more complex than virtio-net. That is why I
suggested using virtio-net as a simpler basis, unless there is a problem
that rules it out.
I would be willing to support allowing AI use in the project for work
like this. I would still need to see these correctness and
maintainability concerns addressed before I could support merging
the series.
Regards,
Akihiko Odaki
[1]
https://lore.kernel.org/qemu-devel/20260727053935.1392269-1-clg@redhat.com/
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [RFC PATCH v2 3/9] igb: Add VF state serialization for live migration
2026-09-08 8:01 ` Akihiko Odaki
@ 2026-09-15 7:47 ` Cédric Le Goater
2026-09-17 18:59 ` Akihiko Odaki
0 siblings, 1 reply; 23+ messages in thread
From: Cédric Le Goater @ 2026-09-15 7:47 UTC (permalink / raw)
To: Akihiko Odaki, qemu-devel
Cc: Sriram Yagnaraman, Jason Wang, Alex Williamson, Peter Xu
Hello Akihiko,
On 9/8/26 10:01, Akihiko Odaki wrote:
>
>
> On 2026/09/03 4:20, Cédric Le Goater wrote:
>> Implement per-VF state serialization and deserialization for the SAVE
>> and LOAD commands. The wire format consists of a header (magic,
>> version, VF number, register count), per-VF register offset/value
>> pairs from a whitelist, RA table entries owned by the VF, and TX
>> queue contexts.
>>
>> Register offsets are relocated on load so a VF can migrate to a
>> different VF number on the destination. RA pool ownership bits are
>> swapped accordingly.
>>
>> AI-used-for: analysis, code (prototype)
>> Signed-off-by: Cédric Le Goater <clg@redhat.com>
>> ---
>> hw/net/igb_core.h | 2 +
>> hw/net/igb_migration.h | 2 +
>> hw/net/igb.c | 5 +
>> hw/net/igb_migration.c | 347 ++++++++++++++++++++++++++++++++++++++++-
>> 4 files changed, 354 insertions(+), 2 deletions(-)
>>
>> diff --git a/hw/net/igb_core.h b/hw/net/igb_core.h
>> index d70b54e318f1..60724e2824ab 100644
>> --- a/hw/net/igb_core.h
>> +++ b/hw/net/igb_core.h
>> @@ -143,4 +143,6 @@ igb_receive_iov(IGBCore *core, const struct iovec *iov, int iovcnt);
>> void
>> igb_start_recv(IGBCore *core);
>> +IGBCore *igb_pf_get_core(void *pf);
>> +
>> #endif
>> diff --git a/hw/net/igb_migration.h b/hw/net/igb_migration.h
>> index ea40ac65c54b..b2f601e74346 100644
>> --- a/hw/net/igb_migration.h
>> +++ b/hw/net/igb_migration.h
>> @@ -73,6 +73,8 @@
>> #define IGB_MIG_ERR_NO_BUFFER 3
>> #define IGB_MIG_ERR_DMA_FAILED 4
>> #define IGB_MIG_ERR_BAD_SIZE 5
>> +#define IGB_MIG_ERR_BAD_MAGIC 6
>> +#define IGB_MIG_ERR_BAD_VERSION 7
>> /* Shared buffer constants */
>> #define IGB_VF_STATE_MAX_SIZE 4096
>> diff --git a/hw/net/igb.c b/hw/net/igb.c
>> index 7268e5473fc3..f39f2bc3a04e 100644
>> --- a/hw/net/igb.c
>> +++ b/hw/net/igb.c
>> @@ -133,6 +133,11 @@ void igb_vf_reset(void *opaque, uint16_t vfn)
>> igb_core_vf_reset(&s->core, vfn);
>> }
>> +IGBCore *igb_pf_get_core(void *pf)
>> +{
>> + return &IGB(pf)->core;
>> +}
>> +
>> static bool
>> igb_io_get_reg_index(IGBState *s, uint32_t *idx)
>> {
>> diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
>> index 8e7e6fac9b9b..c34035974620 100644
>> --- a/hw/net/igb_migration.c
>> +++ b/hw/net/igb_migration.c
>> @@ -10,18 +10,241 @@
>> #include "qemu/log.h"
>> #include "hw/pci/pci_device.h"
>> #include "hw/pci/pcie.h"
>> +#include "net/eth.h"
>> +#include "net/net.h"
>> #include "igb_common.h"
>> +#include "igb_core.h"
>> #include "igb_migration.h"
>> #include "system/address-spaces.h"
>> #include "trace.h"
>> +static IGBCore *igbvf_get_core(IgbVfState *s)
>> +{
>> + return igb_pf_get_core(pcie_sriov_get_pf(PCI_DEVICE(s)));
>> +}
>> +
>> /*
>> * Per-VF state serialization / deserialization
>> */
>> +#define IGB_MIG_BLOB_MAGIC 0x4D494742 /* "MIGB" */
>> +#define IGB_MIG_BLOB_VERSION 1
>> +
>> +typedef struct IgbMigRegPair {
>> + uint32_t offset;
>> + uint32_t value;
>> +} IgbMigRegPair;
>> +
>> +typedef struct IgbMigTxCtx {
>> + uint32_t ctx_desc[8]; /* 2 × adv_tx_context_desc (4 dwords each) */
>> + uint32_t first_cmd_type_len;
>> + uint32_t first_olinfo_status;
>> + uint32_t first;
>> + uint32_t skip_cp;
>> +} IgbMigTxCtx;
>> +
>> +#define IGB_VF_MAX_FIXED_REGS 64
>> +#define IGB_VF_MAX_RA_REGS 48 /* (16 + 8) RA entries × 2 (RAL+RAH) */
>> +
>> +typedef struct IgbMigBlob {
>> + uint32_t magic;
>> + uint32_t version;
>> + uint32_t vfn;
>> + uint32_t num_regs;
>> + IgbMigRegPair regs[IGB_VF_MAX_FIXED_REGS];
>> + uint32_t num_ra;
>> + IgbMigRegPair ra[IGB_VF_MAX_RA_REGS];
>> + uint32_t num_tx_ctx;
>> + IgbMigTxCtx tx_ctx[2];
>> +} IgbMigBlob;
>> +
>> +#define IGB_MIG_BLOB_SIZE sizeof(IgbMigBlob)
>> +
>> +QEMU_BUILD_BUG_ON(IGB_MIG_BLOB_SIZE > IGB_VF_STATE_MAX_SIZE);
>> +
>> +/* Register offsets that constitute a VF's state slice */
>> +static int igb_vf_reg_list(uint16_t vfn, uint32_t *offsets)
>> +{
>> + int n = 0;
>> + int q0 = vfn;
>> + int q1 = vfn + IGB_NUM_VM_POOLS;
>> +
>> + /* Per-VF control and interrupt registers */
>> + offsets[n++] = E1000_PVTCTRL(vfn) >> 2;
>> + offsets[n++] = E1000_PVTEICS(vfn) >> 2;
>> + offsets[n++] = E1000_PVTEIMS(vfn) >> 2;
>> + offsets[n++] = E1000_PVTEIMC(vfn) >> 2;
>> + offsets[n++] = E1000_PVTEIAC(vfn) >> 2;
>> + offsets[n++] = E1000_PVTEIAM(vfn) >> 2;
>> + offsets[n++] = E1000_PVTEICR(vfn) >> 2;
>> +
>> + /* Per-VF statistics */
>> + offsets[n++] = E1000_PVFGPRC(vfn) >> 2;
>> + offsets[n++] = E1000_PVFGPTC(vfn) >> 2;
>> + offsets[n++] = E1000_PVFGORC(vfn) >> 2;
>> + offsets[n++] = E1000_PVFGOTC(vfn) >> 2;
>> + offsets[n++] = E1000_PVFMPRC(vfn) >> 2;
>> + offsets[n++] = E1000_PVFGPRLBC(vfn) >> 2;
>> + offsets[n++] = E1000_PVFGPTLBC(vfn) >> 2;
>> + offsets[n++] = E1000_PVFGORLBC(vfn) >> 2;
>> + offsets[n++] = E1000_PVFGOTLBC(vfn) >> 2;
>> +
>> + /*
>> + * Mailbox control registers only - the 16-dword payload buffer
>> + * (VMBMEM) is transient and drained on quiesce.
>> + */
>> + offsets[n++] = E1000_V2PMAILBOX(vfn) >> 2;
>> + offsets[n++] = E1000_P2VMAILBOX(vfn) >> 2;
>> +
>> + /* Per-VF config */
>> + offsets[n++] = E1000_VMOLR(vfn) >> 2;
>> + offsets[n++] = E1000_VMVIR(vfn) >> 2;
>> + offsets[n++] = E1000_PSRTYPE(vfn) >> 2;
>> +
>> + /*
>> + * VF receive addresses (RA/RA2) are saved dynamically in
>> + * igb_core_vf_save_state by scanning for entries whose pool
>> + * bits match this VF - the PF driver chooses the RA slot.
>> + */
>> +
>> + /* Interrupt routing */
>> + offsets[n++] = (E1000_VTIVAR + vfn * 4) >> 2;
>> + offsets[n++] = (E1000_VTIVAR_MISC + vfn * 4) >> 2;
>> +
>> + /*
>> + * EITR (Extended Interrupt Throttle Register) - 3 vectors per VF.
>> + * Each VF has 3 MSI-X vectors, each with its own EITR controlling
>> + * interrupt coalescing. Without saving these, interrupt
>> + * throttling resets to zero after migration which can cause
>> + * interrupt storms or latency changes. VF N uses PF EITR indices
>> + * (22 - N*3) .. (24 - N*3).
>> + */
>> + {
>> + int eitr_base = 22 - vfn * 3;
>> + offsets[n++] = E1000_EITR(eitr_base) >> 2;
>> + offsets[n++] = E1000_EITR(eitr_base + 1) >> 2;
>> + offsets[n++] = E1000_EITR(eitr_base + 2) >> 2;
>> + }
>> +
>> + /* RX and TX queue registers for queues q0 and q1 */
>> +#define ADD_QUEUE_REGS(q) do { \
>> + offsets[n++] = E1000_RDBAL(q) >> 2; \
>> + offsets[n++] = E1000_RDBAH(q) >> 2; \
>> + offsets[n++] = E1000_RDLEN(q) >> 2; \
>> + offsets[n++] = E1000_SRRCTL(q) >> 2; \
>> + offsets[n++] = E1000_RDH(q) >> 2; \
>> + offsets[n++] = E1000_RDT(q) >> 2; \
>> + offsets[n++] = E1000_RXDCTL(q) >> 2; \
>> + offsets[n++] = E1000_RXCTL(q) >> 2; \
>> + offsets[n++] = E1000_RQDPC(q) >> 2; \
>> + offsets[n++] = E1000_TDBAL(q) >> 2; \
>> + offsets[n++] = E1000_TDBAH(q) >> 2; \
>> + offsets[n++] = E1000_TDLEN(q) >> 2; \
>> + offsets[n++] = E1000_TDH(q) >> 2; \
>> + offsets[n++] = E1000_TDT(q) >> 2; \
>> + offsets[n++] = E1000_TXDCTL(q) >> 2; \
>> + offsets[n++] = E1000_TXCTL(q) >> 2; \
>> + offsets[n++] = E1000_TDWBAL(q) >> 2; \
>> + offsets[n++] = E1000_TDWBAH(q) >> 2; \
>> +} while (0)
>> +
>> + ADD_QUEUE_REGS(q0);
>> + ADD_QUEUE_REGS(q1);
>> +#undef ADD_QUEUE_REGS
>> +
>> + g_assert(n <= IGB_VF_MAX_FIXED_REGS);
>> + return n;
>> +}
>> +
>> +/*
>> + * Scan RA and RA2 arrays for receive address entries assigned to
>> + * this VF. The PF driver picks the RA slot, so we cannot use a
>> + * fixed index - instead check each entry's pool bits.
>> + */
>> +static int igb_core_vf_save_ra(IGBCore *core, uint16_t vfn,
>> + IgbMigRegPair *regs)
>> +{
>> + uint32_t vf_pool_bit = E1000_RAH_POOL_1 << vfn;
>> + int n = 0;
>> + static const struct {
>> + uint32_t base;
>> + int count;
>> + } ra_banks[] = {
>> + { RA, 16 },
>> + { RA2, 8 },
>> + };
>> +
>> + for (int i = 0; i < ARRAY_SIZE(ra_banks); i++) {
>> + for (int j = 0; j < ra_banks[i].count; j++) {
>> + uint32_t ral_off = ra_banks[i].base + j * 2;
>> + uint32_t rah_off = ra_banks[i].base + j * 2 + 1;
>> + uint32_t rah_val = core->mac[rah_off];
>> +
>> + if ((rah_val & E1000_RAH_AV) && (rah_val & vf_pool_bit)) {
>> + regs[n].offset = cpu_to_le32(ral_off);
>> + regs[n].value = cpu_to_le32(core->mac[ral_off]);
>> + n++;
>> + regs[n].offset = cpu_to_le32(rah_off);
>> + regs[n].value = cpu_to_le32(rah_val);
>> + n++;
>> + }
>> + }
>> + }
>> + return n;
>> +}
>> +
>> +static void igb_core_vf_save_tx_ctx(IGBCore *core, int queue,
>> + IgbMigTxCtx *tx)
>> +{
>> + struct igb_tx *src = &core->tx[queue];
>> +
>> + memcpy(tx->ctx_desc, src->ctx, sizeof(tx->ctx_desc));
>> + tx->first_cmd_type_len = cpu_to_le32(src->first_cmd_type_len);
>> + tx->first_olinfo_status = cpu_to_le32(src->first_olinfo_status);
>> + tx->first = cpu_to_le32(src->first);
>> + tx->skip_cp = cpu_to_le32(src->skip_cp);
>> +}
>> +
>> static int igb_core_vf_save_state(IgbVfState *s, void *buf, size_t buf_size)
>> {
>> - int size = 0;
>> + int size = IGB_MIG_BLOB_SIZE;
>> + IGBCore *core = igbvf_get_core(s);
>> + IgbMigBlob *blob = buf;
>> + uint32_t offsets[IGB_VF_MAX_FIXED_REGS];
>> + int num_regs;
>> + int q0 = s->vfn;
>> + int q1 = s->vfn + IGB_NUM_VM_POOLS;
>> +
>> + /*
>> + * Save PVT shadow registers (PVTEIMS/PVTEIAC/PVTEIAM) instead of
>> + * extracting from PF aggregates - the L1 PF driver may have
>> + * transiently cleared EIMS via EIMC. The load path ORs them back.
>> + */
>> + num_regs = igb_vf_reg_list(s->vfn, offsets);
>> +
>> + if (!buf) {
>> + return size;
>> + }
>> +
>> + if (size > buf_size) {
>> + return -IGB_MIG_ERR_BAD_SIZE;
>> + }
>> +
>> + blob->magic = cpu_to_le32(IGB_MIG_BLOB_MAGIC);
>> + blob->version = cpu_to_le32(IGB_MIG_BLOB_VERSION);
>> + blob->vfn = cpu_to_le32(s->vfn);
>> +
>> + blob->num_regs = cpu_to_le32(num_regs);
>> + for (int i = 0; i < num_regs; i++) {
>> + blob->regs[i].offset = cpu_to_le32(offsets[i]);
>> + blob->regs[i].value = cpu_to_le32(core->mac[offsets[i]]);
>> + }
>> +
>> + blob->num_ra = cpu_to_le32(igb_core_vf_save_ra(core, s->vfn, blob->ra));
>
> The serialized state contains RA entries but excludes VFTA/VLVF and MTA/UTA state.
VLVF should not be too complex to add support for. It is the same pattern
as the RA scan.
VFTA/MTA/UTA support is tougher. IIUC, these are PF shared tables and there
are no per-VF bits. So we either need to add per-VF shadow tracking, which
is intrusive or, to begin with, we can ignore: detect and warn.
> A guest-configured VLAN therefore resumes on a destination lacking its filter membership: packets are rejected by the global VLAN check or have their VF queue removed in igb_receive_assign(),. Likewise, restored VMOLR.ROMPE cannot receive subscribed multicast without its MTA hash bits. These are guest-requested settings communicated through the PF mailbox, and the resumed guest does not recreate them automatically.
yes. If we could restore some of the VF state using the PF mailbox, we
would avoid poking the core igb device from the VF but the model doesn't
have enough support yet.
>
>> +
>> + blob->num_tx_ctx = cpu_to_le32(2);
>> + igb_core_vf_save_tx_ctx(core, q0, &blob->tx_ctx[0]);
>> + igb_core_vf_save_tx_ctx(core, q1, &blob->tx_ctx[1]);
>> trace_igbvf_mig_save_state(s->vfn, size);
>> return size;
>> @@ -29,11 +252,131 @@ static int igb_core_vf_save_state(IgbVfState *s, void *buf, size_t buf_size)
>> static int igb_core_vf_max_data_size(IgbVfState *s)
>> {
>> - return sizeof(s->mig.mig_data);
>> + int size = igb_core_vf_save_state(s, NULL, 0);
>> +
>> + g_assert(size > 0 && size <= IGB_VF_STATE_MAX_SIZE);
>> + return size;
>> +}
>> +
>> +static void igb_core_vf_load_tx_ctx(IGBCore *core, int queue,
>> + const IgbMigTxCtx *tx)
>> +{
>> + struct igb_tx *dst = &core->tx[queue];
>> +
>> + /*
>> + * Preserve the destination's tx_pkt - it's a host-side object,
>> + * not guest state
>> + */
>> + memcpy(dst->ctx, tx->ctx_desc, sizeof(dst->ctx));
>> + dst->first_cmd_type_len = le32_to_cpu(tx->first_cmd_type_len);
>> + dst->first_olinfo_status = le32_to_cpu(tx->first_olinfo_status);
>> + dst->first = le32_to_cpu(tx->first);
>> + dst->skip_cp = le32_to_cpu(tx->skip_cp);
>> +}
>> +
>> +static uint32_t igb_vf_relocate_offset(uint32_t offset,
>> + const uint32_t *src_offsets,
>> + const uint32_t *dst_offsets,
>> + int num_offsets)
>> +{
>> + for (int i = 0; i < num_offsets; i++) {
>> + if (src_offsets[i] == offset) {
>> + return dst_offsets[i];
>> + }
>> + }
>> + return 0;
>> }
>> static int igb_core_vf_load_state(IgbVfState *s, const void *buf, size_t size)
>> {
>> + IGBCore *core = igbvf_get_core(s);
>> + uint32_t src_offsets[IGB_VF_MAX_FIXED_REGS];
>> + uint32_t dst_offsets[IGB_VF_MAX_FIXED_REGS];
>> + int q0 = s->vfn;
>> + int q1 = s->vfn + IGB_NUM_VM_POOLS;
>> +
>> + if (size < IGB_MIG_BLOB_SIZE) {
>> + return -IGB_MIG_ERR_BAD_SIZE;
>> + }
>> +
>> + const IgbMigBlob *blob = buf;
>> +
>> + uint32_t magic = le32_to_cpu(blob->magic);
>> + uint32_t version = le32_to_cpu(blob->version);
>> + uint32_t saved_vfn = le32_to_cpu(blob->vfn);
>> + uint32_t num_regs = le32_to_cpu(blob->num_regs);
>> +
>> + if (magic != IGB_MIG_BLOB_MAGIC) {
>> + return -IGB_MIG_ERR_BAD_MAGIC;
>> + }
>> + if (version != IGB_MIG_BLOB_VERSION) {
>> + return -IGB_MIG_ERR_BAD_VERSION;
>> + }
>> + if (num_regs > IGB_VF_MAX_FIXED_REGS) {
>> + return -IGB_MIG_ERR_BAD_SIZE;
>> + }
>> +
>> + uint32_t num_ra = le32_to_cpu(blob->num_ra);
>> + if (num_ra > IGB_VF_MAX_RA_REGS) {
>> + return -IGB_MIG_ERR_BAD_SIZE;
>> + }
>> +
>> + int num_offsets = igb_vf_reg_list(saved_vfn, src_offsets);
>> + igb_vf_reg_list(s->vfn, dst_offsets);
>> +
>> + for (uint32_t i = 0; i < num_regs; i++) {
>> + uint32_t src_off = le32_to_cpu(blob->regs[i].offset);
>> + uint32_t value = le32_to_cpu(blob->regs[i].value);
>> + uint32_t offset = igb_vf_relocate_offset(src_off,
>> + src_offsets, dst_offsets, num_offsets);
>> + if (!offset) {
>> + return -IGB_MIG_ERR_BAD_SIZE;
>> + }
>> +
>> + core->mac[offset] = value;
>> +
>> + /*
>> + * Sync EITR to eitr_guest_value[] shadow array, stripping
>> + * E1000_EITR_CNT_IGNR so guest register readback returns the
>> + * correct value.
>> + */
>> + if (offset >= EITR0 && offset < EITR0 + IGB_INTR_NUM) {
>> + core->eitr_guest_value[offset - EITR0] =
>> + value & ~E1000_EITR_CNT_IGNR;
>> + }
>> + }
>> +
>> + /*
>> + * MSI-X table/PBA is not saved - L1's VFIO reprograms it with
>> + * destination-specific IRTE references after migration.
>> + */
>> +
>> + uint32_t src_pool = E1000_RAH_POOL_1 << saved_vfn;
>> + uint32_t dst_pool = E1000_RAH_POOL_1 << s->vfn;
>> +
>> + for (uint32_t i = 0; i < num_ra; i++) {
>> + uint32_t offset = le32_to_cpu(blob->ra[i].offset);
>> + uint32_t value = le32_to_cpu(blob->ra[i].value);
>> +
>> + /* RAH entries: swap pool ownership bits */
>> + if (offset >= RA && offset < RA + 32 && (offset - RA) % 2 == 1) {
>> + value = (value & ~src_pool) | dst_pool;
>> + }
>> + if (offset >= RA2 && offset < RA2 + 16 && (offset - RA2) % 2 == 1) {
>> + value = (value & ~src_pool) | dst_pool;
>> + }
>> +
>> + core->mac[offset] = value;
>
> Please validate RA offsets before writing core->mac.
Drat. my bad. It should be done and it's easy :
reject offsets not in [RA, RA+32) or [RA2, RA2+16)
> A valid-sized LOAD blob with one RA entry and an out-of-range offset produces an out-of-bounds 32-bit host write. LOAD reads this blob directly from guest RAM. Please validate all offsets and saved_vfn before modifying state.
I think we should drop VF relocation support.
if (saved_vfn != s->vfn) {
return -IGB_MIG_ERR_BAD_VFN;
}
I was too optimistic in v2. This adds too much complexity and we haven't
covered the simple case yet. Let's consider that VF number is part of
the migration state to have more invariants and avoid collisions.
>
> This loop also changes the VF pool bit but preserves the source RA slot. A valid VF0→VF1 migration onto a destination already using VF0 overwrites destination VF0’s primary MAC entry. Linux assigns the primary entry as rar_entry_count - (vf + 1); receive filtering consumes these slots directly in igb_receive_assign(). Thus the unrelated destination VF loses unicast reception. Existing destination entries owned by the migrating VF are also left behind.
There should be no/few entries on the destination for the VF being
migrated, only the fixed primary MAC entries for the VF. So, with the
same VF number enforced, slot assignments should be deterministic
and we can scan and clear any entries with this VF's pool bit at load
time.
Thanks,
C.
>
> Regards,
> Akihiko Odaki
>
>> + }
>> +
>> + uint32_t num_tx = le32_to_cpu(blob->num_tx_ctx);
>> + if (num_tx != 2) {
>> + return -IGB_MIG_ERR_BAD_SIZE;
>> + }
>> +
>> + igb_core_vf_load_tx_ctx(core, q0, &blob->tx_ctx[0]);
>> + igb_core_vf_load_tx_ctx(core, q1, &blob->tx_ctx[1]);
>> +
>> trace_igbvf_mig_load_state(s->vfn, (uint32_t)size);
>> return 0;
>> }
>
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [RFC PATCH v2 4/9] igb: Add VF post-load fixups for live migration
2026-09-08 8:10 ` Akihiko Odaki
@ 2026-09-15 7:58 ` Cédric Le Goater
2026-09-17 19:01 ` Akihiko Odaki
0 siblings, 1 reply; 23+ messages in thread
From: Cédric Le Goater @ 2026-09-15 7:58 UTC (permalink / raw)
To: Akihiko Odaki, qemu-devel
Cc: Sriram Yagnaraman, Jason Wang, Alex Williamson, Peter Xu
On 9/8/26 10:10, Akihiko Odaki wrote:
> On 2026/09/03 4:20, Cédric Le Goater wrote:
>> After restoring per-VF register state, propagate the VF's PVT shadow
>> values back into the PF's aggregate EIMS/EIAC/EIAM registers and
>> re-apply the VTIVAR interrupt vector routing to the shared IVAR0.
>>
>> AI-used-for: analysis, code (prototype)
>> Signed-off-by: Cédric Le Goater <clg@redhat.com>
>> ---
>> hw/net/igb_core.h | 3 ++
>> hw/net/igb_core.c | 66 ++++++++++++++++++++++++++++++++++++++++++
>> hw/net/igb_migration.c | 7 +++++
>> 3 files changed, 76 insertions(+)
>>
>> diff --git a/hw/net/igb_core.h b/hw/net/igb_core.h
>> index 60724e2824ab..22e10e4e6d0b 100644
>> --- a/hw/net/igb_core.h
>> +++ b/hw/net/igb_core.h
>> @@ -145,4 +145,7 @@ igb_start_recv(IGBCore *core);
>> IGBCore *igb_pf_get_core(void *pf);
>> +void igb_core_vf_propagate_irqs(IGBCore *core, uint16_t vfn);
>> +void igb_core_vf_propagate_ivar(IGBCore *core, uint16_t vfn);
>> +
>> #endif
>> diff --git a/hw/net/igb_core.c b/hw/net/igb_core.c
>> index 2a4883907353..01745fe756d0 100644
>> --- a/hw/net/igb_core.c
>> +++ b/hw/net/igb_core.c
>> @@ -4552,3 +4552,69 @@ igb_core_post_load(IGBCore *core)
>> return 0;
>> }
>> +
>> +/*
>> + * Propagate VF interrupt state to PF aggregates after loading VF
>> + * registers. The load path writes directly to mac[] bypassing the
>> + * register handlers that OR VF bits into EIMS/EIAC/EIAM. Also clear
>> + * stale VF bits in EICR that may have been set by packets arriving
>> + * between PF vmstate restore and VF state load.
>> + */
>> +void igb_core_vf_propagate_irqs(IGBCore *core, uint16_t vfn)
>> +{
>> + uint32_t shift = 22 - vfn * IGBVF_MSIX_VEC_NUM;
>> + uint32_t vf_mask = 0x7 << shift;
>> + uint32_t pvt_idx;
>> +
>> + core->mac[EIMS] &= ~vf_mask;
>
> Restore the effective interrupt mask rather than the last EIMS write.
>
> This clears the destination VF mask and copies PVTEIMS, but that shadow records only the last write: igb_set_vteims() assigns it before applying set-bit semantics to aggregate EIMS. For example, source writes EIMS=1, then EIMS=2, leave effective mask 3 and shadow 2; migration restores 2, suppressing vector 0 until explicitly enabled again. Conversely, EIMC updates aggregate EIMS without updating PVTEIMS, so migration reenables explicitly masked vectors. Patch 7 repeats this fixup on unquiesce without repairing the bookkeeping.
ok. I guess we can save and restore the VF's effective mask bits directly
and drop the fixup then. Same for EIAC/EIAM. It should simplify the flow.
How's that ?
Thanks
C.
>
> Regards,
> Akihiko Odaki
>
>> + pvt_idx = PVTEIMS0 + vfn * 0x40;
>> + core->mac[EIMS] |= (core->mac[pvt_idx] & 0x7) << shift;
>> +
>> + core->mac[EIAC] &= ~vf_mask;
>> + pvt_idx = PVTEIAC0 + vfn * 0x40;
>> + core->mac[EIAC] |= (core->mac[pvt_idx] & 0x7) << shift;
>> +
>> + core->mac[EIAM] &= ~vf_mask;
>> + pvt_idx = PVTEIAM0 + vfn * 0x40;
>> + core->mac[EIAM] |= (core->mac[pvt_idx] & 0x7) << shift;
>> +
>> + core->mac[EICR] &= ~vf_mask;
>> +}
>> +
>> +/*
>> + * Re-apply VTIVAR -> IVAR0 interrupt routing. The L1 PF driver
>> + * may have overwritten the shared IVAR0 entries with its own
>> + * queue routing after L0 vmstate restore.
>> + */
>> +void igb_core_vf_propagate_ivar(IGBCore *core, uint16_t vfn)
>> +{
>> + uint32_t vtivar = core->mac[VTIVAR + vfn];
>> + int n;
>> + uint8_t ent;
>> + uint32_t mask;
>> +
>> + n = igb_ivar_entry_rx(vfn);
>> + mask = 0xffU << (8 * (n % 4));
>> + if (vtivar & E1000_IVAR_VALID) {
>> + ent = E1000_IVAR_VALID |
>> + (24 - vfn * IGBVF_MSIX_VEC_NUM - (2 - (vtivar & 0x7)));
>> + core->mac[IVAR0 + n / 4] =
>> + (core->mac[IVAR0 + n / 4] & ~mask) |
>> + ((uint32_t)ent << (8 * (n % 4)));
>> + } else {
>> + core->mac[IVAR0 + n / 4] &= ~mask;
>> + }
>> +
>> + n = igb_ivar_entry_tx(vfn);
>> + mask = 0xffU << (8 * (n % 4));
>> + ent = vtivar >> 8;
>> + if (ent & E1000_IVAR_VALID) {
>> + ent = E1000_IVAR_VALID |
>> + (24 - vfn * IGBVF_MSIX_VEC_NUM - (2 - (ent & 0x7)));
>> + core->mac[IVAR0 + n / 4] =
>> + (core->mac[IVAR0 + n / 4] & ~mask) |
>> + ((uint32_t)ent << (8 * (n % 4)));
>> + } else {
>> + core->mac[IVAR0 + n / 4] &= ~mask;
>> + }
>> +}
>> diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
>> index c34035974620..af7257ccc4fd 100644
>> --- a/hw/net/igb_migration.c
>> +++ b/hw/net/igb_migration.c
>> @@ -384,12 +384,19 @@ static int igb_core_vf_load_state(IgbVfState *s, const void *buf, size_t size)
>> static int igbvf_mig_load(IgbVfState *s, const void *buf, size_t size)
>> {
>> int ret;
>> + IGBCore *core = igbvf_get_core(s);
>> ret = igb_core_vf_load_state(s, buf, size);
>> if (ret < 0) {
>> return ret;
>> }
>> + /*
>> + * Post-load: sync VF interrupt and routing state to PF aggregates
>> + */
>> + igb_core_vf_propagate_irqs(core, s->vfn);
>> + igb_core_vf_propagate_ivar(core, s->vfn);
>> +
>> return 0;
>> }
>
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [RFC PATCH v2 9/9] docs: Add igb VF migration testing setup guide
2026-09-08 8:39 ` Akihiko Odaki
@ 2026-09-15 7:59 ` Cédric Le Goater
0 siblings, 0 replies; 23+ messages in thread
From: Cédric Le Goater @ 2026-09-15 7:59 UTC (permalink / raw)
To: Akihiko Odaki, qemu-devel
Cc: Sriram Yagnaraman, Jason Wang, Alex Williamson, Peter Xu
On 9/8/26 10:39, Akihiko Odaki wrote:
> On 2026/09/03 4:20, Cédric Le Goater wrote:
>> Document the igb VF migration interface: DVSEC register layout, dirty
>> page tracking, testing setup with nested virtualization.
>>
>> AI-used-for: docs
>> Signed-off-by: Cédric Le Goater <clg@redhat.com>
>> ---
>> MAINTAINERS | 1 +
>> docs/system/device-emulation.rst | 1 +
>> docs/system/devices/igb-migration.rst | 417 ++++++++++++++++++++++++++
>> docs/system/devices/igb.rst | 6 +
>> 4 files changed, 425 insertions(+)
>> create mode 100644 docs/system/devices/igb-migration.rst
>>
>> diff --git a/MAINTAINERS b/MAINTAINERS
>> index f88b526be238..4a2f1357e936 100644
>> --- a/MAINTAINERS
>> +++ b/MAINTAINERS
>> @@ -2820,6 +2820,7 @@ igb VF migration
>> M: Cédric Le Goater <clg@redhat.com>
>> S: Maintained
>> F: hw/net/igb_migration.*
>> +F: docs/system/devices/igb-migration.rst
>> eepro100
>> M: Stefan Weil <sw@weilnetz.de>
>> diff --git a/docs/system/device-emulation.rst b/docs/system/device-emulation.rst
>> index 40054bb7dfcc..75f423b795cd 100644
>> --- a/docs/system/device-emulation.rst
>> +++ b/docs/system/device-emulation.rst
>> @@ -90,6 +90,7 @@ Emulated Devices
>> devices/cxl.rst
>> devices/emmc.rst
>> devices/igb.rst
>> + devices/igb-migration.rst
>> devices/ivshmem-flat.rst
>> devices/ivshmem.rst
>> devices/keyboard.rst
>> diff --git a/docs/system/devices/igb-migration.rst b/docs/system/devices/igb-migration.rst
>> new file mode 100644
>> index 000000000000..d72f14f5fe63
>> --- /dev/null
>> +++ b/docs/system/devices/igb-migration.rst
>> @@ -0,0 +1,417 @@
>> +.. SPDX-License-Identifier: GPL-2.0-or-later
>> +.. _igb-migration:
>> +
>> +igb VF Migration
>> +----------------
>> +
>> +Live migration of VFIO-passthrough devices (SR-IOV VFs, vGPUs) is a
>> +growing requirement, but real hardware with migration support is scarce
>> +and hard to debug. An emulated device provides a fully controlled
>> +testbed for developing and validating the entire software stack --
>> +vfio-pci variant drivers, VFIO core migration v2 framework, QEMU,
>> +libvirt -- and for tuning complex migration policies such as downtime
>> +convergence. It also serves as an educational reference for
>> +understanding VFIO migration end-to-end, from device state
>> +serialization to dirty page tracking.
>> +
>> +The igb device supports an experimental VF migration interface that allows
>> +the `igb-vfio-pci`_ variant driver to migrate VF state during live
>> +migration using the standard VFIO migration v2 protocol with stop-copy
>> +and pre-copy support.
>> +
>> +This is enabled with the ``x-vf-migration`` property::
>> +
>> + -device igb,x-vf-migration=on,...
>> +
>> +Each emulated VF then advertises a DVSEC discovered by the
>> +`igb-vfio-pci`_ variant driver at bind time. This feature is
>> +experimental (``x-`` prefix, default off).
>> +
>> +Architecture
>> +~~~~~~~~~~~~
>> +
>> +The target scenario is nested virtualization::
>> +
>> + L0 QEMU
>> + igb PF with x-vf-migration=on
>> + └── VFs with migration DVSEC
>> +
>> + L1 kernel
>> + igb-vfio-pci variant driver
>> + translates VFIO migration v2 ioctls → DVSEC config writes
>> +
>> + L1 QEMU (stock, unmodified)
>> + vfio-pci device model, standard migration fd
>> +
>> + L2 guest
>> + standard igbvf driver, unaware of migration
>> +
>> +The L1 QEMU is completely unmodified -- it sees a standard VFIO
>> +migratable device and uses the normal migration fd path. The
>> +`igb-vfio-pci`_ variant driver handles the translation between
>> +VFIO migration v2 ioctls and DVSEC config writes.
>> +
>> +Design
>> +~~~~~~
>> +
>> +The migration interface is exposed through a DVSEC at offset 0x160
>> +in VF extended config space (see `DVSEC register layout`_ below for
>> +the full register map).
>> +
>> +Device state is serialized as a versioned blob of per-VF register
>> +(offset, value) pairs covering control, interrupt, RX/TX queue,
>> +receive address (RA/RA2), etc. plus TX context descriptors and
>> +VFRE/VFTE enable bits. The buffer address is a guest physical address
>> +(GPA) written by the driver via ``virt_to_phys``; the device accesses
>> +guest RAM directly through the system address space.
>> +
>> +Dirty page tracking is implemented with per-range bitmaps maintained
>> +in IGBCore. All VF DMA paths in ``igb_core.c`` (TX data, RX data,
>> +descriptor writeback) are instrumented to record touched pages. The
>> +`igb-vfio-pci`_ variant driver registers tracked IOVA ranges and
>> +queries dirty bitmaps through a shared buffer. Buffer structures
>> +include len, flags, and reserved fields for future extensibility.
>> +
>> +The dirty bitmaps are maintained inside the device, which is not
>> +realistic for discrete NICs without on-chip DRAM.
>> +
>> +DVSEC register layout
>> +~~~~~~~~~~~~~~~~~~~~~
>> +
>> +The migration DVSEC (36 bytes at offset ``0x160``) uses a command
>> +doorbell model. All commands are synchronous -- the device completes
>> +the operation before the config write returns::
>> +
>> + Offset Name Access Description
>> + +0x00 ExtCap Hdr RO PCIe extended cap (id=0x23, ver=1)
>> + +0x04 DVSEC Hdr 1 RO length[31:20] | rev[19:16] | vendor_id[15:0]
>> + +0x08 DVSEC Hdr 2 RO DVSEC ID (1)
>> + +0x0A Reserved - Padding for DWORD alignment
>> + +0x0C CAPS RO F_STATE[0], F_DIRTY[1], max_ranges[11:8], pgsize[16:12]
>> + +0x10 CTRL WO Doorbell: cmd[7:0], arg[31:8]
>> + +0x14 STATUS RO state[7:0], error_code[15:8], quiesced[16]
>> + +0x18 BUF_ADDR_LO RW Shared DMA buffer GPA (low 32 bits)
>> + +0x1C BUF_ADDR_HI RW Shared DMA buffer GPA (high 32 bits)
>> + +0x20 DATA_SIZE RO State blob size in bytes
>> +
>> +CTRL commands::
>> +
>> + Cmd Name Arg Description
>> + 1 SET_STATE state[31:8] Set migration state
>> + 2 SAVE - DMA-write state to buffer
>> + 3 LOAD size[31:8] DMA-read state from buffer
>> + 4 DIRTY_ENABLE - Enable dirty tracking (params in DMA buffer)
>> + 5 DIRTY_DISABLE - Disable dirty tracking
>> + 6 DIRTY_QUERY - Query dirty bitmap (via DMA buffer)
>> + 7 GET_STATS - Query statistics (via DMA buffer)
>> +
>> +The driver sets ``BUF_ADDR_LO/HI`` before issuing commands that use a
>> +DMA buffer (SAVE, LOAD, DIRTY_ENABLE, DIRTY_QUERY, GET_STATS). The
>> +buffer address is latched per CTRL write, so the driver can use
>> +different buffers for different commands. The buffer address is a
>> +guest physical address (GPA).
>> +
>> +State transitions follow the VFIO migration v2 state machine. The
>> +driver issues ``SET_STATE`` with the target state in the arg field and
>> +reads ``STATUS`` to confirm the transition. Device states::
>> +
>> + 0 ERROR 1 STOP 2 RUNNING
>> + 3 STOP_COPY 4 RESUMING 5 PRE_COPY
>> +
>> +``DATA_SIZE`` reflects the state blob size. At reset and in ``STOP``
>> +state it holds the maximum size the driver should allocate. After
>> +``SET_STATE(STOP_COPY)`` or ``SAVE`` it holds the actual serialized
>> +size. The driver reads it after entering ``STOP_COPY`` to allocate an
>> +exact-sized DMA buffer before issuing ``SAVE``.
>> +
>> +The state blob is a versioned sequence of register (offset, value)
>> +pairs with magic ``0x4D494742`` ("MIGB").
>> +
>> +When ``STATUS`` state is ``ERROR`` (0), bits [15:8] contain an error
>> +code identifying the failure::
>> +
>> + 1 UNK_CMD Unknown CTRL command
>> + 2 BAD_STATE Command issued in wrong migration state
>> + 3 NO_BUFFER Command requires buffer but BUF_ADDR not set
>> + 4 DMA_FAILED DMA transfer to/from buffer failed
>> + 5 BAD_SIZE State blob too large or empty
>> + 6 BAD_MAGIC State blob magic mismatch
>> + 7 BAD_VERSION State blob version mismatch
>> + 8 TOO_MANY_RANGES Exceeds max_ranges from CAPS
>> + 9 BAD_RANGE Invalid range (zero size, misaligned, not contained)
>> + 10 BAD_PGSIZE Invalid or misaligned page size
>> + 11 NOT_ENABLED Dirty query without prior enable
>> +
>> +Dirty page tracking
>> +~~~~~~~~~~~~~~~~~~~
>> +
>> +The migration interface supports per-VF dirty page tracking, advertised
>> +by the ``F_DIRTY`` flag (bit 1) in ``CAPS``. This allows the variant
>> +driver to enter ``PRE_COPY`` state while the VM continues to run,
>> +iterating on dirty pages to reduce the final stop-and-copy window.
>> +
>> +The device maintains one dirty tracking engine per range, each with its
>> +own bitmap scoped to the range boundaries. The ``CAPS`` register
>> +advertises the maximum number of ranges in bits [11:8].
>> +
>> +Dirty tracking is controlled through CTRL commands:
>> +
>> +- **DIRTY_ENABLE** (4): the driver fills an ``igb_mig_dirty_enable_req``
>> + struct in the DMA buffer with the page size, range IOVA, and range
>> + size, then issues the command. The device allocates a bitmap for the
>> + range and begins recording pages touched by DMA. Supported page sizes
>> + are advertised in ``CAPS`` bits [16:12] (bit N = 2^N bytes). The driver
>> + checks ``STATUS`` for errors after the command completes.
>> +- **DIRTY_DISABLE** (5): tears down all ranges and stops tracking.
>> +- **DIRTY_QUERY** (6): the driver writes (iova, size) into the
>> + ``igb_mig_dirty_query`` DMA buffer, then issues the command. The
>> + device validates the range, copies the dirty bitmap into the buffer,
>> + and clears the tracked bits after successful DMA. The driver checks
>> + ``STATUS`` for errors after the command.
>> +
>> +Dirty enable DMA buffer
>> +~~~~~~~~~~~~~~~~~~~~~~~
>> +
>> +The ``DIRTY_ENABLE`` command reads its parameters from the DMA buffer.
>> +The ``len`` field holds the total structure size (including reserved
>> +bytes) so the device can detect newer formats. ``flags`` and
>> +``reserved`` must be zero::
>> +
>> + Offset Field Type Description
>> + 0x00 len uint32 Structure size in bytes
>> + 0x04 flags uint32 Reserved, must be 0
>> + 0x08 pgsize uint64 Page granularity (must match a CAPS pgsize bit)
>> + 0x10 range_iova uint64 Tracked range start address
>> + 0x18 range_size uint64 Tracked range size in bytes
>> + 0x20 reserved[4] uint32 Reserved, must be 0
>> +
>> +Dirty query DMA buffer
>> +~~~~~~~~~~~~~~~~~~~~~~
>> +
>> +The ``DIRTY_QUERY`` command uses a shared DMA buffer for both request
>> +and response. The ``len`` field holds the total buffer size (header +
>> +bitmap). ``flags`` and ``reserved`` must be zero::
>> +
>> + Offset Field Written by Description
>> + 0x00 len driver Total buffer size in bytes
>> + 0x04 flags driver Reserved, must be 0
>> + 0x08 iova driver Query range start
>> + 0x10 size driver Query range size
>> + 0x18 bitmap_size device Bytes written to bitmap
>> + 0x1C dirty_page_count device Number of set bits
>> + 0x20 dma_writes device DMA write count (diagnostic)
>> + 0x28 reserved[6] - Reserved, must be 0
>> + 0x40 bitmap[] device Dirty page bitmap
>> +
>> +Migration statistics
>> +~~~~~~~~~~~~~~~~~~~~
>> +
>> +The ``GET_STATS`` command DMA-writes a statistics response into the
>> +driver-provided buffer. The driver sets ``BUF_ADDR_LO/HI`` and issues
>> +the command; the device writes the response and returns::
>> +
>> + Offset Field Type Description
>> + 0x00 dma_writes uint64 DMA write operations tracked
>> + 0x08 dma_bytes uint64 DMA bytes written
>> + 0x10 dirty_pages_set uint32 Dirty pages marked since enable
>> + 0x14 dirty_pages_cleared uint32 Dirty pages cleared by queries
>> + 0x18 dirty_page_count uint32 Current dirty pages (set - cleared)
>> + 0x1C dirty_query_count uint32 Number of QUERY operations
>> +
>> +The variant driver exposes these via debugfs at
>> +``/sys/kernel/debug/vfio/<device>/migration/dirty/stats``.
>> +
>> +Testing setup
>> +~~~~~~~~~~~~~
>> +
>> +The target scenario is nested virtualization: L0 runs QEMU with an
>> +igb PF (``x-vf-migration=on``), L1 runs the `igb-vfio-pci`_ variant
>> +driver and an unmodified QEMU, and L2 runs a standard igbvf driver.
>> +See `Architecture`_ above for the full stack diagram.
>> +
>> +NetworkManager configuration
>> +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
>> +
>> +In a nested setup, the L1 VMs (source and destination) have emulated
>> +igb PFs connected to the L0 bridge. By default, NetworkManager
>> +acquires DHCP leases on those PF interfaces and on any igb VFs created
>> +later. This causes the VF MAC address to be learned on the L0 bridge,
>> +which can misdirect iperf3 traffic after migration.
>> +
>> +To prevent this, configure NetworkManager on **both L1 VMs** and on the
>> +**L2 guest disk image**.
>> +
>> +L1 VMs (source and destination)
>> +...............................
>> +
>> +1. Prevent NetworkManager from managing igbvf interfaces:
>> +
>> +.. code-block:: bash
>> +
>> + cat > /etc/NetworkManager/conf.d/99-no-igbvf.conf <<EOF
>> + [keyfile]
>> + unmanaged-devices=driver:igbvf
>> + EOF
>> +
>> +2. Disable IP on the igb PF connections (keep the interfaces UP for
>> + bridging, but with no DHCP lease):
>> +
>> +.. code-block:: bash
>> +
>> + # Identify the NM connections for the igb PFs (NOT the virtio management NIC)
>> + nmcli -t -f NAME,DEVICE connection show
>> +
>> + # For each igb PF connection:
>> + nmcli connection modify "<igb-pf-connection>" ipv4.method disabled ipv6.method disabled
>> +
>> +3. Reload NetworkManager:
>> +
>> +.. code-block:: bash
>> +
>> + nmcli general reload
>> +
>> +L2 guest disk image
>> +...................
>> +
>> +Use ``virt-customize`` to add the igbvf unmanaged config to the guest
>> +image (offline, before any test run):
>> +
>> +.. code-block:: bash
>> +
>> + virt-customize -a /srv/migration/rhel10.qcow2 \
>> + --write /etc/NetworkManager/conf.d/99-no-igbvf.conf:'[keyfile]
>> + unmanaged-devices=driver:igbvf'
>> +
>> +Network diagram
>> +^^^^^^^^^^^^^^^
>> +
>> +The diagram below shows the nested setup where the source and
>> +destination hosts are themselves VMs (L1) running on a physical
>> +host (L0) that emulates the igb NIC::
>> +
>> + ┌────────────────────────────────────────────────────────────────────────────┐
>> + │ L0: physical host │
>> + │ │
>> + │ virbr0 192.168.199.1/24 │
>> + │ ├── NFS server: /srv/migration │
>> + │ └── iperf3 client: iperf3 -c 192.168.199.200 -t 60 -i 1 │
>> + │ │ │
>> + │ │ L0 virbr0 bridge (192.168.199.0/24) │
>> + │ ────┼──────────┬──────────────────┬────────────────────────────── │
>> + │ │ │ │ │
>> + │ │ ┌────┴────┐ ┌────┴────┐ │
>> + │ │ │ virtio │ │ emulated│ │
>> + │ │ │c0:ff:ee:│ │ igb PF │ L0 QEMU (vm6) │
>> + │ │ │ :00:06 │ │ + igbvf │ tracks DMA dirty pages │
>> + │ │ └────┬────┘ └────┬────┘ │
>> + │ │ │ │ │
>> + │ ┌───┼──────────┼──────────────────┼──────────────────────────────────┐ │
>> + │ │ │ L1: vm6 (source) │ │ │
>> + │ │ │ enp1s0: 192.168.199.6 │ │ │
>> + │ │ │ (management) │ │ │
>> + │ │ │ enp8s0 (igb PF, no IP) │ │
>> + │ │ │ │ │ │
>> + │ │ │ igb VF0 ──► igb-vfio-pci (VFIO) │ │
>> + │ │ │ │ dirty_sync → L0 igbvf │ │
>> + │ │ │ │ │ │
>> + │ │ │ virbr0 │ VFIO passthrough │ │
>> + │ │ │ 192.168.200.1/24 │ │ │
>> + │ │ │ │ │ │ │
>> + │ │ │ ┌────┼─────────────────────┼───────────────────────────┐ │ │
>> + │ │ │ │ │ L2: rhel10 guest │ │ │ │
>> + │ │ │ │ │ │ │ │ │
>> + │ │ │ │ virtio NIC igb VF (enp7s0) │ │ │
>> + │ │ │ │ 192.168.200.130/24 192.168.199.200/24 │ │ │
>> + │ │ │ │ (SSH login) (iperf3 data path) │ │ │
>> + │ │ │ │ │ │ │ │
>> + │ │ │ │ iperf3 -s -D │ (listens on 0.0.0.0) │ │ │
>> + │ │ │ └──────────────────────────┼───────────────────────────┘ │ │
>> + │ │ │ │ │ │
>> + │ │ │ virsh migrate --live ──────┼──────────────────► vm7 │ │
>> + │ │ │ │ │ │
>> + │ └───┼─────────────────────────────┼──────────────────────────────────┘ │
>> + │ │ │ │
>> + │ │ iperf3 traffic │ │
>> + │ └─────────────────────────────┘ │
>> + │ │
>> + │ ──────────────────────────────────────────────────────────────── │
>> + │ │ │ │
>> + │ │ ┌────────┐ │ ┌─────────┐ │
>> + │ │ │ virtio │ │ │emulated │ L0 QEMU (vm7) │
>> + │ │ │c0:ff:ee│ │ │ igb PF │ │
>> + │ │ │ :00:07 │ │ │ + igbvf │ │
>> + │ │ └────┬───┘ │ └────┬────┘ │
>> + │ ┌──────────────┼───────┼────────┼────────────────────────────────────┐ │
>> + │ │ L1: vm7 (destination) │ │ │
>> + │ │ enp1s0: 192.168.199.7 │ │ │
>> + │ │ (management) enp8s0 (igb PF, no IP) │ │
>> + │ │ │ │ │
>> + │ │ igb VF0 ──► igb-vfio-pci (VFIO) │ │
>> + │ │ │ │ │
>> + │ │ virbr0 │ VFIO passthrough │ │
>> + │ │ 192.168.200.1/24 │ │ │
>> + │ │ │ │ │ │
>> + │ │ ┌────┼───────────────────────┼────────────────────────────┐ │ │
>> + │ │ │ │ L2: rhel10 (after migration) │ │ │
>> + │ │ │ │ │ │ │ │
>> + │ │ │ virtio NIC igb VF (enp7s0) │ │ │
>> + │ │ │ 192.168.200.130/24 192.168.199.200/24 │ │ │
>> + │ │ │ │ │ │ │
>> + │ │ │ iperf3 -s -D │ (connection survives) │ │ │
>> + │ │ └────────────────────────────┼────────────────────────────┘ │ │
>> + │ └───────────────────────────────┼────────────────────────────────────┘ │
>> + │ │ │
>> + │ iperf3 traffic resumes ─────┘ │
>> + │ (same IP, same MAC, same L2 segment → transparent to client) │
>> + └────────────────────────────────────────────────────────────────────────────┘
>> +
>> +Migration under iperf3 load works correctly: dirty page tracking
>> +converges (from ~2000 pages per PRE_COPY iteration down to ~280 at
>> +STOP_COPY), and STOP_COPY stays under 250ms.
>> +
>> +Todo
>> +~~~~
>> +
>> +1. Add migration blocker when ``x-vf-migration=on`` (no VMState yet) or
>> + add VMState support for L0 migration (dirty bitmaps, tracking
>> + engines, DVSEC registers, stats)
>> +2. Add PRE_COPY state transfer to validate device INIT data (magic,
>> + version, etc.)
>> +3. Add qtests for migration state machine transitions, dirty page
>> + tracking
>> +
>> +Ideas
>> +~~~~~
>> +
>> +1. **RX bandwidth throttle** (``x-mig-rx-limit``, uint32, default 0)
>> +
>> + Return false from ``can_receive`` when the per-VF packet count in the
>> + current tracking interval exceeds the limit. Reduces DMA writes and
>> + dirty pages realistically.
>> +
>> +2. **Migration phase timing** (GET_STATS extension)
>> +
>> + Add per-VF timestamps: ``precopy_start_ns``, ``stopcopy_start_ns``,
>> + ``precopy_duration_ns``, ``stopcopy_duration_ns``,
>> + ``state_transition_count``. Expose via GET_STATS.
>> +
>> +3. **Hot page simulation** (``x-mig-hot-pages``, uint32, default 0)
>> +
>> + Re-set the first N bitmap bits after each DIRTY_QUERY, simulating
>> + workloads with hot pages that prevent convergence.
>> +
>> +4. **Error injection** (``x-mig-inject-error``, uint32, default 0)
>> +
>> + One-shot error code injection before command dispatch. A separate
>> + ``x-mig-inject-dma-fail`` (bool) for persistent DMA failure testing.
>> +
>> +AI disclaimer
>> +~~~~~~~~~~~~~
>> +
>> +Claude was used to analyze the IGB PF and VF internal state and
>> +identify the pain points of a working live migration of such devices.
>> +The generated code served as a starting point but *significant* time
>> +was then spent cleaning up, reworking, and shaping it into a clear,
>> +reviewable IGB model extension.
>
> I don't think it's really a good idea to keep this "AI disclaimer" in the documentation.
Agree. No need.
Thanks,
C.
>
>> +
>> +.. _igb-vfio-pci: https://github.com/legoater/vfio-pci-extras
>> diff --git a/docs/system/devices/igb.rst b/docs/system/devices/igb.rst
>> index 50f625fd77e4..00271dbc92c3 100644
>> --- a/docs/system/devices/igb.rst
>> +++ b/docs/system/devices/igb.rst
>> @@ -64,6 +64,12 @@ command:
>> pyvenv/bin/meson test --suite thorough func-x86_64-netdev_ethtool
>> +VF Migration (experimental)
>> +===========================
>> +
>> +See :ref:`igb-migration` for details on the experimental VF live migration
>> +interface.
>> +
>> References
>> ==========
>
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [RFC PATCH v2 0/9] igb: Add experimental VF live migration support
2026-09-09 7:19 ` [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Akihiko Odaki
@ 2026-09-15 8:20 ` Cédric Le Goater
2026-09-17 19:49 ` Akihiko Odaki
0 siblings, 1 reply; 23+ messages in thread
From: Cédric Le Goater @ 2026-09-15 8:20 UTC (permalink / raw)
To: Akihiko Odaki, qemu-devel
Cc: Sriram Yagnaraman, Jason Wang, Alex Williamson, Peter Xu
On 9/9/26 09:19, Akihiko Odaki wrote:
> On 2026/09/03 4:20, Cédric Le Goater wrote:
>> Hello,
>>
>> Live migration of VFIO-passthrough devices - SR-IOV VFs, vGPUs - is a
>> growing requirement, but real hardware with migration support is
>> scarce and hard to debug. An emulated device provides a fully
>> controlled testbed for developing and validating the entire software
>> stack - vfio-pci variant drivers, VFIO core migration v2 framework,
>> QEMU, libvirt - and for tuning complex migration policies such as
>> downtime convergence. It also serves as an educational reference for
>> understanding VFIO migration end-to-end, from device state
>> serialization to dirty page tracking.
>>
>> This series adds an experimental VF live migration interface to the
>> emulated igb (82576) device. It enables a vfio-pci variant driver
>> (igb-vfio-pci) to migrate VFs using the standard VFIO migration v2
>> protocol with stop-copy and pre-copy support.
>>
>> The target scenario is nested virtualization:
>>
>> L0 QEMU (these patches)
>> igb PF with x-vf-migration=on
>> └── VFs with migration DVSEC
>>
>> L1 kernel
>> igb-vfio-pci variant driver [1]
>> translates VFIO migration v2 ioctls → DVSEC config writes
>>
>> L1 QEMU (stock, unmodified)
>> vfio-pci device model, standard migration fd
>>
>> L2 guest
>> standard igbvf driver, unaware of migration
>>
>> The L1 QEMU is completely unmodified -- it sees a standard VFIO
>> migratable device and uses the normal migration fd path.
>>
>> * Design
>>
>> The migration interface is exposed through a DVSEC (Designated
>> Vendor-Specific Extended Capability, PCIe cap id 0x23) at offset
>> 0x160 in VF extended config space. The DVSEC uses a command doorbell
>> model - all commands are synchronous via PCI config space writes.
>>
>> Device state is serialized as a versioned blob of per-VF register
>> (offset, value) pairs covering control, interrupt, RX/TX queue,
>> receive address (RA/RA2), etc. plus TX context descriptors and
>> VFRE/VFTE enable bits. The buffer address is a guest physical
>> address (GPA) written by the driver via virt_to_phys; the device
>> accesses guest RAM directly through the system address space.
>>
>> Dirty page tracking is implemented with per-range bitmaps maintained
>> in IGBCore. All VF DMA paths in igb_core.c (TX data, RX data,
>> descriptor writeback) are instrumented to record touched pages. The
>> variant driver registers tracked IOVA ranges and queries dirty bitmaps
>> through a shared buffer. Buffer structures include len, flags, and
>> reserved fields for future extensibility.
>>
>> * Caveats
>>
>> The x-vf-migration property is experimental (x- prefix, default off).
>>
>> The dirty bitmaps are maintained inside the device, which is not
>> realistic for discrete NICs without on-chip DRAM.
>>
>> * Testing
>>
>> The target scenario is nested virtualization: L0 runs QEMU with an
>> igb PF (x-vf-migration=on), L1 runs the igb-vfio-pci variant driver
>> and an unmodified QEMU, and L2 runs a standard igbvf driver.
>>
>> Migration under iperf3 load works correctly: dirty page tracking
>> converges (from ~2000 pages per PRE_COPY iteration down to ~280 at
>> STOP_COPY), and STOP_COPY stays under 250ms.
>>
>> * Todo
>>
>> 1. Add migration blocker when x-vf-migration=on (no VMState yet) or
>> add VMState support for L0 migration (dirty bitmaps, tracking
>> engines, DVSEC registers, stats)
>> 2. Add PRE_COPY state transfer to validate device INIT data (magic,
>> version, etc.)
>> 3. Add qtests for migration state machine transitions, dirty page
>> tracking ?
>>
>> * Ideas
>>
>> 1. RX bandwidth throttle (x-mig-rx-limit, uint32, default 0)
>>
>> Return false from can_receive when the per-VF packet count in the
>> current tracking interval exceeds the limit. Reduces DMA writes
>> and dirty pages realistically.
>>
>> 2. Migration phase timing (GET_STATS extension)
>>
>> Add per-VF timestamps: precopy_start_ns, stopcopy_start_ns,
>> precopy_duration_ns, stopcopy_duration_ns,
>> state_transition_count. Expose via GET_STATS.
>>
>> 3. Hot page simulation (x-mig-hot-pages, uint32, default 0)
>>
>> Re-set the first N bitmap bits after each DIRTY_QUERY, simulating
>> workloads with hot pages that prevent convergence.
>>
>> 4. Error injection (x-mig-inject-error, uint32, default 0)
>>
>> One-shot error code injection before command dispatch. A separate
>> x-mig-inject-dma-fail (bool) for persistent DMA failure testing.
>>
>> * Credits
>>
>> Alex Williamson suggested the overall approach of a variant driver
>> with the "x-vf-migration" device property to gate the feature. Thanks
>> for the ever ongoing support and valuable discussions throughout these
>> years.
>>
>> * AI disclaimer
>>
>> The lack of a migration-capable device has been a recurring pain point
>> for VFIO development over the years, and we hope this proposal
>> demonstrates the value of having one.
>>
>> Claude was used to analyze the IGB PF and VF internal state and
>> identify the pain points of a working live migration of such devices.
>> The generated code served as a starting point but *significant* time
>> was then spent cleaning up, reworking, and shaping it into a clear,
>> reviewable IGB model extension.
>>
>> As QEMU does not yet accept AI-assisted contributions, this series is
>> submitted as an RFC.
>
> This largely repeats the concerns I raised previously [1]. Your description of the *significant* time spent cleaning up and reworking the generated code suggests that this series is intended for upstream inclusion once the implementation direction is agreed.
Yes. The IGB kernel vfip-pci variant driver is "ready", and so is the
internal test framework for VFIO migration. The emulated IGB device
is also becoming increasingly important as a reference device for
the kernel VFIO self-tests. See:
https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit?id=1583ef1175ba533e326dfeff6d8b9f3aaf761225
> My concerns as a maintainer are therefore correctness and maintainability.
I am trying to avoid too much intrusion in the core igb model :
hw/net/igb_common.h | 16 +
hw/net/igb_core.h | 11 +
hw/net/igb.c | 7 +
hw/net/igb_core.c | 124 ++-
hw/net/igbvf.c | 52 +-
These are mostly code movement or additions. The runtime is hardly
impacted apart from the DMA tracking, which adds one indirection.
I have added myself as Maintainer for the migration part :
+igb VF migration
+M: Cédric Le Goater <clg@redhat.com>
+S: Maintained
+F: hw/net/igb_migration.*
I can add you in the next respin. That would be a natural choice
from my perspective, but I don't want to impose it on you.
> The missing migration blocker is a concrete example: it is a basic correctness issue that should be straightforward to address, yet it remains a TODO.
yes. It's identified and easy to address. It can come at the end once
all other issues have been discussed. What worries me the most is the
control plane : interface with the Linux driver. That is the most
important part on which these changes depend on getting it right :
https://github.com/legoater/linux/commits/vfio/
> My comments on individual patches also raise the same kinds of issues I pointed out in the previous version.
You're being tough on me ! I did all this fixing below ! :)
* Changes since rfc-v1
- Migration BAR replaced with DVSEC at offset 0x160 (no BAR needed)
- Wire-format structs (IgbMigBlob, IgbMigRegPair, IgbMigTxCtx)
replace raw pointer arithmetic and memcpy
- RA entries separated from fixed regs, scanned by pool bit
- VFN relocation support (offset remapping + RA pool bit swap)
- GPA buffer (address_space_read/write) replaces PCI DMA through PF
- State blob validation on load (magic, version, error codes)
- NEED_WORDS macro and igb_vf_offset_valid removed
- Error codes renumbered: removed BAD_VFN, added UNK_CMD (1-11)
Reported by Akihiko Odaki:
- propagate_irqs: clear VF bits before OR (EIMS/EIAC/EIAM)
- propagate_ivar: clear IVAR entry when source VTIVAR is invalid
- rearm_irqs: restore actual PVTEICR causes, not all three
- Dirty bitmap allocation uses BITS_TO_LONGS (heap corruption fix)
- Dirty query validates range before g_malloc0 (memory exhaustion)
- Dirty bits cleared only after successful bitmap DMA write
- Dirty query buffer uses struct offsets (layout mismatch fix)
- Load path: register offsets validated against VF whitelist
- VMBMEM (mailbox payload) documented as transient, not serialized
- dma_writes counter: consistently uint64_t
- Dirty range_size: consistently uint64_t (was truncated to 32 bits)
- ERROR->STOP: quiesce VF (clear VFRE/VFTE) on transition
- rearm_irqs: runs on re || te, not just re (TX-only VF fix)
- Stats DMA-written atomically via GET_STATS (no split MMIO tear)
- Bisectability: DVSEC + state machine introduced together
- Removed NAPI reference and "===" comment decoration
>
> Many of the correctness problems found in v1 and v2 stem from the choice to build on igb, which is more complex than virtio-net. That is why I suggested using virtio-net as a simpler basis, unless there is a problem that rules it out.
virtio-net is a software abstraction designed for virtualization.
Here’s my pitch for igb:
Using an emulated igb device addresses a different problem: validating
migration of hardware-like devices. igb exposes a real PCI device
model, hardware-driver interface, DMA engines, queues, interrupts and
SR-IOV VFs. It allows migration to be implemented at the same
abstraction boundary at which VFIO migrates physical devices, rather
than being reduced to serialization of a paravirtualized protocol
state. Also,
The igb VF uses the inbox igbvf kernel driver, which matches the
hardware VFIO promise: transparent migration with unmodified guest
drivers.
The DVSEC extended capability is a discoverable and standard control
plane with the device. This is a mechanism a hardware vendor could
use.
One cool aspect of an emulated device is that it allows the state to
be inspected, the behavior is deterministic and faults can be injected
at any boundary. That's the QEMU bonus.
> I would be willing to support allowing AI use in the project for work like this. I would still need to see these correctness and maintainability concerns addressed before I could support merging the series.
I agree. The save/load is still in progress and we might need some
more intrusive changes in igb to support migration of all state.
To start with, we could detect and warn.
Thanks for the support,
C.
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [RFC PATCH v2 3/9] igb: Add VF state serialization for live migration
2026-09-15 7:47 ` Cédric Le Goater
@ 2026-09-17 18:59 ` Akihiko Odaki
0 siblings, 0 replies; 23+ messages in thread
From: Akihiko Odaki @ 2026-09-17 18:59 UTC (permalink / raw)
To: Cédric Le Goater, qemu-devel
Cc: Sriram Yagnaraman, Jason Wang, Alex Williamson, Peter Xu
On 2026/09/15 16:47, Cédric Le Goater wrote:
> Hello Akihiko,
>
> On 9/8/26 10:01, Akihiko Odaki wrote:
>>
>>
>> On 2026/09/03 4:20, Cédric Le Goater wrote:
>>> Implement per-VF state serialization and deserialization for the SAVE
>>> and LOAD commands. The wire format consists of a header (magic,
>>> version, VF number, register count), per-VF register offset/value
>>> pairs from a whitelist, RA table entries owned by the VF, and TX
>>> queue contexts.
>>>
>>> Register offsets are relocated on load so a VF can migrate to a
>>> different VF number on the destination. RA pool ownership bits are
>>> swapped accordingly.
>>>
>>> AI-used-for: analysis, code (prototype)
>>> Signed-off-by: Cédric Le Goater <clg@redhat.com>
>>> ---
>>> hw/net/igb_core.h | 2 +
>>> hw/net/igb_migration.h | 2 +
>>> hw/net/igb.c | 5 +
>>> hw/net/igb_migration.c | 347 ++++++++++++++++++++++++++++++++++++++++-
>>> 4 files changed, 354 insertions(+), 2 deletions(-)
>>>
>>> diff --git a/hw/net/igb_core.h b/hw/net/igb_core.h
>>> index d70b54e318f1..60724e2824ab 100644
>>> --- a/hw/net/igb_core.h
>>> +++ b/hw/net/igb_core.h
>>> @@ -143,4 +143,6 @@ igb_receive_iov(IGBCore *core, const struct iovec
>>> *iov, int iovcnt);
>>> void
>>> igb_start_recv(IGBCore *core);
>>> +IGBCore *igb_pf_get_core(void *pf);
>>> +
>>> #endif
>>> diff --git a/hw/net/igb_migration.h b/hw/net/igb_migration.h
>>> index ea40ac65c54b..b2f601e74346 100644
>>> --- a/hw/net/igb_migration.h
>>> +++ b/hw/net/igb_migration.h
>>> @@ -73,6 +73,8 @@
>>> #define IGB_MIG_ERR_NO_BUFFER 3
>>> #define IGB_MIG_ERR_DMA_FAILED 4
>>> #define IGB_MIG_ERR_BAD_SIZE 5
>>> +#define IGB_MIG_ERR_BAD_MAGIC 6
>>> +#define IGB_MIG_ERR_BAD_VERSION 7
>>> /* Shared buffer constants */
>>> #define IGB_VF_STATE_MAX_SIZE 4096
>>> diff --git a/hw/net/igb.c b/hw/net/igb.c
>>> index 7268e5473fc3..f39f2bc3a04e 100644
>>> --- a/hw/net/igb.c
>>> +++ b/hw/net/igb.c
>>> @@ -133,6 +133,11 @@ void igb_vf_reset(void *opaque, uint16_t vfn)
>>> igb_core_vf_reset(&s->core, vfn);
>>> }
>>> +IGBCore *igb_pf_get_core(void *pf)
>>> +{
>>> + return &IGB(pf)->core;
>>> +}
>>> +
>>> static bool
>>> igb_io_get_reg_index(IGBState *s, uint32_t *idx)
>>> {
>>> diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
>>> index 8e7e6fac9b9b..c34035974620 100644
>>> --- a/hw/net/igb_migration.c
>>> +++ b/hw/net/igb_migration.c
>>> @@ -10,18 +10,241 @@
>>> #include "qemu/log.h"
>>> #include "hw/pci/pci_device.h"
>>> #include "hw/pci/pcie.h"
>>> +#include "net/eth.h"
>>> +#include "net/net.h"
>>> #include "igb_common.h"
>>> +#include "igb_core.h"
>>> #include "igb_migration.h"
>>> #include "system/address-spaces.h"
>>> #include "trace.h"
>>> +static IGBCore *igbvf_get_core(IgbVfState *s)
>>> +{
>>> + return igb_pf_get_core(pcie_sriov_get_pf(PCI_DEVICE(s)));
>>> +}
>>> +
>>> /*
>>> * Per-VF state serialization / deserialization
>>> */
>>> +#define IGB_MIG_BLOB_MAGIC 0x4D494742 /* "MIGB" */
>>> +#define IGB_MIG_BLOB_VERSION 1
>>> +
>>> +typedef struct IgbMigRegPair {
>>> + uint32_t offset;
>>> + uint32_t value;
>>> +} IgbMigRegPair;
>>> +
>>> +typedef struct IgbMigTxCtx {
>>> + uint32_t ctx_desc[8]; /* 2 × adv_tx_context_desc (4
>>> dwords each) */
>>> + uint32_t first_cmd_type_len;
>>> + uint32_t first_olinfo_status;
>>> + uint32_t first;
>>> + uint32_t skip_cp;
>>> +} IgbMigTxCtx;
>>> +
>>> +#define IGB_VF_MAX_FIXED_REGS 64
>>> +#define IGB_VF_MAX_RA_REGS 48 /* (16 + 8) RA entries × 2
>>> (RAL+RAH) */
>>> +
>>> +typedef struct IgbMigBlob {
>>> + uint32_t magic;
>>> + uint32_t version;
>>> + uint32_t vfn;
>>> + uint32_t num_regs;
>>> + IgbMigRegPair regs[IGB_VF_MAX_FIXED_REGS];
>>> + uint32_t num_ra;
>>> + IgbMigRegPair ra[IGB_VF_MAX_RA_REGS];
>>> + uint32_t num_tx_ctx;
>>> + IgbMigTxCtx tx_ctx[2];
>>> +} IgbMigBlob;
>>> +
>>> +#define IGB_MIG_BLOB_SIZE sizeof(IgbMigBlob)
>>> +
>>> +QEMU_BUILD_BUG_ON(IGB_MIG_BLOB_SIZE > IGB_VF_STATE_MAX_SIZE);
>>> +
>>> +/* Register offsets that constitute a VF's state slice */
>>> +static int igb_vf_reg_list(uint16_t vfn, uint32_t *offsets)
>>> +{
>>> + int n = 0;
>>> + int q0 = vfn;
>>> + int q1 = vfn + IGB_NUM_VM_POOLS;
>>> +
>>> + /* Per-VF control and interrupt registers */
>>> + offsets[n++] = E1000_PVTCTRL(vfn) >> 2;
>>> + offsets[n++] = E1000_PVTEICS(vfn) >> 2;
>>> + offsets[n++] = E1000_PVTEIMS(vfn) >> 2;
>>> + offsets[n++] = E1000_PVTEIMC(vfn) >> 2;
>>> + offsets[n++] = E1000_PVTEIAC(vfn) >> 2;
>>> + offsets[n++] = E1000_PVTEIAM(vfn) >> 2;
>>> + offsets[n++] = E1000_PVTEICR(vfn) >> 2;
>>> +
>>> + /* Per-VF statistics */
>>> + offsets[n++] = E1000_PVFGPRC(vfn) >> 2;
>>> + offsets[n++] = E1000_PVFGPTC(vfn) >> 2;
>>> + offsets[n++] = E1000_PVFGORC(vfn) >> 2;
>>> + offsets[n++] = E1000_PVFGOTC(vfn) >> 2;
>>> + offsets[n++] = E1000_PVFMPRC(vfn) >> 2;
>>> + offsets[n++] = E1000_PVFGPRLBC(vfn) >> 2;
>>> + offsets[n++] = E1000_PVFGPTLBC(vfn) >> 2;
>>> + offsets[n++] = E1000_PVFGORLBC(vfn) >> 2;
>>> + offsets[n++] = E1000_PVFGOTLBC(vfn) >> 2;
>>> +
>>> + /*
>>> + * Mailbox control registers only - the 16-dword payload buffer
>>> + * (VMBMEM) is transient and drained on quiesce.
>>> + */
>>> + offsets[n++] = E1000_V2PMAILBOX(vfn) >> 2;
>>> + offsets[n++] = E1000_P2VMAILBOX(vfn) >> 2;
>>> +
>>> + /* Per-VF config */
>>> + offsets[n++] = E1000_VMOLR(vfn) >> 2;
>>> + offsets[n++] = E1000_VMVIR(vfn) >> 2;
>>> + offsets[n++] = E1000_PSRTYPE(vfn) >> 2;
>>> +
>>> + /*
>>> + * VF receive addresses (RA/RA2) are saved dynamically in
>>> + * igb_core_vf_save_state by scanning for entries whose pool
>>> + * bits match this VF - the PF driver chooses the RA slot.
>>> + */
>>> +
>>> + /* Interrupt routing */
>>> + offsets[n++] = (E1000_VTIVAR + vfn * 4) >> 2;
>>> + offsets[n++] = (E1000_VTIVAR_MISC + vfn * 4) >> 2;
>>> +
>>> + /*
>>> + * EITR (Extended Interrupt Throttle Register) - 3 vectors per VF.
>>> + * Each VF has 3 MSI-X vectors, each with its own EITR controlling
>>> + * interrupt coalescing. Without saving these, interrupt
>>> + * throttling resets to zero after migration which can cause
>>> + * interrupt storms or latency changes. VF N uses PF EITR indices
>>> + * (22 - N*3) .. (24 - N*3).
>>> + */
>>> + {
>>> + int eitr_base = 22 - vfn * 3;
>>> + offsets[n++] = E1000_EITR(eitr_base) >> 2;
>>> + offsets[n++] = E1000_EITR(eitr_base + 1) >> 2;
>>> + offsets[n++] = E1000_EITR(eitr_base + 2) >> 2;
>>> + }
>>> +
>>> + /* RX and TX queue registers for queues q0 and q1 */
>>> +#define ADD_QUEUE_REGS(q) do { \
>>> + offsets[n++] = E1000_RDBAL(q) >> 2; \
>>> + offsets[n++] = E1000_RDBAH(q) >> 2; \
>>> + offsets[n++] = E1000_RDLEN(q) >> 2; \
>>> + offsets[n++] = E1000_SRRCTL(q) >> 2; \
>>> + offsets[n++] = E1000_RDH(q) >> 2; \
>>> + offsets[n++] = E1000_RDT(q) >> 2; \
>>> + offsets[n++] = E1000_RXDCTL(q) >> 2; \
>>> + offsets[n++] = E1000_RXCTL(q) >> 2; \
>>> + offsets[n++] = E1000_RQDPC(q) >> 2; \
>>> + offsets[n++] = E1000_TDBAL(q) >> 2; \
>>> + offsets[n++] = E1000_TDBAH(q) >> 2; \
>>> + offsets[n++] = E1000_TDLEN(q) >> 2; \
>>> + offsets[n++] = E1000_TDH(q) >> 2; \
>>> + offsets[n++] = E1000_TDT(q) >> 2; \
>>> + offsets[n++] = E1000_TXDCTL(q) >> 2; \
>>> + offsets[n++] = E1000_TXCTL(q) >> 2; \
>>> + offsets[n++] = E1000_TDWBAL(q) >> 2; \
>>> + offsets[n++] = E1000_TDWBAH(q) >> 2; \
>>> +} while (0)
>>> +
>>> + ADD_QUEUE_REGS(q0);
>>> + ADD_QUEUE_REGS(q1);
>>> +#undef ADD_QUEUE_REGS
>>> +
>>> + g_assert(n <= IGB_VF_MAX_FIXED_REGS);
>>> + return n;
>>> +}
>>> +
>>> +/*
>>> + * Scan RA and RA2 arrays for receive address entries assigned to
>>> + * this VF. The PF driver picks the RA slot, so we cannot use a
>>> + * fixed index - instead check each entry's pool bits.
>>> + */
>>> +static int igb_core_vf_save_ra(IGBCore *core, uint16_t vfn,
>>> + IgbMigRegPair *regs)
>>> +{
>>> + uint32_t vf_pool_bit = E1000_RAH_POOL_1 << vfn;
>>> + int n = 0;
>>> + static const struct {
>>> + uint32_t base;
>>> + int count;
>>> + } ra_banks[] = {
>>> + { RA, 16 },
>>> + { RA2, 8 },
>>> + };
>>> +
>>> + for (int i = 0; i < ARRAY_SIZE(ra_banks); i++) {
>>> + for (int j = 0; j < ra_banks[i].count; j++) {
>>> + uint32_t ral_off = ra_banks[i].base + j * 2;
>>> + uint32_t rah_off = ra_banks[i].base + j * 2 + 1;
>>> + uint32_t rah_val = core->mac[rah_off];
>>> +
>>> + if ((rah_val & E1000_RAH_AV) && (rah_val & vf_pool_bit)) {
>>> + regs[n].offset = cpu_to_le32(ral_off);
>>> + regs[n].value = cpu_to_le32(core->mac[ral_off]);
>>> + n++;
>>> + regs[n].offset = cpu_to_le32(rah_off);
>>> + regs[n].value = cpu_to_le32(rah_val);
>>> + n++;
>>> + }
>>> + }
>>> + }
>>> + return n;
>>> +}
>>> +
>>> +static void igb_core_vf_save_tx_ctx(IGBCore *core, int queue,
>>> + IgbMigTxCtx *tx)
>>> +{
>>> + struct igb_tx *src = &core->tx[queue];
>>> +
>>> + memcpy(tx->ctx_desc, src->ctx, sizeof(tx->ctx_desc));
>>> + tx->first_cmd_type_len = cpu_to_le32(src->first_cmd_type_len);
>>> + tx->first_olinfo_status = cpu_to_le32(src->first_olinfo_status);
>>> + tx->first = cpu_to_le32(src->first);
>>> + tx->skip_cp = cpu_to_le32(src->skip_cp);
>>> +}
>>> +
>>> static int igb_core_vf_save_state(IgbVfState *s, void *buf, size_t
>>> buf_size)
>>> {
>>> - int size = 0;
>>> + int size = IGB_MIG_BLOB_SIZE;
>>> + IGBCore *core = igbvf_get_core(s);
>>> + IgbMigBlob *blob = buf;
>>> + uint32_t offsets[IGB_VF_MAX_FIXED_REGS];
>>> + int num_regs;
>>> + int q0 = s->vfn;
>>> + int q1 = s->vfn + IGB_NUM_VM_POOLS;
>>> +
>>> + /*
>>> + * Save PVT shadow registers (PVTEIMS/PVTEIAC/PVTEIAM) instead of
>>> + * extracting from PF aggregates - the L1 PF driver may have
>>> + * transiently cleared EIMS via EIMC. The load path ORs them back.
>>> + */
>>> + num_regs = igb_vf_reg_list(s->vfn, offsets);
>>> +
>>> + if (!buf) {
>>> + return size;
>>> + }
>>> +
>>> + if (size > buf_size) {
>>> + return -IGB_MIG_ERR_BAD_SIZE;
>>> + }
>>> +
>>> + blob->magic = cpu_to_le32(IGB_MIG_BLOB_MAGIC);
>>> + blob->version = cpu_to_le32(IGB_MIG_BLOB_VERSION);
>>> + blob->vfn = cpu_to_le32(s->vfn);
>>> +
>>> + blob->num_regs = cpu_to_le32(num_regs);
>>> + for (int i = 0; i < num_regs; i++) {
>>> + blob->regs[i].offset = cpu_to_le32(offsets[i]);
>>> + blob->regs[i].value = cpu_to_le32(core->mac[offsets[i]]);
>>> + }
>>> +
>>> + blob->num_ra = cpu_to_le32(igb_core_vf_save_ra(core, s->vfn,
>>> blob->ra));
>>
>> The serialized state contains RA entries but excludes VFTA/VLVF and
>> MTA/UTA state.
>
> VLVF should not be too complex to add support for. It is the same pattern
> as the RA scan.
>
> VFTA/MTA/UTA support is tougher. IIUC, these are PF shared tables and there
> are no per-VF bits. So we either need to add per-VF shadow tracking, which
> is intrusive or, to begin with, we can ignore: detect and warn.
>
>
>> A guest-configured VLAN therefore resumes on a destination lacking its
>> filter membership: packets are rejected by the global VLAN check or
>> have their VF queue removed in igb_receive_assign(),. Likewise,
>> restored VMOLR.ROMPE cannot receive subscribed multicast without its
>> MTA hash bits. These are guest-requested settings communicated through
>> the PF mailbox, and the resumed guest does not recreate them
>> automatically.
>
> yes. If we could restore some of the VF state using the PF mailbox, we
> would avoid poking the core igb device from the VF but the model doesn't
> have enough support yet.
I mentioned the PF mailbox to explain how those settings were
established, not to suggest using it to transport or restore migration
state.
The complication is that a VF request through the mailbox can result in
configuration held in shared PF state, outside the VF registers.
Restoring the VF registers alone therefore does not restore everything
the guest previously configured. The resumed guest already considers
those requests complete, so we cannot rely on it to issue them again.
Migration needs to preserve the effects of those requests while
respecting the destination's existing allocations and keeping the PF's
resource management consistent. Mailbox replay might be one way to
coordinate that, but my original point was about the state that should
not be silently discarded, rather than the mechanism used to restore it.
>
>>
>>> +
>>> + blob->num_tx_ctx = cpu_to_le32(2);
>>> + igb_core_vf_save_tx_ctx(core, q0, &blob->tx_ctx[0]);
>>> + igb_core_vf_save_tx_ctx(core, q1, &blob->tx_ctx[1]);
>>> trace_igbvf_mig_save_state(s->vfn, size);
>>> return size;
>>> @@ -29,11 +252,131 @@ static int igb_core_vf_save_state(IgbVfState
>>> *s, void *buf, size_t buf_size)
>>> static int igb_core_vf_max_data_size(IgbVfState *s)
>>> {
>>> - return sizeof(s->mig.mig_data);
>>> + int size = igb_core_vf_save_state(s, NULL, 0);
>>> +
>>> + g_assert(size > 0 && size <= IGB_VF_STATE_MAX_SIZE);
>>> + return size;
>>> +}
>>> +
>>> +static void igb_core_vf_load_tx_ctx(IGBCore *core, int queue,
>>> + const IgbMigTxCtx *tx)
>>> +{
>>> + struct igb_tx *dst = &core->tx[queue];
>>> +
>>> + /*
>>> + * Preserve the destination's tx_pkt - it's a host-side object,
>>> + * not guest state
>>> + */
>>> + memcpy(dst->ctx, tx->ctx_desc, sizeof(dst->ctx));
>>> + dst->first_cmd_type_len = le32_to_cpu(tx->first_cmd_type_len);
>>> + dst->first_olinfo_status = le32_to_cpu(tx->first_olinfo_status);
>>> + dst->first = le32_to_cpu(tx->first);
>>> + dst->skip_cp = le32_to_cpu(tx->skip_cp);
>>> +}
>>> +
>>> +static uint32_t igb_vf_relocate_offset(uint32_t offset,
>>> + const uint32_t *src_offsets,
>>> + const uint32_t *dst_offsets,
>>> + int num_offsets)
>>> +{
>>> + for (int i = 0; i < num_offsets; i++) {
>>> + if (src_offsets[i] == offset) {
>>> + return dst_offsets[i];
>>> + }
>>> + }
>>> + return 0;
>>> }
>>> static int igb_core_vf_load_state(IgbVfState *s, const void *buf,
>>> size_t size)
>>> {
>>> + IGBCore *core = igbvf_get_core(s);
>>> + uint32_t src_offsets[IGB_VF_MAX_FIXED_REGS];
>>> + uint32_t dst_offsets[IGB_VF_MAX_FIXED_REGS];
>>> + int q0 = s->vfn;
>>> + int q1 = s->vfn + IGB_NUM_VM_POOLS;
>>> +
>>> + if (size < IGB_MIG_BLOB_SIZE) {
>>> + return -IGB_MIG_ERR_BAD_SIZE;
>>> + }
>>> +
>>> + const IgbMigBlob *blob = buf;
>>> +
>>> + uint32_t magic = le32_to_cpu(blob->magic);
>>> + uint32_t version = le32_to_cpu(blob->version);
>>> + uint32_t saved_vfn = le32_to_cpu(blob->vfn);
>>> + uint32_t num_regs = le32_to_cpu(blob->num_regs);
>>> +
>>> + if (magic != IGB_MIG_BLOB_MAGIC) {
>>> + return -IGB_MIG_ERR_BAD_MAGIC;
>>> + }
>>> + if (version != IGB_MIG_BLOB_VERSION) {
>>> + return -IGB_MIG_ERR_BAD_VERSION;
>>> + }
>>> + if (num_regs > IGB_VF_MAX_FIXED_REGS) {
>>> + return -IGB_MIG_ERR_BAD_SIZE;
>>> + }
>>> +
>>> + uint32_t num_ra = le32_to_cpu(blob->num_ra);
>>> + if (num_ra > IGB_VF_MAX_RA_REGS) {
>>> + return -IGB_MIG_ERR_BAD_SIZE;
>>> + }
>>> +
>>> + int num_offsets = igb_vf_reg_list(saved_vfn, src_offsets);
>>> + igb_vf_reg_list(s->vfn, dst_offsets);
>>> +
>>> + for (uint32_t i = 0; i < num_regs; i++) {
>>> + uint32_t src_off = le32_to_cpu(blob->regs[i].offset);
>>> + uint32_t value = le32_to_cpu(blob->regs[i].value);
>>> + uint32_t offset = igb_vf_relocate_offset(src_off,
>>> + src_offsets, dst_offsets, num_offsets);
>>> + if (!offset) {
>>> + return -IGB_MIG_ERR_BAD_SIZE;
>>> + }
>>> +
>>> + core->mac[offset] = value;
>>> +
>>> + /*
>>> + * Sync EITR to eitr_guest_value[] shadow array, stripping
>>> + * E1000_EITR_CNT_IGNR so guest register readback returns the
>>> + * correct value.
>>> + */
>>> + if (offset >= EITR0 && offset < EITR0 + IGB_INTR_NUM) {
>>> + core->eitr_guest_value[offset - EITR0] =
>>> + value & ~E1000_EITR_CNT_IGNR;
>>> + }
>>> + }
>>> +
>>> + /*
>>> + * MSI-X table/PBA is not saved - L1's VFIO reprograms it with
>>> + * destination-specific IRTE references after migration.
>>> + */
>>> +
>>> + uint32_t src_pool = E1000_RAH_POOL_1 << saved_vfn;
>>> + uint32_t dst_pool = E1000_RAH_POOL_1 << s->vfn;
>>> +
>>> + for (uint32_t i = 0; i < num_ra; i++) {
>>> + uint32_t offset = le32_to_cpu(blob->ra[i].offset);
>>> + uint32_t value = le32_to_cpu(blob->ra[i].value);
>>> +
>>> + /* RAH entries: swap pool ownership bits */
>>> + if (offset >= RA && offset < RA + 32 && (offset - RA) % 2 ==
>>> 1) {
>>> + value = (value & ~src_pool) | dst_pool;
>>> + }
>>> + if (offset >= RA2 && offset < RA2 + 16 && (offset - RA2) % 2
>>> == 1) {
>>> + value = (value & ~src_pool) | dst_pool;
>>> + }
>>> +
>>> + core->mac[offset] = value;
>>
>> Please validate RA offsets before writing core->mac.
>
> Drat. my bad. It should be done and it's easy :
>
> reject offsets not in [RA, RA+32) or [RA2, RA2+16)
That addresses the out-of-bounds write. Please also validate the
complete blob before modifying any state, including the TX context
count, so a rejected LOAD does not leave partially restored state.
>
>> A valid-sized LOAD blob with one RA entry and an out-of-range offset
>> produces an out-of-bounds 32-bit host write. LOAD reads this blob
>> directly from guest RAM. Please validate all offsets and saved_vfn
>> before modifying state.
>
>
> I think we should drop VF relocation support.
>
> if (saved_vfn != s->vfn) {
> return -IGB_MIG_ERR_BAD_VFN;
> }
>
>
> I was too optimistic in v2. This adds too much complexity and we haven't
> covered the simple case yet. Let's consider that VF number is part of
> the migration state to have more invariants and avoid collisions.
>
>>
>> This loop also changes the VF pool bit but preserves the source RA
>> slot. A valid VF0→VF1 migration onto a destination already using VF0
>> overwrites destination VF0’s primary MAC entry. Linux assigns the
>> primary entry as rar_entry_count - (vf + 1); receive filtering
>> consumes these slots directly in igb_receive_assign(). Thus the
>> unrelated destination VF loses unicast reception. Existing destination
>> entries owned by the migrating VF are also left behind.
>
> There should be no/few entries on the destination for the VF being
> migrated, only the fixed primary MAC entries for the VF. So, with the
> same VF number enforced, slot assignments should be deterministic
> and we can scan and clear any entries with this VF's pool bit at load
> time.
Requiring the same VF number sounds like a reasonable restriction for
the initial implementation. However, it does not guarantee that the
source and destination use the same RA slots. An additional source entry
can occupy a slot used by another VF on the destination, even when the
migrating VF keeps its number.
Scanning and clearing destination entries also needs to preserve other
pools: an RA entry can have multiple pool bits set. We should remove
only this VF's membership and avoid importing unrelated source pool
bits. Conflicting destination slots need to be handled or rejected.
VLVF has a similar issue. Scanning by pool bit identifies the VF's VLAN
memberships, but restoration needs to match VLAN IDs rather than copy
source slots, preserving other destination memberships. The
corresponding VFTA bits also need to be present.
For unsupported VFTA/MTA/UTA state, could we reject migration rather
than just warn? Otherwise migration succeeds while the resumed VF loses
reception. Restricting the supported configurations seems fine, provided
those restrictions are checked.
Regards,
Akihiko Odaki
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [RFC PATCH v2 4/9] igb: Add VF post-load fixups for live migration
2026-09-15 7:58 ` Cédric Le Goater
@ 2026-09-17 19:01 ` Akihiko Odaki
0 siblings, 0 replies; 23+ messages in thread
From: Akihiko Odaki @ 2026-09-17 19:01 UTC (permalink / raw)
To: Cédric Le Goater, qemu-devel
Cc: Sriram Yagnaraman, Jason Wang, Alex Williamson, Peter Xu
On 2026/09/15 16:58, Cédric Le Goater wrote:
> On 9/8/26 10:10, Akihiko Odaki wrote:
>> On 2026/09/03 4:20, Cédric Le Goater wrote:
>>> After restoring per-VF register state, propagate the VF's PVT shadow
>>> values back into the PF's aggregate EIMS/EIAC/EIAM registers and
>>> re-apply the VTIVAR interrupt vector routing to the shared IVAR0.
>>>
>>> AI-used-for: analysis, code (prototype)
>>> Signed-off-by: Cédric Le Goater <clg@redhat.com>
>>> ---
>>> hw/net/igb_core.h | 3 ++
>>> hw/net/igb_core.c | 66 ++++++++++++++++++++++++++++++++++++++++++
>>> hw/net/igb_migration.c | 7 +++++
>>> 3 files changed, 76 insertions(+)
>>>
>>> diff --git a/hw/net/igb_core.h b/hw/net/igb_core.h
>>> index 60724e2824ab..22e10e4e6d0b 100644
>>> --- a/hw/net/igb_core.h
>>> +++ b/hw/net/igb_core.h
>>> @@ -145,4 +145,7 @@ igb_start_recv(IGBCore *core);
>>> IGBCore *igb_pf_get_core(void *pf);
>>> +void igb_core_vf_propagate_irqs(IGBCore *core, uint16_t vfn);
>>> +void igb_core_vf_propagate_ivar(IGBCore *core, uint16_t vfn);
>>> +
>>> #endif
>>> diff --git a/hw/net/igb_core.c b/hw/net/igb_core.c
>>> index 2a4883907353..01745fe756d0 100644
>>> --- a/hw/net/igb_core.c
>>> +++ b/hw/net/igb_core.c
>>> @@ -4552,3 +4552,69 @@ igb_core_post_load(IGBCore *core)
>>> return 0;
>>> }
>>> +
>>> +/*
>>> + * Propagate VF interrupt state to PF aggregates after loading VF
>>> + * registers. The load path writes directly to mac[] bypassing the
>>> + * register handlers that OR VF bits into EIMS/EIAC/EIAM. Also clear
>>> + * stale VF bits in EICR that may have been set by packets arriving
>>> + * between PF vmstate restore and VF state load.
>>> + */
>>> +void igb_core_vf_propagate_irqs(IGBCore *core, uint16_t vfn)
>>> +{
>>> + uint32_t shift = 22 - vfn * IGBVF_MSIX_VEC_NUM;
>>> + uint32_t vf_mask = 0x7 << shift;
>>> + uint32_t pvt_idx;
>>> +
>>> + core->mac[EIMS] &= ~vf_mask;
>>
>> Restore the effective interrupt mask rather than the last EIMS write.
>>
>> This clears the destination VF mask and copies PVTEIMS, but that
>> shadow records only the last write: igb_set_vteims() assigns it before
>> applying set-bit semantics to aggregate EIMS. For example, source
>> writes EIMS=1, then EIMS=2, leave effective mask 3 and shadow 2;
>> migration restores 2, suppressing vector 0 until explicitly enabled
>> again. Conversely, EIMC updates aggregate EIMS without updating
>> PVTEIMS, so migration reenables explicitly masked vectors. Patch 7
>> repeats this fixup on unquiesce without repairing the bookkeeping.
>
> ok. I guess we can save and restore the VF's effective mask bits directly
> and drop the fixup then. Same for EIAC/EIAM. It should simplify the flow.
> How's that ?
Yes, saving and restoring the VF's effective bits in EIMS/EIAC/EIAM
would address my concern and remove the need for that shadow-based
fixup. Restore only this VF's bits, preserving the PF and other VFs
bits in those shared registers.
Please also ensure that the saved EIMS bits reflect the mask before
quiescing disables interrupts, and that unquiesce restores that saved
mask rather than reconstructing it from PVTEIMS.
Regards,
Akihiko Odaki
>
> Thanks
>
> C.
>
>
>>
>> Regards,
>> Akihiko Odaki
>>
>>> + pvt_idx = PVTEIMS0 + vfn * 0x40;
>>> + core->mac[EIMS] |= (core->mac[pvt_idx] & 0x7) << shift;
>>> +
>>> + core->mac[EIAC] &= ~vf_mask;
>>> + pvt_idx = PVTEIAC0 + vfn * 0x40;
>>> + core->mac[EIAC] |= (core->mac[pvt_idx] & 0x7) << shift;
>>> +
>>> + core->mac[EIAM] &= ~vf_mask;
>>> + pvt_idx = PVTEIAM0 + vfn * 0x40;
>>> + core->mac[EIAM] |= (core->mac[pvt_idx] & 0x7) << shift;
>>> +
>>> + core->mac[EICR] &= ~vf_mask;
>>> +}
>>> +
>>> +/*
>>> + * Re-apply VTIVAR -> IVAR0 interrupt routing. The L1 PF driver
>>> + * may have overwritten the shared IVAR0 entries with its own
>>> + * queue routing after L0 vmstate restore.
>>> + */
>>> +void igb_core_vf_propagate_ivar(IGBCore *core, uint16_t vfn)
>>> +{
>>> + uint32_t vtivar = core->mac[VTIVAR + vfn];
>>> + int n;
>>> + uint8_t ent;
>>> + uint32_t mask;
>>> +
>>> + n = igb_ivar_entry_rx(vfn);
>>> + mask = 0xffU << (8 * (n % 4));
>>> + if (vtivar & E1000_IVAR_VALID) {
>>> + ent = E1000_IVAR_VALID |
>>> + (24 - vfn * IGBVF_MSIX_VEC_NUM - (2 - (vtivar & 0x7)));
>>> + core->mac[IVAR0 + n / 4] =
>>> + (core->mac[IVAR0 + n / 4] & ~mask) |
>>> + ((uint32_t)ent << (8 * (n % 4)));
>>> + } else {
>>> + core->mac[IVAR0 + n / 4] &= ~mask;
>>> + }
>>> +
>>> + n = igb_ivar_entry_tx(vfn);
>>> + mask = 0xffU << (8 * (n % 4));
>>> + ent = vtivar >> 8;
>>> + if (ent & E1000_IVAR_VALID) {
>>> + ent = E1000_IVAR_VALID |
>>> + (24 - vfn * IGBVF_MSIX_VEC_NUM - (2 - (ent & 0x7)));
>>> + core->mac[IVAR0 + n / 4] =
>>> + (core->mac[IVAR0 + n / 4] & ~mask) |
>>> + ((uint32_t)ent << (8 * (n % 4)));
>>> + } else {
>>> + core->mac[IVAR0 + n / 4] &= ~mask;
>>> + }
>>> +}
>>> diff --git a/hw/net/igb_migration.c b/hw/net/igb_migration.c
>>> index c34035974620..af7257ccc4fd 100644
>>> --- a/hw/net/igb_migration.c
>>> +++ b/hw/net/igb_migration.c
>>> @@ -384,12 +384,19 @@ static int igb_core_vf_load_state(IgbVfState
>>> *s, const void *buf, size_t size)
>>> static int igbvf_mig_load(IgbVfState *s, const void *buf, size_t size)
>>> {
>>> int ret;
>>> + IGBCore *core = igbvf_get_core(s);
>>> ret = igb_core_vf_load_state(s, buf, size);
>>> if (ret < 0) {
>>> return ret;
>>> }
>>> + /*
>>> + * Post-load: sync VF interrupt and routing state to PF aggregates
>>> + */
>>> + igb_core_vf_propagate_irqs(core, s->vfn);
>>> + igb_core_vf_propagate_ivar(core, s->vfn);
>>> +
>>> return 0;
>>> }
>>
>
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [RFC PATCH v2 0/9] igb: Add experimental VF live migration support
2026-09-15 8:20 ` Cédric Le Goater
@ 2026-09-17 19:49 ` Akihiko Odaki
0 siblings, 0 replies; 23+ messages in thread
From: Akihiko Odaki @ 2026-09-17 19:49 UTC (permalink / raw)
To: Cédric Le Goater, qemu-devel
Cc: Sriram Yagnaraman, Jason Wang, Alex Williamson, Peter Xu
On 2026/09/15 17:20, Cédric Le Goater wrote:
> On 9/9/26 09:19, Akihiko Odaki wrote:
>> On 2026/09/03 4:20, Cédric Le Goater wrote:
>>> Hello,
>>>
>>> Live migration of VFIO-passthrough devices - SR-IOV VFs, vGPUs - is a
>>> growing requirement, but real hardware with migration support is
>>> scarce and hard to debug. An emulated device provides a fully
>>> controlled testbed for developing and validating the entire software
>>> stack - vfio-pci variant drivers, VFIO core migration v2 framework,
>>> QEMU, libvirt - and for tuning complex migration policies such as
>>> downtime convergence. It also serves as an educational reference for
>>> understanding VFIO migration end-to-end, from device state
>>> serialization to dirty page tracking.
>>>
>>> This series adds an experimental VF live migration interface to the
>>> emulated igb (82576) device. It enables a vfio-pci variant driver
>>> (igb-vfio-pci) to migrate VFs using the standard VFIO migration v2
>>> protocol with stop-copy and pre-copy support.
>>>
>>> The target scenario is nested virtualization:
>>>
>>> L0 QEMU (these patches)
>>> igb PF with x-vf-migration=on
>>> └── VFs with migration DVSEC
>>>
>>> L1 kernel
>>> igb-vfio-pci variant driver [1]
>>> translates VFIO migration v2 ioctls → DVSEC config writes
>>>
>>> L1 QEMU (stock, unmodified)
>>> vfio-pci device model, standard migration fd
>>>
>>> L2 guest
>>> standard igbvf driver, unaware of migration
>>>
>>> The L1 QEMU is completely unmodified -- it sees a standard VFIO
>>> migratable device and uses the normal migration fd path.
>>>
>>> * Design
>>>
>>> The migration interface is exposed through a DVSEC (Designated
>>> Vendor-Specific Extended Capability, PCIe cap id 0x23) at offset
>>> 0x160 in VF extended config space. The DVSEC uses a command doorbell
>>> model - all commands are synchronous via PCI config space writes.
>>>
>>> Device state is serialized as a versioned blob of per-VF register
>>> (offset, value) pairs covering control, interrupt, RX/TX queue,
>>> receive address (RA/RA2), etc. plus TX context descriptors and
>>> VFRE/VFTE enable bits. The buffer address is a guest physical
>>> address (GPA) written by the driver via virt_to_phys; the device
>>> accesses guest RAM directly through the system address space.
>>>
>>> Dirty page tracking is implemented with per-range bitmaps maintained
>>> in IGBCore. All VF DMA paths in igb_core.c (TX data, RX data,
>>> descriptor writeback) are instrumented to record touched pages. The
>>> variant driver registers tracked IOVA ranges and queries dirty bitmaps
>>> through a shared buffer. Buffer structures include len, flags, and
>>> reserved fields for future extensibility.
>>>
>>> * Caveats
>>>
>>> The x-vf-migration property is experimental (x- prefix, default off).
>>>
>>> The dirty bitmaps are maintained inside the device, which is not
>>> realistic for discrete NICs without on-chip DRAM.
>>>
>>> * Testing
>>>
>>> The target scenario is nested virtualization: L0 runs QEMU with an
>>> igb PF (x-vf-migration=on), L1 runs the igb-vfio-pci variant driver
>>> and an unmodified QEMU, and L2 runs a standard igbvf driver.
>>>
>>> Migration under iperf3 load works correctly: dirty page tracking
>>> converges (from ~2000 pages per PRE_COPY iteration down to ~280 at
>>> STOP_COPY), and STOP_COPY stays under 250ms.
>>>
>>> * Todo
>>>
>>> 1. Add migration blocker when x-vf-migration=on (no VMState yet) or
>>> add VMState support for L0 migration (dirty bitmaps, tracking
>>> engines, DVSEC registers, stats)
>>> 2. Add PRE_COPY state transfer to validate device INIT data (magic,
>>> version, etc.)
>>> 3. Add qtests for migration state machine transitions, dirty page
>>> tracking ?
>>>
>>> * Ideas
>>>
>>> 1. RX bandwidth throttle (x-mig-rx-limit, uint32, default 0)
>>>
>>> Return false from can_receive when the per-VF packet count in the
>>> current tracking interval exceeds the limit. Reduces DMA writes
>>> and dirty pages realistically.
>>>
>>> 2. Migration phase timing (GET_STATS extension)
>>>
>>> Add per-VF timestamps: precopy_start_ns, stopcopy_start_ns,
>>> precopy_duration_ns, stopcopy_duration_ns,
>>> state_transition_count. Expose via GET_STATS.
>>>
>>> 3. Hot page simulation (x-mig-hot-pages, uint32, default 0)
>>>
>>> Re-set the first N bitmap bits after each DIRTY_QUERY, simulating
>>> workloads with hot pages that prevent convergence.
>>>
>>> 4. Error injection (x-mig-inject-error, uint32, default 0)
>>>
>>> One-shot error code injection before command dispatch. A separate
>>> x-mig-inject-dma-fail (bool) for persistent DMA failure testing.
>>>
>>> * Credits
>>>
>>> Alex Williamson suggested the overall approach of a variant driver
>>> with the "x-vf-migration" device property to gate the feature. Thanks
>>> for the ever ongoing support and valuable discussions throughout these
>>> years.
>>>
>>> * AI disclaimer
>>>
>>> The lack of a migration-capable device has been a recurring pain point
>>> for VFIO development over the years, and we hope this proposal
>>> demonstrates the value of having one.
>>>
>>> Claude was used to analyze the IGB PF and VF internal state and
>>> identify the pain points of a working live migration of such devices.
>>> The generated code served as a starting point but *significant* time
>>> was then spent cleaning up, reworking, and shaping it into a clear,
>>> reviewable IGB model extension.
>>>
>>> As QEMU does not yet accept AI-assisted contributions, this series is
>>> submitted as an RFC.
>>
>> This largely repeats the concerns I raised previously [1]. Your
>> description of the *significant* time spent cleaning up and reworking
>> the generated code suggests that this series is intended for upstream
>> inclusion once the implementation direction is agreed.
>
> Yes. The IGB kernel vfip-pci variant driver is "ready", and so is the
> internal test framework for VFIO migration. The emulated IGB device
> is also becoming increasingly important as a reference device for
> the kernel VFIO self-tests. See:
>
> https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/
> commit?id=1583ef1175ba533e326dfeff6d8b9f3aaf761225
>
>> My concerns as a maintainer are therefore correctness and
>> maintainability.
>
> I am trying to avoid too much intrusion in the core igb model :
>
> hw/net/igb_common.h | 16 +
> hw/net/igb_core.h | 11 +
> hw/net/igb.c | 7 +
> hw/net/igb_core.c | 124 ++-
> hw/net/igbvf.c | 52 +-
>
> These are mostly code movement or additions. The runtime is hardly
> impacted apart from the DMA tracking, which adds one indirection.
>
> I have added myself as Maintainer for the migration part :
>
> +igb VF migration
> +M: Cédric Le Goater <clg@redhat.com>
> +S: Maintained
> +F: hw/net/igb_migration.*
>
> I can add you in the next respin. That would be a natural choice
> from my perspective, but I don't want to impose it on you.
My concern about maintainability is not primarily the line count in
igb_core.c, but the correctness of the architecture as a whole.
The migration layer inherently depends on deep assumptions about the
core device model. For example, my feedback on this version highlighted
dependencies on RA handling, VLAN filtering, and the interrupt state
machine. Moving migration logic into a separate file isolates the file
footprint, but it does not remove these underlying functional dependencies.
>> The missing migration blocker is a concrete example: it is a basic
>> correctness issue that should be straightforward to address, yet it
>> remains a TODO.
>
> yes. It's identified and easy to address. It can come at the end once
> all other issues have been discussed. What worries me the most is the
> control plane : interface with the Linux driver. That is the most
> important part on which these changes depend on getting it right :
>
> https://github.com/legoater/linux/commits/vfio/
That makes sense. If the immediate goal is to review the control plane
interface while save/load state remains incomplete, it would help to
explicitly state that scope in the cover letter, along with the current
limitations in implementation and testing.
>
>> My comments on individual patches also raise the same kinds of issues
>> I pointed out in the previous version.
>
> You're being tough on me ! I did all this fixing below ! :)
>
> * Changes since rfc-v1
>
> - Migration BAR replaced with DVSEC at offset 0x160 (no BAR needed)
> - Wire-format structs (IgbMigBlob, IgbMigRegPair, IgbMigTxCtx)
> replace raw pointer arithmetic and memcpy
> - RA entries separated from fixed regs, scanned by pool bit
> - VFN relocation support (offset remapping + RA pool bit swap)
> - GPA buffer (address_space_read/write) replaces PCI DMA through PF
> - State blob validation on load (magic, version, error codes)
> - NEED_WORDS macro and igb_vf_offset_valid removed
> - Error codes renumbered: removed BAD_VFN, added UNK_CMD (1-11)
>
> Reported by Akihiko Odaki:
>
> - propagate_irqs: clear VF bits before OR (EIMS/EIAC/EIAM)
> - propagate_ivar: clear IVAR entry when source VTIVAR is invalid
> - rearm_irqs: restore actual PVTEICR causes, not all three
> - Dirty bitmap allocation uses BITS_TO_LONGS (heap corruption fix)
> - Dirty query validates range before g_malloc0 (memory exhaustion)
> - Dirty bits cleared only after successful bitmap DMA write
> - Dirty query buffer uses struct offsets (layout mismatch fix)
> - Load path: register offsets validated against VF whitelist
> - VMBMEM (mailbox payload) documented as transient, not serialized
> - dma_writes counter: consistently uint64_t
> - Dirty range_size: consistently uint64_t (was truncated to 32 bits)
> - ERROR->STOP: quiesce VF (clear VFRE/VFTE) on transition
> - rearm_irqs: runs on re || te, not just re (TX-only VF fix)
> - Stats DMA-written atomically via GET_STATS (no split MMIO tear)
> - Bisectability: DVSEC + state machine introduced together
> - Removed NAPI reference and "===" comment decoration
I appreciate the progress, and I certainly don't mean to minimize the
effort put into fixing these issues.
However, many of these bugs stem from the same root cause: high-level
systemic requirements, such as boundary checks, VF renumbering,
interrupt state preservation, and OS-independence, are being addressed
piecemeal as individual bugs arise. I would like to see these
requirements enforced systematically across the entire implementation
rather than patched reactively. Replacing raw pointers with explicit
wire-format structures across the protocol is a great example of the
kind of systematic fix I mean.
>>
>> Many of the correctness problems found in v1 and v2 stem from the
>> choice to build on igb, which is more complex than virtio-net. That is
>> why I suggested using virtio-net as a simpler basis, unless there is a
>> problem that rules it out.
>
> virtio-net is a software abstraction designed for virtualization.
>
> Here’s my pitch for igb:
>
> Using an emulated igb device addresses a different problem: validating
> migration of hardware-like devices. igb exposes a real PCI device
> model, hardware-driver interface, DMA engines, queues, interrupts and
> SR-IOV VFs. It allows migration to be implemented at the same
> abstraction boundary at which VFIO migrates physical devices, rather
> than being reduced to serialization of a paravirtualized protocol
> state. Also,
>
> The igb VF uses the inbox igbvf kernel driver, which matches the
> hardware VFIO promise: transparent migration with unmodified guest
> drivers.
> The DVSEC extended capability is a discoverable and standard control
> plane with the device. This is a mechanism a hardware vendor could
> use.
> One cool aspect of an emulated device is that it allows the state to
> be inspected, the behavior is deterministic and faults can be injected
> at any boundary. That's the QEMU bonus.
I understand the motivation to test hardware-like device migration.
However, modern virtio-net devices are also implemented in physical
SmartNICs/DPUs and expose standard PCI interfaces, BARs, DMA engines,
queues, interrupts, and SR-IOV VFs.
Using virtio-net as a VFIO test device would not require guest driver
changes or bypass hardware-like semantics. Given its cleaner separation
of state, it remains a strong fit for these testing requirements without
carrying igb's extra complexity.
>
>> I would be willing to support allowing AI use in the project for work
>> like this. I would still need to see these correctness and
>> maintainability concerns addressed before I could support merging the
>> series.
>
> I agree. The save/load is still in progress and we might need some
> more intrusive changes in igb to support migration of all state.
> To start with, we could detect and warn.
Starting with a restricted set of supported configurations is a
reasonable way forward. However, for features whose state cannot yet be
restored accurately, we should reject migration rather than merely issue
a warning. Otherwise, a completed migration risks silently corrupting
guest-visible device behavior.
This is another advantage of virtio-net: feature negotiation offers a
clean, standard mechanism to disable unsupported features and define a
tightly constrained initial scope that is fully migratable.
Regards,
Akihiko Odaki
^ permalink raw reply [flat|nested] 23+ messages in thread
end of thread, other threads:[~2026-09-17 19:49 UTC | newest]
Thread overview: 23+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-02 19:20 [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Cédric Le Goater
2026-09-02 19:20 ` [RFC PATCH v2 1/9] igb: Add x-vf-migration property and DVSEC extended capability Cédric Le Goater
2026-09-03 19:57 ` Alex Williamson
2026-09-07 20:58 ` Cédric Le Goater
2026-09-02 19:20 ` [RFC PATCH v2 2/9] igb: Add migration state machine via extended config space Cédric Le Goater
2026-09-02 19:20 ` [RFC PATCH v2 3/9] igb: Add VF state serialization for live migration Cédric Le Goater
2026-09-08 8:01 ` Akihiko Odaki
2026-09-15 7:47 ` Cédric Le Goater
2026-09-17 18:59 ` Akihiko Odaki
2026-09-02 19:20 ` [RFC PATCH v2 4/9] igb: Add VF post-load fixups " Cédric Le Goater
2026-09-08 8:10 ` Akihiko Odaki
2026-09-15 7:58 ` Cédric Le Goater
2026-09-17 19:01 ` Akihiko Odaki
2026-09-02 19:20 ` [RFC PATCH v2 5/9] igb: Add dirty page tracking for IGBVF migration Cédric Le Goater
2026-09-02 19:20 ` [RFC PATCH v2 6/9] igb: Quiesce VFs on STOP and include PF enable state in migration Cédric Le Goater
2026-09-02 19:20 ` [RFC PATCH v2 7/9] igb: Fix post-migration RX ring deadlock Cédric Le Goater
2026-09-02 19:20 ` [RFC PATCH v2 8/9] igb: Add dirty page tracking statistics Cédric Le Goater
2026-09-02 19:20 ` [RFC PATCH v2 9/9] docs: Add igb VF migration testing setup guide Cédric Le Goater
2026-09-08 8:39 ` Akihiko Odaki
2026-09-15 7:59 ` Cédric Le Goater
2026-09-09 7:19 ` [RFC PATCH v2 0/9] igb: Add experimental VF live migration support Akihiko Odaki
2026-09-15 8:20 ` Cédric Le Goater
2026-09-17 19:49 ` Akihiko Odaki
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.