* [PATCH v2 00/10] QEMU: CXL Type-2 device passthrough via vfio-pci
@ 2026-09-16 18:44 mhonap
2026-09-16 18:44 ` [PATCH v2 01/10] linux-headers: Update vfio.h for CXL Type-2 passthrough mhonap
` (9 more replies)
0 siblings, 10 replies; 27+ messages in thread
From: mhonap @ 2026-09-16 18:44 UTC (permalink / raw)
To: alex, ankita, jic23, dave.jiang, alejandro.lucero-palau,
smadhavan, pierrick.bouvier, mst, imammedo, anisinha, pbonzini,
eric.auger, peter.maydell, richard.henderson, clg, cohuck
Cc: kjaju, vsethi, zhiw, mhonap, qemu-devel, qemu-arm, linux-cxl
From: Manish Honap <mhonap@nvidia.com>
This series adds QEMU support for passing a CXL Type-2 device (an
accelerator with host-managed device memory, e.g. a GPU) to a guest via
vfio-pci. The guest drives its own virtual endpoint HDM decoder and QEMU
maps the device memory at the guest physical address the guest commits,
while the host owns the host physical placement.
v1 [1] was reviewed by Junjie Cao and Cedric Le Goater. v2 addresses
Junjie's comments; see "Changes since v1" below.
Base: qemu master, commit 28e7aad522.
Kernel dependency
-----------------
Pairs with the kernel "vfio/cxl: CXL Type-2 device passthrough" series [2].
The kernel exposes the device memory as an HPA-backed VFIO region, traps
the HDM decoder block and runs its lock-on-commit FSM, handles the CXL
DVSEC (including a guest-triggered reset), and reports two things through
VFIO:
- A device flag (VFIO_DEVICE_FLAGS_CXL), and
- The component-register geometry (VFIO_REGION_INFO_CAP_CXL_COMP_REGS).
That series in turn builds on the cxl_reset series [3].
Status of the kernel side: v5 is on-list [2]; rebased onto cxl_reset [3]
and addressed the v4 review comments. This QEMU series depends on the VFIO
uAPI above (patch 1 imports it) and on the kernel servicing fd read/write
on the HDM memory region, which QEMU uses as the fallback when the region
is not mmap'd. Both are part of the same vfio-cxl series and are stable
across those revisions.
Sample supported topology
-------------------------
Guest disk, network, and system-RAM lines are omitted:
-machine virt,accel=kvm,gic-version=3,hmat=on,cxl=on,ras=on, \
highmem-mmio-size=4T
-object iommufd,id=iommufd0
-device pxb-cxl,bus_nr=12,bus=pcie.0,id=cxl.1
-device cxl-rp,port=1,bus=cxl.1,id=rport0.1,chassis=4, \
pref64-reserve=2G,mem-reserve=1G
-M cxl-fmw.0.targets.0=cxl.1,cxl-fmw.0.size=256G
-device arm-smmuv3,primary-bus=cxl.1,id=smmuv3.0,accel=on,ats=on, \
ril=on,ssidsize=8,oas=48
-device vfio-pci-nohotplug,host=<BDF>,bus=rport0.1,id=dev0, \
iommufd=iommufd0
-object acpi-generic-initiator,id=gi0,pci-dev=dev0,node=2
... (one acpi-generic-initiator per guest NUMA node the HDM memory backs)
Address model
-------------
The kernel fixes device memory to a host physical range before the guest
sees the device, and hardware presents a firmware-committed, locked
endpoint HDM decoder whose registers hold that host physical base.
QEMU never exposes the host physical base. It virtualizes the decoder base
registers in the trapped component-register read and returns the base of
the device's CFMWS window, a guest physical address, so the guest only
ever sees a GPA. QEMU maps the RAM-device region (backed by the fixed HPA)
at that CFMWS base. The kernel never learns the GPA and the guest never
learns the HPA.
Because the decoder is already committed at boot, no guest commit write
triggers the mapping. QEMU maps once the guest enables memory decoding
(the Command register Memory-Space bit) and re-checks on any decoder
control write, so the region enters the guest address space and the IOAS
while the device is live.
When the guest clears Memory-Space, QEMU withdraws the mapping, matching
the kernel's revoke of the backing PTEs on the same write, so a guest
access during the disabled interval cannot fault a zapped mapping and stop
the VM; the enable path re-installs it.
The pxb-cxl _DSM is here because an OS may treat PCI configuration as
reassignable, and a BAR move would break the CXL.mem mapping. It applies
only to the host bridge that carries the passed-through CXL device.
Reset
-----
There is no QEMU reset patch. A guest CXL reset is a DVSEC write that
lands in vfio config space and is handled by the host kernel, which runs
the CXL reset sequence and stamps the outcome into DVSEC STATUS2. The
kernel re-commits this firmware-fixed decoder across the reset, and the
guest reaches its memory through the mapping QEMU already installed.
Changes since v1
----------------
All from Junjie Cao's v1 review:
- The _DSM is emitted only when the machine asks OSPM to preserve the
firmware PCI configuration (preserve_config). x86 q35 passes false, so
its DSDT is unchanged and the bios-tables golden files stay valid.
- The endpoint count counts every present function, not one per slot, so
two functions cold-plugged at one slot no longer bind and map the same
window at the same base.
- The unrealize path (vfio_exitfn) drops the CXL mapping, so an ACPI
eject no longer leaves the vfio fd held through the region's owner
reference.
- The decoder-count read uses the shared cxl_decoder_count_dec() helper
with a floor of 1, instead of a local switch that stopped at 4 decoders.
- On commit, a guest decoder base that differs from the CFMWS base is
refused and logged rather than mapped at the window base regardless.
- The hotplug validation branch moved into the patch that registers the
machine-init-done notifier, so no intermediate commit can exit(1) on a
bad device_add.
AI assistance
-------------
Parts of this series were drafted with the help of an AI coding assistant:
the code was AI-prototyped, then reviewed, edited, and rewritten by me,
and the documentation was drafted the same way. The assistant was also
used to research the existing vfio and CXL APIs. I built the series and
functionally tested it against real CXL Type-2 hardware. I take
responsibility for the whole of every patch and certify it under the
DCO via Signed-off-by. The patches carry AI-used-for: trailers,
"code (prototype)" or "docs", per QEMU's code-provenance policy.
This series adds new passthrough code, which is outside the mechanical,
small-bug-fix, docs, and tests categories the policy covers without prior
maintainer agreement. I am flagging that here rather than assuming it is
in scope.
Validation
----------
- Every patch passes scripts/checkpatch.pl --codespell --strict (patch 1
carries the expected imported-from-Linux header warning).
- The series applies cleanly on the stated base.
- Built and functionally tested against a CXL Type-2 device: guest boot,
decoder commit and mapping, and guest-triggered CXL reset.
Pending items
-------------
- Multi-decoder, interleaved, and switched topologies are future work.
- Trapped CXL RAS registers are planned as a new VFIO region subtype that
the existing region-by-subtype detection already handles.
- The bios-tables test refresh is not needed now that the _DSM is gated on
preserve_config.
References
----------
[1] [PATCH 0/10] QEMU: CXL Type-2 device passthrough via vfio-pci
https://lore.kernel.org/qemu-devel/20260813130623.2499506-1-mhonap@nvidia.com/
[2] [PATCH v5 00/27] vfio/pci: Add CXL Type-2 device passthrough support
https://lore.kernel.org/linux-cxl/20260916183540.3813685-1-mhonap@nvidia.com/
[3] [PATCH v12 00/12] PCI/CXL: Add CXL reset support for Type 2 devices:
https://lore.kernel.org/linux-cxl/20260910070808.1444264-1-smadhavan@nvidia.com/
Manish Honap (10):
linux-headers: Update vfio.h for CXL Type-2 passthrough
hw/vfio/region: Add vfio_region_setup_with_ops()
hw/vfio/pci: Detect a CXL Type-2 device and read its geometry
hw/vfio/pci: Enforce the passthrough topology for a CXL device
hw/vfio/pci: Back the CXL memory with a RAM-device region
hw/vfio/pci: Bind a CXL device to its fixed memory window
hw/vfio/pci: Map the CXL memory on the guest decoder commit
docs/cxl: Document CXL Type-2 device passthrough
hw/arm/smmu-common: Allow pxb-cxl as an SMMUv3 primary bus
hw/pci-host: Emit a _DSM on pxb-cxl to preserve firmware PCI config
docs/system/devices/cxl.rst | 45 ++
hw/acpi/Kconfig | 1 +
hw/acpi/cxl-stub.c | 2 +-
hw/acpi/cxl.c | 12 +-
hw/acpi/pci.c | 40 ++
hw/arm/smmu-common.c | 19 +-
hw/cxl/cxl-host-stubs.c | 5 +
hw/i386/acpi-build.c | 2 +-
hw/pci-bridge/pci_expander_bridge_stubs.c | 6 +
hw/pci-host/gpex-acpi.c | 42 +-
hw/vfio/pci.c | 745 ++++++++++++++++++++++
hw/vfio/pci.h | 27 +
hw/vfio/region.c | 27 +-
hw/vfio/vfio-region.h | 3 +
include/hw/acpi/cxl.h | 2 +-
include/hw/acpi/pci.h | 1 +
linux-headers/linux/vfio.h | 24 +
17 files changed, 947 insertions(+), 56 deletions(-)
--
2.25.1
^ permalink raw reply [flat|nested] 27+ messages in thread
* [PATCH v2 01/10] linux-headers: Update vfio.h for CXL Type-2 passthrough
2026-09-16 18:44 [PATCH v2 00/10] QEMU: CXL Type-2 device passthrough via vfio-pci mhonap
@ 2026-09-16 18:44 ` mhonap
2026-09-16 18:44 ` [PATCH v2 02/10] hw/vfio/region: Add vfio_region_setup_with_ops() mhonap
` (8 subsequent siblings)
9 siblings, 0 replies; 27+ messages in thread
From: mhonap @ 2026-09-16 18:44 UTC (permalink / raw)
To: alex, ankita, jic23, dave.jiang, alejandro.lucero-palau,
smadhavan, pierrick.bouvier, mst, imammedo, anisinha, pbonzini,
eric.auger, peter.maydell, richard.henderson, clg, cohuck
Cc: kjaju, vsethi, zhiw, mhonap, qemu-devel, qemu-arm, linux-cxl
From: Manish Honap <mhonap@nvidia.com>
Sync the vfio.h additions from the kernel vfio-cxl series: the
VFIO_DEVICE_FLAGS_CXL device flag, the CXL Type-2 (0x1e98) HDM memory and
component-register sub-types under VFIO_REGION_TYPE_PCI_VENDOR_TYPE, and
the component-register geometry capability that reports the BAR and offset
of the trapped HDM decoder block.
The capability ID is provisional until the kernel series lands, at which
point this becomes a scripted linux-headers update.
AI-used-for: docs
Signed-off-by: Manish Honap <mhonap@nvidia.com>
---
| 24 ++++++++++++++++++++++++
1 file changed, 24 insertions(+)
--git a/linux-headers/linux/vfio.h b/linux-headers/linux/vfio.h
index c85dcbfe30..c12e55426d 100644
--- a/linux-headers/linux/vfio.h
+++ b/linux-headers/linux/vfio.h
@@ -215,6 +215,7 @@ struct vfio_device_info {
#define VFIO_DEVICE_FLAGS_FSL_MC (1 << 6) /* vfio-fsl-mc device */
#define VFIO_DEVICE_FLAGS_CAPS (1 << 7) /* Info supports caps */
#define VFIO_DEVICE_FLAGS_CDX (1 << 8) /* vfio-cdx device */
+#define VFIO_DEVICE_FLAGS_CXL (1 << 9) /* vfio-cxl device */
__u32 num_regions; /* Max region index + 1 */
__u32 num_irqs; /* Max IRQ index + 1 */
__u32 cap_offset; /* Offset within info struct of first cap */
@@ -370,6 +371,12 @@ struct vfio_region_info_cap_type {
*/
#define VFIO_REGION_SUBTYPE_IBM_NVLINK2_ATSD (1)
+/* CXL Type-2 device (0x1e98) sub-types for VFIO_REGION_TYPE_PCI_VENDOR_TYPE */
+/* CXL.mem HDM region of a Type-2 device, mmap-able */
+#define VFIO_REGION_SUBTYPE_CXL_MEM (1)
+/* CXL HDM decoder registers: read live, guest writes are absorbed */
+#define VFIO_REGION_SUBTYPE_CXL_COMP_REGS (2)
+
/* sub-types for VFIO_REGION_TYPE_GFX */
#define VFIO_REGION_SUBTYPE_GFX_EDID (1)
@@ -497,6 +504,23 @@ struct vfio_region_info_cap_nvlink2_lnkspd {
__u32 __pad;
};
+/*
+ * Geometry of a CXL Type-2 device's HDM decoder registers, so a VMM can place
+ * the trapped component register window where the guest expects it. The trapped
+ * region spans the whole HDM decoder block (every decoder), not just decoder 0:
+ * a VMM reads the decoder count and each decoder's committed base from the block
+ * itself. Additional trapped component capabilities, such as CXL RAS, are
+ * exposed as their own region subtypes rather than by extending this cap.
+ */
+#define VFIO_REGION_INFO_CAP_CXL_COMP_REGS 6
+
+struct vfio_region_info_cap_cxl_comp_regs {
+ struct vfio_info_cap_header header;
+ __u32 bar;
+ __u32 __resv;
+ __aligned_u64 offset;
+};
+
/**
* VFIO_DEVICE_GET_IRQ_INFO - _IOWR(VFIO_TYPE, VFIO_BASE + 9,
* struct vfio_irq_info)
--
2.25.1
^ permalink raw reply related [flat|nested] 27+ messages in thread
* [PATCH v2 02/10] hw/vfio/region: Add vfio_region_setup_with_ops()
2026-09-16 18:44 [PATCH v2 00/10] QEMU: CXL Type-2 device passthrough via vfio-pci mhonap
2026-09-16 18:44 ` [PATCH v2 01/10] linux-headers: Update vfio.h for CXL Type-2 passthrough mhonap
@ 2026-09-16 18:44 ` mhonap
2026-09-17 13:27 ` Cédric Le Goater
2026-09-16 18:44 ` [PATCH v2 03/10] hw/vfio/pci: Detect a CXL Type-2 device and read its geometry mhonap
` (7 subsequent siblings)
9 siblings, 1 reply; 27+ messages in thread
From: mhonap @ 2026-09-16 18:44 UTC (permalink / raw)
To: alex, ankita, jic23, dave.jiang, alejandro.lucero-palau,
smadhavan, pierrick.bouvier, mst, imammedo, anisinha, pbonzini,
eric.auger, peter.maydell, richard.henderson, clg, cohuck
Cc: kjaju, vsethi, zhiw, mhonap, qemu-devel, qemu-arm, linux-cxl
From: Manish Honap <mhonap@nvidia.com>
A CXL Type-2 device traps its HDM decoder block, so that region needs
custom MemoryRegionOps rather than the pass-through default. Split the
setup body out and let a caller supply the ops; NULL keeps the existing
behaviour.
AI-used-for: code (prototype)
Signed-off-by: Manish Honap <mhonap@nvidia.com>
---
hw/vfio/region.c | 27 ++++++++++++++++++++++++---
hw/vfio/vfio-region.h | 3 +++
2 files changed, 27 insertions(+), 3 deletions(-)
diff --git a/hw/vfio/region.c b/hw/vfio/region.c
index 54ad11a6c8..21b54978b5 100644
--- a/hw/vfio/region.c
+++ b/hw/vfio/region.c
@@ -228,8 +228,9 @@ static int vfio_setup_region_sparse_mmaps(VFIORegion *region,
return 0;
}
-int vfio_region_setup(Object *obj, VFIODevice *vbasedev, VFIORegion *region,
- int index, const char *name, Error **errp)
+static int vfio_region_do_setup(Object *obj, VFIODevice *vbasedev,
+ VFIORegion *region, int index, const char *name,
+ const MemoryRegionOps *ops, Error **errp)
{
struct vfio_region_info *info = NULL;
int ret;
@@ -249,7 +250,7 @@ int vfio_region_setup(Object *obj, VFIODevice *vbasedev, VFIORegion *region,
if (region->size) {
region->mem = g_new0(MemoryRegion, 1);
- memory_region_init_io(region->mem, obj, &vfio_region_ops,
+ memory_region_init_io(region->mem, obj, ops,
region, name, region->size);
if (!vbasedev->no_mmap &&
@@ -273,6 +274,26 @@ int vfio_region_setup(Object *obj, VFIODevice *vbasedev, VFIORegion *region,
return 0;
}
+int vfio_region_setup(Object *obj, VFIODevice *vbasedev, VFIORegion *region,
+ int index, const char *name, Error **errp)
+{
+ return vfio_region_do_setup(obj, vbasedev, region, index, name,
+ &vfio_region_ops, errp);
+}
+
+/*
+ * Like vfio_region_setup() but traps the region through @ops instead of the
+ * default pass-through, so a caller can intercept accesses (the CXL HDM
+ * decoder block). A NULL @ops keeps the default.
+ */
+int vfio_region_setup_with_ops(Object *obj, VFIODevice *vbasedev,
+ VFIORegion *region, int index, const char *name,
+ const MemoryRegionOps *ops, Error **errp)
+{
+ return vfio_region_do_setup(obj, vbasedev, region, index, name,
+ ops ? ops : &vfio_region_ops, errp);
+}
+
static void vfio_subregion_unmap(VFIORegion *region, int index)
{
trace_vfio_region_unmap(memory_region_name(®ion->mmaps[index].mem),
diff --git a/hw/vfio/vfio-region.h b/hw/vfio/vfio-region.h
index 58b236f113..8d5699013a 100644
--- a/hw/vfio/vfio-region.h
+++ b/hw/vfio/vfio-region.h
@@ -39,6 +39,9 @@ uint64_t vfio_region_read(void *opaque,
hwaddr addr, unsigned size);
int vfio_region_setup(Object *obj, VFIODevice *vbasedev, VFIORegion *region,
int index, const char *name, Error **errp);
+int vfio_region_setup_with_ops(Object *obj, VFIODevice *vbasedev,
+ VFIORegion *region, int index, const char *name,
+ const MemoryRegionOps *ops, Error **errp);
int vfio_region_mmap(VFIORegion *region);
void vfio_region_mmaps_set_enabled(VFIORegion *region, bool enabled);
void vfio_region_exit(VFIORegion *region);
--
2.25.1
^ permalink raw reply related [flat|nested] 27+ messages in thread
* [PATCH v2 03/10] hw/vfio/pci: Detect a CXL Type-2 device and read its geometry
2026-09-16 18:44 [PATCH v2 00/10] QEMU: CXL Type-2 device passthrough via vfio-pci mhonap
2026-09-16 18:44 ` [PATCH v2 01/10] linux-headers: Update vfio.h for CXL Type-2 passthrough mhonap
2026-09-16 18:44 ` [PATCH v2 02/10] hw/vfio/region: Add vfio_region_setup_with_ops() mhonap
@ 2026-09-16 18:44 ` mhonap
2026-09-17 16:04 ` Cédric Le Goater
2026-09-16 18:44 ` [PATCH v2 04/10] hw/vfio/pci: Enforce the passthrough topology for a CXL device mhonap
` (6 subsequent siblings)
9 siblings, 1 reply; 27+ messages in thread
From: mhonap @ 2026-09-16 18:44 UTC (permalink / raw)
To: alex, ankita, jic23, dave.jiang, alejandro.lucero-palau,
smadhavan, pierrick.bouvier, mst, imammedo, anisinha, pbonzini,
eric.auger, peter.maydell, richard.henderson, clg, cohuck
Cc: kjaju, vsethi, zhiw, mhonap, qemu-devel, qemu-arm, linux-cxl
From: Manish Honap <mhonap@nvidia.com>
The kernel marks a passthroughed CXL Type-2 device with a device flag and
exposes two regions:
- HDM memory (host physical)
- the trapped HDM decoder register block.
Read the flag, locate both regions by type, and take the component BAR
and block offset from the geometry capability. Realize only records this;
later patches build the guest mapping on top.
AI-used-for: code (prototype)
Signed-off-by: Manish Honap <mhonap@nvidia.com>
---
hw/vfio/pci.c | 72 +++++++++++++++++++++++++++++++++++++++++++++++++++
hw/vfio/pci.h | 15 +++++++++++
2 files changed, 87 insertions(+)
diff --git a/hw/vfio/pci.c b/hw/vfio/pci.c
index 428ab2f069..2f84af5cf8 100644
--- a/hw/vfio/pci.c
+++ b/hw/vfio/pci.c
@@ -3570,6 +3570,64 @@ bool vfio_pci_interrupt_setup(VFIOPCIDevice *vdev, Error **errp)
return true;
}
+/*
+ * Learn the CXL geometry the kernel reports: the HPA-backed HDM memory region
+ * and the trapped HDM decoder block (which BAR carries it and at what offset).
+ * A non-CXL device leaves cxl.enabled false and takes no CXL paths.
+ */
+static bool vfio_cxl_setup(VFIOPCIDevice *vdev, Error **errp)
+{
+ VFIODevice *vbasedev = &vdev->vbasedev;
+ VFIOCXL *cxl = &vdev->cxl;
+ struct vfio_region_info *mem_info = NULL, *comp_info = NULL;
+ struct vfio_region_info_cap_cxl_comp_regs *cap;
+ struct vfio_info_cap_header *hdr;
+
+ if (!(vbasedev->flags & VFIO_DEVICE_FLAGS_CXL)) {
+ return true;
+ }
+
+ if (vfio_device_get_region_info_type(vbasedev,
+ VFIO_REGION_TYPE_PCI_VENDOR_TYPE |
+ CXL_VENDOR_ID,
+ VFIO_REGION_SUBTYPE_CXL_MEM,
+ &mem_info)) {
+ error_setg(errp, "vfio-cxl: %s: CXL memory region not found",
+ vbasedev->name);
+ return false;
+ }
+
+ if (vfio_device_get_region_info_type(vbasedev,
+ VFIO_REGION_TYPE_PCI_VENDOR_TYPE |
+ CXL_VENDOR_ID,
+ VFIO_REGION_SUBTYPE_CXL_COMP_REGS,
+ &comp_info)) {
+ error_setg(errp,
+ "vfio-cxl: %s: CXL component-register region not found",
+ vbasedev->name);
+ return false;
+ }
+
+ hdr = vfio_get_region_info_cap(comp_info,
+ VFIO_REGION_INFO_CAP_CXL_COMP_REGS);
+ if (!hdr) {
+ error_setg(errp,
+ "vfio-cxl: %s: component-register geometry not reported",
+ vbasedev->name);
+ return false;
+ }
+ cap = container_of(hdr, struct vfio_region_info_cap_cxl_comp_regs, header);
+
+ cxl->mem_region_index = mem_info->index;
+ cxl->comp_regs_region_index = comp_info->index;
+ cxl->dpa_size = mem_info->size;
+ cxl->comp_bar = cap->bar;
+ cxl->hdm_offset = cap->offset;
+ cxl->enabled = true;
+
+ return true;
+}
+
static void vfio_pci_realize(PCIDevice *pdev, Error **errp)
{
ERRP_GUARD();
@@ -3700,6 +3758,20 @@ static void vfio_pci_realize(PCIDevice *pdev, Error **errp)
}
}
+ if (!vfio_cxl_setup(vdev, errp)) {
+ /*
+ * vfio_migration_realize() above installed a migration blocker in auto
+ * mode (generic vfio-pci exposes no migration ops). out_deregister does
+ * not remove it, so a rejected CXL setup would leave VM migration
+ * blocked until QEMU restarts. Drop it here, under the same
+ * failover-pair condition used to install it.
+ */
+ if (!pdev->failover_pair_id) {
+ vfio_migration_exit(vbasedev);
+ }
+ goto out_deregister;
+ }
+
vfio_pci_register_err_notifier(vdev);
vfio_pci_register_req_notifier(vdev);
vfio_setup_resetfn_quirk(vdev);
diff --git a/hw/vfio/pci.h b/hw/vfio/pci.h
index c9ab949870..7fdd695704 100644
--- a/hw/vfio/pci.h
+++ b/hw/vfio/pci.h
@@ -122,10 +122,25 @@ typedef struct VFIOMSIXInfo {
OBJECT_DECLARE_SIMPLE_TYPE(VFIOPCIDevice, VFIO_PCI_DEVICE)
+/*
+ * State for a CXL Type-2 device. The kernel owns the host physical placement
+ * of the device memory; QEMU only maps it at the guest physical address the
+ * guest commits into its endpoint HDM decoder.
+ */
+typedef struct VFIOCXL {
+ bool enabled;
+ uint32_t mem_region_index; /* HPA-backed HDM memory VFIO region */
+ uint32_t comp_regs_region_index; /* trapped HDM decoder register block */
+ uint32_t comp_bar; /* component BAR carrying that block */
+ uint64_t hdm_offset; /* block offset within the component BAR */
+ uint64_t dpa_size; /* size of the HDM memory region */
+} VFIOCXL;
+
struct VFIOPCIDevice {
PCIDevice parent_obj;
VFIODevice vbasedev;
+ VFIOCXL cxl;
VFIOINTx intx;
unsigned int config_size;
uint8_t *emulated_config_bits; /* QEMU emulated bits, little-endian */
--
2.25.1
^ permalink raw reply related [flat|nested] 27+ messages in thread
* [PATCH v2 04/10] hw/vfio/pci: Enforce the passthrough topology for a CXL device
2026-09-16 18:44 [PATCH v2 00/10] QEMU: CXL Type-2 device passthrough via vfio-pci mhonap
` (2 preceding siblings ...)
2026-09-16 18:44 ` [PATCH v2 03/10] hw/vfio/pci: Detect a CXL Type-2 device and read its geometry mhonap
@ 2026-09-16 18:44 ` mhonap
2026-09-17 13:20 ` Cédric Le Goater
2026-09-16 18:44 ` [PATCH v2 05/10] hw/vfio/pci: Back the CXL memory with a RAM-device region mhonap
` (5 subsequent siblings)
9 siblings, 1 reply; 27+ messages in thread
From: mhonap @ 2026-09-16 18:44 UTC (permalink / raw)
To: alex, ankita, jic23, dave.jiang, alejandro.lucero-palau,
smadhavan, pierrick.bouvier, mst, imammedo, anisinha, pbonzini,
eric.auger, peter.maydell, richard.henderson, clg, cohuck
Cc: kjaju, vsethi, zhiw, mhonap, qemu-devel, qemu-arm, linux-cxl
From: Manish Honap <mhonap@nvidia.com>
Only the guest's endpoint HDM decoder is programmed by the guest; the
decoders above it are covered only while the pxb-cxl host bridge stays in
passthrough mode. Reject a switch in the path, or a host bridge taken out
of passthrough, at realize before any mapping exists.
AI-used-for: code (prototype)
Signed-off-by: Manish Honap <mhonap@nvidia.com>
---
hw/vfio/pci.c | 73 +++++++++++++++++++++++++++++++++++++++++++++++++++
1 file changed, 73 insertions(+)
diff --git a/hw/vfio/pci.c b/hw/vfio/pci.c
index 2f84af5cf8..10b13f200d 100644
--- a/hw/vfio/pci.c
+++ b/hw/vfio/pci.c
@@ -27,6 +27,7 @@
#include "hw/pci/msi.h"
#include "hw/pci/msix.h"
#include "hw/pci/pci_bridge.h"
+#include "hw/cxl/cxl.h"
#include "hw/core/qdev-properties.h"
#include "hw/core/qdev-properties-system.h"
#include "hw/vfio/vfio-cpr.h"
@@ -3570,6 +3571,73 @@ bool vfio_pci_interrupt_setup(VFIOPCIDevice *vdev, Error **errp)
return true;
}
+/*
+ * Walk the endpoint's ancestor bridges up to its pxb-cxl host bridge. A CXL
+ * switch in the path is rejected: the guest would then have to program switch
+ * decoders, which the passthrough model does not yet cover. This is the single
+ * place the supported-topology assumption lives, so switch support later
+ * relaxes it here.
+ */
+static PXBCXLDev *vfio_cxl_find_pxb(VFIOPCIDevice *vdev, Error **errp)
+{
+ PCIBus *bus = pci_get_bus(&vdev->parent_obj);
+
+ if (!object_dynamic_cast(OBJECT(bus), TYPE_CXL_BUS)) {
+ error_setg(errp,
+ "vfio-cxl: %s: endpoint is not on a CXL root port (cxl-rp)",
+ vdev->vbasedev.name);
+ return NULL;
+ }
+
+ for (; bus; ) {
+ PCIDevice *bridge = bus->parent_dev;
+
+ if (!bridge) {
+ break;
+ }
+ if (object_dynamic_cast(OBJECT(bridge), TYPE_CXL_USP) ||
+ object_dynamic_cast(OBJECT(bridge), TYPE_CXL_DSP)) {
+ error_setg(errp,
+ "vfio-cxl: %s: switch-attached topology not supported",
+ vdev->vbasedev.name);
+ return NULL;
+ }
+ if (object_dynamic_cast(OBJECT(bridge), TYPE_PXB_CXL_DEV)) {
+ return PXB_CXL_DEV(bridge);
+ }
+ if (pci_bus_is_root(bus)) {
+ break;
+ }
+ bus = pci_get_bus(bridge);
+ }
+
+ error_setg(errp, "vfio-cxl: %s: not attached below a pxb-cxl host bridge",
+ vdev->vbasedev.name);
+ return NULL;
+}
+
+/*
+ * Only guest's endpoint decoder is programmed, so the host bridge must stay in
+ * HDM passthrough mode. Reject anything else at realize, before a mapping is
+ * built. cxl_get_hb_passthrough() reflects reset state and is not
+ * settled yet here, so gate on the static hdm_for_passthrough property.
+ */
+static bool vfio_cxl_check_topology(VFIOPCIDevice *vdev, Error **errp)
+{
+ PXBCXLDev *pxb = vfio_cxl_find_pxb(vdev, errp);
+
+ if (!pxb) {
+ return false;
+ }
+ if (pxb->hdm_for_passthrough) {
+ error_setg(errp,
+ "vfio-cxl: %s: the pxb-cxl must be in HDM passthrough mode "
+ "(do not set hdm_for_passthrough)", vdev->vbasedev.name);
+ return false;
+ }
+ return true;
+}
+
/*
* Learn the CXL geometry the kernel reports: the HPA-backed HDM memory region
* and the trapped HDM decoder block (which BAR carries it and at what offset).
@@ -3623,6 +3691,11 @@ static bool vfio_cxl_setup(VFIOPCIDevice *vdev, Error **errp)
cxl->dpa_size = mem_info->size;
cxl->comp_bar = cap->bar;
cxl->hdm_offset = cap->offset;
+
+ if (!vfio_cxl_check_topology(vdev, errp)) {
+ return false;
+ }
+
cxl->enabled = true;
return true;
--
2.25.1
^ permalink raw reply related [flat|nested] 27+ messages in thread
* [PATCH v2 05/10] hw/vfio/pci: Back the CXL memory with a RAM-device region
2026-09-16 18:44 [PATCH v2 00/10] QEMU: CXL Type-2 device passthrough via vfio-pci mhonap
` (3 preceding siblings ...)
2026-09-16 18:44 ` [PATCH v2 04/10] hw/vfio/pci: Enforce the passthrough topology for a CXL device mhonap
@ 2026-09-16 18:44 ` mhonap
2026-09-17 13:56 ` Cédric Le Goater
2026-09-16 18:44 ` [PATCH v2 06/10] hw/vfio/pci: Bind a CXL device to its fixed memory window mhonap
` (4 subsequent siblings)
9 siblings, 1 reply; 27+ messages in thread
From: mhonap @ 2026-09-16 18:44 UTC (permalink / raw)
To: alex, ankita, jic23, dave.jiang, alejandro.lucero-palau,
smadhavan, pierrick.bouvier, mst, imammedo, anisinha, pbonzini,
eric.auger, peter.maydell, richard.henderson, clg, cohuck
Cc: kjaju, vsethi, zhiw, mhonap, qemu-devel, qemu-arm, linux-cxl
From: Manish Honap <mhonap@nvidia.com>
mmap the kernel's HDM memory region so its host physical pages back a
RAM-device MemoryRegion. It stays out of the guest address space until
the guest commits its decoder and QEMU learns the GPA to place it at.
AI-used-for: code (prototype)
Signed-off-by: Manish Honap <mhonap@nvidia.com>
---
hw/vfio/pci.c | 34 ++++++++++++++++++++++++++++++++++
hw/vfio/pci.h | 1 +
2 files changed, 35 insertions(+)
diff --git a/hw/vfio/pci.c b/hw/vfio/pci.c
index 10b13f200d..4716266595 100644
--- a/hw/vfio/pci.c
+++ b/hw/vfio/pci.c
@@ -3221,8 +3221,11 @@ bool vfio_pci_populate_device(VFIOPCIDevice *vdev, Error **errp)
return true;
}
+static void vfio_cxl_teardown(VFIOPCIDevice *vdev);
+
void vfio_pci_put_device(VFIOPCIDevice *vdev)
{
+ vfio_cxl_teardown(vdev);
vfio_display_finalize(vdev);
vfio_bars_finalize(vdev);
vfio_cpr_pci_unregister_device(vdev);
@@ -3696,11 +3699,42 @@ static bool vfio_cxl_setup(VFIOPCIDevice *vdev, Error **errp)
return false;
}
+ /*
+ * The HDM memory is host physical. Set up the region, which installs the
+ * fd read/write path, and mmap it for direct guest access; it is added to
+ * the guest address space only once the guest commits its endpoint decoder.
+ */
+ if (vfio_region_setup(OBJECT(vdev), vbasedev, &cxl->mem_region,
+ cxl->mem_region_index, "cxl-mem", errp)) {
+ return false;
+ }
+ if (vfio_region_mmap(&cxl->mem_region)) {
+ /*
+ * Without mmap the region falls back to the kernel's fd read/write
+ * path, which works but traps every access. Warn rather than fail.
+ */
+ warn_report("vfio-cxl: %s: failed to mmap the HDM memory region; "
+ "performance may be slow", vbasedev->name);
+ }
+
cxl->enabled = true;
return true;
}
+static void vfio_cxl_teardown(VFIOPCIDevice *vdev)
+{
+ VFIOCXL *cxl = &vdev->cxl;
+
+ if (!cxl->enabled) {
+ return;
+ }
+ if (cxl->mem_region.mem) {
+ vfio_region_exit(&cxl->mem_region);
+ vfio_region_finalize(&cxl->mem_region);
+ }
+}
+
static void vfio_pci_realize(PCIDevice *pdev, Error **errp)
{
ERRP_GUARD();
diff --git a/hw/vfio/pci.h b/hw/vfio/pci.h
index 7fdd695704..06d15e807d 100644
--- a/hw/vfio/pci.h
+++ b/hw/vfio/pci.h
@@ -134,6 +134,7 @@ typedef struct VFIOCXL {
uint32_t comp_bar; /* component BAR carrying that block */
uint64_t hdm_offset; /* block offset within the component BAR */
uint64_t dpa_size; /* size of the HDM memory region */
+ VFIORegion mem_region; /* HDM memory, mapped at committed GPA */
} VFIOCXL;
struct VFIOPCIDevice {
--
2.25.1
^ permalink raw reply related [flat|nested] 27+ messages in thread
* [PATCH v2 06/10] hw/vfio/pci: Bind a CXL device to its fixed memory window
2026-09-16 18:44 [PATCH v2 00/10] QEMU: CXL Type-2 device passthrough via vfio-pci mhonap
` (4 preceding siblings ...)
2026-09-16 18:44 ` [PATCH v2 05/10] hw/vfio/pci: Back the CXL memory with a RAM-device region mhonap
@ 2026-09-16 18:44 ` mhonap
2026-09-17 16:00 ` Cédric Le Goater
2026-09-16 18:44 ` [PATCH v2 07/10] hw/vfio/pci: Map the CXL memory on the guest decoder commit mhonap
` (3 subsequent siblings)
9 siblings, 1 reply; 27+ messages in thread
From: mhonap @ 2026-09-16 18:44 UTC (permalink / raw)
To: alex, ankita, jic23, dave.jiang, alejandro.lucero-palau,
smadhavan, pierrick.bouvier, mst, imammedo, anisinha, pbonzini,
eric.auger, peter.maydell, richard.henderson, clg, cohuck
Cc: kjaju, vsethi, zhiw, mhonap, qemu-devel, qemu-arm, linux-cxl
From: Manish Honap <mhonap@nvidia.com>
The guest programs a GPA inside a CXL fixed memory window and QEMU maps
the HDM memory there, so the device must sit under exactly one
single-target CFMWS. Reject any ambiguous, interleaved, unsized, or
missing window rather than guess. Match windows by target name, since the
resolved target pointers are only filled in by a machine-done notifier
that may run after this one.
Require exactly one endpoint function below the root port: two functions
in the same slot would each match that single-target CFMWS and alias their
HDM memory at one base, so count every present function rather than one per
slot. The CFMWS windows are placed at machine-init-done, so a cold-plugged
device can only be validated from that notifier, where a bad topology is
fatal to startup. Validate a hotplugged device inline from realize instead,
reporting through errp so a bad device_add fails cleanly rather than
aborting the running VM, and arm the notifier only for a cold-plugged
device.
The endpoint's HDM size comes from the kernel region. cxl_fmws_set_memmap()
places the window in the guest PA map before this device is realized, so
the window cannot yet be sized from the device; check the configured window
against the device geometry instead. Reject a window too small to hold the
endpoint and name the size to set, and warn on an oversized window, since
the padding becomes guest CEDT that migration has to preserve.
AI-used-for: code (prototype)
Signed-off-by: Manish Honap <mhonap@nvidia.com>
---
hw/pci-bridge/pci_expander_bridge_stubs.c | 6 +
hw/vfio/pci.c | 192 ++++++++++++++++++++++
hw/vfio/pci.h | 3 +
3 files changed, 201 insertions(+)
diff --git a/hw/pci-bridge/pci_expander_bridge_stubs.c b/hw/pci-bridge/pci_expander_bridge_stubs.c
index b35180311f..c44ad7fab6 100644
--- a/hw/pci-bridge/pci_expander_bridge_stubs.c
+++ b/hw/pci-bridge/pci_expander_bridge_stubs.c
@@ -10,5 +10,11 @@
#include "hw/pci/pci_bus.h"
#include "hw/pci-bridge/pci_expander_bridge.h"
#include "hw/cxl/cxl.h"
+#include "hw/cxl/cxl_component.h"
void pxb_cxl_hook_up_registers(CXLState *state, PCIBus *bus, Error **errp) {};
+
+bool cxl_get_hb_passthrough(PCIHostState *hb)
+{
+ return false;
+}
diff --git a/hw/vfio/pci.c b/hw/vfio/pci.c
index 4716266595..670e0d1da4 100644
--- a/hw/vfio/pci.c
+++ b/hw/vfio/pci.c
@@ -27,7 +27,12 @@
#include "hw/pci/msi.h"
#include "hw/pci/msix.h"
#include "hw/pci/pci_bridge.h"
+#include "hw/pci/pci_host.h"
+#include "hw/pci/pcie_port.h"
#include "hw/cxl/cxl.h"
+#include "hw/cxl/cxl_host.h"
+#include "hw/cxl/cxl_component.h"
+#include "system/system.h"
#include "hw/core/qdev-properties.h"
#include "hw/core/qdev-properties-system.h"
#include "hw/vfio/vfio-cpr.h"
@@ -3641,6 +3646,171 @@ static bool vfio_cxl_check_topology(VFIOPCIDevice *vdev, Error **errp)
return true;
}
+/*
+ * Count the CXL fixed memory windows that target this device's pxb-cxl and
+ * report the matched window's base, size and target count. Match on the target
+ * names: cxl_fmws_link_targets() resolves target_hbs[] only at machine_done,
+ * which may run after this, so the resolved pointers can still be NULL here.
+ */
+static int vfio_cxl_match_fmws(PXBCXLDev *pxb, hwaddr *base, uint64_t *size,
+ int *nwindows, int *ntargets)
+{
+ GSList *list = cxl_fmws_get_all_sorted();
+ GSList *iter;
+ int matches = 0;
+
+ *nwindows = g_slist_length(list);
+ *base = 0;
+ *size = 0;
+ *ntargets = 0;
+
+ for (iter = list; iter; iter = iter->next) {
+ CXLFixedWindow *fw = CXL_FMW(iter->data);
+ int i;
+
+ for (i = 0; i < fw->num_targets; i++) {
+ bool ambiguous = false;
+ Object *t = object_resolve_path_type(fw->targets[i],
+ TYPE_PXB_CXL_DEV, &ambiguous);
+
+ if (t && !ambiguous && PXB_CXL_DEV(t) == pxb) {
+ *base = fw->base;
+ *size = fw->size;
+ *ntargets = fw->num_targets;
+ matches++;
+ break;
+ }
+ }
+ }
+ g_slist_free(list);
+ return matches;
+}
+
+/*
+ * Fix the bounds of the device's memory window once the topology has settled.
+ * Exactly one single-target CFMWS, with an assigned base and room for the HDM
+ * memory, is required; anything else is a misconfiguration the guest cannot
+ * recover from, so reject it rather than guess a window.
+ */
+static bool vfio_cxl_do_bind_fmws(VFIOPCIDevice *vdev, Error **errp)
+{
+ VFIOCXL *cxl = &vdev->cxl;
+ const char *name = vdev->vbasedev.name;
+ int nwindows = 0, ntargets = 0, matches;
+ hwaddr base = 0;
+ uint64_t size = 0, need;
+ PXBCXLDev *pxb;
+
+ pxb = vfio_cxl_find_pxb(vdev, errp);
+ if (!pxb) {
+ return false;
+ }
+
+ if (pcie_count_ds_ports(PCI_HOST_BRIDGE(pxb->cxl_host_bridge)->bus) != 1) {
+ error_setg(errp,
+ "vfio-cxl: %s: pxb-cxl not in HDM passthrough mode "
+ "(use a single cxl-rp)", name);
+ return false;
+ }
+
+ /*
+ * Count every present function, not one per slot, and require exactly one
+ * endpoint below the root port.
+ */
+ {
+ PCIBus *ep_bus = pci_get_bus(&vdev->parent_obj);
+ int slot, fn, nendpoints = 0;
+
+ for (slot = 0; slot < PCI_SLOT_MAX; slot++) {
+ for (fn = 0; fn < PCI_FUNC_MAX; fn++) {
+ if (ep_bus->devices[PCI_DEVFN(slot, fn)]) {
+ nendpoints++;
+ }
+ }
+ }
+ if (nendpoints != 1) {
+ error_setg(errp,
+ "vfio-cxl: %s: %d devices below the cxl-rp; a passed "
+ "through CXL endpoint must be alone below its root port",
+ name, nendpoints);
+ return false;
+ }
+ }
+
+ matches = vfio_cxl_match_fmws(pxb, &base, &size, &nwindows, &ntargets);
+ if (matches == 1 && ntargets == 1) {
+ if (!base) {
+ error_setg(errp,
+ "vfio-cxl: %s: matched CFMWS has no base; reduce its "
+ "size or grow the guest PA space", name);
+ return false;
+ }
+ /*
+ * QEMU reads the endpoint's HDM size from the kernel region, so the
+ * window size is really device information, not a value to make the
+ * caller guess. Sizing the CFMWS from the device the way firmware does
+ * is not possible here: the window is placed in the guest PA map by
+ * cxl_fmws_set_memmap() before this device is realized, so its size is
+ * fixed before dpa_size is known. Until a core change can size the
+ * window from the device, take the device geometry as the source of
+ * truth: reject a window too small to hold the endpoint, and warn when
+ * it is larger than needed, since the padding becomes guest CEDT that
+ * migration has to preserve.
+ */
+ need = ROUND_UP(cxl->dpa_size, 256 * MiB);
+ if (size < need) {
+ error_setg(errp,
+ "vfio-cxl: %s: CFMWS size 0x%" PRIx64 " cannot hold the "
+ "endpoint; set the cxl-fmw size to 0x%" PRIx64,
+ name, size, need);
+ return false;
+ }
+ if (size > need) {
+ warn_report("vfio-cxl: %s: CFMWS size 0x%" PRIx64 " exceeds the "
+ "endpoint's 0x%" PRIx64 "; set the cxl-fmw size to 0x%"
+ PRIx64 " to keep the guest memory map stable across "
+ "migration", name, size, cxl->dpa_size, need);
+ }
+ cxl->fmws_base = base;
+ cxl->fmws_size = size;
+ return true;
+ }
+
+ if (matches == 1) {
+ error_setg(errp,
+ "vfio-cxl: %s: its CFMWS interleaves %d targets; use a "
+ "single-target window", name, ntargets);
+ } else if (matches > 1) {
+ error_setg(errp,
+ "vfio-cxl: %s: pxb-cxl is targeted by %d CFMWS; use one",
+ name, matches);
+ } else {
+ error_setg(errp,
+ "vfio-cxl: %s: no CFMWS targets this device (%d present)",
+ name, nwindows);
+ }
+ return false;
+}
+
+/*
+ * Cold-plug path: the CFMWS windows are placed at machine init done, so the
+ * binding can only be validated from this notifier. The VM has not run yet, so
+ * a configuration error is fatal to startup. A hotplugged device is validated
+ * in realize instead (see vfio_cxl_setup), where the failure fails device_add
+ * without taking down the running VM.
+ */
+static void vfio_cxl_bind_fmws(Notifier *n, void *data)
+{
+ VFIOCXL *cxl = container_of(n, VFIOCXL, machine_done);
+ VFIOPCIDevice *vdev = container_of(cxl, VFIOPCIDevice, cxl);
+ Error *err = NULL;
+
+ if (!vfio_cxl_do_bind_fmws(vdev, &err)) {
+ error_report_err(err);
+ exit(1);
+ }
+}
+
/*
* Learn the CXL geometry the kernel reports: the HPA-backed HDM memory region
* and the trapped HDM decoder block (which BAR carries it and at what offset).
@@ -3717,6 +3887,22 @@ static bool vfio_cxl_setup(VFIOPCIDevice *vdev, Error **errp)
"performance may be slow", vbasedev->name);
}
+ if (DEVICE(vdev)->hotplugged) {
+ /*
+ * The machine is already up, so the CFMWS windows are placed and the
+ * binding can be validated now. Report a failure through errp so a bad
+ * device_add fails cleanly instead of aborting the running VM.
+ */
+ if (!vfio_cxl_do_bind_fmws(vdev, errp)) {
+ vfio_region_exit(&cxl->mem_region);
+ vfio_region_finalize(&cxl->mem_region);
+ return false;
+ }
+ } else {
+ cxl->machine_done.notify = vfio_cxl_bind_fmws;
+ qemu_add_machine_init_done_notifier(&cxl->machine_done);
+ }
+
cxl->enabled = true;
return true;
@@ -3729,6 +3915,12 @@ static void vfio_cxl_teardown(VFIOPCIDevice *vdev)
if (!cxl->enabled) {
return;
}
+
+ if (cxl->machine_done.notify) {
+ qemu_remove_machine_init_done_notifier(&cxl->machine_done);
+ cxl->machine_done.notify = NULL;
+ }
+
if (cxl->mem_region.mem) {
vfio_region_exit(&cxl->mem_region);
vfio_region_finalize(&cxl->mem_region);
diff --git a/hw/vfio/pci.h b/hw/vfio/pci.h
index 06d15e807d..22fe8d7ff7 100644
--- a/hw/vfio/pci.h
+++ b/hw/vfio/pci.h
@@ -135,6 +135,9 @@ typedef struct VFIOCXL {
uint64_t hdm_offset; /* block offset within the component BAR */
uint64_t dpa_size; /* size of the HDM memory region */
VFIORegion mem_region; /* HDM memory, mapped at committed GPA */
+ Notifier machine_done; /* CFMWS validated at machine_done */
+ hwaddr fmws_base; /* base of the memory window */
+ uint64_t fmws_size; /* size of the memory window */
} VFIOCXL;
struct VFIOPCIDevice {
--
2.25.1
^ permalink raw reply related [flat|nested] 27+ messages in thread
* [PATCH v2 07/10] hw/vfio/pci: Map the CXL memory on the guest decoder commit
2026-09-16 18:44 [PATCH v2 00/10] QEMU: CXL Type-2 device passthrough via vfio-pci mhonap
` (5 preceding siblings ...)
2026-09-16 18:44 ` [PATCH v2 06/10] hw/vfio/pci: Bind a CXL device to its fixed memory window mhonap
@ 2026-09-16 18:44 ` mhonap
2026-09-18 16:08 ` Cédric Le Goater
2026-09-20 9:14 ` Junjie Cao
2026-09-16 18:44 ` [PATCH v2 08/10] docs/cxl: Document CXL Type-2 device passthrough mhonap
` (2 subsequent siblings)
9 siblings, 2 replies; 27+ messages in thread
From: mhonap @ 2026-09-16 18:44 UTC (permalink / raw)
To: alex, ankita, jic23, dave.jiang, alejandro.lucero-palau,
smadhavan, pierrick.bouvier, mst, imammedo, anisinha, pbonzini,
eric.auger, peter.maydell, richard.henderson, clg, cohuck
Cc: kjaju, vsethi, zhiw, mhonap, qemu-devel, qemu-arm, linux-cxl
From: Manish Honap <mhonap@nvidia.com>
Overlay the trapped HDM decoder block over the component BAR and forward
its accesses to the kernel. QEMU maps the HDM memory at the base of the
device's CFMWS window and presents that same base back to the guest
through the trapped decoder read
Overlap the CFMWS window MemoryRegion the CXL host bridge maps at that
base, at a higher priority, so precedence is defined by priority rather
than subregion insertion order.
QEMU always maps at the CFMWS base, so a decoder the guest commits at any
other base would leave the guest's view and the mapping diverged. Track
the base the guest writes to the trapped block and refuse to map a commit
at a different base; a firmware-committed decoder the guest never
reprograms keeps the window base.
AI-used-for: code (prototype)
Signed-off-by: Manish Honap <mhonap@nvidia.com>
---
hw/cxl/cxl-host-stubs.c | 5 +
hw/vfio/pci.c | 386 +++++++++++++++++++++++++++++++++++++++-
hw/vfio/pci.h | 8 +
3 files changed, 393 insertions(+), 6 deletions(-)
diff --git a/hw/cxl/cxl-host-stubs.c b/hw/cxl/cxl-host-stubs.c
index 9b515913ea..1067c944e2 100644
--- a/hw/cxl/cxl-host-stubs.c
+++ b/hw/cxl/cxl-host-stubs.c
@@ -23,3 +23,8 @@ GSList *cxl_fmws_get_all_sorted(void)
{
g_assert_not_reached();
}
+
+int cxl_decoder_count_dec(int enc_cnt)
+{
+ g_assert_not_reached();
+}
diff --git a/hw/vfio/pci.c b/hw/vfio/pci.c
index 670e0d1da4..fb3c39d4c6 100644
--- a/hw/vfio/pci.c
+++ b/hw/vfio/pci.c
@@ -1453,6 +1453,39 @@ uint32_t vfio_pci_read_config(PCIDevice *pdev, uint32_t addr, int len)
return val;
}
+static void vfio_cxl_decoder_changed(VFIOPCIDevice *vdev);
+static void vfio_cxl_unmap_mem(VFIOPCIDevice *vdev);
+
+/*
+ * Return true if this config write sets a Function Level Reset bit the kernel
+ * acts on: the PCIe Device Control FLR bit or the Advanced Features FLR bit.
+ * FLR bits are write-1-to-trigger and self-clearing, so inspect the written
+ * value rather than the post-write config.
+ */
+static bool vfio_cxl_flr_write(VFIOPCIDevice *vdev, uint32_t addr,
+ uint32_t val, int len)
+{
+ PCIDevice *pdev = &vdev->parent_obj;
+ uint32_t off;
+
+ if (pdev->exp.exp_cap) {
+ /* PCI_EXP_DEVCTL_BCR_FLR (bit 15) is the high byte of Device Ctrl. */
+ off = pdev->exp.exp_cap + PCI_EXP_DEVCTL + 1;
+ if (off >= addr && off < addr + len &&
+ ((val >> (8 * (off - addr))) & (PCI_EXP_DEVCTL_BCR_FLR >> 8))) {
+ return true;
+ }
+ }
+ if (vdev->cxl.af_offset) {
+ off = vdev->cxl.af_offset + PCI_AF_CTRL;
+ if (off >= addr && off < addr + len &&
+ ((val >> (8 * (off - addr))) & PCI_AF_CTRL_FLR)) {
+ return true;
+ }
+ }
+ return false;
+}
+
void vfio_pci_write_config(PCIDevice *pdev,
uint32_t addr, uint32_t val, int len)
{
@@ -1522,9 +1555,74 @@ void vfio_pci_write_config(PCIDevice *pdev,
vfio_sub_page_bar_update_mapping(pdev, bar);
}
}
+
+ /*
+ * A firmware-committed CXL decoder is already committed at boot, so no
+ * guest control write triggers the HDM mapping. Map it once the guest
+ * enables memory decoding, so the region enters the guest address space
+ * and the IOAS while the device is live. On a clear of memory decoding,
+ * withdraw the overlay: the kernel revokes the HDM PTEs on the same
+ * write, so a retained memslot would fault a guest access during the
+ * disabled interval onto a zapped VMA and stop the VM; the enable path
+ * re-maps it.
+ */
+ if (vdev->cxl.enabled &&
+ range_covers_byte(addr, len, PCI_COMMAND)) {
+ if (pci_get_word(pdev->config + PCI_COMMAND) & PCI_COMMAND_MEMORY) {
+ vfio_cxl_decoder_changed(vdev);
+ } else {
+ vfio_cxl_unmap_mem(vdev);
+ }
+ }
} else {
/* Write everything to QEMU to keep emulated bits correct */
pci_default_write_config(pdev, addr, val, len);
+
+ /*
+ * A guest CXL reset is a write to the CXL Device DVSEC ctrl2 register,
+ * forwarded to the kernel above, which re-commits this firmware-fixed
+ * decoder. If the guest decommitted the decoder before the reset, which
+ * unmapped the HDM memory, that re-commit is not otherwise visible to
+ * QEMU, so rescan the decoder here to restore the mapping. Gated on
+ * memory decoding still being enabled, like the enable path above.
+ */
+ if (vdev->cxl.enabled && vdev->cxl.dvsec_offset &&
+ ranges_overlap(addr, len,
+ vdev->cxl.dvsec_offset +
+ offsetof(CXLDVSECDevice, ctrl2),
+ sizeof_field(CXLDVSECDevice, ctrl2)) &&
+ (pci_get_word(pdev->config + PCI_COMMAND) & PCI_COMMAND_MEMORY)) {
+ vfio_cxl_decoder_changed(vdev);
+ }
+
+ /*
+ * A NoSoftRst- device the guest cycles D3hot->D0 has its physical
+ * decoder restored by the kernel on resume, which re-commits this
+ * firmware-fixed decoder just like a reset. That re-commit is not
+ * otherwise visible to QEMU, so on the transition back to D0 rescan
+ * the decoder to restore a mapping the guest dropped before it
+ * suspended. Gated on memory decoding, like the paths above.
+ */
+ if (vdev->cxl.enabled && pdev->pm_cap &&
+ range_covers_byte(addr, len, pdev->pm_cap + PCI_PM_CTRL) &&
+ (pci_get_word(pdev->config + pdev->pm_cap + PCI_PM_CTRL) &
+ PCI_PM_CTRL_STATE_MASK) == 0 &&
+ (pci_get_word(pdev->config + PCI_COMMAND) & PCI_COMMAND_MEMORY)) {
+ vfio_cxl_decoder_changed(vdev);
+ }
+
+ /*
+ * A Function Level Reset the guest triggers through the PCIe Device
+ * Control or Advanced Features FLR bit is forwarded to the kernel,
+ * which re-commits this firmware-fixed decoder just like a CXL reset.
+ * That re-commit is not otherwise visible to QEMU, so rescan the
+ * decoder to restore a mapping the guest dropped before the FLR. Gated
+ * on memory decoding, like the paths above.
+ */
+ if (vdev->cxl.enabled && vfio_cxl_flr_write(vdev, addr, val, len) &&
+ (pci_get_word(pdev->config + PCI_COMMAND) & PCI_COMMAND_MEMORY)) {
+ vfio_cxl_decoder_changed(vdev);
+ }
}
}
@@ -3811,6 +3909,239 @@ static void vfio_cxl_bind_fmws(Notifier *n, void *data)
}
}
+/*
+ * HDM decoder registers, relative to the trapped decoder block. The block holds
+ * the HDM Decoder Capability register followed by one register set per decoder,
+ * each VFIO_CXL_HDM_DECODER_STRIDE apart, indexed by decoder.
+ */
+#define VFIO_CXL_HDM_CAP 0x00
+#define VFIO_CXL_HDM_DECODER_STRIDE 0x20
+#define VFIO_CXL_HDM_DECODER_BASE_LOW(n) \
+ (0x10 + (n) * VFIO_CXL_HDM_DECODER_STRIDE)
+#define VFIO_CXL_HDM_DECODER_BASE_HIGH(n) \
+ (0x14 + (n) * VFIO_CXL_HDM_DECODER_STRIDE)
+#define VFIO_CXL_HDM_DECODER_CTRL(n) \
+ (0x20 + (n) * VFIO_CXL_HDM_DECODER_STRIDE)
+#define VFIO_CXL_HDM_CTRL_COMMITTED (1 << 10)
+#define VFIO_CXL_HDM_BASE_LOW_MASK 0xf0000000U
+
+static void vfio_cxl_unmap_mem(VFIOPCIDevice *vdev)
+{
+ VFIOCXL *cxl = &vdev->cxl;
+
+ if (!cxl->dpa_mapped) {
+ return;
+ }
+ memory_region_del_subregion(get_system_memory(), cxl->mem_region.mem);
+ cxl->dpa_mapped = false;
+ cxl->mapped_base = 0;
+}
+
+/*
+ * The HDM Decoder Capability register encodes the decoder count in its low
+ * nibble (CXL r3.1 8.2.4.20.1, encodings 0h..Ch for 1..32). Decode it with the
+ * shared cxl_decoder_count_dec() helper so the commit handler walks every
+ * decoder rather than assuming decoder 0. An unreadable register or a reserved
+ * encoding (which the helper decodes to 0) is treated as a single decoder.
+ */
+static unsigned vfio_cxl_decoder_count(VFIOPCIDevice *vdev)
+{
+ off_t off = vdev->cxl.comp_regs_region.fd_offset;
+ uint32_t cap = 0;
+ int count;
+
+ if (pread(vdev->vbasedev.fd, &cap, 4, off + VFIO_CXL_HDM_CAP) != 4) {
+ return 1;
+ }
+ count = cxl_decoder_count_dec(le32_to_cpu(cap) & 0xf);
+ return count > 0 ? count : 1;
+}
+
+/*
+ * The guest committed or tore down an endpoint decoder. Walk each decoder in
+ * the trapped HDM block and (un)map the HDM memory at the CFMWS window base,
+ * not the base the guest programmed (see vfio_cxl_comp_regs_read). This
+ * generation commits one non-interleaved decoder, so the walk stops at the
+ * first committed decoder.
+ */
+static void vfio_cxl_decoder_changed(VFIOPCIDevice *vdev)
+{
+ VFIOCXL *cxl = &vdev->cxl;
+ VFIODevice *vbasedev = &vdev->vbasedev;
+ off_t off = cxl->comp_regs_region.fd_offset;
+ unsigned n, count = vfio_cxl_decoder_count(vdev);
+
+ for (n = 0; n < count; n++) {
+ uint32_t ctrl = 0;
+
+ if (pread(vbasedev->fd, &ctrl, 4,
+ off + VFIO_CXL_HDM_DECODER_CTRL(n)) != 4) {
+ return;
+ }
+ if (!(le32_to_cpu(ctrl) & VFIO_CXL_HDM_CTRL_COMMITTED)) {
+ continue;
+ }
+
+ /*
+ * QEMU always maps at the device's CFMWS base and presents that base
+ * back to the guest, so a decoder the guest committed at any other base
+ * would leave the guest's view and the actual mapping diverged. Reject
+ * it rather than silently relocate. A firmware-committed decoder the
+ * guest never reprogrammed (guest_base_written == false) keeps the
+ * CFMWS base.
+ */
+ if (cxl->guest_base_written) {
+ hwaddr guest_base = ((hwaddr)cxl->guest_base_hi << 32) |
+ cxl->guest_base_lo;
+
+ if (guest_base != cxl->fmws_base) {
+ warn_report("vfio-cxl: %s: guest committed decoder %u at 0x%"
+ HWADDR_PRIx ", not the CFMWS base 0x%" HWADDR_PRIx
+ "; not mapping", vbasedev->name, n, guest_base,
+ cxl->fmws_base);
+ vfio_cxl_unmap_mem(vdev);
+ return;
+ }
+ }
+
+ if (cxl->dpa_mapped && cxl->mapped_base == cxl->fmws_base) {
+ return;
+ }
+ /*
+ * Overlap the CFMWS window MemoryRegion the CXL host bridge already
+ * maps at this base, at a higher priority, so precedence is defined by
+ * priority rather than subregion insertion order.
+ */
+ memory_region_transaction_begin();
+ vfio_cxl_unmap_mem(vdev);
+ memory_region_add_subregion_overlap(get_system_memory(),
+ cxl->fmws_base,
+ cxl->mem_region.mem, 1);
+ memory_region_transaction_commit();
+ cxl->mapped_base = cxl->fmws_base;
+ cxl->dpa_mapped = true;
+ return;
+ }
+
+ /* No committed decoder in the block: tear down any existing mapping. */
+ vfio_cxl_unmap_mem(vdev);
+}
+
+static uint64_t vfio_cxl_comp_regs_read(void *opaque, hwaddr addr,
+ unsigned size)
+{
+ VFIORegion *region = opaque;
+ VFIODevice *vbasedev = region->vbasedev;
+ VFIOPCIDevice *vdev =
+ container_of(region, VFIOPCIDevice, cxl.comp_regs_region);
+ VFIOCXL *cxl = &vdev->cxl;
+ uint32_t val = 0xffffffff;
+
+ if (pread(vbasedev->fd, &val, size, region->fd_offset + addr) != size) {
+ error_report("vfio-cxl: %s: HDM decoder read at 0x%" HWADDR_PRIx
+ " failed", vbasedev->name, addr);
+ }
+ val = le32_to_cpu(val);
+
+ /*
+ * The kernel shadow holds the host physical base for a firmware-committed
+ * decoder. Never expose that to the guest: present the guest physical base
+ * QEMU maps the HDM memory at (the device's CFMWS window). Any decoder's
+ * base registers are virtualized; ctrl, size and the capability header pass
+ * through. The base registers repeat every VFIO_CXL_HDM_DECODER_STRIDE.
+ */
+ if (addr >= VFIO_CXL_HDM_DECODER_BASE_LOW(0) &&
+ (addr - VFIO_CXL_HDM_DECODER_BASE_LOW(0)) %
+ VFIO_CXL_HDM_DECODER_STRIDE == 0) {
+ val = (val & ~VFIO_CXL_HDM_BASE_LOW_MASK) |
+ ((uint32_t)cxl->fmws_base & VFIO_CXL_HDM_BASE_LOW_MASK);
+ } else if (addr >= VFIO_CXL_HDM_DECODER_BASE_HIGH(0) &&
+ (addr - VFIO_CXL_HDM_DECODER_BASE_HIGH(0)) %
+ VFIO_CXL_HDM_DECODER_STRIDE == 0) {
+ val = (uint32_t)(cxl->fmws_base >> 32);
+ }
+
+ return val;
+}
+
+static void vfio_cxl_comp_regs_write(void *opaque, hwaddr addr, uint64_t data,
+ unsigned size)
+{
+ VFIORegion *region = opaque;
+ VFIODevice *vbasedev = region->vbasedev;
+ VFIOPCIDevice *vdev =
+ container_of(region, VFIOPCIDevice, cxl.comp_regs_region);
+ uint32_t val = cpu_to_le32((uint32_t)data);
+
+ if (pwrite(vbasedev->fd, &val, size, region->fd_offset + addr) != size) {
+ error_report("vfio-cxl: %s: HDM decoder write at 0x%" HWADDR_PRIx
+ " failed", vbasedev->name, addr);
+ return;
+ }
+
+ /*
+ * Record the base the guest programs into a decoder so the commit handler
+ * can reject a base other than the device's CFMWS window. The base low
+ * register carries HPA bits [31:28]; the high register carries [63:32].
+ */
+ if (addr >= VFIO_CXL_HDM_DECODER_BASE_LOW(0) &&
+ (addr - VFIO_CXL_HDM_DECODER_BASE_LOW(0)) %
+ VFIO_CXL_HDM_DECODER_STRIDE == 0) {
+ vdev->cxl.guest_base_lo = (uint32_t)data & VFIO_CXL_HDM_BASE_LOW_MASK;
+ vdev->cxl.guest_base_written = true;
+ } else if (addr >= VFIO_CXL_HDM_DECODER_BASE_HIGH(0) &&
+ (addr - VFIO_CXL_HDM_DECODER_BASE_HIGH(0)) %
+ VFIO_CXL_HDM_DECODER_STRIDE == 0) {
+ vdev->cxl.guest_base_hi = (uint32_t)data;
+ vdev->cxl.guest_base_written = true;
+ }
+
+ /*
+ * The kernel runs the lock-on-commit FSM in the write above, so the
+ * committed state is settled by now; a control write on any decoder can
+ * change the mapping.
+ */
+ if (addr >= VFIO_CXL_HDM_DECODER_CTRL(0) &&
+ (addr - VFIO_CXL_HDM_DECODER_CTRL(0)) %
+ VFIO_CXL_HDM_DECODER_STRIDE == 0) {
+ vfio_cxl_decoder_changed(vdev);
+ }
+}
+
+static const MemoryRegionOps vfio_cxl_comp_regs_ops = {
+ .read = vfio_cxl_comp_regs_read,
+ .write = vfio_cxl_comp_regs_write,
+ .endianness = DEVICE_LITTLE_ENDIAN,
+ .valid = { .min_access_size = 4, .max_access_size = 4 },
+ .impl = { .min_access_size = 4, .max_access_size = 4 },
+};
+
+/*
+ * Locate the CXL Device DVSEC (CXL r3.1 8.1.3) in config space. The guest
+ * triggers a CXL reset by writing its ctrl2 register; QEMU rescans the decoder
+ * after that write so the HDM mapping is restored (see vfio_pci_write_config).
+ * Returns the DVSEC config offset, or 0 if the device does not expose it.
+ */
+static uint16_t vfio_cxl_find_device_dvsec(PCIDevice *pdev)
+{
+ uint16_t offset;
+
+ for (offset = PCI_CONFIG_SPACE_SIZE; offset;
+ offset = PCI_EXT_CAP_NEXT(pci_get_long(pdev->config + offset))) {
+ uint32_t hdr = pci_get_long(pdev->config + offset);
+
+ if (PCI_EXT_CAP_ID(hdr) == PCI_EXT_CAP_ID_DVSEC &&
+ (pci_get_long(pdev->config + offset + PCI_DVSEC_HEADER1) & 0xffff)
+ == CXL_VENDOR_ID &&
+ pci_get_word(pdev->config + offset + PCI_DVSEC_HEADER2)
+ == PCIE_CXL_DEVICE_DVSEC) {
+ return offset;
+ }
+ }
+
+ return 0;
+}
+
/*
* Learn the CXL geometry the kernel reports: the HPA-backed HDM memory region
* and the trapped HDM decoder block (which BAR carries it and at what offset).
@@ -3864,6 +4195,8 @@ static bool vfio_cxl_setup(VFIOPCIDevice *vdev, Error **errp)
cxl->dpa_size = mem_info->size;
cxl->comp_bar = cap->bar;
cxl->hdm_offset = cap->offset;
+ cxl->dvsec_offset = vfio_cxl_find_device_dvsec(&vdev->parent_obj);
+ cxl->af_offset = pci_find_capability(&vdev->parent_obj, PCI_CAP_ID_AF);
if (!vfio_cxl_check_topology(vdev, errp)) {
return false;
@@ -3887,6 +4220,26 @@ static bool vfio_cxl_setup(VFIOPCIDevice *vdev, Error **errp)
"performance may be slow", vbasedev->name);
}
+ /*
+ * Trap the HDM decoder block: overlay it, priority 1, over the directly
+ * mapped component BAR, so guest decoder accesses reach the kernel FSM and
+ * the commit becomes visible to QEMU.
+ */
+ if (!vdev->bars[cxl->comp_bar].mr) {
+ error_setg(errp, "vfio-cxl: %s: component BAR %u is not present",
+ vbasedev->name, cxl->comp_bar);
+ goto err;
+ }
+ if (vfio_region_setup_with_ops(OBJECT(vdev), vbasedev,
+ &cxl->comp_regs_region,
+ cxl->comp_regs_region_index, "cxl-comp-regs",
+ &vfio_cxl_comp_regs_ops, errp)) {
+ goto err;
+ }
+ memory_region_add_subregion_overlap(vdev->bars[cxl->comp_bar].mr,
+ cxl->hdm_offset,
+ cxl->comp_regs_region.mem, 1);
+
if (DEVICE(vdev)->hotplugged) {
/*
* The machine is already up, so the CFMWS windows are placed and the
@@ -3894,9 +4247,7 @@ static bool vfio_cxl_setup(VFIOPCIDevice *vdev, Error **errp)
* device_add fails cleanly instead of aborting the running VM.
*/
if (!vfio_cxl_do_bind_fmws(vdev, errp)) {
- vfio_region_exit(&cxl->mem_region);
- vfio_region_finalize(&cxl->mem_region);
- return false;
+ goto err;
}
} else {
cxl->machine_done.notify = vfio_cxl_bind_fmws;
@@ -3906,6 +4257,10 @@ static bool vfio_cxl_setup(VFIOPCIDevice *vdev, Error **errp)
cxl->enabled = true;
return true;
+
+err:
+ vfio_cxl_teardown(vdev);
+ return false;
}
static void vfio_cxl_teardown(VFIOPCIDevice *vdev)
@@ -3916,15 +4271,26 @@ static void vfio_cxl_teardown(VFIOPCIDevice *vdev)
return;
}
- if (cxl->machine_done.notify) {
- qemu_remove_machine_init_done_notifier(&cxl->machine_done);
- cxl->machine_done.notify = NULL;
+ if (cxl->comp_regs_region.mem) {
+ if (vdev->bars[cxl->comp_bar].mr) {
+ memory_region_del_subregion(vdev->bars[cxl->comp_bar].mr,
+ cxl->comp_regs_region.mem);
+ }
+ vfio_region_exit(&cxl->comp_regs_region);
+ vfio_region_finalize(&cxl->comp_regs_region);
}
+ vfio_cxl_unmap_mem(vdev);
+
if (cxl->mem_region.mem) {
vfio_region_exit(&cxl->mem_region);
vfio_region_finalize(&cxl->mem_region);
}
+
+ if (cxl->machine_done.notify) {
+ qemu_remove_machine_init_done_notifier(&cxl->machine_done);
+ cxl->machine_done.notify = NULL;
+ }
}
static void vfio_pci_realize(PCIDevice *pdev, Error **errp)
@@ -4127,6 +4493,14 @@ static void vfio_exitfn(PCIDevice *pdev)
vfio_pci_teardown_msi(vdev);
vfio_pci_disable_rp_atomics(vdev);
vfio_pci_bars_exit(vdev);
+ /*
+ * The committed HDM overlay is a subregion of system memory owned by this
+ * device, so it holds a reference that would keep the object alive past
+ * unrealize and block instance_finalize (where vfio_cxl_teardown otherwise
+ * runs). Drop it here; the call is idempotent for a device that never
+ * mapped or is not CXL.
+ */
+ vfio_cxl_unmap_mem(vdev);
vfio_migration_exit(vbasedev);
if (!vbasedev->mdev) {
pci_device_unset_iommu_device(pdev);
diff --git a/hw/vfio/pci.h b/hw/vfio/pci.h
index 22fe8d7ff7..5bfa976112 100644
--- a/hw/vfio/pci.h
+++ b/hw/vfio/pci.h
@@ -133,11 +133,19 @@ typedef struct VFIOCXL {
uint32_t comp_regs_region_index; /* trapped HDM decoder register block */
uint32_t comp_bar; /* component BAR carrying that block */
uint64_t hdm_offset; /* block offset within the component BAR */
+ uint16_t dvsec_offset; /* CXL Device DVSEC offset, 0 if none */
+ uint16_t af_offset; /* PCI Advanced Features cap, 0 if none */
uint64_t dpa_size; /* size of the HDM memory region */
VFIORegion mem_region; /* HDM memory, mapped at committed GPA */
Notifier machine_done; /* CFMWS validated at machine_done */
hwaddr fmws_base; /* base of the memory window */
uint64_t fmws_size; /* size of the memory window */
+ VFIORegion comp_regs_region; /* trapped HDM decoder block */
+ hwaddr mapped_base; /* GPA the HDM memory is mapped at */
+ bool dpa_mapped; /* HDM memory currently in system memory */
+ uint32_t guest_base_lo; /* decoder base low the guest wrote */
+ uint32_t guest_base_hi; /* decoder base high the guest wrote */
+ bool guest_base_written; /* the guest wrote a decoder base */
} VFIOCXL;
struct VFIOPCIDevice {
--
2.25.1
^ permalink raw reply related [flat|nested] 27+ messages in thread
* [PATCH v2 08/10] docs/cxl: Document CXL Type-2 device passthrough
2026-09-16 18:44 [PATCH v2 00/10] QEMU: CXL Type-2 device passthrough via vfio-pci mhonap
` (6 preceding siblings ...)
2026-09-16 18:44 ` [PATCH v2 07/10] hw/vfio/pci: Map the CXL memory on the guest decoder commit mhonap
@ 2026-09-16 18:44 ` mhonap
2026-09-16 18:44 ` [PATCH v2 09/10] hw/arm/smmu-common: Allow pxb-cxl as an SMMUv3 primary bus mhonap
2026-09-16 18:44 ` [PATCH v2 10/10] hw/pci-host: Emit a _DSM on pxb-cxl to preserve firmware PCI config mhonap
9 siblings, 0 replies; 27+ messages in thread
From: mhonap @ 2026-09-16 18:44 UTC (permalink / raw)
To: alex, ankita, jic23, dave.jiang, alejandro.lucero-palau,
smadhavan, pierrick.bouvier, mst, imammedo, anisinha, pbonzini,
eric.auger, peter.maydell, richard.henderson, clg, cohuck
Cc: kjaju, vsethi, zhiw, mhonap, qemu-devel, qemu-arm, linux-cxl
From: Manish Honap <mhonap@nvidia.com>
Describe the supported single-endpoint passthrough topology, the command
line, and the kernel dependency, and note that a guest CXL reset is
handled by the host kernel rather than QEMU.
AI-used-for: docs
Signed-off-by: Manish Honap <mhonap@nvidia.com>
---
docs/system/devices/cxl.rst | 45 +++++++++++++++++++++++++++++++++++++
1 file changed, 45 insertions(+)
diff --git a/docs/system/devices/cxl.rst b/docs/system/devices/cxl.rst
index 6bee339cea..3cdd320457 100644
--- a/docs/system/devices/cxl.rst
+++ b/docs/system/devices/cxl.rst
@@ -423,6 +423,51 @@ Volatile Memory device::
-device cxl-type3,bus=root_port13,volatile-memdev=vmem0,id=cxl-vmem0 \
-M cxl-fmw.0.targets.0=cxl.1,cxl-fmw.0.size=4G
+Type 2 device passthrough
+-------------------------
+
+A CXL Type 2 device (an accelerator with host-managed device memory) can be
+assigned to a guest with vfio-pci, so the guest reaches the device memory
+through its own CXL stack. This pairs with the kernel vfio-cxl series: the
+host kernel fixes the device memory at a host physical range before the guest
+sees the device, and QEMU maps that range at the guest physical address the
+guest programs into its endpoint HDM decoder.
+
+The device memory reaches the guest as a CXL fixed memory window (``cxl-fmw``),
+advertised through CEDT, exactly like a Type 3 window; there is no separate
+device memory-map slot. Only the endpoint decoder is programmed, by the guest,
+so the host bridge must stay in HDM passthrough mode: a single ``cxl-rp`` under
+the ``pxb-cxl`` and no ``hdm_for_passthrough``. Switch-attached and interleaved
+topologies are rejected.
+
+Example command line::
+
+ -machine q35,cxl=on
+ -device pxb-cxl,bus_nr=12,bus=pcie.0,id=cxl.1
+ -device cxl-rp,port=0,bus=cxl.1,id=rp0,chassis=0,slot=2
+ -device vfio-pci,host=<BDF>,bus=rp0,id=cxl-ep0
+ -M cxl-fmw.0.targets.0=cxl.1,cxl-fmw.0.size=<device-mem-size>
+
+The window must be a single-target ``cxl-fmw`` that targets the device's
+``pxb-cxl`` and is at least the size of the device memory. The guest triggers a
+CXL reset by writing the CXL Device DVSEC; the host kernel runs that sequence,
+so QEMU has no reset handling of its own.
+
+On arm64, pass ``accel=on`` to the ``arm-smmuv3`` when passing a Type 2 device
+through. The accelerated SMMUv3 describes the device MSI doorbell to the guest
+through an IORT Reserved Memory Range (RMR) node, which reserves a fixed guest
+IOVA, so OSPM must preserve the firmware PCI resource assignments rather than
+re-enumerate them. QEMU requests that through PCI Firmware ``_DSM`` function 5
+(preserve firmware PCI configuration), which on arm64 is emitted for the CXL
+host bridge only when the machine requests preserved configuration, that is,
+the accelerated SMMUv3 path. This ``_DSM`` is not about the CXL decoder:
+vfio-pci keeps the host BAR fixed, and QEMU's trapped component-register block
+is a subregion of the guest BAR MemoryRegion, so it follows any guest-visible
+BAR relocation on its own. x86 does not use the accelerated SMMU, so it does
+not advertise or implement preserve-configuration function 5; the CXL host
+bridge still emits the ``_DSM`` method, but its function 0 returns an empty
+support mask.
+
Deprecations
------------
--
2.25.1
^ permalink raw reply related [flat|nested] 27+ messages in thread
* [PATCH v2 09/10] hw/arm/smmu-common: Allow pxb-cxl as an SMMUv3 primary bus
2026-09-16 18:44 [PATCH v2 00/10] QEMU: CXL Type-2 device passthrough via vfio-pci mhonap
` (7 preceding siblings ...)
2026-09-16 18:44 ` [PATCH v2 08/10] docs/cxl: Document CXL Type-2 device passthrough mhonap
@ 2026-09-16 18:44 ` mhonap
2026-09-16 18:44 ` [PATCH v2 10/10] hw/pci-host: Emit a _DSM on pxb-cxl to preserve firmware PCI config mhonap
9 siblings, 0 replies; 27+ messages in thread
From: mhonap @ 2026-09-16 18:44 UTC (permalink / raw)
To: alex, ankita, jic23, dave.jiang, alejandro.lucero-palau,
smadhavan, pierrick.bouvier, mst, imammedo, anisinha, pbonzini,
eric.auger, peter.maydell, richard.henderson, clg, cohuck
Cc: kjaju, vsethi, zhiw, mhonap, qemu-devel, qemu-arm, linux-cxl,
Shameer Kolothum
From: Manish Honap <mhonap@nvidia.com>
The SMMUv3 primary-bus check only accepted pxb-pcie as a valid extra
root complex, so a CXL device behind a pxb-cxl could not reach the
IOMMU: attaching an arm-smmuv3 to a pxb-cxl bus failed at realize with
"SMMU should be attached to a default PCIe root complex (pcie.0) or a
pxb-pcie based root complex".
A pxb-cxl uses the same PCIe-compatible bus as a pxb-pcie, so accept it
too. A CXL Type-2 device passed through with vfio-pci sits behind a
pxb-cxl and needs SMMU translation for its DMA and ATS just like a
pxb-pcie device.
AI-used-for: code (prototype)
Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com>
Signed-off-by: Manish Honap <mhonap@nvidia.com>
---
hw/arm/smmu-common.c | 19 ++++++++++---------
1 file changed, 10 insertions(+), 9 deletions(-)
diff --git a/hw/arm/smmu-common.c b/hw/arm/smmu-common.c
index 8e40ba603d..42fb098ab2 100644
--- a/hw/arm/smmu-common.c
+++ b/hw/arm/smmu-common.c
@@ -963,19 +963,20 @@ static void smmu_base_realize(DeviceState *dev, Error **errp)
s->iommu_ops = &smmu_ops;
}
/*
- * We only allow default PCIe Root Complex(pcie.0) or pxb-pcie based extra
- * root complexes to be associated with SMMU.
+ * We only allow the default PCIe root complex (pcie.0) or pxb-pcie /
+ * pxb-cxl based extra root complexes to be associated with SMMU.
*/
if (pci_bus_is_express(pci_bus) && pci_bus_is_root(pci_bus) &&
object_dynamic_cast(OBJECT(pci_bus)->parent, TYPE_PCI_HOST_BRIDGE)) {
/*
- * This condition matches either the default pcie.0, pxb-pcie, or
- * pxb-cxl. For both pxb-pcie and pxb-cxl, parent_dev will be set.
- * Currently, we don't allow pxb-cxl as it requires further
- * verification. Therefore, make sure this is indeed pxb-pcie.
+ * pcie.0 has no parent_dev; pxb-pcie and pxb-cxl do. Accept both bus
+ * types explicitly so other root complexes are still rejected. A
+ * pxb-cxl uses the same PCIe-compatible bus, so a CXL device behind it
+ * needs the SMMU just like a pxb-pcie one.
*/
if (pci_bus->parent_dev) {
- if (!object_dynamic_cast(OBJECT(pci_bus), TYPE_PXB_PCIE_BUS)) {
+ if (!object_dynamic_cast(OBJECT(pci_bus), TYPE_PXB_PCIE_BUS) &&
+ !object_dynamic_cast(OBJECT(pci_bus), TYPE_PXB_CXL_BUS)) {
goto out_err;
}
}
@@ -990,8 +991,8 @@ static void smmu_base_realize(DeviceState *dev, Error **errp)
return;
}
out_err:
- error_setg(errp, "SMMU should be attached to a default PCIe root complex"
- "(pcie.0) or a pxb-pcie based root complex");
+ error_setg(errp, "SMMU should be attached to a default PCIe root complex "
+ "(pcie.0), a pxb-pcie, or a pxb-cxl based root complex");
}
/*
--
2.25.1
^ permalink raw reply related [flat|nested] 27+ messages in thread
* [PATCH v2 10/10] hw/pci-host: Emit a _DSM on pxb-cxl to preserve firmware PCI config
2026-09-16 18:44 [PATCH v2 00/10] QEMU: CXL Type-2 device passthrough via vfio-pci mhonap
` (8 preceding siblings ...)
2026-09-16 18:44 ` [PATCH v2 09/10] hw/arm/smmu-common: Allow pxb-cxl as an SMMUv3 primary bus mhonap
@ 2026-09-16 18:44 ` mhonap
2026-09-20 9:14 ` Junjie Cao
9 siblings, 1 reply; 27+ messages in thread
From: mhonap @ 2026-09-16 18:44 UTC (permalink / raw)
To: alex, ankita, jic23, dave.jiang, alejandro.lucero-palau,
smadhavan, pierrick.bouvier, mst, imammedo, anisinha, pbonzini,
eric.auger, peter.maydell, richard.henderson, clg, cohuck
Cc: kjaju, vsethi, zhiw, mhonap, qemu-devel, qemu-arm, linux-cxl,
Shameer Kolothum
From: Manish Honap <mhonap@nvidia.com>
A pxb-cxl host bridge had no PCI host-bridge _DSM method, so it could not
emit function 5 (preserve firmware PCI configuration) even when the
machine asks OSPM to keep the firmware resource assignments.
That preservation is required by the accelerated SMMUv3 (accel=on), which
describes MSI 1:1 mappings through IORT RMR nodes. An RMR reserves a fixed
IOVA range, so OSPM must not re-enumerate and reassign the PCI resources
underneath it. virt sets pci_preserve_config for that case and expects
every host bridge to emit the _DSM; the pxb-cxl bridge was the one that
did not, so a CXL topology under an accelerated SMMU lost the guarantee.
The BAR configuration itself is safe without this. vfio-pci marks every BAR
dword ALL_VIRT, so a guest BAR write updates only the virtual config and
cannot move the host resource; QEMU's trapped comp-regs region is a
subregion of the BAR MemoryRegion and follows any guest-visible relocation,
while the kernel's physical decoder mapping stays fixed. The CXL.mem window
comes through CEDT/CFMWS and is likewise unaffected by PCI resource
assignment.
Wire preserve_config through GPEXConfig into the CXL host bridge OSC path
so pxb-cxl bridges emit the _DSM function 5 when the machine preserves the
firmware configuration. Rename build_cxl_osc_method() to
acpi_dsdt_add_cxl_host_bridge_methods() to match the pxb-pcie analogue
acpi_dsdt_add_host_bridge_methods(), since it now appends the _OSC and,
under preserve_config, the _DSM. Gate the _DSM on preserve_config: x86 q35
passes false as it does not use the accelerated SMMU, so it emits no _DSM
and its ACPI tables stay byte for byte what they were.
AI-used-for: code (prototype)
Signed-off-by: Shameer Kolothum <skolothumtho@nvidia.com>
Signed-off-by: Manish Honap <mhonap@nvidia.com>
---
hw/acpi/Kconfig | 1 +
hw/acpi/cxl-stub.c | 2 +-
hw/acpi/cxl.c | 12 +++++++++++-
hw/acpi/pci.c | 40 +++++++++++++++++++++++++++++++++++++++
hw/i386/acpi-build.c | 2 +-
hw/pci-host/gpex-acpi.c | 42 ++---------------------------------------
include/hw/acpi/cxl.h | 2 +-
include/hw/acpi/pci.h | 1 +
8 files changed, 58 insertions(+), 44 deletions(-)
diff --git a/hw/acpi/Kconfig b/hw/acpi/Kconfig
index daabbe6cd1..9490b75c94 100644
--- a/hw/acpi/Kconfig
+++ b/hw/acpi/Kconfig
@@ -88,3 +88,4 @@ config ACPI_ERST
config ACPI_CXL
bool
depends on ACPI
+ select ACPI_PCI
diff --git a/hw/acpi/cxl-stub.c b/hw/acpi/cxl-stub.c
index 15bc21076b..d7c6731975 100644
--- a/hw/acpi/cxl-stub.c
+++ b/hw/acpi/cxl-stub.c
@@ -6,7 +6,7 @@
#include "hw/acpi/aml-build.h"
#include "hw/acpi/cxl.h"
-void build_cxl_osc_method(Aml *dev)
+void acpi_dsdt_add_cxl_host_bridge_methods(Aml *dev, bool preserve_config)
{
g_assert_not_reached();
}
diff --git a/hw/acpi/cxl.c b/hw/acpi/cxl.c
index 77c1db6561..dc2a5fcdfe 100644
--- a/hw/acpi/cxl.c
+++ b/hw/acpi/cxl.c
@@ -23,6 +23,7 @@
#include "hw/pci/pci_host.h"
#include "hw/cxl/cxl.h"
#include "hw/cxl/cxl_host.h"
+#include "hw/acpi/pci.h"
#include "hw/mem/memory-device.h"
#include "hw/acpi/acpi.h"
#include "hw/acpi/aml-build.h"
@@ -320,11 +321,20 @@ static Aml *__build_cxl_osc_method(void)
return method;
}
-void build_cxl_osc_method(Aml *dev)
+void acpi_dsdt_add_cxl_host_bridge_methods(Aml *dev, bool preserve_config)
{
aml_append(dev, aml_name_decl("SUPP", aml_int(0)));
aml_append(dev, aml_name_decl("CTRL", aml_int(0)));
aml_append(dev, aml_name_decl("SUPC", aml_int(0)));
aml_append(dev, aml_name_decl("CTRC", aml_int(0)));
aml_append(dev, __build_cxl_osc_method());
+ /*
+ * Only a machine that asks OSPM to preserve the firmware PCI configuration
+ * needs the _DSM (function 5). Emitting it unconditionally would change the
+ * DSDT of machines that pass preserve_config=false (x86 q35), breaking the
+ * golden-table tests for no functional gain, so gate it on preserve_config.
+ */
+ if (preserve_config) {
+ aml_append(dev, build_pci_host_bridge_dsm_method(preserve_config));
+ }
}
diff --git a/hw/acpi/pci.c b/hw/acpi/pci.c
index c82924be86..f1feb33fc2 100644
--- a/hw/acpi/pci.c
+++ b/hw/acpi/pci.c
@@ -351,3 +351,43 @@ Aml *build_pci_host_bridge_osc_method(bool enable_native_pcie_hotplug)
aml_append(method, aml_return(aml_arg(3)));
return method;
}
+
+Aml *build_pci_host_bridge_dsm_method(bool preserve_config)
+{
+ Aml *method = aml_method("_DSM", 4, AML_NOTSERIALIZED);
+ Aml *UUID, *ifctx, *ifctx1, *buf;
+ uint8_t byte_list[1] = {0};
+
+ /*
+ * PCI Firmware Specification 3.0
+ * 4.6.1. _DSM for PCI Express Slot Information
+ * The UUID in _DSM in this context is
+ * {E5C937D0-3553-4D7A-9117-EA4D19C3434D}
+ */
+ UUID = aml_touuid("E5C937D0-3553-4D7A-9117-EA4D19C3434D");
+ ifctx = aml_if(aml_equal(aml_arg(0), UUID));
+ ifctx1 = aml_if(aml_equal(aml_arg(2), aml_int(0)));
+ if (preserve_config) {
+ /* support functions other than 0, specifically function 5 */
+ byte_list[0] = 0x21;
+ }
+ buf = aml_buffer(1, byte_list);
+ aml_append(ifctx1, aml_return(buf));
+ aml_append(ifctx, ifctx1);
+ if (preserve_config) {
+ Aml *ifctx2 = aml_if(aml_equal(aml_arg(2), aml_int(5)));
+ /*
+ * 0 - The operating system must not ignore the PCI configuration that
+ * firmware has done at boot time.
+ */
+ aml_append(ifctx2, aml_return(aml_int(0)));
+ aml_append(ifctx, ifctx2);
+ }
+
+ aml_append(method, ifctx);
+
+ byte_list[0] = 0;
+ buf = aml_buffer(1, byte_list);
+ aml_append(method, aml_return(buf));
+ return method;
+}
diff --git a/hw/i386/acpi-build.c b/hw/i386/acpi-build.c
index 6d7308b4e4..c8f02ea608 100644
--- a/hw/i386/acpi-build.c
+++ b/hw/i386/acpi-build.c
@@ -1019,7 +1019,7 @@ build_dsdt(GArray *table_data, BIOSLinker *linker,
aml_append(aml_pkg, aml_eisaid("PNP0A08"));
aml_append(aml_pkg, aml_eisaid("PNP0A03"));
aml_append(dev, aml_name_decl("_CID", aml_pkg));
- build_cxl_osc_method(dev);
+ acpi_dsdt_add_cxl_host_bridge_methods(dev, false);
} else if (pci_bus_is_express(bus)) {
aml_append(dev, aml_name_decl("_HID", aml_eisaid("PNP0A08")));
aml_append(dev, aml_name_decl("_CID", aml_eisaid("PNP0A03")));
diff --git a/hw/pci-host/gpex-acpi.c b/hw/pci-host/gpex-acpi.c
index d9820f9b41..014f647c7f 100644
--- a/hw/pci-host/gpex-acpi.c
+++ b/hw/pci-host/gpex-acpi.c
@@ -51,45 +51,6 @@ static void acpi_dsdt_add_pci_route_table(Aml *dev, uint32_t irq,
}
}
-static Aml *build_pci_host_bridge_dsm_method(bool preserve_config)
-{
- Aml *method = aml_method("_DSM", 4, AML_NOTSERIALIZED);
- Aml *UUID, *ifctx, *ifctx1, *buf;
- uint8_t byte_list[1] = {0};
-
- /* PCI Firmware Specification 3.0
- * 4.6.1. _DSM for PCI Express Slot Information
- * The UUID in _DSM in this context is
- * {E5C937D0-3553-4D7A-9117-EA4D19C3434D}
- */
- UUID = aml_touuid("E5C937D0-3553-4D7A-9117-EA4D19C3434D");
- ifctx = aml_if(aml_equal(aml_arg(0), UUID));
- ifctx1 = aml_if(aml_equal(aml_arg(2), aml_int(0)));
- if (preserve_config) {
- /* support functions other than 0, specifically function 5 */
- byte_list[0] = 0x21;
- }
- buf = aml_buffer(1, byte_list);
- aml_append(ifctx1, aml_return(buf));
- aml_append(ifctx, ifctx1);
- if (preserve_config) {
- Aml *ifctx2 = aml_if(aml_equal(aml_arg(2), aml_int(5)));
- /*
- * 0 - The operating system must not ignore the PCI configuration that
- * firmware has done at boot time.
- */
- aml_append(ifctx2, aml_return(aml_int(0)));
- aml_append(ifctx, ifctx2);
- }
-
- aml_append(method, ifctx);
-
- byte_list[0] = 0;
- buf = aml_buffer(1, byte_list);
- aml_append(method, aml_return(buf));
- return method;
-}
-
static void acpi_dsdt_add_host_bridge_methods(Aml *dev,
bool enable_native_pcie_hotplug,
bool preserve_config)
@@ -164,7 +125,8 @@ void acpi_dsdt_add_gpex(Aml *scope, struct GPEXConfig *cfg)
aml_append(dev, aml_name_decl("_CRS", crs));
if (is_cxl) {
- build_cxl_osc_method(dev);
+ acpi_dsdt_add_cxl_host_bridge_methods(dev,
+ cfg->preserve_config);
} else {
/* pxb bridges do not have ACPI PCI Hot-plug enabled */
acpi_dsdt_add_host_bridge_methods(dev, true,
diff --git a/include/hw/acpi/cxl.h b/include/hw/acpi/cxl.h
index 8f22c71530..6fe6c9c58d 100644
--- a/include/hw/acpi/cxl.h
+++ b/include/hw/acpi/cxl.h
@@ -24,7 +24,7 @@
void cxl_build_cedt(GArray *table_offsets, GArray *table_data,
BIOSLinker *linker, const char *oem_id,
const char *oem_table_id, CXLState *cxl_state);
-void build_cxl_osc_method(Aml *dev);
+void acpi_dsdt_add_cxl_host_bridge_methods(Aml *dev, bool preserve_config);
void build_cxl_dsm_method(Aml *dev);
#endif
diff --git a/include/hw/acpi/pci.h b/include/hw/acpi/pci.h
index 20b672575f..c7de33f8cf 100644
--- a/include/hw/acpi/pci.h
+++ b/include/hw/acpi/pci.h
@@ -42,6 +42,7 @@ void build_pci_bridge_aml(AcpiDevAmlIf *adev, Aml *scope);
void build_srat_generic_affinity_structures(GArray *table_data);
Aml *build_pci_host_bridge_osc_method(bool enable_native_pcie_hotplug);
+Aml *build_pci_host_bridge_dsm_method(bool preserve_config);
Aml *build_pci_bridge_edsm(void);
#endif
--
2.25.1
^ permalink raw reply related [flat|nested] 27+ messages in thread
* Re: [PATCH v2 04/10] hw/vfio/pci: Enforce the passthrough topology for a CXL device
2026-09-16 18:44 ` [PATCH v2 04/10] hw/vfio/pci: Enforce the passthrough topology for a CXL device mhonap
@ 2026-09-17 13:20 ` Cédric Le Goater
2026-09-21 10:45 ` Manish Honap
0 siblings, 1 reply; 27+ messages in thread
From: Cédric Le Goater @ 2026-09-17 13:20 UTC (permalink / raw)
To: mhonap, alex, ankita, jic23, dave.jiang, alejandro.lucero-palau,
smadhavan, pierrick.bouvier, mst, imammedo, anisinha, pbonzini,
eric.auger, peter.maydell, richard.henderson, cohuck
Cc: kjaju, vsethi, zhiw, qemu-devel, qemu-arm, linux-cxl
On 9/16/26 20:44, mhonap@nvidia.com wrote:
> From: Manish Honap <mhonap@nvidia.com>
>
> Only the guest's endpoint HDM decoder is programmed by the guest; the
> decoders above it are covered only while the pxb-cxl host bridge stays in
> passthrough mode. Reject a switch in the path, or a host bridge taken out
> of passthrough, at realize before any mapping exists.
>
> AI-used-for: code (prototype)
> Signed-off-by: Manish Honap <mhonap@nvidia.com>
> ---
> hw/vfio/pci.c | 73 +++++++++++++++++++++++++++++++++++++++++++++++++++
> 1 file changed, 73 insertions(+)
>
> diff --git a/hw/vfio/pci.c b/hw/vfio/pci.c
> index 2f84af5cf8..10b13f200d 100644
> --- a/hw/vfio/pci.c
> +++ b/hw/vfio/pci.c
> @@ -27,6 +27,7 @@
> #include "hw/pci/msi.h"
> #include "hw/pci/msix.h"
> #include "hw/pci/pci_bridge.h"
> +#include "hw/cxl/cxl.h"
> #include "hw/core/qdev-properties.h"
> #include "hw/core/qdev-properties-system.h"
> #include "hw/vfio/vfio-cpr.h"
> @@ -3570,6 +3571,73 @@ bool vfio_pci_interrupt_setup(VFIOPCIDevice *vdev, Error **errp)
> return true;
> }
>
> +/*
> + * Walk the endpoint's ancestor bridges up to its pxb-cxl host bridge. A CXL
> + * switch in the path is rejected: the guest would then have to program switch
> + * decoders, which the passthrough model does not yet cover. This is the single
> + * place the supported-topology assumption lives, so switch support later
> + * relaxes it here.
> + */
> +static PXBCXLDev *vfio_cxl_find_pxb(VFIOPCIDevice *vdev, Error **errp)
> +{
> + PCIBus *bus = pci_get_bus(&vdev->parent_obj);
Could this routine be moved to the CXL subsytem component ?
> +
> + if (!object_dynamic_cast(OBJECT(bus), TYPE_CXL_BUS)) {
> + error_setg(errp,
> + "vfio-cxl: %s: endpoint is not on a CXL root port (cxl-rp)",
> + vdev->vbasedev.name);
> + return NULL;
> + }
> +
> + for (; bus; ) {
Why walk, since this is looking for the case "no-switch in between the
bus and the ep" ? the no-switch case is just two hops in the topology.
> + PCIDevice *bridge = bus->parent_dev;
> +
> + if (!bridge) {
> + break;
> + }
> + if (object_dynamic_cast(OBJECT(bridge), TYPE_CXL_USP) ||
> + object_dynamic_cast(OBJECT(bridge), TYPE_CXL_DSP)) {
> + error_setg(errp,
> + "vfio-cxl: %s: switch-attached topology not supported",
> + vdev->vbasedev.name);
> + return NULL;
> + }
> + if (object_dynamic_cast(OBJECT(bridge), TYPE_PXB_CXL_DEV)) {
> + return PXB_CXL_DEV(bridge);
> + }
> + if (pci_bus_is_root(bus)) {
> + break;
> + }
> + bus = pci_get_bus(bridge);
> + }
> +
> + error_setg(errp, "vfio-cxl: %s: not attached below a pxb-cxl host bridge",
> + vdev->vbasedev.name);
> + return NULL;
> +}
> +
> +/*
> + * Only guest's endpoint decoder is programmed, so the host bridge must stay in
> + * HDM passthrough mode. Reject anything else at realize, before a mapping is
> + * built. cxl_get_hb_passthrough() reflects reset state and is not
> + * settled yet here, so gate on the static hdm_for_passthrough property.
> + */
> +static bool vfio_cxl_check_topology(VFIOPCIDevice *vdev, Error **errp)
> +{
> + PXBCXLDev *pxb = vfio_cxl_find_pxb(vdev, errp);
> +
> + if (!pxb) {
> + return false;
> + }
> + if (pxb->hdm_for_passthrough) {
This naming is a bit confusing:
hdm_for_passthrough=false means "use HDM passthrough mode" (no decoders), while
hdm_for_passthrough=true means "keep HDM decoders even in passthrough topology.
I will get used to it.
Thanks,
C.
> + error_setg(errp,
> + "vfio-cxl: %s: the pxb-cxl must be in HDM passthrough mode "
> + "(do not set hdm_for_passthrough)", vdev->vbasedev.name);
> + return false;
> + }
> + return true;
> +}
> +
> /*
> * Learn the CXL geometry the kernel reports: the HPA-backed HDM memory region
> * and the trapped HDM decoder block (which BAR carries it and at what offset).
> @@ -3623,6 +3691,11 @@ static bool vfio_cxl_setup(VFIOPCIDevice *vdev, Error **errp)
> cxl->dpa_size = mem_info->size;
> cxl->comp_bar = cap->bar;
> cxl->hdm_offset = cap->offset;
> +
> + if (!vfio_cxl_check_topology(vdev, errp)) {
> + return false;
> + }
> +
> cxl->enabled = true;
>
> return true;
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v2 02/10] hw/vfio/region: Add vfio_region_setup_with_ops()
2026-09-16 18:44 ` [PATCH v2 02/10] hw/vfio/region: Add vfio_region_setup_with_ops() mhonap
@ 2026-09-17 13:27 ` Cédric Le Goater
2026-09-21 10:48 ` Manish Honap
0 siblings, 1 reply; 27+ messages in thread
From: Cédric Le Goater @ 2026-09-17 13:27 UTC (permalink / raw)
To: mhonap, alex, ankita, jic23, dave.jiang, alejandro.lucero-palau,
smadhavan, pierrick.bouvier, mst, imammedo, anisinha, pbonzini,
eric.auger, peter.maydell, richard.henderson, cohuck
Cc: kjaju, vsethi, zhiw, qemu-devel, qemu-arm, linux-cxl
On 9/16/26 20:44, mhonap@nvidia.com wrote:
> From: Manish Honap <mhonap@nvidia.com>
>
> A CXL Type-2 device traps its HDM decoder block, so that region needs
> custom MemoryRegionOps rather than the pass-through default. Split the
> setup body out and let a caller supply the ops; NULL keeps the existing
> behaviour.
>
> AI-used-for: code (prototype)
> Signed-off-by: Manish Honap <mhonap@nvidia.com>
> ---
> hw/vfio/region.c | 27 ++++++++++++++++++++++++---
> hw/vfio/vfio-region.h | 3 +++
> 2 files changed, 27 insertions(+), 3 deletions(-)
>
> diff --git a/hw/vfio/region.c b/hw/vfio/region.c
> index 54ad11a6c8..21b54978b5 100644
> --- a/hw/vfio/region.c
> +++ b/hw/vfio/region.c
> @@ -228,8 +228,9 @@ static int vfio_setup_region_sparse_mmaps(VFIORegion *region,
> return 0;
> }
>
> -int vfio_region_setup(Object *obj, VFIODevice *vbasedev, VFIORegion *region,
> - int index, const char *name, Error **errp)
> +static int vfio_region_do_setup(Object *obj, VFIODevice *vbasedev,
> + VFIORegion *region, int index, const char *name,
> + const MemoryRegionOps *ops, Error **errp)
> {
> struct vfio_region_info *info = NULL;
> int ret;
> @@ -249,7 +250,7 @@ int vfio_region_setup(Object *obj, VFIODevice *vbasedev, VFIORegion *region,
>
> if (region->size) {
> region->mem = g_new0(MemoryRegion, 1);
> - memory_region_init_io(region->mem, obj, &vfio_region_ops,
> + memory_region_init_io(region->mem, obj, ops,
> region, name, region->size);
>
> if (!vbasedev->no_mmap &&
> @@ -273,6 +274,26 @@ int vfio_region_setup(Object *obj, VFIODevice *vbasedev, VFIORegion *region,
> return 0;
> }
>
> +int vfio_region_setup(Object *obj, VFIODevice *vbasedev, VFIORegion *region,
> + int index, const char *name, Error **errp)
> +{
> + return vfio_region_do_setup(obj, vbasedev, region, index, name,
> + &vfio_region_ops, errp);
> +}
> +
> +/*
> + * Like vfio_region_setup() but traps the region through @ops instead of the
> + * default pass-through, so a caller can intercept accesses (the CXL HDM
> + * decoder block). A NULL @ops keeps the default.
> + */
> +int vfio_region_setup_with_ops(Object *obj, VFIODevice *vbasedev,
> + VFIORegion *region, int index, const char *name,
> + const MemoryRegionOps *ops, Error **errp)
> +{
> + return vfio_region_do_setup(obj, vbasedev, region, index, name,
> + ops ? ops : &vfio_region_ops, errp);
if you're calling _with_ops(), you have custom ops, that's whole point.
Passing NULL is a caller bug, no need to have a fall back. Or assert().
Thanks,
C.
> +}
> +
> static void vfio_subregion_unmap(VFIORegion *region, int index)
> {
> trace_vfio_region_unmap(memory_region_name(®ion->mmaps[index].mem),
> diff --git a/hw/vfio/vfio-region.h b/hw/vfio/vfio-region.h
> index 58b236f113..8d5699013a 100644
> --- a/hw/vfio/vfio-region.h
> +++ b/hw/vfio/vfio-region.h
> @@ -39,6 +39,9 @@ uint64_t vfio_region_read(void *opaque,
> hwaddr addr, unsigned size);
> int vfio_region_setup(Object *obj, VFIODevice *vbasedev, VFIORegion *region,
> int index, const char *name, Error **errp);
> +int vfio_region_setup_with_ops(Object *obj, VFIODevice *vbasedev,
> + VFIORegion *region, int index, const char *name,
> + const MemoryRegionOps *ops, Error **errp);
> int vfio_region_mmap(VFIORegion *region);
> void vfio_region_mmaps_set_enabled(VFIORegion *region, bool enabled);
> void vfio_region_exit(VFIORegion *region);
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v2 05/10] hw/vfio/pci: Back the CXL memory with a RAM-device region
2026-09-16 18:44 ` [PATCH v2 05/10] hw/vfio/pci: Back the CXL memory with a RAM-device region mhonap
@ 2026-09-17 13:56 ` Cédric Le Goater
2026-09-21 10:55 ` Manish Honap
0 siblings, 1 reply; 27+ messages in thread
From: Cédric Le Goater @ 2026-09-17 13:56 UTC (permalink / raw)
To: mhonap, alex, ankita, jic23, dave.jiang, alejandro.lucero-palau,
smadhavan, pierrick.bouvier, mst, imammedo, anisinha, pbonzini,
eric.auger, peter.maydell, richard.henderson, cohuck
Cc: kjaju, vsethi, zhiw, qemu-devel, qemu-arm, linux-cxl
On 9/16/26 20:44, mhonap@nvidia.com wrote:
> From: Manish Honap <mhonap@nvidia.com>
>
> mmap the kernel's HDM memory region so its host physical pages back a
> RAM-device MemoryRegion. It stays out of the guest address space until
> the guest commits its decoder and QEMU learns the GPA to place it at.
>
> AI-used-for: code (prototype)
> Signed-off-by: Manish Honap <mhonap@nvidia.com>
> ---
> hw/vfio/pci.c | 34 ++++++++++++++++++++++++++++++++++
> hw/vfio/pci.h | 1 +
> 2 files changed, 35 insertions(+)
>
> diff --git a/hw/vfio/pci.c b/hw/vfio/pci.c
> index 10b13f200d..4716266595 100644
> --- a/hw/vfio/pci.c
> +++ b/hw/vfio/pci.c
> @@ -3221,8 +3221,11 @@ bool vfio_pci_populate_device(VFIOPCIDevice *vdev, Error **errp)
> return true;
> }
>
> +static void vfio_cxl_teardown(VFIOPCIDevice *vdev);
> +
> void vfio_pci_put_device(VFIOPCIDevice *vdev)
> {
> + vfio_cxl_teardown(vdev);
> vfio_display_finalize(vdev);
> vfio_bars_finalize(vdev);
> vfio_cpr_pci_unregister_device(vdev);
> @@ -3696,11 +3699,42 @@ static bool vfio_cxl_setup(VFIOPCIDevice *vdev, Error **errp)
> return false;
> }
>
> + /*
> + * The HDM memory is host physical. Set up the region, which installs the
> + * fd read/write path, and mmap it for direct guest access; it is added to
> + * the guest address space only once the guest commits its endpoint decoder.
> + */
> + if (vfio_region_setup(OBJECT(vdev), vbasedev, &cxl->mem_region,
> + cxl->mem_region_index, "cxl-mem", errp)) {
> + return false;
> + }
> + if (vfio_region_mmap(&cxl->mem_region)) {
> + /*
> + * Without mmap the region falls back to the kernel's fd read/write
> + * path, which works but traps every access. Warn rather than fail.
> + */
> + warn_report("vfio-cxl: %s: failed to mmap the HDM memory region; "
> + "performance may be slow", vbasedev->name);
> + }
vfio_cxl_setup calls vfio_region_setup() then vfio_region_mmap()
back-to-back when other 'normal' BARs split these calls across
vfio_populate_device() and vfio_bar_register(). I guess it is fine
for CXL since it is not a PCI BAR.
> cxl->enabled = true;
>
> return true;
> }
>
> +static void vfio_cxl_teardown(VFIOPCIDevice *vdev)
> +{
> + VFIOCXL *cxl = &vdev->cxl;
> +
> + if (!cxl->enabled) {
> + return;
> + }
> + if (cxl->mem_region.mem) {
> + vfio_region_exit(&cxl->mem_region);
> + vfio_region_finalize(&cxl->mem_region);
However, the tear down should be split between :
1. vfio_exitfn (unrealize)
drops mmaps, removes subregions and removes references so the
MR refcount can reach zero.
2. vfio_pci_finalize (instance_finalize)
frees the MemoryRegion.
Thanks,
C.
> + }
> +}
> +
> static void vfio_pci_realize(PCIDevice *pdev, Error **errp)
> {
> ERRP_GUARD();
> diff --git a/hw/vfio/pci.h b/hw/vfio/pci.h
> index 7fdd695704..06d15e807d 100644
> --- a/hw/vfio/pci.h
> +++ b/hw/vfio/pci.h
> @@ -134,6 +134,7 @@ typedef struct VFIOCXL {
> uint32_t comp_bar; /* component BAR carrying that block */
> uint64_t hdm_offset; /* block offset within the component BAR */
> uint64_t dpa_size; /* size of the HDM memory region */
> + VFIORegion mem_region; /* HDM memory, mapped at committed GPA */
> } VFIOCXL;
>
> struct VFIOPCIDevice {
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v2 06/10] hw/vfio/pci: Bind a CXL device to its fixed memory window
2026-09-16 18:44 ` [PATCH v2 06/10] hw/vfio/pci: Bind a CXL device to its fixed memory window mhonap
@ 2026-09-17 16:00 ` Cédric Le Goater
2026-09-21 11:03 ` Manish Honap
0 siblings, 1 reply; 27+ messages in thread
From: Cédric Le Goater @ 2026-09-17 16:00 UTC (permalink / raw)
To: mhonap, alex, ankita, jic23, dave.jiang, alejandro.lucero-palau,
smadhavan, pierrick.bouvier, mst, imammedo, anisinha, pbonzini,
eric.auger, peter.maydell, richard.henderson, cohuck,
Nitesh Narayan Lal
Cc: kjaju, vsethi, zhiw, qemu-devel, qemu-arm, linux-cxl
On 9/16/26 20:44, mhonap@nvidia.com wrote:
> From: Manish Honap <mhonap@nvidia.com>
>
> The guest programs a GPA inside a CXL fixed memory window and QEMU maps
> the HDM memory there, so the device must sit under exactly one
> single-target CFMWS.
CFMWS (CXL Fixed Memory Window Structure). Pity that the fields are named
fmws_base, fmws_size.
> Reject any ambiguous, interleaved, unsized, or
> missing window rather than guess. Match windows by target name, since the
> resolved target pointers are only filled in by a machine-done notifier
> that may run after this one.
>
> Require exactly one endpoint function below the root port: two functions
> in the same slot would each match that single-target CFMWS and alias their
> HDM memory at one base, so count every present function rather than one per
> slot. The CFMWS windows are placed at machine-init-done, so a cold-plugged
> device can only be validated from that notifier, where a bad topology is
> fatal to startup. Validate a hotplugged device inline from realize instead,
> reporting through errp so a bad device_add fails cleanly rather than
> aborting the running VM, and arm the notifier only for a cold-plugged
> device.
>
> The endpoint's HDM size comes from the kernel region. cxl_fmws_set_memmap()
> places the window in the guest PA map before this device is realized, so
> the window cannot yet be sized from the device; check the configured window
> against the device geometry instead. Reject a window too small to hold the
> endpoint and name the size to set, and warn on an oversized window, since
> the padding becomes guest CEDT that migration has to preserve.
I am lost ... Too much stuff there.
Each paragraph introduces a new topic/problem, and implementation details,
all mixed up, without first explaining what problem is being solved:
A passed-through CXL Type-2 device needs to know where its memory
window lives in the guest PA space so QEMU can map the device memory
there.
Is that it ?
A maintainer reviewing this patch shouldn't need to know what a CFMWS
is to understand the commit. CXL is still an emerging technology.
Most VFIO and QEMU reviewers will not have CXL spec background and
CXL-specific mechanic knowledge.
> AI-used-for: code (prototype)
The code still reads like a proof-of-concept "prototype". This patch
in particular is very hard to follow. The problem isn't clearly described,
and likely isn't fully understood yet. The series should be broken into
smaller patches.
> Signed-off-by: Manish Honap <mhonap@nvidia.com>
> ---
> hw/pci-bridge/pci_expander_bridge_stubs.c | 6 +
> hw/vfio/pci.c | 192 ++++++++++++++++++++++
> hw/vfio/pci.h | 3 +
> 3 files changed, 201 insertions(+)
>
> diff --git a/hw/pci-bridge/pci_expander_bridge_stubs.c b/hw/pci-bridge/pci_expander_bridge_stubs.c
> index b35180311f..c44ad7fab6 100644
> --- a/hw/pci-bridge/pci_expander_bridge_stubs.c
> +++ b/hw/pci-bridge/pci_expander_bridge_stubs.c
> @@ -10,5 +10,11 @@
> #include "hw/pci/pci_bus.h"
> #include "hw/pci-bridge/pci_expander_bridge.h"
> #include "hw/cxl/cxl.h"
> +#include "hw/cxl/cxl_component.h"
>
> void pxb_cxl_hook_up_registers(CXLState *state, PCIBus *bus, Error **errp) {};
> +
> +bool cxl_get_hb_passthrough(PCIHostState *hb)
> +{
> + return false;
> +}
> diff --git a/hw/vfio/pci.c b/hw/vfio/pci.c
> index 4716266595..670e0d1da4 100644
> --- a/hw/vfio/pci.c
> +++ b/hw/vfio/pci.c
> @@ -27,7 +27,12 @@
> #include "hw/pci/msi.h"
> #include "hw/pci/msix.h"
> #include "hw/pci/pci_bridge.h"
> +#include "hw/pci/pci_host.h"
> +#include "hw/pci/pcie_port.h"
> #include "hw/cxl/cxl.h"
> +#include "hw/cxl/cxl_host.h"
> +#include "hw/cxl/cxl_component.h"
> +#include "system/system.h"
> #include "hw/core/qdev-properties.h"
> #include "hw/core/qdev-properties-system.h"
> #include "hw/vfio/vfio-cpr.h"
> @@ -3641,6 +3646,171 @@ static bool vfio_cxl_check_topology(VFIOPCIDevice *vdev, Error **errp)
> return true;
> }
>
> +/*
> + * Count the CXL fixed memory windows that target this device's pxb-cxl and
> + * report the matched window's base, size and target count. Match on the target
> + * names: cxl_fmws_link_targets() resolves target_hbs[] only at machine_done,
> + * which may run after this, so the resolved pointers can still be NULL here.
> + */
> +static int vfio_cxl_match_fmws(PXBCXLDev *pxb, hwaddr *base, uint64_t *size,
> + int *nwindows, int *ntargets)
I think this should be in the CXL subsystem. That's a *lot* of out
parameters ...
> +{
> + GSList *list = cxl_fmws_get_all_sorted();
> + GSList *iter;
> + int matches = 0;
> +
> + *nwindows = g_slist_length(list);
> + *base = 0;
> + *size = 0;
> + *ntargets = 0;
> +
> + for (iter = list; iter; iter = iter->next) {
> + CXLFixedWindow *fw = CXL_FMW(iter->data);
> + int i;
> +
> + for (i = 0; i < fw->num_targets; i++) {
> + bool ambiguous = false;
> + Object *t = object_resolve_path_type(fw->targets[i],
> + TYPE_PXB_CXL_DEV, &ambiguous);
> +
> + if (t && !ambiguous && PXB_CXL_DEV(t) == pxb) {
> + *base = fw->base;
> + *size = fw->size;
> + *ntargets = fw->num_targets;
> + matches++;
> + break;
> + }
> + }
> + }
> + g_slist_free(list);
> + return matches;
> +}
> +
> +/*
> + * Fix the bounds of the device's memory window once the topology has settled.
> + * Exactly one single-target CFMWS, with an assigned base and room for the HDM
> + * memory, is required; anything else is a misconfiguration the guest cannot
> + * recover from, so reject it rather than guess a window.
> + */
> +static bool vfio_cxl_do_bind_fmws(VFIOPCIDevice *vdev, Error **errp)
> +{
This does too much :
topology validation
CFMWS window matching
size validation
It looks like vfio_cxl_check_topology() and it should be in the CXL subsystem.
> + VFIOCXL *cxl = &vdev->cxl;
> + const char *name = vdev->vbasedev.name;
> + int nwindows = 0, ntargets = 0, matches;
> + hwaddr base = 0;
> + uint64_t size = 0, need;
> + PXBCXLDev *pxb;
> +
> + pxb = vfio_cxl_find_pxb(vdev, errp);
> + if (!pxb) {
> + return false;
> + }
> +
> + if (pcie_count_ds_ports(PCI_HOST_BRIDGE(pxb->cxl_host_bridge)->bus) != 1) {
This won't build on all platforms.
> + error_setg(errp,
> + "vfio-cxl: %s: pxb-cxl not in HDM passthrough mode "
> + "(use a single cxl-rp)", name);
> + return false;
> + }
> +
> + /*
> + * Count every present function, not one per slot, and require exactly one
> + * endpoint below the root port.
> + */> + {
> + PCIBus *ep_bus = pci_get_bus(&vdev->parent_obj);
> + int slot, fn, nendpoints = 0;
This smells like a function.
> + for (slot = 0; slot < PCI_SLOT_MAX; slot++) {
> + for (fn = 0; fn < PCI_FUNC_MAX; fn++) {
> + if (ep_bus->devices[PCI_DEVFN(slot, fn)]) {
> + nendpoints++;
> + }
> + }
> + }
> + if (nendpoints != 1) {
> + error_setg(errp,
> + "vfio-cxl: %s: %d devices below the cxl-rp; a passed "
> + "through CXL endpoint must be alone below its root port",
> + name, nendpoints);
> + return false;
> + }
> + }
> +
> + matches = vfio_cxl_match_fmws(pxb, &base, &size, &nwindows, &ntargets);
> + if (matches == 1 && ntargets == 1) {
> + if (!base) {
> + error_setg(errp,
> + "vfio-cxl: %s: matched CFMWS has no base; reduce its "
> + "size or grow the guest PA space", name);
> + return false;
> + }
> + /*
> + * QEMU reads the endpoint's HDM size from the kernel region, so the
> + * window size is really device information, not a value to make the
> + * caller guess. Sizing the CFMWS from the device the way firmware does
> + * is not possible here: the window is placed in the guest PA map by
> + * cxl_fmws_set_memmap() before this device is realized, so its size is
> + * fixed before dpa_size is known. Until a core change can size the
> + * window from the device, take the device geometry as the source of
> + * truth: reject a window too small to hold the endpoint, and warn when
> + * it is larger than needed, since the padding becomes guest CEDT that
> + * migration has to preserve.
> + */
> + need = ROUND_UP(cxl->dpa_size, 256 * MiB);
> + if (size < need) {
> + error_setg(errp,
> + "vfio-cxl: %s: CFMWS size 0x%" PRIx64 " cannot hold the "
> + "endpoint; set the cxl-fmw size to 0x%" PRIx64,
> + name, size, need);
> + return false;
> + }
> + if (size > need) {
> + warn_report("vfio-cxl: %s: CFMWS size 0x%" PRIx64 " exceeds the "
> + "endpoint's 0x%" PRIx64 "; set the cxl-fmw size to 0x%"
> + PRIx64 " to keep the guest memory map stable across "
> + "migration", name, size, cxl->dpa_size, need);
> + }
> + cxl->fmws_base = base;
> + cxl->fmws_size = size;
> + return true;
> + }
> +
> + if (matches == 1) {
> + error_setg(errp,
> + "vfio-cxl: %s: its CFMWS interleaves %d targets; use a "
> + "single-target window", name, ntargets);
> + } else if (matches > 1) {
> + error_setg(errp,
> + "vfio-cxl: %s: pxb-cxl is targeted by %d CFMWS; use one",
> + name, matches);
> + } else {
> + error_setg(errp,
> + "vfio-cxl: %s: no CFMWS targets this device (%d present)",
> + name, nwindows);
> + }
> + return false;
> +}
> +
> +/*
> + * Cold-plug path: the CFMWS windows are placed at machine init done, so the
are you sure of that ? I think fw->base is already set when vfio_cxl_setup()
runs for cold-plug.
Normal vfio-pci has no hotplug-specific code, vfio_pci_realize() runs
the same path regardless of DEVICE(vdev)->hotplugged. The CXL case
should be the same.
> + * binding can only be validated from this notifier. The VM has not run yet, so
> + * a configuration error is fatal to startup. A hotplugged device is validated
> + * in realize instead (see vfio_cxl_setup), where the failure fails device_add
> + * without taking down the running VM.
> + */
> +static void vfio_cxl_bind_fmws(Notifier *n, void *data)
> +{
> + VFIOCXL *cxl = container_of(n, VFIOCXL, machine_done);
> + VFIOPCIDevice *vdev = container_of(cxl, VFIOPCIDevice, cxl);
> + Error *err = NULL;
> +
> + if (!vfio_cxl_do_bind_fmws(vdev, &err)) {
> + error_report_err(err);
> + exit(1);
> + }
> +}
> +
> /*
> * Learn the CXL geometry the kernel reports: the HPA-backed HDM memory region
> * and the trapped HDM decoder block (which BAR carries it and at what offset).
> @@ -3717,6 +3887,22 @@ static bool vfio_cxl_setup(VFIOPCIDevice *vdev, Error **errp)
> "performance may be slow", vbasedev->name);
> }
>
> + if (DEVICE(vdev)->hotplugged) {
> + /*
> + * The machine is already up, so the CFMWS windows are placed and the
> + * binding can be validated now. Report a failure through errp so a bad
> + * device_add fails cleanly instead of aborting the running VM.
> + */
> + if (!vfio_cxl_do_bind_fmws(vdev, errp)) {
> + vfio_region_exit(&cxl->mem_region);
> + vfio_region_finalize(&cxl->mem_region);
as said before, the tear down has 2 parts. Anyhow, I don't think we need a
CXL-specific hotplug patch.
C.
> + return false;
> + }
> + } else {
> + cxl->machine_done.notify = vfio_cxl_bind_fmws;
> + qemu_add_machine_init_done_notifier(&cxl->machine_done);
> + }
> +
> cxl->enabled = true;
>
> return true;
> @@ -3729,6 +3915,12 @@ static void vfio_cxl_teardown(VFIOPCIDevice *vdev)
> if (!cxl->enabled) {
> return;
> }
> +
> + if (cxl->machine_done.notify) {
> + qemu_remove_machine_init_done_notifier(&cxl->machine_done);
> + cxl->machine_done.notify = NULL;
> + }
> +
> if (cxl->mem_region.mem) {
> vfio_region_exit(&cxl->mem_region);
> vfio_region_finalize(&cxl->mem_region);
> diff --git a/hw/vfio/pci.h b/hw/vfio/pci.h
> index 06d15e807d..22fe8d7ff7 100644
> --- a/hw/vfio/pci.h
> +++ b/hw/vfio/pci.h
> @@ -135,6 +135,9 @@ typedef struct VFIOCXL {
> uint64_t hdm_offset; /* block offset within the component BAR */
> uint64_t dpa_size; /* size of the HDM memory region */
> VFIORegion mem_region; /* HDM memory, mapped at committed GPA */
> + Notifier machine_done; /* CFMWS validated at machine_done */
> + hwaddr fmws_base; /* base of the memory window */
> + uint64_t fmws_size; /* size of the memory window */
> } VFIOCXL;
>
> struct VFIOPCIDevice {
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v2 03/10] hw/vfio/pci: Detect a CXL Type-2 device and read its geometry
2026-09-16 18:44 ` [PATCH v2 03/10] hw/vfio/pci: Detect a CXL Type-2 device and read its geometry mhonap
@ 2026-09-17 16:04 ` Cédric Le Goater
2026-09-21 10:49 ` Manish Honap
0 siblings, 1 reply; 27+ messages in thread
From: Cédric Le Goater @ 2026-09-17 16:04 UTC (permalink / raw)
To: mhonap, alex, ankita, jic23, dave.jiang, alejandro.lucero-palau,
smadhavan, pierrick.bouvier, mst, imammedo, anisinha, pbonzini,
eric.auger, peter.maydell, richard.henderson, cohuck
Cc: kjaju, vsethi, zhiw, qemu-devel, qemu-arm, linux-cxl
On 9/16/26 20:44, mhonap@nvidia.com wrote:
> From: Manish Honap <mhonap@nvidia.com>
>
> The kernel marks a passthroughed CXL Type-2 device with a device flag and
> exposes two regions:
> - HDM memory (host physical)
> - the trapped HDM decoder register block.
>
> Read the flag, locate both regions by type, and take the component BAR
> and block offset from the geometry capability. Realize only records this;
> later patches build the guest mapping on top.
>
> AI-used-for: code (prototype)
> Signed-off-by: Manish Honap <mhonap@nvidia.com>
LGTM.
Thanks,
C.
> ---
> hw/vfio/pci.c | 72 +++++++++++++++++++++++++++++++++++++++++++++++++++
> hw/vfio/pci.h | 15 +++++++++++
> 2 files changed, 87 insertions(+)
>
> diff --git a/hw/vfio/pci.c b/hw/vfio/pci.c
> index 428ab2f069..2f84af5cf8 100644
> --- a/hw/vfio/pci.c
> +++ b/hw/vfio/pci.c
> @@ -3570,6 +3570,64 @@ bool vfio_pci_interrupt_setup(VFIOPCIDevice *vdev, Error **errp)
> return true;
> }
>
> +/*
> + * Learn the CXL geometry the kernel reports: the HPA-backed HDM memory region
> + * and the trapped HDM decoder block (which BAR carries it and at what offset).
> + * A non-CXL device leaves cxl.enabled false and takes no CXL paths.
> + */
> +static bool vfio_cxl_setup(VFIOPCIDevice *vdev, Error **errp)
> +{
> + VFIODevice *vbasedev = &vdev->vbasedev;
> + VFIOCXL *cxl = &vdev->cxl;
> + struct vfio_region_info *mem_info = NULL, *comp_info = NULL;
> + struct vfio_region_info_cap_cxl_comp_regs *cap;
> + struct vfio_info_cap_header *hdr;
> +
> + if (!(vbasedev->flags & VFIO_DEVICE_FLAGS_CXL)) {
> + return true;
> + }
> +
> + if (vfio_device_get_region_info_type(vbasedev,
> + VFIO_REGION_TYPE_PCI_VENDOR_TYPE |
> + CXL_VENDOR_ID,
> + VFIO_REGION_SUBTYPE_CXL_MEM,
> + &mem_info)) {
> + error_setg(errp, "vfio-cxl: %s: CXL memory region not found",
> + vbasedev->name);
> + return false;
> + }
> +
> + if (vfio_device_get_region_info_type(vbasedev,
> + VFIO_REGION_TYPE_PCI_VENDOR_TYPE |
> + CXL_VENDOR_ID,
> + VFIO_REGION_SUBTYPE_CXL_COMP_REGS,
> + &comp_info)) {
> + error_setg(errp,
> + "vfio-cxl: %s: CXL component-register region not found",
> + vbasedev->name);
> + return false;
> + }
> +
> + hdr = vfio_get_region_info_cap(comp_info,
> + VFIO_REGION_INFO_CAP_CXL_COMP_REGS);
> + if (!hdr) {
> + error_setg(errp,
> + "vfio-cxl: %s: component-register geometry not reported",
> + vbasedev->name);
> + return false;
> + }
> + cap = container_of(hdr, struct vfio_region_info_cap_cxl_comp_regs, header);
> +
> + cxl->mem_region_index = mem_info->index;
> + cxl->comp_regs_region_index = comp_info->index;
> + cxl->dpa_size = mem_info->size;
> + cxl->comp_bar = cap->bar;
> + cxl->hdm_offset = cap->offset;
> + cxl->enabled = true;
> +
> + return true;
> +}
> +
> static void vfio_pci_realize(PCIDevice *pdev, Error **errp)
> {
> ERRP_GUARD();
> @@ -3700,6 +3758,20 @@ static void vfio_pci_realize(PCIDevice *pdev, Error **errp)
> }
> }
>
> + if (!vfio_cxl_setup(vdev, errp)) {
> + /*
> + * vfio_migration_realize() above installed a migration blocker in auto
> + * mode (generic vfio-pci exposes no migration ops). out_deregister does
> + * not remove it, so a rejected CXL setup would leave VM migration
> + * blocked until QEMU restarts. Drop it here, under the same
> + * failover-pair condition used to install it.
> + */
> + if (!pdev->failover_pair_id) {
> + vfio_migration_exit(vbasedev);
> + }
> + goto out_deregister;
> + }
> +
> vfio_pci_register_err_notifier(vdev);
> vfio_pci_register_req_notifier(vdev);
> vfio_setup_resetfn_quirk(vdev);
> diff --git a/hw/vfio/pci.h b/hw/vfio/pci.h
> index c9ab949870..7fdd695704 100644
> --- a/hw/vfio/pci.h
> +++ b/hw/vfio/pci.h
> @@ -122,10 +122,25 @@ typedef struct VFIOMSIXInfo {
>
> OBJECT_DECLARE_SIMPLE_TYPE(VFIOPCIDevice, VFIO_PCI_DEVICE)
>
> +/*
> + * State for a CXL Type-2 device. The kernel owns the host physical placement
> + * of the device memory; QEMU only maps it at the guest physical address the
> + * guest commits into its endpoint HDM decoder.
> + */
> +typedef struct VFIOCXL {
> + bool enabled;
> + uint32_t mem_region_index; /* HPA-backed HDM memory VFIO region */
> + uint32_t comp_regs_region_index; /* trapped HDM decoder register block */
> + uint32_t comp_bar; /* component BAR carrying that block */
> + uint64_t hdm_offset; /* block offset within the component BAR */
> + uint64_t dpa_size; /* size of the HDM memory region */
> +} VFIOCXL;
> +
> struct VFIOPCIDevice {
> PCIDevice parent_obj;
>
> VFIODevice vbasedev;
> + VFIOCXL cxl;
> VFIOINTx intx;
> unsigned int config_size;
> uint8_t *emulated_config_bits; /* QEMU emulated bits, little-endian */
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v2 07/10] hw/vfio/pci: Map the CXL memory on the guest decoder commit
2026-09-16 18:44 ` [PATCH v2 07/10] hw/vfio/pci: Map the CXL memory on the guest decoder commit mhonap
@ 2026-09-18 16:08 ` Cédric Le Goater
2026-09-21 11:14 ` Manish Honap
2026-09-20 9:14 ` Junjie Cao
1 sibling, 1 reply; 27+ messages in thread
From: Cédric Le Goater @ 2026-09-18 16:08 UTC (permalink / raw)
To: mhonap, alex, ankita, jic23, dave.jiang, alejandro.lucero-palau,
smadhavan, pierrick.bouvier, mst, imammedo, anisinha, pbonzini,
eric.auger, peter.maydell, richard.henderson, cohuck
Cc: kjaju, vsethi, zhiw, qemu-devel, qemu-arm, linux-cxl
On 9/16/26 20:44, mhonap@nvidia.com wrote:
> From: Manish Honap <mhonap@nvidia.com>
>
> Overlay the trapped HDM decoder block over the component BAR and forward
> its accesses to the kernel. QEMU maps the HDM memory at the base of the
> device's CFMWS window and presents that same base back to the guest
> through the trapped decoder read
>
> Overlap the CFMWS window MemoryRegion the CXL host bridge maps at that
> base, at a higher priority, so precedence is defined by priority rather
> than subregion insertion order.
>
> QEMU always maps at the CFMWS base, so a decoder the guest commits at any
> other base would leave the guest's view and the mapping diverged. Track
> the base the guest writes to the trapped block and refuse to map a commit
> at a different base; a firmware-committed decoder the guest never
> reprograms keeps the window base.
>
> AI-used-for: code (prototype)
> Signed-off-by: Manish Honap <mhonap@nvidia.com>
> ---
> hw/cxl/cxl-host-stubs.c | 5 +
> hw/vfio/pci.c | 386 +++++++++++++++++++++++++++++++++++++++-
> hw/vfio/pci.h | 8 +
> 3 files changed, 393 insertions(+), 6 deletions(-)
This patch introduces too many concepts at once, a reader familiar
with VFIO can't separate the plumbing from the CXL details. This
needs more commits.
So, something like
1. framework
Set up the comp_regs_region
fd pass-through (pread/pwrite), no virtualization.
teardown
2. Virtualization of decoder base registers
3. Map/unmap of HDM memory
if fmws_base never changes, there's no reason to remove and
re-add the subregion. Install it once and toggle with
memory_region_set_enabled(). this should simplify the decoder
handling.
4. Config write hooks
PCI_COMMAND, CXL reset, PM transition, FLR x 2
each could be a separate patch I think
I don't think we need the machine notifier as said in the previous
patch.
More below,
>
> diff --git a/hw/cxl/cxl-host-stubs.c b/hw/cxl/cxl-host-stubs.c
> index 9b515913ea..1067c944e2 100644
> --- a/hw/cxl/cxl-host-stubs.c
> +++ b/hw/cxl/cxl-host-stubs.c
> @@ -23,3 +23,8 @@ GSList *cxl_fmws_get_all_sorted(void)
> {
> g_assert_not_reached();
> }
> +
> +int cxl_decoder_count_dec(int enc_cnt)
> +{
> + g_assert_not_reached();
> +}
> diff --git a/hw/vfio/pci.c b/hw/vfio/pci.c
> index 670e0d1da4..fb3c39d4c6 100644
> --- a/hw/vfio/pci.c
> +++ b/hw/vfio/pci.c
> @@ -1453,6 +1453,39 @@ uint32_t vfio_pci_read_config(PCIDevice *pdev, uint32_t addr, int len)
> return val;
> }
>
> +static void vfio_cxl_decoder_changed(VFIOPCIDevice *vdev);
> +static void vfio_cxl_unmap_mem(VFIOPCIDevice *vdev);
> +
> +/*
> + * Return true if this config write sets a Function Level Reset bit the kernel
> + * acts on: the PCIe Device Control FLR bit or the Advanced Features FLR bit.
> + * FLR bits are write-1-to-trigger and self-clearing, so inspect the written
> + * value rather than the post-write config.
> + */
> +static bool vfio_cxl_flr_write(VFIOPCIDevice *vdev, uint32_t addr,
> + uint32_t val, int len)
> +{
> + PCIDevice *pdev = &vdev->parent_obj;
> + uint32_t off;
> +
> + if (pdev->exp.exp_cap) {
> + /* PCI_EXP_DEVCTL_BCR_FLR (bit 15) is the high byte of Device Ctrl. */
> + off = pdev->exp.exp_cap + PCI_EXP_DEVCTL + 1;
> + if (off >= addr && off < addr + len &&
> + ((val >> (8 * (off - addr))) & (PCI_EXP_DEVCTL_BCR_FLR >> 8))) {
> + return true;
> + }
> + }
> + if (vdev->cxl.af_offset) {
> + off = vdev->cxl.af_offset + PCI_AF_CTRL;
> + if (off >= addr && off < addr + len &&
> + ((val >> (8 * (off - addr))) & PCI_AF_CTRL_FLR)) {
> + return true;
> + }
> + }
> + return false;
> +}
> +
> void vfio_pci_write_config(PCIDevice *pdev,
> uint32_t addr, uint32_t val, int len)
> {
> @@ -1522,9 +1555,74 @@ void vfio_pci_write_config(PCIDevice *pdev,
> vfio_sub_page_bar_update_mapping(pdev, bar);
> }
> }
> +
> + /*
> + * A firmware-committed CXL decoder is already committed at boot, so no
> + * guest control write triggers the HDM mapping. Map it once the guest
> + * enables memory decoding, so the region enters the guest address space
> + * and the IOAS while the device is live. On a clear of memory decoding,
> + * withdraw the overlay: the kernel revokes the HDM PTEs on the same
> + * write, so a retained memslot would fault a guest access during the
> + * disabled interval onto a zapped VMA and stop the VM; the enable path
> + * re-maps it.
> + */
> + if (vdev->cxl.enabled &&
> + range_covers_byte(addr, len, PCI_COMMAND)) {
> + if (pci_get_word(pdev->config + PCI_COMMAND) & PCI_COMMAND_MEMORY) {
> + vfio_cxl_decoder_changed(vdev);
> + } else {
> + vfio_cxl_unmap_mem(vdev);
> + }
> + }
> } else {
> /* Write everything to QEMU to keep emulated bits correct */
> pci_default_write_config(pdev, addr, val, len);
> +
> + /*
> + * A guest CXL reset is a write to the CXL Device DVSEC ctrl2 register,
> + * forwarded to the kernel above, which re-commits this firmware-fixed
> + * decoder. If the guest decommitted the decoder before the reset, which
> + * unmapped the HDM memory, that re-commit is not otherwise visible to
> + * QEMU, so rescan the decoder here to restore the mapping. Gated on
> + * memory decoding still being enabled, like the enable path above.
> + */
> + if (vdev->cxl.enabled && vdev->cxl.dvsec_offset &&
> + ranges_overlap(addr, len,
> + vdev->cxl.dvsec_offset +
> + offsetof(CXLDVSECDevice, ctrl2),
> + sizeof_field(CXLDVSECDevice, ctrl2)) &&
> + (pci_get_word(pdev->config + PCI_COMMAND) & PCI_COMMAND_MEMORY)) {
> + vfio_cxl_decoder_changed(vdev);
> + }
> +
> + /*
> + * A NoSoftRst- device the guest cycles D3hot->D0 has its physical
> + * decoder restored by the kernel on resume, which re-commits this
> + * firmware-fixed decoder just like a reset. That re-commit is not
> + * otherwise visible to QEMU, so on the transition back to D0 rescan
> + * the decoder to restore a mapping the guest dropped before it
> + * suspended. Gated on memory decoding, like the paths above.
> + */
> + if (vdev->cxl.enabled && pdev->pm_cap &&
> + range_covers_byte(addr, len, pdev->pm_cap + PCI_PM_CTRL) &&
> + (pci_get_word(pdev->config + pdev->pm_cap + PCI_PM_CTRL) &
> + PCI_PM_CTRL_STATE_MASK) == 0 &&
> + (pci_get_word(pdev->config + PCI_COMMAND) & PCI_COMMAND_MEMORY)) {
> + vfio_cxl_decoder_changed(vdev);
> + }
> +
> + /*
> + * A Function Level Reset the guest triggers through the PCIe Device
> + * Control or Advanced Features FLR bit is forwarded to the kernel,
> + * which re-commits this firmware-fixed decoder just like a CXL reset.
> + * That re-commit is not otherwise visible to QEMU, so rescan the
> + * decoder to restore a mapping the guest dropped before the FLR. Gated
> + * on memory decoding, like the paths above.
> + */
> + if (vdev->cxl.enabled && vfio_cxl_flr_write(vdev, addr, val, len) &&
> + (pci_get_word(pdev->config + PCI_COMMAND) & PCI_COMMAND_MEMORY)) {
> + vfio_cxl_decoder_changed(vdev);
> + }
> }
> }
That's a lot of CXL-specific code injected into vfio_pci_write_config.
Don't do that. Please add :
if (vdev->cxl.enabled) {
vfio_cxl_config_written(vdev, addr, val, len);
}
I think that we should have a vfio/cxl.c subcomponent to isolate the
CXL-specific code :
vfio_cxl_setup, or better vfio_cxl_realize
vfio_cxl_config_written,
vfio_cxl_exit,
vfio_cxl_finalize,
etc.
> @@ -3811,6 +3909,239 @@ static void vfio_cxl_bind_fmws(Notifier *n, void *data)
> }
> }
>
> +/*
> + * HDM decoder registers, relative to the trapped decoder block. The block holds
> + * the HDM Decoder Capability register followed by one register set per decoder,
> + * each VFIO_CXL_HDM_DECODER_STRIDE apart, indexed by decoder.
> + */
> +#define VFIO_CXL_HDM_CAP 0x00
> +#define VFIO_CXL_HDM_DECODER_STRIDE 0x20
> +#define VFIO_CXL_HDM_DECODER_BASE_LOW(n) \
> + (0x10 + (n) * VFIO_CXL_HDM_DECODER_STRIDE)
> +#define VFIO_CXL_HDM_DECODER_BASE_HIGH(n) \
> + (0x14 + (n) * VFIO_CXL_HDM_DECODER_STRIDE)
> +#define VFIO_CXL_HDM_DECODER_CTRL(n) \
> + (0x20 + (n) * VFIO_CXL_HDM_DECODER_STRIDE)
> +#define VFIO_CXL_HDM_CTRL_COMMITTED (1 << 10)
> +#define VFIO_CXL_HDM_BASE_LOW_MASK 0xf0000000U
> +
> +static void vfio_cxl_unmap_mem(VFIOPCIDevice *vdev)
> +{
> + VFIOCXL *cxl = &vdev->cxl;
> +
> + if (!cxl->dpa_mapped) {
> + return;
> + }
> + memory_region_del_subregion(get_system_memory(), cxl->mem_region.mem);
> + cxl->dpa_mapped = false;
> + cxl->mapped_base = 0;
> +}
> +
> +/*
> + * The HDM Decoder Capability register encodes the decoder count in its low
> + * nibble (CXL r3.1 8.2.4.20.1, encodings 0h..Ch for 1..32). Decode it with the
> + * shared cxl_decoder_count_dec() helper so the commit handler walks every
> + * decoder rather than assuming decoder 0. An unreadable register or a reserved
> + * encoding (which the helper decodes to 0) is treated as a single decoder.
> + */
> +static unsigned vfio_cxl_decoder_count(VFIOPCIDevice *vdev)
> +{
> + off_t off = vdev->cxl.comp_regs_region.fd_offset;
> + uint32_t cap = 0;
> + int count;
> +
> + if (pread(vdev->vbasedev.fd, &cap, 4, off + VFIO_CXL_HDM_CAP) != 4) {
> + return 1;
> + }
> + count = cxl_decoder_count_dec(le32_to_cpu(cap) & 0xf);
> + return count > 0 ? count : 1;
> +}
> +
> +/*
> + * The guest committed or tore down an endpoint decoder. Walk each decoder in
> + * the trapped HDM block and (un)map the HDM memory at the CFMWS window base,
> + * not the base the guest programmed (see vfio_cxl_comp_regs_read). This
> + * generation commits one non-interleaved decoder, so the walk stops at the
> + * first committed decoder.
if only one decoder is supported and the code tracks the guest base
per-device rather than per-decoder, why do the read/write handlers
use VFIO_CXL_HDM_DECODER_BASE_LOW(0) to match all decoders ? This is
inconsistent.
I think we should record the active decoder index 'active_decoder'
instead. If a second committed decoder is found, decoder_changed
should bail out.
> + */
> +static void vfio_cxl_decoder_changed(VFIOPCIDevice *vdev)
> +{
> + VFIOCXL *cxl = &vdev->cxl;
> + VFIODevice *vbasedev = &vdev->vbasedev;
> + off_t off = cxl->comp_regs_region.fd_offset;
> + unsigned n, count = vfio_cxl_decoder_count(vdev);
> +
> + for (n = 0; n < count; n++) {
> + uint32_t ctrl = 0;
> +
> + if (pread(vbasedev->fd, &ctrl, 4,
> + off + VFIO_CXL_HDM_DECODER_CTRL(n)) != 4) {
> + return;
> + }
> + if (!(le32_to_cpu(ctrl) & VFIO_CXL_HDM_CTRL_COMMITTED)) {
> + continue;
> + }
> +
> + /*
> + * QEMU always maps at the device's CFMWS base and presents that base
> + * back to the guest, so a decoder the guest committed at any other base
> + * would leave the guest's view and the actual mapping diverged. Reject
> + * it rather than silently relocate. A firmware-committed decoder the
> + * guest never reprogrammed (guest_base_written == false) keeps the
> + * CFMWS base.
> + */
> + if (cxl->guest_base_written) {
> + hwaddr guest_base = ((hwaddr)cxl->guest_base_hi << 32) |
> + cxl->guest_base_lo;
> +
> + if (guest_base != cxl->fmws_base) {
> + warn_report("vfio-cxl: %s: guest committed decoder %u at 0x%"
> + HWADDR_PRIx ", not the CFMWS base 0x%" HWADDR_PRIx
> + "; not mapping", vbasedev->name, n, guest_base,
> + cxl->fmws_base);
> + vfio_cxl_unmap_mem(vdev);
> + return;
> + }
> + }
> +
> + if (cxl->dpa_mapped && cxl->mapped_base == cxl->fmws_base) {
> + return;
> + }
> + /*
> + * Overlap the CFMWS window MemoryRegion the CXL host bridge already
> + * maps at this base, at a higher priority, so precedence is defined by
> + * priority rather than subregion insertion order.
> + */
> + memory_region_transaction_begin();
> + vfio_cxl_unmap_mem(vdev);
> + memory_region_add_subregion_overlap(get_system_memory(),
> + cxl->fmws_base,
> + cxl->mem_region.mem, 1);
> + memory_region_transaction_commit();
> + cxl->mapped_base = cxl->fmws_base;
> + cxl->dpa_mapped = true;
> + return;
> + }
> +
> + /* No committed decoder in the block: tear down any existing mapping. */
> + vfio_cxl_unmap_mem(vdev);
> +}
> +
> +static uint64_t vfio_cxl_comp_regs_read(void *opaque, hwaddr addr,
> + unsigned size)
> +{
> + VFIORegion *region = opaque;
> + VFIODevice *vbasedev = region->vbasedev;
> + VFIOPCIDevice *vdev =
> + container_of(region, VFIOPCIDevice, cxl.comp_regs_region);
> + VFIOCXL *cxl = &vdev->cxl;
> + uint32_t val = 0xffffffff;
> +
> + if (pread(vbasedev->fd, &val, size, region->fd_offset + addr) != size) {
> + error_report("vfio-cxl: %s: HDM decoder read at 0x%" HWADDR_PRIx
> + " failed", vbasedev->name, addr);
> + }
> + val = le32_to_cpu(val);
> +
> + /*
> + * The kernel shadow holds the host physical base for a firmware-committed
> + * decoder. Never expose that to the guest: present the guest physical base
> + * QEMU maps the HDM memory at (the device's CFMWS window). Any decoder's
> + * base registers are virtualized; ctrl, size and the capability header pass
> + * through. The base registers repeat every VFIO_CXL_HDM_DECODER_STRIDE.
> + */
> + if (addr >= VFIO_CXL_HDM_DECODER_BASE_LOW(0) &&
> + (addr - VFIO_CXL_HDM_DECODER_BASE_LOW(0)) %
> + VFIO_CXL_HDM_DECODER_STRIDE == 0) {
> + val = (val & ~VFIO_CXL_HDM_BASE_LOW_MASK) |
> + ((uint32_t)cxl->fmws_base & VFIO_CXL_HDM_BASE_LOW_MASK);
> + } else if (addr >= VFIO_CXL_HDM_DECODER_BASE_HIGH(0) &&
> + (addr - VFIO_CXL_HDM_DECODER_BASE_HIGH(0)) %
> + VFIO_CXL_HDM_DECODER_STRIDE == 0) {
> + val = (uint32_t)(cxl->fmws_base >> 32);
> + }
> +
> + return val;
> +}
> +
> +static void vfio_cxl_comp_regs_write(void *opaque, hwaddr addr, uint64_t data,
> + unsigned size)
> +{
> + VFIORegion *region = opaque;
> + VFIODevice *vbasedev = region->vbasedev;
> + VFIOPCIDevice *vdev =
> + container_of(region, VFIOPCIDevice, cxl.comp_regs_region);
> + uint32_t val = cpu_to_le32((uint32_t)data);
> +
> + if (pwrite(vbasedev->fd, &val, size, region->fd_offset + addr) != size) {
> + error_report("vfio-cxl: %s: HDM decoder write at 0x%" HWADDR_PRIx
> + " failed", vbasedev->name, addr);
> + return;
> + }
> +
> + /*
> + * Record the base the guest programs into a decoder so the commit handler
> + * can reject a base other than the device's CFMWS window. The base low
> + * register carries HPA bits [31:28]; the high register carries [63:32].
> + */
> + if (addr >= VFIO_CXL_HDM_DECODER_BASE_LOW(0) &&
> + (addr - VFIO_CXL_HDM_DECODER_BASE_LOW(0)) %
> + VFIO_CXL_HDM_DECODER_STRIDE == 0) {
> + vdev->cxl.guest_base_lo = (uint32_t)data & VFIO_CXL_HDM_BASE_LOW_MASK;
> + vdev->cxl.guest_base_written = true;
> + } else if (addr >= VFIO_CXL_HDM_DECODER_BASE_HIGH(0) &&
> + (addr - VFIO_CXL_HDM_DECODER_BASE_HIGH(0)) %
> + VFIO_CXL_HDM_DECODER_STRIDE == 0) {
> + vdev->cxl.guest_base_hi = (uint32_t)data;
> + vdev->cxl.guest_base_written = true;
> + }
> +
> + /*
> + * The kernel runs the lock-on-commit FSM in the write above, so the
> + * committed state is settled by now; a control write on any decoder can
> + * change the mapping.
> + */
> + if (addr >= VFIO_CXL_HDM_DECODER_CTRL(0) &&
> + (addr - VFIO_CXL_HDM_DECODER_CTRL(0)) %
> + VFIO_CXL_HDM_DECODER_STRIDE == 0) {
> + vfio_cxl_decoder_changed(vdev);
> + }
> +}
> +
> +static const MemoryRegionOps vfio_cxl_comp_regs_ops = {
> + .read = vfio_cxl_comp_regs_read,
> + .write = vfio_cxl_comp_regs_write,
> + .endianness = DEVICE_LITTLE_ENDIAN,
> + .valid = { .min_access_size = 4, .max_access_size = 4 },
> + .impl = { .min_access_size = 4, .max_access_size = 4 },
> +};
> +
> +/*
> + * Locate the CXL Device DVSEC (CXL r3.1 8.1.3) in config space. The guest
> + * triggers a CXL reset by writing its ctrl2 register; QEMU rescans the decoder
> + * after that write so the HDM mapping is restored (see vfio_pci_write_config).
> + * Returns the DVSEC config offset, or 0 if the device does not expose it.
> + */
> +static uint16_t vfio_cxl_find_device_dvsec(PCIDevice *pdev)
> +{
> + uint16_t offset;
> +
> + for (offset = PCI_CONFIG_SPACE_SIZE; offset;
> + offset = PCI_EXT_CAP_NEXT(pci_get_long(pdev->config + offset))) {
> + uint32_t hdr = pci_get_long(pdev->config + offset);
> +
> + if (PCI_EXT_CAP_ID(hdr) == PCI_EXT_CAP_ID_DVSEC &&
> + (pci_get_long(pdev->config + offset + PCI_DVSEC_HEADER1) & 0xffff)
> + == CXL_VENDOR_ID &&
> + pci_get_word(pdev->config + offset + PCI_DVSEC_HEADER2)
> + == PCIE_CXL_DEVICE_DVSEC) {
> + return offset;
> + }
> + }
> +
> + return 0;
> +}
This should be a PCIe helper :
uint16_t pcie_find_dvsec(PCIDevice *dev, uint16_t vendor_id, uint16_t dvsec_id);
Thanks,
C.
> /*
> * Learn the CXL geometry the kernel reports: the HPA-backed HDM memory region
> * and the trapped HDM decoder block (which BAR carries it and at what offset).
> @@ -3864,6 +4195,8 @@ static bool vfio_cxl_setup(VFIOPCIDevice *vdev, Error **errp)
> cxl->dpa_size = mem_info->size;
> cxl->comp_bar = cap->bar;
> cxl->hdm_offset = cap->offset;
> + cxl->dvsec_offset = vfio_cxl_find_device_dvsec(&vdev->parent_obj);
> + cxl->af_offset = pci_find_capability(&vdev->parent_obj, PCI_CAP_ID_AF);
>
> if (!vfio_cxl_check_topology(vdev, errp)) {
> return false;
> @@ -3887,6 +4220,26 @@ static bool vfio_cxl_setup(VFIOPCIDevice *vdev, Error **errp)
> "performance may be slow", vbasedev->name);
> }
>
> + /*
> + * Trap the HDM decoder block: overlay it, priority 1, over the directly
> + * mapped component BAR, so guest decoder accesses reach the kernel FSM and
> + * the commit becomes visible to QEMU.
> + */
> + if (!vdev->bars[cxl->comp_bar].mr) {
> + error_setg(errp, "vfio-cxl: %s: component BAR %u is not present",
> + vbasedev->name, cxl->comp_bar);
> + goto err;
> + }
> + if (vfio_region_setup_with_ops(OBJECT(vdev), vbasedev,
> + &cxl->comp_regs_region,
> + cxl->comp_regs_region_index, "cxl-comp-regs",
> + &vfio_cxl_comp_regs_ops, errp)) {
> + goto err;
> + }
> + memory_region_add_subregion_overlap(vdev->bars[cxl->comp_bar].mr,
> + cxl->hdm_offset,
> + cxl->comp_regs_region.mem, 1);
> +
> if (DEVICE(vdev)->hotplugged) {
> /*
> * The machine is already up, so the CFMWS windows are placed and the
> @@ -3894,9 +4247,7 @@ static bool vfio_cxl_setup(VFIOPCIDevice *vdev, Error **errp)
> * device_add fails cleanly instead of aborting the running VM.
> */
> if (!vfio_cxl_do_bind_fmws(vdev, errp)) {
> - vfio_region_exit(&cxl->mem_region);
> - vfio_region_finalize(&cxl->mem_region);
> - return false;
> + goto err;
> }
> } else {
> cxl->machine_done.notify = vfio_cxl_bind_fmws;
> @@ -3906,6 +4257,10 @@ static bool vfio_cxl_setup(VFIOPCIDevice *vdev, Error **errp)
> cxl->enabled = true;
>
> return true;
> +
> +err:
> + vfio_cxl_teardown(vdev);
> + return false;
> }
>
> static void vfio_cxl_teardown(VFIOPCIDevice *vdev)
> @@ -3916,15 +4271,26 @@ static void vfio_cxl_teardown(VFIOPCIDevice *vdev)
> return;
> }
>
> - if (cxl->machine_done.notify) {
> - qemu_remove_machine_init_done_notifier(&cxl->machine_done);
> - cxl->machine_done.notify = NULL;
> + if (cxl->comp_regs_region.mem) {
> + if (vdev->bars[cxl->comp_bar].mr) {
> + memory_region_del_subregion(vdev->bars[cxl->comp_bar].mr,
> + cxl->comp_regs_region.mem);
> + }
> + vfio_region_exit(&cxl->comp_regs_region);
> + vfio_region_finalize(&cxl->comp_regs_region);
> }
>
> + vfio_cxl_unmap_mem(vdev);
> +
> if (cxl->mem_region.mem) {
> vfio_region_exit(&cxl->mem_region);
> vfio_region_finalize(&cxl->mem_region);
> }
> +
> + if (cxl->machine_done.notify) {
> + qemu_remove_machine_init_done_notifier(&cxl->machine_done);
> + cxl->machine_done.notify = NULL;
> + }
> }
>
> static void vfio_pci_realize(PCIDevice *pdev, Error **errp)
> @@ -4127,6 +4493,14 @@ static void vfio_exitfn(PCIDevice *pdev)
> vfio_pci_teardown_msi(vdev);
> vfio_pci_disable_rp_atomics(vdev);
> vfio_pci_bars_exit(vdev);
> + /*
> + * The committed HDM overlay is a subregion of system memory owned by this
> + * device, so it holds a reference that would keep the object alive past
> + * unrealize and block instance_finalize (where vfio_cxl_teardown otherwise
> + * runs). Drop it here; the call is idempotent for a device that never
> + * mapped or is not CXL.
> + */
> + vfio_cxl_unmap_mem(vdev);
> vfio_migration_exit(vbasedev);
> if (!vbasedev->mdev) {
> pci_device_unset_iommu_device(pdev);
> diff --git a/hw/vfio/pci.h b/hw/vfio/pci.h
> index 22fe8d7ff7..5bfa976112 100644
> --- a/hw/vfio/pci.h
> +++ b/hw/vfio/pci.h
> @@ -133,11 +133,19 @@ typedef struct VFIOCXL {
> uint32_t comp_regs_region_index; /* trapped HDM decoder register block */
> uint32_t comp_bar; /* component BAR carrying that block */
> uint64_t hdm_offset; /* block offset within the component BAR */
> + uint16_t dvsec_offset; /* CXL Device DVSEC offset, 0 if none */
> + uint16_t af_offset; /* PCI Advanced Features cap, 0 if none */
> uint64_t dpa_size; /* size of the HDM memory region */
> VFIORegion mem_region; /* HDM memory, mapped at committed GPA */
> Notifier machine_done; /* CFMWS validated at machine_done */
> hwaddr fmws_base; /* base of the memory window */
> uint64_t fmws_size; /* size of the memory window */
> + VFIORegion comp_regs_region; /* trapped HDM decoder block */
> + hwaddr mapped_base; /* GPA the HDM memory is mapped at */
> + bool dpa_mapped; /* HDM memory currently in system memory */
> + uint32_t guest_base_lo; /* decoder base low the guest wrote */
> + uint32_t guest_base_hi; /* decoder base high the guest wrote */
> + bool guest_base_written; /* the guest wrote a decoder base */
> } VFIOCXL;
>
> struct VFIOPCIDevice {
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v2 07/10] hw/vfio/pci: Map the CXL memory on the guest decoder commit
2026-09-16 18:44 ` [PATCH v2 07/10] hw/vfio/pci: Map the CXL memory on the guest decoder commit mhonap
2026-09-18 16:08 ` Cédric Le Goater
@ 2026-09-20 9:14 ` Junjie Cao
2026-09-21 11:17 ` Manish Honap
1 sibling, 1 reply; 27+ messages in thread
From: Junjie Cao @ 2026-09-20 9:14 UTC (permalink / raw)
To: Manish Honap
Cc: alex, ankita, jic23, dave.jiang, alejandro.lucero-palau,
smadhavan, pierrick.bouvier, mst, imammedo, anisinha, pbonzini,
eric.auger, peter.maydell, richard.henderson, clg, cohuck, kjaju,
vsethi, zhiw, qemu-devel, qemu-arm, linux-cxl
On Thu, 17 Sep 2026 00:14:09 +0530, Manish Honap wrote:
> + * Record the base the guest programs into a decoder so the commit handler
> + * can reject a base other than the device's CFMWS window. The base low
[...]
> + vdev->cxl.guest_base_lo = (uint32_t)data & VFIO_CXL_HDM_BASE_LOW_MASK;
> + vdev->cxl.guest_base_written = true;
This answered my v1 question for the kernel v4 FSM model. With v5
(20/27) the kernel gives a live read-only view and leaves any write
virtualization to the VMM, and what QEMU shows the guest is a
committed, locked decoder at the window base. A locked decoder
ignores a base write, but here it is recorded, and if it differs from
the window base the next rescan unmaps while reads still return
Committed at the window base. That is the mismatch the check was
meant to prevent. A current Linux guest never writes a locked
decoder, so it won't trip this.
Is the guest-written base still meant to matter in this model? If
not, guest_base_{lo,hi,written} can go.
> + * The kernel runs the lock-on-commit FSM in the write above, so the
> + * committed state is settled by now; a control write on any decoder can
Stale now that the kernel no longer acts on the write?
Junjie
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v2 10/10] hw/pci-host: Emit a _DSM on pxb-cxl to preserve firmware PCI config
2026-09-16 18:44 ` [PATCH v2 10/10] hw/pci-host: Emit a _DSM on pxb-cxl to preserve firmware PCI config mhonap
@ 2026-09-20 9:14 ` Junjie Cao
2026-09-21 11:20 ` Manish Honap
0 siblings, 1 reply; 27+ messages in thread
From: Junjie Cao @ 2026-09-20 9:14 UTC (permalink / raw)
To: Manish Honap
Cc: alex, ankita, jic23, dave.jiang, alejandro.lucero-palau,
smadhavan, pierrick.bouvier, mst, imammedo, anisinha, pbonzini,
eric.auger, peter.maydell, richard.henderson, clg, cohuck, kjaju,
vsethi, zhiw, qemu-devel, qemu-arm, linux-cxl, skolothumtho
On Thu, 17 Sep 2026 00:14:12 +0530, Manish Honap wrote:
> under preserve_config, the _DSM. Gate the _DSM on preserve_config: x86 q35
> passes false as it does not use the accelerated SMMU, so it emits no _DSM
> and its ACPI tables stay byte for byte what they were.
Confirmed: on 28e7aad522 plus the series bios-tables-test passes on
x86_64 (q35/cxl and q35/acpihmat-genericx, the v1 failures, included),
aarch64, riscv64 and loongarch64.
accel=on is out of reach under TCG, so I forced preserve_config on virt
locally and booted an arm64 guest (7.3-rc1 based, defconfig) with a
cxl-type3 behind a cxl-rp. With this patch Linux keeps the firmware
BARs and bridge window under the pxb-cxl; on 09/10 with the same hack
it releases and reassigns them. No ACPI errors either way.
Tested-by: Junjie Cao <junjie.cao@intel.com>
08/10 still says the x86 bridge "still emits the ``_DSM`` method, but
its function 0 returns an empty support mask"; that was v1.
Junjie
^ permalink raw reply [flat|nested] 27+ messages in thread
* RE: [PATCH v2 04/10] hw/vfio/pci: Enforce the passthrough topology for a CXL device
2026-09-17 13:20 ` Cédric Le Goater
@ 2026-09-21 10:45 ` Manish Honap
0 siblings, 0 replies; 27+ messages in thread
From: Manish Honap @ 2026-09-21 10:45 UTC (permalink / raw)
To: Cédric Le Goater, alex@shazbot.org, Ankit Agrawal,
jic23@kernel.org, dave.jiang@intel.com,
alejandro.lucero-palau@amd.com, Srirangan Madhavan,
pierrick.bouvier@oss.qualcomm.com, mst@redhat.com,
imammedo@redhat.com, anisinha@redhat.com, pbonzini@redhat.com,
eric.auger@redhat.com, peter.maydell@linaro.org,
richard.henderson@linaro.org, cohuck@redhat.com
Cc: Krishnakant Jaju, Vikram Sethi, Zhi Wang, qemu-devel@nongnu.org,
qemu-arm@nongnu.org, linux-cxl@vger.kernel.org, Manish Honap
> -----Original Message-----
> From: Cédric Le Goater <clg@redhat.com>
> Sent: Thursday, September 17, 2026 6:51 PM
> To: Manish Honap <mhonap@nvidia.com>; alex@shazbot.org; Ankit Agrawal
> <ankita@nvidia.com>; jic23@kernel.org; dave.jiang@intel.com;
> alejandro.lucero-palau@amd.com; Srirangan Madhavan
> <smadhavan@nvidia.com>; pierrick.bouvier@oss.qualcomm.com;
> mst@redhat.com; imammedo@redhat.com; anisinha@redhat.com;
> pbonzini@redhat.com; eric.auger@redhat.com; peter.maydell@linaro.org;
> richard.henderson@linaro.org; cohuck@redhat.com
> Cc: Krishnakant Jaju <kjaju@nvidia.com>; Vikram Sethi <vsethi@nvidia.com>;
> Zhi Wang <zhiw@nvidia.com>; qemu-devel@nongnu.org; qemu-
> arm@nongnu.org; linux-cxl@vger.kernel.org
> Subject: Re: [PATCH v2 04/10] hw/vfio/pci: Enforce the passthrough topology
> for a CXL device
>
> External email: Use caution opening links or attachments
>
>
> On 9/16/26 20:44, mhonap@nvidia.com wrote:
> > From: Manish Honap <mhonap@nvidia.com>
> >
> > Only the guest's endpoint HDM decoder is programmed by the guest; the
> > decoders above it are covered only while the pxb-cxl host bridge stays
> > in passthrough mode. Reject a switch in the path, or a host bridge
> > taken out of passthrough, at realize before any mapping exists.
> >
> > AI-used-for: code (prototype)
> > Signed-off-by: Manish Honap <mhonap@nvidia.com>
> > ---
> > hw/vfio/pci.c | 73
> +++++++++++++++++++++++++++++++++++++++++++++++++++
> > 1 file changed, 73 insertions(+)
> >
> > diff --git a/hw/vfio/pci.c b/hw/vfio/pci.c index
> > 2f84af5cf8..10b13f200d 100644
> > --- a/hw/vfio/pci.c
> > +++ b/hw/vfio/pci.c
> > @@ -27,6 +27,7 @@
> > #include "hw/pci/msi.h"
> > #include "hw/pci/msix.h"
> > #include "hw/pci/pci_bridge.h"
> > +#include "hw/cxl/cxl.h"
> > #include "hw/core/qdev-properties.h"
> > #include "hw/core/qdev-properties-system.h"
> > #include "hw/vfio/vfio-cpr.h"
> > @@ -3570,6 +3571,73 @@ bool vfio_pci_interrupt_setup(VFIOPCIDevice
> *vdev, Error **errp)
> > return true;
> > }
> >
> > +/*
> > + * Walk the endpoint's ancestor bridges up to its pxb-cxl host
> > +bridge. A CXL
> > + * switch in the path is rejected: the guest would then have to
> > +program switch
> > + * decoders, which the passthrough model does not yet cover. This is
> > +the single
> > + * place the supported-topology assumption lives, so switch support
> > +later
> > + * relaxes it here.
> > + */
> > +static PXBCXLDev *vfio_cxl_find_pxb(VFIOPCIDevice *vdev, Error
> > +**errp) {
> > + PCIBus *bus = pci_get_bus(&vdev->parent_obj);
>
> Could this routine be moved to the CXL subsytem component ?
okay, this part reads pxb-cxl and CXL bus types that hw/cxl owns, so it fits
there. I am planning that move together with the CFMWS matching from patch 6,
so the vfio-pci side calls a small CXL API.
> > +
> > + if (!object_dynamic_cast(OBJECT(bus), TYPE_CXL_BUS)) {
> > + error_setg(errp,
> > + "vfio-cxl: %s: endpoint is not on a CXL root port (cxl-rp)",
> > + vdev->vbasedev.name);
> > + return NULL;
> > + }
> > +
> > + for (; bus; ) {
>
> Why walk, since this is looking for the case "no-switch in between the bus and
> the ep" ? the no-switch case is just two hops in the topology.
Yes, the supported case is two hops EP <-> cxl-rp <-> pxb-cxl.
I will collapse the loop to this check during movement of the helper.
>
>
> > + PCIDevice *bridge = bus->parent_dev;
> > +
> > + if (!bridge) {
> > + break;
> > + }
> > + if (object_dynamic_cast(OBJECT(bridge), TYPE_CXL_USP) ||
> > + object_dynamic_cast(OBJECT(bridge), TYPE_CXL_DSP)) {
> > + error_setg(errp,
> > + "vfio-cxl: %s: switch-attached topology not supported",
> > + vdev->vbasedev.name);
> > + return NULL;
> > + }
> > + if (object_dynamic_cast(OBJECT(bridge), TYPE_PXB_CXL_DEV)) {
> > + return PXB_CXL_DEV(bridge);
> > + }
> > + if (pci_bus_is_root(bus)) {
> > + break;
> > + }
> > + bus = pci_get_bus(bridge);
> > + }
> > +
> > + error_setg(errp, "vfio-cxl: %s: not attached below a pxb-cxl host bridge",
> > + vdev->vbasedev.name);
> > + return NULL;
> > +}
> > +
> > +/*
> > + * Only guest's endpoint decoder is programmed, so the host bridge
> > +must stay in
> > + * HDM passthrough mode. Reject anything else at realize, before a
> > +mapping is
> > + * built. cxl_get_hb_passthrough() reflects reset state and is not
> > + * settled yet here, so gate on the static hdm_for_passthrough property.
> > + */
> > +static bool vfio_cxl_check_topology(VFIOPCIDevice *vdev, Error
> > +**errp) {
> > + PXBCXLDev *pxb = vfio_cxl_find_pxb(vdev, errp);
> > +
> > + if (!pxb) {
> > + return false;
> > + }
> > + if (pxb->hdm_for_passthrough) {
>
> This naming is a bit confusing:
>
> hdm_for_passthrough=false means "use HDM passthrough mode" (no
> decoders), while
> hdm_for_passthrough=true means "keep HDM decoders even in
> passthrough topology.
>
> I will get used to it.
Yes, I will add a comment at this point for more clarity.
>
> Thanks,
>
> C.
>
>
>
> > + error_setg(errp,
> > + "vfio-cxl: %s: the pxb-cxl must be in HDM passthrough mode "
> > + "(do not set hdm_for_passthrough)", vdev->vbasedev.name);
> > + return false;
> > + }
> > + return true;
> > +}
> > +
> > /*
> > * Learn the CXL geometry the kernel reports: the HPA-backed HDM memory
> region
> > * and the trapped HDM decoder block (which BAR carries it and at what
> offset).
> > @@ -3623,6 +3691,11 @@ static bool vfio_cxl_setup(VFIOPCIDevice *vdev,
> Error **errp)
> > cxl->dpa_size = mem_info->size;
> > cxl->comp_bar = cap->bar;
> > cxl->hdm_offset = cap->offset;
> > +
> > + if (!vfio_cxl_check_topology(vdev, errp)) {
> > + return false;
> > + }
> > +
> > cxl->enabled = true;
> >
> > return true;
^ permalink raw reply [flat|nested] 27+ messages in thread
* RE: [PATCH v2 02/10] hw/vfio/region: Add vfio_region_setup_with_ops()
2026-09-17 13:27 ` Cédric Le Goater
@ 2026-09-21 10:48 ` Manish Honap
0 siblings, 0 replies; 27+ messages in thread
From: Manish Honap @ 2026-09-21 10:48 UTC (permalink / raw)
To: Cédric Le Goater, alex@shazbot.org, Ankit Agrawal,
jic23@kernel.org, dave.jiang@intel.com,
alejandro.lucero-palau@amd.com, Srirangan Madhavan,
pierrick.bouvier@oss.qualcomm.com, mst@redhat.com,
imammedo@redhat.com, anisinha@redhat.com, pbonzini@redhat.com,
eric.auger@redhat.com, peter.maydell@linaro.org,
richard.henderson@linaro.org, cohuck@redhat.com
Cc: Krishnakant Jaju, Vikram Sethi, Zhi Wang, qemu-devel@nongnu.org,
qemu-arm@nongnu.org, linux-cxl@vger.kernel.org, Manish Honap
> -----Original Message-----
> From: Cédric Le Goater <clg@redhat.com>
> Sent: Thursday, September 17, 2026 6:57 PM
> To: Manish Honap <mhonap@nvidia.com>; alex@shazbot.org; Ankit Agrawal
> <ankita@nvidia.com>; jic23@kernel.org; dave.jiang@intel.com;
> alejandro.lucero-palau@amd.com; Srirangan Madhavan
> <smadhavan@nvidia.com>; pierrick.bouvier@oss.qualcomm.com;
> mst@redhat.com; imammedo@redhat.com; anisinha@redhat.com;
> pbonzini@redhat.com; eric.auger@redhat.com; peter.maydell@linaro.org;
> richard.henderson@linaro.org; cohuck@redhat.com
> Cc: Krishnakant Jaju <kjaju@nvidia.com>; Vikram Sethi <vsethi@nvidia.com>;
> Zhi Wang <zhiw@nvidia.com>; qemu-devel@nongnu.org; qemu-
> arm@nongnu.org; linux-cxl@vger.kernel.org
> Subject: Re: [PATCH v2 02/10] hw/vfio/region: Add
> vfio_region_setup_with_ops()
>
> External email: Use caution opening links or attachments
>
>
> On 9/16/26 20:44, mhonap@nvidia.com wrote:
> > From: Manish Honap <mhonap@nvidia.com>
> >
> > A CXL Type-2 device traps its HDM decoder block, so that region needs
> > custom MemoryRegionOps rather than the pass-through default. Split the
> > setup body out and let a caller supply the ops; NULL keeps the
> > existing behaviour.
> >
> > AI-used-for: code (prototype)
> > Signed-off-by: Manish Honap <mhonap@nvidia.com>
> > ---
> > hw/vfio/region.c | 27 ++++++++++++++++++++++++---
> > hw/vfio/vfio-region.h | 3 +++
> > 2 files changed, 27 insertions(+), 3 deletions(-)
> >
> > diff --git a/hw/vfio/region.c b/hw/vfio/region.c index
> > 54ad11a6c8..21b54978b5 100644
> > --- a/hw/vfio/region.c
> > +++ b/hw/vfio/region.c
> > @@ -228,8 +228,9 @@ static int
> vfio_setup_region_sparse_mmaps(VFIORegion *region,
> > return 0;
> > }
> >
> > -int vfio_region_setup(Object *obj, VFIODevice *vbasedev, VFIORegion
> *region,
> > - int index, const char *name, Error **errp)
> > +static int vfio_region_do_setup(Object *obj, VFIODevice *vbasedev,
> > + VFIORegion *region, int index, const char *name,
> > + const MemoryRegionOps *ops, Error
> > +**errp)
> > {
> > struct vfio_region_info *info = NULL;
> > int ret;
> > @@ -249,7 +250,7 @@ int vfio_region_setup(Object *obj, VFIODevice
> > *vbasedev, VFIORegion *region,
> >
> > if (region->size) {
> > region->mem = g_new0(MemoryRegion, 1);
> > - memory_region_init_io(region->mem, obj, &vfio_region_ops,
> > + memory_region_init_io(region->mem, obj, ops,
> > region, name, region->size);
> >
> > if (!vbasedev->no_mmap &&
> > @@ -273,6 +274,26 @@ int vfio_region_setup(Object *obj, VFIODevice
> *vbasedev, VFIORegion *region,
> > return 0;
> > }
> >
> > +int vfio_region_setup(Object *obj, VFIODevice *vbasedev, VFIORegion
> *region,
> > + int index, const char *name, Error **errp) {
> > + return vfio_region_do_setup(obj, vbasedev, region, index, name,
> > + &vfio_region_ops, errp); }
> > +
> > +/*
> > + * Like vfio_region_setup() but traps the region through @ops instead of
> the
> > + * default pass-through, so a caller can intercept accesses (the CXL HDM
> > + * decoder block). A NULL @ops keeps the default.
> > + */
> > +int vfio_region_setup_with_ops(Object *obj, VFIODevice *vbasedev,
> > + VFIORegion *region, int index, const char *name,
> > + const MemoryRegionOps *ops, Error **errp)
> > +{
> > + return vfio_region_do_setup(obj, vbasedev, region, index, name,
> > + ops ? ops : &vfio_region_ops, errp);
>
> if you're calling _with_ops(), you have custom ops, that's whole point.
> Passing NULL is a caller bug, no need to have a fall back. Or assert().
>
Okay, understood. A NULL ops here is a caller bug. I will add an assert on ops
and drop the fallback.
> Thanks,
>
> C.
>
>
>
> > +}
> > +
> > static void vfio_subregion_unmap(VFIORegion *region, int index)
> > {
> > trace_vfio_region_unmap(memory_region_name(®ion-
> >mmaps[index].mem),
> > diff --git a/hw/vfio/vfio-region.h b/hw/vfio/vfio-region.h
> > index 58b236f113..8d5699013a 100644
> > --- a/hw/vfio/vfio-region.h
> > +++ b/hw/vfio/vfio-region.h
> > @@ -39,6 +39,9 @@ uint64_t vfio_region_read(void *opaque,
> > hwaddr addr, unsigned size);
> > int vfio_region_setup(Object *obj, VFIODevice *vbasedev, VFIORegion
> *region,
> > int index, const char *name, Error **errp);
> > +int vfio_region_setup_with_ops(Object *obj, VFIODevice *vbasedev,
> > + VFIORegion *region, int index, const char *name,
> > + const MemoryRegionOps *ops, Error **errp);
> > int vfio_region_mmap(VFIORegion *region);
> > void vfio_region_mmaps_set_enabled(VFIORegion *region, bool enabled);
> > void vfio_region_exit(VFIORegion *region);
^ permalink raw reply [flat|nested] 27+ messages in thread
* RE: [PATCH v2 03/10] hw/vfio/pci: Detect a CXL Type-2 device and read its geometry
2026-09-17 16:04 ` Cédric Le Goater
@ 2026-09-21 10:49 ` Manish Honap
0 siblings, 0 replies; 27+ messages in thread
From: Manish Honap @ 2026-09-21 10:49 UTC (permalink / raw)
To: Cédric Le Goater, alex@shazbot.org, Ankit Agrawal,
jic23@kernel.org, dave.jiang@intel.com,
alejandro.lucero-palau@amd.com, Srirangan Madhavan,
pierrick.bouvier@oss.qualcomm.com, mst@redhat.com,
imammedo@redhat.com, anisinha@redhat.com, pbonzini@redhat.com,
eric.auger@redhat.com, peter.maydell@linaro.org,
richard.henderson@linaro.org, cohuck@redhat.com
Cc: Krishnakant Jaju, Vikram Sethi, Zhi Wang, qemu-devel@nongnu.org,
qemu-arm@nongnu.org, linux-cxl@vger.kernel.org, Manish Honap
> -----Original Message-----
> From: Cédric Le Goater <clg@redhat.com>
> Sent: Thursday, September 17, 2026 9:34 PM
> To: Manish Honap <mhonap@nvidia.com>; alex@shazbot.org; Ankit Agrawal
> <ankita@nvidia.com>; jic23@kernel.org; dave.jiang@intel.com;
> alejandro.lucero-palau@amd.com; Srirangan Madhavan
> <smadhavan@nvidia.com>; pierrick.bouvier@oss.qualcomm.com;
> mst@redhat.com; imammedo@redhat.com; anisinha@redhat.com;
> pbonzini@redhat.com; eric.auger@redhat.com; peter.maydell@linaro.org;
> richard.henderson@linaro.org; cohuck@redhat.com
> Cc: Krishnakant Jaju <kjaju@nvidia.com>; Vikram Sethi <vsethi@nvidia.com>;
> Zhi Wang <zhiw@nvidia.com>; qemu-devel@nongnu.org; qemu-
> arm@nongnu.org; linux-cxl@vger.kernel.org
> Subject: Re: [PATCH v2 03/10] hw/vfio/pci: Detect a CXL Type-2 device and
> read its geometry
>
> External email: Use caution opening links or attachments
>
>
> On 9/16/26 20:44, mhonap@nvidia.com wrote:
> > From: Manish Honap <mhonap@nvidia.com>
> >
> > The kernel marks a passthroughed CXL Type-2 device with a device flag
> > and exposes two regions:
> > - HDM memory (host physical)
> > - the trapped HDM decoder register block.
> >
> > Read the flag, locate both regions by type, and take the component BAR
> > and block offset from the geometry capability. Realize only records
> > this; later patches build the guest mapping on top.
> >
> > AI-used-for: code (prototype)
> > Signed-off-by: Manish Honap <mhonap@nvidia.com>
>
> LGTM.
>
Thank you for the review of this patch. I will carry this forward unchanged in v3.
> Thanks,
>
> C.
>
>
> > ---
> > hw/vfio/pci.c | 72
> +++++++++++++++++++++++++++++++++++++++++++++++++++
> > hw/vfio/pci.h | 15 +++++++++++
> > 2 files changed, 87 insertions(+)
> >
> > diff --git a/hw/vfio/pci.c b/hw/vfio/pci.c index
> > 428ab2f069..2f84af5cf8 100644
> > --- a/hw/vfio/pci.c
> > +++ b/hw/vfio/pci.c
> > @@ -3570,6 +3570,64 @@ bool vfio_pci_interrupt_setup(VFIOPCIDevice
> *vdev, Error **errp)
> > return true;
> > }
> >
> > +/*
> > + * Learn the CXL geometry the kernel reports: the HPA-backed HDM
> > +memory region
> > + * and the trapped HDM decoder block (which BAR carries it and at what
> offset).
> > + * A non-CXL device leaves cxl.enabled false and takes no CXL paths.
> > + */
> > +static bool vfio_cxl_setup(VFIOPCIDevice *vdev, Error **errp) {
> > + VFIODevice *vbasedev = &vdev->vbasedev;
> > + VFIOCXL *cxl = &vdev->cxl;
> > + struct vfio_region_info *mem_info = NULL, *comp_info = NULL;
> > + struct vfio_region_info_cap_cxl_comp_regs *cap;
> > + struct vfio_info_cap_header *hdr;
> > +
> > + if (!(vbasedev->flags & VFIO_DEVICE_FLAGS_CXL)) {
> > + return true;
> > + }
> > +
> > + if (vfio_device_get_region_info_type(vbasedev,
> > + VFIO_REGION_TYPE_PCI_VENDOR_TYPE |
> > + CXL_VENDOR_ID,
> > + VFIO_REGION_SUBTYPE_CXL_MEM,
> > + &mem_info)) {
> > + error_setg(errp, "vfio-cxl: %s: CXL memory region not found",
> > + vbasedev->name);
> > + return false;
> > + }
> > +
> > + if (vfio_device_get_region_info_type(vbasedev,
> > + VFIO_REGION_TYPE_PCI_VENDOR_TYPE |
> > + CXL_VENDOR_ID,
> > + VFIO_REGION_SUBTYPE_CXL_COMP_REGS,
> > + &comp_info)) {
> > + error_setg(errp,
> > + "vfio-cxl: %s: CXL component-register region not found",
> > + vbasedev->name);
> > + return false;
> > + }
> > +
> > + hdr = vfio_get_region_info_cap(comp_info,
> > + VFIO_REGION_INFO_CAP_CXL_COMP_REGS);
> > + if (!hdr) {
> > + error_setg(errp,
> > + "vfio-cxl: %s: component-register geometry not reported",
> > + vbasedev->name);
> > + return false;
> > + }
> > + cap = container_of(hdr, struct
> > + vfio_region_info_cap_cxl_comp_regs, header);
> > +
> > + cxl->mem_region_index = mem_info->index;
> > + cxl->comp_regs_region_index = comp_info->index;
> > + cxl->dpa_size = mem_info->size;
> > + cxl->comp_bar = cap->bar;
> > + cxl->hdm_offset = cap->offset;
> > + cxl->enabled = true;
> > +
> > + return true;
> > +}
> > +
> > static void vfio_pci_realize(PCIDevice *pdev, Error **errp)
> > {
> > ERRP_GUARD();
> > @@ -3700,6 +3758,20 @@ static void vfio_pci_realize(PCIDevice *pdev,
> Error **errp)
> > }
> > }
> >
> > + if (!vfio_cxl_setup(vdev, errp)) {
> > + /*
> > + * vfio_migration_realize() above installed a migration blocker in auto
> > + * mode (generic vfio-pci exposes no migration ops). out_deregister
> does
> > + * not remove it, so a rejected CXL setup would leave VM migration
> > + * blocked until QEMU restarts. Drop it here, under the same
> > + * failover-pair condition used to install it.
> > + */
> > + if (!pdev->failover_pair_id) {
> > + vfio_migration_exit(vbasedev);
> > + }
> > + goto out_deregister;
> > + }
> > +
> > vfio_pci_register_err_notifier(vdev);
> > vfio_pci_register_req_notifier(vdev);
> > vfio_setup_resetfn_quirk(vdev);
> > diff --git a/hw/vfio/pci.h b/hw/vfio/pci.h index
> > c9ab949870..7fdd695704 100644
> > --- a/hw/vfio/pci.h
> > +++ b/hw/vfio/pci.h
> > @@ -122,10 +122,25 @@ typedef struct VFIOMSIXInfo {
> >
> > OBJECT_DECLARE_SIMPLE_TYPE(VFIOPCIDevice, VFIO_PCI_DEVICE)
> >
> > +/*
> > + * State for a CXL Type-2 device. The kernel owns the host physical
> > +placement
> > + * of the device memory; QEMU only maps it at the guest physical
> > +address the
> > + * guest commits into its endpoint HDM decoder.
> > + */
> > +typedef struct VFIOCXL {
> > + bool enabled;
> > + uint32_t mem_region_index; /* HPA-backed HDM memory VFIO
> region */
> > + uint32_t comp_regs_region_index; /* trapped HDM decoder register
> block */
> > + uint32_t comp_bar; /* component BAR carrying that block */
> > + uint64_t hdm_offset; /* block offset within the component BAR */
> > + uint64_t dpa_size; /* size of the HDM memory region */
> > +} VFIOCXL;
> > +
> > struct VFIOPCIDevice {
> > PCIDevice parent_obj;
> >
> > VFIODevice vbasedev;
> > + VFIOCXL cxl;
> > VFIOINTx intx;
> > unsigned int config_size;
> > uint8_t *emulated_config_bits; /* QEMU emulated bits,
> > little-endian */
^ permalink raw reply [flat|nested] 27+ messages in thread
* RE: [PATCH v2 05/10] hw/vfio/pci: Back the CXL memory with a RAM-device region
2026-09-17 13:56 ` Cédric Le Goater
@ 2026-09-21 10:55 ` Manish Honap
0 siblings, 0 replies; 27+ messages in thread
From: Manish Honap @ 2026-09-21 10:55 UTC (permalink / raw)
To: Cédric Le Goater, alex@shazbot.org, Ankit Agrawal,
jic23@kernel.org, dave.jiang@intel.com,
alejandro.lucero-palau@amd.com, Srirangan Madhavan,
pierrick.bouvier@oss.qualcomm.com, mst@redhat.com,
imammedo@redhat.com, anisinha@redhat.com, pbonzini@redhat.com,
eric.auger@redhat.com, peter.maydell@linaro.org,
richard.henderson@linaro.org, cohuck@redhat.com
Cc: Krishnakant Jaju, Vikram Sethi, Zhi Wang, qemu-devel@nongnu.org,
qemu-arm@nongnu.org, linux-cxl@vger.kernel.org, Manish Honap
> -----Original Message-----
> From: Cédric Le Goater <clg@redhat.com>
> Sent: Thursday, September 17, 2026 7:26 PM
> To: Manish Honap <mhonap@nvidia.com>; alex@shazbot.org; Ankit Agrawal
> <ankita@nvidia.com>; jic23@kernel.org; dave.jiang@intel.com;
> alejandro.lucero-palau@amd.com; Srirangan Madhavan
> <smadhavan@nvidia.com>; pierrick.bouvier@oss.qualcomm.com;
> mst@redhat.com; imammedo@redhat.com; anisinha@redhat.com;
> pbonzini@redhat.com; eric.auger@redhat.com; peter.maydell@linaro.org;
> richard.henderson@linaro.org; cohuck@redhat.com
> Cc: Krishnakant Jaju <kjaju@nvidia.com>; Vikram Sethi <vsethi@nvidia.com>;
> Zhi Wang <zhiw@nvidia.com>; qemu-devel@nongnu.org; qemu-
> arm@nongnu.org; linux-cxl@vger.kernel.org
> Subject: Re: [PATCH v2 05/10] hw/vfio/pci: Back the CXL memory with a
> RAM-device region
>
> External email: Use caution opening links or attachments
>
>
> On 9/16/26 20:44, mhonap@nvidia.com wrote:
> > From: Manish Honap <mhonap@nvidia.com>
> >
> > mmap the kernel's HDM memory region so its host physical pages back a
> > RAM-device MemoryRegion. It stays out of the guest address space until
> > the guest commits its decoder and QEMU learns the GPA to place it at.
> >
> > AI-used-for: code (prototype)
> > Signed-off-by: Manish Honap <mhonap@nvidia.com>
> > ---
> > hw/vfio/pci.c | 34 ++++++++++++++++++++++++++++++++++
> > hw/vfio/pci.h | 1 +
> > 2 files changed, 35 insertions(+)
> >
> > diff --git a/hw/vfio/pci.c b/hw/vfio/pci.c index
> > 10b13f200d..4716266595 100644
> > --- a/hw/vfio/pci.c
> > +++ b/hw/vfio/pci.c
> > @@ -3221,8 +3221,11 @@ bool vfio_pci_populate_device(VFIOPCIDevice
> *vdev, Error **errp)
> > return true;
> > }
> >
> > +static void vfio_cxl_teardown(VFIOPCIDevice *vdev);
> > +
> > void vfio_pci_put_device(VFIOPCIDevice *vdev)
> > {
> > + vfio_cxl_teardown(vdev);
> > vfio_display_finalize(vdev);
> > vfio_bars_finalize(vdev);
> > vfio_cpr_pci_unregister_device(vdev);
> > @@ -3696,11 +3699,42 @@ static bool vfio_cxl_setup(VFIOPCIDevice
> *vdev, Error **errp)
> > return false;
> > }
> >
> > + /*
> > + * The HDM memory is host physical. Set up the region, which installs the
> > + * fd read/write path, and mmap it for direct guest access; it is added to
> > + * the guest address space only once the guest commits its endpoint
> decoder.
> > + */
> > + if (vfio_region_setup(OBJECT(vdev), vbasedev, &cxl->mem_region,
> > + cxl->mem_region_index, "cxl-mem", errp)) {
> > + return false;
> > + }
> > + if (vfio_region_mmap(&cxl->mem_region)) {
> > + /*
> > + * Without mmap the region falls back to the kernel's fd read/write
> > + * path, which works but traps every access. Warn rather than fail.
> > + */
> > + warn_report("vfio-cxl: %s: failed to mmap the HDM memory region; "
> > + "performance may be slow", vbasedev->name);
> > + }
>
> vfio_cxl_setup calls vfio_region_setup() then vfio_region_mmap() back-to-
> back when other 'normal' BARs split these calls across
> vfio_populate_device() and vfio_bar_register(). I guess it is fine for CXL since it
> is not a PCI BAR.
>
Yes, I also had the same reasoning for this part.
> > cxl->enabled = true;
> >
> > return true;
> > }
> >
> > +static void vfio_cxl_teardown(VFIOPCIDevice *vdev) {
> > + VFIOCXL *cxl = &vdev->cxl;
> > +
> > + if (!cxl->enabled) {
> > + return;
> > + }
> > + if (cxl->mem_region.mem) {
> > + vfio_region_exit(&cxl->mem_region);
> > + vfio_region_finalize(&cxl->mem_region);
>
> However, the tear down should be split between :
>
> 1. vfio_exitfn (unrealize)
> drops mmaps, removes subregions and removes references so the
> MR refcount can reach zero.
> 2. vfio_pci_finalize (instance_finalize)
> frees the MemoryRegion.
>
okay, I will move this as suggested.
> Thanks,
>
> C.
>
> > + }
> > +}
> > +
> > static void vfio_pci_realize(PCIDevice *pdev, Error **errp)
> > {
> > ERRP_GUARD();
> > diff --git a/hw/vfio/pci.h b/hw/vfio/pci.h index
> > 7fdd695704..06d15e807d 100644
> > --- a/hw/vfio/pci.h
> > +++ b/hw/vfio/pci.h
> > @@ -134,6 +134,7 @@ typedef struct VFIOCXL {
> > uint32_t comp_bar; /* component BAR carrying that block */
> > uint64_t hdm_offset; /* block offset within the component BAR */
> > uint64_t dpa_size; /* size of the HDM memory region */
> > + VFIORegion mem_region; /* HDM memory, mapped at committed
> GPA */
> > } VFIOCXL;
> >
> > struct VFIOPCIDevice {
^ permalink raw reply [flat|nested] 27+ messages in thread
* RE: [PATCH v2 06/10] hw/vfio/pci: Bind a CXL device to its fixed memory window
2026-09-17 16:00 ` Cédric Le Goater
@ 2026-09-21 11:03 ` Manish Honap
0 siblings, 0 replies; 27+ messages in thread
From: Manish Honap @ 2026-09-21 11:03 UTC (permalink / raw)
To: Cédric Le Goater, alex@shazbot.org, Ankit Agrawal,
jic23@kernel.org, dave.jiang@intel.com,
alejandro.lucero-palau@amd.com, Srirangan Madhavan,
pierrick.bouvier@oss.qualcomm.com, mst@redhat.com,
imammedo@redhat.com, anisinha@redhat.com, pbonzini@redhat.com,
eric.auger@redhat.com, peter.maydell@linaro.org,
richard.henderson@linaro.org, cohuck@redhat.com,
Nitesh Narayan Lal
Cc: Krishnakant Jaju, Vikram Sethi, Zhi Wang, qemu-devel@nongnu.org,
qemu-arm@nongnu.org, linux-cxl@vger.kernel.org, Manish Honap
> -----Original Message-----
> From: Cédric Le Goater <clg@redhat.com>
> Sent: Thursday, September 17, 2026 9:30 PM
> To: Manish Honap <mhonap@nvidia.com>; alex@shazbot.org; Ankit Agrawal
> <ankita@nvidia.com>; jic23@kernel.org; dave.jiang@intel.com;
> alejandro.lucero-palau@amd.com; Srirangan Madhavan
> <smadhavan@nvidia.com>; pierrick.bouvier@oss.qualcomm.com;
> mst@redhat.com; imammedo@redhat.com; anisinha@redhat.com;
> pbonzini@redhat.com; eric.auger@redhat.com; peter.maydell@linaro.org;
> richard.henderson@linaro.org; cohuck@redhat.com; Nitesh Narayan Lal
> <nilal@redhat.com>
> Cc: Krishnakant Jaju <kjaju@nvidia.com>; Vikram Sethi <vsethi@nvidia.com>;
> Zhi Wang <zhiw@nvidia.com>; qemu-devel@nongnu.org; qemu-
> arm@nongnu.org; linux-cxl@vger.kernel.org
> Subject: Re: [PATCH v2 06/10] hw/vfio/pci: Bind a CXL device to its fixed
> memory window
>
> External email: Use caution opening links or attachments
>
>
> On 9/16/26 20:44, mhonap@nvidia.com wrote:
> > From: Manish Honap <mhonap@nvidia.com>
> >
> > The guest programs a GPA inside a CXL fixed memory window and QEMU
> > maps the HDM memory there, so the device must sit under exactly one
> > single-target CFMWS.
>
> CFMWS (CXL Fixed Memory Window Structure). Pity that the fields are named
> fmws_base, fmws_size.
>
> > Reject any ambiguous, interleaved, unsized, or missing window rather
> > than guess. Match windows by target name, since the resolved target
> > pointers are only filled in by a machine-done notifier that may run
> > after this one.
> >
> > Require exactly one endpoint function below the root port: two
> > functions in the same slot would each match that single-target CFMWS
> > and alias their HDM memory at one base, so count every present
> > function rather than one per slot. The CFMWS windows are placed at
> > machine-init-done, so a cold-plugged device can only be validated from
> > that notifier, where a bad topology is fatal to startup. Validate a
> > hotplugged device inline from realize instead, reporting through errp
> > so a bad device_add fails cleanly rather than aborting the running VM,
> > and arm the notifier only for a cold-plugged device.
> >
> > The endpoint's HDM size comes from the kernel region.
> > cxl_fmws_set_memmap() places the window in the guest PA map before
> > this device is realized, so the window cannot yet be sized from the
> > device; check the configured window against the device geometry
> > instead. Reject a window too small to hold the endpoint and name the
> > size to set, and warn on an oversized window, since the padding becomes
> guest CEDT that migration has to preserve.
>
> I am lost ... Too much stuff there.
>
> Each paragraph introduces a new topic/problem, and implementation details,
> all mixed up, without first explaining what problem is being solved:
>
> A passed-through CXL Type-2 device needs to know where its memory
> window lives in the guest PA space so QEMU can map the device memory
> there.
>
> Is that it ?
>
> A maintainer reviewing this patch shouldn't need to know what a CFMWS is to
> understand the commit. CXL is still an emerging technology.
> Most VFIO and QEMU reviewers will not have CXL spec background and CXL-
> specific mechanic knowledge.
>
> > AI-used-for: code (prototype)
>
> The code still reads like a proof-of-concept "prototype". This patch in
> particular is very hard to follow. The problem isn't clearly described, and likely
> isn't fully understood yet. The series should be broken into smaller patches.
>
> > Signed-off-by: Manish Honap <mhonap@nvidia.com>
> > ---
> > hw/pci-bridge/pci_expander_bridge_stubs.c | 6 +
> > hw/vfio/pci.c | 192 ++++++++++++++++++++++
> > hw/vfio/pci.h | 3 +
> > 3 files changed, 201 insertions(+)
> >
> > diff --git a/hw/pci-bridge/pci_expander_bridge_stubs.c
> > b/hw/pci-bridge/pci_expander_bridge_stubs.c
> > index b35180311f..c44ad7fab6 100644
> > --- a/hw/pci-bridge/pci_expander_bridge_stubs.c
> > +++ b/hw/pci-bridge/pci_expander_bridge_stubs.c
> > @@ -10,5 +10,11 @@
> > #include "hw/pci/pci_bus.h"
> > #include "hw/pci-bridge/pci_expander_bridge.h"
> > #include "hw/cxl/cxl.h"
> > +#include "hw/cxl/cxl_component.h"
> >
> > void pxb_cxl_hook_up_registers(CXLState *state, PCIBus *bus, Error
> > **errp) {};
> > +
> > +bool cxl_get_hb_passthrough(PCIHostState *hb) {
> > + return false;
> > +}
> > diff --git a/hw/vfio/pci.c b/hw/vfio/pci.c index
> > 4716266595..670e0d1da4 100644
> > --- a/hw/vfio/pci.c
> > +++ b/hw/vfio/pci.c
> > @@ -27,7 +27,12 @@
> > #include "hw/pci/msi.h"
> > #include "hw/pci/msix.h"
> > #include "hw/pci/pci_bridge.h"
> > +#include "hw/pci/pci_host.h"
> > +#include "hw/pci/pcie_port.h"
> > #include "hw/cxl/cxl.h"
> > +#include "hw/cxl/cxl_host.h"
> > +#include "hw/cxl/cxl_component.h"
> > +#include "system/system.h"
> > #include "hw/core/qdev-properties.h"
> > #include "hw/core/qdev-properties-system.h"
> > #include "hw/vfio/vfio-cpr.h"
> > @@ -3641,6 +3646,171 @@ static bool
> vfio_cxl_check_topology(VFIOPCIDevice *vdev, Error **errp)
> > return true;
> > }
> >
> > +/*
> > + * Count the CXL fixed memory windows that target this device's
> > +pxb-cxl and
> > + * report the matched window's base, size and target count. Match on
> > +the target
> > + * names: cxl_fmws_link_targets() resolves target_hbs[] only at
> > +machine_done,
> > + * which may run after this, so the resolved pointers can still be NULL here.
> > + */
> > +static int vfio_cxl_match_fmws(PXBCXLDev *pxb, hwaddr *base, uint64_t
> *size,
> > + int *nwindows, int *ntargets)
>
> I think this should be in the CXL subsystem. That's a *lot* of out parameters ...
>
> > +{
> > + GSList *list = cxl_fmws_get_all_sorted();
> > + GSList *iter;
> > + int matches = 0;
> > +
> > + *nwindows = g_slist_length(list);
> > + *base = 0;
> > + *size = 0;
> > + *ntargets = 0;
> > +
> > + for (iter = list; iter; iter = iter->next) {
> > + CXLFixedWindow *fw = CXL_FMW(iter->data);
> > + int i;
> > +
> > + for (i = 0; i < fw->num_targets; i++) {
> > + bool ambiguous = false;
> > + Object *t = object_resolve_path_type(fw->targets[i],
> > + TYPE_PXB_CXL_DEV,
> > + &ambiguous);
> > +
> > + if (t && !ambiguous && PXB_CXL_DEV(t) == pxb) {
> > + *base = fw->base;
> > + *size = fw->size;
> > + *ntargets = fw->num_targets;
> > + matches++;
> > + break;
> > + }
> > + }
> > + }
> > + g_slist_free(list);
> > + return matches;
> > +}
> > +
> > +/*
> > + * Fix the bounds of the device's memory window once the topology has
> settled.
> > + * Exactly one single-target CFMWS, with an assigned base and room
> > +for the HDM
> > + * memory, is required; anything else is a misconfiguration the guest
> > +cannot
> > + * recover from, so reject it rather than guess a window.
> > + */
> > +static bool vfio_cxl_do_bind_fmws(VFIOPCIDevice *vdev, Error **errp)
> > +{
>
> This does too much :
>
> topology validation
> CFMWS window matching
> size validation
>
> It looks like vfio_cxl_check_topology() and it should be in the CXL subsystem.
>
> > + VFIOCXL *cxl = &vdev->cxl;
> > + const char *name = vdev->vbasedev.name;
> > + int nwindows = 0, ntargets = 0, matches;
> > + hwaddr base = 0;
> > + uint64_t size = 0, need;
> > + PXBCXLDev *pxb;
> > +
> > + pxb = vfio_cxl_find_pxb(vdev, errp);
> > + if (!pxb) {
> > + return false;
> > + }
> > +
> > + if
> > + (pcie_count_ds_ports(PCI_HOST_BRIDGE(pxb->cxl_host_bridge)->bus) !=
> > + 1) {
>
> This won't build on all platforms.
>
> > + error_setg(errp,
> > + "vfio-cxl: %s: pxb-cxl not in HDM passthrough mode "
> > + "(use a single cxl-rp)", name);
> > + return false;
> > + }
> > +
> > + /*
> > + * Count every present function, not one per slot, and require exactly one
> > + * endpoint below the root port.
> > + */> + {
> > + PCIBus *ep_bus = pci_get_bus(&vdev->parent_obj);
> > + int slot, fn, nendpoints = 0;
>
> This smells like a function.
>
> > + for (slot = 0; slot < PCI_SLOT_MAX; slot++) {
> > + for (fn = 0; fn < PCI_FUNC_MAX; fn++) {
> > + if (ep_bus->devices[PCI_DEVFN(slot, fn)]) {
> > + nendpoints++;
> > + }
> > + }
> > + }
> > + if (nendpoints != 1) {
> > + error_setg(errp,
> > + "vfio-cxl: %s: %d devices below the cxl-rp; a passed "
> > + "through CXL endpoint must be alone below its root port",
> > + name, nendpoints);
> > + return false;
> > + }
> > + }
> > +
> > + matches = vfio_cxl_match_fmws(pxb, &base, &size, &nwindows,
> &ntargets);
> > + if (matches == 1 && ntargets == 1) {
> > + if (!base) {
> > + error_setg(errp,
> > + "vfio-cxl: %s: matched CFMWS has no base; reduce its "
> > + "size or grow the guest PA space", name);
> > + return false;
> > + }
> > + /*
> > + * QEMU reads the endpoint's HDM size from the kernel region, so the
> > + * window size is really device information, not a value to make the
> > + * caller guess. Sizing the CFMWS from the device the way firmware
> does
> > + * is not possible here: the window is placed in the guest PA map by
> > + * cxl_fmws_set_memmap() before this device is realized, so its size is
> > + * fixed before dpa_size is known. Until a core change can size the
> > + * window from the device, take the device geometry as the source of
> > + * truth: reject a window too small to hold the endpoint, and warn
> when
> > + * it is larger than needed, since the padding becomes guest CEDT that
> > + * migration has to preserve.
> > + */
> > + need = ROUND_UP(cxl->dpa_size, 256 * MiB);
> > + if (size < need) {
> > + error_setg(errp,
> > + "vfio-cxl: %s: CFMWS size 0x%" PRIx64 " cannot hold the "
> > + "endpoint; set the cxl-fmw size to 0x%" PRIx64,
> > + name, size, need);
> > + return false;
> > + }
> > + if (size > need) {
> > + warn_report("vfio-cxl: %s: CFMWS size 0x%" PRIx64 " exceeds the "
> > + "endpoint's 0x%" PRIx64 "; set the cxl-fmw size to 0x%"
> > + PRIx64 " to keep the guest memory map stable across "
> > + "migration", name, size, cxl->dpa_size, need);
> > + }
> > + cxl->fmws_base = base;
> > + cxl->fmws_size = size;
> > + return true;
> > + }
> > +
> > + if (matches == 1) {
> > + error_setg(errp,
> > + "vfio-cxl: %s: its CFMWS interleaves %d targets; use a "
> > + "single-target window", name, ntargets);
> > + } else if (matches > 1) {
> > + error_setg(errp,
> > + "vfio-cxl: %s: pxb-cxl is targeted by %d CFMWS; use one",
> > + name, matches);
> > + } else {
> > + error_setg(errp,
> > + "vfio-cxl: %s: no CFMWS targets this device (%d present)",
> > + name, nwindows);
> > + }
> > + return false;
> > +}
> > +
> > +/*
> > + * Cold-plug path: the CFMWS windows are placed at machine init done,
> > +so the
>
> are you sure of that ? I think fw->base is already set when vfio_cxl_setup()
> runs for cold-plug.
>
> Normal vfio-pci has no hotplug-specific code, vfio_pci_realize() runs the same
> path regardless of DEVICE(vdev)->hotplugged. The CXL case should be the
> same.
>
> > + * binding can only be validated from this notifier. The VM has not
> > +run yet, so
> > + * a configuration error is fatal to startup. A hotplugged device is
> > +validated
> > + * in realize instead (see vfio_cxl_setup), where the failure fails
> > +device_add
> > + * without taking down the running VM.
> > + */
> > +static void vfio_cxl_bind_fmws(Notifier *n, void *data) {
> > + VFIOCXL *cxl = container_of(n, VFIOCXL, machine_done);
> > + VFIOPCIDevice *vdev = container_of(cxl, VFIOPCIDevice, cxl);
> > + Error *err = NULL;
> > +
> > + if (!vfio_cxl_do_bind_fmws(vdev, &err)) {
> > + error_report_err(err);
> > + exit(1);
> > + }
> > +}
> > +
> > /*
> > * Learn the CXL geometry the kernel reports: the HPA-backed HDM memory
> region
> > * and the trapped HDM decoder block (which BAR carries it and at what
> offset).
> > @@ -3717,6 +3887,22 @@ static bool vfio_cxl_setup(VFIOPCIDevice *vdev,
> Error **errp)
> > "performance may be slow", vbasedev->name);
> > }
> >
> > + if (DEVICE(vdev)->hotplugged) {
> > + /*
> > + * The machine is already up, so the CFMWS windows are placed and
> the
> > + * binding can be validated now. Report a failure through errp so a bad
> > + * device_add fails cleanly instead of aborting the running VM.
> > + */
> > + if (!vfio_cxl_do_bind_fmws(vdev, errp)) {
> > + vfio_region_exit(&cxl->mem_region);
> > + vfio_region_finalize(&cxl->mem_region);
>
> as said before, the tear down has 2 parts. Anyhow, I don't think we need a
> CXL-specific hotplug patch.
>
I Agree this patch is doing too much. I will break this patch into topology
check, window match, size validation, and the suggested fixes in this
thread. I will start the commit message with the problem description and define
the relevant terms (e.g. CFMWS etc.), so a reviewer without CXL background can
follow it.
> C.
>
>
>
> > + return false;
> > + }
> > + } else {
> > + cxl->machine_done.notify = vfio_cxl_bind_fmws;
> > + qemu_add_machine_init_done_notifier(&cxl->machine_done);
> > + }
> > +
> > cxl->enabled = true;
> >
> > return true;
> > @@ -3729,6 +3915,12 @@ static void vfio_cxl_teardown(VFIOPCIDevice
> *vdev)
> > if (!cxl->enabled) {
> > return;
> > }
> > +
> > + if (cxl->machine_done.notify) {
> > + qemu_remove_machine_init_done_notifier(&cxl->machine_done);
> > + cxl->machine_done.notify = NULL;
> > + }
> > +
> > if (cxl->mem_region.mem) {
> > vfio_region_exit(&cxl->mem_region);
> > vfio_region_finalize(&cxl->mem_region);
> > diff --git a/hw/vfio/pci.h b/hw/vfio/pci.h index
> > 06d15e807d..22fe8d7ff7 100644
> > --- a/hw/vfio/pci.h
> > +++ b/hw/vfio/pci.h
> > @@ -135,6 +135,9 @@ typedef struct VFIOCXL {
> > uint64_t hdm_offset; /* block offset within the component BAR */
> > uint64_t dpa_size; /* size of the HDM memory region */
> > VFIORegion mem_region; /* HDM memory, mapped at committed
> GPA */
> > + Notifier machine_done; /* CFMWS validated at machine_done */
> > + hwaddr fmws_base; /* base of the memory window */
> > + uint64_t fmws_size; /* size of the memory window */
> > } VFIOCXL;
> >
> > struct VFIOPCIDevice {
^ permalink raw reply [flat|nested] 27+ messages in thread
* RE: [PATCH v2 07/10] hw/vfio/pci: Map the CXL memory on the guest decoder commit
2026-09-18 16:08 ` Cédric Le Goater
@ 2026-09-21 11:14 ` Manish Honap
0 siblings, 0 replies; 27+ messages in thread
From: Manish Honap @ 2026-09-21 11:14 UTC (permalink / raw)
To: Cédric Le Goater, alex@shazbot.org, Ankit Agrawal,
jic23@kernel.org, dave.jiang@intel.com,
alejandro.lucero-palau@amd.com, Srirangan Madhavan,
pierrick.bouvier@oss.qualcomm.com, mst@redhat.com,
imammedo@redhat.com, anisinha@redhat.com, pbonzini@redhat.com,
eric.auger@redhat.com, peter.maydell@linaro.org,
richard.henderson@linaro.org, cohuck@redhat.com
Cc: Krishnakant Jaju, Vikram Sethi, Zhi Wang, qemu-devel@nongnu.org,
qemu-arm@nongnu.org, linux-cxl@vger.kernel.org, Manish Honap
> -----Original Message-----
> From: Cédric Le Goater <clg@redhat.com>
> Sent: Friday, September 18, 2026 9:39 PM
> To: Manish Honap <mhonap@nvidia.com>; alex@shazbot.org; Ankit Agrawal
> <ankita@nvidia.com>; jic23@kernel.org; dave.jiang@intel.com;
> alejandro.lucero-palau@amd.com; Srirangan Madhavan
> <smadhavan@nvidia.com>; pierrick.bouvier@oss.qualcomm.com;
> mst@redhat.com; imammedo@redhat.com; anisinha@redhat.com;
> pbonzini@redhat.com; eric.auger@redhat.com; peter.maydell@linaro.org;
> richard.henderson@linaro.org; cohuck@redhat.com
> Cc: Krishnakant Jaju <kjaju@nvidia.com>; Vikram Sethi <vsethi@nvidia.com>;
> Zhi Wang <zhiw@nvidia.com>; qemu-devel@nongnu.org; qemu-
> arm@nongnu.org; linux-cxl@vger.kernel.org
> Subject: Re: [PATCH v2 07/10] hw/vfio/pci: Map the CXL memory on the
> guest decoder commit
>
> External email: Use caution opening links or attachments
>
>
> On 9/16/26 20:44, mhonap@nvidia.com wrote:
> > From: Manish Honap <mhonap@nvidia.com>
> >
> > Overlay the trapped HDM decoder block over the component BAR and
> > forward its accesses to the kernel. QEMU maps the HDM memory at the
> > base of the device's CFMWS window and presents that same base back to
> > the guest through the trapped decoder read
> >
> > Overlap the CFMWS window MemoryRegion the CXL host bridge maps at
> that
> > base, at a higher priority, so precedence is defined by priority
> > rather than subregion insertion order.
> >
> > QEMU always maps at the CFMWS base, so a decoder the guest commits at
> > any other base would leave the guest's view and the mapping diverged.
> > Track the base the guest writes to the trapped block and refuse to map
> > a commit at a different base; a firmware-committed decoder the guest
> > never reprograms keeps the window base.
> >
> > AI-used-for: code (prototype)
> > Signed-off-by: Manish Honap <mhonap@nvidia.com>
> > ---
> > hw/cxl/cxl-host-stubs.c | 5 +
> > hw/vfio/pci.c | 386
> +++++++++++++++++++++++++++++++++++++++-
> > hw/vfio/pci.h | 8 +
> > 3 files changed, 393 insertions(+), 6 deletions(-)
>
>
> This patch introduces too many concepts at once, a reader familiar with VFIO
> can't separate the plumbing from the CXL details. This needs more commits.
>
> So, something like
>
> 1. framework
> Set up the comp_regs_region
> fd pass-through (pread/pwrite), no virtualization.
> teardown
>
> 2. Virtualization of decoder base registers
>
> 3. Map/unmap of HDM memory
>
> if fmws_base never changes, there's no reason to remove and
> re-add the subregion. Install it once and toggle with
> memory_region_set_enabled(). this should simplify the decoder
> handling.
>
> 4. Config write hooks
> PCI_COMMAND, CXL reset, PM transition, FLR x 2
> each could be a separate patch I think
>
> I don't think we need the machine notifier as said in the previous patch.
>
> More below,
>
> >
> > diff --git a/hw/cxl/cxl-host-stubs.c b/hw/cxl/cxl-host-stubs.c index
> > 9b515913ea..1067c944e2 100644
> > --- a/hw/cxl/cxl-host-stubs.c
> > +++ b/hw/cxl/cxl-host-stubs.c
> > @@ -23,3 +23,8 @@ GSList *cxl_fmws_get_all_sorted(void)
> > {
> > g_assert_not_reached();
> > }
> > +
> > +int cxl_decoder_count_dec(int enc_cnt) {
> > + g_assert_not_reached();
> > +}
> > diff --git a/hw/vfio/pci.c b/hw/vfio/pci.c index
> > 670e0d1da4..fb3c39d4c6 100644
> > --- a/hw/vfio/pci.c
> > +++ b/hw/vfio/pci.c
> > @@ -1453,6 +1453,39 @@ uint32_t vfio_pci_read_config(PCIDevice *pdev,
> uint32_t addr, int len)
> > return val;
> > }
> >
> > +static void vfio_cxl_decoder_changed(VFIOPCIDevice *vdev); static
> > +void vfio_cxl_unmap_mem(VFIOPCIDevice *vdev);
> > +
> > +/*
> > + * Return true if this config write sets a Function Level Reset bit
> > +the kernel
> > + * acts on: the PCIe Device Control FLR bit or the Advanced Features FLR bit.
> > + * FLR bits are write-1-to-trigger and self-clearing, so inspect the
> > +written
> > + * value rather than the post-write config.
> > + */
> > +static bool vfio_cxl_flr_write(VFIOPCIDevice *vdev, uint32_t addr,
> > + uint32_t val, int len) {
> > + PCIDevice *pdev = &vdev->parent_obj;
> > + uint32_t off;
> > +
> > + if (pdev->exp.exp_cap) {
> > + /* PCI_EXP_DEVCTL_BCR_FLR (bit 15) is the high byte of Device Ctrl. */
> > + off = pdev->exp.exp_cap + PCI_EXP_DEVCTL + 1;
> > + if (off >= addr && off < addr + len &&
> > + ((val >> (8 * (off - addr))) & (PCI_EXP_DEVCTL_BCR_FLR >> 8))) {
> > + return true;
> > + }
> > + }
> > + if (vdev->cxl.af_offset) {
> > + off = vdev->cxl.af_offset + PCI_AF_CTRL;
> > + if (off >= addr && off < addr + len &&
> > + ((val >> (8 * (off - addr))) & PCI_AF_CTRL_FLR)) {
> > + return true;
> > + }
> > + }
> > + return false;
> > +}
> > +
> > void vfio_pci_write_config(PCIDevice *pdev,
> > uint32_t addr, uint32_t val, int len)
> > {
> > @@ -1522,9 +1555,74 @@ void vfio_pci_write_config(PCIDevice *pdev,
> > vfio_sub_page_bar_update_mapping(pdev, bar);
> > }
> > }
> > +
> > + /*
> > + * A firmware-committed CXL decoder is already committed at boot, so
> no
> > + * guest control write triggers the HDM mapping. Map it once the guest
> > + * enables memory decoding, so the region enters the guest address
> space
> > + * and the IOAS while the device is live. On a clear of memory decoding,
> > + * withdraw the overlay: the kernel revokes the HDM PTEs on the same
> > + * write, so a retained memslot would fault a guest access during the
> > + * disabled interval onto a zapped VMA and stop the VM; the enable
> path
> > + * re-maps it.
> > + */
> > + if (vdev->cxl.enabled &&
> > + range_covers_byte(addr, len, PCI_COMMAND)) {
> > + if (pci_get_word(pdev->config + PCI_COMMAND) &
> PCI_COMMAND_MEMORY) {
> > + vfio_cxl_decoder_changed(vdev);
> > + } else {
> > + vfio_cxl_unmap_mem(vdev);
> > + }
> > + }
> > } else {
> > /* Write everything to QEMU to keep emulated bits correct */
> > pci_default_write_config(pdev, addr, val, len);
> > +
> > + /*
> > + * A guest CXL reset is a write to the CXL Device DVSEC ctrl2 register,
> > + * forwarded to the kernel above, which re-commits this firmware-fixed
> > + * decoder. If the guest decommitted the decoder before the reset,
> which
> > + * unmapped the HDM memory, that re-commit is not otherwise visible
> to
> > + * QEMU, so rescan the decoder here to restore the mapping. Gated on
> > + * memory decoding still being enabled, like the enable path above.
> > + */
> > + if (vdev->cxl.enabled && vdev->cxl.dvsec_offset &&
> > + ranges_overlap(addr, len,
> > + vdev->cxl.dvsec_offset +
> > + offsetof(CXLDVSECDevice, ctrl2),
> > + sizeof_field(CXLDVSECDevice, ctrl2)) &&
> > + (pci_get_word(pdev->config + PCI_COMMAND) &
> PCI_COMMAND_MEMORY)) {
> > + vfio_cxl_decoder_changed(vdev);
> > + }
> > +
> > + /*
> > + * A NoSoftRst- device the guest cycles D3hot->D0 has its physical
> > + * decoder restored by the kernel on resume, which re-commits this
> > + * firmware-fixed decoder just like a reset. That re-commit is not
> > + * otherwise visible to QEMU, so on the transition back to D0 rescan
> > + * the decoder to restore a mapping the guest dropped before it
> > + * suspended. Gated on memory decoding, like the paths above.
> > + */
> > + if (vdev->cxl.enabled && pdev->pm_cap &&
> > + range_covers_byte(addr, len, pdev->pm_cap + PCI_PM_CTRL) &&
> > + (pci_get_word(pdev->config + pdev->pm_cap + PCI_PM_CTRL) &
> > + PCI_PM_CTRL_STATE_MASK) == 0 &&
> > + (pci_get_word(pdev->config + PCI_COMMAND) &
> PCI_COMMAND_MEMORY)) {
> > + vfio_cxl_decoder_changed(vdev);
> > + }
> > +
> > + /*
> > + * A Function Level Reset the guest triggers through the PCIe Device
> > + * Control or Advanced Features FLR bit is forwarded to the kernel,
> > + * which re-commits this firmware-fixed decoder just like a CXL reset.
> > + * That re-commit is not otherwise visible to QEMU, so rescan the
> > + * decoder to restore a mapping the guest dropped before the FLR.
> Gated
> > + * on memory decoding, like the paths above.
> > + */
> > + if (vdev->cxl.enabled && vfio_cxl_flr_write(vdev, addr, val, len) &&
> > + (pci_get_word(pdev->config + PCI_COMMAND) &
> PCI_COMMAND_MEMORY)) {
> > + vfio_cxl_decoder_changed(vdev);
> > + }
> > }
> > }
>
>
> That's a lot of CXL-specific code injected into vfio_pci_write_config.
> Don't do that. Please add :
>
> if (vdev->cxl.enabled) {
> vfio_cxl_config_written(vdev, addr, val, len);
> }
>
> I think that we should have a vfio/cxl.c subcomponent to isolate the CXL-
> specific code :
>
> vfio_cxl_setup, or better vfio_cxl_realize
> vfio_cxl_config_written,
> vfio_cxl_exit,
> vfio_cxl_finalize,
> etc.
>
>
>
> > @@ -3811,6 +3909,239 @@ static void vfio_cxl_bind_fmws(Notifier *n,
> void *data)
> > }
> > }
> >
> > +/*
> > + * HDM decoder registers, relative to the trapped decoder block. The
> > +block holds
> > + * the HDM Decoder Capability register followed by one register set
> > +per decoder,
> > + * each VFIO_CXL_HDM_DECODER_STRIDE apart, indexed by decoder.
> > + */
> > +#define VFIO_CXL_HDM_CAP 0x00
> > +#define VFIO_CXL_HDM_DECODER_STRIDE 0x20
> > +#define VFIO_CXL_HDM_DECODER_BASE_LOW(n) \
> > + (0x10 + (n) * VFIO_CXL_HDM_DECODER_STRIDE) #define
> > +VFIO_CXL_HDM_DECODER_BASE_HIGH(n) \
> > + (0x14 + (n) * VFIO_CXL_HDM_DECODER_STRIDE) #define
> > +VFIO_CXL_HDM_DECODER_CTRL(n) \
> > + (0x20 + (n) * VFIO_CXL_HDM_DECODER_STRIDE)
> > +#define VFIO_CXL_HDM_CTRL_COMMITTED (1 << 10)
> > +#define VFIO_CXL_HDM_BASE_LOW_MASK 0xf0000000U
> > +
> > +static void vfio_cxl_unmap_mem(VFIOPCIDevice *vdev) {
> > + VFIOCXL *cxl = &vdev->cxl;
> > +
> > + if (!cxl->dpa_mapped) {
> > + return;
> > + }
> > + memory_region_del_subregion(get_system_memory(), cxl-
> >mem_region.mem);
> > + cxl->dpa_mapped = false;
> > + cxl->mapped_base = 0;
> > +}
> > +
> > +/*
> > + * The HDM Decoder Capability register encodes the decoder count in
> > +its low
> > + * nibble (CXL r3.1 8.2.4.20.1, encodings 0h..Ch for 1..32). Decode
> > +it with the
> > + * shared cxl_decoder_count_dec() helper so the commit handler walks
> > +every
> > + * decoder rather than assuming decoder 0. An unreadable register or
> > +a reserved
> > + * encoding (which the helper decodes to 0) is treated as a single decoder.
> > + */
> > +static unsigned vfio_cxl_decoder_count(VFIOPCIDevice *vdev) {
> > + off_t off = vdev->cxl.comp_regs_region.fd_offset;
> > + uint32_t cap = 0;
> > + int count;
> > +
> > + if (pread(vdev->vbasedev.fd, &cap, 4, off + VFIO_CXL_HDM_CAP) != 4) {
> > + return 1;
> > + }
> > + count = cxl_decoder_count_dec(le32_to_cpu(cap) & 0xf);
> > + return count > 0 ? count : 1;
> > +}
> > +
> > +/*
> > + * The guest committed or tore down an endpoint decoder. Walk each
> > +decoder in
> > + * the trapped HDM block and (un)map the HDM memory at the CFMWS
> > +window base,
> > + * not the base the guest programmed (see vfio_cxl_comp_regs_read).
> > +This
> > + * generation commits one non-interleaved decoder, so the walk stops
> > +at the
> > + * first committed decoder.
>
> if only one decoder is supported and the code tracks the guest base per-device
> rather than per-decoder, why do the read/write handlers use
> VFIO_CXL_HDM_DECODER_BASE_LOW(0) to match all decoders ? This is
> inconsistent.
>
> I think we should record the active decoder index 'active_decoder'
> instead. If a second committed decoder is found, decoder_changed should
> bail out.
>
>
> > + */
> > +static void vfio_cxl_decoder_changed(VFIOPCIDevice *vdev) {
> > + VFIOCXL *cxl = &vdev->cxl;
> > + VFIODevice *vbasedev = &vdev->vbasedev;
> > + off_t off = cxl->comp_regs_region.fd_offset;
> > + unsigned n, count = vfio_cxl_decoder_count(vdev);
> > +
> > + for (n = 0; n < count; n++) {
> > + uint32_t ctrl = 0;
> > +
> > + if (pread(vbasedev->fd, &ctrl, 4,
> > + off + VFIO_CXL_HDM_DECODER_CTRL(n)) != 4) {
> > + return;
> > + }
> > + if (!(le32_to_cpu(ctrl) & VFIO_CXL_HDM_CTRL_COMMITTED)) {
> > + continue;
> > + }
> > +
> > + /*
> > + * QEMU always maps at the device's CFMWS base and presents that
> base
> > + * back to the guest, so a decoder the guest committed at any other
> base
> > + * would leave the guest's view and the actual mapping diverged. Reject
> > + * it rather than silently relocate. A firmware-committed decoder the
> > + * guest never reprogrammed (guest_base_written == false) keeps the
> > + * CFMWS base.
> > + */
> > + if (cxl->guest_base_written) {
> > + hwaddr guest_base = ((hwaddr)cxl->guest_base_hi << 32) |
> > + cxl->guest_base_lo;
> > +
> > + if (guest_base != cxl->fmws_base) {
> > + warn_report("vfio-cxl: %s: guest committed decoder %u at 0x%"
> > + HWADDR_PRIx ", not the CFMWS base 0x%" HWADDR_PRIx
> > + "; not mapping", vbasedev->name, n, guest_base,
> > + cxl->fmws_base);
> > + vfio_cxl_unmap_mem(vdev);
> > + return;
> > + }
> > + }
> > +
> > + if (cxl->dpa_mapped && cxl->mapped_base == cxl->fmws_base) {
> > + return;
> > + }
> > + /*
> > + * Overlap the CFMWS window MemoryRegion the CXL host bridge
> already
> > + * maps at this base, at a higher priority, so precedence is defined by
> > + * priority rather than subregion insertion order.
> > + */
> > + memory_region_transaction_begin();
> > + vfio_cxl_unmap_mem(vdev);
> > + memory_region_add_subregion_overlap(get_system_memory(),
> > + cxl->fmws_base,
> > + cxl->mem_region.mem, 1);
> > + memory_region_transaction_commit();
> > + cxl->mapped_base = cxl->fmws_base;
> > + cxl->dpa_mapped = true;
> > + return;
> > + }
> > +
> > + /* No committed decoder in the block: tear down any existing mapping.
> */
> > + vfio_cxl_unmap_mem(vdev);
> > +}
> > +
> > +static uint64_t vfio_cxl_comp_regs_read(void *opaque, hwaddr addr,
> > + unsigned size) {
> > + VFIORegion *region = opaque;
> > + VFIODevice *vbasedev = region->vbasedev;
> > + VFIOPCIDevice *vdev =
> > + container_of(region, VFIOPCIDevice, cxl.comp_regs_region);
> > + VFIOCXL *cxl = &vdev->cxl;
> > + uint32_t val = 0xffffffff;
> > +
> > + if (pread(vbasedev->fd, &val, size, region->fd_offset + addr) != size) {
> > + error_report("vfio-cxl: %s: HDM decoder read at 0x%" HWADDR_PRIx
> > + " failed", vbasedev->name, addr);
> > + }
> > + val = le32_to_cpu(val);
> > +
> > + /*
> > + * The kernel shadow holds the host physical base for a firmware-
> committed
> > + * decoder. Never expose that to the guest: present the guest physical
> base
> > + * QEMU maps the HDM memory at (the device's CFMWS window). Any
> decoder's
> > + * base registers are virtualized; ctrl, size and the capability header pass
> > + * through. The base registers repeat every
> VFIO_CXL_HDM_DECODER_STRIDE.
> > + */
> > + if (addr >= VFIO_CXL_HDM_DECODER_BASE_LOW(0) &&
> > + (addr - VFIO_CXL_HDM_DECODER_BASE_LOW(0)) %
> > + VFIO_CXL_HDM_DECODER_STRIDE == 0) {
> > + val = (val & ~VFIO_CXL_HDM_BASE_LOW_MASK) |
> > + ((uint32_t)cxl->fmws_base & VFIO_CXL_HDM_BASE_LOW_MASK);
> > + } else if (addr >= VFIO_CXL_HDM_DECODER_BASE_HIGH(0) &&
> > + (addr - VFIO_CXL_HDM_DECODER_BASE_HIGH(0)) %
> > + VFIO_CXL_HDM_DECODER_STRIDE == 0) {
> > + val = (uint32_t)(cxl->fmws_base >> 32);
> > + }
> > +
> > + return val;
> > +}
> > +
> > +static void vfio_cxl_comp_regs_write(void *opaque, hwaddr addr, uint64_t
> data,
> > + unsigned size) {
> > + VFIORegion *region = opaque;
> > + VFIODevice *vbasedev = region->vbasedev;
> > + VFIOPCIDevice *vdev =
> > + container_of(region, VFIOPCIDevice, cxl.comp_regs_region);
> > + uint32_t val = cpu_to_le32((uint32_t)data);
> > +
> > + if (pwrite(vbasedev->fd, &val, size, region->fd_offset + addr) != size) {
> > + error_report("vfio-cxl: %s: HDM decoder write at 0x%" HWADDR_PRIx
> > + " failed", vbasedev->name, addr);
> > + return;
> > + }
> > +
> > + /*
> > + * Record the base the guest programs into a decoder so the commit
> handler
> > + * can reject a base other than the device's CFMWS window. The base low
> > + * register carries HPA bits [31:28]; the high register carries [63:32].
> > + */
> > + if (addr >= VFIO_CXL_HDM_DECODER_BASE_LOW(0) &&
> > + (addr - VFIO_CXL_HDM_DECODER_BASE_LOW(0)) %
> > + VFIO_CXL_HDM_DECODER_STRIDE == 0) {
> > + vdev->cxl.guest_base_lo = (uint32_t)data &
> VFIO_CXL_HDM_BASE_LOW_MASK;
> > + vdev->cxl.guest_base_written = true;
> > + } else if (addr >= VFIO_CXL_HDM_DECODER_BASE_HIGH(0) &&
> > + (addr - VFIO_CXL_HDM_DECODER_BASE_HIGH(0)) %
> > + VFIO_CXL_HDM_DECODER_STRIDE == 0) {
> > + vdev->cxl.guest_base_hi = (uint32_t)data;
> > + vdev->cxl.guest_base_written = true;
> > + }
> > +
> > + /*
> > + * The kernel runs the lock-on-commit FSM in the write above, so the
> > + * committed state is settled by now; a control write on any decoder can
> > + * change the mapping.
> > + */
> > + if (addr >= VFIO_CXL_HDM_DECODER_CTRL(0) &&
> > + (addr - VFIO_CXL_HDM_DECODER_CTRL(0)) %
> > + VFIO_CXL_HDM_DECODER_STRIDE == 0) {
> > + vfio_cxl_decoder_changed(vdev);
> > + }
> > +}
> > +
> > +static const MemoryRegionOps vfio_cxl_comp_regs_ops = {
> > + .read = vfio_cxl_comp_regs_read,
> > + .write = vfio_cxl_comp_regs_write,
> > + .endianness = DEVICE_LITTLE_ENDIAN,
> > + .valid = { .min_access_size = 4, .max_access_size = 4 },
> > + .impl = { .min_access_size = 4, .max_access_size = 4 }, };
> > +
> > +/*
> > + * Locate the CXL Device DVSEC (CXL r3.1 8.1.3) in config space. The
> > +guest
> > + * triggers a CXL reset by writing its ctrl2 register; QEMU rescans
> > +the decoder
> > + * after that write so the HDM mapping is restored (see
> vfio_pci_write_config).
> > + * Returns the DVSEC config offset, or 0 if the device does not expose it.
> > + */
> > +static uint16_t vfio_cxl_find_device_dvsec(PCIDevice *pdev) {
> > + uint16_t offset;
> > +
> > + for (offset = PCI_CONFIG_SPACE_SIZE; offset;
> > + offset = PCI_EXT_CAP_NEXT(pci_get_long(pdev->config + offset))) {
> > + uint32_t hdr = pci_get_long(pdev->config + offset);
> > +
> > + if (PCI_EXT_CAP_ID(hdr) == PCI_EXT_CAP_ID_DVSEC &&
> > + (pci_get_long(pdev->config + offset + PCI_DVSEC_HEADER1) & 0xffff)
> > + == CXL_VENDOR_ID &&
> > + pci_get_word(pdev->config + offset + PCI_DVSEC_HEADER2)
> > + == PCIE_CXL_DEVICE_DVSEC) {
> > + return offset;
> > + }
> > + }
> > +
> > + return 0;
> > +}
>
> This should be a PCIe helper :
>
> uint16_t pcie_find_dvsec(PCIDevice *dev, uint16_t vendor_id, uint16_t
> dvsec_id);
>
Okay, I will refactor this patch along the lines you suggested:
- framework (comp_regs_region setup, fd pass-through, teardown),
- base register virtualization,
- map and unmap of HDM memory,
- config write hooks for PCI_COMMAND, CXL reset, PM transition, and FLR.
map and unmap: I will install the subregion once and toggle it with
memory_region_set_enabled() instead of add and delete on every change.
Decoder handling: I will record the active decoder index and virtualize and
map for that decoder, and have decoder_changed bail out if it finds a second
committed decoder, since this patch series supports only one.
Rename vfio_cxl_find_device_dvsec() as pcie_find_dvsec(dev, vendor_id, dvsec_id)
in the PCIe code, so it is reusable.
Remove the machine_done notifier as suggested.
> Thanks,
>
> C.
>
>
> > /*
> > * Learn the CXL geometry the kernel reports: the HPA-backed HDM memory
> region
> > * and the trapped HDM decoder block (which BAR carries it and at what
> offset).
> > @@ -3864,6 +4195,8 @@ static bool vfio_cxl_setup(VFIOPCIDevice *vdev,
> Error **errp)
> > cxl->dpa_size = mem_info->size;
> > cxl->comp_bar = cap->bar;
> > cxl->hdm_offset = cap->offset;
> > + cxl->dvsec_offset = vfio_cxl_find_device_dvsec(&vdev->parent_obj);
> > + cxl->af_offset = pci_find_capability(&vdev->parent_obj,
> > + PCI_CAP_ID_AF);
> >
> > if (!vfio_cxl_check_topology(vdev, errp)) {
> > return false;
> > @@ -3887,6 +4220,26 @@ static bool vfio_cxl_setup(VFIOPCIDevice *vdev,
> Error **errp)
> > "performance may be slow", vbasedev->name);
> > }
> >
> > + /*
> > + * Trap the HDM decoder block: overlay it, priority 1, over the directly
> > + * mapped component BAR, so guest decoder accesses reach the kernel
> FSM and
> > + * the commit becomes visible to QEMU.
> > + */
> > + if (!vdev->bars[cxl->comp_bar].mr) {
> > + error_setg(errp, "vfio-cxl: %s: component BAR %u is not present",
> > + vbasedev->name, cxl->comp_bar);
> > + goto err;
> > + }
> > + if (vfio_region_setup_with_ops(OBJECT(vdev), vbasedev,
> > + &cxl->comp_regs_region,
> > + cxl->comp_regs_region_index, "cxl-comp-regs",
> > + &vfio_cxl_comp_regs_ops, errp)) {
> > + goto err;
> > + }
> > + memory_region_add_subregion_overlap(vdev->bars[cxl->comp_bar].mr,
> > + cxl->hdm_offset,
> > + cxl->comp_regs_region.mem,
> > + 1);
> > +
> > if (DEVICE(vdev)->hotplugged) {
> > /*
> > * The machine is already up, so the CFMWS windows are
> > placed and the @@ -3894,9 +4247,7 @@ static bool
> vfio_cxl_setup(VFIOPCIDevice *vdev, Error **errp)
> > * device_add fails cleanly instead of aborting the running VM.
> > */
> > if (!vfio_cxl_do_bind_fmws(vdev, errp)) {
> > - vfio_region_exit(&cxl->mem_region);
> > - vfio_region_finalize(&cxl->mem_region);
> > - return false;
> > + goto err;
> > }
> > } else {
> > cxl->machine_done.notify = vfio_cxl_bind_fmws; @@ -3906,6
> > +4257,10 @@ static bool vfio_cxl_setup(VFIOPCIDevice *vdev, Error **errp)
> > cxl->enabled = true;
> >
> > return true;
> > +
> > +err:
> > + vfio_cxl_teardown(vdev);
> > + return false;
> > }
> >
> > static void vfio_cxl_teardown(VFIOPCIDevice *vdev) @@ -3916,15
> > +4271,26 @@ static void vfio_cxl_teardown(VFIOPCIDevice *vdev)
> > return;
> > }
> >
> > - if (cxl->machine_done.notify) {
> > - qemu_remove_machine_init_done_notifier(&cxl->machine_done);
> > - cxl->machine_done.notify = NULL;
> > + if (cxl->comp_regs_region.mem) {
> > + if (vdev->bars[cxl->comp_bar].mr) {
> > + memory_region_del_subregion(vdev->bars[cxl->comp_bar].mr,
> > + cxl->comp_regs_region.mem);
> > + }
> > + vfio_region_exit(&cxl->comp_regs_region);
> > + vfio_region_finalize(&cxl->comp_regs_region);
> > }
> >
> > + vfio_cxl_unmap_mem(vdev);
> > +
> > if (cxl->mem_region.mem) {
> > vfio_region_exit(&cxl->mem_region);
> > vfio_region_finalize(&cxl->mem_region);
> > }
> > +
> > + if (cxl->machine_done.notify) {
> > + qemu_remove_machine_init_done_notifier(&cxl->machine_done);
> > + cxl->machine_done.notify = NULL;
> > + }
> > }
> >
> > static void vfio_pci_realize(PCIDevice *pdev, Error **errp) @@
> > -4127,6 +4493,14 @@ static void vfio_exitfn(PCIDevice *pdev)
> > vfio_pci_teardown_msi(vdev);
> > vfio_pci_disable_rp_atomics(vdev);
> > vfio_pci_bars_exit(vdev);
> > + /*
> > + * The committed HDM overlay is a subregion of system memory owned
> by this
> > + * device, so it holds a reference that would keep the object alive past
> > + * unrealize and block instance_finalize (where vfio_cxl_teardown
> otherwise
> > + * runs). Drop it here; the call is idempotent for a device that never
> > + * mapped or is not CXL.
> > + */
> > + vfio_cxl_unmap_mem(vdev);
> > vfio_migration_exit(vbasedev);
> > if (!vbasedev->mdev) {
> > pci_device_unset_iommu_device(pdev);
> > diff --git a/hw/vfio/pci.h b/hw/vfio/pci.h index
> > 22fe8d7ff7..5bfa976112 100644
> > --- a/hw/vfio/pci.h
> > +++ b/hw/vfio/pci.h
> > @@ -133,11 +133,19 @@ typedef struct VFIOCXL {
> > uint32_t comp_regs_region_index; /* trapped HDM decoder register
> block */
> > uint32_t comp_bar; /* component BAR carrying that block */
> > uint64_t hdm_offset; /* block offset within the component BAR */
> > + uint16_t dvsec_offset; /* CXL Device DVSEC offset, 0 if none */
> > + uint16_t af_offset; /* PCI Advanced Features cap, 0 if none */
> > uint64_t dpa_size; /* size of the HDM memory region */
> > VFIORegion mem_region; /* HDM memory, mapped at committed
> GPA */
> > Notifier machine_done; /* CFMWS validated at machine_done */
> > hwaddr fmws_base; /* base of the memory window */
> > uint64_t fmws_size; /* size of the memory window */
> > + VFIORegion comp_regs_region; /* trapped HDM decoder block */
> > + hwaddr mapped_base; /* GPA the HDM memory is mapped at */
> > + bool dpa_mapped; /* HDM memory currently in system memory
> */
> > + uint32_t guest_base_lo; /* decoder base low the guest wrote */
> > + uint32_t guest_base_hi; /* decoder base high the guest wrote */
> > + bool guest_base_written; /* the guest wrote a decoder base */
> > } VFIOCXL;
> >
> > struct VFIOPCIDevice {
^ permalink raw reply [flat|nested] 27+ messages in thread
* RE: [PATCH v2 07/10] hw/vfio/pci: Map the CXL memory on the guest decoder commit
2026-09-20 9:14 ` Junjie Cao
@ 2026-09-21 11:17 ` Manish Honap
0 siblings, 0 replies; 27+ messages in thread
From: Manish Honap @ 2026-09-21 11:17 UTC (permalink / raw)
To: Junjie Cao
Cc: alex@shazbot.org, Ankit Agrawal, jic23@kernel.org,
dave.jiang@intel.com, alejandro.lucero-palau@amd.com,
Srirangan Madhavan, pierrick.bouvier@oss.qualcomm.com,
mst@redhat.com, imammedo@redhat.com, anisinha@redhat.com,
pbonzini@redhat.com, eric.auger@redhat.com,
peter.maydell@linaro.org, richard.henderson@linaro.org,
clg@redhat.com, cohuck@redhat.com, Krishnakant Jaju, Vikram Sethi,
Zhi Wang, qemu-devel@nongnu.org, qemu-arm@nongnu.org,
linux-cxl@vger.kernel.org, Manish Honap
> -----Original Message-----
> From: Junjie Cao <junjie.cao@intel.com>
> Sent: Sunday, September 20, 2026 2:44 PM
> To: Manish Honap <mhonap@nvidia.com>
> Cc: alex@shazbot.org; Ankit Agrawal <ankita@nvidia.com>; jic23@kernel.org;
> dave.jiang@intel.com; alejandro.lucero-palau@amd.com; Srirangan
> Madhavan <smadhavan@nvidia.com>; pierrick.bouvier@oss.qualcomm.com;
> mst@redhat.com; imammedo@redhat.com; anisinha@redhat.com;
> pbonzini@redhat.com; eric.auger@redhat.com; peter.maydell@linaro.org;
> richard.henderson@linaro.org; clg@redhat.com; cohuck@redhat.com;
> Krishnakant Jaju <kjaju@nvidia.com>; Vikram Sethi <vsethi@nvidia.com>; Zhi
> Wang <zhiw@nvidia.com>; qemu-devel@nongnu.org; qemu-
> arm@nongnu.org; linux-cxl@vger.kernel.org
> Subject: Re: [PATCH v2 07/10] hw/vfio/pci: Map the CXL memory on the
> guest decoder commit
>
> External email: Use caution opening links or attachments
>
>
> On Thu, 17 Sep 2026 00:14:09 +0530, Manish Honap wrote:
> > + * Record the base the guest programs into a decoder so the commit
> handler
> > + * can reject a base other than the device's CFMWS window. The
> > + base low
> [...]
> > + vdev->cxl.guest_base_lo = (uint32_t)data &
> VFIO_CXL_HDM_BASE_LOW_MASK;
> > + vdev->cxl.guest_base_written = true;
>
> This answered my v1 question for the kernel v4 FSM model. With v5
> (20/27) the kernel gives a live read-only view and leaves any write
> virtualization to the VMM, and what QEMU shows the guest is a committed,
> locked decoder at the window base. A locked decoder ignores a base write,
> but here it is recorded, and if it differs from the window base the next rescan
> unmaps while reads still return Committed at the window base. That is the
> mismatch the check was meant to prevent. A current Linux guest never writes
> a locked decoder, so it won't trip this.
>
> Is the guest-written base still meant to matter in this model? If not,
> guest_base_{lo,hi,written} can go.
>
> > + * The kernel runs the lock-on-commit FSM in the write above, so the
> > + * committed state is settled by now; a control write on any
> > + decoder can
>
> Stale now that the kernel no longer acts on the write?
Yes, you are right on both parts. I will drop guest_base_lo, guest_base_hi,
and guest_base_written together with the mismatch check and, reword
the control write comment to correctly match the code.
>
> Junjie
^ permalink raw reply [flat|nested] 27+ messages in thread
* RE: [PATCH v2 10/10] hw/pci-host: Emit a _DSM on pxb-cxl to preserve firmware PCI config
2026-09-20 9:14 ` Junjie Cao
@ 2026-09-21 11:20 ` Manish Honap
0 siblings, 0 replies; 27+ messages in thread
From: Manish Honap @ 2026-09-21 11:20 UTC (permalink / raw)
To: Junjie Cao
Cc: alex@shazbot.org, Ankit Agrawal, jic23@kernel.org,
dave.jiang@intel.com, alejandro.lucero-palau@amd.com,
Srirangan Madhavan, pierrick.bouvier@oss.qualcomm.com,
mst@redhat.com, imammedo@redhat.com, anisinha@redhat.com,
pbonzini@redhat.com, eric.auger@redhat.com,
peter.maydell@linaro.org, richard.henderson@linaro.org,
clg@redhat.com, cohuck@redhat.com, Krishnakant Jaju, Vikram Sethi,
Zhi Wang, qemu-devel@nongnu.org, qemu-arm@nongnu.org,
linux-cxl@vger.kernel.org, Shameer Kolothum Thodi, Manish Honap
> -----Original Message-----
> From: Junjie Cao <junjie.cao@intel.com>
> Sent: Sunday, September 20, 2026 2:44 PM
> To: Manish Honap <mhonap@nvidia.com>
> Cc: alex@shazbot.org; Ankit Agrawal <ankita@nvidia.com>; jic23@kernel.org;
> dave.jiang@intel.com; alejandro.lucero-palau@amd.com; Srirangan
> Madhavan <smadhavan@nvidia.com>; pierrick.bouvier@oss.qualcomm.com;
> mst@redhat.com; imammedo@redhat.com; anisinha@redhat.com;
> pbonzini@redhat.com; eric.auger@redhat.com; peter.maydell@linaro.org;
> richard.henderson@linaro.org; clg@redhat.com; cohuck@redhat.com;
> Krishnakant Jaju <kjaju@nvidia.com>; Vikram Sethi <vsethi@nvidia.com>; Zhi
> Wang <zhiw@nvidia.com>; qemu-devel@nongnu.org; qemu-
> arm@nongnu.org; linux-cxl@vger.kernel.org; Shameer Kolothum Thodi
> <skolothumtho@nvidia.com>
> Subject: Re: [PATCH v2 10/10] hw/pci-host: Emit a _DSM on pxb-cxl to
> preserve firmware PCI config
>
> External email: Use caution opening links or attachments
>
>
> On Thu, 17 Sep 2026 00:14:12 +0530, Manish Honap wrote:
> > under preserve_config, the _DSM. Gate the _DSM on preserve_config: x86
> > q35 passes false as it does not use the accelerated SMMU, so it emits
> > no _DSM and its ACPI tables stay byte for byte what they were.
>
> Confirmed: on 28e7aad522 plus the series bios-tables-test passes on
> x86_64 (q35/cxl and q35/acpihmat-genericx, the v1 failures, included),
> aarch64, riscv64 and loongarch64.
>
> accel=on is out of reach under TCG, so I forced preserve_config on virt locally
> and booted an arm64 guest (7.3-rc1 based, defconfig) with a
> cxl-type3 behind a cxl-rp. With this patch Linux keeps the firmware BARs and
> bridge window under the pxb-cxl; on 09/10 with the same hack it releases and
> reassigns them. No ACPI errors either way.
>
> Tested-by: Junjie Cao <junjie.cao@intel.com>
>
> 08/10 still says the x86 bridge "still emits the ``_DSM`` method, but its
> function 0 returns an empty support mask"; that was v1.
Oops, I missed this line. In current v2, _DSM is gated on preserve_config, so x86
q35 emits no _DSM at all rather than an empty support mask.
I will update this.
>
> Junjie
^ permalink raw reply [flat|nested] 27+ messages in thread
end of thread, other threads:[~2026-09-21 11:20 UTC | newest]
Thread overview: 27+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-16 18:44 [PATCH v2 00/10] QEMU: CXL Type-2 device passthrough via vfio-pci mhonap
2026-09-16 18:44 ` [PATCH v2 01/10] linux-headers: Update vfio.h for CXL Type-2 passthrough mhonap
2026-09-16 18:44 ` [PATCH v2 02/10] hw/vfio/region: Add vfio_region_setup_with_ops() mhonap
2026-09-17 13:27 ` Cédric Le Goater
2026-09-21 10:48 ` Manish Honap
2026-09-16 18:44 ` [PATCH v2 03/10] hw/vfio/pci: Detect a CXL Type-2 device and read its geometry mhonap
2026-09-17 16:04 ` Cédric Le Goater
2026-09-21 10:49 ` Manish Honap
2026-09-16 18:44 ` [PATCH v2 04/10] hw/vfio/pci: Enforce the passthrough topology for a CXL device mhonap
2026-09-17 13:20 ` Cédric Le Goater
2026-09-21 10:45 ` Manish Honap
2026-09-16 18:44 ` [PATCH v2 05/10] hw/vfio/pci: Back the CXL memory with a RAM-device region mhonap
2026-09-17 13:56 ` Cédric Le Goater
2026-09-21 10:55 ` Manish Honap
2026-09-16 18:44 ` [PATCH v2 06/10] hw/vfio/pci: Bind a CXL device to its fixed memory window mhonap
2026-09-17 16:00 ` Cédric Le Goater
2026-09-21 11:03 ` Manish Honap
2026-09-16 18:44 ` [PATCH v2 07/10] hw/vfio/pci: Map the CXL memory on the guest decoder commit mhonap
2026-09-18 16:08 ` Cédric Le Goater
2026-09-21 11:14 ` Manish Honap
2026-09-20 9:14 ` Junjie Cao
2026-09-21 11:17 ` Manish Honap
2026-09-16 18:44 ` [PATCH v2 08/10] docs/cxl: Document CXL Type-2 device passthrough mhonap
2026-09-16 18:44 ` [PATCH v2 09/10] hw/arm/smmu-common: Allow pxb-cxl as an SMMUv3 primary bus mhonap
2026-09-16 18:44 ` [PATCH v2 10/10] hw/pci-host: Emit a _DSM on pxb-cxl to preserve firmware PCI config mhonap
2026-09-20 9:14 ` Junjie Cao
2026-09-21 11:20 ` Manish Honap
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox