Linux CXL
 help / color / mirror / Atom feed
From: Chen Pei <cp0613@linux.alibaba.com>
To: palmer@dabbelt.com, alistair.francis@wdc.com, mst@redhat.com,
	imammedo@redhat.com, sunilvl@ventanamicro.com, jic23@kernel.org
Cc: pbonzini@redhat.com, liwei1518@gmail.com,
	daniel.barboza@oss.qualcomm.com, zhiwei_liu@linux.alibaba.com,
	chao.liu@processmission.com, anisinha@redhat.com,
	dave.jiang@intel.com, alison.schofield@intel.com,
	junjie.cao@intel.com, guoren@kernel.org, qemu-riscv@nongnu.org,
	qemu-devel@nongnu.org, linux-cxl@vger.kernel.org
Subject: [PATCH 3/4] hw/riscv/virt: Provide a 32-bit MMIO window for CXL host bridges
Date: Fri, 21 Aug 2026 16:19:53 +0800	[thread overview]
Message-ID: <20260821081954.1171-4-cp0613@linux.alibaba.com> (raw)
In-Reply-To: <20260821081954.1171-1-cp0613@linux.alibaba.com>

CXL component register BAR (BAR0 on CXL Root Port and Type3 device)
and the CXL device register BAR (BAR2 on Type3 device) are declared
as 64-bit non-prefetchable memory.  Per the PCI-to-PCI Bridge
Architecture Specification Rev 1.2 (PCI-SIG, 2003):

  - §3.2.5.8 (Memory Base/Limit): the non-prefetchable window covers
    only 32-bit addresses (AD[31:20]); the Type 1 header defines no
    upper-32-bit extension for it.
  - §3.2.5.9 (Prefetchable Memory Base/Limit): the bottom 4 bits
    encode 64-bit support (01h), but this applies exclusively to the
    *prefetchable* window.
  - §3.2.5.10 (Prefetchable Base/Limit Upper 32 Bits): optional
    registers for AD[63:32] of the prefetchable range only.

The architecture therefore allows a 64-bit window only when it is also
prefetchable; there is no 64-bit non-prefetchable form.  PCIe inherits
this Type 1 header layout unchanged.  Linux thus places 64-bit
non-prefetchable BARs in the 32-bit non-prefetchable bridge window,
which requires the bridge to own enough address space below 4 GiB.

On RISC-V virt the 32-bit PCIe MMIO range (1 GiB at 0x40000000) is
currently consumed entirely by PCI0, so CXL host bridges (ACPI0016)
have no non-prefetchable window and Linux fails to assign these BARs.

Reserve the top 256 MiB of the 32-bit MMIO window exclusively for CXL
host bridges, keeping all of the logic in the riscv code:

- Reduce the FDT 'ranges' for PCI0 by 256 MiB so that device-tree
  driven firmware does not hand that range to PCI0.
- The CXL host bridge _CRS is produced by the generic build_crs()
  path.  EDK2's PciBusDxe only recurses into PCI-to-PCI bridges, while
  the pxb-cxl expander bridge presents as a class 0x0600 host bridge
  with a type-0 header, so firmware never enumerates behind it and
  leaves the CXL root port's window and bus-number registers unset;
  build_crs() would therefore return an empty _CRS.  Simulate the
  firmware PCI initialization in virt_cxl_init_bridge_windows(): assign
  the reserved window to the root port's memory window, and allocate
  primary/secondary/subordinate bus numbers depth-first.  build_crs()
  then emits both the memory range and a correct bus-range resource
  (without a bus range Linux only claims the host-bridge bus and cannot
  enumerate the devices behind the root port), and the range is
  excluded from PCI0's _CRS.  A reset handler re-applies the state, as
  a PCI reset clears those registers.

This targets a single pxb-cxl (one CXL host bridge), which matches the
bios-tables test added later in the series.  With more than one
pxb-cxl the single reserved window would be claimed by every ACPI0016
bridge, so it would need to be partitioned per bridge; that is left as
future work (see the TODO in virt_cxl_init_bridge_windows()).

Verified by booting an RVA22 kernel: the CXL root port and Type3 device
are enumerated (0000:0c:00.0 / 0000:0d:00.0) and 'cxl list' reports the
memdev and the CFMWS root decoder.

Signed-off-by: Chen Pei <cp0613@linux.alibaba.com>
---
 hw/riscv/virt.c | 154 +++++++++++++++++++++++++++++++++++++++++++++++-
 1 file changed, 153 insertions(+), 1 deletion(-)

diff --git a/hw/riscv/virt.c b/hw/riscv/virt.c
index e08c76af4d..af83603293 100644
--- a/hw/riscv/virt.c
+++ b/hw/riscv/virt.c
@@ -48,11 +48,13 @@
 #include "chardev/char.h"
 #include "system/device_tree.h"
 #include "system/system.h"
+#include "system/reset.h"
 #include "system/tcg.h"
 #include "system/kvm.h"
 #include "system/tpm.h"
 #include "system/qtest.h"
 #include "hw/pci/pci.h"
+#include "hw/pci/pci_bridge.h"
 #include "hw/pci-host/gpex.h"
 #include "hw/display/ramfb.h"
 #include "hw/cxl/cxl.h"
@@ -115,6 +117,9 @@ static const MemMapEntry virt_memmap[] = {
 /* PCIe high mmio for RV64, size is fixed but base depends on top of RAM */
 #define VIRT64_HIGH_PCIE_MMIO_SIZE  (16 * GiB)
 
+/* 32-bit MMIO range carved out of VIRT_PCIE_MMIO for CXL host bridges */
+#define VIRT_CXL_MMIO32_SIZE        (256 * MiB)
+
 static MemMapEntry virt_high_pcie_memmap;
 
 #define VIRT_FLASH_SECTOR_SIZE (256 * KiB)
@@ -727,6 +732,18 @@ static void create_fdt_pcie(RISCVVirtState *s,
 {
     g_autofree char *name = NULL;
     MachineState *ms = MACHINE(s);
+    /*
+     * When CXL is enabled, reserve the last 256 MiB of the 32-bit MMIO
+     * window for CXL host bridges and exclude it from the main PCIe host
+     * bridge's FDT 'ranges' so UEFI's PciHostBridgeDxe does not allocate
+     * that range to PCI0.  The CXL host bridge _CRS declares this range
+     * independently.
+     */
+    hwaddr mmio32_size = s->memmap[VIRT_PCIE_MMIO].size;
+
+    if (s->cxl_devices_state.is_enabled) {
+        mmio32_size -= VIRT_CXL_MMIO32_SIZE;
+    }
 
     name = g_strdup_printf("/soc/pci@%"HWADDR_PRIx,
                            s->memmap[VIRT_PCIE_ECAM].base);
@@ -752,7 +769,7 @@ static void create_fdt_pcie(RISCVVirtState *s,
         2, s->memmap[VIRT_PCIE_PIO].base, 2, s->memmap[VIRT_PCIE_PIO].size,
         1, FDT_PCI_RANGE_MMIO,
         2, s->memmap[VIRT_PCIE_MMIO].base,
-        2, s->memmap[VIRT_PCIE_MMIO].base, 2, s->memmap[VIRT_PCIE_MMIO].size,
+        2, s->memmap[VIRT_PCIE_MMIO].base, 2, mmio32_size,
         1, FDT_PCI_RANGE_MMIO_64BIT,
         2, virt_high_pcie_memmap.base,
         2, virt_high_pcie_memmap.base, 2, virt_high_pcie_memmap.size);
@@ -1130,6 +1147,132 @@ static void cxl_host_state_init(RISCVVirtState *s)
     cxl_fmws_update_mmio();
 }
 
+/*
+ * Assign PCI bus numbers to the bridges under @bus in depth-first order,
+ * the way firmware does during enumeration.  QEMU leaves the secondary and
+ * subordinate bus registers at 0 until firmware programs them, so without
+ * this build_crs() would see an empty bus range for the CXL host bridge.
+ * @bus itself is numbered @bus_num.  Returns the highest bus number used.
+ */
+static int virt_cxl_assign_bus_numbers(PCIBus *bus, int bus_num)
+{
+    int max_bus = bus_num;
+    int next_bus = bus_num + 1;
+    int devfn;
+
+    for (devfn = 0; devfn < ARRAY_SIZE(bus->devices); devfn++) {
+        PCIDevice *dev = bus->devices[devfn];
+        PCIBus *sec_bus;
+        int subordinate;
+
+        if (!dev) {
+            continue;
+        }
+        if ((dev->config[PCI_HEADER_TYPE] &
+             ~PCI_HEADER_TYPE_MULTI_FUNCTION) != PCI_HEADER_TYPE_BRIDGE) {
+            continue;
+        }
+
+        sec_bus = pci_bridge_get_sec_bus(PCI_BRIDGE(dev));
+        subordinate = virt_cxl_assign_bus_numbers(sec_bus, next_bus);
+
+        pci_set_byte(dev->config + PCI_PRIMARY_BUS, bus_num);
+        pci_set_byte(dev->config + PCI_SECONDARY_BUS, next_bus);
+        pci_set_byte(dev->config + PCI_SUBORDINATE_BUS, subordinate);
+
+        if (subordinate > max_bus) {
+            max_bus = subordinate;
+        }
+        next_bus = subordinate + 1;
+    }
+
+    return max_bus;
+}
+
+/*
+ * Simulate the PCI resource initialization that firmware (UEFI) would
+ * normally perform for the CXL host bridge.
+ *
+ * EDK2's PciBusDxe only recurses into devices that look like PCI-to-PCI
+ * bridges, while the pxb-cxl expander bridge presents as a class 0x0600
+ * host bridge with a standard (type 0) header, so the firmware neither
+ * enumerates behind it nor assigns it a memory window or bus numbers.  As a
+ * result the generic build_crs() path would produce an empty _CRS for the
+ * CXL host bridge (ACPI0016).
+ *
+ * Program each CXL root port the way firmware would:
+ *  - assign the reserved 32-bit MMIO range (see create_fdt_pcie()) to its
+ *    memory window, so build_crs() emits that range in the host bridge _CRS
+ *    and excludes it from PCI0's _CRS;
+ *  - set the primary/secondary/subordinate bus numbers, so build_crs()
+ *    emits a correct bus-range resource.  Without a bus range Linux only
+ *    claims the single host-bridge bus and cannot enumerate the devices
+ *    behind the root port.
+ *
+ * TODO: this handles a single pxb-cxl with a single root port, which is
+ * the supported topology for now.  With multiple CXL host bridges or root
+ * ports the reserved window would need to be partitioned per bridge.
+ */
+static void virt_cxl_init_bridge_windows(RISCVVirtState *s)
+{
+    hwaddr cxl_mmio_base, cxl_mmio_limit;
+    PCIBus *bus;
+
+    if (!s->cxl_devices_state.is_enabled) {
+        return;
+    }
+
+    cxl_mmio_base = s->memmap[VIRT_PCIE_MMIO].base +
+                    s->memmap[VIRT_PCIE_MMIO].size - VIRT_CXL_MMIO32_SIZE;
+    cxl_mmio_limit = cxl_mmio_base + VIRT_CXL_MMIO32_SIZE - 1;
+
+    QLIST_FOREACH(bus, &s->pci_bus->child, sibling) {
+        int devfn;
+
+        if (!pci_bus_is_root(bus) || !pci_bus_is_cxl(bus)) {
+            continue;
+        }
+
+        /* Assign bus numbers for the whole CXL subtree (firmware would). */
+        virt_cxl_assign_bus_numbers(bus, pci_bus_num(bus));
+
+        /* Assign the reserved MMIO window to the root-port bridge. */
+        for (devfn = 0; devfn < ARRAY_SIZE(bus->devices); devfn++) {
+            PCIDevice *dev = bus->devices[devfn];
+            PCIBridge *br;
+
+            if (!dev) {
+                continue;
+            }
+            if ((dev->config[PCI_HEADER_TYPE] &
+                 ~PCI_HEADER_TYPE_MULTI_FUNCTION) != PCI_HEADER_TYPE_BRIDGE) {
+                continue;
+            }
+
+            br = PCI_BRIDGE(dev);
+
+            pci_set_word(dev->config + PCI_MEMORY_BASE,
+                         (cxl_mmio_base >> 16) & PCI_MEMORY_RANGE_MASK);
+            pci_set_word(dev->config + PCI_MEMORY_LIMIT,
+                         (cxl_mmio_limit >> 16) & PCI_MEMORY_RANGE_MASK);
+            pci_word_test_and_set_mask(dev->config + PCI_COMMAND,
+                                       PCI_COMMAND_MEMORY);
+            pci_bridge_update_mappings(br);
+            break;
+        }
+    }
+}
+
+/*
+ * A PCI reset clears the bridge window and bus-number registers programmed
+ * above, so re-apply them after every reset, the way firmware would during
+ * its PCI enumeration.
+ */
+static void virt_cxl_reset_bridge_windows(void *opaque)
+{
+    virt_cxl_init_bridge_windows(RISCV_VIRT_MACHINE(opaque));
+}
+
 static FWCfgState *create_fw_cfg(const MachineState *ms, hwaddr base)
 {
     FWCfgState *fw_cfg;
@@ -1238,6 +1381,15 @@ static void virt_machine_done(Notifier *notifier, void *data)
     if (s->cxl_devices_state.is_enabled) {
         cxl_fmws_link_targets(&error_fatal);
     }
+
+    /*
+     * Assign the CXL bridge memory window (simulating firmware) so the
+     * ACPI _CRS built for the CXL host bridge is populated.  Do it here for
+     * the initial build, and register a reset handler so the window is
+     * re-applied after any PCI reset (which clears the bridge registers).
+     */
+    virt_cxl_init_bridge_windows(s);
+    qemu_register_reset(virt_cxl_reset_bridge_windows, s);
     hwaddr firmware_end_addr;
     vaddr kernel_start_addr;
     const char *firmware_name = riscv_default_firmware_name(&s->soc[0]);
-- 
2.50.1


  parent reply	other threads:[~2026-08-21  8:20 UTC|newest]

Thread overview: 11+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-21  8:19 [PATCH 0/4] hw/riscv/virt: Add CXL support to the RISC-V virt machine Chen Pei
2026-08-21  8:19 ` [PATCH 1/4] " Chen Pei
2026-08-21  8:19 ` [PATCH 2/4] hw/riscv/virt-acpi-build: Add _DEP to ACPI0017 for CXL host bridge dependency Chen Pei
2026-08-21  8:19 ` Chen Pei [this message]
2026-08-26 13:37   ` [PATCH 3/4] hw/riscv/virt: Provide a 32-bit MMIO window for CXL host bridges Junjie Cao
2026-08-27 12:17     ` Chen Pei
2026-08-21  8:19 ` [PATCH 4/4] tests/qtest: Add RISC-V ACPI bios tables test for CXL Chen Pei
2026-08-26 13:38   ` Junjie Cao
2026-08-27 12:24     ` Chen Pei
2026-08-21  8:43 ` [PATCH 0/4] hw/riscv/virt: Add CXL support to the RISC-V virt machine Chen Pei
2026-08-26 13:37 ` Junjie Cao

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260821081954.1171-4-cp0613@linux.alibaba.com \
    --to=cp0613@linux.alibaba.com \
    --cc=alison.schofield@intel.com \
    --cc=alistair.francis@wdc.com \
    --cc=anisinha@redhat.com \
    --cc=chao.liu@processmission.com \
    --cc=daniel.barboza@oss.qualcomm.com \
    --cc=dave.jiang@intel.com \
    --cc=guoren@kernel.org \
    --cc=imammedo@redhat.com \
    --cc=jic23@kernel.org \
    --cc=junjie.cao@intel.com \
    --cc=linux-cxl@vger.kernel.org \
    --cc=liwei1518@gmail.com \
    --cc=mst@redhat.com \
    --cc=palmer@dabbelt.com \
    --cc=pbonzini@redhat.com \
    --cc=qemu-devel@nongnu.org \
    --cc=qemu-riscv@nongnu.org \
    --cc=sunilvl@ventanamicro.com \
    --cc=zhiwei_liu@linux.alibaba.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox