* [PATCH v9 01/10] cxl: Add BI register probing and port initialization
2026-09-22 23:38 [PATCH v9 0/10] cxl: Support Back-Invalidate Davidlohr Bueso
@ 2026-09-22 23:38 ` Davidlohr Bueso
2026-09-29 4:06 ` Richard Cheng
2026-09-22 23:38 ` [PATCH v9 02/10] cxl/pci: Add BI topology enable/disable Davidlohr Bueso
` (9 subsequent siblings)
10 siblings, 1 reply; 33+ messages in thread
From: Davidlohr Bueso @ 2026-09-22 23:38 UTC (permalink / raw)
To: dave.jiang
Cc: jic23, alison.schofield, icheng, ming.li, benjamin.cheatham,
alucerop, dave, linux-cxl, Jonathan Cameron
Add register probing for BI Route Table and BI Decoder capability
structures in cxl_probe_component_regs(), and helpers to map them.
cxl_dport_map_bi() maps the BI Decoder of a downstream port (root
port or switch DSP) at dport-creation time via cxl_port_add_dport();
devm_cxl_port_bi_setup() maps a port's own BI capability when the
upstream link is in 256B Flit operation - the BI Decoder during
endpoint port probe, the BI RT of a switch USP on first-dport setup.
Reviewed-by: Dave Jiang <dave.jiang@intel.com>
Reviewed-by: Alison Schofield <alison.schofield@intel.com>
Reviewed-by: Li Ming <ming.li@zohomail.com>
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
Signed-off-by: Davidlohr Bueso <dave@stgolabs.net>
---
drivers/cxl/core/core.h | 1 +
drivers/cxl/core/pci.c | 38 ++++++++++++++++++++++++++++++++++
drivers/cxl/core/port.c | 4 +++-
drivers/cxl/core/regs.c | 14 +++++++++++++
drivers/cxl/cxl.h | 7 +++++++
drivers/cxl/port.c | 45 +++++++++++++++++++++++++++++++++++++++++
include/cxl/cxl.h | 6 ++++++
7 files changed, 114 insertions(+), 1 deletion(-)
diff --git a/drivers/cxl/core/core.h b/drivers/cxl/core/core.h
index 35eaf636adc9..384bad8cb70f 100644
--- a/drivers/cxl/core/core.h
+++ b/drivers/cxl/core/core.h
@@ -208,6 +208,7 @@ static inline void devm_cxl_dport_ras_setup(struct cxl_dport *dport) { }
#endif /* CONFIG_CXL_RAS */
int cxl_gpf_port_setup(struct cxl_dport *dport);
+void devm_cxl_dport_bi_setup(struct cxl_dport *dport);
struct cxl_hdm;
int cxl_hdm_decode_init(struct cxl_dev_state *cxlds, struct cxl_hdm *cxlhdm,
diff --git a/drivers/cxl/core/pci.c b/drivers/cxl/core/pci.c
index 9d807c1a002c..b8676a3d6ec9 100644
--- a/drivers/cxl/core/pci.c
+++ b/drivers/cxl/core/pci.c
@@ -927,3 +927,41 @@ int cxl_port_get_possible_dports(struct cxl_port *port)
return ctx.count;
}
+
+static void cxl_dport_map_bi(struct cxl_dport *dport)
+{
+ struct cxl_register_map *map = &dport->reg_map;
+ struct device *dev = dport->dport_dev;
+
+ if (!map->component_map.bi_decoder.valid) {
+ dev_dbg(dev, "BI Decoder registers not found\n");
+ return;
+ }
+
+ if (cxl_map_component_regs(map, &dport->regs.component,
+ BIT(CXL_CM_CAP_CAP_ID_BI_DECODER)))
+ dev_dbg(dev, "Failed to map BI Decoder capability\n");
+}
+
+/**
+ * devm_cxl_dport_bi_setup - Map BI Decoder registers on a CXL dport
+ * @dport: the cxl_dport that needs to be initialized
+ *
+ * Must be called while the dport's devres group is open so iomap
+ * allocations are released on dport removal.
+ */
+void devm_cxl_dport_bi_setup(struct cxl_dport *dport)
+{
+ if (!dev_is_pci(dport->dport_dev))
+ return;
+
+ switch (pci_pcie_type(to_pci_dev(dport->dport_dev))) {
+ case PCI_EXP_TYPE_ROOT_PORT:
+ case PCI_EXP_TYPE_DOWNSTREAM:
+ dport->reg_map.host = dport_to_host(dport);
+ cxl_dport_map_bi(dport);
+ break;
+ default:
+ break;
+ }
+}
diff --git a/drivers/cxl/core/port.c b/drivers/cxl/core/port.c
index 625e4aa427db..131ed62e8db3 100644
--- a/drivers/cxl/core/port.c
+++ b/drivers/cxl/core/port.c
@@ -1242,8 +1242,10 @@ __devm_cxl_add_dport(struct cxl_port *port, struct device *dport_dev,
cxl_debugfs_create_dport_dir(dport);
- if (!dport->rch)
+ if (!dport->rch) {
devm_cxl_dport_ras_setup(dport);
+ devm_cxl_dport_bi_setup(dport);
+ }
/* keep the group, and mark the end of devm actions */
cxl_dport_close_dr_group(dport, no_free_ptr(dport_dr_group));
diff --git a/drivers/cxl/core/regs.c b/drivers/cxl/core/regs.c
index 20c2d9fbcfe7..f2c424d129f7 100644
--- a/drivers/cxl/core/regs.c
+++ b/drivers/cxl/core/regs.c
@@ -93,6 +93,18 @@ void cxl_probe_component_regs(struct device *dev, void __iomem *base,
length = CXL_RAS_CAPABILITY_LENGTH;
rmap = &map->ras;
break;
+ case CXL_CM_CAP_CAP_ID_BI_RT:
+ dev_dbg(dev, "found BI RT capability (0x%x)\n",
+ offset);
+ length = CXL_BI_RT_CAPABILITY_LENGTH;
+ rmap = &map->bi_rt;
+ break;
+ case CXL_CM_CAP_CAP_ID_BI_DECODER:
+ dev_dbg(dev, "found BI Decoder capability (0x%x)\n",
+ offset);
+ length = CXL_BI_DECODER_CAPABILITY_LENGTH;
+ rmap = &map->bi_decoder;
+ break;
default:
dev_dbg(dev, "Unknown CM cap ID: %d (0x%x)\n", cap_id,
offset);
@@ -212,6 +224,8 @@ int cxl_map_component_regs(const struct cxl_register_map *map,
} mapinfo[] = {
{ &map->component_map.hdm_decoder, ®s->hdm_decoder },
{ &map->component_map.ras, ®s->ras },
+ { &map->component_map.bi_rt, ®s->bi_rt },
+ { &map->component_map.bi_decoder, ®s->bi_decoder },
};
int i;
diff --git a/drivers/cxl/cxl.h b/drivers/cxl/cxl.h
index cab8ce39f465..65be3b91259a 100644
--- a/drivers/cxl/cxl.h
+++ b/drivers/cxl/cxl.h
@@ -39,9 +39,16 @@ extern const struct nvdimm_security_ops *cxl_security_ops;
#define CXL_CM_CAP_HDR_ARRAY_SIZE_MASK GENMASK(31, 24)
#define CXL_CM_CAP_PTR_MASK GENMASK(31, 20)
+/* CXL 4.0 8.2.4 Table 8-74 */
#define CXL_CM_CAP_CAP_ID_RAS 0x2
#define CXL_CM_CAP_CAP_ID_HDM 0x5
#define CXL_CM_CAP_CAP_HDM_VERSION 1
+#define CXL_CM_CAP_CAP_ID_BI_RT 0xB
+#define CXL_CM_CAP_CAP_ID_BI_DECODER 0xC
+
+/* CXL 4.0 8.2.4.26 / 8.2.4.27 BI Capability Structures */
+#define CXL_BI_RT_CAPABILITY_LENGTH 0xC
+#define CXL_BI_DECODER_CAPABILITY_LENGTH 0xC
/* HDM decoders CXL 2.0 8.2.5.12 CXL HDM Decoder Capability Structure */
#define CXL_HDM_DECODER_CAP_OFFSET 0x0
diff --git a/drivers/cxl/port.c b/drivers/cxl/port.c
index 99cf77b6b699..eab935eca9b1 100644
--- a/drivers/cxl/port.c
+++ b/drivers/cxl/port.c
@@ -58,6 +58,47 @@ static int discover_region(struct device *dev, void *unused)
return 0;
}
+static void devm_cxl_port_bi_setup(struct cxl_port *port)
+{
+ struct cxl_register_map *map = &port->reg_map;
+ struct cxl_dport *parent_dport = port->parent_dport;
+ struct device *udev;
+ int cap_id;
+
+ /* no upstream BI registers above host bridges or the cxl_root */
+ if (!parent_dport || is_cxl_root(parent_dport->port))
+ return;
+
+ udev = is_cxl_endpoint(port) ?
+ port->uport_dev->parent : port->uport_dev;
+ if (!dev_is_pci(udev))
+ return;
+
+ /* BI requires 256B Flit on the upstream link */
+ if (!cxl_pci_flit_256(to_pci_dev(udev)))
+ return;
+
+ /* map this port's own BI capability */
+ if (is_cxl_endpoint(port)) {
+ if (!map->component_map.bi_decoder.valid) {
+ dev_dbg(&port->dev, "BI Decoder registers not found\n");
+ return;
+ }
+ cap_id = CXL_CM_CAP_CAP_ID_BI_DECODER;
+ } else {
+ if (!map->component_map.bi_rt.valid) {
+ dev_dbg(&port->dev, "BI RT registers not found\n");
+ return;
+ }
+ cap_id = CXL_CM_CAP_CAP_ID_BI_RT;
+ }
+
+ map->host = &port->dev;
+ if (cxl_map_component_regs(map, &port->regs, BIT(cap_id)))
+ dev_dbg(&port->dev, "Failed to map BI capability 0x%x\n",
+ cap_id);
+}
+
static int cxl_switch_port_probe(struct cxl_port *port)
{
/* Reset nr_dports for rebind of driver */
@@ -128,6 +169,8 @@ static int cxl_endpoint_port_probe(struct cxl_port *port)
read_cdat_data(port);
cxl_endpoint_parse_cdat(port);
+ devm_cxl_port_bi_setup(port);
+
get_device(&cxlmd->dev);
rc = devm_add_action_or_reset(&port->dev, schedule_detach, cxlmd);
if (rc)
@@ -252,6 +295,8 @@ static struct cxl_dport *cxl_port_add_dport(struct cxl_port *port,
* on failure, or the device does not implement RAS registers.
*/
devm_cxl_port_ras_setup(port);
+
+ devm_cxl_port_bi_setup(port);
}
dport = devm_cxl_add_dport_by_dev(port, dport_dev);
diff --git a/include/cxl/cxl.h b/include/cxl/cxl.h
index 802b143de83d..278b84b08c83 100644
--- a/include/cxl/cxl.h
+++ b/include/cxl/cxl.h
@@ -34,10 +34,14 @@ struct cxl_regs {
* Common set of CXL Component register block base pointers
* @hdm_decoder: CXL 2.0 8.2.5.12 CXL HDM Decoder Capability Structure
* @ras: CXL 2.0 8.2.5.9 CXL RAS Capability Structure
+ * @bi_rt: CXL 4.0 8.2.4.26 CXL BI Route Table Capability Structure
+ * @bi_decoder: CXL 4.0 8.2.4.27 CXL BI Decoder Capability Structure
*/
struct_group_tagged(cxl_component_regs, component,
void __iomem *hdm_decoder;
void __iomem *ras;
+ void __iomem *bi_rt;
+ void __iomem *bi_decoder;
);
/*
* Common set of CXL Device register block base pointers
@@ -80,6 +84,8 @@ struct cxl_reg_map {
struct cxl_component_reg_map {
struct cxl_reg_map hdm_decoder;
struct cxl_reg_map ras;
+ struct cxl_reg_map bi_rt;
+ struct cxl_reg_map bi_decoder;
};
struct cxl_device_reg_map {
--
2.39.5
^ permalink raw reply related [flat|nested] 33+ messages in thread* Re: [PATCH v9 01/10] cxl: Add BI register probing and port initialization
2026-09-22 23:38 ` [PATCH v9 01/10] cxl: Add BI register probing and port initialization Davidlohr Bueso
@ 2026-09-29 4:06 ` Richard Cheng
0 siblings, 0 replies; 33+ messages in thread
From: Richard Cheng @ 2026-09-29 4:06 UTC (permalink / raw)
To: Davidlohr Bueso
Cc: dave.jiang, jic23, alison.schofield, ming.li, benjamin.cheatham,
alucerop, linux-cxl, Jonathan Cameron
On Tue, Sep 22, 2026 at 04:38:38PM +0800, Davidlohr Bueso wrote:
> Add register probing for BI Route Table and BI Decoder capability
> structures in cxl_probe_component_regs(), and helpers to map them.
>
> cxl_dport_map_bi() maps the BI Decoder of a downstream port (root
> port or switch DSP) at dport-creation time via cxl_port_add_dport();
> devm_cxl_port_bi_setup() maps a port's own BI capability when the
> upstream link is in 256B Flit operation - the BI Decoder during
> endpoint port probe, the BI RT of a switch USP on first-dport setup.
>
> Reviewed-by: Dave Jiang <dave.jiang@intel.com>
> Reviewed-by: Alison Schofield <alison.schofield@intel.com>
> Reviewed-by: Li Ming <ming.li@zohomail.com>
> Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
> Signed-off-by: Davidlohr Bueso <dave@stgolabs.net>
Reviewed-by: Richard Cheng <icheng@nvidia.com>
Just a nit below, no blocker.
> ---
> drivers/cxl/core/core.h | 1 +
> drivers/cxl/core/pci.c | 38 ++++++++++++++++++++++++++++++++++
> drivers/cxl/core/port.c | 4 +++-
> drivers/cxl/core/regs.c | 14 +++++++++++++
> drivers/cxl/cxl.h | 7 +++++++
> drivers/cxl/port.c | 45 +++++++++++++++++++++++++++++++++++++++++
> include/cxl/cxl.h | 6 ++++++
> 7 files changed, 114 insertions(+), 1 deletion(-)
>
> diff --git a/drivers/cxl/core/core.h b/drivers/cxl/core/core.h
> index 35eaf636adc9..384bad8cb70f 100644
> --- a/drivers/cxl/core/core.h
> +++ b/drivers/cxl/core/core.h
> @@ -208,6 +208,7 @@ static inline void devm_cxl_dport_ras_setup(struct cxl_dport *dport) { }
> #endif /* CONFIG_CXL_RAS */
>
> int cxl_gpf_port_setup(struct cxl_dport *dport);
> +void devm_cxl_dport_bi_setup(struct cxl_dport *dport);
>
> struct cxl_hdm;
> int cxl_hdm_decode_init(struct cxl_dev_state *cxlds, struct cxl_hdm *cxlhdm,
> diff --git a/drivers/cxl/core/pci.c b/drivers/cxl/core/pci.c
> index 9d807c1a002c..b8676a3d6ec9 100644
> --- a/drivers/cxl/core/pci.c
> +++ b/drivers/cxl/core/pci.c
> @@ -927,3 +927,41 @@ int cxl_port_get_possible_dports(struct cxl_port *port)
>
> return ctx.count;
> }
> +
> +static void cxl_dport_map_bi(struct cxl_dport *dport)
> +{
> + struct cxl_register_map *map = &dport->reg_map;
> + struct device *dev = dport->dport_dev;
> +
> + if (!map->component_map.bi_decoder.valid) {
> + dev_dbg(dev, "BI Decoder registers not found\n");
> + return;
> + }
> +
> + if (cxl_map_component_regs(map, &dport->regs.component,
> + BIT(CXL_CM_CAP_CAP_ID_BI_DECODER)))
> + dev_dbg(dev, "Failed to map BI Decoder capability\n");
> +}
> +
> +/**
> + * devm_cxl_dport_bi_setup - Map BI Decoder registers on a CXL dport
> + * @dport: the cxl_dport that needs to be initialized
> + *
> + * Must be called while the dport's devres group is open so iomap
> + * allocations are released on dport removal.
> + */
> +void devm_cxl_dport_bi_setup(struct cxl_dport *dport)
> +{
> + if (!dev_is_pci(dport->dport_dev))
> + return;
> +
> + switch (pci_pcie_type(to_pci_dev(dport->dport_dev))) {
> + case PCI_EXP_TYPE_ROOT_PORT:
> + case PCI_EXP_TYPE_DOWNSTREAM:
> + dport->reg_map.host = dport_to_host(dport);
> + cxl_dport_map_bi(dport);
> + break;
> + default:
> + break;
> + }
> +}
> diff --git a/drivers/cxl/core/port.c b/drivers/cxl/core/port.c
> index 625e4aa427db..131ed62e8db3 100644
> --- a/drivers/cxl/core/port.c
> +++ b/drivers/cxl/core/port.c
> @@ -1242,8 +1242,10 @@ __devm_cxl_add_dport(struct cxl_port *port, struct device *dport_dev,
>
> cxl_debugfs_create_dport_dir(dport);
>
> - if (!dport->rch)
> + if (!dport->rch) {
> devm_cxl_dport_ras_setup(dport);
> + devm_cxl_dport_bi_setup(dport);
> + }
>
> /* keep the group, and mark the end of devm actions */
> cxl_dport_close_dr_group(dport, no_free_ptr(dport_dr_group));
> diff --git a/drivers/cxl/core/regs.c b/drivers/cxl/core/regs.c
> index 20c2d9fbcfe7..f2c424d129f7 100644
> --- a/drivers/cxl/core/regs.c
> +++ b/drivers/cxl/core/regs.c
> @@ -93,6 +93,18 @@ void cxl_probe_component_regs(struct device *dev, void __iomem *base,
> length = CXL_RAS_CAPABILITY_LENGTH;
> rmap = &map->ras;
> break;
> + case CXL_CM_CAP_CAP_ID_BI_RT:
> + dev_dbg(dev, "found BI RT capability (0x%x)\n",
> + offset);
> + length = CXL_BI_RT_CAPABILITY_LENGTH;
> + rmap = &map->bi_rt;
> + break;
> + case CXL_CM_CAP_CAP_ID_BI_DECODER:
> + dev_dbg(dev, "found BI Decoder capability (0x%x)\n",
> + offset);
> + length = CXL_BI_DECODER_CAPABILITY_LENGTH;
> + rmap = &map->bi_decoder;
> + break;
> default:
> dev_dbg(dev, "Unknown CM cap ID: %d (0x%x)\n", cap_id,
> offset);
> @@ -212,6 +224,8 @@ int cxl_map_component_regs(const struct cxl_register_map *map,
> } mapinfo[] = {
> { &map->component_map.hdm_decoder, ®s->hdm_decoder },
> { &map->component_map.ras, ®s->ras },
> + { &map->component_map.bi_rt, ®s->bi_rt },
> + { &map->component_map.bi_decoder, ®s->bi_decoder },
> };
> int i;
>
> diff --git a/drivers/cxl/cxl.h b/drivers/cxl/cxl.h
> index cab8ce39f465..65be3b91259a 100644
> --- a/drivers/cxl/cxl.h
> +++ b/drivers/cxl/cxl.h
> @@ -39,9 +39,16 @@ extern const struct nvdimm_security_ops *cxl_security_ops;
> #define CXL_CM_CAP_HDR_ARRAY_SIZE_MASK GENMASK(31, 24)
> #define CXL_CM_CAP_PTR_MASK GENMASK(31, 20)
>
> +/* CXL 4.0 8.2.4 Table 8-74 */
> #define CXL_CM_CAP_CAP_ID_RAS 0x2
> #define CXL_CM_CAP_CAP_ID_HDM 0x5
> #define CXL_CM_CAP_CAP_HDM_VERSION 1
> +#define CXL_CM_CAP_CAP_ID_BI_RT 0xB
> +#define CXL_CM_CAP_CAP_ID_BI_DECODER 0xC
> +
> +/* CXL 4.0 8.2.4.26 / 8.2.4.27 BI Capability Structures */
> +#define CXL_BI_RT_CAPABILITY_LENGTH 0xC
> +#define CXL_BI_DECODER_CAPABILITY_LENGTH 0xC
>
These 3 macro all refers to 0xC , should we add comment that they're
independent of each other so if in the future someone wants to refactor,
they won't accidently treat them as dependent value ?
Best regards,
Richard Cheng.
> /* HDM decoders CXL 2.0 8.2.5.12 CXL HDM Decoder Capability Structure */
> #define CXL_HDM_DECODER_CAP_OFFSET 0x0
> diff --git a/drivers/cxl/port.c b/drivers/cxl/port.c
> index 99cf77b6b699..eab935eca9b1 100644
> --- a/drivers/cxl/port.c
> +++ b/drivers/cxl/port.c
> @@ -58,6 +58,47 @@ static int discover_region(struct device *dev, void *unused)
> return 0;
> }
>
> +static void devm_cxl_port_bi_setup(struct cxl_port *port)
> +{
> + struct cxl_register_map *map = &port->reg_map;
> + struct cxl_dport *parent_dport = port->parent_dport;
> + struct device *udev;
> + int cap_id;
> +
> + /* no upstream BI registers above host bridges or the cxl_root */
> + if (!parent_dport || is_cxl_root(parent_dport->port))
> + return;
> +
> + udev = is_cxl_endpoint(port) ?
> + port->uport_dev->parent : port->uport_dev;
> + if (!dev_is_pci(udev))
> + return;
> +
> + /* BI requires 256B Flit on the upstream link */
> + if (!cxl_pci_flit_256(to_pci_dev(udev)))
> + return;
> +
> + /* map this port's own BI capability */
> + if (is_cxl_endpoint(port)) {
> + if (!map->component_map.bi_decoder.valid) {
> + dev_dbg(&port->dev, "BI Decoder registers not found\n");
> + return;
> + }
> + cap_id = CXL_CM_CAP_CAP_ID_BI_DECODER;
> + } else {
> + if (!map->component_map.bi_rt.valid) {
> + dev_dbg(&port->dev, "BI RT registers not found\n");
> + return;
> + }
> + cap_id = CXL_CM_CAP_CAP_ID_BI_RT;
> + }
> +
> + map->host = &port->dev;
> + if (cxl_map_component_regs(map, &port->regs, BIT(cap_id)))
> + dev_dbg(&port->dev, "Failed to map BI capability 0x%x\n",
> + cap_id);
> +}
> +
> static int cxl_switch_port_probe(struct cxl_port *port)
> {
> /* Reset nr_dports for rebind of driver */
> @@ -128,6 +169,8 @@ static int cxl_endpoint_port_probe(struct cxl_port *port)
> read_cdat_data(port);
> cxl_endpoint_parse_cdat(port);
>
> + devm_cxl_port_bi_setup(port);
> +
> get_device(&cxlmd->dev);
> rc = devm_add_action_or_reset(&port->dev, schedule_detach, cxlmd);
> if (rc)
> @@ -252,6 +295,8 @@ static struct cxl_dport *cxl_port_add_dport(struct cxl_port *port,
> * on failure, or the device does not implement RAS registers.
> */
> devm_cxl_port_ras_setup(port);
> +
> + devm_cxl_port_bi_setup(port);
> }
>
> dport = devm_cxl_add_dport_by_dev(port, dport_dev);
> diff --git a/include/cxl/cxl.h b/include/cxl/cxl.h
> index 802b143de83d..278b84b08c83 100644
> --- a/include/cxl/cxl.h
> +++ b/include/cxl/cxl.h
> @@ -34,10 +34,14 @@ struct cxl_regs {
> * Common set of CXL Component register block base pointers
> * @hdm_decoder: CXL 2.0 8.2.5.12 CXL HDM Decoder Capability Structure
> * @ras: CXL 2.0 8.2.5.9 CXL RAS Capability Structure
> + * @bi_rt: CXL 4.0 8.2.4.26 CXL BI Route Table Capability Structure
> + * @bi_decoder: CXL 4.0 8.2.4.27 CXL BI Decoder Capability Structure
> */
> struct_group_tagged(cxl_component_regs, component,
> void __iomem *hdm_decoder;
> void __iomem *ras;
> + void __iomem *bi_rt;
> + void __iomem *bi_decoder;
> );
> /*
> * Common set of CXL Device register block base pointers
> @@ -80,6 +84,8 @@ struct cxl_reg_map {
> struct cxl_component_reg_map {
> struct cxl_reg_map hdm_decoder;
> struct cxl_reg_map ras;
> + struct cxl_reg_map bi_rt;
> + struct cxl_reg_map bi_decoder;
> };
>
> struct cxl_device_reg_map {
> --
> 2.39.5
>
^ permalink raw reply [flat|nested] 33+ messages in thread
* [PATCH v9 02/10] cxl/pci: Add BI topology enable/disable
2026-09-22 23:38 [PATCH v9 0/10] cxl: Support Back-Invalidate Davidlohr Bueso
2026-09-22 23:38 ` [PATCH v9 01/10] cxl: Add BI register probing and port initialization Davidlohr Bueso
@ 2026-09-22 23:38 ` Davidlohr Bueso
2026-09-23 1:12 ` sashiko-bot
` (4 more replies)
2026-09-22 23:38 ` [PATCH v9 03/10] cxl/hdm: Add BI coherency support for endpoint decoders Davidlohr Bueso
` (8 subsequent siblings)
10 siblings, 5 replies; 33+ messages in thread
From: Davidlohr Bueso @ 2026-09-22 23:38 UTC (permalink / raw)
To: dave.jiang
Cc: jic23, alison.schofield, icheng, ming.li, benjamin.cheatham,
alucerop, dave, linux-cxl
Implement cxl_bi_setup() to enable BI flows on the device and every
component in the path, and its teardown counterpart cxl_bi_dealloc().
Setup runs from devm_cxl_endpoint_decoders_setup(), between the
port's HDM state and its decoders. Registered there, its devres
teardown brings BI down after the decoders quiesce and before the
HDM state is freed, and BI is settled before the decoders, and later
the regions, are looked at. The BI-ID and path enablement belong to
the endpoint port's lifetime.
Setup is safe in endpoint port probe context. The port probes
synchronously from cxl_mem_probe(), pinning the memdev state the
walk consumes, and the whole ancestor path already exists with BI
registers mapped (dports at dport-add time, the switch USP RT at
first-dport setup) because devm_cxl_enumerate_ports() completes
before the endpoint is created.
Dealloc is safe in endpoint devres context. Both setup and dealloc
walk the endpoint's parent_dport topology rather than getting the
port by bus lookup - an ancestor teardown delists the parent port
before the endpoint's devres runs.
The topology walk is stable as parent_dport pointers are fixed at
port creation; ancestors cannot be reaped while holding this
memdev's cxl_ep; and their own teardown frees dports only after
the endpoint is gone.
Likewise, the device state outlives the walk - cxlmd->cxlds is
nulled only after cxl_memdev_unregister() has torn the endpoint
down, and delete_endpoint() clears cxlmd->endpoint only after the
endpoint devres has run.
Each dport is programmed by its position - the one immediately above
the device takes BI Enable, every dport above it takes BI Forward
(Table 8-157, Table 9-13), at any switch depth (Table 7-97). Any
level can be shared, so nr_bi refcounts endpoints at every dport.
Registers are written on the first endpoint and cleared on the last,
but only downstream ports commit (Table 8-156), once per endpoint
(Table 8-152), and a failed commit undoes its write and commits the
undo. A USP advertising a BI Route Table that failed to map is
refused rather than treated as absent. nr_bi counts only the
endpoints this driver enabled, so a level can be cleared while
firmware still has an unbound device on it.
A reset may wipe the device's BI Enable, whose reset default is 0
(Table 8-157). .reset_done reads the hardware rather than assume
which reset ran, and invalidates cxlds->bi, failing closed with
recovery by rebind as for decoder loss; dealloc unwinds the dport
refcounts regardless. It also clears cxlds->bi unconditionally - the
endpoint disable fails when the hardware already shows BI Enable
clear, from a reset .reset_done never saw, and the flag must not
outlive the BI Decoder mapping it describes, which the endpoint port
releases moments later.
With dealloc in the endpoint's devres, delete_endpoint() already
holds the parent port's device lock, so to avoid deadlocking, add a
per-port bi_lock, serializing the dports that share state (nr_bi and
the control register at any shared level, the switch USP's BI RT).
Reviewed-by: Ben Cheatham <benjamin.cheatham@amd.com>
Signed-off-by: Davidlohr Bueso <dave@stgolabs.net>
---
drivers/cxl/core/core.h | 7 +
drivers/cxl/core/hdm.c | 10 +
drivers/cxl/core/pci.c | 427 ++++++++++++++++++++++++++++++++++++++++
drivers/cxl/core/port.c | 1 +
drivers/cxl/cxl.h | 31 +++
drivers/cxl/pci.c | 8 +-
include/cxl/cxl.h | 2 +
7 files changed, 484 insertions(+), 2 deletions(-)
diff --git a/drivers/cxl/core/core.h b/drivers/cxl/core/core.h
index 384bad8cb70f..ef41632d7b7d 100644
--- a/drivers/cxl/core/core.h
+++ b/drivers/cxl/core/core.h
@@ -210,6 +210,13 @@ static inline void devm_cxl_dport_ras_setup(struct cxl_dport *dport) { }
int cxl_gpf_port_setup(struct cxl_dport *dport);
void devm_cxl_dport_bi_setup(struct cxl_dport *dport);
+static inline bool cxl_bi_decoder_enabled(struct cxl_port *port)
+{
+ return FIELD_GET(CXL_BI_DECODER_CTRL_BI_ENABLE,
+ readl(port->regs.bi_decoder +
+ CXL_BI_DECODER_CTRL_OFFSET));
+}
+
struct cxl_hdm;
int cxl_hdm_decode_init(struct cxl_dev_state *cxlds, struct cxl_hdm *cxlhdm,
struct cxl_endpoint_dvsec_info *info);
diff --git a/drivers/cxl/core/hdm.c b/drivers/cxl/core/hdm.c
index 0c80b76a5f9b..0e9d652b568e 100644
--- a/drivers/cxl/core/hdm.c
+++ b/drivers/cxl/core/hdm.c
@@ -1274,6 +1274,16 @@ int devm_cxl_endpoint_decoders_setup(struct cxl_port *port)
if (rc)
return rc;
+ /*
+ * Between the port's HDM state and its decoders: devres,
+ * unwinding in reverse, brings BI down only after the decoders
+ * quiesce, while its slow walk still precedes the HDM state
+ * free.
+ */
+ rc = cxl_bi_setup(port);
+ if (rc)
+ dev_dbg(&port->dev, "BI setup failed rc=%d\n", rc);
+
return devm_cxl_enumerate_decoders(cxlhdm, &info);
}
EXPORT_SYMBOL_NS_GPL(devm_cxl_endpoint_decoders_setup, "CXL");
diff --git a/drivers/cxl/core/pci.c b/drivers/cxl/core/pci.c
index b8676a3d6ec9..f3ca9861a7ba 100644
--- a/drivers/cxl/core/pci.c
+++ b/drivers/cxl/core/pci.c
@@ -2,12 +2,14 @@
/* Copyright(c) 2021 Intel Corporation. All rights reserved. */
#include <linux/units.h>
#include <linux/io-64-nonatomic-lo-hi.h>
+#include <linux/iopoll.h>
#include <linux/device.h>
#include <linux/delay.h>
#include <linux/pci.h>
#include <linux/pci-doe.h>
#include <cxl/pci.h>
#include <linux/aer.h>
+#include <linux/string_choices.h>
#include <cxlpci.h>
#include <cxlmem.h>
#include <cxl.h>
@@ -965,3 +967,428 @@ void devm_cxl_dport_bi_setup(struct cxl_dport *dport)
break;
}
}
+
+/*
+ * BI requires 256B Flit operation on the link. RP/DSP/endpoint must
+ * also have the BI Decoder cap mapped (@bi); for USPs the BI RT cap
+ * is optional per CXL 4.0 8.2.4.26, so absent @bi is allowed.
+ */
+static bool cxl_is_bi_capable(struct pci_dev *pdev, void __iomem *bi)
+{
+ if (!cxl_pci_flit_256(pdev))
+ return false;
+
+ if (pci_pcie_type(pdev) != PCI_EXP_TYPE_UPSTREAM && !bi) {
+ dev_dbg(&pdev->dev, "No BI Decoder registers.\n");
+ return false;
+ }
+
+ return true;
+}
+
+/* limit any insane timeouts from hw */
+#define CXL_BI_COMMIT_MAXTMO_US (20 * USEC_PER_SEC)
+
+static unsigned long __cxl_bi_get_timeout_us(struct device *dev,
+ unsigned int scale,
+ unsigned int base)
+{
+ static const unsigned long scale_tbl[] = {
+ 1, 10, 100, 1000, 10000, 100000, 1000000, 10000000,
+ };
+
+ if (scale >= ARRAY_SIZE(scale_tbl) || !base) {
+ dev_dbg(dev, "Invalid BI commit timeout: scale=%u base=%u\n",
+ scale, base);
+ return CXL_BI_COMMIT_MAXTMO_US;
+ }
+
+ return scale_tbl[scale] * base;
+}
+
+static int __cxl_bi_wait_commit(struct device *dev, void __iomem *status_reg,
+ u32 committed_bit, u32 err_bit,
+ unsigned int scale, unsigned int base)
+{
+ unsigned long tmo_us, poll_us;
+ ktime_t start;
+ u32 status;
+ int rc;
+
+ tmo_us = min_t(unsigned long, CXL_BI_COMMIT_MAXTMO_US,
+ __cxl_bi_get_timeout_us(dev, scale, base));
+ poll_us = max_t(unsigned long, tmo_us / 10, 1); /* ~10% */
+ start = ktime_get();
+
+ rc = readx_poll_timeout(readl, status_reg, status,
+ status & (committed_bit | err_bit),
+ poll_us, tmo_us);
+ if (rc) {
+ dev_err(dev, "BI-ID commit timed out (%luus)\n", tmo_us);
+ return rc; /* -ETIMEDOUT */
+ }
+
+ if (status & err_bit) {
+ dev_err(dev, "BI-ID commit rejected by hardware\n");
+ return -EIO;
+ }
+
+ dev_dbg(dev, "BI-ID commit wait took %lluus\n",
+ ktime_to_us(ktime_sub(ktime_get(), start)));
+ return 0;
+}
+
+/* BI RT only exists on switch upstream ports. */
+static int __cxl_bi_commit_rt(struct device *dev, void __iomem *bi)
+{
+ u32 status, ctrl;
+ unsigned int scale, base;
+
+ if (!FIELD_GET(CXL_BI_RT_CAPS_EXPLICIT_COMMIT_REQ,
+ readl(bi + CXL_BI_RT_CAPS_OFFSET)))
+ return 0;
+
+ ctrl = readl(bi + CXL_BI_RT_CTRL_OFFSET);
+ writel(ctrl & ~CXL_BI_RT_CTRL_BI_COMMIT, bi + CXL_BI_RT_CTRL_OFFSET);
+ writel(ctrl | CXL_BI_RT_CTRL_BI_COMMIT, bi + CXL_BI_RT_CTRL_OFFSET);
+
+ status = readl(bi + CXL_BI_RT_STATUS_OFFSET);
+ scale = FIELD_GET(CXL_BI_RT_STATUS_BI_COMMIT_TM_SCALE, status);
+ base = FIELD_GET(CXL_BI_RT_STATUS_BI_COMMIT_TM_BASE, status);
+
+ return __cxl_bi_wait_commit(dev, bi + CXL_BI_RT_STATUS_OFFSET,
+ CXL_BI_RT_STATUS_BI_COMMITTED,
+ CXL_BI_RT_STATUS_BI_ERR_NOT_COMMITTED,
+ scale, base);
+}
+
+static int __cxl_bi_commit_decoder(struct device *dev, void __iomem *bi)
+{
+ u32 status, ctrl;
+ unsigned int scale, base;
+
+ if (!FIELD_GET(CXL_BI_DECODER_CAPS_EXPLICIT_COMMIT_REQ,
+ readl(bi + CXL_BI_DECODER_CAPS_OFFSET)))
+ return 0;
+
+ ctrl = readl(bi + CXL_BI_DECODER_CTRL_OFFSET);
+ writel(ctrl & ~CXL_BI_DECODER_CTRL_BI_COMMIT,
+ bi + CXL_BI_DECODER_CTRL_OFFSET);
+ writel(ctrl | CXL_BI_DECODER_CTRL_BI_COMMIT,
+ bi + CXL_BI_DECODER_CTRL_OFFSET);
+
+ status = readl(bi + CXL_BI_DECODER_STATUS_OFFSET);
+ scale = FIELD_GET(CXL_BI_DECODER_STATUS_BI_COMMIT_TM_SCALE, status);
+ base = FIELD_GET(CXL_BI_DECODER_STATUS_BI_COMMIT_TM_BASE, status);
+
+ return __cxl_bi_wait_commit(dev, bi + CXL_BI_DECODER_STATUS_OFFSET,
+ CXL_BI_DECODER_STATUS_BI_COMMITTED,
+ CXL_BI_DECODER_STATUS_BI_ERR_NOT_COMMITTED,
+ scale, base);
+}
+
+static int cxl_bi_commit_dport(struct cxl_dport *dport)
+{
+ struct cxl_port *port = dport->port;
+ int rc;
+
+ lockdep_assert_held(&port->bi_lock);
+
+ /* root ports never require the explicit commit */
+ if (pci_pcie_type(to_pci_dev(dport->dport_dev)) !=
+ PCI_EXP_TYPE_DOWNSTREAM)
+ return 0;
+
+ rc = __cxl_bi_commit_decoder(dport->dport_dev, dport->regs.bi_decoder);
+ if (rc)
+ return rc;
+
+ if (port->regs.bi_rt)
+ rc = __cxl_bi_commit_rt(&port->dev, port->regs.bi_rt);
+
+ return rc;
+}
+
+/*
+ * Enable BI-ID changes in the given level of the topology.
+ * @direct says the device is connected directly to this dport, which
+ * takes BI Enable; every dport above it takes BI Forward.
+ */
+static int cxl_bi_ctrl_dport_enable(struct cxl_dport *dport, bool direct)
+{
+ struct cxl_port *port = dport->port;
+ u32 ctrl, value, set, clr;
+ void __iomem *bi;
+ int rc;
+
+ guard(mutex)(&port->bi_lock);
+
+ bi = dport->regs.bi_decoder;
+ if (!bi)
+ return -EINVAL;
+
+ ctrl = readl(bi + CXL_BI_DECODER_CTRL_OFFSET);
+
+ set = direct ? CXL_BI_DECODER_CTRL_BI_ENABLE :
+ CXL_BI_DECODER_CTRL_BI_FW;
+ clr = direct ? CXL_BI_DECODER_CTRL_BI_FW :
+ CXL_BI_DECODER_CTRL_BI_ENABLE;
+
+ value = (ctrl | set) & ~clr;
+ if (value != ctrl)
+ writel(value, bi + CXL_BI_DECODER_CTRL_OFFSET);
+
+ /* owed per new device below, not per register change */
+ rc = cxl_bi_commit_dport(dport);
+ if (rc) {
+ if (value != ctrl) {
+ /* the undo is a BI-ID change owing its own commit */
+ writel(ctrl, bi + CXL_BI_DECODER_CTRL_OFFSET);
+ cxl_bi_commit_dport(dport);
+ }
+ return rc;
+ }
+ dport->nr_bi++;
+
+ return 0;
+}
+
+/*
+ * Dealloc BI-ID changes in the given level of the topology. Called
+ * once per endpoint that enabled this level: on teardown, or to
+ * unwind a path that failed partway up.
+ */
+static int cxl_bi_ctrl_dport_disable(struct cxl_dport *dport)
+{
+ struct cxl_port *port = dport->port;
+ void __iomem *bi;
+ u32 ctrl;
+
+ guard(mutex)(&port->bi_lock);
+
+ bi = dport->regs.bi_decoder;
+ if (!bi)
+ return -EINVAL;
+
+ if (WARN_ON_ONCE(dport->nr_bi == 0))
+ return -EINVAL;
+
+ /* others below still need this level */
+ if (--dport->nr_bi > 0)
+ return 0;
+
+ ctrl = readl(bi + CXL_BI_DECODER_CTRL_OFFSET);
+ writel(ctrl & ~(CXL_BI_DECODER_CTRL_BI_FW |
+ CXL_BI_DECODER_CTRL_BI_ENABLE),
+ bi + CXL_BI_DECODER_CTRL_OFFSET);
+
+ return cxl_bi_commit_dport(dport);
+}
+
+static int __cxl_bi_ctrl_endpoint(struct cxl_dev_state *cxlds, bool enable)
+{
+ struct cxl_port *endpoint = cxlds->cxlmd->endpoint;
+ void __iomem *bi = endpoint->regs.bi_decoder;
+ u32 ctrl;
+
+ if (!bi)
+ return -EINVAL;
+
+ ctrl = readl(bi + CXL_BI_DECODER_CTRL_OFFSET);
+
+ if (enable) {
+ if (FIELD_GET(CXL_BI_DECODER_CTRL_BI_ENABLE, ctrl)) {
+ if (cxlds->bi)
+ return 0;
+ dev_err(cxlds->dev,
+ "BI already enabled in hardware\n");
+ return -EBUSY;
+ }
+ } else {
+ if (!FIELD_GET(CXL_BI_DECODER_CTRL_BI_ENABLE, ctrl)) {
+ if (!cxlds->bi)
+ return 0;
+ dev_err(cxlds->dev,
+ "BI already disabled in hardware\n");
+ return -EBUSY;
+ }
+ }
+
+ FIELD_MODIFY(CXL_BI_DECODER_CTRL_BI_ENABLE, &ctrl, enable);
+ writel(ctrl, bi + CXL_BI_DECODER_CTRL_OFFSET);
+ cxlds->bi = enable;
+
+ dev_dbg(cxlds->dev, "BI requests %s\n", str_enabled_disabled(enable));
+
+ return 0;
+}
+
+static int cxl_bi_ctrl_endpoint_enable(struct cxl_dev_state *cxlds)
+{
+ return __cxl_bi_ctrl_endpoint(cxlds, true);
+}
+
+static int cxl_bi_ctrl_endpoint_disable(struct cxl_dev_state *cxlds)
+{
+ return __cxl_bi_ctrl_endpoint(cxlds, false);
+}
+
+/*
+ * devm teardown on endpoint port destruction. Registered before the
+ * decoders, so devres runs it after them: regions are detached and
+ * decoders unregistered by the time BI comes down.
+ */
+static void cxl_bi_dealloc(void *data)
+{
+ struct cxl_port *endpoint = data;
+ struct cxl_memdev *cxlmd = to_cxl_memdev(endpoint->uport_dev);
+ struct cxl_dev_state *cxlds = cxlmd->cxlds;
+ struct cxl_dport *dport_iter;
+ struct cxl_port *port_iter;
+
+ cxl_bi_ctrl_endpoint_disable(cxlds);
+ cxlds->bi = false;
+
+ /*
+ * Walk the same parent_dport chain that enabled the path. A bus
+ * lookup cannot stand in for it: an ancestor-driven teardown
+ * delists the parent port before this devres action runs.
+ */
+ dport_iter = endpoint->parent_dport;
+ port_iter = dport_iter->port;
+ while (!is_cxl_root(port_iter)) {
+ int rc = cxl_bi_ctrl_dport_disable(dport_iter);
+
+ /* best effort */
+ if (rc)
+ dev_dbg(&port_iter->dev,
+ "BI dport disable failed: %d\n", rc);
+
+ dport_iter = port_iter->parent_dport;
+ port_iter = dport_iter->port;
+ }
+}
+
+/*
+ * Enable BI on every dport in the path, then on the device itself.
+ * On failure, unwind only the dports that enabled.
+ */
+static int cxl_bi_enable_path(struct cxl_dev_state *cxlds,
+ struct cxl_port *port, struct cxl_dport *dport)
+{
+ struct cxl_dport *dport_iter, *failed;
+ struct cxl_port *port_iter;
+ int rc;
+
+ port_iter = port;
+ dport_iter = dport;
+ while (!is_cxl_root(port_iter)) {
+ rc = cxl_bi_ctrl_dport_enable(dport_iter, dport_iter == dport);
+ if (rc)
+ goto err_rollback;
+
+ dport_iter = port_iter->parent_dport;
+ port_iter = dport_iter->port;
+ }
+
+ rc = cxl_bi_ctrl_endpoint_enable(cxlds);
+ if (rc)
+ goto err_rollback;
+
+ return 0;
+
+err_rollback:
+ failed = dport_iter;
+ dport_iter = dport;
+ port_iter = port;
+ while (!is_cxl_root(port_iter) && dport_iter != failed) {
+ cxl_bi_ctrl_dport_disable(dport_iter);
+ dport_iter = port_iter->parent_dport;
+ port_iter = dport_iter->port;
+ }
+ return rc;
+}
+
+/*
+ * An SBR wipes the device's BI Enable; an FLR leaves it alone.
+ * The check is against the hardware, not decoder state: BI is
+ * enabled at probe, so it can be wiped with no decoder ever
+ * committed. A wipe invalidates the software state; BI is never
+ * re-enabled here.
+ */
+void cxl_bi_reset_detected(struct cxl_port *endpoint)
+{
+ struct cxl_memdev *cxlmd = to_cxl_memdev(endpoint->uport_dev);
+ struct cxl_dev_state *cxlds = cxlmd->cxlds;
+
+ if (!cxlds->bi)
+ return;
+
+ if (cxl_bi_decoder_enabled(endpoint))
+ return;
+
+ dev_warn(cxlds->dev, "BI disabled by reset\n");
+ cxlds->bi = false;
+}
+EXPORT_SYMBOL_NS_GPL(cxl_bi_reset_detected, "CXL");
+
+int cxl_bi_setup(struct cxl_port *endpoint)
+{
+ struct cxl_memdev *cxlmd = to_cxl_memdev(endpoint->uport_dev);
+ struct cxl_dev_state *cxlds = cxlmd->cxlds;
+ struct cxl_dport *dport = endpoint->parent_dport;
+ struct cxl_dport *dport_iter;
+ struct cxl_port *port_iter;
+ int rc;
+
+ if (!dev_is_pci(cxlds->dev))
+ return 0;
+
+ /* BI is VH-only */
+ if (cxlds->rcd)
+ return 0;
+
+ if (!cxl_is_bi_capable(to_pci_dev(cxlds->dev),
+ endpoint->regs.bi_decoder))
+ return 0;
+
+ port_iter = dport->port;
+ dport_iter = dport;
+ while (!is_cxl_root(port_iter)) {
+ /* check rp, dsp */
+ if (!cxl_is_bi_capable(to_pci_dev(dport_iter->dport_dev),
+ dport_iter->regs.bi_decoder)) {
+ dev_dbg(cxlds->dev, "BI not supported by topology\n");
+ return 0;
+ }
+
+ /* check usp */
+ if (dev_is_pci(port_iter->uport_dev) &&
+ pci_pcie_type(to_pci_dev(port_iter->uport_dev)) ==
+ PCI_EXP_TYPE_UPSTREAM) {
+ if (!cxl_is_bi_capable(to_pci_dev(port_iter->uport_dev),
+ NULL)) {
+ dev_dbg(cxlds->dev,
+ "BI not supported by USP\n");
+ return 0;
+ }
+ if (port_iter->reg_map.component_map.bi_rt.valid &&
+ !port_iter->regs.bi_rt) {
+ dev_dbg(cxlds->dev,
+ "BI RT advertised but unmapped\n");
+ return 0;
+ }
+ }
+
+ dport_iter = port_iter->parent_dport;
+ port_iter = dport_iter->port;
+ }
+
+ rc = cxl_bi_enable_path(cxlds, dport->port, dport);
+ if (rc)
+ return rc;
+
+ return devm_add_action_or_reset(&endpoint->dev, cxl_bi_dealloc,
+ endpoint);
+}
+EXPORT_SYMBOL_NS_GPL(cxl_bi_setup, "CXL");
diff --git a/drivers/cxl/core/port.c b/drivers/cxl/core/port.c
index 131ed62e8db3..b81fd680d18a 100644
--- a/drivers/cxl/core/port.c
+++ b/drivers/cxl/core/port.c
@@ -742,6 +742,7 @@ static struct cxl_port *cxl_port_alloc(struct device *uport_dev,
xa_init(&port->dports);
xa_init(&port->endpoints);
xa_init(&port->regions);
+ mutex_init(&port->bi_lock);
port->component_reg_phys = CXL_RESOURCE_NONE;
device_initialize(dev);
diff --git a/drivers/cxl/cxl.h b/drivers/cxl/cxl.h
index 65be3b91259a..ad7c991f0f9b 100644
--- a/drivers/cxl/cxl.h
+++ b/drivers/cxl/cxl.h
@@ -180,6 +180,31 @@ static inline int ways_to_eiw(unsigned int ways, u8 *eiw)
#define CXL_HEADERLOG_TRACE_SIZE SZ_512
#define CXL_HEADERLOG_TRACE_SIZE_U32 (CXL_HEADERLOG_TRACE_SIZE / sizeof(u32))
+/* CXL 4.0 8.2.4.26 CXL BI Route Table Capability Structure */
+#define CXL_BI_RT_CAPS_OFFSET 0x0
+#define CXL_BI_RT_CAPS_EXPLICIT_COMMIT_REQ BIT(0)
+#define CXL_BI_RT_CTRL_OFFSET 0x4
+#define CXL_BI_RT_CTRL_BI_COMMIT BIT(0)
+#define CXL_BI_RT_STATUS_OFFSET 0x8
+#define CXL_BI_RT_STATUS_BI_COMMITTED BIT(0)
+#define CXL_BI_RT_STATUS_BI_ERR_NOT_COMMITTED BIT(1)
+#define CXL_BI_RT_STATUS_BI_COMMIT_TM_SCALE GENMASK(11, 8)
+#define CXL_BI_RT_STATUS_BI_COMMIT_TM_BASE GENMASK(15, 12)
+
+/* CXL 4.0 8.2.4.27 CXL BI Decoder Capability Structure */
+#define CXL_BI_DECODER_CAPS_OFFSET 0x0
+#define CXL_BI_DECODER_CAPS_HDMD_CAP BIT(0)
+#define CXL_BI_DECODER_CAPS_EXPLICIT_COMMIT_REQ BIT(1)
+#define CXL_BI_DECODER_CTRL_OFFSET 0x4
+#define CXL_BI_DECODER_CTRL_BI_FW BIT(0)
+#define CXL_BI_DECODER_CTRL_BI_ENABLE BIT(1)
+#define CXL_BI_DECODER_CTRL_BI_COMMIT BIT(2)
+#define CXL_BI_DECODER_STATUS_OFFSET 0x8
+#define CXL_BI_DECODER_STATUS_BI_COMMITTED BIT(0)
+#define CXL_BI_DECODER_STATUS_BI_ERR_NOT_COMMITTED BIT(1)
+#define CXL_BI_DECODER_STATUS_BI_COMMIT_TM_SCALE GENMASK(11, 8)
+#define CXL_BI_DECODER_STATUS_BI_COMMIT_TM_BASE GENMASK(15, 12)
+
/* CXL 2.0 8.2.8.1 Device Capabilities Array Register */
#define CXLDEV_CAP_ARRAY_OFFSET 0x0
#define CXLDEV_CAP_ARRAY_CAP_ID 0
@@ -565,6 +590,7 @@ struct cxl_dax_region {
* @decoder_ida: allocator for decoder ids
* @reg_map: component and ras register mapping parameters
* @regs: mapped component registers
+ * @bi_lock: serializes BI Decoder/RT state of this port's dports
* @nr_dports: number of entries in @dports
* @hdm_end: track last allocated HDM decoder instance for allocation ordering
* @commit_end: cursor to track highest committed decoder for commit ordering
@@ -587,6 +613,7 @@ struct cxl_port {
struct ida decoder_ida;
struct cxl_register_map reg_map;
struct cxl_component_regs regs;
+ struct mutex bi_lock; /* dport BI state shared below this port */
int nr_dports;
int hdm_end;
int commit_end;
@@ -650,6 +677,7 @@ struct cxl_rcrb_info {
* @coord: access coordinates (bandwidth and latency performance attributes)
* @link_latency: calculated PCIe downstream latency
* @gpf_dvsec: Cached GPF port DVSEC
+ * @nr_bi: number of BI-enabled endpoints below this dport
*/
struct cxl_dport {
struct device *dport_dev;
@@ -662,6 +690,7 @@ struct cxl_dport {
struct access_coordinate coord[ACCESS_COORDINATE_MAX];
long link_latency;
int gpf_dvsec;
+ int nr_bi;
};
/**
@@ -920,6 +949,8 @@ void cxl_coordinates_combine(struct access_coordinate *out,
struct access_coordinate *c2);
bool cxl_endpoint_decoder_reset_detected(struct cxl_port *port);
+int cxl_bi_setup(struct cxl_port *endpoint);
+void cxl_bi_reset_detected(struct cxl_port *endpoint);
struct cxl_dport *devm_cxl_add_dport_by_dev(struct cxl_port *port,
struct device *dport_dev);
diff --git a/drivers/cxl/pci.c b/drivers/cxl/pci.c
index c7c91e8dc51d..cdab4804dba7 100644
--- a/drivers/cxl/pci.c
+++ b/drivers/cxl/pci.c
@@ -987,8 +987,12 @@ static void cxl_reset_done(struct pci_dev *pdev)
if (!cxlmd->dev.driver)
return;
- if (cxlmd->endpoint &&
- cxl_endpoint_decoder_reset_detected(cxlmd->endpoint)) {
+ if (!cxlmd->endpoint)
+ return;
+
+ cxl_bi_reset_detected(cxlmd->endpoint);
+
+ if (cxl_endpoint_decoder_reset_detected(cxlmd->endpoint)) {
device_for_each_child(&cxlmd->endpoint->dev, NULL,
cxl_endpoint_decoder_clear_reset_flags);
diff --git a/include/cxl/cxl.h b/include/cxl/cxl.h
index 278b84b08c83..718eb4353887 100644
--- a/include/cxl/cxl.h
+++ b/include/cxl/cxl.h
@@ -168,6 +168,7 @@ struct cxl_dpa_partition {
* @regs: Parsed register blocks
* @cxl_dvsec: Offset to the PCIe device DVSEC
* @rcd: operating in RCD mode (CXL 3.0 9.11.8 CXL Devices Attached to an RCH)
+ * @bi: device is BI (Back-Invalidate) enabled
* @media_ready: Indicate whether the device media is usable
* @dpa_res: Overall DPA resource tree for the device
* @part: DPA partition array
@@ -187,6 +188,7 @@ struct cxl_dev_state {
struct cxl_device_regs regs;
int cxl_dvsec;
bool rcd;
+ bool bi;
bool media_ready;
struct resource dpa_res;
struct cxl_dpa_partition part[CXL_NR_PARTITIONS_MAX];
--
2.39.5
^ permalink raw reply related [flat|nested] 33+ messages in thread* Re: [PATCH v9 02/10] cxl/pci: Add BI topology enable/disable
2026-09-22 23:38 ` [PATCH v9 02/10] cxl/pci: Add BI topology enable/disable Davidlohr Bueso
@ 2026-09-23 1:12 ` sashiko-bot
2026-09-23 17:00 ` Davidlohr Bueso
2026-09-23 5:41 ` Li Ming
` (3 subsequent siblings)
4 siblings, 1 reply; 33+ messages in thread
From: sashiko-bot @ 2026-09-23 1:12 UTC (permalink / raw)
To: Davidlohr Bueso; +Cc: linux-cxl
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] cxl/pci: TOCTOU race in BI setup against concurrent AER reset
--
commit 1daa4037cbac0ebc16630c3d4f45c358bd7890f9
Author: Davidlohr Bueso <dave@stgolabs.net>
cxl/pci: Add BI topology enable/disable
This patch implements setup and teardown of the Back-Invalidate (BI) topology
on CXL endpoints and upstream components. It also adds a reset detection
mechanism to verify whether a reset cleared the hardware BI Enable state.
> diff --git a/drivers/cxl/core/pci.c b/drivers/cxl/core/pci.c
> index b8676a3d6ec90..f3ca9861a7ba1 100644
> --- a/drivers/cxl/core/pci.c
> +++ b/drivers/cxl/core/pci.c
[ ... ]
> +static int __cxl_bi_ctrl_endpoint(struct cxl_dev_state *cxlds, bool enable)
> +{
[ ... ]
> + FIELD_MODIFY(CXL_BI_DECODER_CTRL_BI_ENABLE, &ctrl, enable);
> + writel(ctrl, bi + CXL_BI_DECODER_CTRL_OFFSET);
> + cxlds->bi = enable;
[Severity: Medium]
Could this lockless update of cxlds->bi race with a concurrent AER reset
handled by cxl_bi_reset_detected()?
If a user manually binds the cxl_port driver via sysfs concurrently with a
PCI AER reset (SBR), the manual sysfs bind takes the endpoint->dev lock,
while the AER recovery (cxl_reset_done()) takes the cxlmd->dev lock, allowing
them to execute concurrently.
There is a window between the writel() that sets the hardware bit and the
assignment of cxlds->bi. If the AER SBR wipes the hardware bit in this window,
could cxl_bi_reset_detected() check the still-false cxlds->bi and ignore the
reset?
> +
> + dev_dbg(cxlds->dev, "BI requests %s\n", str_enabled_disabled(enable));
> +
> + return 0;
> +}
[ ... ]
> +void cxl_bi_reset_detected(struct cxl_port *endpoint)
> +{
> + struct cxl_memdev *cxlmd = to_cxl_memdev(endpoint->uport_dev);
> + struct cxl_dev_state *cxlds = cxlmd->cxlds;
> +
> + if (!cxlds->bi)
> + return;
If cxl_bi_reset_detected() returns early here because cxlds->bi is false, and
then __cxl_bi_ctrl_endpoint() resumes and sets cxlds->bi = true, the software
state becomes permanently desynchronized from the hardware state
(BI_ENABLE = 0), breaking the fail-closed coherency recovery model.
> +
> + if (cxl_bi_decoder_enabled(endpoint))
> + return;
> +
> + dev_warn(cxlds->dev, "BI disabled by reset\n");
> + cxlds->bi = false;
> +}
--
Sashiko AI review · https://sashiko.dev/#/patchset/cover.1790103847.git.dave@stgolabs.net?part=2
^ permalink raw reply [flat|nested] 33+ messages in thread* Re: [PATCH v9 02/10] cxl/pci: Add BI topology enable/disable
2026-09-23 1:12 ` sashiko-bot
@ 2026-09-23 17:00 ` Davidlohr Bueso
0 siblings, 0 replies; 33+ messages in thread
From: Davidlohr Bueso @ 2026-09-23 17:00 UTC (permalink / raw)
To: sashiko-reviews, dave.jiang; +Cc: linux-cxl
On Wed, 23 Sep 2026, sashiko-bot@kernel.org wrote:
>Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
>- [Medium] cxl/pci: TOCTOU race in BI setup against concurrent AER reset
False positive, see below.
>--
>
>commit 1daa4037cbac0ebc16630c3d4f45c358bd7890f9
>Author: Davidlohr Bueso <dave@stgolabs.net>
>
>cxl/pci: Add BI topology enable/disable
>
>This patch implements setup and teardown of the Back-Invalidate (BI) topology
>on CXL endpoints and upstream components. It also adds a reset detection
>mechanism to verify whether a reset cleared the hardware BI Enable state.
>
>> diff --git a/drivers/cxl/core/pci.c b/drivers/cxl/core/pci.c
>> index b8676a3d6ec90..f3ca9861a7ba1 100644
>> --- a/drivers/cxl/core/pci.c
>> +++ b/drivers/cxl/core/pci.c
>[ ... ]
>> +static int __cxl_bi_ctrl_endpoint(struct cxl_dev_state *cxlds, bool enable)
>> +{
>[ ... ]
>> + FIELD_MODIFY(CXL_BI_DECODER_CTRL_BI_ENABLE, &ctrl, enable);
>> + writel(ctrl, bi + CXL_BI_DECODER_CTRL_OFFSET);
>> + cxlds->bi = enable;
>
>[Severity: Medium]
>Could this lockless update of cxlds->bi race with a concurrent AER reset
>handled by cxl_bi_reset_detected()?
>
>If a user manually binds the cxl_port driver via sysfs concurrently with a
>PCI AER reset (SBR), the manual sysfs bind takes the endpoint->dev lock,
>while the AER recovery (cxl_reset_done()) takes the cxlmd->dev lock, allowing
>them to execute concurrently.
The endpoint port cannot be bound on its own. Unbinding it from cxl_port
tears it down through cxl_detach_ep() and a subsequent sysfs bind fails
with -ENODEV. The only path that reaches __cxl_bi_ctrl_endpoint(enable) is
the endpoint port probe, and that runs synchronously from
devm_cxl_add_endpoint().
The driver core holds device_lock(&cxlmd->dev) across cxl_mem_probe(), and
cxl_reset_done() takes that same lock before calling
cxl_bi_reset_detected(). The reset side therefore cannot observe the
window between the writel() and the cxlds->bi store; it waits for the
probe to return and only then compares the flag against the hardware bit.
>
>There is a window between the writel() that sets the hardware bit and the
>assignment of cxlds->bi. If the AER SBR wipes the hardware bit in this window,
>could cxl_bi_reset_detected() check the still-false cxlds->bi and ignore the
>reset?
I widened the window with a 4s msleep() between the writel() and the store,
rebinded mem0 and issued an SBR through sysfs while the probe was sleeping
in that window:
[16.846] cxl_pci 0000:0d:00.0: reset via cxl_bus
[19.290] cxl_pci 0000:0d:00.0: BI requests enabled
[19.335] cxl_mem mem0: probe: 0
[19.337] cxl_pci 0000:0d:00.0: BI disabled by reset
The SBR itself lands in the window and clears BI Enable, but the reset
write blocks on the memdev lock until the probe completes. .reset_done then
runs, sees cxlds->bi set with the hardware bit clear, and clears the flag.
A subsequent HDM-DB attach is refused with "BI not enabled on device". The
software state does not desynchronize.
Independently of this, patch 4 re-reads BI Enable after committing each
endpoint decoder, so a stale flag could not commit a decoder with BI on a
device whose BI Enable is clear.
Thanks,
Davidlohr
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: [PATCH v9 02/10] cxl/pci: Add BI topology enable/disable
2026-09-22 23:38 ` [PATCH v9 02/10] cxl/pci: Add BI topology enable/disable Davidlohr Bueso
2026-09-23 1:12 ` sashiko-bot
@ 2026-09-23 5:41 ` Li Ming
2026-09-25 23:32 ` Jonathan Cameron
` (2 subsequent siblings)
4 siblings, 0 replies; 33+ messages in thread
From: Li Ming @ 2026-09-23 5:41 UTC (permalink / raw)
To: Davidlohr Bueso, dave.jiang
Cc: jic23, alison.schofield, icheng, benjamin.cheatham, alucerop,
linux-cxl
在 2026/9/23 07:38, Davidlohr Bueso 写道:
> Implement cxl_bi_setup() to enable BI flows on the device and every
> component in the path, and its teardown counterpart cxl_bi_dealloc().
> Setup runs from devm_cxl_endpoint_decoders_setup(), between the
> port's HDM state and its decoders. Registered there, its devres
> teardown brings BI down after the decoders quiesce and before the
> HDM state is freed, and BI is settled before the decoders, and later
> the regions, are looked at. The BI-ID and path enablement belong to
> the endpoint port's lifetime.
>
> Setup is safe in endpoint port probe context. The port probes
> synchronously from cxl_mem_probe(), pinning the memdev state the
> walk consumes, and the whole ancestor path already exists with BI
> registers mapped (dports at dport-add time, the switch USP RT at
> first-dport setup) because devm_cxl_enumerate_ports() completes
> before the endpoint is created.
>
> Dealloc is safe in endpoint devres context. Both setup and dealloc
> walk the endpoint's parent_dport topology rather than getting the
> port by bus lookup - an ancestor teardown delists the parent port
> before the endpoint's devres runs.
>
> The topology walk is stable as parent_dport pointers are fixed at
> port creation; ancestors cannot be reaped while holding this
> memdev's cxl_ep; and their own teardown frees dports only after
> the endpoint is gone.
>
> Likewise, the device state outlives the walk - cxlmd->cxlds is
> nulled only after cxl_memdev_unregister() has torn the endpoint
> down, and delete_endpoint() clears cxlmd->endpoint only after the
> endpoint devres has run.
>
> Each dport is programmed by its position - the one immediately above
> the device takes BI Enable, every dport above it takes BI Forward
> (Table 8-157, Table 9-13), at any switch depth (Table 7-97). Any
> level can be shared, so nr_bi refcounts endpoints at every dport.
> Registers are written on the first endpoint and cleared on the last,
> but only downstream ports commit (Table 8-156), once per endpoint
> (Table 8-152), and a failed commit undoes its write and commits the
> undo. A USP advertising a BI Route Table that failed to map is
> refused rather than treated as absent. nr_bi counts only the
> endpoints this driver enabled, so a level can be cleared while
> firmware still has an unbound device on it.
>
> A reset may wipe the device's BI Enable, whose reset default is 0
> (Table 8-157). .reset_done reads the hardware rather than assume
> which reset ran, and invalidates cxlds->bi, failing closed with
> recovery by rebind as for decoder loss; dealloc unwinds the dport
> refcounts regardless. It also clears cxlds->bi unconditionally - the
> endpoint disable fails when the hardware already shows BI Enable
> clear, from a reset .reset_done never saw, and the flag must not
> outlive the BI Decoder mapping it describes, which the endpoint port
> releases moments later.
>
> With dealloc in the endpoint's devres, delete_endpoint() already
> holds the parent port's device lock, so to avoid deadlocking, add a
> per-port bi_lock, serializing the dports that share state (nr_bi and
> the control register at any shared level, the switch USP's BI RT).
>
> Reviewed-by: Ben Cheatham <benjamin.cheatham@amd.com>
> Signed-off-by: Davidlohr Bueso <dave@stgolabs.net>
Reviewed-by: Li Ming <ming.li@zohomail.com>
Just one minor comment below
> ---
> drivers/cxl/core/core.h | 7 +
> drivers/cxl/core/hdm.c | 10 +
> drivers/cxl/core/pci.c | 427 ++++++++++++++++++++++++++++++++++++++++
> drivers/cxl/core/port.c | 1 +
> drivers/cxl/cxl.h | 31 +++
> drivers/cxl/pci.c | 8 +-
> include/cxl/cxl.h | 2 +
> 7 files changed, 484 insertions(+), 2 deletions(-)
>
> diff --git a/drivers/cxl/core/core.h b/drivers/cxl/core/core.h
> index 384bad8cb70f..ef41632d7b7d 100644
> --- a/drivers/cxl/core/core.h
> +++ b/drivers/cxl/core/core.h
> @@ -210,6 +210,13 @@ static inline void devm_cxl_dport_ras_setup(struct cxl_dport *dport) { }
> int cxl_gpf_port_setup(struct cxl_dport *dport);
> void devm_cxl_dport_bi_setup(struct cxl_dport *dport);
>
> +static inline bool cxl_bi_decoder_enabled(struct cxl_port *port)
> +{
> + return FIELD_GET(CXL_BI_DECODER_CTRL_BI_ENABLE,
> + readl(port->regs.bi_decoder +
> + CXL_BI_DECODER_CTRL_OFFSET));
> +}
> +
> struct cxl_hdm;
> int cxl_hdm_decode_init(struct cxl_dev_state *cxlds, struct cxl_hdm *cxlhdm,
> struct cxl_endpoint_dvsec_info *info);
> diff --git a/drivers/cxl/core/hdm.c b/drivers/cxl/core/hdm.c
> index 0c80b76a5f9b..0e9d652b568e 100644
> --- a/drivers/cxl/core/hdm.c
> +++ b/drivers/cxl/core/hdm.c
> @@ -1274,6 +1274,16 @@ int devm_cxl_endpoint_decoders_setup(struct cxl_port *port)
> if (rc)
> return rc;
>
> + /*
> + * Between the port's HDM state and its decoders: devres,
> + * unwinding in reverse, brings BI down only after the decoders
> + * quiesce, while its slow walk still precedes the HDM state
> + * free.
> + */
> + rc = cxl_bi_setup(port);
> + if (rc)
> + dev_dbg(&port->dev, "BI setup failed rc=%d\n", rc);
> +
> return devm_cxl_enumerate_decoders(cxlhdm, &info);
> }
> EXPORT_SYMBOL_NS_GPL(devm_cxl_endpoint_decoders_setup, "CXL");
> diff --git a/drivers/cxl/core/pci.c b/drivers/cxl/core/pci.c
> index b8676a3d6ec9..f3ca9861a7ba 100644
> --- a/drivers/cxl/core/pci.c
> +++ b/drivers/cxl/core/pci.c
> @@ -2,12 +2,14 @@
> /* Copyright(c) 2021 Intel Corporation. All rights reserved. */
> #include <linux/units.h>
> #include <linux/io-64-nonatomic-lo-hi.h>
> +#include <linux/iopoll.h>
> #include <linux/device.h>
> #include <linux/delay.h>
> #include <linux/pci.h>
> #include <linux/pci-doe.h>
> #include <cxl/pci.h>
> #include <linux/aer.h>
> +#include <linux/string_choices.h>
> #include <cxlpci.h>
> #include <cxlmem.h>
> #include <cxl.h>
> @@ -965,3 +967,428 @@ void devm_cxl_dport_bi_setup(struct cxl_dport *dport)
> break;
> }
> }
> +
> +/*
> + * BI requires 256B Flit operation on the link. RP/DSP/endpoint must
> + * also have the BI Decoder cap mapped (@bi); for USPs the BI RT cap
> + * is optional per CXL 4.0 8.2.4.26, so absent @bi is allowed.
> + */
> +static bool cxl_is_bi_capable(struct pci_dev *pdev, void __iomem *bi)
> +{
> + if (!cxl_pci_flit_256(pdev))
> + return false;
> +
> + if (pci_pcie_type(pdev) != PCI_EXP_TYPE_UPSTREAM && !bi) {
> + dev_dbg(&pdev->dev, "No BI Decoder registers.\n");
> + return false;
> + }
> +
> + return true;
> +}
> +
> +/* limit any insane timeouts from hw */
> +#define CXL_BI_COMMIT_MAXTMO_US (20 * USEC_PER_SEC)
> +
> +static unsigned long __cxl_bi_get_timeout_us(struct device *dev,
> + unsigned int scale,
> + unsigned int base)
> +{
> + static const unsigned long scale_tbl[] = {
> + 1, 10, 100, 1000, 10000, 100000, 1000000, 10000000,
> + };
> +
> + if (scale >= ARRAY_SIZE(scale_tbl) || !base) {
> + dev_dbg(dev, "Invalid BI commit timeout: scale=%u base=%u\n",
> + scale, base);
> + return CXL_BI_COMMIT_MAXTMO_US;
> + }
> +
> + return scale_tbl[scale] * base;
> +}
> +
> +static int __cxl_bi_wait_commit(struct device *dev, void __iomem *status_reg,
> + u32 committed_bit, u32 err_bit,
> + unsigned int scale, unsigned int base)
> +{
> + unsigned long tmo_us, poll_us;
> + ktime_t start;
> + u32 status;
> + int rc;
> +
> + tmo_us = min_t(unsigned long, CXL_BI_COMMIT_MAXTMO_US,
> + __cxl_bi_get_timeout_us(dev, scale, base));
> + poll_us = max_t(unsigned long, tmo_us / 10, 1); /* ~10% */
> + start = ktime_get();
> +
> + rc = readx_poll_timeout(readl, status_reg, status,
> + status & (committed_bit | err_bit),
> + poll_us, tmo_us);
> + if (rc) {
> + dev_err(dev, "BI-ID commit timed out (%luus)\n", tmo_us);
> + return rc; /* -ETIMEDOUT */
> + }
> +
> + if (status & err_bit) {
> + dev_err(dev, "BI-ID commit rejected by hardware\n");
> + return -EIO;
> + }
> +
> + dev_dbg(dev, "BI-ID commit wait took %lluus\n",
> + ktime_to_us(ktime_sub(ktime_get(), start)));
> + return 0;
> +}
> +
> +/* BI RT only exists on switch upstream ports. */
> +static int __cxl_bi_commit_rt(struct device *dev, void __iomem *bi)
> +{
> + u32 status, ctrl;
> + unsigned int scale, base;
> +
> + if (!FIELD_GET(CXL_BI_RT_CAPS_EXPLICIT_COMMIT_REQ,
> + readl(bi + CXL_BI_RT_CAPS_OFFSET)))
> + return 0;
> +
> + ctrl = readl(bi + CXL_BI_RT_CTRL_OFFSET);
> + writel(ctrl & ~CXL_BI_RT_CTRL_BI_COMMIT, bi + CXL_BI_RT_CTRL_OFFSET);
> + writel(ctrl | CXL_BI_RT_CTRL_BI_COMMIT, bi + CXL_BI_RT_CTRL_OFFSET);
> +
> + status = readl(bi + CXL_BI_RT_STATUS_OFFSET);
> + scale = FIELD_GET(CXL_BI_RT_STATUS_BI_COMMIT_TM_SCALE, status);
> + base = FIELD_GET(CXL_BI_RT_STATUS_BI_COMMIT_TM_BASE, status);
> +
> + return __cxl_bi_wait_commit(dev, bi + CXL_BI_RT_STATUS_OFFSET,
> + CXL_BI_RT_STATUS_BI_COMMITTED,
> + CXL_BI_RT_STATUS_BI_ERR_NOT_COMMITTED,
> + scale, base);
> +}
> +
> +static int __cxl_bi_commit_decoder(struct device *dev, void __iomem *bi)
> +{
> + u32 status, ctrl;
> + unsigned int scale, base;
> +
> + if (!FIELD_GET(CXL_BI_DECODER_CAPS_EXPLICIT_COMMIT_REQ,
> + readl(bi + CXL_BI_DECODER_CAPS_OFFSET)))
> + return 0;
> +
> + ctrl = readl(bi + CXL_BI_DECODER_CTRL_OFFSET);
> + writel(ctrl & ~CXL_BI_DECODER_CTRL_BI_COMMIT,
> + bi + CXL_BI_DECODER_CTRL_OFFSET);
> + writel(ctrl | CXL_BI_DECODER_CTRL_BI_COMMIT,
> + bi + CXL_BI_DECODER_CTRL_OFFSET);
> +
> + status = readl(bi + CXL_BI_DECODER_STATUS_OFFSET);
> + scale = FIELD_GET(CXL_BI_DECODER_STATUS_BI_COMMIT_TM_SCALE, status);
> + base = FIELD_GET(CXL_BI_DECODER_STATUS_BI_COMMIT_TM_BASE, status);
> +
> + return __cxl_bi_wait_commit(dev, bi + CXL_BI_DECODER_STATUS_OFFSET,
> + CXL_BI_DECODER_STATUS_BI_COMMITTED,
> + CXL_BI_DECODER_STATUS_BI_ERR_NOT_COMMITTED,
> + scale, base);
> +}
> +
> +static int cxl_bi_commit_dport(struct cxl_dport *dport)
> +{
> + struct cxl_port *port = dport->port;
> + int rc;
> +
> + lockdep_assert_held(&port->bi_lock);
> +
> + /* root ports never require the explicit commit */
> + if (pci_pcie_type(to_pci_dev(dport->dport_dev)) !=
> + PCI_EXP_TYPE_DOWNSTREAM)
> + return 0;
> +
> + rc = __cxl_bi_commit_decoder(dport->dport_dev, dport->regs.bi_decoder);
> + if (rc)
> + return rc;
> +
> + if (port->regs.bi_rt)
> + rc = __cxl_bi_commit_rt(&port->dev, port->regs.bi_rt);
> +
> + return rc;
> +}
> +
> +/*
> + * Enable BI-ID changes in the given level of the topology.
> + * @direct says the device is connected directly to this dport, which
> + * takes BI Enable; every dport above it takes BI Forward.
> + */
> +static int cxl_bi_ctrl_dport_enable(struct cxl_dport *dport, bool direct)
> +{
> + struct cxl_port *port = dport->port;
> + u32 ctrl, value, set, clr;
> + void __iomem *bi;
> + int rc;
> +
> + guard(mutex)(&port->bi_lock);
> +
> + bi = dport->regs.bi_decoder;
> + if (!bi)
> + return -EINVAL;
> +
> + ctrl = readl(bi + CXL_BI_DECODER_CTRL_OFFSET);
> +
> + set = direct ? CXL_BI_DECODER_CTRL_BI_ENABLE :
> + CXL_BI_DECODER_CTRL_BI_FW;
> + clr = direct ? CXL_BI_DECODER_CTRL_BI_FW :
> + CXL_BI_DECODER_CTRL_BI_ENABLE;
> +
> + value = (ctrl | set) & ~clr;
> + if (value != ctrl)
> + writel(value, bi + CXL_BI_DECODER_CTRL_OFFSET);
> +
> + /* owed per new device below, not per register change */
> + rc = cxl_bi_commit_dport(dport);
> + if (rc) {
> + if (value != ctrl) {
> + /* the undo is a BI-ID change owing its own commit */
> + writel(ctrl, bi + CXL_BI_DECODER_CTRL_OFFSET);
> + cxl_bi_commit_dport(dport);
> + }
> + return rc;
> + }
> + dport->nr_bi++;
> +
> + return 0;
> +}
> +
> +/*
> + * Dealloc BI-ID changes in the given level of the topology. Called
> + * once per endpoint that enabled this level: on teardown, or to
> + * unwind a path that failed partway up.
> + */
> +static int cxl_bi_ctrl_dport_disable(struct cxl_dport *dport)
> +{
> + struct cxl_port *port = dport->port;
> + void __iomem *bi;
> + u32 ctrl;
> +
> + guard(mutex)(&port->bi_lock);
> +
> + bi = dport->regs.bi_decoder;
> + if (!bi)
> + return -EINVAL;
> +
> + if (WARN_ON_ONCE(dport->nr_bi == 0))
> + return -EINVAL;
> +
> + /* others below still need this level */
> + if (--dport->nr_bi > 0)
> + return 0;
> +
> + ctrl = readl(bi + CXL_BI_DECODER_CTRL_OFFSET);
> + writel(ctrl & ~(CXL_BI_DECODER_CTRL_BI_FW |
> + CXL_BI_DECODER_CTRL_BI_ENABLE),
> + bi + CXL_BI_DECODER_CTRL_OFFSET);
> +
> + return cxl_bi_commit_dport(dport);
> +}
> +
> +static int __cxl_bi_ctrl_endpoint(struct cxl_dev_state *cxlds, bool enable)
> +{
> + struct cxl_port *endpoint = cxlds->cxlmd->endpoint;
> + void __iomem *bi = endpoint->regs.bi_decoder;
> + u32 ctrl;
> +
> + if (!bi)
> + return -EINVAL;
> +
> + ctrl = readl(bi + CXL_BI_DECODER_CTRL_OFFSET);
> +
> + if (enable) {
> + if (FIELD_GET(CXL_BI_DECODER_CTRL_BI_ENABLE, ctrl)) {
> + if (cxlds->bi)
> + return 0;
> + dev_err(cxlds->dev,
> + "BI already enabled in hardware\n");
> + return -EBUSY;
> + }
> + } else {
> + if (!FIELD_GET(CXL_BI_DECODER_CTRL_BI_ENABLE, ctrl)) {
> + if (!cxlds->bi)
> + return 0;
> + dev_err(cxlds->dev,
> + "BI already disabled in hardware\n");
> + return -EBUSY;
> + }
> + }
> +
> + FIELD_MODIFY(CXL_BI_DECODER_CTRL_BI_ENABLE, &ctrl, enable);
> + writel(ctrl, bi + CXL_BI_DECODER_CTRL_OFFSET);
> + cxlds->bi = enable;
> +
> + dev_dbg(cxlds->dev, "BI requests %s\n", str_enabled_disabled(enable));
> +
> + return 0;
> +}
> +
> +static int cxl_bi_ctrl_endpoint_enable(struct cxl_dev_state *cxlds)
> +{
> + return __cxl_bi_ctrl_endpoint(cxlds, true);
> +}
> +
> +static int cxl_bi_ctrl_endpoint_disable(struct cxl_dev_state *cxlds)
> +{
> + return __cxl_bi_ctrl_endpoint(cxlds, false);
> +}
> +
> +/*
> + * devm teardown on endpoint port destruction. Registered before the
> + * decoders, so devres runs it after them: regions are detached and
> + * decoders unregistered by the time BI comes down.
> + */
> +static void cxl_bi_dealloc(void *data)
> +{
> + struct cxl_port *endpoint = data;
> + struct cxl_memdev *cxlmd = to_cxl_memdev(endpoint->uport_dev);
> + struct cxl_dev_state *cxlds = cxlmd->cxlds;
> + struct cxl_dport *dport_iter;
> + struct cxl_port *port_iter;
> +
> + cxl_bi_ctrl_endpoint_disable(cxlds);
> + cxlds->bi = false;
> +
> + /*
> + * Walk the same parent_dport chain that enabled the path. A bus
> + * lookup cannot stand in for it: an ancestor-driven teardown
> + * delists the parent port before this devres action runs.
> + */
> + dport_iter = endpoint->parent_dport;
> + port_iter = dport_iter->port;
> + while (!is_cxl_root(port_iter)) {
> + int rc = cxl_bi_ctrl_dport_disable(dport_iter);
> +
> + /* best effort */
> + if (rc)
> + dev_dbg(&port_iter->dev,
> + "BI dport disable failed: %d\n", rc);
> +
> + dport_iter = port_iter->parent_dport;
> + port_iter = dport_iter->port;
> + }
> +}
> +
> +/*
> + * Enable BI on every dport in the path, then on the device itself.
> + * On failure, unwind only the dports that enabled.
> + */
I think it is better to mention the enable/disable flow implementation
based on the description of CXL r4.0 Section 9.14 Back-Invalidate
Configuration.
> +static int cxl_bi_enable_path(struct cxl_dev_state *cxlds,
> + struct cxl_port *port, struct cxl_dport *dport)
> +{
> + struct cxl_dport *dport_iter, *failed;
> + struct cxl_port *port_iter;
> + int rc;
> +
> + port_iter = port;
> + dport_iter = dport;
> + while (!is_cxl_root(port_iter)) {
> + rc = cxl_bi_ctrl_dport_enable(dport_iter, dport_iter == dport);
> + if (rc)
> + goto err_rollback;
> +
> + dport_iter = port_iter->parent_dport;
> + port_iter = dport_iter->port;
> + }
> +
> + rc = cxl_bi_ctrl_endpoint_enable(cxlds);
> + if (rc)
> + goto err_rollback;
> +
> + return 0;
> +
> +err_rollback:
> + failed = dport_iter;
> + dport_iter = dport;
> + port_iter = port;
> + while (!is_cxl_root(port_iter) && dport_iter != failed) {
> + cxl_bi_ctrl_dport_disable(dport_iter);
> + dport_iter = port_iter->parent_dport;
> + port_iter = dport_iter->port;
> + }
> + return rc;
> +}
> +
> +/*
> + * An SBR wipes the device's BI Enable; an FLR leaves it alone.
> + * The check is against the hardware, not decoder state: BI is
> + * enabled at probe, so it can be wiped with no decoder ever
> + * committed. A wipe invalidates the software state; BI is never
> + * re-enabled here.
> + */
> +void cxl_bi_reset_detected(struct cxl_port *endpoint)
> +{
> + struct cxl_memdev *cxlmd = to_cxl_memdev(endpoint->uport_dev);
> + struct cxl_dev_state *cxlds = cxlmd->cxlds;
> +
> + if (!cxlds->bi)
> + return;
> +
> + if (cxl_bi_decoder_enabled(endpoint))
> + return;
> +
> + dev_warn(cxlds->dev, "BI disabled by reset\n");
> + cxlds->bi = false;
> +}
> +EXPORT_SYMBOL_NS_GPL(cxl_bi_reset_detected, "CXL");
> +
> +int cxl_bi_setup(struct cxl_port *endpoint)
> +{
> + struct cxl_memdev *cxlmd = to_cxl_memdev(endpoint->uport_dev);
> + struct cxl_dev_state *cxlds = cxlmd->cxlds;
> + struct cxl_dport *dport = endpoint->parent_dport;
> + struct cxl_dport *dport_iter;
> + struct cxl_port *port_iter;
> + int rc;
> +
> + if (!dev_is_pci(cxlds->dev))
> + return 0;
> +
> + /* BI is VH-only */
> + if (cxlds->rcd)
> + return 0;
> +
> + if (!cxl_is_bi_capable(to_pci_dev(cxlds->dev),
> + endpoint->regs.bi_decoder))
> + return 0;
> +
> + port_iter = dport->port;
> + dport_iter = dport;
> + while (!is_cxl_root(port_iter)) {
> + /* check rp, dsp */
> + if (!cxl_is_bi_capable(to_pci_dev(dport_iter->dport_dev),
> + dport_iter->regs.bi_decoder)) {
> + dev_dbg(cxlds->dev, "BI not supported by topology\n");
> + return 0;
> + }
> +
> + /* check usp */
> + if (dev_is_pci(port_iter->uport_dev) &&
> + pci_pcie_type(to_pci_dev(port_iter->uport_dev)) ==
> + PCI_EXP_TYPE_UPSTREAM) {
> + if (!cxl_is_bi_capable(to_pci_dev(port_iter->uport_dev),
> + NULL)) {
> + dev_dbg(cxlds->dev,
> + "BI not supported by USP\n");
> + return 0;
> + }
> + if (port_iter->reg_map.component_map.bi_rt.valid &&
> + !port_iter->regs.bi_rt) {
> + dev_dbg(cxlds->dev,
> + "BI RT advertised but unmapped\n");
> + return 0;
> + }
> + }
> +
> + dport_iter = port_iter->parent_dport;
> + port_iter = dport_iter->port;
> + }
> +
> + rc = cxl_bi_enable_path(cxlds, dport->port, dport);
> + if (rc)
> + return rc;
> +
> + return devm_add_action_or_reset(&endpoint->dev, cxl_bi_dealloc,
> + endpoint);
> +}
> +EXPORT_SYMBOL_NS_GPL(cxl_bi_setup, "CXL");
> diff --git a/drivers/cxl/core/port.c b/drivers/cxl/core/port.c
> index 131ed62e8db3..b81fd680d18a 100644
> --- a/drivers/cxl/core/port.c
> +++ b/drivers/cxl/core/port.c
> @@ -742,6 +742,7 @@ static struct cxl_port *cxl_port_alloc(struct device *uport_dev,
> xa_init(&port->dports);
> xa_init(&port->endpoints);
> xa_init(&port->regions);
> + mutex_init(&port->bi_lock);
> port->component_reg_phys = CXL_RESOURCE_NONE;
>
> device_initialize(dev);
> diff --git a/drivers/cxl/cxl.h b/drivers/cxl/cxl.h
> index 65be3b91259a..ad7c991f0f9b 100644
> --- a/drivers/cxl/cxl.h
> +++ b/drivers/cxl/cxl.h
> @@ -180,6 +180,31 @@ static inline int ways_to_eiw(unsigned int ways, u8 *eiw)
> #define CXL_HEADERLOG_TRACE_SIZE SZ_512
> #define CXL_HEADERLOG_TRACE_SIZE_U32 (CXL_HEADERLOG_TRACE_SIZE / sizeof(u32))
>
> +/* CXL 4.0 8.2.4.26 CXL BI Route Table Capability Structure */
> +#define CXL_BI_RT_CAPS_OFFSET 0x0
> +#define CXL_BI_RT_CAPS_EXPLICIT_COMMIT_REQ BIT(0)
> +#define CXL_BI_RT_CTRL_OFFSET 0x4
> +#define CXL_BI_RT_CTRL_BI_COMMIT BIT(0)
> +#define CXL_BI_RT_STATUS_OFFSET 0x8
> +#define CXL_BI_RT_STATUS_BI_COMMITTED BIT(0)
> +#define CXL_BI_RT_STATUS_BI_ERR_NOT_COMMITTED BIT(1)
> +#define CXL_BI_RT_STATUS_BI_COMMIT_TM_SCALE GENMASK(11, 8)
> +#define CXL_BI_RT_STATUS_BI_COMMIT_TM_BASE GENMASK(15, 12)
> +
> +/* CXL 4.0 8.2.4.27 CXL BI Decoder Capability Structure */
> +#define CXL_BI_DECODER_CAPS_OFFSET 0x0
> +#define CXL_BI_DECODER_CAPS_HDMD_CAP BIT(0)
> +#define CXL_BI_DECODER_CAPS_EXPLICIT_COMMIT_REQ BIT(1)
> +#define CXL_BI_DECODER_CTRL_OFFSET 0x4
> +#define CXL_BI_DECODER_CTRL_BI_FW BIT(0)
> +#define CXL_BI_DECODER_CTRL_BI_ENABLE BIT(1)
> +#define CXL_BI_DECODER_CTRL_BI_COMMIT BIT(2)
> +#define CXL_BI_DECODER_STATUS_OFFSET 0x8
> +#define CXL_BI_DECODER_STATUS_BI_COMMITTED BIT(0)
> +#define CXL_BI_DECODER_STATUS_BI_ERR_NOT_COMMITTED BIT(1)
> +#define CXL_BI_DECODER_STATUS_BI_COMMIT_TM_SCALE GENMASK(11, 8)
> +#define CXL_BI_DECODER_STATUS_BI_COMMIT_TM_BASE GENMASK(15, 12)
> +
> /* CXL 2.0 8.2.8.1 Device Capabilities Array Register */
> #define CXLDEV_CAP_ARRAY_OFFSET 0x0
> #define CXLDEV_CAP_ARRAY_CAP_ID 0
> @@ -565,6 +590,7 @@ struct cxl_dax_region {
> * @decoder_ida: allocator for decoder ids
> * @reg_map: component and ras register mapping parameters
> * @regs: mapped component registers
> + * @bi_lock: serializes BI Decoder/RT state of this port's dports
> * @nr_dports: number of entries in @dports
> * @hdm_end: track last allocated HDM decoder instance for allocation ordering
> * @commit_end: cursor to track highest committed decoder for commit ordering
> @@ -587,6 +613,7 @@ struct cxl_port {
> struct ida decoder_ida;
> struct cxl_register_map reg_map;
> struct cxl_component_regs regs;
> + struct mutex bi_lock; /* dport BI state shared below this port */
> int nr_dports;
> int hdm_end;
> int commit_end;
> @@ -650,6 +677,7 @@ struct cxl_rcrb_info {
> * @coord: access coordinates (bandwidth and latency performance attributes)
> * @link_latency: calculated PCIe downstream latency
> * @gpf_dvsec: Cached GPF port DVSEC
> + * @nr_bi: number of BI-enabled endpoints below this dport
> */
> struct cxl_dport {
> struct device *dport_dev;
> @@ -662,6 +690,7 @@ struct cxl_dport {
> struct access_coordinate coord[ACCESS_COORDINATE_MAX];
> long link_latency;
> int gpf_dvsec;
> + int nr_bi;
> };
>
> /**
> @@ -920,6 +949,8 @@ void cxl_coordinates_combine(struct access_coordinate *out,
> struct access_coordinate *c2);
>
> bool cxl_endpoint_decoder_reset_detected(struct cxl_port *port);
> +int cxl_bi_setup(struct cxl_port *endpoint);
> +void cxl_bi_reset_detected(struct cxl_port *endpoint);
> struct cxl_dport *devm_cxl_add_dport_by_dev(struct cxl_port *port,
> struct device *dport_dev);
>
> diff --git a/drivers/cxl/pci.c b/drivers/cxl/pci.c
> index c7c91e8dc51d..cdab4804dba7 100644
> --- a/drivers/cxl/pci.c
> +++ b/drivers/cxl/pci.c
> @@ -987,8 +987,12 @@ static void cxl_reset_done(struct pci_dev *pdev)
> if (!cxlmd->dev.driver)
> return;
>
> - if (cxlmd->endpoint &&
> - cxl_endpoint_decoder_reset_detected(cxlmd->endpoint)) {
> + if (!cxlmd->endpoint)
> + return;
> +
> + cxl_bi_reset_detected(cxlmd->endpoint);
> +
> + if (cxl_endpoint_decoder_reset_detected(cxlmd->endpoint)) {
> device_for_each_child(&cxlmd->endpoint->dev, NULL,
> cxl_endpoint_decoder_clear_reset_flags);
>
> diff --git a/include/cxl/cxl.h b/include/cxl/cxl.h
> index 278b84b08c83..718eb4353887 100644
> --- a/include/cxl/cxl.h
> +++ b/include/cxl/cxl.h
> @@ -168,6 +168,7 @@ struct cxl_dpa_partition {
> * @regs: Parsed register blocks
> * @cxl_dvsec: Offset to the PCIe device DVSEC
> * @rcd: operating in RCD mode (CXL 3.0 9.11.8 CXL Devices Attached to an RCH)
> + * @bi: device is BI (Back-Invalidate) enabled
> * @media_ready: Indicate whether the device media is usable
> * @dpa_res: Overall DPA resource tree for the device
> * @part: DPA partition array
> @@ -187,6 +188,7 @@ struct cxl_dev_state {
> struct cxl_device_regs regs;
> int cxl_dvsec;
> bool rcd;
> + bool bi;
> bool media_ready;
> struct resource dpa_res;
> struct cxl_dpa_partition part[CXL_NR_PARTITIONS_MAX];
^ permalink raw reply [flat|nested] 33+ messages in thread* Re: [PATCH v9 02/10] cxl/pci: Add BI topology enable/disable
2026-09-22 23:38 ` [PATCH v9 02/10] cxl/pci: Add BI topology enable/disable Davidlohr Bueso
2026-09-23 1:12 ` sashiko-bot
2026-09-23 5:41 ` Li Ming
@ 2026-09-25 23:32 ` Jonathan Cameron
2026-09-29 8:52 ` Richard Cheng
2026-09-30 23:01 ` Alison Schofield
4 siblings, 0 replies; 33+ messages in thread
From: Jonathan Cameron @ 2026-09-25 23:32 UTC (permalink / raw)
To: Davidlohr Bueso
Cc: dave.jiang, alison.schofield, icheng, ming.li, benjamin.cheatham,
alucerop, linux-cxl
On Tue, 22 Sep 2026 16:38:39 -0700
Davidlohr Bueso <dave@stgolabs.net> wrote:
> Implement cxl_bi_setup() to enable BI flows on the device and every
> component in the path, and its teardown counterpart cxl_bi_dealloc().
> Setup runs from devm_cxl_endpoint_decoders_setup(), between the
> port's HDM state and its decoders. Registered there, its devres
> teardown brings BI down after the decoders quiesce and before the
> HDM state is freed, and BI is settled before the decoders, and later
> the regions, are looked at. The BI-ID and path enablement belong to
> the endpoint port's lifetime.
>
> Setup is safe in endpoint port probe context. The port probes
> synchronously from cxl_mem_probe(), pinning the memdev state the
> walk consumes, and the whole ancestor path already exists with BI
> registers mapped (dports at dport-add time, the switch USP RT at
> first-dport setup) because devm_cxl_enumerate_ports() completes
> before the endpoint is created.
>
> Dealloc is safe in endpoint devres context. Both setup and dealloc
> walk the endpoint's parent_dport topology rather than getting the
> port by bus lookup - an ancestor teardown delists the parent port
> before the endpoint's devres runs.
>
> The topology walk is stable as parent_dport pointers are fixed at
> port creation; ancestors cannot be reaped while holding this
> memdev's cxl_ep; and their own teardown frees dports only after
> the endpoint is gone.
>
> Likewise, the device state outlives the walk - cxlmd->cxlds is
> nulled only after cxl_memdev_unregister() has torn the endpoint
> down, and delete_endpoint() clears cxlmd->endpoint only after the
> endpoint devres has run.
>
> Each dport is programmed by its position - the one immediately above
> the device takes BI Enable, every dport above it takes BI Forward
> (Table 8-157, Table 9-13), at any switch depth (Table 7-97). Any
> level can be shared, so nr_bi refcounts endpoints at every dport.
> Registers are written on the first endpoint and cleared on the last,
> but only downstream ports commit (Table 8-156), once per endpoint
> (Table 8-152), and a failed commit undoes its write and commits the
> undo. A USP advertising a BI Route Table that failed to map is
> refused rather than treated as absent. nr_bi counts only the
> endpoints this driver enabled, so a level can be cleared while
> firmware still has an unbound device on it.
>
> A reset may wipe the device's BI Enable, whose reset default is 0
> (Table 8-157). .reset_done reads the hardware rather than assume
> which reset ran, and invalidates cxlds->bi, failing closed with
> recovery by rebind as for decoder loss; dealloc unwinds the dport
> refcounts regardless. It also clears cxlds->bi unconditionally - the
> endpoint disable fails when the hardware already shows BI Enable
> clear, from a reset .reset_done never saw, and the flag must not
> outlive the BI Decoder mapping it describes, which the endpoint port
> releases moments later.
>
> With dealloc in the endpoint's devres, delete_endpoint() already
> holds the parent port's device lock, so to avoid deadlocking, add a
> per-port bi_lock, serializing the dports that share state (nr_bi and
> the control register at any shared level, the switch USP's BI RT).
>
> Reviewed-by: Ben Cheatham <benjamin.cheatham@amd.com>
> Signed-off-by: Davidlohr Bueso <dave@stgolabs.net>
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: [PATCH v9 02/10] cxl/pci: Add BI topology enable/disable
2026-09-22 23:38 ` [PATCH v9 02/10] cxl/pci: Add BI topology enable/disable Davidlohr Bueso
` (2 preceding siblings ...)
2026-09-25 23:32 ` Jonathan Cameron
@ 2026-09-29 8:52 ` Richard Cheng
2026-09-30 23:01 ` Alison Schofield
4 siblings, 0 replies; 33+ messages in thread
From: Richard Cheng @ 2026-09-29 8:52 UTC (permalink / raw)
To: Davidlohr Bueso
Cc: dave.jiang, jic23, alison.schofield, ming.li, benjamin.cheatham,
alucerop, linux-cxl
On Tue, Sep 22, 2026 at 04:38:39PM +0800, Davidlohr Bueso wrote:
> Implement cxl_bi_setup() to enable BI flows on the device and every
> component in the path, and its teardown counterpart cxl_bi_dealloc().
> Setup runs from devm_cxl_endpoint_decoders_setup(), between the
> port's HDM state and its decoders. Registered there, its devres
> teardown brings BI down after the decoders quiesce and before the
> HDM state is freed, and BI is settled before the decoders, and later
> the regions, are looked at. The BI-ID and path enablement belong to
> the endpoint port's lifetime.
>
> Setup is safe in endpoint port probe context. The port probes
> synchronously from cxl_mem_probe(), pinning the memdev state the
> walk consumes, and the whole ancestor path already exists with BI
> registers mapped (dports at dport-add time, the switch USP RT at
> first-dport setup) because devm_cxl_enumerate_ports() completes
> before the endpoint is created.
>
> Dealloc is safe in endpoint devres context. Both setup and dealloc
> walk the endpoint's parent_dport topology rather than getting the
> port by bus lookup - an ancestor teardown delists the parent port
> before the endpoint's devres runs.
>
> The topology walk is stable as parent_dport pointers are fixed at
> port creation; ancestors cannot be reaped while holding this
> memdev's cxl_ep; and their own teardown frees dports only after
> the endpoint is gone.
>
> Likewise, the device state outlives the walk - cxlmd->cxlds is
> nulled only after cxl_memdev_unregister() has torn the endpoint
> down, and delete_endpoint() clears cxlmd->endpoint only after the
> endpoint devres has run.
>
> Each dport is programmed by its position - the one immediately above
> the device takes BI Enable, every dport above it takes BI Forward
> (Table 8-157, Table 9-13), at any switch depth (Table 7-97). Any
> level can be shared, so nr_bi refcounts endpoints at every dport.
> Registers are written on the first endpoint and cleared on the last,
> but only downstream ports commit (Table 8-156), once per endpoint
> (Table 8-152), and a failed commit undoes its write and commits the
> undo. A USP advertising a BI Route Table that failed to map is
> refused rather than treated as absent. nr_bi counts only the
> endpoints this driver enabled, so a level can be cleared while
> firmware still has an unbound device on it.
>
> A reset may wipe the device's BI Enable, whose reset default is 0
> (Table 8-157). .reset_done reads the hardware rather than assume
> which reset ran, and invalidates cxlds->bi, failing closed with
> recovery by rebind as for decoder loss; dealloc unwinds the dport
> refcounts regardless. It also clears cxlds->bi unconditionally - the
> endpoint disable fails when the hardware already shows BI Enable
> clear, from a reset .reset_done never saw, and the flag must not
> outlive the BI Decoder mapping it describes, which the endpoint port
> releases moments later.
>
> With dealloc in the endpoint's devres, delete_endpoint() already
> holds the parent port's device lock, so to avoid deadlocking, add a
> per-port bi_lock, serializing the dports that share state (nr_bi and
> the control register at any shared level, the switch USP's BI RT).
>
> Reviewed-by: Ben Cheatham <benjamin.cheatham@amd.com>
> Signed-off-by: Davidlohr Bueso <dave@stgolabs.net>
> ---
> drivers/cxl/core/core.h | 7 +
> drivers/cxl/core/hdm.c | 10 +
> drivers/cxl/core/pci.c | 427 ++++++++++++++++++++++++++++++++++++++++
> drivers/cxl/core/port.c | 1 +
> drivers/cxl/cxl.h | 31 +++
> drivers/cxl/pci.c | 8 +-
> include/cxl/cxl.h | 2 +
> 7 files changed, 484 insertions(+), 2 deletions(-)
>
> diff --git a/drivers/cxl/core/core.h b/drivers/cxl/core/core.h
> index 384bad8cb70f..ef41632d7b7d 100644
> --- a/drivers/cxl/core/core.h
> +++ b/drivers/cxl/core/core.h
> @@ -210,6 +210,13 @@ static inline void devm_cxl_dport_ras_setup(struct cxl_dport *dport) { }
> int cxl_gpf_port_setup(struct cxl_dport *dport);
> void devm_cxl_dport_bi_setup(struct cxl_dport *dport);
>
> +static inline bool cxl_bi_decoder_enabled(struct cxl_port *port)
> +{
> + return FIELD_GET(CXL_BI_DECODER_CTRL_BI_ENABLE,
> + readl(port->regs.bi_decoder +
> + CXL_BI_DECODER_CTRL_OFFSET));
> +}
> +
> struct cxl_hdm;
> int cxl_hdm_decode_init(struct cxl_dev_state *cxlds, struct cxl_hdm *cxlhdm,
> struct cxl_endpoint_dvsec_info *info);
> diff --git a/drivers/cxl/core/hdm.c b/drivers/cxl/core/hdm.c
> index 0c80b76a5f9b..0e9d652b568e 100644
> --- a/drivers/cxl/core/hdm.c
> +++ b/drivers/cxl/core/hdm.c
> @@ -1274,6 +1274,16 @@ int devm_cxl_endpoint_decoders_setup(struct cxl_port *port)
> if (rc)
> return rc;
>
> + /*
> + * Between the port's HDM state and its decoders: devres,
> + * unwinding in reverse, brings BI down only after the decoders
> + * quiesce, while its slow walk still precedes the HDM state
> + * free.
> + */
> + rc = cxl_bi_setup(port);
> + if (rc)
> + dev_dbg(&port->dev, "BI setup failed rc=%d\n", rc);
> +
> return devm_cxl_enumerate_decoders(cxlhdm, &info);
> }
> EXPORT_SYMBOL_NS_GPL(devm_cxl_endpoint_decoders_setup, "CXL");
> diff --git a/drivers/cxl/core/pci.c b/drivers/cxl/core/pci.c
> index b8676a3d6ec9..f3ca9861a7ba 100644
> --- a/drivers/cxl/core/pci.c
> +++ b/drivers/cxl/core/pci.c
> @@ -2,12 +2,14 @@
> /* Copyright(c) 2021 Intel Corporation. All rights reserved. */
> #include <linux/units.h>
> #include <linux/io-64-nonatomic-lo-hi.h>
> +#include <linux/iopoll.h>
> #include <linux/device.h>
> #include <linux/delay.h>
> #include <linux/pci.h>
> #include <linux/pci-doe.h>
> #include <cxl/pci.h>
> #include <linux/aer.h>
> +#include <linux/string_choices.h>
> #include <cxlpci.h>
> #include <cxlmem.h>
> #include <cxl.h>
> @@ -965,3 +967,428 @@ void devm_cxl_dport_bi_setup(struct cxl_dport *dport)
> break;
> }
> }
> +
> +/*
> + * BI requires 256B Flit operation on the link. RP/DSP/endpoint must
> + * also have the BI Decoder cap mapped (@bi); for USPs the BI RT cap
> + * is optional per CXL 4.0 8.2.4.26, so absent @bi is allowed.
> + */
> +static bool cxl_is_bi_capable(struct pci_dev *pdev, void __iomem *bi)
> +{
> + if (!cxl_pci_flit_256(pdev))
> + return false;
> +
> + if (pci_pcie_type(pdev) != PCI_EXP_TYPE_UPSTREAM && !bi) {
> + dev_dbg(&pdev->dev, "No BI Decoder registers.\n");
> + return false;
> + }
> +
> + return true;
> +}
> +
> +/* limit any insane timeouts from hw */
> +#define CXL_BI_COMMIT_MAXTMO_US (20 * USEC_PER_SEC)
> +
> +static unsigned long __cxl_bi_get_timeout_us(struct device *dev,
> + unsigned int scale,
> + unsigned int base)
> +{
> + static const unsigned long scale_tbl[] = {
> + 1, 10, 100, 1000, 10000, 100000, 1000000, 10000000,
> + };
> +
> + if (scale >= ARRAY_SIZE(scale_tbl) || !base) {
> + dev_dbg(dev, "Invalid BI commit timeout: scale=%u base=%u\n",
> + scale, base);
> + return CXL_BI_COMMIT_MAXTMO_US;
> + }
> +
> + return scale_tbl[scale] * base;
> +}
> +
> +static int __cxl_bi_wait_commit(struct device *dev, void __iomem *status_reg,
> + u32 committed_bit, u32 err_bit,
> + unsigned int scale, unsigned int base)
> +{
> + unsigned long tmo_us, poll_us;
> + ktime_t start;
> + u32 status;
> + int rc;
> +
> + tmo_us = min_t(unsigned long, CXL_BI_COMMIT_MAXTMO_US,
> + __cxl_bi_get_timeout_us(dev, scale, base));
> + poll_us = max_t(unsigned long, tmo_us / 10, 1); /* ~10% */
> + start = ktime_get();
> +
> + rc = readx_poll_timeout(readl, status_reg, status,
> + status & (committed_bit | err_bit),
> + poll_us, tmo_us);
> + if (rc) {
> + dev_err(dev, "BI-ID commit timed out (%luus)\n", tmo_us);
> + return rc; /* -ETIMEDOUT */
> + }
> +
> + if (status & err_bit) {
> + dev_err(dev, "BI-ID commit rejected by hardware\n");
> + return -EIO;
> + }
> +
> + dev_dbg(dev, "BI-ID commit wait took %lluus\n",
> + ktime_to_us(ktime_sub(ktime_get(), start)));
> + return 0;
> +}
> +
> +/* BI RT only exists on switch upstream ports. */
> +static int __cxl_bi_commit_rt(struct device *dev, void __iomem *bi)
> +{
> + u32 status, ctrl;
> + unsigned int scale, base;
> +
> + if (!FIELD_GET(CXL_BI_RT_CAPS_EXPLICIT_COMMIT_REQ,
> + readl(bi + CXL_BI_RT_CAPS_OFFSET)))
> + return 0;
> +
> + ctrl = readl(bi + CXL_BI_RT_CTRL_OFFSET);
> + writel(ctrl & ~CXL_BI_RT_CTRL_BI_COMMIT, bi + CXL_BI_RT_CTRL_OFFSET);
> + writel(ctrl | CXL_BI_RT_CTRL_BI_COMMIT, bi + CXL_BI_RT_CTRL_OFFSET);
> +
> + status = readl(bi + CXL_BI_RT_STATUS_OFFSET);
> + scale = FIELD_GET(CXL_BI_RT_STATUS_BI_COMMIT_TM_SCALE, status);
> + base = FIELD_GET(CXL_BI_RT_STATUS_BI_COMMIT_TM_BASE, status);
> +
> + return __cxl_bi_wait_commit(dev, bi + CXL_BI_RT_STATUS_OFFSET,
> + CXL_BI_RT_STATUS_BI_COMMITTED,
> + CXL_BI_RT_STATUS_BI_ERR_NOT_COMMITTED,
> + scale, base);
> +}
> +
> +static int __cxl_bi_commit_decoder(struct device *dev, void __iomem *bi)
> +{
> + u32 status, ctrl;
> + unsigned int scale, base;
> +
> + if (!FIELD_GET(CXL_BI_DECODER_CAPS_EXPLICIT_COMMIT_REQ,
> + readl(bi + CXL_BI_DECODER_CAPS_OFFSET)))
> + return 0;
> +
> + ctrl = readl(bi + CXL_BI_DECODER_CTRL_OFFSET);
> + writel(ctrl & ~CXL_BI_DECODER_CTRL_BI_COMMIT,
> + bi + CXL_BI_DECODER_CTRL_OFFSET);
> + writel(ctrl | CXL_BI_DECODER_CTRL_BI_COMMIT,
> + bi + CXL_BI_DECODER_CTRL_OFFSET);
> +
> + status = readl(bi + CXL_BI_DECODER_STATUS_OFFSET);
> + scale = FIELD_GET(CXL_BI_DECODER_STATUS_BI_COMMIT_TM_SCALE, status);
> + base = FIELD_GET(CXL_BI_DECODER_STATUS_BI_COMMIT_TM_BASE, status);
> +
> + return __cxl_bi_wait_commit(dev, bi + CXL_BI_DECODER_STATUS_OFFSET,
> + CXL_BI_DECODER_STATUS_BI_COMMITTED,
> + CXL_BI_DECODER_STATUS_BI_ERR_NOT_COMMITTED,
> + scale, base);
> +}
> +
> +static int cxl_bi_commit_dport(struct cxl_dport *dport)
> +{
> + struct cxl_port *port = dport->port;
> + int rc;
> +
> + lockdep_assert_held(&port->bi_lock);
> +
> + /* root ports never require the explicit commit */
> + if (pci_pcie_type(to_pci_dev(dport->dport_dev)) !=
> + PCI_EXP_TYPE_DOWNSTREAM)
> + return 0;
> +
> + rc = __cxl_bi_commit_decoder(dport->dport_dev, dport->regs.bi_decoder);
> + if (rc)
> + return rc;
> +
> + if (port->regs.bi_rt)
> + rc = __cxl_bi_commit_rt(&port->dev, port->regs.bi_rt);
> +
> + return rc;
> +}
> +
> +/*
> + * Enable BI-ID changes in the given level of the topology.
> + * @direct says the device is connected directly to this dport, which
> + * takes BI Enable; every dport above it takes BI Forward.
> + */
> +static int cxl_bi_ctrl_dport_enable(struct cxl_dport *dport, bool direct)
> +{
> + struct cxl_port *port = dport->port;
> + u32 ctrl, value, set, clr;
> + void __iomem *bi;
> + int rc;
> +
> + guard(mutex)(&port->bi_lock);
> +
> + bi = dport->regs.bi_decoder;
> + if (!bi)
> + return -EINVAL;
> +
> + ctrl = readl(bi + CXL_BI_DECODER_CTRL_OFFSET);
> +
> + set = direct ? CXL_BI_DECODER_CTRL_BI_ENABLE :
> + CXL_BI_DECODER_CTRL_BI_FW;
> + clr = direct ? CXL_BI_DECODER_CTRL_BI_FW :
> + CXL_BI_DECODER_CTRL_BI_ENABLE;
> +
> + value = (ctrl | set) & ~clr;
> + if (value != ctrl)
> + writel(value, bi + CXL_BI_DECODER_CTRL_OFFSET);
> +
> + /* owed per new device below, not per register change */
> + rc = cxl_bi_commit_dport(dport);
> + if (rc) {
> + if (value != ctrl) {
> + /* the undo is a BI-ID change owing its own commit */
> + writel(ctrl, bi + CXL_BI_DECODER_CTRL_OFFSET);
> + cxl_bi_commit_dport(dport);
> + }
> + return rc;
> + }
> + dport->nr_bi++;
> +
> + return 0;
> +}
> +
> +/*
> + * Dealloc BI-ID changes in the given level of the topology. Called
> + * once per endpoint that enabled this level: on teardown, or to
> + * unwind a path that failed partway up.
> + */
> +static int cxl_bi_ctrl_dport_disable(struct cxl_dport *dport)
> +{
> + struct cxl_port *port = dport->port;
> + void __iomem *bi;
> + u32 ctrl;
> +
> + guard(mutex)(&port->bi_lock);
> +
> + bi = dport->regs.bi_decoder;
> + if (!bi)
> + return -EINVAL;
> +
> + if (WARN_ON_ONCE(dport->nr_bi == 0))
> + return -EINVAL;
> +
> + /* others below still need this level */
> + if (--dport->nr_bi > 0)
> + return 0;
> +
> + ctrl = readl(bi + CXL_BI_DECODER_CTRL_OFFSET);
> + writel(ctrl & ~(CXL_BI_DECODER_CTRL_BI_FW |
> + CXL_BI_DECODER_CTRL_BI_ENABLE),
> + bi + CXL_BI_DECODER_CTRL_OFFSET);
> +
> + return cxl_bi_commit_dport(dport);
> +}
> +
> +static int __cxl_bi_ctrl_endpoint(struct cxl_dev_state *cxlds, bool enable)
> +{
> + struct cxl_port *endpoint = cxlds->cxlmd->endpoint;
> + void __iomem *bi = endpoint->regs.bi_decoder;
> + u32 ctrl;
> +
> + if (!bi)
> + return -EINVAL;
> +
> + ctrl = readl(bi + CXL_BI_DECODER_CTRL_OFFSET);
> +
> + if (enable) {
> + if (FIELD_GET(CXL_BI_DECODER_CTRL_BI_ENABLE, ctrl)) {
> + if (cxlds->bi)
> + return 0;
> + dev_err(cxlds->dev,
> + "BI already enabled in hardware\n");
> + return -EBUSY;
> + }
> + } else {
> + if (!FIELD_GET(CXL_BI_DECODER_CTRL_BI_ENABLE, ctrl)) {
> + if (!cxlds->bi)
> + return 0;
> + dev_err(cxlds->dev,
> + "BI already disabled in hardware\n");
> + return -EBUSY;
> + }
> + }
> +
> + FIELD_MODIFY(CXL_BI_DECODER_CTRL_BI_ENABLE, &ctrl, enable);
> + writel(ctrl, bi + CXL_BI_DECODER_CTRL_OFFSET);
> + cxlds->bi = enable;
> +
> + dev_dbg(cxlds->dev, "BI requests %s\n", str_enabled_disabled(enable));
> +
> + return 0;
> +}
> +
> +static int cxl_bi_ctrl_endpoint_enable(struct cxl_dev_state *cxlds)
> +{
> + return __cxl_bi_ctrl_endpoint(cxlds, true);
> +}
> +
> +static int cxl_bi_ctrl_endpoint_disable(struct cxl_dev_state *cxlds)
> +{
> + return __cxl_bi_ctrl_endpoint(cxlds, false);
> +}
> +
> +/*
> + * devm teardown on endpoint port destruction. Registered before the
> + * decoders, so devres runs it after them: regions are detached and
> + * decoders unregistered by the time BI comes down.
> + */
> +static void cxl_bi_dealloc(void *data)
> +{
> + struct cxl_port *endpoint = data;
> + struct cxl_memdev *cxlmd = to_cxl_memdev(endpoint->uport_dev);
> + struct cxl_dev_state *cxlds = cxlmd->cxlds;
> + struct cxl_dport *dport_iter;
> + struct cxl_port *port_iter;
> +
> + cxl_bi_ctrl_endpoint_disable(cxlds);
> + cxlds->bi = false;
> +
Is it safe to disable BI unconditionally ?
For a FW-configured HDM-DB region with locked decoders, resetting decoders can
leave the HW mapping active. If that memory is boot System RAM, it can remain
in use after driver unbind, while the cleanup here disable BI.
I think one of the scenario is when decoders are boot-programmed and locked.
FW could configure such a range with BI enabled and expose it
Should we preserve BI when those locked mapping remain active ?
Best regards,
Richard Cheng.
> + /*
> + * Walk the same parent_dport chain that enabled the path. A bus
> + * lookup cannot stand in for it: an ancestor-driven teardown
> + * delists the parent port before this devres action runs.
> + */
> + dport_iter = endpoint->parent_dport;
> + port_iter = dport_iter->port;
> + while (!is_cxl_root(port_iter)) {
> + int rc = cxl_bi_ctrl_dport_disable(dport_iter);
> +
> + /* best effort */
> + if (rc)
> + dev_dbg(&port_iter->dev,
> + "BI dport disable failed: %d\n", rc);
> +
> + dport_iter = port_iter->parent_dport;
> + port_iter = dport_iter->port;
> + }
> +}
> +
> +/*
> + * Enable BI on every dport in the path, then on the device itself.
> + * On failure, unwind only the dports that enabled.
> + */
> +static int cxl_bi_enable_path(struct cxl_dev_state *cxlds,
> + struct cxl_port *port, struct cxl_dport *dport)
> +{
> + struct cxl_dport *dport_iter, *failed;
> + struct cxl_port *port_iter;
> + int rc;
> +
> + port_iter = port;
> + dport_iter = dport;
> + while (!is_cxl_root(port_iter)) {
> + rc = cxl_bi_ctrl_dport_enable(dport_iter, dport_iter == dport);
> + if (rc)
> + goto err_rollback;
> +
> + dport_iter = port_iter->parent_dport;
> + port_iter = dport_iter->port;
> + }
> +
> + rc = cxl_bi_ctrl_endpoint_enable(cxlds);
> + if (rc)
> + goto err_rollback;
> +
> + return 0;
> +
> +err_rollback:
> + failed = dport_iter;
> + dport_iter = dport;
> + port_iter = port;
> + while (!is_cxl_root(port_iter) && dport_iter != failed) {
> + cxl_bi_ctrl_dport_disable(dport_iter);
> + dport_iter = port_iter->parent_dport;
> + port_iter = dport_iter->port;
> + }
> + return rc;
> +}
> +
> +/*
> + * An SBR wipes the device's BI Enable; an FLR leaves it alone.
> + * The check is against the hardware, not decoder state: BI is
> + * enabled at probe, so it can be wiped with no decoder ever
> + * committed. A wipe invalidates the software state; BI is never
> + * re-enabled here.
> + */
> +void cxl_bi_reset_detected(struct cxl_port *endpoint)
> +{
> + struct cxl_memdev *cxlmd = to_cxl_memdev(endpoint->uport_dev);
> + struct cxl_dev_state *cxlds = cxlmd->cxlds;
> +
> + if (!cxlds->bi)
> + return;
> +
> + if (cxl_bi_decoder_enabled(endpoint))
> + return;
> +
> + dev_warn(cxlds->dev, "BI disabled by reset\n");
> + cxlds->bi = false;
> +}
> +EXPORT_SYMBOL_NS_GPL(cxl_bi_reset_detected, "CXL");
> +
> +int cxl_bi_setup(struct cxl_port *endpoint)
> +{
> + struct cxl_memdev *cxlmd = to_cxl_memdev(endpoint->uport_dev);
> + struct cxl_dev_state *cxlds = cxlmd->cxlds;
> + struct cxl_dport *dport = endpoint->parent_dport;
> + struct cxl_dport *dport_iter;
> + struct cxl_port *port_iter;
> + int rc;
> +
> + if (!dev_is_pci(cxlds->dev))
> + return 0;
> +
> + /* BI is VH-only */
> + if (cxlds->rcd)
> + return 0;
> +
> + if (!cxl_is_bi_capable(to_pci_dev(cxlds->dev),
> + endpoint->regs.bi_decoder))
> + return 0;
> +
> + port_iter = dport->port;
> + dport_iter = dport;
> + while (!is_cxl_root(port_iter)) {
> + /* check rp, dsp */
> + if (!cxl_is_bi_capable(to_pci_dev(dport_iter->dport_dev),
> + dport_iter->regs.bi_decoder)) {
> + dev_dbg(cxlds->dev, "BI not supported by topology\n");
> + return 0;
> + }
> +
> + /* check usp */
> + if (dev_is_pci(port_iter->uport_dev) &&
> + pci_pcie_type(to_pci_dev(port_iter->uport_dev)) ==
> + PCI_EXP_TYPE_UPSTREAM) {
> + if (!cxl_is_bi_capable(to_pci_dev(port_iter->uport_dev),
> + NULL)) {
> + dev_dbg(cxlds->dev,
> + "BI not supported by USP\n");
> + return 0;
> + }
> + if (port_iter->reg_map.component_map.bi_rt.valid &&
> + !port_iter->regs.bi_rt) {
> + dev_dbg(cxlds->dev,
> + "BI RT advertised but unmapped\n");
> + return 0;
> + }
> + }
> +
> + dport_iter = port_iter->parent_dport;
> + port_iter = dport_iter->port;
> + }
> +
> + rc = cxl_bi_enable_path(cxlds, dport->port, dport);
> + if (rc)
> + return rc;
> +
> + return devm_add_action_or_reset(&endpoint->dev, cxl_bi_dealloc,
> + endpoint);
> +}
> +EXPORT_SYMBOL_NS_GPL(cxl_bi_setup, "CXL");
> diff --git a/drivers/cxl/core/port.c b/drivers/cxl/core/port.c
> index 131ed62e8db3..b81fd680d18a 100644
> --- a/drivers/cxl/core/port.c
> +++ b/drivers/cxl/core/port.c
> @@ -742,6 +742,7 @@ static struct cxl_port *cxl_port_alloc(struct device *uport_dev,
> xa_init(&port->dports);
> xa_init(&port->endpoints);
> xa_init(&port->regions);
> + mutex_init(&port->bi_lock);
> port->component_reg_phys = CXL_RESOURCE_NONE;
>
> device_initialize(dev);
> diff --git a/drivers/cxl/cxl.h b/drivers/cxl/cxl.h
> index 65be3b91259a..ad7c991f0f9b 100644
> --- a/drivers/cxl/cxl.h
> +++ b/drivers/cxl/cxl.h
> @@ -180,6 +180,31 @@ static inline int ways_to_eiw(unsigned int ways, u8 *eiw)
> #define CXL_HEADERLOG_TRACE_SIZE SZ_512
> #define CXL_HEADERLOG_TRACE_SIZE_U32 (CXL_HEADERLOG_TRACE_SIZE / sizeof(u32))
>
> +/* CXL 4.0 8.2.4.26 CXL BI Route Table Capability Structure */
> +#define CXL_BI_RT_CAPS_OFFSET 0x0
> +#define CXL_BI_RT_CAPS_EXPLICIT_COMMIT_REQ BIT(0)
> +#define CXL_BI_RT_CTRL_OFFSET 0x4
> +#define CXL_BI_RT_CTRL_BI_COMMIT BIT(0)
> +#define CXL_BI_RT_STATUS_OFFSET 0x8
> +#define CXL_BI_RT_STATUS_BI_COMMITTED BIT(0)
> +#define CXL_BI_RT_STATUS_BI_ERR_NOT_COMMITTED BIT(1)
> +#define CXL_BI_RT_STATUS_BI_COMMIT_TM_SCALE GENMASK(11, 8)
> +#define CXL_BI_RT_STATUS_BI_COMMIT_TM_BASE GENMASK(15, 12)
> +
> +/* CXL 4.0 8.2.4.27 CXL BI Decoder Capability Structure */
> +#define CXL_BI_DECODER_CAPS_OFFSET 0x0
> +#define CXL_BI_DECODER_CAPS_HDMD_CAP BIT(0)
> +#define CXL_BI_DECODER_CAPS_EXPLICIT_COMMIT_REQ BIT(1)
> +#define CXL_BI_DECODER_CTRL_OFFSET 0x4
> +#define CXL_BI_DECODER_CTRL_BI_FW BIT(0)
> +#define CXL_BI_DECODER_CTRL_BI_ENABLE BIT(1)
> +#define CXL_BI_DECODER_CTRL_BI_COMMIT BIT(2)
> +#define CXL_BI_DECODER_STATUS_OFFSET 0x8
> +#define CXL_BI_DECODER_STATUS_BI_COMMITTED BIT(0)
> +#define CXL_BI_DECODER_STATUS_BI_ERR_NOT_COMMITTED BIT(1)
> +#define CXL_BI_DECODER_STATUS_BI_COMMIT_TM_SCALE GENMASK(11, 8)
> +#define CXL_BI_DECODER_STATUS_BI_COMMIT_TM_BASE GENMASK(15, 12)
> +
> /* CXL 2.0 8.2.8.1 Device Capabilities Array Register */
> #define CXLDEV_CAP_ARRAY_OFFSET 0x0
> #define CXLDEV_CAP_ARRAY_CAP_ID 0
> @@ -565,6 +590,7 @@ struct cxl_dax_region {
> * @decoder_ida: allocator for decoder ids
> * @reg_map: component and ras register mapping parameters
> * @regs: mapped component registers
> + * @bi_lock: serializes BI Decoder/RT state of this port's dports
> * @nr_dports: number of entries in @dports
> * @hdm_end: track last allocated HDM decoder instance for allocation ordering
> * @commit_end: cursor to track highest committed decoder for commit ordering
> @@ -587,6 +613,7 @@ struct cxl_port {
> struct ida decoder_ida;
> struct cxl_register_map reg_map;
> struct cxl_component_regs regs;
> + struct mutex bi_lock; /* dport BI state shared below this port */
> int nr_dports;
> int hdm_end;
> int commit_end;
> @@ -650,6 +677,7 @@ struct cxl_rcrb_info {
> * @coord: access coordinates (bandwidth and latency performance attributes)
> * @link_latency: calculated PCIe downstream latency
> * @gpf_dvsec: Cached GPF port DVSEC
> + * @nr_bi: number of BI-enabled endpoints below this dport
> */
> struct cxl_dport {
> struct device *dport_dev;
> @@ -662,6 +690,7 @@ struct cxl_dport {
> struct access_coordinate coord[ACCESS_COORDINATE_MAX];
> long link_latency;
> int gpf_dvsec;
> + int nr_bi;
> };
>
> /**
> @@ -920,6 +949,8 @@ void cxl_coordinates_combine(struct access_coordinate *out,
> struct access_coordinate *c2);
>
> bool cxl_endpoint_decoder_reset_detected(struct cxl_port *port);
> +int cxl_bi_setup(struct cxl_port *endpoint);
> +void cxl_bi_reset_detected(struct cxl_port *endpoint);
> struct cxl_dport *devm_cxl_add_dport_by_dev(struct cxl_port *port,
> struct device *dport_dev);
>
> diff --git a/drivers/cxl/pci.c b/drivers/cxl/pci.c
> index c7c91e8dc51d..cdab4804dba7 100644
> --- a/drivers/cxl/pci.c
> +++ b/drivers/cxl/pci.c
> @@ -987,8 +987,12 @@ static void cxl_reset_done(struct pci_dev *pdev)
> if (!cxlmd->dev.driver)
> return;
>
> - if (cxlmd->endpoint &&
> - cxl_endpoint_decoder_reset_detected(cxlmd->endpoint)) {
> + if (!cxlmd->endpoint)
> + return;
> +
> + cxl_bi_reset_detected(cxlmd->endpoint);
> +
> + if (cxl_endpoint_decoder_reset_detected(cxlmd->endpoint)) {
> device_for_each_child(&cxlmd->endpoint->dev, NULL,
> cxl_endpoint_decoder_clear_reset_flags);
>
> diff --git a/include/cxl/cxl.h b/include/cxl/cxl.h
> index 278b84b08c83..718eb4353887 100644
> --- a/include/cxl/cxl.h
> +++ b/include/cxl/cxl.h
> @@ -168,6 +168,7 @@ struct cxl_dpa_partition {
> * @regs: Parsed register blocks
> * @cxl_dvsec: Offset to the PCIe device DVSEC
> * @rcd: operating in RCD mode (CXL 3.0 9.11.8 CXL Devices Attached to an RCH)
> + * @bi: device is BI (Back-Invalidate) enabled
> * @media_ready: Indicate whether the device media is usable
> * @dpa_res: Overall DPA resource tree for the device
> * @part: DPA partition array
> @@ -187,6 +188,7 @@ struct cxl_dev_state {
> struct cxl_device_regs regs;
> int cxl_dvsec;
> bool rcd;
> + bool bi;
> bool media_ready;
> struct resource dpa_res;
> struct cxl_dpa_partition part[CXL_NR_PARTITIONS_MAX];
> --
> 2.39.5
>
^ permalink raw reply [flat|nested] 33+ messages in thread* Re: [PATCH v9 02/10] cxl/pci: Add BI topology enable/disable
2026-09-22 23:38 ` [PATCH v9 02/10] cxl/pci: Add BI topology enable/disable Davidlohr Bueso
` (3 preceding siblings ...)
2026-09-29 8:52 ` Richard Cheng
@ 2026-09-30 23:01 ` Alison Schofield
4 siblings, 0 replies; 33+ messages in thread
From: Alison Schofield @ 2026-09-30 23:01 UTC (permalink / raw)
To: Davidlohr Bueso
Cc: dave.jiang, jic23, icheng, ming.li, benjamin.cheatham, alucerop,
linux-cxl
On Tue, Sep 22, 2026 at 04:38:39PM -0700, Davidlohr Bueso wrote:
> Implement cxl_bi_setup() to enable BI flows on the device and every
> component in the path, and its teardown counterpart cxl_bi_dealloc().
> Setup runs from devm_cxl_endpoint_decoders_setup(), between the
> port's HDM state and its decoders. Registered there, its devres
> teardown brings BI down after the decoders quiesce and before the
> HDM state is freed, and BI is settled before the decoders, and later
> the regions, are looked at. The BI-ID and path enablement belong to
> the endpoint port's lifetime.
>
> Setup is safe in endpoint port probe context. The port probes
> synchronously from cxl_mem_probe(), pinning the memdev state the
> walk consumes, and the whole ancestor path already exists with BI
> registers mapped (dports at dport-add time, the switch USP RT at
> first-dport setup) because devm_cxl_enumerate_ports() completes
> before the endpoint is created.
>
> Dealloc is safe in endpoint devres context. Both setup and dealloc
> walk the endpoint's parent_dport topology rather than getting the
> port by bus lookup - an ancestor teardown delists the parent port
> before the endpoint's devres runs.
>
> The topology walk is stable as parent_dport pointers are fixed at
> port creation; ancestors cannot be reaped while holding this
> memdev's cxl_ep; and their own teardown frees dports only after
> the endpoint is gone.
>
> Likewise, the device state outlives the walk - cxlmd->cxlds is
> nulled only after cxl_memdev_unregister() has torn the endpoint
> down, and delete_endpoint() clears cxlmd->endpoint only after the
> endpoint devres has run.
>
> Each dport is programmed by its position - the one immediately above
> the device takes BI Enable, every dport above it takes BI Forward
> (Table 8-157, Table 9-13), at any switch depth (Table 7-97). Any
> level can be shared, so nr_bi refcounts endpoints at every dport.
> Registers are written on the first endpoint and cleared on the last,
> but only downstream ports commit (Table 8-156), once per endpoint
> (Table 8-152), and a failed commit undoes its write and commits the
> undo. A USP advertising a BI Route Table that failed to map is
> refused rather than treated as absent. nr_bi counts only the
> endpoints this driver enabled, so a level can be cleared while
> firmware still has an unbound device on it.
>
> A reset may wipe the device's BI Enable, whose reset default is 0
> (Table 8-157). .reset_done reads the hardware rather than assume
> which reset ran, and invalidates cxlds->bi, failing closed with
> recovery by rebind as for decoder loss; dealloc unwinds the dport
> refcounts regardless. It also clears cxlds->bi unconditionally - the
> endpoint disable fails when the hardware already shows BI Enable
> clear, from a reset .reset_done never saw, and the flag must not
> outlive the BI Decoder mapping it describes, which the endpoint port
> releases moments later.
>
> With dealloc in the endpoint's devres, delete_endpoint() already
> holds the parent port's device lock, so to avoid deadlocking, add a
> per-port bi_lock, serializing the dports that share state (nr_bi and
> the control register at any shared level, the switch USP's BI RT).
Reviewed-by: Alison Schofield <alison.schofield@intel.com>
^ permalink raw reply [flat|nested] 33+ messages in thread
* [PATCH v9 03/10] cxl/hdm: Add BI coherency support for endpoint decoders
2026-09-22 23:38 [PATCH v9 0/10] cxl: Support Back-Invalidate Davidlohr Bueso
2026-09-22 23:38 ` [PATCH v9 01/10] cxl: Add BI register probing and port initialization Davidlohr Bueso
2026-09-22 23:38 ` [PATCH v9 02/10] cxl/pci: Add BI topology enable/disable Davidlohr Bueso
@ 2026-09-22 23:38 ` Davidlohr Bueso
2026-09-23 6:06 ` Li Ming
2026-09-30 23:11 ` Alison Schofield
2026-09-22 23:38 ` [PATCH v9 04/10] cxl: Add HDM-DB region creation Davidlohr Bueso
` (7 subsequent siblings)
10 siblings, 2 replies; 33+ messages in thread
From: Davidlohr Bueso @ 2026-09-22 23:38 UTC (permalink / raw)
To: dave.jiang
Cc: jic23, alison.schofield, icheng, ming.li, benjamin.cheatham,
alucerop, dave, linux-cxl, Jonathan Cameron
Cache the HDM decoder's Supported Coherency Models on struct cxl_hdm.
A later patch has region attach consult it to verify the HDM supports
the region's coherency type.
For uncommitted endpoint decoders, init_hdm_decoder() defaults
target_type from supported_coherency: Type 3 devices default to
HDM-DB when the HDM is device-coherent-only, HDM-H otherwise.
Pre-committed decoders with the BI bit set are refused. cxl_bi_setup()
does not yet take over a path firmware already enabled, so nothing
verifies that such a decoder's topology can route BISnp.
Reviewed-by: Dave Jiang <dave.jiang@intel.com>
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
Signed-off-by: Davidlohr Bueso <dave@stgolabs.net>
---
drivers/cxl/core/core.h | 1 +
drivers/cxl/core/hdm.c | 35 ++++++++++++++++++++++++-----------
drivers/cxl/cxl.h | 8 +++++++-
drivers/cxl/cxlmem.h | 2 ++
4 files changed, 34 insertions(+), 12 deletions(-)
diff --git a/drivers/cxl/core/core.h b/drivers/cxl/core/core.h
index ef41632d7b7d..d0f352899903 100644
--- a/drivers/cxl/core/core.h
+++ b/drivers/cxl/core/core.h
@@ -221,6 +221,7 @@ struct cxl_hdm;
int cxl_hdm_decode_init(struct cxl_dev_state *cxlds, struct cxl_hdm *cxlhdm,
struct cxl_endpoint_dvsec_info *info);
int cxl_port_get_possible_dports(struct cxl_port *port);
+enum cxl_decoder_type cxled_default_type(struct cxl_endpoint_decoder *cxled);
#ifdef CONFIG_CXL_FEATURES
struct cxl_feat_entry *
diff --git a/drivers/cxl/core/hdm.c b/drivers/cxl/core/hdm.c
index 0e9d652b568e..18200a8f3f72 100644
--- a/drivers/cxl/core/hdm.c
+++ b/drivers/cxl/core/hdm.c
@@ -87,6 +87,8 @@ static void parse_hdm_decoder_caps(struct cxl_hdm *cxlhdm)
cxlhdm->iw_cap_mask |= BIT(3) | BIT(6) | BIT(12);
if (FIELD_GET(CXL_HDM_DECODER_INTERLEAVE_16_WAY, hdm_cap))
cxlhdm->iw_cap_mask |= BIT(16);
+ cxlhdm->supported_coherency =
+ FIELD_GET(CXL_HDM_DECODER_SUPPORTED_COHERENCY_MASK, hdm_cap);
}
static bool should_emulate_decoders(struct cxl_endpoint_dvsec_info *info)
@@ -968,6 +970,19 @@ static int cxl_setup_hdm_decoder_from_dvsec(
return 0;
}
+enum cxl_decoder_type cxled_default_type(struct cxl_endpoint_decoder *cxled)
+{
+ struct cxl_dev_state *cxlds = cxled_to_memdev(cxled)->cxlds;
+ struct cxl_port *port = cxled_to_port(cxled);
+ struct cxl_hdm *cxlhdm = dev_get_drvdata(&port->dev);
+
+ if (cxlds->type == CXL_DEVTYPE_CLASSMEM &&
+ cxlhdm->supported_coherency != CXL_HDM_DECODER_COHERENCY_DEV)
+ return CXL_DECODER_HOSTONLYMEM;
+
+ return CXL_DECODER_DEVMEM;
+}
+
static int init_hdm_decoder(struct cxl_port *port, struct cxl_decoder *cxld,
void __iomem *hdm, int which,
u64 *dpa_base, struct cxl_endpoint_dvsec_info *info)
@@ -1023,6 +1038,14 @@ static int init_hdm_decoder(struct cxl_port *port, struct cxl_decoder *cxld,
else
cxld->target_type = CXL_DECODER_DEVMEM;
+ /*
+ * Autocommit BI-enabled decoders is not supported. A path
+ * firmware enabled is not taken over, so nothing verifies
+ * it can route BISnp.
+ */
+ if (FIELD_GET(CXL_HDM_DECODER0_CTRL_BI, ctrl))
+ return -ENXIO;
+
guard(rwsem_write)(&cxl_rwsem.region);
if (cxld->id != cxl_num_decoders_committed(port)) {
dev_warn(&port->dev,
@@ -1040,17 +1063,7 @@ static int init_hdm_decoder(struct cxl_port *port, struct cxl_decoder *cxld,
port->commit_end = cxld->id;
} else {
if (cxled) {
- struct cxl_memdev *cxlmd = cxled_to_memdev(cxled);
- struct cxl_dev_state *cxlds = cxlmd->cxlds;
-
- /*
- * Default by devtype until a device arrives that needs
- * more precision.
- */
- if (cxlds->type == CXL_DEVTYPE_CLASSMEM)
- cxld->target_type = CXL_DECODER_HOSTONLYMEM;
- else
- cxld->target_type = CXL_DECODER_DEVMEM;
+ cxld->target_type = cxled_default_type(cxled);
} else {
/* To be overridden by region type at commit time */
cxld->target_type = CXL_DECODER_HOSTONLYMEM;
diff --git a/drivers/cxl/cxl.h b/drivers/cxl/cxl.h
index ad7c991f0f9b..a990ed41edef 100644
--- a/drivers/cxl/cxl.h
+++ b/drivers/cxl/cxl.h
@@ -50,7 +50,7 @@ extern const struct nvdimm_security_ops *cxl_security_ops;
#define CXL_BI_RT_CAPABILITY_LENGTH 0xC
#define CXL_BI_DECODER_CAPABILITY_LENGTH 0xC
-/* HDM decoders CXL 2.0 8.2.5.12 CXL HDM Decoder Capability Structure */
+/* HDM decoders CXL 4.0 8.2.4.20 CXL HDM Decoder Capability Structure */
#define CXL_HDM_DECODER_CAP_OFFSET 0x0
#define CXL_HDM_DECODER_COUNT_MASK GENMASK(3, 0)
#define CXL_HDM_DECODER_TARGET_COUNT_MASK GENMASK(7, 4)
@@ -58,6 +58,11 @@ extern const struct nvdimm_security_ops *cxl_security_ops;
#define CXL_HDM_DECODER_INTERLEAVE_14_12 BIT(9)
#define CXL_HDM_DECODER_INTERLEAVE_3_6_12_WAY BIT(11)
#define CXL_HDM_DECODER_INTERLEAVE_16_WAY BIT(12)
+#define CXL_HDM_DECODER_SUPPORTED_COHERENCY_MASK GENMASK(22, 21)
+#define CXL_HDM_DECODER_COHERENCY_UNKNOWN 0x0
+#define CXL_HDM_DECODER_COHERENCY_DEV 0x1
+#define CXL_HDM_DECODER_COHERENCY_HOST 0x2
+#define CXL_HDM_DECODER_COHERENCY_BOTH 0x3
#define CXL_HDM_DECODER_CTRL_OFFSET 0x4
#define CXL_HDM_DECODER_ENABLE BIT(1)
#define CXL_HDM_DECODER0_BASE_LOW_OFFSET(i) (0x20 * (i) + 0x10)
@@ -72,6 +77,7 @@ extern const struct nvdimm_security_ops *cxl_security_ops;
#define CXL_HDM_DECODER0_CTRL_COMMITTED BIT(10)
#define CXL_HDM_DECODER0_CTRL_COMMIT_ERROR BIT(11)
#define CXL_HDM_DECODER0_CTRL_HOSTONLY BIT(12)
+#define CXL_HDM_DECODER0_CTRL_BI BIT(13)
#define CXL_HDM_DECODER0_TL_LOW(i) (0x20 * (i) + 0x24)
#define CXL_HDM_DECODER0_TL_HIGH(i) (0x20 * (i) + 0x28)
#define CXL_HDM_DECODER0_SKIP_LOW(i) CXL_HDM_DECODER0_TL_LOW(i)
diff --git a/drivers/cxl/cxlmem.h b/drivers/cxl/cxlmem.h
index c401e3a1af06..33bac6bbdb59 100644
--- a/drivers/cxl/cxlmem.h
+++ b/drivers/cxl/cxlmem.h
@@ -857,6 +857,7 @@ int cxl_mem_sanitize(struct cxl_memdev *cxlmd, u16 cmd);
* @target_count: for switch decoders, max downstream port targets
* @interleave_mask: interleave granularity capability, see check_interleave_cap()
* @iw_cap_mask: bitmask of supported interleave ways, see check_interleave_cap()
+ * @supported_coherency: HDM Decoder Capability supported coherency models
* @port: mapped cxl_port, see devm_cxl_setup_hdm()
*/
struct cxl_hdm {
@@ -865,6 +866,7 @@ struct cxl_hdm {
unsigned int target_count;
unsigned int interleave_mask;
unsigned long iw_cap_mask;
+ unsigned int supported_coherency;
struct cxl_port *port;
};
--
2.39.5
^ permalink raw reply related [flat|nested] 33+ messages in thread* Re: [PATCH v9 03/10] cxl/hdm: Add BI coherency support for endpoint decoders
2026-09-22 23:38 ` [PATCH v9 03/10] cxl/hdm: Add BI coherency support for endpoint decoders Davidlohr Bueso
@ 2026-09-23 6:06 ` Li Ming
2026-09-30 23:11 ` Alison Schofield
1 sibling, 0 replies; 33+ messages in thread
From: Li Ming @ 2026-09-23 6:06 UTC (permalink / raw)
To: Davidlohr Bueso, dave.jiang
Cc: jic23, alison.schofield, icheng, benjamin.cheatham, alucerop,
linux-cxl, Jonathan Cameron
在 2026/9/23 07:38, Davidlohr Bueso 写道:
> Cache the HDM decoder's Supported Coherency Models on struct cxl_hdm.
> A later patch has region attach consult it to verify the HDM supports
> the region's coherency type.
>
> For uncommitted endpoint decoders, init_hdm_decoder() defaults
> target_type from supported_coherency: Type 3 devices default to
> HDM-DB when the HDM is device-coherent-only, HDM-H otherwise.
>
> Pre-committed decoders with the BI bit set are refused. cxl_bi_setup()
> does not yet take over a path firmware already enabled, so nothing
> verifies that such a decoder's topology can route BISnp.
>
> Reviewed-by: Dave Jiang <dave.jiang@intel.com>
> Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
> Signed-off-by: Davidlohr Bueso <dave@stgolabs.net>
Reviewed-by: Li Ming <ming.li@zohomail.com>
> ---
> drivers/cxl/core/core.h | 1 +
> drivers/cxl/core/hdm.c | 35 ++++++++++++++++++++++++-----------
> drivers/cxl/cxl.h | 8 +++++++-
> drivers/cxl/cxlmem.h | 2 ++
> 4 files changed, 34 insertions(+), 12 deletions(-)
>
> diff --git a/drivers/cxl/core/core.h b/drivers/cxl/core/core.h
> index ef41632d7b7d..d0f352899903 100644
> --- a/drivers/cxl/core/core.h
> +++ b/drivers/cxl/core/core.h
> @@ -221,6 +221,7 @@ struct cxl_hdm;
> int cxl_hdm_decode_init(struct cxl_dev_state *cxlds, struct cxl_hdm *cxlhdm,
> struct cxl_endpoint_dvsec_info *info);
> int cxl_port_get_possible_dports(struct cxl_port *port);
> +enum cxl_decoder_type cxled_default_type(struct cxl_endpoint_decoder *cxled);
>
> #ifdef CONFIG_CXL_FEATURES
> struct cxl_feat_entry *
> diff --git a/drivers/cxl/core/hdm.c b/drivers/cxl/core/hdm.c
> index 0e9d652b568e..18200a8f3f72 100644
> --- a/drivers/cxl/core/hdm.c
> +++ b/drivers/cxl/core/hdm.c
> @@ -87,6 +87,8 @@ static void parse_hdm_decoder_caps(struct cxl_hdm *cxlhdm)
> cxlhdm->iw_cap_mask |= BIT(3) | BIT(6) | BIT(12);
> if (FIELD_GET(CXL_HDM_DECODER_INTERLEAVE_16_WAY, hdm_cap))
> cxlhdm->iw_cap_mask |= BIT(16);
> + cxlhdm->supported_coherency =
> + FIELD_GET(CXL_HDM_DECODER_SUPPORTED_COHERENCY_MASK, hdm_cap);
> }
>
> static bool should_emulate_decoders(struct cxl_endpoint_dvsec_info *info)
> @@ -968,6 +970,19 @@ static int cxl_setup_hdm_decoder_from_dvsec(
> return 0;
> }
>
> +enum cxl_decoder_type cxled_default_type(struct cxl_endpoint_decoder *cxled)
> +{
> + struct cxl_dev_state *cxlds = cxled_to_memdev(cxled)->cxlds;
> + struct cxl_port *port = cxled_to_port(cxled);
> + struct cxl_hdm *cxlhdm = dev_get_drvdata(&port->dev);
> +
> + if (cxlds->type == CXL_DEVTYPE_CLASSMEM &&
> + cxlhdm->supported_coherency != CXL_HDM_DECODER_COHERENCY_DEV)
> + return CXL_DECODER_HOSTONLYMEM;
> +
> + return CXL_DECODER_DEVMEM;
> +}
> +
> static int init_hdm_decoder(struct cxl_port *port, struct cxl_decoder *cxld,
> void __iomem *hdm, int which,
> u64 *dpa_base, struct cxl_endpoint_dvsec_info *info)
> @@ -1023,6 +1038,14 @@ static int init_hdm_decoder(struct cxl_port *port, struct cxl_decoder *cxld,
> else
> cxld->target_type = CXL_DECODER_DEVMEM;
>
> + /*
> + * Autocommit BI-enabled decoders is not supported. A path
> + * firmware enabled is not taken over, so nothing verifies
> + * it can route BISnp.
> + */
> + if (FIELD_GET(CXL_HDM_DECODER0_CTRL_BI, ctrl))
> + return -ENXIO;
> +
> guard(rwsem_write)(&cxl_rwsem.region);
> if (cxld->id != cxl_num_decoders_committed(port)) {
> dev_warn(&port->dev,
> @@ -1040,17 +1063,7 @@ static int init_hdm_decoder(struct cxl_port *port, struct cxl_decoder *cxld,
> port->commit_end = cxld->id;
> } else {
> if (cxled) {
> - struct cxl_memdev *cxlmd = cxled_to_memdev(cxled);
> - struct cxl_dev_state *cxlds = cxlmd->cxlds;
> -
> - /*
> - * Default by devtype until a device arrives that needs
> - * more precision.
> - */
> - if (cxlds->type == CXL_DEVTYPE_CLASSMEM)
> - cxld->target_type = CXL_DECODER_HOSTONLYMEM;
> - else
> - cxld->target_type = CXL_DECODER_DEVMEM;
> + cxld->target_type = cxled_default_type(cxled);
> } else {
> /* To be overridden by region type at commit time */
> cxld->target_type = CXL_DECODER_HOSTONLYMEM;
> diff --git a/drivers/cxl/cxl.h b/drivers/cxl/cxl.h
> index ad7c991f0f9b..a990ed41edef 100644
> --- a/drivers/cxl/cxl.h
> +++ b/drivers/cxl/cxl.h
> @@ -50,7 +50,7 @@ extern const struct nvdimm_security_ops *cxl_security_ops;
> #define CXL_BI_RT_CAPABILITY_LENGTH 0xC
> #define CXL_BI_DECODER_CAPABILITY_LENGTH 0xC
>
> -/* HDM decoders CXL 2.0 8.2.5.12 CXL HDM Decoder Capability Structure */
> +/* HDM decoders CXL 4.0 8.2.4.20 CXL HDM Decoder Capability Structure */
> #define CXL_HDM_DECODER_CAP_OFFSET 0x0
> #define CXL_HDM_DECODER_COUNT_MASK GENMASK(3, 0)
> #define CXL_HDM_DECODER_TARGET_COUNT_MASK GENMASK(7, 4)
> @@ -58,6 +58,11 @@ extern const struct nvdimm_security_ops *cxl_security_ops;
> #define CXL_HDM_DECODER_INTERLEAVE_14_12 BIT(9)
> #define CXL_HDM_DECODER_INTERLEAVE_3_6_12_WAY BIT(11)
> #define CXL_HDM_DECODER_INTERLEAVE_16_WAY BIT(12)
> +#define CXL_HDM_DECODER_SUPPORTED_COHERENCY_MASK GENMASK(22, 21)
> +#define CXL_HDM_DECODER_COHERENCY_UNKNOWN 0x0
> +#define CXL_HDM_DECODER_COHERENCY_DEV 0x1
> +#define CXL_HDM_DECODER_COHERENCY_HOST 0x2
> +#define CXL_HDM_DECODER_COHERENCY_BOTH 0x3
> #define CXL_HDM_DECODER_CTRL_OFFSET 0x4
> #define CXL_HDM_DECODER_ENABLE BIT(1)
> #define CXL_HDM_DECODER0_BASE_LOW_OFFSET(i) (0x20 * (i) + 0x10)
> @@ -72,6 +77,7 @@ extern const struct nvdimm_security_ops *cxl_security_ops;
> #define CXL_HDM_DECODER0_CTRL_COMMITTED BIT(10)
> #define CXL_HDM_DECODER0_CTRL_COMMIT_ERROR BIT(11)
> #define CXL_HDM_DECODER0_CTRL_HOSTONLY BIT(12)
> +#define CXL_HDM_DECODER0_CTRL_BI BIT(13)
> #define CXL_HDM_DECODER0_TL_LOW(i) (0x20 * (i) + 0x24)
> #define CXL_HDM_DECODER0_TL_HIGH(i) (0x20 * (i) + 0x28)
> #define CXL_HDM_DECODER0_SKIP_LOW(i) CXL_HDM_DECODER0_TL_LOW(i)
> diff --git a/drivers/cxl/cxlmem.h b/drivers/cxl/cxlmem.h
> index c401e3a1af06..33bac6bbdb59 100644
> --- a/drivers/cxl/cxlmem.h
> +++ b/drivers/cxl/cxlmem.h
> @@ -857,6 +857,7 @@ int cxl_mem_sanitize(struct cxl_memdev *cxlmd, u16 cmd);
> * @target_count: for switch decoders, max downstream port targets
> * @interleave_mask: interleave granularity capability, see check_interleave_cap()
> * @iw_cap_mask: bitmask of supported interleave ways, see check_interleave_cap()
> + * @supported_coherency: HDM Decoder Capability supported coherency models
> * @port: mapped cxl_port, see devm_cxl_setup_hdm()
> */
> struct cxl_hdm {
> @@ -865,6 +866,7 @@ struct cxl_hdm {
> unsigned int target_count;
> unsigned int interleave_mask;
> unsigned long iw_cap_mask;
> + unsigned int supported_coherency;
> struct cxl_port *port;
> };
>
^ permalink raw reply [flat|nested] 33+ messages in thread* Re: [PATCH v9 03/10] cxl/hdm: Add BI coherency support for endpoint decoders
2026-09-22 23:38 ` [PATCH v9 03/10] cxl/hdm: Add BI coherency support for endpoint decoders Davidlohr Bueso
2026-09-23 6:06 ` Li Ming
@ 2026-09-30 23:11 ` Alison Schofield
1 sibling, 0 replies; 33+ messages in thread
From: Alison Schofield @ 2026-09-30 23:11 UTC (permalink / raw)
To: Davidlohr Bueso
Cc: dave.jiang, jic23, icheng, ming.li, benjamin.cheatham, alucerop,
linux-cxl, Jonathan Cameron
On Tue, Sep 22, 2026 at 04:38:40PM -0700, Davidlohr Bueso wrote:
> Cache the HDM decoder's Supported Coherency Models on struct cxl_hdm.
> A later patch has region attach consult it to verify the HDM supports
> the region's coherency type.
>
> For uncommitted endpoint decoders, init_hdm_decoder() defaults
> target_type from supported_coherency: Type 3 devices default to
> HDM-DB when the HDM is device-coherent-only, HDM-H otherwise.
>
> Pre-committed decoders with the BI bit set are refused. cxl_bi_setup()
> does not yet take over a path firmware already enabled, so nothing
> verifies that such a decoder's topology can route BISnp.
Reviewed-by: Alison Schofield <alison.schofield@intel.com>
^ permalink raw reply [flat|nested] 33+ messages in thread
* [PATCH v9 04/10] cxl: Add HDM-DB region creation
2026-09-22 23:38 [PATCH v9 0/10] cxl: Support Back-Invalidate Davidlohr Bueso
` (2 preceding siblings ...)
2026-09-22 23:38 ` [PATCH v9 03/10] cxl/hdm: Add BI coherency support for endpoint decoders Davidlohr Bueso
@ 2026-09-22 23:38 ` Davidlohr Bueso
2026-10-01 3:28 ` Alison Schofield
2026-09-22 23:38 ` [PATCH v9 05/10] cxl/hdm: Rename decoder coherency flags Davidlohr Bueso
` (6 subsequent siblings)
10 siblings, 1 reply; 33+ messages in thread
From: Davidlohr Bueso @ 2026-09-22 23:38 UTC (permalink / raw)
To: dave.jiang
Cc: jic23, alison.schofield, icheng, ming.li, benjamin.cheatham,
alucerop, dave, linux-cxl, Jonathan Cameron
A region inherits its coherency from the chosen root decoder - HDM-DB
if the root has CXL_DECODER_F_BI, otherwise HDM-H.
cxl_acpi_cfmws_verify() rejects a Window that declares no coherency
model at all (neither Device Coherent nor Host-only Coherent), one that
sets BI together with Host-only Coherent, which the CFMWS definition
calls undefined behavior, and one that sets BI without Device Coherent,
since HDM-DB is defined only as bit[0] and bit[5] together. A BI Window
therefore always exposes device-coherent memory and nothing else.
The root decoder's target_type follows the Window as well -
device-coherent when only Device Coherent is set, host-only otherwise.
Introduce the following read-only sysfs ABI.
- decoderX.Y/cap_back_invalidate (root) reports the CFMWS BI
restriction.
- decoderX.Y/back_invalidate (endpoint) reads '1' when configured
for HDM-DB.
cxl_region_attach() rejects endpoints whose device or HDM cannot
serve the region's type; target_type is inherited from cxlr->type
in cxl_rr_assign_decoder(), restored to the endpoint default on
detach, and to whatever it was on a failed attach - a refusal may
come before any inheritance, for a decoder another region owns or
one firmware committed, and must not relabel it.
An HDM that reports Unknown coherency support is not rejected.
Supported Coherency Models is how a device declares whether Target
Range Type is writable - Host-only+Device Coherent means RW, a single
model means the bit may be hardwired to it, Unknown declares neither
- so refusing Unknown would also exclude devices that support both
models without saying so.
A Type 3 decoder defaults to host-only, unless its HDM supports only
device-coherent, and inherits device-coherent when it joins an HDM-DB
region, which requires cxlds->bi. A Type 2 decoder defaults to
device-coherent and an HDM-H region refuses it outright, since the
spec reserves HDM-H for Type 3 devices. An HDM-D region is only
assembled from a committed decoder, taking that decoder's type, so no
inheritance is involved.
The HDM Decoder Control BI bit is set at commit time for a
device-coherent decoder in a region under a BI root, endpoint and
switch decoders alike. Whether the device has BI enabled (cxlds->bi)
is checked when an endpoint attaches and again for every target
before a commit programs anything, since a reset in between
invalidates it; the refusal is the same in both places.
Reviewed-by: Dave Jiang <dave.jiang@intel.com>
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
Signed-off-by: Davidlohr Bueso <dave@stgolabs.net>
---
Documentation/ABI/testing/sysfs-bus-cxl | 22 ++++++
drivers/cxl/acpi.c | 27 +++++++
drivers/cxl/core/hdm.c | 19 +++++
drivers/cxl/core/port.c | 39 +++++++++-
drivers/cxl/core/region.c | 99 +++++++++++++++++++++----
drivers/cxl/cxl.h | 7 ++
include/cxl/cxl.h | 3 +-
7 files changed, 197 insertions(+), 19 deletions(-)
diff --git a/Documentation/ABI/testing/sysfs-bus-cxl b/Documentation/ABI/testing/sysfs-bus-cxl
index 7352dbd70bc7..e9b3e2863704 100644
--- a/Documentation/ABI/testing/sysfs-bus-cxl
+++ b/Documentation/ABI/testing/sysfs-bus-cxl
@@ -432,6 +432,28 @@ Description:
current cached value.
+What: /sys/bus/cxl/devices/decoderX.Y/cap_back_invalidate
+Date: September, 2026
+KernelVersion: v7.4
+Contact: linux-cxl@vger.kernel.org
+Description:
+ (RO) When a CXL decoder is of devtype "cxl_decoder_root", '1'
+ indicates the fixed memory window carries the CFMWS
+ Back-Invalidate restriction, so HDM-DB memory may be mapped
+ behind it.
+
+
+What: /sys/bus/cxl/devices/decoderX.Y/back_invalidate
+Date: September, 2026
+KernelVersion: v7.4
+Contact: linux-cxl@vger.kernel.org
+Description:
+ (RO) Shows '1' if this endpoint decoder is currently configured
+ for HDM-DB (device-managed coherency with back-invalidate).
+ The HDM-DB state is inherited from the region the decoder is
+ attached to, which is in turn set from the chosen root
+ decoder's CFMWS BI restriction (see cap_back_invalidate).
+
What: /sys/bus/cxl/devices/decoderX.Y/delete_region
Date: May, 2022
KernelVersion: v6.0
diff --git a/drivers/cxl/acpi.c b/drivers/cxl/acpi.c
index 3b818adbd38b..2e8e31544a5f 100644
--- a/drivers/cxl/acpi.c
+++ b/drivers/cxl/acpi.c
@@ -152,6 +152,8 @@ static unsigned long cfmws_to_decoder_flags(int restrictions)
flags |= CXL_DECODER_F_PMEM;
if (restrictions & ACPI_CEDT_CFMWS_RESTRICT_FIXED)
flags |= CXL_DECODER_F_LOCK;
+ if (restrictions & ACPI_CEDT_CFMWS_RESTRICT_BI)
+ flags |= CXL_DECODER_F_BI;
return flags;
}
@@ -198,6 +200,24 @@ static int cxl_acpi_cfmws_verify(struct device *dev,
dev_dbg(dev, "CFMWS length %d greater than expected %d\n",
cfmws->header.length, expected_len);
+ if ((cfmws->restrictions & ACPI_CEDT_CFMWS_RESTRICT_HOSTONLYMEM) &&
+ (cfmws->restrictions & ACPI_CEDT_CFMWS_RESTRICT_BI)) {
+ dev_err(dev, "CFMWS cannot have both HDM-H and HDM-DB\n");
+ return -EINVAL;
+ }
+
+ if (!(cfmws->restrictions & (ACPI_CEDT_CFMWS_RESTRICT_DEVMEM |
+ ACPI_CEDT_CFMWS_RESTRICT_HOSTONLYMEM))) {
+ dev_err(dev, "CFMWS has no coherency model\n");
+ return -EINVAL;
+ }
+
+ if ((cfmws->restrictions & ACPI_CEDT_CFMWS_RESTRICT_BI) &&
+ !(cfmws->restrictions & ACPI_CEDT_CFMWS_RESTRICT_DEVMEM)) {
+ dev_err(dev, "CFMWS BI requires device-coherent\n");
+ return -EINVAL;
+ }
+
return 0;
}
@@ -437,7 +457,14 @@ static int __cxl_parse_cfmws(struct acpi_cedt_cfmws *cfmws,
cxld = &cxlrd->cxlsd.cxld;
cxld->flags = cfmws_to_decoder_flags(cfmws->restrictions);
+ /* host-only wins if firmware sets both coherency restrictions */
cxld->target_type = CXL_DECODER_HOSTONLYMEM;
+ if (cxld->flags & CXL_DECODER_F_TYPE2) {
+ if (cxld->flags & CXL_DECODER_F_TYPE3)
+ dev_dbg(dev, "CFMWS has both HDM-H and HDM-D\n");
+ else
+ cxld->target_type = CXL_DECODER_DEVMEM;
+ }
cxld->hpa_range = (struct range) {
.start = cfmws->base_hpa,
.end = cfmws->base_hpa + cfmws->window_size - 1,
diff --git a/drivers/cxl/core/hdm.c b/drivers/cxl/core/hdm.c
index 18200a8f3f72..44ff64b1a9df 100644
--- a/drivers/cxl/core/hdm.c
+++ b/drivers/cxl/core/hdm.c
@@ -705,9 +705,21 @@ static void cxld_set_interleave(struct cxl_decoder *cxld, u32 *ctrl)
static void cxld_set_type(struct cxl_decoder *cxld, u32 *ctrl)
{
+ bool bi = cxld->target_type == CXL_DECODER_DEVMEM &&
+ cxld->region && cxl_root_decoder_is_bi(cxld->region->cxlrd);
+
u32p_replace_bits(ctrl,
!!(cxld->target_type == CXL_DECODER_HOSTONLYMEM),
CXL_HDM_DECODER0_CTRL_HOSTONLY);
+ u32p_replace_bits(ctrl, bi, CXL_HDM_DECODER0_CTRL_BI);
+
+ if (bi && is_endpoint_decoder(&cxld->dev)) {
+ struct cxl_endpoint_decoder *cxled =
+ to_cxl_endpoint_decoder(&cxld->dev);
+
+ u32p_replace_bits(ctrl, cxled->pos,
+ CXL_HDM_DECODER0_CTRL_ISP_MASK);
+ }
}
static void cxlsd_set_targets(struct cxl_switch_decoder *cxlsd, u64 *tgt)
@@ -970,6 +982,13 @@ static int cxl_setup_hdm_decoder_from_dvsec(
return 0;
}
+/*
+ * HDMs that advertise support for both coherency modes
+ * (CXL_HDM_DECODER_COHERENCY_BOTH) default to host-only; the region
+ * attach path switches target_type to device-coherent if the region's
+ * root decoder has the CFMWS BI bit set. Only HDMs that strictly
+ * support device-coherent mode default to HDM-DB.
+ */
enum cxl_decoder_type cxled_default_type(struct cxl_endpoint_decoder *cxled)
{
struct cxl_dev_state *cxlds = cxled_to_memdev(cxled)->cxlds;
diff --git a/drivers/cxl/core/port.c b/drivers/cxl/core/port.c
index b81fd680d18a..1d26cd7c88a1 100644
--- a/drivers/cxl/core/port.c
+++ b/drivers/cxl/core/port.c
@@ -132,6 +132,7 @@ CXL_DECODER_FLAG_ATTR(cap_ram, CXL_DECODER_F_RAM);
CXL_DECODER_FLAG_ATTR(cap_type2, CXL_DECODER_F_TYPE2);
CXL_DECODER_FLAG_ATTR(cap_type3, CXL_DECODER_F_TYPE3);
CXL_DECODER_FLAG_ATTR(locked, CXL_DECODER_F_LOCK);
+CXL_DECODER_FLAG_ATTR(cap_back_invalidate, CXL_DECODER_F_BI);
static ssize_t target_type_show(struct device *dev,
struct device_attribute *attr, char *buf)
@@ -234,6 +235,26 @@ static ssize_t mode_store(struct device *dev, struct device_attribute *attr,
}
static DEVICE_ATTR_RW(mode);
+static ssize_t back_invalidate_show(struct device *dev,
+ struct device_attribute *attr, char *buf)
+{
+ struct cxl_endpoint_decoder *cxled = to_cxl_endpoint_decoder(dev);
+ struct cxl_dev_state *cxlds = cxled_to_memdev(cxled)->cxlds;
+ struct cxl_region *cxlr;
+
+ guard(rwsem_read)(&cxl_rwsem.region);
+ /*
+ * An endpoint decoder is HDM-DB when the device has BI enabled
+ * (cxlds->bi) and it is attached to a device-coherent (DEVMEM)
+ * region whose root decoder advertises the CFMWS BI restriction.
+ */
+ cxlr = cxled->cxld.region;
+ return sysfs_emit(buf, "%d\n", cxlds->bi && cxlr &&
+ cxled->cxld.target_type == CXL_DECODER_DEVMEM &&
+ cxl_root_decoder_is_bi(cxlr->cxlrd));
+}
+static DEVICE_ATTR_RO(back_invalidate);
+
static ssize_t dpa_resource_show(struct device *dev, struct device_attribute *attr,
char *buf)
{
@@ -330,6 +351,7 @@ static struct attribute *cxl_decoder_root_attrs[] = {
&dev_attr_cap_ram.attr,
&dev_attr_cap_type2.attr,
&dev_attr_cap_type3.attr,
+ &dev_attr_cap_back_invalidate.attr,
&dev_attr_target_list.attr,
&dev_attr_qos_class.attr,
SET_CXL_REGION_ATTR(create_pmem_region)
@@ -340,16 +362,24 @@ static struct attribute *cxl_decoder_root_attrs[] = {
static bool can_create_pmem(struct cxl_root_decoder *cxlrd)
{
- unsigned long flags = CXL_DECODER_F_TYPE3 | CXL_DECODER_F_PMEM;
+ unsigned long flags = cxlrd->cxlsd.cxld.flags;
+ unsigned long hdm_h, hdm_db;
- return (cxlrd->cxlsd.cxld.flags & flags) == flags;
+ hdm_h = CXL_DECODER_F_TYPE3 | CXL_DECODER_F_PMEM;
+ hdm_db = CXL_DECODER_F_TYPE2 | CXL_DECODER_F_BI | CXL_DECODER_F_PMEM;
+
+ return (flags & hdm_h) == hdm_h || (flags & hdm_db) == hdm_db;
}
static bool can_create_ram(struct cxl_root_decoder *cxlrd)
{
- unsigned long flags = CXL_DECODER_F_TYPE3 | CXL_DECODER_F_RAM;
+ unsigned long flags = cxlrd->cxlsd.cxld.flags;
+ unsigned long hdm_h, hdm_db;
+
+ hdm_h = CXL_DECODER_F_TYPE3 | CXL_DECODER_F_RAM;
+ hdm_db = CXL_DECODER_F_TYPE2 | CXL_DECODER_F_BI | CXL_DECODER_F_RAM;
- return (cxlrd->cxlsd.cxld.flags & flags) == flags;
+ return (flags & hdm_h) == hdm_h || (flags & hdm_db) == hdm_db;
}
static umode_t cxl_root_decoder_visible(struct kobject *kobj, struct attribute *a, int n)
@@ -403,6 +433,7 @@ static const struct attribute_group *cxl_decoder_switch_attribute_groups[] = {
static struct attribute *cxl_decoder_endpoint_attrs[] = {
&dev_attr_target_type.attr,
&dev_attr_mode.attr,
+ &dev_attr_back_invalidate.attr,
&dev_attr_dpa_size.attr,
&dev_attr_dpa_resource.attr,
SET_CXL_REGION_ATTR(region)
diff --git a/drivers/cxl/core/region.c b/drivers/cxl/core/region.c
index 27e63e6dab7c..0a204aea64b5 100644
--- a/drivers/cxl/core/region.c
+++ b/drivers/cxl/core/region.c
@@ -312,8 +312,25 @@ static int commit_decoder(struct cxl_decoder *cxld)
static int cxl_region_decode_commit(struct cxl_region *cxlr)
{
struct cxl_region_params *p = &cxlr->params;
+ bool hdm_db;
int i, rc = 0;
+ hdm_db = cxlr->type == CXL_DECODER_DEVMEM &&
+ cxl_root_decoder_is_bi(cxlr->cxlrd);
+
+ /* a reset since attach may have invalidated cxlds->bi */
+ for (i = 0; i < p->nr_targets; i++) {
+ struct cxl_endpoint_decoder *cxled = p->targets[i];
+ struct cxl_memdev *cxlmd = cxled_to_memdev(cxled);
+
+ if (hdm_db && !cxlmd->cxlds->bi) {
+ dev_err(&cxlr->dev, "%s:%s BI not enabled on device\n",
+ dev_name(&cxlmd->dev),
+ dev_name(&cxled->cxld.dev));
+ return -ENXIO;
+ }
+ }
+
for (i = 0; i < p->nr_targets; i++) {
struct cxl_endpoint_decoder *cxled = p->targets[i];
struct cxl_memdev *cxlmd = cxled_to_memdev(cxled);
@@ -345,6 +362,21 @@ static int cxl_region_decode_commit(struct cxl_region *cxlr)
cxled->cxld.reset(&cxled->cxld);
goto err;
}
+
+ /*
+ * A reset after the check above wipes BI Enable, but
+ * cxlds->bi is only refreshed from .reset_done, which may
+ * not have run. Ask the hardware.
+ */
+ if (hdm_db && !cxl_bi_decoder_enabled(cxled_to_port(cxled))) {
+ dev_err(&cxlr->dev,
+ "%s:%s BI disabled by reset during commit\n",
+ dev_name(&cxlmd->dev),
+ dev_name(&cxled->cxld.dev));
+ rc = -ENXIO;
+ i++; /* this target committed, so undo it too */
+ goto err;
+ }
}
return 0;
@@ -1131,16 +1163,11 @@ static int cxl_rr_assign_decoder(struct cxl_port *port, struct cxl_region *cxlr,
}
/*
- * Endpoints should already match the region type, but backstop that
- * assumption with an assertion. Switch-decoders change mapping-type
- * based on what is mapped when they are assigned to a region.
+ * Endpoint decoders inherit their type from cxlr->type; broken
+ * pairings were already rejected by the coherency checks in
+ * cxl_region_attach(). Switch-decoders change mapping-type based
+ * on what is mapped when they are assigned to a region.
*/
- dev_WARN_ONCE(&cxlr->dev,
- port == cxled_to_port(cxled) &&
- cxld->target_type != cxlr->type,
- "%s:%s mismatch decoder type %d -> %d\n",
- dev_name(&cxled_to_memdev(cxled)->dev),
- dev_name(&cxld->dev), cxld->target_type, cxlr->type);
cxld->target_type = cxlr->type;
cxl_rr->decoder = cxld;
return 0;
@@ -1803,6 +1830,7 @@ static int cxl_region_attach_position(struct cxl_region *cxlr,
struct cxl_root_decoder *cxlrd = cxlr->cxlrd;
struct cxl_memdev *cxlmd = cxled_to_memdev(cxled);
struct cxl_switch_decoder *cxlsd = &cxlrd->cxlsd;
+ enum cxl_decoder_type type = cxled->cxld.target_type;
struct cxl_decoder *cxld = &cxlsd->cxld;
int iw = cxld->interleave_ways;
struct cxl_port *iter;
@@ -1828,6 +1856,8 @@ static int cxl_region_attach_position(struct cxl_region *cxlr,
for (iter = cxled_to_port(cxled); !is_cxl_root(iter);
iter = to_cxl_port(iter->dev.parent))
cxl_port_detach_region(iter, cxlr, cxled);
+ /* undo cxl_rr_assign_decoder() type inheritance */
+ cxled->cxld.target_type = type;
return rc;
}
@@ -2056,6 +2086,7 @@ static int cxl_region_attach(struct cxl_region *cxlr,
struct cxl_region_params *p = &cxlr->params;
struct cxl_port *ep_port, *root_port;
struct cxl_dport *dport;
+ struct cxl_hdm *cxlhdm;
int rc = -ENXIO;
rc = check_interleave_cap(&cxled->cxld, p->interleave_ways,
@@ -2105,10 +2136,39 @@ static int cxl_region_attach(struct cxl_region *cxlr,
return -ENXIO;
}
- if (cxled->cxld.target_type != cxlr->type) {
- dev_dbg(&cxlr->dev, "%s:%s type mismatch: %d vs %d\n",
- dev_name(&cxlmd->dev), dev_name(&cxled->cxld.dev),
- cxled->cxld.target_type, cxlr->type);
+ /*
+ * Verify the device and HDM are capable of the region's flavor before
+ * proceeding. The endpoint decoder's target_type is then inherited
+ * from cxlr->type later in cxl_rr_assign_decoder().
+ */
+ if (cxlr->type == CXL_DECODER_DEVMEM &&
+ cxl_root_decoder_is_bi(cxlrd) && !cxlds->bi) {
+ dev_err(&cxlr->dev, "%s:%s BI not enabled on device\n",
+ dev_name(&cxlmd->dev), dev_name(&cxled->cxld.dev));
+ return -ENXIO;
+ }
+
+ if (cxled->state != CXL_DECODER_STATE_AUTO &&
+ cxlr->type == CXL_DECODER_HOSTONLYMEM &&
+ cxlds->type == CXL_DEVTYPE_DEVMEM) {
+ dev_warn(&cxlr->dev, "%s:%s HDM-H requires a Type 3 device\n",
+ dev_name(&cxlmd->dev), dev_name(&cxled->cxld.dev));
+ return -ENXIO;
+ }
+
+ cxlhdm = dev_get_drvdata(&ep_port->dev);
+ if (!cxlhdm)
+ return -ENXIO;
+ if (cxlr->type == CXL_DECODER_HOSTONLYMEM &&
+ cxlhdm->supported_coherency == CXL_HDM_DECODER_COHERENCY_DEV) {
+ dev_warn(&cxlr->dev, "%s:%s HDM is device-coherent only\n",
+ dev_name(&cxlmd->dev), dev_name(&cxled->cxld.dev));
+ return -ENXIO;
+ }
+ if (cxlr->type == CXL_DECODER_DEVMEM &&
+ cxlhdm->supported_coherency == CXL_HDM_DECODER_COHERENCY_HOST) {
+ dev_warn(&cxlr->dev, "%s:%s HDM is host-only coherent\n",
+ dev_name(&cxlmd->dev), dev_name(&cxled->cxld.dev));
return -ENXIO;
}
@@ -2324,6 +2384,8 @@ __cxl_decoder_detach(struct cxl_region *cxlr,
.start = 0,
.end = -1,
};
+ /* undo cxl_rr_assign_decoder() type inheritance */
+ cxled->cxld.target_type = cxled_default_type(cxled);
get_device(&cxlr->dev);
return cxlr;
@@ -2820,6 +2882,7 @@ static ssize_t create_region_store(struct device *dev, const char *buf,
size_t len, enum cxl_partition_mode mode)
{
struct cxl_root_decoder *cxlrd = to_cxl_root_decoder(dev);
+ enum cxl_decoder_type target_type;
struct cxl_region *cxlr;
int rc, id;
@@ -2831,7 +2894,15 @@ static ssize_t create_region_store(struct device *dev, const char *buf,
if ((rc = ACQUIRE_ERR(mutex_intr, ®ions_lock)))
return rc;
- cxlr = __create_region(cxlrd, mode, id, CXL_DECODER_HOSTONLYMEM);
+ /*
+ * The CFMWS dictates endpoint coherency: a BI-restricted Window
+ * produces an HDM-DB region; otherwise HDM-H. HDM-D is not an
+ * option here, type2 regions are created by their driver.
+ */
+ target_type = cxl_root_decoder_is_bi(cxlrd) ?
+ CXL_DECODER_DEVMEM : CXL_DECODER_HOSTONLYMEM;
+
+ cxlr = __create_region(cxlrd, mode, id, target_type);
if (IS_ERR(cxlr))
return PTR_ERR(cxlr);
diff --git a/drivers/cxl/cxl.h b/drivers/cxl/cxl.h
index a990ed41edef..9fb0278e2ce4 100644
--- a/drivers/cxl/cxl.h
+++ b/drivers/cxl/cxl.h
@@ -78,6 +78,7 @@ extern const struct nvdimm_security_ops *cxl_security_ops;
#define CXL_HDM_DECODER0_CTRL_COMMIT_ERROR BIT(11)
#define CXL_HDM_DECODER0_CTRL_HOSTONLY BIT(12)
#define CXL_HDM_DECODER0_CTRL_BI BIT(13)
+#define CXL_HDM_DECODER0_CTRL_ISP_MASK GENMASK(27, 24)
#define CXL_HDM_DECODER0_TL_LOW(i) (0x20 * (i) + 0x24)
#define CXL_HDM_DECODER0_TL_HIGH(i) (0x20 * (i) + 0x28)
#define CXL_HDM_DECODER0_SKIP_LOW(i) CXL_HDM_DECODER0_TL_LOW(i)
@@ -300,6 +301,7 @@ int cxl_dport_map_rcd_linkcap(struct pci_dev *pdev, struct cxl_dport *dport);
#define CXL_DECODER_F_LOCK BIT(4)
#define CXL_DECODER_F_ENABLE BIT(5)
#define CXL_DECODER_F_NORMALIZED_ADDRESSING BIT(6)
+#define CXL_DECODER_F_BI BIT(7)
#define CXL_DECODER_F_RESET_MASK (CXL_DECODER_F_ENABLE | CXL_DECODER_F_LOCK)
enum cxl_decoder_type {
@@ -829,6 +831,11 @@ static inline int cxl_root_decoder_autoremove(struct device *host,
{
return cxl_decoder_autoremove(host, &cxlrd->cxlsd.cxld);
}
+
+static inline bool cxl_root_decoder_is_bi(struct cxl_root_decoder *cxlrd)
+{
+ return cxlrd->cxlsd.cxld.flags & CXL_DECODER_F_BI;
+}
int cxl_endpoint_autoremove(struct cxl_memdev *cxlmd, struct cxl_port *endpoint);
/**
diff --git a/include/cxl/cxl.h b/include/cxl/cxl.h
index 718eb4353887..f9e9058a85da 100644
--- a/include/cxl/cxl.h
+++ b/include/cxl/cxl.h
@@ -16,7 +16,8 @@
* mailbox, or other memory-device-standard manageability
* flows.
* @CXL_DEVTYPE_CLASSMEM: Common class definition of a CXL Type-3 device with
- * HDM-H and class-mandatory memory device registers
+ * HDM-H or HDM-DB, and class-mandatory memory device
+ * registers
*/
enum cxl_devtype {
CXL_DEVTYPE_DEVMEM,
--
2.39.5
^ permalink raw reply related [flat|nested] 33+ messages in thread* Re: [PATCH v9 04/10] cxl: Add HDM-DB region creation
2026-09-22 23:38 ` [PATCH v9 04/10] cxl: Add HDM-DB region creation Davidlohr Bueso
@ 2026-10-01 3:28 ` Alison Schofield
2026-10-01 8:57 ` Davidlohr Bueso
0 siblings, 1 reply; 33+ messages in thread
From: Alison Schofield @ 2026-10-01 3:28 UTC (permalink / raw)
To: Davidlohr Bueso
Cc: dave.jiang, jic23, icheng, ming.li, benjamin.cheatham, alucerop,
linux-cxl, Jonathan Cameron
On Tue, Sep 22, 2026 at 04:38:41PM -0700, Davidlohr Bueso wrote:
> A region inherits its coherency from the chosen root decoder - HDM-DB
> if the root has CXL_DECODER_F_BI, otherwise HDM-H.
Above holds for regions created from sysfs. An auto region takes its type
from the first committed endpoint decoder in construct_region(), without
looking at the root decoder. A committed host-only decoder in a BI window
becomes an HDM-H region under a BI root, and an HDM-D region is neither
HDM-DB nor HDM-H. Can this say which regions it applies to?
>
> cxl_acpi_cfmws_verify() rejects a Window that declares no coherency
> model at all (neither Device Coherent nor Host-only Coherent), one that
> sets BI together with Host-only Coherent, which the CFMWS definition
> calls undefined behavior, and one that sets BI without Device Coherent,
> since HDM-DB is defined only as bit[0] and bit[5] together. A BI Window
> therefore always exposes device-coherent memory and nothing else.
These rejections drop windows that are accepted and usable today. See
the cxl_acpi_cfmws_verify() comment inline below.
> The root decoder's target_type follows the Window as well -
> device-coherent when only Device Coherent is set, host-only otherwise.
I'll comment inline on this one too, acpi.c hunk below. I think this
describes dead code and opportunity for cleanup.
>
> Introduce the following read-only sysfs ABI.
>
> - decoderX.Y/cap_back_invalidate (root) reports the CFMWS BI
> restriction.
> - decoderX.Y/back_invalidate (endpoint) reads '1' when configured
> for HDM-DB.
>
> cxl_region_attach() rejects endpoints whose device or HDM cannot
> serve the region's type; target_type is inherited from cxlr->type
> in cxl_rr_assign_decoder(), restored to the endpoint default on
> detach, and to whatever it was on a failed attach - a refusal may
> come before any inheritance, for a decoder another region owns or
> one firmware committed, and must not relabel it.
A decoder firmware committed is still relabelled when its attach succeeds,
since the type mismatch check is gone. More inline at
cxl_region_attach()
snip
> +
> +
> +What: /sys/bus/cxl/devices/decoderX.Y/back_invalidate
> +Date: September, 2026
> +KernelVersion: v7.4
> +Contact: linux-cxl@vger.kernel.org
> +Description:
> + (RO) Shows '1' if this endpoint decoder is currently configured
> + for HDM-DB (device-managed coherency with back-invalidate).
> + The HDM-DB state is inherited from the region the decoder is
> + attached to, which is in turn set from the chosen root
> + decoder's CFMWS BI restriction (see cap_back_invalidate).
> +
> What: /sys/bus/cxl/devices/decoderX.Y/delete_region
> Date: May, 2022
> KernelVersion: v6.0
> diff --git a/drivers/cxl/acpi.c b/drivers/cxl/acpi.c
> index 3b818adbd38b..2e8e31544a5f 100644
> --- a/drivers/cxl/acpi.c
> +++ b/drivers/cxl/acpi.c
> @@ -152,6 +152,8 @@ static unsigned long cfmws_to_decoder_flags(int restrictions)
> flags |= CXL_DECODER_F_PMEM;
> if (restrictions & ACPI_CEDT_CFMWS_RESTRICT_FIXED)
> flags |= CXL_DECODER_F_LOCK;
> + if (restrictions & ACPI_CEDT_CFMWS_RESTRICT_BI)
> + flags |= CXL_DECODER_F_BI;
>
> return flags;
> }
> @@ -198,6 +200,24 @@ static int cxl_acpi_cfmws_verify(struct device *dev,
> dev_dbg(dev, "CFMWS length %d greater than expected %d\n",
> cfmws->header.length, expected_len);
>
> + if ((cfmws->restrictions & ACPI_CEDT_CFMWS_RESTRICT_HOSTONLYMEM) &&
> + (cfmws->restrictions & ACPI_CEDT_CFMWS_RESTRICT_BI)) {
> + dev_err(dev, "CFMWS cannot have both HDM-H and HDM-DB\n");
> + return -EINVAL;
> + }
> +
> + if (!(cfmws->restrictions & (ACPI_CEDT_CFMWS_RESTRICT_DEVMEM |
> + ACPI_CEDT_CFMWS_RESTRICT_HOSTONLYMEM))) {
> + dev_err(dev, "CFMWS has no coherency model\n");
> + return -EINVAL;
> + }
> +
> + if ((cfmws->restrictions & ACPI_CEDT_CFMWS_RESTRICT_BI) &&
> + !(cfmws->restrictions & ACPI_CEDT_CFMWS_RESTRICT_DEVMEM)) {
> + dev_err(dev, "CFMWS BI requires device-coherent\n");
> + return -EINVAL;
> + }
> +
Can these checks break windows that are accepted and usable today? A
rejected window gets no root decoder, and auto-assembly then fails with
"no CXL window for range".
A window that sets neither coherency bit is accepted today and can hold
auto regions, since assembly never looks at the restriction flags.
A window that sets BI together with Host-only, or BI without Device Coherent,
is accepted today as a non-BI window, because BI is ignored.
Manual HDM-H region creation in it is lost as well.
Even if no shipping firmware does this, is cxl_acpi_cfmws_verify() the right
place for these checks? Til now it has only rejected windows that cannot be
decoded (arithmetic, alignment, ways, length). What a window may be used for
is decided later, by cfmws_to_decoder_flags() and can_create_ram|pmem().
Might it be enough to withhold CXL_DECODER_F_BI for an invalid combination,
so that only HDM-DB is refused?
> return 0;
> }
>
> @@ -437,7 +457,14 @@ static int __cxl_parse_cfmws(struct acpi_cedt_cfmws *cfmws,
>
> cxld = &cxlrd->cxlsd.cxld;
> cxld->flags = cfmws_to_decoder_flags(cfmws->restrictions);
> + /* host-only wins if firmware sets both coherency restrictions */
> cxld->target_type = CXL_DECODER_HOSTONLYMEM;
> + if (cxld->flags & CXL_DECODER_F_TYPE2) {
> + if (cxld->flags & CXL_DECODER_F_TYPE3)
> + dev_dbg(dev, "CFMWS has both HDM-H and HDM-D\n");
> + else
> + cxld->target_type = CXL_DECODER_DEVMEM;
> + }
Mentioned in commit msg, I think this is dead code needing cleanup.
I don't see a consumer of a root decoder's target_type.
Building window-derived logic and a dev_dbg() on top of it makes it
look as if the root decoder's type matters when it doesn't. Can this
hunk be dropped? Removing the old line can be a separate cleanup.
> cxld->hpa_range = (struct range) {
> .start = cfmws->base_hpa,
> .end = cfmws->base_hpa + cfmws->window_size - 1,
> diff --git a/drivers/cxl/core/hdm.c b/drivers/cxl/core/hdm.c
> index 18200a8f3f72..44ff64b1a9df 100644
> --- a/drivers/cxl/core/hdm.c
> +++ b/drivers/cxl/core/hdm.c
> @@ -705,9 +705,21 @@ static void cxld_set_interleave(struct cxl_decoder *cxld, u32 *ctrl)
>
> static void cxld_set_type(struct cxl_decoder *cxld, u32 *ctrl)
> {
> + bool bi = cxld->target_type == CXL_DECODER_DEVMEM &&
> + cxld->region && cxl_root_decoder_is_bi(cxld->region->cxlrd);
> +
> u32p_replace_bits(ctrl,
> !!(cxld->target_type == CXL_DECODER_HOSTONLYMEM),
> CXL_HDM_DECODER0_CTRL_HOSTONLY);
> + u32p_replace_bits(ctrl, bi, CXL_HDM_DECODER0_CTRL_BI);
> +
> + if (bi && is_endpoint_decoder(&cxld->dev)) {
> + struct cxl_endpoint_decoder *cxled =
> + to_cxl_endpoint_decoder(&cxld->dev);
> +
> + u32p_replace_bits(ctrl, cxled->pos,
> + CXL_HDM_DECODER0_CTRL_ISP_MASK);
> + }
The commit message does not mention ISP programming. Can it get a
sentence with the spec reference for ISP being required when BI is set?
snip
> diff --git a/drivers/cxl/core/region.c b/drivers/cxl/core/region.c
snip
> @@ -1131,16 +1163,11 @@ static int cxl_rr_assign_decoder(struct cxl_port *port, struct cxl_region *cxlr,
> }
>
> /*
> - * Endpoints should already match the region type, but backstop that
> - * assumption with an assertion. Switch-decoders change mapping-type
> - * based on what is mapped when they are assigned to a region.
> + * Endpoint decoders inherit their type from cxlr->type; broken
> + * pairings were already rejected by the coherency checks in
> + * cxl_region_attach(). Switch-decoders change mapping-type based
> + * on what is mapped when they are assigned to a region.
> */
> - dev_WARN_ONCE(&cxlr->dev,
> - port == cxled_to_port(cxled) &&
> - cxld->target_type != cxlr->type,
> - "%s:%s mismatch decoder type %d -> %d\n",
> - dev_name(&cxled_to_memdev(cxled)->dev),
> - dev_name(&cxld->dev), cxld->target_type, cxlr->type);
> cxld->target_type = cxlr->type;
A Type 3 endpoint decoder in an HDM-DB region now reads "accelerator" from
decoderX.Y/target_type, and the ndctl test on the branch linked in the cover
letter checks for that. So I take it target_type is now being used to report
the decoder's coherency model, device-coherent vs host-only, rather than the
Type-2/Type-3 as documented by the existing ABI.
Does this need sysfs ABI update?
cxl list doesn't show target_type, but libcxl exports it through
cxl_decoder_get_target_type(). If that is intentional, libcxl.txt will
need the same semantic update with the ndctl patches.
> cxl_rr->decoder = cxld;
> return 0;
> @@ -1803,6 +1830,7 @@ static int cxl_region_attach_position(struct cxl_region *cxlr,
> struct cxl_root_decoder *cxlrd = cxlr->cxlrd;
> struct cxl_memdev *cxlmd = cxled_to_memdev(cxled);
> struct cxl_switch_decoder *cxlsd = &cxlrd->cxlsd;
> + enum cxl_decoder_type type = cxled->cxld.target_type;
> struct cxl_decoder *cxld = &cxlsd->cxld;
> int iw = cxld->interleave_ways;
> struct cxl_port *iter;
> @@ -1828,6 +1856,8 @@ static int cxl_region_attach_position(struct cxl_region *cxlr,
> for (iter = cxled_to_port(cxled); !is_cxl_root(iter);
> iter = to_cxl_port(iter->dev.parent))
> cxl_port_detach_region(iter, cxlr, cxled);
> + /* undo cxl_rr_assign_decoder() type inheritance */
> + cxled->cxld.target_type = type;
> return rc;
> }
>
> @@ -2056,6 +2086,7 @@ static int cxl_region_attach(struct cxl_region *cxlr,
> struct cxl_region_params *p = &cxlr->params;
> struct cxl_port *ep_port, *root_port;
> struct cxl_dport *dport;
> + struct cxl_hdm *cxlhdm;
> int rc = -ENXIO;
>
> rc = check_interleave_cap(&cxled->cxld, p->interleave_ways,
> @@ -2105,10 +2136,39 @@ static int cxl_region_attach(struct cxl_region *cxlr,
> return -ENXIO;
> }
>
> - if (cxled->cxld.target_type != cxlr->type) {
> - dev_dbg(&cxlr->dev, "%s:%s type mismatch: %d vs %d\n",
> - dev_name(&cxlmd->dev), dev_name(&cxled->cxld.dev),
> - cxled->cxld.target_type, cxlr->type);
Should the type check stay for committed decoders? With it gone, a
second auto-assembled decoder whose committed Target Range Type differs
from the region's is accepted, and cxl_rr_assign_decoder() then
overwrites its target_type while the hardware keeps the committed value.
I think this changes in Patch 8 - which really makes patch 8 required,
not optional. (more on that in patch 8)
Maybe the CXL_DECODER_STATE_AUTO mismatch check fits better in here,
since this patch that removes the old one?
> + /*
> + * Verify the device and HDM are capable of the region's flavor before
> + * proceeding. The endpoint decoder's target_type is then inherited
> + * from cxlr->type later in cxl_rr_assign_decoder().
> + */
> + if (cxlr->type == CXL_DECODER_DEVMEM &&
> + cxl_root_decoder_is_bi(cxlrd) && !cxlds->bi) {
> + dev_err(&cxlr->dev, "%s:%s BI not enabled on device\n",
> + dev_name(&cxlmd->dev), dev_name(&cxled->cxld.dev));
> + return -ENXIO;
> + }
> +
> + if (cxled->state != CXL_DECODER_STATE_AUTO &&
> + cxlr->type == CXL_DECODER_HOSTONLYMEM &&
> + cxlds->type == CXL_DEVTYPE_DEVMEM) {
> + dev_warn(&cxlr->dev, "%s:%s HDM-H requires a Type 3 device\n",
> + dev_name(&cxlmd->dev), dev_name(&cxled->cxld.dev));
> + return -ENXIO;
> + }
> +
> + cxlhdm = dev_get_drvdata(&ep_port->dev);
> + if (!cxlhdm)
> + return -ENXIO;
> + if (cxlr->type == CXL_DECODER_HOSTONLYMEM &&
> + cxlhdm->supported_coherency == CXL_HDM_DECODER_COHERENCY_DEV) {
> + dev_warn(&cxlr->dev, "%s:%s HDM is device-coherent only\n",
> + dev_name(&cxlmd->dev), dev_name(&cxled->cxld.dev));
> + return -ENXIO;
> + }
> + if (cxlr->type == CXL_DECODER_DEVMEM &&
> + cxlhdm->supported_coherency == CXL_HDM_DECODER_COHERENCY_HOST) {
> + dev_warn(&cxlr->dev, "%s:%s HDM is host-only coherent\n",
> + dev_name(&cxlmd->dev), dev_name(&cxled->cxld.dev));
> return -ENXIO;
> }
Why dev_err()/dev_warn() here? The type mismatch refusal this replaces was
a dev_dbg(). The attach fail here would be reported as -ENXIO.
Tidy up opportunity, maybe move the 4 coherency checks into a helper so
cxl_region_attach() stays cleaner.
Another cleanup - the 'is this reion HDM-DB test is now coded in a few
places, my look may have been limited:
cxld_set_type(): target_type == DEVMEM && cxl_root_decoder_is_bi()
back_invalidate_show(): same, plus cxlds->bi
cxl_region_decode_commit(): cxlr->type == DEVMEM && cxl_root_decoder_is_bi()
cxl_region_attach(): same
Maybe a single cxl_region_is_hdm_db() would be clearer and cleaner.
snip to end
^ permalink raw reply [flat|nested] 33+ messages in thread* Re: [PATCH v9 04/10] cxl: Add HDM-DB region creation
2026-10-01 3:28 ` Alison Schofield
@ 2026-10-01 8:57 ` Davidlohr Bueso
0 siblings, 0 replies; 33+ messages in thread
From: Davidlohr Bueso @ 2026-10-01 8:57 UTC (permalink / raw)
To: Alison Schofield
Cc: dave.jiang, jic23, icheng, ming.li, benjamin.cheatham, alucerop,
linux-cxl, Jonathan Cameron
On Wed, 30 Sep 2026, Alison Schofield wrote:
>On Tue, Sep 22, 2026 at 04:38:41PM -0700, Davidlohr Bueso wrote:
>> A region inherits its coherency from the chosen root decoder - HDM-DB
>> if the root has CXL_DECODER_F_BI, otherwise HDM-H.
>
>Above holds for regions created from sysfs. An auto region takes its type
>from the first committed endpoint decoder in construct_region(), without
>looking at the root decoder. A committed host-only decoder in a BI window
>becomes an HDM-H region under a BI root, and an HDM-D region is neither
>HDM-DB nor HDM-H. Can this say which regions it applies to?
Right, that only describes regions created through sysfs. An
auto-assembled region still takes its type from the committed decoder
it is built from, so under a BI root a committed host-only decoder
gives an HDM-H region and an HDM-D region is neither, as you say.
Nothing here consults the Window for a committed decoder; that
starts with patch 8, which also removes the refusal of committed BI
decoders at enumeration, so auto HDM-DB regions do not exist before
it.
I'll scope the sentence to sysfs regions.
>
>>
>> cxl_acpi_cfmws_verify() rejects a Window that declares no coherency
>> model at all (neither Device Coherent nor Host-only Coherent), one that
>> sets BI together with Host-only Coherent, which the CFMWS definition
>> calls undefined behavior, and one that sets BI without Device Coherent,
>> since HDM-DB is defined only as bit[0] and bit[5] together. A BI Window
>> therefore always exposes device-coherent memory and nothing else.
>
>These rejections drop windows that are accepted and usable today. See
>the cxl_acpi_cfmws_verify() comment inline below.
>
>> The root decoder's target_type follows the Window as well -
>> device-coherent when only Device Coherent is set, host-only otherwise.
>
>I'll comment inline on this one too, acpi.c hunk below. I think this
>describes dead code and opportunity for cleanup.
>
>>
>> Introduce the following read-only sysfs ABI.
>>
>> - decoderX.Y/cap_back_invalidate (root) reports the CFMWS BI
>> restriction.
>> - decoderX.Y/back_invalidate (endpoint) reads '1' when configured
>> for HDM-DB.
>>
>> cxl_region_attach() rejects endpoints whose device or HDM cannot
>> serve the region's type; target_type is inherited from cxlr->type
>> in cxl_rr_assign_decoder(), restored to the endpoint default on
>> detach, and to whatever it was on a failed attach - a refusal may
>> come before any inheritance, for a decoder another region owns or
>> one firmware committed, and must not relabel it.
>
>A decoder firmware committed is still relabelled when its attach succeeds,
>since the type mismatch check is gone. More inline at
>cxl_region_attach()
>
>snip
>> +
>> +
>> +What: /sys/bus/cxl/devices/decoderX.Y/back_invalidate
>> +Date: September, 2026
>> +KernelVersion: v7.4
>> +Contact: linux-cxl@vger.kernel.org
>> +Description:
>> + (RO) Shows '1' if this endpoint decoder is currently configured
>> + for HDM-DB (device-managed coherency with back-invalidate).
>> + The HDM-DB state is inherited from the region the decoder is
>> + attached to, which is in turn set from the chosen root
>> + decoder's CFMWS BI restriction (see cap_back_invalidate).
>> +
>> What: /sys/bus/cxl/devices/decoderX.Y/delete_region
>> Date: May, 2022
>> KernelVersion: v6.0
>> diff --git a/drivers/cxl/acpi.c b/drivers/cxl/acpi.c
>> index 3b818adbd38b..2e8e31544a5f 100644
>> --- a/drivers/cxl/acpi.c
>> +++ b/drivers/cxl/acpi.c
>> @@ -152,6 +152,8 @@ static unsigned long cfmws_to_decoder_flags(int restrictions)
>> flags |= CXL_DECODER_F_PMEM;
>> if (restrictions & ACPI_CEDT_CFMWS_RESTRICT_FIXED)
>> flags |= CXL_DECODER_F_LOCK;
>> + if (restrictions & ACPI_CEDT_CFMWS_RESTRICT_BI)
>> + flags |= CXL_DECODER_F_BI;
>>
>> return flags;
>> }
>> @@ -198,6 +200,24 @@ static int cxl_acpi_cfmws_verify(struct device *dev,
>> dev_dbg(dev, "CFMWS length %d greater than expected %d\n",
>> cfmws->header.length, expected_len);
>>
>> + if ((cfmws->restrictions & ACPI_CEDT_CFMWS_RESTRICT_HOSTONLYMEM) &&
>> + (cfmws->restrictions & ACPI_CEDT_CFMWS_RESTRICT_BI)) {
>> + dev_err(dev, "CFMWS cannot have both HDM-H and HDM-DB\n");
>> + return -EINVAL;
>> + }
>> +
>> + if (!(cfmws->restrictions & (ACPI_CEDT_CFMWS_RESTRICT_DEVMEM |
>> + ACPI_CEDT_CFMWS_RESTRICT_HOSTONLYMEM))) {
>> + dev_err(dev, "CFMWS has no coherency model\n");
>> + return -EINVAL;
>> + }
>> +
>> + if ((cfmws->restrictions & ACPI_CEDT_CFMWS_RESTRICT_BI) &&
>> + !(cfmws->restrictions & ACPI_CEDT_CFMWS_RESTRICT_DEVMEM)) {
>> + dev_err(dev, "CFMWS BI requires device-coherent\n");
>> + return -EINVAL;
>> + }
>> +
>
>
>Can these checks break windows that are accepted and usable today? A
>rejected window gets no root decoder, and auto-assembly then fails with
>"no CXL window for range".
>
>A window that sets neither coherency bit is accepted today and can hold
>auto regions, since assembly never looks at the restriction flags.
I could remove the check for no ACPI_CEDT_CFMWS_RESTRICT_DEVMEM/HOSTONLYMEM,
but would rather break bogus firmware.
>
>A window that sets BI together with Host-only, or BI without Device Coherent,
>is accepted today as a non-BI window, because BI is ignored.
>Manual HDM-H region creation in it is lost as well.
We cannot accept UB scenarios, I think these should be rejected regardless.
>
>Even if no shipping firmware does this, is cxl_acpi_cfmws_verify() the right
>place for these checks? Til now it has only rejected windows that cannot be
>decoded (arithmetic, alignment, ways, length). What a window may be used for
>is decided later, by cfmws_to_decoder_flags() and can_create_ram|pmem().
>Might it be enough to withhold CXL_DECODER_F_BI for an invalid combination,
>so that only HDM-DB is refused?
Yes, I think this is the right place for these checks, aligned with its
name cxl_acpi_cfmws_verify(). These are minimum sanity checks, just as you
point out with those that cannot be decoded.
I will leave these checks as is.
>
>
>> return 0;
>> }
>>
>> @@ -437,7 +457,14 @@ static int __cxl_parse_cfmws(struct acpi_cedt_cfmws *cfmws,
>>
>> cxld = &cxlrd->cxlsd.cxld;
>> cxld->flags = cfmws_to_decoder_flags(cfmws->restrictions);
>> + /* host-only wins if firmware sets both coherency restrictions */
>> cxld->target_type = CXL_DECODER_HOSTONLYMEM;
>> + if (cxld->flags & CXL_DECODER_F_TYPE2) {
>> + if (cxld->flags & CXL_DECODER_F_TYPE3)
>> + dev_dbg(dev, "CFMWS has both HDM-H and HDM-D\n");
>> + else
>> + cxld->target_type = CXL_DECODER_DEVMEM;
>> + }
>
>
>Mentioned in commit msg, I think this is dead code needing cleanup.
>I don't see a consumer of a root decoder's target_type.
>
>Building window-derived logic and a dev_dbg() on top of it makes it
>look as if the root decoder's type matters when it doesn't. Can this
>hunk be dropped? Removing the old line can be a separate cleanup.
Agreed. I justified this because we already had the old line doing
the host-only assignment, but this can be removed.
>
>
>> cxld->hpa_range = (struct range) {
>> .start = cfmws->base_hpa,
>> .end = cfmws->base_hpa + cfmws->window_size - 1,
>> diff --git a/drivers/cxl/core/hdm.c b/drivers/cxl/core/hdm.c
>> index 18200a8f3f72..44ff64b1a9df 100644
>> --- a/drivers/cxl/core/hdm.c
>> +++ b/drivers/cxl/core/hdm.c
>> @@ -705,9 +705,21 @@ static void cxld_set_interleave(struct cxl_decoder *cxld, u32 *ctrl)
>>
>> static void cxld_set_type(struct cxl_decoder *cxld, u32 *ctrl)
>> {
>> + bool bi = cxld->target_type == CXL_DECODER_DEVMEM &&
>> + cxld->region && cxl_root_decoder_is_bi(cxld->region->cxlrd);
>> +
>> u32p_replace_bits(ctrl,
>> !!(cxld->target_type == CXL_DECODER_HOSTONLYMEM),
>> CXL_HDM_DECODER0_CTRL_HOSTONLY);
>> + u32p_replace_bits(ctrl, bi, CXL_HDM_DECODER0_CTRL_BI);
>> +
>> + if (bi && is_endpoint_decoder(&cxld->dev)) {
>> + struct cxl_endpoint_decoder *cxled =
>> + to_cxl_endpoint_decoder(&cxld->dev);
>> +
>> + u32p_replace_bits(ctrl, cxled->pos,
>> + CXL_HDM_DECODER0_CTRL_ISP_MASK);
>> + }
>
>
>The commit message does not mention ISP programming. Can it get a
>sentence with the spec reference for ISP being required when BI is set?
It's per Table 8-123, will add a comment.
>
>snip
>
>> diff --git a/drivers/cxl/core/region.c b/drivers/cxl/core/region.c
>
>snip
>
>> @@ -1131,16 +1163,11 @@ static int cxl_rr_assign_decoder(struct cxl_port *port, struct cxl_region *cxlr,
>> }
>>
>> /*
>> - * Endpoints should already match the region type, but backstop that
>> - * assumption with an assertion. Switch-decoders change mapping-type
>> - * based on what is mapped when they are assigned to a region.
>> + * Endpoint decoders inherit their type from cxlr->type; broken
>> + * pairings were already rejected by the coherency checks in
>> + * cxl_region_attach(). Switch-decoders change mapping-type based
>> + * on what is mapped when they are assigned to a region.
>> */
>> - dev_WARN_ONCE(&cxlr->dev,
>> - port == cxled_to_port(cxled) &&
>> - cxld->target_type != cxlr->type,
>> - "%s:%s mismatch decoder type %d -> %d\n",
>> - dev_name(&cxled_to_memdev(cxled)->dev),
>> - dev_name(&cxld->dev), cxld->target_type, cxlr->type);
>> cxld->target_type = cxlr->type;
>
>
>A Type 3 endpoint decoder in an HDM-DB region now reads "accelerator" from
>decoderX.Y/target_type, and the ndctl test on the branch linked in the cover
>letter checks for that. So I take it target_type is now being used to report
>the decoder's coherency model, device-coherent vs host-only, rather than the
>Type-2/Type-3 as documented by the existing ABI.
Type2 and Type3 are names from the old pre-BI CXL (and were later renamed by
the spec, per patch 5), but the intent behind was always about coherency type.
>
>Does this need sysfs ABI update?
I don't think so, the doc already mentions:
""
The 'target_type' attribute indicates the current setting which may
dynamically change based on what memory regions are activated in this
decode hierarchy.
""
>
>cxl list doesn't show target_type, but libcxl exports it through
>cxl_decoder_get_target_type(). If that is intentional, libcxl.txt will
>need the same semantic update with the ndctl patches.
>
>
>> cxl_rr->decoder = cxld;
>> return 0;
>> @@ -1803,6 +1830,7 @@ static int cxl_region_attach_position(struct cxl_region *cxlr,
>> struct cxl_root_decoder *cxlrd = cxlr->cxlrd;
>> struct cxl_memdev *cxlmd = cxled_to_memdev(cxled);
>> struct cxl_switch_decoder *cxlsd = &cxlrd->cxlsd;
>> + enum cxl_decoder_type type = cxled->cxld.target_type;
>> struct cxl_decoder *cxld = &cxlsd->cxld;
>> int iw = cxld->interleave_ways;
>> struct cxl_port *iter;
>> @@ -1828,6 +1856,8 @@ static int cxl_region_attach_position(struct cxl_region *cxlr,
>> for (iter = cxled_to_port(cxled); !is_cxl_root(iter);
>> iter = to_cxl_port(iter->dev.parent))
>> cxl_port_detach_region(iter, cxlr, cxled);
>> + /* undo cxl_rr_assign_decoder() type inheritance */
>> + cxled->cxld.target_type = type;
>> return rc;
>> }
>>
>> @@ -2056,6 +2086,7 @@ static int cxl_region_attach(struct cxl_region *cxlr,
>> struct cxl_region_params *p = &cxlr->params;
>> struct cxl_port *ep_port, *root_port;
>> struct cxl_dport *dport;
>> + struct cxl_hdm *cxlhdm;
>> int rc = -ENXIO;
>>
>> rc = check_interleave_cap(&cxled->cxld, p->interleave_ways,
>> @@ -2105,10 +2136,39 @@ static int cxl_region_attach(struct cxl_region *cxlr,
>> return -ENXIO;
>> }
>>
>> - if (cxled->cxld.target_type != cxlr->type) {
>> - dev_dbg(&cxlr->dev, "%s:%s type mismatch: %d vs %d\n",
>> - dev_name(&cxlmd->dev), dev_name(&cxled->cxld.dev),
>> - cxled->cxld.target_type, cxlr->type);
>
>
>Should the type check stay for committed decoders? With it gone, a
>second auto-assembled decoder whose committed Target Range Type differs
>from the region's is accepted, and cxl_rr_assign_decoder() then
>overwrites its target_type while the hardware keeps the committed value.
>I think this changes in Patch 8 - which really makes patch 8 required,
>not optional. (more on that in patch 8)
Yeah I see what you mean, I will move that logic from cxl_region_attach()
in patch 8 in here.
>Maybe the CXL_DECODER_STATE_AUTO mismatch check fits better in here,
>since this patch that removes the old one?
Agreed.
>
>
>> + /*
>> + * Verify the device and HDM are capable of the region's flavor before
>> + * proceeding. The endpoint decoder's target_type is then inherited
>> + * from cxlr->type later in cxl_rr_assign_decoder().
>> + */
>> + if (cxlr->type == CXL_DECODER_DEVMEM &&
>> + cxl_root_decoder_is_bi(cxlrd) && !cxlds->bi) {
>> + dev_err(&cxlr->dev, "%s:%s BI not enabled on device\n",
>> + dev_name(&cxlmd->dev), dev_name(&cxled->cxld.dev));
>> + return -ENXIO;
>> + }
>> +
>> + if (cxled->state != CXL_DECODER_STATE_AUTO &&
>> + cxlr->type == CXL_DECODER_HOSTONLYMEM &&
>> + cxlds->type == CXL_DEVTYPE_DEVMEM) {
>> + dev_warn(&cxlr->dev, "%s:%s HDM-H requires a Type 3 device\n",
>> + dev_name(&cxlmd->dev), dev_name(&cxled->cxld.dev));
>> + return -ENXIO;
>> + }
>> +
>> + cxlhdm = dev_get_drvdata(&ep_port->dev);
>> + if (!cxlhdm)
>> + return -ENXIO;
>> + if (cxlr->type == CXL_DECODER_HOSTONLYMEM &&
>> + cxlhdm->supported_coherency == CXL_HDM_DECODER_COHERENCY_DEV) {
>> + dev_warn(&cxlr->dev, "%s:%s HDM is device-coherent only\n",
>> + dev_name(&cxlmd->dev), dev_name(&cxled->cxld.dev));
>> + return -ENXIO;
>> + }
>> + if (cxlr->type == CXL_DECODER_DEVMEM &&
>> + cxlhdm->supported_coherency == CXL_HDM_DECODER_COHERENCY_HOST) {
>> + dev_warn(&cxlr->dev, "%s:%s HDM is host-only coherent\n",
>> + dev_name(&cxlmd->dev), dev_name(&cxled->cxld.dev));
>> return -ENXIO;
>> }
>
>
>Why dev_err()/dev_warn() here? The type mismatch refusal this replaces was
>a dev_dbg(). The attach fail here would be reported as -ENXIO.
Ok, will change to dev_dbg().
>
>Tidy up opportunity, maybe move the 4 coherency checks into a helper so
>cxl_region_attach() stays cleaner.
Ok
>
>Another cleanup - the 'is this reion HDM-DB test is now coded in a few
>places, my look may have been limited:
> cxld_set_type(): target_type == DEVMEM && cxl_root_decoder_is_bi()
> back_invalidate_show(): same, plus cxlds->bi
> cxl_region_decode_commit(): cxlr->type == DEVMEM && cxl_root_decoder_is_bi()
> cxl_region_attach(): same
>
>Maybe a single cxl_region_is_hdm_db() would be clearer and cleaner.
Ok, makes sense.
Thanks,
Davidlohr
^ permalink raw reply [flat|nested] 33+ messages in thread
* [PATCH v9 05/10] cxl/hdm: Rename decoder coherency flags
2026-09-22 23:38 [PATCH v9 0/10] cxl: Support Back-Invalidate Davidlohr Bueso
` (3 preceding siblings ...)
2026-09-22 23:38 ` [PATCH v9 04/10] cxl: Add HDM-DB region creation Davidlohr Bueso
@ 2026-09-22 23:38 ` Davidlohr Bueso
2026-09-24 0:49 ` Li Ming
2026-09-22 23:38 ` [PATCH v9 06/10] cxl/region: Log the coherency model at region creation Davidlohr Bueso
` (5 subsequent siblings)
10 siblings, 1 reply; 33+ messages in thread
From: Davidlohr Bueso @ 2026-09-22 23:38 UTC (permalink / raw)
To: dave.jiang
Cc: jic23, alison.schofield, icheng, ming.li, benjamin.cheatham,
alucerop, dave, linux-cxl, Jonathan Cameron
Align with the ACPI CXL Window restriction naming and convert
CXL_DECODER_F_TYPE2/F_TYPE3 to F_DEVMEM/F_HOSTONLY. Type2 and
Type3 coherency models were named prior to Back-Invalidate.
Reviewed-by: Ben Cheatham <benjamin.cheatham@amd.com>
Reviewed-by: Dave Jiang <dave.jiang@intel.com>
Reviewed-by: Alison Schofield <alison.schofield@intel.com>
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
Signed-off-by: Davidlohr Bueso <dave@stgolabs.net>
---
drivers/cxl/acpi.c | 8 ++++----
drivers/cxl/core/port.c | 12 ++++++------
drivers/cxl/cxl.h | 4 ++--
3 files changed, 12 insertions(+), 12 deletions(-)
diff --git a/drivers/cxl/acpi.c b/drivers/cxl/acpi.c
index 2e8e31544a5f..f04ac5275af8 100644
--- a/drivers/cxl/acpi.c
+++ b/drivers/cxl/acpi.c
@@ -143,9 +143,9 @@ static unsigned long cfmws_to_decoder_flags(int restrictions)
unsigned long flags = CXL_DECODER_F_ENABLE;
if (restrictions & ACPI_CEDT_CFMWS_RESTRICT_DEVMEM)
- flags |= CXL_DECODER_F_TYPE2;
+ flags |= CXL_DECODER_F_DEVMEM;
if (restrictions & ACPI_CEDT_CFMWS_RESTRICT_HOSTONLYMEM)
- flags |= CXL_DECODER_F_TYPE3;
+ flags |= CXL_DECODER_F_HOSTONLY;
if (restrictions & ACPI_CEDT_CFMWS_RESTRICT_VOLATILE)
flags |= CXL_DECODER_F_RAM;
if (restrictions & ACPI_CEDT_CFMWS_RESTRICT_PMEM)
@@ -459,8 +459,8 @@ static int __cxl_parse_cfmws(struct acpi_cedt_cfmws *cfmws,
cxld->flags = cfmws_to_decoder_flags(cfmws->restrictions);
/* host-only wins if firmware sets both coherency restrictions */
cxld->target_type = CXL_DECODER_HOSTONLYMEM;
- if (cxld->flags & CXL_DECODER_F_TYPE2) {
- if (cxld->flags & CXL_DECODER_F_TYPE3)
+ if (cxld->flags & CXL_DECODER_F_DEVMEM) {
+ if (cxld->flags & CXL_DECODER_F_HOSTONLY)
dev_dbg(dev, "CFMWS has both HDM-H and HDM-D\n");
else
cxld->target_type = CXL_DECODER_DEVMEM;
diff --git a/drivers/cxl/core/port.c b/drivers/cxl/core/port.c
index 1d26cd7c88a1..fa1574b46b6f 100644
--- a/drivers/cxl/core/port.c
+++ b/drivers/cxl/core/port.c
@@ -129,8 +129,8 @@ static DEVICE_ATTR_RO(name)
CXL_DECODER_FLAG_ATTR(cap_pmem, CXL_DECODER_F_PMEM);
CXL_DECODER_FLAG_ATTR(cap_ram, CXL_DECODER_F_RAM);
-CXL_DECODER_FLAG_ATTR(cap_type2, CXL_DECODER_F_TYPE2);
-CXL_DECODER_FLAG_ATTR(cap_type3, CXL_DECODER_F_TYPE3);
+CXL_DECODER_FLAG_ATTR(cap_type2, CXL_DECODER_F_DEVMEM);
+CXL_DECODER_FLAG_ATTR(cap_type3, CXL_DECODER_F_HOSTONLY);
CXL_DECODER_FLAG_ATTR(locked, CXL_DECODER_F_LOCK);
CXL_DECODER_FLAG_ATTR(cap_back_invalidate, CXL_DECODER_F_BI);
@@ -365,8 +365,8 @@ static bool can_create_pmem(struct cxl_root_decoder *cxlrd)
unsigned long flags = cxlrd->cxlsd.cxld.flags;
unsigned long hdm_h, hdm_db;
- hdm_h = CXL_DECODER_F_TYPE3 | CXL_DECODER_F_PMEM;
- hdm_db = CXL_DECODER_F_TYPE2 | CXL_DECODER_F_BI | CXL_DECODER_F_PMEM;
+ hdm_h = CXL_DECODER_F_HOSTONLY | CXL_DECODER_F_PMEM;
+ hdm_db = CXL_DECODER_F_DEVMEM | CXL_DECODER_F_BI | CXL_DECODER_F_PMEM;
return (flags & hdm_h) == hdm_h || (flags & hdm_db) == hdm_db;
}
@@ -376,8 +376,8 @@ static bool can_create_ram(struct cxl_root_decoder *cxlrd)
unsigned long flags = cxlrd->cxlsd.cxld.flags;
unsigned long hdm_h, hdm_db;
- hdm_h = CXL_DECODER_F_TYPE3 | CXL_DECODER_F_RAM;
- hdm_db = CXL_DECODER_F_TYPE2 | CXL_DECODER_F_BI | CXL_DECODER_F_RAM;
+ hdm_h = CXL_DECODER_F_HOSTONLY | CXL_DECODER_F_RAM;
+ hdm_db = CXL_DECODER_F_DEVMEM | CXL_DECODER_F_BI | CXL_DECODER_F_RAM;
return (flags & hdm_h) == hdm_h || (flags & hdm_db) == hdm_db;
}
diff --git a/drivers/cxl/cxl.h b/drivers/cxl/cxl.h
index 9fb0278e2ce4..3639cd2daf0c 100644
--- a/drivers/cxl/cxl.h
+++ b/drivers/cxl/cxl.h
@@ -296,8 +296,8 @@ int cxl_dport_map_rcd_linkcap(struct pci_dev *pdev, struct cxl_dport *dport);
*/
#define CXL_DECODER_F_RAM BIT(0)
#define CXL_DECODER_F_PMEM BIT(1)
-#define CXL_DECODER_F_TYPE2 BIT(2)
-#define CXL_DECODER_F_TYPE3 BIT(3)
+#define CXL_DECODER_F_DEVMEM BIT(2)
+#define CXL_DECODER_F_HOSTONLY BIT(3)
#define CXL_DECODER_F_LOCK BIT(4)
#define CXL_DECODER_F_ENABLE BIT(5)
#define CXL_DECODER_F_NORMALIZED_ADDRESSING BIT(6)
--
2.39.5
^ permalink raw reply related [flat|nested] 33+ messages in thread* Re: [PATCH v9 05/10] cxl/hdm: Rename decoder coherency flags
2026-09-22 23:38 ` [PATCH v9 05/10] cxl/hdm: Rename decoder coherency flags Davidlohr Bueso
@ 2026-09-24 0:49 ` Li Ming
0 siblings, 0 replies; 33+ messages in thread
From: Li Ming @ 2026-09-24 0:49 UTC (permalink / raw)
To: Davidlohr Bueso, dave.jiang
Cc: jic23, alison.schofield, icheng, benjamin.cheatham, alucerop,
linux-cxl, Jonathan Cameron
On 9/23/2026 7:38 AM, Davidlohr Bueso wrote:
> Align with the ACPI CXL Window restriction naming and convert
> CXL_DECODER_F_TYPE2/F_TYPE3 to F_DEVMEM/F_HOSTONLY. Type2 and
> Type3 coherency models were named prior to Back-Invalidate.
>
> Reviewed-by: Ben Cheatham <benjamin.cheatham@amd.com>
> Reviewed-by: Dave Jiang <dave.jiang@intel.com>
> Reviewed-by: Alison Schofield <alison.schofield@intel.com>
> Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
> Signed-off-by: Davidlohr Bueso <dave@stgolabs.net>
Reviewed-by: Li Ming <ming.li@zohomail.com>
> ---
> drivers/cxl/acpi.c | 8 ++++----
> drivers/cxl/core/port.c | 12 ++++++------
> drivers/cxl/cxl.h | 4 ++--
> 3 files changed, 12 insertions(+), 12 deletions(-)
>
> diff --git a/drivers/cxl/acpi.c b/drivers/cxl/acpi.c
> index 2e8e31544a5f..f04ac5275af8 100644
> --- a/drivers/cxl/acpi.c
> +++ b/drivers/cxl/acpi.c
> @@ -143,9 +143,9 @@ static unsigned long cfmws_to_decoder_flags(int restrictions)
> unsigned long flags = CXL_DECODER_F_ENABLE;
>
> if (restrictions & ACPI_CEDT_CFMWS_RESTRICT_DEVMEM)
> - flags |= CXL_DECODER_F_TYPE2;
> + flags |= CXL_DECODER_F_DEVMEM;
> if (restrictions & ACPI_CEDT_CFMWS_RESTRICT_HOSTONLYMEM)
> - flags |= CXL_DECODER_F_TYPE3;
> + flags |= CXL_DECODER_F_HOSTONLY;
> if (restrictions & ACPI_CEDT_CFMWS_RESTRICT_VOLATILE)
> flags |= CXL_DECODER_F_RAM;
> if (restrictions & ACPI_CEDT_CFMWS_RESTRICT_PMEM)
> @@ -459,8 +459,8 @@ static int __cxl_parse_cfmws(struct acpi_cedt_cfmws *cfmws,
> cxld->flags = cfmws_to_decoder_flags(cfmws->restrictions);
> /* host-only wins if firmware sets both coherency restrictions */
> cxld->target_type = CXL_DECODER_HOSTONLYMEM;
> - if (cxld->flags & CXL_DECODER_F_TYPE2) {
> - if (cxld->flags & CXL_DECODER_F_TYPE3)
> + if (cxld->flags & CXL_DECODER_F_DEVMEM) {
> + if (cxld->flags & CXL_DECODER_F_HOSTONLY)
> dev_dbg(dev, "CFMWS has both HDM-H and HDM-D\n");
> else
> cxld->target_type = CXL_DECODER_DEVMEM;
> diff --git a/drivers/cxl/core/port.c b/drivers/cxl/core/port.c
> index 1d26cd7c88a1..fa1574b46b6f 100644
> --- a/drivers/cxl/core/port.c
> +++ b/drivers/cxl/core/port.c
> @@ -129,8 +129,8 @@ static DEVICE_ATTR_RO(name)
>
> CXL_DECODER_FLAG_ATTR(cap_pmem, CXL_DECODER_F_PMEM);
> CXL_DECODER_FLAG_ATTR(cap_ram, CXL_DECODER_F_RAM);
> -CXL_DECODER_FLAG_ATTR(cap_type2, CXL_DECODER_F_TYPE2);
> -CXL_DECODER_FLAG_ATTR(cap_type3, CXL_DECODER_F_TYPE3);
> +CXL_DECODER_FLAG_ATTR(cap_type2, CXL_DECODER_F_DEVMEM);
> +CXL_DECODER_FLAG_ATTR(cap_type3, CXL_DECODER_F_HOSTONLY);
> CXL_DECODER_FLAG_ATTR(locked, CXL_DECODER_F_LOCK);
> CXL_DECODER_FLAG_ATTR(cap_back_invalidate, CXL_DECODER_F_BI);
>
> @@ -365,8 +365,8 @@ static bool can_create_pmem(struct cxl_root_decoder *cxlrd)
> unsigned long flags = cxlrd->cxlsd.cxld.flags;
> unsigned long hdm_h, hdm_db;
>
> - hdm_h = CXL_DECODER_F_TYPE3 | CXL_DECODER_F_PMEM;
> - hdm_db = CXL_DECODER_F_TYPE2 | CXL_DECODER_F_BI | CXL_DECODER_F_PMEM;
> + hdm_h = CXL_DECODER_F_HOSTONLY | CXL_DECODER_F_PMEM;
> + hdm_db = CXL_DECODER_F_DEVMEM | CXL_DECODER_F_BI | CXL_DECODER_F_PMEM;
>
> return (flags & hdm_h) == hdm_h || (flags & hdm_db) == hdm_db;
> }
> @@ -376,8 +376,8 @@ static bool can_create_ram(struct cxl_root_decoder *cxlrd)
> unsigned long flags = cxlrd->cxlsd.cxld.flags;
> unsigned long hdm_h, hdm_db;
>
> - hdm_h = CXL_DECODER_F_TYPE3 | CXL_DECODER_F_RAM;
> - hdm_db = CXL_DECODER_F_TYPE2 | CXL_DECODER_F_BI | CXL_DECODER_F_RAM;
> + hdm_h = CXL_DECODER_F_HOSTONLY | CXL_DECODER_F_RAM;
> + hdm_db = CXL_DECODER_F_DEVMEM | CXL_DECODER_F_BI | CXL_DECODER_F_RAM;
>
> return (flags & hdm_h) == hdm_h || (flags & hdm_db) == hdm_db;
> }
> diff --git a/drivers/cxl/cxl.h b/drivers/cxl/cxl.h
> index 9fb0278e2ce4..3639cd2daf0c 100644
> --- a/drivers/cxl/cxl.h
> +++ b/drivers/cxl/cxl.h
> @@ -296,8 +296,8 @@ int cxl_dport_map_rcd_linkcap(struct pci_dev *pdev, struct cxl_dport *dport);
> */
> #define CXL_DECODER_F_RAM BIT(0)
> #define CXL_DECODER_F_PMEM BIT(1)
> -#define CXL_DECODER_F_TYPE2 BIT(2)
> -#define CXL_DECODER_F_TYPE3 BIT(3)
> +#define CXL_DECODER_F_DEVMEM BIT(2)
> +#define CXL_DECODER_F_HOSTONLY BIT(3)
> #define CXL_DECODER_F_LOCK BIT(4)
> #define CXL_DECODER_F_ENABLE BIT(5)
> #define CXL_DECODER_F_NORMALIZED_ADDRESSING BIT(6)
^ permalink raw reply [flat|nested] 33+ messages in thread
* [PATCH v9 06/10] cxl/region: Log the coherency model at region creation
2026-09-22 23:38 [PATCH v9 0/10] cxl: Support Back-Invalidate Davidlohr Bueso
` (4 preceding siblings ...)
2026-09-22 23:38 ` [PATCH v9 05/10] cxl/hdm: Rename decoder coherency flags Davidlohr Bueso
@ 2026-09-22 23:38 ` Davidlohr Bueso
2026-09-24 0:49 ` Li Ming
2026-09-30 23:10 ` Alison Schofield
2026-09-22 23:38 ` [PATCH v9 07/10] cxl/pci: Split BI capability probe from setup Davidlohr Bueso
` (4 subsequent siblings)
10 siblings, 2 replies; 33+ messages in thread
From: Davidlohr Bueso @ 2026-09-22 23:38 UTC (permalink / raw)
To: dave.jiang
Cc: jic23, alison.schofield, icheng, ming.li, benjamin.cheatham,
alucerop, dave, linux-cxl, Jonathan Cameron
Region assembly emits plenty of debug information - resources,
interleave geometry, target placement - but not the coherency model
the region operates in.
Reviewed-by: Dave Jiang <dave.jiang@intel.com>
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
Signed-off-by: Davidlohr Bueso <dave@stgolabs.net>
---
drivers/cxl/core/region.c | 13 +++++++++++--
1 file changed, 11 insertions(+), 2 deletions(-)
diff --git a/drivers/cxl/core/region.c b/drivers/cxl/core/region.c
index 0a204aea64b5..f322a8a4a66d 100644
--- a/drivers/cxl/core/region.c
+++ b/drivers/cxl/core/region.c
@@ -1861,6 +1861,13 @@ static int cxl_region_attach_position(struct cxl_region *cxlr,
return rc;
}
+static const char *cxl_coherency_name(enum cxl_decoder_type type, bool bi)
+{
+ if (type == CXL_DECODER_HOSTONLYMEM)
+ return "HDM-H";
+ return bi ? "HDM-DB" : "HDM-D";
+}
+
static int cxl_region_attach_auto(struct cxl_region *cxlr,
struct cxl_endpoint_decoder *cxled, int pos)
{
@@ -2823,8 +2830,10 @@ static struct cxl_region *devm_cxl_add_region(struct cxl_root_decoder *cxlrd,
return ERR_PTR(rc);
}
- dev_dbg(port->uport_dev, "%s: created %s\n",
- dev_name(&cxlrd->cxlsd.cxld.dev), dev_name(dev));
+ dev_dbg(port->uport_dev, "%s: created %s %s\n",
+ dev_name(&cxlrd->cxlsd.cxld.dev),
+ cxl_coherency_name(cxlr->type, cxl_root_decoder_is_bi(cxlrd)),
+ dev_name(dev));
return cxlr;
err:
put_device(dev);
--
2.39.5
^ permalink raw reply related [flat|nested] 33+ messages in thread* Re: [PATCH v9 06/10] cxl/region: Log the coherency model at region creation
2026-09-22 23:38 ` [PATCH v9 06/10] cxl/region: Log the coherency model at region creation Davidlohr Bueso
@ 2026-09-24 0:49 ` Li Ming
2026-09-30 23:10 ` Alison Schofield
1 sibling, 0 replies; 33+ messages in thread
From: Li Ming @ 2026-09-24 0:49 UTC (permalink / raw)
To: Davidlohr Bueso, dave.jiang
Cc: jic23, alison.schofield, icheng, benjamin.cheatham, alucerop,
linux-cxl, Jonathan Cameron
On 9/23/2026 7:38 AM, Davidlohr Bueso wrote:
> Region assembly emits plenty of debug information - resources,
> interleave geometry, target placement - but not the coherency model
> the region operates in.
>
> Reviewed-by: Dave Jiang <dave.jiang@intel.com>
> Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
> Signed-off-by: Davidlohr Bueso <dave@stgolabs.net>
Reviewed-by: Li Ming <ming.li@zohomail.com>
> ---
> drivers/cxl/core/region.c | 13 +++++++++++--
> 1 file changed, 11 insertions(+), 2 deletions(-)
>
> diff --git a/drivers/cxl/core/region.c b/drivers/cxl/core/region.c
> index 0a204aea64b5..f322a8a4a66d 100644
> --- a/drivers/cxl/core/region.c
> +++ b/drivers/cxl/core/region.c
> @@ -1861,6 +1861,13 @@ static int cxl_region_attach_position(struct cxl_region *cxlr,
> return rc;
> }
>
> +static const char *cxl_coherency_name(enum cxl_decoder_type type, bool bi)
> +{
> + if (type == CXL_DECODER_HOSTONLYMEM)
> + return "HDM-H";
> + return bi ? "HDM-DB" : "HDM-D";
> +}
> +
> static int cxl_region_attach_auto(struct cxl_region *cxlr,
> struct cxl_endpoint_decoder *cxled, int pos)
> {
> @@ -2823,8 +2830,10 @@ static struct cxl_region *devm_cxl_add_region(struct cxl_root_decoder *cxlrd,
> return ERR_PTR(rc);
> }
>
> - dev_dbg(port->uport_dev, "%s: created %s\n",
> - dev_name(&cxlrd->cxlsd.cxld.dev), dev_name(dev));
> + dev_dbg(port->uport_dev, "%s: created %s %s\n",
> + dev_name(&cxlrd->cxlsd.cxld.dev),
> + cxl_coherency_name(cxlr->type, cxl_root_decoder_is_bi(cxlrd)),
> + dev_name(dev));
> return cxlr;
> err:
> put_device(dev);
^ permalink raw reply [flat|nested] 33+ messages in thread* Re: [PATCH v9 06/10] cxl/region: Log the coherency model at region creation
2026-09-22 23:38 ` [PATCH v9 06/10] cxl/region: Log the coherency model at region creation Davidlohr Bueso
2026-09-24 0:49 ` Li Ming
@ 2026-09-30 23:10 ` Alison Schofield
1 sibling, 0 replies; 33+ messages in thread
From: Alison Schofield @ 2026-09-30 23:10 UTC (permalink / raw)
To: Davidlohr Bueso
Cc: dave.jiang, jic23, icheng, ming.li, benjamin.cheatham, alucerop,
linux-cxl, Jonathan Cameron
On Tue, Sep 22, 2026 at 04:38:43PM -0700, Davidlohr Bueso wrote:
> Region assembly emits plenty of debug information - resources,
> interleave geometry, target placement - but not the coherency model
> the region operates in.
Reviewed-by: Alison Schofield <alison.schofield@intel.com>
^ permalink raw reply [flat|nested] 33+ messages in thread
* [PATCH v9 07/10] cxl/pci: Split BI capability probe from setup
2026-09-22 23:38 [PATCH v9 0/10] cxl: Support Back-Invalidate Davidlohr Bueso
` (5 preceding siblings ...)
2026-09-22 23:38 ` [PATCH v9 06/10] cxl/region: Log the coherency model at region creation Davidlohr Bueso
@ 2026-09-22 23:38 ` Davidlohr Bueso
2026-09-24 0:49 ` Li Ming
2026-09-30 23:09 ` Alison Schofield
2026-09-22 23:38 ` [PATCH v9 08/10] cxl: Allow auto-committed BI hdm decoders Davidlohr Bueso
` (3 subsequent siblings)
10 siblings, 2 replies; 33+ messages in thread
From: Davidlohr Bueso @ 2026-09-22 23:38 UTC (permalink / raw)
To: dave.jiang
Cc: jic23, alison.schofield, icheng, ming.li, benjamin.cheatham,
alucerop, dave, linux-cxl, Jonathan Cameron
Decouple the topology read-only capability verification phase from
cxl_bi_setup() into cxl_bi_probe_capable(), recording the result in
cxlds->bi_capable.
This allows further dealing with auto-committed BI hdm decoders;
having such knowledge upon decoder enumeration time.
No functional change intended.
Reviewed-by: Dave Jiang <dave.jiang@intel.com>
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
Signed-off-by: Davidlohr Bueso <dave@stgolabs.net>
---
drivers/cxl/core/pci.c | 40 +++++++++++++++++++++++++++++-----------
drivers/cxl/cxl.h | 1 +
drivers/cxl/port.c | 1 +
include/cxl/cxl.h | 2 ++
4 files changed, 33 insertions(+), 11 deletions(-)
diff --git a/drivers/cxl/core/pci.c b/drivers/cxl/core/pci.c
index f3ca9861a7ba..9bf05820c142 100644
--- a/drivers/cxl/core/pci.c
+++ b/drivers/cxl/core/pci.c
@@ -1332,34 +1332,38 @@ void cxl_bi_reset_detected(struct cxl_port *endpoint)
}
EXPORT_SYMBOL_NS_GPL(cxl_bi_reset_detected, "CXL");
-int cxl_bi_setup(struct cxl_port *endpoint)
+/*
+ * Verify the device and every component in the path up to the root
+ * are BI capable.
+ */
+void cxl_bi_probe_capable(struct cxl_port *endpoint)
{
struct cxl_memdev *cxlmd = to_cxl_memdev(endpoint->uport_dev);
struct cxl_dev_state *cxlds = cxlmd->cxlds;
- struct cxl_dport *dport = endpoint->parent_dport;
struct cxl_dport *dport_iter;
struct cxl_port *port_iter;
- int rc;
+
+ cxlds->bi_capable = false;
if (!dev_is_pci(cxlds->dev))
- return 0;
+ return;
/* BI is VH-only */
if (cxlds->rcd)
- return 0;
+ return;
if (!cxl_is_bi_capable(to_pci_dev(cxlds->dev),
endpoint->regs.bi_decoder))
- return 0;
+ return;
- port_iter = dport->port;
- dport_iter = dport;
+ dport_iter = endpoint->parent_dport;
+ port_iter = dport_iter->port;
while (!is_cxl_root(port_iter)) {
/* check rp, dsp */
if (!cxl_is_bi_capable(to_pci_dev(dport_iter->dport_dev),
dport_iter->regs.bi_decoder)) {
dev_dbg(cxlds->dev, "BI not supported by topology\n");
- return 0;
+ return;
}
/* check usp */
@@ -1370,13 +1374,13 @@ int cxl_bi_setup(struct cxl_port *endpoint)
NULL)) {
dev_dbg(cxlds->dev,
"BI not supported by USP\n");
- return 0;
+ return;
}
if (port_iter->reg_map.component_map.bi_rt.valid &&
!port_iter->regs.bi_rt) {
dev_dbg(cxlds->dev,
"BI RT advertised but unmapped\n");
- return 0;
+ return;
}
}
@@ -1384,6 +1388,20 @@ int cxl_bi_setup(struct cxl_port *endpoint)
port_iter = dport_iter->port;
}
+ cxlds->bi_capable = true;
+}
+EXPORT_SYMBOL_NS_GPL(cxl_bi_probe_capable, "CXL");
+
+int cxl_bi_setup(struct cxl_port *endpoint)
+{
+ struct cxl_memdev *cxlmd = to_cxl_memdev(endpoint->uport_dev);
+ struct cxl_dport *dport = endpoint->parent_dport;
+ struct cxl_dev_state *cxlds = cxlmd->cxlds;
+ int rc;
+
+ if (!cxlds->bi_capable)
+ return 0;
+
rc = cxl_bi_enable_path(cxlds, dport->port, dport);
if (rc)
return rc;
diff --git a/drivers/cxl/cxl.h b/drivers/cxl/cxl.h
index 3639cd2daf0c..5ab2f9c39f5c 100644
--- a/drivers/cxl/cxl.h
+++ b/drivers/cxl/cxl.h
@@ -962,6 +962,7 @@ void cxl_coordinates_combine(struct access_coordinate *out,
struct access_coordinate *c2);
bool cxl_endpoint_decoder_reset_detected(struct cxl_port *port);
+void cxl_bi_probe_capable(struct cxl_port *endpoint);
int cxl_bi_setup(struct cxl_port *endpoint);
void cxl_bi_reset_detected(struct cxl_port *endpoint);
struct cxl_dport *devm_cxl_add_dport_by_dev(struct cxl_port *port,
diff --git a/drivers/cxl/port.c b/drivers/cxl/port.c
index eab935eca9b1..8e58adbc2e29 100644
--- a/drivers/cxl/port.c
+++ b/drivers/cxl/port.c
@@ -170,6 +170,7 @@ static int cxl_endpoint_port_probe(struct cxl_port *port)
cxl_endpoint_parse_cdat(port);
devm_cxl_port_bi_setup(port);
+ cxl_bi_probe_capable(port);
get_device(&cxlmd->dev);
rc = devm_add_action_or_reset(&port->dev, schedule_detach, cxlmd);
diff --git a/include/cxl/cxl.h b/include/cxl/cxl.h
index f9e9058a85da..d0f03f305072 100644
--- a/include/cxl/cxl.h
+++ b/include/cxl/cxl.h
@@ -170,6 +170,7 @@ struct cxl_dpa_partition {
* @cxl_dvsec: Offset to the PCIe device DVSEC
* @rcd: operating in RCD mode (CXL 3.0 9.11.8 CXL Devices Attached to an RCH)
* @bi: device is BI (Back-Invalidate) enabled
+ * @bi_capable: device and topology path are BI capable
* @media_ready: Indicate whether the device media is usable
* @dpa_res: Overall DPA resource tree for the device
* @part: DPA partition array
@@ -190,6 +191,7 @@ struct cxl_dev_state {
int cxl_dvsec;
bool rcd;
bool bi;
+ bool bi_capable;
bool media_ready;
struct resource dpa_res;
struct cxl_dpa_partition part[CXL_NR_PARTITIONS_MAX];
--
2.39.5
^ permalink raw reply related [flat|nested] 33+ messages in thread* Re: [PATCH v9 07/10] cxl/pci: Split BI capability probe from setup
2026-09-22 23:38 ` [PATCH v9 07/10] cxl/pci: Split BI capability probe from setup Davidlohr Bueso
@ 2026-09-24 0:49 ` Li Ming
2026-09-30 23:09 ` Alison Schofield
1 sibling, 0 replies; 33+ messages in thread
From: Li Ming @ 2026-09-24 0:49 UTC (permalink / raw)
To: Davidlohr Bueso, dave.jiang
Cc: jic23, alison.schofield, icheng, benjamin.cheatham, alucerop,
linux-cxl, Jonathan Cameron
On 9/23/2026 7:38 AM, Davidlohr Bueso wrote:
> Decouple the topology read-only capability verification phase from
> cxl_bi_setup() into cxl_bi_probe_capable(), recording the result in
> cxlds->bi_capable.
>
> This allows further dealing with auto-committed BI hdm decoders;
> having such knowledge upon decoder enumeration time.
>
> No functional change intended.
>
> Reviewed-by: Dave Jiang <dave.jiang@intel.com>
> Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
> Signed-off-by: Davidlohr Bueso <dave@stgolabs.net>
Reviewed-by: Li Ming <ming.li@zohomail.com>
> ---
> drivers/cxl/core/pci.c | 40 +++++++++++++++++++++++++++++-----------
> drivers/cxl/cxl.h | 1 +
> drivers/cxl/port.c | 1 +
> include/cxl/cxl.h | 2 ++
> 4 files changed, 33 insertions(+), 11 deletions(-)
>
> diff --git a/drivers/cxl/core/pci.c b/drivers/cxl/core/pci.c
> index f3ca9861a7ba..9bf05820c142 100644
> --- a/drivers/cxl/core/pci.c
> +++ b/drivers/cxl/core/pci.c
> @@ -1332,34 +1332,38 @@ void cxl_bi_reset_detected(struct cxl_port *endpoint)
> }
> EXPORT_SYMBOL_NS_GPL(cxl_bi_reset_detected, "CXL");
>
> -int cxl_bi_setup(struct cxl_port *endpoint)
> +/*
> + * Verify the device and every component in the path up to the root
> + * are BI capable.
> + */
> +void cxl_bi_probe_capable(struct cxl_port *endpoint)
> {
> struct cxl_memdev *cxlmd = to_cxl_memdev(endpoint->uport_dev);
> struct cxl_dev_state *cxlds = cxlmd->cxlds;
> - struct cxl_dport *dport = endpoint->parent_dport;
> struct cxl_dport *dport_iter;
> struct cxl_port *port_iter;
> - int rc;
> +
> + cxlds->bi_capable = false;
>
> if (!dev_is_pci(cxlds->dev))
> - return 0;
> + return;
>
> /* BI is VH-only */
> if (cxlds->rcd)
> - return 0;
> + return;
>
> if (!cxl_is_bi_capable(to_pci_dev(cxlds->dev),
> endpoint->regs.bi_decoder))
> - return 0;
> + return;
>
> - port_iter = dport->port;
> - dport_iter = dport;
> + dport_iter = endpoint->parent_dport;
> + port_iter = dport_iter->port;
> while (!is_cxl_root(port_iter)) {
> /* check rp, dsp */
> if (!cxl_is_bi_capable(to_pci_dev(dport_iter->dport_dev),
> dport_iter->regs.bi_decoder)) {
> dev_dbg(cxlds->dev, "BI not supported by topology\n");
> - return 0;
> + return;
> }
>
> /* check usp */
> @@ -1370,13 +1374,13 @@ int cxl_bi_setup(struct cxl_port *endpoint)
> NULL)) {
> dev_dbg(cxlds->dev,
> "BI not supported by USP\n");
> - return 0;
> + return;
> }
> if (port_iter->reg_map.component_map.bi_rt.valid &&
> !port_iter->regs.bi_rt) {
> dev_dbg(cxlds->dev,
> "BI RT advertised but unmapped\n");
> - return 0;
> + return;
> }
> }
>
> @@ -1384,6 +1388,20 @@ int cxl_bi_setup(struct cxl_port *endpoint)
> port_iter = dport_iter->port;
> }
>
> + cxlds->bi_capable = true;
> +}
> +EXPORT_SYMBOL_NS_GPL(cxl_bi_probe_capable, "CXL");
> +
> +int cxl_bi_setup(struct cxl_port *endpoint)
> +{
> + struct cxl_memdev *cxlmd = to_cxl_memdev(endpoint->uport_dev);
> + struct cxl_dport *dport = endpoint->parent_dport;
> + struct cxl_dev_state *cxlds = cxlmd->cxlds;
> + int rc;
> +
> + if (!cxlds->bi_capable)
> + return 0;
> +
> rc = cxl_bi_enable_path(cxlds, dport->port, dport);
> if (rc)
> return rc;
> diff --git a/drivers/cxl/cxl.h b/drivers/cxl/cxl.h
> index 3639cd2daf0c..5ab2f9c39f5c 100644
> --- a/drivers/cxl/cxl.h
> +++ b/drivers/cxl/cxl.h
> @@ -962,6 +962,7 @@ void cxl_coordinates_combine(struct access_coordinate *out,
> struct access_coordinate *c2);
>
> bool cxl_endpoint_decoder_reset_detected(struct cxl_port *port);
> +void cxl_bi_probe_capable(struct cxl_port *endpoint);
> int cxl_bi_setup(struct cxl_port *endpoint);
> void cxl_bi_reset_detected(struct cxl_port *endpoint);
> struct cxl_dport *devm_cxl_add_dport_by_dev(struct cxl_port *port,
> diff --git a/drivers/cxl/port.c b/drivers/cxl/port.c
> index eab935eca9b1..8e58adbc2e29 100644
> --- a/drivers/cxl/port.c
> +++ b/drivers/cxl/port.c
> @@ -170,6 +170,7 @@ static int cxl_endpoint_port_probe(struct cxl_port *port)
> cxl_endpoint_parse_cdat(port);
>
> devm_cxl_port_bi_setup(port);
> + cxl_bi_probe_capable(port);
>
> get_device(&cxlmd->dev);
> rc = devm_add_action_or_reset(&port->dev, schedule_detach, cxlmd);
> diff --git a/include/cxl/cxl.h b/include/cxl/cxl.h
> index f9e9058a85da..d0f03f305072 100644
> --- a/include/cxl/cxl.h
> +++ b/include/cxl/cxl.h
> @@ -170,6 +170,7 @@ struct cxl_dpa_partition {
> * @cxl_dvsec: Offset to the PCIe device DVSEC
> * @rcd: operating in RCD mode (CXL 3.0 9.11.8 CXL Devices Attached to an RCH)
> * @bi: device is BI (Back-Invalidate) enabled
> + * @bi_capable: device and topology path are BI capable
> * @media_ready: Indicate whether the device media is usable
> * @dpa_res: Overall DPA resource tree for the device
> * @part: DPA partition array
> @@ -190,6 +191,7 @@ struct cxl_dev_state {
> int cxl_dvsec;
> bool rcd;
> bool bi;
> + bool bi_capable;
> bool media_ready;
> struct resource dpa_res;
> struct cxl_dpa_partition part[CXL_NR_PARTITIONS_MAX];
^ permalink raw reply [flat|nested] 33+ messages in thread* Re: [PATCH v9 07/10] cxl/pci: Split BI capability probe from setup
2026-09-22 23:38 ` [PATCH v9 07/10] cxl/pci: Split BI capability probe from setup Davidlohr Bueso
2026-09-24 0:49 ` Li Ming
@ 2026-09-30 23:09 ` Alison Schofield
1 sibling, 0 replies; 33+ messages in thread
From: Alison Schofield @ 2026-09-30 23:09 UTC (permalink / raw)
To: Davidlohr Bueso
Cc: dave.jiang, jic23, icheng, ming.li, benjamin.cheatham, alucerop,
linux-cxl, Jonathan Cameron
On Tue, Sep 22, 2026 at 04:38:44PM -0700, Davidlohr Bueso wrote:
> Decouple the topology read-only capability verification phase from
> cxl_bi_setup() into cxl_bi_probe_capable(), recording the result in
> cxlds->bi_capable.
>
> This allows further dealing with auto-committed BI hdm decoders;
> having such knowledge upon decoder enumeration time.
>
> No functional change intended.
Reviewed-by: Alison Schofield <alison.schofield@intel.com>
^ permalink raw reply [flat|nested] 33+ messages in thread
* [PATCH v9 08/10] cxl: Allow auto-committed BI hdm decoders
2026-09-22 23:38 [PATCH v9 0/10] cxl: Support Back-Invalidate Davidlohr Bueso
` (6 preceding siblings ...)
2026-09-22 23:38 ` [PATCH v9 07/10] cxl/pci: Split BI capability probe from setup Davidlohr Bueso
@ 2026-09-22 23:38 ` Davidlohr Bueso
2026-10-01 3:49 ` Alison Schofield
2026-09-22 23:38 ` [PATCH v9 09/10] cxl/test: Add mock BI topology support Davidlohr Bueso
` (2 subsequent siblings)
10 siblings, 1 reply; 33+ messages in thread
From: Davidlohr Bueso @ 2026-09-22 23:38 UTC (permalink / raw)
To: dave.jiang
Cc: jic23, alison.schofield, icheng, ming.li, benjamin.cheatham,
alucerop, dave, linux-cxl, Jonathan Cameron
An HDM decoder that firmware committed with the BI bit set was
refused outright. Accept it, and refuse only what is broken - BI
paired with a host-only target range type, or for an endpoint
decoder BI on a path that cannot route BISnp.
Taking over such a path must not re-program what firmware already
did, and a dport's committed state does not say which device it was
committed for, only that firmware committed for something below it.
Two endpoints sharing a level show why that is not enough to go on.
RP Forward set, but for mem0
|
USP
/ \
DSP0 DSP1
| |
mem0 mem1
BI Enable BI Enable
set clear
Adoption is therefore seeded from the endpoint's own BI Enable, since
only a device firmware itself enabled can have had its BI-ID
accounted for above, and carried up the walk. A level is taken as
found, with no register write and no commit, when nothing this driver
routed sits below it, the value this driver would write is already
there, and both the decoder and the switch's route table are
committed - the table carries a commit of its own that firmware may
not have performed. The first level to fall short ends adoption, and
from it up the driver programs and commits as for any other device.
RP Forward, written but never committed
|
USP0 route table committed by the driver
|
DSP0 Forward, written and committed
| ..... adoption ended here .....
USP1 route table already committed
|
DSP1 Enable, taken as found
|
mem0 BI Enable already set by firmware
An adopted level owes no commit either way. Table 8-152 and Table
8-156 owe one per new BI device enabled below the port, and an
adopted level is not seeing one.
The endpoint is adopted the same way. BI Enable already set was an
error before, since only this driver set it; it is now kept and
cxlds->bi taken up without a write.
On failure the walk unwinds only the levels it did not adopt. Since
adoption is a prefix, the unwind drops the refcount below the first
non-adopted level and leaves those registers alone. Teardown is
unchanged and still deallocates every level.
A committed decoder is refused when its coherency model does not fit
where it lands. The window must permit it, and unlike a decoder this
driver programs it cannot inherit a region's model, so its committed
Target Range Type and BI bit must already match. Assembly never
validated these bits, so setting them wrongly costs a platform the
auto-assembly of the affected decoders.
Reviewed-by: Dave Jiang <dave.jiang@intel.com>
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
Signed-off-by: Davidlohr Bueso <dave@stgolabs.net>
---
drivers/cxl/core/hdm.c | 27 ++++++++----
drivers/cxl/core/pci.c | 90 +++++++++++++++++++++++++++++++++------
drivers/cxl/core/region.c | 48 +++++++++++++++++++++
3 files changed, 143 insertions(+), 22 deletions(-)
diff --git a/drivers/cxl/core/hdm.c b/drivers/cxl/core/hdm.c
index 44ff64b1a9df..0d9ef785e25f 100644
--- a/drivers/cxl/core/hdm.c
+++ b/drivers/cxl/core/hdm.c
@@ -1057,13 +1057,23 @@ static int init_hdm_decoder(struct cxl_port *port, struct cxl_decoder *cxld,
else
cxld->target_type = CXL_DECODER_DEVMEM;
- /*
- * Autocommit BI-enabled decoders is not supported. A path
- * firmware enabled is not taken over, so nothing verifies
- * it can route BISnp.
- */
- if (FIELD_GET(CXL_HDM_DECODER0_CTRL_BI, ctrl))
- return -ENXIO;
+ if (FIELD_GET(CXL_HDM_DECODER0_CTRL_BI, ctrl)) {
+ struct cxl_dev_state *cxlds = cxled ?
+ cxled_to_memdev(cxled)->cxlds : NULL;
+
+ if (cxld->target_type == CXL_DECODER_HOSTONLYMEM) {
+ dev_warn(&port->dev,
+ "decoder%d.%d: BI with host-only\n",
+ port->id, cxld->id);
+ return -ENXIO;
+ }
+ if (cxlds && !cxlds->bi_capable) {
+ dev_warn(&port->dev,
+ "decoder%d.%d: path not BI capable\n",
+ port->id, cxld->id);
+ return -ENXIO;
+ }
+ }
guard(rwsem_write)(&cxl_rwsem.region);
if (cxld->id != cxl_num_decoders_committed(port)) {
@@ -1310,7 +1320,8 @@ int devm_cxl_endpoint_decoders_setup(struct cxl_port *port)
* Between the port's HDM state and its decoders: devres,
* unwinding in reverse, brings BI down only after the decoders
* quiesce, while its slow walk still precedes the HDM state
- * free.
+ * free. Must also precede region discovery, where HDM-DB
+ * assembly requires cxlds->bi.
*/
rc = cxl_bi_setup(port);
if (rc)
diff --git a/drivers/cxl/core/pci.c b/drivers/cxl/core/pci.c
index 9bf05820c142..f83f5d1d8255 100644
--- a/drivers/cxl/core/pci.c
+++ b/drivers/cxl/core/pci.c
@@ -1087,6 +1087,29 @@ static int __cxl_bi_commit_decoder(struct device *dev, void __iomem *bi)
scale, base);
}
+/* Committed, or no explicit commit required */
+static bool cxl_bi_decoder_committed(void __iomem *bi)
+{
+ u32 caps = readl(bi + CXL_BI_DECODER_CAPS_OFFSET);
+ u32 sts = readl(bi + CXL_BI_DECODER_STATUS_OFFSET);
+
+ if (!FIELD_GET(CXL_BI_DECODER_CAPS_EXPLICIT_COMMIT_REQ, caps))
+ return true;
+
+ return FIELD_GET(CXL_BI_DECODER_STATUS_BI_COMMITTED, sts);
+}
+
+static bool cxl_bi_rt_committed(void __iomem *bi)
+{
+ u32 caps = readl(bi + CXL_BI_RT_CAPS_OFFSET);
+ u32 sts = readl(bi + CXL_BI_RT_STATUS_OFFSET);
+
+ if (!FIELD_GET(CXL_BI_RT_CAPS_EXPLICIT_COMMIT_REQ, caps))
+ return true;
+
+ return FIELD_GET(CXL_BI_RT_STATUS_BI_COMMITTED, sts);
+}
+
static int cxl_bi_commit_dport(struct cxl_dport *dport)
{
struct cxl_port *port = dport->port;
@@ -1114,7 +1137,8 @@ static int cxl_bi_commit_dport(struct cxl_dport *dport)
* @direct says the device is connected directly to this dport, which
* takes BI Enable; every dport above it takes BI Forward.
*/
-static int cxl_bi_ctrl_dport_enable(struct cxl_dport *dport, bool direct)
+static int cxl_bi_ctrl_dport_enable(struct cxl_dport *dport, bool direct,
+ bool *adopt)
{
struct cxl_port *port = dport->port;
u32 ctrl, value, set, clr;
@@ -1135,6 +1159,20 @@ static int cxl_bi_ctrl_dport_enable(struct cxl_dport *dport, bool direct)
CXL_BI_DECODER_CTRL_BI_ENABLE;
value = (ctrl | set) & ~clr;
+
+ /*
+ * Adopt this level as firmware left it: nothing below it that
+ * firmware did not account for, nothing this driver has routed
+ * through it, the value this walk would write already in place,
+ * and the decoder (and route table) committed.
+ */
+ if (*adopt && !dport->nr_bi && value == ctrl &&
+ cxl_bi_decoder_committed(bi) &&
+ (!port->regs.bi_rt || cxl_bi_rt_committed(port->regs.bi_rt)))
+ goto done;
+ /* firmware did not bring this level up, so nothing above may */
+ *adopt = false;
+
if (value != ctrl)
writel(value, bi + CXL_BI_DECODER_CTRL_OFFSET);
@@ -1148,6 +1186,7 @@ static int cxl_bi_ctrl_dport_enable(struct cxl_dport *dport, bool direct)
}
return rc;
}
+done:
dport->nr_bi++;
return 0;
@@ -1156,9 +1195,11 @@ static int cxl_bi_ctrl_dport_enable(struct cxl_dport *dport, bool direct)
/*
* Dealloc BI-ID changes in the given level of the topology. Called
* once per endpoint that enabled this level: on teardown, or to
- * unwind a path that failed partway up.
+ * unwind a path that failed partway up. @clear false drops the
+ * reference without touching the register, for a level this driver
+ * adopted and therefore never wrote.
*/
-static int cxl_bi_ctrl_dport_disable(struct cxl_dport *dport)
+static int cxl_bi_ctrl_dport_disable(struct cxl_dport *dport, bool clear)
{
struct cxl_port *port = dport->port;
void __iomem *bi;
@@ -1177,6 +1218,9 @@ static int cxl_bi_ctrl_dport_disable(struct cxl_dport *dport)
if (--dport->nr_bi > 0)
return 0;
+ if (!clear)
+ return 0;
+
ctrl = readl(bi + CXL_BI_DECODER_CTRL_OFFSET);
writel(ctrl & ~(CXL_BI_DECODER_CTRL_BI_FW |
CXL_BI_DECODER_CTRL_BI_ENABLE),
@@ -1198,11 +1242,10 @@ static int __cxl_bi_ctrl_endpoint(struct cxl_dev_state *cxlds, bool enable)
if (enable) {
if (FIELD_GET(CXL_BI_DECODER_CTRL_BI_ENABLE, ctrl)) {
- if (cxlds->bi)
- return 0;
- dev_err(cxlds->dev,
- "BI already enabled in hardware\n");
- return -EBUSY;
+ if (!cxlds->bi) /* adopt firmware enabled */
+ dev_dbg(cxlds->dev,
+ "adopting firmware-enabled BI\n");
+ goto done;
}
} else {
if (!FIELD_GET(CXL_BI_DECODER_CTRL_BI_ENABLE, ctrl)) {
@@ -1216,10 +1259,10 @@ static int __cxl_bi_ctrl_endpoint(struct cxl_dev_state *cxlds, bool enable)
FIELD_MODIFY(CXL_BI_DECODER_CTRL_BI_ENABLE, &ctrl, enable);
writel(ctrl, bi + CXL_BI_DECODER_CTRL_OFFSET);
- cxlds->bi = enable;
dev_dbg(cxlds->dev, "BI requests %s\n", str_enabled_disabled(enable));
-
+done:
+ cxlds->bi = enable;
return 0;
}
@@ -1257,7 +1300,7 @@ static void cxl_bi_dealloc(void *data)
dport_iter = endpoint->parent_dport;
port_iter = dport_iter->port;
while (!is_cxl_root(port_iter)) {
- int rc = cxl_bi_ctrl_dport_disable(dport_iter);
+ int rc = cxl_bi_ctrl_dport_disable(dport_iter, true);
/* best effort */
if (rc)
@@ -1276,17 +1319,33 @@ static void cxl_bi_dealloc(void *data)
static int cxl_bi_enable_path(struct cxl_dev_state *cxlds,
struct cxl_port *port, struct cxl_dport *dport)
{
- struct cxl_dport *dport_iter, *failed;
+ struct cxl_dport *dport_iter, *failed, *programmed = NULL;
+ struct cxl_port *endpoint = cxlds->cxlmd->endpoint;
struct cxl_port *port_iter;
+ bool adopt, clear;
int rc;
+ /*
+ * Adoption is a path property, not a per-dport one: only when
+ * firmware enabled the device itself does a dport's committed
+ * state cover this device's BI-ID.
+ */
+ adopt = FIELD_GET(CXL_BI_DECODER_CTRL_BI_ENABLE,
+ readl(endpoint->regs.bi_decoder +
+ CXL_BI_DECODER_CTRL_OFFSET));
+
port_iter = port;
dport_iter = dport;
while (!is_cxl_root(port_iter)) {
- rc = cxl_bi_ctrl_dport_enable(dport_iter, dport_iter == dport);
+ rc = cxl_bi_ctrl_dport_enable(dport_iter, dport_iter == dport,
+ &adopt);
if (rc)
goto err_rollback;
+ /* adoption is a prefix: this is the first level not adopted */
+ if (!adopt && !programmed)
+ programmed = dport_iter;
+
dport_iter = port_iter->parent_dport;
port_iter = dport_iter->port;
}
@@ -1301,8 +1360,11 @@ static int cxl_bi_enable_path(struct cxl_dev_state *cxlds,
failed = dport_iter;
dport_iter = dport;
port_iter = port;
+ clear = false;
while (!is_cxl_root(port_iter) && dport_iter != failed) {
- cxl_bi_ctrl_dport_disable(dport_iter);
+ if (dport_iter == programmed)
+ clear = true;
+ cxl_bi_ctrl_dport_disable(dport_iter, clear);
dport_iter = port_iter->parent_dport;
port_iter = dport_iter->port;
}
diff --git a/drivers/cxl/core/region.c b/drivers/cxl/core/region.c
index f322a8a4a66d..0e3ba7c79a75 100644
--- a/drivers/cxl/core/region.c
+++ b/drivers/cxl/core/region.c
@@ -1861,6 +1861,20 @@ static int cxl_region_attach_position(struct cxl_region *cxlr,
return rc;
}
+static bool cxled_committed_bi(struct cxl_endpoint_decoder *cxled)
+{
+ struct cxl_port *port = cxled_to_port(cxled);
+ struct cxl_hdm *cxlhdm = dev_get_drvdata(&port->dev);
+ u32 ctrl;
+
+ if (!cxlhdm || !cxlhdm->regs.hdm_decoder)
+ return false;
+
+ ctrl = readl(cxlhdm->regs.hdm_decoder +
+ CXL_HDM_DECODER0_CTRL_OFFSET(cxled->cxld.id));
+ return FIELD_GET(CXL_HDM_DECODER0_CTRL_BI, ctrl);
+}
+
static const char *cxl_coherency_name(enum cxl_decoder_type type, bool bi)
{
if (type == CXL_DECODER_HOSTONLYMEM)
@@ -2179,6 +2193,24 @@ static int cxl_region_attach(struct cxl_region *cxlr,
return -ENXIO;
}
+ /* a committed decoder cannot inherit the region's flavor */
+ if (cxled->state == CXL_DECODER_STATE_AUTO) {
+ bool bi = cxled_committed_bi(cxled);
+ const char *have, *want;
+
+ have = cxl_coherency_name(cxled->cxld.target_type, bi);
+ want = cxl_coherency_name(cxlr->type,
+ cxl_root_decoder_is_bi(cxlrd));
+ if (cxled->cxld.target_type != cxlr->type ||
+ bi != cxl_root_decoder_is_bi(cxlrd)) {
+ dev_err(&cxlr->dev,
+ "%s:%s coherency model mismatch: %s vs %s\n",
+ dev_name(&cxlmd->dev),
+ dev_name(&cxled->cxld.dev), have, want);
+ return -ENXIO;
+ }
+ }
+
if (!cxled->dpa_res) {
dev_dbg(&cxlr->dev, "%s:%s: missing DPA allocation.\n",
dev_name(&cxlmd->dev), dev_name(&cxled->cxld.dev));
@@ -3838,10 +3870,26 @@ static struct cxl_region *construct_region(struct cxl_root_decoder *cxlrd,
struct cxl_dev_state *cxlds = cxlmd->cxlds;
int rc, part = READ_ONCE(cxled->part);
struct cxl_region *cxlr;
+ unsigned long need;
if (part < 0)
return ERR_PTR(-EBUSY);
+ /*
+ * A committed decoder defines the region built from it, so no
+ * attach check can find its coherency model wrong. Only the
+ * window's restrictions can, on the BI and range type axes.
+ */
+ need = cxled->cxld.target_type == CXL_DECODER_DEVMEM ?
+ CXL_DECODER_F_DEVMEM : CXL_DECODER_F_HOSTONLY;
+ if (cxled_committed_bi(cxled) != cxl_root_decoder_is_bi(cxlrd) ||
+ !(cxlrd->cxlsd.cxld.flags & need)) {
+ dev_err(&cxlrd->cxlsd.cxld.dev,
+ "%s:%s coherency model not permitted by the window\n",
+ dev_name(&cxlmd->dev), dev_name(&cxled->cxld.dev));
+ return ERR_PTR(-ENXIO);
+ }
+
do {
cxlr = __create_region(cxlrd, cxlds->part[part].mode,
atomic_read(&cxlrd->region_id),
--
2.39.5
^ permalink raw reply related [flat|nested] 33+ messages in thread* Re: [PATCH v9 08/10] cxl: Allow auto-committed BI hdm decoders
2026-09-22 23:38 ` [PATCH v9 08/10] cxl: Allow auto-committed BI hdm decoders Davidlohr Bueso
@ 2026-10-01 3:49 ` Alison Schofield
0 siblings, 0 replies; 33+ messages in thread
From: Alison Schofield @ 2026-10-01 3:49 UTC (permalink / raw)
To: Davidlohr Bueso
Cc: dave.jiang, jic23, icheng, ming.li, benjamin.cheatham, alucerop,
linux-cxl, Jonathan Cameron
On Tue, Sep 22, 2026 at 04:38:45PM -0700, Davidlohr Bueso wrote:
The cover letter says the series could be picked up with or without this
patch. I'm guessing that is a stale statement. Today a committed BI decoder
is enumerated like any other. After patch 3 it fails port enumeration,
and this patch is what fixes that. So not really optional?
> An HDM decoder that firmware committed with the BI bit set was
> refused outright. Accept it, and refuse only what is broken - BI
> paired with a host-only target range type, or for an endpoint
> decoder BI on a path that cannot route BISnp.
>
> Taking over such a path must not re-program what firmware already
> did, and a dport's committed state does not say which device it was
> committed for, only that firmware committed for something below it.
> Two endpoints sharing a level show why that is not enough to go on.
>
> RP Forward set, but for mem0
> |
> USP
> / \
> DSP0 DSP1
> | |
> mem0 mem1
> BI Enable BI Enable
> set clear
>
> Adoption is therefore seeded from the endpoint's own BI Enable, since
> only a device firmware itself enabled can have had its BI-ID
> accounted for above, and carried up the walk. A level is taken as
> found, with no register write and no commit, when nothing this driver
> routed sits below it, the value this driver would write is already
> there, and both the decoder and the switch's route table are
> committed - the table carries a commit of its own that firmware may
> not have performed. The first level to fall short ends adoption, and
> from it up the driver programs and commits as for any other device.
>
> RP Forward, written but never committed
> |
> USP0 route table committed by the driver
> |
> DSP0 Forward, written and committed
> | ..... adoption ended here .....
> USP1 route table already committed
> |
> DSP1 Enable, taken as found
> |
> mem0 BI Enable already set by firmware
>
> An adopted level owes no commit either way. Table 8-152 and Table
> 8-156 owe one per new BI device enabled below the port, and an
> adopted level is not seeing one.
>
> The endpoint is adopted the same way. BI Enable already set was an
> error before, since only this driver set it; it is now kept and
> cxlds->bi taken up without a write.
>
> On failure the walk unwinds only the levels it did not adopt. Since
> adoption is a prefix, the unwind drops the refcount below the first
> non-adopted level and leaves those registers alone. Teardown is
> unchanged and still deallocates every level.
>
> A committed decoder is refused when its coherency model does not fit
> where it lands. The window must permit it, and unlike a decoder this
> driver programs it cannot inherit a region's model, so its committed
> Target Range Type and BI bit must already match. Assembly never
> validated these bits, so setting them wrongly costs a platform the
> auto-assembly of the affected decoders.
Assembly did validate Target Range Type until patch 4.
cxl_region_attach() refused a decoder whose type differed from the
region's. The window check this adds also refuses non-BI auto regions
that assemble today. More inline at construct_region()...
snip
> diff --git a/drivers/cxl/core/pci.c b/drivers/cxl/core/pci.c
> index 9bf05820c142..f83f5d1d8255 100644
> --- a/drivers/cxl/core/pci.c
> +++ b/drivers/cxl/core/pci.c
> @@ -1087,6 +1087,29 @@ static int __cxl_bi_commit_decoder(struct device *dev, void __iomem *bi)
> scale, base);
> }
>
> +/* Committed, or no explicit commit required */
> +static bool cxl_bi_decoder_committed(void __iomem *bi)
> +{
> + u32 caps = readl(bi + CXL_BI_DECODER_CAPS_OFFSET);
> + u32 sts = readl(bi + CXL_BI_DECODER_STATUS_OFFSET);
> +
> + if (!FIELD_GET(CXL_BI_DECODER_CAPS_EXPLICIT_COMMIT_REQ, caps))
> + return true;
> +
> + return FIELD_GET(CXL_BI_DECODER_STATUS_BI_COMMITTED, sts);
> +}
> +
> +static bool cxl_bi_rt_committed(void __iomem *bi)
> +{
> + u32 caps = readl(bi + CXL_BI_RT_CAPS_OFFSET);
> + u32 sts = readl(bi + CXL_BI_RT_STATUS_OFFSET);
> +
> + if (!FIELD_GET(CXL_BI_RT_CAPS_EXPLICIT_COMMIT_REQ, caps))
> + return true;
> +
> + return FIELD_GET(CXL_BI_RT_STATUS_BI_COMMITTED, sts);
> +}
> +
Above 2 helpers differ only in register offsets and masks, and repeat
the EXPLICIT_COMMIT_REQ caps test already in __cxl_bi_commit_rt() and
__cxl_bi_commit_decoder().
Could they be one helper taking offsets and masks, like __cxl_bi_wait_commit()
does?
snip
> static int cxl_bi_enable_path(struct cxl_dev_state *cxlds,
> struct cxl_port *port, struct cxl_dport *dport)
> {
> - struct cxl_dport *dport_iter, *failed;
> + struct cxl_dport *dport_iter, *failed, *programmed = NULL;
> + struct cxl_port *endpoint = cxlds->cxlmd->endpoint;
> struct cxl_port *port_iter;
> + bool adopt, clear;
> int rc;
>
> + /*
> + * Adoption is a path property, not a per-dport one: only when
> + * firmware enabled the device itself does a dport's committed
> + * state cover this device's BI-ID.
> + */
> + adopt = FIELD_GET(CXL_BI_DECODER_CTRL_BI_ENABLE,
> + readl(endpoint->regs.bi_decoder +
> + CXL_BI_DECODER_CTRL_OFFSET));
> +
Is this cxl_bi_decoder_enabled(endpoint) from core.h?
> port_iter = port;
> dport_iter = dport;
> while (!is_cxl_root(port_iter)) {
> - rc = cxl_bi_ctrl_dport_enable(dport_iter, dport_iter == dport);
> + rc = cxl_bi_ctrl_dport_enable(dport_iter, dport_iter == dport,
> + &adopt);
> if (rc)
> goto err_rollback;
>
> + /* adoption is a prefix: this is the first level not adopted */
> + if (!adopt && !programmed)
> + programmed = dport_iter;
> +
> dport_iter = port_iter->parent_dport;
> port_iter = dport_iter->port;
> }
> @@ -1301,8 +1360,11 @@ static int cxl_bi_enable_path(struct cxl_dev_state *cxlds,
> failed = dport_iter;
> dport_iter = dport;
> port_iter = port;
> + clear = false;
> while (!is_cxl_root(port_iter) && dport_iter != failed) {
> - cxl_bi_ctrl_dport_disable(dport_iter);
> + if (dport_iter == programmed)
> + clear = true;
> + cxl_bi_ctrl_dport_disable(dport_iter, clear);
> dport_iter = port_iter->parent_dport;
> port_iter = dport_iter->port;
> }
> diff --git a/drivers/cxl/core/region.c b/drivers/cxl/core/region.c
> index f322a8a4a66d..0e3ba7c79a75 100644
> --- a/drivers/cxl/core/region.c
> +++ b/drivers/cxl/core/region.c
> @@ -1861,6 +1861,20 @@ static int cxl_region_attach_position(struct cxl_region *cxlr,
> return rc;
> }
>
> +static bool cxled_committed_bi(struct cxl_endpoint_decoder *cxled)
> +{
> + struct cxl_port *port = cxled_to_port(cxled);
> + struct cxl_hdm *cxlhdm = dev_get_drvdata(&port->dev);
> + u32 ctrl;
> +
> + if (!cxlhdm || !cxlhdm->regs.hdm_decoder)
> + return false;
> +
> + ctrl = readl(cxlhdm->regs.hdm_decoder +
> + CXL_HDM_DECODER0_CTRL_OFFSET(cxled->cxld.id));
I'm having some deja vu here. I may have asked and you answered this on
an earlier version, but now that I see Patch 9 it makes me ask again.
Patch 9 wouldn't need to mock this if we stored it as committed BI
state instead. Maybe in cxld->flags using CXL_DECODE_F_BI.
> + return FIELD_GET(CXL_HDM_DECODER0_CTRL_BI, ctrl);
> +}
> +
> static const char *cxl_coherency_name(enum cxl_decoder_type type, bool bi)
> {
> if (type == CXL_DECODER_HOSTONLYMEM)
> @@ -2179,6 +2193,24 @@ static int cxl_region_attach(struct cxl_region *cxlr,
> return -ENXIO;
> }
>
> + /* a committed decoder cannot inherit the region's flavor */
> + if (cxled->state == CXL_DECODER_STATE_AUTO) {
> + bool bi = cxled_committed_bi(cxled);
> + const char *have, *want;
> +
> + have = cxl_coherency_name(cxled->cxld.target_type, bi);
> + want = cxl_coherency_name(cxlr->type,
> + cxl_root_decoder_is_bi(cxlrd));
> + if (cxled->cxld.target_type != cxlr->type ||
> + bi != cxl_root_decoder_is_bi(cxlrd)) {
Looks like only need have and want computed when mismatch is detected.
cxl_root_decoder_is_bi could also be cached once.
> + dev_err(&cxlr->dev,
> + "%s:%s coherency model mismatch: %s vs %s\n",
> + dev_name(&cxlmd->dev),
> + dev_name(&cxled->cxld.dev), have, want);
> + return -ENXIO;
> + }
> + }
> +
> if (!cxled->dpa_res) {
> dev_dbg(&cxlr->dev, "%s:%s: missing DPA allocation.\n",
> dev_name(&cxlmd->dev), dev_name(&cxled->cxld.dev));
> @@ -3838,10 +3870,26 @@ static struct cxl_region *construct_region(struct cxl_root_decoder *cxlrd,
> struct cxl_dev_state *cxlds = cxlmd->cxlds;
> int rc, part = READ_ONCE(cxled->part);
> struct cxl_region *cxlr;
> + unsigned long need;
>
> if (part < 0)
> return ERR_PTR(-EBUSY);
>
> + /*
> + * A committed decoder defines the region built from it, so no
> + * attach check can find its coherency model wrong. Only the
> + * window's restrictions can, on the BI and range type axes.
> + */
> + need = cxled->cxld.target_type == CXL_DECODER_DEVMEM ?
> + CXL_DECODER_F_DEVMEM : CXL_DECODER_F_HOSTONLY;
> + if (cxled_committed_bi(cxled) != cxl_root_decoder_is_bi(cxlrd) ||
> + !(cxlrd->cxlsd.cxld.flags & need)) {
This refuses non BI auto regions that assemble before this patch.
Why is this in a patch called 'Allow auto-committed BI hdm decoders'.
> + dev_err(&cxlrd->cxlsd.cxld.dev,
> + "%s:%s coherency model not permitted by the window\n",
> + dev_name(&cxlmd->dev), dev_name(&cxled->cxld.dev));
> + return ERR_PTR(-ENXIO);
> + }
> +
> do {
> cxlr = __create_region(cxlrd, cxlds->part[part].mode,
> atomic_read(&cxlrd->region_id),
> --
> 2.39.5
>
^ permalink raw reply [flat|nested] 33+ messages in thread
* [PATCH v9 09/10] cxl/test: Add mock BI topology support
2026-09-22 23:38 [PATCH v9 0/10] cxl: Support Back-Invalidate Davidlohr Bueso
` (7 preceding siblings ...)
2026-09-22 23:38 ` [PATCH v9 08/10] cxl: Allow auto-committed BI hdm decoders Davidlohr Bueso
@ 2026-09-22 23:38 ` Davidlohr Bueso
2026-09-23 1:11 ` sashiko-bot
2026-09-23 0:43 ` [PATCH v9 10/10] cxl/doc: Update maturity map with BI support Davidlohr Bueso
2026-09-30 16:47 ` [PATCH v9 0/10] cxl: Support Back-Invalidate Alison Schofield
10 siblings, 1 reply; 33+ messages in thread
From: Davidlohr Bueso @ 2026-09-22 23:38 UTC (permalink / raw)
To: dave.jiang
Cc: jic23, alison.schofield, icheng, ming.li, benjamin.cheatham,
alucerop, dave, linux-cxl, Jonathan Cameron
Extend the mock topology with a Back-Invalidate path
covering both the type3 memdevs and the type2 accelerator.
Following the framework's convention of substituting software state
for register programming, cxl_bi_probe_capable() gains a --wrap shim
dispatching through cxl_mock_ops, and the mock decoder setup enables
BI in software in place of cxl_bi_setup(). The mock setup mirrors
cxl_bi_enable_path()'s walk - dport nr_bi accounting up to the root,
unwound by a devm action - and the capability check keeps the
VH-only rule.
A single HDM-DB window (BI | DEVMEM | VOLATILE, targeting host
bridge 0) is emitted in every topology mode and parses into a
CXL_DECODER_F_BI root decoder; DEVMEM satisfies can_create_ram().
In type2 mode it coexists with the accelerator's HDM-D window.
cxled_committed_bi() reads the BI bit from the HDM decoder
registers, so each mock port's cxl_hdm carries a page of plain
memory as that register block, maintained on decoder commit/reset
and restored on saved-decoder replay - letting a committed HDM-DB
region survive a cxl_acpi rebind. Endpoints get a second such page
standing in for BI Decoder Control, with BI Enable set, since
committing an HDM-DB region reads it back. The accelerator grows to
1G of capacity so the boot-time auto region leaves DPA for a BI
region.
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
Signed-off-by: Davidlohr Bueso <dave@stgolabs.net>
---
tools/testing/cxl/Kbuild | 1 +
tools/testing/cxl/test/accel.c | 2 +-
tools/testing/cxl/test/cxl.c | 171 ++++++++++++++++++++++++++++++++-
tools/testing/cxl/test/mock.c | 13 +++
tools/testing/cxl/test/mock.h | 1 +
5 files changed, 184 insertions(+), 4 deletions(-)
diff --git a/tools/testing/cxl/Kbuild b/tools/testing/cxl/Kbuild
index 2be1df80fcc9..811fb43225b8 100644
--- a/tools/testing/cxl/Kbuild
+++ b/tools/testing/cxl/Kbuild
@@ -14,6 +14,7 @@ ldflags-y += --wrap=devm_cxl_switch_port_decoders_setup
ldflags-y += --wrap=walk_hmem_resources
ldflags-y += --wrap=region_intersects
ldflags-y += --wrap=region_intersects_soft_reserve
+ldflags-y += --wrap=cxl_bi_probe_capable
DRIVERS := ../../../drivers
DAX_HMEM_SRC := $(DRIVERS)/dax/hmem
diff --git a/tools/testing/cxl/test/accel.c b/tools/testing/cxl/test/accel.c
index 8e6f4687ca02..f3800083ea27 100644
--- a/tools/testing/cxl/test/accel.c
+++ b/tools/testing/cxl/test/accel.c
@@ -31,7 +31,7 @@ static int cxl_mock_accel_probe(struct platform_device *pdev)
cxlds = &cxl_accel->cxlds;
cxlds->media_ready = true;
- rc = cxl_set_capacity(cxlds, SZ_512M);
+ rc = cxl_set_capacity(cxlds, SZ_1G);
if (rc)
return rc;
diff --git a/tools/testing/cxl/test/cxl.c b/tools/testing/cxl/test/cxl.c
index 62bd92b3be45..cddea44e780a 100644
--- a/tools/testing/cxl/test/cxl.c
+++ b/tools/testing/cxl/test/cxl.c
@@ -190,6 +190,10 @@ static struct {
struct acpi_cedt_cfmws cfmws;
u32 target[3];
} cfmws8;
+ struct {
+ struct acpi_cedt_cfmws cfmws;
+ u32 target[1];
+ } cfmws_bi;
struct {
struct acpi_cedt_cxims cxims;
u64 xormap_list[2];
@@ -373,6 +377,22 @@ static struct {
},
.target = { 0, 1, 2, },
},
+ .cfmws_bi = {
+ .cfmws = {
+ .header = {
+ .type = ACPI_CEDT_TYPE_CFMWS,
+ .length = sizeof(mock_cedt.cfmws_bi),
+ },
+ .interleave_ways = 0,
+ .granularity = 4,
+ .restrictions = ACPI_CEDT_CFMWS_RESTRICT_BI |
+ ACPI_CEDT_CFMWS_RESTRICT_DEVMEM |
+ ACPI_CEDT_CFMWS_RESTRICT_VOLATILE,
+ .qtg_id = FAKE_QTG_ID,
+ .window_size = SZ_256M * 4UL,
+ },
+ .target = { 0 },
+ },
.cxims0 = {
.cxims = {
.header = {
@@ -534,6 +554,12 @@ static int populate_cedt(void)
window->base_hpa = res->range.start;
}
+ res = alloc_mock_res(mock_cedt.cfmws_bi.cfmws.window_size,
+ max_t(int, SZ_256M, PMD_SIZE));
+ if (!res)
+ return -ENOMEM;
+ mock_cedt.cfmws_bi.cfmws.base_hpa = res->range.start;
+
return 0;
}
@@ -569,12 +595,17 @@ static int mock_acpi_table_parse_cedt(enum acpi_cedt_type id,
handler_arg(h, arg, end);
}
- if (id == ACPI_CEDT_TYPE_CFMWS)
+ if (id == ACPI_CEDT_TYPE_CFMWS) {
for (i = cfmws_start; i <= cfmws_end; i++) {
h = (union acpi_subtable_headers *) mock_cfmws[i];
end = (unsigned long) h + mock_cfmws[i]->header.length;
handler_arg(h, arg, end);
}
+ /* one HDM-DB window in every topology */
+ h = (union acpi_subtable_headers *)&mock_cedt.cfmws_bi.cfmws;
+ end = (unsigned long)h + mock_cedt.cfmws_bi.cfmws.header.length;
+ handler_arg(h, arg, end);
+ }
if (id == ACPI_CEDT_TYPE_CXIMS)
for (i = 0; i < ARRAY_SIZE(mock_cxims); i++) {
@@ -736,10 +767,65 @@ static struct cxl_hdm *mock_cxl_setup_hdm(struct cxl_port *port,
cxlhdm->port = port;
cxlhdm->interleave_mask = ~0U;
cxlhdm->iw_cap_mask = ~0UL;
+
+ /*
+ * A page of plain memory stands in for the HDM decoder register
+ * block: cxled_committed_bi() reads the per-decoder BI bit from
+ * it, which mock_decoder_commit()/reset() maintain below. All
+ * other consumers of these registers are bypassed by the mocked
+ * decoder setup and commit paths.
+ */
+ cxlhdm->regs.hdm_decoder =
+ (void __iomem *)devm_get_free_pages(dev,
+ GFP_KERNEL | __GFP_ZERO, 0);
+ if (!cxlhdm->regs.hdm_decoder)
+ return ERR_PTR(-ENOMEM);
+
+ /* likewise for the endpoint's BI Decoder block, BI Enable set */
+ if (is_cxl_endpoint(port) && !port->regs.bi_decoder) {
+ void __iomem *bi = (void __iomem *)
+ devm_get_free_pages(dev, GFP_KERNEL | __GFP_ZERO, 0);
+
+ if (!bi)
+ return ERR_PTR(-ENOMEM);
+ writel(CXL_BI_DECODER_CTRL_BI_ENABLE,
+ bi + CXL_BI_DECODER_CTRL_OFFSET);
+ port->regs.bi_decoder = bi;
+ }
+
dev_set_drvdata(dev, cxlhdm);
return cxlhdm;
}
+/* HPA-based, so replay after cxl_acpi rebind can re-derive it */
+static bool mock_hpa_is_bi(u64 hpa)
+{
+ struct acpi_cedt_cfmws *bi = &mock_cedt.cfmws_bi.cfmws;
+
+ return hpa >= bi->base_hpa && hpa < bi->base_hpa + bi->window_size;
+}
+
+static void mock_decoder_set_bi(struct cxl_decoder *cxld, bool bi)
+{
+ struct cxl_port *port = to_cxl_port(cxld->dev.parent);
+ struct cxl_hdm *cxlhdm = dev_get_drvdata(&port->dev);
+ void __iomem *ctrl;
+ u32 val;
+
+ if (!is_endpoint_decoder(&cxld->dev) || !cxlhdm ||
+ !cxlhdm->regs.hdm_decoder)
+ return;
+
+ ctrl = cxlhdm->regs.hdm_decoder +
+ CXL_HDM_DECODER0_CTRL_OFFSET(cxld->id);
+ val = readl(ctrl);
+ if (bi)
+ val |= CXL_HDM_DECODER0_CTRL_BI;
+ else
+ val &= ~CXL_HDM_DECODER0_CTRL_BI;
+ writel(val, ctrl);
+}
+
struct target_map_ctx {
u32 *target_map;
int index;
@@ -974,6 +1060,7 @@ static int mock_decoder_commit(struct cxl_decoder *cxld)
cxled->state = CXL_DECODER_STATE_AUTO;
}
+ mock_decoder_set_bi(cxld, mock_hpa_is_bi(cxld->hpa_range.start));
cxld_registry_update(cxld);
return 0;
@@ -1003,6 +1090,7 @@ static void mock_decoder_reset(struct cxl_decoder *cxld)
cxled->state = CXL_DECODER_STATE_MANUAL;
cxled->skip = 0;
}
+ mock_decoder_set_bi(cxld, false);
if (decoder_reset_preserve_registry)
dev_dbg(port->uport_dev, "decoder%d: skip registry update\n",
cxld->id);
@@ -1130,8 +1218,13 @@ static bool mock_decoder_handle_saved(struct cxl_decoder *cxld, struct cxl_test_
else
enabled = td->cxled.cxld.flags & CXL_DECODER_F_ENABLE;
- if (enabled)
- return !cxld_registry_restore(cxld, td);
+ if (enabled) {
+ if (cxld_registry_restore(cxld, td))
+ return false;
+ mock_decoder_set_bi(cxld,
+ mock_hpa_is_bi(cxld->hpa_range.start));
+ return true;
+ }
init_disabled_mock_decoder(cxld);
return false;
@@ -1453,9 +1546,12 @@ static int mock_cxl_enumerate_decoders(struct cxl_hdm *cxlhdm,
return 0;
}
+static int mock_cxl_bi_setup(struct cxl_port *endpoint);
+
static int __mock_cxl_decoders_setup(struct cxl_port *port)
{
struct cxl_hdm *cxlhdm;
+ int rc;
cxlhdm = mock_cxl_setup_hdm(port, NULL);
if (IS_ERR(cxlhdm)) {
@@ -1464,6 +1560,13 @@ static int __mock_cxl_decoders_setup(struct cxl_port *port)
return PTR_ERR(cxlhdm);
}
+ /* as the real setup: BI between the HDM state and the decoders */
+ if (is_cxl_endpoint(port)) {
+ rc = mock_cxl_bi_setup(port);
+ if (rc)
+ dev_dbg(&port->dev, "BI setup failed rc=%d\n", rc);
+ }
+
return mock_cxl_enumerate_decoders(cxlhdm, NULL);
}
@@ -1652,6 +1755,67 @@ mock_region_intersects_soft_reserve(resource_size_t start, size_t size)
return -1;
}
+/*
+ * All-software mirror of the BI enable path: no BI RT registers exist
+ * on mock devices, so capability and enablement are asserted directly
+ * while the dport nr_bi accounting - the part with driver-visible
+ * semantics (shared transit dports, teardown order) - follows the same
+ * walk the real cxl_bi_enable_path() takes.
+ */
+static void mock_cxl_bi_probe_capable(struct cxl_port *endpoint)
+{
+ struct cxl_memdev *cxlmd = to_cxl_memdev(endpoint->uport_dev);
+ struct cxl_dev_state *cxlds = cxlmd->cxlds;
+
+ /* BI is VH-only, mirroring cxl_bi_probe_capable() */
+ cxlds->bi_capable = !cxlds->rcd;
+}
+
+static void mock_cxl_bi_dealloc(void *data)
+{
+ struct cxl_port *endpoint = data;
+ struct cxl_memdev *cxlmd = to_cxl_memdev(endpoint->uport_dev);
+ struct cxl_dev_state *cxlds = cxlmd->cxlds;
+ struct cxl_dport *dport_iter;
+ struct cxl_port *port_iter;
+
+ cxlds->bi = false;
+ dport_iter = endpoint->parent_dport;
+ port_iter = dport_iter->port;
+ while (port_iter->parent_dport) {
+ scoped_guard(mutex, &port_iter->bi_lock) {
+ if (!WARN_ON_ONCE(dport_iter->nr_bi == 0))
+ dport_iter->nr_bi--;
+ }
+ dport_iter = port_iter->parent_dport;
+ port_iter = dport_iter->port;
+ }
+}
+
+static int mock_cxl_bi_setup(struct cxl_port *endpoint)
+{
+ struct cxl_memdev *cxlmd = to_cxl_memdev(endpoint->uport_dev);
+ struct cxl_dev_state *cxlds = cxlmd->cxlds;
+ struct cxl_dport *dport_iter;
+ struct cxl_port *port_iter;
+
+ if (!cxlds->bi_capable)
+ return 0;
+
+ dport_iter = endpoint->parent_dport;
+ port_iter = dport_iter->port;
+ while (port_iter->parent_dport) {
+ scoped_guard(mutex, &port_iter->bi_lock)
+ dport_iter->nr_bi++;
+ dport_iter = port_iter->parent_dport;
+ port_iter = dport_iter->port;
+ }
+ cxlds->bi = true;
+
+ return devm_add_action_or_reset(&endpoint->dev, mock_cxl_bi_dealloc,
+ endpoint);
+}
+
static struct cxl_mock_ops cxl_mock_ops = {
.is_mock_adev = is_mock_adev,
.is_mock_bridge = is_mock_bridge,
@@ -1670,6 +1834,7 @@ static struct cxl_mock_ops cxl_mock_ops = {
.walk_hmem_resources = mock_walk_hmem_resources,
.region_intersects = mock_region_intersects,
.region_intersects_soft_reserve = mock_region_intersects_soft_reserve,
+ .cxl_bi_probe_capable = mock_cxl_bi_probe_capable,
.list = LIST_HEAD_INIT(cxl_mock_ops.list),
};
diff --git a/tools/testing/cxl/test/mock.c b/tools/testing/cxl/test/mock.c
index 6454b868b122..eabc9bba0fac 100644
--- a/tools/testing/cxl/test/mock.c
+++ b/tools/testing/cxl/test/mock.c
@@ -181,6 +181,19 @@ int __wrap_devm_cxl_endpoint_decoders_setup(struct cxl_port *port)
}
EXPORT_SYMBOL_NS_GPL(__wrap_devm_cxl_endpoint_decoders_setup, "CXL");
+void __wrap_cxl_bi_probe_capable(struct cxl_port *endpoint)
+{
+ int index;
+ struct cxl_mock_ops *ops = get_cxl_mock_ops(&index);
+
+ if (ops && ops->is_mock_port(endpoint->uport_dev))
+ ops->cxl_bi_probe_capable(endpoint);
+ else
+ cxl_bi_probe_capable(endpoint);
+ put_cxl_mock_ops(index);
+}
+EXPORT_SYMBOL_NS_GPL(__wrap_cxl_bi_probe_capable, "CXL");
+
int __wrap_cxl_await_media_ready(struct cxl_dev_state *cxlds)
{
int rc, index;
diff --git a/tools/testing/cxl/test/mock.h b/tools/testing/cxl/test/mock.h
index 846d7c5d6eaa..c7919c3eb875 100644
--- a/tools/testing/cxl/test/mock.h
+++ b/tools/testing/cxl/test/mock.h
@@ -35,6 +35,7 @@ struct cxl_mock_ops {
unsigned long flags, unsigned long desc);
int (*region_intersects_soft_reserve)(resource_size_t start,
size_t size);
+ void (*cxl_bi_probe_capable)(struct cxl_port *endpoint);
};
int hmem_test_init(void);
--
2.39.5
^ permalink raw reply related [flat|nested] 33+ messages in thread* Re: [PATCH v9 09/10] cxl/test: Add mock BI topology support
2026-09-22 23:38 ` [PATCH v9 09/10] cxl/test: Add mock BI topology support Davidlohr Bueso
@ 2026-09-23 1:11 ` sashiko-bot
2026-09-23 20:57 ` Davidlohr Bueso
0 siblings, 1 reply; 33+ messages in thread
From: sashiko-bot @ 2026-09-23 1:11 UTC (permalink / raw)
To: Davidlohr Bueso; +Cc: linux-cxl
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] cxl/test: Use-After-Free in mock_cxl_setup_hdm() on unbind/rebind
--
commit 943579b999ae7670b28d433527c8563008ab779f
Author: Davidlohr Bueso <dave@stgolabs.net>
cxl/test: Add mock BI topology support
Extend the mock topology with a Back-Invalidate path
covering both the type3 memdevs and the type2 accelerator.
[ ... ]
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
Signed-off-by: Davidlohr Bueso <dave@stgolabs.net>
> diff --git a/tools/testing/cxl/test/cxl.c b/tools/testing/cxl/test/cxl.c
> index 62bd92b3be45f..cddea44e780aa 100644
> --- a/tools/testing/cxl/test/cxl.c
> +++ b/tools/testing/cxl/test/cxl.c
[ ... ]
> @@ -736,10 +767,65 @@ static struct cxl_hdm *mock_cxl_setup_hdm(struct cxl_port *port,
> cxlhdm->port = port;
> cxlhdm->interleave_mask = ~0U;
> cxlhdm->iw_cap_mask = ~0UL;
> +
> + /*
> + * A page of plain memory stands in for the HDM decoder register
> + * block: cxled_committed_bi() reads the per-decoder BI bit from
> + * it, which mock_decoder_commit()/reset() maintain below. All
> + * other consumers of these registers are bypassed by the mocked
> + * decoder setup and commit paths.
> + */
> + cxlhdm->regs.hdm_decoder =
> + (void __iomem *)devm_get_free_pages(dev,
> + GFP_KERNEL | __GFP_ZERO, 0);
> + if (!cxlhdm->regs.hdm_decoder)
> + return ERR_PTR(-ENOMEM);
> +
> + /* likewise for the endpoint's BI Decoder block, BI Enable set */
> + if (is_cxl_endpoint(port) && !port->regs.bi_decoder) {
> + void __iomem *bi = (void __iomem *)
> + devm_get_free_pages(dev, GFP_KERNEL | __GFP_ZERO, 0);
> +
> + if (!bi)
> + return ERR_PTR(-ENOMEM);
> + writel(CXL_BI_DECODER_CTRL_BI_ENABLE,
> + bi + CXL_BI_DECODER_CTRL_OFFSET);
> + port->regs.bi_decoder = bi;
> + }
[Severity: Medium]
Could this lead to a use-after-free during a driver unbind/rebind cycle?
In mock_cxl_setup_hdm(), devm_get_free_pages() is used to allocate memory
for port->regs.bi_decoder. During a cxl_test unbind, devres will free this
memory, but the pointer remains set in the long-lived struct cxl_port.
On a subsequent rebind, the condition !port->regs.bi_decoder will evaluate to
false, skipping the allocation of new memory. This leaves a dangling pointer
that could be dereferenced later by cxl_bi_decoder_enabled().
--
Sashiko AI review · https://sashiko.dev/#/patchset/cover.1790103847.git.dave@stgolabs.net?part=9
^ permalink raw reply [flat|nested] 33+ messages in thread* Re: [PATCH v9 09/10] cxl/test: Add mock BI topology support
2026-09-23 1:11 ` sashiko-bot
@ 2026-09-23 20:57 ` Davidlohr Bueso
0 siblings, 0 replies; 33+ messages in thread
From: Davidlohr Bueso @ 2026-09-23 20:57 UTC (permalink / raw)
To: sashiko-reviews, dave.jiang; +Cc: linux-cxl
On Wed, 23 Sep 2026, sashiko-bot@kernel.org wrote:
>Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
>- [Medium] cxl/test: Use-After-Free in mock_cxl_setup_hdm() on unbind/rebind
This is valid, but just affects the mock device, I have fixed it per the below.
>--
>
>commit 943579b999ae7670b28d433527c8563008ab779f
>Author: Davidlohr Bueso <dave@stgolabs.net>
>
> cxl/test: Add mock BI topology support
>
> Extend the mock topology with a Back-Invalidate path
> covering both the type3 memdevs and the type2 accelerator.
>
>[ ... ]
>
> Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
> Signed-off-by: Davidlohr Bueso <dave@stgolabs.net>
>
>> diff --git a/tools/testing/cxl/test/cxl.c b/tools/testing/cxl/test/cxl.c
>> index 62bd92b3be45f..cddea44e780aa 100644
>> --- a/tools/testing/cxl/test/cxl.c
>> +++ b/tools/testing/cxl/test/cxl.c
>[ ... ]
>> @@ -736,10 +767,65 @@ static struct cxl_hdm *mock_cxl_setup_hdm(struct cxl_port *port,
>> cxlhdm->port = port;
>> cxlhdm->interleave_mask = ~0U;
>> cxlhdm->iw_cap_mask = ~0UL;
>> +
>> + /*
>> + * A page of plain memory stands in for the HDM decoder register
>> + * block: cxled_committed_bi() reads the per-decoder BI bit from
>> + * it, which mock_decoder_commit()/reset() maintain below. All
>> + * other consumers of these registers are bypassed by the mocked
>> + * decoder setup and commit paths.
>> + */
>> + cxlhdm->regs.hdm_decoder =
>> + (void __iomem *)devm_get_free_pages(dev,
>> + GFP_KERNEL | __GFP_ZERO, 0);
>> + if (!cxlhdm->regs.hdm_decoder)
>> + return ERR_PTR(-ENOMEM);
>> +
>> + /* likewise for the endpoint's BI Decoder block, BI Enable set */
>> + if (is_cxl_endpoint(port) && !port->regs.bi_decoder) {
>> + void __iomem *bi = (void __iomem *)
>> + devm_get_free_pages(dev, GFP_KERNEL | __GFP_ZERO, 0);
>> +
>> + if (!bi)
>> + return ERR_PTR(-ENOMEM);
>> + writel(CXL_BI_DECODER_CTRL_BI_ENABLE,
>> + bi + CXL_BI_DECODER_CTRL_OFFSET);
>> + port->regs.bi_decoder = bi;
>> + }
>
>[Severity: Medium]
>Could this lead to a use-after-free during a driver unbind/rebind cycle?
>
>In mock_cxl_setup_hdm(), devm_get_free_pages() is used to allocate memory
>for port->regs.bi_decoder. During a cxl_test unbind, devres will free this
>memory, but the pointer remains set in the long-lived struct cxl_port.
>
>On a subsequent rebind, the condition !port->regs.bi_decoder will evaluate to
>false, skipping the allocation of new memory. This leaves a dangling pointer
>that could be dereferenced later by cxl_bi_decoder_enabled().
I was able to reproduce this pounding with the mock device with:
echo endpoint16 > /sys/bus/cxl/drivers/cxl_port/unbind
echo endpoint16 > /sys/bus/cxl/drivers/cxl_port/bind
<create an HDM-DB region on decoder16.0 and commit it>
The bind mostly fails with ENODEV, but for a window where that bind beats
the the unbind's queued detach work, that reprobes the same cxl_port, so
we run into:
[ 25.154927] cxl_port endpoint14: probe: 0
[ 25.229702] cxl region6: mem7:decoder14.0 BI disabled by reset during commit
Nothing had cleared BI Enable on that endpoint.
I will fold this fix fix with:
- if (is_cxl_endpoint(port) && !port->regs.bi_decoder) {
+ if (is_cxl_endpoint(port)) {
(mock device only, real driver not affected since devm_cxl_port_bi_setup()
re-maps on every probe and cxl_map_component_regs() assigns
unconditionally).
^ permalink raw reply [flat|nested] 33+ messages in thread
* [PATCH v9 10/10] cxl/doc: Update maturity map with BI support
2026-09-22 23:38 [PATCH v9 0/10] cxl: Support Back-Invalidate Davidlohr Bueso
` (8 preceding siblings ...)
2026-09-22 23:38 ` [PATCH v9 09/10] cxl/test: Add mock BI topology support Davidlohr Bueso
@ 2026-09-23 0:43 ` Davidlohr Bueso
2026-09-24 0:50 ` Li Ming
2026-09-30 23:08 ` Alison Schofield
2026-09-30 16:47 ` [PATCH v9 0/10] cxl: Support Back-Invalidate Alison Schofield
10 siblings, 2 replies; 33+ messages in thread
From: Davidlohr Bueso @ 2026-09-23 0:43 UTC (permalink / raw)
To: dave.jiang
Cc: jic23, alison.schofield, icheng, ming.li, benjamin.cheatham,
alucerop, dave, linux-cxl, Jonathan Cameron
Add the respective Back-Invalidate info to the maturity map document.
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
Signed-off-by: Davidlohr Bueso <dave@stgolabs.net>
---
Documentation/driver-api/cxl/maturity-map.rst | 16 +++++++++++++++-
1 file changed, 15 insertions(+), 1 deletion(-)
diff --git a/Documentation/driver-api/cxl/maturity-map.rst b/Documentation/driver-api/cxl/maturity-map.rst
index 282c1102dd81..0b857476cb47 100644
--- a/Documentation/driver-api/cxl/maturity-map.rst
+++ b/Documentation/driver-api/cxl/maturity-map.rst
@@ -63,6 +63,11 @@ in place, but there are several corner cases that are pending closure.
* [0] Decoder target and granularity constraints
+* [1] :ref:`Back-Invalidate (HDM-DB) <back-invalidate>`
+
+ * [3] BI topology enable / disable (endpoint, DSP, USP route table, RP)
+ * [2] Firmware-enabled path adoption
+
* [2] Performance enumeration
* [3] Endpoint CDAT
@@ -166,7 +171,7 @@ Accelerator
-----------
* [0] Accelerator memory enumeration HDM-D (CXL 1.1/2.0 Type-2)
-* [0] Accelerator memory enumeration HDM-DB (CXL 3.0 Type-2)
+* [1] Accelerator memory enumeration HDM-DB (CXL 3.0 Type-2)
* [0] CXL.cache 68b (CXL 2.0)
* [0] CXL.cache 256b Cache IDs (CXL 3.0)
@@ -192,6 +197,15 @@ Details
hiding some standard registers like PCIe Link Status / Capabilities in
the CXL RCRB (Root Complex Register Block).
+.. _back-invalidate:
+
+* **Back-Invalidate**: HDM-DB lets a device snoop the host over the
+ CXL.mem BISnp channel instead of CXL.cache. The kernel enables BI on
+ every port between a device and its root port, requires 256B Flit
+ mode on the path, adopts paths firmware already enabled, and creates
+ HDM-DB regions under CFMWS windows carrying the Back-Invalidate
+ restriction for Type 3 and Type 2 devices alike.
+
.. _background-commands:
* **Background commands**: The CXL background command mechanism is
--
2.39.5
^ permalink raw reply related [flat|nested] 33+ messages in thread* Re: [PATCH v9 10/10] cxl/doc: Update maturity map with BI support
2026-09-23 0:43 ` [PATCH v9 10/10] cxl/doc: Update maturity map with BI support Davidlohr Bueso
@ 2026-09-24 0:50 ` Li Ming
2026-09-30 23:08 ` Alison Schofield
1 sibling, 0 replies; 33+ messages in thread
From: Li Ming @ 2026-09-24 0:50 UTC (permalink / raw)
To: Davidlohr Bueso, dave.jiang
Cc: jic23, alison.schofield, icheng, benjamin.cheatham, alucerop,
linux-cxl, Jonathan Cameron
On 9/23/2026 8:43 AM, Davidlohr Bueso wrote:
> Add the respective Back-Invalidate info to the maturity map document.
>
> Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
> Signed-off-by: Davidlohr Bueso <dave@stgolabs.net>
Reviewed-by: Li Ming <ming.li@zohomail.com>
> ---
> Documentation/driver-api/cxl/maturity-map.rst | 16 +++++++++++++++-
> 1 file changed, 15 insertions(+), 1 deletion(-)
>
> diff --git a/Documentation/driver-api/cxl/maturity-map.rst b/Documentation/driver-api/cxl/maturity-map.rst
> index 282c1102dd81..0b857476cb47 100644
> --- a/Documentation/driver-api/cxl/maturity-map.rst
> +++ b/Documentation/driver-api/cxl/maturity-map.rst
> @@ -63,6 +63,11 @@ in place, but there are several corner cases that are pending closure.
>
> * [0] Decoder target and granularity constraints
>
> +* [1] :ref:`Back-Invalidate (HDM-DB) <back-invalidate>`
> +
> + * [3] BI topology enable / disable (endpoint, DSP, USP route table, RP)
> + * [2] Firmware-enabled path adoption
> +
> * [2] Performance enumeration
>
> * [3] Endpoint CDAT
> @@ -166,7 +171,7 @@ Accelerator
> -----------
>
> * [0] Accelerator memory enumeration HDM-D (CXL 1.1/2.0 Type-2)
> -* [0] Accelerator memory enumeration HDM-DB (CXL 3.0 Type-2)
> +* [1] Accelerator memory enumeration HDM-DB (CXL 3.0 Type-2)
> * [0] CXL.cache 68b (CXL 2.0)
> * [0] CXL.cache 256b Cache IDs (CXL 3.0)
>
> @@ -192,6 +197,15 @@ Details
> hiding some standard registers like PCIe Link Status / Capabilities in
> the CXL RCRB (Root Complex Register Block).
>
> +.. _back-invalidate:
> +
> +* **Back-Invalidate**: HDM-DB lets a device snoop the host over the
> + CXL.mem BISnp channel instead of CXL.cache. The kernel enables BI on
> + every port between a device and its root port, requires 256B Flit
> + mode on the path, adopts paths firmware already enabled, and creates
> + HDM-DB regions under CFMWS windows carrying the Back-Invalidate
> + restriction for Type 3 and Type 2 devices alike.
> +
> .. _background-commands:
>
> * **Background commands**: The CXL background command mechanism is
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: [PATCH v9 10/10] cxl/doc: Update maturity map with BI support
2026-09-23 0:43 ` [PATCH v9 10/10] cxl/doc: Update maturity map with BI support Davidlohr Bueso
2026-09-24 0:50 ` Li Ming
@ 2026-09-30 23:08 ` Alison Schofield
1 sibling, 0 replies; 33+ messages in thread
From: Alison Schofield @ 2026-09-30 23:08 UTC (permalink / raw)
To: Davidlohr Bueso
Cc: dave.jiang, jic23, icheng, ming.li, benjamin.cheatham, alucerop,
linux-cxl, Jonathan Cameron
On Tue, Sep 22, 2026 at 05:43:25PM -0700, Davidlohr Bueso wrote:
> Add the respective Back-Invalidate info to the maturity map document.
Looks like this actually updates BI and Type2 status, which is good,
just not stated in the commit subject or message.
If you're reving anyways, please fix that up.
Reviewed-by: Alison Schofield <alison.schofield@intel.com>
>
> Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
> Signed-off-by: Davidlohr Bueso <dave@stgolabs.net>
> ---
> Documentation/driver-api/cxl/maturity-map.rst | 16 +++++++++++++++-
> 1 file changed, 15 insertions(+), 1 deletion(-)
>
> diff --git a/Documentation/driver-api/cxl/maturity-map.rst b/Documentation/driver-api/cxl/maturity-map.rst
> index 282c1102dd81..0b857476cb47 100644
> --- a/Documentation/driver-api/cxl/maturity-map.rst
> +++ b/Documentation/driver-api/cxl/maturity-map.rst
> @@ -63,6 +63,11 @@ in place, but there are several corner cases that are pending closure.
>
> * [0] Decoder target and granularity constraints
>
> +* [1] :ref:`Back-Invalidate (HDM-DB) <back-invalidate>`
> +
> + * [3] BI topology enable / disable (endpoint, DSP, USP route table, RP)
> + * [2] Firmware-enabled path adoption
> +
> * [2] Performance enumeration
>
> * [3] Endpoint CDAT
> @@ -166,7 +171,7 @@ Accelerator
> -----------
>
> * [0] Accelerator memory enumeration HDM-D (CXL 1.1/2.0 Type-2)
> -* [0] Accelerator memory enumeration HDM-DB (CXL 3.0 Type-2)
> +* [1] Accelerator memory enumeration HDM-DB (CXL 3.0 Type-2)
> * [0] CXL.cache 68b (CXL 2.0)
> * [0] CXL.cache 256b Cache IDs (CXL 3.0)
>
> @@ -192,6 +197,15 @@ Details
> hiding some standard registers like PCIe Link Status / Capabilities in
> the CXL RCRB (Root Complex Register Block).
>
> +.. _back-invalidate:
> +
> +* **Back-Invalidate**: HDM-DB lets a device snoop the host over the
> + CXL.mem BISnp channel instead of CXL.cache. The kernel enables BI on
> + every port between a device and its root port, requires 256B Flit
> + mode on the path, adopts paths firmware already enabled, and creates
> + HDM-DB regions under CFMWS windows carrying the Back-Invalidate
> + restriction for Type 3 and Type 2 devices alike.
> +
> .. _background-commands:
>
> * **Background commands**: The CXL background command mechanism is
> --
> 2.39.5
>
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: [PATCH v9 0/10] cxl: Support Back-Invalidate
2026-09-22 23:38 [PATCH v9 0/10] cxl: Support Back-Invalidate Davidlohr Bueso
` (9 preceding siblings ...)
2026-09-23 0:43 ` [PATCH v9 10/10] cxl/doc: Update maturity map with BI support Davidlohr Bueso
@ 2026-09-30 16:47 ` Alison Schofield
10 siblings, 0 replies; 33+ messages in thread
From: Alison Schofield @ 2026-09-30 16:47 UTC (permalink / raw)
To: Davidlohr Bueso
Cc: dave.jiang, jic23, icheng, ming.li, benjamin.cheatham, alucerop,
linux-cxl
On Tue, Sep 22, 2026 at 04:38:37PM -0700, Davidlohr Bueso wrote:
snip
>
> This passes regression testing (nothing breaks) ndctl suite via cxl_test,
> with the mock HDM-DB window of patch 9 and the BI phase covering HDM-DB
> assembly, cxl_acpi rebind replay and the Type 2 accelerator. The BI
> mock device test script along with tooling to consume the new sysfs ABI
> from this series is in the ndctl tree here:
>
> https://github.com/davidlohr/ndctl/tree/cxl-back-invalidate-test
Please post this to linux-cxl and nvdimm@lists.linux.dev
Thanks!
^ permalink raw reply [flat|nested] 33+ messages in thread