Linux CXL
 help / color / mirror / Atom feed
* [PATCH 00/15] Add initial CXL.cache support
@ 2026-09-23 17:33 Ben Cheatham
  2026-09-23 17:33 ` [PATCH 01/15] cxl/core: Add CXL.cache device struct Ben Cheatham
                   ` (15 more replies)
  0 siblings, 16 replies; 28+ messages in thread
From: Ben Cheatham @ 2026-09-23 17:33 UTC (permalink / raw)
  To: linux-cxl, dave, jic23, dave.jiang, alison.schofield
  Cc: linux-iommu, benjamin.cheatham, terry.bowman, robert.richter

This series adds support for CXL.cache devices to the CXL subsystem. The
set includes the following:
	- Updates to the existing CXL infrastructure to enable adding
	  CXL.cache devices (struct cxl_cachedev) to the driver
	- Support for cache id and snoop filter capabilities
	- IOMMU updates to allow CXL.cache ATS requests
	- Driver for struct cxl_cachedev devices that calls all of the above

The vast majority of the support is gated behind CONFIG_CXL_CACHE since
CXL.cache devices aren't commonplace. This set is untested since I don't
(currently) have access to a CXL.cache-capable device. Using multiple
devices requires cache id capabilties, which may not be available
depending on the platform (for AMD it's Venice or later).

There is a major missing piece in this set: There's no endpoint driver that
adds a cxl_cachedev device, so this code is currently unused. I was
planning on only sending this out once I had a device, but the plans for
the device I was planning to support fell through and I didn't want to
just sit on this.

Another thing that's missing is mapping host memory for CXL.cache use. I
think the DMA API *should* just work after these changes, but it's not
tested. I originally had the CXL core driver set up the DMA range for
the endpoint driver, but it didn't end up making much sense. So, I've
left the implementation for the (eventual) endpoint driver.

Thanks,
Ben

base-commit: cee9395acd8043be0644b25c34bfa86623f2b935 (v7.3-rc1)

Ben Cheatham (15):
  cxl/core: Add CXL.cache device struct
  cxl/cache: Add cxl_cache driver
  cxl/core: Change cxl_ep_load() to use device pointer parameter
  cxl/core: Update devm_cxl_enumerate_ports() for cxl_cachedevs
  cxl/port: Split endpoint port probe on device type
  cxl/core: Update devm_cxl_add_endpoint() for cxl_cachedevs
  cxl/cache: Verify port hierarchy has CXL.cache enabled
  cxl/core, cache: Add Cache ID register probing and init
  cxl/core: Add Cache ID verification
  cxl/core: Add Cache ID allocation
  cxl/core: Add support for HDM-D cache id programming
  cxl/cache: Add snoop filter creation and set up
  cxl/cache: Add snoop filter allocation
  iommu, cxl: Configure IOMMU for CXL.cache
  cxl/cache: Enable CXL.cache on successful probe

 drivers/cxl/Kconfig                 |  14 +
 drivers/cxl/Makefile                |   6 +-
 drivers/cxl/cache.c                 | 285 ++++++++++
 drivers/cxl/core/Makefile           |   2 +
 drivers/cxl/core/cache.c            | 851 ++++++++++++++++++++++++++++
 drivers/cxl/core/cachedev.c         | 131 +++++
 drivers/cxl/core/core.h             |   7 +
 drivers/cxl/core/memdev.c           |   1 +
 drivers/cxl/core/pci.c              |  78 ++-
 drivers/cxl/core/port.c             | 149 +++--
 drivers/cxl/core/region.c           |  24 +-
 drivers/cxl/core/regs.c             |  30 +
 drivers/cxl/cxl.h                   |  74 ++-
 drivers/cxl/cxlcache.h              |  50 ++
 drivers/cxl/cxlmem.h                |   4 +-
 drivers/cxl/mem.c                   |   4 +-
 drivers/cxl/port.c                  |  97 +++-
 drivers/iommu/amd/amd_iommu_types.h |   4 +
 drivers/iommu/amd/init.c            |  14 +
 drivers/iommu/amd/iommu.c           |  56 +-
 drivers/iommu/iommu.c               |  28 +
 include/cxl/cxl.h                   |  78 ++-
 include/linux/iommu.h               |   8 +
 include/uapi/linux/pci_regs.h       |   8 +
 24 files changed, 1911 insertions(+), 92 deletions(-)
 create mode 100644 drivers/cxl/cache.c
 create mode 100644 drivers/cxl/core/cache.c
 create mode 100644 drivers/cxl/core/cachedev.c
 create mode 100644 drivers/cxl/cxlcache.h

-- 
2.53.0


^ permalink raw reply	[flat|nested] 28+ messages in thread

* [PATCH 01/15] cxl/core: Add CXL.cache device struct
  2026-09-23 17:33 [PATCH 00/15] Add initial CXL.cache support Ben Cheatham
@ 2026-09-23 17:33 ` Ben Cheatham
  2026-09-23 17:41   ` sashiko-bot
  2026-09-23 17:33 ` [PATCH 02/15] cxl/cache: Add cxl_cache driver Ben Cheatham
                   ` (14 subsequent siblings)
  15 siblings, 1 reply; 28+ messages in thread
From: Ben Cheatham @ 2026-09-23 17:33 UTC (permalink / raw)
  To: linux-cxl, dave, jic23, dave.jiang, alison.schofield
  Cc: linux-iommu, benjamin.cheatham, terry.bowman, robert.richter

Add a new CXL.cache device (struct cxl_cachedev) that is the cache
analogue to struct cxl_memdev. This device will be created by CXL type
1 & 2 endpoint drivers to enable and manage the cache capabilities of
the underlying PCIe device via the CXL core.

Signed-off-by: Ben Cheatham <Benjamin.Cheatham@amd.com>
---
 drivers/cxl/core/Makefile   |   1 +
 drivers/cxl/core/cachedev.c | 119 ++++++++++++++++++++++++++++++++++++
 drivers/cxl/core/port.c     |   3 +
 drivers/cxl/cxl.h           |   1 +
 drivers/cxl/cxlcache.h      |  37 +++++++++++
 include/cxl/cxl.h           |   2 +
 6 files changed, 163 insertions(+)
 create mode 100644 drivers/cxl/core/cachedev.c
 create mode 100644 drivers/cxl/cxlcache.h

diff --git a/drivers/cxl/core/Makefile b/drivers/cxl/core/Makefile
index ce7213818d3c..f51575abe5af 100644
--- a/drivers/cxl/core/Makefile
+++ b/drivers/cxl/core/Makefile
@@ -9,6 +9,7 @@ cxl_core-y := port.o
 cxl_core-y += pmem.o
 cxl_core-y += regs.o
 cxl_core-y += memdev.o
+cxl_core-y += cachedev.o
 cxl_core-y += mbox.o
 cxl_core-y += pci.o
 cxl_core-y += hdm.o
diff --git a/drivers/cxl/core/cachedev.c b/drivers/cxl/core/cachedev.c
new file mode 100644
index 000000000000..3a7a60165f62
--- /dev/null
+++ b/drivers/cxl/core/cachedev.c
@@ -0,0 +1,119 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Copyright (C) 2026 Advanced Micro Devices, Inc. */
+#include <linux/device.h>
+#include <linux/pci.h>
+#include <cxlcache.h>
+
+static DEFINE_IDA(cxl_cachedev_ida);
+
+static void cxl_cachedev_release(struct device *dev)
+{
+	struct cxl_cachedev *cxlcd = to_cxl_cachedev(dev);
+
+	ida_free(&cxl_cachedev_ida, cxlcd->id);
+	kfree(cxlcd);
+}
+
+static void cxl_cachedev_unregister(void *dev)
+{
+	struct cxl_cachedev *cxlcd = dev;
+	struct cxl_dev_state *cxlds = cxlcd->cxlds;
+
+	cxlcd->cxlds = NULL;
+	cxlds->cxlcd = NULL;
+	device_del(&cxlcd->dev);
+	put_device(&cxlcd->dev);
+}
+
+static char *cxl_cachedev_devnode(const struct device *dev, umode_t *mode,
+				  kuid_t *uid, kgid_t *gid)
+{
+	return kasprintf(GFP_KERNEL, "cxl/%s", dev_name(dev));
+}
+
+static const struct device_type cxl_cachedev_type = {
+	.name = "cxl_cachedev",
+	.release = cxl_cachedev_release,
+	.devnode = cxl_cachedev_devnode,
+};
+
+bool is_cxl_cachedev(const struct device *dev)
+{
+	return dev->type == &cxl_cachedev_type;
+}
+EXPORT_SYMBOL_NS_GPL(is_cxl_cachedev, "CXL");
+
+static struct lock_class_key cxl_cachedev_key;
+
+static struct cxl_cachedev *cxl_cachedev_alloc(struct cxl_dev_state *cxlds)
+{
+	struct device *dev;
+	int rc;
+
+	struct cxl_cachedev *cxlcd __free(kfree) =
+		kzalloc(sizeof(*cxlcd), GFP_KERNEL);
+	if (!cxlcd)
+		return ERR_PTR(-ENOMEM);
+
+	rc = ida_alloc(&cxl_cachedev_ida, GFP_KERNEL);
+	if (rc < 0)
+		return ERR_PTR(rc);
+
+	cxlcd->id = rc;
+	cxlcd->depth = -1;
+	cxlcd->endpoint = ERR_PTR(-ENODEV);
+
+	dev = &cxlcd->dev;
+	device_initialize(dev);
+	lockdep_set_class(&dev->mutex, &cxl_cachedev_key);
+	dev->parent = cxlds->dev;
+	dev->bus = &cxl_bus_type;
+	dev->type = &cxl_cachedev_type;
+	device_set_pm_not_required(dev);
+
+	return_ptr(cxlcd);
+}
+
+DEFINE_FREE(put_cxlcd, struct cxl_cachedev *,
+	    if (!IS_ERR_OR_NULL(_T)) put_device(&_T->dev));
+
+static struct cxl_cachedev *cxl_cachedev_autoremove(struct cxl_cachedev *cxlcd)
+{
+	int rc;
+
+	rc = devm_add_action_or_reset(cxlcd->cxlds->dev,
+				      cxl_cachedev_unregister, cxlcd);
+	if (rc)
+		return ERR_PTR(rc);
+
+	return cxlcd;
+}
+
+struct cxl_cachedev *__devm_cxl_add_cachedev(struct cxl_dev_state *cxlds,
+					     void *attach)
+{
+	struct device *dev;
+	int rc;
+
+	struct cxl_cachedev *cxlcd __free(put_cxlcd) =
+		cxl_cachedev_alloc(cxlds);
+	if (IS_ERR(cxlcd))
+		return cxlcd;
+
+	dev = &cxlcd->dev;
+	rc = dev_set_name(dev, "cache%d", cxlcd->id);
+	if (rc)
+		return ERR_PTR(rc);
+
+	cxlcd->cxlds = cxlds;
+	cxlds->cxlcd = cxlcd;
+
+	rc = device_add(dev);
+	if (rc) {
+		cxlds->cxlcd = NULL;
+		return ERR_PTR(rc);
+	}
+
+	return cxl_cachedev_autoremove(no_free_ptr(cxlcd));
+}
+EXPORT_SYMBOL_NS_GPL(__devm_cxl_add_cachedev, "CXL");
diff --git a/drivers/cxl/core/port.c b/drivers/cxl/core/port.c
index 625e4aa427db..c0f066c0d3d1 100644
--- a/drivers/cxl/core/port.c
+++ b/drivers/cxl/core/port.c
@@ -12,6 +12,7 @@
 #include <linux/node.h>
 #include <cxl/einj.h>
 #include <cxl/pci.h>
+#include <cxlcache.h>
 #include <cxlmem.h>
 #include <cxlpci.h>
 #include <cxl.h>
@@ -78,6 +79,8 @@ static int cxl_device_id(const struct device *dev)
 		return CXL_DEVICE_REGION;
 	if (dev->type == &cxl_pmu_type)
 		return CXL_DEVICE_PMU;
+	if (is_cxl_cachedev(dev))
+		return CXL_DEVICE_ACCELERATOR;
 	return 0;
 }
 
diff --git a/drivers/cxl/cxl.h b/drivers/cxl/cxl.h
index cab8ce39f465..72d2686a408b 100644
--- a/drivers/cxl/cxl.h
+++ b/drivers/cxl/cxl.h
@@ -853,6 +853,7 @@ void cxl_driver_unregister(struct cxl_driver *cxl_drv);
 #define CXL_DEVICE_PMEM_REGION		7
 #define CXL_DEVICE_DAX_REGION		8
 #define CXL_DEVICE_PMU			9
+#define CXL_DEVICE_ACCELERATOR		10
 
 #define MODULE_ALIAS_CXL(type) MODULE_ALIAS("cxl:t" __stringify(type) "*")
 #define CXL_MODALIAS_FMT "cxl:t%d"
diff --git a/drivers/cxl/cxlcache.h b/drivers/cxl/cxlcache.h
new file mode 100644
index 000000000000..6ca887e078a1
--- /dev/null
+++ b/drivers/cxl/cxlcache.h
@@ -0,0 +1,37 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+#ifndef __CXL_CACHE_H__
+#define __CXL_CACHE_H__
+
+#include "cxl.h"
+
+/**
+ * struct cxl_cachedev - CXL bus object representing the cache capabilities of
+ * a CXL device
+ * @dev: driver core device object
+ * @cxlds: device state backing this device
+ * @endpoint: connection to the CXL port topology for this device
+ * @id: id number of this cachedev instance
+ * @depth: endpoint port depth in hierarchy
+ */
+struct cxl_cachedev {
+	struct device dev;
+	struct cxl_dev_state *cxlds;
+	struct cxl_port *endpoint;
+	int id;
+	int depth;
+};
+
+bool is_cxl_cachedev(const struct device *dev);
+
+static inline struct cxl_cachedev *to_cxl_cachedev(struct device *dev)
+{
+	if (dev_WARN_ONCE(dev, !is_cxl_cachedev(dev),
+			  "to_cxl_cachedev() called on non cxl_cachedev device"))
+		return NULL;
+
+	return container_of(dev, struct cxl_cachedev, dev);
+}
+
+struct cxl_cachedev *__devm_cxl_add_cachedev(struct cxl_dev_state *cxlds,
+					     void *attach);
+#endif
diff --git a/include/cxl/cxl.h b/include/cxl/cxl.h
index 802b143de83d..f0077ffa0d91 100644
--- a/include/cxl/cxl.h
+++ b/include/cxl/cxl.h
@@ -158,6 +158,7 @@ struct cxl_dpa_partition {
  *
  * @dev: The device associated with this CXL state
  * @cxlmd: The device representing the CXL.mem capabilities of @dev
+ * @cxlcd: The device representing the CXL.cache capabilities of @dev
  * @reg_map: component and ras register mapping parameters
  * @regs: Parsed register blocks
  * @cxl_dvsec: Offset to the PCIe device DVSEC
@@ -175,6 +176,7 @@ struct cxl_dev_state {
 	/* public for Type2 drivers */
 	struct device *dev;
 	struct cxl_memdev *cxlmd;
+	struct cxl_cachedev *cxlcd;
 
 	/* private for Type2 drivers */
 	struct cxl_register_map reg_map;
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 28+ messages in thread

* [PATCH 02/15] cxl/cache: Add cxl_cache driver
  2026-09-23 17:33 [PATCH 00/15] Add initial CXL.cache support Ben Cheatham
  2026-09-23 17:33 ` [PATCH 01/15] cxl/core: Add CXL.cache device struct Ben Cheatham
@ 2026-09-23 17:33 ` Ben Cheatham
  2026-09-23 17:49   ` sashiko-bot
  2026-09-23 17:33 ` [PATCH 03/15] cxl/core: Change cxl_ep_load() to use device pointer parameter Ben Cheatham
                   ` (13 subsequent siblings)
  15 siblings, 1 reply; 28+ messages in thread
From: Ben Cheatham @ 2026-09-23 17:33 UTC (permalink / raw)
  To: linux-cxl, dave, jic23, dave.jiang, alison.schofield
  Cc: linux-iommu, benjamin.cheatham, terry.bowman, robert.richter

Add the cxl_cache driver. This driver will provide management functions
for common portions of CXL.cache capable endpoints, manage struct
cxl_cachedev devices, and validate the system's CXL.cache
configuration.

Add a function for getting the CXL.cache information from the CXL
device DVSEC and store it in the new struct cxl_cache_state member of
cxl_dev_state for use by endpoint drivers.

Add another set of functions to enable/disable CXL.cache on a device.
The cxl_cache driver disables CXL.cache on a device during probe until
the device's configuration can be validated. Validation of the
configuration will come in later commits, so the driver just disables
CXL.cache for now.

Signed-off-by: Ben Cheatham <Benjamin.Cheatham@amd.com>
---
 drivers/cxl/Kconfig           | 14 +++++++
 drivers/cxl/Makefile          |  6 ++-
 drivers/cxl/cache.c           | 61 +++++++++++++++++++++++++++
 drivers/cxl/core/cachedev.c   |  2 +-
 drivers/cxl/core/pci.c        | 78 ++++++++++++++++++++++++++++++++---
 drivers/cxl/cxlcache.h        |  3 ++
 include/cxl/cxl.h             | 27 +++++++++++-
 include/uapi/linux/pci_regs.h |  4 ++
 8 files changed, 186 insertions(+), 9 deletions(-)
 create mode 100644 drivers/cxl/cache.c

diff --git a/drivers/cxl/Kconfig b/drivers/cxl/Kconfig
index 80aeb0d556bd..9c3d3cb6e391 100644
--- a/drivers/cxl/Kconfig
+++ b/drivers/cxl/Kconfig
@@ -243,4 +243,18 @@ config CXL_ATL
 	depends on CXL_REGION
 	depends on ACPI_PRMT && AMD_NB
 
+config CXL_CACHE
+	tristate "CXL: Cache Management Support"
+	depends on CXL_BUS
+	help
+	  Enables a driver that manages the CXL.cache capabilities of a CXL.cache
+	  capable device. This driver validates and provides support for
+	  programming the CXL cache device topology. This driver is required for
+	  using multiple CXL.cache devices (Type 1 or 2) below a given CXL 3.0+
+	  capable PCIe Root Port. This driver only provides CXL cache management and
+	  reporting capabilities, a vendor-specific device driver is expected to
+	  use the device's CXL cache and enable any other capabilities.
+
+	  If unsure, say 'm'.
+
 endif
diff --git a/drivers/cxl/Makefile b/drivers/cxl/Makefile
index 2caa90fa4bf2..6539fe19c08b 100644
--- a/drivers/cxl/Makefile
+++ b/drivers/cxl/Makefile
@@ -4,18 +4,20 @@
 # - 'core' first for fundamental init
 # - 'port' before platform root drivers like 'acpi' so that CXL-root ports
 #   are immediately enabled
-# - 'mem' and 'pmem' before endpoint drivers so that memdevs are
-#   immediately enabled
+# - 'mem', 'pmem', and 'cache' before endpoint drivers so that memdevs and
+#   cachedevs are immediately enabled
 # - 'pci' last, also mirrors the hardware enumeration hierarchy
 obj-y += core/
 obj-$(CONFIG_CXL_PORT) += cxl_port.o
 obj-$(CONFIG_CXL_ACPI) += cxl_acpi.o
 obj-$(CONFIG_CXL_PMEM) += cxl_pmem.o
 obj-$(CONFIG_CXL_MEM) += cxl_mem.o
+obj-$(CONFIG_CXL_CACHE) += cxl_cache.o
 obj-$(CONFIG_CXL_PCI) += cxl_pci.o
 
 cxl_port-y := port.o
 cxl_acpi-y := acpi.o
 cxl_pmem-y := pmem.o security.o
 cxl_mem-y := mem.o
+cxl_cache-y := cache.o
 cxl_pci-y := pci.o
diff --git a/drivers/cxl/cache.c b/drivers/cxl/cache.c
new file mode 100644
index 000000000000..1815ae3a37d5
--- /dev/null
+++ b/drivers/cxl/cache.c
@@ -0,0 +1,61 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Copyright (C) 2026 Advanced Micro Devices, Inc. */
+
+#include <cxl/cxl.h>
+
+#include "cxlcache.h"
+
+/**
+ * DOC: cxl cache
+ *
+ * The cxl_cache driver is responsible for validating the CXL.cache system
+ * configuration and providing a common management interface for the CXL cache
+ * of CXL.cache enabled devices. This driver does not discover devices; a
+ * device-specific driver is required for discovery and portions of set up.
+ */
+
+/**
+ * devm_cxl_add_cachedev - Add a CXL cache device
+ * @cxlds: CXL device state to associate with the cachedev
+ */
+struct cxl_cachedev *devm_cxl_add_cachedev(struct cxl_dev_state *cxlds)
+{
+	return __devm_cxl_add_cachedev(cxlds, NULL);
+}
+EXPORT_SYMBOL_NS_GPL(devm_cxl_add_cachedev, "CXL");
+
+static int cxl_cache_probe(struct device *dev)
+{
+	struct cxl_cachedev *cxlcd = to_cxl_cachedev(dev);
+	struct cxl_dev_state *cxlds = cxlcd->cxlds;
+	int rc;
+
+	/* Disable CXL.cache until we can validate the device configuration */
+	cxl_clear_cache_enable(cxlds);
+
+	rc = cxl_accel_read_cache_info(cxlds);
+	if (rc)
+		return rc;
+
+	return 0;
+}
+
+static struct cxl_driver cxl_cache_driver = {
+	.name = "cxl_cache",
+	.probe = cxl_cache_probe,
+	.drv = {
+		/*
+		 * Needed to guarantee probe and set up order for
+		 * endpoint drivers.
+		 */
+		.probe_type = PROBE_FORCE_SYNCHRONOUS,
+	},
+	.id = CXL_DEVICE_ACCELERATOR,
+};
+
+module_cxl_driver(cxl_cache_driver);
+
+MODULE_DESCRIPTION("CXL: Cache Management");
+MODULE_LICENSE("GPL");
+MODULE_IMPORT_NS("CXL");
+MODULE_ALIAS_CXL(CXL_DEVICE_ACCELERATOR);
diff --git a/drivers/cxl/core/cachedev.c b/drivers/cxl/core/cachedev.c
index 3a7a60165f62..25a24c3bd817 100644
--- a/drivers/cxl/core/cachedev.c
+++ b/drivers/cxl/core/cachedev.c
@@ -116,4 +116,4 @@ struct cxl_cachedev *__devm_cxl_add_cachedev(struct cxl_dev_state *cxlds,
 
 	return cxl_cachedev_autoremove(no_free_ptr(cxlcd));
 }
-EXPORT_SYMBOL_NS_GPL(__devm_cxl_add_cachedev, "CXL");
+EXPORT_SYMBOL_FOR_MODULES(__devm_cxl_add_cachedev, "cxl_cache");
diff --git a/drivers/cxl/core/pci.c b/drivers/cxl/core/pci.c
index 9d807c1a002c..796c4582d2f3 100644
--- a/drivers/cxl/core/pci.c
+++ b/drivers/cxl/core/pci.c
@@ -8,6 +8,7 @@
 #include <linux/pci-doe.h>
 #include <cxl/pci.h>
 #include <linux/aer.h>
+#include <cxlcache.h>
 #include <cxlpci.h>
 #include <cxlmem.h>
 #include <cxl.h>
@@ -180,7 +181,8 @@ int cxl_await_media_ready(struct cxl_dev_state *cxlds)
 }
 EXPORT_SYMBOL_NS_GPL(cxl_await_media_ready, "CXL");
 
-static int cxl_set_mem_enable(struct cxl_dev_state *cxlds, u16 val)
+static int cxl_set_protocol_enable(struct cxl_dev_state *cxlds, u16 val,
+				   u16 enable_bit)
 {
 	struct pci_dev *pdev = to_pci_dev(cxlds->dev);
 	int d = cxlds->cxl_dvsec;
@@ -191,9 +193,9 @@ static int cxl_set_mem_enable(struct cxl_dev_state *cxlds, u16 val)
 	if (rc)
 		return pcibios_err_to_errno(rc);
 
-	if ((ctrl & PCI_DVSEC_CXL_MEM_ENABLE) == val)
+	if ((ctrl & enable_bit) == val)
 		return 1;
-	ctrl &= ~PCI_DVSEC_CXL_MEM_ENABLE;
+	ctrl &= ~enable_bit;
 	ctrl |= val;
 
 	rc = pci_write_config_word(pdev, d + PCI_DVSEC_CXL_CTRL, ctrl);
@@ -205,14 +207,15 @@ static int cxl_set_mem_enable(struct cxl_dev_state *cxlds, u16 val)
 
 static void clear_mem_enable(void *cxlds)
 {
-	cxl_set_mem_enable(cxlds, 0);
+	cxl_set_protocol_enable(cxlds, 0, PCI_DVSEC_CXL_MEM_ENABLE);
 }
 
 static int devm_cxl_enable_mem(struct device *host, struct cxl_dev_state *cxlds)
 {
 	int rc;
 
-	rc = cxl_set_mem_enable(cxlds, PCI_DVSEC_CXL_MEM_ENABLE);
+	rc = cxl_set_protocol_enable(cxlds, PCI_DVSEC_CXL_MEM_ENABLE,
+				     PCI_DVSEC_CXL_MEM_ENABLE);
 	if (rc < 0)
 		return rc;
 	if (rc > 0)
@@ -220,6 +223,31 @@ static int devm_cxl_enable_mem(struct device *host, struct cxl_dev_state *cxlds)
 	return devm_add_action_or_reset(host, clear_mem_enable, cxlds);
 }
 
+static void __clear_cache_enable(void *cxlds)
+{
+	cxl_set_protocol_enable(cxlds, 0, PCI_DVSEC_CXL_CACHE_ENABLE);
+}
+
+void cxl_clear_cache_enable(struct cxl_dev_state *cxlds)
+{
+	__clear_cache_enable(cxlds);
+}
+EXPORT_SYMBOL_FOR_MODULES(cxl_clear_cache_enable, "cxl_cache");
+
+int devm_cxl_enable_cache(struct device *host, struct cxl_dev_state *cxlds)
+{
+	int rc;
+
+	rc = cxl_set_protocol_enable(cxlds, PCI_DVSEC_CXL_CACHE_ENABLE,
+				     PCI_DVSEC_CXL_CACHE_ENABLE);
+	if (rc < 0)
+		return rc;
+	if (rc > 0)
+		return 0;
+	return devm_add_action_or_reset(host, __clear_cache_enable, cxlds);
+}
+EXPORT_SYMBOL_FOR_MODULES(devm_cxl_enable_cache, "cxl_cache");
+
 /* require dvsec ranges to be covered by a locked platform window */
 static int dvsec_range_allowed(struct device *dev, const void *arg)
 {
@@ -927,3 +955,43 @@ int cxl_port_get_possible_dports(struct cxl_port *port)
 
 	return ctx.count;
 }
+
+int cxl_accel_read_cache_info(struct cxl_dev_state *cxlds)
+{
+	struct cxl_cache_state *cstate = &cxlds->cstate;
+	struct pci_dev *pdev = to_pci_dev(cxlds->dev);
+	int dvsec = cxlds->cxl_dvsec;
+	u16 cap, cap2;
+	u32 unit;
+	int rc;
+
+	if (!dev_is_pci(cxlds->dev))
+		return -EINVAL;
+
+	rc = pci_read_config_word(pdev, dvsec + PCI_DVSEC_CXL_CAP, &cap);
+	if (rc)
+		return pcibios_err_to_errno(rc);
+
+	if (!FIELD_GET(PCI_DVSEC_CXL_CACHE_CAPABLE, cap))
+		return -ENXIO;
+
+	rc = pci_read_config_word(pdev, dvsec + PCI_DVSEC_CXL_CAP2, &cap2);
+	if (rc)
+		return pcibios_err_to_errno(rc);
+
+	switch (FIELD_GET(PCI_DVSEC_CXL_CACHE_UNIT, cap2)) {
+	case 1:
+		unit = SZ_64K;
+		break;
+	case 2:
+		unit = SZ_1M;
+		break;
+	default:
+		return -ENXIO;
+	}
+
+	cstate->size = FIELD_GET(PCI_DVSEC_CXL_CACHE_SIZE, cap2) * unit;
+	cstate->unit = unit;
+	return 0;
+}
+EXPORT_SYMBOL_NS_GPL(cxl_accel_read_cache_info, "CXL");
diff --git a/drivers/cxl/cxlcache.h b/drivers/cxl/cxlcache.h
index 6ca887e078a1..be6dac93b91e 100644
--- a/drivers/cxl/cxlcache.h
+++ b/drivers/cxl/cxlcache.h
@@ -34,4 +34,7 @@ static inline struct cxl_cachedev *to_cxl_cachedev(struct device *dev)
 
 struct cxl_cachedev *__devm_cxl_add_cachedev(struct cxl_dev_state *cxlds,
 					     void *attach);
+int cxl_accel_read_cache_info(struct cxl_dev_state *cxlds);
+void cxl_clear_cache_enable(struct cxl_dev_state *cxlds);
+int devm_cxl_enable_cache(struct device *host, struct cxl_dev_state *cxlds);
 #endif
diff --git a/include/cxl/cxl.h b/include/cxl/cxl.h
index f0077ffa0d91..29c0687e4ac2 100644
--- a/include/cxl/cxl.h
+++ b/include/cxl/cxl.h
@@ -149,6 +149,22 @@ struct cxl_dpa_partition {
 
 #define CXL_NR_PARTITIONS_MAX 2
 
+/**
+ * struct cxl_cache_state - CXL cache device state for use by external drivers
+ * @size: Size of device's cache
+ * @unit: Unit of device's cache in bytes
+ */
+struct cxl_cache_state {
+	/* Public for endpoint drivers */
+
+	/* 
+	 * Populated by cxl_cache driver, will be overwritten on 
+	 * cxl_cache::probe()
+	 */
+	u64 size;
+	u64 unit;
+};
+
 /**
  * struct cxl_dev_state - The driver device state
  *
@@ -159,6 +175,7 @@ struct cxl_dpa_partition {
  * @dev: The device associated with this CXL state
  * @cxlmd: The device representing the CXL.mem capabilities of @dev
  * @cxlcd: The device representing the CXL.cache capabilities of @dev
+ * @cstate: CXL.cache information for @dev
  * @reg_map: component and ras register mapping parameters
  * @regs: Parsed register blocks
  * @cxl_dvsec: Offset to the PCIe device DVSEC
@@ -177,6 +194,7 @@ struct cxl_dev_state {
 	struct device *dev;
 	struct cxl_memdev *cxlmd;
 	struct cxl_cachedev *cxlcd;
+ 	struct cxl_cache_state cstate;
 
 	/* private for Type2 drivers */
 	struct cxl_register_map reg_map;
@@ -228,6 +246,13 @@ struct cxl_dev_state *_devm_cxl_dev_state_create(struct device *dev,
 
 struct cxl_memdev *devm_cxl_probe_mem(struct cxl_dev_state *cxlds,
 				      struct range *range);
-
 int cxl_set_capacity(struct cxl_dev_state *cxlds, u64 capacity);
+
+#if IS_ENABLED(CONFIG_CXL_CACHE)
+struct cxl_cachedev *devm_cxl_add_cachedev(struct cxl_dev_state *cxlds);
+#else
+static inline struct cxl_cachedev *
+devm_cxl_add_cachedev(struct cxl_dev_state *cxlds)
+{ return ERR_PTR(-ENXIO); }
+#endif /* CONFIG_CXL_CACHE */
 #endif /* __CXL_CXL_H__ */
diff --git a/include/uapi/linux/pci_regs.h b/include/uapi/linux/pci_regs.h
index facaa324bd86..5383b54ede96 100644
--- a/include/uapi/linux/pci_regs.h
+++ b/include/uapi/linux/pci_regs.h
@@ -1353,7 +1353,11 @@
 #define   PCI_DVSEC_CXL_MEM_CAPABLE			_BITUL(2)
 #define   PCI_DVSEC_CXL_HDM_COUNT			__GENMASK(5, 4)
 #define  PCI_DVSEC_CXL_CTRL				0xC
+#define   PCI_DVSEC_CXL_CACHE_ENABLE			_BITUL(0)
 #define   PCI_DVSEC_CXL_MEM_ENABLE			_BITUL(2)
+#define  PCI_DVSEC_CXL_CAP2				0x16
+#define   PCI_DVSEC_CXL_CACHE_UNIT			__GENMASK(3, 0)
+#define   PCI_DVSEC_CXL_CACHE_SIZE			__GENMASK(15, 8)
 #define  PCI_DVSEC_CXL_RANGE_SIZE_HIGH(i)		(0x18 + (i * 0x10))
 #define  PCI_DVSEC_CXL_RANGE_SIZE_LOW(i)		(0x1C + (i * 0x10))
 #define   PCI_DVSEC_CXL_MEM_INFO_VALID			_BITUL(0)
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 28+ messages in thread

* [PATCH 03/15] cxl/core: Change cxl_ep_load() to use device pointer parameter
  2026-09-23 17:33 [PATCH 00/15] Add initial CXL.cache support Ben Cheatham
  2026-09-23 17:33 ` [PATCH 01/15] cxl/core: Add CXL.cache device struct Ben Cheatham
  2026-09-23 17:33 ` [PATCH 02/15] cxl/cache: Add cxl_cache driver Ben Cheatham
@ 2026-09-23 17:33 ` Ben Cheatham
  2026-09-23 17:33 ` [PATCH 04/15] cxl/core: Update devm_cxl_enumerate_ports() for cxl_cachedevs Ben Cheatham
                   ` (12 subsequent siblings)
  15 siblings, 0 replies; 28+ messages in thread
From: Ben Cheatham @ 2026-09-23 17:33 UTC (permalink / raw)
  To: linux-cxl, dave, jic23, dave.jiang, alison.schofield
  Cc: linux-iommu, benjamin.cheatham, terry.bowman, robert.richter

cxl_ep_load() currently takes a struct cxl_memdev pointer parameter to
find endpoint devices under a port. Change this paramater to use a
device pointer in preparation for adding cxl_cachedev devices to the
CXL port hierarchy.

No functional change intended.

Signed-off-by: Ben Cheatham <Benjamin.Cheatham@amd.com>
---
 drivers/cxl/core/port.c   |  4 ++--
 drivers/cxl/core/region.c | 24 ++++++++++++------------
 drivers/cxl/cxlmem.h      |  4 ++--
 drivers/cxl/port.c        |  2 +-
 4 files changed, 17 insertions(+), 17 deletions(-)

diff --git a/drivers/cxl/core/port.c b/drivers/cxl/core/port.c
index c0f066c0d3d1..6181ec730ec2 100644
--- a/drivers/cxl/core/port.c
+++ b/drivers/cxl/core/port.c
@@ -1504,7 +1504,7 @@ static int port_has_memdev(struct device *dev, const void *data)
 	if (port->depth != ctx->depth)
 		return 0;
 
-	return !!cxl_ep_load(port, ctx->cxlmd);
+	return !!cxl_ep_load(port, &ctx->cxlmd->dev);
 }
 
 static void cxl_detach_ep(void *data)
@@ -1529,7 +1529,7 @@ static void cxl_detach_ep(void *data)
 		parent_port = to_cxl_port(port->dev.parent);
 		device_lock(&parent_port->dev);
 		device_lock(&port->dev);
-		ep = cxl_ep_load(port, cxlmd);
+		ep = cxl_ep_load(port, &cxlmd->dev);
 		dev_dbg(&cxlmd->dev, "disconnect %s from %s\n",
 			ep ? dev_name(ep->ep) : "", dev_name(&port->dev));
 		cxl_ep_remove(port, ep);
diff --git a/drivers/cxl/core/region.c b/drivers/cxl/core/region.c
index 27e63e6dab7c..a8ed7161296c 100644
--- a/drivers/cxl/core/region.c
+++ b/drivers/cxl/core/region.c
@@ -272,8 +272,8 @@ static void cxl_region_decode_reset(struct cxl_region *cxlr, int count)
 		while (!is_cxl_root(to_cxl_port(iter->dev.parent)))
 			iter = to_cxl_port(iter->dev.parent);
 
-		for (ep = cxl_ep_load(iter, cxlmd); iter;
-		     iter = ep->next, ep = cxl_ep_load(iter, cxlmd)) {
+		for (ep = cxl_ep_load(iter, &cxlmd->dev); iter;
+		     iter = ep->next, ep = cxl_ep_load(iter, &cxlmd->dev)) {
 			struct cxl_region_ref *cxl_rr;
 			struct cxl_decoder *cxld;
 
@@ -334,8 +334,8 @@ static int cxl_region_decode_commit(struct cxl_region *cxlr)
 
 		if (rc) {
 			/* programming @iter failed, teardown */
-			for (ep = cxl_ep_load(iter, cxlmd); ep && iter;
-			     iter = ep->next, ep = cxl_ep_load(iter, cxlmd)) {
+			for (ep = cxl_ep_load(iter, &cxlmd->dev); ep && iter;
+			     iter = ep->next, ep = cxl_ep_load(iter, &cxlmd->dev)) {
 				cxl_rr = cxl_rr_load(iter, cxlr);
 				cxld = cxl_rr->decoder;
 				if (cxld->reset)
@@ -1092,7 +1092,7 @@ static int cxl_rr_ep_add(struct cxl_region_ref *cxl_rr,
 	struct cxl_port *port = cxl_rr->port;
 	struct cxl_region *cxlr = cxl_rr->region;
 	struct cxl_decoder *cxld = cxl_rr->decoder;
-	struct cxl_ep *ep = cxl_ep_load(port, cxled_to_memdev(cxled));
+	struct cxl_ep *ep = cxl_ep_load(port, &cxled_to_memdev(cxled)->dev);
 
 	if (ep) {
 		rc = xa_insert(&cxl_rr->endpoints, (unsigned long)cxled, ep,
@@ -1200,7 +1200,7 @@ static int cxl_port_attach_region(struct cxl_port *port,
 				  struct cxl_endpoint_decoder *cxled, int pos)
 {
 	struct cxl_memdev *cxlmd = cxled_to_memdev(cxled);
-	struct cxl_ep *ep = cxl_ep_load(port, cxlmd);
+	struct cxl_ep *ep = cxl_ep_load(port, &cxlmd->dev);
 	struct cxl_region_ref *cxl_rr;
 	bool nr_targets_inc = false;
 	struct cxl_decoder *cxld;
@@ -1377,7 +1377,7 @@ static int check_last_peer(struct cxl_endpoint_decoder *cxled,
 	}
 	cxled_peer = p->targets[pos - distance];
 	cxlmd_peer = cxled_to_memdev(cxled_peer);
-	ep_peer = cxl_ep_load(port, cxlmd_peer);
+	ep_peer = cxl_ep_load(port, &cxlmd_peer->dev);
 	if (ep->dport != ep_peer->dport) {
 		dev_dbg(&cxlr->dev,
 			"%s:%s: %s:%s pos %d mismatched peer %s:%s\n",
@@ -1444,7 +1444,7 @@ static int cxl_port_setup_targets(struct cxl_port *port,
 	struct cxl_port *parent_port = to_cxl_port(port->dev.parent);
 	struct cxl_region_ref *cxl_rr = cxl_rr_load(port, cxlr);
 	struct cxl_memdev *cxlmd = cxled_to_memdev(cxled);
-	struct cxl_ep *ep = cxl_ep_load(port, cxlmd);
+	struct cxl_ep *ep = cxl_ep_load(port, &cxlmd->dev);
 	struct cxl_region_params *p = &cxlr->params;
 	struct cxl_decoder *cxld = cxl_rr->decoder;
 	struct cxl_switch_decoder *cxlsd;
@@ -1692,8 +1692,8 @@ static void cxl_region_teardown_targets(struct cxl_region *cxlr)
 		while (!is_cxl_root(to_cxl_port(iter->dev.parent)))
 			iter = to_cxl_port(iter->dev.parent);
 
-		for (ep = cxl_ep_load(iter, cxlmd); iter;
-		     iter = ep->next, ep = cxl_ep_load(iter, cxlmd))
+		for (ep = cxl_ep_load(iter, &cxlmd->dev); iter;
+		     iter = ep->next, ep = cxl_ep_load(iter, &cxlmd->dev))
 			cxl_port_reset_targets(iter, cxlr);
 	}
 }
@@ -1729,8 +1729,8 @@ static int cxl_region_setup_targets(struct cxl_region *cxlr)
 		 * Descend the topology tree programming / validating
 		 * targets while looking for conflicts.
 		 */
-		for (ep = cxl_ep_load(iter, cxlmd); iter;
-		     iter = ep->next, ep = cxl_ep_load(iter, cxlmd)) {
+		for (ep = cxl_ep_load(iter, &cxlmd->dev); iter;
+		     iter = ep->next, ep = cxl_ep_load(iter, &cxlmd->dev)) {
 			rc = cxl_port_setup_targets(iter, cxlr, cxled);
 			if (rc) {
 				cxl_region_teardown_targets(cxlr);
diff --git a/drivers/cxl/cxlmem.h b/drivers/cxl/cxlmem.h
index c401e3a1af06..a102b8eba5cd 100644
--- a/drivers/cxl/cxlmem.h
+++ b/drivers/cxl/cxlmem.h
@@ -147,12 +147,12 @@ struct cxl_dpa_info {
 int cxl_dpa_setup(struct cxl_dev_state *cxlds, const struct cxl_dpa_info *info);
 
 static inline struct cxl_ep *cxl_ep_load(struct cxl_port *port,
-					 struct cxl_memdev *cxlmd)
+					 struct device *ep_dev)
 {
 	if (!port)
 		return NULL;
 
-	return xa_load(&port->endpoints, (unsigned long)&cxlmd->dev);
+	return xa_load(&port->endpoints, (unsigned long)ep_dev);
 }
 
 /*
diff --git a/drivers/cxl/port.c b/drivers/cxl/port.c
index 99cf77b6b699..edf0ff759fbf 100644
--- a/drivers/cxl/port.c
+++ b/drivers/cxl/port.c
@@ -298,7 +298,7 @@ int devm_cxl_add_endpoint(struct device *host, struct cxl_memdev *cxlmd,
 	     down = iter, iter = to_cxl_port(iter->dev.parent)) {
 		struct cxl_ep *ep;
 
-		ep = cxl_ep_load(iter, cxlmd);
+		ep = cxl_ep_load(iter, &cxlmd->dev);
 		ep->next = down;
 	}
 
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 28+ messages in thread

* [PATCH 04/15] cxl/core: Update devm_cxl_enumerate_ports() for cxl_cachedevs
  2026-09-23 17:33 [PATCH 00/15] Add initial CXL.cache support Ben Cheatham
                   ` (2 preceding siblings ...)
  2026-09-23 17:33 ` [PATCH 03/15] cxl/core: Change cxl_ep_load() to use device pointer parameter Ben Cheatham
@ 2026-09-23 17:33 ` Ben Cheatham
  2026-09-23 17:33 ` [PATCH 05/15] cxl/port: Split endpoint port probe on device type Ben Cheatham
                   ` (11 subsequent siblings)
  15 siblings, 0 replies; 28+ messages in thread
From: Ben Cheatham @ 2026-09-23 17:33 UTC (permalink / raw)
  To: linux-cxl, dave, jic23, dave.jiang, alison.schofield
  Cc: linux-iommu, benjamin.cheatham, terry.bowman, robert.richter

Update devm_cxl_enumerate_ports() to use a struct device pointer instead
of struct cxl_memdev pointer to prepare for adding cxl_cachedevs to the
port hierarchy.

No functional change intended.

Signed-off-by: Ben Cheatham <Benjamin.Cheatham@amd.com>
---
 drivers/cxl/core/port.c | 95 ++++++++++++++++++++++++++++-------------
 drivers/cxl/cxl.h       |  2 +-
 drivers/cxl/mem.c       |  2 +-
 3 files changed, 67 insertions(+), 32 deletions(-)

diff --git a/drivers/cxl/core/port.c b/drivers/cxl/core/port.c
index 6181ec730ec2..506461d5db22 100644
--- a/drivers/cxl/core/port.c
+++ b/drivers/cxl/core/port.c
@@ -839,6 +839,28 @@ static void cxl_debugfs_create_dport_dir(struct cxl_dport *dport)
 			    &cxl_einj_inject_fops);
 }
 
+static struct cxl_dev_state *cxl_ep_to_cxlds(struct device *ep_dev)
+{
+	if (is_cxl_memdev(ep_dev))
+		return to_cxl_memdev(ep_dev)->cxlds;
+
+	if (is_cxl_cachedev(ep_dev))
+		return to_cxl_cachedev(ep_dev)->cxlds;
+
+	return NULL;
+}
+
+static struct cxl_port **cxl_ep_get_endpoint_port(struct device *ep_dev)
+{
+	if (is_cxl_memdev(ep_dev))
+		return &to_cxl_memdev(ep_dev)->endpoint;
+
+	if (is_cxl_cachedev(ep_dev))
+		return &to_cxl_cachedev(ep_dev)->endpoint;
+
+	return NULL;
+}
+
 static int cxl_port_add(struct cxl_port *port,
 			resource_size_t component_reg_phys,
 			struct cxl_dport *parent_dport)
@@ -846,9 +868,10 @@ static int cxl_port_add(struct cxl_port *port,
 	struct device *dev __free(put_device) = &port->dev;
 	int rc;
 
-	if (is_cxl_memdev(port->uport_dev)) {
-		struct cxl_memdev *cxlmd = to_cxl_memdev(port->uport_dev);
-		struct cxl_dev_state *cxlds = cxlmd->cxlds;
+	if (is_cxl_memdev(port->uport_dev) ||
+	    is_cxl_cachedev(port->uport_dev)) {
+		struct cxl_dev_state *cxlds = cxl_ep_to_cxlds(port->uport_dev);
+		struct cxl_port **endpoint = cxl_ep_get_endpoint_port(port->uport_dev);
 
 		rc = dev_set_name(dev, "endpoint%d", port->id);
 		if (rc)
@@ -861,7 +884,7 @@ static int cxl_port_add(struct cxl_port *port,
 		 */
 		port->reg_map = cxlds->reg_map;
 		port->reg_map.host = &port->dev;
-		cxlmd->endpoint = port;
+		*endpoint = port;
 	} else if (parent_dport) {
 		rc = dev_set_name(dev, "port%d", port->id);
 		if (rc)
@@ -1488,11 +1511,11 @@ static void del_dports(struct cxl_port *port)
 }
 
 struct detach_ctx {
-	struct cxl_memdev *cxlmd;
+	struct device *ep_dev;
 	int depth;
 };
 
-static int port_has_memdev(struct device *dev, const void *data)
+static int port_has_ep_dev(struct device *dev, const void *data)
 {
 	const struct detach_ctx *ctx = data;
 	struct cxl_port *port;
@@ -1504,24 +1527,33 @@ static int port_has_memdev(struct device *dev, const void *data)
 	if (port->depth != ctx->depth)
 		return 0;
 
-	return !!cxl_ep_load(port, &ctx->cxlmd->dev);
+	return !!cxl_ep_load(port, ctx->ep_dev);
+}
+
+static int cxl_ep_get_depth(struct device *ep_dev)
+{
+	if (is_cxl_memdev(ep_dev))
+		return to_cxl_memdev(ep_dev)->depth;
+	else
+		return to_cxl_cachedev(ep_dev)->depth;
 }
 
 static void cxl_detach_ep(void *data)
 {
-	struct cxl_memdev *cxlmd = data;
+	struct device *ep_dev = data;
+	int depth = cxl_ep_get_depth(ep_dev);
 
-	for (int i = cxlmd->depth - 1; i >= 1; i--) {
+	for (int i = depth - 1; i >= 1; i--) {
 		struct cxl_port *port, *parent_port;
 		struct detach_ctx ctx = {
-			.cxlmd = cxlmd,
+			.ep_dev = ep_dev,
 			.depth = i,
 		};
 		struct cxl_ep *ep;
 		bool died = false;
 
 		struct device *dev __free(put_device) =
-			bus_find_device(&cxl_bus_type, NULL, &ctx, port_has_memdev);
+			bus_find_device(&cxl_bus_type, NULL, &ctx, port_has_ep_dev);
 		if (!dev)
 			continue;
 		port = to_cxl_port(dev);
@@ -1529,8 +1561,8 @@ static void cxl_detach_ep(void *data)
 		parent_port = to_cxl_port(port->dev.parent);
 		device_lock(&parent_port->dev);
 		device_lock(&port->dev);
-		ep = cxl_ep_load(port, &cxlmd->dev);
-		dev_dbg(&cxlmd->dev, "disconnect %s from %s\n",
+		ep = cxl_ep_load(port, ep_dev);
+		dev_dbg(ep_dev, "disconnect %s from %s\n",
 			ep ? dev_name(ep->ep) : "", dev_name(&port->dev));
 		cxl_ep_remove(port, ep);
 		if (ep && !port->dead && xa_empty(&port->endpoints) &&
@@ -1547,7 +1579,7 @@ static void cxl_detach_ep(void *data)
 		device_unlock(&port->dev);
 
 		if (died) {
-			dev_dbg(&cxlmd->dev, "delete %s\n",
+			dev_dbg(ep_dev, "delete %s\n",
 				dev_name(&port->dev));
 			delete_switch_port(port);
 		}
@@ -1582,8 +1614,8 @@ static int match_port_by_uport(struct device *dev, const void *data)
 		return 0;
 
 	port = to_cxl_port(dev);
-	/* Endpoint ports are hosted by memdevs */
-	if (is_cxl_memdev(port->uport_dev))
+	/* Endpoint ports are hosted by memdevs or cachedevs */
+	if (is_cxl_memdev(port->uport_dev) || is_cxl_cachedev(port->uport_dev))
 		return uport_dev == port->uport_dev->parent;
 	return uport_dev == port->uport_dev;
 }
@@ -1722,7 +1754,7 @@ static struct cxl_dport *devm_cxl_create_port(struct device *ep_dev,
 	return probe_dport(port, dport_dev);
 }
 
-static int add_port_attach_ep(struct cxl_memdev *cxlmd,
+static int add_port_attach_ep(struct device *ep_dev,
 			      struct device *uport_dev,
 			      struct device *dport_dev)
 {
@@ -1736,7 +1768,7 @@ static int add_port_attach_ep(struct cxl_memdev *cxlmd,
 		 * CXL-root 'cxl_port' on a previous iteration, fail for now to
 		 * be re-probed after platform driver attaches.
 		 */
-		dev_dbg(&cxlmd->dev, "%s is a root dport\n",
+		dev_dbg(ep_dev, "%s is a root dport\n",
 			dev_name(dport_dev));
 		return -ENXIO;
 	}
@@ -1756,7 +1788,7 @@ static int add_port_attach_ep(struct cxl_memdev *cxlmd,
 				return PTR_ERR(parent_dport);
 		}
 
-		dport = devm_cxl_create_port(&cxlmd->dev, parent_port,
+		dport = devm_cxl_create_port(ep_dev, parent_port,
 					     parent_dport, uport_dev,
 					     dport_dev);
 		if (IS_ERR(dport)) {
@@ -1767,7 +1799,7 @@ static int add_port_attach_ep(struct cxl_memdev *cxlmd,
 		}
 	}
 
-	rc = cxl_add_ep(dport, &cxlmd->dev);
+	rc = cxl_add_ep(dport, ep_dev);
 	if (rc == -EBUSY) {
 		/*
 		 * "can't" happen, but this error code means
@@ -1807,20 +1839,23 @@ static struct cxl_dport *find_or_add_dport(struct cxl_port *port,
 	return dport;
 }
 
-int devm_cxl_enumerate_ports(struct cxl_memdev *cxlmd)
+int devm_cxl_enumerate_ports(struct device *ep_dev)
 {
-	struct device *dev = &cxlmd->dev;
+	struct cxl_dev_state *cxlds = cxl_ep_to_cxlds(ep_dev);
 	struct device *iter;
 	int rc;
 
+	if (!cxlds)
+		return -EINVAL;
+
 	/*
 	 * Skip intermediate port enumeration in the RCH case, there
 	 * are no ports in between a host bridge and an endpoint.
 	 */
-	if (cxlmd->cxlds->rcd)
+	if (cxlds->rcd)
 		return 0;
 
-	rc = devm_add_action_or_reset(&cxlmd->dev, cxl_detach_ep, cxlmd);
+	rc = devm_add_action_or_reset(ep_dev, cxl_detach_ep, ep_dev);
 	if (rc)
 		return rc;
 
@@ -1830,7 +1865,7 @@ int devm_cxl_enumerate_ports(struct cxl_memdev *cxlmd)
 	 * attempt fails.
 	 */
 retry:
-	for (iter = dev; iter; iter = grandparent(iter)) {
+	for (iter = ep_dev; iter; iter = grandparent(iter)) {
 		struct device *dport_dev = grandparent(iter);
 		struct device *uport_dev;
 		struct cxl_dport *dport;
@@ -1840,18 +1875,18 @@ int devm_cxl_enumerate_ports(struct cxl_memdev *cxlmd)
 
 		uport_dev = dport_dev->parent;
 		if (!uport_dev) {
-			dev_warn(dev, "at %s no parent for dport: %s\n",
+			dev_warn(ep_dev, "at %s no parent for dport: %s\n",
 				 dev_name(iter), dev_name(dport_dev));
 			return -ENXIO;
 		}
 
-		dev_dbg(dev, "scan: iter: %s dport_dev: %s parent: %s\n",
+		dev_dbg(ep_dev, "scan: iter: %s dport_dev: %s parent: %s\n",
 			dev_name(iter), dev_name(dport_dev),
 			dev_name(uport_dev));
 		struct cxl_port *port __free(put_cxl_port) =
 			find_cxl_port_by_uport(uport_dev);
 		if (port) {
-			dev_dbg(&cxlmd->dev,
+			dev_dbg(ep_dev,
 				"found already registered port %s:%s\n",
 				dev_name(&port->dev),
 				dev_name(port->uport_dev));
@@ -1867,7 +1902,7 @@ int devm_cxl_enumerate_ports(struct cxl_memdev *cxlmd)
 				return PTR_ERR(dport);
 			}
 
-			rc = cxl_add_ep(dport, &cxlmd->dev);
+			rc = cxl_add_ep(dport, ep_dev);
 
 			/*
 			 * If the endpoint already exists in the port's list,
@@ -1888,7 +1923,7 @@ int devm_cxl_enumerate_ports(struct cxl_memdev *cxlmd)
 			return 0;
 		}
 
-		rc = add_port_attach_ep(cxlmd, uport_dev, dport_dev);
+		rc = add_port_attach_ep(ep_dev, uport_dev, dport_dev);
 		/* port missing, try to add parent */
 		if (rc == -EAGAIN)
 			continue;
diff --git a/drivers/cxl/cxl.h b/drivers/cxl/cxl.h
index 72d2686a408b..7d369030198a 100644
--- a/drivers/cxl/cxl.h
+++ b/drivers/cxl/cxl.h
@@ -744,7 +744,7 @@ DEFINE_FREE(put_cxl_root_decoder, struct cxl_root_decoder *, if (!IS_ERR_OR_NULL
 DEFINE_FREE(put_cxl_region, struct cxl_region *, if (!IS_ERR_OR_NULL(_T)) put_device(&_T->dev))
 DEFINE_FREE(put_cxl_dax_region, struct cxl_dax_region *, if (!IS_ERR_OR_NULL(_T)) put_device(&_T->dev))
 
-int devm_cxl_enumerate_ports(struct cxl_memdev *cxlmd);
+int devm_cxl_enumerate_ports(struct device *epdev);
 void cxl_bus_rescan(void);
 void cxl_bus_drain(void);
 struct cxl_port *cxl_pci_find_port(struct pci_dev *pdev,
diff --git a/drivers/cxl/mem.c b/drivers/cxl/mem.c
index 798e5c369cfc..b18dcce2e8a6 100644
--- a/drivers/cxl/mem.c
+++ b/drivers/cxl/mem.c
@@ -128,7 +128,7 @@ static int cxl_mem_probe(struct device *dev)
 	if (rc)
 		return rc;
 
-	rc = devm_cxl_enumerate_ports(cxlmd);
+	rc = devm_cxl_enumerate_ports(&cxlmd->dev);
 	if (rc)
 		return rc;
 
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 28+ messages in thread

* [PATCH 05/15] cxl/port: Split endpoint port probe on device type
  2026-09-23 17:33 [PATCH 00/15] Add initial CXL.cache support Ben Cheatham
                   ` (3 preceding siblings ...)
  2026-09-23 17:33 ` [PATCH 04/15] cxl/core: Update devm_cxl_enumerate_ports() for cxl_cachedevs Ben Cheatham
@ 2026-09-23 17:33 ` Ben Cheatham
  2026-09-23 17:46   ` sashiko-bot
  2026-09-23 17:33 ` [PATCH 06/15] cxl/core: Update devm_cxl_add_endpoint() for cxl_cachedevs Ben Cheatham
                   ` (10 subsequent siblings)
  15 siblings, 1 reply; 28+ messages in thread
From: Ben Cheatham @ 2026-09-23 17:33 UTC (permalink / raw)
  To: linux-cxl, dave, jic23, dave.jiang, alison.schofield
  Cc: linux-iommu, benjamin.cheatham, terry.bowman, robert.richter

CXL.cache devices (struct cxl_cachedev) don't support or need all of the
set up done by endpoint port probe for CXL.mem devices (struct
cxl_memdev). Split endpoint port probe on device type and refactor dport
RAS set up into a common routine. The CXL.cache endpoint probe function
will be used when cxl_cachedevs are added to the port heirarchy in a
later commit.

No functional change intended for cxl_memdev path.

Signed-off-by: Ben Cheatham <Benjamin.Cheatham@amd.com>
---
 drivers/cxl/cache.c         |  4 +++
 drivers/cxl/core/cachedev.c | 11 ++++++
 drivers/cxl/core/port.c     |  6 ++++
 drivers/cxl/cxl.h           |  1 +
 drivers/cxl/cxlcache.h      |  2 ++
 drivers/cxl/port.c          | 69 ++++++++++++++++++++++++++-----------
 6 files changed, 73 insertions(+), 20 deletions(-)

diff --git a/drivers/cxl/cache.c b/drivers/cxl/cache.c
index 1815ae3a37d5..2ab783f7c365 100644
--- a/drivers/cxl/cache.c
+++ b/drivers/cxl/cache.c
@@ -33,6 +33,10 @@ static int cxl_cache_probe(struct device *dev)
 	/* Disable CXL.cache until we can validate the device configuration */
 	cxl_clear_cache_enable(cxlds);
 
+	/* See comment in cxl_mem_probe() */
+	if (work_pending(&cxlcd->detach_work))
+		return -EBUSY;
+
 	rc = cxl_accel_read_cache_info(cxlds);
 	if (rc)
 		return rc;
diff --git a/drivers/cxl/core/cachedev.c b/drivers/cxl/core/cachedev.c
index 25a24c3bd817..2bd69133c631 100644
--- a/drivers/cxl/core/cachedev.c
+++ b/drivers/cxl/core/cachedev.c
@@ -43,6 +43,16 @@ bool is_cxl_cachedev(const struct device *dev)
 }
 EXPORT_SYMBOL_NS_GPL(is_cxl_cachedev, "CXL");
 
+static void detach_cachedev(struct work_struct *work)
+{
+	struct cxl_cachedev *cxlcd;
+
+	cxlcd = container_of(work, typeof(*cxlcd), detach_work);
+
+	device_release_driver(&cxlcd->dev);
+	put_device(&cxlcd->dev);
+}
+
 static struct lock_class_key cxl_cachedev_key;
 
 static struct cxl_cachedev *cxl_cachedev_alloc(struct cxl_dev_state *cxlds)
@@ -70,6 +80,7 @@ static struct cxl_cachedev *cxl_cachedev_alloc(struct cxl_dev_state *cxlds)
 	dev->bus = &cxl_bus_type;
 	dev->type = &cxl_cachedev_type;
 	device_set_pm_not_required(dev);
+	INIT_WORK(&cxlcd->detach_work, detach_cachedev);
 
 	return_ptr(cxlcd);
 }
diff --git a/drivers/cxl/core/port.c b/drivers/cxl/core/port.c
index 506461d5db22..125e175b1c0d 100644
--- a/drivers/cxl/core/port.c
+++ b/drivers/cxl/core/port.c
@@ -2356,6 +2356,12 @@ bool schedule_cxl_memdev_detach(struct cxl_memdev *cxlmd)
 }
 EXPORT_SYMBOL_NS_GPL(schedule_cxl_memdev_detach, "CXL");
 
+bool schedule_cxl_cachedev_detach(struct cxl_cachedev *cxlcd)
+{
+	return queue_work(cxl_bus_wq, &cxlcd->detach_work);
+}
+EXPORT_SYMBOL_NS_GPL(schedule_cxl_cachedev_detach, "CXL");
+
 static void add_latency(struct access_coordinate *c, long latency)
 {
 	for (int i = 0; i < ACCESS_COORDINATE_MAX; i++) {
diff --git a/drivers/cxl/cxl.h b/drivers/cxl/cxl.h
index 7d369030198a..5cc2fe844396 100644
--- a/drivers/cxl/cxl.h
+++ b/drivers/cxl/cxl.h
@@ -752,6 +752,7 @@ struct cxl_port *cxl_pci_find_port(struct pci_dev *pdev,
 struct cxl_port *cxl_mem_find_port(struct cxl_memdev *cxlmd,
 				   struct cxl_dport **dport);
 bool schedule_cxl_memdev_detach(struct cxl_memdev *cxlmd);
+bool schedule_cxl_cachedev_detach(struct cxl_cachedev *cxlcd);
 
 struct cxl_dport *devm_cxl_add_dport(struct cxl_port *port,
 				     struct device *dport, int port_id,
diff --git a/drivers/cxl/cxlcache.h b/drivers/cxl/cxlcache.h
index be6dac93b91e..e9a4567b7f2c 100644
--- a/drivers/cxl/cxlcache.h
+++ b/drivers/cxl/cxlcache.h
@@ -9,6 +9,7 @@
  * a CXL device
  * @dev: driver core device object
  * @cxlds: device state backing this device
+ * @detach_work: active cachedev lost a port in its ancestry
  * @endpoint: connection to the CXL port topology for this device
  * @id: id number of this cachedev instance
  * @depth: endpoint port depth in hierarchy
@@ -16,6 +17,7 @@
 struct cxl_cachedev {
 	struct device dev;
 	struct cxl_dev_state *cxlds;
+	struct work_struct detach_work;
 	struct cxl_port *endpoint;
 	int id;
 	int depth;
diff --git a/drivers/cxl/port.c b/drivers/cxl/port.c
index edf0ff759fbf..7c93fabfb095 100644
--- a/drivers/cxl/port.c
+++ b/drivers/cxl/port.c
@@ -5,6 +5,7 @@
 #include <linux/module.h>
 #include <linux/slab.h>
 
+#include "cxlcache.h"
 #include "cxlmem.h"
 #include "cxlpci.h"
 
@@ -26,9 +27,13 @@
  * PCIe topology.
  */
 
-static void schedule_detach(void *cxlmd)
+static void schedule_detach(void *ep_dev)
 {
-	schedule_cxl_memdev_detach(cxlmd);
+	if (is_cxl_memdev(ep_dev))
+		schedule_cxl_memdev_detach(ep_dev);
+
+	if (is_cxl_cachedev(ep_dev))
+		schedule_cxl_cachedev_detach(ep_dev);
 }
 
 static int discover_region(struct device *dev, void *unused)
@@ -118,24 +123,9 @@ static int cxl_ras_unmask(struct cxl_port *port)
 	return 0;
 }
 
-static int cxl_endpoint_port_probe(struct cxl_port *port)
+static void cxl_endpoint_setup_dport_ras(struct cxl_port *port)
 {
-	struct cxl_memdev *cxlmd = to_cxl_memdev(port->uport_dev);
 	struct cxl_dport *dport = port->parent_dport;
-	int rc;
-
-	/* Cache the data early to ensure is_visible() works */
-	read_cdat_data(port);
-	cxl_endpoint_parse_cdat(port);
-
-	get_device(&cxlmd->dev);
-	rc = devm_add_action_or_reset(&port->dev, schedule_detach, cxlmd);
-	if (rc)
-		return rc;
-
-	rc = devm_cxl_endpoint_decoders_setup(port);
-	if (rc)
-		return rc;
 
 	/*
 	 * With VH (CXL Virtual Host) topology the cxl_port::add_dport() method
@@ -151,6 +141,27 @@ static int cxl_endpoint_port_probe(struct cxl_port *port)
 	devm_cxl_port_ras_setup(port);
 	if (cxl_ras_unmask(port))
 		dev_dbg(&port->dev, "failed to unmask RAS interrupts\n");
+}
+
+static int cxl_mem_endpoint_port_probe(struct cxl_port *port)
+{
+	struct cxl_memdev *cxlmd = to_cxl_memdev(port->uport_dev);
+	int rc;
+
+	/* Cache the data early to ensure is_visible() works */
+	read_cdat_data(port);
+	cxl_endpoint_parse_cdat(port);
+
+	get_device(&cxlmd->dev);
+	rc = devm_add_action_or_reset(&port->dev, schedule_detach, &cxlmd->dev);
+	if (rc)
+		return rc;
+
+	rc = devm_cxl_endpoint_decoders_setup(port);
+	if (rc)
+		return rc;
+
+	cxl_endpoint_setup_dport_ras(port);
 
 	/*
 	 * Now that all endpoint decoders are successfully enumerated, try to
@@ -161,12 +172,30 @@ static int cxl_endpoint_port_probe(struct cxl_port *port)
 	return 0;
 }
 
+static int cxl_cache_endpoint_port_probe(struct cxl_port *port)
+{
+	struct cxl_cachedev *cxlcd = to_cxl_cachedev(port->uport_dev);
+	int rc;
+
+	get_device(&cxlcd->dev);
+	rc = devm_add_action_or_reset(&port->dev, schedule_detach,
+				      &cxlcd->dev);
+	if (rc)
+		return rc;
+
+	cxl_endpoint_setup_dport_ras(port);
+
+	return rc;
+}
+
 static int cxl_port_probe(struct device *dev)
 {
 	struct cxl_port *port = to_cxl_port(dev);
 
-	if (is_cxl_endpoint(port))
-		return cxl_endpoint_port_probe(port);
+	if (is_cxl_memdev(port->uport_dev))
+		return cxl_mem_endpoint_port_probe(port);
+	else if (is_cxl_cachedev(port->uport_dev))
+		return cxl_cache_endpoint_port_probe(port);
 	return cxl_switch_port_probe(port);
 }
 
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 28+ messages in thread

* [PATCH 06/15] cxl/core: Update devm_cxl_add_endpoint() for cxl_cachedevs
  2026-09-23 17:33 [PATCH 00/15] Add initial CXL.cache support Ben Cheatham
                   ` (4 preceding siblings ...)
  2026-09-23 17:33 ` [PATCH 05/15] cxl/port: Split endpoint port probe on device type Ben Cheatham
@ 2026-09-23 17:33 ` Ben Cheatham
  2026-09-23 17:51   ` sashiko-bot
  2026-09-23 17:33 ` [PATCH 07/15] cxl/cache: Verify port hierarchy has CXL.cache enabled Ben Cheatham
                   ` (9 subsequent siblings)
  15 siblings, 1 reply; 28+ messages in thread
From: Ben Cheatham @ 2026-09-23 17:33 UTC (permalink / raw)
  To: linux-cxl, dave, jic23, dave.jiang, alison.schofield
  Cc: linux-iommu, benjamin.cheatham, terry.bowman, robert.richter

Update devm_cxl_add_endpoint() to allow for cxl_cachedevs as well as
cxl_memdevs. Add cxl_cachedevs to the port heirarchy.

Signed-off-by: Ben Cheatham <Benjamin.Cheatham@amd.com>
---
 drivers/cxl/cache.c     | 30 ++++++++++++++++++++++++++++++
 drivers/cxl/core/port.c | 36 ++++++++++++++++++++++++++++--------
 drivers/cxl/cxl.h       |  6 ++++--
 drivers/cxl/mem.c       |  2 +-
 drivers/cxl/port.c      | 12 ++++++------
 5 files changed, 69 insertions(+), 17 deletions(-)

diff --git a/drivers/cxl/cache.c b/drivers/cxl/cache.c
index 2ab783f7c365..dea5af7b2d3f 100644
--- a/drivers/cxl/cache.c
+++ b/drivers/cxl/cache.c
@@ -28,6 +28,8 @@ static int cxl_cache_probe(struct device *dev)
 {
 	struct cxl_cachedev *cxlcd = to_cxl_cachedev(dev);
 	struct cxl_dev_state *cxlds = cxlcd->cxlds;
+	struct device *endpoint_parent;
+	struct cxl_dport *dport;
 	int rc;
 
 	/* Disable CXL.cache until we can validate the device configuration */
@@ -41,6 +43,34 @@ static int cxl_cache_probe(struct device *dev)
 	if (rc)
 		return rc;
 
+	rc = devm_cxl_enumerate_ports(&cxlcd->dev);
+	if (rc)
+		return rc;
+
+	struct cxl_port *parent_port __free(put_cxl_port) =
+		cxl_cache_find_port(cxlcd, &dport);
+	if (!parent_port) {
+		dev_err(dev, "CXL port topology not found\n");
+		return -ENXIO;
+	}
+
+	if (dport->rch)
+		endpoint_parent = parent_port->uport_dev;
+	else
+		endpoint_parent = &parent_port->dev;
+
+	scoped_guard(device, endpoint_parent) {
+		if (!endpoint_parent->driver) {
+			dev_err(dev, "CXL port topology %s not enabled\n",
+				dev_name(endpoint_parent));
+			return -ENXIO;
+		}
+
+		rc = devm_cxl_add_endpoint(endpoint_parent, &cxlcd->dev, dport);
+		if (rc)
+			return rc;
+	}
+
 	return 0;
 }
 
diff --git a/drivers/cxl/core/port.c b/drivers/cxl/core/port.c
index 125e175b1c0d..f6e981d87088 100644
--- a/drivers/cxl/core/port.c
+++ b/drivers/cxl/core/port.c
@@ -1453,10 +1453,9 @@ static struct device *grandparent(struct device *dev)
 	return NULL;
 }
 
-static void delete_endpoint(void *data)
+static void __delete_endpoint(struct cxl_port **ep_port)
 {
-	struct cxl_memdev *cxlmd = data;
-	struct cxl_port *endpoint = cxlmd->endpoint;
+	struct cxl_port *endpoint = *ep_port;
 	struct device *host = port_to_host(endpoint);
 
 	scoped_guard(device, host) {
@@ -1465,21 +1464,35 @@ static void delete_endpoint(void *data)
 			devm_release_action(host, cxl_unlink_uport, endpoint);
 			devm_release_action(host, unregister_port, endpoint);
 		}
-		cxlmd->endpoint = NULL;
+		*ep_port = NULL;
 	}
 	put_device(&endpoint->dev);
 	put_device(host);
 }
 
-int cxl_endpoint_autoremove(struct cxl_memdev *cxlmd, struct cxl_port *endpoint)
+static void delete_endpoint(void *data)
+{
+	struct device *ep_dev = data;
+
+	if (is_cxl_memdev(ep_dev))
+		__delete_endpoint(&to_cxl_memdev(ep_dev)->endpoint);
+	else
+		__delete_endpoint(&to_cxl_cachedev(ep_dev)->endpoint);
+}
+
+int cxl_endpoint_autoremove(struct device *ep_dev, struct cxl_port *endpoint)
 {
 	struct device *host = port_to_host(endpoint);
-	struct device *dev = &cxlmd->dev;
 
 	get_device(host);
 	get_device(&endpoint->dev);
-	cxlmd->depth = endpoint->depth;
-	return devm_add_action_or_reset(dev, delete_endpoint, cxlmd);
+
+	if (is_cxl_memdev(ep_dev))
+		to_cxl_memdev(ep_dev)->depth = endpoint->depth;
+	else
+		to_cxl_cachedev(ep_dev)->depth = endpoint->depth;
+
+	return devm_add_action_or_reset(ep_dev, delete_endpoint, ep_dev);
 }
 EXPORT_SYMBOL_NS_GPL(cxl_endpoint_autoremove, "CXL");
 
@@ -1952,6 +1965,13 @@ struct cxl_port *cxl_mem_find_port(struct cxl_memdev *cxlmd,
 }
 EXPORT_SYMBOL_NS_GPL(cxl_mem_find_port, "CXL");
 
+struct cxl_port *cxl_cache_find_port(struct cxl_cachedev *cxlcd,
+				     struct cxl_dport **dport)
+{
+	return find_cxl_port_by_dport(grandparent(&cxlcd->dev), dport);
+}
+EXPORT_SYMBOL_NS_GPL(cxl_cache_find_port, "CXL");
+
 static int decoder_populate_targets(struct cxl_switch_decoder *cxlsd,
 				    struct cxl_port *port)
 {
diff --git a/drivers/cxl/cxl.h b/drivers/cxl/cxl.h
index 5cc2fe844396..9711b8499508 100644
--- a/drivers/cxl/cxl.h
+++ b/drivers/cxl/cxl.h
@@ -734,7 +734,7 @@ struct cxl_port *devm_cxl_add_port(struct device *host,
 				   resource_size_t component_reg_phys,
 				   struct cxl_dport *parent_dport);
 struct cxl_root *devm_cxl_add_root(struct device *host);
-int devm_cxl_add_endpoint(struct device *host, struct cxl_memdev *cxlmd,
+int devm_cxl_add_endpoint(struct device *host, struct device *ep_dev,
 			  struct cxl_dport *parent_dport);
 struct cxl_root *find_cxl_root(struct cxl_port *port);
 
@@ -751,6 +751,8 @@ struct cxl_port *cxl_pci_find_port(struct pci_dev *pdev,
 				   struct cxl_dport **dport);
 struct cxl_port *cxl_mem_find_port(struct cxl_memdev *cxlmd,
 				   struct cxl_dport **dport);
+struct cxl_port *cxl_cache_find_port(struct cxl_cachedev *cxlcd,
+				     struct cxl_dport **dport);
 bool schedule_cxl_memdev_detach(struct cxl_memdev *cxlmd);
 bool schedule_cxl_cachedev_detach(struct cxl_cachedev *cxlcd);
 
@@ -788,7 +790,7 @@ static inline int cxl_root_decoder_autoremove(struct device *host,
 {
 	return cxl_decoder_autoremove(host, &cxlrd->cxlsd.cxld);
 }
-int cxl_endpoint_autoremove(struct cxl_memdev *cxlmd, struct cxl_port *endpoint);
+int cxl_endpoint_autoremove(struct device *ep_dev, struct cxl_port *endpoint);
 
 /**
  * struct cxl_endpoint_dvsec_info - Cached DVSEC info
diff --git a/drivers/cxl/mem.c b/drivers/cxl/mem.c
index b18dcce2e8a6..178ca937ee82 100644
--- a/drivers/cxl/mem.c
+++ b/drivers/cxl/mem.c
@@ -160,7 +160,7 @@ static int cxl_mem_probe(struct device *dev)
 			return -ENXIO;
 		}
 
-		rc = devm_cxl_add_endpoint(endpoint_parent, cxlmd, dport);
+		rc = devm_cxl_add_endpoint(endpoint_parent, &cxlmd->dev, dport);
 		if (rc)
 			return rc;
 	}
diff --git a/drivers/cxl/port.c b/drivers/cxl/port.c
index 7c93fabfb095..44d978d4205a 100644
--- a/drivers/cxl/port.c
+++ b/drivers/cxl/port.c
@@ -312,7 +312,7 @@ static struct cxl_driver cxl_port_driver = {
 	},
 };
 
-int devm_cxl_add_endpoint(struct device *host, struct cxl_memdev *cxlmd,
+int devm_cxl_add_endpoint(struct device *host, struct device *ep_dev,
 			  struct cxl_dport *parent_dport)
 {
 	struct cxl_port *parent_port = parent_dport->port;
@@ -327,29 +327,29 @@ int devm_cxl_add_endpoint(struct device *host, struct cxl_memdev *cxlmd,
 	     down = iter, iter = to_cxl_port(iter->dev.parent)) {
 		struct cxl_ep *ep;
 
-		ep = cxl_ep_load(iter, &cxlmd->dev);
+		ep = cxl_ep_load(iter, ep_dev);
 		ep->next = down;
 	}
 
 	/* Note: endpoint port component registers are derived from @cxlds */
-	endpoint = devm_cxl_add_port(host, &cxlmd->dev, CXL_RESOURCE_NONE,
+	endpoint = devm_cxl_add_port(host, ep_dev, CXL_RESOURCE_NONE,
 				     parent_dport);
 	if (IS_ERR(endpoint))
 		return PTR_ERR(endpoint);
 
-	rc = cxl_endpoint_autoremove(cxlmd, endpoint);
+	rc = cxl_endpoint_autoremove(ep_dev, endpoint);
 	if (rc)
 		return rc;
 
 	if (!endpoint->dev.driver) {
-		dev_err(&cxlmd->dev, "%s failed probe\n",
+		dev_err(ep_dev, "%s failed probe\n",
 			dev_name(&endpoint->dev));
 		return -ENXIO;
 	}
 
 	return 0;
 }
-EXPORT_SYMBOL_FOR_MODULES(devm_cxl_add_endpoint, "cxl_mem");
+EXPORT_SYMBOL_FOR_MODULES(devm_cxl_add_endpoint, "cxl_mem,cxl_cache");
 
 static int __init cxl_port_init(void)
 {
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 28+ messages in thread

* [PATCH 07/15] cxl/cache: Verify port hierarchy has CXL.cache enabled
  2026-09-23 17:33 [PATCH 00/15] Add initial CXL.cache support Ben Cheatham
                   ` (5 preceding siblings ...)
  2026-09-23 17:33 ` [PATCH 06/15] cxl/core: Update devm_cxl_add_endpoint() for cxl_cachedevs Ben Cheatham
@ 2026-09-23 17:33 ` Ben Cheatham
  2026-09-23 17:33 ` [PATCH 08/15] cxl/core, cache: Add Cache ID register probing and init Ben Cheatham
                   ` (8 subsequent siblings)
  15 siblings, 0 replies; 28+ messages in thread
From: Ben Cheatham @ 2026-09-23 17:33 UTC (permalink / raw)
  To: linux-cxl, dave, jic23, dave.jiang, alison.schofield
  Cc: linux-iommu, benjamin.cheatham, terry.bowman, robert.richter

Check that the ports above a CXL.cache device support CXL.cache and have
the protocol enabled. The device's CXL.cache enable bit will be set in a
later commit.

CXL.cache should be enabled as part of link negotiation by firmware
during boot. Upstream ports can't be programmed after boot, but
downstream ports can. Reprogramming downstream ports requires a link
reset and is may have implications for the underlying PCI device, so
fail probe instead of reprogramming the downstream port.

Signed-off-by: Ben Cheatham <Benjamin.Cheatham@amd.com>
---
 drivers/cxl/cache.c           | 76 ++++++++++++++++++++++++++++++++++-
 include/uapi/linux/pci_regs.h |  4 ++
 2 files changed, 79 insertions(+), 1 deletion(-)

diff --git a/drivers/cxl/cache.c b/drivers/cxl/cache.c
index dea5af7b2d3f..dfefd301696a 100644
--- a/drivers/cxl/cache.c
+++ b/drivers/cxl/cache.c
@@ -1,6 +1,7 @@
 // SPDX-License-Identifier: GPL-2.0-only
 /* Copyright (C) 2026 Advanced Micro Devices, Inc. */
 
+#include <linux/pci.h>
 #include <cxl/cxl.h>
 
 #include "cxlcache.h"
@@ -14,6 +15,75 @@
  * device-specific driver is required for discovery and portions of set up.
  */
 
+static bool cxl_flexbus_cache_enabled(struct device *host)
+{
+
+	u16 dvsec, cap, status;
+	struct pci_dev *pdev;
+	int rc;
+
+	if (!dev_is_pci(host))
+		return false;
+	pdev = to_pci_dev(host);
+
+	dvsec = pci_find_dvsec_capability(pdev, PCI_VENDOR_ID_CXL,
+					  PCI_DVSEC_CXL_FLEXBUS_PORT);
+	if (!dvsec)
+		return false;
+
+	rc = pci_read_config_word(pdev,
+				  dvsec + PCI_DVSEC_CXL_FLEXBUS_PORT_CAPABILITY,
+				  &cap);
+	if (rc)
+		return false;
+
+	if (!FIELD_GET(PCI_DVSEC_CXL_FLEXBUS_PORT_CAP_CACHE, cap))
+		return false;
+
+	rc = pci_read_config_word(pdev,
+				  dvsec + PCI_DVSEC_CXL_FLEXBUS_PORT_STATUS,
+				  &status);
+	if (rc)
+		return false;
+
+	return FIELD_GET(PCI_DVSEC_CXL_FLEXBUS_PORT_STATUS_CACHE, status);
+}
+
+static int cxl_endpoint_cache_enabled(struct cxl_port *endpoint)
+{
+	struct cxl_dport *dport_iter = endpoint->parent_dport;
+	struct cxl_port *port_iter = dport_iter->port;
+
+	while (!is_cxl_root(port_iter)) {
+		/* 
+		 * CXL host bridge isn't a PCI device, but we still need to
+		 * check the PCIe root port
+		 */
+		if (dev_is_pci(port_iter->uport_dev) &&
+		    !cxl_flexbus_cache_enabled(port_iter->uport_dev)) {
+			dev_dbg(port_iter->uport_dev,
+				"CXL.cache not supported or enabled\n");
+			return -ENXIO;
+		}
+
+		/*
+		 * Cache can be enabled for dports, but it requires a link
+		 * reset. Could be done here, but should probably be done
+		 * by the endpoint's driver.
+		 */
+		if (!cxl_flexbus_cache_enabled(dport_iter->dport_dev)) {
+			dev_dbg(dport_iter->dport_dev,
+				"CXL.cache not supported or enabled\n");
+			return -ENXIO;
+		}
+
+		dport_iter = port_iter->parent_dport;
+		port_iter = dport_iter->port;
+	}
+
+	return 0;
+}
+
 /**
  * devm_cxl_add_cachedev - Add a CXL cache device
  * @cxlds: CXL device state to associate with the cachedev
@@ -71,7 +141,11 @@ static int cxl_cache_probe(struct device *dev)
 			return rc;
 	}
 
-	return 0;
+	rc = cxl_endpoint_cache_enabled(cxlcd->endpoint);
+	if (rc)
+		dev_err(dev, "CXL.cache not enabled on parent port(s)");
+
+	return rc;
 }
 
 static struct cxl_driver cxl_cache_driver = {
diff --git a/include/uapi/linux/pci_regs.h b/include/uapi/linux/pci_regs.h
index 5383b54ede96..9146ee574de8 100644
--- a/include/uapi/linux/pci_regs.h
+++ b/include/uapi/linux/pci_regs.h
@@ -1392,6 +1392,10 @@
 
 /* CXL r4.0, 8.1.8: Flex Bus DVSEC */
 #define PCI_DVSEC_CXL_FLEXBUS_PORT			7
+#define  PCI_DVSEC_CXL_FLEXBUS_PORT_CAPABILITY		0xA
+#define   PCI_DVSEC_CXL_FLEXBUS_PORT_CAP_CACHE		_BITUL(0)
+#define  PCI_DVSEC_CXL_FLEXBUS_PORT_CONTROL		0xC
+#define   PCI_DVSEC_CXL_FLEXBUS_PORT_CTRL_CACHE_EN	_BITUL(0)
 #define  PCI_DVSEC_CXL_FLEXBUS_PORT_STATUS		0xE
 #define   PCI_DVSEC_CXL_FLEXBUS_PORT_STATUS_CACHE	_BITUL(0)
 #define   PCI_DVSEC_CXL_FLEXBUS_PORT_STATUS_MEM		_BITUL(2)
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 28+ messages in thread

* [PATCH 08/15] cxl/core, cache: Add Cache ID register probing and init
  2026-09-23 17:33 [PATCH 00/15] Add initial CXL.cache support Ben Cheatham
                   ` (6 preceding siblings ...)
  2026-09-23 17:33 ` [PATCH 07/15] cxl/cache: Verify port hierarchy has CXL.cache enabled Ben Cheatham
@ 2026-09-23 17:33 ` Ben Cheatham
  2026-09-23 17:46   ` sashiko-bot
  2026-09-23 17:33 ` [PATCH 09/15] cxl/core: Add Cache ID verification Ben Cheatham
                   ` (7 subsequent siblings)
  15 siblings, 1 reply; 28+ messages in thread
From: Ben Cheatham @ 2026-09-23 17:33 UTC (permalink / raw)
  To: linux-cxl, dave, jic23, dave.jiang, alison.schofield
  Cc: linux-iommu, benjamin.cheatham, terry.bowman, robert.richter

The CXL Cache ID Route Table capability (CXL 4.0 8.2.4.28) and CXL Cache
ID Decoder capability (CXL 4.0 8.2.4.29) are optional capabilities that
are required when more than one CXL.cache device is enabled under a host
bridge.

Probe for cache id route table and decoder registers as part of
cxl_probe_component_regs(). Map these registers, if found, as part of
cxl_port probe.

Signed-off-by: Ben Cheatham <Benjamin.Cheatham@amd.com>
---
 drivers/cxl/core/Makefile |  1 +
 drivers/cxl/core/cache.c  | 41 +++++++++++++++++++++++++++++++++++++++
 drivers/cxl/core/port.c   |  5 +++++
 drivers/cxl/core/regs.c   | 23 ++++++++++++++++++++++
 drivers/cxl/cxl.h         | 21 ++++++++++++++++++++
 drivers/cxl/port.c        |  6 ++++++
 include/cxl/cxl.h         |  6 ++++++
 7 files changed, 103 insertions(+)
 create mode 100644 drivers/cxl/core/cache.c

diff --git a/drivers/cxl/core/Makefile b/drivers/cxl/core/Makefile
index f51575abe5af..33420aaf40db 100644
--- a/drivers/cxl/core/Makefile
+++ b/drivers/cxl/core/Makefile
@@ -15,6 +15,7 @@ cxl_core-y += pci.o
 cxl_core-y += hdm.o
 cxl_core-y += pmu.o
 cxl_core-y += cdat.o
+cxl_core-y += cache.o
 cxl_core-$(CONFIG_TRACING) += trace.o
 cxl_core-$(CONFIG_CXL_REGION) += region.o region_pmem.o region_dax.o
 cxl_core-$(CONFIG_CXL_MCE) += mce.o
diff --git a/drivers/cxl/core/cache.c b/drivers/cxl/core/cache.c
new file mode 100644
index 000000000000..17df0e603b68
--- /dev/null
+++ b/drivers/cxl/core/cache.c
@@ -0,0 +1,41 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Copyright (C) Advanced Micro Device, Inc. */
+
+#include <linux/pci.h>
+#include <cxlcache.h>
+
+#include "cxlpci.h"
+
+int cxl_port_map_cache_id_rt(struct cxl_port *port)
+{
+	struct cxl_register_map *map = &port->reg_map;
+
+	if (is_cxl_root(port) || is_cxl_cachedev(port->uport_dev))
+		return -EINVAL;
+
+	if (!dev_is_pci(port->uport_dev) ||
+	    !cxl_pci_flit_256(to_pci_dev(port->uport_dev)))
+		return -EINVAL;
+
+	if (!map->component_map.cidrt.valid)
+		return -ENXIO;
+
+	return cxl_map_component_regs(map, &port->regs,
+				      BIT(CXL_CM_CAP_CAP_ID_CACHE_ID_RT));
+}
+EXPORT_SYMBOL_FOR_MODULES(cxl_port_map_cache_id_rt, "cxl_port");
+
+int cxl_dport_map_cache_id_dc(struct cxl_dport *dport)
+{
+	struct cxl_register_map *map = &dport->reg_map;
+	struct cxl_port *port = dport->port;
+
+	if (is_cxl_root(port) || is_cxl_cachedev(port->uport_dev))
+		return -EINVAL;
+
+	if (!map->component_map.ciddc.valid)
+		return -ENXIO;
+
+	return cxl_map_component_regs(map, &dport->regs.component,
+				      BIT(CXL_CM_CAP_CAP_ID_CACHE_ID_DC));
+}
diff --git a/drivers/cxl/core/port.c b/drivers/cxl/core/port.c
index f6e981d87088..6da11d28bb5e 100644
--- a/drivers/cxl/core/port.c
+++ b/drivers/cxl/core/port.c
@@ -1271,6 +1271,11 @@ __devm_cxl_add_dport(struct cxl_port *port, struct device *dport_dev,
 	if (!dport->rch)
 		devm_cxl_dport_ras_setup(dport);
 
+	rc = cxl_dport_map_cache_id_dc(dport);
+	if (rc)
+		dev_dbg(dport->dport_dev,
+			"Failed to map cache id decoder capability: %d\n", rc);
+
 	/* keep the group, and mark the end of devm actions */
 	cxl_dport_close_dr_group(dport, no_free_ptr(dport_dr_group));
 
diff --git a/drivers/cxl/core/regs.c b/drivers/cxl/core/regs.c
index 20c2d9fbcfe7..2809c72fb2eb 100644
--- a/drivers/cxl/core/regs.c
+++ b/drivers/cxl/core/regs.c
@@ -93,6 +93,27 @@ void cxl_probe_component_regs(struct device *dev, void __iomem *base,
 			length = CXL_RAS_CAPABILITY_LENGTH;
 			rmap = &map->ras;
 			break;
+		case CXL_CM_CAP_CAP_ID_CACHE_ID_RT: {
+			int target_cnt;
+
+			dev_dbg(dev,
+				"found Cache ID Route Table capability (0x%x)\n",
+				offset);
+
+			target_cnt = FIELD_GET(CXL_CACHE_ID_RT_CAP_TARGET_CNT,
+					       hdr);
+			length = 0x10 + 2 * target_cnt;
+			rmap = &map->cidrt;
+			break;
+		}
+		case CXL_CM_CAP_CAP_ID_CACHE_ID_DC:
+			dev_dbg(dev,
+				"found Cache ID Decoder capability (0x%x)\n",
+				offset);
+
+			length = CXL_CACHE_ID_DC_CAPABILITY_LENGTH;
+			rmap = &map->ciddc;
+			break;
 		default:
 			dev_dbg(dev, "Unknown CM cap ID: %d (0x%x)\n", cap_id,
 				offset);
@@ -212,6 +233,8 @@ int cxl_map_component_regs(const struct cxl_register_map *map,
 	} mapinfo[] = {
 		{ &map->component_map.hdm_decoder, &regs->hdm_decoder },
 		{ &map->component_map.ras, &regs->ras },
+		{ &map->component_map.cidrt, &regs->cidrt },
+		{ &map->component_map.ciddc, &regs->ciddc },
 	};
 	int i;
 
diff --git a/drivers/cxl/cxl.h b/drivers/cxl/cxl.h
index 9711b8499508..9f0c1a39aaeb 100644
--- a/drivers/cxl/cxl.h
+++ b/drivers/cxl/cxl.h
@@ -41,6 +41,8 @@ extern const struct nvdimm_security_ops *cxl_security_ops;
 
 #define   CXL_CM_CAP_CAP_ID_RAS 0x2
 #define   CXL_CM_CAP_CAP_ID_HDM 0x5
+#define   CXL_CM_CAP_CAP_ID_CACHE_ID_RT 0xD
+#define   CXL_CM_CAP_CAP_ID_CACHE_ID_DC 0xE
 #define   CXL_CM_CAP_CAP_HDM_VERSION 1
 
 /* HDM decoders CXL 2.0 8.2.5.12 CXL HDM Decoder Capability Structure */
@@ -223,6 +225,15 @@ static inline int ways_to_eiw(unsigned int ways, u8 *eiw)
 #define   CXLDEV_MBOX_BG_CMD_COMMAND_VENDOR_MASK GENMASK_ULL(63, 48)
 #define CXLDEV_MBOX_PAYLOAD_OFFSET 0x20
 
+
+/* CXL 4.0 8.2.4.28.1 CXL Cache ID Route Table Capability Structure */
+#define CXL_CACHE_ID_RT_CAP_OFFSET 0x0
+#define   CXL_CACHE_ID_RT_CAP_TARGET_CNT GENMASK(4, 0)
+#define CXL_CACHE_ID_RT_TARGETN_OFFSET(n) (0x10 + (2 * (n)))
+
+/* CXL 4.0 8.2.4.29.1 CXL Cache ID Decoder Capability Structure */
+#define CXL_CACHE_ID_DC_CAPABILITY_LENGTH 0xC
+
 void cxl_probe_component_regs(struct device *dev, void __iomem *base,
 			      struct cxl_component_reg_map *map);
 void cxl_probe_device_regs(struct device *dev, void __iomem *base,
@@ -930,4 +941,14 @@ struct cxl_dport *devm_cxl_add_dport_by_dev(struct cxl_port *port,
 
 u16 cxl_gpf_get_dvsec(struct device *dev);
 
+#if IS_ENABLED(CONFIG_CXL_CACHE)
+int cxl_port_map_cache_id_rt(struct cxl_port *port);
+int cxl_dport_map_cache_id_dc(struct cxl_dport *dport);
+#else
+static inline int cxl_port_map_cache_id_rt(struct cxl_port *port)
+{ return -ENXIO; }
+static inline int cxl_dport_map_cache_id_dc(struct cxl_dport *dport)
+{ return -ENXIO; }
+#endif
+ 
 #endif /* __CXL_H__ */
diff --git a/drivers/cxl/port.c b/drivers/cxl/port.c
index 44d978d4205a..942b92af55d9 100644
--- a/drivers/cxl/port.c
+++ b/drivers/cxl/port.c
@@ -281,6 +281,12 @@ static struct cxl_dport *cxl_port_add_dport(struct cxl_port *port,
 		 * on failure, or the device does not implement RAS registers.
 		 */
 		devm_cxl_port_ras_setup(port);
+
+		rc = cxl_port_map_cache_id_rt(port);
+		if (rc)
+			dev_dbg(&port->dev,
+				"Failed to map cache id route table capability: %d\n",
+				rc);
 	}
 
 	dport = devm_cxl_add_dport_by_dev(port, dport_dev);
diff --git a/include/cxl/cxl.h b/include/cxl/cxl.h
index 29c0687e4ac2..7a5ed64904a6 100644
--- a/include/cxl/cxl.h
+++ b/include/cxl/cxl.h
@@ -34,10 +34,14 @@ struct cxl_regs {
 	 * Common set of CXL Component register block base pointers
 	 * @hdm_decoder: CXL 2.0 8.2.5.12 CXL HDM Decoder Capability Structure
 	 * @ras: CXL 2.0 8.2.5.9 CXL RAS Capability Structure
+	 * @cidrt: CXL 4.0 8.2.4.28 CXL Cache ID Route Table Capability Structure
+	 * @ciddc: CXL 4.0 8.2.4.29 CXL Cache ID Decoder Capability Structure
 	 */
 	struct_group_tagged(cxl_component_regs, component,
 		void __iomem *hdm_decoder;
 		void __iomem *ras;
+		void __iomem *cidrt;
+		void __iomem *ciddc;
 	);
 	/*
 	 * Common set of CXL Device register block base pointers
@@ -80,6 +84,8 @@ struct cxl_reg_map {
 struct cxl_component_reg_map {
 	struct cxl_reg_map hdm_decoder;
 	struct cxl_reg_map ras;
+	struct cxl_reg_map cidrt;
+	struct cxl_reg_map ciddc;
 };
 
 struct cxl_device_reg_map {
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 28+ messages in thread

* [PATCH 09/15] cxl/core: Add Cache ID verification
  2026-09-23 17:33 [PATCH 00/15] Add initial CXL.cache support Ben Cheatham
                   ` (7 preceding siblings ...)
  2026-09-23 17:33 ` [PATCH 08/15] cxl/core, cache: Add Cache ID register probing and init Ben Cheatham
@ 2026-09-23 17:33 ` Ben Cheatham
  2026-09-23 17:49   ` sashiko-bot
  2026-09-23 17:33 ` [PATCH 10/15] cxl/core: Add Cache ID allocation Ben Cheatham
                   ` (6 subsequent siblings)
  15 siblings, 1 reply; 28+ messages in thread
From: Ben Cheatham @ 2026-09-23 17:33 UTC (permalink / raw)
  To: linux-cxl, dave, jic23, dave.jiang, alison.schofield
  Cc: linux-iommu, benjamin.cheatham, terry.bowman, robert.richter

Check if system firmware has pre-programmed a CXL.cache device's cache
id as part of cxl_cache::probe(). A cache id is required when multiple
CXL.cache devices are present under a host bridge, so fail probe if the
id has not been programmed and another device is present.

Also add @hdmd to struct cxl_dev_state for endpoint drivers to indicate
whether their device is using HDM-D flows. This is required for
correctly validating cache id programming.

Signed-off-by: Ben Cheatham <Benjamin.Cheatham@amd.com>
---
 drivers/cxl/cache.c         |  94 ++++++++++++++++++++-
 drivers/cxl/core/cache.c    | 158 ++++++++++++++++++++++++++++++++++++
 drivers/cxl/core/cachedev.c |   1 +
 drivers/cxl/core/port.c     |   1 +
 drivers/cxl/cxl.h           |  16 ++++
 drivers/cxl/cxlcache.h      |   6 ++
 include/cxl/cxl.h           |   4 +
 7 files changed, 278 insertions(+), 2 deletions(-)

diff --git a/drivers/cxl/cache.c b/drivers/cxl/cache.c
index dfefd301696a..6c098010149a 100644
--- a/drivers/cxl/cache.c
+++ b/drivers/cxl/cache.c
@@ -84,6 +84,90 @@ static int cxl_endpoint_cache_enabled(struct cxl_port *endpoint)
 	return 0;
 }
 
+static struct cxl_port *find_host_bridge(struct cxl_port *endpoint)
+{
+	struct cxl_port *parent = endpoint->parent_dport->port;
+	struct cxl_port *hb = endpoint;
+
+	if (is_cxl_root(parent))
+		return NULL;
+
+	while (!is_cxl_root(parent)) {
+		hb = parent;
+		parent = hb->parent_dport->port;
+	}
+
+	return hb;
+}
+
+static int get_num_cachedevs_present(struct cxl_port *host_bridge, u32 *num)
+{
+	u32 num_cachedevs = 0;
+	unsigned long index;
+	struct cxl_ep *ep;
+
+	xa_for_each(&host_bridge->endpoints, index, ep) {
+		if (is_cxl_cachedev(ep->ep))
+			num_cachedevs++;
+	}
+
+	*num = num_cachedevs;
+	return 0;
+}
+
+static void deprogram_cache_id(void *_cxlcd)
+{
+	struct cxl_cachedev *cxlcd = _cxlcd;
+	struct cxl_port *hb = find_host_bridge(cxlcd->endpoint);
+
+	if (cxlcd->cache_id == CXL_CACHE_ID_NO_ID)
+		return;
+
+	guard(device)(&hb->dev);
+	cxl_free_cache_id(cxlcd);
+}
+
+static int program_cache_id(struct cxl_cachedev *cxlcd)
+{
+	struct cxl_port *endpoint = cxlcd->endpoint;
+	struct cxl_port *hb = find_host_bridge(endpoint);
+	struct device *dev = &cxlcd->dev;
+	u32 num_cachedevs;
+	int rc;
+
+	if (!hb)
+		return -ENODEV;
+
+	/*
+	 * Cache ids are unique to the host bridge, so synchronize programming
+	 * around the bridge's device lock
+	 */
+
+	guard(device)(&hb->dev);
+	rc = get_num_cachedevs_present(hb, &num_cachedevs);
+	if (rc)
+		return rc;
+
+	if (!cxl_cache_id_supported(cxlcd))
+		return num_cachedevs > 1 ? -ENXIO : 0;
+
+	rc = cxl_cachedev_validate_cache_id(cxlcd);
+	if (rc && num_cachedevs > 1) {
+		dev_err(dev,
+			"Cache id not programmed with other CXL.cache devices present: %d\n",
+			rc);
+		return rc;
+	} else if (rc) {
+		cxlcd->cache_id = 0;
+	}
+
+	rc = cxl_allocate_cache_id(cxlcd);
+	if (rc)
+		return rc;
+
+	return 0;
+}
+
 /**
  * devm_cxl_add_cachedev - Add a CXL cache device
  * @cxlds: CXL device state to associate with the cachedev
@@ -142,10 +226,16 @@ static int cxl_cache_probe(struct device *dev)
 	}
 
 	rc = cxl_endpoint_cache_enabled(cxlcd->endpoint);
-	if (rc)
+	if (rc) {
 		dev_err(dev, "CXL.cache not enabled on parent port(s)");
+		return rc;
+	}
+
+	rc = program_cache_id(cxlcd);
+	if (rc)
+		return rc;
 
-	return rc;
+	return devm_add_action_or_reset(dev, deprogram_cache_id, cxlcd);
 }
 
 static struct cxl_driver cxl_cache_driver = {
diff --git a/drivers/cxl/core/cache.c b/drivers/cxl/core/cache.c
index 17df0e603b68..be570d3dc609 100644
--- a/drivers/cxl/core/cache.c
+++ b/drivers/cxl/core/cache.c
@@ -39,3 +39,161 @@ int cxl_dport_map_cache_id_dc(struct cxl_dport *dport)
 	return cxl_map_component_regs(map, &dport->regs.component,
 				      BIT(CXL_CM_CAP_CAP_ID_CACHE_ID_DC));
 }
+
+static int cache_decoder_committed(struct cxl_dport *dport)
+{
+	u32 cap, stat;
+
+	cap = readl(dport->regs.ciddc + CXL_CACHE_ID_DC_CAP_OFFSET);
+	if (!FIELD_GET(CXL_CACHE_ID_DC_CAP_COMMIT_REQ, cap))
+		return -ENXIO;
+
+	stat = readl(dport->regs.ciddc + CXL_CACHE_ID_DC_STATUS_OFFSET);
+	return FIELD_GET(CXL_CACHE_ID_DC_STATUS_COMMITTED, stat);
+}
+
+static int cache_decoder_get_id(struct cxl_dport *dport)
+{
+	u32 ctrl;
+
+	ctrl = readl(dport->regs.ciddc + CXL_CACHE_ID_DC_CTRL_OFFSET);
+	if (FIELD_GET(CXL_CACHE_ID_DC_CTRL_HDMD_PRESENT, ctrl))
+		 return FIELD_GET(CXL_CACHE_ID_DC_CTRL_HDMD_ID, ctrl);
+
+	return FIELD_GET(CXL_CACHE_ID_DC_CTRL_LOCAL_ID, ctrl);
+}
+
+static bool cache_decoder_valid(struct cxl_dport *dport, int id, bool endpoint)
+{
+	struct pci_dev *pdev = to_pci_dev(dport->dport_dev);
+	bool flit_256b;
+	u32 ctrl;
+
+	if (!dev_is_pci(dport->dport_dev))
+		return false;
+
+	flit_256b = cxl_pci_flit_256(pdev);
+	if (id && !flit_256b)
+		return false;
+
+	ctrl = readl(dport->regs.ciddc + CXL_CACHE_ID_DC_CTRL_OFFSET);
+	if (endpoint || !flit_256b) {
+		if (!FIELD_GET(CXL_CACHE_ID_DC_CTRL_ASGN_ID, ctrl))
+			return false;
+	} else if (!FIELD_GET(CXL_CACHE_ID_DC_CTRL_FWD_ID, ctrl)) {
+		return false;
+	}
+
+	return cache_decoder_get_id(dport) == id;
+}
+
+static int cache_idrt_committed(struct cxl_port *port)
+{
+	u32 cap, stat;
+
+	cap = readl(port->regs.cidrt + CXL_CACHE_ID_RT_CAP_OFFSET);
+	if (!FIELD_GET(CXL_CACHE_ID_RT_CAP_COMMIT_REQ, cap))
+		return true;
+
+	stat = readl(port->regs.cidrt + CXL_CACHE_ID_RT_STATUS_OFFSET);
+	return FIELD_GET(CXL_CACHE_ID_RT_STATUS_COMMITTED, stat);
+}
+
+static int cache_idrt_entry_valid(struct cxl_port *port, int id)
+{
+	u16 entry;
+	u32 cap;
+
+	cap = readl(port->regs.cidrt + CXL_CACHE_ID_RT_CAP_OFFSET);
+	if (id >= FIELD_GET(CXL_CACHE_ID_RT_CAP_TARGET_CNT, cap))
+		return false;
+
+	entry = readw(port->regs.cidrt + CXL_CACHE_ID_RT_TARGETN_OFFSET(id));
+	return FIELD_GET(CXL_CACHE_ID_RT_TARGETN_VALID, entry);
+}
+
+static struct ida *find_cache_id_ida(struct cxl_port *port)
+{
+	struct cxl_port *parent = parent_port_of(port);
+
+	while (parent) {
+		if (is_cxl_root(parent))
+			return &port->cache_ida;
+
+		port = parent;
+		parent = parent_port_of(port);
+	}
+
+	return NULL;
+}
+
+int cxl_allocate_cache_id(struct cxl_cachedev *cxlcd)
+{
+	struct ida *ida = find_cache_id_ida(cxlcd->endpoint);
+	int id;
+
+	if (!ida)
+		return -ENOSPC;
+
+	id = ida_alloc_range(ida, cxlcd->cache_id, cxlcd->cache_id, GFP_KERNEL);
+	if (id != cxlcd->cache_id)
+		return -EINVAL;
+
+	return 0;
+}
+EXPORT_SYMBOL_FOR_MODULES(cxl_allocate_cache_id, "cxl_cache");
+
+void cxl_free_cache_id(struct cxl_cachedev *cxlcd)
+{
+	struct ida *ida = find_cache_id_ida(cxlcd->endpoint);
+
+	if (ida)
+		ida_free(ida, cxlcd->cache_id);
+}
+EXPORT_SYMBOL_FOR_MODULES(cxl_free_cache_id, "cxl_cache");
+
+bool cxl_cache_id_supported(struct cxl_cachedev *cxlcd)
+{
+	struct cxl_dport *dport = cxlcd->endpoint->parent_dport;
+	struct cxl_port *port = dport->port;
+
+	for (; !is_cxl_root(port);
+	     dport = port->parent_dport, port = dport->port) {
+		if (!port->regs.cidrt || !dport->regs.ciddc)
+			return false;
+	}
+
+	return true;
+}
+EXPORT_SYMBOL_FOR_MODULES(cxl_cache_id_supported, "cxl_cache");
+
+int cxl_cachedev_validate_cache_id(struct cxl_cachedev *cxlcd)
+{
+	struct cxl_dport *dport = cxlcd->endpoint->parent_dport;
+	struct cxl_port *port = dport->port;
+	int id = CXL_CACHE_ID_NO_ID;
+	bool endpoint = true;
+
+	cxlcd->cache_id = CXL_CACHE_ID_NO_ID;
+	id = cache_decoder_get_id(dport);
+
+	while (!is_cxl_root(port)) {
+		if (!cache_decoder_valid(dport, id, endpoint) ||
+		    !cache_decoder_committed(dport))
+			return -EINVAL;
+
+		endpoint = false;
+
+		if (!cache_idrt_entry_valid(port, id) ||
+		    !cache_idrt_committed(port))
+			return -EINVAL;
+
+		dport = port->parent_dport;
+		port = dport->port;
+	}
+
+	cxlcd->cache_id = id;
+	return 0;
+}
+EXPORT_SYMBOL_FOR_MODULES(cxl_cachedev_validate_cache_id, "cxl_cache");
+>>>>>>> conflict 1 of 1 ends
diff --git a/drivers/cxl/core/cachedev.c b/drivers/cxl/core/cachedev.c
index 2bd69133c631..38f0adf2b197 100644
--- a/drivers/cxl/core/cachedev.c
+++ b/drivers/cxl/core/cachedev.c
@@ -72,6 +72,7 @@ static struct cxl_cachedev *cxl_cachedev_alloc(struct cxl_dev_state *cxlds)
 	cxlcd->id = rc;
 	cxlcd->depth = -1;
 	cxlcd->endpoint = ERR_PTR(-ENODEV);
+	cxlcd->cache_id = CXL_CACHE_ID_NO_ID;
 
 	dev = &cxlcd->dev;
 	device_initialize(dev);
diff --git a/drivers/cxl/core/port.c b/drivers/cxl/core/port.c
index 6da11d28bb5e..047801faa020 100644
--- a/drivers/cxl/core/port.c
+++ b/drivers/cxl/core/port.c
@@ -740,6 +740,7 @@ static struct cxl_port *cxl_port_alloc(struct device *uport_dev,
 		dev->parent = uport_dev;
 
 	ida_init(&port->decoder_ida);
+	ida_init(&port->cache_ida);
 	port->hdm_end = -1;
 	port->commit_end = -1;
 	xa_init(&port->dports);
diff --git a/drivers/cxl/cxl.h b/drivers/cxl/cxl.h
index 9f0c1a39aaeb..ed6b56a269ee 100644
--- a/drivers/cxl/cxl.h
+++ b/drivers/cxl/cxl.h
@@ -229,9 +229,23 @@ static inline int ways_to_eiw(unsigned int ways, u8 *eiw)
 /* CXL 4.0 8.2.4.28.1 CXL Cache ID Route Table Capability Structure */
 #define CXL_CACHE_ID_RT_CAP_OFFSET 0x0
 #define   CXL_CACHE_ID_RT_CAP_TARGET_CNT GENMASK(4, 0)
+#define   CXL_CACHE_ID_RT_CAP_COMMIT_REQ BIT(16)
+#define CXL_CACHE_ID_RT_STATUS_OFFSET 0x8
+#define   CXL_CACHE_ID_RT_STATUS_COMMITTED BIT(0)
 #define CXL_CACHE_ID_RT_TARGETN_OFFSET(n) (0x10 + (2 * (n)))
+#define   CXL_CACHE_ID_RT_TARGETN_VALID BIT(0)
 
 /* CXL 4.0 8.2.4.29.1 CXL Cache ID Decoder Capability Structure */
+#define CXL_CACHE_ID_DC_CAP_OFFSET 0x0
+#define   CXL_CACHE_ID_DC_CAP_COMMIT_REQ BIT(0)
+#define CXL_CACHE_ID_DC_CTRL_OFFSET 0x4
+#define   CXL_CACHE_ID_DC_CTRL_FWD_ID BIT(0)
+#define   CXL_CACHE_ID_DC_CTRL_ASGN_ID BIT(1)
+#define   CXL_CACHE_ID_DC_CTRL_HDMD_PRESENT BIT(2)
+#define   CXL_CACHE_ID_DC_CTRL_HDMD_ID GENMASK(11, 8)
+#define   CXL_CACHE_ID_DC_CTRL_LOCAL_ID GENMASK(19, 16)
+#define CXL_CACHE_ID_DC_STATUS_OFFSET 0x8
+#define   CXL_CACHE_ID_DC_STATUS_COMMITTED BIT(0)
 #define CXL_CACHE_ID_DC_CAPABILITY_LENGTH 0xC
 
 void cxl_probe_component_regs(struct device *dev, void __iomem *base,
@@ -578,6 +592,7 @@ struct cxl_dax_region {
  * @cdat_available: Should a CDAT attribute be available in sysfs
  * @pci_latency: Upstream latency in picoseconds
  * @component_reg_phys: Physical address of component register
+ * @cache_ida: ida used for cache id allocations on host bridges
  */
 struct cxl_port {
 	struct device dev;
@@ -603,6 +618,7 @@ struct cxl_port {
 	bool cdat_available;
 	long pci_latency;
 	resource_size_t component_reg_phys;
+	struct ida cache_ida;
 };
 
 struct cxl_root;
diff --git a/drivers/cxl/cxlcache.h b/drivers/cxl/cxlcache.h
index e9a4567b7f2c..1e4a1e61ad7b 100644
--- a/drivers/cxl/cxlcache.h
+++ b/drivers/cxl/cxlcache.h
@@ -11,6 +11,7 @@
  * @cxlds: device state backing this device
  * @detach_work: active cachedev lost a port in its ancestry
  * @endpoint: connection to the CXL port topology for this device
+ * @cache_id: id used for CXL cache routing capabilities
  * @id: id number of this cachedev instance
  * @depth: endpoint port depth in hierarchy
  */
@@ -19,6 +20,7 @@ struct cxl_cachedev {
 	struct cxl_dev_state *cxlds;
 	struct work_struct detach_work;
 	struct cxl_port *endpoint;
+	int cache_id;
 	int id;
 	int depth;
 };
@@ -39,4 +41,8 @@ struct cxl_cachedev *__devm_cxl_add_cachedev(struct cxl_dev_state *cxlds,
 int cxl_accel_read_cache_info(struct cxl_dev_state *cxlds);
 void cxl_clear_cache_enable(struct cxl_dev_state *cxlds);
 int devm_cxl_enable_cache(struct device *host, struct cxl_dev_state *cxlds);
+bool cxl_cache_id_supported(struct cxl_cachedev *cxlcd);
+int cxl_allocate_cache_id(struct cxl_cachedev *cxlcd);
+void cxl_free_cache_id(struct cxl_cachedev *cxlcd);
+int cxl_cachedev_validate_cache_id(struct cxl_cachedev *cxlcd);
 #endif
diff --git a/include/cxl/cxl.h b/include/cxl/cxl.h
index 7a5ed64904a6..b99a9727e284 100644
--- a/include/cxl/cxl.h
+++ b/include/cxl/cxl.h
@@ -155,6 +155,8 @@ struct cxl_dpa_partition {
 
 #define CXL_NR_PARTITIONS_MAX 2
 
+#define CXL_CACHE_ID_NO_ID (-1)
+ 
 /**
  * struct cxl_cache_state - CXL cache device state for use by external drivers
  * @size: Size of device's cache
@@ -182,6 +184,7 @@ struct cxl_cache_state {
  * @cxlmd: The device representing the CXL.mem capabilities of @dev
  * @cxlcd: The device representing the CXL.cache capabilities of @dev
  * @cstate: CXL.cache information for @dev
+ * @hdmd: Whether the device is using HDM-D flows for CXL.cache
  * @reg_map: component and ras register mapping parameters
  * @regs: Parsed register blocks
  * @cxl_dvsec: Offset to the PCIe device DVSEC
@@ -201,6 +204,7 @@ struct cxl_dev_state {
 	struct cxl_memdev *cxlmd;
 	struct cxl_cachedev *cxlcd;
  	struct cxl_cache_state cstate;
+	bool hdmd;
 
 	/* private for Type2 drivers */
 	struct cxl_register_map reg_map;
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 28+ messages in thread

* [PATCH 10/15] cxl/core: Add Cache ID allocation
  2026-09-23 17:33 [PATCH 00/15] Add initial CXL.cache support Ben Cheatham
                   ` (8 preceding siblings ...)
  2026-09-23 17:33 ` [PATCH 09/15] cxl/core: Add Cache ID verification Ben Cheatham
@ 2026-09-23 17:33 ` Ben Cheatham
  2026-09-23 17:49   ` sashiko-bot
  2026-09-23 17:33 ` [PATCH 11/15] cxl/core: Add support for HDM-D cache id programming Ben Cheatham
                   ` (5 subsequent siblings)
  15 siblings, 1 reply; 28+ messages in thread
From: Ben Cheatham @ 2026-09-23 17:33 UTC (permalink / raw)
  To: linux-cxl, dave, jic23, dave.jiang, alison.schofield
  Cc: linux-iommu, benjamin.cheatham, terry.bowman, robert.richter

Add allocation and programming of CXL cache ids for struct
cxl_cachedevs as part of cxl_cache::probe(). Programming only occurs
when system firmware has not already programmed an id and the device is
*not* using HDM-D flows. Cache id allocation for HDM-D devices will be
added in a later commit ("cxl/core": Add support for HDM-D cache id
programming").

Signed-off-by: Ben Cheatham <Benjamin.Cheatham@amd.com>
---
 drivers/cxl/cache.c      |  21 ++-
 drivers/cxl/core/cache.c | 337 ++++++++++++++++++++++++++++++++++++++-
 drivers/cxl/cxl.h        |  10 ++
 drivers/cxl/cxlcache.h   |   2 +
 4 files changed, 359 insertions(+), 11 deletions(-)

diff --git a/drivers/cxl/cache.c b/drivers/cxl/cache.c
index 6c098010149a..0fc8f413c638 100644
--- a/drivers/cxl/cache.c
+++ b/drivers/cxl/cache.c
@@ -124,6 +124,7 @@ static void deprogram_cache_id(void *_cxlcd)
 		return;
 
 	guard(device)(&hb->dev);
+	cxl_cachedev_deprogram_cache_id(cxlcd);
 	cxl_free_cache_id(cxlcd);
 }
 
@@ -152,20 +153,26 @@ static int program_cache_id(struct cxl_cachedev *cxlcd)
 		return num_cachedevs > 1 ? -ENXIO : 0;
 
 	rc = cxl_cachedev_validate_cache_id(cxlcd);
-	if (rc && num_cachedevs > 1) {
+	if (!rc)
+		return cxl_allocate_cache_id(cxlcd);
+
+	if (cxlcd->cxlds->hdmd) {
 		dev_err(dev,
-			"Cache id not programmed with other CXL.cache devices present: %d\n",
-			rc);
-		return rc;
-	} else if (rc) {
-		cxlcd->cache_id = 0;
+			"Cache id programming not supported for HDM-D devices\n");
+		return -ENXIO;
 	}
 
 	rc = cxl_allocate_cache_id(cxlcd);
 	if (rc)
 		return rc;
 
-	return 0;
+	rc = cxl_cachedev_program_cache_id(cxlcd);
+	if (rc) {
+		dev_err(dev, "Failed to program cache id: %d\n", rc);
+		cxl_free_cache_id(cxlcd);
+	}
+
+	return rc;
 }
 
 /**
diff --git a/drivers/cxl/core/cache.c b/drivers/cxl/core/cache.c
index be570d3dc609..9c6c8ea713ef 100644
--- a/drivers/cxl/core/cache.c
+++ b/drivers/cxl/core/cache.c
@@ -1,11 +1,15 @@
 // SPDX-License-Identifier: GPL-2.0-only
 /* Copyright (C) Advanced Micro Device, Inc. */
 
+#include <linux/iopoll.h>
 #include <linux/pci.h>
 #include <cxlcache.h>
 
 #include "cxlpci.h"
 
+#define CXL_CACHE_ID_COMMIT_MAXTMO_US (5 * USEC_PER_SEC)
+#define CACHE_DECODER_MAX_CACHE_ID (15)
+
 int cxl_port_map_cache_id_rt(struct cxl_port *port)
 {
 	struct cxl_register_map *map = &port->reg_map;
@@ -112,6 +116,260 @@ static int cache_idrt_entry_valid(struct cxl_port *port, int id)
 	return FIELD_GET(CXL_CACHE_ID_RT_TARGETN_VALID, entry);
 }
 
+static unsigned long __cxl_cid_get_timeout_us(struct device *dev,
+					      unsigned int scale,
+					      unsigned int base)
+{
+	static const unsigned long scale_tbl[] = {
+		1, 10, 100, 1000, 10000, 100000, 1000000, 10000000,
+	};
+
+	if (scale >= ARRAY_SIZE(scale_tbl) || !base) {
+		dev_dbg(dev,
+			"Invalid Cache ID commit timeout: scale=%u base=%u\n",
+			scale, base);
+		return CXL_CACHE_ID_COMMIT_MAXTMO_US;
+	}
+
+	return scale_tbl[scale] * base;
+}
+
+static int __cxl_cid_wait_commit(struct device *dev, void __iomem *status_reg,
+				 u32 commit_bit, u32 err_bit,
+				 unsigned int scale, unsigned int base)
+{
+	unsigned long tmo_us, poll_us;
+	ktime_t start;
+	u32 status;
+	int rc;
+
+	tmo_us = min_t(unsigned long, CXL_CACHE_ID_COMMIT_MAXTMO_US,
+		       __cxl_cid_get_timeout_us(dev, scale, base));
+	poll_us = max_t(unsigned long, tmo_us / 10, 1); /* ~10% */
+	start = ktime_get();
+
+	rc = readx_poll_timeout(readl, status_reg, status,
+				status & (commit_bit | err_bit), poll_us,
+				tmo_us);
+	if (rc) {
+		dev_err(dev, "Cache ID commit timed out\n");
+		return rc;
+	}
+
+	if (status & err_bit) {
+		dev_err(dev, "Cache ID commit rejected by hardware\n");
+		return -EIO;
+	}
+
+	dev_dbg(dev, "Cache ID commit took %lluus\n",
+		ktime_to_us(ktime_sub(ktime_get(), start)));
+	return 0;
+}
+
+static int cxl_cid_commit_decoder(struct cxl_dport *dport)
+{
+	void __iomem *ciddc = dport->regs.ciddc;
+	u32 cap, ctrl, status;
+	u8 scale, base;
+	int rc;
+
+	cap = readl(ciddc + CXL_CACHE_ID_DC_CAP_OFFSET);
+	if (!FIELD_GET(CXL_CACHE_ID_DC_CAP_COMMIT_REQ, cap))
+		return 0;
+
+	status = readl(ciddc + CXL_CACHE_ID_DC_STATUS_OFFSET);
+	scale = FIELD_GET(CXL_CACHE_ID_DC_STATUS_TM_SCALE, status);
+	base = FIELD_GET(CXL_CACHE_ID_DC_STATUS_TM_BASE, status);
+
+	ctrl = readl(ciddc + CXL_CACHE_ID_DC_CTRL_OFFSET);
+	if (FIELD_GET(CXL_CACHE_ID_DC_CTRL_COMMIT, ctrl)) {
+		ctrl &= ~CXL_CACHE_ID_DC_CTRL_COMMIT;
+		writel(ctrl, ciddc + CXL_CACHE_ID_DC_CTRL_OFFSET);
+	}
+
+	ctrl |= CXL_CACHE_ID_DC_CTRL_COMMIT;
+	writel(ctrl, ciddc + CXL_CACHE_ID_DC_CTRL_OFFSET);
+
+	rc = __cxl_cid_wait_commit(dport->dport_dev,
+				   ciddc + CXL_CACHE_ID_DC_STATUS_OFFSET,
+				   CXL_CACHE_ID_DC_STATUS_COMMITTED,
+				   CXL_CACHE_ID_DC_STATUS_COMMIT_ERR, scale,
+				   base);
+	if (rc) {
+		ctrl &= ~CXL_CACHE_ID_DC_CTRL_COMMIT;
+		writel(ctrl, ciddc + CXL_CACHE_ID_DC_CTRL_OFFSET);
+	}
+
+	return rc;
+}
+
+static int cxl_cid_program_decoder(struct cxl_dport *dport, int cid,
+				   bool endpoint)
+{
+	void __iomem *ciddc = dport->regs.ciddc;
+	u32 ctrl;
+
+	if (!ciddc)
+		return -EINVAL;
+
+	ctrl = readl(ciddc + CXL_CACHE_ID_DC_CTRL_OFFSET);
+
+	/*
+	 * The decoder may have been programmed before, so we zero out
+	 * all the fields before writing to them
+	 */
+	ctrl &= ~(CXL_CACHE_ID_DC_CTRL_ASGN_ID | CXL_CACHE_ID_DC_CTRL_FWD_ID);
+	if (endpoint) {
+		ctrl |= CXL_CACHE_ID_DC_CTRL_ASGN_ID;
+	} else {
+		ctrl |= CXL_CACHE_ID_DC_CTRL_FWD_ID;
+	}
+
+	FIELD_MODIFY(CXL_CACHE_ID_DC_CTRL_LOCAL_ID, &ctrl, cid);
+	writel(ctrl, ciddc + CXL_CACHE_ID_DC_CTRL_OFFSET);
+
+	return cxl_cid_commit_decoder(dport);
+}
+
+static int cxl_cid_commit_table(struct cxl_port *port)
+{
+	void __iomem *cidrt = port->regs.cidrt;
+	u32 cap, ctrl, status;
+	u8 scale, base;
+	int rc;
+
+	cap = readl(cidrt + CXL_CACHE_ID_RT_CAP_OFFSET);
+	if (!FIELD_GET(CXL_CACHE_ID_RT_CAP_COMMIT_REQ, cap))
+		return 0;
+
+	status = readl(cidrt + CXL_CACHE_ID_RT_STATUS_OFFSET);
+	scale = FIELD_GET(CXL_CACHE_ID_RT_STATUS_TM_SCALE, status);
+	base = FIELD_GET(CXL_CACHE_ID_RT_STATUS_TM_BASE, status);
+
+	ctrl = readl(cidrt + CXL_CACHE_ID_RT_CTRL_OFFSET);
+	if (FIELD_GET(CXL_CACHE_ID_RT_CTRL_COMMIT, ctrl)) {
+		ctrl &= ~CXL_CACHE_ID_RT_CTRL_COMMIT;
+		writel(ctrl, cidrt + CXL_CACHE_ID_RT_CTRL_OFFSET);
+	}
+
+	ctrl |= CXL_CACHE_ID_RT_CTRL_COMMIT;
+	writel(ctrl, cidrt + CXL_CACHE_ID_RT_CTRL_OFFSET);
+
+	rc = __cxl_cid_wait_commit(&port->dev,
+				   cidrt + CXL_CACHE_ID_RT_STATUS_OFFSET,
+				   CXL_CACHE_ID_RT_STATUS_COMMITTED,
+				   CXL_CACHE_ID_RT_STATUS_COMMIT_ERR, scale,
+				   base);
+	if (rc) {
+		ctrl &= ~CXL_CACHE_ID_RT_CTRL_COMMIT;
+		writel(ctrl, cidrt + CXL_CACHE_ID_RT_CTRL_OFFSET);
+	}
+
+	return rc;
+}
+
+static int cxl_cid_program_table_entry(struct cxl_port *port, int cid,
+				       unsigned int dport_id)
+{
+	void __iomem *cidrt = port->regs.cidrt;
+	u8 target_cnt, portn;
+	u16 target_n;
+	u32 cap;
+
+	if (!cidrt)
+		return -EINVAL;
+
+	cap = readl(cidrt + CXL_CACHE_ID_RT_CAP_OFFSET);
+	target_cnt = FIELD_GET(CXL_CACHE_ID_RT_CAP_TARGET_CNT, cap);
+
+	/* Shouldn't be possible, but better to be safe */
+	if (cid >= target_cnt) {
+		dev_err(&port->dev,
+			"Tried to allocate cache ID (%d) larger than table size (%d)\n",
+			cid, target_cnt);
+		return -EINVAL;
+	}
+
+	target_n = readw(cidrt + CXL_CACHE_ID_RT_TARGETN_OFFSET(cid));
+	if (FIELD_GET(CXL_CACHE_ID_RT_TARGETN_VALID, target_n)) {
+		portn = FIELD_GET(CXL_CACHE_ID_RT_TARGETN_PORTN, target_n);
+		return dport_id == portn ? 0 : -EINVAL;
+	}
+
+	target_n = CXL_CACHE_ID_RT_TARGETN_VALID;
+	target_n |= FIELD_PREP(CXL_CACHE_ID_RT_TARGETN_PORTN, dport_id);
+	writew(target_n, cidrt + CXL_CACHE_ID_RT_TARGETN_OFFSET(cid));
+
+	return cxl_cid_commit_table(port);
+}
+
+static void cxl_cid_invalidate_table_entry(struct cxl_port *port, int cid)
+{
+	void __iomem *cidrt = port->regs.cidrt;
+	u8 target_cnt;
+	u16 target_n;
+	u32 cap;
+	int rc;
+
+	if (!cidrt)
+		return;
+
+	cap = readl(cidrt + CXL_CACHE_ID_RT_CAP_OFFSET);
+	target_cnt = FIELD_GET(CXL_CACHE_ID_RT_CAP_TARGET_CNT, cap);
+
+	/* Shouldn't be possible, but better to be safe */
+	if (cid >= target_cnt) {
+		dev_err(&port->dev,
+			"Tried to free cache ID (%d) larger than table size (%d)\n",
+			cid, target_cnt);
+		return;
+	}
+
+	target_n = readw(cidrt + CXL_CACHE_ID_RT_TARGETN_OFFSET(cid));
+	if (!FIELD_GET(CXL_CACHE_ID_RT_TARGETN_VALID, target_n))
+		return;
+
+	target_n &= ~CXL_CACHE_ID_RT_TARGETN_VALID;
+	writew(target_n, cidrt + CXL_CACHE_ID_RT_TARGETN_OFFSET(cid));
+
+	rc = cxl_cid_commit_table(port);
+	if (rc)
+		dev_warn(
+			&port->dev,
+			"Failed to commit invalidation of cache id table entry %d: %d\n",
+			cid, rc);
+}
+
+static int get_max_cid(struct cxl_port *endpoint)
+{
+	struct cxl_port *port = parent_port_of(endpoint);
+	void __iomem *cidrt;
+	u32 cap;
+	u8 cnt;
+
+	if (!port)
+		return -EINVAL;
+
+	while (!is_cxl_root(port) && !is_cxl_root(parent_port_of(port)))
+		port = parent_port_of(port);
+
+	cidrt = port->regs.cidrt;
+	if (!cidrt)
+		return -EINVAL;
+
+	cap = readl(cidrt + CXL_CACHE_ID_RT_CAP_OFFSET);
+	cnt = FIELD_GET(CXL_CACHE_ID_RT_CAP_TARGET_CNT, cap);
+	if (cnt == 0)
+		return 0;
+
+	/*
+	 * Cache id decoders have 4 bits for the cache id, while target count
+	 * is a 5 bit long field. Limit cache ids to the maximum value allowed
+	 * by cache id decoders.
+	 */
+	return min(CACHE_DECODER_MAX_CACHE_ID, cnt - 1);
+}
+
 static struct ida *find_cache_id_ida(struct cxl_port *port)
 {
 	struct cxl_port *parent = parent_port_of(port);
@@ -130,15 +388,26 @@ static struct ida *find_cache_id_ida(struct cxl_port *port)
 int cxl_allocate_cache_id(struct cxl_cachedev *cxlcd)
 {
 	struct ida *ida = find_cache_id_ida(cxlcd->endpoint);
-	int id;
+	int min, max, id;
 
 	if (!ida)
 		return -ENOSPC;
 
-	id = ida_alloc_range(ida, cxlcd->cache_id, cxlcd->cache_id, GFP_KERNEL);
-	if (id != cxlcd->cache_id)
+	max = get_max_cid(cxlcd->endpoint);
+	if (max < 0 || cxlcd->cache_id > max)
 		return -EINVAL;
 
+	if (cxlcd->cache_id == CXL_CACHE_ID_NO_ID) {
+		min = 0;
+	} else {
+		min = max = cxlcd->cache_id;
+	}
+
+	id = ida_alloc_range(ida, min, max, GFP_KERNEL);
+	if (id < 0)
+		return id;
+
+	cxlcd->cache_id = id;
 	return 0;
 }
 EXPORT_SYMBOL_FOR_MODULES(cxl_allocate_cache_id, "cxl_cache");
@@ -196,4 +465,64 @@ int cxl_cachedev_validate_cache_id(struct cxl_cachedev *cxlcd)
 	return 0;
 }
 EXPORT_SYMBOL_FOR_MODULES(cxl_cachedev_validate_cache_id, "cxl_cache");
->>>>>>> conflict 1 of 1 ends
+
+static void __deprogram_cache_id(struct cxl_cachedev *cxlcd,
+				 struct cxl_port *stop)
+{
+	struct cxl_port *iter = cxlcd->endpoint->parent_dport->port;
+
+	/*
+	 * Leave cache id decoders programmed; the table entry being invalidated
+	 * should be enough
+	 */
+	while (!is_cxl_root(iter)) {
+		cxl_cid_invalidate_table_entry(iter, cxlcd->cache_id);
+
+		if (iter == stop)
+			return;
+
+		iter = parent_port_of(iter);
+	}
+}
+
+int cxl_cachedev_program_cache_id(struct cxl_cachedev *cxlcd)
+{
+	struct cxl_dport *dport = cxlcd->endpoint->parent_dport;
+	struct cxl_port *port = dport->port;
+	bool endpoint = true;
+	int rc;
+
+	while (!is_cxl_root(port)) {
+		rc = cxl_cid_program_decoder(dport, cxlcd->cache_id,
+					     endpoint);
+		if (rc)
+			goto err;
+
+		rc = cxl_cid_program_table_entry(port, cxlcd->cache_id,
+						 dport->port_id);
+		if (rc)
+			goto err;
+
+		endpoint = false;
+		dport = port->parent_dport;
+		port = dport->port;
+	}
+
+	return 0;
+
+err:
+	__deprogram_cache_id(cxlcd, port);
+	return rc;
+}
+EXPORT_SYMBOL_FOR_MODULES(cxl_cachedev_program_cache_id, "cxl_cache");
+
+void cxl_cachedev_deprogram_cache_id(struct cxl_cachedev *cxlcd)
+{
+	struct cxl_root *root = find_cxl_root(cxlcd->endpoint);
+
+	if (!root)
+		return;
+
+	__deprogram_cache_id(cxlcd, &root->port);
+}
+EXPORT_SYMBOL_FOR_MODULES(cxl_cachedev_deprogram_cache_id, "cxl_cache");
diff --git a/drivers/cxl/cxl.h b/drivers/cxl/cxl.h
index ed6b56a269ee..4cdc26dbdb6a 100644
--- a/drivers/cxl/cxl.h
+++ b/drivers/cxl/cxl.h
@@ -230,10 +230,16 @@ static inline int ways_to_eiw(unsigned int ways, u8 *eiw)
 #define CXL_CACHE_ID_RT_CAP_OFFSET 0x0
 #define   CXL_CACHE_ID_RT_CAP_TARGET_CNT GENMASK(4, 0)
 #define   CXL_CACHE_ID_RT_CAP_COMMIT_REQ BIT(16)
+#define CXL_CACHE_ID_RT_CTRL_OFFSET 0x4
+#define   CXL_CACHE_ID_RT_CTRL_COMMIT BIT(0)
 #define CXL_CACHE_ID_RT_STATUS_OFFSET 0x8
 #define   CXL_CACHE_ID_RT_STATUS_COMMITTED BIT(0)
+#define   CXL_CACHE_ID_RT_STATUS_COMMIT_ERR BIT(1)
+#define   CXL_CACHE_ID_RT_STATUS_TM_SCALE GENMASK(11, 8)
+#define   CXL_CACHE_ID_RT_STATUS_TM_BASE GENMASK(15, 12)
 #define CXL_CACHE_ID_RT_TARGETN_OFFSET(n) (0x10 + (2 * (n)))
 #define   CXL_CACHE_ID_RT_TARGETN_VALID BIT(0)
+#define   CXL_CACHE_ID_RT_TARGETN_PORTN GENMASK(15, 8)
 
 /* CXL 4.0 8.2.4.29.1 CXL Cache ID Decoder Capability Structure */
 #define CXL_CACHE_ID_DC_CAP_OFFSET 0x0
@@ -242,10 +248,14 @@ static inline int ways_to_eiw(unsigned int ways, u8 *eiw)
 #define   CXL_CACHE_ID_DC_CTRL_FWD_ID BIT(0)
 #define   CXL_CACHE_ID_DC_CTRL_ASGN_ID BIT(1)
 #define   CXL_CACHE_ID_DC_CTRL_HDMD_PRESENT BIT(2)
+#define   CXL_CACHE_ID_DC_CTRL_COMMIT BIT(3)
 #define   CXL_CACHE_ID_DC_CTRL_HDMD_ID GENMASK(11, 8)
 #define   CXL_CACHE_ID_DC_CTRL_LOCAL_ID GENMASK(19, 16)
 #define CXL_CACHE_ID_DC_STATUS_OFFSET 0x8
 #define   CXL_CACHE_ID_DC_STATUS_COMMITTED BIT(0)
+#define   CXL_CACHE_ID_DC_STATUS_COMMIT_ERR BIT(1)
+#define   CXL_CACHE_ID_DC_STATUS_TM_SCALE GENMASK(11, 8)
+#define   CXL_CACHE_ID_DC_STATUS_TM_BASE GENMASK(15, 12)
 #define CXL_CACHE_ID_DC_CAPABILITY_LENGTH 0xC
 
 void cxl_probe_component_regs(struct device *dev, void __iomem *base,
diff --git a/drivers/cxl/cxlcache.h b/drivers/cxl/cxlcache.h
index 1e4a1e61ad7b..72d58eded292 100644
--- a/drivers/cxl/cxlcache.h
+++ b/drivers/cxl/cxlcache.h
@@ -45,4 +45,6 @@ bool cxl_cache_id_supported(struct cxl_cachedev *cxlcd);
 int cxl_allocate_cache_id(struct cxl_cachedev *cxlcd);
 void cxl_free_cache_id(struct cxl_cachedev *cxlcd);
 int cxl_cachedev_validate_cache_id(struct cxl_cachedev *cxlcd);
+int cxl_cachedev_program_cache_id(struct cxl_cachedev *cxlcd);
+void cxl_cachedev_deprogram_cache_id(struct cxl_cachedev *cxlcd);
 #endif
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 28+ messages in thread

* [PATCH 11/15] cxl/core: Add support for HDM-D cache id programming
  2026-09-23 17:33 [PATCH 00/15] Add initial CXL.cache support Ben Cheatham
                   ` (9 preceding siblings ...)
  2026-09-23 17:33 ` [PATCH 10/15] cxl/core: Add Cache ID allocation Ben Cheatham
@ 2026-09-23 17:33 ` Ben Cheatham
  2026-09-23 17:51   ` sashiko-bot
  2026-09-23 17:33 ` [PATCH 12/15] cxl/cache: Add snoop filter creation and set up Ben Cheatham
                   ` (4 subsequent siblings)
  15 siblings, 1 reply; 28+ messages in thread
From: Ben Cheatham @ 2026-09-23 17:33 UTC (permalink / raw)
  To: linux-cxl, dave, jic23, dave.jiang, alison.schofield
  Cc: linux-iommu, benjamin.cheatham, terry.bowman, robert.richter

Add cache id allocation and programming to the pre-existing cache id
programming routines.

Signed-off-by: Ben Cheatham <Benjamin.Cheatham@amd.com>
---
 drivers/cxl/cache.c      |  6 -----
 drivers/cxl/core/cache.c | 52 +++++++++++++++++++++++++++++++++++-----
 drivers/cxl/cxl.h        |  3 +++
 3 files changed, 49 insertions(+), 12 deletions(-)

diff --git a/drivers/cxl/cache.c b/drivers/cxl/cache.c
index 0fc8f413c638..40d1e8330df7 100644
--- a/drivers/cxl/cache.c
+++ b/drivers/cxl/cache.c
@@ -156,12 +156,6 @@ static int program_cache_id(struct cxl_cachedev *cxlcd)
 	if (!rc)
 		return cxl_allocate_cache_id(cxlcd);
 
-	if (cxlcd->cxlds->hdmd) {
-		dev_err(dev,
-			"Cache id programming not supported for HDM-D devices\n");
-		return -ENXIO;
-	}
-
 	rc = cxl_allocate_cache_id(cxlcd);
 	if (rc)
 		return rc;
diff --git a/drivers/cxl/core/cache.c b/drivers/cxl/core/cache.c
index 9c6c8ea713ef..7810772326e9 100644
--- a/drivers/cxl/core/cache.c
+++ b/drivers/cxl/core/cache.c
@@ -204,7 +204,7 @@ static int cxl_cid_commit_decoder(struct cxl_dport *dport)
 }
 
 static int cxl_cid_program_decoder(struct cxl_dport *dport, int cid,
-				   bool endpoint)
+				   bool endpoint, bool hdmd)
 {
 	void __iomem *ciddc = dport->regs.ciddc;
 	u32 ctrl;
@@ -225,7 +225,14 @@ static int cxl_cid_program_decoder(struct cxl_dport *dport, int cid,
 		ctrl |= CXL_CACHE_ID_DC_CTRL_FWD_ID;
 	}
 
-	FIELD_MODIFY(CXL_CACHE_ID_DC_CTRL_LOCAL_ID, &ctrl, cid);
+	if (hdmd) {
+		ctrl |= CXL_CACHE_ID_DC_CTRL_HDMD_PRESENT;
+		FIELD_MODIFY(CXL_CACHE_ID_DC_CTRL_HDMD_ID, &ctrl, cid);
+	} else {
+		ctrl &= ~CXL_CACHE_ID_DC_CTRL_HDMD_PRESENT;
+		FIELD_MODIFY(CXL_CACHE_ID_DC_CTRL_LOCAL_ID, &ctrl, cid);
+	}
+
 	writel(ctrl, ciddc + CXL_CACHE_ID_DC_CTRL_OFFSET);
 
 	return cxl_cid_commit_decoder(dport);
@@ -269,10 +276,10 @@ static int cxl_cid_commit_table(struct cxl_port *port)
 }
 
 static int cxl_cid_program_table_entry(struct cxl_port *port, int cid,
-				       unsigned int dport_id)
+				       unsigned int dport_id, bool hdmd)
 {
 	void __iomem *cidrt = port->regs.cidrt;
-	u8 target_cnt, portn;
+	u8 target_cnt, hdmd_max, portn;
 	u16 target_n;
 	u32 cap;
 
@@ -290,6 +297,16 @@ static int cxl_cid_program_table_entry(struct cxl_port *port, int cid,
 		return -EINVAL;
 	}
 
+	if (is_cxl_root(parent_port_of(port)) && hdmd) {
+		hdmd_max = FIELD_GET(CXL_CACHE_ID_RT_CAP_HDMD_MAX, cap);
+
+		if (port->num_hdmd > hdmd_max) {
+			dev_err(&port->dev,
+				"Maximum number of devices using HDM-D reached\n");
+			return -EINVAL;
+		}
+	}
+
 	target_n = readw(cidrt + CXL_CACHE_ID_RT_TARGETN_OFFSET(cid));
 	if (FIELD_GET(CXL_CACHE_ID_RT_TARGETN_VALID, target_n)) {
 		portn = FIELD_GET(CXL_CACHE_ID_RT_TARGETN_PORTN, target_n);
@@ -385,6 +402,22 @@ static struct ida *find_cache_id_ida(struct cxl_port *port)
 	return NULL;
 }
 
+static void cxl_port_add_hdmd(struct cxl_port *endpoint, int val)
+{
+	struct cxl_port *parent = parent_port_of(endpoint);
+	struct cxl_port *port = endpoint;
+
+	if (!parent || !is_cxl_cachedev(endpoint->uport_dev))
+		return;
+
+	while (parent && !is_cxl_root(parent)) {
+		port = parent;
+		parent = parent_port_of(port);
+	}
+
+	port->num_hdmd += val;
+}
+
 int cxl_allocate_cache_id(struct cxl_cachedev *cxlcd)
 {
 	struct ida *ida = find_cache_id_ida(cxlcd->endpoint);
@@ -407,6 +440,9 @@ int cxl_allocate_cache_id(struct cxl_cachedev *cxlcd)
 	if (id < 0)
 		return id;
 
+	if (cxlcd->cxlds->hdmd)
+		cxl_port_add_hdmd(cxlcd->endpoint, 1);
+
 	cxlcd->cache_id = id;
 	return 0;
 }
@@ -418,6 +454,9 @@ void cxl_free_cache_id(struct cxl_cachedev *cxlcd)
 
 	if (ida)
 		ida_free(ida, cxlcd->cache_id);
+
+	if (cxlcd->cxlds->hdmd)
+		cxl_port_add_hdmd(cxlcd->endpoint, -1);
 }
 EXPORT_SYMBOL_FOR_MODULES(cxl_free_cache_id, "cxl_cache");
 
@@ -489,17 +528,18 @@ int cxl_cachedev_program_cache_id(struct cxl_cachedev *cxlcd)
 {
 	struct cxl_dport *dport = cxlcd->endpoint->parent_dport;
 	struct cxl_port *port = dport->port;
+	bool hdmd = cxlcd->cxlds->hdmd;
 	bool endpoint = true;
 	int rc;
 
 	while (!is_cxl_root(port)) {
 		rc = cxl_cid_program_decoder(dport, cxlcd->cache_id,
-					     endpoint);
+					     endpoint, hdmd);
 		if (rc)
 			goto err;
 
 		rc = cxl_cid_program_table_entry(port, cxlcd->cache_id,
-						 dport->port_id);
+						 dport->port_id, hdmd);
 		if (rc)
 			goto err;
 
diff --git a/drivers/cxl/cxl.h b/drivers/cxl/cxl.h
index 4cdc26dbdb6a..63efe98d1c89 100644
--- a/drivers/cxl/cxl.h
+++ b/drivers/cxl/cxl.h
@@ -229,6 +229,7 @@ static inline int ways_to_eiw(unsigned int ways, u8 *eiw)
 /* CXL 4.0 8.2.4.28.1 CXL Cache ID Route Table Capability Structure */
 #define CXL_CACHE_ID_RT_CAP_OFFSET 0x0
 #define   CXL_CACHE_ID_RT_CAP_TARGET_CNT GENMASK(4, 0)
+#define   CXL_CACHE_ID_RT_CAP_HDMD_MAX GENMASK(11, 8)
 #define   CXL_CACHE_ID_RT_CAP_COMMIT_REQ BIT(16)
 #define CXL_CACHE_ID_RT_CTRL_OFFSET 0x4
 #define   CXL_CACHE_ID_RT_CTRL_COMMIT BIT(0)
@@ -603,6 +604,7 @@ struct cxl_dax_region {
  * @pci_latency: Upstream latency in picoseconds
  * @component_reg_phys: Physical address of component register
  * @cache_ida: ida used for cache id allocations on host bridges
+ * @num_hdmd: Number of devices using HDM-D flows below this port
  */
 struct cxl_port {
 	struct device dev;
@@ -629,6 +631,7 @@ struct cxl_port {
 	long pci_latency;
 	resource_size_t component_reg_phys;
 	struct ida cache_ida;
+	u32 num_hdmd;
 };
 
 struct cxl_root;
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 28+ messages in thread

* [PATCH 12/15] cxl/cache: Add snoop filter creation and set up
  2026-09-23 17:33 [PATCH 00/15] Add initial CXL.cache support Ben Cheatham
                   ` (10 preceding siblings ...)
  2026-09-23 17:33 ` [PATCH 11/15] cxl/core: Add support for HDM-D cache id programming Ben Cheatham
@ 2026-09-23 17:33 ` Ben Cheatham
  2026-09-23 17:50   ` sashiko-bot
  2026-09-23 17:33 ` [PATCH 13/15] cxl/cache: Add snoop filter allocation Ben Cheatham
                   ` (3 subsequent siblings)
  15 siblings, 1 reply; 28+ messages in thread
From: Ben Cheatham @ 2026-09-23 17:33 UTC (permalink / raw)
  To: linux-cxl, dave, jic23, dave.jiang, alison.schofield
  Cc: linux-iommu, benjamin.cheatham, terry.bowman, robert.richter

The CXL Snoop Filter capability (CXL 4.0 8.2.4.23) is used by the host
to track CXL.cache device transactions. CXL-enabled PCIe root ports
belong to one CXL snoop filter, indicated by the root port's group id
programmed by firmware. There may be multiple groups and multiple snoop
filters in a system.

Add the capability to track the system's snoop filters to the CXL core
and point dports to these filters as part of probe. Allocation of snoop
filter capacity will come in a later commit.

Signed-off-by: Ben Cheatham <Benjamin.Cheatham@amd.com>
---
 drivers/cxl/core/cache.c | 119 +++++++++++++++++++++++++++++++++++++++
 drivers/cxl/core/core.h  |   7 +++
 drivers/cxl/core/port.c  |   3 +
 drivers/cxl/core/regs.c  |   7 +++
 drivers/cxl/cxl.h        |  16 +++++-
 drivers/cxl/port.c       |  10 ++++
 include/cxl/cxl.h        |   6 +-
 7 files changed, 166 insertions(+), 2 deletions(-)

diff --git a/drivers/cxl/core/cache.c b/drivers/cxl/core/cache.c
index 7810772326e9..f07c49cfacec 100644
--- a/drivers/cxl/core/cache.c
+++ b/drivers/cxl/core/cache.c
@@ -6,10 +6,14 @@
 #include <cxlcache.h>
 
 #include "cxlpci.h"
+#include "core.h"
 
 #define CXL_CACHE_ID_COMMIT_MAXTMO_US (5 * USEC_PER_SEC)
 #define CACHE_DECODER_MAX_CACHE_ID (15)
 
+static DEFINE_XARRAY(snoop_filters);
+static DECLARE_RWSEM(snoop_rwsem);
+
 int cxl_port_map_cache_id_rt(struct cxl_port *port)
 {
 	struct cxl_register_map *map = &port->reg_map;
@@ -566,3 +570,118 @@ void cxl_cachedev_deprogram_cache_id(struct cxl_cachedev *cxlcd)
 	__deprogram_cache_id(cxlcd, &root->port);
 }
 EXPORT_SYMBOL_FOR_MODULES(cxl_cachedev_deprogram_cache_id, "cxl_cache");
+
+/**
+ * struct cxl_snoop_filter - CXL snoop filter instance for tracking CXL.cache
+ * devices below a dport
+ *
+ * @lock: Used for allocations
+ * @avail: Available capacity left in the filter
+ * @size: Size of the filter
+ * @id: Group ID of the filter
+ */
+struct cxl_snoop_filter {
+	struct mutex lock;
+	u64 avail;
+	u64 size;
+	int id;
+};
+
+static struct cxl_snoop_filter *create_snoop_filter(u64 size, int id)
+{
+	struct cxl_snoop_filter *sf;
+
+	sf = kzalloc_obj(*sf, GFP_KERNEL);
+	if (!sf)
+		return ERR_PTR(-ENOMEM);
+
+	sf->id = id;
+	sf->size = sf->avail = size;
+	mutex_init(&sf->lock);
+
+	return sf;
+}
+
+static void destroy_snoop_filter(struct cxl_snoop_filter *sf)
+{
+	lockdep_assert_held_write(&snoop_rwsem);
+
+	mutex_destroy(&sf->lock);
+	kfree(sf);
+}
+
+static struct cxl_snoop_filter *find_or_add_snoop_filter(u64 size, int id)
+{
+	struct cxl_snoop_filter *sf;
+	int rc;
+
+	guard(rwsem_write)(&snoop_rwsem);
+	sf = xa_load(&snoop_filters, id);
+	if (sf) {
+		if (sf->size != size)
+			pr_warn("Mismatched snoop filter (gid: %d) size: found %llu, expected %llu",
+				id, sf->size, size);
+
+		return sf;
+	}
+
+	sf = create_snoop_filter(size, id);
+	if (IS_ERR_OR_NULL(sf))
+		return sf;
+
+	rc = xa_insert(&snoop_filters, id, sf, GFP_KERNEL);
+	if (rc) {
+		destroy_snoop_filter(sf);
+		return ERR_PTR(rc);
+	}
+
+	return sf;
+}
+
+int cxl_dport_probe_snoop_filter(struct cxl_dport *dport)
+{
+	struct cxl_snoop_filter *sf;
+	u32 group, size;
+	int id, rc;
+
+	if (!dport->reg_map.component_map.snoop.valid) {
+		dev_dbg(dport->dport_dev, "missing snoop filter capability\n");
+		return 0;
+	}
+
+	rc = cxl_map_component_regs(&dport->reg_map, &dport->regs.component,
+				    BIT(CXL_CM_CAP_CAP_ID_SNOOP));
+	if (rc)
+		return rc;
+
+	group = readl(dport->regs.snoop + CXL_SNOOP_FILTER_GROUP_ID_OFFSET);
+	id = FIELD_GET(CXL_SNOOP_FILTER_GROUP_ID_MASK, group);
+
+	size = readl(dport->regs.snoop + CXL_SNOOP_FILTER_SIZE_OFFSET);
+	if (!size) {
+		dev_dbg(dport->dport_dev, "CXL snoop filter has no capacity\n");
+		return 0;
+	}
+
+	sf = find_or_add_snoop_filter(size, id);
+	if (IS_ERR_OR_NULL(sf)) {
+		return PTR_ERR(sf);
+	}
+
+	dev_dbg(dport->dport_dev, "assigned snoop filter gid %d\n", sf->id);
+	dport->snoop = id;
+	return 0;
+}
+EXPORT_SYMBOL_NS_GPL(cxl_dport_probe_snoop_filter, "CXL");
+
+void cxl_destroy_snoop_filters(void)
+{
+	struct cxl_snoop_filter *sf;
+	unsigned long id;
+
+	guard(rwsem_write)(&snoop_rwsem);
+	xa_for_each(&snoop_filters, id, sf)
+		destroy_snoop_filter(sf);
+
+	xa_destroy(&snoop_filters);
+}
diff --git a/drivers/cxl/core/core.h b/drivers/cxl/core/core.h
index 35eaf636adc9..7a6a496a5017 100644
--- a/drivers/cxl/core/core.h
+++ b/drivers/cxl/core/core.h
@@ -229,4 +229,11 @@ int cxl_set_feature(struct cxl_mailbox *cxl_mbox, const uuid_t *feat_uuid,
 
 resource_size_t cxl_rcd_component_reg_phys(struct device *dev,
 					   struct cxl_dport *dport);
+
+#if IS_ENABLED(CONFIG_CXL_CACHE)
+void cxl_destroy_snoop_filters(void);
+#else
+static inline void cxl_destroy_snoop_filters(void) {}
+#endif
+
 #endif /* __CXL_CORE_H__ */
diff --git a/drivers/cxl/core/port.c b/drivers/cxl/core/port.c
index 047801faa020..f97ece7cddbe 100644
--- a/drivers/cxl/core/port.c
+++ b/drivers/cxl/core/port.c
@@ -1272,6 +1272,8 @@ __devm_cxl_add_dport(struct cxl_port *port, struct device *dport_dev,
 	if (!dport->rch)
 		devm_cxl_dport_ras_setup(dport);
 
+	dport->snoop = CXL_SNOOP_FILTER_NO_GROUP_ID;
+
 	rc = cxl_dport_map_cache_id_dc(dport);
 	if (rc)
 		dev_dbg(dport->dport_dev,
@@ -2629,6 +2631,7 @@ static void cxl_core_exit(void)
 	bus_unregister(&cxl_bus_type);
 	destroy_workqueue(cxl_bus_wq);
 	cxl_memdev_exit();
+	cxl_destroy_snoop_filters();
 	debugfs_remove_recursive(cxl_debugfs);
 }
 
diff --git a/drivers/cxl/core/regs.c b/drivers/cxl/core/regs.c
index 2809c72fb2eb..07c3a9926b49 100644
--- a/drivers/cxl/core/regs.c
+++ b/drivers/cxl/core/regs.c
@@ -93,6 +93,12 @@ void cxl_probe_component_regs(struct device *dev, void __iomem *base,
 			length = CXL_RAS_CAPABILITY_LENGTH;
 			rmap = &map->ras;
 			break;
+		case CXL_CM_CAP_CAP_ID_SNOOP:
+			dev_dbg(dev, "found Snoop capability (0x%x)\n",
+				offset);
+			length = CXL_SNOOP_FILTER_CAPABILITY_LENGTH;
+			rmap = &map->snoop;
+			break;
 		case CXL_CM_CAP_CAP_ID_CACHE_ID_RT: {
 			int target_cnt;
 
@@ -233,6 +239,7 @@ int cxl_map_component_regs(const struct cxl_register_map *map,
 	} mapinfo[] = {
 		{ &map->component_map.hdm_decoder, &regs->hdm_decoder },
 		{ &map->component_map.ras, &regs->ras },
+		{ &map->component_map.snoop, &regs->snoop },
 		{ &map->component_map.cidrt, &regs->cidrt },
 		{ &map->component_map.ciddc, &regs->ciddc },
 	};
diff --git a/drivers/cxl/cxl.h b/drivers/cxl/cxl.h
index 63efe98d1c89..4c4ed6123b4a 100644
--- a/drivers/cxl/cxl.h
+++ b/drivers/cxl/cxl.h
@@ -41,6 +41,7 @@ extern const struct nvdimm_security_ops *cxl_security_ops;
 
 #define   CXL_CM_CAP_CAP_ID_RAS 0x2
 #define   CXL_CM_CAP_CAP_ID_HDM 0x5
+#define   CXL_CM_CAP_CAP_ID_SNOOP 0x8
 #define   CXL_CM_CAP_CAP_ID_CACHE_ID_RT 0xD
 #define   CXL_CM_CAP_CAP_ID_CACHE_ID_DC 0xE
 #define   CXL_CM_CAP_CAP_HDM_VERSION 1
@@ -225,6 +226,11 @@ static inline int ways_to_eiw(unsigned int ways, u8 *eiw)
 #define   CXLDEV_MBOX_BG_CMD_COMMAND_VENDOR_MASK GENMASK_ULL(63, 48)
 #define CXLDEV_MBOX_PAYLOAD_OFFSET 0x20
 
+/* CXL 4.0 8.2.4.23 CXL Snoop Filter Capability Structure */
+#define CXL_SNOOP_FILTER_GROUP_ID_OFFSET 0x0
+#define   CXL_SNOOP_FILTER_GROUP_ID_MASK GENMASK(15, 0)
+#define CXL_SNOOP_FILTER_SIZE_OFFSET 0x4
+#define CXL_SNOOP_FILTER_CAPABILITY_LENGTH 0x8
 
 /* CXL 4.0 8.2.4.28.1 CXL Cache ID Route Table Capability Structure */
 #define CXL_CACHE_ID_RT_CAP_OFFSET 0x0
@@ -683,6 +689,7 @@ struct cxl_rcrb_info {
  * @coord: access coordinates (bandwidth and latency performance attributes)
  * @link_latency: calculated PCIe downstream latency
  * @gpf_dvsec: Cached GPF port DVSEC
+ * @snoop: Group id of snoop filter this dport belongs to
  */
 struct cxl_dport {
 	struct device *dport_dev;
@@ -695,6 +702,7 @@ struct cxl_dport {
 	struct access_coordinate coord[ACCESS_COORDINATE_MAX];
 	long link_latency;
 	int gpf_dvsec;
+	int snoop;
 };
 
 /**
@@ -973,11 +981,17 @@ u16 cxl_gpf_get_dvsec(struct device *dev);
 #if IS_ENABLED(CONFIG_CXL_CACHE)
 int cxl_port_map_cache_id_rt(struct cxl_port *port);
 int cxl_dport_map_cache_id_dc(struct cxl_dport *dport);
+int cxl_dport_probe_snoop_filter(struct cxl_dport *dport);
 #else
 static inline int cxl_port_map_cache_id_rt(struct cxl_port *port)
 { return -ENXIO; }
 static inline int cxl_dport_map_cache_id_dc(struct cxl_dport *dport)
 { return -ENXIO; }
+static inline int cxl_dport_probe_snoop_filter(struct cxl_dport *dport)
+{
+	dport->snoop = CXL_SNOOP_FILTER_NO_GROUP_ID;
+	return 0;
+}
 #endif
- 
+
 #endif /* __CXL_H__ */
diff --git a/drivers/cxl/port.c b/drivers/cxl/port.c
index 942b92af55d9..2293ab27729d 100644
--- a/drivers/cxl/port.c
+++ b/drivers/cxl/port.c
@@ -301,6 +301,16 @@ static struct cxl_dport *cxl_port_add_dport(struct cxl_port *port,
 	/* New dport added, update the decoder targets */
 	cxl_port_update_decoder_targets(port, dport);
 
+	/* 
+	 * cxl_cachedevs won't probe if this fails, but it's not an error for
+	 * cxl_memdevs
+	 */
+	rc = cxl_dport_probe_snoop_filter(dport);
+	if (rc)
+		dev_info(dport->dport_dev,
+			 "Failed to find or create a CXL snoop filter: %d\n",
+			 rc);
+
 	dev_dbg(&port->dev, "dport%d:%s added\n", dport->port_id,
 		dev_name(dport_dev));
 
diff --git a/include/cxl/cxl.h b/include/cxl/cxl.h
index b99a9727e284..22b9c8c9c06e 100644
--- a/include/cxl/cxl.h
+++ b/include/cxl/cxl.h
@@ -34,12 +34,14 @@ struct cxl_regs {
 	 * Common set of CXL Component register block base pointers
 	 * @hdm_decoder: CXL 2.0 8.2.5.12 CXL HDM Decoder Capability Structure
 	 * @ras: CXL 2.0 8.2.5.9 CXL RAS Capability Structure
+	 * @snoop: CXL 4.0 8.2.4.23 CXL Snoop Filter Capability Structure
 	 * @cidrt: CXL 4.0 8.2.4.28 CXL Cache ID Route Table Capability Structure
 	 * @ciddc: CXL 4.0 8.2.4.29 CXL Cache ID Decoder Capability Structure
 	 */
 	struct_group_tagged(cxl_component_regs, component,
 		void __iomem *hdm_decoder;
 		void __iomem *ras;
+		void __iomem *snoop;
 		void __iomem *cidrt;
 		void __iomem *ciddc;
 	);
@@ -84,6 +86,7 @@ struct cxl_reg_map {
 struct cxl_component_reg_map {
 	struct cxl_reg_map hdm_decoder;
 	struct cxl_reg_map ras;
+	struct cxl_reg_map snoop;
 	struct cxl_reg_map cidrt;
 	struct cxl_reg_map ciddc;
 };
@@ -156,7 +159,8 @@ struct cxl_dpa_partition {
 #define CXL_NR_PARTITIONS_MAX 2
 
 #define CXL_CACHE_ID_NO_ID (-1)
- 
+#define CXL_SNOOP_FILTER_NO_GROUP_ID (-1)
+
 /**
  * struct cxl_cache_state - CXL cache device state for use by external drivers
  * @size: Size of device's cache
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 28+ messages in thread

* [PATCH 13/15] cxl/cache: Add snoop filter allocation
  2026-09-23 17:33 [PATCH 00/15] Add initial CXL.cache support Ben Cheatham
                   ` (11 preceding siblings ...)
  2026-09-23 17:33 ` [PATCH 12/15] cxl/cache: Add snoop filter creation and set up Ben Cheatham
@ 2026-09-23 17:33 ` Ben Cheatham
  2026-09-23 17:58   ` sashiko-bot
  2026-09-23 17:33 ` [PATCH 14/15] iommu, cxl: Configure IOMMU for CXL.cache Ben Cheatham
                   ` (2 subsequent siblings)
  15 siblings, 1 reply; 28+ messages in thread
From: Ben Cheatham @ 2026-09-23 17:33 UTC (permalink / raw)
  To: linux-cxl, dave, jic23, dave.jiang, alison.schofield
  Cc: linux-iommu, benjamin.cheatham, terry.bowman, robert.richter

Add snoop filter capacity allocation for CXL.cache devices. CXL.cache
devices are expected to allocate snoop filter capacity equal to the size
of the address range they will cache using CXL.cache before using the
protocol (i.e. before allocating DMA space).

Devices may also indicate whether a CXL snoop filter being full is a
failure condition using the @strict_snoop member of struct
cxl_cache_state in struct cxl_dev_state. If any allocations on a filter
are made with @strict_snoop enabled, any allocations afterwards fail if
the filter is full regardless of @strict_snoop's setting.

A full snoop filter may cause performance degradation or unpredictable
behavior depending on the host's snoop filter implementation.

Signed-off-by: Ben Cheatham <Benjamin.Cheatham@amd.com>
---
 drivers/cxl/cache.c       |  26 +++++++-
 drivers/cxl/core/cache.c  | 125 +++++++++++++++++++++++++++++++++++++-
 drivers/cxl/core/memdev.c |   1 +
 include/cxl/cxl.h         |  13 ++++
 4 files changed, 162 insertions(+), 3 deletions(-)

diff --git a/drivers/cxl/cache.c b/drivers/cxl/cache.c
index 40d1e8330df7..b4574a76ac0b 100644
--- a/drivers/cxl/cache.c
+++ b/drivers/cxl/cache.c
@@ -179,6 +179,20 @@ struct cxl_cachedev *devm_cxl_add_cachedev(struct cxl_dev_state *cxlds)
 }
 EXPORT_SYMBOL_NS_GPL(devm_cxl_add_cachedev, "CXL");
 
+static int cxl_cachedev_find_snoop_gid(struct cxl_cachedev *cxlcd)
+{
+	struct cxl_dport *iter;
+
+	for (iter = cxlcd->endpoint->parent_dport;
+	     iter && !is_cxl_root(iter->port);
+	     iter = iter->port->parent_dport) {
+		if (iter->snoop != CXL_SNOOP_FILTER_NO_GROUP_ID)
+			return iter->snoop;
+	}
+
+	return -ENXIO;
+}
+
 static int cxl_cache_probe(struct device *dev)
 {
 	struct cxl_cachedev *cxlcd = to_cxl_cachedev(dev);
@@ -236,7 +250,17 @@ static int cxl_cache_probe(struct device *dev)
 	if (rc)
 		return rc;
 
-	return devm_add_action_or_reset(dev, deprogram_cache_id, cxlcd);
+	rc = devm_add_action_or_reset(dev, deprogram_cache_id, cxlcd);
+ 	if (rc)
+ 		return rc;
+
+	rc = cxl_cachedev_find_snoop_gid(cxlcd);
+	if (rc < 0)
+		return rc;
+
+	cxlcd->cxlds->cstate.gid = rc;
+
+ 	return 0;
 }
 
 static struct cxl_driver cxl_cache_driver = {
diff --git a/drivers/cxl/core/cache.c b/drivers/cxl/core/cache.c
index f07c49cfacec..429c997b2ada 100644
--- a/drivers/cxl/core/cache.c
+++ b/drivers/cxl/core/cache.c
@@ -571,17 +571,28 @@ void cxl_cachedev_deprogram_cache_id(struct cxl_cachedev *cxlcd)
 }
 EXPORT_SYMBOL_FOR_MODULES(cxl_cachedev_deprogram_cache_id, "cxl_cache");
 
+struct snoop_allocation {
+	struct cxl_cachedev *cxlcd;
+	bool strict;
+	u64 size;
+	int id;
+};
+
 /**
  * struct cxl_snoop_filter - CXL snoop filter instance for tracking CXL.cache
  * devices below a dport
  *
+ * @allocations: Allocations on this filter
  * @lock: Used for allocations
+ * @strict: Number of allocations from devices that disallow oversubscription
  * @avail: Available capacity left in the filter
  * @size: Size of the filter
  * @id: Group ID of the filter
  */
 struct cxl_snoop_filter {
+	struct xarray allocations;
 	struct mutex lock;
+	u32 strict;
 	u64 avail;
 	u64 size;
 	int id;
@@ -597,15 +608,37 @@ static struct cxl_snoop_filter *create_snoop_filter(u64 size, int id)
 
 	sf->id = id;
 	sf->size = sf->avail = size;
+	xa_init(&sf->allocations);
 	mutex_init(&sf->lock);
 
 	return sf;
 }
 
-static void destroy_snoop_filter(struct cxl_snoop_filter *sf)
+static void free_snoop_capacity(void *);
+
+static void __free_snoop_capacity(struct snoop_allocation *alloc)
 {
-	lockdep_assert_held_write(&snoop_rwsem);
+	struct cxl_snoop_filter *sf;
+
+	sf = xa_load(&snoop_filters, alloc->id);
+	if (sf) {
+		scoped_guard(mutex, &sf->lock) {
+			sf->avail += alloc->size;
+
+			if (alloc->strict)
+				sf->strict--;
+
+			xa_erase(&sf->allocations, (unsigned long)alloc);
+		}
+	}
 
+	put_device(&alloc->cxlcd->dev);
+	kfree(alloc);
+}
+
+static void destroy_snoop_filter(struct cxl_snoop_filter *sf)
+{
+	xa_destroy(&sf->allocations);
 	mutex_destroy(&sf->lock);
 	kfree(sf);
 }
@@ -638,6 +671,94 @@ static struct cxl_snoop_filter *find_or_add_snoop_filter(u64 size, int id)
 	return sf;
 }
 
+static struct snoop_allocation *
+snoop_alloc_capacity(struct cxl_snoop_filter *sf, struct cxl_cachedev *cxlcd,
+		     u64 size)
+{
+	struct cxl_cache_state *cstate = &cxlcd->cxlds->cstate;
+	struct snoop_allocation *alloc;
+	int rc;
+
+	guard(mutex)(&sf->lock);
+	if (sf->avail < size) {
+		/*
+		 * Either this device or a device already using the filter
+		 * can't use an oversubscribed filter
+		 */
+		if (cstate->strict_snoop || sf->strict > 0)
+			return ERR_PTR(-ENOSPC);
+
+		size = sf->avail;
+	}
+
+	alloc = kmalloc_obj(*alloc);
+	if (!alloc)
+		return ERR_PTR(-ENOMEM);
+
+	*alloc = (struct snoop_allocation) {
+		.cxlcd = cxlcd,
+		.size = size,
+		.strict = cstate->strict_snoop,
+		.id = sf->id,
+	};
+
+	rc = xa_insert(&sf->allocations, (unsigned long)alloc, alloc,
+		       GFP_KERNEL);
+	if (rc) {
+		kfree(alloc);
+		return ERR_PTR(rc);
+	}
+
+	if (cstate->strict_snoop)
+		sf->strict++;
+
+	sf->avail -= size;
+	return alloc;
+}
+
+static void free_snoop_capacity(void *alloc)
+{
+	guard(rwsem_read)(&snoop_rwsem);
+	__free_snoop_capacity(alloc);
+}
+
+/**
+ * devm_cxl_cachedev_alloc_snoop_capacity - Allocate space in the system's snoop
+ * filter for a given CXL.cache device
+ * @cxlcd: Cache device to allocate capacity for
+ * @size: Size of allocation
+ *
+ * NOTE: CXL accelerator drivers are expected to call this function before
+ * using CXL.cache to access host memory. Failure to do so may result in
+ * unpredictable or undesirable behavior depending on the host's snoop filter
+ * implementation.
+ */
+int devm_cxl_cachedev_alloc_snoop_capacity(struct cxl_cachedev *cxlcd, u64 size)
+{
+	struct cxl_cache_state *cstate = &cxlcd->cxlds->cstate;
+	struct snoop_allocation *alloc;
+	struct cxl_snoop_filter *sf;
+
+	if (cstate->gid < 0)
+		return -EINVAL;
+
+	scoped_guard(rwsem_read, &snoop_rwsem) {
+		sf = xa_load(&snoop_filters, cstate->gid);
+		if (!sf)
+			return -ENODEV;
+
+		alloc = snoop_alloc_capacity(sf, cxlcd, size);
+		if (IS_ERR_OR_NULL(alloc))
+			return !alloc ? -ENOMEM : PTR_ERR(alloc);
+
+		get_device(&cxlcd->dev);
+	}
+
+	return devm_add_action_or_reset(&cxlcd->dev, free_snoop_capacity,
+					alloc);
+}
+EXPORT_SYMBOL_NS_GPL(devm_cxl_cachedev_alloc_snoop_capacity, "CXL");
+
 int cxl_dport_probe_snoop_filter(struct cxl_dport *dport)
 {
 	struct cxl_snoop_filter *sf;
diff --git a/drivers/cxl/core/memdev.c b/drivers/cxl/core/memdev.c
index b3419df586b9..e169301bbfcf 100644
--- a/drivers/cxl/core/memdev.c
+++ b/drivers/cxl/core/memdev.c
@@ -751,6 +751,7 @@ struct cxl_dev_state *_devm_cxl_dev_state_create(struct device *dev,
 	cxlds->cxl_dvsec = dvsec;
 	cxlds->reg_map.host = dev;
 	cxlds->reg_map.resource = CXL_RESOURCE_NONE;
+	cxlds->cstate.gid = CXL_SNOOP_FILTER_NO_GROUP_ID;
 
 	if (has_mbox)
 		cxlds->cxl_mbox.host = dev;
diff --git a/include/cxl/cxl.h b/include/cxl/cxl.h
index 22b9c8c9c06e..5d550dd70870 100644
--- a/include/cxl/cxl.h
+++ b/include/cxl/cxl.h
@@ -165,6 +165,8 @@ struct cxl_dpa_partition {
  * struct cxl_cache_state - CXL cache device state for use by external drivers
  * @size: Size of device's cache
  * @unit: Unit of device's cache in bytes
+ * @strict_snoop: Whether a full snoop filter should fail allocations
+ * @gid: Group ID of the snoop filter this device belongs to
  */
 struct cxl_cache_state {
 	/* Public for endpoint drivers */
@@ -175,6 +177,12 @@ struct cxl_cache_state {
 	 */
 	u64 size;
 	u64 unit;
+
+	/* Should be set by endpoint drivers */
+	bool strict_snoop;
+
+	/* Private for endpoint drivers */
+	int gid;
 };
 
 /**
@@ -264,9 +272,14 @@ int cxl_set_capacity(struct cxl_dev_state *cxlds, u64 capacity);
 
 #if IS_ENABLED(CONFIG_CXL_CACHE)
 struct cxl_cachedev *devm_cxl_add_cachedev(struct cxl_dev_state *cxlds);
+int devm_cxl_cachedev_alloc_snoop_capacity(struct cxl_cachedev *cxlcd,
+					   u64 size);
 #else
 static inline struct cxl_cachedev *
 devm_cxl_add_cachedev(struct cxl_dev_state *cxlds)
 { return ERR_PTR(-ENXIO); }
+static inline int
+devm_cxl_cachedev_alloc_snoop_capacity(struct cxl_cachedev *cxlcd, u64 size)
+{ return -ENXIO; }
 #endif /* CONFIG_CXL_CACHE */
 #endif /* __CXL_CXL_H__ */
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 28+ messages in thread

* [PATCH 14/15] iommu, cxl: Configure IOMMU for CXL.cache
  2026-09-23 17:33 [PATCH 00/15] Add initial CXL.cache support Ben Cheatham
                   ` (12 preceding siblings ...)
  2026-09-23 17:33 ` [PATCH 13/15] cxl/cache: Add snoop filter allocation Ben Cheatham
@ 2026-09-23 17:33 ` Ben Cheatham
  2026-09-23 17:57   ` sashiko-bot
  2026-09-23 17:33 ` [PATCH 15/15] cxl/cache: Enable CXL.cache on successful probe Ben Cheatham
  2026-09-23 17:35 ` [PATCH 00/15] Add initial CXL.cache support Cheatham, Benjamin
  15 siblings, 1 reply; 28+ messages in thread
From: Ben Cheatham @ 2026-09-23 17:33 UTC (permalink / raw)
  To: linux-cxl, dave, jic23, dave.jiang, alison.schofield
  Cc: linux-iommu, benjamin.cheatham, terry.bowman, robert.richter

Some IOMMU implementations require additional set up for enabling
ATS requests past enabling the base PCI ATS support. Create a callback
in the IOMMU core to be used by the CXL driver during device set up
that, when called, configures the IOMMU to handle ATS requests with the
CXL source bit set. For CXL.cache devices that don't support ATS, the
expectation is for the device to be attached to an identity domain, or
the iommu be disabled.

Update the AMD IOMMU driver with an implementation of the CXL ATS
callback and defaulting CXL.cache devices without ATS support to an
identity domain.

Signed-off-by: Ben Cheatham <Benjamin.Cheatham@amd.com>
---

I have a note next to the definition of FEATURE_CXLMEMATTR that it's
missing in the AMD IOMMU spec. The spec I'm referring to is the "AMD I/O
Virtualization Technology (IOMMU) Specification, Rev. 3.11, April 2026"
spec available on AMD's website.

I *think* the bit is actually supposed to be 0 based on the figures for
the register it's located in, but I had to guess since it's missing from
the register field table and isn't mentioned elsewhere AFAIK. I figured
it's fine while this is in RFC, but I'll update with the actual bit
when/if this comes out of RFC. Thanks!

---
 drivers/cxl/cache.c                 |  1 +
 drivers/cxl/core/cache.c            | 43 ++++++++++++++++++++++
 drivers/iommu/amd/amd_iommu_types.h |  4 +++
 drivers/iommu/amd/init.c            | 14 ++++++++
 drivers/iommu/amd/iommu.c           | 56 ++++++++++++++++++++++++++++-
 drivers/iommu/iommu.c               | 28 +++++++++++++++
 include/cxl/cxl.h                   | 22 ++++++++++++
 include/linux/iommu.h               |  8 +++++
 8 files changed, 175 insertions(+), 1 deletion(-)

diff --git a/drivers/cxl/cache.c b/drivers/cxl/cache.c
index b4574a76ac0b..f73bf96d231e 100644
--- a/drivers/cxl/cache.c
+++ b/drivers/cxl/cache.c
@@ -3,6 +3,7 @@
 
 #include <linux/pci.h>
 #include <cxl/cxl.h>
+#include <linux/iommu.h>
 
 #include "cxlcache.h"
 
diff --git a/drivers/cxl/core/cache.c b/drivers/cxl/core/cache.c
index 429c997b2ada..44323dfdb759 100644
--- a/drivers/cxl/core/cache.c
+++ b/drivers/cxl/core/cache.c
@@ -2,6 +2,7 @@
 /* Copyright (C) Advanced Micro Device, Inc. */
 
 #include <linux/iopoll.h>
+#include <linux/iommu.h>
 #include <linux/pci.h>
 #include <cxlcache.h>
 
@@ -806,3 +807,45 @@ void cxl_destroy_snoop_filters(void)
 
 	xa_destroy(&snoop_filters);
 }
+
+/**
+ * cxl_cache_configure_iommu() - Configure a device's IOMMU for CXL.cache
+ * @cxlds: struct cxl_dev_state of a cxl_cachedev that has been through
+ *	   CXL.cache probe
+ *
+ * Fails if the underlying PCI device supports ATS and IOMMU can't be
+ * configured, or if the device doesn't support ATS and is not attached to an
+ * identity IOMMU domain.
+ */
+int cxl_cache_configure_iommu(struct cxl_dev_state *cxlds)
+{
+	struct device *dev = cxlds->dev;
+	struct iommu_domain *domain;
+	int rc;
+
+	lockdep_assert_held(&dev->mutex);
+
+	if (!dev->iommu || !dev->iommu->iommu_dev)
+		return 0;
+
+	if (!device_iommu_capable(dev, IOMMU_CAP_PCI_ATS_SUPPORTED)) {
+		domain = iommu_get_domain_for_dev(dev);
+		if (!domain || domain->type != IOMMU_DOMAIN_IDENTITY)
+			return -EINVAL;
+
+		return 0;
+	}
+
+	rc = iommu_enable_cxl_ats(cxlds->dev);
+	if (rc == -EOPNOTSUPP) {
+		dev_warn(cxlds->dev,
+			"IOMMU doesn't support enabling CXL ATS requests; CXL.cache may not function properly.");
+		rc = 0;
+	} else if (rc) {
+		dev_err(cxlds->dev, "Failed to enable CXL ATS requests: %d\n",
+			rc);
+	}
+
+	return rc;
+}
+EXPORT_SYMBOL_NS_GPL(cxl_cache_configure_iommu, "CXL");
diff --git a/drivers/iommu/amd/amd_iommu_types.h b/drivers/iommu/amd/amd_iommu_types.h
index 3dbe20023456..696d9879ce9b 100644
--- a/drivers/iommu/amd/amd_iommu_types.h
+++ b/drivers/iommu/amd/amd_iommu_types.h
@@ -107,6 +107,7 @@
 
 
 /* Extended Feature 2 Bits */
+#define FEATURE_CXLMEMATTR	BIT_ULL(0) /* WARNING: This bit isn't in the spec as of 09/26 */
 #define FEATURE_SEVSNPIO_SUP	BIT_ULL(1)
 #define FEATURE_GCR3TRPMODE	BIT_ULL(3)
 #define FEATURE_SNPAVICSUP	GENMASK_ULL(7, 5)
@@ -189,6 +190,7 @@
 #define CONTROL_EPH_EN		45
 #define CONTROL_XT_EN		50
 #define CONTROL_INTCAPXT_EN	51
+#define CONTROL_CXLMEMATTR_EN	57
 #define CONTROL_GCR3TRPMODE	58
 #define CONTROL_IRTCACHEDIS	59
 #define CONTROL_SNPAVIC_EN	61
@@ -355,6 +357,7 @@
  */
 #define DTE_FLAG_V	BIT_ULL(0)
 #define DTE_FLAG_TV	BIT_ULL(1)
+#define DTE_CXLMEM_MASK	GENMASK_ULL(6, 4)
 #define DTE_FLAG_HAD	(3ULL << 7)
 #define DTE_MODE_MASK	GENMASK_ULL(11, 9)
 #define DTE_HOST_TRP	GENMASK_ULL(51, 12)
@@ -831,6 +834,7 @@ struct iommu_dev_data {
 	u8 ppr          :1;		  /* Enable device PPR support */
 	bool use_vapic;			  /* Enable device to use vapic mode */
 	bool defer_attach;
+	u8 cxl_memattr;			  /* CXLMemAttr DTE setting */
 
 	struct ratelimit_state rs;        /* Ratelimit IOPF messages */
 };
diff --git a/drivers/iommu/amd/init.c b/drivers/iommu/amd/init.c
index 40726dfef273..e11f98db826e 100644
--- a/drivers/iommu/amd/init.c
+++ b/drivers/iommu/amd/init.c
@@ -1125,6 +1125,14 @@ static void iommu_enable_gt(struct amd_iommu *iommu)
 		iommu_feature_enable(iommu, CONTROL_GCR3TRPMODE);
 }
 
+static void iommu_enable_cxlmemattr(struct amd_iommu *iommu)
+{
+	if (!check_feature2(FEATURE_CXLMEMATTR) || !amd_iommu_iotlb_sup)
+		return;
+
+	iommu_feature_enable(iommu, CONTROL_CXLMEMATTR_EN);
+}
+
 /* sets a specific bit in the device table entry. */
 static void set_dte_bit(struct dev_table_entry *dte, u8 bit)
 {
@@ -2245,6 +2253,12 @@ static int __init iommu_init_pci(struct amd_iommu *iommu)
 			return ret;
 	}
 
+	/* 
+	 * Set the CXLMemAttr control bit and do the Device Table Entry parts
+	 * of set up when a CXL.cache device shows up.
+	 */
+	iommu_enable_cxlmemattr(iommu);
+
 	ret = iommu_device_register(&iommu->iommu, &amd_iommu_ops, NULL);
 	if (ret || amd_iommu_pgtable == PD_MODE_NONE) {
 		/*
diff --git a/drivers/iommu/amd/iommu.c b/drivers/iommu/amd/iommu.c
index 4dc306a4b5c6..78d34b907450 100644
--- a/drivers/iommu/amd/iommu.c
+++ b/drivers/iommu/amd/iommu.c
@@ -41,6 +41,7 @@
 #include <asm/dma.h>
 #include <uapi/linux/iommufd.h>
 #include <linux/generic_pt/iommu.h>
+#include <cxl/cxl.h>
 
 #include "amd_iommu.h"
 #include "iommufd.h"
@@ -2199,6 +2200,13 @@ static void set_dte_passthrough(struct iommu_dev_data *dev_data,
 
 }
 
+static void set_dte_cxl_mem_attr(struct iommu_dev_data *dev_data,
+				 struct dev_table_entry *new)
+{
+	new->data[0] &= ~DTE_CXLMEM_MASK;
+	new->data[0] |= FIELD_PREP(DTE_CXLMEM_MASK, dev_data->cxl_memattr);
+}
+
 static void set_dte_entry(struct amd_iommu *iommu,
 			  struct iommu_dev_data *dev_data,
 			  phys_addr_t top_paddr, unsigned int top_level)
@@ -2217,11 +2225,14 @@ static void set_dte_entry(struct amd_iommu *iommu,
 	else if (domain->domain.type == IOMMU_DOMAIN_IDENTITY)
 		set_dte_passthrough(dev_data, domain, &new);
 	else if ((domain->domain.type & __IOMMU_DOMAIN_PAGING) &&
-		 domain->pd_mode == PD_MODE_V1)
+		   domain->pd_mode == PD_MODE_V1)
 		set_dte_v1(dev_data, domain, domain->id, top_paddr, top_level, &new);
 	else
 		WARN_ON(true);
 
+	if (amd_iommu_iotlb_sup && check_feature2(FEATURE_CXLMEMATTR))
+		set_dte_cxl_mem_attr(dev_data, &new);
+
 	amd_iommu_update_dte(iommu, dev_data, &new);
 
 	/*
@@ -3190,6 +3201,14 @@ static int amd_iommu_def_domain_type(struct device *dev)
 		return IOMMU_DOMAIN_IDENTITY;
 	}
 
+	/*
+	 * CXL.cache devices that have no ATS capability need a passthrough
+	 * domain to function correctly.
+	 */
+	if (dev_is_pci(dev) && !pci_ats_supported(to_pci_dev(dev)) &&
+	    cxl_cache_supported(to_pci_dev(dev)))
+		return IOMMU_DOMAIN_IDENTITY;
+
 	return 0;
 }
 
@@ -3199,6 +3218,40 @@ static bool amd_iommu_enforce_cache_coherency(struct iommu_domain *domain)
 	return true;
 }
 
+static int amd_iommu_enable_cxl_ats(struct device *dev)
+{
+	struct iommu_dev_data *dev_data = dev_iommu_priv_get(dev);
+	int ret = 0;
+
+	if (!dev_data || !amd_iommu_iotlb_sup)
+		return -EINVAL;
+
+	if (!check_feature2(FEATURE_CXLMEMATTR))
+		return -ENXIO;
+
+	mutex_lock(&dev_data->mutex);
+
+	/* Already set up */
+	if (dev_data->cxl_memattr != 0)
+		goto out;
+
+	if (!dev_data->ats_enabled) {
+		ret = -EINVAL;
+		goto out;
+	}
+
+	/*
+	 * Give device the choice of whether to use CXL.io or CXL.cache and
+	 * whether accesses are snooped
+	 */
+	dev_data->cxl_memattr = 1;
+	dev_update_dte(dev_data, true);
+
+out:
+	mutex_unlock(&dev_data->mutex);
+	return ret;
+}
+
 const struct iommu_ops amd_iommu_ops = {
 	.capable = amd_iommu_capable,
 	.hw_info = amd_iommufd_hw_info,
@@ -3216,6 +3269,7 @@ const struct iommu_ops amd_iommu_ops = {
 	.page_response = amd_iommu_page_response,
 	.get_viommu_size = amd_iommufd_get_viommu_size,
 	.viommu_init = amd_iommufd_viommu_init,
+	.enable_cxl_ats = amd_iommu_enable_cxl_ats,
 };
 
 #ifdef CONFIG_IRQ_REMAP
diff --git a/drivers/iommu/iommu.c b/drivers/iommu/iommu.c
index cd1bca7ede9a..5d82b5d09a3f 100644
--- a/drivers/iommu/iommu.c
+++ b/drivers/iommu/iommu.c
@@ -4222,6 +4222,34 @@ void pci_dev_reset_iommu_done(struct pci_dev *pdev)
 }
 EXPORT_SYMBOL_GPL(pci_dev_reset_iommu_done);
 
+/*
+ * iommu_enable_cxl_ats() - Enable CXL source bit extension of ATS for the
+ * given device.
+ * @dev: CXL.cache-capable PCIe device to enable capability for
+ *
+ * Returns:
+ * + -ENODEV if no IOMMU present
+ * + -EOPNOTSUPP if IOMMU isn't CXL-aware or doesn't need extra set up
+ * + Result of enablement callback otherwise
+ *
+ * Required by devices looking to use CXL.cache with non-identity IOMMU domais:
+ * see CXL 4.0 specification, section 3.1.6 "Memory Type Indication on ATS"
+ */
+int iommu_enable_cxl_ats(struct device *dev)
+{
+	const struct iommu_ops *ops;
+
+	if (!dev_has_iommu(dev))
+		return -ENODEV;
+
+	ops = dev_iommu_ops(dev);
+	if (!ops->enable_cxl_ats)
+		return -EOPNOTSUPP;
+
+	return ops->enable_cxl_ats(dev);
+}
+EXPORT_SYMBOL_GPL(iommu_enable_cxl_ats);
+
 #if IS_ENABLED(CONFIG_IRQ_MSI_IOMMU)
 /**
  * iommu_dma_prepare_msi() - Map the MSI page in the IOMMU domain
diff --git a/include/cxl/cxl.h b/include/cxl/cxl.h
index 5d550dd70870..0471d5c00144 100644
--- a/include/cxl/cxl.h
+++ b/include/cxl/cxl.h
@@ -8,6 +8,7 @@
 #include <linux/node.h>
 #include <linux/ioport.h>
 #include <cxl/mailbox.h>
+#include <linux/pci.h>
 
 /**
  * enum cxl_devtype - delineate type-2 from a generic type-3 device
@@ -274,6 +275,23 @@ int cxl_set_capacity(struct cxl_dev_state *cxlds, u64 capacity);
 struct cxl_cachedev *devm_cxl_add_cachedev(struct cxl_dev_state *cxlds);
 int devm_cxl_cachedev_alloc_snoop_capacity(struct cxl_cachedev *cxlcd,
 					   u64 size);
+int cxl_cache_configure_iommu(struct cxl_dev_state *cxlds);
+
+static inline bool cxl_cache_supported(struct pci_dev *pdev)
+{
+	int offset;
+	u16 cap;
+
+	offset = pci_find_dvsec_capability(pdev, PCI_VENDOR_ID_CXL,
+					   PCI_DVSEC_CXL_DEVICE);
+	if (!offset)
+		return false;
+
+	if (pci_read_config_word(pdev, offset + PCI_DVSEC_CXL_CAP, &cap))
+		return false;
+
+	return cap & PCI_DVSEC_CXL_CACHE_CAPABLE;
+}
 #else
 static inline struct cxl_cachedev *
 devm_cxl_add_cachedev(struct cxl_dev_state *cxlds)
@@ -281,5 +299,9 @@ devm_cxl_add_cachedev(struct cxl_dev_state *cxlds)
 static inline int
 devm_cxl_cachedev_alloc_snoop_capacity(struct cxl_cachedev *cxlcd, u64 size)
 { return -ENXIO; }
+static inline int cxl_cache_configure_iommu(struct cxl_dev_state *cxlds)
+{ return -ENXIO; }
+static inline bool cxl_cache_supported(struct pci_dev *pdev)
+{ return false; }
 #endif /* CONFIG_CXL_CACHE */
 #endif /* __CXL_CXL_H__ */
diff --git a/include/linux/iommu.h b/include/linux/iommu.h
index ac43b8b93f14..aa1e42ec5e38 100644
--- a/include/linux/iommu.h
+++ b/include/linux/iommu.h
@@ -734,6 +734,7 @@ struct iommu_ops {
 	int (*viommu_init)(struct iommufd_viommu *viommu,
 			   struct iommu_domain *parent_domain,
 			   const struct iommu_user_data *user_data);
+	int (*enable_cxl_ats)(struct device *dev);
 
 	const struct iommu_domain_ops *default_domain_ops;
 	struct module *owner;
@@ -1222,6 +1223,8 @@ void iommu_detach_device_pasid(struct iommu_domain *domain,
 ioasid_t iommu_alloc_global_pasid(struct device *dev);
 void iommu_free_global_pasid(ioasid_t pasid);
 
+int iommu_enable_cxl_ats(struct device *dev);
+
 /* PCI device reset functions */
 int pci_dev_reset_iommu_prepare(struct pci_dev *pdev);
 void pci_dev_reset_iommu_done(struct pci_dev *pdev);
@@ -1549,6 +1552,11 @@ static inline ioasid_t iommu_alloc_global_pasid(struct device *dev)
 
 static inline void iommu_free_global_pasid(ioasid_t pasid) {}
 
+static inline int iommu_enable_cxl_ats(struct device *dev)
+{
+	return -ENODEV;
+}
+
 static inline int pci_dev_reset_iommu_prepare(struct pci_dev *pdev)
 {
 	return 0;
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 28+ messages in thread

* [PATCH 15/15] cxl/cache: Enable CXL.cache on successful probe
  2026-09-23 17:33 [PATCH 00/15] Add initial CXL.cache support Ben Cheatham
                   ` (13 preceding siblings ...)
  2026-09-23 17:33 ` [PATCH 14/15] iommu, cxl: Configure IOMMU for CXL.cache Ben Cheatham
@ 2026-09-23 17:33 ` Ben Cheatham
  2026-09-23 17:35 ` [PATCH 00/15] Add initial CXL.cache support Cheatham, Benjamin
  15 siblings, 0 replies; 28+ messages in thread
From: Ben Cheatham @ 2026-09-23 17:33 UTC (permalink / raw)
  To: linux-cxl, dave, jic23, dave.jiang, alison.schofield
  Cc: linux-iommu, benjamin.cheatham, terry.bowman, robert.richter

Enable CXL.cache as the final action of cxl_cache::probe(). After this,
devices may allocate snoop filter capacity and set up DMA regions for
using CXL.cache.

Signed-off-by: Ben Cheatham <Benjamin.Cheatham@amd.com>
---
 drivers/cxl/cache.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/drivers/cxl/cache.c b/drivers/cxl/cache.c
index f73bf96d231e..a321c4a55f82 100644
--- a/drivers/cxl/cache.c
+++ b/drivers/cxl/cache.c
@@ -261,7 +261,7 @@ static int cxl_cache_probe(struct device *dev)
 
 	cxlcd->cxlds->cstate.gid = rc;
 
- 	return 0;
+	return devm_cxl_enable_cache(&cxlcd->dev, cxlcd->cxlds);
 }
 
 static struct cxl_driver cxl_cache_driver = {
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 28+ messages in thread

* Re: [PATCH 00/15] Add initial CXL.cache support
  2026-09-23 17:33 [PATCH 00/15] Add initial CXL.cache support Ben Cheatham
                   ` (14 preceding siblings ...)
  2026-09-23 17:33 ` [PATCH 15/15] cxl/cache: Enable CXL.cache on successful probe Ben Cheatham
@ 2026-09-23 17:35 ` Cheatham, Benjamin
  15 siblings, 0 replies; 28+ messages in thread
From: Cheatham, Benjamin @ 2026-09-23 17:35 UTC (permalink / raw)
  To: linux-cxl, dave, jic23, dave.jiang, alison.schofield
  Cc: linux-iommu, terry.bowman, robert.richter

This set is RFC, I realized I forgot to mark the cover letter right as I hit send...

Sorry about that!

Thanks,
Ben

On 9/23/2026 12:33 PM, Ben Cheatham wrote:
> This series adds support for CXL.cache devices to the CXL subsystem. The
> set includes the following:
> 	- Updates to the existing CXL infrastructure to enable adding
> 	  CXL.cache devices (struct cxl_cachedev) to the driver
> 	- Support for cache id and snoop filter capabilities
> 	- IOMMU updates to allow CXL.cache ATS requests
> 	- Driver for struct cxl_cachedev devices that calls all of the above
> 
> The vast majority of the support is gated behind CONFIG_CXL_CACHE since
> CXL.cache devices aren't commonplace. This set is untested since I don't
> (currently) have access to a CXL.cache-capable device. Using multiple
> devices requires cache id capabilties, which may not be available
> depending on the platform (for AMD it's Venice or later).
> 
> There is a major missing piece in this set: There's no endpoint driver that
> adds a cxl_cachedev device, so this code is currently unused. I was
> planning on only sending this out once I had a device, but the plans for
> the device I was planning to support fell through and I didn't want to
> just sit on this.
> 
> Another thing that's missing is mapping host memory for CXL.cache use. I
> think the DMA API *should* just work after these changes, but it's not
> tested. I originally had the CXL core driver set up the DMA range for
> the endpoint driver, but it didn't end up making much sense. So, I've
> left the implementation for the (eventual) endpoint driver.
> 
> Thanks,
> Ben
> 
> base-commit: cee9395acd8043be0644b25c34bfa86623f2b935 (v7.3-rc1)
> 
> Ben Cheatham (15):
>   cxl/core: Add CXL.cache device struct
>   cxl/cache: Add cxl_cache driver
>   cxl/core: Change cxl_ep_load() to use device pointer parameter
>   cxl/core: Update devm_cxl_enumerate_ports() for cxl_cachedevs
>   cxl/port: Split endpoint port probe on device type
>   cxl/core: Update devm_cxl_add_endpoint() for cxl_cachedevs
>   cxl/cache: Verify port hierarchy has CXL.cache enabled
>   cxl/core, cache: Add Cache ID register probing and init
>   cxl/core: Add Cache ID verification
>   cxl/core: Add Cache ID allocation
>   cxl/core: Add support for HDM-D cache id programming
>   cxl/cache: Add snoop filter creation and set up
>   cxl/cache: Add snoop filter allocation
>   iommu, cxl: Configure IOMMU for CXL.cache
>   cxl/cache: Enable CXL.cache on successful probe
> 
>  drivers/cxl/Kconfig                 |  14 +
>  drivers/cxl/Makefile                |   6 +-
>  drivers/cxl/cache.c                 | 285 ++++++++++
>  drivers/cxl/core/Makefile           |   2 +
>  drivers/cxl/core/cache.c            | 851 ++++++++++++++++++++++++++++
>  drivers/cxl/core/cachedev.c         | 131 +++++
>  drivers/cxl/core/core.h             |   7 +
>  drivers/cxl/core/memdev.c           |   1 +
>  drivers/cxl/core/pci.c              |  78 ++-
>  drivers/cxl/core/port.c             | 149 +++--
>  drivers/cxl/core/region.c           |  24 +-
>  drivers/cxl/core/regs.c             |  30 +
>  drivers/cxl/cxl.h                   |  74 ++-
>  drivers/cxl/cxlcache.h              |  50 ++
>  drivers/cxl/cxlmem.h                |   4 +-
>  drivers/cxl/mem.c                   |   4 +-
>  drivers/cxl/port.c                  |  97 +++-
>  drivers/iommu/amd/amd_iommu_types.h |   4 +
>  drivers/iommu/amd/init.c            |  14 +
>  drivers/iommu/amd/iommu.c           |  56 +-
>  drivers/iommu/iommu.c               |  28 +
>  include/cxl/cxl.h                   |  78 ++-
>  include/linux/iommu.h               |   8 +
>  include/uapi/linux/pci_regs.h       |   8 +
>  24 files changed, 1911 insertions(+), 92 deletions(-)
>  create mode 100644 drivers/cxl/cache.c
>  create mode 100644 drivers/cxl/core/cache.c
>  create mode 100644 drivers/cxl/core/cachedev.c
>  create mode 100644 drivers/cxl/cxlcache.h
> 


^ permalink raw reply	[flat|nested] 28+ messages in thread

* Re: [PATCH 01/15] cxl/core: Add CXL.cache device struct
  2026-09-23 17:33 ` [PATCH 01/15] cxl/core: Add CXL.cache device struct Ben Cheatham
@ 2026-09-23 17:41   ` sashiko-bot
  0 siblings, 0 replies; 28+ messages in thread
From: sashiko-bot @ 2026-09-23 17:41 UTC (permalink / raw)
  To: Ben Cheatham; +Cc: linux-cxl

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Low] cxl: unused 'attach' parameter in __devm_cxl_add_cachedev()
--

commit 16ca3df82e8ba8b03985c5445d288c9759b9ba65
Author: Ben Cheatham <Benjamin.Cheatham@amd.com>

cxl/core: Add CXL.cache device struct

Add a new CXL.cache device (struct cxl_cachedev) that is the cache
analogue to struct cxl_memdev. This device will be created by CXL type
1 & 2 endpoint drivers to enable and manage the cache capabilities of
the underlying PCIe device via the CXL core.

> diff --git a/drivers/cxl/core/cachedev.c b/drivers/cxl/core/cachedev.c
> new file mode 100644
> index 0000000000000..3a7a60165f621
> --- /dev/null
> +++ b/drivers/cxl/core/cachedev.c
[ ... ]
> +struct cxl_cachedev *__devm_cxl_add_cachedev(struct cxl_dev_state *cxlds,
> +					     void *attach)

[Severity: Low]
Is it intentional that the attach parameter is completely ignored in the
implementation of __devm_cxl_add_cachedev() in drivers/cxl/core/cachedev.c?
Any data passed to it is silently discarded, which might cause API confusion.
Does it perhaps stem from an incomplete copy of __devm_cxl_add_memdev()?

> +{
> +	struct device *dev;
> +	int rc;
> +
> +	struct cxl_cachedev *cxlcd __free(put_cxlcd) =
> +		cxl_cachedev_alloc(cxlds);
[ ... ]
> diff --git a/drivers/cxl/cxlcache.h b/drivers/cxl/cxlcache.h
> new file mode 100644
> index 0000000000000..6ca887e078a1e
> --- /dev/null
> +++ b/drivers/cxl/cxlcache.h
[ ... ]
> +struct cxl_cachedev {
> +	struct device dev;
> +	struct cxl_dev_state *cxlds;
> +	struct cxl_port *endpoint;
> +	int id;
> +	int depth;
> +};

[Severity: Low]
Should struct cxl_cachedev in drivers/cxl/cxlcache.h have an attach member
to store the attachment data provided to __devm_cxl_add_cachedev()?

> +
> +bool is_cxl_cachedev(const struct device *dev);
[ ... ]

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260923173342.5584-1-Benjamin.Cheatham@amd.com?part=1

^ permalink raw reply	[flat|nested] 28+ messages in thread

* Re: [PATCH 08/15] cxl/core, cache: Add Cache ID register probing and init
  2026-09-23 17:33 ` [PATCH 08/15] cxl/core, cache: Add Cache ID register probing and init Ben Cheatham
@ 2026-09-23 17:46   ` sashiko-bot
  0 siblings, 0 replies; 28+ messages in thread
From: sashiko-bot @ 2026-09-23 17:46 UTC (permalink / raw)
  To: Ben Cheatham; +Cc: linux-cxl

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Low] build: function redefinition in cache.c when CONFIG_CXL_CACHE=n
--

commit 666bdbfb794604b7884fcb19d8260bded28d941b
Author: Ben Cheatham <Benjamin.Cheatham@amd.com>

cxl/core, cache: Add Cache ID register probing and init

The CXL Cache ID Route Table capability and CXL Cache ID Decoder
capability are optional capabilities required when more than one CXL.cache
device is enabled under a host bridge. This commit probes for these
registers and maps them if found during the cxl_port probe.

> diff --git a/drivers/cxl/core/Makefile b/drivers/cxl/core/Makefile
> index f51575abe5af4..33420aaf40db9 100644
> --- a/drivers/cxl/core/Makefile
> +++ b/drivers/cxl/core/Makefile
> @@ -15,6 +15,7 @@ cxl_core-y += pci.o
>  cxl_core-y += hdm.o
>  cxl_core-y += pmu.o
>  cxl_core-y += cdat.o
> +cxl_core-y += cache.o

[Severity: Low]
Will this cause a build failure when CONFIG_CXL_CACHE is disabled?

Because cache.o is added unconditionally here, drivers/cxl/core/cache.c is
compiled even when CONFIG_CXL_CACHE=n. This causes a function redefinition
error for cxl_port_map_cache_id_rt() and cxl_dport_map_cache_id_dc() since
they are also defined as static inline stubs in drivers/cxl/cxl.h:

#if IS_ENABLED(CONFIG_CXL_CACHE)
int cxl_port_map_cache_id_rt(struct cxl_port *port);
int cxl_dport_map_cache_id_dc(struct cxl_dport *dport);
#else
static inline int cxl_port_map_cache_id_rt(struct cxl_port *port)
{ return -ENXIO; }
static inline int cxl_dport_map_cache_id_dc(struct cxl_dport *dport)
{ return -ENXIO; }
#endif

Should cache.o be added to cxl_core-$(CONFIG_CXL_CACHE) instead of
cxl_core-y?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260923173342.5584-1-Benjamin.Cheatham@amd.com?part=8

^ permalink raw reply	[flat|nested] 28+ messages in thread

* Re: [PATCH 05/15] cxl/port: Split endpoint port probe on device type
  2026-09-23 17:33 ` [PATCH 05/15] cxl/port: Split endpoint port probe on device type Ben Cheatham
@ 2026-09-23 17:46   ` sashiko-bot
  0 siblings, 0 replies; 28+ messages in thread
From: sashiko-bot @ 2026-09-23 17:46 UTC (permalink / raw)
  To: Ben Cheatham; +Cc: linux-cxl

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Low] cxl/port: type mismatch due to implicit pointer cast in schedule_detach()
--

commit 4ffbc8dadee1abe8b3c2deec219a795c0d6ff353
Author: Ben Cheatham <Benjamin.Cheatham@amd.com>

cxl/port: Split endpoint port probe on device type

This commit splits endpoint port probe on device type and refactors
dport RAS set up into a common routine. CXL.cache devices do not support
or need all of the set up done by endpoint port probe for CXL.mem devices.

> diff --git a/drivers/cxl/port.c b/drivers/cxl/port.c
> index edf0ff759fbf6..7c93fabfb0956 100644
> --- a/drivers/cxl/port.c
> +++ b/drivers/cxl/port.c
[ ... ]
> @@ -26,9 +27,13 @@
>   * PCIe topology.
>   */
>  
> -static void schedule_detach(void *cxlmd)
> +static void schedule_detach(void *ep_dev)
>  {
> -	schedule_cxl_memdev_detach(cxlmd);
> +	if (is_cxl_memdev(ep_dev))
> +		schedule_cxl_memdev_detach(ep_dev);
> +
> +	if (is_cxl_cachedev(ep_dev))
> +		schedule_cxl_cachedev_detach(ep_dev);
>  }

[Severity: Low]
Does this code inadvertently bypass type safety by passing ep_dev (which holds
a struct device *) directly to schedule_cxl_memdev_detach() and
schedule_cxl_cachedev_detach()?

These functions expect struct cxl_memdev * and struct cxl_cachedev *
respectively. While this avoids memory corruption right now because the dev
member is located at offset 0 in both container structures, it relies on an
implicit pointer cast.

Would it be better to use to_cxl_memdev() and to_cxl_cachedev() to explicitly
resolve the container structs before passing them?

[ ... ]
> @@ -151,6 +141,27 @@ static int cxl_endpoint_port_probe(struct cxl_port *port)
[ ... ]
> +static int cxl_mem_endpoint_port_probe(struct cxl_port *port)
> +{
> +	struct cxl_memdev *cxlmd = to_cxl_memdev(port->uport_dev);
> +	int rc;
> +
> +	/* Cache the data early to ensure is_visible() works */
> +	read_cdat_data(port);
> +	cxl_endpoint_parse_cdat(port);
> +
> +	get_device(&cxlmd->dev);
> +	rc = devm_add_action_or_reset(&port->dev, schedule_detach, &cxlmd->dev);

[Severity: Low]
Is this where the mismatch originates? The devres data queued for
schedule_detach is explicitly set to a struct device * here, confirming
what is passed when the callback executes.

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260923173342.5584-1-Benjamin.Cheatham@amd.com?part=5

^ permalink raw reply	[flat|nested] 28+ messages in thread

* Re: [PATCH 02/15] cxl/cache: Add cxl_cache driver
  2026-09-23 17:33 ` [PATCH 02/15] cxl/cache: Add cxl_cache driver Ben Cheatham
@ 2026-09-23 17:49   ` sashiko-bot
  0 siblings, 0 replies; 28+ messages in thread
From: sashiko-bot @ 2026-09-23 17:49 UTC (permalink / raw)
  To: Ben Cheatham; +Cc: linux-cxl

Thank you for your contribution! Sashiko AI review found 5 potential issue(s) to consider:
- [High] cxl/cache: kernel crash in cxl_clear_cache_enable() on non-PCI mock devices
- [High] cxl/core: race condition in RMW of PCI_DVSEC_CXL_CTRL without locking
- [High] cxl/cache: NULL pointer dereference in probe due to early clear in unregister
- [Medium] cxl/core: PCI config space corruption due to missing cxl_dvsec capability check
- [Low] cxl/cache: build link failure when built-in driver calls devm_cxl_add_cachedev
--

commit bd84836a0f10aed3ee9d665572f6ba9474a6ea11
Author: Ben Cheatham <Benjamin.Cheatham@amd.com>
cxl/cache: Add cxl_cache driver

This commit adds the cxl_cache driver for managing CXL.cache capable endpoints
and validating system configurations. It provides functionality to get cache
information from the CXL device DVSEC and enables or disables CXL.cache during
device probe.

> diff --git a/drivers/cxl/cache.c b/drivers/cxl/cache.c
> --- /dev/null
> +++ b/drivers/cxl/cache.c

[ ... ]

> +static int cxl_cache_probe(struct device *dev)
> +{
> +	struct cxl_cachedev *cxlcd = to_cxl_cachedev(dev);
> +	struct cxl_dev_state *cxlds = cxlcd->cxlds;
> +	int rc;
> +
> +	/* Disable CXL.cache until we can validate the device configuration */
> +	cxl_clear_cache_enable(cxlds);

[Severity: High]
Could cxl_clear_cache_enable() panic if a non-PCI device (such as a mock 
device used in testing) is bound to the driver? It appears we blindly cast 
to a struct pci_dev pointer inside cxl_set_protocol_enable() before any 
dev_is_pci() check occurs.

[Severity: High]
Additionally, could this probe function race with device unregistration? If a 
sysfs unbind triggers cxl_cachedev_unregister(), it sets cxlcd->cxlds to 
NULL before calling device_del(). If this probe runs concurrently, could it 
retrieve the NULL pointer and pass it to cxl_clear_cache_enable(), resulting 
in a NULL pointer dereference?

> +
> +	rc = cxl_accel_read_cache_info(cxlds);
> +	if (rc)
> +		return rc;
> +
> +	return 0;
> +}

[ ... ]

> diff --git a/drivers/cxl/core/pci.c b/drivers/cxl/core/pci.c
> --- a/drivers/cxl/core/pci.c
> +++ b/drivers/cxl/core/pci.c

[ ... ]

> -static int cxl_set_mem_enable(struct cxl_dev_state *cxlds, u16 val)
> +static int cxl_set_protocol_enable(struct cxl_dev_state *cxlds, u16 val,
> +				   u16 enable_bit)
>  {
>  	struct pci_dev *pdev = to_pci_dev(cxlds->dev);
>  	int d = cxlds->cxl_dvsec;
> @@ -191,9 +193,9 @@ static int cxl_set_mem_enable(struct cxl_dev_state *cxlds, u16 val)
>  	if (rc)
>  		return pcibios_err_to_errno(rc);
>  
> -	if ((ctrl & PCI_DVSEC_CXL_MEM_ENABLE) == val)
> +	if ((ctrl & enable_bit) == val)
>  		return 1;
> -	ctrl &= ~PCI_DVSEC_CXL_MEM_ENABLE;
> +	ctrl &= ~enable_bit;
>  	ctrl |= val;
>  
>  	rc = pci_write_config_word(pdev, d + PCI_DVSEC_CXL_CTRL, ctrl);

[Severity: High]
Is this read-modify-write cycle protected against concurrent modifications? If
the cxl_mem and cxl_cache drivers probe concurrently, could one thread
overwrite the modifications of the other in the shared PCI_DVSEC_CXL_CTRL
register, potentially dropping the MEM_ENABLE bit while setting CACHE_ENABLE?

[Severity: Medium]
What happens if an endpoint driver adds a cache device on hardware lacking a 
valid CXL DVSEC capability? If cxlds->cxl_dvsec is 0, this will calculate 
d + PCI_DVSEC_CXL_CTRL (which is 0x0C) and overwrite the standard PCI Cache
Line Size register. Does this need a cxlds->cxl_dvsec > 0 check?

[ ... ]

> +int cxl_accel_read_cache_info(struct cxl_dev_state *cxlds)
> +{
> +	struct cxl_cache_state *cstate = &cxlds->cstate;
> +	struct pci_dev *pdev = to_pci_dev(cxlds->dev);
> +	int dvsec = cxlds->cxl_dvsec;
> +	u16 cap, cap2;
> +	u32 unit;
> +	int rc;
> +
> +	if (!dev_is_pci(cxlds->dev))
> +		return -EINVAL;
> +
> +	rc = pci_read_config_word(pdev, dvsec + PCI_DVSEC_CXL_CAP, &cap);

[Severity: Medium]
Similarly, could this read garbage capability data from the standard PCI 
Class Code register (0x0A) if cxlds->cxl_dvsec is 0?

> +	if (rc)
> +		return pcibios_err_to_errno(rc);

[ ... ]

> diff --git a/include/cxl/cxl.h b/include/cxl/cxl.h
> --- a/include/cxl/cxl.h
> +++ b/include/cxl/cxl.h

[ ... ]

>  struct cxl_memdev *devm_cxl_probe_mem(struct cxl_dev_state *cxlds,
>  				      struct range *range);
> -
>  int cxl_set_capacity(struct cxl_dev_state *cxlds, u64 capacity);
> +
> +#if IS_ENABLED(CONFIG_CXL_CACHE)
> +struct cxl_cachedev *devm_cxl_add_cachedev(struct cxl_dev_state *cxlds);
> +#else
> +static inline struct cxl_cachedev *
> +devm_cxl_add_cachedev(struct cxl_dev_state *cxlds)
> +{ return ERR_PTR(-ENXIO); }
> +#endif /* CONFIG_CXL_CACHE */

[Severity: Low]
Can this cause a build link failure if CONFIG_CXL_CACHE is built as a module 
(=m) but the endpoint driver calling devm_cxl_add_cachedev() is built-in (=y)? 

IS_ENABLED(CONFIG_CXL_CACHE) evaluates to true in this case, providing the 
external declaration, but the symbol won't be available during vmlinux linking. 
Should this use IS_REACHABLE(CONFIG_CXL_CACHE) instead?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260923173342.5584-1-Benjamin.Cheatham@amd.com?part=2

^ permalink raw reply	[flat|nested] 28+ messages in thread

* Re: [PATCH 10/15] cxl/core: Add Cache ID allocation
  2026-09-23 17:33 ` [PATCH 10/15] cxl/core: Add Cache ID allocation Ben Cheatham
@ 2026-09-23 17:49   ` sashiko-bot
  0 siblings, 0 replies; 28+ messages in thread
From: sashiko-bot @ 2026-09-23 17:49 UTC (permalink / raw)
  To: Ben Cheatham; +Cc: linux-cxl

Thank you for your contribution! Sashiko AI review found 6 potential issue(s) to consider:
- [High] cxl: reference leak of CXL root device in cxl_cachedev_deprogram_cache_id()
- [High] cxl: cxl_cid_program_decoder() fails to clear HDM-D fields
- [High] cxl: __deprogram_cache_id() clobbers valid routing table entries of other devices
- [High] cxl: modifying decoder and routing registers while COMMIT bit is set
- [High] cxl: toggling COMMIT bit without polling STATUS_COMMITTED causes race condition
- [Medium] cxl: get_max_cid() improperly returns 0 for zero-capacity routing tables
--

commit 3ec914a55f349bd29cbb06f7c4d3468cbd56c388
Author: Ben Cheatham <Benjamin.Cheatham@amd.com>

    cxl/core: Add Cache ID allocation
    
    Add allocation and programming of CXL cache ids for struct
    cxl_cachedevs as part of cxl_cache::probe(). Programming only occurs
    when system firmware has not already programmed an id and the device is
    *not* using HDM-D flows.

> diff --git a/drivers/cxl/core/cache.c b/drivers/cxl/core/cache.c
> index be570d3dc6090..9c6c8ea713ef3 100644
> --- a/drivers/cxl/core/cache.c
> +++ b/drivers/cxl/core/cache.c
[ ... ]
> @@ -112,6 +116,260 @@ static int cache_idrt_entry_valid(struct cxl_port *port, int id)
[ ... ]
> +static int cxl_cid_commit_decoder(struct cxl_dport *dport)
> +{
> +	void __iomem *ciddc = dport->regs.ciddc;
> +	u32 cap, ctrl, status;
> +	u8 scale, base;
> +	int rc;
> +
> +	cap = readl(ciddc + CXL_CACHE_ID_DC_CAP_OFFSET);
> +	if (!FIELD_GET(CXL_CACHE_ID_DC_CAP_COMMIT_REQ, cap))
> +		return 0;
> +
> +	status = readl(ciddc + CXL_CACHE_ID_DC_STATUS_OFFSET);
> +	scale = FIELD_GET(CXL_CACHE_ID_DC_STATUS_TM_SCALE, status);
> +	base = FIELD_GET(CXL_CACHE_ID_DC_STATUS_TM_BASE, status);
> +
> +	ctrl = readl(ciddc + CXL_CACHE_ID_DC_CTRL_OFFSET);
> +	if (FIELD_GET(CXL_CACHE_ID_DC_CTRL_COMMIT, ctrl)) {
> +		ctrl &= ~CXL_CACHE_ID_DC_CTRL_COMMIT;
> +		writel(ctrl, ciddc + CXL_CACHE_ID_DC_CTRL_OFFSET);
> +	}
> +
> +	ctrl |= CXL_CACHE_ID_DC_CTRL_COMMIT;
> +	writel(ctrl, ciddc + CXL_CACHE_ID_DC_CTRL_OFFSET);

[Severity: High]
Does toggling the COMMIT bit from 1 to 0 to 1 without polling the status
register cause a race condition? If the hardware hasn't cleared the
STATUS_COMMITTED bit from the uncommit, the subsequent poll might immediately
succeed on the stale state before hardware processes the cycle.

> +
> +	rc = __cxl_cid_wait_commit(dport->dport_dev,
> +				   ciddc + CXL_CACHE_ID_DC_STATUS_OFFSET,
> +				   CXL_CACHE_ID_DC_STATUS_COMMITTED,
> +				   CXL_CACHE_ID_DC_STATUS_COMMIT_ERR, scale,
> +				   base);
[ ... ]
> +static int cxl_cid_program_decoder(struct cxl_dport *dport, int cid,
> +				   bool endpoint)
> +{
> +	void __iomem *ciddc = dport->regs.ciddc;
> +	u32 ctrl;
> +
> +	if (!ciddc)
> +		return -EINVAL;
> +
> +	ctrl = readl(ciddc + CXL_CACHE_ID_DC_CTRL_OFFSET);
> +
> +	/*
> +	 * The decoder may have been programmed before, so we zero out
> +	 * all the fields before writing to them
> +	 */
> +	ctrl &= ~(CXL_CACHE_ID_DC_CTRL_ASGN_ID | CXL_CACHE_ID_DC_CTRL_FWD_ID);

[Severity: High]
Does this actually zero out all the fields as the comment suggests? It looks
like CXL_CACHE_ID_DC_CTRL_HDMD_PRESENT and HDMD_ID are left intact, which
could result in hardware routing using a stale HDMD_ID instead of the
programmed LOCAL_ID if firmware previously set this up in HDM-D mode.

> +	if (endpoint) {
> +		ctrl |= CXL_CACHE_ID_DC_CTRL_ASGN_ID;
> +	} else {
> +		ctrl |= CXL_CACHE_ID_DC_CTRL_FWD_ID;
> +	}
> +
> +	FIELD_MODIFY(CXL_CACHE_ID_DC_CTRL_LOCAL_ID, &ctrl, cid);
> +	writel(ctrl, ciddc + CXL_CACHE_ID_DC_CTRL_OFFSET);

[Severity: High]
Could writing to the control register fail to take effect if the COMMIT bit
is still set (e.g. by firmware)? The register might ignore writes while
committed, which would cause the new configuration to be lost when
cxl_cid_commit_decoder() subsequently reads and commits the old value.

> +
> +	return cxl_cid_commit_decoder(dport);
> +}
> +
> +static int cxl_cid_commit_table(struct cxl_port *port)
> +{
> +	void __iomem *cidrt = port->regs.cidrt;
> +	u32 cap, ctrl, status;
> +	u8 scale, base;
> +	int rc;
> +
> +	cap = readl(cidrt + CXL_CACHE_ID_RT_CAP_OFFSET);
> +	if (!FIELD_GET(CXL_CACHE_ID_RT_CAP_COMMIT_REQ, cap))
> +		return 0;
> +
> +	status = readl(cidrt + CXL_CACHE_ID_RT_STATUS_OFFSET);
> +	scale = FIELD_GET(CXL_CACHE_ID_RT_STATUS_TM_SCALE, status);
> +	base = FIELD_GET(CXL_CACHE_ID_RT_STATUS_TM_BASE, status);
> +
> +	ctrl = readl(cidrt + CXL_CACHE_ID_RT_CTRL_OFFSET);
> +	if (FIELD_GET(CXL_CACHE_ID_RT_CTRL_COMMIT, ctrl)) {
> +		ctrl &= ~CXL_CACHE_ID_RT_CTRL_COMMIT;
> +		writel(ctrl, cidrt + CXL_CACHE_ID_RT_CTRL_OFFSET);
> +	}
> +
> +	ctrl |= CXL_CACHE_ID_RT_CTRL_COMMIT;
> +	writel(ctrl, cidrt + CXL_CACHE_ID_RT_CTRL_OFFSET);

[Severity: High]
Similar to cxl_cid_commit_decoder(), does this sequence risk a race condition
by clearing and setting the COMMIT bit without waiting for the status register
to clear?

[ ... ]
> +static int cxl_cid_program_table_entry(struct cxl_port *port, int cid,
> +				       unsigned int dport_id)
> +{
> +	void __iomem *cidrt = port->regs.cidrt;
> +	u8 target_cnt, portn;
> +	u16 target_n;
> +	u32 cap;
[ ... ]
> +	target_n = CXL_CACHE_ID_RT_TARGETN_VALID;
> +	target_n |= FIELD_PREP(CXL_CACHE_ID_RT_TARGETN_PORTN, dport_id);
> +	writew(target_n, cidrt + CXL_CACHE_ID_RT_TARGETN_OFFSET(cid));

[Severity: High]
Can writing the new target to the routing table fail if the table is still
committed? Since the table is uncommitted in cxl_cid_commit_table() after
this write, hardware might ignore the write here.

> +
> +	return cxl_cid_commit_table(port);
> +}
> +
[ ... ]
> +static int get_max_cid(struct cxl_port *endpoint)
> +{
> +	struct cxl_port *port = parent_port_of(endpoint);
> +	void __iomem *cidrt;
> +	u32 cap;
> +	u8 cnt;
> +
> +	if (!port)
> +		return -EINVAL;
> +
> +	while (!is_cxl_root(port) && !is_cxl_root(parent_port_of(port)))
> +		port = parent_port_of(port);
> +
> +	cidrt = port->regs.cidrt;
> +	if (!cidrt)
> +		return -EINVAL;
> +
> +	cap = readl(cidrt + CXL_CACHE_ID_RT_CAP_OFFSET);
> +	cnt = FIELD_GET(CXL_CACHE_ID_RT_CAP_TARGET_CNT, cap);
> +	if (cnt == 0)
> +		return 0;

[Severity: Medium]
Will returning 0 here for a zero-capacity routing table mislead the caller?
This seems to conflate a table with no capacity (0 entries) with a table
that has 1 entry (max ID 0), potentially allowing the IDA to allocate ID 0
and fail later. Should this return an error like -EINVAL instead?

[ ... ]
> @@ -196,4 +465,64 @@ int cxl_cachedev_validate_cache_id(struct cxl_cachedev *cxlcd)
>  	return 0;
>  }
>  EXPORT_SYMBOL_FOR_MODULES(cxl_cachedev_validate_cache_id, "cxl_cache");
> +
> +static void __deprogram_cache_id(struct cxl_cachedev *cxlcd,
> +				 struct cxl_port *stop)
> +{
> +	struct cxl_port *iter = cxlcd->endpoint->parent_dport->port;
> +
> +	/*
> +	 * Leave cache id decoders programmed; the table entry being invalidated
> +	 * should be enough
> +	 */
> +	while (!is_cxl_root(iter)) {
> +		cxl_cid_invalidate_table_entry(iter, cxlcd->cache_id);
> +
> +		if (iter == stop)
> +			return;

[Severity: High]
Does invalidating the entry on the stop port before checking the loop
condition clobber valid routing table entries belonging to other devices?
When cxl_cid_program_table_entry() fails and triggers rollback, running
the invalidation on the failing stop port might destroy a legitimate
configuration owned by another port.

> +
> +		iter = parent_port_of(iter);
> +	}
> +}
[ ... ]
> +void cxl_cachedev_deprogram_cache_id(struct cxl_cachedev *cxlcd)
> +{
> +	struct cxl_root *root = find_cxl_root(cxlcd->endpoint);
> +
> +	if (!root)
> +		return;
> +
> +	__deprogram_cache_id(cxlcd, &root->port);
> +}

[Severity: High]
Does this leak the reference to the CXL root device? find_cxl_root()
returns a reference-counted device, but it appears to be missing a call
to put_device() before returning, which would pin the root device
in memory and prevent proper teardown.

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260923173342.5584-1-Benjamin.Cheatham@amd.com?part=10

^ permalink raw reply	[flat|nested] 28+ messages in thread

* Re: [PATCH 09/15] cxl/core: Add Cache ID verification
  2026-09-23 17:33 ` [PATCH 09/15] cxl/core: Add Cache ID verification Ben Cheatham
@ 2026-09-23 17:49   ` sashiko-bot
  0 siblings, 0 replies; 28+ messages in thread
From: sashiko-bot @ 2026-09-23 17:49 UTC (permalink / raw)
  To: Ben Cheatham; +Cc: linux-cxl

Thank you for your contribution! Sashiko AI review found 5 potential issue(s) to consider:
- [High] cxl/cache: missing rcu_read_lock during xa_for_each iteration
- [High] cxl/cache: unconditional probe failure for Restricted CXL Host (RCH) topologies
- [High] cxl/cache: use-after-free of parent_dport during asynchronous teardown
- [Medium] cxl/core: missing validation of hdmd state against hardware cache ID
- [Low] cxl/core: returning negative error code for truthy boolean condition
--

commit fb180ed6317b83962ae34e684038bcea834c1f39
Author: Ben Cheatham <Benjamin.Cheatham@amd.com>

cxl/core: Add Cache ID verification
    
Check if system firmware has pre-programmed a CXL.cache device's cache
id as part of cxl_cache::probe(). A cache id is required when multiple
CXL.cache devices are present under a host bridge, so fail probe if the
id has not been programmed and another device is present.

Also add @hdmd to struct cxl_dev_state for endpoint drivers to indicate
whether their device is using HDM-D flows. This is required for
correctly validating cache id programming.

> diff --git a/drivers/cxl/cache.c b/drivers/cxl/cache.c
> index dfefd301696a2..6c098010149ae 100644
> --- a/drivers/cxl/cache.c
> +++ b/drivers/cxl/cache.c
> @@ -84,6 +84,90 @@ static int cxl_endpoint_cache_enabled(struct cxl_port *endpoint)
>  	return 0;
>  }
>  
> +static struct cxl_port *find_host_bridge(struct cxl_port *endpoint)
> +{
> +	struct cxl_port *parent = endpoint->parent_dport->port;

[Severity: High]
Can this result in a use-after-free during asynchronous teardown?

When a CXL host bridge is hot-removed, the upstream CXL port topology can be
torn down, and devres can free parent_dport, before the PCI device is
unbound.

During PCI unbind, the cachedev's devres will call deprogram_cache_id()
which dereferences cxlcd->endpoint->parent_dport->port here, accessing
freed memory.

> +	struct cxl_port *hb = endpoint;
> +
> +	if (is_cxl_root(parent))
> +		return NULL;

[Severity: High]
Will this unconditionally fail probes for Restricted CXL Host (RCH)
topologies?

In an RCH topology, the endpoint's parent is the root port, so
is_cxl_root(parent) will evaluate to true immediately, causing
find_host_bridge() to return NULL.

Later in program_cache_id(), this NULL return causes the entire cache
programming and probe sequence to fail with -ENODEV.

> +
> +	while (!is_cxl_root(parent)) {
> +		hb = parent;
> +		parent = hb->parent_dport->port;
> +	}
> +
> +	return hb;
> +}
> +
> +static int get_num_cachedevs_present(struct cxl_port *host_bridge, u32 *num)
> +{
> +	u32 num_cachedevs = 0;
> +	unsigned long index;
> +	struct cxl_ep *ep;
> +
> +	xa_for_each(&host_bridge->endpoints, index, ep) {
> +		if (is_cxl_cachedev(ep->ep))
> +			num_cachedevs++;
> +	}

[Severity: High]
Is an rcu_read_lock() or xa_lock() required here when iterating the
host_bridge->endpoints array?

While program_cache_id() holds the host bridge device lock (hb->dev),
concurrent endpoint probes in cxl_add_ep() insert into this array without
holding the higher-level hb->dev lock. 

Iterating xa_for_each() without RCU protection or the array's lock risks
data races and use-after-free, especially since cxl_ep_remove() can free the
endpoint structures concurrently.

[ ... ]
> diff --git a/drivers/cxl/core/cache.c b/drivers/cxl/core/cache.c
> index 17df0e603b68f..be570d3dc6090 100644
> --- a/drivers/cxl/core/cache.c
> +++ b/drivers/cxl/core/cache.c
> @@ -39,3 +39,161 @@ int cxl_dport_map_cache_id_dc(struct cxl_dport *dport)
>  	return cxl_map_component_regs(map, &dport->regs.component,
>  				      BIT(CXL_CM_CAP_CAP_ID_CACHE_ID_DC));
>  }
> +
> +static int cache_decoder_committed(struct cxl_dport *dport)
> +{
> +	u32 cap, stat;
> +
> +	cap = readl(dport->regs.ciddc + CXL_CACHE_ID_DC_CAP_OFFSET);
> +	if (!FIELD_GET(CXL_CACHE_ID_DC_CAP_COMMIT_REQ, cap))
> +		return -ENXIO;

[Severity: Low]
Is it safe to return a negative error code for a truthy boolean condition?

When the COMMIT_REQ bit is not set, this returns -ENXIO, which the caller
evaluates as truthy via negation (e.g., !cache_decoder_committed(dport)
evaluates to false since !(-ENXIO) is false).

While this happens to work in C, returning a negative integer to signify a
successful or truthy condition creates a silent refactoring hazard.

> +
> +	stat = readl(dport->regs.ciddc + CXL_CACHE_ID_DC_STATUS_OFFSET);
> +	return FIELD_GET(CXL_CACHE_ID_DC_STATUS_COMMITTED, stat);
> +}
> +
> +static int cache_decoder_get_id(struct cxl_dport *dport)
> +{
> +	u32 ctrl;
> +
> +	ctrl = readl(dport->regs.ciddc + CXL_CACHE_ID_DC_CTRL_OFFSET);
> +	if (FIELD_GET(CXL_CACHE_ID_DC_CTRL_HDMD_PRESENT, ctrl))
> +		 return FIELD_GET(CXL_CACHE_ID_DC_CTRL_HDMD_ID, ctrl);
> +
> +	return FIELD_GET(CXL_CACHE_ID_DC_CTRL_LOCAL_ID, ctrl);
> +}

[Severity: Medium]
Does this miss validating the hardware's HDM-D capability against the
endpoint driver's expectation?

The commit message notes that @hdmd is added to struct cxl_dev_state to
indicate whether the device uses HDM-D flows, and that this is required for
correctly validating cache id programming. 

However, cache_decoder_get_id() parses the ID purely based on the hardware
HDMD_PRESENT bit, and cxl_cachedev_validate_cache_id() never
cross-references this with the software cxlds->hdmd state. Firmware
misconfigurations where the cache ID flow type doesn't match the endpoint
driver's requirement will go undetected.

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260923173342.5584-1-Benjamin.Cheatham@amd.com?part=9

^ permalink raw reply	[flat|nested] 28+ messages in thread

* Re: [PATCH 12/15] cxl/cache: Add snoop filter creation and set up
  2026-09-23 17:33 ` [PATCH 12/15] cxl/cache: Add snoop filter creation and set up Ben Cheatham
@ 2026-09-23 17:50   ` sashiko-bot
  0 siblings, 0 replies; 28+ messages in thread
From: sashiko-bot @ 2026-09-23 17:50 UTC (permalink / raw)
  To: Ben Cheatham; +Cc: linux-cxl

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] cxl/cache: devres leak on dport hot-remove due to devm group mismatch
--

commit a51df73f15bcf70227ccbfa7fe083fcb3c6ed6e3
Author: Ben Cheatham <Benjamin.Cheatham@amd.com>

cxl/cache: Add snoop filter creation and set up

This commit adds the capability to track the system's CXL snoop filters and
probes them during dport addition. Snoop filter capacity allocation will be
handled in a later patch.

> diff --git a/drivers/cxl/core/cache.c b/drivers/cxl/core/cache.c
> --- a/drivers/cxl/core/cache.c
> +++ b/drivers/cxl/core/cache.c
[ ... ]
> +int cxl_dport_probe_snoop_filter(struct cxl_dport *dport)
> +{
[ ... ]
> +	rc = cxl_map_component_regs(&dport->reg_map, &dport->regs.component,
> +				    BIT(CXL_CM_CAP_CAP_ID_SNOOP));
[ ... ]
> diff --git a/drivers/cxl/port.c b/drivers/cxl/port.c
> --- a/drivers/cxl/port.c
> +++ b/drivers/cxl/port.c
> @@ -301,6 +301,16 @@ static struct cxl_dport *cxl_port_add_dport(struct cxl_port *port,
>  	/* New dport added, update the decoder targets */
>  	cxl_port_update_decoder_targets(port, dport);
>  
> +	/* 
> +	 * cxl_cachedevs won't probe if this fails, but it's not an error for
> +	 * cxl_memdevs
> +	 */
> +	rc = cxl_dport_probe_snoop_filter(dport);

[Severity: Medium]
Does this code leak devres memory and virtual address space on dport
hot-remove?

This function call happens outside of the dport devres group, which was
already closed when __devm_cxl_add_dport() returned (and the temporary
port_dr_group is removed earlier in this function).

When cxl_dport_probe_snoop_filter() is called, it triggers a devm
allocation:

cxl_dport_probe_snoop_filter()
  -> cxl_map_component_regs()
    -> devm_cxl_iomap_block()

Because there is no active devres group specifically scoped to the dport
at this point, will this devm allocation attach directly to the parent
port instead? If so, the allocation won't be freed when the dport is
hot-removed, leading to a resource leak.

> +	if (rc)
> +		dev_info(dport->dport_dev,
> +			 "Failed to find or create a CXL snoop filter: %d\n",
> +			 rc);
> +
>  	dev_dbg(&port->dev, "dport%d:%s added\n", dport->port_id,
>  		dev_name(dport_dev));
>

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260923173342.5584-1-Benjamin.Cheatham@amd.com?part=12

^ permalink raw reply	[flat|nested] 28+ messages in thread

* Re: [PATCH 11/15] cxl/core: Add support for HDM-D cache id programming
  2026-09-23 17:33 ` [PATCH 11/15] cxl/core: Add support for HDM-D cache id programming Ben Cheatham
@ 2026-09-23 17:51   ` sashiko-bot
  0 siblings, 0 replies; 28+ messages in thread
From: sashiko-bot @ 2026-09-23 17:51 UTC (permalink / raw)
  To: Ben Cheatham; +Cc: linux-cxl

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] cxl/cache: HDM-D capability limit ignored for intermediate switches
- [Medium] cxl/cache: off-by-one error in HDMDMax capability limit check
--

commit bdfe2b9c776e24b9a52c8e7a8594ae19a2e9c210
Author: Ben Cheatham <Benjamin.Cheatham@amd.com>

cxl/core: Add support for HDM-D cache id programming

This commit adds cache id allocation and programming support to the
pre-existing cache id programming routines for HDM-D devices.

> diff --git a/drivers/cxl/core/cache.c b/drivers/cxl/core/cache.c
> --- a/drivers/cxl/core/cache.c
> +++ b/drivers/cxl/core/cache.c
[ ... ]
> @@ -290,6 +297,16 @@ static int cxl_cid_program_table_entry(struct cxl_port *port, int cid,
>  		return -EINVAL;
>  	}
>  
> +	if (is_cxl_root(parent_port_of(port)) && hdmd) {
> +		hdmd_max = FIELD_GET(CXL_CACHE_ID_RT_CAP_HDMD_MAX, cap);
> +
> +		if (port->num_hdmd > hdmd_max) {

[Severity: Medium]
Will this strict inequality reject valid configurations when the hardware
capability register is 0-based?

In the CXL capability registers, CXL_CACHE_ID_RT_CAP_HDMD_MAX is only 4 bits.
To support 16 HDM-D devices (the maximum), it is highly likely to be 0-based.
The strict inequality port->num_hdmd > hdmd_max incorrectly evaluates to 1 > 0
for the first device, which could erroneously reject valid HDM-D devices during
initialization.

[Severity: High]
Does this capability limit check silently ignore intermediate switches?

Because cxl_cid_program_table_entry() explicitly restricts the limit validation
only to the top-level switch (direct child of the root), intermediate cascaded
switches completely skip validation. In a nested CXL topology, probing multiple
HDM-D endpoints connected to an intermediate switch that has a lower HDMDMax
capability limit than the top-level switch could silently over-program the
intermediate switch. This could potentially cause incorrect hardware routing,
dropped cache snoops, or fatal errors.

> +			dev_err(&port->dev,
> +				"Maximum number of devices using HDM-D reached\n");
> +			return -EINVAL;
> +		}
> +	}
[ ... ]
> @@ -385,6 +402,22 @@ static struct ida *find_cache_id_ida(struct cxl_port *port)
>  	return NULL;
>  }
>  
> +static void cxl_port_add_hdmd(struct cxl_port *endpoint, int val)
> +{
> +	struct cxl_port *parent = parent_port_of(endpoint);
> +	struct cxl_port *port = endpoint;
> +
> +	if (!parent || !is_cxl_cachedev(endpoint->uport_dev))
> +		return;
> +
> +	while (parent && !is_cxl_root(parent)) {
> +		port = parent;
> +		parent = parent_port_of(port);
> +	}
> +
> +	port->num_hdmd += val;

[Severity: High]
Can this loop fail to track limits for nested topologies by only aggregating
the count on the top-level switch?

By skipping intermediate switches, the HDM-D device is attributed entirely
to the top-level switch. This causes the capability limit checks to be
bypassed for intermediate cascaded switches.

> +}

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260923173342.5584-1-Benjamin.Cheatham@amd.com?part=11

^ permalink raw reply	[flat|nested] 28+ messages in thread

* Re: [PATCH 06/15] cxl/core: Update devm_cxl_add_endpoint() for cxl_cachedevs
  2026-09-23 17:33 ` [PATCH 06/15] cxl/core: Update devm_cxl_add_endpoint() for cxl_cachedevs Ben Cheatham
@ 2026-09-23 17:51   ` sashiko-bot
  0 siblings, 0 replies; 28+ messages in thread
From: sashiko-bot @ 2026-09-23 17:51 UTC (permalink / raw)
  To: Ben Cheatham; +Cc: linux-cxl

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Low] cxl: implicit cast from void* to struct bypassing container_of in schedule_detach
--

commit e4fe9279629e13947567c92e9190b5eeaa65a71a
Author: Ben Cheatham <Benjamin.Cheatham@amd.com>

cxl/core: Update devm_cxl_add_endpoint() for cxl_cachedevs

Update devm_cxl_add_endpoint() to allow for cxl_cachedevs as well as
cxl_memdevs. Add cxl_cachedevs to the port heirarchy.

> diff --git a/drivers/cxl/port.c b/drivers/cxl/port.c
> index 99cf77b6b6990..44d978d4205a2 100644
> --- a/drivers/cxl/port.c
> +++ b/drivers/cxl/port.c
> @@ -26,9 +27,13 @@
>   * PCIe topology.
>   */
>  
> -static void schedule_detach(void *cxlmd)
> +static void schedule_detach(void *ep_dev)
>  {
> -	schedule_cxl_memdev_detach(cxlmd);
> +	if (is_cxl_memdev(ep_dev))
> +		schedule_cxl_memdev_detach(ep_dev);
> +
> +	if (is_cxl_cachedev(ep_dev))
> +		schedule_cxl_cachedev_detach(ep_dev);

[Severity: Low]
In schedule_detach(), does this code implicitly cast the void pointer
representing a struct device directly to incompatible struct types?

Because ep_dev is a void pointer, passing it to schedule_cxl_memdev_detach()
and schedule_cxl_cachedev_detach() casts it to struct cxl_memdev * and
struct cxl_cachedev * without using the proper container_of() wrappers
like to_cxl_memdev() and to_cxl_cachedev().

While this works currently because struct device is the first member of these
structures, could this regression cause memory corruption if the struct
layouts ever change?

>  }

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260923173342.5584-1-Benjamin.Cheatham@amd.com?part=6

^ permalink raw reply	[flat|nested] 28+ messages in thread

* Re: [PATCH 14/15] iommu, cxl: Configure IOMMU for CXL.cache
  2026-09-23 17:33 ` [PATCH 14/15] iommu, cxl: Configure IOMMU for CXL.cache Ben Cheatham
@ 2026-09-23 17:57   ` sashiko-bot
  0 siblings, 0 replies; 28+ messages in thread
From: sashiko-bot @ 2026-09-23 17:57 UTC (permalink / raw)
  To: Ben Cheatham; +Cc: linux-cxl

Thank you for your contribution! Sashiko AI review found 5 potential issue(s) to consider:
- [High] iommu/amd: undefined behavior from using undocumented hardware control bits
- [High] iommu/amd: SEV-SNP hardware restriction bypass for CXL.cache devices
- [Medium] iommu/amd: cxl_cache_configure_iommu() fails due to incorrect -ENXIO return code
- [Medium] iommu/amd: CXL.cache devices miss identity domain fallback when IOMMU lacks ATS
- [Medium] cxl/cache: missing fallback to identity domain when ATS enablement fails
--

commit fde91018699b16a94e54a114be24afe059b28c07
Author: Ben Cheatham <Benjamin.Cheatham@amd.com>

iommu, cxl: Configure IOMMU for CXL.cache

Some IOMMU implementations require additional set up for enabling
ATS requests past enabling the base PCI ATS support. Create a callback
in the IOMMU core to be used by the CXL driver during device set up
that configures the IOMMU to handle ATS requests. Update the AMD IOMMU
driver with an implementation of this callback.

> diff --git a/drivers/cxl/core/cache.c b/drivers/cxl/core/cache.c
> index 429c997b2ada6..44323dfdb759d 100644
> --- a/drivers/cxl/core/cache.c
> +++ b/drivers/cxl/core/cache.c

[ ... ]

> @@ -806,3 +807,45 @@ void cxl_destroy_snoop_filters(void)
>  
>  	xa_destroy(&snoop_filters);
>  }
> +
> +/**
> + * cxl_cache_configure_iommu() - Configure a device's IOMMU for CXL.cache
> + * @cxlds: struct cxl_dev_state of a cxl_cachedev that has been through
> + *	   CXL.cache probe
> + *
> + * Fails if the underlying PCI device supports ATS and IOMMU can't be
> + * configured, or if the device doesn't support ATS and is not attached to an
> + * identity IOMMU domain.
> + */
> +int cxl_cache_configure_iommu(struct cxl_dev_state *cxlds)
> +{
> +	struct device *dev = cxlds->dev;
> +	struct iommu_domain *domain;
> +	int rc;
> +
> +	lockdep_assert_held(&dev->mutex);
> +
> +	if (!dev->iommu || !dev->iommu->iommu_dev)
> +		return 0;
> +
> +	if (!device_iommu_capable(dev, IOMMU_CAP_PCI_ATS_SUPPORTED)) {
> +		domain = iommu_get_domain_for_dev(dev);
> +		if (!domain || domain->type != IOMMU_DOMAIN_IDENTITY)
> +			return -EINVAL;
> +
> +		return 0;
> +	}
> +
> +	rc = iommu_enable_cxl_ats(cxlds->dev);
> +	if (rc == -EOPNOTSUPP) {
> +		dev_warn(cxlds->dev,
> +			"IOMMU doesn't support enabling CXL ATS requests; CXL.cache may not function properly.");
> +		rc = 0;
> +	} else if (rc) {
> +		dev_err(cxlds->dev, "Failed to enable CXL ATS requests: %d\n",
> +			rc);
> +	}
> +
> +	return rc;
> +}

[Severity: Medium]
If iommu_enable_cxl_ats() fails (for example, returning -EINVAL because ATS
is disabled at runtime), initialization unconditionally fails here. Since ATS
is not strictly required when operating in an identity domain, should this
fall back to checking if the device is already attached to an
IOMMU_DOMAIN_IDENTITY before aborting?

> diff --git a/drivers/iommu/amd/amd_iommu_types.h b/drivers/iommu/amd/amd_iommu_types.h
> index 3dbe20023456b..696d9879ce9b7 100644
> --- a/drivers/iommu/amd/amd_iommu_types.h
> +++ b/drivers/iommu/amd/amd_iommu_types.h

[ ... ]

> @@ -107,6 +107,7 @@
>  
>  
>  /* Extended Feature 2 Bits */
> +#define FEATURE_CXLMEMATTR	BIT_ULL(0) /* WARNING: This bit isn't in the spec as of 09/26 */

[Severity: High]
Is it safe to define and use an undocumented hardware control bit? The commit
explicitly states that FEATURE_CXLMEMATTR is guessed and not present in the
AMD IOMMU specification. Modifying undocumented hardware control bits based on
guesses could trigger undefined behavior, memory corruption, or system panics
on real silicon.

>  #define FEATURE_SEVSNPIO_SUP	BIT_ULL(1)
>  #define FEATURE_GCR3TRPMODE	BIT_ULL(3)

[ ... ]

> diff --git a/drivers/iommu/amd/init.c b/drivers/iommu/amd/init.c
> index 40726dfef2733..e11f98db826e0 100644
> --- a/drivers/iommu/amd/init.c
> +++ b/drivers/iommu/amd/init.c

[ ... ]

> @@ -1125,6 +1125,14 @@ static void iommu_enable_gt(struct amd_iommu *iommu)
>  		iommu_feature_enable(iommu, CONTROL_GCR3TRPMODE);
>  }
>  
> +static void iommu_enable_cxlmemattr(struct amd_iommu *iommu)
> +{
> +	if (!check_feature2(FEATURE_CXLMEMATTR) || !amd_iommu_iotlb_sup)
> +		return;
> +
> +	iommu_feature_enable(iommu, CONTROL_CXLMEMATTR_EN);
> +}

[Severity: High]
This code evaluates the guessed FEATURE_CXLMEMATTR bit and writes to the
hardware control register. Could this cause unpredictable hardware behavior
since the bit's functionality is not yet defined in the specification?

> +
>  /* sets a specific bit in the device table entry. */
>  static void set_dte_bit(struct dev_table_entry *dte, u8 bit)

[ ... ]

> diff --git a/drivers/iommu/amd/iommu.c b/drivers/iommu/amd/iommu.c
> index 4dc306a4b5c62..78d34b9074503 100644
> --- a/drivers/iommu/amd/iommu.c
> +++ b/drivers/iommu/amd/iommu.c

[ ... ]

> @@ -3190,6 +3201,14 @@ static int amd_iommu_def_domain_type(struct device *dev)
>  		return IOMMU_DOMAIN_IDENTITY;
>  	}
>  
> +	/*
> +	 * CXL.cache devices that have no ATS capability need a passthrough
> +	 * domain to function correctly.
> +	 */
> +	if (dev_is_pci(dev) && !pci_ats_supported(to_pci_dev(dev)) &&
> +	    cxl_cache_supported(to_pci_dev(dev)))
> +		return IOMMU_DOMAIN_IDENTITY;

[Severity: High]
Hardware explicitly prohibits identity mapping (DTE[Mode]=0) when SEV-SNP is
active. By unconditionally returning IOMMU_DOMAIN_IDENTITY here, does this
bypass the global encryption checks that prevent this state? Forcing this
prohibited hardware state could cause hardware faults, MCEs, or memory
encryption bypass on SEV-SNP enabled systems.

[Severity: Medium]
Does this properly handle the case where the endpoint supports ATS but the
IOMMU globally lacks ATS capability (!amd_iommu_iotlb_sup)?

If the device has ATS, it is assigned a DMA domain here. But later,
cxl_cache_configure_iommu() will realize the IOMMU cannot support ATS and
will require an identity domain, causing initialization to fail with -EINVAL.
Should this also return an identity domain if the IOMMU itself lacks ATS
support?

> +
>  	return 0;
>  }
>  
> @@ -3199,6 +3218,40 @@ static bool amd_iommu_enforce_cache_coherency(struct iommu_domain *domain)
>  	return true;
>  }
>  
> +static int amd_iommu_enable_cxl_ats(struct device *dev)
> +{
> +	struct iommu_dev_data *dev_data = dev_iommu_priv_get(dev);
> +	int ret = 0;
> +
> +	if (!dev_data || !amd_iommu_iotlb_sup)
> +		return -EINVAL;
> +
> +	if (!check_feature2(FEATURE_CXLMEMATTR))
> +		return -ENXIO;

[Severity: Medium]
The core API contract expects -EOPNOTSUPP if the IOMMU does not support the
required CXL features, but this returns -ENXIO.

Since the caller in cxl_cache_configure_iommu() handles -EOPNOTSUPP
gracefully, will returning -ENXIO here cause CXL.cache device probe to fail
unconditionally on older AMD hardware instead of gracefully falling back?

> +
> +	mutex_lock(&dev_data->mutex);
> +

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260923173342.5584-1-Benjamin.Cheatham@amd.com?part=14

^ permalink raw reply	[flat|nested] 28+ messages in thread

* Re: [PATCH 13/15] cxl/cache: Add snoop filter allocation
  2026-09-23 17:33 ` [PATCH 13/15] cxl/cache: Add snoop filter allocation Ben Cheatham
@ 2026-09-23 17:58   ` sashiko-bot
  0 siblings, 0 replies; 28+ messages in thread
From: sashiko-bot @ 2026-09-23 17:58 UTC (permalink / raw)
  To: Ben Cheatham; +Cc: linux-cxl

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] cxl: cxl_cachedev_find_snoop_gid() loop prematurely ignores root dports
- [Medium] cxl: memory bloat from XArray with kernel pointer as index
--

commit 0d6bed69d638badd54cfffcfb59d4229114673a5
Author: Ben Cheatham <Benjamin.Cheatham@amd.com>

cxl/cache: Add snoop filter allocation

This commit adds snoop filter capacity allocation for CXL.cache devices.
Devices are expected to allocate snoop filter capacity equal to the size of
the address range they will cache before using the protocol. It also allows
devices to strictly fail allocations if the filter is full.

> diff --git a/drivers/cxl/cache.c b/drivers/cxl/cache.c
> index 40d1e8330df78..b4574a76ac0b8 100644
> --- a/drivers/cxl/cache.c
> +++ b/drivers/cxl/cache.c
> @@ -179,6 +179,20 @@ struct cxl_cachedev *devm_cxl_add_cachedev(struct cxl_dev_state *cxlds)
>  }
>  EXPORT_SYMBOL_NS_GPL(devm_cxl_add_cachedev, "CXL");
>  
> +static int cxl_cachedev_find_snoop_gid(struct cxl_cachedev *cxlcd)
> +{
> +	struct cxl_dport *iter;
> +
> +	for (iter = cxlcd->endpoint->parent_dport;
> +	     iter && !is_cxl_root(iter->port);
> +	     iter = iter->port->parent_dport) {

[Severity: High]
Does this loop prematurely ignore snoop filters located on root ports?

The termination condition !is_cxl_root(iter->port) is evaluated before the
loop body. When iter reaches the cxl_dport of a host bridge (which is a
downstream port of the cxl_root port), iter->port points to the cxl_root
port. This makes is_cxl_root(iter->port) evaluate to true, causing the loop
to terminate immediately before checking iter->snoop.

Consequently, snoop filters located on host bridges are completely ignored.
If no intermediate switch has a snoop filter, will this cause
cxl_cache_probe() to unconditionally fail with -ENXIO for direct attached
devices?

> +		if (iter->snoop != CXL_SNOOP_FILTER_NO_GROUP_ID)
> +			return iter->snoop;
> +	}
> +
> +	return -ENXIO;
> +}

[ ... ]

> diff --git a/drivers/cxl/core/cache.c b/drivers/cxl/core/cache.c
> index f07c49cfacec2..429c997b2ada6 100644
> --- a/drivers/cxl/core/cache.c
> +++ b/drivers/cxl/core/cache.c

[ ... ]

> @@ -638,6 +671,94 @@ static struct cxl_snoop_filter *find_or_add_snoop_filter(u64 size, int id)
>  	return sf;
>  }
>  
> +static struct snoop_allocation *
> +snoop_alloc_capacity(struct cxl_snoop_filter *sf, struct cxl_cachedev *cxlcd,
> +		     u64 size)
> +{
> +	struct cxl_cache_state *cstate = &cxlcd->cxlds->cstate;
> +	struct snoop_allocation *alloc;
> +	int rc;
> +
> +	guard(mutex)(&sf->lock);
> +	if (sf->avail < size) {
> +		/*
> +		 * Either this device or a device already using the filter
> +		 * can't use an oversubscribed filter
> +		 */
> +		if (cstate->strict_snoop || sf->strict > 0)
> +			return ERR_PTR(-ENOSPC);
> +
> +		size = sf->avail;
> +	}
> +
> +	alloc = kmalloc_obj(*alloc);
> +	if (!alloc)
> +		return ERR_PTR(-ENOMEM);
> +
> +	*alloc = (struct snoop_allocation) {
> +		.cxlcd = cxlcd,
> +		.size = size,
> +		.strict = cstate->strict_snoop,
> +		.id = sf->id,
> +	};
> +
> +	rc = xa_insert(&sf->allocations, (unsigned long)alloc, alloc,
> +		       GFP_KERNEL);

[Severity: Medium]
Does using a kernel pointer as an XArray index cause memory bloat here?

On 64-bit architectures, kernel pointers have their high bits set. Using a
dense array structure like XArray for a highly sparse, massive index forces
it to allocate its maximum depth of internal xa_node radix tree levels
(typically 11 levels). This wastes thousands of bytes of memory per
allocation to track a single pointer.

Since the XArray is never used for index-based lookups (elements are only
erased by pointer during cleanup), would a standard struct list_head provide
the same functionality with better performance and no memory bloat?

> +	if (rc) {
> +		kfree(alloc);
> +		return ERR_PTR(rc);
> +	}
> +
> +	if (cstate->strict_snoop)
> +		sf->strict++;
> +
> +	sf->avail -= size;
> +	return alloc;
> +}

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260923173342.5584-1-Benjamin.Cheatham@amd.com?part=13

^ permalink raw reply	[flat|nested] 28+ messages in thread

end of thread, other threads:[~2026-09-23 17:58 UTC | newest]

Thread overview: 28+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-23 17:33 [PATCH 00/15] Add initial CXL.cache support Ben Cheatham
2026-09-23 17:33 ` [PATCH 01/15] cxl/core: Add CXL.cache device struct Ben Cheatham
2026-09-23 17:41   ` sashiko-bot
2026-09-23 17:33 ` [PATCH 02/15] cxl/cache: Add cxl_cache driver Ben Cheatham
2026-09-23 17:49   ` sashiko-bot
2026-09-23 17:33 ` [PATCH 03/15] cxl/core: Change cxl_ep_load() to use device pointer parameter Ben Cheatham
2026-09-23 17:33 ` [PATCH 04/15] cxl/core: Update devm_cxl_enumerate_ports() for cxl_cachedevs Ben Cheatham
2026-09-23 17:33 ` [PATCH 05/15] cxl/port: Split endpoint port probe on device type Ben Cheatham
2026-09-23 17:46   ` sashiko-bot
2026-09-23 17:33 ` [PATCH 06/15] cxl/core: Update devm_cxl_add_endpoint() for cxl_cachedevs Ben Cheatham
2026-09-23 17:51   ` sashiko-bot
2026-09-23 17:33 ` [PATCH 07/15] cxl/cache: Verify port hierarchy has CXL.cache enabled Ben Cheatham
2026-09-23 17:33 ` [PATCH 08/15] cxl/core, cache: Add Cache ID register probing and init Ben Cheatham
2026-09-23 17:46   ` sashiko-bot
2026-09-23 17:33 ` [PATCH 09/15] cxl/core: Add Cache ID verification Ben Cheatham
2026-09-23 17:49   ` sashiko-bot
2026-09-23 17:33 ` [PATCH 10/15] cxl/core: Add Cache ID allocation Ben Cheatham
2026-09-23 17:49   ` sashiko-bot
2026-09-23 17:33 ` [PATCH 11/15] cxl/core: Add support for HDM-D cache id programming Ben Cheatham
2026-09-23 17:51   ` sashiko-bot
2026-09-23 17:33 ` [PATCH 12/15] cxl/cache: Add snoop filter creation and set up Ben Cheatham
2026-09-23 17:50   ` sashiko-bot
2026-09-23 17:33 ` [PATCH 13/15] cxl/cache: Add snoop filter allocation Ben Cheatham
2026-09-23 17:58   ` sashiko-bot
2026-09-23 17:33 ` [PATCH 14/15] iommu, cxl: Configure IOMMU for CXL.cache Ben Cheatham
2026-09-23 17:57   ` sashiko-bot
2026-09-23 17:33 ` [PATCH 15/15] cxl/cache: Enable CXL.cache on successful probe Ben Cheatham
2026-09-23 17:35 ` [PATCH 00/15] Add initial CXL.cache support Cheatham, Benjamin

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox