Devicetree
 help / color / mirror / Atom feed
* [PATCH v4 0/2] pci: AMD: Add Versal2 CPM6 PCIe host controller support
@ 2026-08-08 10:52 Sai Krishna Musham
  2026-08-08 10:52 ` [PATCH v4 1/2] dt-bindings: PCI: amd-mdb: Add CPM6 support Sai Krishna Musham
  2026-08-08 10:52 ` [PATCH v4 2/2] PCI: amd-mdb: Add CPM6 host controller support Sai Krishna Musham
  0 siblings, 2 replies; 5+ messages in thread
From: Sai Krishna Musham @ 2026-08-08 10:52 UTC (permalink / raw)
  To: bhelgaas, lpieralisi, kw, mani, robh, krzk+dt, conor+dt, cassel
  Cc: linux-pci, devicetree, linux-kernel, michal.simek,
	bharat.kumar.gogada, thippeswamy.havalige, sai.krishna.musham,
	pranav.sanwal

Add support for the AMD Versal2 CPM6 PCIe host controller to the
AMD MDB PCIe driver.

Sai Krishna Musham (2):
  dt-bindings: PCI: amd-mdb: Add CPM6 support
  PCI: amd-mdb: Add CPM6 host controller support

 .../bindings/pci/amd,versal2-mdb-host.yaml    |  45 +-
 .../devicetree/bindings/pci/snps,dw-pcie.yaml |   2 +
 drivers/pci/controller/dwc/pcie-amd-mdb.c     | 424 +++++++++++++++---
 3 files changed, 405 insertions(+), 66 deletions(-)


base-commit: dc59e4fea9d83f03bad6bddf3fa2e52491777482
-- 
2.44.4


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH v4 1/2] dt-bindings: PCI: amd-mdb: Add CPM6 support
  2026-08-08 10:52 [PATCH v4 0/2] pci: AMD: Add Versal2 CPM6 PCIe host controller support Sai Krishna Musham
@ 2026-08-08 10:52 ` Sai Krishna Musham
  2026-08-08 10:57   ` sashiko-bot
  2026-08-08 10:52 ` [PATCH v4 2/2] PCI: amd-mdb: Add CPM6 host controller support Sai Krishna Musham
  1 sibling, 1 reply; 5+ messages in thread
From: Sai Krishna Musham @ 2026-08-08 10:52 UTC (permalink / raw)
  To: bhelgaas, lpieralisi, kw, mani, robh, krzk+dt, conor+dt, cassel
  Cc: linux-pci, devicetree, linux-kernel, michal.simek,
	bharat.kumar.gogada, thippeswamy.havalige, sai.krishna.musham,
	pranav.sanwal

The AMD CPM6 PCIe controller is based on the Synopsys DesignWare PCIe IP.
Add "intr" to snps,dw-pcie.yaml vendor-specific reg-names for the
per-controller interrupt register region used by CPM6.

Update amd,versal2-mdb-host.yaml with separate register definitions:
- MDB5: 4 regions (slcr, config, dbi, atu)
- CPM6: 5 regions (slcr, config, dbi, atu, intr)

Signed-off-by: Sai Krishna Musham <sai.krishna.musham@amd.com>
---
Changes in v4:
- None

Changes in v3:
- Update subject to match history.
- Move allOf to the end, after required block.
- Drop the CPM6 example.

Changes in v2:
- Update the CPM6 device tree binding and example.

v1 https://lore.kernel.org/all/20260402180006.486229-2-sai.krishna.musham@amd.com/
v2 https://lore.kernel.org/all/20260728202044.1785986-2-sai.krishna.musham@amd.com/
v3 https://lore.kernel.org/all/20260803144412.713639-2-sai.krishna.musham@amd.com/
---
 .../bindings/pci/amd,versal2-mdb-host.yaml    | 45 ++++++++++++++++---
 .../devicetree/bindings/pci/snps,dw-pcie.yaml |  2 +
 2 files changed, 42 insertions(+), 5 deletions(-)

diff --git a/Documentation/devicetree/bindings/pci/amd,versal2-mdb-host.yaml b/Documentation/devicetree/bindings/pci/amd,versal2-mdb-host.yaml
index 406c15e1dee1..672589066854 100644
--- a/Documentation/devicetree/bindings/pci/amd,versal2-mdb-host.yaml
+++ b/Documentation/devicetree/bindings/pci/amd,versal2-mdb-host.yaml
@@ -9,27 +9,30 @@ title: AMD Versal2 MDB(Multimedia DMA Bridge) Host Controller
 maintainers:
   - Thippeswamy Havalige <thippeswamy.havalige@amd.com>
 
-allOf:
-  - $ref: /schemas/pci/pci-host-bridge.yaml#
-  - $ref: /schemas/pci/snps,dw-pcie.yaml#
-
 properties:
   compatible:
-    const: amd,versal2-mdb-host
+    enum:
+      - amd,versal2-mdb-host
+      - amd,versal2-cpm6-host
+      - amd,versal2-cpm6-host1
 
   reg:
+    minItems: 4
     items:
       - description: MDB System Level Control and Status Register (SLCR) Base
       - description: configuration region
       - description: data bus interface
       - description: address translation unit register
+      - description: CPM6 PCIe error and event interrupt registers
 
   reg-names:
+    minItems: 4
     items:
       - const: slcr
       - const: config
       - const: dbi
       - const: atu
+      - const: intr
 
   ranges:
     maxItems: 2
@@ -92,6 +95,38 @@ required:
   - "#interrupt-cells"
   - interrupt-controller
 
+allOf:
+  - $ref: /schemas/pci/pci-host-bridge.yaml#
+  - $ref: /schemas/pci/snps,dw-pcie.yaml#
+  - if:
+      properties:
+        compatible:
+          contains:
+            const: amd,versal2-mdb-host
+    then:
+      properties:
+        reg:
+          minItems: 4
+          maxItems: 4
+        reg-names:
+          minItems: 4
+          maxItems: 4
+  - if:
+      properties:
+        compatible:
+          contains:
+            enum:
+              - amd,versal2-cpm6-host
+              - amd,versal2-cpm6-host1
+    then:
+      properties:
+        reg:
+          minItems: 5
+          maxItems: 5
+        reg-names:
+          minItems: 5
+          maxItems: 5
+
 unevaluatedProperties: false
 
 examples:
diff --git a/Documentation/devicetree/bindings/pci/snps,dw-pcie.yaml b/Documentation/devicetree/bindings/pci/snps,dw-pcie.yaml
index b3216141881c..21f86609ddb6 100644
--- a/Documentation/devicetree/bindings/pci/snps,dw-pcie.yaml
+++ b/Documentation/devicetree/bindings/pci/snps,dw-pcie.yaml
@@ -117,6 +117,8 @@ properties:
               enum: [ ecam ]
             - description: AMD MDB PCIe SLCR region
               const: slcr
+            - description: AMD CPM6 PCIe error and event interrupt registers
+              const: intr
     allOf:
       - contains:
           enum: [ dbi, ctrl ]
-- 
2.44.4


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* [PATCH v4 2/2] PCI: amd-mdb: Add CPM6 host controller support
  2026-08-08 10:52 [PATCH v4 0/2] pci: AMD: Add Versal2 CPM6 PCIe host controller support Sai Krishna Musham
  2026-08-08 10:52 ` [PATCH v4 1/2] dt-bindings: PCI: amd-mdb: Add CPM6 support Sai Krishna Musham
@ 2026-08-08 10:52 ` Sai Krishna Musham
  2026-08-08 11:06   ` sashiko-bot
  1 sibling, 1 reply; 5+ messages in thread
From: Sai Krishna Musham @ 2026-08-08 10:52 UTC (permalink / raw)
  To: bhelgaas, lpieralisi, kw, mani, robh, krzk+dt, conor+dt, cassel
  Cc: linux-pci, devicetree, linux-kernel, michal.simek,
	bharat.kumar.gogada, thippeswamy.havalige, sai.krishna.musham,
	pranav.sanwal

Extend the AMD MDB PCIe driver to support AMD Versal2 CPM6 variants
(amd,versal2-cpm6-host and amd,versal2-cpm6-host1). Like MDB5, CPM6 is
based on the Synopsys DesignWare PCIe controller and reuses the existing
DesignWare host support; only the interrupt architecture differs:
- The per-controller interrupt registers (MISC_EVENT, PCIE_ERR) are in a
  dedicated "intr" region instead of the shared SLCR
- The aggregation layer (MERGED, PS_MISC) is shared and sticky; it
  requires an explicit W1C-clear after servicing to prevent IRQ storms
- The completion timeout is reported via PCIE_ERR instead of MISC_EVENT

Add variant-specific register layout and implement aggregator clearing
that, after every interrupt pass, clears this host's contribution from
the MERGED register and the corresponding bit in the PS_MISC severity
register.

Signed-off-by: Sai Krishna Musham <sai.krishna.musham@amd.com>
---
Changes in v4:
- Use IRQF_SHARED for the CPM6 event IRQ since the misc/error GIC line
  is shared between both host controllers.
- Read PCIE_ERR before servicing MISC_EVENT to snapshot the error state
  before it is cleared.

Changes in v3:
- Add missing completion-timeout handling via PCIE_ERR.
- Add CPM6 INTx shared-enable refcount handling (mask/unmask via intx_refmask).
- Fix probe-failure cleanup to avoid double-free/use-after-free risk.
- Fix CPM6 slcr mapping to use non-exclusive ioremap for shared SLCR.
- Clear unhandled asserted bits in AMD_CPM6_PCIE_ERR_STATUS.

Changes in v2:
- Add CPM6 host controller interrupt support.

v1 https://lore.kernel.org/all/20260402180006.486229-3-sai.krishna.musham@amd.com/
v2 https://lore.kernel.org/all/20260728202044.1785986-3-sai.krishna.musham@amd.com/
v3 https://lore.kernel.org/all/20260803144412.713639-3-sai.krishna.musham@amd.com/
---
 drivers/pci/controller/dwc/pcie-amd-mdb.c | 424 ++++++++++++++++++----
 1 file changed, 363 insertions(+), 61 deletions(-)

diff --git a/drivers/pci/controller/dwc/pcie-amd-mdb.c b/drivers/pci/controller/dwc/pcie-amd-mdb.c
index dee295f281a3..3e8e02ab96bc 100644
--- a/drivers/pci/controller/dwc/pcie-amd-mdb.c
+++ b/drivers/pci/controller/dwc/pcie-amd-mdb.c
@@ -21,6 +21,23 @@
 #include "../../pci.h"
 #include "pcie-designware.h"
 
+/*
+ * On CPM6 the per-controller PCIe interrupt registers (MISC_EVENT and
+ * PCIE_ERR) live in a dedicated region ("intr"), separate from the CPM SLCR
+ * region ("slcr") that holds the MERGED and PS severity registers they feed
+ * into. Each has a sticky W1C STATUS, a read-only MASK, and write-1
+ * ENABLE/DISABLE register.
+ */
+#define AMD_CPM6_PCIE_ERR_STATUS		0x500
+#define AMD_CPM6_PCIE_ERR_MASK			0x504
+#define AMD_CPM6_PCIE_ERR_ENABLE		0x508
+#define AMD_CPM6_PCIE_ERR_DISABLE		0x50C
+
+#define AMD_CPM6_MISC_EVENT_STATUS		0x514
+#define AMD_CPM6_MISC_EVENT_MASK		0x518
+#define AMD_CPM6_MISC_EVENT_ENABLE		0x51C
+#define AMD_CPM6_MISC_EVENT_DISABLE		0x520
+
 #define AMD_MDB_TLP_IR_STATUS_MISC		0x4C0
 #define AMD_MDB_TLP_IR_MASK_MISC		0x4C4
 #define AMD_MDB_TLP_IR_ENABLE_MISC		0x4C8
@@ -30,7 +47,23 @@
 
 #define AMD_MDB_PCIE_INTR_INTX_ASSERT(x)	BIT((x) * 2)
 
-/* Interrupt registers definitions. */
+#define AMD_CPM6_MERGED_STATUS			0x648
+#define AMD_CPM6_MERGED_ENABLE			0x650
+
+/* MERGED input bits for the MISC_EVENT/PCIE_ERR sources this driver handles. */
+#define AMD_CPM6_MERGED_PCIE_ERR_HOST0		13
+#define AMD_CPM6_MERGED_MISC_EVENT_HOST0	14
+#define AMD_CPM6_MERGED_PCIE_ERR_HOST1		16
+#define AMD_CPM6_MERGED_MISC_EVENT_HOST1	17
+
+/*
+ * The PS_MISC severity register feeds the misc/OR GIC line. The MERGED
+ * aggregator appears as bit 21 within it.
+ */
+#define AMD_CPM6_PS_MISC_IR_STATUS		0x340
+#define AMD_CPM6_PS_IR_MERGED		BIT(21)
+
+/* MDB5 interrupt register definitions. */
 #define AMD_MDB_PCIE_INTR_CMPL_TIMEOUT		15
 #define AMD_MDB_PCIE_INTR_INTX			16
 #define AMD_MDB_PCIE_INTR_PM_PME_RCVD		24
@@ -39,6 +72,9 @@
 #define AMD_MDB_PCIE_INTR_NONFATAL		27
 #define AMD_MDB_PCIE_INTR_FATAL			28
 
+/* Completion timeout lives in PCIE_ERR_STATUS; give it a dedicated hwirq. */
+#define AMD_CPM6_PCIE_ERR_CMPL_RADM		20
+
 #define IMR(x) BIT(AMD_MDB_PCIE_INTR_ ##x)
 #define AMD_MDB_PCIE_IMR_ALL_MASK			\
 	(						\
@@ -51,24 +87,133 @@
 		AMD_MDB_TLP_PCIE_INTX_MASK		\
 	)
 
+/* CPM6 hwirq mapping (hwirq == MISC_EVENT status bit). */
+#define AMD_CPM6_PCIE_INTR_FATAL		17
+#define AMD_CPM6_PCIE_INTR_NONFATAL		18
+#define AMD_CPM6_PCIE_INTR_MISC_CORRECTABLE	19
+#define AMD_CPM6_PCIE_INTR_PME_TO_ACK_RCVD	20
+#define AMD_CPM6_PCIE_INTR_PM_PME_RCVD		21
+#define AMD_CPM6_PCIE_INTR_INTX			22
+
+/* Map the PCIE_ERR completion timeout onto an event-domain hwirq. */
+#define AMD_CPM6_PCIE_INTR_CMPL_TIMEOUT		15
+
+#define AMD_CPM6_MISC_EVENT_MASK_ALL					\
+	(								\
+		BIT(AMD_CPM6_PCIE_INTR_FATAL)			|	\
+		BIT(AMD_CPM6_PCIE_INTR_NONFATAL)		|	\
+		BIT(AMD_CPM6_PCIE_INTR_MISC_CORRECTABLE)	|	\
+		BIT(AMD_CPM6_PCIE_INTR_PME_TO_ACK_RCVD)		|	\
+		BIT(AMD_CPM6_PCIE_INTR_PM_PME_RCVD)		|	\
+		BIT(AMD_CPM6_PCIE_INTR_INTX)				\
+	)
+
+/* Sources handled in the PCIE_ERR register. */
+#define AMD_CPM6_PCIE_ERR_MASK_ALL	BIT(AMD_CPM6_PCIE_ERR_CMPL_RADM)
+
+enum amd_mdb_pcie_version {
+	MDB5,
+	CPM6,
+	CPM6_HOST1,
+};
+
+struct amd_mdb_intr_cause {
+	const char	*sym;
+	const char	*str;
+};
+
+struct amd_mdb_pcie_variant {
+	enum	amd_mdb_pcie_version version;
+	u32	misc_status_reg;
+	u32	misc_mask_reg;
+	u32	misc_enable_reg;
+	u32	misc_disable_reg;
+	u32	misc_mask_all;
+	u32	intx_hwirq;
+	u32	intx_mask;
+};
+
 /**
  * struct amd_mdb_pcie - PCIe port information
  * @pci: DesignWare PCIe controller structure
  * @slcr: MDB System Level Control and Status Register (SLCR) base
+ * @intr_base: Per-controller interrupt register base. On CPM6 this maps the
+ *             "intr" region holding the MISC_EVENT/PCIE_ERR registers; on MDB5
+ *             the interrupt registers live in the SLCR block, so it aliases
+ *             @slcr.
+ * @variant: Interrupt layout data for the matched platform compatible
  * @intx_domain: INTx IRQ domain pointer
  * @mdb_domain: MDB IRQ domain pointer
  * @perst_gpio: GPIO descriptor for PERST# signal handling
  * @intx_irq: INTx IRQ interrupt number
+ * @intx_refmask: CPM6 mask of unmasked INTx lines; gates the shared aggregate
  */
 struct amd_mdb_pcie {
 	struct dw_pcie			pci;
 	void __iomem			*slcr;
+	void __iomem			*intr_base;
+	const struct amd_mdb_pcie_variant	*variant;
 	struct irq_domain		*intx_domain;
 	struct irq_domain		*mdb_domain;
 	struct gpio_desc		*perst_gpio;
 	int				intx_irq;
+	u32				intx_refmask;
+};
+
+#define _IC(x, s)[AMD_MDB_PCIE_INTR_ ## x] = { __stringify(x), s }
+
+static const struct amd_mdb_intr_cause mdb5_intr_cause[32] = {
+	_IC(CMPL_TIMEOUT,	"Completion timeout"),
+	_IC(PM_PME_RCVD,	"PM_PME message received"),
+	_IC(PME_TO_ACK_RCVD,	"PME_TO_ACK message received"),
+	_IC(MISC_CORRECTABLE,	"Correctable error message"),
+	_IC(NONFATAL,		"Non fatal error message"),
+	_IC(FATAL,		"Fatal error message"),
+};
+
+#define _IC6(x, s)[AMD_CPM6_PCIE_INTR_ ## x] = { __stringify(x), s }
+
+static const struct amd_mdb_intr_cause cpm6_intr_cause[32] = {
+	_IC6(CMPL_TIMEOUT,	"Completion timeout"),
+	_IC6(PM_PME_RCVD,	"PM_PME message received"),
+	_IC6(PME_TO_ACK_RCVD,	"PME_TO_ACK message received"),
+	_IC6(MISC_CORRECTABLE,	"Correctable error message"),
+	_IC6(NONFATAL,		"Non fatal error message"),
+	_IC6(FATAL,		"Fatal error message"),
 };
 
+/* Both cause tables are indexed by hwirq, so they must be the same size. */
+static_assert(ARRAY_SIZE(mdb5_intr_cause) == ARRAY_SIZE(cpm6_intr_cause));
+
+static u32 amd_mdb_pcie_merged_host_mask(struct amd_mdb_pcie *pcie)
+{
+	return pcie->variant->version == CPM6 ?
+	       BIT(AMD_CPM6_MERGED_MISC_EVENT_HOST0) |
+	       BIT(AMD_CPM6_MERGED_PCIE_ERR_HOST0) :
+	       BIT(AMD_CPM6_MERGED_MISC_EVENT_HOST1) |
+	       BIT(AMD_CPM6_MERGED_PCIE_ERR_HOST1);
+}
+
+static void amd_mdb_pcie_clear_aggregators(struct amd_mdb_pcie *pcie)
+{
+	if (pcie->variant->version == MDB5)
+		return;
+
+	/*
+	 * Clear this host's serviced contributions (MISC_EVENT and PCIE_ERR)
+	 * from MERGED.
+	 */
+	writel_relaxed(amd_mdb_pcie_merged_host_mask(pcie),
+		       pcie->slcr + AMD_CPM6_MERGED_STATUS);
+
+	/*
+	 * Clear MERGED in the PS_MISC severity register so the misc GIC line
+	 * de-asserts.
+	 */
+	writel_relaxed(AMD_CPM6_PS_IR_MERGED,
+		       pcie->slcr + AMD_CPM6_PS_MISC_IR_STATUS);
+}
+
 static const struct dw_pcie_host_ops amd_mdb_pcie_host_ops = {
 };
 
@@ -81,14 +226,21 @@ static void amd_mdb_intx_irq_mask(struct irq_data *data)
 	u32 val;
 
 	raw_spin_lock_irqsave(&port->lock, flags);
-	val = FIELD_PREP(AMD_MDB_TLP_PCIE_INTX_MASK,
-			 AMD_MDB_PCIE_INTR_INTX_ASSERT(data->hwirq));
-
-	/*
-	 * Writing '1' to a bit in AMD_MDB_TLP_IR_DISABLE_MISC disables that
-	 * interrupt, writing '0' has no effect.
-	 */
-	writel_relaxed(val, pcie->slcr + AMD_MDB_TLP_IR_DISABLE_MISC);
+	if (pcie->variant->version == MDB5) {
+		val = FIELD_PREP(AMD_MDB_TLP_PCIE_INTX_MASK,
+				 AMD_MDB_PCIE_INTR_INTX_ASSERT(data->hwirq));
+		/*
+		 * Writing '1' to a bit in AMD_MDB_TLP_IR_DISABLE_MISC disables
+		 * that interrupt, writing '0' has no effect.
+		 */
+		writel_relaxed(val, pcie->intr_base + pcie->variant->misc_disable_reg);
+	} else {
+		/* CPM6 shares one INTx enable; drop it on the last mask. */
+		pcie->intx_refmask &= ~BIT(data->hwirq);
+		if (!pcie->intx_refmask)
+			writel_relaxed(pcie->variant->intx_mask,
+				       pcie->intr_base + pcie->variant->misc_disable_reg);
+	}
 	raw_spin_unlock_irqrestore(&port->lock, flags);
 }
 
@@ -101,14 +253,21 @@ static void amd_mdb_intx_irq_unmask(struct irq_data *data)
 	u32 val;
 
 	raw_spin_lock_irqsave(&port->lock, flags);
-	val = FIELD_PREP(AMD_MDB_TLP_PCIE_INTX_MASK,
-			 AMD_MDB_PCIE_INTR_INTX_ASSERT(data->hwirq));
-
-	/*
-	 * Writing '1' to a bit in AMD_MDB_TLP_IR_ENABLE_MISC enables that
-	 * interrupt, writing '0' has no effect.
-	 */
-	writel_relaxed(val, pcie->slcr + AMD_MDB_TLP_IR_ENABLE_MISC);
+	if (pcie->variant->version == MDB5) {
+		val = FIELD_PREP(AMD_MDB_TLP_PCIE_INTX_MASK,
+				 AMD_MDB_PCIE_INTR_INTX_ASSERT(data->hwirq));
+		/*
+		 * Writing '1' to a bit in AMD_MDB_TLP_IR_ENABLE_MISC enables
+		 * that interrupt, writing '0' has no effect.
+		 */
+		writel_relaxed(val, pcie->intr_base + pcie->variant->misc_enable_reg);
+	} else {
+		/* CPM6 shares one INTx enable; raise it on the first unmask. */
+		if (!pcie->intx_refmask)
+			writel_relaxed(pcie->variant->intx_mask,
+				       pcie->intr_base + pcie->variant->misc_enable_reg);
+		pcie->intx_refmask |= BIT(data->hwirq);
+	}
 	raw_spin_unlock_irqrestore(&port->lock, flags);
 }
 
@@ -148,42 +307,40 @@ static irqreturn_t dw_pcie_rp_intx(int irq, void *args)
 	unsigned long val;
 	int i, int_status;
 
-	val = readl_relaxed(pcie->slcr + AMD_MDB_TLP_IR_STATUS_MISC);
-	int_status = FIELD_GET(AMD_MDB_TLP_PCIE_INTX_MASK, val);
+	val = readl_relaxed(pcie->intr_base + pcie->variant->misc_status_reg);
 
-	for (i = 0; i < PCI_NUM_INTX; i++) {
-		if (int_status & AMD_MDB_PCIE_INTR_INTX_ASSERT(i))
+	if (pcie->variant->version == MDB5) {
+		int_status = FIELD_GET(AMD_MDB_TLP_PCIE_INTX_MASK, val);
+		for (i = 0; i < PCI_NUM_INTX; i++) {
+			if (int_status & AMD_MDB_PCIE_INTR_INTX_ASSERT(i))
+				generic_handle_domain_irq(pcie->intx_domain, i);
+		}
+	} else {
+		/* CPM6 exposes only an aggregate INTx indication */
+		if (!(val & pcie->variant->intx_mask))
+			return IRQ_NONE;
+		for (i = 0; i < PCI_NUM_INTX; i++)
 			generic_handle_domain_irq(pcie->intx_domain, i);
 	}
 
 	return IRQ_HANDLED;
 }
 
-#define _IC(x, s)[AMD_MDB_PCIE_INTR_ ## x] = { __stringify(x), s }
-
-static const struct {
-	const char	*sym;
-	const char	*str;
-} intr_cause[32] = {
-	_IC(CMPL_TIMEOUT,	"Completion timeout"),
-	_IC(PM_PME_RCVD,	"PM_PME message received"),
-	_IC(PME_TO_ACK_RCVD,	"PME_TO_ACK message received"),
-	_IC(MISC_CORRECTABLE,	"Correctable error message"),
-	_IC(NONFATAL,		"Non fatal error message"),
-	_IC(FATAL,		"Fatal error message"),
-};
-
 static void amd_mdb_event_irq_mask(struct irq_data *d)
 {
 	struct amd_mdb_pcie *pcie = irq_data_get_irq_chip_data(d);
 	struct dw_pcie *pci = &pcie->pci;
 	struct dw_pcie_rp *port = &pci->pp;
 	unsigned long flags;
-	u32 val;
 
 	raw_spin_lock_irqsave(&port->lock, flags);
-	val = BIT(d->hwirq);
-	writel_relaxed(val, pcie->slcr + AMD_MDB_TLP_IR_DISABLE_MISC);
+	if (pcie->variant->version != MDB5 &&
+	    d->hwirq == AMD_CPM6_PCIE_INTR_CMPL_TIMEOUT)
+		writel_relaxed(AMD_CPM6_PCIE_ERR_MASK_ALL,
+			       pcie->intr_base + AMD_CPM6_PCIE_ERR_DISABLE);
+	else
+		writel_relaxed(BIT(d->hwirq),
+			       pcie->intr_base + pcie->variant->misc_disable_reg);
 	raw_spin_unlock_irqrestore(&port->lock, flags);
 }
 
@@ -193,11 +350,15 @@ static void amd_mdb_event_irq_unmask(struct irq_data *d)
 	struct dw_pcie *pci = &pcie->pci;
 	struct dw_pcie_rp *port = &pci->pp;
 	unsigned long flags;
-	u32 val;
 
 	raw_spin_lock_irqsave(&port->lock, flags);
-	val = BIT(d->hwirq);
-	writel_relaxed(val, pcie->slcr + AMD_MDB_TLP_IR_ENABLE_MISC);
+	if (pcie->variant->version != MDB5 &&
+	    d->hwirq == AMD_CPM6_PCIE_INTR_CMPL_TIMEOUT)
+		writel_relaxed(AMD_CPM6_PCIE_ERR_MASK_ALL,
+			       pcie->intr_base + AMD_CPM6_PCIE_ERR_ENABLE);
+	else
+		writel_relaxed(BIT(d->hwirq),
+			       pcie->intr_base + pcie->variant->misc_enable_reg);
 	raw_spin_unlock_irqrestore(&port->lock, flags);
 }
 
@@ -226,13 +387,44 @@ static irqreturn_t amd_mdb_pcie_event(int irq, void *args)
 {
 	struct amd_mdb_pcie *pcie = args;
 	unsigned long val;
+	u32 ev_raw, err;
 	int i;
 
-	val = readl_relaxed(pcie->slcr + AMD_MDB_TLP_IR_STATUS_MISC);
-	val &= ~readl_relaxed(pcie->slcr + AMD_MDB_TLP_IR_MASK_MISC);
+	ev_raw = readl_relaxed(pcie->intr_base + pcie->variant->misc_status_reg);
+	val = ev_raw;
+	val &= ~readl_relaxed(pcie->intr_base + pcie->variant->misc_mask_reg);
+
+	if (pcie->variant->version == MDB5) {
+		for_each_set_bit(i, &val, 32)
+			generic_handle_domain_irq(pcie->mdb_domain, i);
+		writel_relaxed(val, pcie->intr_base + pcie->variant->misc_status_reg);
+		return IRQ_HANDLED;
+	}
+
+	err = readl_relaxed(pcie->intr_base + AMD_CPM6_PCIE_ERR_STATUS);
+
+	val &= pcie->variant->misc_mask_all;
+
 	for_each_set_bit(i, &val, 32)
 		generic_handle_domain_irq(pcie->mdb_domain, i);
-	writel_relaxed(val, pcie->slcr + AMD_MDB_TLP_IR_STATUS_MISC);
+
+	/* Clear handled + any unhandled sticky bits to avoid IRQ storms. */
+	writel_relaxed(ev_raw, pcie->intr_base + pcie->variant->misc_status_reg);
+
+	/* On CPM6 completion timeout is reported via PCIE_ERR. */
+	if (err) {
+		u32 pending = err & ~readl_relaxed(pcie->intr_base +
+						   AMD_CPM6_PCIE_ERR_MASK);
+
+		if (pending & AMD_CPM6_PCIE_ERR_MASK_ALL)
+			generic_handle_domain_irq(pcie->mdb_domain,
+						  AMD_CPM6_PCIE_INTR_CMPL_TIMEOUT);
+		/* Clear every asserted bit so the leaf and MERGED de-assert. */
+		writel_relaxed(err, pcie->intr_base + AMD_CPM6_PCIE_ERR_STATUS);
+	}
+
+	/* Sticky aggregation bits; clear each pass or the IRQ re-fires */
+	amd_mdb_pcie_clear_aggregators(pcie);
 
 	return IRQ_HANDLED;
 }
@@ -250,24 +442,47 @@ static void amd_mdb_pcie_free_irq_domains(struct amd_mdb_pcie *pcie)
 	}
 }
 
-static int amd_mdb_pcie_init_port(struct amd_mdb_pcie *pcie)
+static void amd_mdb_pcie_init_port(struct amd_mdb_pcie *pcie)
 {
-	unsigned long val;
+	u32 misc_mask_all;
+	u32 val;
+
+	misc_mask_all = pcie->variant->misc_mask_all;
 
 	/* Disable all TLP interrupts. */
-	writel_relaxed(AMD_MDB_PCIE_IMR_ALL_MASK,
-		       pcie->slcr + AMD_MDB_TLP_IR_DISABLE_MISC);
+	writel_relaxed(misc_mask_all,
+		       pcie->intr_base + pcie->variant->misc_disable_reg);
+
+	if (pcie->variant->version != MDB5)
+		writel_relaxed(AMD_CPM6_PCIE_ERR_MASK_ALL,
+			       pcie->intr_base + AMD_CPM6_PCIE_ERR_DISABLE);
 
 	/* Clear pending TLP interrupts. */
-	val = readl_relaxed(pcie->slcr + AMD_MDB_TLP_IR_STATUS_MISC);
-	val &= AMD_MDB_PCIE_IMR_ALL_MASK;
-	writel_relaxed(val, pcie->slcr + AMD_MDB_TLP_IR_STATUS_MISC);
+	val = readl_relaxed(pcie->intr_base + pcie->variant->misc_status_reg) &
+	      misc_mask_all;
+	writel_relaxed(val, pcie->intr_base + pcie->variant->misc_status_reg);
+
+	if (pcie->variant->version != MDB5) {
+		val = readl_relaxed(pcie->intr_base + AMD_CPM6_PCIE_ERR_STATUS) &
+		      AMD_CPM6_PCIE_ERR_MASK_ALL;
+		writel_relaxed(val, pcie->intr_base + AMD_CPM6_PCIE_ERR_STATUS);
+	}
 
 	/* Enable all TLP interrupts. */
-	writel_relaxed(AMD_MDB_PCIE_IMR_ALL_MASK,
-		       pcie->slcr + AMD_MDB_TLP_IR_ENABLE_MISC);
-
-	return 0;
+	writel_relaxed(misc_mask_all,
+		       pcie->intr_base + pcie->variant->misc_enable_reg);
+
+	if (pcie->variant->version != MDB5) {
+		writel_relaxed(AMD_CPM6_PCIE_ERR_MASK_ALL,
+			       pcie->intr_base + AMD_CPM6_PCIE_ERR_ENABLE);
+
+		/*
+		 * Unmask this host's MISC_EVENT and PCIE_ERR inputs in the
+		 * shared MERGED aggregator so they reach the GIC.
+		 */
+		writel_relaxed(amd_mdb_pcie_merged_host_mask(pcie),
+			       pcie->slcr + AMD_CPM6_MERGED_ENABLE);
+	}
 }
 
 /**
@@ -327,10 +542,13 @@ static int amd_mdb_pcie_init_irq_domains(struct amd_mdb_pcie *pcie,
 static irqreturn_t amd_mdb_pcie_intr_handler(int irq, void *args)
 {
 	struct amd_mdb_pcie *pcie = args;
+	const struct amd_mdb_intr_cause *intr_cause;
 	struct device *dev;
 	struct irq_data *d;
 
 	dev = pcie->pci.dev;
+	intr_cause = pcie->variant->version == MDB5 ?
+		     mdb5_intr_cause : cpm6_intr_cause;
 
 	/*
 	 * In the future, error reporting will be hooked to the AER subsystem.
@@ -351,15 +569,20 @@ static int amd_mdb_setup_irq(struct amd_mdb_pcie *pcie,
 	struct dw_pcie *pci = &pcie->pci;
 	struct dw_pcie_rp *pp = &pci->pp;
 	struct device *dev = &pdev->dev;
+	const struct amd_mdb_intr_cause *intr_cause;
+	unsigned long event_flags = IRQF_NO_THREAD;
 	int i, irq, err;
 
+	intr_cause = pcie->variant->version == MDB5 ?
+		     mdb5_intr_cause : cpm6_intr_cause;
+
 	amd_mdb_pcie_init_port(pcie);
 
 	pp->irq = platform_get_irq(pdev, 0);
 	if (pp->irq < 0)
 		return pp->irq;
 
-	for (i = 0; i < ARRAY_SIZE(intr_cause); i++) {
+	for (i = 0; i < ARRAY_SIZE(mdb5_intr_cause); i++) {
 		if (!intr_cause[i].str)
 			continue;
 
@@ -379,7 +602,7 @@ static int amd_mdb_setup_irq(struct amd_mdb_pcie *pcie,
 	}
 
 	pcie->intx_irq = irq_create_mapping(pcie->mdb_domain,
-					    AMD_MDB_PCIE_INTR_INTX);
+				    pcie->variant->intx_hwirq);
 	if (!pcie->intx_irq) {
 		dev_err(dev, "Failed to map INTx interrupt\n");
 		return -ENXIO;
@@ -393,8 +616,15 @@ static int amd_mdb_setup_irq(struct amd_mdb_pcie *pcie,
 		return err;
 	}
 
+	/*
+	 * On CPM6 the misc GIC line is shared between both host controllers,
+	 * so the event IRQ must allow sharing.
+	 */
+	if (pcie->variant->version != MDB5)
+		event_flags |= IRQF_SHARED;
+
 	/* Plug the main event handler. */
-	err = devm_request_irq(dev, pp->irq, amd_mdb_pcie_event, IRQF_NO_THREAD,
+	err = devm_request_irq(dev, pp->irq, amd_mdb_pcie_event, event_flags,
 			       "amd_mdb pcie_irq", pcie);
 	if (err) {
 		dev_err(dev, "Failed to request event IRQ %d, err=%d\n",
@@ -435,9 +665,36 @@ static int amd_mdb_add_pcie_port(struct amd_mdb_pcie *pcie,
 	struct device *dev = &pdev->dev;
 	int err;
 
-	pcie->slcr = devm_platform_ioremap_resource_byname(pdev, "slcr");
-	if (IS_ERR(pcie->slcr))
-		return PTR_ERR(pcie->slcr);
+	if (pcie->variant->version == MDB5) {
+		/*
+		 * On MDB5 all interrupt registers live in the SLCR block, so
+		 * the interrupt-register base simply aliases @slcr.
+		 */
+		pcie->slcr = devm_platform_ioremap_resource_byname(pdev, "slcr");
+		if (IS_ERR(pcie->slcr))
+			return PTR_ERR(pcie->slcr);
+		pcie->intr_base = pcie->slcr;
+	} else {
+		struct resource *res;
+
+		/*
+		 * CPM6 moves the per-controller MISC_EVENT/PCIE_ERR registers
+		 * into a separate "intr" region. The SLCR block, which holds
+		 * the shared MERGED/PS_MISC aggregators, is shared by both CPM6
+		 * host controllers, so map it without requesting exclusive
+		 * ownership; otherwise the second controller fails to probe.
+		 */
+		res = platform_get_resource_byname(pdev, IORESOURCE_MEM, "slcr");
+		if (!res)
+			return -EINVAL;
+		pcie->slcr = devm_ioremap(dev, res->start, resource_size(res));
+		if (!pcie->slcr)
+			return -ENOMEM;
+
+		pcie->intr_base = devm_platform_ioremap_resource_byname(pdev, "intr");
+		if (IS_ERR(pcie->intr_base))
+			return PTR_ERR(pcie->intr_base);
+	}
 
 	err = amd_mdb_pcie_init_irq_domains(pcie, pdev);
 	if (err)
@@ -483,6 +740,9 @@ static int amd_mdb_pcie_probe(struct platform_device *pdev)
 
 	pci = &pcie->pci;
 	pci->dev = dev;
+	pcie->variant = of_device_get_match_data(dev);
+	if (!pcie->variant)
+		return -EINVAL;
 
 	platform_set_drvdata(pdev, pcie);
 
@@ -514,9 +774,51 @@ static void amd_mdb_pcie_shutdown(struct platform_device *pdev)
 	gpiod_set_value_cansleep(pcie->perst_gpio, 1);
 }
 
+static const struct amd_mdb_pcie_variant cpm6_host = {
+	.version = CPM6,
+	.misc_status_reg = AMD_CPM6_MISC_EVENT_STATUS,
+	.misc_mask_reg = AMD_CPM6_MISC_EVENT_MASK,
+	.misc_enable_reg = AMD_CPM6_MISC_EVENT_ENABLE,
+	.misc_disable_reg = AMD_CPM6_MISC_EVENT_DISABLE,
+	.misc_mask_all = AMD_CPM6_MISC_EVENT_MASK_ALL,
+	.intx_hwirq = AMD_CPM6_PCIE_INTR_INTX,
+	.intx_mask = BIT(AMD_CPM6_PCIE_INTR_INTX),
+};
+
+static const struct amd_mdb_pcie_variant cpm6_host1 = {
+	.version = CPM6_HOST1,
+	.misc_status_reg = AMD_CPM6_MISC_EVENT_STATUS,
+	.misc_mask_reg = AMD_CPM6_MISC_EVENT_MASK,
+	.misc_enable_reg = AMD_CPM6_MISC_EVENT_ENABLE,
+	.misc_disable_reg = AMD_CPM6_MISC_EVENT_DISABLE,
+	.misc_mask_all = AMD_CPM6_MISC_EVENT_MASK_ALL,
+	.intx_hwirq = AMD_CPM6_PCIE_INTR_INTX,
+	.intx_mask = BIT(AMD_CPM6_PCIE_INTR_INTX),
+};
+
+static const struct amd_mdb_pcie_variant mdb5_host = {
+	.version = MDB5,
+	.misc_status_reg = AMD_MDB_TLP_IR_STATUS_MISC,
+	.misc_mask_reg = AMD_MDB_TLP_IR_MASK_MISC,
+	.misc_enable_reg = AMD_MDB_TLP_IR_ENABLE_MISC,
+	.misc_disable_reg = AMD_MDB_TLP_IR_DISABLE_MISC,
+	.misc_mask_all = AMD_MDB_PCIE_IMR_ALL_MASK,
+	.intx_hwirq = AMD_MDB_PCIE_INTR_INTX,
+	.intx_mask = AMD_MDB_TLP_PCIE_INTX_MASK,
+};
+
 static const struct of_device_id amd_mdb_pcie_of_match[] = {
 	{
 		.compatible = "amd,versal2-mdb-host",
+		.data = &mdb5_host,
+	},
+	{
+		.compatible = "amd,versal2-cpm6-host",
+		.data = &cpm6_host,
+	},
+	{
+		.compatible = "amd,versal2-cpm6-host1",
+		.data = &cpm6_host1,
 	},
 	{},
 };
-- 
2.44.4


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* Re: [PATCH v4 1/2] dt-bindings: PCI: amd-mdb: Add CPM6 support
  2026-08-08 10:52 ` [PATCH v4 1/2] dt-bindings: PCI: amd-mdb: Add CPM6 support Sai Krishna Musham
@ 2026-08-08 10:57   ` sashiko-bot
  0 siblings, 0 replies; 5+ messages in thread
From: sashiko-bot @ 2026-08-08 10:57 UTC (permalink / raw)
  To: Sai Krishna Musham; +Cc: linux-pci, conor+dt, devicetree, robh

> The AMD CPM6 PCIe controller is based on the Synopsys DesignWare PCIe IP.
> Add "intr" to snps,dw-pcie.yaml vendor-specific reg-names for the
> per-controller interrupt register region used by CPM6.
> 
> Update amd,versal2-mdb-host.yaml with separate register definitions:
> - MDB5: 4 regions (slcr, config, dbi, atu)
> - CPM6: 5 regions (slcr, config, dbi, atu, intr)
> 
> Signed-off-by: Sai Krishna Musham <sai.krishna.musham@amd.com>

Sashiko has reviewed this patch and found no issues. It looks great!

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260808105256.3705889-1-sai.krishna.musham@amd.com?part=1


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH v4 2/2] PCI: amd-mdb: Add CPM6 host controller support
  2026-08-08 10:52 ` [PATCH v4 2/2] PCI: amd-mdb: Add CPM6 host controller support Sai Krishna Musham
@ 2026-08-08 11:06   ` sashiko-bot
  0 siblings, 0 replies; 5+ messages in thread
From: sashiko-bot @ 2026-08-08 11:06 UTC (permalink / raw)
  To: Sai Krishna Musham; +Cc: conor+dt, robh, devicetree, linux-pci

Thank you for your contribution! Sashiko AI review found 5 potential issue(s) to consider:

New issues:
- [High] The shared event IRQ handler unconditionally returns IRQ_HANDLED and clears shared aggregators, breaking spurious interrupt detection.
- [High] Probe error paths fail to disable hardware interrupts, leading to unrecoverable IRQ storms on the shared interrupt line.
- [High] Hardware interrupts are unmasked before the interrupt handler is registered, creating a race window for an IRQ storm.

Pre-existing issues:
- [High] Probe failure cleanup frees IRQ domains while devm-managed handlers are still active, leading to Use-After-Free.
- [High] MDB5 INTB, INTC, and INTD interrupts trigger a WARN_ON_ONCE due to attempting to dispatch unmapped hwirqs.
--

--- Patch [2]: [PATCH v4 2/2] PCI: amd-mdb: Add CPM6 host controller support ---
commit b10c5fc8b35403027eb649f10e64e4ddee1b0920
Author: Sai Krishna Musham <sai.krishna.musham@amd.com>
Subject: PCI: amd-mdb: Add CPM6 host controller support

This commit extends the AMD MDB PCIe driver to support AMD Versal2 CPM6
variants. It adds variant-specific register layouts and implements
aggregator clearing for the new shared interrupt architecture.

Link: https://lore.kernel.org/all/20260402180006.486229-3-sai.krishna.musham@amd.com/
Link: https://lore.kernel.org/all/20260728202044.1785986-3-sai.krishna.musham@amd.com/
Link: https://lore.kernel.org/all/20260803144412.713639-3-sai.krishna.musham@amd.com/

> diff --git a/drivers/pci/controller/dwc/pcie-amd-mdb.c b/drivers/pci/controller/dwc/pcie-amd-mdb.c
> index dee295f281a30..3e8e02ab96bce 100644
> --- a/drivers/pci/controller/dwc/pcie-amd-mdb.c
> +++ b/drivers/pci/controller/dwc/pcie-amd-mdb.c
[ ... ]
> @@ -226,13 +387,44 @@ static irqreturn_t amd_mdb_pcie_event(int irq, void *args)
>  {
>  	struct amd_mdb_pcie *pcie = args;
>  	unsigned long val;
> +	u32 ev_raw, err;
>  	int i;
>  
> -	val = readl_relaxed(pcie->slcr + AMD_MDB_TLP_IR_STATUS_MISC);
> -	val &= ~readl_relaxed(pcie->slcr + AMD_MDB_TLP_IR_MASK_MISC);
> +	ev_raw = readl_relaxed(pcie->intr_base + pcie->variant->misc_status_reg);
> +	val = ev_raw;
> +	val &= ~readl_relaxed(pcie->intr_base + pcie->variant->misc_mask_reg);
> +
> +	if (pcie->variant->version == MDB5) {
> +		for_each_set_bit(i, &val, 32)
> +			generic_handle_domain_irq(pcie->mdb_domain, i);

[Severity: High]
This isn't a bug introduced by this patch, but does this code cause a
WARN_ON_ONCE if INTB, INTC, or INTD is asserted on MDB5?

The mask AMD_MDB_TLP_PCIE_INTX_MASK covers bits 16 to 23 of misc_status_reg.
This loop dispatches all set bits to generic_handle_domain_irq(). However,
in amd_mdb_setup_irq(), only INTA (hwirq 16) is mapped via
irq_create_mapping().

If INTB (bit 18) asserts, the handler will dispatch hwirq 18, which is
unmapped, failing to invoke the handler and triggering a warning in the IRQ
core.

> +		writel_relaxed(val, pcie->intr_base + pcie->variant->misc_status_reg);
> +		return IRQ_HANDLED;
> +	}
> +
> +	err = readl_relaxed(pcie->intr_base + AMD_CPM6_PCIE_ERR_STATUS);
> +
> +	val &= pcie->variant->misc_mask_all;
> +
>  	for_each_set_bit(i, &val, 32)
>  		generic_handle_domain_irq(pcie->mdb_domain, i);
> -	writel_relaxed(val, pcie->slcr + AMD_MDB_TLP_IR_STATUS_MISC);
> +
> +	/* Clear handled + any unhandled sticky bits to avoid IRQ storms. */
> +	writel_relaxed(ev_raw, pcie->intr_base + pcie->variant->misc_status_reg);
> +
> +	/* On CPM6 completion timeout is reported via PCIE_ERR. */
> +	if (err) {
> +		u32 pending = err & ~readl_relaxed(pcie->intr_base +
> +						   AMD_CPM6_PCIE_ERR_MASK);
> +
> +		if (pending & AMD_CPM6_PCIE_ERR_MASK_ALL)
> +			generic_handle_domain_irq(pcie->mdb_domain,
> +						  AMD_CPM6_PCIE_INTR_CMPL_TIMEOUT);
> +		/* Clear every asserted bit so the leaf and MERGED de-assert. */
> +		writel_relaxed(err, pcie->intr_base + AMD_CPM6_PCIE_ERR_STATUS);
> +	}
> +
> +	/* Sticky aggregation bits; clear each pass or the IRQ re-fires */
> +	amd_mdb_pcie_clear_aggregators(pcie);
>  
>  	return IRQ_HANDLED;
>  }

[Severity: High]
Could returning IRQ_HANDLED unconditionally here break the kernel's spurious
interrupt detection?

Since this IRQ is shared (as seen by the IRQF_SHARED flag), any interrupt on
the shared line invokes this handler. If the interrupt was generated by the
other host controller and this one is idle (ev_raw == 0 and err == 0), this
code returns IRQ_HANDLED without doing any work.

Additionally, unconditionally calling amd_mdb_pcie_clear_aggregators() clears
the shared PS_MISC_IR_STATUS aggregator, which could drop interrupts for the
other host controller.

[ ... ]
> @@ -351,15 +569,20 @@ static int amd_mdb_setup_irq(struct amd_mdb_pcie *pcie,
>  	struct dw_pcie *pci = &pcie->pci;
>  	struct dw_pcie_rp *pp = &pci->pp;
>  	struct device *dev = &pdev->dev;
> +	const struct amd_mdb_intr_cause *intr_cause;
> +	unsigned long event_flags = IRQF_NO_THREAD;
>  	int i, irq, err;
>  
> +	intr_cause = pcie->variant->version == MDB5 ?
> +		     mdb5_intr_cause : cpm6_intr_cause;
> +
>  	amd_mdb_pcie_init_port(pcie);

[Severity: High]
Is there a race condition leading to an IRQ storm by enabling hardware
interrupts before registering the handler?

amd_mdb_pcie_init_port() enables and unmasks hardware interrupts and shared
aggregators. The devm_request_irq() for the handler happens later in this
function.

For CPM6's shared IRQ, if the hardware asserts an interrupt during this window,
the shared line fires. The other host's handler will spin endlessly because it
cannot clear this un-probed host's internal hardware state.

>  
>  	pp->irq = platform_get_irq(pdev, 0);
>  	if (pp->irq < 0)
>  		return pp->irq;
[ ... ]
> +	/*
> +	 * On CPM6 the misc GIC line is shared between both host controllers,
> +	 * so the event IRQ must allow sharing.
> +	 */
> +	if (pcie->variant->version != MDB5)
> +		event_flags |= IRQF_SHARED;
> +
>  	/* Plug the main event handler. */
> -	err = devm_request_irq(dev, pp->irq, amd_mdb_pcie_event, IRQF_NO_THREAD,
> +	err = devm_request_irq(dev, pp->irq, amd_mdb_pcie_event, event_flags,
>  			       "amd_mdb pcie_irq", pcie);
>  	if (err) {
>  		dev_err(dev, "Failed to request event IRQ %d, err=%d\n",
> -			pp->irq, err);
> +			pp->irq, err);
>  		return err;
>  	}

[Severity: High]
Will the probe error path leave hardware interrupts enabled and unmasked?

If a failure occurs after this point (e.g., in dw_pcie_host_init()), the error
path in amd_mdb_add_pcie_port() only cleans up the IRQ domains. It fails to
mask the hardware interrupts and shared aggregators that were just enabled in
amd_mdb_pcie_init_port().

Because the IRQ is shared on CPM6, if the uninitialized hardware asserts an
interrupt, the other host controller's handler will be continuously invoked but
unable to clear the source, causing an unrecoverable IRQ storm.

[Severity: High]
This isn't a bug introduced by this patch, but can the probe error path cause
a use-after-free of the IRQ domain?

If a failure occurs after devm_request_irq() succeeds, amd_mdb_add_pcie_port()
jumps to its error path and calls amd_mdb_pcie_free_irq_domains(), which
immediately destroys and frees pcie->mdb_domain. However, the IRQ handlers
were registered with devm-managed functions and remain active until devres
cleanup runs after the probe function returns.

If an interrupt fires in this window (which is likely if hardware interrupts
are left enabled), amd_mdb_pcie_event() will dereference the freed domain via:

amd_mdb_pcie_event() {
    ...
    generic_handle_domain_irq(pcie->mdb_domain, i);
    ...
}

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260808105256.3705889-1-sai.krishna.musham@amd.com?part=2

^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-08-08 11:06 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-08 10:52 [PATCH v4 0/2] pci: AMD: Add Versal2 CPM6 PCIe host controller support Sai Krishna Musham
2026-08-08 10:52 ` [PATCH v4 1/2] dt-bindings: PCI: amd-mdb: Add CPM6 support Sai Krishna Musham
2026-08-08 10:57   ` sashiko-bot
2026-08-08 10:52 ` [PATCH v4 2/2] PCI: amd-mdb: Add CPM6 host controller support Sai Krishna Musham
2026-08-08 11:06   ` sashiko-bot

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox