* [PATCH v20 1/9] PCI/AER: Introduce AER-CXL protocol error kfifo
2026-09-02 13:39 [PATCH v20 0/9] Enable CXL PCIe Port Protocol Error handling and logging Terry Bowman
@ 2026-09-02 13:39 ` Terry Bowman
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-02 13:39 ` [PATCH v20 2/9] PCI: Establish common CXL Port protocol error flow Terry Bowman
` (7 subsequent siblings)
8 siblings, 1 reply; 17+ messages in thread
From: Terry Bowman @ 2026-09-02 13:39 UTC (permalink / raw)
To: Jonathan Cameron, Dave Jiang, Alison Schofield, Vishal Verma,
Davidlohr Bueso, Bjorn Helgaas, Dan Williams, Rafael J . Wysocki,
Jonathan Corbet, linux-cxl
Cc: Tony Luck, Borislav Petkov, Hanjun Guo, Mauro Carvalho Chehab,
Shuai Xue, Len Brown, Ira Weiny, Li Ming, Shuah Khan,
Ben Cheatham, Richard Cheng, Robert Richter, Lukas Wunner,
linux-pci, linux-acpi, linux-doc, linux-kernel
CXL VH RAS handling requires the AER driver to hand off CXL protocol
errors to cxl_core for logging and recovery before PCIe AER recovery
tears down the device. Introduce pci/pcie/aer_cxl_vh.c to implement
this handoff via a kfifo-backed work item.
The producer, cxl_forward_error(), is gated by is_cxl_error() and
enqueues the error source PCI device and severity. cxl_core registers a
consumer via cxl_register_proto_err_work(); the consumer drains the
kfifo with for_each_cxl_proto_err(). For uncorrectable errors,
cxl_proto_err_wait_for_empty() lets the AER path block until the CXL
plane has finished so recovery does not race device teardown.
A rwsem serializes registration, deregistration, enqueue, and dequeue
against concurrent AER IRQ threads; a spinlock serializes concurrent
kfifo writers. is_aer_internal_error() moves into this file and now
evaluates info->status & ~info->mask rather than the raw info->status,
so a masked internal-error bit is treated as not-set. For the RCH RCEC
path this is equivalent because cxl_rch_enable_rcec() first calls
pci_aer_unmask_internal_errors(), which clears those mask bits in
hardware before the AER status is read back.
A subsequent patch wires cxl_forward_error() into handle_error_source().
Add MAINTAINERS entries for aer_cxl_vh.c and aer_cxl_rch.c under the CXL
entry.
Co-developed-by: Dan Williams <djbw@kernel.org>
Signed-off-by: Dan Williams <djbw@kernel.org>
Signed-off-by: Terry Bowman <terry.bowman@amd.com>
---
Changes in v19->v20:
- Drop a correctable error on full kfifo instead of panicking
- Add explicit ATOMIC_INIT(0) for flush_inflight
- Change is_aer_internal_error() to use mask in evaluation.
- Condense commit message (Jonathan)
- Document constraints for_each_cxl_proto_err()
- Document the @wd and @fn parameters of for_each_cxl_proto_err() (kernel-doc)
Changes in v18 -> v19:
- Rename cxl_proto_err_flush() to cxl_proto_err_wait_for_empty() to
better reflect that it waits for the kfifo to drain (Jonathan).
Changes in v17->v18:
- Remove correctable status clear from cxl_forward_error(); the AER core
clears all status bits via pci_aer_handle_error() info->status writeback
- Schedule consumer on kfifo overflow so existing entries can be drained
Changes in v16->v17:
- Reword "kfifo semaphore" to "kfifo spinlock" to match fifo_lock.
- Defer the handle_error_source() is_cxl_error() switch to the patch that
registers the kfifo consumer to keep each commit bisect-safe.
- Rename rwsema to rwsem
- Change CPER exports to use EXPORT_SYMBOL_FOR_MODULES.
- Add work cancel function.
- Replace kfifo_put() with kfifo_in_spinlocked() for multiple producers
- Add fifo_lock spinlock for concurrent producer serialisation
- Initialize the embedded kfifo with INIT_KFIFO() in a subsys_initcall so
kfifo->mask, ->esize and ->data are set before first use.
- Clear PCI_ERR_COR_STATUS in cxl_forward_error() after enqueue so the
device is acked for correctable events even when the consumer drops the
event. Uncorrectable status is left for cxl_do_recovery() to clear after
recovery completes, mirroring the AER core convention.
- WARN on double-registration in cxl_register_proto_err_work() to make an
unintended second consumer visible at runtime.
- Add direct rwsem.h, cleanup.h and workqueue.h includes for symbols used
in aer_cxl_vh.c
- Add MAINTAINERS entries for drivers/pci/pcie/aer_cxl_*.c
- Update message
---
MAINTAINERS | 2 +
drivers/pci/pcie/Makefile | 1 +
drivers/pci/pcie/aer.c | 10 --
drivers/pci/pcie/aer_cxl_vh.c | 243 ++++++++++++++++++++++++++++++++++
drivers/pci/pcie/portdrv.h | 6 +
include/linux/aer.h | 24 ++++
6 files changed, 276 insertions(+), 10 deletions(-)
create mode 100644 drivers/pci/pcie/aer_cxl_vh.c
diff --git a/MAINTAINERS b/MAINTAINERS
index 3a19da74d00c9..f6ca37995ff84 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -6572,6 +6572,8 @@ S: Maintained
F: Documentation/driver-api/cxl
F: Documentation/userspace-api/fwctl/fwctl-cxl.rst
F: drivers/cxl/
+F: drivers/pci/pcie/aer_cxl_rch.c
+F: drivers/pci/pcie/aer_cxl_vh.c
F: include/cxl/
F: include/uapi/linux/cxl_mem.h
F: tools/testing/cxl/
diff --git a/drivers/pci/pcie/Makefile b/drivers/pci/pcie/Makefile
index b0b43a18c304b..62d3d3c69a5df 100644
--- a/drivers/pci/pcie/Makefile
+++ b/drivers/pci/pcie/Makefile
@@ -9,6 +9,7 @@ obj-$(CONFIG_PCIEPORTBUS) += pcieportdrv.o bwctrl.o
obj-y += aspm.o
obj-$(CONFIG_PCIEAER) += aer.o err.o tlp.o
obj-$(CONFIG_CXL_RAS) += aer_cxl_rch.o
+obj-$(CONFIG_CXL_RAS) += aer_cxl_vh.o
obj-$(CONFIG_PCIEAER_INJECT) += aer_inject.o
obj-$(CONFIG_PCIE_PME) += pme.o
obj-$(CONFIG_PCIE_DPC) += dpc.o
diff --git a/drivers/pci/pcie/aer.c b/drivers/pci/pcie/aer.c
index d8dcd238fda1f..21dfc9c933d77 100644
--- a/drivers/pci/pcie/aer.c
+++ b/drivers/pci/pcie/aer.c
@@ -1288,16 +1288,6 @@ void pci_aer_unmask_internal_errors(struct pci_dev *dev)
*/
EXPORT_SYMBOL_FOR_MODULES(pci_aer_unmask_internal_errors, "cxl_core");
-#ifdef CONFIG_CXL_RAS
-bool is_aer_internal_error(struct aer_err_info *info)
-{
- if (info->severity == AER_CORRECTABLE)
- return info->status & PCI_ERR_COR_INTERNAL;
-
- return info->status & PCI_ERR_UNC_INTN;
-}
-#endif
-
/**
* pci_aer_handle_error - handle logging error into an event log
* @dev: pointer to pci_dev data structure of error source device
diff --git a/drivers/pci/pcie/aer_cxl_vh.c b/drivers/pci/pcie/aer_cxl_vh.c
new file mode 100644
index 0000000000000..9fc12d4e644bc
--- /dev/null
+++ b/drivers/pci/pcie/aer_cxl_vh.c
@@ -0,0 +1,243 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Copyright(c) 2026 AMD Corporation. All rights reserved. */
+
+#include <linux/aer.h>
+#include <linux/atomic.h>
+#include <linux/cleanup.h>
+#include <linux/init.h>
+#include <linux/kfifo.h>
+#include <linux/lockdep.h>
+#include <linux/rwsem.h>
+#include <linux/spinlock.h>
+#include <linux/wait_bit.h>
+#include <linux/workqueue.h>
+#include "../pci.h"
+#include "portdrv.h"
+
+#define CXL_ERROR_SOURCES_MAX 128
+
+struct cxl_proto_err_kfifo {
+ struct work_struct *work;
+ void (*flush)(void);
+ struct rw_semaphore rwsem;
+ spinlock_t fifo_lock; /* Serializes kfifo writers */
+ atomic_t flush_inflight;
+ DECLARE_KFIFO(fifo, struct cxl_proto_err_work_data,
+ CXL_ERROR_SOURCES_MAX);
+};
+
+static struct cxl_proto_err_kfifo cxl_proto_err_kfifo = {
+ .rwsem = __RWSEM_INITIALIZER(cxl_proto_err_kfifo.rwsem),
+ .fifo_lock = __SPIN_LOCK_UNLOCKED(cxl_proto_err_kfifo.fifo_lock),
+ .flush_inflight = ATOMIC_INIT(0),
+};
+
+static int __init cxl_proto_err_kfifo_init(void)
+{
+ INIT_KFIFO(cxl_proto_err_kfifo.fifo);
+ return 0;
+}
+subsys_initcall(cxl_proto_err_kfifo_init);
+
+bool is_aer_internal_error(struct aer_err_info *info)
+{
+ u32 status = info->status & ~info->mask;
+
+ if (info->severity == AER_CORRECTABLE)
+ return status & PCI_ERR_COR_INTERNAL;
+
+ return status & PCI_ERR_UNC_INTN;
+}
+
+bool is_cxl_error(struct pci_dev *pdev, struct aer_err_info *info)
+{
+ if (!info || !info->is_cxl)
+ return false;
+
+ if (pci_pcie_type(pdev) != PCI_EXP_TYPE_ENDPOINT)
+ return false;
+
+ return is_aer_internal_error(info);
+}
+
+/**
+ * cxl_forward_error - Forward a CXL protocol error to the CXL subsystem via kfifo
+ * @pdev: PCI device that reported the AER error
+ * @info: AER error info containing severity and status
+ *
+ * Producer side of the AER-CXL kfifo. Enqueues a CXL protocol error work
+ * item and schedules the consumer workqueue. Takes a reference on @pdev
+ * that the consumer releases after handling.
+ *
+ * Return: true if the consumer workqueue was scheduled and the caller may
+ * need to drain the kfifo before AER recovery; false if no CXL error
+ * handling was initiated due to an early return on error (e.g. no kfifo
+ * consumer registered). Note that on a full kfifo a correctable error is
+ * dropped but true is still returned; this is harmless because the caller
+ * only drains the kfifo for non-correctable events.
+ */
+bool cxl_forward_error(struct pci_dev *pdev, struct aer_err_info *info)
+{
+ struct cxl_proto_err_work_data wd = {
+ .severity = info->severity,
+ .pdev = pdev,
+ };
+
+ guard(rwsem_read)(&cxl_proto_err_kfifo.rwsem);
+
+ if (!cxl_proto_err_kfifo.work) {
+ dev_err_ratelimited(&pdev->dev, "AER-CXL kfifo reader not registered\n");
+ return false;
+ }
+
+ /*
+ * Reference discipline: the AER caller (handle_error_source()) holds
+ * a ref on @pdev for the duration of this call and releases it on
+ * return. Take a fresh ref here so the pdev stays live while queued
+ * in the kfifo; the corresponding consumer is for_each_cxl_proto_err()
+ * and will drop that ref after handling. On enqueue failure below,
+ * drop the ref we just took to avoid a leak.
+ */
+ pci_dev_get(pdev);
+
+ /* Serialize concurrent kfifo writers: multiple AER threaded IRQs */
+ if (!kfifo_in_spinlocked(&cxl_proto_err_kfifo.fifo, &wd, 1,
+ &cxl_proto_err_kfifo.fifo_lock)) {
+
+ pci_dev_put(pdev);
+
+ /*
+ * A correctable error is device-local and recoverable, so a
+ * dropped CE is safe to log and discard. Never panic on CE
+ * pressure: a CE storm must not fill the fifo and escalate a
+ * later event.
+ */
+ if (info->severity == AER_CORRECTABLE) {
+ dev_err_ratelimited(&pdev->dev,
+ "AER-CXL kfifo full, CE dropped\n");
+ return true;
+ }
+
+ /*
+ * Unlike PCIe AER, a dropped CXL.mem uncorrectable error
+ * cannot be treated as device-local: it may signal lost cache
+ * coherency over HDM memory in active use. The error can no
+ * longer be confirmed via CXL RAS, and reaching here means the
+ * fifo is saturated with pending protocol errors. Collapse the
+ * unknown state to the same conservative outcome as a confirmed
+ * UCE.
+ */
+ panic("CXL: dropped uncorrectable protocol error\n");
+ }
+
+ schedule_work(cxl_proto_err_kfifo.work);
+ return true;
+}
+
+void cxl_register_proto_err_work(struct work_struct *work,
+ void (*flush)(void))
+{
+ guard(rwsem_write)(&cxl_proto_err_kfifo.rwsem);
+
+ /*
+ * Warn on double-registration to surface driver bugs (e.g. missing
+ * cxl_unregister_proto_err_work() on module exit)
+ */
+ if (WARN(cxl_proto_err_kfifo.work,
+ "AER-CXL kfifo consumer already registered\n"))
+ return;
+ cxl_proto_err_kfifo.work = work;
+ cxl_proto_err_kfifo.flush = flush;
+}
+EXPORT_SYMBOL_FOR_MODULES(cxl_register_proto_err_work, "cxl_core");
+
+static struct work_struct *cancel_cxl_proto_err(void)
+{
+ struct work_struct *work;
+ struct cxl_proto_err_work_data wd;
+
+ guard(rwsem_write)(&cxl_proto_err_kfifo.rwsem);
+ work = cxl_proto_err_kfifo.work;
+ cxl_proto_err_kfifo.work = NULL;
+ cxl_proto_err_kfifo.flush = NULL;
+
+ /* rwsem_write excludes all producers; fifo_lock not needed */
+ while (kfifo_get(&cxl_proto_err_kfifo.fifo, &wd)) {
+ dev_err_ratelimited(&wd.pdev->dev,
+ "AER-CXL error report canceled\n");
+ pci_dev_put(wd.pdev);
+ }
+ return work;
+}
+
+void cxl_unregister_proto_err_work(void)
+{
+ struct work_struct *work;
+
+ lockdep_assert_not_held(&cxl_proto_err_kfifo.rwsem);
+
+ work = cancel_cxl_proto_err();
+
+ /* Wait for any in-flight cxl_proto_err_wait_for_empty() calls to complete */
+ wait_var_event(&cxl_proto_err_kfifo.flush_inflight,
+ atomic_read(&cxl_proto_err_kfifo.flush_inflight) == 0);
+
+ if (work)
+ cancel_work_sync(work);
+}
+EXPORT_SYMBOL_FOR_MODULES(cxl_unregister_proto_err_work, "cxl_core");
+
+/**
+ * for_each_cxl_proto_err - Call a function for each kfifo work item
+ * @wd: caller-provided work-data scratch buffer, populated per kfifo entry
+ * @fn: callback invoked for each dequeued &struct cxl_proto_err_work_data
+ *
+ * Single-consumer invariant: this function is only called from
+ * cxl_proto_err_work_fn() via a single DECLARE_WORK.
+ *
+ * Holds rwsem_read internally; fn() must not call cxl_register_proto_err_work(),
+ * cxl_unregister_proto_err_work(), or cxl_proto_err_wait_for_empty() (all take
+ * or wait on the same rwsem).
+ */
+void for_each_cxl_proto_err(struct cxl_proto_err_work_data *wd,
+ cxl_proto_err_fn_t fn)
+{
+ guard(rwsem_read)(&cxl_proto_err_kfifo.rwsem);
+ while (kfifo_get(&cxl_proto_err_kfifo.fifo, wd)) {
+ fn(wd);
+
+ /* Corresponding ref incr taken in cxl_forward_error() */
+ pci_dev_put(wd->pdev);
+ }
+}
+EXPORT_SYMBOL_FOR_MODULES(for_each_cxl_proto_err, "cxl_core");
+
+/**
+ * cxl_proto_err_wait_for_empty - drain pending AER-CXL kfifo work synchronously
+ *
+ * Ensures CXL RAS handling and panic policy complete before AER
+ * recovery proceeds. Only needed for UCE; CE runs asynchronously.
+ *
+ * Snapshots the flush callback under rwsem_read, then releases the
+ * rwsem before calling it to avoid deadlock with a concurrent
+ * rwsem_write from cxl_unregister_proto_err_work().
+ *
+ * The flush_inflight counter (typically 0 or 1) prevents module
+ * unload while a flush is in progress outside the rwsem.
+ */
+void cxl_proto_err_wait_for_empty(void)
+{
+ void (*flush)(void);
+
+ scoped_guard(rwsem_read, &cxl_proto_err_kfifo.rwsem) {
+ flush = cxl_proto_err_kfifo.flush;
+ if (flush)
+ atomic_inc(&cxl_proto_err_kfifo.flush_inflight);
+ }
+
+ if (flush) {
+ flush();
+ if (atomic_dec_and_test(&cxl_proto_err_kfifo.flush_inflight))
+ wake_up_var(&cxl_proto_err_kfifo.flush_inflight);
+ }
+}
diff --git a/drivers/pci/pcie/portdrv.h b/drivers/pci/pcie/portdrv.h
index cc58bf2f2c844..357310916088f 100644
--- a/drivers/pci/pcie/portdrv.h
+++ b/drivers/pci/pcie/portdrv.h
@@ -130,9 +130,15 @@ struct aer_err_info;
bool is_aer_internal_error(struct aer_err_info *info);
void cxl_rch_handle_error(struct pci_dev *dev, struct aer_err_info *info);
void cxl_rch_enable_rcec(struct pci_dev *rcec);
+bool is_cxl_error(struct pci_dev *pdev, struct aer_err_info *info);
+bool cxl_forward_error(struct pci_dev *pdev, struct aer_err_info *info);
+void cxl_proto_err_wait_for_empty(void);
#else
static inline bool is_aer_internal_error(struct aer_err_info *info) { return false; }
static inline void cxl_rch_handle_error(struct pci_dev *dev, struct aer_err_info *info) { }
static inline void cxl_rch_enable_rcec(struct pci_dev *rcec) { }
+static inline bool is_cxl_error(struct pci_dev *pdev, struct aer_err_info *info) { return false; }
+static inline bool cxl_forward_error(struct pci_dev *pdev, struct aer_err_info *info) { return false; }
+static inline void cxl_proto_err_wait_for_empty(void) { }
#endif /* CONFIG_CXL_RAS */
#endif /* _PORTDRV_H_ */
diff --git a/include/linux/aer.h b/include/linux/aer.h
index df0f5c382286f..8eba3192e2d15 100644
--- a/include/linux/aer.h
+++ b/include/linux/aer.h
@@ -25,6 +25,7 @@
#define PCIE_STD_MAX_TLP_HEADERLOG (PCIE_STD_NUM_TLP_HEADERLOG + 10)
struct pci_dev;
+struct work_struct;
struct pcie_tlp_log {
union {
@@ -66,6 +67,29 @@ static inline int pcie_aer_is_native(struct pci_dev *dev) { return 0; }
static inline void pci_aer_unmask_internal_errors(struct pci_dev *dev) { }
#endif
+#ifdef CONFIG_CXL_RAS
+/**
+ * struct cxl_proto_err_work_data - Error information used in CXL error handling
+ * @pdev: PCI device detecting the error
+ * @severity: AER severity
+ */
+struct cxl_proto_err_work_data {
+ struct pci_dev *pdev;
+ int severity;
+};
+
+/**
+ * Callback for processing a CXL protocol error from the AER-CXL kfifo.
+ */
+typedef void (*cxl_proto_err_fn_t)(struct cxl_proto_err_work_data *wd);
+
+void cxl_register_proto_err_work(struct work_struct *work,
+ void (*flush)(void));
+void for_each_cxl_proto_err(struct cxl_proto_err_work_data *wd,
+ cxl_proto_err_fn_t fn);
+void cxl_unregister_proto_err_work(void);
+#endif
+
void pci_print_aer(struct pci_dev *dev, int aer_severity,
struct aer_capability_regs *aer);
int cper_severity_to_aer(int cper_severity);
base-commit: cee9395acd8043be0644b25c34bfa86623f2b935
--
2.34.1
^ permalink raw reply related [flat|nested] 17+ messages in thread* Re: [PATCH v20 1/9] PCI/AER: Introduce AER-CXL protocol error kfifo
2026-09-02 13:39 ` [PATCH v20 1/9] PCI/AER: Introduce AER-CXL protocol error kfifo Terry Bowman
@ 2026-09-02 20:57 ` Cheatham, Benjamin
0 siblings, 0 replies; 17+ messages in thread
From: Cheatham, Benjamin @ 2026-09-02 20:57 UTC (permalink / raw)
To: Terry Bowman, Jonathan Cameron, Dave Jiang, Alison Schofield,
Vishal Verma, Davidlohr Bueso, Bjorn Helgaas, Dan Williams,
Rafael J . Wysocki, Jonathan Corbet, linux-cxl
Cc: Tony Luck, Borislav Petkov, Hanjun Guo, Mauro Carvalho Chehab,
Shuai Xue, Len Brown, Ira Weiny, Li Ming, Shuah Khan,
Richard Cheng, Robert Richter, Lukas Wunner, linux-pci,
linux-acpi, linux-doc, linux-kernel
On 9/2/2026 8:39 AM, Terry Bowman wrote:
> CXL VH RAS handling requires the AER driver to hand off CXL protocol
> errors to cxl_core for logging and recovery before PCIe AER recovery
> tears down the device. Introduce pci/pcie/aer_cxl_vh.c to implement
> this handoff via a kfifo-backed work item.
>
> The producer, cxl_forward_error(), is gated by is_cxl_error() and
> enqueues the error source PCI device and severity. cxl_core registers a
> consumer via cxl_register_proto_err_work(); the consumer drains the
> kfifo with for_each_cxl_proto_err(). For uncorrectable errors,
> cxl_proto_err_wait_for_empty() lets the AER path block until the CXL
> plane has finished so recovery does not race device teardown.
>
> A rwsem serializes registration, deregistration, enqueue, and dequeue
> against concurrent AER IRQ threads; a spinlock serializes concurrent
> kfifo writers. is_aer_internal_error() moves into this file and now
> evaluates info->status & ~info->mask rather than the raw info->status,
> so a masked internal-error bit is treated as not-set. For the RCH RCEC
> path this is equivalent because cxl_rch_enable_rcec() first calls
> pci_aer_unmask_internal_errors(), which clears those mask bits in
> hardware before the AER status is read back.
>
> A subsequent patch wires cxl_forward_error() into handle_error_source().
>
> Add MAINTAINERS entries for aer_cxl_vh.c and aer_cxl_rch.c under the CXL
> entry.
>
> Co-developed-by: Dan Williams <djbw@kernel.org>
> Signed-off-by: Dan Williams <djbw@kernel.org>
> Signed-off-by: Terry Bowman <terry.bowman@amd.com>
>
> ---
Just a few nits, but nothing critical so:
Revewied-by: Ben Cheatham <benjamin.cheatham@amd.com>
>
> Changes in v19->v20:
> - Drop a correctable error on full kfifo instead of panicking
> - Add explicit ATOMIC_INIT(0) for flush_inflight
> - Change is_aer_internal_error() to use mask in evaluation.
> - Condense commit message (Jonathan)
> - Document constraints for_each_cxl_proto_err()
> - Document the @wd and @fn parameters of for_each_cxl_proto_err() (kernel-doc)
>
> Changes in v18 -> v19:
> - Rename cxl_proto_err_flush() to cxl_proto_err_wait_for_empty() to
> better reflect that it waits for the kfifo to drain (Jonathan).
>
> Changes in v17->v18:
> - Remove correctable status clear from cxl_forward_error(); the AER core
> clears all status bits via pci_aer_handle_error() info->status writeback
> - Schedule consumer on kfifo overflow so existing entries can be drained
>
> Changes in v16->v17:
> - Reword "kfifo semaphore" to "kfifo spinlock" to match fifo_lock.
> - Defer the handle_error_source() is_cxl_error() switch to the patch that
> registers the kfifo consumer to keep each commit bisect-safe.
> - Rename rwsema to rwsem
> - Change CPER exports to use EXPORT_SYMBOL_FOR_MODULES.
> - Add work cancel function.
> - Replace kfifo_put() with kfifo_in_spinlocked() for multiple producers
> - Add fifo_lock spinlock for concurrent producer serialisation
> - Initialize the embedded kfifo with INIT_KFIFO() in a subsys_initcall so
> kfifo->mask, ->esize and ->data are set before first use.
> - Clear PCI_ERR_COR_STATUS in cxl_forward_error() after enqueue so the
> device is acked for correctable events even when the consumer drops the
> event. Uncorrectable status is left for cxl_do_recovery() to clear after
> recovery completes, mirroring the AER core convention.
> - WARN on double-registration in cxl_register_proto_err_work() to make an
> unintended second consumer visible at runtime.
> - Add direct rwsem.h, cleanup.h and workqueue.h includes for symbols used
> in aer_cxl_vh.c
> - Add MAINTAINERS entries for drivers/pci/pcie/aer_cxl_*.c
> - Update message
> ---
> MAINTAINERS | 2 +
> drivers/pci/pcie/Makefile | 1 +
> drivers/pci/pcie/aer.c | 10 --
> drivers/pci/pcie/aer_cxl_vh.c | 243 ++++++++++++++++++++++++++++++++++
> drivers/pci/pcie/portdrv.h | 6 +
> include/linux/aer.h | 24 ++++
> 6 files changed, 276 insertions(+), 10 deletions(-)
> create mode 100644 drivers/pci/pcie/aer_cxl_vh.c
>
> diff --git a/MAINTAINERS b/MAINTAINERS
> index 3a19da74d00c9..f6ca37995ff84 100644
> --- a/MAINTAINERS
> +++ b/MAINTAINERS
> @@ -6572,6 +6572,8 @@ S: Maintained
> F: Documentation/driver-api/cxl
> F: Documentation/userspace-api/fwctl/fwctl-cxl.rst
> F: drivers/cxl/
> +F: drivers/pci/pcie/aer_cxl_rch.c
> +F: drivers/pci/pcie/aer_cxl_vh.c
> F: include/cxl/
> F: include/uapi/linux/cxl_mem.h
> F: tools/testing/cxl/
> diff --git a/drivers/pci/pcie/Makefile b/drivers/pci/pcie/Makefile
> index b0b43a18c304b..62d3d3c69a5df 100644
> --- a/drivers/pci/pcie/Makefile
> +++ b/drivers/pci/pcie/Makefile
> @@ -9,6 +9,7 @@ obj-$(CONFIG_PCIEPORTBUS) += pcieportdrv.o bwctrl.o
> obj-y += aspm.o
> obj-$(CONFIG_PCIEAER) += aer.o err.o tlp.o
> obj-$(CONFIG_CXL_RAS) += aer_cxl_rch.o
> +obj-$(CONFIG_CXL_RAS) += aer_cxl_vh.o
This can go on the same line as the obj-$(CONFIG_CXL_RAS) above, i.e.:
obj-$(CONFIG_CXL_RAS) += aer_cxl_rch.o aer_cxl_vh.o
> obj-$(CONFIG_PCIEAER_INJECT) += aer_inject.o
> obj-$(CONFIG_PCIE_PME) += pme.o
> obj-$(CONFIG_PCIE_DPC) += dpc.o
> diff --git a/drivers/pci/pcie/aer.c b/drivers/pci/pcie/aer.c
> index d8dcd238fda1f..21dfc9c933d77 100644
> --- a/drivers/pci/pcie/aer.c
> +++ b/drivers/pci/pcie/aer.c
> @@ -1288,16 +1288,6 @@ void pci_aer_unmask_internal_errors(struct pci_dev *dev)
> */
> EXPORT_SYMBOL_FOR_MODULES(pci_aer_unmask_internal_errors, "cxl_core");
>
> -#ifdef CONFIG_CXL_RAS
> -bool is_aer_internal_error(struct aer_err_info *info)
> -{
> - if (info->severity == AER_CORRECTABLE)
> - return info->status & PCI_ERR_COR_INTERNAL;
> -
> - return info->status & PCI_ERR_UNC_INTN;
> -}
> -#endif
> -
> /**
> * pci_aer_handle_error - handle logging error into an event log
> * @dev: pointer to pci_dev data structure of error source device
> diff --git a/drivers/pci/pcie/aer_cxl_vh.c b/drivers/pci/pcie/aer_cxl_vh.c
> new file mode 100644
> index 0000000000000..9fc12d4e644bc
> --- /dev/null
> +++ b/drivers/pci/pcie/aer_cxl_vh.c
> @@ -0,0 +1,243 @@
> +// SPDX-License-Identifier: GPL-2.0-only
> +/* Copyright(c) 2026 AMD Corporation. All rights reserved. */
> +
> +#include <linux/aer.h>
> +#include <linux/atomic.h>
> +#include <linux/cleanup.h>
> +#include <linux/init.h>
> +#include <linux/kfifo.h>
> +#include <linux/lockdep.h>
> +#include <linux/rwsem.h>
> +#include <linux/spinlock.h>
> +#include <linux/wait_bit.h>
> +#include <linux/workqueue.h>
> +#include "../pci.h"
> +#include "portdrv.h"
> +
> +#define CXL_ERROR_SOURCES_MAX 128
> +
> +struct cxl_proto_err_kfifo {
> + struct work_struct *work;
> + void (*flush)(void);
> + struct rw_semaphore rwsem;
> + spinlock_t fifo_lock; /* Serializes kfifo writers */
> + atomic_t flush_inflight;
> + DECLARE_KFIFO(fifo, struct cxl_proto_err_work_data,
> + CXL_ERROR_SOURCES_MAX);
> +};
> +
> +static struct cxl_proto_err_kfifo cxl_proto_err_kfifo = {
> + .rwsem = __RWSEM_INITIALIZER(cxl_proto_err_kfifo.rwsem),
> + .fifo_lock = __SPIN_LOCK_UNLOCKED(cxl_proto_err_kfifo.fifo_lock),
> + .flush_inflight = ATOMIC_INIT(0),
> +};
I don't know what the style is for pcie, but I think the above can get shortened to:
static struct cxl_proto_err_kfifo {
...
} cxl_proto_err_kfifo = {
.rwsem = ...
};
> +
> +static int __init cxl_proto_err_kfifo_init(void)
> +{
> + INIT_KFIFO(cxl_proto_err_kfifo.fifo);
> + return 0;
> +}
> +subsys_initcall(cxl_proto_err_kfifo_init);
> +
> +bool is_aer_internal_error(struct aer_err_info *info)
> +{
> + u32 status = info->status & ~info->mask;
> +
> + if (info->severity == AER_CORRECTABLE)
> + return status & PCI_ERR_COR_INTERNAL;
> +
> + return status & PCI_ERR_UNC_INTN;
> +}
> +
> +bool is_cxl_error(struct pci_dev *pdev, struct aer_err_info *info)
> +{
> + if (!info || !info->is_cxl)
> + return false;
> +
> + if (pci_pcie_type(pdev) != PCI_EXP_TYPE_ENDPOINT)
> + return false;
> +
> + return is_aer_internal_error(info);
> +}
> +
> +/**
> + * cxl_forward_error - Forward a CXL protocol error to the CXL subsystem via kfifo
> + * @pdev: PCI device that reported the AER error
> + * @info: AER error info containing severity and status
> + *
> + * Producer side of the AER-CXL kfifo. Enqueues a CXL protocol error work
> + * item and schedules the consumer workqueue. Takes a reference on @pdev
> + * that the consumer releases after handling.
> + *
> + * Return: true if the consumer workqueue was scheduled and the caller may
> + * need to drain the kfifo before AER recovery; false if no CXL error
> + * handling was initiated due to an early return on error (e.g. no kfifo
> + * consumer registered). Note that on a full kfifo a correctable error is
> + * dropped but true is still returned; this is harmless because the caller
> + * only drains the kfifo for non-correctable events.
> + */
> +bool cxl_forward_error(struct pci_dev *pdev, struct aer_err_info *info)
> +{
> + struct cxl_proto_err_work_data wd = {
> + .severity = info->severity,
> + .pdev = pdev,
> + };
> +
> + guard(rwsem_read)(&cxl_proto_err_kfifo.rwsem);
> +
> + if (!cxl_proto_err_kfifo.work) {
> + dev_err_ratelimited(&pdev->dev, "AER-CXL kfifo reader not registered\n");
> + return false;
> + }
> +
> + /*
> + * Reference discipline: the AER caller (handle_error_source()) holds
> + * a ref on @pdev for the duration of this call and releases it on
> + * return. Take a fresh ref here so the pdev stays live while queued
> + * in the kfifo; the corresponding consumer is for_each_cxl_proto_err()
> + * and will drop that ref after handling. On enqueue failure below,
> + * drop the ref we just took to avoid a leak.
> + */
> + pci_dev_get(pdev);
> +
> + /* Serialize concurrent kfifo writers: multiple AER threaded IRQs */
> + if (!kfifo_in_spinlocked(&cxl_proto_err_kfifo.fifo, &wd, 1,
> + &cxl_proto_err_kfifo.fifo_lock)) {
> +
> + pci_dev_put(pdev);
> +
> + /*
> + * A correctable error is device-local and recoverable, so a
> + * dropped CE is safe to log and discard. Never panic on CE
> + * pressure: a CE storm must not fill the fifo and escalate a
> + * later event.
> + */
I don't think you need this comment when the same info is mentioned in the doc comment
above the function. I'd probably keep the doc comment since it's more visible, but that's
up to you.
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH v20 2/9] PCI: Establish common CXL Port protocol error flow
2026-09-02 13:39 [PATCH v20 0/9] Enable CXL PCIe Port Protocol Error handling and logging Terry Bowman
2026-09-02 13:39 ` [PATCH v20 1/9] PCI/AER: Introduce AER-CXL protocol error kfifo Terry Bowman
@ 2026-09-02 13:39 ` Terry Bowman
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-02 13:39 ` [PATCH v20 3/9] cxl/ras: Handle RCH correctable and uncorrectable errors in one pass Terry Bowman
` (6 subsequent siblings)
8 siblings, 1 reply; 17+ messages in thread
From: Terry Bowman @ 2026-09-02 13:39 UTC (permalink / raw)
To: Jonathan Cameron, Dave Jiang, Alison Schofield, Vishal Verma,
Davidlohr Bueso, Bjorn Helgaas, Dan Williams, Rafael J . Wysocki,
Jonathan Corbet, linux-cxl
Cc: Tony Luck, Borislav Petkov, Hanjun Guo, Mauro Carvalho Chehab,
Shuai Xue, Len Brown, Ira Weiny, Li Ming, Shuah Khan,
Ben Cheatham, Richard Cheng, Robert Richter, Lukas Wunner,
linux-pci, linux-acpi, linux-doc, linux-kernel
Establish a single CXL protocol error path shared by CXL Virtual
Hierarchy (VH) and Restricted CXL Host (RCH) topologies. AER dispatch in
handle_error_source() routes CXL protocol errors, gated by
is_cxl_error(), through the AER-CXL kfifo to a cxl_core consumer for
logging and recovery. Producer and consumer go live together so no CXL
error is silently dropped across a bisect.
is_cxl_error() expands from Endpoint-only to also cover Root Port,
Upstream Port, and Downstream Port. RCDs report on behalf of an upstream
RCH Downstream Port and instead reach the kfifo via
cxl_rch_handle_error().
For uncorrectable errors, cxl_proto_err_wait_for_empty() drains the CXL
plane (RAS read, panic policy, state clear) before pci_aer_handle_error()
drives PCIe recovery, so recovery does not tear down RAS iomaps while the
consumer is still reading them. Correctable errors run asynchronously.
Panic policy: cxl_do_recovery() panics on a confirmed UCE, and also when
the RAS registers cannot be mapped -- an unconfirmable UCE is treated
conservatively as fatal since CXL.mem coherency may be lost. A
mapped-but-clear status is logged as spurious with no panic.
to_ras_base() centralizes RAS base lookup (dport->regs.ras for
Root/Downstream Ports, port->regs.ras otherwise) and provides an
injection point for RAS status simulation during testing. The
cxl_cor_error_detected() AER callback is removed; correctable Endpoint
errors now route through the kfifo like every other CXL protocol error.
Update cxl_handle_rdport_errors() with locking to prevent dport from
being freed and RAS from being unmapped.
At this step cxl_handle_rdport_errors() still dispatches a single
severity per pass (matching the pre-series baseline). The following
patch, "cxl/ras: Handle RCH correctable and uncorrectable errors in one
pass", processes a simultaneously signalled CE and UCE together.
Co-developed-by: Dan Williams <djbw@kernel.org>
Signed-off-by: Dan Williams <djbw@kernel.org>
Signed-off-by: Terry Bowman <terry.bowman@amd.com>
---
Changes in v19->v20:
- Condense commit message (Jonathan)
- Document simultaneous RCH CE+UCE handling in next patch
- Document the cxl_handle_rdport() dport lock
Changes in v18->v19:
- Merge "Establish common CXL Port protocol error flow" and "PCI/CXL: Add
RCH support to CXL handlers" into a single patch; route RCD correctable
errors through the AER-CXL kfifo and remove cxl_cor_error_detected().
- Rename cxl_proto_err_flush() to cxl_proto_err_wait_for_empty()
- Use guard(device) in __cxl_proto_err_work_fn() instead of scoped_guard
- Reword "silent data corruption" comment; drop "interleaved HDM
regions"
- Add blank line after the port lookup in __cxl_proto_err_work_fn()
- Change cxl_do_recovery() to panic if to_ras_base() returns NULL
- Clarify Port driver bound gates RAS access
- Add review-by for DaveJ
Changes in v17->v18:
- Fix pre-existing race: hold memdev device lock around
cxl_handle_rdport_errors(), release before port lock
- Fix handle_error_source() to call pci_aer_handle_error() unconditionally
so AER handling always runs after cxl_forward_error()
- Add cxl_proto_err_flush() call for CXL UCE to drain kfifo before AER
recovery tears down the device
- Fix NULL dereference of dport->dport_dev in cxl_handle_cor_ras() and
cxl_handle_ras() for UPSTREAM/ENDPOINT port types: use dport->dport_dev
when dport is non-NULL, else fall back to port->uport_dev
- Remove duplicate pcie_clear_device_status() call from
cxl_handle_proto_error() CE path; pci_aer_handle_error() already clears it
- Clarify panic policy: panic only on confirmed UCE via RAS status read
- Document kfifo consumer serialization against driver unbind via
guard(device)(&port->dev) and port->dev.driver check
Changes in v16->v17:
- get_cxl_port() -> find_cxl_port_by_dev()
- Simplified find_cxl_port_by_dev()
- Replace and remove cxl_serial_number() w/ pci_get_dsn()
- cxl_get_ras_base() -> to_ras_base()
- Drop dependency on PCI_ERS_RESULT_PANIC; cxl_do_recovery() panics
directly. (PANIC enum patch dropped from series.)
- Clarify panic semantics: panic on any uncorrectable CXL RAS error, not
only AER-FATAL severities.
- Add is_cxl_error() switch in handle_error_source() here, paired with the
kfifo consumer registration, to keep each commit bisect-safe.
- Drop pcie_aer_is_native() guard in cxl_do_recovery() (always native).
- Swap order with the "Limit" patch for bisectability w/ cxl_ras_exit()
- Reword for "any uncorrectable" CXL RAS error panics.
- Restore log messages for port-not-found and port-unbound cases.
- Whitespace cleanup (Jonathan)
- Update to get_cxl_port() documentation (Terry)
- Fix __cxl_proto_err_work_fn() to return 0 for transient errors.
- Drop !port check in cxl_do_recovery(), caller already validated
- Fix kerneldoc @pdev -> @dev in find_cxl_port_by_dev()
- Fix missing space in pr_err_ratelimited()
- Made pcie_clear_device_status() and pci_aer_clear_fatal_status()
EXPORT_SYMBOL_FOR_MODULES("cxl_core") (Dan)
- Move find_cxl_port_by_dport() and find_cxl_port_by_uport()
de-staticisation and core.h declarations from the rename patch to
here, where the first cross-file callers in find_cxl_port_by_dev()
land.
Changes in v15->v16:
- get_ras_base(), initialize dport to NULL (Jonathan)
- Remove guard(device)(&cxlmd->dev) (Jonathan)
- Fix dev_warns() (Jonathan)
- Remove comment in cxl_port_error_detected() (Dan)
- Update switch-case brackets to follow clang-format (Dan)
- Add PCI_EXP_TYPE_RC_END for cxl_get_ras_base() (Terry)
- Add NULL port check in cxl_serial_number() (Terry)
Changes in v14->v15:
- Update commit message and title. Added Bjorn's ack.
- Move CE and UCE handling logic here
Changes in v13->v14:
- Add Dave Jiang's review-by
- Update commit message & headline (Bjorn)
- Refactor cxl_port_error_detected()/cxl_port_cor_error_detected() to
one line (Jonathan)
- Remove cxl_walk_port() (Dan)
- Remove cxl_pci_drv_bound(). Check for 'is_cxl' parent port is
sufficient (Dan)
- Remove device_lock_if()
- Combined CE and UCE here (Terry)
Changes in v12->v13:
- Move get_pci_cxl_host_dev() and cxl_handle_proto_error() to Dequeue
patch (Terry)
- Remove EP case in cxl_get_ras_base(), not used. (Terry)
- Remove check for dport->dport_dev (Dave)
- Remove whitespace (Terry)
Changes in v11->v12:
- Add call to cxl_pci_drv_bound() in cxl_handle_proto_error() and
pci_to_cxl_dev()
- Change cxl_error_detected() -> cxl_cor_error_detected()
- Remove NULL variable assignments
- Replace bus_find_device() with find_cxl_port_by_uport() for upstream
port searches.
Changes in v10->v11:
- None
---
drivers/cxl/core/core.h | 17 ++-
drivers/cxl/core/port.c | 6 +-
drivers/cxl/core/ras.c | 225 +++++++++++++++++++++++++--------
drivers/cxl/core/ras_rch.c | 22 +++-
drivers/cxl/cxlpci.h | 3 -
drivers/cxl/pci.c | 1 -
drivers/pci/pci.h | 1 -
drivers/pci/pcie/aer.c | 13 +-
drivers/pci/pcie/aer_cxl_rch.c | 39 +++---
drivers/pci/pcie/aer_cxl_vh.c | 16 ++-
drivers/pci/pcie/portdrv.h | 4 +-
11 files changed, 249 insertions(+), 98 deletions(-)
diff --git a/drivers/cxl/core/core.h b/drivers/cxl/core/core.h
index 35eaf636adc9c..645824167f788 100644
--- a/drivers/cxl/core/core.h
+++ b/drivers/cxl/core/core.h
@@ -188,10 +188,13 @@ static inline struct device *dport_to_host(struct cxl_dport *dport)
void cxl_ras_init(void);
void cxl_ras_exit(void);
bool cxl_handle_ras(struct device *dev, void __iomem *ras_base);
+void cxl_do_recovery(struct pci_dev *pdev, struct cxl_port *port,
+ struct cxl_dport *dport);
void cxl_handle_cor_ras(struct device *dev, void __iomem *ras_base);
void cxl_dport_map_rch_aer(struct cxl_dport *dport);
void cxl_disable_rch_root_ints(struct cxl_dport *dport);
-void cxl_handle_rdport_errors(struct cxl_dev_state *cxlds);
+void cxl_handle_rdport_errors(struct pci_dev *pdev);
+void __iomem *to_ras_base(struct cxl_port *port, struct cxl_dport *dport);
void devm_cxl_dport_ras_setup(struct cxl_dport *dport);
#else
static inline void cxl_ras_init(void) { }
@@ -200,14 +203,24 @@ static inline bool cxl_handle_ras(struct device *dev, void __iomem *ras_base)
{
return false;
}
+static inline void cxl_do_recovery(struct pci_dev *pdev, struct cxl_port *port,
+ struct cxl_dport *dport) { }
static inline void cxl_handle_cor_ras(struct device *dev, void __iomem *ras_base) { }
static inline void cxl_dport_map_rch_aer(struct cxl_dport *dport) { }
static inline void cxl_disable_rch_root_ints(struct cxl_dport *dport) { }
-static inline void cxl_handle_rdport_errors(struct cxl_dev_state *cxlds) { }
+static inline void cxl_handle_rdport_errors(struct pci_dev *pdev) { }
+static inline void __iomem *to_ras_base(struct cxl_port *port,
+ struct cxl_dport *dport)
+{
+ return NULL;
+}
static inline void devm_cxl_dport_ras_setup(struct cxl_dport *dport) { }
#endif /* CONFIG_CXL_RAS */
int cxl_gpf_port_setup(struct cxl_dport *dport);
+struct cxl_port *find_cxl_port_by_dport(struct device *dport_dev,
+ struct cxl_dport **dport);
+struct cxl_port *find_cxl_port_by_uport(struct device *uport_dev);
struct cxl_hdm;
int cxl_hdm_decode_init(struct cxl_dev_state *cxlds, struct cxl_hdm *cxlhdm,
diff --git a/drivers/cxl/core/port.c b/drivers/cxl/core/port.c
index 625e4aa427db0..9a746be6f1967 100644
--- a/drivers/cxl/core/port.c
+++ b/drivers/cxl/core/port.c
@@ -1400,8 +1400,8 @@ static struct cxl_port *__find_cxl_port_by_dport(struct cxl_find_port_ctx *ctx)
* Return a 'struct cxl_port' with an elevated reference if found. Use
* __free(put_cxl_port) to release.
*/
-static struct cxl_port *find_cxl_port_by_dport(struct device *dport_dev,
- struct cxl_dport **dport)
+struct cxl_port *find_cxl_port_by_dport(struct device *dport_dev,
+ struct cxl_dport **dport)
{
struct cxl_find_port_ctx ctx = {
.dport_dev = dport_dev,
@@ -1596,7 +1596,7 @@ static int match_port_by_uport(struct device *dev, const void *data)
* Function takes a device reference on the port device. Caller should do a
* put_device() when done.
*/
-static struct cxl_port *find_cxl_port_by_uport(struct device *uport_dev)
+struct cxl_port *find_cxl_port_by_uport(struct device *uport_dev)
{
struct device *dev;
diff --git a/drivers/cxl/core/ras.c b/drivers/cxl/core/ras.c
index e307361bb39e4..02c791f149270 100644
--- a/drivers/cxl/core/ras.c
+++ b/drivers/cxl/core/ras.c
@@ -77,6 +77,35 @@ static int match_memdev_by_parent(struct device *dev, const void *uport)
return 0;
}
+/**
+ * find_cxl_port_by_dev - Use @dev as hint to do a _by_dport or _by_uport lookup
+ * @dev: generic device that may either be a companion of port or target dport
+ * @dport: optional output; if non-NULL, set to the matched dport for
+ * Root Port and Downstream Port lookups, NULL for all other types.
+ *
+ * Return a 'struct cxl_port' with an elevated reference if found. Use
+ * __free(put_cxl_port) to release.
+ */
+static struct cxl_port *find_cxl_port_by_dev(struct device *dev, struct cxl_dport **dport)
+{
+ if (dport)
+ *dport = NULL;
+ if (!dev_is_pci(dev))
+ return NULL;
+
+ switch (pci_pcie_type(to_pci_dev(dev))) {
+ case PCI_EXP_TYPE_ROOT_PORT:
+ case PCI_EXP_TYPE_DOWNSTREAM:
+ return find_cxl_port_by_dport(dev, dport);
+ case PCI_EXP_TYPE_UPSTREAM:
+ case PCI_EXP_TYPE_ENDPOINT:
+ case PCI_EXP_TYPE_RC_END:
+ return find_cxl_port_by_uport(dev);
+ }
+
+ return NULL;
+}
+
void cxl_cper_handle_prot_err(struct cxl_cper_prot_err_work_data *data)
{
unsigned int devfn = PCI_DEVFN(data->prot_err.agent_addr.device,
@@ -129,16 +158,6 @@ static void cxl_cper_prot_err_work_fn(struct work_struct *work)
}
static DECLARE_WORK(cxl_cper_prot_err_work, cxl_cper_prot_err_work_fn);
-void cxl_ras_init(void)
-{
- cxl_cper_register_prot_err_work(&cxl_cper_prot_err_work);
-}
-
-void cxl_ras_exit(void)
-{
- cxl_cper_unregister_prot_err_work();
-}
-
static void cxl_dport_map_ras(struct cxl_dport *dport)
{
struct cxl_register_map *map = &dport->reg_map;
@@ -195,6 +214,32 @@ void devm_cxl_port_ras_setup(struct cxl_port *port)
}
EXPORT_SYMBOL_NS_GPL(devm_cxl_port_ras_setup, "CXL");
+void __iomem *to_ras_base(struct cxl_port *port, struct cxl_dport *dport)
+{
+ if (!port)
+ return NULL;
+
+ if (dport)
+ return dport->regs.ras;
+
+ return port->regs.ras;
+}
+
+void cxl_do_recovery(struct pci_dev *pdev, struct cxl_port *port, struct cxl_dport *dport)
+{
+ struct device *dev = dport ? dport->dport_dev : port->uport_dev;
+ void __iomem *ras_base = to_ras_base(port, dport);
+
+ if (!ras_base)
+ panic("CXL: UCE with unmapped RAS registers");
+
+ if (cxl_handle_ras(dev, ras_base))
+ panic("CXL cachemem error");
+
+ dev_dbg(&pdev->dev,
+ "CXL UCE signaled but no CXL RAS status bits set\n");
+}
+
void cxl_handle_cor_ras(struct device *dev, void __iomem *ras_base)
{
void __iomem *addr;
@@ -207,7 +252,10 @@ void cxl_handle_cor_ras(struct device *dev, void __iomem *ras_base)
status = readl(addr);
if (status & CXL_RAS_CORRECTABLE_STATUS_MASK) {
writel(status & CXL_RAS_CORRECTABLE_STATUS_MASK, addr);
- trace_cxl_aer_correctable_error(to_cxl_memdev(dev), status);
+ if (is_cxl_memdev(dev))
+ trace_cxl_aer_correctable_error(to_cxl_memdev(dev), status);
+ else
+ trace_cxl_port_aer_correctable_error(dev, status);
}
}
@@ -259,73 +307,61 @@ bool cxl_handle_ras(struct device *dev, void __iomem *ras_base)
}
header_log_copy(ras_base, hl);
- trace_cxl_aer_uncorrectable_error(to_cxl_memdev(dev), status, fe, hl);
+ if (is_cxl_memdev(dev))
+ trace_cxl_aer_uncorrectable_error(to_cxl_memdev(dev), status, fe, hl);
+ else
+ trace_cxl_port_aer_uncorrectable_error(dev, status, fe, hl);
+
writel(status & CXL_RAS_UNCORRECTABLE_STATUS_MASK, addr);
return true;
}
-void cxl_cor_error_detected(struct pci_dev *pdev)
-{
- struct cxl_dev_state *cxlds = pci_get_drvdata(pdev);
- struct cxl_memdev *cxlmd = cxlds->cxlmd;
- struct device *dev = &cxlds->cxlmd->dev;
-
- scoped_guard(device, dev) {
- if (!dev->driver) {
- dev_warn(&pdev->dev,
- "%s: memdev disabled, abort error handling\n",
- dev_name(dev));
- return;
- }
-
- if (cxlds->rcd)
- cxl_handle_rdport_errors(cxlds);
-
- cxl_handle_cor_ras(&cxlds->cxlmd->dev, cxlmd->endpoint->regs.ras);
- }
-}
-EXPORT_SYMBOL_NS_GPL(cxl_cor_error_detected, "CXL");
-
pci_ers_result_t cxl_error_detected(struct pci_dev *pdev,
pci_channel_state_t state)
{
- struct cxl_dev_state *cxlds = pci_get_drvdata(pdev);
- struct cxl_memdev *cxlmd = cxlds->cxlmd;
- struct device *dev = &cxlmd->dev;
- bool ue;
+ struct cxl_port *port __free(put_cxl_port) = find_cxl_port_by_uport(&pdev->dev);
+ bool ue = false;
+
+ if (!port)
+ return PCI_ERS_RESULT_DISCONNECT;
+
+ if (is_cxl_restricted(pdev))
+ cxl_handle_rdport_errors(pdev);
- scoped_guard(device, dev) {
- if (!dev->driver) {
+ scoped_guard(device, &port->dev) {
+ if (!port->dev.driver) {
dev_warn(&pdev->dev,
- "%s: memdev disabled, abort error handling\n",
- dev_name(dev));
+ "%s: port disabled, abort error handling\n",
+ dev_name(&port->dev));
return PCI_ERS_RESULT_DISCONNECT;
}
- if (cxlds->rcd)
- cxl_handle_rdport_errors(cxlds);
/*
- * A frozen channel indicates an impending reset which is fatal to
- * CXL.mem operation, and will likely crash the system. On the off
- * chance the situation is recoverable dump the status of the RAS
- * capability registers and bounce the active state of the memdev.
+ * The CXL RAS read is unconditional regardless of channel
+ * state. Any uncorrectable error bit set in the CXL RAS
+ * status register triggers a panic below because CXL.mem
+ * cache coherency is already lost; continuing risks silent
+ * data corruption.
*/
- ue = cxl_handle_ras(&cxlds->cxlmd->dev, cxlmd->endpoint->regs.ras);
+ ue = cxl_handle_ras(port->uport_dev, to_ras_base(port, NULL));
}
+ /*
+ * CXL.mem UCE means cache coherency is lost. Continuing risks
+ * silent data corruption.
+ */
+ if (ue)
+ panic("CXL cachemem error");
+
switch (state) {
case pci_channel_io_normal:
- if (ue) {
- device_release_driver(dev);
- return PCI_ERS_RESULT_NEED_RESET;
- }
return PCI_ERS_RESULT_CAN_RECOVER;
case pci_channel_io_frozen:
dev_warn(&pdev->dev,
"%s: frozen state error detected, disable CXL.mem\n",
- dev_name(dev));
- device_release_driver(dev);
+ dev_name(port->uport_dev));
+ device_release_driver(port->uport_dev);
return PCI_ERS_RESULT_NEED_RESET;
case pci_channel_io_perm_failure:
dev_warn(&pdev->dev,
@@ -335,3 +371,82 @@ pci_ers_result_t cxl_error_detected(struct pci_dev *pdev,
return PCI_ERS_RESULT_NEED_RESET;
}
EXPORT_SYMBOL_NS_GPL(cxl_error_detected, "CXL");
+
+static void cxl_handle_proto_error(struct pci_dev *pdev, struct cxl_port *port,
+ struct cxl_dport *dport, int severity)
+{
+ struct device *dev = dport ? dport->dport_dev : port->uport_dev;
+
+ if (severity == AER_CORRECTABLE)
+ cxl_handle_cor_ras(dev, to_ras_base(port, dport));
+ else
+ cxl_do_recovery(pdev, port, dport);
+}
+
+static void __cxl_proto_err_work_fn(struct cxl_proto_err_work_data *wd)
+{
+ struct cxl_dport *dport;
+ struct device *host;
+
+ /*
+ * For RCD devices, handle RCH Downstream Port errors first.
+ * cxl_handle_rdport_errors() does its own port lookup and locking,
+ * keeping the Downstream Port lock separate from the Endpoint Port
+ * lock taken below.
+ */
+ if (is_cxl_restricted(wd->pdev))
+ cxl_handle_rdport_errors(wd->pdev);
+
+ struct cxl_port *port __free(put_cxl_port) = find_cxl_port_by_dev(&wd->pdev->dev, NULL);
+ if (!port) {
+ dev_err_ratelimited(&wd->pdev->dev,
+ "Failed to find parent port device in CXL topology\n");
+ return;
+ }
+
+ host = is_cxl_root(port) ? port->uport_dev : &port->dev;
+
+ guard(device)(host);
+ if (!host->driver) {
+ dev_err_ratelimited(host,
+ "Port host device is unbound, abort error handling\n");
+ return;
+ }
+
+ dport = cxl_find_dport_by_dev(port, &wd->pdev->dev);
+ if (!dport && (pci_pcie_type(wd->pdev) == PCI_EXP_TYPE_ROOT_PORT ||
+ pci_pcie_type(wd->pdev) == PCI_EXP_TYPE_DOWNSTREAM)) {
+ dev_err_ratelimited(&wd->pdev->dev,
+ "Failed to find dport device in CXL topology\n");
+ return;
+ }
+
+ cxl_handle_proto_error(wd->pdev, port, dport, wd->severity);
+}
+
+static void cxl_proto_err_work_fn(struct work_struct *work)
+{
+ struct cxl_proto_err_work_data wd;
+
+ for_each_cxl_proto_err(&wd, __cxl_proto_err_work_fn);
+}
+
+static DECLARE_WORK(cxl_proto_err_work, cxl_proto_err_work_fn);
+
+static void cxl_proto_err_do_flush(void)
+{
+ flush_work(&cxl_proto_err_work);
+}
+
+void cxl_ras_init(void)
+{
+ cxl_cper_register_prot_err_work(&cxl_cper_prot_err_work);
+ cxl_register_proto_err_work(&cxl_proto_err_work,
+ cxl_proto_err_do_flush);
+}
+
+void cxl_ras_exit(void)
+{
+ cxl_unregister_proto_err_work();
+ cxl_cper_unregister_prot_err_work();
+}
diff --git a/drivers/cxl/core/ras_rch.c b/drivers/cxl/core/ras_rch.c
index e0e01aa5eba6c..ddaa3d7781678 100644
--- a/drivers/cxl/core/ras_rch.c
+++ b/drivers/cxl/core/ras_rch.c
@@ -1,7 +1,6 @@
// SPDX-License-Identifier: GPL-2.0-only
/* Copyright(c) 2025 AMD Corporation. All rights reserved. */
-#include <linux/types.h>
#include <linux/aer.h>
#include "cxl.h"
#include "core.h"
@@ -110,18 +109,27 @@ static bool cxl_rch_get_aer_severity(struct aer_capability_regs *aer_regs,
return false;
}
-void cxl_handle_rdport_errors(struct cxl_dev_state *cxlds)
+void cxl_handle_rdport_errors(struct pci_dev *pdev)
{
- struct pci_dev *pdev = to_pci_dev(cxlds->dev);
struct aer_capability_regs aer_regs;
struct cxl_dport *dport;
int severity;
- struct cxl_port *port __free(put_cxl_port) =
- cxl_pci_find_port(pdev, &dport);
+ struct cxl_port *port __free(put_cxl_port) = cxl_pci_find_port(pdev, NULL);
if (!port)
return;
+ /*
+ * The RCH Downstream Port is the Root Port's dport
+ * (dport free and RAS iomap) is hosted on the CXL Host Bridge
+ * (port->uport_dev), not &port->dev. Hold that device's lock so the
+ * dport cannot be freed and its registers unmapped while in use here.
+ */
+ guard(device)(port->uport_dev);
+ dport = cxl_find_dport_by_dev(port, pdev->dev.parent);
+ if (!dport)
+ return;
+
if (!cxl_rch_get_aer_info(dport->regs.dport_aer, &aer_regs))
return;
@@ -130,7 +138,7 @@ void cxl_handle_rdport_errors(struct cxl_dev_state *cxlds)
pci_print_aer(pdev, severity, &aer_regs);
if (severity == AER_CORRECTABLE)
- cxl_handle_cor_ras(&cxlds->cxlmd->dev, dport->regs.ras);
+ cxl_handle_cor_ras(dport->dport_dev, to_ras_base(port, dport));
else
- cxl_handle_ras(&cxlds->cxlmd->dev, dport->regs.ras);
+ cxl_do_recovery(pdev, dport->port, dport);
}
diff --git a/drivers/cxl/cxlpci.h b/drivers/cxl/cxlpci.h
index 110ec9c44f09f..606fadd2476f3 100644
--- a/drivers/cxl/cxlpci.h
+++ b/drivers/cxl/cxlpci.h
@@ -79,14 +79,11 @@ struct cxl_dev_state;
void read_cdat_data(struct cxl_port *port);
#ifdef CONFIG_CXL_RAS
-void cxl_cor_error_detected(struct pci_dev *pdev);
pci_ers_result_t cxl_error_detected(struct pci_dev *pdev,
pci_channel_state_t state);
void devm_cxl_dport_rch_ras_setup(struct cxl_dport *dport);
void devm_cxl_port_ras_setup(struct cxl_port *port);
#else
-static inline void cxl_cor_error_detected(struct pci_dev *pdev) { }
-
static inline pci_ers_result_t cxl_error_detected(struct pci_dev *pdev,
pci_channel_state_t state)
{
diff --git a/drivers/cxl/pci.c b/drivers/cxl/pci.c
index c7c91e8dc51dc..fe11be7fd9aab 100644
--- a/drivers/cxl/pci.c
+++ b/drivers/cxl/pci.c
@@ -1002,7 +1002,6 @@ static const struct pci_error_handlers cxl_error_handlers = {
.error_detected = cxl_error_detected,
.slot_reset = cxl_slot_reset,
.resume = cxl_error_resume,
- .cor_error_detected = cxl_cor_error_detected,
.reset_done = cxl_reset_done,
};
diff --git a/drivers/pci/pci.h b/drivers/pci/pci.h
index ba3c3fddddc23..6c7decfb171b0 100644
--- a/drivers/pci/pci.h
+++ b/drivers/pci/pci.h
@@ -1344,7 +1344,6 @@ void pci_restore_aer_state(struct pci_dev *dev);
static inline void pci_no_aer(void) { }
static inline void pci_aer_init(struct pci_dev *d) { }
static inline void pci_aer_exit(struct pci_dev *d) { }
-static inline void pci_aer_clear_fatal_status(struct pci_dev *dev) { }
static inline int pci_aer_clear_status(struct pci_dev *dev) { return -EINVAL; }
static inline int pci_aer_raw_clear_status(struct pci_dev *dev) { return -EINVAL; }
static inline void pci_save_aer_state(struct pci_dev *dev) { }
diff --git a/drivers/pci/pcie/aer.c b/drivers/pci/pcie/aer.c
index 21dfc9c933d77..8c998cffa89e4 100644
--- a/drivers/pci/pcie/aer.c
+++ b/drivers/pci/pcie/aer.c
@@ -1328,7 +1328,18 @@ static void pci_aer_handle_error(struct pci_dev *dev, struct aer_err_info *info)
static void handle_error_source(struct pci_dev *dev, struct aer_err_info *info)
{
- cxl_rch_handle_error(dev, info);
+ bool cxl_pending = cxl_rch_handle_error(dev, info);
+
+ if (is_cxl_error(dev, info))
+ cxl_pending |= cxl_forward_error(dev, info);
+
+ /*
+ * Wait for UCE CXL work to complete before AER recovery
+ * tears down the device. CE can run asynchronously.
+ */
+ if (cxl_pending && info->severity != AER_CORRECTABLE)
+ cxl_proto_err_wait_for_empty();
+
pci_aer_handle_error(dev, info);
pci_dev_put(dev);
}
diff --git a/drivers/pci/pcie/aer_cxl_rch.c b/drivers/pci/pcie/aer_cxl_rch.c
index e471eefec9c40..ab31c4281b483 100644
--- a/drivers/pci/pcie/aer_cxl_rch.c
+++ b/drivers/pci/pcie/aer_cxl_rch.c
@@ -34,42 +34,37 @@ static bool cxl_error_is_native(struct pci_dev *dev)
return (pcie_ports_native || host->native_aer);
}
+struct cxl_rch_error_ctx {
+ struct aer_err_info *info;
+ bool enqueued;
+};
+
static int cxl_rch_handle_error_iter(struct pci_dev *dev, void *data)
{
- struct aer_err_info *info = (struct aer_err_info *)data;
- const struct pci_error_handlers *err_handler;
+ struct cxl_rch_error_ctx *ctx = data;
if (!is_cxl_mem_dev(dev) || !cxl_error_is_native(dev))
return 0;
- guard(device)(&dev->dev);
-
- err_handler = dev->driver ? dev->driver->err_handler : NULL;
- if (!err_handler)
- return 0;
-
- if (info->severity == AER_CORRECTABLE) {
- if (err_handler->cor_error_detected)
- err_handler->cor_error_detected(dev);
- } else if (err_handler->error_detected) {
- if (info->severity == AER_NONFATAL)
- err_handler->error_detected(dev, pci_channel_io_normal);
- else if (info->severity == AER_FATAL)
- err_handler->error_detected(dev, pci_channel_io_frozen);
- }
+ if (cxl_forward_error(dev, ctx->info))
+ ctx->enqueued = true;
return 0;
}
-void cxl_rch_handle_error(struct pci_dev *dev, struct aer_err_info *info)
+bool cxl_rch_handle_error(struct pci_dev *dev, struct aer_err_info *info)
{
+ struct cxl_rch_error_ctx ctx = { .info = info };
+
/*
- * Internal errors of an RCEC indicate an AER error in an
- * RCH's downstream port. Check and handle them in the CXL.mem
- * device driver.
+ * An RCEC AER internal error indicates an error in an
+ * associated RCH Downstream Port or RCD device or both.
+ * Forward to the cxl_core module for handling.
*/
if (pci_pcie_type(dev) == PCI_EXP_TYPE_RC_EC &&
is_aer_internal_error(info))
- pcie_walk_rcec(dev, cxl_rch_handle_error_iter, info);
+ pcie_walk_rcec(dev, cxl_rch_handle_error_iter, &ctx);
+
+ return ctx.enqueued;
}
static int handles_cxl_error_iter(struct pci_dev *dev, void *data)
diff --git a/drivers/pci/pcie/aer_cxl_vh.c b/drivers/pci/pcie/aer_cxl_vh.c
index 9fc12d4e644bc..d04296c01e446 100644
--- a/drivers/pci/pcie/aer_cxl_vh.c
+++ b/drivers/pci/pcie/aer_cxl_vh.c
@@ -54,8 +54,22 @@ bool is_cxl_error(struct pci_dev *pdev, struct aer_err_info *info)
if (!info || !info->is_cxl)
return false;
- if (pci_pcie_type(pdev) != PCI_EXP_TYPE_ENDPOINT)
+ /*
+ * RCD (PCI_EXP_TYPE_RC_END) is not included here because RCDs
+ * report errors on behalf of upstream RCH Downstream Port and thus
+ * require a unique discovery detailed in CXL4.0 spec (12.2.1.1).
+ * The RCH device error discovery and RCD forwarding flow begins
+ * in cxl_rch_handle_error().
+ */
+ switch (pci_pcie_type(pdev)) {
+ case PCI_EXP_TYPE_ENDPOINT:
+ case PCI_EXP_TYPE_ROOT_PORT:
+ case PCI_EXP_TYPE_UPSTREAM:
+ case PCI_EXP_TYPE_DOWNSTREAM:
+ break;
+ default:
return false;
+ }
return is_aer_internal_error(info);
}
diff --git a/drivers/pci/pcie/portdrv.h b/drivers/pci/pcie/portdrv.h
index 357310916088f..b13b9ac9571dc 100644
--- a/drivers/pci/pcie/portdrv.h
+++ b/drivers/pci/pcie/portdrv.h
@@ -128,14 +128,14 @@ struct aer_err_info;
#ifdef CONFIG_CXL_RAS
bool is_aer_internal_error(struct aer_err_info *info);
-void cxl_rch_handle_error(struct pci_dev *dev, struct aer_err_info *info);
+bool cxl_rch_handle_error(struct pci_dev *dev, struct aer_err_info *info);
void cxl_rch_enable_rcec(struct pci_dev *rcec);
bool is_cxl_error(struct pci_dev *pdev, struct aer_err_info *info);
bool cxl_forward_error(struct pci_dev *pdev, struct aer_err_info *info);
void cxl_proto_err_wait_for_empty(void);
#else
static inline bool is_aer_internal_error(struct aer_err_info *info) { return false; }
-static inline void cxl_rch_handle_error(struct pci_dev *dev, struct aer_err_info *info) { }
+static inline bool cxl_rch_handle_error(struct pci_dev *dev, struct aer_err_info *info) { return false; }
static inline void cxl_rch_enable_rcec(struct pci_dev *rcec) { }
static inline bool is_cxl_error(struct pci_dev *pdev, struct aer_err_info *info) { return false; }
static inline bool cxl_forward_error(struct pci_dev *pdev, struct aer_err_info *info) { return false; }
--
2.34.1
^ permalink raw reply related [flat|nested] 17+ messages in thread* Re: [PATCH v20 2/9] PCI: Establish common CXL Port protocol error flow
2026-09-02 13:39 ` [PATCH v20 2/9] PCI: Establish common CXL Port protocol error flow Terry Bowman
@ 2026-09-02 20:57 ` Cheatham, Benjamin
0 siblings, 0 replies; 17+ messages in thread
From: Cheatham, Benjamin @ 2026-09-02 20:57 UTC (permalink / raw)
To: Terry Bowman, Jonathan Cameron, Dave Jiang, Alison Schofield,
Vishal Verma, Davidlohr Bueso, Bjorn Helgaas, Dan Williams,
Rafael J . Wysocki, Jonathan Corbet, linux-cxl
Cc: Tony Luck, Borislav Petkov, Hanjun Guo, Mauro Carvalho Chehab,
Shuai Xue, Len Brown, Ira Weiny, Li Ming, Shuah Khan,
Richard Cheng, Robert Richter, Lukas Wunner, linux-pci,
linux-acpi, linux-doc, linux-kernel
On 9/2/2026 8:39 AM, Terry Bowman wrote:
> Establish a single CXL protocol error path shared by CXL Virtual
> Hierarchy (VH) and Restricted CXL Host (RCH) topologies. AER dispatch in
> handle_error_source() routes CXL protocol errors, gated by
> is_cxl_error(), through the AER-CXL kfifo to a cxl_core consumer for
> logging and recovery. Producer and consumer go live together so no CXL
> error is silently dropped across a bisect.
>
> is_cxl_error() expands from Endpoint-only to also cover Root Port,
> Upstream Port, and Downstream Port. RCDs report on behalf of an upstream
> RCH Downstream Port and instead reach the kfifo via
> cxl_rch_handle_error().
>
> For uncorrectable errors, cxl_proto_err_wait_for_empty() drains the CXL
> plane (RAS read, panic policy, state clear) before pci_aer_handle_error()
> drives PCIe recovery, so recovery does not tear down RAS iomaps while the
> consumer is still reading them. Correctable errors run asynchronously.
>
> Panic policy: cxl_do_recovery() panics on a confirmed UCE, and also when
> the RAS registers cannot be mapped -- an unconfirmable UCE is treated
> conservatively as fatal since CXL.mem coherency may be lost. A
> mapped-but-clear status is logged as spurious with no panic.
>
> to_ras_base() centralizes RAS base lookup (dport->regs.ras for
> Root/Downstream Ports, port->regs.ras otherwise) and provides an
> injection point for RAS status simulation during testing. The
> cxl_cor_error_detected() AER callback is removed; correctable Endpoint
> errors now route through the kfifo like every other CXL protocol error.
>
> Update cxl_handle_rdport_errors() with locking to prevent dport from
> being freed and RAS from being unmapped.
>
> At this step cxl_handle_rdport_errors() still dispatches a single
> severity per pass (matching the pre-series baseline). The following
> patch, "cxl/ras: Handle RCH correctable and uncorrectable errors in one
> pass", processes a simultaneously signalled CE and UCE together.
>
> Co-developed-by: Dan Williams <djbw@kernel.org>
> Signed-off-by: Dan Williams <djbw@kernel.org>
> Signed-off-by: Terry Bowman <terry.bowman@amd.com>
>
> ---
>
One small nit, but otherwise LGTM:
Reviewed-by: Ben Cheatham <benjamin.cheatham@amd.com>
...
> pci_ers_result_t cxl_error_detected(struct pci_dev *pdev,
> pci_channel_state_t state)
> {
> - struct cxl_dev_state *cxlds = pci_get_drvdata(pdev);
> - struct cxl_memdev *cxlmd = cxlds->cxlmd;
> - struct device *dev = &cxlmd->dev;
> - bool ue;
> + struct cxl_port *port __free(put_cxl_port) = find_cxl_port_by_uport(&pdev->dev);
> + bool ue = false;
> +
> + if (!port)
> + return PCI_ERS_RESULT_DISCONNECT;
> +
> + if (is_cxl_restricted(pdev))
> + cxl_handle_rdport_errors(pdev);
>
> - scoped_guard(device, dev) {
> - if (!dev->driver) {
> + scoped_guard(device, &port->dev) {
> + if (!port->dev.driver) {
> dev_warn(&pdev->dev,
> - "%s: memdev disabled, abort error handling\n",
> - dev_name(dev));
> + "%s: port disabled, abort error handling\n",
> + dev_name(&port->dev));
> return PCI_ERS_RESULT_DISCONNECT;
> }
>
> - if (cxlds->rcd)
> - cxl_handle_rdport_errors(cxlds);
> /*
> - * A frozen channel indicates an impending reset which is fatal to
> - * CXL.mem operation, and will likely crash the system. On the off
> - * chance the situation is recoverable dump the status of the RAS
> - * capability registers and bounce the active state of the memdev.
> + * The CXL RAS read is unconditional regardless of channel
> + * state. Any uncorrectable error bit set in the CXL RAS
> + * status register triggers a panic below because CXL.mem
> + * cache coherency is already lost; continuing risks silent
> + * data corruption.
> */
> - ue = cxl_handle_ras(&cxlds->cxlmd->dev, cxlmd->endpoint->regs.ras);
> + ue = cxl_handle_ras(port->uport_dev, to_ras_base(port, NULL));
> }
>
> + /*
> + * CXL.mem UCE means cache coherency is lost. Continuing risks
> + * silent data corruption.
> + */
Don't need this comment and the last sentence in the comment above.
> + if (ue)
> + panic("CXL cachemem error");
> +
> switch (state) {
> case pci_channel_io_normal:
> - if (ue) {
> - device_release_driver(dev);
> - return PCI_ERS_RESULT_NEED_RESET;
> - }
> return PCI_ERS_RESULT_CAN_RECOVER;
> case pci_channel_io_frozen:
> dev_warn(&pdev->dev,
> "%s: frozen state error detected, disable CXL.mem\n",
> - dev_name(dev));
> - device_release_driver(dev);
> + dev_name(port->uport_dev));
> + device_release_driver(port->uport_dev);
> return PCI_ERS_RESULT_NEED_RESET;
> case pci_channel_io_perm_failure:
> dev_warn(&pdev->dev,
> @@ -335,3 +371,82 @@ pci_ers_result_t cxl_error_detected(struct pci_dev *pdev,
> return PCI_ERS_RESULT_NEED_RESET;
> }
> EXPORT_SYMBOL_NS_GPL(cxl_error_detected, "CXL");
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH v20 3/9] cxl/ras: Handle RCH correctable and uncorrectable errors in one pass
2026-09-02 13:39 [PATCH v20 0/9] Enable CXL PCIe Port Protocol Error handling and logging Terry Bowman
2026-09-02 13:39 ` [PATCH v20 1/9] PCI/AER: Introduce AER-CXL protocol error kfifo Terry Bowman
2026-09-02 13:39 ` [PATCH v20 2/9] PCI: Establish common CXL Port protocol error flow Terry Bowman
@ 2026-09-02 13:39 ` Terry Bowman
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-02 13:39 ` [PATCH v20 4/9] cxl/pci: Thread port and dport through RAS handling helpers Terry Bowman
` (5 subsequent siblings)
8 siblings, 1 reply; 17+ messages in thread
From: Terry Bowman @ 2026-09-02 13:39 UTC (permalink / raw)
To: Jonathan Cameron, Dave Jiang, Alison Schofield, Vishal Verma,
Davidlohr Bueso, Bjorn Helgaas, Dan Williams, Rafael J . Wysocki,
Jonathan Corbet, linux-cxl
Cc: Tony Luck, Borislav Petkov, Hanjun Guo, Mauro Carvalho Chehab,
Shuai Xue, Len Brown, Ira Weiny, Li Ming, Shuah Khan,
Ben Cheatham, Richard Cheng, Robert Richter, Lukas Wunner,
linux-pci, linux-acpi, linux-doc, linux-kernel
cxl_rch_get_aer_info() reads and clears both the correctable and
uncorrectable AER status registers in a single pass. The previous
severity decode returned after the first matching class, so when a
correctable and an uncorrectable error were logged simultaneously the
correctable event was cleared in hardware but never traced or handled.
Handle both classes independently: dispatch cxl_handle_cor_ras() when
correctable status is set and cxl_do_recovery() when uncorrectable
status is set. Remove the now-unused cxl_rch_get_aer_severity() helper
and decode the uncorrectable severity inline.
Log the uncorrectable status unconditionally. __pci_print_aer() only
logs it via ANFE recursion when aer_compute_anfe_status() is non-zero.
A conditional skip could drop a fatal or non-ANFE record once the
hardware status is cleared.
Reported-by: Sashiko <sashiko-bot@kernel.org>
Link: https://lore.kernel.org/linux-cxl/20260803222923.517B11F00A3A@smtp.kernel.org/
Signed-off-by: Terry Bowman <terry.bowman@amd.com>
---
Changes in v19 -> v20:
- Add snapshot of aer_regs before the correctable pci_print_aer(). (Sashiko)
- Log the UCE status unconditionally instead of skipping on PCI_ERR_COR_ADV_NFAT
Changes in v18 -> v19:
- New patch to process correctable and uncorrectable RCH errors in the
same call so a co-logged correctable event is not lost.
---
drivers/cxl/core/ras_rch.c | 64 +++++++++++++++++++-------------------
1 file changed, 32 insertions(+), 32 deletions(-)
diff --git a/drivers/cxl/core/ras_rch.c b/drivers/cxl/core/ras_rch.c
index ddaa3d7781678..285eed5b828b9 100644
--- a/drivers/cxl/core/ras_rch.c
+++ b/drivers/cxl/core/ras_rch.c
@@ -89,42 +89,15 @@ static bool cxl_rch_get_aer_info(void __iomem *aer_base,
return true;
}
-/* Get AER severity. Return false if there is no error. */
-static bool cxl_rch_get_aer_severity(struct aer_capability_regs *aer_regs,
- int *severity)
-{
- u32 uncor_status = aer_regs->uncor_status & ~aer_regs->uncor_mask;
-
- if (uncor_status) {
- *severity = (uncor_status & aer_regs->uncor_severity) ?
- AER_FATAL : AER_NONFATAL;
- return true;
- }
-
- if (aer_regs->cor_status & ~aer_regs->cor_mask) {
- *severity = AER_CORRECTABLE;
- return true;
- }
-
- return false;
-}
-
void cxl_handle_rdport_errors(struct pci_dev *pdev)
{
struct aer_capability_regs aer_regs;
struct cxl_dport *dport;
- int severity;
struct cxl_port *port __free(put_cxl_port) = cxl_pci_find_port(pdev, NULL);
if (!port)
return;
- /*
- * The RCH Downstream Port is the Root Port's dport
- * (dport free and RAS iomap) is hosted on the CXL Host Bridge
- * (port->uport_dev), not &port->dev. Hold that device's lock so the
- * dport cannot be freed and its registers unmapped while in use here.
- */
guard(device)(port->uport_dev);
dport = cxl_find_dport_by_dev(port, pdev->dev.parent);
if (!dport)
@@ -133,12 +106,39 @@ void cxl_handle_rdport_errors(struct pci_dev *pdev)
if (!cxl_rch_get_aer_info(dport->regs.dport_aer, &aer_regs))
return;
- if (!cxl_rch_get_aer_severity(&aer_regs, &severity))
- return;
+ /*
+ * Snapshot aer_regs before handling the correctable error:
+ * pci_print_aer() takes it by pointer and rewrites uncor_status/
+ * uncor_mask in place on the Advisory Non-Fatal Error path. Use the
+ * copy for the uncorrectable dispatch and logging below so neither is
+ * corrupted by the correctable pci_print_aer() call.
+ */
+ struct aer_capability_regs uncor_regs = aer_regs;
+ u32 uncor_status = uncor_regs.uncor_status & ~uncor_regs.uncor_mask;
- pci_print_aer(pdev, severity, &aer_regs);
- if (severity == AER_CORRECTABLE)
+ /*
+ * Handle correctable and uncorrectable errors independently; both
+ * may be set in the same pass and cxl_rch_get_aer_info() has already
+ * cleared both status registers.
+ */
+ if (aer_regs.cor_status & ~aer_regs.cor_mask) {
+ pci_print_aer(pdev, AER_CORRECTABLE, &aer_regs);
cxl_handle_cor_ras(dport->dport_dev, to_ras_base(port, dport));
- else
+ }
+
+ if (uncor_status) {
+ int severity = (uncor_status & uncor_regs.uncor_severity) ?
+ AER_FATAL : AER_NONFATAL;
+
+ /*
+ * Log unconditionally. The correctable pci_print_aer() only
+ * logs this via ANFE recursion when aer_compute_anfe_status()
+ * is non-zero. Testing PCI_ERR_COR_ADV_NFAT alone cannot tell
+ * whether it fired, and the HW status is already cleared. A
+ * duplicate line beats a lost error.
+ */
+ pci_print_aer(pdev, severity, &uncor_regs);
+
cxl_do_recovery(pdev, dport->port, dport);
+ }
}
--
2.34.1
^ permalink raw reply related [flat|nested] 17+ messages in thread* Re: [PATCH v20 3/9] cxl/ras: Handle RCH correctable and uncorrectable errors in one pass
2026-09-02 13:39 ` [PATCH v20 3/9] cxl/ras: Handle RCH correctable and uncorrectable errors in one pass Terry Bowman
@ 2026-09-02 20:57 ` Cheatham, Benjamin
0 siblings, 0 replies; 17+ messages in thread
From: Cheatham, Benjamin @ 2026-09-02 20:57 UTC (permalink / raw)
To: Terry Bowman, Jonathan Cameron, Dave Jiang, Alison Schofield,
Vishal Verma, Davidlohr Bueso, Bjorn Helgaas, Dan Williams,
Rafael J . Wysocki, Jonathan Corbet, linux-cxl
Cc: Tony Luck, Borislav Petkov, Hanjun Guo, Mauro Carvalho Chehab,
Shuai Xue, Len Brown, Ira Weiny, Li Ming, Shuah Khan,
Richard Cheng, Robert Richter, Lukas Wunner, linux-pci,
linux-acpi, linux-doc, linux-kernel
On 9/2/2026 8:39 AM, Terry Bowman wrote:
> cxl_rch_get_aer_info() reads and clears both the correctable and
> uncorrectable AER status registers in a single pass. The previous
> severity decode returned after the first matching class, so when a
> correctable and an uncorrectable error were logged simultaneously the
> correctable event was cleared in hardware but never traced or handled.
>
> Handle both classes independently: dispatch cxl_handle_cor_ras() when
> correctable status is set and cxl_do_recovery() when uncorrectable
> status is set. Remove the now-unused cxl_rch_get_aer_severity() helper
> and decode the uncorrectable severity inline.
>
> Log the uncorrectable status unconditionally. __pci_print_aer() only
> logs it via ANFE recursion when aer_compute_anfe_status() is non-zero.
> A conditional skip could drop a fatal or non-ANFE record once the
> hardware status is cleared.
>
> Reported-by: Sashiko <sashiko-bot@kernel.org>
> Link: https://lore.kernel.org/linux-cxl/20260803222923.517B11F00A3A@smtp.kernel.org/
> Signed-off-by: Terry Bowman <terry.bowman@amd.com>
>
> ---
LGTM:
Reviewed-by: Ben Cheatham <benjamin.cheatham@amd.com>
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH v20 4/9] cxl/pci: Thread port and dport through RAS handling helpers
2026-09-02 13:39 [PATCH v20 0/9] Enable CXL PCIe Port Protocol Error handling and logging Terry Bowman
` (2 preceding siblings ...)
2026-09-02 13:39 ` [PATCH v20 3/9] cxl/ras: Handle RCH correctable and uncorrectable errors in one pass Terry Bowman
@ 2026-09-02 13:39 ` Terry Bowman
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-02 13:39 ` [PATCH v20 5/9] cxl: Update CXL Endpoint AER handler Terry Bowman
` (4 subsequent siblings)
8 siblings, 1 reply; 17+ messages in thread
From: Terry Bowman @ 2026-09-02 13:39 UTC (permalink / raw)
To: Jonathan Cameron, Dave Jiang, Alison Schofield, Vishal Verma,
Davidlohr Bueso, Bjorn Helgaas, Dan Williams, Rafael J . Wysocki,
Jonathan Corbet, linux-cxl
Cc: Tony Luck, Borislav Petkov, Hanjun Guo, Mauro Carvalho Chehab,
Shuai Xue, Len Brown, Ira Weiny, Li Ming, Shuah Khan,
Ben Cheatham, Richard Cheng, Robert Richter, Lukas Wunner,
linux-pci, linux-acpi, linux-doc, linux-kernel
From: Dan Williams <djbw@kernel.org>
The callers of cxl_handle_ras() and cxl_handle_cor_ras() already hold
a struct cxl_port * and optionally a struct cxl_dport * for the device
being handled. Passing a generic struct device * requires is_cxl_memdev()
to distinguish Endpoints from Ports at trace emission time. Threading
Port and Downstream Port directly enables is_cxl_endpoint() and explicit
dport/port branching for cleaner trace dispatch.
Refactor cxl_handle_ras() and cxl_handle_cor_ras() to accept struct
cxl_port * and struct cxl_dport * directly. The CXL RAS trace event
emission logic is split into three branches: Endpoint events are
identified via is_cxl_endpoint() and emit with the memdev, dport events
emit with dport->dport_dev, and Upstream Port events fall back to
port->uport_dev. This branching is transitional: the follow-on patch
("cxl: Add port and dport identifiers to CXL AER trace events") unifies
the trace events on port/dport and removes it.
Update cxl_handle_rdport_errors() and cxl_handle_proto_error() to pass
Port and Downstream Port to the refactored functions.
RCH Downstream Port correctable trace events now report the dport device
(dport->dport_dev) as a consequence of threading Port and Downstream
Port through the RAS helpers. The following trace event rework ("cxl: Add
port and dport identifiers to CXL AER trace events") adds explicit memdev,
Port, Downstream Port, and host fields that provide full context for all
device types.
Co-developed-by: Terry Bowman <terry.bowman@amd.com>
Signed-off-by: Terry Bowman <terry.bowman@amd.com>
Signed-off-by: Dan Williams <djbw@kernel.org>
Reviewed-by: Dave Jiang <dave.jiang@intel.com>
Reviewed-by: Alison Schofield <alison.schofield@intel.com>
---
Changes in v19 -> v20:
- None
Changes in v18 -> v19:
- Added review-by for DaveJ
- Update commit message that three-way trace branching is transitional
and removed by the following trace-event unification patch
Changes in v17 -> v18:
- New patch.
---
drivers/cxl/core/core.h | 12 ++++++++----
drivers/cxl/core/ras.c | 29 +++++++++++++++--------------
drivers/cxl/core/ras_rch.c | 2 +-
3 files changed, 24 insertions(+), 19 deletions(-)
diff --git a/drivers/cxl/core/core.h b/drivers/cxl/core/core.h
index 645824167f788..4f6b4702deb38 100644
--- a/drivers/cxl/core/core.h
+++ b/drivers/cxl/core/core.h
@@ -187,10 +187,12 @@ static inline struct device *dport_to_host(struct cxl_dport *dport)
#ifdef CONFIG_CXL_RAS
void cxl_ras_init(void);
void cxl_ras_exit(void);
-bool cxl_handle_ras(struct device *dev, void __iomem *ras_base);
+bool cxl_handle_ras(struct cxl_port *port, struct cxl_dport *dport,
+ void __iomem *ras_base);
void cxl_do_recovery(struct pci_dev *pdev, struct cxl_port *port,
struct cxl_dport *dport);
-void cxl_handle_cor_ras(struct device *dev, void __iomem *ras_base);
+void cxl_handle_cor_ras(struct cxl_port *port, struct cxl_dport *dport,
+ void __iomem *ras_base);
void cxl_dport_map_rch_aer(struct cxl_dport *dport);
void cxl_disable_rch_root_ints(struct cxl_dport *dport);
void cxl_handle_rdport_errors(struct pci_dev *pdev);
@@ -199,13 +201,15 @@ void devm_cxl_dport_ras_setup(struct cxl_dport *dport);
#else
static inline void cxl_ras_init(void) { }
static inline void cxl_ras_exit(void) { }
-static inline bool cxl_handle_ras(struct device *dev, void __iomem *ras_base)
+static inline bool cxl_handle_ras(struct cxl_port *port, struct cxl_dport *dport,
+ void __iomem *ras_base)
{
return false;
}
static inline void cxl_do_recovery(struct pci_dev *pdev, struct cxl_port *port,
struct cxl_dport *dport) { }
-static inline void cxl_handle_cor_ras(struct device *dev, void __iomem *ras_base) { }
+static inline void cxl_handle_cor_ras(struct cxl_port *port, struct cxl_dport *dport,
+ void __iomem *ras_base) { }
static inline void cxl_dport_map_rch_aer(struct cxl_dport *dport) { }
static inline void cxl_disable_rch_root_ints(struct cxl_dport *dport) { }
static inline void cxl_handle_rdport_errors(struct pci_dev *pdev) { }
diff --git a/drivers/cxl/core/ras.c b/drivers/cxl/core/ras.c
index 02c791f149270..e99dcd028b738 100644
--- a/drivers/cxl/core/ras.c
+++ b/drivers/cxl/core/ras.c
@@ -227,20 +227,19 @@ void __iomem *to_ras_base(struct cxl_port *port, struct cxl_dport *dport)
void cxl_do_recovery(struct pci_dev *pdev, struct cxl_port *port, struct cxl_dport *dport)
{
- struct device *dev = dport ? dport->dport_dev : port->uport_dev;
void __iomem *ras_base = to_ras_base(port, dport);
if (!ras_base)
panic("CXL: UCE with unmapped RAS registers");
- if (cxl_handle_ras(dev, ras_base))
+ if (cxl_handle_ras(port, dport, ras_base))
panic("CXL cachemem error");
dev_dbg(&pdev->dev,
"CXL UCE signaled but no CXL RAS status bits set\n");
}
-void cxl_handle_cor_ras(struct device *dev, void __iomem *ras_base)
+void cxl_handle_cor_ras(struct cxl_port *port, struct cxl_dport *dport, void __iomem *ras_base)
{
void __iomem *addr;
u32 status;
@@ -252,10 +251,12 @@ void cxl_handle_cor_ras(struct device *dev, void __iomem *ras_base)
status = readl(addr);
if (status & CXL_RAS_CORRECTABLE_STATUS_MASK) {
writel(status & CXL_RAS_CORRECTABLE_STATUS_MASK, addr);
- if (is_cxl_memdev(dev))
- trace_cxl_aer_correctable_error(to_cxl_memdev(dev), status);
+ if (is_cxl_endpoint(port))
+ trace_cxl_aer_correctable_error(to_cxl_memdev(port->uport_dev), status);
+ else if (dport)
+ trace_cxl_port_aer_correctable_error(dport->dport_dev, status);
else
- trace_cxl_port_aer_correctable_error(dev, status);
+ trace_cxl_port_aer_correctable_error(port->uport_dev, status);
}
}
@@ -280,7 +281,7 @@ static void header_log_copy(void __iomem *ras_base, u32 *log)
* Log the state of the RAS status registers and prepare them to log the
* next error status. Return 1 if reset needed.
*/
-bool cxl_handle_ras(struct device *dev, void __iomem *ras_base)
+bool cxl_handle_ras(struct cxl_port *port, struct cxl_dport *dport, void __iomem *ras_base)
{
u32 hl[CXL_HEADERLOG_TRACE_SIZE_U32] = {};
void __iomem *addr;
@@ -307,10 +308,12 @@ bool cxl_handle_ras(struct device *dev, void __iomem *ras_base)
}
header_log_copy(ras_base, hl);
- if (is_cxl_memdev(dev))
- trace_cxl_aer_uncorrectable_error(to_cxl_memdev(dev), status, fe, hl);
+ if (is_cxl_endpoint(port))
+ trace_cxl_aer_uncorrectable_error(to_cxl_memdev(port->uport_dev), status, fe, hl);
+ else if (dport)
+ trace_cxl_port_aer_uncorrectable_error(dport->dport_dev, status, fe, hl);
else
- trace_cxl_port_aer_uncorrectable_error(dev, status, fe, hl);
+ trace_cxl_port_aer_uncorrectable_error(port->uport_dev, status, fe, hl);
writel(status & CXL_RAS_UNCORRECTABLE_STATUS_MASK, addr);
@@ -344,7 +347,7 @@ pci_ers_result_t cxl_error_detected(struct pci_dev *pdev,
* cache coherency is already lost; continuing risks silent
* data corruption.
*/
- ue = cxl_handle_ras(port->uport_dev, to_ras_base(port, NULL));
+ ue = cxl_handle_ras(port, NULL, to_ras_base(port, NULL));
}
/*
@@ -375,10 +378,8 @@ EXPORT_SYMBOL_NS_GPL(cxl_error_detected, "CXL");
static void cxl_handle_proto_error(struct pci_dev *pdev, struct cxl_port *port,
struct cxl_dport *dport, int severity)
{
- struct device *dev = dport ? dport->dport_dev : port->uport_dev;
-
if (severity == AER_CORRECTABLE)
- cxl_handle_cor_ras(dev, to_ras_base(port, dport));
+ cxl_handle_cor_ras(port, dport, to_ras_base(port, dport));
else
cxl_do_recovery(pdev, port, dport);
}
diff --git a/drivers/cxl/core/ras_rch.c b/drivers/cxl/core/ras_rch.c
index 285eed5b828b9..bc6cf50fcc926 100644
--- a/drivers/cxl/core/ras_rch.c
+++ b/drivers/cxl/core/ras_rch.c
@@ -123,7 +123,7 @@ void cxl_handle_rdport_errors(struct pci_dev *pdev)
*/
if (aer_regs.cor_status & ~aer_regs.cor_mask) {
pci_print_aer(pdev, AER_CORRECTABLE, &aer_regs);
- cxl_handle_cor_ras(dport->dport_dev, to_ras_base(port, dport));
+ cxl_handle_cor_ras(port, dport, to_ras_base(port, dport));
}
if (uncor_status) {
--
2.34.1
^ permalink raw reply related [flat|nested] 17+ messages in thread* Re: [PATCH v20 4/9] cxl/pci: Thread port and dport through RAS handling helpers
2026-09-02 13:39 ` [PATCH v20 4/9] cxl/pci: Thread port and dport through RAS handling helpers Terry Bowman
@ 2026-09-02 20:57 ` Cheatham, Benjamin
0 siblings, 0 replies; 17+ messages in thread
From: Cheatham, Benjamin @ 2026-09-02 20:57 UTC (permalink / raw)
To: Terry Bowman, Jonathan Cameron, Dave Jiang, Alison Schofield,
Vishal Verma, Davidlohr Bueso, Bjorn Helgaas, Dan Williams,
Rafael J . Wysocki, Jonathan Corbet, linux-cxl
Cc: Tony Luck, Borislav Petkov, Hanjun Guo, Mauro Carvalho Chehab,
Shuai Xue, Len Brown, Ira Weiny, Li Ming, Shuah Khan,
Richard Cheng, Robert Richter, Lukas Wunner, linux-pci,
linux-acpi, linux-doc, linux-kernel
On 9/2/2026 8:39 AM, Terry Bowman wrote:
> From: Dan Williams <djbw@kernel.org>
>
> The callers of cxl_handle_ras() and cxl_handle_cor_ras() already hold
> a struct cxl_port * and optionally a struct cxl_dport * for the device
> being handled. Passing a generic struct device * requires is_cxl_memdev()
> to distinguish Endpoints from Ports at trace emission time. Threading
> Port and Downstream Port directly enables is_cxl_endpoint() and explicit
> dport/port branching for cleaner trace dispatch.
>
> Refactor cxl_handle_ras() and cxl_handle_cor_ras() to accept struct
> cxl_port * and struct cxl_dport * directly. The CXL RAS trace event
> emission logic is split into three branches: Endpoint events are
> identified via is_cxl_endpoint() and emit with the memdev, dport events
> emit with dport->dport_dev, and Upstream Port events fall back to
> port->uport_dev. This branching is transitional: the follow-on patch
> ("cxl: Add port and dport identifiers to CXL AER trace events") unifies
> the trace events on port/dport and removes it.
>
> Update cxl_handle_rdport_errors() and cxl_handle_proto_error() to pass
> Port and Downstream Port to the refactored functions.
>
> RCH Downstream Port correctable trace events now report the dport device
> (dport->dport_dev) as a consequence of threading Port and Downstream
> Port through the RAS helpers. The following trace event rework ("cxl: Add
> port and dport identifiers to CXL AER trace events") adds explicit memdev,
> Port, Downstream Port, and host fields that provide full context for all
> device types.
>
> Co-developed-by: Terry Bowman <terry.bowman@amd.com>
> Signed-off-by: Terry Bowman <terry.bowman@amd.com>
> Signed-off-by: Dan Williams <djbw@kernel.org>
> Reviewed-by: Dave Jiang <dave.jiang@intel.com>
> Reviewed-by: Alison Schofield <alison.schofield@intel.com>
>
> ---
>
> Changes in v19 -> v20:
> - None
>
> Changes in v18 -> v19:
> - Added review-by for DaveJ
> - Update commit message that three-way trace branching is transitional
> and removed by the following trace-event unification patch
>
> Changes in v17 -> v18:
> - New patch.
> ---
> drivers/cxl/core/core.h | 12 ++++++++----
> drivers/cxl/core/ras.c | 29 +++++++++++++++--------------
> drivers/cxl/core/ras_rch.c | 2 +-
> 3 files changed, 24 insertions(+), 19 deletions(-)
>
> diff --git a/drivers/cxl/core/core.h b/drivers/cxl/core/core.h
> index 645824167f788..4f6b4702deb38 100644
> --- a/drivers/cxl/core/core.h
> +++ b/drivers/cxl/core/core.h
> @@ -187,10 +187,12 @@ static inline struct device *dport_to_host(struct cxl_dport *dport)
> #ifdef CONFIG_CXL_RAS
> void cxl_ras_init(void);
> void cxl_ras_exit(void);
> -bool cxl_handle_ras(struct device *dev, void __iomem *ras_base);
> +bool cxl_handle_ras(struct cxl_port *port, struct cxl_dport *dport,
> + void __iomem *ras_base);
> void cxl_do_recovery(struct pci_dev *pdev, struct cxl_port *port,
> struct cxl_dport *dport);
> -void cxl_handle_cor_ras(struct device *dev, void __iomem *ras_base);
> +void cxl_handle_cor_ras(struct cxl_port *port, struct cxl_dport *dport,
> + void __iomem *ras_base);
> void cxl_dport_map_rch_aer(struct cxl_dport *dport);
> void cxl_disable_rch_root_ints(struct cxl_dport *dport);
> void cxl_handle_rdport_errors(struct pci_dev *pdev);
> @@ -199,13 +201,15 @@ void devm_cxl_dport_ras_setup(struct cxl_dport *dport);
> #else
> static inline void cxl_ras_init(void) { }
> static inline void cxl_ras_exit(void) { }
> -static inline bool cxl_handle_ras(struct device *dev, void __iomem *ras_base)
> +static inline bool cxl_handle_ras(struct cxl_port *port, struct cxl_dport *dport,
> + void __iomem *ras_base)
> {
> return false;
> }
> static inline void cxl_do_recovery(struct pci_dev *pdev, struct cxl_port *port,
> struct cxl_dport *dport) { }
> -static inline void cxl_handle_cor_ras(struct device *dev, void __iomem *ras_base) { }
> +static inline void cxl_handle_cor_ras(struct cxl_port *port, struct cxl_dport *dport,
> + void __iomem *ras_base) { }
> static inline void cxl_dport_map_rch_aer(struct cxl_dport *dport) { }
> static inline void cxl_disable_rch_root_ints(struct cxl_dport *dport) { }
> static inline void cxl_handle_rdport_errors(struct pci_dev *pdev) { }
> diff --git a/drivers/cxl/core/ras.c b/drivers/cxl/core/ras.c
> index 02c791f149270..e99dcd028b738 100644
> --- a/drivers/cxl/core/ras.c
> +++ b/drivers/cxl/core/ras.c
> @@ -227,20 +227,19 @@ void __iomem *to_ras_base(struct cxl_port *port, struct cxl_dport *dport)
>
> void cxl_do_recovery(struct pci_dev *pdev, struct cxl_port *port, struct cxl_dport *dport)
> {
> - struct device *dev = dport ? dport->dport_dev : port->uport_dev;
> void __iomem *ras_base = to_ras_base(port, dport);
>
> if (!ras_base)
> panic("CXL: UCE with unmapped RAS registers");
>
> - if (cxl_handle_ras(dev, ras_base))
> + if (cxl_handle_ras(port, dport, ras_base))
I haven't looked ahead, and maybe I'm missing something, but why not have cxl_handle_ras() just
take the port & dport and call to_ras_base() internally? It's not a big deal, but it would
make the call in cxl_error_detected() a lot prettier.
Either way:
Reviewed-by: Ben Cheatham <benjamin.cheatham@amd.com>
> panic("CXL cachemem error");
>
> dev_dbg(&pdev->dev,
> "CXL UCE signaled but no CXL RAS status bits set\n");
> }
>
> -void cxl_handle_cor_ras(struct device *dev, void __iomem *ras_base)
> +void cxl_handle_cor_ras(struct cxl_port *port, struct cxl_dport *dport, void __iomem *ras_base)
> {
> void __iomem *addr;
> u32 status;
> @@ -252,10 +251,12 @@ void cxl_handle_cor_ras(struct device *dev, void __iomem *ras_base)
> status = readl(addr);
> if (status & CXL_RAS_CORRECTABLE_STATUS_MASK) {
> writel(status & CXL_RAS_CORRECTABLE_STATUS_MASK, addr);
> - if (is_cxl_memdev(dev))
> - trace_cxl_aer_correctable_error(to_cxl_memdev(dev), status);
> + if (is_cxl_endpoint(port))
> + trace_cxl_aer_correctable_error(to_cxl_memdev(port->uport_dev), status);
> + else if (dport)
> + trace_cxl_port_aer_correctable_error(dport->dport_dev, status);
> else
> - trace_cxl_port_aer_correctable_error(dev, status);
> + trace_cxl_port_aer_correctable_error(port->uport_dev, status);
> }
> }
>
> @@ -280,7 +281,7 @@ static void header_log_copy(void __iomem *ras_base, u32 *log)
> * Log the state of the RAS status registers and prepare them to log the
> * next error status. Return 1 if reset needed.
> */
> -bool cxl_handle_ras(struct device *dev, void __iomem *ras_base)
> +bool cxl_handle_ras(struct cxl_port *port, struct cxl_dport *dport, void __iomem *ras_base)
> {
> u32 hl[CXL_HEADERLOG_TRACE_SIZE_U32] = {};
> void __iomem *addr;
> @@ -307,10 +308,12 @@ bool cxl_handle_ras(struct device *dev, void __iomem *ras_base)
> }
>
> header_log_copy(ras_base, hl);
> - if (is_cxl_memdev(dev))
> - trace_cxl_aer_uncorrectable_error(to_cxl_memdev(dev), status, fe, hl);
> + if (is_cxl_endpoint(port))
> + trace_cxl_aer_uncorrectable_error(to_cxl_memdev(port->uport_dev), status, fe, hl);
> + else if (dport)
> + trace_cxl_port_aer_uncorrectable_error(dport->dport_dev, status, fe, hl);
> else
> - trace_cxl_port_aer_uncorrectable_error(dev, status, fe, hl);
> + trace_cxl_port_aer_uncorrectable_error(port->uport_dev, status, fe, hl);
>
> writel(status & CXL_RAS_UNCORRECTABLE_STATUS_MASK, addr);
>
> @@ -344,7 +347,7 @@ pci_ers_result_t cxl_error_detected(struct pci_dev *pdev,
> * cache coherency is already lost; continuing risks silent
> * data corruption.
> */
> - ue = cxl_handle_ras(port->uport_dev, to_ras_base(port, NULL));
> + ue = cxl_handle_ras(port, NULL, to_ras_base(port, NULL));
> }
>
> /*
> @@ -375,10 +378,8 @@ EXPORT_SYMBOL_NS_GPL(cxl_error_detected, "CXL");
> static void cxl_handle_proto_error(struct pci_dev *pdev, struct cxl_port *port,
> struct cxl_dport *dport, int severity)
> {
> - struct device *dev = dport ? dport->dport_dev : port->uport_dev;
> -
> if (severity == AER_CORRECTABLE)
> - cxl_handle_cor_ras(dev, to_ras_base(port, dport));
> + cxl_handle_cor_ras(port, dport, to_ras_base(port, dport));
> else
> cxl_do_recovery(pdev, port, dport);
> }
> diff --git a/drivers/cxl/core/ras_rch.c b/drivers/cxl/core/ras_rch.c
> index 285eed5b828b9..bc6cf50fcc926 100644
> --- a/drivers/cxl/core/ras_rch.c
> +++ b/drivers/cxl/core/ras_rch.c
> @@ -123,7 +123,7 @@ void cxl_handle_rdport_errors(struct pci_dev *pdev)
> */
> if (aer_regs.cor_status & ~aer_regs.cor_mask) {
> pci_print_aer(pdev, AER_CORRECTABLE, &aer_regs);
> - cxl_handle_cor_ras(dport->dport_dev, to_ras_base(port, dport));
> + cxl_handle_cor_ras(port, dport, to_ras_base(port, dport));
> }
>
> if (uncor_status) {
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH v20 5/9] cxl: Update CXL Endpoint AER handler
2026-09-02 13:39 [PATCH v20 0/9] Enable CXL PCIe Port Protocol Error handling and logging Terry Bowman
` (3 preceding siblings ...)
2026-09-02 13:39 ` [PATCH v20 4/9] cxl/pci: Thread port and dport through RAS handling helpers Terry Bowman
@ 2026-09-02 13:39 ` Terry Bowman
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-02 13:39 ` [PATCH v20 6/9] PCI: Cache PCI DSN into pci_dev->dsn during probe Terry Bowman
` (3 subsequent siblings)
8 siblings, 1 reply; 17+ messages in thread
From: Terry Bowman @ 2026-09-02 13:39 UTC (permalink / raw)
To: Jonathan Cameron, Dave Jiang, Alison Schofield, Vishal Verma,
Davidlohr Bueso, Bjorn Helgaas, Dan Williams, Rafael J . Wysocki,
Jonathan Corbet, linux-cxl
Cc: Tony Luck, Borislav Petkov, Hanjun Guo, Mauro Carvalho Chehab,
Shuai Xue, Len Brown, Ira Weiny, Li Ming, Shuah Khan,
Ben Cheatham, Richard Cheng, Robert Richter, Lukas Wunner,
linux-pci, linux-acpi, linux-doc, linux-kernel
Rename cxl_error_detected() to cxl_pci_error_detected() and rename the
struct pci_error_handlers instance from cxl_error_handlers to
cxl_pci_error_handlers for consistency with the cxl_pci_ prefix used by
the renamed handler and the rest of the PCI-facing entry points.
Document the unconditional CXL RAS read policy: on a dead link, readl()
returns 0xFFFFFFFF which is interpreted as UCE bits set and triggers a
panic. If RAS registers are not mapped the read is skipped and the
frozen/perm_failure switch cases defer to AER recovery for devices
without active CXL.mem traffic.
Signed-off-by: Terry Bowman <terry.bowman@amd.com>
Reviewed-by: Dave Jiang <dave.jiang@intel.com>
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
Reviewed-by: Alison Schofield <alison.schofield@intel.com>
---
Changes in v19->v20:
- Expand the comment block in cxl_pci_error_detected().
- Clarify commit message's first paragraph. cxl_error_handlers is a variable
- Clarify AER handling in cxl_pci_error_detected() block comment
Changes in v18->v19:
- Add review-by for DaveJ and Jonathan
- Remove reindent introduced in v18 at cxl_error_handlers definition.
Changes in v17->v18:
- Fix cxl_pci_error_detected() to use find_cxl_port_by_uport() and port->uport_dev
- Read CXL RAS unconditionally; panic on UCE regardless of channel state
- Document unconditional read policy and 0xFFFFFFFF behavior in comment
- Drop guard removal paragraph from commit message (not in this diff)
- Drop Reviewed-by tags pending re-review after message change
Changes in v16->v17:
- Rename pci_error_handlers struct instance to cxl_pci_error_handlers for
cxl_pci_ naming consistency.
- Restore scoped_guard(device) and dev->driver check around AER read.
- NULL-check find_cxl_port_by_dev() before deref of port->uport_dev.
- Updated commit message. (Terry)
- Add scope cleanup for port variable in cxl_pci_error_detected() (Terry)
- Drop cxl_uncor_aer_present(), rely on AER state
Changes in v15->v16:
- Update commit message (DaveJ)
- s/cxl_handle_aer()/cxl_uncor_aer_present()/g (Jonathan)
- cxl_uncor_aer_present(): Leave original result calculation based on
if a UCE is present and the provided state (Terry)
- Add call to pci_print_aer(). AER fails to log because is upstream
link (Terry)
Changes in v14->v15:
- Update commit message and title. Added Bjorn's ack.
- Move CE and UCE handling logic here
Changes in v13->v14:
- Add Dave Jiang's review-by
- Update commit message & headline (Bjorn)
- Refactor cxl_port_error_detected()/cxl_port_cor_error_detected() to
one line (Jonathan)
- Remove cxl_walk_port() (Dan)
- Remove cxl_pci_drv_bound(). Check for 'is_cxl' parent port is
sufficient (Dan)
- Remove device_lock_if()
- Combined CE and UCE here (Terry)
Changes in v12->v13:
- Move get_pci_cxl_host_dev() and cxl_handle_proto_error() to Dequeue
patch (Terry)
- Remove EP case in cxl_get_ras_base(), not used. (Terry)
- Remove check for dport->dport_dev (Dave)
- Remove whitespace (Terry)
Changes in v11->v12:
- Add call to cxl_pci_drv_bound() in cxl_handle_proto_error() and
pci_to_cxl_dev()
- Change cxl_error_detected() -> cxl_cor_error_detected()
- Remove NULL variable assignments
- Replace bus_find_device() with find_cxl_port_by_uport() for upstream
port searches.
Changes in v10->v11:
- None
---
drivers/cxl/core/ras.c | 25 ++++++++++++++++---------
drivers/cxl/cxlpci.h | 8 ++++----
drivers/cxl/pci.c | 6 +++---
3 files changed, 23 insertions(+), 16 deletions(-)
diff --git a/drivers/cxl/core/ras.c b/drivers/cxl/core/ras.c
index e99dcd028b738..bf479fac08565 100644
--- a/drivers/cxl/core/ras.c
+++ b/drivers/cxl/core/ras.c
@@ -61,7 +61,7 @@ cxl_cper_trace_uncorr_prot_err(struct cxl_memdev *cxlmd,
/*
* ras_cap.header_log[] holds CXL_HEADERLOG_SIZE_U32 (16) hardware
- * dwords. Copy them into the front of a zero-filled
+ * dwords. Copy them into the front of a zero-filled
* CXL_HEADERLOG_TRACE_SIZE_U32 (128) u32 staging buffer so the trace
* event memcpy sees a full 512-byte source and the userspace ABI
* (rasdaemon) is preserved.
@@ -320,8 +320,8 @@ bool cxl_handle_ras(struct cxl_port *port, struct cxl_dport *dport, void __iomem
return true;
}
-pci_ers_result_t cxl_error_detected(struct pci_dev *pdev,
- pci_channel_state_t state)
+pci_ers_result_t cxl_pci_error_detected(struct pci_dev *pdev,
+ pci_channel_state_t state)
{
struct cxl_port *port __free(put_cxl_port) = find_cxl_port_by_uport(&pdev->dev);
bool ue = false;
@@ -341,11 +341,18 @@ pci_ers_result_t cxl_error_detected(struct pci_dev *pdev,
}
/*
- * The CXL RAS read is unconditional regardless of channel
- * state. Any uncorrectable error bit set in the CXL RAS
- * status register triggers a panic below because CXL.mem
- * cache coherency is already lost; continuing risks silent
- * data corruption.
+ * The CXL RAS uncorrectable status is the only signal here
+ * that the error is a CXL internal (protocol) error. A set
+ * UCE bit confirms it and triggers the panic below. On a dead
+ * link readl() returns 0xFFFFFFFF, which sets all UCE bits and
+ * triggers the panic intentionally.
+ *
+ * If RAS is not mapped the read is skipped. Unlike
+ * cxl_do_recovery(), which is reached only after
+ * is_aer_internal_error() has already confirmed a CXL internal
+ * UCE, this path has no such confirmation, so an unmapped RAS
+ * block cannot attribute the error to CXL and must not panic.
+ * The switch cases below then handle AER recovery.
*/
ue = cxl_handle_ras(port, NULL, to_ras_base(port, NULL));
}
@@ -373,7 +380,7 @@ pci_ers_result_t cxl_error_detected(struct pci_dev *pdev,
}
return PCI_ERS_RESULT_NEED_RESET;
}
-EXPORT_SYMBOL_NS_GPL(cxl_error_detected, "CXL");
+EXPORT_SYMBOL_NS_GPL(cxl_pci_error_detected, "CXL");
static void cxl_handle_proto_error(struct pci_dev *pdev, struct cxl_port *port,
struct cxl_dport *dport, int severity)
diff --git a/drivers/cxl/cxlpci.h b/drivers/cxl/cxlpci.h
index 606fadd2476f3..7421132899fc4 100644
--- a/drivers/cxl/cxlpci.h
+++ b/drivers/cxl/cxlpci.h
@@ -79,13 +79,13 @@ struct cxl_dev_state;
void read_cdat_data(struct cxl_port *port);
#ifdef CONFIG_CXL_RAS
-pci_ers_result_t cxl_error_detected(struct pci_dev *pdev,
- pci_channel_state_t state);
+pci_ers_result_t cxl_pci_error_detected(struct pci_dev *pdev,
+ pci_channel_state_t state);
void devm_cxl_dport_rch_ras_setup(struct cxl_dport *dport);
void devm_cxl_port_ras_setup(struct cxl_port *port);
#else
-static inline pci_ers_result_t cxl_error_detected(struct pci_dev *pdev,
- pci_channel_state_t state)
+static inline pci_ers_result_t cxl_pci_error_detected(struct pci_dev *pdev,
+ pci_channel_state_t state)
{
return PCI_ERS_RESULT_NONE;
}
diff --git a/drivers/cxl/pci.c b/drivers/cxl/pci.c
index fe11be7fd9aab..1e7be77ded634 100644
--- a/drivers/cxl/pci.c
+++ b/drivers/cxl/pci.c
@@ -998,8 +998,8 @@ static void cxl_reset_done(struct pci_dev *pdev)
}
}
-static const struct pci_error_handlers cxl_error_handlers = {
- .error_detected = cxl_error_detected,
+static const struct pci_error_handlers cxl_pci_error_handlers = {
+ .error_detected = cxl_pci_error_detected,
.slot_reset = cxl_slot_reset,
.resume = cxl_error_resume,
.reset_done = cxl_reset_done,
@@ -1009,7 +1009,7 @@ static struct pci_driver cxl_pci_driver = {
.name = KBUILD_MODNAME,
.id_table = cxl_mem_pci_tbl,
.probe = cxl_pci_probe,
- .err_handler = &cxl_error_handlers,
+ .err_handler = &cxl_pci_error_handlers,
.dev_groups = cxl_rcd_groups,
.driver = {
.probe_type = PROBE_PREFER_ASYNCHRONOUS,
--
2.34.1
^ permalink raw reply related [flat|nested] 17+ messages in thread* Re: [PATCH v20 5/9] cxl: Update CXL Endpoint AER handler
2026-09-02 13:39 ` [PATCH v20 5/9] cxl: Update CXL Endpoint AER handler Terry Bowman
@ 2026-09-02 20:57 ` Cheatham, Benjamin
0 siblings, 0 replies; 17+ messages in thread
From: Cheatham, Benjamin @ 2026-09-02 20:57 UTC (permalink / raw)
To: Terry Bowman, Jonathan Cameron, Dave Jiang, Alison Schofield,
Vishal Verma, Davidlohr Bueso, Bjorn Helgaas, Dan Williams,
Rafael J . Wysocki, Jonathan Corbet, linux-cxl
Cc: Tony Luck, Borislav Petkov, Hanjun Guo, Mauro Carvalho Chehab,
Shuai Xue, Len Brown, Ira Weiny, Li Ming, Shuah Khan,
Richard Cheng, Robert Richter, Lukas Wunner, linux-pci,
linux-acpi, linux-doc, linux-kernel
On 9/2/2026 8:39 AM, Terry Bowman wrote:
> Rename cxl_error_detected() to cxl_pci_error_detected() and rename the
> struct pci_error_handlers instance from cxl_error_handlers to
> cxl_pci_error_handlers for consistency with the cxl_pci_ prefix used by
> the renamed handler and the rest of the PCI-facing entry points.
>
> Document the unconditional CXL RAS read policy: on a dead link, readl()
> returns 0xFFFFFFFF which is interpreted as UCE bits set and triggers a
> panic. If RAS registers are not mapped the read is skipped and the
> frozen/perm_failure switch cases defer to AER recovery for devices
> without active CXL.mem traffic.
>
> Signed-off-by: Terry Bowman <terry.bowman@amd.com>
> Reviewed-by: Dave Jiang <dave.jiang@intel.com>
> Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
> Reviewed-by: Alison Schofield <alison.schofield@intel.com>
>
> ---
>
> Changes in v19->v20:
> - Expand the comment block in cxl_pci_error_detected().
> - Clarify commit message's first paragraph. cxl_error_handlers is a variable
> - Clarify AER handling in cxl_pci_error_detected() block comment
>
> Changes in v18->v19:
> - Add review-by for DaveJ and Jonathan
> - Remove reindent introduced in v18 at cxl_error_handlers definition.
>
> Changes in v17->v18:
> - Fix cxl_pci_error_detected() to use find_cxl_port_by_uport() and port->uport_dev
> - Read CXL RAS unconditionally; panic on UCE regardless of channel state
> - Document unconditional read policy and 0xFFFFFFFF behavior in comment
> - Drop guard removal paragraph from commit message (not in this diff)
> - Drop Reviewed-by tags pending re-review after message change
>
> Changes in v16->v17:
> - Rename pci_error_handlers struct instance to cxl_pci_error_handlers for
> cxl_pci_ naming consistency.
> - Restore scoped_guard(device) and dev->driver check around AER read.
> - NULL-check find_cxl_port_by_dev() before deref of port->uport_dev.
> - Updated commit message. (Terry)
> - Add scope cleanup for port variable in cxl_pci_error_detected() (Terry)
> - Drop cxl_uncor_aer_present(), rely on AER state
>
> Changes in v15->v16:
> - Update commit message (DaveJ)
> - s/cxl_handle_aer()/cxl_uncor_aer_present()/g (Jonathan)
> - cxl_uncor_aer_present(): Leave original result calculation based on
> if a UCE is present and the provided state (Terry)
> - Add call to pci_print_aer(). AER fails to log because is upstream
> link (Terry)
>
> Changes in v14->v15:
> - Update commit message and title. Added Bjorn's ack.
> - Move CE and UCE handling logic here
>
> Changes in v13->v14:
> - Add Dave Jiang's review-by
> - Update commit message & headline (Bjorn)
> - Refactor cxl_port_error_detected()/cxl_port_cor_error_detected() to
> one line (Jonathan)
> - Remove cxl_walk_port() (Dan)
> - Remove cxl_pci_drv_bound(). Check for 'is_cxl' parent port is
> sufficient (Dan)
> - Remove device_lock_if()
> - Combined CE and UCE here (Terry)
>
> Changes in v12->v13:
> - Move get_pci_cxl_host_dev() and cxl_handle_proto_error() to Dequeue
> patch (Terry)
> - Remove EP case in cxl_get_ras_base(), not used. (Terry)
> - Remove check for dport->dport_dev (Dave)
> - Remove whitespace (Terry)
>
> Changes in v11->v12:
> - Add call to cxl_pci_drv_bound() in cxl_handle_proto_error() and
> pci_to_cxl_dev()
> - Change cxl_error_detected() -> cxl_cor_error_detected()
> - Remove NULL variable assignments
> - Replace bus_find_device() with find_cxl_port_by_uport() for upstream
> port searches.
>
> Changes in v10->v11:
> - None
> ---
> drivers/cxl/core/ras.c | 25 ++++++++++++++++---------
> drivers/cxl/cxlpci.h | 8 ++++----
> drivers/cxl/pci.c | 6 +++---
> 3 files changed, 23 insertions(+), 16 deletions(-)
>
> diff --git a/drivers/cxl/core/ras.c b/drivers/cxl/core/ras.c
> index e99dcd028b738..bf479fac08565 100644
> --- a/drivers/cxl/core/ras.c
> +++ b/drivers/cxl/core/ras.c
> @@ -61,7 +61,7 @@ cxl_cper_trace_uncorr_prot_err(struct cxl_memdev *cxlmd,
>
> /*
> * ras_cap.header_log[] holds CXL_HEADERLOG_SIZE_U32 (16) hardware
> - * dwords. Copy them into the front of a zero-filled
> + * dwords. Copy them into the front of a zero-filled
Stray whitespace fix? Doesn't matter much to me, so:
Reviewed-by: Ben Cheatham <benjamin.cheatham@amd.com>
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH v20 6/9] PCI: Cache PCI DSN into pci_dev->dsn during probe
2026-09-02 13:39 [PATCH v20 0/9] Enable CXL PCIe Port Protocol Error handling and logging Terry Bowman
` (4 preceding siblings ...)
2026-09-02 13:39 ` [PATCH v20 5/9] cxl: Update CXL Endpoint AER handler Terry Bowman
@ 2026-09-02 13:39 ` Terry Bowman
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-02 13:39 ` [PATCH v20 7/9] cxl: Add port and dport identifiers to CXL AER trace events Terry Bowman
` (2 subsequent siblings)
8 siblings, 1 reply; 17+ messages in thread
From: Terry Bowman @ 2026-09-02 13:39 UTC (permalink / raw)
To: Jonathan Cameron, Dave Jiang, Alison Schofield, Vishal Verma,
Davidlohr Bueso, Bjorn Helgaas, Dan Williams, Rafael J . Wysocki,
Jonathan Corbet, linux-cxl
Cc: Tony Luck, Borislav Petkov, Hanjun Guo, Mauro Carvalho Chehab,
Shuai Xue, Len Brown, Ira Weiny, Li Ming, Shuah Khan,
Ben Cheatham, Richard Cheng, Robert Richter, Lukas Wunner,
linux-pci, linux-acpi, linux-doc, linux-kernel
Subsequent CXL error-reporting code paths need to log the PCI Device
Serial Number (DSN) as part of trace events emitted from interrupt or
panic context. Computing the DSN there via pci_get_dsn() requires PCI
configuration space reads, which are slow, can fail when the link is
down or frozen, and may not be safe in some contexts.
Add a u64 dsn field to struct pci_dev and populate it from pci_get_dsn()
during pci_init_capabilities() at probe time via pci_dsn_init(). Only
write dev->dsn when the read succeeds. The zero initial value from
pci_dev allocation already represents 'no DSN available.'
Remove the now-redundant dsn member from the pciehp struct controller
along with its kernel-doc. Drop the pci_get_slot()/pci_dev_put() pairs
in pciehp_configure_device() and pcie_init() that existed solely to
read the DSN into ctrl->dsn. Use pdev->dsn in pciehp_device_replaced()
for the device-replacement comparison.
pci_get_dsn() is not modified because it remains a pure config-space read
with no side effects on pci_dev. The cache is written exclusively by
pci_dsn_init() at probe time.
Signed-off-by: Terry Bowman <terry.bowman@amd.com>
Reviewed-by: Dave Jiang <dave.jiang@intel.com>
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
Reviewed-by: Alison Schofield <alison.schofield@intel.com>
---
Changes in v19 -> v20:
- Change 2 space to 1 space in comment block in pci_dsn_init
- Add Reviewed-by from Alison Schofield
Changes in v18 -> v19:
- Update pciehp hotplug to use cached DSN.
- Swapped order in series with
("cxl: Add port and dport identifiers to CXL AER trace events ")
- Update function documentation for pci_dsn_init()
Changes in v17->v18:
- New commit.
---
drivers/cxl/pci.c | 2 +-
drivers/pci/hotplug/pciehp.h | 4 ----
drivers/pci/hotplug/pciehp_hpc.c | 7 +------
drivers/pci/hotplug/pciehp_pci.c | 4 ----
drivers/pci/probe.c | 17 +++++++++++++++++
include/linux/pci.h | 1 +
6 files changed, 20 insertions(+), 15 deletions(-)
diff --git a/drivers/cxl/pci.c b/drivers/cxl/pci.c
index 1e7be77ded634..c4cd5fa960362 100644
--- a/drivers/cxl/pci.c
+++ b/drivers/cxl/pci.c
@@ -802,7 +802,7 @@ static int cxl_pci_probe(struct pci_dev *pdev, const struct pci_device_id *id)
if (!dvsec)
pci_warn(pdev, "Device DVSEC not present, skip CXL.mem init\n");
- mds = cxl_memdev_state_create(&pdev->dev, pci_get_dsn(pdev), dvsec);
+ mds = cxl_memdev_state_create(&pdev->dev, pdev->dsn, dvsec);
if (IS_ERR(mds))
return PTR_ERR(mds);
cxlds = &mds->cxlds;
diff --git a/drivers/pci/hotplug/pciehp.h b/drivers/pci/hotplug/pciehp.h
index debc79b0adfb2..30aad3b66bab9 100644
--- a/drivers/pci/hotplug/pciehp.h
+++ b/drivers/pci/hotplug/pciehp.h
@@ -46,9 +46,6 @@ extern int pciehp_poll_time;
/**
* struct controller - PCIe hotplug controller
* @pcie: pointer to the controller's PCIe port service device
- * @dsn: cached copy of Device Serial Number of Function 0 in the hotplug slot
- * (PCIe r6.2 sec 7.9.3); used to determine whether a hotplugged device
- * was replaced with a different one during system sleep
* @slot_cap: cached copy of the Slot Capabilities register
* @inband_presence_disabled: In-Band Presence Detect Disable supported by
* controller and disabled per spec recommendation (PCIe r5.0, appendix I
@@ -90,7 +87,6 @@ extern int pciehp_poll_time;
*/
struct controller {
struct pcie_device *pcie;
- u64 dsn;
u32 slot_cap; /* capabilities and quirks */
unsigned int inband_presence_disabled:1;
diff --git a/drivers/pci/hotplug/pciehp_hpc.c b/drivers/pci/hotplug/pciehp_hpc.c
index 4c62140a3cb44..5336e9c9003ae 100644
--- a/drivers/pci/hotplug/pciehp_hpc.c
+++ b/drivers/pci/hotplug/pciehp_hpc.c
@@ -587,7 +587,7 @@ bool pciehp_device_replaced(struct controller *ctrl)
reg != (pdev->subsystem_vendor | (pdev->subsystem_device << 16))))
return true;
- if (pci_get_dsn(pdev) != ctrl->dsn)
+ if (pci_get_dsn(pdev) != pdev->dsn)
return true;
return false;
@@ -1085,11 +1085,6 @@ struct controller *pcie_init(struct pcie_device *dev)
}
}
- pdev = pci_get_slot(subordinate, PCI_DEVFN(0, 0));
- if (pdev)
- ctrl->dsn = pci_get_dsn(pdev);
- pci_dev_put(pdev);
-
return ctrl;
}
diff --git a/drivers/pci/hotplug/pciehp_pci.c b/drivers/pci/hotplug/pciehp_pci.c
index 65e50bee1a8c0..ad12515a4a121 100644
--- a/drivers/pci/hotplug/pciehp_pci.c
+++ b/drivers/pci/hotplug/pciehp_pci.c
@@ -72,10 +72,6 @@ int pciehp_configure_device(struct controller *ctrl)
pci_bus_add_devices(parent);
down_read_nested(&ctrl->reset_lock, ctrl->depth);
- dev = pci_get_slot(parent, PCI_DEVFN(0, 0));
- ctrl->dsn = pci_get_dsn(dev);
- pci_dev_put(dev);
-
out:
pci_unlock_rescan_remove();
return ret;
diff --git a/drivers/pci/probe.c b/drivers/pci/probe.c
index 27008e2ea5afc..2f71bee732c36 100644
--- a/drivers/pci/probe.c
+++ b/drivers/pci/probe.c
@@ -2638,6 +2638,22 @@ void pcie_report_downtraining(struct pci_dev *dev)
__pcie_print_link_status(dev, false);
}
+/*
+ * Cache the Device Serial Number for use in contexts where config-space reads
+ * are unsafe (interrupt, panic). Process-context callers that need a fresh
+ * value (e.g. hotplug device replacement) call pci_get_dsn() and compare it
+ * against this cached pdev->dsn to detect a changed device. Note pdev->dsn
+ * is 0 for devices without the DSN capability, so such a comparison cannot
+ * distinguish a replacement.
+ */
+static void pci_dsn_init(struct pci_dev *dev)
+{
+ u64 dsn = pci_get_dsn(dev);
+
+ if (dsn)
+ dev->dsn = dsn;
+}
+
static void pci_imm_ready_init(struct pci_dev *dev)
{
u16 status;
@@ -2674,6 +2690,7 @@ static void pci_init_capabilities(struct pci_dev *dev)
pci_rebar_init(dev); /* Resizable BAR */
pci_dev3_init(dev); /* Device 3 capabilities */
pci_ide_init(dev); /* Link Integrity and Data Encryption */
+ pci_dsn_init(dev); /* Serial number */
pcie_report_downtraining(dev);
pci_init_reset_methods(dev);
diff --git a/include/linux/pci.h b/include/linux/pci.h
index d31a8d107b1ef..a7fea52c8f46e 100644
--- a/include/linux/pci.h
+++ b/include/linux/pci.h
@@ -390,6 +390,7 @@ struct pci_dev {
unsigned long *dma_alias_mask;/* Mask of enabled devfn aliases */
struct pci_driver *driver; /* Driver bound to this device */
+ u64 dsn; /* PCI Device Serial Number */
u64 dma_mask; /* Mask of the bits of bus address this
device implements. Normally this is
0xffffffff. You only need to change
--
2.34.1
^ permalink raw reply related [flat|nested] 17+ messages in thread* Re: [PATCH v20 6/9] PCI: Cache PCI DSN into pci_dev->dsn during probe
2026-09-02 13:39 ` [PATCH v20 6/9] PCI: Cache PCI DSN into pci_dev->dsn during probe Terry Bowman
@ 2026-09-02 20:57 ` Cheatham, Benjamin
0 siblings, 0 replies; 17+ messages in thread
From: Cheatham, Benjamin @ 2026-09-02 20:57 UTC (permalink / raw)
To: Terry Bowman, Jonathan Cameron, Dave Jiang, Alison Schofield,
Vishal Verma, Davidlohr Bueso, Bjorn Helgaas, Dan Williams,
Rafael J . Wysocki, Jonathan Corbet, linux-cxl
Cc: Tony Luck, Borislav Petkov, Hanjun Guo, Mauro Carvalho Chehab,
Shuai Xue, Len Brown, Ira Weiny, Li Ming, Shuah Khan,
Richard Cheng, Robert Richter, Lukas Wunner, linux-pci,
linux-acpi, linux-doc, linux-kernel
On 9/2/2026 8:39 AM, Terry Bowman wrote:
> Subsequent CXL error-reporting code paths need to log the PCI Device
> Serial Number (DSN) as part of trace events emitted from interrupt or
> panic context. Computing the DSN there via pci_get_dsn() requires PCI
> configuration space reads, which are slow, can fail when the link is
> down or frozen, and may not be safe in some contexts.
>
> Add a u64 dsn field to struct pci_dev and populate it from pci_get_dsn()
> during pci_init_capabilities() at probe time via pci_dsn_init(). Only
> write dev->dsn when the read succeeds. The zero initial value from
> pci_dev allocation already represents 'no DSN available.'
>
> Remove the now-redundant dsn member from the pciehp struct controller
> along with its kernel-doc. Drop the pci_get_slot()/pci_dev_put() pairs
> in pciehp_configure_device() and pcie_init() that existed solely to
> read the DSN into ctrl->dsn. Use pdev->dsn in pciehp_device_replaced()
> for the device-replacement comparison.
>
> pci_get_dsn() is not modified because it remains a pure config-space read
> with no side effects on pci_dev. The cache is written exclusively by
> pci_dsn_init() at probe time.
>
> Signed-off-by: Terry Bowman <terry.bowman@amd.com>
> Reviewed-by: Dave Jiang <dave.jiang@intel.com>
> Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
> Reviewed-by: Alison Schofield <alison.schofield@intel.com>
>
> ---
Reviewed-by: Ben Cheatham <benjamin.cheatham@amd.com>
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH v20 7/9] cxl: Add port and dport identifiers to CXL AER trace events
2026-09-02 13:39 [PATCH v20 0/9] Enable CXL PCIe Port Protocol Error handling and logging Terry Bowman
` (5 preceding siblings ...)
2026-09-02 13:39 ` [PATCH v20 6/9] PCI: Cache PCI DSN into pci_dev->dsn during probe Terry Bowman
@ 2026-09-02 13:39 ` Terry Bowman
2026-09-02 13:39 ` [PATCH v20 8/9] PCI/CXL: Mask/Unmask CXL protocol errors Terry Bowman
2026-09-02 13:39 ` [PATCH v20 9/9] Documentation: cxl: Document CXL protocol error handling Terry Bowman
8 siblings, 0 replies; 17+ messages in thread
From: Terry Bowman @ 2026-09-02 13:39 UTC (permalink / raw)
To: Jonathan Cameron, Dave Jiang, Alison Schofield, Vishal Verma,
Davidlohr Bueso, Bjorn Helgaas, Dan Williams, Rafael J . Wysocki,
Jonathan Corbet, linux-cxl
Cc: Tony Luck, Borislav Petkov, Hanjun Guo, Mauro Carvalho Chehab,
Shuai Xue, Len Brown, Ira Weiny, Li Ming, Shuah Khan,
Ben Cheatham, Richard Cheng, Robert Richter, Lukas Wunner,
linux-pci, linux-acpi, linux-doc, linux-kernel
From: Dan Williams <djbw@kernel.org>
Pass struct cxl_port * and struct cxl_dport * to the cxl_aer_* trace events
instead of a plain struct device * derived at the caller. The trace event
helpers then derive the right strings for Endpoints, Switch Ports, Root
Ports, and RCH Downstream Ports consistently across the CPER and native AER
paths.
The unified cxl_aer_* events keep "memdev" as the legacy field (Endpoint
events populate it with the memdev name; non-Endpoint events emit
memdev="") and add new "port" and "dport" string fields populated for all
CXL device classes. Updated userspace can key off "port" and "dport"
without a parallel set of events.
Remove the separate cxl_port_aer_uncorrectable_error and
cxl_port_aer_correctable_error trace events. All CXL AER events now use the
unified cxl_aer_* events with port and dport fields.
Rework cxl_cper_handle_prot_err() to use find_cxl_port_by_dev() and the
unified trace helpers, replacing the per-port-type branching and
bus_find_device() memdev lookup.
The TP_printk format string places "port=%s dport=%s" between "memdev=%s"
and "host=%s", changing the text-mode field order from the pre-patch output.
This does not affect consumers such as rasdaemon that use libtraceevent to
parse fields by name rather than by fixed text position.
For non-Endpoint events (Switch Port, Root Port, RCH Dport), "memdev" is
empty and "port"/"dport" carry the topology information.
CPER: trace firmware-supplied protocol errors even when the host device is
unbound; the record is self-contained in ras_cap and reads no MMIO. Keep
the host lock only to serialize the dport lookup against teardown.
Below are examples of the different CXL devices' error trace logs
after this patch:
---------------------
| CXL RP - 0C:00.0 |
---------------------
|
---------------------
| CXL USP - 0D:00.0 |
---------------------
|
--------------------
| CXL DSP - 0E:00.0 |
--------------------
|
---------------------
| CXL EP - 0F:00.0 |
---------------------
Root Port:
cxl_aer_correctable_error: memdev= port=port1 dport=0000:0c:00.0 \
host=pci0000:0c serial=0: status: 'Memory Data ECC Error'
cxl_aer_uncorrectable_error: memdev= port=port1 dport=0000:0c:00.0 \
host=pci0000:0c serial=0: status: 'Cache Address Parity Error' \
first_error: 'Cache Address Parity Error'
Upstream Switch Port:
cxl_aer_correctable_error: memdev= port=port2 dport= host=0000:0d:00.0 \
serial=0: status: 'Memory Data ECC Error'
UCE NA - Upstream Switch Port UCE's are handled in the portdrv driver's
PCI AER callbacks that are not CXL aware.
Downstream Switch Port:
cxl_aer_correctable_error: memdev= port=port2 dport=0000:0e:00.0 \
host=0000:0d:00.0 serial=0: status: 'Memory Data ECC Error'
cxl_aer_uncorrectable_error: memdev= port=port2 dport=0000:0e:00.0 \
host=0000:0d:00.0 serial=0: status: 'Cache Address Parity Error' \
first_error: 'Cache Address Parity Error'
RCH Downstream Port (RCD attached under a host bridge, no switch):
cxl_aer_correctable_error: memdev= port=root0 dport=pci0000:0c \
host=pci0000:0c serial=0: status: 'Memory Data ECC Error'
cxl_aer_uncorrectable_error: memdev= port=root0 dport=pci0000:0c \
host=pci0000:0c serial=0: status: 'Cache Address Parity Error' \
first_error: 'Cache Address Parity Error'
For RCH topologies, both correctable and uncorrectable protocol errors
were previously traced against the memdev via the cxl_aer_* events with
memdev populated. They now emit memdev="" with the RCH Downstream Port
carried in the "dport" field (dport->dport_dev, the host bridge) and the
host bridge in "host". Consumers that keyed RCH errors off "memdev" must
key off "dport" instead.
Endpoint:
cxl_aer_uncorrectable_error: memdev=mem1 port=endpoint4 dport= \
host=0000:0f:00.0 serial=0: status: 'Cache Address Parity Error' \
first_error: 'Cache Address Parity Error'
cxl_aer_correctable_error: memdev=mem1 port=endpoint4 dport= host=0000:0f:00.0 \
serial=0: status: 'Memory Data ECC Error'
Co-developed-by: Terry Bowman <terry.bowman@amd.com>
Signed-off-by: Terry Bowman <terry.bowman@amd.com>
Signed-off-by: Dan Williams <djbw@kernel.org>
---
Changes in v19->v20:
- Remove host->driver gate. Not required for dev_name() call
Changes in v18->v19:
- Drop redundant device lock in cxl_cper_handle_prot_err(); the port
reference already keeps the object alive and no RAS iomap is accessed.
- Swap order in series with ("PCI: Cache PCI DSN into pci_dev->dsn during
probe")
- Add review-by for DaveJ
Changes in v17->v18:
- Consolidate double find_cxl_port_by_dev() in cxl_cper_handle_prot_err()
- Add comment noting dport is NULL for Endpoint and Upstream Port devices
- Add cxl_trace_* helpers
- Add CPER refactor
Changes in v16->v17:
- Replace cxlds->serial with pci_get_dsn()
- Change 'memdev' to 'device' (Dan)
- Updated Commit message
Changes in v15->v16:
- Add Dan's review-by
- Incorporate Dan's comment into commit message:
"Add the serial number at the end to preserve compatibility with
libtraceevent parsing of the parameters."
Changes in v14->v15:
- Update commit message.
- Moved cxl_handle_ras/cxl_handle_cor_ras() changes to future patch (terry)
Changes in v13->v14:
- Update commit headline (Bjorn)
Changes in v12->v13:
- Added Dave Jiang's review-by
Changes in v11 -> v12:
- Correct parameters to call trace_cxl_aer_correctable_error()
- Add reviewed-by for Jonathan and Shiju
Changes in v10->v11:
- Updated CE and UCE trace routines to maintain consistent TP_Struct ABI
and unchanged TP_printk() logging.
---
drivers/cxl/core/core.h | 8 +--
drivers/cxl/core/ras.c | 137 +++++++++++++------------------------
drivers/cxl/core/ras_rch.c | 3 +-
drivers/cxl/core/trace.c | 35 ++++++++++
drivers/cxl/core/trace.h | 91 +++++++-----------------
drivers/cxl/cxlmem.h | 7 ++
6 files changed, 122 insertions(+), 159 deletions(-)
diff --git a/drivers/cxl/core/core.h b/drivers/cxl/core/core.h
index 4f6b4702deb38..5509521a2034d 100644
--- a/drivers/cxl/core/core.h
+++ b/drivers/cxl/core/core.h
@@ -188,11 +188,11 @@ static inline struct device *dport_to_host(struct cxl_dport *dport)
void cxl_ras_init(void);
void cxl_ras_exit(void);
bool cxl_handle_ras(struct cxl_port *port, struct cxl_dport *dport,
- void __iomem *ras_base);
+ void __iomem *ras_base, u64 serial);
void cxl_do_recovery(struct pci_dev *pdev, struct cxl_port *port,
struct cxl_dport *dport);
void cxl_handle_cor_ras(struct cxl_port *port, struct cxl_dport *dport,
- void __iomem *ras_base);
+ void __iomem *ras_base, u64 serial);
void cxl_dport_map_rch_aer(struct cxl_dport *dport);
void cxl_disable_rch_root_ints(struct cxl_dport *dport);
void cxl_handle_rdport_errors(struct pci_dev *pdev);
@@ -202,14 +202,14 @@ void devm_cxl_dport_ras_setup(struct cxl_dport *dport);
static inline void cxl_ras_init(void) { }
static inline void cxl_ras_exit(void) { }
static inline bool cxl_handle_ras(struct cxl_port *port, struct cxl_dport *dport,
- void __iomem *ras_base)
+ void __iomem *ras_base, u64 serial)
{
return false;
}
static inline void cxl_do_recovery(struct pci_dev *pdev, struct cxl_port *port,
struct cxl_dport *dport) { }
static inline void cxl_handle_cor_ras(struct cxl_port *port, struct cxl_dport *dport,
- void __iomem *ras_base) { }
+ void __iomem *ras_base, u64 serial) { }
static inline void cxl_dport_map_rch_aer(struct cxl_dport *dport) { }
static inline void cxl_disable_rch_root_ints(struct cxl_dport *dport) { }
static inline void cxl_handle_rdport_errors(struct pci_dev *pdev) { }
diff --git a/drivers/cxl/core/ras.c b/drivers/cxl/core/ras.c
index bf479fac08565..fac37b6fd882f 100644
--- a/drivers/cxl/core/ras.c
+++ b/drivers/cxl/core/ras.c
@@ -12,69 +12,37 @@
static_assert(CXL_HEADERLOG_TRACE_SIZE_U32 == 128,
"rasdaemon ABI requires exactly 128 u32s");
-static void cxl_cper_trace_corr_port_prot_err(struct pci_dev *pdev,
- struct cxl_ras_capability_regs ras_cap)
-{
- u32 status = ras_cap.cor_status & ~ras_cap.cor_mask;
-
- trace_cxl_port_aer_correctable_error(&pdev->dev, status);
-}
-
-static void cxl_cper_trace_uncorr_port_prot_err(struct pci_dev *pdev,
- struct cxl_ras_capability_regs ras_cap)
-{
- u32 hl[CXL_HEADERLOG_TRACE_SIZE_U32] = {};
- u32 status = ras_cap.uncor_status & ~ras_cap.uncor_mask;
- u32 fe;
-
- if (hweight32(status) > 1)
- fe = BIT(FIELD_GET(CXL_RAS_CAP_CONTROL_FE_MASK,
- ras_cap.cap_control));
- else
- fe = status;
-
- memcpy(hl, ras_cap.header_log, CXL_HEADERLOG_SIZE);
- trace_cxl_port_aer_uncorrectable_error(&pdev->dev, status, fe, hl);
-}
-
-static void cxl_cper_trace_corr_prot_err(struct cxl_memdev *cxlmd,
- struct cxl_ras_capability_regs ras_cap)
-{
- u32 status = ras_cap.cor_status & ~ras_cap.cor_mask;
-
- trace_cxl_aer_correctable_error(cxlmd, status);
-}
-
-static void
-cxl_cper_trace_uncorr_prot_err(struct cxl_memdev *cxlmd,
- struct cxl_ras_capability_regs ras_cap)
+static void cxl_cper_trace_uncorr_prot_err(struct cxl_port *port, struct cxl_dport *dport,
+ u64 serial, struct cxl_ras_capability_regs *ras_cap)
{
u32 hl[CXL_HEADERLOG_TRACE_SIZE_U32] = {};
- u32 status = ras_cap.uncor_status & ~ras_cap.uncor_mask;
+ u32 status = ras_cap->uncor_status & ~ras_cap->uncor_mask;
u32 fe;
if (hweight32(status) > 1)
fe = BIT(FIELD_GET(CXL_RAS_CAP_CONTROL_FE_MASK,
- ras_cap.cap_control));
+ ras_cap->cap_control));
else
fe = status;
/*
- * ras_cap.header_log[] holds CXL_HEADERLOG_SIZE_U32 (16) hardware
+ * ras_cap->header_log[] holds CXL_HEADERLOG_SIZE_U32 (16) hardware
* dwords. Copy them into the front of a zero-filled
* CXL_HEADERLOG_TRACE_SIZE_U32 (128) u32 staging buffer so the trace
* event memcpy sees a full 512-byte source and the userspace ABI
* (rasdaemon) is preserved.
*/
- memcpy(hl, ras_cap.header_log, CXL_HEADERLOG_SIZE);
- trace_cxl_aer_uncorrectable_error(cxlmd, status, fe, hl);
+ memcpy(hl, ras_cap->header_log, CXL_HEADERLOG_SIZE);
+ trace_cxl_aer_uncorrectable_error(port, dport, status, fe,
+ hl, serial);
}
-static int match_memdev_by_parent(struct device *dev, const void *uport)
+static void cxl_cper_trace_corr_prot_err(struct cxl_port *port, struct cxl_dport *dport,
+ u64 serial, struct cxl_ras_capability_regs *ras_cap)
{
- if (is_cxl_memdev(dev) && dev->parent == uport)
- return 1;
- return 0;
+ u32 status = ras_cap->cor_status & ~ras_cap->cor_mask;
+
+ trace_cxl_aer_correctable_error(port, dport, status, serial);
}
/**
@@ -108,44 +76,42 @@ static struct cxl_port *find_cxl_port_by_dev(struct device *dev, struct cxl_dpor
void cxl_cper_handle_prot_err(struct cxl_cper_prot_err_work_data *data)
{
+ struct cxl_dport *dport;
+ struct device *host;
unsigned int devfn = PCI_DEVFN(data->prot_err.agent_addr.device,
data->prot_err.agent_addr.function);
- struct pci_dev *pdev __free(pci_dev_put) =
- pci_get_domain_bus_and_slot(data->prot_err.agent_addr.segment,
- data->prot_err.agent_addr.bus,
- devfn);
- struct cxl_memdev *cxlmd;
- int port_type;
-
- if (!pdev)
+ struct pci_dev *pdev __free(pci_dev_put) = pci_get_domain_bus_and_slot(
+ data->prot_err.agent_addr.segment, data->prot_err.agent_addr.bus, devfn);
+ if (!pdev) {
+ pr_err_ratelimited("Failed to find CPER device in CXL topology\n");
return;
+ }
- port_type = pci_pcie_type(pdev);
- if (port_type == PCI_EXP_TYPE_ROOT_PORT ||
- port_type == PCI_EXP_TYPE_DOWNSTREAM ||
- port_type == PCI_EXP_TYPE_UPSTREAM) {
- if (data->severity == AER_CORRECTABLE)
- cxl_cper_trace_corr_port_prot_err(pdev, data->ras_cap);
- else
- cxl_cper_trace_uncorr_port_prot_err(pdev, data->ras_cap);
-
+ struct cxl_port *port __free(put_cxl_port) = find_cxl_port_by_dev(&pdev->dev, NULL);
+ if (!port) {
+ dev_err_ratelimited(&pdev->dev,
+ "Failed to find parent port device in CXL topology\n");
return;
}
- guard(device)(&pdev->dev);
- if (!pdev->dev.driver)
- return;
+ host = is_cxl_root(port) ? port->uport_dev : &port->dev;
- struct device *mem_dev __free(put_device) = bus_find_device(
- &cxl_bus_type, NULL, pdev, match_memdev_by_parent);
- if (!mem_dev)
- return;
+ /*
+ * Lock host to serialize the dport lookup/trace against dport teardown.
+ * Don't gate on host->driver as the CPER record is FW provided and needs
+ * no device MMIO.
+ */
+ guard(device)(host);
+
+ /* dport is NULL for Endpoint and Upstream Port devices */
+ dport = cxl_find_dport_by_dev(port, &pdev->dev);
- cxlmd = to_cxl_memdev(mem_dev);
if (data->severity == AER_CORRECTABLE)
- cxl_cper_trace_corr_prot_err(cxlmd, data->ras_cap);
+ cxl_cper_trace_corr_prot_err(port, dport, pdev->dsn,
+ &data->ras_cap);
else
- cxl_cper_trace_uncorr_prot_err(cxlmd, data->ras_cap);
+ cxl_cper_trace_uncorr_prot_err(port, dport, pdev->dsn,
+ &data->ras_cap);
}
EXPORT_SYMBOL_GPL(cxl_cper_handle_prot_err);
@@ -232,14 +198,15 @@ void cxl_do_recovery(struct pci_dev *pdev, struct cxl_port *port, struct cxl_dpo
if (!ras_base)
panic("CXL: UCE with unmapped RAS registers");
- if (cxl_handle_ras(port, dport, ras_base))
+ if (cxl_handle_ras(port, dport, ras_base, pdev->dsn))
panic("CXL cachemem error");
dev_dbg(&pdev->dev,
"CXL UCE signaled but no CXL RAS status bits set\n");
}
-void cxl_handle_cor_ras(struct cxl_port *port, struct cxl_dport *dport, void __iomem *ras_base)
+void cxl_handle_cor_ras(struct cxl_port *port, struct cxl_dport *dport,
+ void __iomem *ras_base, u64 serial)
{
void __iomem *addr;
u32 status;
@@ -251,12 +218,7 @@ void cxl_handle_cor_ras(struct cxl_port *port, struct cxl_dport *dport, void __i
status = readl(addr);
if (status & CXL_RAS_CORRECTABLE_STATUS_MASK) {
writel(status & CXL_RAS_CORRECTABLE_STATUS_MASK, addr);
- if (is_cxl_endpoint(port))
- trace_cxl_aer_correctable_error(to_cxl_memdev(port->uport_dev), status);
- else if (dport)
- trace_cxl_port_aer_correctable_error(dport->dport_dev, status);
- else
- trace_cxl_port_aer_correctable_error(port->uport_dev, status);
+ trace_cxl_aer_correctable_error(port, dport, status, serial);
}
}
@@ -281,7 +243,8 @@ static void header_log_copy(void __iomem *ras_base, u32 *log)
* Log the state of the RAS status registers and prepare them to log the
* next error status. Return 1 if reset needed.
*/
-bool cxl_handle_ras(struct cxl_port *port, struct cxl_dport *dport, void __iomem *ras_base)
+bool cxl_handle_ras(struct cxl_port *port, struct cxl_dport *dport,
+ void __iomem *ras_base, u64 serial)
{
u32 hl[CXL_HEADERLOG_TRACE_SIZE_U32] = {};
void __iomem *addr;
@@ -308,12 +271,7 @@ bool cxl_handle_ras(struct cxl_port *port, struct cxl_dport *dport, void __iomem
}
header_log_copy(ras_base, hl);
- if (is_cxl_endpoint(port))
- trace_cxl_aer_uncorrectable_error(to_cxl_memdev(port->uport_dev), status, fe, hl);
- else if (dport)
- trace_cxl_port_aer_uncorrectable_error(dport->dport_dev, status, fe, hl);
- else
- trace_cxl_port_aer_uncorrectable_error(port->uport_dev, status, fe, hl);
+ trace_cxl_aer_uncorrectable_error(port, dport, status, fe, hl, serial);
writel(status & CXL_RAS_UNCORRECTABLE_STATUS_MASK, addr);
@@ -354,7 +312,8 @@ pci_ers_result_t cxl_pci_error_detected(struct pci_dev *pdev,
* block cannot attribute the error to CXL and must not panic.
* The switch cases below then handle AER recovery.
*/
- ue = cxl_handle_ras(port, NULL, to_ras_base(port, NULL));
+ ue = cxl_handle_ras(port, NULL, to_ras_base(port, NULL),
+ pdev->dsn);
}
/*
@@ -386,7 +345,7 @@ static void cxl_handle_proto_error(struct pci_dev *pdev, struct cxl_port *port,
struct cxl_dport *dport, int severity)
{
if (severity == AER_CORRECTABLE)
- cxl_handle_cor_ras(port, dport, to_ras_base(port, dport));
+ cxl_handle_cor_ras(port, dport, to_ras_base(port, dport), pdev->dsn);
else
cxl_do_recovery(pdev, port, dport);
}
diff --git a/drivers/cxl/core/ras_rch.c b/drivers/cxl/core/ras_rch.c
index bc6cf50fcc926..3684d1d829e1e 100644
--- a/drivers/cxl/core/ras_rch.c
+++ b/drivers/cxl/core/ras_rch.c
@@ -123,7 +123,8 @@ void cxl_handle_rdport_errors(struct pci_dev *pdev)
*/
if (aer_regs.cor_status & ~aer_regs.cor_mask) {
pci_print_aer(pdev, AER_CORRECTABLE, &aer_regs);
- cxl_handle_cor_ras(port, dport, to_ras_base(port, dport));
+ cxl_handle_cor_ras(port, dport, to_ras_base(port, dport),
+ pdev->dsn);
}
if (uncor_status) {
diff --git a/drivers/cxl/core/trace.c b/drivers/cxl/core/trace.c
index 7f2a9dd0d0e3f..df42d119c53dd 100644
--- a/drivers/cxl/core/trace.c
+++ b/drivers/cxl/core/trace.c
@@ -2,7 +2,42 @@
/* Copyright(c) 2022 Intel Corporation. All rights reserved. */
#include <cxl.h>
+#include <cxlmem.h>
#include "core.h"
+const char *cxl_trace_memdev_name(struct cxl_port *port)
+{
+ if (is_cxl_endpoint(port)) {
+ struct cxl_memdev *cxlmd = to_cxl_memdev(port->uport_dev);
+
+ return dev_name(&cxlmd->dev);
+ }
+
+ return "";
+}
+
+const char *cxl_trace_host_name(struct cxl_port *port)
+{
+ if (is_cxl_endpoint(port)) {
+ struct cxl_memdev *cxlmd = to_cxl_memdev(port->uport_dev);
+
+ return dev_name(cxlmd->dev.parent);
+ }
+
+ return dev_name(port->uport_dev);
+}
+
+const char *cxl_trace_port_name(struct cxl_port *port)
+{
+ return dev_name(&port->dev);
+}
+
+const char *cxl_trace_dport_name(struct cxl_dport *dport)
+{
+ if (dport)
+ return dev_name(dport->dport_dev);
+ return "";
+}
+
#define CREATE_TRACE_POINTS
#include "trace.h"
diff --git a/drivers/cxl/core/trace.h b/drivers/cxl/core/trace.h
index c379d60047fc8..9f7fad8d77a54 100644
--- a/drivers/cxl/core/trace.h
+++ b/drivers/cxl/core/trace.h
@@ -48,44 +48,15 @@
{ CXL_RAS_UC_IDE_RX_ERR, "IDE Rx Error" } \
)
-TRACE_EVENT(cxl_port_aer_uncorrectable_error,
- TP_PROTO(struct device *dev, u32 status, u32 fe, u32 *hl),
- TP_ARGS(dev, status, fe, hl),
- TP_STRUCT__entry(
- __string(device, dev_name(dev))
- __string(host, dev_name(dev->parent))
- __field(u32, status)
- __field(u32, first_error)
- __array(u32, header_log, CXL_HEADERLOG_TRACE_SIZE_U32)
- ),
- TP_fast_assign(
- __assign_str(device);
- __assign_str(host);
- __entry->status = status;
- __entry->first_error = fe;
- /*
- * Embed headerlog data for user app retrieval and parsing,
- * but no need to print in the trace buffer. Only
- * CXL_HEADERLOG_SIZE_U32 (16) dwords are hardware data;
- * the remaining entries preserve the 512-byte ABI layout
- * rasdaemon depends on and are zero-filled by the caller.
- */
- memcpy(__entry->header_log, hl,
- CXL_HEADERLOG_TRACE_SIZE_U32 * sizeof(u32));
- ),
- TP_printk("device=%s host=%s status: '%s' first_error: '%s'",
- __get_str(device), __get_str(host),
- show_uc_errs(__entry->status),
- show_uc_errs(__entry->first_error)
- )
-);
-
TRACE_EVENT(cxl_aer_uncorrectable_error,
- TP_PROTO(const struct cxl_memdev *cxlmd, u32 status, u32 fe, u32 *hl),
- TP_ARGS(cxlmd, status, fe, hl),
+ TP_PROTO(struct cxl_port *port, struct cxl_dport *dport,
+ u32 status, u32 fe, u32 *hl, u64 serial),
+ TP_ARGS(port, dport, status, fe, hl, serial),
TP_STRUCT__entry(
- __string(memdev, dev_name(&cxlmd->dev))
- __string(host, dev_name(cxlmd->dev.parent))
+ __string(memdev, cxl_trace_memdev_name(port))
+ __string(port, cxl_trace_port_name(port))
+ __string(dport, cxl_trace_dport_name(dport))
+ __string(host, cxl_trace_host_name(port))
__field(u64, serial)
__field(u32, status)
__field(u32, first_error)
@@ -93,8 +64,10 @@ TRACE_EVENT(cxl_aer_uncorrectable_error,
),
TP_fast_assign(
__assign_str(memdev);
+ __assign_str(port);
+ __assign_str(dport);
__assign_str(host);
- __entry->serial = cxlmd->cxlds->serial;
+ __entry->serial = serial;
__entry->status = status;
__entry->first_error = fe;
/*
@@ -107,8 +80,9 @@ TRACE_EVENT(cxl_aer_uncorrectable_error,
memcpy(__entry->header_log, hl,
CXL_HEADERLOG_TRACE_SIZE_U32 * sizeof(u32));
),
- TP_printk("memdev=%s host=%s serial=%llu: status: '%s' first_error: '%s'",
- __get_str(memdev), __get_str(host), __entry->serial,
+ TP_printk("memdev=%s port=%s dport=%s host=%s serial=%llu: status: '%s' first_error: '%s'",
+ __get_str(memdev), __get_str(port), __get_str(dport),
+ __get_str(host), __entry->serial,
show_uc_errs(__entry->status),
show_uc_errs(__entry->first_error)
)
@@ -132,42 +106,29 @@ TRACE_EVENT(cxl_aer_uncorrectable_error,
{ CXL_RAS_CE_PHYS_LAYER_ERR, "Received Error From Physical Layer" } \
)
-TRACE_EVENT(cxl_port_aer_correctable_error,
- TP_PROTO(struct device *dev, u32 status),
- TP_ARGS(dev, status),
- TP_STRUCT__entry(
- __string(device, dev_name(dev))
- __string(host, dev_name(dev->parent))
- __field(u32, status)
- ),
- TP_fast_assign(
- __assign_str(device);
- __assign_str(host);
- __entry->status = status;
- ),
- TP_printk("device=%s host=%s status='%s'",
- __get_str(device), __get_str(host),
- show_ce_errs(__entry->status)
- )
-);
-
TRACE_EVENT(cxl_aer_correctable_error,
- TP_PROTO(const struct cxl_memdev *cxlmd, u32 status),
- TP_ARGS(cxlmd, status),
+ TP_PROTO(struct cxl_port *port, struct cxl_dport *dport,
+ u32 status, u64 serial),
+ TP_ARGS(port, dport, status, serial),
TP_STRUCT__entry(
- __string(memdev, dev_name(&cxlmd->dev))
- __string(host, dev_name(cxlmd->dev.parent))
+ __string(memdev, cxl_trace_memdev_name(port))
+ __string(port, cxl_trace_port_name(port))
+ __string(dport, cxl_trace_dport_name(dport))
+ __string(host, cxl_trace_host_name(port))
__field(u64, serial)
__field(u32, status)
),
TP_fast_assign(
__assign_str(memdev);
+ __assign_str(port);
+ __assign_str(dport);
__assign_str(host);
- __entry->serial = cxlmd->cxlds->serial;
+ __entry->serial = serial;
__entry->status = status;
),
- TP_printk("memdev=%s host=%s serial=%llu: status: '%s'",
- __get_str(memdev), __get_str(host), __entry->serial,
+ TP_printk("memdev=%s port=%s dport=%s host=%s serial=%llu: status: '%s'",
+ __get_str(memdev), __get_str(port), __get_str(dport),
+ __get_str(host), __entry->serial,
show_ce_errs(__entry->status)
)
);
diff --git a/drivers/cxl/cxlmem.h b/drivers/cxl/cxlmem.h
index c401e3a1af06f..c6368ddae4a21 100644
--- a/drivers/cxl/cxlmem.h
+++ b/drivers/cxl/cxlmem.h
@@ -125,6 +125,13 @@ static inline int cxl_memdev_attach_region(struct cxl_memdev *cxlmd)
#endif
struct cxl_memdev *devm_cxl_add_classdev(struct cxl_dev_state *cxlds);
+
+/* trace-event helpers */
+const char *cxl_trace_memdev_name(struct cxl_port *port);
+const char *cxl_trace_host_name(struct cxl_port *port);
+const char *cxl_trace_port_name(struct cxl_port *port);
+const char *cxl_trace_dport_name(struct cxl_dport *dport);
+
struct cxl_memdev *__devm_cxl_add_memdev(struct cxl_dev_state *cxlds,
const struct cxl_memdev_attach *attach);
int devm_cxl_sanitize_setup_notifier(struct device *host,
--
2.34.1
^ permalink raw reply related [flat|nested] 17+ messages in thread* [PATCH v20 8/9] PCI/CXL: Mask/Unmask CXL protocol errors
2026-09-02 13:39 [PATCH v20 0/9] Enable CXL PCIe Port Protocol Error handling and logging Terry Bowman
` (6 preceding siblings ...)
2026-09-02 13:39 ` [PATCH v20 7/9] cxl: Add port and dport identifiers to CXL AER trace events Terry Bowman
@ 2026-09-02 13:39 ` Terry Bowman
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-02 13:39 ` [PATCH v20 9/9] Documentation: cxl: Document CXL protocol error handling Terry Bowman
8 siblings, 1 reply; 17+ messages in thread
From: Terry Bowman @ 2026-09-02 13:39 UTC (permalink / raw)
To: Jonathan Cameron, Dave Jiang, Alison Schofield, Vishal Verma,
Davidlohr Bueso, Bjorn Helgaas, Dan Williams, Rafael J . Wysocki,
Jonathan Corbet, linux-cxl
Cc: Tony Luck, Borislav Petkov, Hanjun Guo, Mauro Carvalho Chehab,
Shuai Xue, Len Brown, Ira Weiny, Li Ming, Shuah Khan,
Ben Cheatham, Richard Cheng, Robert Richter, Lukas Wunner,
linux-pci, linux-acpi, linux-doc, linux-kernel
CXL protocol errors must be unmasked to be reported. Add
pci_aer_mask_internal_errors() as the symmetric counterpart to
pci_aer_unmask_internal_errors() and export both for cxl_core.
Unmask CXL internal errors in cxl_dport_map_ras() and
devm_cxl_port_ras_setup() after the RAS register block is
successfully mapped. Register a devm action to restore the mask
on teardown. The unmask/mask helpers gate on dev_is_pci() and
pcie_aer_is_native() internally so callers need no special-casing.
Remove the dev_is_pci(dport->dport_dev) guard in
devm_cxl_dport_rch_ras_setup(). On RCH systems dport->dport_dev
is the pci_host_bridge device which is not on pci_bus_type, so
this guard blocked real hardware. The caller already gates on
dport->rch.
Co-developed-by: Dan Williams <djbw@kernel.org>
Signed-off-by: Dan Williams <djbw@kernel.org>
Signed-off-by: Terry Bowman <terry.bowman@amd.com>
Reviewed-by: Dave Jiang <dave.jiang@intel.com>
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
Reviewed-by: Alison Schofield <alison.schofield@intel.com>
---
Changes in v19 -> v20:
- None
Changes in v18 -> v19:
- Reworded commit message to be more concise.
- Use pci_clear_and_set_config_dword() in pci_aer_mask_internal_errors()
- Add review-by for DaveJ and Jonathan
Changes in v17->v18:
- Make cxl_unmask_proto_interrupts() and cxl_mask_proto_interrupts() static
- Remove dev_is_pci() guard from devm_cxl_dport_rch_ras_setup(); the guard
blocked real RCH hardware because pci_host_bridge is not on pci_bus_type
Changes in v16->v17:
- Drop redundant cxl_mask_proto_interrupts() calls from unregister_port()
and cxl_dport_remove(); the devres action registered alongside the unmask
is the sole mask path.
- Update title
- Remove unnecessary check for aer_capabilities
- Gate cxl_unmask_proto_interrupts() on pcie_aer_is_native()
- Add pci_aer_mask_internal_errors() and cxl_mask_proto_interrupts()
- Only unmask on successful cxl_map_component_regs()
- NULL-check @dev in cxl_{un,}mask_proto_interrupts()
- Drop static and declare in core/core.h
Change in v15 -> v16:
- None
Change in v14 -> v15:
- None
Changes in v13->v14:
- Update commit title's prefix (Bjorn)
Changes in v12->v13:
- Add dev and dev_is_pci() NULL checks in cxl_unmask_proto_interrupts() (Terry)
- Add Dave Jiang's and Ben's review-by
Changes in v11->v12:
- None
---
drivers/cxl/core/ras.c | 73 +++++++++++++++++++++++++++++++----
drivers/pci/pcie/aer.c | 22 +++++++++--
include/linux/aer.h | 2 +
tools/testing/cxl/Kbuild | 1 +
tools/testing/cxl/test/mock.c | 12 ++++++
5 files changed, 99 insertions(+), 11 deletions(-)
diff --git a/drivers/cxl/core/ras.c b/drivers/cxl/core/ras.c
index fac37b6fd882f..f1e05d240059b 100644
--- a/drivers/cxl/core/ras.c
+++ b/drivers/cxl/core/ras.c
@@ -124,16 +124,64 @@ static void cxl_cper_prot_err_work_fn(struct work_struct *work)
}
static DECLARE_WORK(cxl_cper_prot_err_work, cxl_cper_prot_err_work_fn);
+static void cxl_unmask_proto_interrupts(struct device *dev)
+{
+ struct pci_dev *pdev;
+
+ if (!dev || !dev_is_pci(dev))
+ return;
+
+ pdev = to_pci_dev(dev);
+ if (!pcie_aer_is_native(pdev))
+ return;
+
+ pci_aer_unmask_internal_errors(pdev);
+}
+
+static void cxl_mask_proto_interrupts(struct device *dev)
+{
+ struct pci_dev *pdev;
+
+ if (!dev || !dev_is_pci(dev))
+ return;
+
+ pdev = to_pci_dev(dev);
+ if (!pcie_aer_is_native(pdev))
+ return;
+
+ pci_aer_mask_internal_errors(pdev);
+}
+
+static void cxl_mask_proto_irqs(void *dev)
+{
+ cxl_mask_proto_interrupts(dev);
+}
+
static void cxl_dport_map_ras(struct cxl_dport *dport)
{
struct cxl_register_map *map = &dport->reg_map;
struct device *dev = dport->dport_dev;
- if (!map->component_map.ras.valid)
+ if (!map->component_map.ras.valid) {
dev_dbg(dev, "RAS registers not found\n");
- else if (cxl_map_component_regs(map, &dport->regs.component,
- BIT(CXL_CM_CAP_CAP_ID_RAS)))
+ return;
+ }
+
+ if (cxl_map_component_regs(map, &dport->regs.component,
+ BIT(CXL_CM_CAP_CAP_ID_RAS))) {
dev_dbg(dev, "Failed to map RAS capability.\n");
+ return;
+ }
+
+ if (!dev_is_pci(dev))
+ return;
+
+ cxl_unmask_proto_interrupts(dev);
+ if (devm_add_action_or_reset(dport_to_host(dport),
+ cxl_mask_proto_irqs, dev)) {
+ dev_warn(dev, "failed to defer CXL proto-irq mask; CXL protocol error reporting disabled\n");
+ dport->regs.component.ras = NULL;
+ }
}
/**
@@ -150,9 +198,6 @@ void devm_cxl_dport_rch_ras_setup(struct cxl_dport *dport)
{
struct pci_host_bridge *host_bridge;
- if (!dev_is_pci(dport->dport_dev))
- return;
-
devm_cxl_dport_ras_setup(dport);
host_bridge = to_pci_host_bridge(dport->dport_dev);
@@ -167,6 +212,7 @@ EXPORT_SYMBOL_NS_GPL(devm_cxl_dport_rch_ras_setup, "CXL");
void devm_cxl_port_ras_setup(struct cxl_port *port)
{
struct cxl_register_map *map = &port->reg_map;
+ struct device *dev;
if (!map->component_map.ras.valid) {
dev_dbg(&port->dev, "RAS registers not found\n");
@@ -175,8 +221,21 @@ void devm_cxl_port_ras_setup(struct cxl_port *port)
map->host = &port->dev;
if (cxl_map_component_regs(map, &port->regs,
- BIT(CXL_CM_CAP_CAP_ID_RAS)))
+ BIT(CXL_CM_CAP_CAP_ID_RAS))) {
dev_dbg(&port->dev, "Failed to map RAS capability\n");
+ return;
+ }
+
+ dev = is_cxl_endpoint(port) ? port->uport_dev->parent : port->uport_dev;
+ if (!dev_is_pci(dev))
+ return;
+
+ cxl_unmask_proto_interrupts(dev);
+ if (devm_add_action_or_reset(&port->dev, cxl_mask_proto_irqs, dev)) {
+ dev_warn(&port->dev,
+ "failed to defer CXL proto-irq mask; CXL protocol error reporting disabled\n");
+ port->regs.ras = NULL;
+ }
}
EXPORT_SYMBOL_NS_GPL(devm_cxl_port_ras_setup, "CXL");
diff --git a/drivers/pci/pcie/aer.c b/drivers/pci/pcie/aer.c
index 8c998cffa89e4..1b182b9cc9553 100644
--- a/drivers/pci/pcie/aer.c
+++ b/drivers/pci/pcie/aer.c
@@ -1281,12 +1281,26 @@ void pci_aer_unmask_internal_errors(struct pci_dev *dev)
mask &= ~PCI_ERR_COR_INTERNAL;
pci_write_config_dword(dev, aer + PCI_ERR_COR_MASK, mask);
}
+EXPORT_SYMBOL_FOR_MODULES(pci_aer_unmask_internal_errors, "cxl_core");
-/*
- * Internal errors are too device-specific to enable generally, however for CXL
- * their behavior is standardized for conveying CXL protocol errors.
+/**
+ * pci_aer_mask_internal_errors - mask internal errors
+ * @dev: pointer to the pci_dev data structure
+ *
+ * Mask internal errors in the Uncorrectable and Correctable Error
+ * Mask registers.
+ *
+ * Note: AER must be enabled and supported by the device which must be
+ * checked in advance, e.g. with pcie_aer_is_native().
*/
-EXPORT_SYMBOL_FOR_MODULES(pci_aer_unmask_internal_errors, "cxl_core");
+void pci_aer_mask_internal_errors(struct pci_dev *dev)
+{
+ pci_clear_and_set_config_dword(dev, dev->aer_cap + PCI_ERR_UNCOR_MASK,
+ 0, PCI_ERR_UNC_INTN);
+ pci_clear_and_set_config_dword(dev, dev->aer_cap + PCI_ERR_COR_MASK,
+ 0, PCI_ERR_COR_INTERNAL);
+}
+EXPORT_SYMBOL_FOR_MODULES(pci_aer_mask_internal_errors, "cxl_core");
/**
* pci_aer_handle_error - handle logging error into an event log
diff --git a/include/linux/aer.h b/include/linux/aer.h
index 8eba3192e2d15..b3657b80564b9 100644
--- a/include/linux/aer.h
+++ b/include/linux/aer.h
@@ -58,6 +58,7 @@ struct aer_capability_regs {
int pci_aer_clear_nonfatal_status(struct pci_dev *dev);
int pcie_aer_is_native(struct pci_dev *dev);
void pci_aer_unmask_internal_errors(struct pci_dev *dev);
+void pci_aer_mask_internal_errors(struct pci_dev *dev);
#else
static inline int pci_aer_clear_nonfatal_status(struct pci_dev *dev)
{
@@ -65,6 +66,7 @@ static inline int pci_aer_clear_nonfatal_status(struct pci_dev *dev)
}
static inline int pcie_aer_is_native(struct pci_dev *dev) { return 0; }
static inline void pci_aer_unmask_internal_errors(struct pci_dev *dev) { }
+static inline void pci_aer_mask_internal_errors(struct pci_dev *dev) { }
#endif
#ifdef CONFIG_CXL_RAS
diff --git a/tools/testing/cxl/Kbuild b/tools/testing/cxl/Kbuild
index 2be1df80fcc93..957945201f04d 100644
--- a/tools/testing/cxl/Kbuild
+++ b/tools/testing/cxl/Kbuild
@@ -6,6 +6,7 @@ ldflags-y += --wrap=acpi_pci_find_root
ldflags-y += --wrap=nvdimm_bus_register
ldflags-y += --wrap=cxl_await_media_ready
ldflags-y += --wrap=devm_cxl_add_rch_dport
+ldflags-y += --wrap=devm_cxl_dport_rch_ras_setup
ldflags-y += --wrap=cxl_endpoint_parse_cdat
ldflags-y += --wrap=devm_cxl_endpoint_decoders_setup
ldflags-y += --wrap=hmat_get_extended_linear_cache_size
diff --git a/tools/testing/cxl/test/mock.c b/tools/testing/cxl/test/mock.c
index 6454b868b122c..5ad3243da8d29 100644
--- a/tools/testing/cxl/test/mock.c
+++ b/tools/testing/cxl/test/mock.c
@@ -220,6 +220,18 @@ struct cxl_dport *__wrap_devm_cxl_add_rch_dport(struct cxl_port *port,
}
EXPORT_SYMBOL_NS_GPL(__wrap_devm_cxl_add_rch_dport, "CXL");
+void __wrap_devm_cxl_dport_rch_ras_setup(struct cxl_dport *dport)
+{
+ int index;
+ struct cxl_mock_ops *ops = get_cxl_mock_ops(&index);
+
+ if (!ops || !ops->is_mock_port(dport->dport_dev))
+ devm_cxl_dport_rch_ras_setup(dport);
+
+ put_cxl_mock_ops(index);
+}
+EXPORT_SYMBOL_NS_GPL(__wrap_devm_cxl_dport_rch_ras_setup, "CXL");
+
void __wrap_cxl_endpoint_parse_cdat(struct cxl_port *port)
{
int index;
--
2.34.1
^ permalink raw reply related [flat|nested] 17+ messages in thread* Re: [PATCH v20 8/9] PCI/CXL: Mask/Unmask CXL protocol errors
2026-09-02 13:39 ` [PATCH v20 8/9] PCI/CXL: Mask/Unmask CXL protocol errors Terry Bowman
@ 2026-09-02 20:57 ` Cheatham, Benjamin
0 siblings, 0 replies; 17+ messages in thread
From: Cheatham, Benjamin @ 2026-09-02 20:57 UTC (permalink / raw)
To: Terry Bowman, Jonathan Cameron, Dave Jiang, Alison Schofield,
Vishal Verma, Davidlohr Bueso, Bjorn Helgaas, Dan Williams,
Rafael J . Wysocki, Jonathan Corbet, linux-cxl
Cc: Tony Luck, Borislav Petkov, Hanjun Guo, Mauro Carvalho Chehab,
Shuai Xue, Len Brown, Ira Weiny, Li Ming, Shuah Khan,
Richard Cheng, Robert Richter, Lukas Wunner, linux-pci,
linux-acpi, linux-doc, linux-kernel
On 9/2/2026 8:39 AM, Terry Bowman wrote:
> CXL protocol errors must be unmasked to be reported. Add
> pci_aer_mask_internal_errors() as the symmetric counterpart to
> pci_aer_unmask_internal_errors() and export both for cxl_core.
>
> Unmask CXL internal errors in cxl_dport_map_ras() and
> devm_cxl_port_ras_setup() after the RAS register block is
> successfully mapped. Register a devm action to restore the mask
> on teardown. The unmask/mask helpers gate on dev_is_pci() and
> pcie_aer_is_native() internally so callers need no special-casing.
>
> Remove the dev_is_pci(dport->dport_dev) guard in
> devm_cxl_dport_rch_ras_setup(). On RCH systems dport->dport_dev
> is the pci_host_bridge device which is not on pci_bus_type, so
> this guard blocked real hardware. The caller already gates on
> dport->rch.
>
> Co-developed-by: Dan Williams <djbw@kernel.org>
> Signed-off-by: Dan Williams <djbw@kernel.org>
> Signed-off-by: Terry Bowman <terry.bowman@amd.com>
> Reviewed-by: Dave Jiang <dave.jiang@intel.com>
> Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
> Reviewed-by: Alison Schofield <alison.schofield@intel.com>
>
> ---
Reviewed-by: Ben Cheatham <benjamin.cheatham@amd.com>
Thanks,
Ben
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH v20 9/9] Documentation: cxl: Document CXL protocol error handling
2026-09-02 13:39 [PATCH v20 0/9] Enable CXL PCIe Port Protocol Error handling and logging Terry Bowman
` (7 preceding siblings ...)
2026-09-02 13:39 ` [PATCH v20 8/9] PCI/CXL: Mask/Unmask CXL protocol errors Terry Bowman
@ 2026-09-02 13:39 ` Terry Bowman
8 siblings, 0 replies; 17+ messages in thread
From: Terry Bowman @ 2026-09-02 13:39 UTC (permalink / raw)
To: Jonathan Cameron, Dave Jiang, Alison Schofield, Vishal Verma,
Davidlohr Bueso, Bjorn Helgaas, Dan Williams, Rafael J . Wysocki,
Jonathan Corbet, linux-cxl
Cc: Tony Luck, Borislav Petkov, Hanjun Guo, Mauro Carvalho Chehab,
Shuai Xue, Len Brown, Ira Weiny, Li Ming, Shuah Khan,
Ben Cheatham, Richard Cheng, Robert Richter, Lukas Wunner,
linux-pci, linux-acpi, linux-doc, linux-kernel
Add Documentation/driver-api/cxl/linux/protocol-error-handling.rst
describing the end-to-end CXL protocol error path: AER ingress, the
AER-CXL kfifo handoff, the cxl_core consumer worker, RCD/RCH special
cases, severity policy, trace events, and a source code map.
This documents the architecture introduced by the preceding patches in
this series.
Assisted-by: Claude:claude-opus-4.8
Signed-off-by: Terry Bowman <terry.bowman@amd.com>
Reviewed-by: Dave Jiang <dave.jiang@intel.com>
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
Reviewed-by: Alison Schofield <alison.schofield@intel.com>
---
Changes in v19 -> v20:
- Document the kfifo-overflow panic policy
- Name the CXL Host Bridge device lock (port->uport_dev) in RCH case
- Reword the cxl_do_recovery() no-UE-bits path
- Fix CPER path description: firmware-first is trace-only, no panic
- Diagram: EP RAS read is channel-state independent, not "unconditional"
- Fix cxl_handle_ras() argument count in the fatal-UCE diagram
- Clarify .error_detected skips RAS read without panic when unmapped
Changes in v18->v19:
- Alignment fixes
- Add RC_END/RCiEP to diagram
- Wrap past 80 columns.
- Update cxl_forward_error() return-value behavior
- Add review-by for Jonathan Cameron
- Document the CPER/firmware-first flow (GHES + ext-log producers,
CPER-CXL kfifo, trace-only consumer)
- Document the fatal EP/RCD UCE flow via the .error_detected
(cxl_pci_error_detected) RAS path, including channel-state handling
- Add an "AER handlers vs RAS handlers" section clarifying the two
handler layers and their relationship
Changes in v17->v18:
- Simplify document for readability (Jonathan)
- Drop historical context that goes stale (Jonathan)
- Shorten ASCII flow diagram (Jonathan)
- Drop manual backtick markup, use automarkup (Jonathan)
- Clarify USP/DSP as single switch component (Dave)
- Fix line wrapping to 80 chars (Jonathan)
---
Documentation/driver-api/cxl/index.rst | 1 +
.../cxl/linux/protocol-error-handling.rst | 441 ++++++++++++++++++
2 files changed, 442 insertions(+)
create mode 100644 Documentation/driver-api/cxl/linux/protocol-error-handling.rst
diff --git a/Documentation/driver-api/cxl/index.rst b/Documentation/driver-api/cxl/index.rst
index 3dfae1d310ca5..6861b2e5726a3 100644
--- a/Documentation/driver-api/cxl/index.rst
+++ b/Documentation/driver-api/cxl/index.rst
@@ -42,6 +42,7 @@ that have impacts on each other. The docs here break up configurations steps.
linux/dax-driver
linux/memory-hotplug
linux/access-coordinates
+ linux/protocol-error-handling
.. toctree::
:maxdepth: 2
diff --git a/Documentation/driver-api/cxl/linux/protocol-error-handling.rst b/Documentation/driver-api/cxl/linux/protocol-error-handling.rst
new file mode 100644
index 0000000000000..1da71d0409a05
--- /dev/null
+++ b/Documentation/driver-api/cxl/linux/protocol-error-handling.rst
@@ -0,0 +1,441 @@
+.. SPDX-License-Identifier: GPL-2.0
+
+==============================
+CXL Protocol Error Handling
+==============================
+
+CXL devices report protocol-layer failures (CXL.cachemem RAS) as PCIe AER
+Internal Errors: PCI_ERR_COR_INTERNAL for correctable events and
+PCI_ERR_UNC_INTN for uncorrectable events. The actual fault information
+lives in CXL RAS capability registers, not in the PCIe AER status registers.
+
+The kernel routes every CXL Internal Error through a producer/consumer
+pipeline shared by all CXL device types: Root Ports, Upstream/Downstream
+Switch Ports, Endpoints, and Restricted CXL Devices (RCDs).
+
+Errors are delivered by one of two mechanisms. On native-AER platforms the
+kernel takes the AER interrupt and reads the CXL RAS registers itself. On
+firmware-first (CPER/GHES) platforms, platform firmware handles the error
+and hands the kernel a CPER record; that path is trace-only. Both converge
+on the same cxl_core RAS handlers.
+
+
+Architecture
+============
+
+Two error planes run side by side:
+
+* The **PCIe AER plane** handles native PCIe errors (receiver overflows,
+ malformed TLPs, completion timeouts, etc.). This includes CXL.io, which
+ is functionally PCIe and reports through native AER status registers.
+* The **CXL protocol error plane** handles CXL.cachemem (CXL.cache and
+ CXL.mem) protocol errors. These have no native AER status; they are
+ signaled as AER Internal Errors, with the fault detail held in the CXL
+ RAS capability registers. The AER core forwards them to cxl_core via a
+ dedicated kfifo; cxl_core reads the CXL RAS registers, emits trace
+ events, and applies recovery/panic policy.
+
+The boundary between the two planes is enforced by is_cxl_error() in
+aer_cxl_vh.c. It checks info->is_cxl, the PCIe device type (Endpoint,
+Root Port, Upstream, or Downstream), and whether the AER status word
+indicates an internal error. RC_END devices are excluded from
+is_cxl_error() because they reach the kfifo via the separate
+cxl_rch_handle_error() path instead.
+
+The pipeline:
+
+1. **Producer** (aer_cxl_vh.c, aer_cxl_rch.c) - AER threaded handler
+ context. Classifies and enqueues a struct cxl_proto_err_work_data
+ into the kfifo.
+2. **Queue** - the AER-CXL kfifo plus a backing work_struct.
+3. **Consumer** (cxl_core/ras.c) - workqueue context. Resolves the CXL
+ port topology and dispatches to CE/UE handlers.
+
+
+AER handlers vs RAS handlers
+============================
+
+Two distinct handler layers cooperate; keeping them separate is central to
+the design:
+
+* **AER handlers** run in PCIe AER context (aer.c, aer_cxl_vh.c,
+ aer_cxl_rch.c). They own the PCIe side: they observe the Internal Error,
+ classify it with is_cxl_error(), and act only as the *producer* - they
+ enqueue a work item into the AER-CXL kfifo. AER handlers never touch the
+ CXL RAS capability registers and, in the normal path, make no
+ recovery/panic decision. The one exception is kfifo overflow: if the
+ AER-CXL kfifo is full, cxl_forward_error() drops a correctable error but
+ panics on a non-correctable one, since a dropped uncorrectable protocol
+ error can no longer be confirmed via CXL RAS (see "Severity policy").
+
+* **RAS handlers** run in cxl_core (cxl_core/ras.c, cxl_core/ras_rch.c).
+ They are the *consumers*: they read the CXL RAS capability registers,
+ emit the CXL trace events, and apply the CE/UCE severity policy (clear
+ correctable status, or panic on an uncorrectable error). A RAS handler
+ is where the actual CXL fault information is decoded, because that
+ information lives in the RAS registers, not in the PCIe AER status word.
+
+The AER handler and the RAS handler are decoupled by the kfifo: the AER
+handler cannot acquire the sleeping port device lock, and the RAS handler
+runs in workqueue context where it can safely take port locks. The
+same RAS decode is reached through three different entry paths:
+
+* the AER-CXL kfifo consumer (native AER, VH and RCH),
+* the pci_error_handlers .error_detected callback (fatal EP/RCD UCE, where
+ no AER status is available), and
+* the CPER-CXL kfifo consumer (firmware-first CPER/GHES, trace-only).
+
+The first two paths share the same decode and severity policy (read the
+RAS registers, clear correctable status, panic on an uncorrectable error).
+They differ only when the RAS registers are unmapped: the AER kfifo path
+reaches cxl_do_recovery() only after a CXL internal error is confirmed and
+so panics on an unmapped RAS block, while the .error_detected path has no
+such prior confirmation and instead skips the read without panicking.
+The CPER path differs: it is firmware-first and trace-only - it emits the
+CXL trace events but does not call cxl_do_recovery() and never panics (see
+the "CPER / firmware-first flow" and "Severity policy" sections below).
+
+
+Topologies
+==========
+
+Virtual Hierarchy (VH)
+----------------------
+
+Standard PCIe topology: Root Port, optional switch (Upstream Port with one
+or more Downstream Ports), and Endpoints. Each component raises Internal
+Errors directly via the Root Port's AER interrupt.
+
+Producer: cxl_forward_error() in aer_cxl_vh.c.
+
+Restricted CXL Host (RCH)
+--------------------------
+
+A Root Complex Event Collector (RCEC) aggregates errors from RCDs attached
+as Root Complex Integrated Endpoints. The AER driver iterates RCDs beneath
+the RCEC via pcie_walk_rcec() and forwards each qualifying device through
+cxl_forward_error() into the same kfifo.
+
+Producer: cxl_forward_error() in aer_cxl_vh.c, called from
+cxl_rch_handle_error_iter() via pcie_walk_rcec().
+
+
+Error flow
+==========
+
+.. code-block:: text
+
+ CXL device raises AER Internal Error
+ (PCI_ERR_COR_INTERNAL or PCI_ERR_UNC_INTN)
+ |
+ v
+ +--------------------------------------+
+ | AER core (aer.c) |
+ | aer_irq() -> aer_isr() |
+ | -> find_source_device() |
+ | -> handle_error_source(dev, info) |
+ +--------------------------------------+
+ |
+ v
+ +--------------------------------------+
+ | handle_error_source() dispatch |
+ | |
+ | 1. cxl_rch_handle_error() |
+ | [always; filters internally. |
+ | RC_END enters the kfifo here |
+ | via pcie_walk_rcec(), NOT via |
+ | is_cxl_error() below] |
+ | |
+ | 2. if is_cxl_error(): |
+ | cxl_forward_error() |
+ | [enqueue to kfifo; EP/RP/USP/ |
+ | DSP only, RC_END excluded] |
+ | |
+ | 3. if cxl_pending && non-CE: |
+ | cxl_proto_err_wait_for_empty() |
+ | [sync drain before recovery] |
+ | |
+ | 4. pci_aer_handle_error() [always] |
+ +--------------------------------------+
+ |
+ (kfifo -> workqueue)
+ |
+ v
+ +--------------------------------------+
+ | __cxl_proto_err_work_fn() consumer |
+ | |
+ | if is_cxl_restricted(pdev): |
+ | cxl_handle_rdport_errors() |
+ | [RCH dport RAS first] |
+ | |
+ | cxl_handle_proto_error() |
+ +--------------------------------------+
+ | |
+ v v
+ +-----------------+ +--------------------+
+ | CE | | UCE |
+ | cxl_handle_ | | cxl_do_recovery() |
+ | cor_ras() | | read RAS status |
+ | trace + clear | | trace + panic |
+ +-----------------+ +--------------------+
+
+cxl_do_recovery() first checks whether the CXL RAS register block is
+mapped. If it is not (to_ras_base() returns NULL), the kernel panics
+immediately without reading any register or emitting a trace event,
+because a signaled UCE cannot be confirmed or cleared. Otherwise it
+reads the CXL RAS uncorrectable status register. If UE bits are set, it
+emits the trace event and panics. If no bits are set (e.g. RAS mapped but
+error already cleared), it logs a debug diagnostic and returns without
+action, allowing the caller to proceed with normal AER recovery.
+
+The consumer reaches cxl_handle_proto_error() only after the topology
+lookups succeed: it returns early (dropping the event) if the parent port
+is not found, if the port host device is unbound, or if the dport is not
+found for a Root Port or Downstream Port. A UCE enqueued to the kfifo can
+therefore be discarded without a panic if the topology is torn down or not
+yet registered between enqueue and consumer execution.
+
+
+Fatal UCE flow for Endpoints and RCDs
+=====================================
+
+For a fatal (AER_FATAL) uncorrectable error, aer_get_device_error_info()
+reads the AER uncorrectable status register only for Root Ports, RC Event
+Collectors, and Downstream Ports; it skips the read for Endpoints and
+Upstream Ports because their link is presumed down. With info->status left
+zero, is_cxl_error() cannot classify the event as a CXL protocol error, so
+it never enters the AER-CXL kfifo. This is a severity/device-type property,
+not an RCH-specific one: it affects every Endpoint (VH Endpoint and RCD
+alike) and every Upstream Port.
+
+Endpoints instead reach the RAS handler through the pci_error_handlers
+.error_detected callback (cxl_pci_error_detected()), which is registered by
+the CXL memdev driver and fires for both VH Endpoints and RCDs. The only
+RCD-specific step is the leading cxl_handle_rdport_errors() call, which
+processes the RCH Downstream Port's RAS registers first; the Endpoint RAS
+read and panic policy that follow are identical for VH and RCH:
+
+.. code-block:: text
+
+ Fatal UCE on Endpoint (VH Endpoint or RCD; link down, no AER status)
+ |
+ v
+ +--------------------------------------+
+ | PCIe core error recovery |
+ | pcie_do_recovery() |
+ | -> report_error_detected() |
+ | -> cxl_pci_error_detected() |
+ | [pci_error_handlers callback in |
+ | cxl_core/ras.c; the RAS handler,|
+ | NOT the AER kfifo path] |
+ +--------------------------------------+
+ |
+ v
+ +--------------------------------------+
+ | cxl_pci_error_detected() |
+ | |
+ | if is_cxl_restricted(pdev): |
+ | cxl_handle_rdport_errors() |
+ | [RCD-only: RCH Dport RAS first] |
+ | |
+ | if port->dev.driver == NULL: |
+ | return DISCONNECT [port unbound] |
+ | |
+ | cxl_handle_ras(port, NULL, |
+ | to_ras_base(...), |
+ | pdev->dsn) |
+ | [EP RAS read, independent of |
+ | channel state (not skipped for |
+ | io_normal); dead link |
+ | readl()==0xFFFFFFFF sets all UE |
+ | bits -> panic] |
+ | |
+ | if ue: panic("CXL cachemem error") |
+ | |
+ | else switch (channel state): |
+ | io_normal -> CAN_RECOVER |
+ | io_frozen -> release driver, |
+ | NEED_RESET |
+ | perm_failure -> DISCONNECT |
+ +--------------------------------------+
+
+This path handles both severities: a non-fatal EP UCE arrives as
+pci_channel_io_normal and a fatal EP UCE as pci_channel_io_frozen. Either
+way the CXL RAS read runs first, so a real CXL.mem UCE always panics; only
+when no CXL UE bit is set (or RAS is unmapped) does the channel state drive
+ordinary AER recovery. Endpoint unbind therefore does not depend on the
+AER-status-to-RAS coupling that the kfifo path relies on.
+
+Upstream Ports bound to portdrv have no such .error_detected callback and
+fall back to standard AER recovery - this is a known limitation.
+
+
+CPER / firmware-first flow
+==========================
+
+On firmware-first platforms, CXL protocol errors are delivered by platform
+firmware as an ACPI CPER record (CPER_SEC_CXL_PROT_ERR) instead of a native
+AER interrupt. These records already contain a snapshot of the CXL RAS
+capability registers, so the RAS handler does not read hardware; it only
+emits trace events. Firmware-first is therefore trace-only and never
+panics or drives recovery - the platform owns the recovery decision.
+
+CPER protocol-error records are delivered by the GHES/APEI firmware-first
+path:
+
+* **GHES/APEI** (ghes.c) - the common firmware-first path. Its producer,
+ cxl_cper_post_prot_err(), enqueues a struct cxl_cper_prot_err_work_data
+ into a dedicated CPER-CXL kfifo (cxl_cper_prot_err_fifo, depth 8) and
+ schedules the cxl_core consumer work item.
+
+.. code-block:: text
+
+ Platform firmware CPER record (CPER_SEC_CXL_PROT_ERR)
+ |
+ v
+ +----------------------+
+ | GHES/APEI (ghes.c) |
+ | ghes_do_proc() |
+ | cxl_cper_post_ |
+ | prot_err() |
+ | kfifo_put(CPER-CXL) |
+ | schedule_work() |
+ +----------------------+
+ |
+ v
+ +----------------------+
+ | CPER-CXL kfifo |
+ | + work_struct |
+ +----------------------+
+ |
+ v
+ +----------------------+
+ | cxl_cper_prot_err_ |
+ | work_fn() consumer |
+ | (cxl_core/ras.c) |
+ | drain kfifo -> |
+ +----------------------+
+ |
+ v
+ +--------------------------------+
+ | cxl_cper_handle_prot_err() |
+ | pci_get_domain_bus_and_slot() |
+ | find_cxl_port_by_dev() |
+ | cxl_find_dport_by_dev() |
+ | |
+ | if CE: trace correctable |
+ | else: trace uncorrectable |
+ | [trace-only; no panic, |
+ | no cxl_do_recovery()] |
+ +--------------------------------+
+
+The consumer work item is registered with GHES via
+cxl_cper_register_prot_err_work() when cxl_core loads and torn down with
+cxl_cper_unregister_prot_err_work(), which cancels any pending work and
+resets the kfifo so stale records are not replayed on the next module load.
+
+
+Severity policy
+===============
+
+**CE** - cxl_handle_cor_ras() reads the CXL RAS correctable status register,
+clears set bits, and emits a cxl_aer_correctable_error trace event. No
+recovery action.
+
+**UCE (non-fatal, and fatal on Root Port/Downstream Port)** -
+cxl_do_recovery() reads the CXL RAS uncorrectable status register. If UE
+bits are set, the kernel panics. If the CXL RAS register block is not
+mapped (to_ras_base() returns NULL), cxl_do_recovery() panics before any
+register read and emits no trace event, since the UCE cannot be confirmed.
+CXL.cachemem traffic cannot be safely recovered once an uncorrectable error
+is signaled; continuing risks silent data corruption. This panic policy
+applies to the native AER path. On firmware-first (CPER/GHES) platforms the
+CPER handler emits trace events only and does not call cxl_do_recovery().
+
+**kfifo overflow** - if the AER-CXL kfifo is full when the producer
+(cxl_forward_error()) tries to enqueue, a correctable error is dropped with
+a ratelimited log message, but a non-correctable error triggers an
+immediate panic() from the AER handler itself. A dropped uncorrectable
+protocol error can no longer be confirmed or recovered via CXL RAS and may
+signal lost cache coherency over HDM memory in active use, so it is
+collapsed to the same conservative outcome as a confirmed UCE. This is the
+only case where an AER handler, rather than a RAS handler, makes the panic
+decision.
+
+**Fatal UCE on EP/USP** - A fatal event brings the link down, so the AER
+core reads no AER status and is_cxl_error() cannot enqueue the event to the
+kfifo. Endpoints and RCDs are instead handled through the
+pci_error_handlers .error_detected callback (cxl_pci_error_detected()),
+which reads the CXL RAS registers when they are mapped and panics on any UE
+bit. If the RAS registers are unmapped the read is skipped without a panic,
+because this path has no prior confirmation that the error is CXL internal.
+Upstream Ports bound to portdrv fall back to standard AER recovery - a known
+limitation. See "Fatal UCE flow for Endpoints and RCDs" above for the full
+path and channel-state handling.
+
+
+RCH special case
+================
+
+When the consumer sees is_cxl_restricted(pdev), it calls
+cxl_handle_rdport_errors() first to process the RCH Downstream Port's RAS
+registers (accessed via RCRB, not standard config space). It then
+continues to process the RCD Endpoint's own RAS registers via the common
+path. Both register blocks are checked because errors can appear in either
+independently.
+
+cxl_handle_rdport_errors() acquires the device lock on port->uport_dev (the
+CXL Host Bridge) internally. Callers must not hold it.
+
+
+Trace events
+============
+
+Two trace events cover all device types and both the native AER and
+CPER/GHES firmware-first paths:
+
+* cxl_aer_correctable_error
+* cxl_aer_uncorrectable_error
+
+Fields:
+
+* ``memdev`` - memdev name for Endpoints; empty for non-Endpoints.
+* ``port`` - CXL port device name.
+* ``dport`` - Downstream Port device name; empty when not applicable.
+* ``host`` - parent host bridge or uport device name.
+* ``serial`` - PCI Device Serial Number from pdev->dsn (cached at
+ enumeration; no config-space read in the error path).
+
+
+Interrupt masking
+=================
+
+CXL Internal Error bits (PCI_ERR_UNC_INTN and PCI_ERR_COR_INTERNAL) are
+unmasked in the AER capability only after the CXL RAS register block is
+successfully mapped. A devm teardown action restores the mask when the
+port or dport is removed, ensuring clean state after driver removal.
+
+
+Source files
+============
+
+.. list-table::
+ :header-rows: 1
+
+ * - File
+ - Role
+ * - drivers/pci/pcie/aer.c
+ - AER core; IRQ, dispatch
+ * - drivers/pci/pcie/aer_cxl_vh.c
+ - VH AER producer; AER-CXL kfifo
+ * - drivers/pci/pcie/aer_cxl_rch.c
+ - RCH AER dispatch; RCEC walk
+ * - drivers/cxl/core/ras.c
+ - RAS handlers; AER-CXL and CPER-CXL kfifo consumers;
+ .error_detected callback (cxl_pci_error_detected)
+ * - drivers/cxl/core/ras_rch.c
+ - RCH Downstream Port RAS handling
+ * - drivers/acpi/apei/ghes.c
+ - CPER/GHES producer; CPER-CXL kfifo
+ * - drivers/acpi/apei/ghes_helpers.c
+ - CPER record validation and work-data setup helpers
--
2.34.1
^ permalink raw reply related [flat|nested] 17+ messages in thread