From: "Bowman, Terry" <terry.bowman@amd.com>
To: Richard Cheng <icheng@nvidia.com>
Cc: Jonathan Cameron <jic23@kernel.org>,
Dave Jiang <dave.jiang@intel.com>,
Alison Schofield <alison.schofield@intel.com>,
Vishal Verma <vishal.l.verma@intel.com>,
Davidlohr Bueso <dave@stgolabs.net>,
Bjorn Helgaas <bhelgaas@google.com>,
"Rafael J . Wysocki" <rafael@kernel.org>,
Jonathan Corbet <corbet@lwn.net>,
linux-cxl@vger.kernel.org, Tony Luck <tony.luck@intel.com>,
Borislav Petkov <bp@alien8.de>, Hanjun Guo <guohanjun@huawei.com>,
Mauro Carvalho Chehab <mchehab@kernel.org>,
Shuai Xue <xueshuai@linux.alibaba.com>,
Len Brown <lenb@kernel.org>, Ira Weiny <iweiny@kernel.org>,
Li Ming <ming.li@zohomail.com>,
Shuah Khan <skhan@linuxfoundation.org>,
Ben Cheatham <Benjamin.Cheatham@amd.com>,
Robert Richter <rrichter@amd.com>,
linux-pci@vger.kernel.org, linux-acpi@vger.kernel.org,
linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH v19 12/14] cxl: Add port and dport identifiers to CXL AER trace events
Date: Tue, 4 Aug 2026 08:46:37 -0500 [thread overview]
Message-ID: <c8129007-3fda-44a9-a43b-e8e9ab84f2c4@amd.com> (raw)
In-Reply-To: <anGYFW9WC40gaAjp@MWDK4CY14F>
On 8/4/2026 2:56 AM, Richard Cheng wrote:
> On Mon, Aug 03, 2026 at 05:18:08PM +0800, Terry Bowman wrote:
>> From: Dan Williams <djbw@kernel.org>
>>
>> Pass struct cxl_port * and struct cxl_dport * to the cxl_aer_*
>> trace events instead of a plain struct device * derived at the
>> caller. The trace event helpers then derive the right strings for
>> Endpoints, Switch Ports, Root Ports, and RCH Downstream Ports
>> consistently across the CPER and native AER paths.
>>
>> The unified cxl_aer_* events keep "memdev" as the legacy field
>> (endpoint events populate it with the memdev name; non-endpoint
>> events emit memdev="") and add new "port" and "dport" string fields
>> populated for all CXL device classes. Updated userspace can key
>> off "port" and "dport" without a parallel set of events.
>>
>> Remove the separate cxl_port_aer_uncorrectable_error and
>> cxl_port_aer_correctable_error trace events. All CXL AER events now
>> use the unified cxl_aer_* events with port and dport fields.
>>
>> Rework cxl_cper_handle_prot_err() to use find_cxl_port_by_dev() and
>> the unified trace helpers, replacing the per-port-type branching and
>> bus_find_device() memdev lookup.
>>
>> The TP_printk format string places "port=%s dport=%s" between
>> "memdev=%s" and "host=%s", changing the text-mode field order from
>> the pre-patch output. This does not affect consumers such as
>> rasdaemon that use libtraceevent to parse fields by name rather than
>> by fixed text position.
>>
>> For non-Endpoint events (Switch Port, Root Port, RCH Dport),
>> "memdev" is empty and "port"/"dport" carry the topology information.
>>
>> Below are examples of the different CXL devices' error trace logs
>> after this patch:
>>
>> ---------------------
>> | CXL RP - 0C:00.0 |
>> ---------------------
>> |
>> ---------------------
>> | CXL USP - 0D:00.0 |
>> ---------------------
>> |
>> --------------------
>> | CXL DSP - 0E:00.0 |
>> --------------------
>> |
>> ---------------------
>> | CXL EP - 0F:00.0 |
>> ---------------------
>>
>> Root Port:
>> cxl_aer_correctable_error: memdev= port=port1 dport=0000:0c:00.0 \
>> host=pci0000:0c serial=0: status: 'Memory Data ECC Error'
>>
>> cxl_aer_uncorrectable_error: memdev= port=port1 dport=0000:0c:00.0 \
>> host=pci0000:0c serial=0: status: 'Cache Address Parity Error' \
>> first_error: 'Cache Address Parity Error'
>>
>> Upstream Switch Port:
>> cxl_aer_correctable_error: memdev= port=port2 dport= host=0000:0d:00.0 \
>> serial=0: status: 'Memory Data ECC Error'
>>
>> UCE NA - Upstream Switch Port UCE's are handled in the portdrv driver's
>> PCI AER callbacks that are not CXL aware.
>>
>> Downstream Switch Port:
>> cxl_aer_correctable_error: memdev= port=port2 dport=0000:0e:00.0 \
>> host=0000:0d:00.0 serial=0: status: 'Memory Data ECC Error'
>>
>> cxl_aer_uncorrectable_error: memdev= port=port2 dport=0000:0e:00.0 \
>> host=0000:0d:00.0 serial=0: status: 'Cache Address Parity Error' \
>> first_error: 'Cache Address Parity Error'
>>
>> Endpoint:
>> cxl_aer_uncorrectable_error: memdev=mem1 port=endpoint4 dport= \
>> host=0000:0f:00.0 serial=0: status: 'Cache Address Parity Error' \
>> first_error: 'Cache Address Parity Error'
>>
>> cxl_aer_correctable_error: memdev=mem1 port=endpoint4 dport= host=0000:0f:00.0 \
>> serial=0: status: 'Memory Data ECC Error'
>>
>> Co-developed-by: Terry Bowman <terry.bowman@amd.com>
>> Signed-off-by: Terry Bowman <terry.bowman@amd.com>
>> Signed-off-by: Dan Williams <djbw@kernel.org>
>> Reviewed-by: Dave Jiang <dave.jiang@intel.com>
>>
>> ---
>>
>> Changes in v18->v19:
>> - Drop redundant device lock in cxl_cper_handle_prot_err(); the port
>> reference already keeps the object alive and no RAS iomap is accessed.
>> - Swap order in series with ("PCI: Cache PCI DSN into pci_dev->dsn during
>> probe")
>> - Add review-by for DaveJ
>>
>> Changes in v17->v18:
>> - Consolidate double find_cxl_port_by_dev() in cxl_cper_handle_prot_err()
>> - Add comment noting dport is NULL for Endpoint and Upstream Port devices
>> - Add cxl_trace_* helpers
>> - Add CPER refactor
>>
>> Changes in v16->v17:
>> - Replace cxlds->serial with pci_get_dsn()
>> - Change 'memdev' to 'device' (Dan)
>> - Updated Commit message
>>
>> Changes in v15->v16:
>> - Add Dan's review-by
>> - Incorporate Dan's comment into commit message:
>> "Add the serial number at the end to preserve compatibility with
>> libtraceevent parsing of the parameters."
>>
>> Changes in v14->v15:
>> - Update commit message.
>> - Moved cxl_handle_ras/cxl_handle_cor_ras() changes to future patch (terry)
>>
>> Changes in v13->v14:
>> - Update commit headline (Bjorn)
>>
>> Changes in v12->v13:
>> - Added Dave Jiang's review-by
>>
>> Changes in v11 -> v12:
>> - Correct parameters to call trace_cxl_aer_correctable_error()
>> - Add reviewed-by for Jonathan and Shiju
>>
>> Changes in v10->v11:
>> - Updated CE and UCE trace routines to maintain consistent TP_Struct ABI
>> and unchanged TP_printk() logging.
>> ---
>> drivers/cxl/core/core.h | 8 +--
>> drivers/cxl/core/ras.c | 129 +++++++++++--------------------------
>> drivers/cxl/core/ras_rch.c | 3 +-
>> drivers/cxl/core/trace.c | 35 ++++++++++
>> drivers/cxl/core/trace.h | 91 ++++++++------------------
>> drivers/cxl/cxlmem.h | 7 ++
>> 6 files changed, 113 insertions(+), 160 deletions(-)
>>
>> diff --git a/drivers/cxl/core/core.h b/drivers/cxl/core/core.h
>> index 5ca1275fd8f35..a55a4e409feda 100644
>> --- a/drivers/cxl/core/core.h
>> +++ b/drivers/cxl/core/core.h
>> @@ -186,11 +186,11 @@ static inline struct device *dport_to_host(struct cxl_dport *dport)
>> void cxl_ras_init(void);
>> void cxl_ras_exit(void);
>> bool cxl_handle_ras(struct cxl_port *port, struct cxl_dport *dport,
>> - void __iomem *ras_base);
>> + void __iomem *ras_base, u64 serial);
>> void cxl_do_recovery(struct pci_dev *pdev, struct cxl_port *port,
>> struct cxl_dport *dport);
>> void cxl_handle_cor_ras(struct cxl_port *port, struct cxl_dport *dport,
>> - void __iomem *ras_base);
>> + void __iomem *ras_base, u64 serial);
>> void cxl_dport_map_rch_aer(struct cxl_dport *dport);
>> void cxl_disable_rch_root_ints(struct cxl_dport *dport);
>> void cxl_handle_rdport_errors(struct pci_dev *pdev);
>> @@ -200,14 +200,14 @@ void devm_cxl_dport_ras_setup(struct cxl_dport *dport);
>> static inline void cxl_ras_init(void) { }
>> static inline void cxl_ras_exit(void) { }
>> static inline bool cxl_handle_ras(struct cxl_port *port, struct cxl_dport *dport,
>> - void __iomem *ras_base)
>> + void __iomem *ras_base, u64 serial)
>> {
>> return false;
>> }
>> static inline void cxl_do_recovery(struct pci_dev *pdev, struct cxl_port *port,
>> struct cxl_dport *dport) { }
>> static inline void cxl_handle_cor_ras(struct cxl_port *port, struct cxl_dport *dport,
>> - void __iomem *ras_base) { }
>> + void __iomem *ras_base, u64 serial) { }
>> static inline void cxl_dport_map_rch_aer(struct cxl_dport *dport) { }
>> static inline void cxl_disable_rch_root_ints(struct cxl_dport *dport) { }
>> static inline void cxl_handle_rdport_errors(struct pci_dev *pdev) { }
>> diff --git a/drivers/cxl/core/ras.c b/drivers/cxl/core/ras.c
>> index 5183b3c532952..0254b7ec64c30 100644
>> --- a/drivers/cxl/core/ras.c
>> +++ b/drivers/cxl/core/ras.c
>> @@ -12,69 +12,37 @@
>> static_assert(CXL_HEADERLOG_TRACE_SIZE_U32 == 128,
>> "rasdaemon ABI requires exactly 128 u32s");
>>
>> -static void cxl_cper_trace_corr_port_prot_err(struct pci_dev *pdev,
>> - struct cxl_ras_capability_regs ras_cap)
>> -{
>> - u32 status = ras_cap.cor_status & ~ras_cap.cor_mask;
>> -
>> - trace_cxl_port_aer_correctable_error(&pdev->dev, status);
>> -}
>> -
>> -static void cxl_cper_trace_uncorr_port_prot_err(struct pci_dev *pdev,
>> - struct cxl_ras_capability_regs ras_cap)
>> +static void cxl_cper_trace_uncorr_prot_err(struct cxl_port *port, struct cxl_dport *dport,
>> + u64 serial, struct cxl_ras_capability_regs *ras_cap)
>> {
>> u32 hl[CXL_HEADERLOG_TRACE_SIZE_U32] = {};
>> - u32 status = ras_cap.uncor_status & ~ras_cap.uncor_mask;
>> + u32 status = ras_cap->uncor_status & ~ras_cap->uncor_mask;
>> u32 fe;
>>
>> if (hweight32(status) > 1)
>> fe = BIT(FIELD_GET(CXL_RAS_CAP_CONTROL_FE_MASK,
>> - ras_cap.cap_control));
>> - else
>> - fe = status;
>> -
>> - memcpy(hl, ras_cap.header_log, CXL_HEADERLOG_SIZE);
>> - trace_cxl_port_aer_uncorrectable_error(&pdev->dev, status, fe, hl);
>> -}
>> -
>> -static void cxl_cper_trace_corr_prot_err(struct cxl_memdev *cxlmd,
>> - struct cxl_ras_capability_regs ras_cap)
>> -{
>> - u32 status = ras_cap.cor_status & ~ras_cap.cor_mask;
>> -
>> - trace_cxl_aer_correctable_error(cxlmd, status);
>> -}
>> -
>> -static void
>> -cxl_cper_trace_uncorr_prot_err(struct cxl_memdev *cxlmd,
>> - struct cxl_ras_capability_regs ras_cap)
>> -{
>> - u32 hl[CXL_HEADERLOG_TRACE_SIZE_U32] = {};
>> - u32 status = ras_cap.uncor_status & ~ras_cap.uncor_mask;
>> - u32 fe;
>> -
>> - if (hweight32(status) > 1)
>> - fe = BIT(FIELD_GET(CXL_RAS_CAP_CONTROL_FE_MASK,
>> - ras_cap.cap_control));
>> + ras_cap->cap_control));
>> else
>> fe = status;
>>
>> /*
>> - * ras_cap.header_log[] holds CXL_HEADERLOG_SIZE_U32 (16) hardware
>> + * ras_cap->header_log[] holds CXL_HEADERLOG_SIZE_U32 (16) hardware
>> * dwords. Copy them into the front of a zero-filled
>> * CXL_HEADERLOG_TRACE_SIZE_U32 (128) u32 staging buffer so the trace
>> * event memcpy sees a full 512-byte source and the userspace ABI
>> * (rasdaemon) is preserved.
>> */
>> - memcpy(hl, ras_cap.header_log, CXL_HEADERLOG_SIZE);
>> - trace_cxl_aer_uncorrectable_error(cxlmd, status, fe, hl);
>> + memcpy(hl, ras_cap->header_log, CXL_HEADERLOG_SIZE);
>> + trace_cxl_aer_uncorrectable_error(port, dport, status, fe,
>> + hl, serial);
>> }
>>
>> -static int match_memdev_by_parent(struct device *dev, const void *uport)
>> +static void cxl_cper_trace_corr_prot_err(struct cxl_port *port, struct cxl_dport *dport,
>> + u64 serial, struct cxl_ras_capability_regs *ras_cap)
>> {
>> - if (is_cxl_memdev(dev) && dev->parent == uport)
>> - return 1;
>> - return 0;
>> + u32 status = ras_cap->cor_status & ~ras_cap->cor_mask;
>> +
>> + trace_cxl_aer_correctable_error(port, dport, status, serial);
>> }
>>
>> /**
>> @@ -108,44 +76,32 @@ static struct cxl_port *find_cxl_port_by_dev(struct device *dev, struct cxl_dpor
>>
>> void cxl_cper_handle_prot_err(struct cxl_cper_prot_err_work_data *data)
>> {
>> + struct cxl_dport *dport;
>> unsigned int devfn = PCI_DEVFN(data->prot_err.agent_addr.device,
>> data->prot_err.agent_addr.function);
>> - struct pci_dev *pdev __free(pci_dev_put) =
>> - pci_get_domain_bus_and_slot(data->prot_err.agent_addr.segment,
>> - data->prot_err.agent_addr.bus,
>> - devfn);
>> - struct cxl_memdev *cxlmd;
>> - int port_type;
>> -
>> - if (!pdev)
>> - return;
>> -
>> - port_type = pci_pcie_type(pdev);
>> - if (port_type == PCI_EXP_TYPE_ROOT_PORT ||
>> - port_type == PCI_EXP_TYPE_DOWNSTREAM ||
>> - port_type == PCI_EXP_TYPE_UPSTREAM) {
>> - if (data->severity == AER_CORRECTABLE)
>> - cxl_cper_trace_corr_port_prot_err(pdev, data->ras_cap);
>> - else
>> - cxl_cper_trace_uncorr_port_prot_err(pdev, data->ras_cap);
>> -
>> + struct pci_dev *pdev __free(pci_dev_put) = pci_get_domain_bus_and_slot(
>> + data->prot_err.agent_addr.segment, data->prot_err.agent_addr.bus, devfn);
>> + if (!pdev) {
>> + pr_err_ratelimited("Failed to find CPER device in CXL topology\n");
>> return;
>> }
>>
>> - guard(device)(&pdev->dev);
>> - if (!pdev->dev.driver)
>> + struct cxl_port *port __free(put_cxl_port) = find_cxl_port_by_dev(&pdev->dev, NULL);
>> + if (!port) {
>> + dev_err_ratelimited(&pdev->dev,
>> + "Failed to find parent port device in CXL topology\n");
>> return;
>> + }
>>
>> - struct device *mem_dev __free(put_device) = bus_find_device(
>> - &cxl_bus_type, NULL, pdev, match_memdev_by_parent);
>> - if (!mem_dev)
>> - return;
>> + /* dport is NULL for Endpoint and Upstream Port devices */
>> + dport = cxl_find_dport_by_dev(port, &pdev->dev);
>>
>
> Hi Terry,
>
> I have a question here.
>
> Do we need the port device lock here ?
>
> cxl_find_dport_by_dev() is xa_load(), and the free side is serialized by that
> lock. del_dports() has device_lock_assert(&port->dev) and frees the dport which
> in the end resolves to kfree().
>
> __free(put_cxl_port) doesn't cover it, that's a kobject ref on &port->dev,
> so it pins the struct cxl_port but not the dports.
>
> I think cxl_cper_handle_prot_err() might race cxl_detach_ep() calling
> del_dports(), and cxl_trace_dport_name() then does dev_name(dport->dport_dev)
> on freed memory.
>
> I see __cxl_proto_err_work_fn(), cxl_handle_rdport_errors() and
> cxl_pci_error_detected() all take the guard first, should we do the same here?
>
>
> Though I didn't poke this issue out in runtime, I guess it needs
> FW-first, cxl_aer_* enabled, and a concurrent teardown. Not an expert of FW,
> let me knowo if something already rules it out.
>
> Best regards,
> Richard Cheng.
>
>
Hi Richard,
Yes, a port lock is needed for preventing dport being freed here. This could be a
use after free in the path's trace dev_name(). Thanks for pointing out. I'll fix
this in v20 with adding: guard(device)(&port->dev)
-Terry
>> - cxlmd = to_cxl_memdev(mem_dev);
>> if (data->severity == AER_CORRECTABLE)
>> - cxl_cper_trace_corr_prot_err(cxlmd, data->ras_cap);
>> + cxl_cper_trace_corr_prot_err(port, dport, pdev->dsn,
>> + &data->ras_cap);
>> else
>> - cxl_cper_trace_uncorr_prot_err(cxlmd, data->ras_cap);
>> + cxl_cper_trace_uncorr_prot_err(port, dport, pdev->dsn,
>> + &data->ras_cap);
>> }
>> EXPORT_SYMBOL_GPL(cxl_cper_handle_prot_err);
>>
>> @@ -232,14 +188,15 @@ void cxl_do_recovery(struct pci_dev *pdev, struct cxl_port *port, struct cxl_dpo
>> if (!ras_base)
>> panic("CXL: UCE with unmapped RAS registers");
>>
>> - if (cxl_handle_ras(port, dport, ras_base))
>> + if (cxl_handle_ras(port, dport, ras_base, pdev->dsn))
>> panic("CXL cachemem error");
>>
>> dev_dbg(&pdev->dev,
>> "CXL UCE signaled but no CXL RAS status bits set\n");
>> }
>>
>> -void cxl_handle_cor_ras(struct cxl_port *port, struct cxl_dport *dport, void __iomem *ras_base)
>> +void cxl_handle_cor_ras(struct cxl_port *port, struct cxl_dport *dport,
>> + void __iomem *ras_base, u64 serial)
>> {
>> void __iomem *addr;
>> u32 status;
>> @@ -251,12 +208,7 @@ void cxl_handle_cor_ras(struct cxl_port *port, struct cxl_dport *dport, void __i
>> status = readl(addr);
>> if (status & CXL_RAS_CORRECTABLE_STATUS_MASK) {
>> writel(status & CXL_RAS_CORRECTABLE_STATUS_MASK, addr);
>> - if (is_cxl_endpoint(port))
>> - trace_cxl_aer_correctable_error(to_cxl_memdev(port->uport_dev), status);
>> - else if (dport)
>> - trace_cxl_port_aer_correctable_error(dport->dport_dev, status);
>> - else
>> - trace_cxl_port_aer_correctable_error(port->uport_dev, status);
>> + trace_cxl_aer_correctable_error(port, dport, status, serial);
>> }
>> }
>>
>> @@ -281,7 +233,8 @@ static void header_log_copy(void __iomem *ras_base, u32 *log)
>> * Log the state of the RAS status registers and prepare them to log the
>> * next error status. Return 1 if reset needed.
>> */
>> -bool cxl_handle_ras(struct cxl_port *port, struct cxl_dport *dport, void __iomem *ras_base)
>> +bool cxl_handle_ras(struct cxl_port *port, struct cxl_dport *dport,
>> + void __iomem *ras_base, u64 serial)
>> {
>> u32 hl[CXL_HEADERLOG_TRACE_SIZE_U32] = {};
>> void __iomem *addr;
>> @@ -308,12 +261,7 @@ bool cxl_handle_ras(struct cxl_port *port, struct cxl_dport *dport, void __iomem
>> }
>>
>> header_log_copy(ras_base, hl);
>> - if (is_cxl_endpoint(port))
>> - trace_cxl_aer_uncorrectable_error(to_cxl_memdev(port->uport_dev), status, fe, hl);
>> - else if (dport)
>> - trace_cxl_port_aer_uncorrectable_error(dport->dport_dev, status, fe, hl);
>> - else
>> - trace_cxl_port_aer_uncorrectable_error(port->uport_dev, status, fe, hl);
>> + trace_cxl_aer_uncorrectable_error(port, dport, status, fe, hl, serial);
>>
>> writel(status & CXL_RAS_UNCORRECTABLE_STATUS_MASK, addr);
>>
>> @@ -354,7 +302,8 @@ pci_ers_result_t cxl_pci_error_detected(struct pci_dev *pdev,
>> * cases below handle AER recovery for devices without active
>> * CXL.mem traffic.
>> */
>> - ue = cxl_handle_ras(port, NULL, to_ras_base(port, NULL));
>> + ue = cxl_handle_ras(port, NULL, to_ras_base(port, NULL),
>> + pdev->dsn);
>> }
>>
>> /*
>> @@ -386,7 +335,7 @@ static void cxl_handle_proto_error(struct pci_dev *pdev, struct cxl_port *port,
>> struct cxl_dport *dport, int severity)
>> {
>> if (severity == AER_CORRECTABLE)
>> - cxl_handle_cor_ras(port, dport, to_ras_base(port, dport));
>> + cxl_handle_cor_ras(port, dport, to_ras_base(port, dport), pdev->dsn);
>> else
>> cxl_do_recovery(pdev, port, dport);
>> }
>> diff --git a/drivers/cxl/core/ras_rch.c b/drivers/cxl/core/ras_rch.c
>> index a5c62c71060d9..a371174536a8b 100644
>> --- a/drivers/cxl/core/ras_rch.c
>> +++ b/drivers/cxl/core/ras_rch.c
>> @@ -113,7 +113,8 @@ void cxl_handle_rdport_errors(struct pci_dev *pdev)
>> */
>> if (aer_regs.cor_status & ~aer_regs.cor_mask) {
>> pci_print_aer(pdev, AER_CORRECTABLE, &aer_regs);
>> - cxl_handle_cor_ras(port, dport, to_ras_base(port, dport));
>> + cxl_handle_cor_ras(port, dport, to_ras_base(port, dport),
>> + pdev->dsn);
>> }
>>
>> if (aer_regs.uncor_status & ~aer_regs.uncor_mask) {
>> diff --git a/drivers/cxl/core/trace.c b/drivers/cxl/core/trace.c
>> index 7f2a9dd0d0e3f..df42d119c53dd 100644
>> --- a/drivers/cxl/core/trace.c
>> +++ b/drivers/cxl/core/trace.c
>> @@ -2,7 +2,42 @@
>> /* Copyright(c) 2022 Intel Corporation. All rights reserved. */
>>
>> #include <cxl.h>
>> +#include <cxlmem.h>
>> #include "core.h"
>>
>> +const char *cxl_trace_memdev_name(struct cxl_port *port)
>> +{
>> + if (is_cxl_endpoint(port)) {
>> + struct cxl_memdev *cxlmd = to_cxl_memdev(port->uport_dev);
>> +
>> + return dev_name(&cxlmd->dev);
>> + }
>> +
>> + return "";
>> +}
>> +
>> +const char *cxl_trace_host_name(struct cxl_port *port)
>> +{
>> + if (is_cxl_endpoint(port)) {
>> + struct cxl_memdev *cxlmd = to_cxl_memdev(port->uport_dev);
>> +
>> + return dev_name(cxlmd->dev.parent);
>> + }
>> +
>> + return dev_name(port->uport_dev);
>> +}
>> +
>> +const char *cxl_trace_port_name(struct cxl_port *port)
>> +{
>> + return dev_name(&port->dev);
>> +}
>> +
>> +const char *cxl_trace_dport_name(struct cxl_dport *dport)
>> +{
>> + if (dport)
>> + return dev_name(dport->dport_dev);
>> + return "";
>> +}
>> +
>> #define CREATE_TRACE_POINTS
>> #include "trace.h"
>> diff --git a/drivers/cxl/core/trace.h b/drivers/cxl/core/trace.h
>> index d37876096dd7c..910aceb2ca3ab 100644
>> --- a/drivers/cxl/core/trace.h
>> +++ b/drivers/cxl/core/trace.h
>> @@ -48,44 +48,15 @@
>> { CXL_RAS_UC_IDE_RX_ERR, "IDE Rx Error" } \
>> )
>>
>> -TRACE_EVENT(cxl_port_aer_uncorrectable_error,
>> - TP_PROTO(struct device *dev, u32 status, u32 fe, u32 *hl),
>> - TP_ARGS(dev, status, fe, hl),
>> - TP_STRUCT__entry(
>> - __string(device, dev_name(dev))
>> - __string(host, dev_name(dev->parent))
>> - __field(u32, status)
>> - __field(u32, first_error)
>> - __array(u32, header_log, CXL_HEADERLOG_TRACE_SIZE_U32)
>> - ),
>> - TP_fast_assign(
>> - __assign_str(device);
>> - __assign_str(host);
>> - __entry->status = status;
>> - __entry->first_error = fe;
>> - /*
>> - * Embed headerlog data for user app retrieval and parsing,
>> - * but no need to print in the trace buffer. Only
>> - * CXL_HEADERLOG_SIZE_U32 (16) dwords are hardware data;
>> - * the remaining entries preserve the 512-byte ABI layout
>> - * rasdaemon depends on and are zero-filled by the caller.
>> - */
>> - memcpy(__entry->header_log, hl,
>> - CXL_HEADERLOG_TRACE_SIZE_U32 * sizeof(u32));
>> - ),
>> - TP_printk("device=%s host=%s status: '%s' first_error: '%s'",
>> - __get_str(device), __get_str(host),
>> - show_uc_errs(__entry->status),
>> - show_uc_errs(__entry->first_error)
>> - )
>> -);
>> -
>> TRACE_EVENT(cxl_aer_uncorrectable_error,
>> - TP_PROTO(const struct cxl_memdev *cxlmd, u32 status, u32 fe, u32 *hl),
>> - TP_ARGS(cxlmd, status, fe, hl),
>> + TP_PROTO(struct cxl_port *port, struct cxl_dport *dport,
>> + u32 status, u32 fe, u32 *hl, u64 serial),
>> + TP_ARGS(port, dport, status, fe, hl, serial),
>> TP_STRUCT__entry(
>> - __string(memdev, dev_name(&cxlmd->dev))
>> - __string(host, dev_name(cxlmd->dev.parent))
>> + __string(memdev, cxl_trace_memdev_name(port))
>> + __string(port, cxl_trace_port_name(port))
>> + __string(dport, cxl_trace_dport_name(dport))
>> + __string(host, cxl_trace_host_name(port))
>> __field(u64, serial)
>> __field(u32, status)
>> __field(u32, first_error)
>> @@ -93,8 +64,10 @@ TRACE_EVENT(cxl_aer_uncorrectable_error,
>> ),
>> TP_fast_assign(
>> __assign_str(memdev);
>> + __assign_str(port);
>> + __assign_str(dport);
>> __assign_str(host);
>> - __entry->serial = cxlmd->cxlds->serial;
>> + __entry->serial = serial;
>> __entry->status = status;
>> __entry->first_error = fe;
>> /*
>> @@ -107,8 +80,9 @@ TRACE_EVENT(cxl_aer_uncorrectable_error,
>> memcpy(__entry->header_log, hl,
>> CXL_HEADERLOG_TRACE_SIZE_U32 * sizeof(u32));
>> ),
>> - TP_printk("memdev=%s host=%s serial=%lld: status: '%s' first_error: '%s'",
>> - __get_str(memdev), __get_str(host), __entry->serial,
>> + TP_printk("memdev=%s port=%s dport=%s host=%s serial=%lld: status: '%s' first_error: '%s'",
>> + __get_str(memdev), __get_str(port), __get_str(dport),
>> + __get_str(host), __entry->serial,
>> show_uc_errs(__entry->status),
>> show_uc_errs(__entry->first_error)
>> )
>> @@ -132,42 +106,29 @@ TRACE_EVENT(cxl_aer_uncorrectable_error,
>> { CXL_RAS_CE_PHYS_LAYER_ERR, "Received Error From Physical Layer" } \
>> )
>>
>> -TRACE_EVENT(cxl_port_aer_correctable_error,
>> - TP_PROTO(struct device *dev, u32 status),
>> - TP_ARGS(dev, status),
>> - TP_STRUCT__entry(
>> - __string(device, dev_name(dev))
>> - __string(host, dev_name(dev->parent))
>> - __field(u32, status)
>> - ),
>> - TP_fast_assign(
>> - __assign_str(device);
>> - __assign_str(host);
>> - __entry->status = status;
>> - ),
>> - TP_printk("device=%s host=%s status='%s'",
>> - __get_str(device), __get_str(host),
>> - show_ce_errs(__entry->status)
>> - )
>> -);
>> -
>> TRACE_EVENT(cxl_aer_correctable_error,
>> - TP_PROTO(const struct cxl_memdev *cxlmd, u32 status),
>> - TP_ARGS(cxlmd, status),
>> + TP_PROTO(struct cxl_port *port, struct cxl_dport *dport,
>> + u32 status, u64 serial),
>> + TP_ARGS(port, dport, status, serial),
>> TP_STRUCT__entry(
>> - __string(memdev, dev_name(&cxlmd->dev))
>> - __string(host, dev_name(cxlmd->dev.parent))
>> + __string(memdev, cxl_trace_memdev_name(port))
>> + __string(port, cxl_trace_port_name(port))
>> + __string(dport, cxl_trace_dport_name(dport))
>> + __string(host, cxl_trace_host_name(port))
>> __field(u64, serial)
>> __field(u32, status)
>> ),
>> TP_fast_assign(
>> __assign_str(memdev);
>> + __assign_str(port);
>> + __assign_str(dport);
>> __assign_str(host);
>> - __entry->serial = cxlmd->cxlds->serial;
>> + __entry->serial = serial;
>> __entry->status = status;
>> ),
>> - TP_printk("memdev=%s host=%s serial=%lld: status: '%s'",
>> - __get_str(memdev), __get_str(host), __entry->serial,
>> + TP_printk("memdev=%s port=%s dport=%s host=%s serial=%lld: status: '%s'",
>> + __get_str(memdev), __get_str(port), __get_str(dport),
>> + __get_str(host), __entry->serial,
>> show_ce_errs(__entry->status)
>> )
>> );
>> diff --git a/drivers/cxl/cxlmem.h b/drivers/cxl/cxlmem.h
>> index ed419d0c59f2f..f1ef8b78db18a 100644
>> --- a/drivers/cxl/cxlmem.h
>> +++ b/drivers/cxl/cxlmem.h
>> @@ -125,6 +125,13 @@ static inline int cxl_memdev_attach_region(struct cxl_memdev *cxlmd)
>> #endif
>>
>> struct cxl_memdev *devm_cxl_add_classdev(struct cxl_dev_state *cxlds);
>> +
>> +/* trace-event helpers */
>> +const char *cxl_trace_memdev_name(struct cxl_port *port);
>> +const char *cxl_trace_host_name(struct cxl_port *port);
>> +const char *cxl_trace_port_name(struct cxl_port *port);
>> +const char *cxl_trace_dport_name(struct cxl_dport *dport);
>> +
>> struct cxl_memdev *__devm_cxl_add_memdev(struct cxl_dev_state *cxlds,
>> const struct cxl_memdev_attach *attach);
>> int devm_cxl_sanitize_setup_notifier(struct device *host,
>> --
>> 2.34.1
>>
next prev parent reply other threads:[~2026-08-04 13:46 UTC|newest]
Thread overview: 46+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-03 22:17 [PATCH v19 00/14] Enable CXL PCIe Port Protocol Error handling and logging Terry Bowman
2026-08-03 22:17 ` [PATCH v19 01/14] cxl/ras: Fix cxl_rch_get_aer_info() out-of-bounds AER register read Terry Bowman
2026-08-03 22:42 ` sashiko-bot
2026-08-04 16:20 ` Bowman, Terry
2026-08-04 2:10 ` Alison Schofield
2026-08-03 22:17 ` [PATCH v19 02/14] cxl/ras: Fix cxl_rch_get_aer_severity() wrong severity register Terry Bowman
2026-08-03 22:35 ` sashiko-bot
2026-08-04 2:11 ` Alison Schofield
2026-08-03 22:17 ` [PATCH v19 03/14] acpi/apei/ghes: Use raw_spinlock_t for CXL CPER work locks Terry Bowman
2026-08-03 22:39 ` sashiko-bot
2026-08-05 18:41 ` Luck, Tony
2026-08-03 22:18 ` [PATCH v19 04/14] cxl: Tighten CPER kfifo registration API and symbol visibility Terry Bowman
2026-08-03 22:30 ` sashiko-bot
2026-08-04 2:13 ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 05/14] cxl: Rename find_cxl_port() to find_cxl_port_by_dport() Terry Bowman
2026-08-03 22:29 ` sashiko-bot
2026-08-04 2:14 ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 06/14] PCI/AER: Introduce AER-CXL protocol error kfifo Terry Bowman
2026-08-03 22:28 ` sashiko-bot
2026-08-04 8:15 ` Richard Cheng
2026-08-04 14:10 ` Bowman, Terry
2026-08-03 22:18 ` [PATCH v19 07/14] PCI: Establish common CXL Port protocol error flow Terry Bowman
2026-08-03 22:56 ` sashiko-bot
2026-08-03 22:18 ` [PATCH v19 08/14] cxl/ras: Handle RCH correctable and uncorrectable errors in one pass Terry Bowman
2026-08-03 22:29 ` sashiko-bot
2026-08-04 2:16 ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 09/14] cxl/pci: Thread port and dport through RAS handling helpers Terry Bowman
2026-08-03 22:33 ` sashiko-bot
2026-08-04 2:16 ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 10/14] cxl: Update CXL Endpoint AER handler Terry Bowman
2026-08-03 22:40 ` sashiko-bot
2026-08-04 2:17 ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 11/14] PCI: Cache PCI DSN into pci_dev->dsn during probe Terry Bowman
2026-08-03 22:29 ` sashiko-bot
2026-08-04 2:26 ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 12/14] cxl: Add port and dport identifiers to CXL AER trace events Terry Bowman
2026-08-03 22:42 ` sashiko-bot
2026-08-04 2:27 ` Alison Schofield
2026-08-04 7:56 ` Richard Cheng
2026-08-04 13:46 ` Bowman, Terry [this message]
2026-08-03 22:18 ` [PATCH v19 13/14] PCI/CXL: Mask/Unmask CXL protocol errors Terry Bowman
2026-08-03 22:55 ` sashiko-bot
2026-08-04 2:29 ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 14/14] Documentation: cxl: Document CXL protocol error handling Terry Bowman
2026-08-03 22:31 ` sashiko-bot
2026-08-04 2:30 ` Alison Schofield
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=c8129007-3fda-44a9-a43b-e8e9ab84f2c4@amd.com \
--to=terry.bowman@amd.com \
--cc=Benjamin.Cheatham@amd.com \
--cc=alison.schofield@intel.com \
--cc=bhelgaas@google.com \
--cc=bp@alien8.de \
--cc=corbet@lwn.net \
--cc=dave.jiang@intel.com \
--cc=dave@stgolabs.net \
--cc=guohanjun@huawei.com \
--cc=icheng@nvidia.com \
--cc=iweiny@kernel.org \
--cc=jic23@kernel.org \
--cc=lenb@kernel.org \
--cc=linux-acpi@vger.kernel.org \
--cc=linux-cxl@vger.kernel.org \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-pci@vger.kernel.org \
--cc=mchehab@kernel.org \
--cc=ming.li@zohomail.com \
--cc=rafael@kernel.org \
--cc=rrichter@amd.com \
--cc=skhan@linuxfoundation.org \
--cc=tony.luck@intel.com \
--cc=vishal.l.verma@intel.com \
--cc=xueshuai@linux.alibaba.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox