From: "Bowman, Terry" <terry.bowman@amd.com>
To: "Cheatham, Benjamin" <benjamin.cheatham@amd.com>,
Jonathan Cameron <jic23@kernel.org>,
Dave Jiang <dave.jiang@intel.com>,
Alison Schofield <alison.schofield@intel.com>,
Vishal Verma <vishal.l.verma@intel.com>,
Davidlohr Bueso <dave@stgolabs.net>,
Bjorn Helgaas <bhelgaas@google.com>,
Dan Williams <djbw@kernel.org>,
"Rafael J . Wysocki" <rafael@kernel.org>,
Jonathan Corbet <corbet@lwn.net>,
linux-cxl@vger.kernel.org
Cc: Tony Luck <tony.luck@intel.com>, Borislav Petkov <bp@alien8.de>,
Hanjun Guo <guohanjun@huawei.com>,
Mauro Carvalho Chehab <mchehab@kernel.org>,
Shuai Xue <xueshuai@linux.alibaba.com>,
Len Brown <lenb@kernel.org>, Ira Weiny <iweiny@kernel.org>,
Li Ming <ming.li@zohomail.com>,
Shuah Khan <skhan@linuxfoundation.org>,
Richard Cheng <icheng@nvidia.com>,
Robert Richter <rrichter@amd.com>, Lukas Wunner <lukas@wunner.de>,
linux-pci@vger.kernel.org, linux-acpi@vger.kernel.org,
linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH v20 4/9] cxl/pci: Thread port and dport through RAS handling helpers
Date: Wed, 9 Sep 2026 09:42:20 -0500 [thread overview]
Message-ID: <51154495-cddd-4ab2-ae37-2f085fb058df@amd.com> (raw)
In-Reply-To: <2a1ddf3f-64fc-426c-9453-93b9e034dbf8@amd.com>
On 9/2/2026 3:57 PM, Cheatham, Benjamin wrote:
> On 9/2/2026 8:39 AM, Terry Bowman wrote:
>> From: Dan Williams <djbw@kernel.org>
>>
>> The callers of cxl_handle_ras() and cxl_handle_cor_ras() already hold
>> a struct cxl_port * and optionally a struct cxl_dport * for the device
>> being handled. Passing a generic struct device * requires is_cxl_memdev()
>> to distinguish Endpoints from Ports at trace emission time. Threading
>> Port and Downstream Port directly enables is_cxl_endpoint() and explicit
>> dport/port branching for cleaner trace dispatch.
>>
>> Refactor cxl_handle_ras() and cxl_handle_cor_ras() to accept struct
>> cxl_port * and struct cxl_dport * directly. The CXL RAS trace event
>> emission logic is split into three branches: Endpoint events are
>> identified via is_cxl_endpoint() and emit with the memdev, dport events
>> emit with dport->dport_dev, and Upstream Port events fall back to
>> port->uport_dev. This branching is transitional: the follow-on patch
>> ("cxl: Add port and dport identifiers to CXL AER trace events") unifies
>> the trace events on port/dport and removes it.
>>
>> Update cxl_handle_rdport_errors() and cxl_handle_proto_error() to pass
>> Port and Downstream Port to the refactored functions.
>>
>> RCH Downstream Port correctable trace events now report the dport device
>> (dport->dport_dev) as a consequence of threading Port and Downstream
>> Port through the RAS helpers. The following trace event rework ("cxl: Add
>> port and dport identifiers to CXL AER trace events") adds explicit memdev,
>> Port, Downstream Port, and host fields that provide full context for all
>> device types.
>>
>> Co-developed-by: Terry Bowman <terry.bowman@amd.com>
>> Signed-off-by: Terry Bowman <terry.bowman@amd.com>
>> Signed-off-by: Dan Williams <djbw@kernel.org>
>> Reviewed-by: Dave Jiang <dave.jiang@intel.com>
>> Reviewed-by: Alison Schofield <alison.schofield@intel.com>
>>
>> ---
>>
>> Changes in v19 -> v20:
>> - None
>>
>> Changes in v18 -> v19:
>> - Added review-by for DaveJ
>> - Update commit message that three-way trace branching is transitional
>> and removed by the following trace-event unification patch
>>
>> Changes in v17 -> v18:
>> - New patch.
>> ---
>> drivers/cxl/core/core.h | 12 ++++++++----
>> drivers/cxl/core/ras.c | 29 +++++++++++++++--------------
>> drivers/cxl/core/ras_rch.c | 2 +-
>> 3 files changed, 24 insertions(+), 19 deletions(-)
>>
>> diff --git a/drivers/cxl/core/core.h b/drivers/cxl/core/core.h
>> index 645824167f788..4f6b4702deb38 100644
>> --- a/drivers/cxl/core/core.h
>> +++ b/drivers/cxl/core/core.h
>> @@ -187,10 +187,12 @@ static inline struct device *dport_to_host(struct cxl_dport *dport)
>> #ifdef CONFIG_CXL_RAS
>> void cxl_ras_init(void);
>> void cxl_ras_exit(void);
>> -bool cxl_handle_ras(struct device *dev, void __iomem *ras_base);
>> +bool cxl_handle_ras(struct cxl_port *port, struct cxl_dport *dport,
>> + void __iomem *ras_base);
>> void cxl_do_recovery(struct pci_dev *pdev, struct cxl_port *port,
>> struct cxl_dport *dport);
>> -void cxl_handle_cor_ras(struct device *dev, void __iomem *ras_base);
>> +void cxl_handle_cor_ras(struct cxl_port *port, struct cxl_dport *dport,
>> + void __iomem *ras_base);
>> void cxl_dport_map_rch_aer(struct cxl_dport *dport);
>> void cxl_disable_rch_root_ints(struct cxl_dport *dport);
>> void cxl_handle_rdport_errors(struct pci_dev *pdev);
>> @@ -199,13 +201,15 @@ void devm_cxl_dport_ras_setup(struct cxl_dport *dport);
>> #else
>> static inline void cxl_ras_init(void) { }
>> static inline void cxl_ras_exit(void) { }
>> -static inline bool cxl_handle_ras(struct device *dev, void __iomem *ras_base)
>> +static inline bool cxl_handle_ras(struct cxl_port *port, struct cxl_dport *dport,
>> + void __iomem *ras_base)
>> {
>> return false;
>> }
>> static inline void cxl_do_recovery(struct pci_dev *pdev, struct cxl_port *port,
>> struct cxl_dport *dport) { }
>> -static inline void cxl_handle_cor_ras(struct device *dev, void __iomem *ras_base) { }
>> +static inline void cxl_handle_cor_ras(struct cxl_port *port, struct cxl_dport *dport,
>> + void __iomem *ras_base) { }
>> static inline void cxl_dport_map_rch_aer(struct cxl_dport *dport) { }
>> static inline void cxl_disable_rch_root_ints(struct cxl_dport *dport) { }
>> static inline void cxl_handle_rdport_errors(struct pci_dev *pdev) { }
>> diff --git a/drivers/cxl/core/ras.c b/drivers/cxl/core/ras.c
>> index 02c791f149270..e99dcd028b738 100644
>> --- a/drivers/cxl/core/ras.c
>> +++ b/drivers/cxl/core/ras.c
>> @@ -227,20 +227,19 @@ void __iomem *to_ras_base(struct cxl_port *port, struct cxl_dport *dport)
>>
>> void cxl_do_recovery(struct pci_dev *pdev, struct cxl_port *port, struct cxl_dport *dport)
>> {
>> - struct device *dev = dport ? dport->dport_dev : port->uport_dev;
>> void __iomem *ras_base = to_ras_base(port, dport);
>>
>> if (!ras_base)
>> panic("CXL: UCE with unmapped RAS registers");
>>
>> - if (cxl_handle_ras(dev, ras_base))
>> + if (cxl_handle_ras(port, dport, ras_base))
>
> I haven't looked ahead, and maybe I'm missing something, but why not have cxl_handle_ras() just
> take the port & dport and call to_ras_base() internally? It's not a big deal, but it would
> make the call in cxl_error_detected() a lot prettier.
>
This would break for the CPER case which uses GHES/CPER buffer instead of a MMIO RAS block.
> Either way:
> Reviewed-by: Ben Cheatham <benjamin.cheatham@amd.com>
Thanks Ben.
-Terry
>> panic("CXL cachemem error");
>>
>> dev_dbg(&pdev->dev,
>> "CXL UCE signaled but no CXL RAS status bits set\n");
>> }
>>
>> -void cxl_handle_cor_ras(struct device *dev, void __iomem *ras_base)
>> +void cxl_handle_cor_ras(struct cxl_port *port, struct cxl_dport *dport, void __iomem *ras_base)
>> {
>> void __iomem *addr;
>> u32 status;
>> @@ -252,10 +251,12 @@ void cxl_handle_cor_ras(struct device *dev, void __iomem *ras_base)
>> status = readl(addr);
>> if (status & CXL_RAS_CORRECTABLE_STATUS_MASK) {
>> writel(status & CXL_RAS_CORRECTABLE_STATUS_MASK, addr);
>> - if (is_cxl_memdev(dev))
>> - trace_cxl_aer_correctable_error(to_cxl_memdev(dev), status);
>> + if (is_cxl_endpoint(port))
>> + trace_cxl_aer_correctable_error(to_cxl_memdev(port->uport_dev), status);
>> + else if (dport)
>> + trace_cxl_port_aer_correctable_error(dport->dport_dev, status);
>> else
>> - trace_cxl_port_aer_correctable_error(dev, status);
>> + trace_cxl_port_aer_correctable_error(port->uport_dev, status);
>> }
>> }
>>
>> @@ -280,7 +281,7 @@ static void header_log_copy(void __iomem *ras_base, u32 *log)
>> * Log the state of the RAS status registers and prepare them to log the
>> * next error status. Return 1 if reset needed.
>> */
>> -bool cxl_handle_ras(struct device *dev, void __iomem *ras_base)
>> +bool cxl_handle_ras(struct cxl_port *port, struct cxl_dport *dport, void __iomem *ras_base)
>> {
>> u32 hl[CXL_HEADERLOG_TRACE_SIZE_U32] = {};
>> void __iomem *addr;
>> @@ -307,10 +308,12 @@ bool cxl_handle_ras(struct device *dev, void __iomem *ras_base)
>> }
>>
>> header_log_copy(ras_base, hl);
>> - if (is_cxl_memdev(dev))
>> - trace_cxl_aer_uncorrectable_error(to_cxl_memdev(dev), status, fe, hl);
>> + if (is_cxl_endpoint(port))
>> + trace_cxl_aer_uncorrectable_error(to_cxl_memdev(port->uport_dev), status, fe, hl);
>> + else if (dport)
>> + trace_cxl_port_aer_uncorrectable_error(dport->dport_dev, status, fe, hl);
>> else
>> - trace_cxl_port_aer_uncorrectable_error(dev, status, fe, hl);
>> + trace_cxl_port_aer_uncorrectable_error(port->uport_dev, status, fe, hl);
>>
>> writel(status & CXL_RAS_UNCORRECTABLE_STATUS_MASK, addr);
>>
>> @@ -344,7 +347,7 @@ pci_ers_result_t cxl_error_detected(struct pci_dev *pdev,
>> * cache coherency is already lost; continuing risks silent
>> * data corruption.
>> */
>> - ue = cxl_handle_ras(port->uport_dev, to_ras_base(port, NULL));
>> + ue = cxl_handle_ras(port, NULL, to_ras_base(port, NULL));
>> }
>>
>> /*
>> @@ -375,10 +378,8 @@ EXPORT_SYMBOL_NS_GPL(cxl_error_detected, "CXL");
>> static void cxl_handle_proto_error(struct pci_dev *pdev, struct cxl_port *port,
>> struct cxl_dport *dport, int severity)
>> {
>> - struct device *dev = dport ? dport->dport_dev : port->uport_dev;
>> -
>> if (severity == AER_CORRECTABLE)
>> - cxl_handle_cor_ras(dev, to_ras_base(port, dport));
>> + cxl_handle_cor_ras(port, dport, to_ras_base(port, dport));
>> else
>> cxl_do_recovery(pdev, port, dport);
>> }
>> diff --git a/drivers/cxl/core/ras_rch.c b/drivers/cxl/core/ras_rch.c
>> index 285eed5b828b9..bc6cf50fcc926 100644
>> --- a/drivers/cxl/core/ras_rch.c
>> +++ b/drivers/cxl/core/ras_rch.c
>> @@ -123,7 +123,7 @@ void cxl_handle_rdport_errors(struct pci_dev *pdev)
>> */
>> if (aer_regs.cor_status & ~aer_regs.cor_mask) {
>> pci_print_aer(pdev, AER_CORRECTABLE, &aer_regs);
>> - cxl_handle_cor_ras(dport->dport_dev, to_ras_base(port, dport));
>> + cxl_handle_cor_ras(port, dport, to_ras_base(port, dport));
>> }
>>
>> if (uncor_status) {
>
next prev parent reply other threads:[~2026-09-09 14:42 UTC|newest]
Thread overview: 45+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-02 13:39 [PATCH v20 0/9] Enable CXL PCIe Port Protocol Error handling and logging Terry Bowman
2026-09-02 13:39 ` [PATCH v20 1/9] PCI/AER: Introduce AER-CXL protocol error kfifo Terry Bowman
2026-09-02 13:48 ` sashiko-bot
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-08 0:51 ` Jonathan Cameron
2026-09-09 15:38 ` Bowman, Terry
2026-09-09 22:02 ` Jonathan Cameron
2026-09-10 14:57 ` Bowman, Terry
2026-09-02 13:39 ` [PATCH v20 2/9] PCI: Establish common CXL Port protocol error flow Terry Bowman
2026-09-02 13:57 ` sashiko-bot
2026-09-02 16:11 ` Bowman, Terry
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-10 16:55 ` Bowman, Terry
2026-09-08 0:57 ` Jonathan Cameron
2026-09-02 13:39 ` [PATCH v20 3/9] cxl/ras: Handle RCH correctable and uncorrectable errors in one pass Terry Bowman
2026-09-02 13:50 ` sashiko-bot
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-08 1:06 ` Jonathan Cameron
2026-09-02 13:39 ` [PATCH v20 4/9] cxl/pci: Thread port and dport through RAS handling helpers Terry Bowman
2026-09-02 13:51 ` sashiko-bot
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-09 14:42 ` Bowman, Terry [this message]
2026-09-09 15:21 ` Bowman, Terry
2026-09-08 17:47 ` Jonathan Cameron
2026-09-02 13:39 ` [PATCH v20 5/9] cxl: Update CXL Endpoint AER handler Terry Bowman
2026-09-02 14:05 ` sashiko-bot
2026-09-02 18:19 ` Bowman, Terry
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-02 13:39 ` [PATCH v20 6/9] PCI: Cache PCI DSN into pci_dev->dsn during probe Terry Bowman
2026-09-02 13:47 ` sashiko-bot
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-09 15:16 ` Lukas Wunner
2026-09-02 13:39 ` [PATCH v20 7/9] cxl: Add port and dport identifiers to CXL AER trace events Terry Bowman
2026-09-02 13:50 ` sashiko-bot
2026-09-08 18:13 ` Jonathan Cameron
2026-09-02 13:39 ` [PATCH v20 8/9] PCI/CXL: Mask/Unmask CXL protocol errors Terry Bowman
2026-09-02 14:03 ` sashiko-bot
2026-09-02 15:49 ` Bowman, Terry
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-02 13:39 ` [PATCH v20 9/9] Documentation: cxl: Document CXL protocol error handling Terry Bowman
2026-09-02 13:50 ` sashiko-bot
2026-09-08 18:39 ` Jonathan Cameron
2026-09-10 15:19 ` Bowman, Terry
2026-09-09 16:03 ` [PATCH v20 0/9] Enable CXL PCIe Port Protocol Error handling and logging Lukas Wunner
2026-09-09 20:31 ` Bowman, Terry
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=51154495-cddd-4ab2-ae37-2f085fb058df@amd.com \
--to=terry.bowman@amd.com \
--cc=alison.schofield@intel.com \
--cc=benjamin.cheatham@amd.com \
--cc=bhelgaas@google.com \
--cc=bp@alien8.de \
--cc=corbet@lwn.net \
--cc=dave.jiang@intel.com \
--cc=dave@stgolabs.net \
--cc=djbw@kernel.org \
--cc=guohanjun@huawei.com \
--cc=icheng@nvidia.com \
--cc=iweiny@kernel.org \
--cc=jic23@kernel.org \
--cc=lenb@kernel.org \
--cc=linux-acpi@vger.kernel.org \
--cc=linux-cxl@vger.kernel.org \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-pci@vger.kernel.org \
--cc=lukas@wunner.de \
--cc=mchehab@kernel.org \
--cc=ming.li@zohomail.com \
--cc=rafael@kernel.org \
--cc=rrichter@amd.com \
--cc=skhan@linuxfoundation.org \
--cc=tony.luck@intel.com \
--cc=vishal.l.verma@intel.com \
--cc=xueshuai@linux.alibaba.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.