From: "Cheatham, Benjamin" <benjamin.cheatham@amd.com>
To: Terry Bowman <terry.bowman@amd.com>,
Jonathan Cameron <jic23@kernel.org>,
Dave Jiang <dave.jiang@intel.com>,
Alison Schofield <alison.schofield@intel.com>,
Vishal Verma <vishal.l.verma@intel.com>,
Davidlohr Bueso <dave@stgolabs.net>,
Bjorn Helgaas <bhelgaas@google.com>,
"Dan Williams" <djbw@kernel.org>,
"Rafael J . Wysocki" <rafael@kernel.org>,
Jonathan Corbet <corbet@lwn.net>, <linux-cxl@vger.kernel.org>
Cc: Tony Luck <tony.luck@intel.com>, Borislav Petkov <bp@alien8.de>,
"Hanjun Guo" <guohanjun@huawei.com>,
Mauro Carvalho Chehab <mchehab@kernel.org>,
"Shuai Xue" <xueshuai@linux.alibaba.com>,
Len Brown <lenb@kernel.org>, Ira Weiny <iweiny@kernel.org>,
Li Ming <ming.li@zohomail.com>,
Shuah Khan <skhan@linuxfoundation.org>,
Richard Cheng <icheng@nvidia.com>,
"Robert Richter" <rrichter@amd.com>,
Lukas Wunner <lukas@wunner.de>, <linux-pci@vger.kernel.org>,
<linux-acpi@vger.kernel.org>, <linux-doc@vger.kernel.org>,
<linux-kernel@vger.kernel.org>
Subject: Re: [PATCH v20 2/9] PCI: Establish common CXL Port protocol error flow
Date: Wed, 2 Sep 2026 15:57:12 -0500 [thread overview]
Message-ID: <a8ea024b-a608-4eea-bbdc-cba37b2bf2af@amd.com> (raw)
In-Reply-To: <20260902133933.2992457-3-terry.bowman@amd.com>
On 9/2/2026 8:39 AM, Terry Bowman wrote:
> Establish a single CXL protocol error path shared by CXL Virtual
> Hierarchy (VH) and Restricted CXL Host (RCH) topologies. AER dispatch in
> handle_error_source() routes CXL protocol errors, gated by
> is_cxl_error(), through the AER-CXL kfifo to a cxl_core consumer for
> logging and recovery. Producer and consumer go live together so no CXL
> error is silently dropped across a bisect.
>
> is_cxl_error() expands from Endpoint-only to also cover Root Port,
> Upstream Port, and Downstream Port. RCDs report on behalf of an upstream
> RCH Downstream Port and instead reach the kfifo via
> cxl_rch_handle_error().
>
> For uncorrectable errors, cxl_proto_err_wait_for_empty() drains the CXL
> plane (RAS read, panic policy, state clear) before pci_aer_handle_error()
> drives PCIe recovery, so recovery does not tear down RAS iomaps while the
> consumer is still reading them. Correctable errors run asynchronously.
>
> Panic policy: cxl_do_recovery() panics on a confirmed UCE, and also when
> the RAS registers cannot be mapped -- an unconfirmable UCE is treated
> conservatively as fatal since CXL.mem coherency may be lost. A
> mapped-but-clear status is logged as spurious with no panic.
>
> to_ras_base() centralizes RAS base lookup (dport->regs.ras for
> Root/Downstream Ports, port->regs.ras otherwise) and provides an
> injection point for RAS status simulation during testing. The
> cxl_cor_error_detected() AER callback is removed; correctable Endpoint
> errors now route through the kfifo like every other CXL protocol error.
>
> Update cxl_handle_rdport_errors() with locking to prevent dport from
> being freed and RAS from being unmapped.
>
> At this step cxl_handle_rdport_errors() still dispatches a single
> severity per pass (matching the pre-series baseline). The following
> patch, "cxl/ras: Handle RCH correctable and uncorrectable errors in one
> pass", processes a simultaneously signalled CE and UCE together.
>
> Co-developed-by: Dan Williams <djbw@kernel.org>
> Signed-off-by: Dan Williams <djbw@kernel.org>
> Signed-off-by: Terry Bowman <terry.bowman@amd.com>
>
> ---
>
One small nit, but otherwise LGTM:
Reviewed-by: Ben Cheatham <benjamin.cheatham@amd.com>
...
> pci_ers_result_t cxl_error_detected(struct pci_dev *pdev,
> pci_channel_state_t state)
> {
> - struct cxl_dev_state *cxlds = pci_get_drvdata(pdev);
> - struct cxl_memdev *cxlmd = cxlds->cxlmd;
> - struct device *dev = &cxlmd->dev;
> - bool ue;
> + struct cxl_port *port __free(put_cxl_port) = find_cxl_port_by_uport(&pdev->dev);
> + bool ue = false;
> +
> + if (!port)
> + return PCI_ERS_RESULT_DISCONNECT;
> +
> + if (is_cxl_restricted(pdev))
> + cxl_handle_rdport_errors(pdev);
>
> - scoped_guard(device, dev) {
> - if (!dev->driver) {
> + scoped_guard(device, &port->dev) {
> + if (!port->dev.driver) {
> dev_warn(&pdev->dev,
> - "%s: memdev disabled, abort error handling\n",
> - dev_name(dev));
> + "%s: port disabled, abort error handling\n",
> + dev_name(&port->dev));
> return PCI_ERS_RESULT_DISCONNECT;
> }
>
> - if (cxlds->rcd)
> - cxl_handle_rdport_errors(cxlds);
> /*
> - * A frozen channel indicates an impending reset which is fatal to
> - * CXL.mem operation, and will likely crash the system. On the off
> - * chance the situation is recoverable dump the status of the RAS
> - * capability registers and bounce the active state of the memdev.
> + * The CXL RAS read is unconditional regardless of channel
> + * state. Any uncorrectable error bit set in the CXL RAS
> + * status register triggers a panic below because CXL.mem
> + * cache coherency is already lost; continuing risks silent
> + * data corruption.
> */
> - ue = cxl_handle_ras(&cxlds->cxlmd->dev, cxlmd->endpoint->regs.ras);
> + ue = cxl_handle_ras(port->uport_dev, to_ras_base(port, NULL));
> }
>
> + /*
> + * CXL.mem UCE means cache coherency is lost. Continuing risks
> + * silent data corruption.
> + */
Don't need this comment and the last sentence in the comment above.
> + if (ue)
> + panic("CXL cachemem error");
> +
> switch (state) {
> case pci_channel_io_normal:
> - if (ue) {
> - device_release_driver(dev);
> - return PCI_ERS_RESULT_NEED_RESET;
> - }
> return PCI_ERS_RESULT_CAN_RECOVER;
> case pci_channel_io_frozen:
> dev_warn(&pdev->dev,
> "%s: frozen state error detected, disable CXL.mem\n",
> - dev_name(dev));
> - device_release_driver(dev);
> + dev_name(port->uport_dev));
> + device_release_driver(port->uport_dev);
> return PCI_ERS_RESULT_NEED_RESET;
> case pci_channel_io_perm_failure:
> dev_warn(&pdev->dev,
> @@ -335,3 +371,82 @@ pci_ers_result_t cxl_error_detected(struct pci_dev *pdev,
> return PCI_ERS_RESULT_NEED_RESET;
> }
> EXPORT_SYMBOL_NS_GPL(cxl_error_detected, "CXL");
next prev parent reply other threads:[~2026-09-02 20:57 UTC|newest]
Thread overview: 17+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-02 13:39 [PATCH v20 0/9] Enable CXL PCIe Port Protocol Error handling and logging Terry Bowman
2026-09-02 13:39 ` [PATCH v20 1/9] PCI/AER: Introduce AER-CXL protocol error kfifo Terry Bowman
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-02 13:39 ` [PATCH v20 2/9] PCI: Establish common CXL Port protocol error flow Terry Bowman
2026-09-02 20:57 ` Cheatham, Benjamin [this message]
2026-09-02 13:39 ` [PATCH v20 3/9] cxl/ras: Handle RCH correctable and uncorrectable errors in one pass Terry Bowman
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-02 13:39 ` [PATCH v20 4/9] cxl/pci: Thread port and dport through RAS handling helpers Terry Bowman
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-02 13:39 ` [PATCH v20 5/9] cxl: Update CXL Endpoint AER handler Terry Bowman
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-02 13:39 ` [PATCH v20 6/9] PCI: Cache PCI DSN into pci_dev->dsn during probe Terry Bowman
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-02 13:39 ` [PATCH v20 7/9] cxl: Add port and dport identifiers to CXL AER trace events Terry Bowman
2026-09-02 13:39 ` [PATCH v20 8/9] PCI/CXL: Mask/Unmask CXL protocol errors Terry Bowman
2026-09-02 20:57 ` Cheatham, Benjamin
2026-09-02 13:39 ` [PATCH v20 9/9] Documentation: cxl: Document CXL protocol error handling Terry Bowman
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=a8ea024b-a608-4eea-bbdc-cba37b2bf2af@amd.com \
--to=benjamin.cheatham@amd.com \
--cc=alison.schofield@intel.com \
--cc=bhelgaas@google.com \
--cc=bp@alien8.de \
--cc=corbet@lwn.net \
--cc=dave.jiang@intel.com \
--cc=dave@stgolabs.net \
--cc=djbw@kernel.org \
--cc=guohanjun@huawei.com \
--cc=icheng@nvidia.com \
--cc=iweiny@kernel.org \
--cc=jic23@kernel.org \
--cc=lenb@kernel.org \
--cc=linux-acpi@vger.kernel.org \
--cc=linux-cxl@vger.kernel.org \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-pci@vger.kernel.org \
--cc=lukas@wunner.de \
--cc=mchehab@kernel.org \
--cc=ming.li@zohomail.com \
--cc=rafael@kernel.org \
--cc=rrichter@amd.com \
--cc=skhan@linuxfoundation.org \
--cc=terry.bowman@amd.com \
--cc=tony.luck@intel.com \
--cc=vishal.l.verma@intel.com \
--cc=xueshuai@linux.alibaba.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox