From: Jonathan Cameron <jic23@kernel.org>
To: Terry Bowman <terry.bowman@amd.com>
Cc: Bjorn Helgaas <bhelgaas@google.com>,
Dan Williams <djbw@kernel.org>,
"Dave Jiang" <dave.jiang@intel.com>,
Ira Weiny <iweiny@kernel.org>, Len Brown <lenb@kernel.org>,
"Rafael J . Wysocki" <rafael@kernel.org>,
Robert Richter <rrichter@amd.com>, <linux-acpi@vger.kernel.org>,
<linux-cxl@vger.kernel.org>, <linux-doc@vger.kernel.org>,
<linux-kernel@vger.kernel.org>, <linux-pci@vger.kernel.org>,
<linuxppc-dev@lists.ozlabs.org>,
"Alejandro Lucero" <alucerop@amd.com>,
Alison Schofield <alison.schofield@intel.com>,
Ankit Agrawal <ankita@nvidia.com>,
Ard Biesheuvel <ardb@kernel.org>,
"Ben Cheatham" <Benjamin.Cheatham@amd.com>,
Borislav Petkov <bp@alien8.de>,
"Breno Leitao" <leitao@debian.org>,
Davidlohr Bueso <dave@stgolabs.net>,
"Fabio M . De Francesco" <fabio.m.de.francesco@linux.intel.com>,
Gregory Price <gourry@gourry.net>,
Hanjun Guo <guohanjun@huawei.com>,
Jonathan Corbet <corbet@lwn.net>, Kees Cook <kees@kernel.org>,
Kuppuswamy Sathyanarayanan
<sathyanarayanan.kuppuswamy@linux.intel.com>,
Li Ming <ming.li@zohomail.com>,
Mahesh J Salgaonkar <mahesh@linux.ibm.com>,
Mauro Carvalho Chehab <mchehab@kernel.org>,
Oliver O'Halloran <oohall@gmail.com>,
Shiju Jose <shiju.jose@huawei.com>,
Shuah Khan <skhan@linuxfoundation.org>,
Shuai Xue <xueshuai@linux.alibaba.com>,
Smita Koralahalli <Smita.KoralahalliChannabasappa@amd.com>,
Tony Luck <tony.luck@intel.com>,
Vishal Verma <vishal.l.verma@intel.com>
Subject: Re: [PATCH v18 13/13] Documentation: cxl: Document CXL protocol error handling
Date: Tue, 21 Jul 2026 01:19:02 +0100 [thread overview]
Message-ID: <20260721011902.0a93da2d@jic23-huawei> (raw)
In-Reply-To: <20260717222706.3540281-14-terry.bowman@amd.com>
On Fri, 17 Jul 2026 17:27:06 -0500
Terry Bowman <terry.bowman@amd.com> wrote:
> Add Documentation/driver-api/cxl/linux/protocol-error-handling.rst
> describing the end-to-end CXL protocol error path: AER ingress, the
> AER-CXL kfifo handoff, the cxl_core consumer worker, RCD/RCH special
> cases, severity policy, trace events, and a source code map.
>
> This documents the architecture introduced by the preceding patches in
> this series.
>
> Assisted-by: Claude:claude-opus-4.7
> Signed-off-by: Terry Bowman <terry.bowman@amd.com>
A couple of trivial things inline to tidy up. Actual text seems good
to me.
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
>
> ---
> Changes in v17->v18:
> - Simplify document for readability (Jonathan)
> - Drop historical context that goes stale (Jonathan)
> - Shorten ASCII flow diagram (Jonathan)
> - Drop manual backtick markup, use automarkup (Jonathan)
> - Clarify USP/DSP as single switch component (Dave)
> - Fix line wrapping to 80 chars (Jonathan)
> ---
> Documentation/driver-api/cxl/index.rst | 1 +
> .../cxl/linux/protocol-error-handling.rst | 222 ++++++++++++++++++
> 2 files changed, 223 insertions(+)
> create mode 100644 Documentation/driver-api/cxl/linux/protocol-error-handling.rst
>
> diff --git a/Documentation/driver-api/cxl/index.rst b/Documentation/driver-api/cxl/index.rst
> index 3dfae1d310ca5..6861b2e5726a3 100644
> --- a/Documentation/driver-api/cxl/index.rst
> +++ b/Documentation/driver-api/cxl/index.rst
> @@ -42,6 +42,7 @@ that have impacts on each other. The docs here break up configurations steps.
> linux/dax-driver
> linux/memory-hotplug
> linux/access-coordinates
> + linux/protocol-error-handling
>
> .. toctree::
> :maxdepth: 2
> diff --git a/Documentation/driver-api/cxl/linux/protocol-error-handling.rst b/Documentation/driver-api/cxl/linux/protocol-error-handling.rst
> new file mode 100644
> index 0000000000000..67f0492e56702
> --- /dev/null
> +++ b/Documentation/driver-api/cxl/linux/protocol-error-handling.rst
> @@ -0,0 +1,222 @@
> +.. SPDX-License-Identifier: GPL-2.0
> +
> +==============================
> +CXL Protocol Error Handling
> +==============================
> +
> +CXL devices report protocol-layer failures (CXL.cachemem RAS) as PCIe
Why short wrap? Docs are 80 chars I think.
> +AER Internal Errors: PCI_ERR_COR_INTERNAL for correctable events and
> +PCI_ERR_UNC_INTN for uncorrectable events. The actual fault
> +information lives in CXL RAS capability registers, not in the PCIe AER
> +status registers.
> +Error flow
> +==========
> +
> +.. code-block:: text
> +
> + CXL device raises AER Internal Error
> + (PCI_ERR_COR_INTERNAL or PCI_ERR_UNC_INTN)
> + |
> + v
> + +--------------------------------------+
> + | AER core (aer.c) |
> + | aer_irq() -> aer_isr() |
> + | -> find_source_device() |
> + | -> handle_error_source(dev, info) |
> + +--------------------------------------+
> + |
> + v
> + +--------------------------------------+
> + | handle_error_source() dispatch |
> + | |
> + | 1. cxl_rch_handle_error() |
> + | [always; filters internally] |
> + | |
> + | 2. if is_cxl_error(): |
Smells like a missing space.
> + | cxl_forward_error() |
> + | [enqueue to kfifo] |
> + | |
> + | 3. if cxl_pending && non-CE: |
> + | cxl_proto_err_flush() |
> + | [sync drain before recovery] |
> + | |
> + | 4. pci_aer_handle_error() [always] |
> + +--------------------------------------+
> + |
> + (kfifo -> workqueue)
> + |
> + v
> + +--------------------------------------+
> + | __cxl_proto_err_work_fn() consumer |
> + | |
> + | if is_cxl_restricted(pdev): |
> + | cxl_handle_rdport_errors() |
> + | [RCH dport RAS first] |
> + | |
> + | port = find_cxl_port_by_dev( |
> + | &pdev->dev, NULL) |
> + | dport = cxl_find_dport_by_dev( |
> + | port, &pdev->dev) |
> + | [dport NULL for EP/USP; set RP/DSP] |
check the alignment here as well. A couple of extra spaces in the
lines above I think.
> + | |
> + | cxl_handle_proto_error() |
> + +--------------------------------------+
> + | |
> + v v
> + +-----------------+ +--------------------+
> + | CE | | UCE |
> + | cxl_handle_ | | cxl_do_recovery() |
> + | cor_ras() | | read RAS status |
> + | trace + clear | | trace + panic |
> + +-----------------+ +--------------------+
> +
> +cxl_do_recovery() reads the CXL RAS uncorrectable status register.
> +If UE bits are set, it emits the trace event and panics. If no bits
> +are set (e.g. RAS mapped but error already cleared), it logs a
> +diagnostic and defers to AER recovery.
> +
> +
> +Severity policy
> +===============
> +
> +**CE** - cxl_handle_cor_ras() reads the CXL RAS correctable status
> +register, clears set bits, and emits a cxl_aer_correctable_error
> +trace event. No recovery action.
> +
> +**UCE (non-fatal, and fatal on Root Port/Downstream Port)** - cxl_do_recovery() reads the CXL RAS
Wrap needs an update here.
> +uncorrectable status register. If UE bits are set, the kernel panics.
> +CXL.cachemem traffic cannot be safely recovered once an uncorrectable
> +error is signaled; continuing risks silent data corruption across
> +interleaved HDM regions. This panic policy applies to the native AER
> +path. On firmware-first (CPER/GHES) platforms the CPER handler emits
> +trace events only and does not call cxl_do_recovery().
prev parent reply other threads:[~2026-07-21 0:19 UTC|newest]
Thread overview: 63+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-17 22:26 [PATCH v18 00/13] Enable CXL PCIe Port Protocol Error handling and logging Terry Bowman
2026-07-17 22:26 ` [PATCH v18 01/13] cxl/ras: Fix cxl_rch_get_aer_severity() wrong severity register Terry Bowman
2026-07-17 22:43 ` sashiko-bot
2026-07-20 20:09 ` Dave Jiang
2026-07-20 20:36 ` Bowman, Terry
2026-07-20 21:26 ` Jonathan Cameron
2026-07-17 22:26 ` [PATCH v18 02/13] acpi/apei/ghes: Use raw_spinlock_t for CXL CPER work locks Terry Bowman
2026-07-17 22:49 ` sashiko-bot
2026-07-20 15:02 ` Bowman, Terry
2026-07-20 21:36 ` Jonathan Cameron
2026-07-20 20:12 ` Dave Jiang
2026-07-20 20:38 ` Bowman, Terry
2026-07-20 21:41 ` Jonathan Cameron
2026-07-17 22:26 ` [PATCH v18 03/13] cxl: Tighten CPER kfifo registration API and symbol visibility Terry Bowman
2026-07-17 22:37 ` sashiko-bot
2026-07-20 20:15 ` Dave Jiang
2026-07-20 21:59 ` Jonathan Cameron
2026-07-17 22:26 ` [PATCH v18 04/13] cxl: Rename find_cxl_port() to find_cxl_port_by_dport() Terry Bowman
2026-07-17 22:34 ` sashiko-bot
2026-07-20 20:25 ` Dave Jiang
2026-07-20 22:02 ` Jonathan Cameron
2026-07-17 22:26 ` [PATCH v18 05/13] PCI/AER: Introduce AER-CXL protocol error kfifo Terry Bowman
2026-07-17 22:35 ` sashiko-bot
2026-07-20 20:29 ` Dave Jiang
2026-07-20 22:41 ` Jonathan Cameron
2026-07-17 22:26 ` [PATCH v18 06/13] PCI: Establish common CXL Port protocol error flow Terry Bowman
2026-07-17 22:43 ` sashiko-bot
2026-07-20 20:44 ` Dave Jiang
2026-07-20 23:05 ` Jonathan Cameron
2026-07-17 22:27 ` [PATCH v18 07/13] PCI/CXL: Add RCH support to CXL handlers Terry Bowman
2026-07-17 22:43 ` sashiko-bot
2026-07-20 15:06 ` Bowman, Terry
2026-07-20 21:47 ` Dave Jiang
2026-07-20 23:12 ` Jonathan Cameron
2026-07-17 22:27 ` [PATCH v18 08/13] cxl/pci: Thread port and dport through RAS handling helpers Terry Bowman
2026-07-17 22:40 ` sashiko-bot
2026-07-20 22:15 ` Dave Jiang
2026-07-20 23:17 ` Jonathan Cameron
2026-07-17 22:27 ` [PATCH v18 09/13] cxl: Update CXL Endpoint AER handler Terry Bowman
2026-07-17 22:53 ` sashiko-bot
2026-07-20 15:09 ` Bowman, Terry
2026-07-20 22:25 ` Dave Jiang
2026-07-20 23:29 ` Jonathan Cameron
2026-07-17 22:27 ` [PATCH v18 10/13] cxl: Add port and dport identifiers to CXL AER trace events Terry Bowman
2026-07-17 22:53 ` sashiko-bot
2026-07-20 15:14 ` Bowman, Terry
2026-07-20 23:53 ` Jonathan Cameron
2026-07-20 22:44 ` Dave Jiang
2026-07-21 0:00 ` Jonathan Cameron
2026-07-21 20:59 ` Bowman, Terry
2026-07-17 22:27 ` [PATCH v18 11/13] PCI: Cache PCI DSN into pci_dev->dsn during probe Terry Bowman
2026-07-17 22:44 ` sashiko-bot
2026-07-18 7:02 ` Lukas Wunner
2026-07-20 15:48 ` Bowman, Terry
2026-07-21 8:37 ` Lukas Wunner
2026-07-17 22:27 ` [PATCH v18 12/13] PCI/CXL: Mask/Unmask CXL protocol errors Terry Bowman
2026-07-17 22:58 ` sashiko-bot
2026-07-20 22:52 ` Dave Jiang
2026-07-21 0:10 ` Jonathan Cameron
2026-07-17 22:27 ` [PATCH v18 13/13] Documentation: cxl: Document CXL protocol error handling Terry Bowman
2026-07-17 22:43 ` sashiko-bot
2026-07-20 23:40 ` Dave Jiang
2026-07-21 0:19 ` Jonathan Cameron [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260721011902.0a93da2d@jic23-huawei \
--to=jic23@kernel.org \
--cc=Benjamin.Cheatham@amd.com \
--cc=Smita.KoralahalliChannabasappa@amd.com \
--cc=alison.schofield@intel.com \
--cc=alucerop@amd.com \
--cc=ankita@nvidia.com \
--cc=ardb@kernel.org \
--cc=bhelgaas@google.com \
--cc=bp@alien8.de \
--cc=corbet@lwn.net \
--cc=dave.jiang@intel.com \
--cc=dave@stgolabs.net \
--cc=djbw@kernel.org \
--cc=fabio.m.de.francesco@linux.intel.com \
--cc=gourry@gourry.net \
--cc=guohanjun@huawei.com \
--cc=iweiny@kernel.org \
--cc=kees@kernel.org \
--cc=leitao@debian.org \
--cc=lenb@kernel.org \
--cc=linux-acpi@vger.kernel.org \
--cc=linux-cxl@vger.kernel.org \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-pci@vger.kernel.org \
--cc=linuxppc-dev@lists.ozlabs.org \
--cc=mahesh@linux.ibm.com \
--cc=mchehab@kernel.org \
--cc=ming.li@zohomail.com \
--cc=oohall@gmail.com \
--cc=rafael@kernel.org \
--cc=rrichter@amd.com \
--cc=sathyanarayanan.kuppuswamy@linux.intel.com \
--cc=shiju.jose@huawei.com \
--cc=skhan@linuxfoundation.org \
--cc=terry.bowman@amd.com \
--cc=tony.luck@intel.com \
--cc=vishal.l.verma@intel.com \
--cc=xueshuai@linux.alibaba.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.