From: Lukas Wunner <lukas@wunner.de>
To: Terry Bowman <terry.bowman@amd.com>
Cc: Jonathan Cameron <jic23@kernel.org>,
Dave Jiang <dave.jiang@intel.com>,
Alison Schofield <alison.schofield@intel.com>,
Vishal Verma <vishal.l.verma@intel.com>,
Davidlohr Bueso <dave@stgolabs.net>,
Bjorn Helgaas <bhelgaas@google.com>,
"Rafael J . Wysocki" <rafael@kernel.org>,
Jonathan Corbet <corbet@lwn.net>,
linux-cxl@vger.kernel.org, Tony Luck <tony.luck@intel.com>,
Borislav Petkov <bp@alien8.de>, Hanjun Guo <guohanjun@huawei.com>,
Mauro Carvalho Chehab <mchehab@kernel.org>,
Shuai Xue <xueshuai@linux.alibaba.com>,
Len Brown <lenb@kernel.org>, Ira Weiny <iweiny@kernel.org>,
Li Ming <ming.li@zohomail.com>,
Shuah Khan <skhan@linuxfoundation.org>,
Ben Cheatham <Benjamin.Cheatham@amd.com>,
Richard Cheng <icheng@nvidia.com>,
Robert Richter <rrichter@amd.com>,
linux-pci@vger.kernel.org, linux-acpi@vger.kernel.org,
linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH v19 02/14] cxl/ras: Fix cxl_rch_get_aer_severity() wrong severity register
Date: Sun, 9 Aug 2026 17:57:23 +0200 [thread overview]
Message-ID: <anijY3eTVWSqHLxh@wunner.de> (raw)
In-Reply-To: <20260803221810.3685703-3-terry.bowman@amd.com>
On Mon, Aug 03, 2026 at 05:17:58PM -0500, Terry Bowman wrote:
> +++ b/drivers/cxl/core/ras_rch.c
> @@ -94,11 +94,11 @@ static bool cxl_rch_get_aer_info(void __iomem *aer_base,
> static bool cxl_rch_get_aer_severity(struct aer_capability_regs *aer_regs,
> int *severity)
> {
> - if (aer_regs->uncor_status & ~aer_regs->uncor_mask) {
> - if (aer_regs->uncor_status & PCI_ERR_ROOT_FATAL_RCV)
> - *severity = AER_FATAL;
> - else
> - *severity = AER_NONFATAL;
> + u32 uncor_status = aer_regs->uncor_status & ~aer_regs->uncor_mask;
> +
> + if (uncor_status) {
> + *severity = (uncor_status & aer_regs->uncor_severity) ?
> + AER_FATAL : AER_NONFATAL;
> return true;
> }
>
Independently of this patch, I'm wondering why the severity is inferred
from the AER registers. I would assume that the severity always equals
the message received by the RCEC (ERR_COR, ERR_NONFATAL or ERR_FATAL).
So the severity could be passed to cxl_handle_rdport_errors() from its
callers: cxl_cor_err_detected() would pass AER_CORRECTABLE and
cxl_error_detected() would pass ERR_NONFATAL or ERR_FATAL (depending
on the "state" variable).
cxl_handle_rdport_errors() would no longer need to call
cxl_rch_get_aer_severity(), so the latter could be removed.
Am I missing something? Is a scenario ever conceivable where the RCEC
receives a message with different severity than what is inferred from
the registers by cxl_rch_get_aer_severity()?
And a related question: linux-next commit 21963e6e4e04 ("PCI/AER:
Support Advisory Non-Fatal Errors") enables support for Non-Fatal
Errors which are signaled with an ERR_COR message.
These so-called Advisory Non-Fatal Errors set one bit in the Correctable
Error Status Register (Advisory Non-Fatal Error Status, bit 13) and
additionally one or more bits in the Uncorrectable Error Status Register
(see the commit for details).
cxl_rch_get_aer_severity() is not able to cope with such errors and
will incorrectly infer that the severity is AER_NONFATAL.
Now I *think* this is not a problem because Advisory Non-Fatal Errors
are masked by default and the commit only unmasks them on regular PCI
devices, not in the RCRB of a CXL device. Only once bit 13 in the
Correctable Error Mask Register is cleared in the RCRB will
cxl_rch_get_aer_severity() fail. Right?
Thanks,
Lukas
next prev parent reply other threads:[~2026-08-09 15:57 UTC|newest]
Thread overview: 48+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-03 22:17 [PATCH v19 00/14] Enable CXL PCIe Port Protocol Error handling and logging Terry Bowman
2026-08-03 22:17 ` [PATCH v19 01/14] cxl/ras: Fix cxl_rch_get_aer_info() out-of-bounds AER register read Terry Bowman
2026-08-03 22:42 ` sashiko-bot
2026-08-04 16:20 ` Bowman, Terry
2026-08-04 2:10 ` Alison Schofield
2026-08-03 22:17 ` [PATCH v19 02/14] cxl/ras: Fix cxl_rch_get_aer_severity() wrong severity register Terry Bowman
2026-08-03 22:35 ` sashiko-bot
2026-08-04 2:11 ` Alison Schofield
2026-08-09 15:57 ` Lukas Wunner [this message]
2026-08-03 22:17 ` [PATCH v19 03/14] acpi/apei/ghes: Use raw_spinlock_t for CXL CPER work locks Terry Bowman
2026-08-03 22:39 ` sashiko-bot
2026-08-05 18:41 ` Luck, Tony
2026-08-03 22:18 ` [PATCH v19 04/14] cxl: Tighten CPER kfifo registration API and symbol visibility Terry Bowman
2026-08-03 22:30 ` sashiko-bot
2026-08-04 2:13 ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 05/14] cxl: Rename find_cxl_port() to find_cxl_port_by_dport() Terry Bowman
2026-08-03 22:29 ` sashiko-bot
2026-08-04 2:14 ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 06/14] PCI/AER: Introduce AER-CXL protocol error kfifo Terry Bowman
2026-08-03 22:28 ` sashiko-bot
2026-08-04 8:15 ` Richard Cheng
2026-08-04 14:10 ` Bowman, Terry
2026-08-03 22:18 ` [PATCH v19 07/14] PCI: Establish common CXL Port protocol error flow Terry Bowman
2026-08-03 22:56 ` sashiko-bot
2026-08-03 22:18 ` [PATCH v19 08/14] cxl/ras: Handle RCH correctable and uncorrectable errors in one pass Terry Bowman
2026-08-03 22:29 ` sashiko-bot
2026-08-04 2:16 ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 09/14] cxl/pci: Thread port and dport through RAS handling helpers Terry Bowman
2026-08-03 22:33 ` sashiko-bot
2026-08-04 2:16 ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 10/14] cxl: Update CXL Endpoint AER handler Terry Bowman
2026-08-03 22:40 ` sashiko-bot
2026-08-04 2:17 ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 11/14] PCI: Cache PCI DSN into pci_dev->dsn during probe Terry Bowman
2026-08-03 22:29 ` sashiko-bot
2026-08-04 2:26 ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 12/14] cxl: Add port and dport identifiers to CXL AER trace events Terry Bowman
2026-08-03 22:42 ` sashiko-bot
2026-08-04 2:27 ` Alison Schofield
2026-08-04 7:56 ` Richard Cheng
2026-08-04 13:46 ` Bowman, Terry
2026-08-03 22:18 ` [PATCH v19 13/14] PCI/CXL: Mask/Unmask CXL protocol errors Terry Bowman
2026-08-03 22:55 ` sashiko-bot
2026-08-04 2:29 ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 14/14] Documentation: cxl: Document CXL protocol error handling Terry Bowman
2026-08-03 22:31 ` sashiko-bot
2026-08-04 2:30 ` Alison Schofield
2026-08-05 21:19 ` [PATCH v19 00/14] Enable CXL PCIe Port Protocol Error handling and logging Dave Jiang
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=anijY3eTVWSqHLxh@wunner.de \
--to=lukas@wunner.de \
--cc=Benjamin.Cheatham@amd.com \
--cc=alison.schofield@intel.com \
--cc=bhelgaas@google.com \
--cc=bp@alien8.de \
--cc=corbet@lwn.net \
--cc=dave.jiang@intel.com \
--cc=dave@stgolabs.net \
--cc=guohanjun@huawei.com \
--cc=icheng@nvidia.com \
--cc=iweiny@kernel.org \
--cc=jic23@kernel.org \
--cc=lenb@kernel.org \
--cc=linux-acpi@vger.kernel.org \
--cc=linux-cxl@vger.kernel.org \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-pci@vger.kernel.org \
--cc=mchehab@kernel.org \
--cc=ming.li@zohomail.com \
--cc=rafael@kernel.org \
--cc=rrichter@amd.com \
--cc=skhan@linuxfoundation.org \
--cc=terry.bowman@amd.com \
--cc=tony.luck@intel.com \
--cc=vishal.l.verma@intel.com \
--cc=xueshuai@linux.alibaba.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.