Linux PCI subsystem development
 help / color / mirror / Atom feed
* [PATCH 0/7] Error reporting for AER-incapable devices
@ 2026-09-27 18:02 Lukas Wunner
  2026-09-27 18:02 ` [PATCH 1/7] PCI/DPC: Avoid access to non-existent AER capability Lukas Wunner
                   ` (6 more replies)
  0 siblings, 7 replies; 27+ messages in thread
From: Lukas Wunner @ 2026-09-27 18:02 UTC (permalink / raw)
  To: Bjorn Helgaas, Raag Jadav, Riana Tauro, Yury Murashka,
	Matthew W Carlis, linux-pci
  Cc: Mahesh J Salgaonkar, Oliver OHalloran, linuxppc-dev,
	Aravind Iddamsetty, Srinivasa Adatrao, Sathyanarayanan Kuppuswamy,
	Terry Bowman, Keith Busch

Bjorn suggests "the PCI core should do pci_enable_pcie_error_reporting()
independent of whether the device has an AER Capability":

https://lore.kernel.org/r/20260826212619.GA1566339@bhelgaas/

This is done in patch [7/7].  The preceding patches contain prep work.

For context, PCIe r7.1 sec 6.2.1 defines two error reporting paradigms:
AER and baseline capability.  The latter relies (solely) on the error
enable/status bits in the Device Control and Device Status registers,
which exist on every PCIe device.  The kernel has supported the AER
paradigm for 20 years.  With the present series, it gains support for
the baseline capability paradigm.

Patches [1/7] and [2/7] reinstate DPC on AER-incapable Downstream Ports.
The kernel used to support this until 2020.  It is hereby fixed because
it eases introduction of baseline capability error reporting in the
subsequent patches.

Lukas Wunner (6):
  PCI/DPC: Avoid access to non-existent AER capability
  PCI/DPC: Reinstate support for AER-incapable ports
  PCI/AER: Drop AER native check from handles_cxl_errors()
  PCI/AER: Move AER capability check out of pcie_aer_is_native()
  PCI/AER: Renumber severity constants
  PCI/AER: Enable baseline capability error reporting

Yury Murashka (1):
  PCI/ERR: Avoid stale error status bits on recovery failure

 .../ABI/testing/sysfs-bus-pci-devices-aer     | 26 +++++++---
 Documentation/PCI/pcieaer-howto.rst           |  9 ++--
 drivers/pci/pcie/aer.c                        | 51 +++++++++++--------
 drivers/pci/pcie/aer_cxl_rch.c                |  3 +-
 drivers/pci/pcie/dpc.c                        | 24 ++++-----
 drivers/pci/pcie/err.c                        |  5 ++
 include/linux/aer.h                           |  8 +--
 include/ras/ras_event.h                       |  2 +-
 8 files changed, 72 insertions(+), 56 deletions(-)

-- 
2.53.0


^ permalink raw reply	[flat|nested] 27+ messages in thread

* [PATCH 1/7] PCI/DPC: Avoid access to non-existent AER capability
  2026-09-27 18:02 [PATCH 0/7] Error reporting for AER-incapable devices Lukas Wunner
@ 2026-09-27 18:02 ` Lukas Wunner
  2026-09-27 18:27   ` sashiko-bot
  2026-09-29 18:49   ` Kuppuswamy Sathyanarayanan
  2026-09-27 18:02 ` [PATCH 2/7] PCI/DPC: Reinstate support for AER-incapable ports Lukas Wunner
                   ` (5 subsequent siblings)
  6 siblings, 2 replies; 27+ messages in thread
From: Lukas Wunner @ 2026-09-27 18:02 UTC (permalink / raw)
  To: Bjorn Helgaas, Raag Jadav, Riana Tauro, Yury Murashka,
	Matthew W Carlis, linux-pci
  Cc: Mahesh J Salgaonkar, Oliver OHalloran, linuxppc-dev,
	Aravind Iddamsetty, Srinivasa Adatrao, Sathyanarayanan Kuppuswamy,
	Terry Bowman, Keith Busch

Downstream Port Containment does not mandate presence of an Advanced Error
Reporting capability, so a Downstream Port may support DPC, but not AER
(PCIe r7.1 sec 6.2.11.2).

In February 2019, commit 9f08a5d896ce ("PCI/DPC: Fix print AER status in
DPC event handling") amended the DPC driver to access the AER capability
without checking for its presence.

In May 2020, commit 708b20003624 ("PCI/AER: Remove HEST/FIRMWARE_FIRST
parsing for AER ownership") fixed it by inserting a call to
pcie_aer_is_native() in dpc_probe(), which implicitly checks for presence
of an AER capability.

However already in October 2019, commit 35a0b2378c19 ("PCI/DPC: Add
"pcie_ports=dpc-native" to allow DPC without AER control") made it
possible to override the check:  The DPC driver may access a non-existent
AER capability if "pcie_ports=dpc-native" is passed on the command line.

Fix it by making the DPC driver cope with AER-unsupporting Downstream
Ports.

There are two places where the AER capability is accessed:

- dpc_get_aer_uncorrect_severity() uses it to discern whether a Fatal or
  Non-Fatal Error triggered DPC.  Access the Device Status Register
  instead, in accordance with PCIe r7.1 sec 6.2.5.

- dpc_is_surprise_removal() uses it to detect whether a Surprise Down
  Error triggered DPC.  Return false on AER-unsupporting devices.  The
  function works around an AMD-specific quirk and it seems reasonable to
  assume that all affected products are AER-supporting.  In any case the
  detection is not possible without AER capability.

Insert a temporary check for an AER capability after the call to
aer_get_device_error_info() because the function currently returns false
for AER-unsupporting devices.  The check will become obsolete and will be
removed with the imminent baseline capability error reporting.

Fixes: 35a0b2378c19 ("PCI/DPC: Add "pcie_ports=dpc-native" to allow DPC without AER control")
Signed-off-by: Lukas Wunner <lukas@wunner.de>
Cc: stable@vger.kernel.org # v5.5+
---
 drivers/pci/pcie/dpc.c | 23 ++++++++++-------------
 1 file changed, 10 insertions(+), 13 deletions(-)

diff --git a/drivers/pci/pcie/dpc.c b/drivers/pci/pcie/dpc.c
index 2b779bd1d861..793a799053f1 100644
--- a/drivers/pci/pcie/dpc.c
+++ b/drivers/pci/pcie/dpc.c
@@ -236,21 +236,15 @@ static void dpc_process_rp_pio_error(struct pci_dev *pdev)
 static int dpc_get_aer_uncorrect_severity(struct pci_dev *dev,
 					  struct aer_err_info *info)
 {
-	int pos = dev->aer_cap;
-	u32 status, mask, sev;
+	u16 devsta;
 
-	pci_read_config_dword(dev, pos + PCI_ERR_UNCOR_STATUS, &status);
-	pci_read_config_dword(dev, pos + PCI_ERR_UNCOR_MASK, &mask);
-	status &= ~mask;
-	if (!status)
-		return 0;
-
-	pci_read_config_dword(dev, pos + PCI_ERR_UNCOR_SEVER, &sev);
-	status &= sev;
-	if (status)
+	pcie_capability_read_word(dev, PCI_EXP_DEVSTA, &devsta);
+	if (devsta & PCI_EXP_DEVSTA_FED)
 		info->severity = AER_FATAL;
-	else
+	else if (devsta & PCI_EXP_DEVSTA_NFED)
 		info->severity = AER_NONFATAL;
+	else
+		return 0;
 
 	info->level = KERN_ERR;
 
@@ -275,7 +269,7 @@ void dpc_process_error(struct pci_dev *pdev)
 		pci_warn(pdev, "containment event, status:%#06x: unmasked uncorrectable error detected\n",
 			 status);
 		if (dpc_get_aer_uncorrect_severity(pdev, &info) &&
-		    aer_get_device_error_info(&info, 0)) {
+		    (aer_get_device_error_info(&info, 0) || !pdev->aer_cap)) {
 			aer_print_error(&info, 0);
 			pci_aer_clear_nonfatal_status(pdev);
 			pci_aer_clear_fatal_status(pdev);
@@ -353,6 +347,9 @@ static bool dpc_is_surprise_removal(struct pci_dev *pdev)
 	if (!pdev->is_hotplug_bridge)
 		return false;
 
+	if (!pdev->aer_cap)
+		return false;
+
 	if (pci_read_config_word(pdev, pdev->aer_cap + PCI_ERR_UNCOR_STATUS,
 				 &status))
 		return false;
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [PATCH 2/7] PCI/DPC: Reinstate support for AER-incapable ports
  2026-09-27 18:02 [PATCH 0/7] Error reporting for AER-incapable devices Lukas Wunner
  2026-09-27 18:02 ` [PATCH 1/7] PCI/DPC: Avoid access to non-existent AER capability Lukas Wunner
@ 2026-09-27 18:02 ` Lukas Wunner
  2026-09-27 18:26   ` sashiko-bot
  2026-09-29 19:04   ` Kuppuswamy Sathyanarayanan
  2026-09-27 18:02 ` [PATCH 3/7] PCI/ERR: Avoid stale error status bits on recovery failure Lukas Wunner
                   ` (4 subsequent siblings)
  6 siblings, 2 replies; 27+ messages in thread
From: Lukas Wunner @ 2026-09-27 18:02 UTC (permalink / raw)
  To: Bjorn Helgaas, Raag Jadav, Riana Tauro, Yury Murashka,
	Matthew W Carlis, linux-pci
  Cc: Mahesh J Salgaonkar, Oliver OHalloran, linuxppc-dev,
	Aravind Iddamsetty, Srinivasa Adatrao, Sathyanarayanan Kuppuswamy,
	Terry Bowman, Keith Busch

Downstream Port Containment does not mandate presence of an Advanced Error
Reporting capability, so a Downstream Port may support DPC, but not AER
(PCIe r7.1 sec 6.2.11.2).

Commit 708b20003624 ("PCI/AER: Remove HEST/FIRMWARE_FIRST parsing for AER
ownership") seemingly inadvertently constrained DPC to AER-supporting
Downstream Ports by inserting a call to pcie_aer_is_native() in
dpc_probe(), which implicitly checks for presence of an AER capability.

Drop the call.  It is superfluous because it duplicates the conditions
applied by the PCIe port service driver when it decides whether to
instantiate a DPC port service (see get_port_device_capability()).
If the conditions are not met, no port service is instantiated and
dpc_probe() isn't executed.

The only difference between the conditions in dpc_probe() and the ones in
get_port_device_capability() is the check for the AER capability, so by
removing the checks from dpc_probe(), DPC is supported on AER-incapable
Downstream Ports again.

Fixes: 708b20003624 ("PCI/AER: Remove HEST/FIRMWARE_FIRST parsing for AER ownership")
Signed-off-by: Lukas Wunner <lukas@wunner.de>
Cc: stable@vger.kernel.org # v5.8+
---
 drivers/pci/pcie/dpc.c | 3 ---
 1 file changed, 3 deletions(-)

diff --git a/drivers/pci/pcie/dpc.c b/drivers/pci/pcie/dpc.c
index 793a799053f1..104ff2b91f1f 100644
--- a/drivers/pci/pcie/dpc.c
+++ b/drivers/pci/pcie/dpc.c
@@ -474,9 +474,6 @@ static int dpc_probe(struct pcie_device *dev)
 	int status;
 	u16 cap;
 
-	if (!pcie_aer_is_native(pdev) && !pcie_ports_dpc_native)
-		return -ENOTSUPP;
-
 	status = devm_request_threaded_irq(device, dev->irq, dpc_irq,
 					   dpc_handler, IRQF_SHARED,
 					   "pcie-dpc", pdev);
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [PATCH 3/7] PCI/ERR: Avoid stale error status bits on recovery failure
  2026-09-27 18:02 [PATCH 0/7] Error reporting for AER-incapable devices Lukas Wunner
  2026-09-27 18:02 ` [PATCH 1/7] PCI/DPC: Avoid access to non-existent AER capability Lukas Wunner
  2026-09-27 18:02 ` [PATCH 2/7] PCI/DPC: Reinstate support for AER-incapable ports Lukas Wunner
@ 2026-09-27 18:02 ` Lukas Wunner
  2026-09-27 18:31   ` sashiko-bot
  2026-09-29 19:20   ` Kuppuswamy Sathyanarayanan
  2026-09-27 18:02 ` [PATCH 4/7] PCI/AER: Drop AER native check from handles_cxl_errors() Lukas Wunner
                   ` (3 subsequent siblings)
  6 siblings, 2 replies; 27+ messages in thread
From: Lukas Wunner @ 2026-09-27 18:02 UTC (permalink / raw)
  To: Bjorn Helgaas, Raag Jadav, Riana Tauro, Yury Murashka,
	Matthew W Carlis, linux-pci
  Cc: Mahesh J Salgaonkar, Oliver OHalloran, linuxppc-dev,
	Aravind Iddamsetty, Srinivasa Adatrao, Sathyanarayanan Kuppuswamy,
	Terry Bowman, Keith Busch

From: Yury Murashka <yurypm@arista.com>

pcie_do_recovery() only clears the error status bits of an error-reporting
device if recovery succeeds, but not if it fails.  (E.g. because reset
failed, drivers are missing pci_error_handlers or those handlers failed.)

One undesirable consequence is that the AER driver may subsequently
identify a wrong device when searching for error-reporting devices in
is_error_source(), due to stale bits on an unaffected device.

Avoid by clearing error status bits on recovery failure.

Link: https://lore.kernel.org/r/CAPzpGcRCTCZtaX1EVaJNZ103THZKsoszZduY7=gwfYdcrMo-SQ@mail.gmail.com/
Signed-off-by: Yury Murashka <yurypm@arista.com>
[lukas: drop cmdline param, rewrite commit msg, tag for stable]
Signed-off-by: Lukas Wunner <lukas@wunner.de>
Cc: stable@vger.kernel.org
---
 drivers/pci/pcie/err.c | 5 +++++
 1 file changed, 5 insertions(+)

diff --git a/drivers/pci/pcie/err.c b/drivers/pci/pcie/err.c
index d77403d8855b..cfcfc7e4f10d 100644
--- a/drivers/pci/pcie/err.c
+++ b/drivers/pci/pcie/err.c
@@ -284,6 +284,11 @@ pci_ers_result_t pcie_do_recovery(struct pci_dev *dev,
 	return status;
 
 failed:
+	if (host->native_aer || pcie_ports_native) {
+		pcie_clear_device_status(dev);
+		pci_aer_clear_nonfatal_status(dev);
+	}
+
 	pci_walk_bridge(bridge, pci_pm_runtime_put, NULL);
 
 	pci_walk_bridge(bridge, report_perm_failure_detected, NULL);
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [PATCH 4/7] PCI/AER: Drop AER native check from handles_cxl_errors()
  2026-09-27 18:02 [PATCH 0/7] Error reporting for AER-incapable devices Lukas Wunner
                   ` (2 preceding siblings ...)
  2026-09-27 18:02 ` [PATCH 3/7] PCI/ERR: Avoid stale error status bits on recovery failure Lukas Wunner
@ 2026-09-27 18:02 ` Lukas Wunner
  2026-09-27 18:26   ` sashiko-bot
                     ` (2 more replies)
  2026-09-27 18:02 ` [PATCH 5/7] PCI/AER: Move AER capability check out of pcie_aer_is_native() Lukas Wunner
                   ` (2 subsequent siblings)
  6 siblings, 3 replies; 27+ messages in thread
From: Lukas Wunner @ 2026-09-27 18:02 UTC (permalink / raw)
  To: Bjorn Helgaas, Raag Jadav, Riana Tauro, Yury Murashka,
	Matthew W Carlis, linux-pci
  Cc: Mahesh J Salgaonkar, Oliver OHalloran, linuxppc-dev,
	Aravind Iddamsetty, Srinivasa Adatrao, Sathyanarayanan Kuppuswamy,
	Terry Bowman, Keith Busch

Drop a superfluous call to pcie_aer_is_native() from handles_cxl_errors().

The call is superfluous because the function is only called from:

  aer_probe()
    cxl_rch_enable_rcec()
      handles_cxl_errors()

...and aer_probe() is only invoked if an AER port service was instantiated
in get_port_device_capability(), which is conditional on:

  (pci_pcie_type(dev) == PCI_EXP_TYPE_ROOT_PORT ||
   pci_pcie_type(dev) == PCI_EXP_TYPE_RC_EC) &&
  pci_aer_available() &&
  dev->aer_cap && (pcie_ports_native || host->native_aer))

...and the last line of those conditions is equivalent to
pcie_aer_is_native().

Signed-off-by: Lukas Wunner <lukas@wunner.de>
---
 drivers/pci/pcie/aer_cxl_rch.c | 3 +--
 1 file changed, 1 insertion(+), 2 deletions(-)

diff --git a/drivers/pci/pcie/aer_cxl_rch.c b/drivers/pci/pcie/aer_cxl_rch.c
index e471eefec9c4..9cc2d5749b9d 100644
--- a/drivers/pci/pcie/aer_cxl_rch.c
+++ b/drivers/pci/pcie/aer_cxl_rch.c
@@ -87,8 +87,7 @@ static bool handles_cxl_errors(struct pci_dev *rcec)
 {
 	bool handles_cxl = false;
 
-	if (pci_pcie_type(rcec) == PCI_EXP_TYPE_RC_EC &&
-	    pcie_aer_is_native(rcec))
+	if (pci_pcie_type(rcec) == PCI_EXP_TYPE_RC_EC)
 		pcie_walk_rcec(rcec, handles_cxl_error_iter, &handles_cxl);
 
 	return handles_cxl;
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [PATCH 5/7] PCI/AER: Move AER capability check out of pcie_aer_is_native()
  2026-09-27 18:02 [PATCH 0/7] Error reporting for AER-incapable devices Lukas Wunner
                   ` (3 preceding siblings ...)
  2026-09-27 18:02 ` [PATCH 4/7] PCI/AER: Drop AER native check from handles_cxl_errors() Lukas Wunner
@ 2026-09-27 18:02 ` Lukas Wunner
  2026-09-27 18:30   ` sashiko-bot
  2026-09-29 19:47   ` Kuppuswamy Sathyanarayanan
  2026-09-27 18:02 ` [PATCH 6/7] PCI/AER: Renumber severity constants Lukas Wunner
  2026-09-27 18:02 ` [PATCH 7/7] PCI/AER: Enable baseline capability error reporting Lukas Wunner
  6 siblings, 2 replies; 27+ messages in thread
From: Lukas Wunner @ 2026-09-27 18:02 UTC (permalink / raw)
  To: Bjorn Helgaas, Raag Jadav, Riana Tauro, Yury Murashka,
	Matthew W Carlis, linux-pci
  Cc: Mahesh J Salgaonkar, Oliver OHalloran, linuxppc-dev,
	Aravind Iddamsetty, Srinivasa Adatrao, Sathyanarayanan Kuppuswamy,
	Terry Bowman, Keith Busch

When firmware grants control of Advanced Error Reporting to the operating
system, that not only encompasses the AER capability, but also error
enable/status bits in the Device Control and Device Status registers
(PCI Firmware r3.3 table 4-6 bit 3).

PCIe devices without AER capability still support baseline capability
error reporting through these enable/status bits (PCIe r7.1 sec 6.2.1),
but the bits must not be modified unless AER control was granted.

pcie_aer_is_native() is unsuitable to check for control of AER-incapable
devices because it implicitly checks for presence of an AER capability.

Move that check to its callers (where needed) to allow using the function
for the imminent baseline capability error reporting.

Signed-off-by: Lukas Wunner <lukas@wunner.de>
---
 drivers/pci/pcie/aer.c | 7 ++-----
 1 file changed, 2 insertions(+), 5 deletions(-)

diff --git a/drivers/pci/pcie/aer.c b/drivers/pci/pcie/aer.c
index d8dcd238fda1..34a8eddc427a 100644
--- a/drivers/pci/pcie/aer.c
+++ b/drivers/pci/pcie/aer.c
@@ -257,9 +257,6 @@ int pcie_aer_is_native(struct pci_dev *dev)
 {
 	struct pci_host_bridge *host = pci_find_host_bridge(dev->bus);
 
-	if (!dev->aer_cap)
-		return 0;
-
 	return pcie_ports_native || host->native_aer;
 }
 EXPORT_SYMBOL_NS_GPL(pcie_aer_is_native, "CXL");
@@ -280,7 +277,7 @@ int pci_aer_clear_nonfatal_status(struct pci_dev *dev)
 	int aer = dev->aer_cap;
 	u32 status, sev;
 
-	if (!pcie_aer_is_native(dev))
+	if (!aer || !pcie_aer_is_native(dev))
 		return -EIO;
 
 	/* Clear status bits for ERR_NONFATAL errors only */
@@ -299,7 +296,7 @@ void pci_aer_clear_fatal_status(struct pci_dev *dev)
 	int aer = dev->aer_cap;
 	u32 status, sev;
 
-	if (!pcie_aer_is_native(dev))
+	if (!aer || !pcie_aer_is_native(dev))
 		return;
 
 	/* Clear status bits for ERR_FATAL errors only */
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [PATCH 6/7] PCI/AER: Renumber severity constants
  2026-09-27 18:02 [PATCH 0/7] Error reporting for AER-incapable devices Lukas Wunner
                   ` (4 preceding siblings ...)
  2026-09-27 18:02 ` [PATCH 5/7] PCI/AER: Move AER capability check out of pcie_aer_is_native() Lukas Wunner
@ 2026-09-27 18:02 ` Lukas Wunner
  2026-09-27 18:29   ` sashiko-bot
  2026-09-29 19:55   ` Kuppuswamy Sathyanarayanan
  2026-09-27 18:02 ` [PATCH 7/7] PCI/AER: Enable baseline capability error reporting Lukas Wunner
  6 siblings, 2 replies; 27+ messages in thread
From: Lukas Wunner @ 2026-09-27 18:02 UTC (permalink / raw)
  To: Bjorn Helgaas, Raag Jadav, Riana Tauro, Yury Murashka,
	Matthew W Carlis, linux-pci
  Cc: Mahesh J Salgaonkar, Oliver OHalloran, linuxppc-dev,
	Aravind Iddamsetty, Srinivasa Adatrao, Sathyanarayanan Kuppuswamy,
	Terry Bowman, Keith Busch

The AER driver uses constants for Correctable, Non-Fatal and Fatal Error
severity which look as if they match something in the spec, but are
actually just made-up numbers.

Use the bit number in the Device Status Register instead (PCIe r7.1 sec
7.5.3.5).  A subsequent commit takes advantage of this by using BIT() to
conveniently compute the register bit corresponding to a given severity.

While at it, drop the DPC_FATAL constant which was introduced by commit
b09803b5e546 ("PCI/DPC: Use the generic pcie_do_fatal_recovery() path")
but never saw any use in 8 years.

Signed-off-by: Lukas Wunner <lukas@wunner.de>
---
 drivers/pci/pcie/aer.c  | 2 +-
 include/linux/aer.h     | 8 ++++----
 include/ras/ras_event.h | 2 +-
 3 files changed, 6 insertions(+), 6 deletions(-)

diff --git a/drivers/pci/pcie/aer.c b/drivers/pci/pcie/aer.c
index 34a8eddc427a..6bc843ab9b37 100644
--- a/drivers/pci/pcie/aer.c
+++ b/drivers/pci/pcie/aer.c
@@ -504,9 +504,9 @@ void pci_aer_exit(struct pci_dev *dev)
  * AER error strings
  */
 static const char * const aer_error_severity_string[] = {
+	"Correctable",
 	"Uncorrectable (Non-Fatal)",
 	"Uncorrectable (Fatal)",
-	"Correctable"
 };
 
 static const char *aer_error_layer[] = {
diff --git a/include/linux/aer.h b/include/linux/aer.h
index df0f5c382286..795c55132008 100644
--- a/include/linux/aer.h
+++ b/include/linux/aer.h
@@ -11,10 +11,10 @@
 #include <linux/errno.h>
 #include <linux/types.h>
 
-#define AER_NONFATAL			0
-#define AER_FATAL			1
-#define AER_CORRECTABLE			2
-#define DPC_FATAL			3
+/* Must match bit number in Device Status Register (PCIe r7.1, sec 7.5.3.5) */
+#define AER_CORRECTABLE			0
+#define AER_NONFATAL			1
+#define AER_FATAL			2
 
 /*
  * AER and DPC capabilities TLP Logging register sizes (PCIe r6.2, sec 7.8.4
diff --git a/include/ras/ras_event.h b/include/ras/ras_event.h
index fdb785fa4613..1e03ade99386 100644
--- a/include/ras/ras_event.h
+++ b/include/ras/ras_event.h
@@ -302,7 +302,7 @@ TRACE_EVENT(non_standard_event,
  *			([domain:]bus:device.function).
  * u32 status -		Either the correctable or uncorrectable register
  *			indicating what error or errors have been seen
- * u8 severity -	error severity 0:NONFATAL 1:FATAL 2:CORRECTED
+ * u8 severity -	error severity 0:CORRECTED 1:NONFATAL 2:FATAL
  */
 
 #define aer_correctable_errors					\
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [PATCH 7/7] PCI/AER: Enable baseline capability error reporting
  2026-09-27 18:02 [PATCH 0/7] Error reporting for AER-incapable devices Lukas Wunner
                   ` (5 preceding siblings ...)
  2026-09-27 18:02 ` [PATCH 6/7] PCI/AER: Renumber severity constants Lukas Wunner
@ 2026-09-27 18:02 ` Lukas Wunner
  2026-09-27 18:33   ` sashiko-bot
  2026-09-29 21:01   ` Kuppuswamy Sathyanarayanan
  6 siblings, 2 replies; 27+ messages in thread
From: Lukas Wunner @ 2026-09-27 18:02 UTC (permalink / raw)
  To: Bjorn Helgaas, Raag Jadav, Riana Tauro, Yury Murashka,
	Matthew W Carlis, linux-pci
  Cc: Mahesh J Salgaonkar, Oliver OHalloran, linuxppc-dev,
	Aravind Iddamsetty, Srinivasa Adatrao, Sathyanarayanan Kuppuswamy,
	Terry Bowman, Keith Busch

PCIe r7.1 sec 6.2.1 defines two error reporting paradigms:  Baseline
capability and Advanced Error Reporting Extended Capability (AER).

So far the kernel only supports the AER paradigm.  Enable the baseline
capability paradigm to complete the kernel's support for error reporting.

Baseline capability (solely) relies on the error enable/status bits in the
Device Control and Device Status registers, which exist on every PCIe
device.  One difference between baseline capability and AER is that the
severity of Uncorrectable Errors cannot be controlled.  Another difference
is that Advisory Non-Fatal Errors are not reported at all (PCIe r7.1 sec
6.2.5).

But the main difference to AER is that no detailed error information is
available:  It is only known that an error occurred and which severity it
had, but not its type or the Prefix / Header of the TLP that caused it.
An exception are Unsupported Request Errors which are signaled with a
dedicated bit in the Device Status register.  Report these as if the
Unsupported Request Error Status bit in the Uncorrectable Error Status
register was set, for consistency with the AER paradigm.

Error recovery works the same for baseline capability as it already does
for AER:  Drivers are informed about errors through the pci_error_handlers
callbacks.  Recovery from Fatal Errors (and from Non-Fatal Errors, if
chosen by drivers) is attempted through a Secondary Bus Reset.

Similarly to commit f26e58bf6f54 ("PCI/AER: Enable error reporting when
AER is native"), which universally enabled error reporting on AER-capable
devices, the present commit is invasive because it universally enables
error reporting on non-AER-capable devices.  Previously those errors were
neither reported nor recovered.  It may be necessary to amend more drivers
with pci_error_handlers callbacks to recover from newly reported errors.

Baseline capability is only enabled if firmware grants AER control to the
operating system.  Otherwise firmware owns the error enable/status bits in
the Device Control and Device Status registers (PCI Firmware r3.3 table
4-6 bit 3).

Baseline capability support is initially only implemented for native error
handling, not Firmware First error handling:  ghes_handle_aer() would have
to be amended to pass the Device Control and Device Status registers to
aer_recover_queue(), and to only pass an AER capability structure if it is
present in the CPER record.  It's not clear whether any firmware actually
generates such CPER records and whether implementing support for it is
worthwhile, so postpone that for now.

Suggested-by: Bjorn Helgaas <bhelgaas@google.com>
Link: https://lore.kernel.org/r/20260826212619.GA1566339@bhelgaas/
Signed-off-by: Lukas Wunner <lukas@wunner.de>
---
 .../ABI/testing/sysfs-bus-pci-devices-aer     | 26 ++++++++----
 Documentation/PCI/pcieaer-howto.rst           |  9 ++--
 drivers/pci/pcie/aer.c                        | 42 ++++++++++++-------
 drivers/pci/pcie/dpc.c                        |  2 +-
 4 files changed, 50 insertions(+), 29 deletions(-)

diff --git a/Documentation/ABI/testing/sysfs-bus-pci-devices-aer b/Documentation/ABI/testing/sysfs-bus-pci-devices-aer
index 5ed284523956..9e772e1ed9b3 100644
--- a/Documentation/ABI/testing/sysfs-bus-pci-devices-aer
+++ b/Documentation/ABI/testing/sysfs-bus-pci-devices-aer
@@ -1,9 +1,9 @@
 PCIe Device AER statistics
 --------------------------
 
-These attributes show up under all the devices that are AER capable. These
+These attributes show up under all PCIe devices (AER capable or not). These
 statistical counters indicate the errors "as seen/reported by the device".
-Note that this may mean that if an endpoint is causing problems, the AER
+Note that this may mean that if an endpoint is causing problems, the error
 counters may increment at its link partner (e.g. root port) because the
 errors may be "seen" / reported by the link partner and not the
 problematic endpoint itself (which may report all counters as 0 as it never
@@ -17,7 +17,10 @@ Description:	List of correctable errors seen and reported by this
 		PCI device using ERR_COR. Note that since multiple errors may
 		be reported using a single ERR_COR message, thus
 		TOTAL_ERR_COR at the end of the file may not match the actual
-		total of all the errors in the file. Sample output::
+		total of all the errors in the file.
+		For PCIe devices without AER Extended Capability, only the
+		total counter is meaningful while all other counters remain 0.
+		Sample output::
 
 		    localhost /sys/devices/pci0000:00/0000:00:1c.0 # cat aer_dev_correctable
 		    Receiver Error 2
@@ -38,7 +41,10 @@ Description:	List of uncorrectable fatal errors seen and reported by this
 		PCI device using ERR_FATAL. Note that since multiple errors may
 		be reported using a single ERR_FATAL message, thus
 		TOTAL_ERR_FATAL at the end of the file may not match the actual
-		total of all the errors in the file. Sample output::
+		total of all the errors in the file.
+		For PCIe devices without AER Extended Capability, only the
+		total counter is meaningful while all other counters remain 0.
+		Sample output::
 
 		    localhost /sys/devices/pci0000:00/0000:00:1c.0 # cat aer_dev_fatal
 		    Undefined 0
@@ -68,7 +74,11 @@ Description:	List of uncorrectable nonfatal errors seen and reported by this
 		PCI device using ERR_NONFATAL. Note that since multiple errors
 		may be reported using a single ERR_FATAL message, thus
 		TOTAL_ERR_NONFATAL at the end of the file may not match the
-		actual total of all the errors in the file. Sample output::
+		actual total of all the errors in the file.
+		For PCIe devices without AER Extended Capability, only the
+		total counter and the Unsupported Request counter is meaningful
+		while all other counters remain 0.
+		Sample output::
 
 		    localhost /sys/devices/pci0000:00/0000:00:1c.0 # cat aer_dev_nonfatal
 		    Undefined 0
@@ -121,7 +131,7 @@ Description:	Total number of ERR_NONFATAL messages reported to rootport.
 PCIe AER ratelimits
 -------------------
 
-These attributes show up under all the devices that are AER capable.
+These attributes show up under all PCIe devices (AER capable or not).
 They represent configurable ratelimits of logs per error type.
 
 See Documentation/PCI/pcieaer-howto.rst for more info on ratelimits.
@@ -130,7 +140,7 @@ What:		/sys/bus/pci/devices/<dev>/aer/correctable_ratelimit_interval_ms
 Date:		May 2025
 KernelVersion:	6.16.0
 Contact:	linux-pci@vger.kernel.org
-Description:	Writing 0 disables AER correctable error log ratelimiting.
+Description:	Writing 0 disables Correctable Error log ratelimiting.
 		Writing a positive value sets the ratelimit interval in ms.
 		Default is DEFAULT_RATELIMIT_INTERVAL (5000 ms).
 
@@ -147,7 +157,7 @@ What:		/sys/bus/pci/devices/<dev>/aer/nonfatal_ratelimit_interval_ms
 Date:		May 2025
 KernelVersion:	6.16.0
 Contact:	linux-pci@vger.kernel.org
-Description:	Writing 0 disables AER non-fatal uncorrectable error log
+Description:	Writing 0 disables Non-Fatal Uncorrectable Error log
 		ratelimiting. Writing a positive value sets the ratelimit
 		interval in ms. Default is DEFAULT_RATELIMIT_INTERVAL
 		(5000 ms).
diff --git a/Documentation/PCI/pcieaer-howto.rst b/Documentation/PCI/pcieaer-howto.rst
index 90fdfddd3ae5..e969e563b61d 100644
--- a/Documentation/PCI/pcieaer-howto.rst
+++ b/Documentation/PCI/pcieaer-howto.rst
@@ -34,9 +34,8 @@ set of error reporting requirements. Advanced Error Reporting
 capability is implemented with a PCIe Advanced Error Reporting
 extended capability structure providing more robust error reporting.
 
-The PCIe AER driver provides the infrastructure to support PCIe Advanced
-Error Reporting capability. The PCIe AER driver provides three basic
-functions:
+The PCIe AER driver provides the infrastructure to support both paradigms.
+It provides three basic functions:
 
   - Gathers the comprehensive error information if errors occurred.
   - Reports error to the users.
@@ -69,7 +68,7 @@ Specification for details regarding _OSC usage.
 AER error output
 ----------------
 
-When a PCIe AER error is captured, an error message will be output to
+When a PCIe error is captured, an error message will be output to
 console. If it's a correctable error, it is output as a warning message.
 Otherwise, it is printed as an error. So users could choose different
 log level to filter out correctable error messages.
@@ -113,7 +112,7 @@ See Documentation/ABI/testing/sysfs-bus-pci-devices-aer.
 AER Statistics / Counters
 -------------------------
 
-When PCIe AER errors are captured, the counters / statistics are also exposed
+When PCIe errors are captured, the counters / statistics are also exposed
 in the form of sysfs attributes which are documented at
 Documentation/ABI/testing/sysfs-bus-pci-devices-aer.
 
diff --git a/drivers/pci/pcie/aer.c b/drivers/pci/pcie/aer.c
index 6bc843ab9b37..62376aab4b6f 100644
--- a/drivers/pci/pcie/aer.c
+++ b/drivers/pci/pcie/aer.c
@@ -397,21 +397,22 @@ void pci_aer_init(struct pci_dev *dev)
 {
 	int n;
 
-	dev->aer_cap = pci_find_ext_capability(dev, PCI_EXT_CAP_ID_ERR);
-	if (!dev->aer_cap)
+	if (!pci_is_pcie(dev))
 		return;
 
 	dev->aer_info = kzalloc_obj(*dev->aer_info);
-	if (!dev->aer_info) {
-		dev->aer_cap = 0;
+	if (!dev->aer_info)
 		return;
-	}
 
 	ratelimit_state_init(&dev->aer_info->correctable_ratelimit,
 			     DEFAULT_RATELIMIT_INTERVAL, DEFAULT_RATELIMIT_BURST);
 	ratelimit_state_init(&dev->aer_info->nonfatal_ratelimit,
 			     DEFAULT_RATELIMIT_INTERVAL, DEFAULT_RATELIMIT_BURST);
 
+	dev->aer_cap = pci_find_ext_capability(dev, PCI_EXT_CAP_ID_ERR);
+	if (!dev->aer_cap)
+		goto enable;
+
 	/*
 	 * We save/restore PCI_ERR_UNCOR_MASK, PCI_ERR_UNCOR_SEVER,
 	 * PCI_ERR_COR_MASK, and PCI_ERR_CAP.  Root and Root Complex Event
@@ -431,6 +432,9 @@ void pci_aer_init(struct pci_dev *dev)
 					       PCI_ERR_COR_ADV_NFAT, 0);
 
 	pci_aer_clear_status(dev);
+enable:
+	if (pcie_aer_is_native(dev))
+		pcie_clear_device_status(dev);
 
 	if (pci_aer_available())
 		pci_enable_pcie_error_reporting(dev);
@@ -980,15 +984,12 @@ void aer_print_error(struct aer_err_info *info, int i)
 	if (!info->ratelimit_print[i])
 		goto anfe;
 
-	if (!info->status) {
-		pci_err(dev, "%s Bus Error: severity=%s (Inaccessible)\n",
-			bus_type, aer_error_severity_string[info->severity]);
-		return;
-	}
-
 	aer_printk(level, dev, "%s Bus Error: severity=%s\n",
 		   bus_type, aer_error_severity_string[info->severity]);
 
+	if (!info->status)
+		return;
+
 	aer_printk(level, dev, "  device [%04x:%04x] error status/mask=%08x/%08x\n",
 		   dev->vendor, dev->device, info->status, info->mask);
 
@@ -1182,8 +1183,11 @@ static bool is_error_source(struct pci_dev *dev, struct aer_err_info *e_info)
 	if (!(reg16 & PCI_EXP_AER_FLAGS))
 		return false;
 
-	if (!aer)
-		return false;
+	if (!aer) {
+		pcie_capability_read_word(dev, PCI_EXP_DEVSTA, &reg16);
+		return reg16 & BIT(e_info->severity) &&
+		       !PCI_POSSIBLE_ERROR(reg16);
+	}
 
 	/* Check if error is recorded */
 	if (e_info->severity == AER_CORRECTABLE) {
@@ -1455,6 +1459,7 @@ int aer_get_device_error_info(struct aer_err_info *info, int i)
 {
 	struct pci_dev *dev;
 	int type, aer;
+	u16 devsta;
 
 	if (i >= AER_MAX_MULTI_ERR_DEVICES)
 		return 0;
@@ -1470,8 +1475,15 @@ int aer_get_device_error_info(struct aer_err_info *info, int i)
 	info->is_cxl = pcie_is_cxl(dev);
 
 	/* The device might not support AER */
-	if (!aer)
-		return 0;
+	if (!aer) {
+		if (info->severity != AER_CORRECTABLE) {
+			pcie_capability_read_word(dev, PCI_EXP_DEVSTA, &devsta);
+			if (devsta & PCI_EXP_DEVSTA_URD &&
+			    !PCI_POSSIBLE_ERROR(devsta))
+				info->status = PCI_ERR_UNC_UNSUP;
+		}
+		return 1;
+	}
 
 	if (info->severity == AER_CORRECTABLE) {
 		pci_read_config_dword(dev, aer + PCI_ERR_COR_STATUS,
diff --git a/drivers/pci/pcie/dpc.c b/drivers/pci/pcie/dpc.c
index 104ff2b91f1f..fdea3db61a4a 100644
--- a/drivers/pci/pcie/dpc.c
+++ b/drivers/pci/pcie/dpc.c
@@ -269,7 +269,7 @@ void dpc_process_error(struct pci_dev *pdev)
 		pci_warn(pdev, "containment event, status:%#06x: unmasked uncorrectable error detected\n",
 			 status);
 		if (dpc_get_aer_uncorrect_severity(pdev, &info) &&
-		    (aer_get_device_error_info(&info, 0) || !pdev->aer_cap)) {
+		    aer_get_device_error_info(&info, 0)) {
 			aer_print_error(&info, 0);
 			pci_aer_clear_nonfatal_status(pdev);
 			pci_aer_clear_fatal_status(pdev);
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* Re: [PATCH 4/7] PCI/AER: Drop AER native check from handles_cxl_errors()
  2026-09-27 18:02 ` [PATCH 4/7] PCI/AER: Drop AER native check from handles_cxl_errors() Lukas Wunner
@ 2026-09-27 18:26   ` sashiko-bot
  2026-09-28 20:58   ` Bowman, Terry
  2026-09-29 19:24   ` Kuppuswamy Sathyanarayanan
  2 siblings, 0 replies; 27+ messages in thread
From: sashiko-bot @ 2026-09-27 18:26 UTC (permalink / raw)
  To: Lukas Wunner; +Cc: linux-pci

> Drop a superfluous call to pcie_aer_is_native() from handles_cxl_errors().
> 
> The call is superfluous because the function is only called from:
> 
>   aer_probe()
>     cxl_rch_enable_rcec()
>       handles_cxl_errors()
> 
> ...and aer_probe() is only invoked if an AER port service was instantiated
> in get_port_device_capability(), which is conditional on:
> 
>   (pci_pcie_type(dev) == PCI_EXP_TYPE_ROOT_PORT ||
>    pci_pcie_type(dev) == PCI_EXP_TYPE_RC_EC) &&
>   pci_aer_available() &&
>   dev->aer_cap && (pcie_ports_native || host->native_aer))
> 
> ...and the last line of those conditions is equivalent to
> pcie_aer_is_native().
> 
> Signed-off-by: Lukas Wunner <lukas@wunner.de>

Sashiko has reviewed this patch and found no issues. It looks great!

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/cover.1790531238.git.lukas@wunner.de?part=4


^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH 2/7] PCI/DPC: Reinstate support for AER-incapable ports
  2026-09-27 18:02 ` [PATCH 2/7] PCI/DPC: Reinstate support for AER-incapable ports Lukas Wunner
@ 2026-09-27 18:26   ` sashiko-bot
  2026-09-29 19:04   ` Kuppuswamy Sathyanarayanan
  1 sibling, 0 replies; 27+ messages in thread
From: sashiko-bot @ 2026-09-27 18:26 UTC (permalink / raw)
  To: Lukas Wunner; +Cc: linux-pci

> Downstream Port Containment does not mandate presence of an Advanced Error
> Reporting capability, so a Downstream Port may support DPC, but not AER
> (PCIe r7.1 sec 6.2.11.2).
> 
> Commit 708b20003624 ("PCI/AER: Remove HEST/FIRMWARE_FIRST parsing for AER
> ownership") seemingly inadvertently constrained DPC to AER-supporting
> Downstream Ports by inserting a call to pcie_aer_is_native() in
> dpc_probe(), which implicitly checks for presence of an AER capability.
> 
> Drop the call.  It is superfluous because it duplicates the conditions
> applied by the PCIe port service driver when it decides whether to
> instantiate a DPC port service (see get_port_device_capability()).
> If the conditions are not met, no port service is instantiated and
> dpc_probe() isn't executed.
> 
> The only difference between the conditions in dpc_probe() and the ones in
> get_port_device_capability() is the check for the AER capability, so by
> removing the checks from dpc_probe(), DPC is supported on AER-incapable
> Downstream Ports again.
> 
> Fixes: 708b20003624 ("PCI/AER: Remove HEST/FIRMWARE_FIRST parsing for AER ownership")
> Signed-off-by: Lukas Wunner <lukas@wunner.de>

Sashiko has reviewed this patch and found no issues. It looks great!

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/cover.1790531238.git.lukas@wunner.de?part=2


^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH 1/7] PCI/DPC: Avoid access to non-existent AER capability
  2026-09-27 18:02 ` [PATCH 1/7] PCI/DPC: Avoid access to non-existent AER capability Lukas Wunner
@ 2026-09-27 18:27   ` sashiko-bot
  2026-09-29 18:49   ` Kuppuswamy Sathyanarayanan
  1 sibling, 0 replies; 27+ messages in thread
From: sashiko-bot @ 2026-09-27 18:27 UTC (permalink / raw)
  To: Lukas Wunner; +Cc: linux-pci

> Downstream Port Containment does not mandate presence of an Advanced Error
> Reporting capability, so a Downstream Port may support DPC, but not AER
> (PCIe r7.1 sec 6.2.11.2).
> 
> In February 2019, commit 9f08a5d896ce ("PCI/DPC: Fix print AER status in
> DPC event handling") amended the DPC driver to access the AER capability
> without checking for its presence.
> 
> In May 2020, commit 708b20003624 ("PCI/AER: Remove HEST/FIRMWARE_FIRST
> parsing for AER ownership") fixed it by inserting a call to
> pcie_aer_is_native() in dpc_probe(), which implicitly checks for presence
> of an AER capability.
> 
> However already in October 2019, commit 35a0b2378c19 ("PCI/DPC: Add
> "pcie_ports=dpc-native" to allow DPC without AER control") made it
> [ ... ]
> for AER-unsupporting devices.  The check will become obsolete and will be
> removed with the imminent baseline capability error reporting.
> 
> Fixes: 35a0b2378c19 ("PCI/DPC: Add "pcie_ports=dpc-native" to allow DPC without AER control")
> Signed-off-by: Lukas Wunner <lukas@wunner.de>

Sashiko has reviewed this patch and found no issues. It looks great!

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/cover.1790531238.git.lukas@wunner.de?part=1


^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH 6/7] PCI/AER: Renumber severity constants
  2026-09-27 18:02 ` [PATCH 6/7] PCI/AER: Renumber severity constants Lukas Wunner
@ 2026-09-27 18:29   ` sashiko-bot
  2026-09-27 19:23     ` Lukas Wunner
  2026-09-29 19:55   ` Kuppuswamy Sathyanarayanan
  1 sibling, 1 reply; 27+ messages in thread
From: sashiko-bot @ 2026-09-27 18:29 UTC (permalink / raw)
  To: Lukas Wunner; +Cc: linux-pci

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Critical] ras/aer: ABI breakage in aer_event tracepoint severity values
--

commit 1b45a2c88d40942123bcb76f3b3bbf8eb9ec2cac
Author: Lukas Wunner <lukas@wunner.de>

PCI/AER: Renumber severity constants

This commit modifies the AER severity constants to align with the Device Status
Register bit numbers defined in the PCIe specification. It updates the severity
strings and the ras_event documentation to reflect the new constant values.

> diff --git a/include/linux/aer.h b/include/linux/aer.h
> index df0f5c382286f..795c551320088 100644
> --- a/include/linux/aer.h
> +++ b/include/linux/aer.h
> @@ -11,10 +11,10 @@
>  #include <linux/errno.h>
>  #include <linux/types.h>
>  
> -#define AER_NONFATAL			0
> -#define AER_FATAL			1
> -#define AER_CORRECTABLE			2
> -#define DPC_FATAL			3
> +/* Must match bit number in Device Status Register (PCIe r7.1, sec 7.5.3.5) */
> +#define AER_CORRECTABLE			0
> +#define AER_NONFATAL			1
> +#define AER_FATAL			2

[ ... ]

> diff --git a/include/ras/ras_event.h b/include/ras/ras_event.h
> index fdb785fa4613a..1e03ade99386b 100644
> --- a/include/ras/ras_event.h
> +++ b/include/ras/ras_event.h
> @@ -302,7 +302,7 @@ TRACE_EVENT(non_standard_event,
>   *			([domain:]bus:device.function).
>   * u32 status -		Either the correctable or uncorrectable register
>   *			indicating what error or errors have been seen
> - * u8 severity -	error severity 0:NONFATAL 1:FATAL 2:CORRECTED
> + * u8 severity -	error severity 0:CORRECTED 1:NONFATAL 2:FATAL
>   */

[Severity: Critical]
Does this change unintentionally break the userspace ABI for the aer_event
tracepoint?

The tracepoint exports the severity field as a raw u8 integer in the binary
trace ring buffer. By renumbering the underlying AER constants in
include/linux/aer.h, the tracepoint now emits different integers for hardware
error events without an ABI translation layer.

Userspace tools like rasdaemon that rely on the historical raw integer values
will now misclassify hardware error severities. For example, a correctable
error (now 0) would be incorrectly interpreted as non-fatal by older tools
expecting the previous ABI layout.

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/cover.1790531238.git.lukas@wunner.de?part=6

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH 5/7] PCI/AER: Move AER capability check out of pcie_aer_is_native()
  2026-09-27 18:02 ` [PATCH 5/7] PCI/AER: Move AER capability check out of pcie_aer_is_native() Lukas Wunner
@ 2026-09-27 18:30   ` sashiko-bot
  2026-09-29 19:47   ` Kuppuswamy Sathyanarayanan
  1 sibling, 0 replies; 27+ messages in thread
From: sashiko-bot @ 2026-09-27 18:30 UTC (permalink / raw)
  To: Lukas Wunner; +Cc: linux-pci

> When firmware grants control of Advanced Error Reporting to the operating
> system, that not only encompasses the AER capability, but also error
> enable/status bits in the Device Control and Device Status registers
> (PCI Firmware r3.3 table 4-6 bit 3).
> 
> PCIe devices without AER capability still support baseline capability
> error reporting through these enable/status bits (PCIe r7.1 sec 6.2.1),
> but the bits must not be modified unless AER control was granted.
> 
> pcie_aer_is_native() is unsuitable to check for control of AER-incapable
> devices because it implicitly checks for presence of an AER capability.
> 
> Move that check to its callers (where needed) to allow using the function
> for the imminent baseline capability error reporting.
> 
> Signed-off-by: Lukas Wunner <lukas@wunner.de>

Sashiko has reviewed this patch and found no issues. It looks great!

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/cover.1790531238.git.lukas@wunner.de?part=5


^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH 3/7] PCI/ERR: Avoid stale error status bits on recovery failure
  2026-09-27 18:02 ` [PATCH 3/7] PCI/ERR: Avoid stale error status bits on recovery failure Lukas Wunner
@ 2026-09-27 18:31   ` sashiko-bot
  2026-09-27 18:57     ` Lukas Wunner
  2026-09-29 19:20   ` Kuppuswamy Sathyanarayanan
  1 sibling, 1 reply; 27+ messages in thread
From: sashiko-bot @ 2026-09-27 18:31 UTC (permalink / raw)
  To: Lukas Wunner; +Cc: linux-pci

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] pci: stale fatal AER status bits left uncleared on recovery failure path
--

commit f65c6a5d555f5001060d8335ce4ca73c0130b547
Author: Yury Murashka <yurypm@arista.com>
PCI/ERR: Avoid stale error status bits on recovery failure

This commit clears error status bits on recovery failure to prevent the
AER driver from falsely identifying unaffected devices as the source of
new errors. It does this by calling pcie_clear_device_status() and
pci_aer_clear_nonfatal_status() on the failure path in pcie_do_recovery().

Link: https://lore.kernel.org/r/CAPzpGcRCTCZtaX1EVaJNZ103THZKsoszZduY7=gwfYdcrMo-SQ@mail.gmail.com/

> diff --git a/drivers/pci/pcie/err.c b/drivers/pci/pcie/err.c
> --- a/drivers/pci/pcie/err.c
> +++ b/drivers/pci/pcie/err.c
> @@ -284,6 +284,11 @@ pci_ers_result_t pcie_do_recovery(struct pci_dev *dev,
>  	return status;
>  
>  failed:
> +	if (host->native_aer || pcie_ports_native) {
> +		pcie_clear_device_status(dev);
> +		pci_aer_clear_nonfatal_status(dev);

[Severity: Medium]
When pcie_do_recovery() is invoked for a fatal error (pci_channel_io_frozen)
and recovery fails, will this leave the fatal error bits uncleared?

Looking at pci_aer_clear_nonfatal_status() in drivers/pci/pcie/aer.c, it
explicitly masks out fatal errors before writing back to the status register:

    pci_read_config_dword(dev, aer + PCI_ERR_UNCOR_STATUS, &status);
    pci_read_config_dword(dev, aer + PCI_ERR_UNCOR_SEVER, &sev);
    status &= ~sev;
    if (status)
        pci_write_config_dword(dev, aer + PCI_ERR_UNCOR_STATUS, status);

If a fatal error occurred, the severity bit is 1, so the status bit is never
cleared. Could this leave stale fatal error bits that cause subsequent
is_error_source() checks to falsely identify this device?

Should this failure path use pci_aer_clear_status() or explicitly clear all
uncorrectable errors instead?

> +	}
> +
>  	pci_walk_bridge(bridge, pci_pm_runtime_put, NULL);
>  
>  	pci_walk_bridge(bridge, report_perm_failure_detected, NULL);

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/cover.1790531238.git.lukas@wunner.de?part=3

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH 7/7] PCI/AER: Enable baseline capability error reporting
  2026-09-27 18:02 ` [PATCH 7/7] PCI/AER: Enable baseline capability error reporting Lukas Wunner
@ 2026-09-27 18:33   ` sashiko-bot
  2026-09-29 21:01   ` Kuppuswamy Sathyanarayanan
  1 sibling, 0 replies; 27+ messages in thread
From: sashiko-bot @ 2026-09-27 18:33 UTC (permalink / raw)
  To: Lukas Wunner; +Cc: linux-pci

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] pci/aer: uncorrectable errors dropped for non-AER devices due to DEVSTA clear
- [Medium] pci/aer: leaked mask/first_error state to non-AER devices
--

commit 29da5e077bdaa4990a1072f7b27ccf2bb9cae518
Author: Lukas Wunner <lukas@wunner.de>

    PCI/AER: Enable baseline capability error reporting
    
    This commit enables PCIe baseline capability error reporting for devices
    that do not support AER. It relies on the error enable/status bits in the
    Device Control and Device Status registers. This completes the kernel's
    support for native error reporting.

Link: https://lore.kernel.org/r/20260826212619.GA1566339@bhelgaas/

> diff --git a/drivers/pci/pcie/aer.c b/drivers/pci/pcie/aer.c
> index 6bc843ab9b37f..62376aab4b6f1 100644
> --- a/drivers/pci/pcie/aer.c
> +++ b/drivers/pci/pcie/aer.c
[ ... ]
> @@ -1182,8 +1183,11 @@ static bool is_error_source(struct pci_dev *dev, struct aer_err_info *e_info)
>  	if (!(reg16 & PCI_EXP_AER_FLAGS))
>  		return false;
>  
> -	if (!aer)
> -		return false;
> +	if (!aer) {
> +		pcie_capability_read_word(dev, PCI_EXP_DEVSTA, &reg16);
> +		return reg16 & BIT(e_info->severity) &&
> +		       !PCI_POSSIBLE_ERROR(reg16);
> +	}

[Severity: High]
When a Root Port receives both a correctable and an uncorrectable error,
aer_isr_one_error() processes the correctable error first.

During this, pci_aer_handle_error() ends by unconditionally calling
pcie_clear_device_status():

pci_aer_handle_error() {
    ...
        if (pdrv && pdrv->err_handler &&
            pdrv->err_handler->cor_error_detected)
            pdrv->err_handler->cor_error_detected(dev);
        pcie_clear_device_status(dev);
}

This clears all error bits in PCI_EXP_DEVSTA (CED, NFED, FED, and URD).

Because this new logic in is_error_source() relies entirely on the NFED or
FED bit in PCI_EXP_DEVSTA to locate uncorrectable errors on non-AER devices,
won't this cause pending uncorrectable errors to be silently dropped?

[ ... ]
> @@ -1455,6 +1459,7 @@ int aer_get_device_error_info(struct aer_err_info *info, int i)
>  {
>  	struct pci_dev *dev;
>  	int type, aer;
> +	u16 devsta;
>  
>  	if (i >= AER_MAX_MULTI_ERR_DEVICES)
>  		return 0;
> @@ -1470,8 +1475,15 @@ int aer_get_device_error_info(struct aer_err_info *info, int i)
>  	info->is_cxl = pcie_is_cxl(dev);
>  
>  	/* The device might not support AER */
> -	if (!aer)
> -		return 0;
> +	if (!aer) {
> +		if (info->severity != AER_CORRECTABLE) {
> +			pcie_capability_read_word(dev, PCI_EXP_DEVSTA, &devsta);
> +			if (devsta & PCI_EXP_DEVSTA_URD &&
> +			    !PCI_POSSIBLE_ERROR(devsta))
> +				info->status = PCI_ERR_UNC_UNSUP;
> +		}
> +		return 1;
> +	}

[Severity: Medium]
When aer_process_err_devices() processes multiple devices in a loop, it uses
a shared 'info' structure. Since this early return for non-AER devices
doesn't reset info->mask or info->first_error, will it leak state from the
previously processed AER device?

This leaked mask could be used by pci_dev_aer_stats_incr() to filter errors,
potentially dropping statistics for valid errors if the leaked mask covers
the PCI_ERR_UNC_UNSUP bit.

Should info->mask and info->first_error be explicitly cleared here?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/cover.1790531238.git.lukas@wunner.de?part=7

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH 3/7] PCI/ERR: Avoid stale error status bits on recovery failure
  2026-09-27 18:31   ` sashiko-bot
@ 2026-09-27 18:57     ` Lukas Wunner
  0 siblings, 0 replies; 27+ messages in thread
From: Lukas Wunner @ 2026-09-27 18:57 UTC (permalink / raw)
  To: sashiko-reviews; +Cc: linux-pci

On Sun, Sep 27, 2026 at 06:31:22PM +0000, sashiko-bot@kernel.org wrote:
> > +++ b/drivers/pci/pcie/err.c
> > @@ -284,6 +284,11 @@ pci_ers_result_t pcie_do_recovery(struct pci_dev *dev,
> >  	return status;
> >  
> >  failed:
> > +	if (host->native_aer || pcie_ports_native) {
> > +		pcie_clear_device_status(dev);
> > +		pci_aer_clear_nonfatal_status(dev);
> 
> [Severity: Medium]
> When pcie_do_recovery() is invoked for a fatal error (pci_channel_io_frozen)
> and recovery fails, will this leave the fatal error bits uncleared?

Generally not.

First of all, the error status bits in the Device Status register are
cleared because in the "pci_channel_io_frozen" case, a reset is performed
unconditionally.  (They are not sticky, i.e. they do not survive a reset.)

Second, the Uncorrectable Error Status register is cleared as well because
the first thing that a driver's ->slot_reset() callback does is generally
a call to pci_restore_state(), which calls pci_aer_clear_status().

Hence explicitly clearing error status bits in the Fatal Error case has
historically been considered unnecessary.  That said, I'm not happy with
how the bits are cleared and consider this patch an interim solution.

In the long run, I would like to clear error bits immediately when a
device is examined in is_error_source() and aer_get_device_error_info(),
and I would like to clear only those bits that have been read.

Right now we may lose error bits that occur in-between reading the register
and later-on clearing it, so this is all somewhat suboptimal.

Thanks,

Lukas

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH 6/7] PCI/AER: Renumber severity constants
  2026-09-27 18:29   ` sashiko-bot
@ 2026-09-27 19:23     ` Lukas Wunner
  0 siblings, 0 replies; 27+ messages in thread
From: Lukas Wunner @ 2026-09-27 19:23 UTC (permalink / raw)
  To: sashiko-reviews, Tony Luck, Steven Rostedt; +Cc: linux-pci

[cc += Tony Luck, Steven Rostedt, start of thread is here:
https://lore.kernel.org/r/244713dea3786099e9ad4eb10ff9f2f430cd7ad3.1790531238.git.lukas@wunner.de
]

On Sun, Sep 27, 2026 at 06:29:45PM +0000, sashiko-bot@kernel.org wrote:
> PCI/AER: Renumber severity constants
> 
> This commit modifies the AER severity constants to align with the Device Status
> Register bit numbers defined in the PCIe specification. It updates the severity
> strings and the ras_event documentation to reflect the new constant values.
> 
> > +++ b/include/linux/aer.h
> > @@ -11,10 +11,10 @@
> >  #include <linux/errno.h>
> >  #include <linux/types.h>
> >  
> > -#define AER_NONFATAL			0
> > -#define AER_FATAL			1
> > -#define AER_CORRECTABLE			2
> > -#define DPC_FATAL			3
> > +/* Must match bit number in Device Status Register (PCIe r7.1, sec 7.5.3.5) */
> > +#define AER_CORRECTABLE			0
> > +#define AER_NONFATAL			1
> > +#define AER_FATAL			2
> 
> [ ... ]
> 
> > +++ b/include/ras/ras_event.h
> > @@ -302,7 +302,7 @@ TRACE_EVENT(non_standard_event,
> >   *			([domain:]bus:device.function).
> >   * u32 status -		Either the correctable or uncorrectable register
> >   *			indicating what error or errors have been seen
> > - * u8 severity -	error severity 0:NONFATAL 1:FATAL 2:CORRECTED
> > + * u8 severity -	error severity 0:CORRECTED 1:NONFATAL 2:FATAL
> >   */
> 
> [Severity: Critical]
> Does this change unintentionally break the userspace ABI for the aer_event
> tracepoint?
> 
> The tracepoint exports the severity field as a raw u8 integer in the binary
> trace ring buffer. By renumbering the underlying AER constants in
> include/linux/aer.h, the tracepoint now emits different integers for hardware
> error events without an ABI translation layer.
> 
> Userspace tools like rasdaemon that rely on the historical raw integer values
> will now misclassify hardware error severities. For example, a correctable
> error (now 0) would be incorrectly interpreted as non-fatal by older tools
> expecting the previous ABI layout.

Hm, my understanding was that the TP_PROTO() arguments to a TRACE_EVENT()
do not constitute ABI, but rather (only) the string emitted by
TP_printk().

My patch changes the value of a TP_PROTO() argument but not the
TP_printk() output.

However, looking at rasdaemon source code, it's defining constants with
values identical to the kernel:

https://github.com/mchehab/rasdaemon/blob/master/core/ras-events.h#L248

So it looks like this is uAPI without being marked as such and we're stuck
with it. And so apparently patch [6/7] is not applicable and patch [7/7]
needs a fixup. :(

Thanks,

Lukas

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH 4/7] PCI/AER: Drop AER native check from handles_cxl_errors()
  2026-09-27 18:02 ` [PATCH 4/7] PCI/AER: Drop AER native check from handles_cxl_errors() Lukas Wunner
  2026-09-27 18:26   ` sashiko-bot
@ 2026-09-28 20:58   ` Bowman, Terry
  2026-09-29 19:24   ` Kuppuswamy Sathyanarayanan
  2 siblings, 0 replies; 27+ messages in thread
From: Bowman, Terry @ 2026-09-28 20:58 UTC (permalink / raw)
  To: Lukas Wunner, Bjorn Helgaas, Raag Jadav, Riana Tauro,
	Yury Murashka, Matthew W Carlis, linux-pci
  Cc: Mahesh J Salgaonkar, Oliver OHalloran, linuxppc-dev,
	Aravind Iddamsetty, Srinivasa Adatrao, Sathyanarayanan Kuppuswamy,
	Keith Busch

On 9/27/2026 1:02 PM, Lukas Wunner wrote:
> Drop a superfluous call to pcie_aer_is_native() from handles_cxl_errors().
> 
> The call is superfluous because the function is only called from:
> 
>   aer_probe()
>     cxl_rch_enable_rcec()
>       handles_cxl_errors()
> 
> ...and aer_probe() is only invoked if an AER port service was instantiated
> in get_port_device_capability(), which is conditional on:
> 
>   (pci_pcie_type(dev) == PCI_EXP_TYPE_ROOT_PORT ||
>    pci_pcie_type(dev) == PCI_EXP_TYPE_RC_EC) &&
>   pci_aer_available() &&
>   dev->aer_cap && (pcie_ports_native || host->native_aer))
> 
> ...and the last line of those conditions is equivalent to
> pcie_aer_is_native().
> 
> Signed-off-by: Lukas Wunner <lukas@wunner.de>
> ---
>  drivers/pci/pcie/aer_cxl_rch.c | 3 +--
>  1 file changed, 1 insertion(+), 2 deletions(-)
> 
> diff --git a/drivers/pci/pcie/aer_cxl_rch.c b/drivers/pci/pcie/aer_cxl_rch.c
> index e471eefec9c4..9cc2d5749b9d 100644
> --- a/drivers/pci/pcie/aer_cxl_rch.c
> +++ b/drivers/pci/pcie/aer_cxl_rch.c
> @@ -87,8 +87,7 @@ static bool handles_cxl_errors(struct pci_dev *rcec)
>  {
>  	bool handles_cxl = false;
>  
> -	if (pci_pcie_type(rcec) == PCI_EXP_TYPE_RC_EC &&
> -	    pcie_aer_is_native(rcec))
> +	if (pci_pcie_type(rcec) == PCI_EXP_TYPE_RC_EC)
>  		pcie_walk_rcec(rcec, handles_cxl_error_iter, &handles_cxl);
>  
>  	return handles_cxl;


Nice cleanup. The pcie_aer_is_native() check isn't necessary. You can add my RB.

Reviewed-by: Terry Bowman <terry.bowman@amd.com>

-Terry 


^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH 1/7] PCI/DPC: Avoid access to non-existent AER capability
  2026-09-27 18:02 ` [PATCH 1/7] PCI/DPC: Avoid access to non-existent AER capability Lukas Wunner
  2026-09-27 18:27   ` sashiko-bot
@ 2026-09-29 18:49   ` Kuppuswamy Sathyanarayanan
  1 sibling, 0 replies; 27+ messages in thread
From: Kuppuswamy Sathyanarayanan @ 2026-09-29 18:49 UTC (permalink / raw)
  To: Lukas Wunner, Bjorn Helgaas, Raag Jadav, Riana Tauro,
	Yury Murashka, Matthew W Carlis, linux-pci
  Cc: Mahesh J Salgaonkar, Oliver OHalloran, linuxppc-dev,
	Aravind Iddamsetty, Srinivasa Adatrao, Terry Bowman, Keith Busch



On 9/27/2026 11:02 AM, Lukas Wunner wrote:
> Downstream Port Containment does not mandate presence of an Advanced Error
> Reporting capability, so a Downstream Port may support DPC, but not AER
> (PCIe r7.1 sec 6.2.11.2).
> 
> In February 2019, commit 9f08a5d896ce ("PCI/DPC: Fix print AER status in
> DPC event handling") amended the DPC driver to access the AER capability
> without checking for its presence.
> 
> In May 2020, commit 708b20003624 ("PCI/AER: Remove HEST/FIRMWARE_FIRST
> parsing for AER ownership") fixed it by inserting a call to
> pcie_aer_is_native() in dpc_probe(), which implicitly checks for presence
> of an AER capability.
> 
> However already in October 2019, commit 35a0b2378c19 ("PCI/DPC: Add
> "pcie_ports=dpc-native" to allow DPC without AER control") made it
> possible to override the check:  The DPC driver may access a non-existent
> AER capability if "pcie_ports=dpc-native" is passed on the command line.
> 
> Fix it by making the DPC driver cope with AER-unsupporting Downstream
> Ports.
> 
> There are two places where the AER capability is accessed:
> 
> - dpc_get_aer_uncorrect_severity() uses it to discern whether a Fatal or
>   Non-Fatal Error triggered DPC.  Access the Device Status Register
>   instead, in accordance with PCIe r7.1 sec 6.2.5.

This changes behavior not only for AER-incapable ports, but also for
the AER-capable ones.  PCIe r7.1 sec 7.5.3.5 says for the Fatal/
Non-Fatal Error Detected bits:

  "For Functions supporting Advanced Error Handling, errors are logged
   in this register regardless of the settings of the Uncorrectable
   Error Mask register."

The old code only considered unmasked errors, the new code also picks
up masked ones.  So a masked Fatal error (e.g. Surprise Down) alongside
the unmasked Non-Fatal error which triggered DPC is now reported as
Fatal.  Stale bits from earlier masked errors can have the same effect.

Since this is tagged for stable, how about keeping the AER-based logic
when dev->aer_cap is present and using DEVSTA only as a fallback?

> 
> - dpc_is_surprise_removal() uses it to detect whether a Surprise Down
>   Error triggered DPC.  Return false on AER-unsupporting devices.  The
>   function works around an AMD-specific quirk and it seems reasonable to
>   assume that all affected products are AER-supporting.  In any case the
>   detection is not possible without AER capability.
> 
> Insert a temporary check for an AER capability after the call to
> aer_get_device_error_info() because the function currently returns false
> for AER-unsupporting devices.  The check will become obsolete and will be
> removed with the imminent baseline capability error reporting.
> 
> Fixes: 35a0b2378c19 ("PCI/DPC: Add "pcie_ports=dpc-native" to allow DPC without AER control")
> Signed-off-by: Lukas Wunner <lukas@wunner.de>
> Cc: stable@vger.kernel.org # v5.5+
> ---
>  drivers/pci/pcie/dpc.c | 23 ++++++++++-------------
>  1 file changed, 10 insertions(+), 13 deletions(-)
> 
> diff --git a/drivers/pci/pcie/dpc.c b/drivers/pci/pcie/dpc.c
> index 2b779bd1d861..793a799053f1 100644
> --- a/drivers/pci/pcie/dpc.c
> +++ b/drivers/pci/pcie/dpc.c
> @@ -236,21 +236,15 @@ static void dpc_process_rp_pio_error(struct pci_dev *pdev)
>  static int dpc_get_aer_uncorrect_severity(struct pci_dev *dev,
>  					  struct aer_err_info *info)
>  {
> -	int pos = dev->aer_cap;
> -	u32 status, mask, sev;
> +	u16 devsta;
>  
> -	pci_read_config_dword(dev, pos + PCI_ERR_UNCOR_STATUS, &status);
> -	pci_read_config_dword(dev, pos + PCI_ERR_UNCOR_MASK, &mask);
> -	status &= ~mask;
> -	if (!status)
> -		return 0;
> -
> -	pci_read_config_dword(dev, pos + PCI_ERR_UNCOR_SEVER, &sev);
> -	status &= sev;
> -	if (status)
> +	pcie_capability_read_word(dev, PCI_EXP_DEVSTA, &devsta);
> +	if (devsta & PCI_EXP_DEVSTA_FED)
>  		info->severity = AER_FATAL;
> -	else
> +	else if (devsta & PCI_EXP_DEVSTA_NFED)
>  		info->severity = AER_NONFATAL;
> +	else
> +		return 0;
>  
>  	info->level = KERN_ERR;
>  
> @@ -275,7 +269,7 @@ void dpc_process_error(struct pci_dev *pdev)
>  		pci_warn(pdev, "containment event, status:%#06x: unmasked uncorrectable error detected\n",
>  			 status);
>  		if (dpc_get_aer_uncorrect_severity(pdev, &info) &&
> -		    aer_get_device_error_info(&info, 0)) {
> +		    (aer_get_device_error_info(&info, 0) || !pdev->aer_cap)) {
>  			aer_print_error(&info, 0);
>  			pci_aer_clear_nonfatal_status(pdev);
>  			pci_aer_clear_fatal_status(pdev);
> @@ -353,6 +347,9 @@ static bool dpc_is_surprise_removal(struct pci_dev *pdev)
>  	if (!pdev->is_hotplug_bridge)
>  		return false;
>  
> +	if (!pdev->aer_cap)
> +		return false;
> +
>  	if (pci_read_config_word(pdev, pdev->aer_cap + PCI_ERR_UNCOR_STATUS,
>  				 &status))
>  		return false;

-- 
Sathyanarayanan Kuppuswamy
Linux Kernel Developer


^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH 2/7] PCI/DPC: Reinstate support for AER-incapable ports
  2026-09-27 18:02 ` [PATCH 2/7] PCI/DPC: Reinstate support for AER-incapable ports Lukas Wunner
  2026-09-27 18:26   ` sashiko-bot
@ 2026-09-29 19:04   ` Kuppuswamy Sathyanarayanan
  1 sibling, 0 replies; 27+ messages in thread
From: Kuppuswamy Sathyanarayanan @ 2026-09-29 19:04 UTC (permalink / raw)
  To: Lukas Wunner, Bjorn Helgaas, Raag Jadav, Riana Tauro,
	Yury Murashka, Matthew W Carlis, linux-pci
  Cc: Mahesh J Salgaonkar, Oliver OHalloran, linuxppc-dev,
	Aravind Iddamsetty, Srinivasa Adatrao, Terry Bowman, Keith Busch

Hi,

On 9/27/2026 11:02 AM, Lukas Wunner wrote:
> Downstream Port Containment does not mandate presence of an Advanced Error
> Reporting capability, so a Downstream Port may support DPC, but not AER
> (PCIe r7.1 sec 6.2.11.2).
> 
> Commit 708b20003624 ("PCI/AER: Remove HEST/FIRMWARE_FIRST parsing for AER
> ownership") seemingly inadvertently constrained DPC to AER-supporting
> Downstream Ports by inserting a call to pcie_aer_is_native() in
> dpc_probe(), which implicitly checks for presence of an AER capability.
> 
> Drop the call.  It is superfluous because it duplicates the conditions
> applied by the PCIe port service driver when it decides whether to
> instantiate a DPC port service (see get_port_device_capability()).
> If the conditions are not met, no port service is instantiated and
> dpc_probe() isn't executed.
> 
> The only difference between the conditions in dpc_probe() and the ones in
> get_port_device_capability() is the check for the AER capability, so by
> removing the checks from dpc_probe(), DPC is supported on AER-incapable
> Downstream Ports again.
> 
> Fixes: 708b20003624 ("PCI/AER: Remove HEST/FIRMWARE_FIRST parsing for AER ownership")

I don't think 708b20003624 changed behavior here.  At the time,
get_port_device_capability() only instantiated the DPC service if
pcie_ports_dpc_native was set or the AER service was instantiated,
and the latter required dev->aer_cap and native AER control:

	if (pci_find_ext_capability(dev, PCI_EXT_CAP_ID_DPC) &&
	    pci_aer_available() &&
	    (pcie_ports_dpc_native || (services & PCIE_PORT_SERVICE_AER)))
		services |= PCIE_PORT_SERVICE_DPC;

So without dpc-native, the pcie_aer_is_native() check in dpc_probe()
could never fail, and with dpc-native it was bypassed.

The check only became effective with 97ca178c899d ("PCI/DPC: Allow DPC
on all Downstream Ports when OS controls AER"), which replaced
"services & PCIE_PORT_SERVICE_AER" with "host->native_aer".  Since then
a DPC service is instantiated on AER-incapable ports, but dpc_probe()
rejects it.

So I think this should rather be:

  Fixes: 97ca178c899d ("PCI/DPC: Allow DPC on all Downstream Ports when OS controls AER")

Also, for stable kernels, is there any way to note that this patch needs
to be picked with patch 1. Otherwise DPC driver might access AER config
without aer_cap.

Otherwise it looks good.

Reviewed-by: Kuppuswamy Sathyanarayanan <sathyanarayanan.kuppuswamy@linux.intel.com>


> Signed-off-by: Lukas Wunner <lukas@wunner.de>
> Cc: stable@vger.kernel.org # v5.8+
> ---
>  drivers/pci/pcie/dpc.c | 3 ---
>  1 file changed, 3 deletions(-)
> 
> diff --git a/drivers/pci/pcie/dpc.c b/drivers/pci/pcie/dpc.c
> index 793a799053f1..104ff2b91f1f 100644
> --- a/drivers/pci/pcie/dpc.c
> +++ b/drivers/pci/pcie/dpc.c
> @@ -474,9 +474,6 @@ static int dpc_probe(struct pcie_device *dev)
>  	int status;
>  	u16 cap;
>  
> -	if (!pcie_aer_is_native(pdev) && !pcie_ports_dpc_native)
> -		return -ENOTSUPP;
> -
>  	status = devm_request_threaded_irq(device, dev->irq, dpc_irq,
>  					   dpc_handler, IRQF_SHARED,
>  					   "pcie-dpc", pdev);

-- 
Sathyanarayanan Kuppuswamy
Linux Kernel Developer


^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH 3/7] PCI/ERR: Avoid stale error status bits on recovery failure
  2026-09-27 18:02 ` [PATCH 3/7] PCI/ERR: Avoid stale error status bits on recovery failure Lukas Wunner
  2026-09-27 18:31   ` sashiko-bot
@ 2026-09-29 19:20   ` Kuppuswamy Sathyanarayanan
  1 sibling, 0 replies; 27+ messages in thread
From: Kuppuswamy Sathyanarayanan @ 2026-09-29 19:20 UTC (permalink / raw)
  To: Lukas Wunner, Bjorn Helgaas, Raag Jadav, Riana Tauro,
	Yury Murashka, Matthew W Carlis, linux-pci
  Cc: Mahesh J Salgaonkar, Oliver OHalloran, linuxppc-dev,
	Aravind Iddamsetty, Srinivasa Adatrao, Terry Bowman, Keith Busch

Hi,

On 9/27/2026 11:02 AM, Lukas Wunner wrote:
> From: Yury Murashka <yurypm@arista.com>
> 
> pcie_do_recovery() only clears the error status bits of an error-reporting
> device if recovery succeeds, but not if it fails.  (E.g. because reset
> failed, drivers are missing pci_error_handlers or those handlers failed.)
> 
> One undesirable consequence is that the AER driver may subsequently
> identify a wrong device when searching for error-reporting devices in
> is_error_source(), due to stale bits on an unaffected device.
> 
> Avoid by clearing error status bits on recovery failure.
> 
> Link: https://lore.kernel.org/r/CAPzpGcRCTCZtaX1EVaJNZ103THZKsoszZduY7=gwfYdcrMo-SQ@mail.gmail.com/
> Signed-off-by: Yury Murashka <yurypm@arista.com>
> [lukas: drop cmdline param, rewrite commit msg, tag for stable]
> Signed-off-by: Lukas Wunner <lukas@wunner.de>
> Cc: stable@vger.kernel.org
> ---
>  drivers/pci/pcie/err.c | 5 +++++
>  1 file changed, 5 insertions(+)
> 
> diff --git a/drivers/pci/pcie/err.c b/drivers/pci/pcie/err.c
> index d77403d8855b..cfcfc7e4f10d 100644
> --- a/drivers/pci/pcie/err.c
> +++ b/drivers/pci/pcie/err.c
> @@ -284,6 +284,11 @@ pci_ers_result_t pcie_do_recovery(struct pci_dev *dev,
>  	return status;
>  
>  failed:
> +	if (host->native_aer || pcie_ports_native) {

If Bjorn takes AER cleanup patch before this series, we can avoid sprinkling
pcie_ports_native checks everywhere.

https://lore.kernel.org/linux-pci/20260922204548.3884906-1-sathyanarayanan.kuppuswamy@linux.intel.com/


> +		pcie_clear_device_status(dev);
> +		pci_aer_clear_nonfatal_status(dev);

This only clears Non-Fatal status bits.  On the success path, Fatal
status bits are cleared by pci_restore_state() -> pci_aer_clear_status(),
which drivers call from their ->slot_reset() callback.  On the failure
path, that may not happen (e.g. if the reset failed or the driver lacks
pci_error_handlers).

> +	}
> +
>  	pci_walk_bridge(bridge, pci_pm_runtime_put, NULL);
>  
>  	pci_walk_bridge(bridge, report_perm_failure_detected, NULL);

-- 
Sathyanarayanan Kuppuswamy
Linux Kernel Developer


^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH 4/7] PCI/AER: Drop AER native check from handles_cxl_errors()
  2026-09-27 18:02 ` [PATCH 4/7] PCI/AER: Drop AER native check from handles_cxl_errors() Lukas Wunner
  2026-09-27 18:26   ` sashiko-bot
  2026-09-28 20:58   ` Bowman, Terry
@ 2026-09-29 19:24   ` Kuppuswamy Sathyanarayanan
  2 siblings, 0 replies; 27+ messages in thread
From: Kuppuswamy Sathyanarayanan @ 2026-09-29 19:24 UTC (permalink / raw)
  To: Lukas Wunner, Bjorn Helgaas, Raag Jadav, Riana Tauro,
	Yury Murashka, Matthew W Carlis, linux-pci
  Cc: Mahesh J Salgaonkar, Oliver OHalloran, linuxppc-dev,
	Aravind Iddamsetty, Srinivasa Adatrao, Terry Bowman, Keith Busch



On 9/27/2026 11:02 AM, Lukas Wunner wrote:
> Drop a superfluous call to pcie_aer_is_native() from handles_cxl_errors().
> 
> The call is superfluous because the function is only called from:
> 
>   aer_probe()
>     cxl_rch_enable_rcec()
>       handles_cxl_errors()
> 
> ...and aer_probe() is only invoked if an AER port service was instantiated
> in get_port_device_capability(), which is conditional on:
> 
>   (pci_pcie_type(dev) == PCI_EXP_TYPE_ROOT_PORT ||
>    pci_pcie_type(dev) == PCI_EXP_TYPE_RC_EC) &&
>   pci_aer_available() &&
>   dev->aer_cap && (pcie_ports_native || host->native_aer))
> 
> ...and the last line of those conditions is equivalent to
> pcie_aer_is_native().
> 
> Signed-off-by: Lukas Wunner <lukas@wunner.de>
> ---

Looks good to me.

Reviewed-by: Kuppuswamy Sathyanarayanan <sathyanarayanan.kuppuswamy@linux.intel.com>


>  drivers/pci/pcie/aer_cxl_rch.c | 3 +--
>  1 file changed, 1 insertion(+), 2 deletions(-)
> 
> diff --git a/drivers/pci/pcie/aer_cxl_rch.c b/drivers/pci/pcie/aer_cxl_rch.c
> index e471eefec9c4..9cc2d5749b9d 100644
> --- a/drivers/pci/pcie/aer_cxl_rch.c
> +++ b/drivers/pci/pcie/aer_cxl_rch.c
> @@ -87,8 +87,7 @@ static bool handles_cxl_errors(struct pci_dev *rcec)
>  {
>  	bool handles_cxl = false;
>  
> -	if (pci_pcie_type(rcec) == PCI_EXP_TYPE_RC_EC &&
> -	    pcie_aer_is_native(rcec))
> +	if (pci_pcie_type(rcec) == PCI_EXP_TYPE_RC_EC)
>  		pcie_walk_rcec(rcec, handles_cxl_error_iter, &handles_cxl);
>  
>  	return handles_cxl;

-- 
Sathyanarayanan Kuppuswamy
Linux Kernel Developer


^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH 5/7] PCI/AER: Move AER capability check out of pcie_aer_is_native()
  2026-09-27 18:02 ` [PATCH 5/7] PCI/AER: Move AER capability check out of pcie_aer_is_native() Lukas Wunner
  2026-09-27 18:30   ` sashiko-bot
@ 2026-09-29 19:47   ` Kuppuswamy Sathyanarayanan
  1 sibling, 0 replies; 27+ messages in thread
From: Kuppuswamy Sathyanarayanan @ 2026-09-29 19:47 UTC (permalink / raw)
  To: Lukas Wunner, Bjorn Helgaas, Raag Jadav, Riana Tauro,
	Yury Murashka, Matthew W Carlis, linux-pci
  Cc: Mahesh J Salgaonkar, Oliver OHalloran, linuxppc-dev,
	Aravind Iddamsetty, Srinivasa Adatrao, Terry Bowman, Keith Busch



On 9/27/2026 11:02 AM, Lukas Wunner wrote:
> When firmware grants control of Advanced Error Reporting to the operating
> system, that not only encompasses the AER capability, but also error
> enable/status bits in the Device Control and Device Status registers
> (PCI Firmware r3.3 table 4-6 bit 3).
> 
> PCIe devices without AER capability still support baseline capability
> error reporting through these enable/status bits (PCIe r7.1 sec 6.2.1),
> but the bits must not be modified unless AER control was granted.
> 
> pcie_aer_is_native() is unsuitable to check for control of AER-incapable
> devices because it implicitly checks for presence of an AER capability.
> 
> Move that check to its callers (where needed) to allow using the function
> for the imminent baseline capability error reporting.

Agreed that ownership and AER presence are separate questions, but
changing the semantics while keeping the name may trip up callers that
assume "native" implies "present".  Would a separate ownership-only
helper (e.g. pcie_err_is_native()) be cleaner?  It could also replace
cxl_error_is_native().

I think you also need to fix kernel-doc of pci_aer_unmask_internal_errors().
it says to check AER support with pcie_aer_is_native().  That's no longer
sufficient, and the function has no aer_cap check, Please update the comment
and ideally add an "if (!aer) return;".


> 
> Signed-off-by: Lukas Wunner <lukas@wunner.de>
> ---
>  drivers/pci/pcie/aer.c | 7 ++-----
>  1 file changed, 2 insertions(+), 5 deletions(-)
> 
> diff --git a/drivers/pci/pcie/aer.c b/drivers/pci/pcie/aer.c
> index d8dcd238fda1..34a8eddc427a 100644
> --- a/drivers/pci/pcie/aer.c
> +++ b/drivers/pci/pcie/aer.c
> @@ -257,9 +257,6 @@ int pcie_aer_is_native(struct pci_dev *dev)
>  {
>  	struct pci_host_bridge *host = pci_find_host_bridge(dev->bus);
>  
> -	if (!dev->aer_cap)
> -		return 0;
> -
>  	return pcie_ports_native || host->native_aer;
>  }
>  EXPORT_SYMBOL_NS_GPL(pcie_aer_is_native, "CXL");
> @@ -280,7 +277,7 @@ int pci_aer_clear_nonfatal_status(struct pci_dev *dev)
>  	int aer = dev->aer_cap;
>  	u32 status, sev;
>  
> -	if (!pcie_aer_is_native(dev))
> +	if (!aer || !pcie_aer_is_native(dev))
>  		return -EIO;
>  
>  	/* Clear status bits for ERR_NONFATAL errors only */
> @@ -299,7 +296,7 @@ void pci_aer_clear_fatal_status(struct pci_dev *dev)
>  	int aer = dev->aer_cap;
>  	u32 status, sev;
>  
> -	if (!pcie_aer_is_native(dev))
> +	if (!aer || !pcie_aer_is_native(dev))
>  		return;
>  
>  	/* Clear status bits for ERR_FATAL errors only */

-- 
Sathyanarayanan Kuppuswamy
Linux Kernel Developer


^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH 6/7] PCI/AER: Renumber severity constants
  2026-09-27 18:02 ` [PATCH 6/7] PCI/AER: Renumber severity constants Lukas Wunner
  2026-09-27 18:29   ` sashiko-bot
@ 2026-09-29 19:55   ` Kuppuswamy Sathyanarayanan
  1 sibling, 0 replies; 27+ messages in thread
From: Kuppuswamy Sathyanarayanan @ 2026-09-29 19:55 UTC (permalink / raw)
  To: Lukas Wunner, Bjorn Helgaas, Raag Jadav, Riana Tauro,
	Yury Murashka, Matthew W Carlis, linux-pci
  Cc: Mahesh J Salgaonkar, Oliver OHalloran, linuxppc-dev,
	Aravind Iddamsetty, Srinivasa Adatrao, Terry Bowman, Keith Busch

Hi Lukas,

On 9/27/2026 11:02 AM, Lukas Wunner wrote:
> The AER driver uses constants for Correctable, Non-Fatal and Fatal Error
> severity which look as if they match something in the spec, but are
> actually just made-up numbers.
> 
> Use the bit number in the Device Status Register instead (PCIe r7.1 sec
> 7.5.3.5).  A subsequent commit takes advantage of this by using BIT() to
> conveniently compute the register bit corresponding to a given severity.
> 
> While at it, drop the DPC_FATAL constant which was introduced by commit
> b09803b5e546 ("PCI/DPC: Use the generic pcie_do_fatal_recovery() path")
> but never saw any use in 8 years.
> 

This changes the raw severity value in the aer_event tracepoint, which
may break userspace.  rasdaemon relies on the current numbering
(enum hw_event_aer_err_type in core/ras-events.h), so it would
misclassify error.


> Signed-off-by: Lukas Wunner <lukas@wunner.de>
> ---
>  drivers/pci/pcie/aer.c  | 2 +-
>  include/linux/aer.h     | 8 ++++----
>  include/ras/ras_event.h | 2 +-
>  3 files changed, 6 insertions(+), 6 deletions(-)
> 
> diff --git a/drivers/pci/pcie/aer.c b/drivers/pci/pcie/aer.c
> index 34a8eddc427a..6bc843ab9b37 100644
> --- a/drivers/pci/pcie/aer.c
> +++ b/drivers/pci/pcie/aer.c
> @@ -504,9 +504,9 @@ void pci_aer_exit(struct pci_dev *dev)
>   * AER error strings
>   */
>  static const char * const aer_error_severity_string[] = {
> +	"Correctable",
>  	"Uncorrectable (Non-Fatal)",
>  	"Uncorrectable (Fatal)",
> -	"Correctable"
>  };
>  
>  static const char *aer_error_layer[] = {
> diff --git a/include/linux/aer.h b/include/linux/aer.h
> index df0f5c382286..795c55132008 100644
> --- a/include/linux/aer.h
> +++ b/include/linux/aer.h
> @@ -11,10 +11,10 @@
>  #include <linux/errno.h>
>  #include <linux/types.h>
>  
> -#define AER_NONFATAL			0
> -#define AER_FATAL			1
> -#define AER_CORRECTABLE			2
> -#define DPC_FATAL			3
> +/* Must match bit number in Device Status Register (PCIe r7.1, sec 7.5.3.5) */
> +#define AER_CORRECTABLE			0
> +#define AER_NONFATAL			1
> +#define AER_FATAL			2
>  
>  /*
>   * AER and DPC capabilities TLP Logging register sizes (PCIe r6.2, sec 7.8.4
> diff --git a/include/ras/ras_event.h b/include/ras/ras_event.h
> index fdb785fa4613..1e03ade99386 100644
> --- a/include/ras/ras_event.h
> +++ b/include/ras/ras_event.h
> @@ -302,7 +302,7 @@ TRACE_EVENT(non_standard_event,
>   *			([domain:]bus:device.function).
>   * u32 status -		Either the correctable or uncorrectable register
>   *			indicating what error or errors have been seen
> - * u8 severity -	error severity 0:NONFATAL 1:FATAL 2:CORRECTED
> + * u8 severity -	error severity 0:CORRECTED 1:NONFATAL 2:FATAL
>   */
>  
>  #define aer_correctable_errors					\

-- 
Sathyanarayanan Kuppuswamy
Linux Kernel Developer


^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH 7/7] PCI/AER: Enable baseline capability error reporting
  2026-09-27 18:02 ` [PATCH 7/7] PCI/AER: Enable baseline capability error reporting Lukas Wunner
  2026-09-27 18:33   ` sashiko-bot
@ 2026-09-29 21:01   ` Kuppuswamy Sathyanarayanan
  2026-09-30  7:08     ` Lukas Wunner
  1 sibling, 1 reply; 27+ messages in thread
From: Kuppuswamy Sathyanarayanan @ 2026-09-29 21:01 UTC (permalink / raw)
  To: Lukas Wunner, Bjorn Helgaas, Raag Jadav, Riana Tauro,
	Yury Murashka, Matthew W Carlis, linux-pci
  Cc: Mahesh J Salgaonkar, Oliver OHalloran, linuxppc-dev,
	Aravind Iddamsetty, Srinivasa Adatrao, Terry Bowman, Keith Busch



On 9/27/2026 11:02 AM, Lukas Wunner wrote:
> PCIe r7.1 sec 6.2.1 defines two error reporting paradigms:  Baseline
> capability and Advanced Error Reporting Extended Capability (AER).
> 
> So far the kernel only supports the AER paradigm.  Enable the baseline
> capability paradigm to complete the kernel's support for error reporting.
> 
> Baseline capability (solely) relies on the error enable/status bits in the
> Device Control and Device Status registers, which exist on every PCIe
> device.  One difference between baseline capability and AER is that the
> severity of Uncorrectable Errors cannot be controlled.  Another difference
> is that Advisory Non-Fatal Errors are not reported at all (PCIe r7.1 sec
> 6.2.5).
> 
> But the main difference to AER is that no detailed error information is
> available:  It is only known that an error occurred and which severity it
> had, but not its type or the Prefix / Header of the TLP that caused it.
> An exception are Unsupported Request Errors which are signaled with a
> dedicated bit in the Device Status register.  Report these as if the
> Unsupported Request Error Status bit in the Uncorrectable Error Status
> register was set, for consistency with the AER paradigm.
> 
> Error recovery works the same for baseline capability as it already does
> for AER:  Drivers are informed about errors through the pci_error_handlers
> callbacks.  Recovery from Fatal Errors (and from Non-Fatal Errors, if
> chosen by drivers) is attempted through a Secondary Bus Reset.
> 
> Similarly to commit f26e58bf6f54 ("PCI/AER: Enable error reporting when
> AER is native"), which universally enabled error reporting on AER-capable
> devices, the present commit is invasive because it universally enables
> error reporting on non-AER-capable devices.  Previously those errors were
> neither reported nor recovered.  It may be necessary to amend more drivers
> with pci_error_handlers callbacks to recover from newly reported errors.
> 
> Baseline capability is only enabled if firmware grants AER control to the
> operating system.  Otherwise firmware owns the error enable/status bits in
> the Device Control and Device Status registers (PCI Firmware r3.3 table
> 4-6 bit 3).
> 
> Baseline capability support is initially only implemented for native error
> handling, not Firmware First error handling:  ghes_handle_aer() would have
> to be amended to pass the Device Control and Device Status registers to
> aer_recover_queue(), and to only pass an AER capability structure if it is
> present in the CPER record.  It's not clear whether any firmware actually
> generates such CPER records and whether implementing support for it is
> worthwhile, so postpone that for now.
> 
> Suggested-by: Bjorn Helgaas <bhelgaas@google.com>
> Link: https://lore.kernel.org/r/20260826212619.GA1566339@bhelgaas/
> Signed-off-by: Lukas Wunner <lukas@wunner.de>
> ---
>  .../ABI/testing/sysfs-bus-pci-devices-aer     | 26 ++++++++----
>  Documentation/PCI/pcieaer-howto.rst           |  9 ++--
>  drivers/pci/pcie/aer.c                        | 42 ++++++++++++-------
>  drivers/pci/pcie/dpc.c                        |  2 +-
>  4 files changed, 50 insertions(+), 29 deletions(-)
> 
> diff --git a/Documentation/ABI/testing/sysfs-bus-pci-devices-aer b/Documentation/ABI/testing/sysfs-bus-pci-devices-aer
> index 5ed284523956..9e772e1ed9b3 100644
> --- a/Documentation/ABI/testing/sysfs-bus-pci-devices-aer
> +++ b/Documentation/ABI/testing/sysfs-bus-pci-devices-aer
> @@ -1,9 +1,9 @@
>  PCIe Device AER statistics
>  --------------------------
>  
> -These attributes show up under all the devices that are AER capable. These
> +These attributes show up under all PCIe devices (AER capable or not). These
>  statistical counters indicate the errors "as seen/reported by the device".
> -Note that this may mean that if an endpoint is causing problems, the AER
> +Note that this may mean that if an endpoint is causing problems, the error
>  counters may increment at its link partner (e.g. root port) because the
>  errors may be "seen" / reported by the link partner and not the
>  problematic endpoint itself (which may report all counters as 0 as it never
> @@ -17,7 +17,10 @@ Description:	List of correctable errors seen and reported by this
>  		PCI device using ERR_COR. Note that since multiple errors may
>  		be reported using a single ERR_COR message, thus
>  		TOTAL_ERR_COR at the end of the file may not match the actual
> -		total of all the errors in the file. Sample output::
> +		total of all the errors in the file.
> +		For PCIe devices without AER Extended Capability, only the
> +		total counter is meaningful while all other counters remain 0.
> +		Sample output::
>  
>  		    localhost /sys/devices/pci0000:00/0000:00:1c.0 # cat aer_dev_correctable
>  		    Receiver Error 2
> @@ -38,7 +41,10 @@ Description:	List of uncorrectable fatal errors seen and reported by this
>  		PCI device using ERR_FATAL. Note that since multiple errors may
>  		be reported using a single ERR_FATAL message, thus
>  		TOTAL_ERR_FATAL at the end of the file may not match the actual
> -		total of all the errors in the file. Sample output::
> +		total of all the errors in the file.
> +		For PCIe devices without AER Extended Capability, only the
> +		total counter is meaningful while all other counters remain 0.
> +		Sample output::
>  
>  		    localhost /sys/devices/pci0000:00/0000:00:1c.0 # cat aer_dev_fatal
>  		    Undefined 0
> @@ -68,7 +74,11 @@ Description:	List of uncorrectable nonfatal errors seen and reported by this
>  		PCI device using ERR_NONFATAL. Note that since multiple errors
>  		may be reported using a single ERR_FATAL message, thus
>  		TOTAL_ERR_NONFATAL at the end of the file may not match the
> -		actual total of all the errors in the file. Sample output::
> +		actual total of all the errors in the file.
> +		For PCIe devices without AER Extended Capability, only the
> +		total counter and the Unsupported Request counter is meaningful
> +		while all other counters remain 0.
> +		Sample output::
>  
>  		    localhost /sys/devices/pci0000:00/0000:00:1c.0 # cat aer_dev_nonfatal
>  		    Undefined 0
> @@ -121,7 +131,7 @@ Description:	Total number of ERR_NONFATAL messages reported to rootport.
>  PCIe AER ratelimits
>  -------------------
>  
> -These attributes show up under all the devices that are AER capable.
> +These attributes show up under all PCIe devices (AER capable or not).
>  They represent configurable ratelimits of logs per error type.
>  
>  See Documentation/PCI/pcieaer-howto.rst for more info on ratelimits.
> @@ -130,7 +140,7 @@ What:		/sys/bus/pci/devices/<dev>/aer/correctable_ratelimit_interval_ms
>  Date:		May 2025
>  KernelVersion:	6.16.0
>  Contact:	linux-pci@vger.kernel.org
> -Description:	Writing 0 disables AER correctable error log ratelimiting.
> +Description:	Writing 0 disables Correctable Error log ratelimiting.
>  		Writing a positive value sets the ratelimit interval in ms.
>  		Default is DEFAULT_RATELIMIT_INTERVAL (5000 ms).
>  
> @@ -147,7 +157,7 @@ What:		/sys/bus/pci/devices/<dev>/aer/nonfatal_ratelimit_interval_ms
>  Date:		May 2025
>  KernelVersion:	6.16.0
>  Contact:	linux-pci@vger.kernel.org
> -Description:	Writing 0 disables AER non-fatal uncorrectable error log
> +Description:	Writing 0 disables Non-Fatal Uncorrectable Error log
>  		ratelimiting. Writing a positive value sets the ratelimit
>  		interval in ms. Default is DEFAULT_RATELIMIT_INTERVAL
>  		(5000 ms).
> diff --git a/Documentation/PCI/pcieaer-howto.rst b/Documentation/PCI/pcieaer-howto.rst
> index 90fdfddd3ae5..e969e563b61d 100644
> --- a/Documentation/PCI/pcieaer-howto.rst
> +++ b/Documentation/PCI/pcieaer-howto.rst
> @@ -34,9 +34,8 @@ set of error reporting requirements. Advanced Error Reporting
>  capability is implemented with a PCIe Advanced Error Reporting
>  extended capability structure providing more robust error reporting.
>  
> -The PCIe AER driver provides the infrastructure to support PCIe Advanced
> -Error Reporting capability. The PCIe AER driver provides three basic
> -functions:
> +The PCIe AER driver provides the infrastructure to support both paradigms.
> +It provides three basic functions:
>  
>    - Gathers the comprehensive error information if errors occurred.
>    - Reports error to the users.
> @@ -69,7 +68,7 @@ Specification for details regarding _OSC usage.
>  AER error output
>  ----------------
>  
> -When a PCIe AER error is captured, an error message will be output to
> +When a PCIe error is captured, an error message will be output to
>  console. If it's a correctable error, it is output as a warning message.
>  Otherwise, it is printed as an error. So users could choose different
>  log level to filter out correctable error messages.
> @@ -113,7 +112,7 @@ See Documentation/ABI/testing/sysfs-bus-pci-devices-aer.
>  AER Statistics / Counters
>  -------------------------
>  
> -When PCIe AER errors are captured, the counters / statistics are also exposed
> +When PCIe errors are captured, the counters / statistics are also exposed
>  in the form of sysfs attributes which are documented at
>  Documentation/ABI/testing/sysfs-bus-pci-devices-aer.
>  
> diff --git a/drivers/pci/pcie/aer.c b/drivers/pci/pcie/aer.c
> index 6bc843ab9b37..62376aab4b6f 100644
> --- a/drivers/pci/pcie/aer.c
> +++ b/drivers/pci/pcie/aer.c
> @@ -397,21 +397,22 @@ void pci_aer_init(struct pci_dev *dev)
>  {
>  	int n;
>  
> -	dev->aer_cap = pci_find_ext_capability(dev, PCI_EXT_CAP_ID_ERR);
> -	if (!dev->aer_cap)
> +	if (!pci_is_pcie(dev))
>  		return;
>  
>  	dev->aer_info = kzalloc_obj(*dev->aer_info);
> -	if (!dev->aer_info) {
> -		dev->aer_cap = 0;
> +	if (!dev->aer_info)
>  		return;
> -	}
>  
>  	ratelimit_state_init(&dev->aer_info->correctable_ratelimit,
>  			     DEFAULT_RATELIMIT_INTERVAL, DEFAULT_RATELIMIT_BURST);
>  	ratelimit_state_init(&dev->aer_info->nonfatal_ratelimit,
>  			     DEFAULT_RATELIMIT_INTERVAL, DEFAULT_RATELIMIT_BURST);
>  
> +	dev->aer_cap = pci_find_ext_capability(dev, PCI_EXT_CAP_ID_ERR);
> +	if (!dev->aer_cap)
> +		goto enable;
> +
>  	/*
>  	 * We save/restore PCI_ERR_UNCOR_MASK, PCI_ERR_UNCOR_SEVER,
>  	 * PCI_ERR_COR_MASK, and PCI_ERR_CAP.  Root and Root Complex Event
> @@ -431,6 +432,9 @@ void pci_aer_init(struct pci_dev *dev)
>  					       PCI_ERR_COR_ADV_NFAT, 0);
>  
>  	pci_aer_clear_status(dev);
> +enable:
> +	if (pcie_aer_is_native(dev))
> +		pcie_clear_device_status(dev);
>  
>  	if (pci_aer_available())
>  		pci_enable_pcie_error_reporting(dev);

This also enables error reporting below AER-incapable Root Ports, where
no AER service handles the ERR_* Messages.  The Root Control System
Error enable bits are only cleared by aer_enable_rootport(), which
doesn't run on such ports.  If firmware left them set, the newly
enabled Messages could result in System Errors.

Should reporting be enabled only if an AER service (or DPC) is above
the device?  Or alternatively, clear the Root Control System Error
enable bits on AER-incapable Root Ports?

> @@ -980,15 +984,12 @@ void aer_print_error(struct aer_err_info *info, int i)
>  	if (!info->ratelimit_print[i])
>  		goto anfe;
>  
> -	if (!info->status) {
> -		pci_err(dev, "%s Bus Error: severity=%s (Inaccessible)\n",
> -			bus_type, aer_error_severity_string[info->severity]);
> -		return;
> -	}
> -
>  	aer_printk(level, dev, "%s Bus Error: severity=%s\n",
>  		   bus_type, aer_error_severity_string[info->severity]);
>  
> +	if (!info->status)
> +		return;

For AER-capable devices, info->status == 0 mostly means ERR_FATAL on an
Endpoint or Upstream Port whose registers we deliberately didn't read.
"(Inaccessible)" explained why no details follow; maybe keep it for
dev->aer_cap, or print "(no details available)" in both cases?

> +
>  	aer_printk(level, dev, "  device [%04x:%04x] error status/mask=%08x/%08x\n",
>  		   dev->vendor, dev->device, info->status, info->mask);
>  
> @@ -1182,8 +1183,11 @@ static bool is_error_source(struct pci_dev *dev, struct aer_err_info *e_info)
>  	if (!(reg16 & PCI_EXP_AER_FLAGS))
>  		return false;
>  
> -	if (!aer)
> -		return false;
> +	if (!aer) {
> +		pcie_capability_read_word(dev, PCI_EXP_DEVSTA, &reg16);
> +		return reg16 & BIT(e_info->severity) &&
> +		       !PCI_POSSIBLE_ERROR(reg16);
> +	}
>  
>  	/* Check if error is recorded */
>  	if (e_info->severity == AER_CORRECTABLE) {
> @@ -1455,6 +1459,7 @@ int aer_get_device_error_info(struct aer_err_info *info, int i)
>  {
>  	struct pci_dev *dev;
>  	int type, aer;
> +	u16 devsta;
>  
>  	if (i >= AER_MAX_MULTI_ERR_DEVICES)
>  		return 0;
> @@ -1470,8 +1475,15 @@ int aer_get_device_error_info(struct aer_err_info *info, int i)
>  	info->is_cxl = pcie_is_cxl(dev);
>  
>  	/* The device might not support AER */
> -	if (!aer)
> -		return 0;
> +	if (!aer) {
> +		if (info->severity != AER_CORRECTABLE) {
> +			pcie_capability_read_word(dev, PCI_EXP_DEVSTA, &devsta);
> +			if (devsta & PCI_EXP_DEVSTA_URD &&
> +			    !PCI_POSSIBLE_ERROR(devsta))
> +				info->status = PCI_ERR_UNC_UNSUP;
> +		}
> +		return 1;

I think info->mask needs to be reset here to 0.

> +	}
>  
>  	if (info->severity == AER_CORRECTABLE) {
>  		pci_read_config_dword(dev, aer + PCI_ERR_COR_STATUS,
> diff --git a/drivers/pci/pcie/dpc.c b/drivers/pci/pcie/dpc.c
> index 104ff2b91f1f..fdea3db61a4a 100644
> --- a/drivers/pci/pcie/dpc.c
> +++ b/drivers/pci/pcie/dpc.c
> @@ -269,7 +269,7 @@ void dpc_process_error(struct pci_dev *pdev)
>  		pci_warn(pdev, "containment event, status:%#06x: unmasked uncorrectable error detected\n",
>  			 status);
>  		if (dpc_get_aer_uncorrect_severity(pdev, &info) &&
> -		    (aer_get_device_error_info(&info, 0) || !pdev->aer_cap)) {
> +		    aer_get_device_error_info(&info, 0)) {
>  			aer_print_error(&info, 0);
>  			pci_aer_clear_nonfatal_status(pdev);
>  			pci_aer_clear_fatal_status(pdev);

-- 
Sathyanarayanan Kuppuswamy
Linux Kernel Developer


^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH 7/7] PCI/AER: Enable baseline capability error reporting
  2026-09-29 21:01   ` Kuppuswamy Sathyanarayanan
@ 2026-09-30  7:08     ` Lukas Wunner
  2026-10-01 17:44       ` Kuppuswamy Sathyanarayanan
  0 siblings, 1 reply; 27+ messages in thread
From: Lukas Wunner @ 2026-09-30  7:08 UTC (permalink / raw)
  To: Kuppuswamy Sathyanarayanan
  Cc: Bjorn Helgaas, Raag Jadav, Riana Tauro, Yury Murashka,
	Matthew W Carlis, linux-pci, Mahesh J Salgaonkar,
	Oliver OHalloran, linuxppc-dev, Aravind Iddamsetty,
	Srinivasa Adatrao, Terry Bowman, Keith Busch

On Tue, Sep 29, 2026 at 02:01:42PM -0700, Kuppuswamy Sathyanarayanan wrote:
> On 9/27/2026 11:02 AM, Lukas Wunner wrote:
> > @@ -431,6 +432,9 @@ void pci_aer_init(struct pci_dev *dev)
> >  					       PCI_ERR_COR_ADV_NFAT, 0);
> >  
> >  	pci_aer_clear_status(dev);
> > +enable:
> > +	if (pcie_aer_is_native(dev))
> > +		pcie_clear_device_status(dev);
> >  
> >  	if (pci_aer_available())
> >  		pci_enable_pcie_error_reporting(dev);
> 
> This also enables error reporting below AER-incapable Root Ports, where
> no AER service handles the ERR_* Messages.  The Root Control System
> Error enable bits are only cleared by aer_enable_rootport(), which
> doesn't run on such ports.  If firmware left them set, the newly
> enabled Messages could result in System Errors.
> 
> Should reporting be enabled only if an AER service (or DPC) is above
> the device?

Excellent observation.  This is a pre-existing issue but I think you're
right.  However it's non-trivial to fix because just checking for DPC
capability in the ancestry or AER capability at the Root Port isn't
sufficient:

For RCiEPs, we'd need to check whether an RCEC exists which has
AER capability.  There's an "rcec" pointer in struct pci_dev which
allows discovering the RCEC responsible for an RCiEP.  But the pointer
is only set when portdrv binds to the RCEC (pcie_link_rcec()).
That's much later than when the RCiEP and its capabilities are enumerated.

When enumerating an RCiEP, we'd need to walk the entire set of PCI devices,
check if it's an RCEC, check if it's responsible for this RCiEP and assign
the rcec pointer.  We could try to avoid that by running pcie_link_rcec()
already on enumeration of the RCEC (and not on probing of portdrv),
but the RCEC may be enumerated after the RCiEP.  User space could also
force an unset rcec pointer by issuing remove/rescan of the RCiEP.

Also, right now when firmware does keep System Error Enable bits in the
Root Control register set, there's a window between endpoints being
enumerated (which enables sending of ERR_* messages) and Root Ports
being bound to portdrv (which clears System Error Enable bits).

Any errors that occur during that window will cause a System Error
right now.

> Or alternatively, clear the Root Control System Error
> enable bits on AER-incapable Root Ports?

I'm worried that users may deliberately enable System Error bits in
BIOS on such systems precisely because there's no other way to catch
them.

Thanks for the thoughtful review, much appreciated!

Lukas

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH 7/7] PCI/AER: Enable baseline capability error reporting
  2026-09-30  7:08     ` Lukas Wunner
@ 2026-10-01 17:44       ` Kuppuswamy Sathyanarayanan
  0 siblings, 0 replies; 27+ messages in thread
From: Kuppuswamy Sathyanarayanan @ 2026-10-01 17:44 UTC (permalink / raw)
  To: Lukas Wunner
  Cc: Bjorn Helgaas, Raag Jadav, Riana Tauro, Yury Murashka,
	Matthew W Carlis, linux-pci, Mahesh J Salgaonkar,
	Oliver OHalloran, linuxppc-dev, Aravind Iddamsetty,
	Srinivasa Adatrao, Terry Bowman, Keith Busch

Hi Lukas,

On 9/30/2026 12:08 AM, Lukas Wunner wrote:
> On Tue, Sep 29, 2026 at 02:01:42PM -0700, Kuppuswamy Sathyanarayanan wrote:
>> On 9/27/2026 11:02 AM, Lukas Wunner wrote:
>>> @@ -431,6 +432,9 @@ void pci_aer_init(struct pci_dev *dev)
>>>  					       PCI_ERR_COR_ADV_NFAT, 0);
>>>  
>>>  	pci_aer_clear_status(dev);
>>> +enable:
>>> +	if (pcie_aer_is_native(dev))
>>> +		pcie_clear_device_status(dev);
>>>  
>>>  	if (pci_aer_available())
>>>  		pci_enable_pcie_error_reporting(dev);
>>
>> This also enables error reporting below AER-incapable Root Ports, where
>> no AER service handles the ERR_* Messages.  The Root Control System
>> Error enable bits are only cleared by aer_enable_rootport(), which
>> doesn't run on such ports.  If firmware left them set, the newly
>> enabled Messages could result in System Errors.
>>
>> Should reporting be enabled only if an AER service (or DPC) is above
>> the device?
> 
> Excellent observation.  This is a pre-existing issue but I think you're
> right.  However it's non-trivial to fix because just checking for DPC
> capability in the ancestry or AER capability at the Root Port isn't
> sufficient:
> 
> For RCiEPs, we'd need to check whether an RCEC exists which has
> AER capability.  There's an "rcec" pointer in struct pci_dev which
> allows discovering the RCEC responsible for an RCiEP.  But the pointer
> is only set when portdrv binds to the RCEC (pcie_link_rcec()).
> That's much later than when the RCiEP and its capabilities are enumerated.
> 
> When enumerating an RCiEP, we'd need to walk the entire set of PCI devices,
> check if it's an RCEC, check if it's responsible for this RCiEP and assign
> the rcec pointer.  We could try to avoid that by running pcie_link_rcec()
> already on enumeration of the RCEC (and not on probing of portdrv),
> but the RCEC may be enumerated after the RCiEP.  User space could also
> force an unset rcec pointer by issuing remove/rescan of the RCiEP.
> 
> Also, right now when firmware does keep System Error Enable bits in the
> Root Control register set, there's a window between endpoints being
> enumerated (which enables sending of ERR_* messages) and Root Ports
> being bound to portdrv (which clears System Error Enable bits).
> 

Agreed, the RCiEP case makes this hard, and the window already exists.

Could you add a sentence to the commit message noting that reporting
is now also enabled on AER-incapable devices below AER-incapable Root
Ports?  That way, if someone bisects a new System Error to this
commit, the reason is obvious.


> Any errors that occur during that window will cause a System Error
> right now.
> 
>> Or alternatively, clear the Root Control System Error
>> enable bits on AER-incapable Root Ports?
> 
> I'm worried that users may deliberately enable System Error bits in
> BIOS on such systems precisely because there's no other way to catch
> them.
> 

Fair point, let's leave them alone.


> Thanks for the thoughtful review, much appreciated!
> 
> Lukas

-- 
Sathyanarayanan Kuppuswamy
Linux Kernel Developer


^ permalink raw reply	[flat|nested] 27+ messages in thread

end of thread, other threads:[~2026-10-01 17:44 UTC | newest]

Thread overview: 27+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-27 18:02 [PATCH 0/7] Error reporting for AER-incapable devices Lukas Wunner
2026-09-27 18:02 ` [PATCH 1/7] PCI/DPC: Avoid access to non-existent AER capability Lukas Wunner
2026-09-27 18:27   ` sashiko-bot
2026-09-29 18:49   ` Kuppuswamy Sathyanarayanan
2026-09-27 18:02 ` [PATCH 2/7] PCI/DPC: Reinstate support for AER-incapable ports Lukas Wunner
2026-09-27 18:26   ` sashiko-bot
2026-09-29 19:04   ` Kuppuswamy Sathyanarayanan
2026-09-27 18:02 ` [PATCH 3/7] PCI/ERR: Avoid stale error status bits on recovery failure Lukas Wunner
2026-09-27 18:31   ` sashiko-bot
2026-09-27 18:57     ` Lukas Wunner
2026-09-29 19:20   ` Kuppuswamy Sathyanarayanan
2026-09-27 18:02 ` [PATCH 4/7] PCI/AER: Drop AER native check from handles_cxl_errors() Lukas Wunner
2026-09-27 18:26   ` sashiko-bot
2026-09-28 20:58   ` Bowman, Terry
2026-09-29 19:24   ` Kuppuswamy Sathyanarayanan
2026-09-27 18:02 ` [PATCH 5/7] PCI/AER: Move AER capability check out of pcie_aer_is_native() Lukas Wunner
2026-09-27 18:30   ` sashiko-bot
2026-09-29 19:47   ` Kuppuswamy Sathyanarayanan
2026-09-27 18:02 ` [PATCH 6/7] PCI/AER: Renumber severity constants Lukas Wunner
2026-09-27 18:29   ` sashiko-bot
2026-09-27 19:23     ` Lukas Wunner
2026-09-29 19:55   ` Kuppuswamy Sathyanarayanan
2026-09-27 18:02 ` [PATCH 7/7] PCI/AER: Enable baseline capability error reporting Lukas Wunner
2026-09-27 18:33   ` sashiko-bot
2026-09-29 21:01   ` Kuppuswamy Sathyanarayanan
2026-09-30  7:08     ` Lukas Wunner
2026-10-01 17:44       ` Kuppuswamy Sathyanarayanan

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox