* [PATCH 1/4] PCI: Reject all-ones responses in pci_dev_wait()
2026-09-29 2:04 [PATCH 0/4] PCI/CXL: Guard reset recovery against inaccessible devices Richard Cheng
@ 2026-09-29 2:04 ` Richard Cheng
2026-09-29 2:04 ` [PATCH 2/4] PCI: Return -ETIMEOUT when reset readiness polling expires Richard Cheng
` (2 subsequent siblings)
3 siblings, 0 replies; 6+ messages in thread
From: Richard Cheng @ 2026-09-29 2:04 UTC (permalink / raw)
To: jic23, dave, dave.jiang, alison.schofield, vishal.l.verma, iweiny
Cc: ming.li, kaihengf, kobak, newtonl, kristinc, mochs, linux-cxl,
linux-kernel, Richard Cheng
With RRS software visibility enabled, pci_dev_wait() treats an all-ones
Vendor ID response as reset completion, even though the device is
inaccessible.
Reject all-ones responses and continue polling. Use the existing
Command-register check for VFs, whose Vendor ID can legitimately read as
all ones.
Signed-off-by: Richard Cheng <icheng@nvidia.com>
---
drivers/pci/pci.c | 10 +++++-----
1 file changed, 5 insertions(+), 5 deletions(-)
diff --git a/drivers/pci/pci.c b/drivers/pci/pci.c
index b2879a6be5f8..6461274bdbcf 100644
--- a/drivers/pci/pci.c
+++ b/drivers/pci/pci.c
@@ -1248,10 +1248,10 @@ static int pci_dev_wait(struct pci_dev *dev, char *reset_type, int timeout)
* If the device is below a Root Port with Configuration RRS
* Software Visibility enabled, reading the Vendor ID returns a
* special data value if the device responded with RRS. Read the
- * Vendor ID until we get non-RRS status.
+ * Vendor ID until we get neither RRS nor an error response.
*
- * If there's no Root Port or Configuration RRS Software Visibility
- * is not enabled, the device may still respond with RRS, but
+ * For VFs, or if there's no Root Port or Configuration RRS Software
+ * Visibility is not enabled, the device may still respond with RRS, but
* hardware may retry the config request. If no retries receive
* Successful Completion, hardware generally synthesizes ~0
* (PCI_ERROR_RESPONSE) data to complete the read. Reading Vendor
@@ -1266,9 +1266,9 @@ static int pci_dev_wait(struct pci_dev *dev, char *reset_type, int timeout)
return -ENOTTY;
}
- if (root && root->config_rrs_sv) {
+ if (root && root->config_rrs_sv && !dev->is_virtfn) {
pci_read_config_dword(dev, PCI_VENDOR_ID, &id);
- if (!pci_bus_rrs_vendor_id(id))
+ if (!PCI_POSSIBLE_ERROR(id) && !pci_bus_rrs_vendor_id(id))
break;
} else {
pci_read_config_dword(dev, PCI_COMMAND, &id);
--
2.43.0
^ permalink raw reply related [flat|nested] 6+ messages in thread* [PATCH 2/4] PCI: Return -ETIMEOUT when reset readiness polling expires
2026-09-29 2:04 [PATCH 0/4] PCI/CXL: Guard reset recovery against inaccessible devices Richard Cheng
2026-09-29 2:04 ` [PATCH 1/4] PCI: Reject all-ones responses in pci_dev_wait() Richard Cheng
@ 2026-09-29 2:04 ` Richard Cheng
2026-09-29 2:19 ` sashiko-bot
2026-09-29 2:04 ` [PATCH 3/4] PCI: Skip state restoration when configuration is inaccessible Richard Cheng
2026-09-29 2:04 ` [PATCH 4/4] cxl/pci: Guard decoder accesses after reset Richard Cheng
3 siblings, 1 reply; 6+ messages in thread
From: Richard Cheng @ 2026-09-29 2:04 UTC (permalink / raw)
To: jic23, dave, dave.jiang, alison.schofield, vishal.l.verma, iweiny
Cc: ming.li, kaihengf, kobak, newtonl, kristinc, mochs, linux-cxl,
linux-kernel, Richard Cheng
pci_dev_wait() return -ENOTTY when the device fails to become ready. The
reset framekwork interprets this as an unsupported reset method and may
try another method.
Return -ETIMEOUT to report the timeout and stop reset-method fallback.
Signed-off-by: Richard Cheng <icheng@nvidia.com>
---
drivers/pci/pci.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/drivers/pci/pci.c b/drivers/pci/pci.c
index 6461274bdbcf..d59a374a350d 100644
--- a/drivers/pci/pci.c
+++ b/drivers/pci/pci.c
@@ -1279,7 +1279,7 @@ static int pci_dev_wait(struct pci_dev *dev, char *reset_type, int timeout)
if (delay > timeout) {
pci_err(dev, "not ready %dms after %s; giving up\n",
delay - 1, reset_type);
- return -ENOTTY;
+ return -ETIMEDOUT;
}
if (delay > PCI_RESET_WAIT) {
--
2.43.0
^ permalink raw reply related [flat|nested] 6+ messages in thread* Re: [PATCH 2/4] PCI: Return -ETIMEOUT when reset readiness polling expires
2026-09-29 2:04 ` [PATCH 2/4] PCI: Return -ETIMEOUT when reset readiness polling expires Richard Cheng
@ 2026-09-29 2:19 ` sashiko-bot
0 siblings, 0 replies; 6+ messages in thread
From: sashiko-bot @ 2026-09-29 2:19 UTC (permalink / raw)
To: Richard Cheng; +Cc: linux-cxl
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] pci: breaking reset fallback leaves devices permanently wedged on timeout
--
commit 4617b1cca51c668aae8c9b4809552b80cdac13dd
Author: Richard Cheng <icheng@nvidia.com>
PCI: Return -ETIMEOUT when reset readiness polling expires
This patch modifies pci_dev_wait() to return -ETIMEDOUT instead of
-ENOTTY when a device fails to become ready. The commit message notes
this stops the reset-method fallback.
> diff --git a/drivers/pci/pci.c b/drivers/pci/pci.c
> index 6461274bdbcf9..d59a374a350d8 100644
> --- a/drivers/pci/pci.c
> +++ b/drivers/pci/pci.c
> @@ -1279,7 +1279,7 @@ static int pci_dev_wait(struct pci_dev *dev, char *reset_type, int timeout)
> if (delay > timeout) {
> pci_err(dev, "not ready %dms after %s; giving up\n",
> delay - 1, reset_type);
> - return -ENOTTY;
> + return -ETIMEDOUT;
> }
[Severity: High]
Will this regression leave devices permanently wedged if they hang during a
soft reset like a function level reset?
Looking at __pci_reset_function_locked() in drivers/pci/pci.c:
for (i = 0; i < ARRAY_SIZE(pci_reset_fn_methods); i++) {
...
rc = method->reset_fn(dev, false);
...
if (rc != -ENOTTY)
return rc;
}
The fallback loop aborts on any error other than -ENOTTY.
Previously, if a device hung during a function level reset and
pci_dev_wait() timed out, it would return -ENOTTY. This allowed the
fallback mechanism to try a stronger reset method like a secondary bus
reset to recover the device.
By stopping this fallback, could userspace applications and virtual machine
guests using vfio permanently wedge an assigned device that has a buggy
reset implementation, instead of falling back to a secondary bus reset?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260929020405.9724-1-icheng@nvidia.com?part=2
^ permalink raw reply [flat|nested] 6+ messages in thread
* [PATCH 3/4] PCI: Skip state restoration when configuration is inaccessible
2026-09-29 2:04 [PATCH 0/4] PCI/CXL: Guard reset recovery against inaccessible devices Richard Cheng
2026-09-29 2:04 ` [PATCH 1/4] PCI: Reject all-ones responses in pci_dev_wait() Richard Cheng
2026-09-29 2:04 ` [PATCH 2/4] PCI: Return -ETIMEOUT when reset readiness polling expires Richard Cheng
@ 2026-09-29 2:04 ` Richard Cheng
2026-09-29 2:04 ` [PATCH 4/4] cxl/pci: Guard decoder accesses after reset Richard Cheng
3 siblings, 0 replies; 6+ messages in thread
From: Richard Cheng @ 2026-09-29 2:04 UTC (permalink / raw)
To: jic23, dave, dave.jiang, alison.schofield, vishal.l.verma, iweiny
Cc: ming.li, kaihengf, kobak, newtonl, kristinc, mochs, linux-cxl,
linux-kernel, Richard Cheng
pci_dev_restore() restores state even when the device is inaccessible.
MSI-X restoration can then issue MMIO accesses that trigger DPC.
Check configuration accessibility before restoring state. Keep driver
reset callbacks and reset return values unchanged.
Signed-off-by: Richard Cheng <icheng@nvidia.com>
---
drivers/pci/pci.c | 4 +++-
1 file changed, 3 insertions(+), 1 deletion(-)
diff --git a/drivers/pci/pci.c b/drivers/pci/pci.c
index d59a374a350d..9a343f8b5429 100644
--- a/drivers/pci/pci.c
+++ b/drivers/pci/pci.c
@@ -5127,7 +5127,9 @@ static void pci_dev_restore(struct pci_dev *dev)
const struct pci_error_handlers *err_handler =
dev->driver ? dev->driver->err_handler : NULL;
- pci_restore_state(dev);
+ /* MSI-X restoration may access device MMIO. */
+ if (pci_dev_config_accessible(dev, "restore state"))
+ pci_restore_state(dev);
/*
* dev->driver->err_handler->reset_done() is protected against
--
2.43.0
^ permalink raw reply related [flat|nested] 6+ messages in thread
* [PATCH 4/4] cxl/pci: Guard decoder accesses after reset
2026-09-29 2:04 [PATCH 0/4] PCI/CXL: Guard reset recovery against inaccessible devices Richard Cheng
` (2 preceding siblings ...)
2026-09-29 2:04 ` [PATCH 3/4] PCI: Skip state restoration when configuration is inaccessible Richard Cheng
@ 2026-09-29 2:04 ` Richard Cheng
3 siblings, 0 replies; 6+ messages in thread
From: Richard Cheng @ 2026-09-29 2:04 UTC (permalink / raw)
To: jic23, dave, dave.jiang, alison.schofield, vishal.l.verma, iweiny
Cc: ming.li, kaihengf, kobak, newtonl, kristinc, mochs, linux-cxl,
linux-kernel, Richard Cheng
cxl_reset_done() runs even when PCI state restoration is skipped.
Reading HDM decoder registers while the device is inaccessible can
trigger another error.
Skip decoder checks if the PCI Command read fails, returns all ones, or
Memory Space Enable is clear.
Signed-off-by: Richard Cheng <icheng@nvidia.com>
---
drivers/cxl/pci.c | 6 ++++++
1 file changed, 6 insertions(+)
diff --git a/drivers/cxl/pci.c b/drivers/cxl/pci.c
index c7c91e8dc51d..2294f7761e50 100644
--- a/drivers/cxl/pci.c
+++ b/drivers/cxl/pci.c
@@ -976,6 +976,12 @@ static void cxl_reset_done(struct pci_dev *pdev)
struct cxl_dev_state *cxlds = pci_get_drvdata(pdev);
struct cxl_memdev *cxlmd = cxlds->cxlmd;
struct device *dev = &pdev->dev;
+ u32 command;
+
+ /* PCI state restoration may have been skipped after a failed reset. */
+ if (pci_read_config_dword(pdev, PCI_COMMAND, &command) ||
+ PCI_POSSIBLE_ERROR(command) || !(command & PCI_COMMAND_MEMORY))
+ return;
/*
* FLR does not expect to touch the HDM decoders and related
--
2.43.0
^ permalink raw reply related [flat|nested] 6+ messages in thread