From: sashiko-bot@kernel.org
To: "Srirangan Madhavan" <smadhavan@nvidia.com>
Cc: linux-pci@vger.kernel.org, linux-cxl@vger.kernel.org
Subject: Re: [PATCH v13 10/15] cxl: Add CXL Device Reset sequencing
Date: Tue, 22 Sep 2026 08:49:37 +0000 [thread overview]
Message-ID: <20260922084937.8C6001F000FF@smtp.kernel.org> (raw)
In-Reply-To: <20260922083924.2451158-11-smadhavan@nvidia.com>
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Medium] cxl: 100s thread stall due to ignoring PCI errors in cxl_reset_wait_done()
- [Medium] cxl: spurious reset timeout for devices with <= 100ms budget
--
commit 0940b93cca76ebfd20726ff13344213cdb258009
Author: Srirangan Madhavan <smadhavan@nvidia.com>
cxl: Add CXL Device Reset sequencing
This commit implements CXL Device Reset for eligible Type 2 devices. It
disables caching, prepares the IOMMU, initiates reset with Memory Clear
enabled, and waits for completion.
> diff --git a/drivers/cxl/core/resource.c b/drivers/cxl/core/resource.c
> index b5f8c28e613fe..57eec77cbb213 100644
> --- a/drivers/cxl/core/resource.c
> +++ b/drivers/cxl/core/resource.c
[ ... ]
> +static int cxl_reset_wait_done(struct pci_dev *pdev, int dvsec, u16 cap)
> +{
> + unsigned long deadline;
> + u32 timeout_ms;
> + u16 status2;
> + int idx, rc;
> +
> + idx = FIELD_GET(PCI_DVSEC_CXL_RST_TIMEOUT, cap);
> + if (idx >= ARRAY_SIZE(cxl_reset_timeout_ms)) {
> + int last = ARRAY_SIZE(cxl_reset_timeout_ms) - 1;
> +
> + pci_warn(pdev,
> + "unknown CXL reset timeout encoding %d; using %u ms\n",
> + idx, cxl_reset_timeout_ms[last]);
> + idx = last;
> + }
> +
> + timeout_ms = max_t(u32, cxl_reset_timeout_ms[idx],
> + CXL_RESET_RRS_WAIT_MS);
> + msleep(CXL_RESET_RRS_WAIT_MS);
> + deadline = jiffies + msecs_to_jiffies(timeout_ms -
> + CXL_RESET_RRS_WAIT_MS);
[Severity: Medium]
Does this logic yield a 0-jiffy budget and immediately timeout if the device
is not ready on the first poll?
Because timeout_ms is clamped to a minimum of 100ms (CXL_RESET_RRS_WAIT_MS),
the calculation of timeout_ms - CXL_RESET_RRS_WAIT_MS results in 0 for
devices advertising 10ms or 100ms timeouts.
After the initial 100ms sleep, any U16_MAX response (which is common while
the PCIe link is recovering) will immediately trigger -ETIMEDOUT without
retrying, providing zero grace period for link recovery.
> +
> + for (;;) {
> + rc = pci_read_config_word(pdev, dvsec + PCI_DVSEC_CXL_STATUS2,
> + &status2);
> + if (rc || status2 == U16_MAX)
> + goto not_ready;
[Severity: Medium]
Could this cause a kernel thread to stall for up to 100 seconds if the device
experiences a permanent error?
By grouping a non-zero rc (which pci_read_config_word can return for
permanent errors like device disconnection, e.g., PCIBIOS_DEVICE_NOT_FOUND)
with the transient U16_MAX state, the code falls through to not_ready.
If a device advertises a 100,000 ms timeout, this causes the kernel thread
to sleep and poll uselessly for up to 100 seconds when the device is entirely
gone, delaying recovery and blocking operations.
> + if (status2 & PCI_DVSEC_CXL_RST_ERR)
> + return -EIO;
> + if (status2 & PCI_DVSEC_CXL_RST_DONE)
> + return 0;
> +
> +not_ready:
> + if (time_after_eq(jiffies, deadline))
> + return -ETIMEDOUT;
> +
> + msleep(CXL_RESET_STATUS_POLL_MS);
> + }
> +}
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260922083924.2451158-1-smadhavan@nvidia.com?part=10
next prev parent reply other threads:[~2026-09-22 8:49 UTC|newest]
Thread overview: 71+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-22 8:39 [PATCH v13 00/15] PCI/CXL: Add CXL reset support for Type 2 devices Srirangan Madhavan
2026-09-22 8:39 ` [PATCH v13 01/15] cxl: Drop stale decoder interleave limit comment Srirangan Madhavan
2026-09-22 8:41 ` sashiko-bot
2026-09-24 1:16 ` Jonathan Cameron
2026-09-24 16:23 ` Dave Jiang
2026-09-22 8:39 ` [PATCH v13 02/15] cxl: Share CXL port upstream PCI device lookup Srirangan Madhavan
2026-09-22 8:47 ` sashiko-bot
2026-09-23 21:39 ` Cheatham, Benjamin
2026-09-24 1:21 ` Jonathan Cameron
2026-09-24 16:55 ` Dave Jiang
2026-10-01 22:33 ` Srirangan Madhavan
2026-09-24 1:22 ` Jonathan Cameron
2026-10-01 22:28 ` Srirangan Madhavan
2026-09-24 17:01 ` Dave Jiang
2026-09-22 8:39 ` [PATCH v13 03/15] cxl: Move HDM decoder programming helpers Srirangan Madhavan
2026-09-22 8:50 ` sashiko-bot
2026-09-24 1:29 ` Jonathan Cameron
2026-10-01 23:40 ` Srirangan Madhavan
2026-09-24 17:02 ` Dave Jiang
2026-09-22 8:39 ` [PATCH v13 04/15] cxl: Move decoder declarations to shared header Srirangan Madhavan
2026-09-22 8:48 ` sashiko-bot
2026-09-24 1:31 ` Jonathan Cameron
2026-09-24 17:03 ` Dave Jiang
2026-09-22 8:39 ` [PATCH v13 05/15] cxl: Introduce reusable HDM decoder settings Srirangan Madhavan
2026-09-22 8:47 ` sashiko-bot
2026-09-23 21:39 ` Cheatham, Benjamin
2026-09-24 1:35 ` Jonathan Cameron
2026-10-01 22:43 ` Srirangan Madhavan
2026-09-24 2:45 ` Jonathan Cameron
2026-10-01 23:42 ` Srirangan Madhavan
2026-09-22 8:39 ` [PATCH v13 06/15] cxl: Make HDM reset helpers available to built-in PCI code Srirangan Madhavan
2026-09-22 8:54 ` sashiko-bot
2026-09-23 21:40 ` Cheatham, Benjamin
2026-09-24 2:49 ` Jonathan Cameron
2026-10-01 22:46 ` Srirangan Madhavan
2026-09-22 8:39 ` [PATCH v13 07/15] cxl: Share HDM decoder register unpacking Srirangan Madhavan
2026-09-22 8:55 ` sashiko-bot
2026-09-24 3:05 ` Jonathan Cameron
2026-10-01 23:52 ` Srirangan Madhavan
2026-09-22 8:39 ` [PATCH v13 08/15] cxl: Refresh cached PCI HDM decoder settings Srirangan Madhavan
2026-09-22 8:47 ` sashiko-bot
2026-09-23 21:40 ` Cheatham, Benjamin
2026-10-01 22:58 ` Srirangan Madhavan
2026-09-24 3:08 ` Jonathan Cameron
2026-09-22 8:39 ` [PATCH v13 09/15] cxl: Cache endpoint HDM state during PCI enumeration Srirangan Madhavan
2026-09-22 8:54 ` sashiko-bot
2026-09-23 21:40 ` Cheatham, Benjamin
2026-10-01 23:14 ` Srirangan Madhavan
2026-09-24 3:36 ` Jonathan Cameron
2026-10-01 23:55 ` Srirangan Madhavan
2026-09-22 8:39 ` [PATCH v13 10/15] cxl: Add CXL Device Reset sequencing Srirangan Madhavan
2026-09-22 8:49 ` sashiko-bot [this message]
2026-09-23 21:40 ` Cheatham, Benjamin
2026-09-24 17:29 ` Dave Jiang
2026-10-01 23:36 ` Srirangan Madhavan
2026-09-22 8:39 ` [PATCH v13 11/15] cxl: Validate and synchronize HDM ranges around reset Srirangan Madhavan
2026-09-22 8:51 ` sashiko-bot
2026-09-23 21:40 ` Cheatham, Benjamin
2026-10-01 23:25 ` Srirangan Madhavan
2026-09-22 8:39 ` [PATCH v13 12/15] PCI/CXL: Reject reset with unsafe function scope Srirangan Madhavan
2026-09-22 8:56 ` sashiko-bot
2026-09-23 21:41 ` Cheatham, Benjamin
2026-09-24 17:33 ` Dave Jiang
2026-09-22 8:39 ` [PATCH v13 13/15] cxl: Restore CXL state after PCI reset Srirangan Madhavan
2026-09-22 8:55 ` sashiko-bot
2026-09-24 3:50 ` Jonathan Cameron
2026-10-01 23:58 ` Srirangan Madhavan
2026-09-22 8:39 ` [PATCH v13 14/15] PCI/CXL: Expose CXL Reset as a PCI reset method Srirangan Madhavan
2026-09-22 9:02 ` sashiko-bot
2026-09-22 8:39 ` [PATCH v13 15/15] PCI/CXL: Restore CXL state after CXL bus reset Srirangan Madhavan
2026-09-22 9:01 ` sashiko-bot
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260922084937.8C6001F000FF@smtp.kernel.org \
--to=sashiko-bot@kernel.org \
--cc=linux-cxl@vger.kernel.org \
--cc=linux-pci@vger.kernel.org \
--cc=sashiko-reviews@lists.linux.dev \
--cc=smadhavan@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox