Linux USB
 help / color / mirror / Atom feed
From: Henry Tseng <henrytseng@qnap.com>
To: Mathias Nyman <mathias.nyman@intel.com>
Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>,
	linux-usb@vger.kernel.org, Henry Tseng <henrytseng@qnap.com>
Subject: [PATCH 1/2] xhci: make xhci_handshake() timeout wall-clock based again
Date: Wed, 30 Sep 2026 18:17:51 +0800	[thread overview]
Message-ID: <20260930101752.15794-2-henrytseng@qnap.com> (raw)
In-Reply-To: <20260930101752.15794-1-henrytseng@qnap.com>

xhci_abort_cmd_ring() waits for the command ring to stop with
xhci_handshake() and a 5 s timeout. On an AMD Raven USB 3.1 xHCI
(1022:15e0), after a configure endpoint command failed to complete
during device disconnect, that handshake returned -ETIMEDOUT 15.8 s
later:

  [   86.086412] xhci_hcd 0000:0c:00.3: Command timeout, USBSTS: 0x00000010 PCD
  [   86.086450] xhci_hcd 0000:0c:00.3: Abort command ring
  [  101.924383] xhci_hcd 0000:0c:00.3: Abort failed to stop command ring: -110

xhci_abort_cmd_ring() is called from xhci_handle_command_timeout() with
xhci->lock held, so interrupts were disabled on that CPU for the whole
15.8 s.

The timeout budget is no longer tied to elapsed time. xhci_handshake()
originally polled with readl(), udelay(1) and a decrementing
microsecond counter. Commit f7fac17ca925 ("xhci: Convert
xhci_handshake() to use readl_poll_timeout_atomic()") replaced that loop
with readl_poll_timeout_atomic(), which at the time computed its
deadline with ktime_get(). Its commit message notes that this also
fixed a bug on AMD Stoneyridge, where udelay(1) sometimes took over
10 ms and a 5 s timeout ran for over 15 s, triggering the watchdog.
Commit 7349a69cf312 ("iopoll: Do not use timekeeping in
read_poll_timeout_atomic()") then replaced that deadline with a locally
estimated one, which does not account for the readl(), however long it
takes.

To fix this, restore the readl() and udelay() loop, with the
decrementing microsecond counter replaced by a ktime deadline.

Signed-off-by: Henry Tseng <henrytseng@qnap.com>
---
The overrun was measured on v7.3-rc5 by calling xhci_handshake() on
USBSTS bit 5, which is reserved and zeroed, so the poll always runs
to the timeout:

  ret = xhci_handshake(&xhci->op_regs->status, BIT(5), BIT(5),
                       XHCI_RESET_LONG_USEC);

XHCI_RESET_LONG_USEC is 10000000 us. Elapsed time was taken around the call
with ktime_get(). All runs returned -ETIMEDOUT.

  Intel Core 9 273PE, Alder Lake-S PCH USB 3.2 (8086:7ae0)
    unpatched 19121395 us (1.91x)   patched 10000001 us
  Intel Core Ultra 7 155H, Meteor Lake-P USB 3.2 (8086:7e7d)
    unpatched 17161328 us (1.71x)   patched 10000001 us

 drivers/usb/host/xhci.c | 20 ++++++++++++--------
 1 file changed, 12 insertions(+), 8 deletions(-)

diff --git a/drivers/usb/host/xhci.c b/drivers/usb/host/xhci.c
index a9e47e178c28..8169d30dd40e 100644
--- a/drivers/usb/host/xhci.c
+++ b/drivers/usb/host/xhci.c
@@ -86,16 +86,20 @@ static bool td_on_ring(struct xhci_td *td, struct xhci_ring *ring)
 int xhci_handshake(void __iomem *ptr, u32 mask, u32 done, u64 timeout_us)
 {
 	u32	result;
-	int	ret;
+	ktime_t deadline = ktime_add_us(ktime_get(), timeout_us);
 
-	ret = readl_poll_timeout_atomic(ptr, result,
-					(result & mask) == done ||
-					result == U32_MAX,
-					1, timeout_us);
-	if (result == U32_MAX)		/* card removed */
-		return -ENODEV;
+	for (;;) {
+		result = readl(ptr);
+		if (result == U32_MAX)			/* card removed */
+			return -ENODEV;
+		if ((result & mask) == done)
+			return 0;
 
-	return ret;
+		if (ktime_compare(ktime_get(), deadline) > 0)
+			return -ETIMEDOUT;
+
+		udelay(1);
+	}
 }
 
 /*
-- 
2.43.0


  reply	other threads:[~2026-09-30 10:18 UTC|newest]

Thread overview: 8+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-30 10:17 [PATCH 0/2] xhci: handshake timeout overrun and configure endpoint hang on device disconnect Henry Tseng
2026-09-30 10:17 ` Henry Tseng [this message]
2026-09-30 10:17 ` [PATCH 2/2] xhci: skip configure endpoint when dropping endpoints of a disconnected device Henry Tseng
2026-10-02  9:30 ` [PATCH 0/2] xhci: handshake timeout overrun and configure endpoint hang on device disconnect Michal Pecio
2026-10-07  9:52   ` Henry Tseng
2026-10-08  9:06     ` Michal Pecio
2026-10-08  9:13       ` Michal Pecio
2026-10-08 10:25       ` Henry Tseng

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260930101752.15794-2-henrytseng@qnap.com \
    --to=henrytseng@qnap.com \
    --cc=gregkh@linuxfoundation.org \
    --cc=linux-usb@vger.kernel.org \
    --cc=mathias.nyman@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox