Linux USB
 help / color / mirror / Atom feed
* [PATCH 0/2] xhci: handshake timeout overrun and configure endpoint hang on device disconnect
@ 2026-09-30 10:17 Henry Tseng
  2026-09-30 10:17 ` [PATCH 1/2] xhci: make xhci_handshake() timeout wall-clock based again Henry Tseng
                   ` (2 more replies)
  0 siblings, 3 replies; 8+ messages in thread
From: Henry Tseng @ 2026-09-30 10:17 UTC (permalink / raw)
  To: Mathias Nyman; +Cc: Greg Kroah-Hartman, linux-usb, Henry Tseng

This series addresses two problems found while debugging an AMD Raven
USB 3.1 xHCI (1022:15e0) that gets declared dead when a USB storage
enclosure is unplugged. During teardown a command fails to complete,
aborting the command ring also fails, and the whole host is declared
dead, taking unrelated devices on other root ports down with it:

  xhci_hcd 0000:0c:00.3: Command timeout, USBSTS: 0x00000010 PCD
  xhci_hcd 0000:0c:00.3: Abort command ring
  xhci_hcd 0000:0c:00.3: Abort failed to stop command ring: -110
  xhci_hcd 0000:0c:00.3: xHCI host controller not responding, assume dead

Reproduced on mainline v7.3-rc5:

  CPU:        AMD Ryzen Embedded V1500B
  xHCI:       AMD Raven USB 3.1, 0000:0c:00.3 [1022:15e0]
  Enclosure:  QNAP TL-D800C

I have only seen this with this enclosure on this controller. Other USB
devices unplug cleanly from the same host, and this enclosure unplugs
cleanly from a Meteor Lake-P xHCI. Similar messages were reported in
[1][2] without a resolution. I could not confirm the same command was
involved.

Patch 2 fixes the host death on disconnect. The command that never
completed is a configure endpoint command issued during teardown of a
device behind the disconnected root port, right before disable slot for
the same slot. Skip it when the roothub port is gone, as
xhci_check_bandwidth() already does when the host is being removed.

Patch 1 is an independent handshake overrun noticed during the same
debugging, and does not fix the disconnect hang on its own. The command
abort handshake has a 5 s timeout but took 15.8 s with interrupts
disabled. This is not specific to this controller. When forced to time
out, a 10 s xhci_handshake() returns after 19.1 s on Alder Lake-S and
17.2 s on Meteor Lake-P. With patch 1 applied it returns after 10 s on
both.

[1] https://bugs.launchpad.net/ubuntu/+source/linux/+bug/1857385
[2] https://marc.info/?l=linux-usb&m=155120372130870&w=2

Henry Tseng (2):
  xhci: make xhci_handshake() timeout wall-clock based again
  xhci: skip configure endpoint when dropping endpoints of a
    disconnected device

 drivers/usb/host/xhci.c | 30 ++++++++++++++++++++++--------
 1 file changed, 22 insertions(+), 8 deletions(-)

-- 
2.43.0


^ permalink raw reply	[flat|nested] 8+ messages in thread

* [PATCH 1/2] xhci: make xhci_handshake() timeout wall-clock based again
  2026-09-30 10:17 [PATCH 0/2] xhci: handshake timeout overrun and configure endpoint hang on device disconnect Henry Tseng
@ 2026-09-30 10:17 ` Henry Tseng
  2026-09-30 10:17 ` [PATCH 2/2] xhci: skip configure endpoint when dropping endpoints of a disconnected device Henry Tseng
  2026-10-02  9:30 ` [PATCH 0/2] xhci: handshake timeout overrun and configure endpoint hang on device disconnect Michal Pecio
  2 siblings, 0 replies; 8+ messages in thread
From: Henry Tseng @ 2026-09-30 10:17 UTC (permalink / raw)
  To: Mathias Nyman; +Cc: Greg Kroah-Hartman, linux-usb, Henry Tseng

xhci_abort_cmd_ring() waits for the command ring to stop with
xhci_handshake() and a 5 s timeout. On an AMD Raven USB 3.1 xHCI
(1022:15e0), after a configure endpoint command failed to complete
during device disconnect, that handshake returned -ETIMEDOUT 15.8 s
later:

  [   86.086412] xhci_hcd 0000:0c:00.3: Command timeout, USBSTS: 0x00000010 PCD
  [   86.086450] xhci_hcd 0000:0c:00.3: Abort command ring
  [  101.924383] xhci_hcd 0000:0c:00.3: Abort failed to stop command ring: -110

xhci_abort_cmd_ring() is called from xhci_handle_command_timeout() with
xhci->lock held, so interrupts were disabled on that CPU for the whole
15.8 s.

The timeout budget is no longer tied to elapsed time. xhci_handshake()
originally polled with readl(), udelay(1) and a decrementing
microsecond counter. Commit f7fac17ca925 ("xhci: Convert
xhci_handshake() to use readl_poll_timeout_atomic()") replaced that loop
with readl_poll_timeout_atomic(), which at the time computed its
deadline with ktime_get(). Its commit message notes that this also
fixed a bug on AMD Stoneyridge, where udelay(1) sometimes took over
10 ms and a 5 s timeout ran for over 15 s, triggering the watchdog.
Commit 7349a69cf312 ("iopoll: Do not use timekeeping in
read_poll_timeout_atomic()") then replaced that deadline with a locally
estimated one, which does not account for the readl(), however long it
takes.

To fix this, restore the readl() and udelay() loop, with the
decrementing microsecond counter replaced by a ktime deadline.

Signed-off-by: Henry Tseng <henrytseng@qnap.com>
---
The overrun was measured on v7.3-rc5 by calling xhci_handshake() on
USBSTS bit 5, which is reserved and zeroed, so the poll always runs
to the timeout:

  ret = xhci_handshake(&xhci->op_regs->status, BIT(5), BIT(5),
                       XHCI_RESET_LONG_USEC);

XHCI_RESET_LONG_USEC is 10000000 us. Elapsed time was taken around the call
with ktime_get(). All runs returned -ETIMEDOUT.

  Intel Core 9 273PE, Alder Lake-S PCH USB 3.2 (8086:7ae0)
    unpatched 19121395 us (1.91x)   patched 10000001 us
  Intel Core Ultra 7 155H, Meteor Lake-P USB 3.2 (8086:7e7d)
    unpatched 17161328 us (1.71x)   patched 10000001 us

 drivers/usb/host/xhci.c | 20 ++++++++++++--------
 1 file changed, 12 insertions(+), 8 deletions(-)

diff --git a/drivers/usb/host/xhci.c b/drivers/usb/host/xhci.c
index a9e47e178c28..8169d30dd40e 100644
--- a/drivers/usb/host/xhci.c
+++ b/drivers/usb/host/xhci.c
@@ -86,16 +86,20 @@ static bool td_on_ring(struct xhci_td *td, struct xhci_ring *ring)
 int xhci_handshake(void __iomem *ptr, u32 mask, u32 done, u64 timeout_us)
 {
 	u32	result;
-	int	ret;
+	ktime_t deadline = ktime_add_us(ktime_get(), timeout_us);
 
-	ret = readl_poll_timeout_atomic(ptr, result,
-					(result & mask) == done ||
-					result == U32_MAX,
-					1, timeout_us);
-	if (result == U32_MAX)		/* card removed */
-		return -ENODEV;
+	for (;;) {
+		result = readl(ptr);
+		if (result == U32_MAX)			/* card removed */
+			return -ENODEV;
+		if ((result & mask) == done)
+			return 0;
 
-	return ret;
+		if (ktime_compare(ktime_get(), deadline) > 0)
+			return -ETIMEDOUT;
+
+		udelay(1);
+	}
 }
 
 /*
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 8+ messages in thread

* [PATCH 2/2] xhci: skip configure endpoint when dropping endpoints of a disconnected device
  2026-09-30 10:17 [PATCH 0/2] xhci: handshake timeout overrun and configure endpoint hang on device disconnect Henry Tseng
  2026-09-30 10:17 ` [PATCH 1/2] xhci: make xhci_handshake() timeout wall-clock based again Henry Tseng
@ 2026-09-30 10:17 ` Henry Tseng
  2026-10-02  9:30 ` [PATCH 0/2] xhci: handshake timeout overrun and configure endpoint hang on device disconnect Michal Pecio
  2 siblings, 0 replies; 8+ messages in thread
From: Henry Tseng @ 2026-09-30 10:17 UTC (permalink / raw)
  To: Mathias Nyman; +Cc: Greg Kroah-Hartman, linux-usb, Henry Tseng

When a USB device is disconnected, its endpoints are dropped with a
configure endpoint command, immediately followed by a disable slot
command for the same slot.

Per xHCI 1.2b section 4.8.3 that configure endpoint transitions the
dropped endpoints to the Disabled state, and the following disable slot
transitions all endpoints of the slot to Disabled. The dropped
endpoints end up Disabled either way, so skipping the configure
endpoint does not change the outcome.

xhci_check_bandwidth() already returns -ENODEV without issuing the
command when the host is dying or being removed, and teardown then
continues to disable slot in xhci_free_dev(). Add the case of the
device's roothub port being disconnected or its link being inactive to
that condition, using the port state tracked since commit 042aad8d0db6
("xhci: prevent endpoint recovery after roothub disconnect"). Device
disconnect then follows the same path as host removal. The caller runs
xhci_reset_bandwidth() on -ENODEV, and the rings of the dropped
endpoints are freed by xhci_free_virt_device() after disable slot.

On an AMD Raven USB 3.1 xHCI (1022:15e0), unplugging a USB device
sometimes leaves the configure endpoint command incomplete. The command
ring abort then fails too, and the host is declared dead, taking
unrelated devices on other root ports with it:

  [   80.928319] usb 2-1: USB disconnect, device number 3
  [   80.939654] xhci_hcd 0000:0c:00.3: drop ep 0x83, slot id 5, new drop flags = 0x80, new add flags = 0x0
  [   86.086412] xhci_hcd 0000:0c:00.3: Command timeout, USBSTS: 0x00000010 PCD
  [  101.924383] xhci_hcd 0000:0c:00.3: Abort failed to stop command ring: -110
  [  101.936213] xhci_hcd 0000:0c:00.3: xHCI host controller not responding, assume dead
  [  101.937662] xhci_hcd 0000:0c:00.3: Timeout while waiting for configure endpoint command

With this change the same slot 5 drop returns immediately. Teardown
continues and the port 2-1 device is unregistered with the host still
alive:

  [  114.356084] usb 2-1: USB disconnect, device number 3
  [  114.356982] xhci_hcd 0000:0c:00.3: drop ep 0x83, slot id 5, new drop flags = 0x80, new add flags = 0x0
  [  114.357675] xhci_hcd 0000:0c:00.3: drop ep 0x81, slot id 6, new drop flags = 0x8, new add flags = 0x0
  ...
  [  114.564127] usb 2-1: unregistering device

Signed-off-by: Henry Tseng <henrytseng@qnap.com>
---
 drivers/usb/host/xhci.c | 10 ++++++++++
 1 file changed, 10 insertions(+)

diff --git a/drivers/usb/host/xhci.c b/drivers/usb/host/xhci.c
index 8169d30dd40e..3d701d0fd3df 100644
--- a/drivers/usb/host/xhci.c
+++ b/drivers/usb/host/xhci.c
@@ -3088,6 +3088,7 @@ int xhci_check_bandwidth(struct usb_hcd *hcd, struct usb_device *udev)
 	struct xhci_input_control_ctx *ctrl_ctx;
 	struct xhci_slot_ctx *slot_ctx;
 	struct xhci_command *command;
+	struct xhci_port *rhub_port;
 
 	ret = xhci_check_args(hcd, udev, NULL, 0, true, __func__);
 	if (ret <= 0)
@@ -3099,6 +3100,15 @@ int xhci_check_bandwidth(struct usb_hcd *hcd, struct usb_device *udev)
 
 	xhci_dbg(xhci, "%s called for udev %p\n", __func__, udev);
 	virt_dev = xhci->devs[udev->slot_id];
+	rhub_port = virt_dev->rhub_port;
+
+	/*
+	 * Skip configure endpoint if the roothub port is gone or its link
+	 * is inactive. Some hosts stop completing endpoint commands after
+	 * disconnect, wedging the command ring.
+	 */
+	if (rhub_port->link_inactive || !rhub_port->connected)
+		return -ENODEV;
 
 	command = xhci_alloc_command(xhci, true, GFP_KERNEL);
 	if (!command)
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 8+ messages in thread

* Re: [PATCH 0/2] xhci: handshake timeout overrun and configure endpoint hang on device disconnect
  2026-09-30 10:17 [PATCH 0/2] xhci: handshake timeout overrun and configure endpoint hang on device disconnect Henry Tseng
  2026-09-30 10:17 ` [PATCH 1/2] xhci: make xhci_handshake() timeout wall-clock based again Henry Tseng
  2026-09-30 10:17 ` [PATCH 2/2] xhci: skip configure endpoint when dropping endpoints of a disconnected device Henry Tseng
@ 2026-10-02  9:30 ` Michal Pecio
  2026-10-07  9:52   ` Henry Tseng
  2 siblings, 1 reply; 8+ messages in thread
From: Michal Pecio @ 2026-10-02  9:30 UTC (permalink / raw)
  To: Henry Tseng; +Cc: Mathias Nyman, Greg Kroah-Hartman, linux-usb

On Wed, 30 Sep 2026 18:17:50 +0800, Henry Tseng wrote:
> This series addresses two problems found while debugging an AMD Raven
> USB 3.1 xHCI (1022:15e0) that gets declared dead when a USB storage
> enclosure is unplugged. During teardown a command fails to complete,
> aborting the command ring also fails, and the whole host is declared
> dead, taking unrelated devices on other root ports down with it:
> 
>   xhci_hcd 0000:0c:00.3: Command timeout, USBSTS: 0x00000010 PCD
>   xhci_hcd 0000:0c:00.3: Abort command ring
>   xhci_hcd 0000:0c:00.3: Abort failed to stop command ring: -110
>   xhci_hcd 0000:0c:00.3: xHCI host controller not responding, assume dead

Sounds like UAS, because I doubt that similar problems with ordinary
bulk endpoints could remain unknown for long.

I wonder if the kernel may be doing something crazy or out of spec
to deserve "undefined xHC behavior". See notes in xHCI 4.6.6, similar
requirements are also spelled in 4.6.4.

Is this easily reproducible? Could you send debugfs of this failure,
preferably before the "assume dead" message? See also my patch below.

zip -r debugfs.zip /sys/kernel/debug/usb/xhci/0000:0c:00.3/

> Reproduced on mainline v7.3-rc5:

Any other known affected or unaffected releases?

> Patch 2 fixes the host death on disconnect. The command that never
> completed is a configure endpoint command issued during teardown of a
> device behind the disconnected root port, right before disable slot for
> the same slot. Skip it when the roothub port is gone, as
> xhci_check_bandwidth() already does when the host is being removed.

That's not exactly the same, becasue with the xHC gone, we need not
worry what happens later. You found that Disable Slot works, so that's
OK, at least with this HC. Not sure about Reset Device, in case it's
not a disconnection but SS.Inactive due to link error.

> Patch 1 is an independent handshake overrun noticed during the same
> debugging, and does not fix the disconnect hang on its own. The
> command abort handshake has a 5 s timeout but took 15.8 s with
> interrupts disabled.

This patch should fix the "interrupts disabled" part:
https://lore.kernel.org/linux-usb/20260824095944.1c8335fa.michal.pecio@gmail.com/

And yes, the timeout is actually longer than intended. Interesting that
this is apparently a regression due to core changes, not an xhci-hcd
bug. It's possible that other drivers were similarly affected.

Regards,
Michal

^ permalink raw reply	[flat|nested] 8+ messages in thread

* Re: [PATCH 0/2] xhci: handshake timeout overrun and configure endpoint hang on device disconnect
  2026-10-02  9:30 ` [PATCH 0/2] xhci: handshake timeout overrun and configure endpoint hang on device disconnect Michal Pecio
@ 2026-10-07  9:52   ` Henry Tseng
  2026-10-08  9:06     ` Michal Pecio
  0 siblings, 1 reply; 8+ messages in thread
From: Henry Tseng @ 2026-10-07  9:52 UTC (permalink / raw)
  To: Michal Pecio; +Cc: Mathias Nyman, Greg Kroah-Hartman, linux-usb

[-- Attachment #1: Type: text/plain, Size: 4973 bytes --]

Hi Michal,

On Fri, 2 Oct 2026 11:30:12 +0200, Michal Pecio <michal.pecio@gmail.com> wrote:
> On Wed, 30 Sep 2026 18:17:50 +0800, Henry Tseng wrote:
> > This series addresses two problems found while debugging an AMD Raven
> > USB 3.1 xHCI (1022:15e0) that gets declared dead when a USB storage
> > enclosure is unplugged. During teardown a command fails to complete,
> > aborting the command ring also fails, and the whole host is declared
> > dead, taking unrelated devices on other root ports down with it:
> > 
> >   xhci_hcd 0000:0c:00.3: Command timeout, USBSTS: 0x00000010 PCD
> >   xhci_hcd 0000:0c:00.3: Abort command ring
> >   xhci_hcd 0000:0c:00.3: Abort failed to stop command ring: -110
> >   xhci_hcd 0000:0c:00.3: xHCI host controller not responding, assume dead
> 
> Sounds like UAS, because I doubt that similar problems with ordinary
> bulk endpoints could remain unknown for long.
> 

Yes, the disk is attached through uas. However, the host is still
declared dead when the enclosure is unplugged with no disk installed.

> I wonder if the kernel may be doing something crazy or out of spec
> to deserve "undefined xHC behavior". See notes in xHCI 4.6.6, similar
> requirements are also spelled in 4.6.4.
> 

The debugfs was checked against those notes and nothing suspicious
stood out, although something may well have been missed.

> Is this easily reproducible? Could you send debugfs of this failure,
> preferably before the "assume dead" message? See also my patch below.
> 

Yes, it's easy to reproduce. Plug in the enclosure, wait for
enumeration to complete, then unplug it, and the host ends up declared
dead every time. No fast replug is involved.

The debugfs is attached. I captured it while the command abort was in
progress, after the "Abort command ring" message and before "Abort
failed to stop command ring". port_bandwidth/ is left out, since
reading it takes xhci->lock and queues a get port bandwidth command.
In an earlier attempt those reads blocked until the host was declared
dead and then failed with -ESHUTDOWN.

> Any other known affected or unaffected releases?
> 

We first saw this on our v5.2.10-based vendor kernel, and Ubuntu 20.04
showed the same symptom at that time. The exact Ubuntu kernel version
wasn't recorded, but the 20.04 GA kernel is v5.4. Other mainline
releases haven't been tested yet, and I don't know of an unaffected
one.

> > Patch 2 fixes the host death on disconnect. The command that never
> > completed is a configure endpoint command issued during teardown of a
> > device behind the disconnected root port, right before disable slot for
> > the same slot. Skip it when the roothub port is gone, as
> > xhci_check_bandwidth() already does when the host is being removed.
> 
> That's not exactly the same, becasue with the xHC gone, we need not
> worry what happens later. You found that Disable Slot works, so that's
> OK, at least with this HC. Not sure about Reset Device, in case it's
> not a disconnection but SS.Inactive due to link error.
> 

I haven't found a way to test that case. Could you share a way to
trigger it, or the command sequence you have in mind?

Later tests suggest the link_inactive part of the check may not be
needed. With a debug print added to xhci_check_bandwidth(), the link
was never inactive whenever it was called during unplugging. Dropping
it would also leave a port in SS.Inactive with the device still 
attached on the current path.

> > Patch 1 is an independent handshake overrun noticed during the same
> > debugging, and does not fix the disconnect hang on its own. The
> > command abort handshake has a 5 s timeout but took 15.8 s with
> > interrupts disabled.
> 
> This patch should fix the "interrupts disabled" part:
> https://lore.kernel.org/linux-usb/20260824095944.1c8335fa.michal.pecio@gmail.com/
> 
> And yes, the timeout is actually longer than intended. Interesting that
> this is apparently a regression due to core changes, not an xhci-hcd
> bug. It's possible that other drivers were similarly affected.
> 

Thanks. Your patch takes care of holding the lock across the abort
handshake, but it doesn't change how long xhci_handshake() itself
takes, and xhci_handshake() is still called with xhci->lock held and
interrupts disabled in other places. For example, xhci_resume() waits
for STS_CNR with a 10 s timeout and calls xhci_reset() with
XHCI_RESET_LONG_USEC (10 s), both under spin_lock_irq(). On the two
hosts I measured, a 10 s xhci_handshake() that runs to its timeout
returned after 17-19 s, so I think patch 1 is still worth having.

As for the core change, the late timeout appears to be by design. When
this was raised after commit 7349a69cf312 [1], the advice was to pick
a delay_us larger than the time op() takes. Since xhci_handshake() is
shared by many different hosts, it would be hard to pick a single
value that suits all of them.

[1] https://lore.kernel.org/all/20240326013119.10591-1-zong.li@sifive.com/

Thanks,
Henry

[-- Attachment #2: debugfs.zip --]
[-- Type: application/zip, Size: 74138 bytes --]

^ permalink raw reply	[flat|nested] 8+ messages in thread

* Re: [PATCH 0/2] xhci: handshake timeout overrun and configure endpoint hang on device disconnect
  2026-10-07  9:52   ` Henry Tseng
@ 2026-10-08  9:06     ` Michal Pecio
  2026-10-08  9:13       ` Michal Pecio
  2026-10-08 10:25       ` Henry Tseng
  0 siblings, 2 replies; 8+ messages in thread
From: Michal Pecio @ 2026-10-08  9:06 UTC (permalink / raw)
  To: Henry Tseng; +Cc: Mathias Nyman, Greg Kroah-Hartman, linux-usb

On Wed, 07 Oct 2026 17:52:13 +0800, Henry Tseng wrote:
> port_bandwidth/ is left out, since reading it takes xhci->lock and
> queues a get port bandwidth command.

Thanks for pointing it out, we do need to skip this in such cases.

> xhci_handshake() is still called with xhci->lock held and interrupts
> disabled in other places.

Yes, we know. They need to be individually fixed to drop the lock.

Your debugfs shows the following tree of devices still existing:

# USB3 hub
devices/01/name:1-1
devices/02/name:2-1

	# unknown USB3 device
	devices/04/name:2-1.3

	# another USB3 hub
	devices/06/name:2-1.4
	devices/07/name:1-1.4

		# USB3 UAS device
		devices/08/name:2-1.4.4

Event ring dump (--> shows corresponding commands):

 1 0x00000000ffffc230: TRB 00000000bb0324f0 status 'Short Packet' len 96 slot 8 ep 7 type 'Transfer Event' flags e:c
 1 0x00000000ffffc240: TRB 0000000001000000 status 'Success' len 0 slot 0 ep 0 type 'Port Status Change Event' flags e:c
 1 0x00000000ffffc250: TRB 00000000ffffe5e0 status 'Success' len 0 slot 3 ep 0 type 'Command Completion Event' flags e:c
-->  0 0x00000000ffffe5e0: Configure Endpoint Command: ctx 00000000ba6ac000 slot 3 flags d:C 
 1 0x00000000ffffc260: TRB 00000000ffffe5f0 status 'Success' len 0 slot 3 ep 0 type 'Command Completion Event' flags e:c
-->  0 0x00000000ffffe5f0: Disable Slot Command: slot 3 flags C 
 1 0x00000000ffffc270: TRB 00000000bb067000 status 'Stopped' len 1 slot 5 ep 3 type 'Transfer Event' flags e:c
 1 0x00000000ffffc280: TRB 00000000ffffe600 status 'Success' len 0 slot 5 ep 0 type 'Command Completion Event' flags e:c
-->  0 0x00000000ffffe600: Stop Ring Command: slot 5 sp 0 ep 3 flags C 
 1 0x00000000ffffc290: TRB 0000000005000000 status 'Success' len 0 slot 0 ep 0 type 'Port Status Change Event' flags e:c
 1 0x00000000ffffc2a0: TRB 00000000ffffe610 status 'Success' len 0 slot 5 ep 0 type 'Command Completion Event' flags e:c
-->  0 0x00000000ffffe610: Configure Endpoint Command: ctx 00000000bb025000 slot 5 flags d:C 
 1 0x00000000ffffc2b0: TRB 00000000bb030000 status 'Stopped' len 2 slot 4 ep 7 type 'Transfer Event' flags e:c
 1 0x00000000ffffc2c0: TRB 00000000ffffe620 status 'Success' len 0 slot 4 ep 0 type 'Command Completion Event' flags e:c
-->  0 0x00000000ffffe620: Stop Ring Command: slot 4 sp 0 ep 7 flags C 
 1 0x00000000ffffc2d0: TRB 00000000fd0c7000 status 'USB Transaction Error' len 1 slot 7 ep 3 type 'Transfer Event' flags e:c
 1 0x00000000ffffc2e0: TRB 00000000ba600010 status 'USB Transaction Error' len 1 slot 1 ep 3 type 'Transfer Event' flags e:c

Some UAS transfer completes normally. Then the 1-1/2-1 hub is
disconnected from the root port 1/5. USB2 disconnection is reported
first and causes removal of unknown devices on slot 3 and 5.

USB3 disconnection is reported later and causes 2-1.3 endpoint 0x83
to be stopped (probably URB unlink?) and the device deconfigured.
Stop completes normally at bb030000 (possibly no transfer has ever
completed on this endpoint) and Configure Endpoint gets stuck.

The last sign of life from the xHC are Trannsaction Errors for USB2
status endpoints of those hubs. They get halted and never reset, so
technically it's a spec violation to disable these endpoints, but we
aren't even thinking about doing it yet.

Can you identify the 2-1.3 device? Can it be separated or is all 
this stuff inside one physical box?

Does it help to blacklist "uas" driver before connecting the whole
tree? My guess at this point: probably not, it's not a streams bug.

What happens if deconfigure manually before disconnection:

  echo 0 > /sys/bus/usb/devices/4-6/bConfigurationValue

Will it hang on deconfiguration, disconnection, or not at all?

Regards,
Michal

^ permalink raw reply	[flat|nested] 8+ messages in thread

* Re: [PATCH 0/2] xhci: handshake timeout overrun and configure endpoint hang on device disconnect
  2026-10-08  9:06     ` Michal Pecio
@ 2026-10-08  9:13       ` Michal Pecio
  2026-10-08 10:25       ` Henry Tseng
  1 sibling, 0 replies; 8+ messages in thread
From: Michal Pecio @ 2026-10-08  9:13 UTC (permalink / raw)
  To: Henry Tseng; +Cc: Mathias Nyman, Greg Kroah-Hartman, linux-usb

On Thu, 8 Oct 2026 11:06:42 +0200, Michal Pecio wrote:
> What happens if deconfigure manually before disconnection:
> 
>   echo 0 > /sys/bus/usb/devices/4-6/bConfigurationValue

Sorry, I pasted this from my own terminal. Of course, it needs to be:

  echo 0 > /sys/bus/usb/devices/2-1.3/bConfigurationValue

This assumes that device names stay the same. They will stay the same
if you connect the same devices to the same ports as before. If unsure,
see 'dmesg' and 'lsusb -t'.

^ permalink raw reply	[flat|nested] 8+ messages in thread

* Re: [PATCH 0/2] xhci: handshake timeout overrun and configure endpoint hang on device disconnect
  2026-10-08  9:06     ` Michal Pecio
  2026-10-08  9:13       ` Michal Pecio
@ 2026-10-08 10:25       ` Henry Tseng
  1 sibling, 0 replies; 8+ messages in thread
From: Henry Tseng @ 2026-10-08 10:25 UTC (permalink / raw)
  To: Michal Pecio; +Cc: Henry Tseng, Mathias Nyman, Greg Kroah-Hartman, linux-usb

Hi Michal,

On Thu, 8 Oct 2026 11:06:42 +0200, Michal Pecio <michal.pecio@gmail.com> wrote:
> Can you identify the 2-1.3 device? Can it be separated or is all 
> this stuff inside one physical box?
> 

2-1.3 is a hub, 1c04:0018 "QNAP Systems, Inc. USB3.2 Hub". 2-1 and
2-1.4 have the same VID:PID. They are all inside the TL-D800C and
can't be separated.

/:  Bus 001.Port 001: Dev 001, Class=root_hub, Driver=xhci_hcd/4p, 480M
    ID 1d6b:0002 Linux Foundation 2.0 root hub
    |__ Port 001: Dev 002, If 0, Class=Hub, Driver=hub/4p, 480M
        ID 1c04:0018 QNAP System Inc.
        |__ Port 001: Dev 003, If 0, Class=Communications, Driver=cdc_acm, 12M
            ID 04d8:000a Microchip Technology, Inc. CDC RS-232 Emulation Demo
        |__ Port 001: Dev 003, If 1, Class=CDC Data, Driver=cdc_acm, 12M
            ID 04d8:000a Microchip Technology, Inc. CDC RS-232 Emulation Demo
        |__ Port 003: Dev 004, If 0, Class=Hub, Driver=hub/4p, 480M
            ID 1c04:0610 QNAP System Inc.
        |__ Port 004: Dev 005, If 0, Class=Hub, Driver=hub/4p, 480M
            ID 1c04:0610 QNAP System Inc.
/:  Bus 002.Port 001: Dev 001, Class=root_hub, Driver=xhci_hcd/4p, 10000M
    ID 1d6b:0003 Linux Foundation 3.0 root hub
    |__ Port 001: Dev 002, If 0, Class=Hub, Driver=hub/4p, 10000M
        ID 1c04:0018 QNAP System Inc.
        |__ Port 003: Dev 003, If 0, Class=Hub, Driver=hub/4p, 10000M
            ID 1c04:0018 QNAP System Inc.
        |__ Port 004: Dev 004, If 0, Class=Hub, Driver=hub/4p, 10000M
            ID 1c04:0018 QNAP System Inc.
            |__ Port 004: Dev 005, If 0, Class=Mass Storage, Driver=uas, 10000M
                ID 174c:55aa ASMedia Technology Inc. ASM1051E SATA 6Gb/s bridge, ASM1053E SATA 6Gb/s bridge, ASM1153 SATA 3Gb/s bridge, ASM1153E SATA 6Gb/s bridge

> Does it help to blacklist "uas" driver before connecting the whole
> tree? My guess at this point: probably not, it's not a streams bug.
> 

No, the host still dies.

> What happens if deconfigure manually before disconnection:
> 
>   echo 0 > /sys/bus/usb/devices/4-6/bConfigurationValue
> 
> Will it hang on deconfiguration, disconnection, or not at all?
> 

I used 2-1.3 as in your follow-up. The whole tree enumerates the same
way in every run. The logs in patch 2 were taken with an extra USB
stick on 2-2 of the same host, which shifted slot IDs by one, so 2-1.3
was slot 5 there and slot 4 here.

Deconfiguring 2-1.3 completes fine:

  [   87.929570] xhci_hcd 0000:0c:00.3: Cancel URB 0000000064a1e51d, dev 1.3, ep 0x83, starting at offset 0xfff49000
  [   87.929675] xhci_hcd 0000:0c:00.3: Stopped on Transfer TRB for slot 4 ep 6
  [   87.929756] xhci_hcd 0000:0c:00.3: Successful Set TR Deq Ptr cmd, deq = @fff49010
  [   87.930564] xhci_hcd 0000:0c:00.3: drop ep 0x83, slot id 4, new drop flags = 0x80, new add flags = 0x0
  [   87.937756] xhci_hcd 0000:0c:00.3: Successful Endpoint Configure command

The host still dies on disconnection, but the stuck command is now the
configure endpoint dropping 0x83 of the 2-1.4 hub (slot 6):

  [  164.793408] usb 2-1: USB disconnect, device number 2
  [  164.793419] usb 2-1.3: USB disconnect, device number 3
  ...
  [  164.803390] usb 2-1.4: USB disconnect, device number 4
  ...
  [  164.834173] xhci_hcd 0000:0c:00.3: Cancel URB 000000009ac4b95a, dev 1.4, ep 0x83, starting at offset 0xfff15010
  [  164.834219] xhci_hcd 0000:0c:00.3: Stopped on Transfer TRB for slot 6 ep 6
  [  164.834819] xhci_hcd 0000:0c:00.3: drop ep 0x83, slot id 6, new drop flags = 0x80, new add flags = 0x0
  [  170.058142] xhci_hcd 0000:0c:00.3: Command timeout, USBSTS: 0x00000010 PCD
  [  185.909602] xhci_hcd 0000:0c:00.3: Abort failed to stop command ring: -110
  [  185.921446] xhci_hcd 0000:0c:00.3: xHCI host controller not responding, assume dead

If 2-1 is deconfigured instead, the same command completes for all
three hubs (2-1.3, 2-1.4 and 2-1 on slots 4, 6 and 2) while still
connected, and the host survives the unplug:

  [  165.209890] xhci_hcd 0000:0c:00.3: drop ep 0x83, slot id 4, new drop flags = 0x80, new add flags = 0x0
  [  165.210644] xhci_hcd 0000:0c:00.3: Successful Endpoint Configure command
  ...
  [  165.397383] xhci_hcd 0000:0c:00.3: drop ep 0x83, slot id 6, new drop flags = 0x80, new add flags = 0x0
  [  165.402651] xhci_hcd 0000:0c:00.3: Successful Endpoint Configure command
  ...
  [  165.404152] xhci_hcd 0000:0c:00.3: drop ep 0x83, slot id 2, new drop flags = 0x80, new add flags = 0x0
  [  165.418650] xhci_hcd 0000:0c:00.3: Successful Endpoint Configure command
  ...
  [  173.245013] usb 2-1: USB disconnect, device number 2

So far this command has only hung when issued after the disconnect.

Thanks,
Henry

^ permalink raw reply	[flat|nested] 8+ messages in thread

end of thread, other threads:[~2026-10-08 10:25 UTC | newest]

Thread overview: 8+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-30 10:17 [PATCH 0/2] xhci: handshake timeout overrun and configure endpoint hang on device disconnect Henry Tseng
2026-09-30 10:17 ` [PATCH 1/2] xhci: make xhci_handshake() timeout wall-clock based again Henry Tseng
2026-09-30 10:17 ` [PATCH 2/2] xhci: skip configure endpoint when dropping endpoints of a disconnected device Henry Tseng
2026-10-02  9:30 ` [PATCH 0/2] xhci: handshake timeout overrun and configure endpoint hang on device disconnect Michal Pecio
2026-10-07  9:52   ` Henry Tseng
2026-10-08  9:06     ` Michal Pecio
2026-10-08  9:13       ` Michal Pecio
2026-10-08 10:25       ` Henry Tseng

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox