Linux USB
 help / color / mirror / Atom feed
From: Henry Tseng <henrytseng@qnap.com>
To: Michal Pecio <michal.pecio@gmail.com>
Cc: Mathias Nyman <mathias.nyman@intel.com>,
	Greg Kroah-Hartman <gregkh@linuxfoundation.org>,
	linux-usb@vger.kernel.org
Subject: Re: [PATCH 0/2] xhci: handshake timeout overrun and configure endpoint hang on device disconnect
Date: Wed, 07 Oct 2026 17:52:13 +0800	[thread overview]
Message-ID: <179136673328.386669.9410030239460428155@qnap.com> (raw)
In-Reply-To: <20261002113012.271ceead.michal.pecio@gmail.com>

[-- Attachment #1: Type: text/plain, Size: 4973 bytes --]

Hi Michal,

On Fri, 2 Oct 2026 11:30:12 +0200, Michal Pecio <michal.pecio@gmail.com> wrote:
> On Wed, 30 Sep 2026 18:17:50 +0800, Henry Tseng wrote:
> > This series addresses two problems found while debugging an AMD Raven
> > USB 3.1 xHCI (1022:15e0) that gets declared dead when a USB storage
> > enclosure is unplugged. During teardown a command fails to complete,
> > aborting the command ring also fails, and the whole host is declared
> > dead, taking unrelated devices on other root ports down with it:
> > 
> >   xhci_hcd 0000:0c:00.3: Command timeout, USBSTS: 0x00000010 PCD
> >   xhci_hcd 0000:0c:00.3: Abort command ring
> >   xhci_hcd 0000:0c:00.3: Abort failed to stop command ring: -110
> >   xhci_hcd 0000:0c:00.3: xHCI host controller not responding, assume dead
> 
> Sounds like UAS, because I doubt that similar problems with ordinary
> bulk endpoints could remain unknown for long.
> 

Yes, the disk is attached through uas. However, the host is still
declared dead when the enclosure is unplugged with no disk installed.

> I wonder if the kernel may be doing something crazy or out of spec
> to deserve "undefined xHC behavior". See notes in xHCI 4.6.6, similar
> requirements are also spelled in 4.6.4.
> 

The debugfs was checked against those notes and nothing suspicious
stood out, although something may well have been missed.

> Is this easily reproducible? Could you send debugfs of this failure,
> preferably before the "assume dead" message? See also my patch below.
> 

Yes, it's easy to reproduce. Plug in the enclosure, wait for
enumeration to complete, then unplug it, and the host ends up declared
dead every time. No fast replug is involved.

The debugfs is attached. I captured it while the command abort was in
progress, after the "Abort command ring" message and before "Abort
failed to stop command ring". port_bandwidth/ is left out, since
reading it takes xhci->lock and queues a get port bandwidth command.
In an earlier attempt those reads blocked until the host was declared
dead and then failed with -ESHUTDOWN.

> Any other known affected or unaffected releases?
> 

We first saw this on our v5.2.10-based vendor kernel, and Ubuntu 20.04
showed the same symptom at that time. The exact Ubuntu kernel version
wasn't recorded, but the 20.04 GA kernel is v5.4. Other mainline
releases haven't been tested yet, and I don't know of an unaffected
one.

> > Patch 2 fixes the host death on disconnect. The command that never
> > completed is a configure endpoint command issued during teardown of a
> > device behind the disconnected root port, right before disable slot for
> > the same slot. Skip it when the roothub port is gone, as
> > xhci_check_bandwidth() already does when the host is being removed.
> 
> That's not exactly the same, becasue with the xHC gone, we need not
> worry what happens later. You found that Disable Slot works, so that's
> OK, at least with this HC. Not sure about Reset Device, in case it's
> not a disconnection but SS.Inactive due to link error.
> 

I haven't found a way to test that case. Could you share a way to
trigger it, or the command sequence you have in mind?

Later tests suggest the link_inactive part of the check may not be
needed. With a debug print added to xhci_check_bandwidth(), the link
was never inactive whenever it was called during unplugging. Dropping
it would also leave a port in SS.Inactive with the device still 
attached on the current path.

> > Patch 1 is an independent handshake overrun noticed during the same
> > debugging, and does not fix the disconnect hang on its own. The
> > command abort handshake has a 5 s timeout but took 15.8 s with
> > interrupts disabled.
> 
> This patch should fix the "interrupts disabled" part:
> https://lore.kernel.org/linux-usb/20260824095944.1c8335fa.michal.pecio@gmail.com/
> 
> And yes, the timeout is actually longer than intended. Interesting that
> this is apparently a regression due to core changes, not an xhci-hcd
> bug. It's possible that other drivers were similarly affected.
> 

Thanks. Your patch takes care of holding the lock across the abort
handshake, but it doesn't change how long xhci_handshake() itself
takes, and xhci_handshake() is still called with xhci->lock held and
interrupts disabled in other places. For example, xhci_resume() waits
for STS_CNR with a 10 s timeout and calls xhci_reset() with
XHCI_RESET_LONG_USEC (10 s), both under spin_lock_irq(). On the two
hosts I measured, a 10 s xhci_handshake() that runs to its timeout
returned after 17-19 s, so I think patch 1 is still worth having.

As for the core change, the late timeout appears to be by design. When
this was raised after commit 7349a69cf312 [1], the advice was to pick
a delay_us larger than the time op() takes. Since xhci_handshake() is
shared by many different hosts, it would be hard to pick a single
value that suits all of them.

[1] https://lore.kernel.org/all/20240326013119.10591-1-zong.li@sifive.com/

Thanks,
Henry

[-- Attachment #2: debugfs.zip --]
[-- Type: application/zip, Size: 74138 bytes --]

  reply	other threads:[~2026-10-07 10:04 UTC|newest]

Thread overview: 9+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-30 10:17 [PATCH 0/2] xhci: handshake timeout overrun and configure endpoint hang on device disconnect Henry Tseng
2026-09-30 10:17 ` [PATCH 1/2] xhci: make xhci_handshake() timeout wall-clock based again Henry Tseng
2026-09-30 10:17 ` [PATCH 2/2] xhci: skip configure endpoint when dropping endpoints of a disconnected device Henry Tseng
2026-10-02  9:30 ` [PATCH 0/2] xhci: handshake timeout overrun and configure endpoint hang on device disconnect Michal Pecio
2026-10-07  9:52   ` Henry Tseng [this message]
2026-10-08  9:06     ` Michal Pecio
2026-10-08  9:13       ` Michal Pecio
2026-10-08 10:25       ` Henry Tseng
2026-10-09 15:35         ` Michal Pecio

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=179136673328.386669.9410030239460428155@qnap.com \
    --to=henrytseng@qnap.com \
    --cc=gregkh@linuxfoundation.org \
    --cc=linux-usb@vger.kernel.org \
    --cc=mathias.nyman@intel.com \
    --cc=michal.pecio@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox