Linux USB
 help / color / mirror / Atom feed
* AMD Strix Halo USB4: 20/40 Gbit/s negotiation with RJ45 and original host-reset patch regression
@ 2026-09-15 11:21 Nils Arnold
  2026-09-15 12:49 ` Mika Westerberg
  0 siblings, 1 reply; 7+ messages in thread
From: Nils Arnold @ 2026-09-15 11:21 UTC (permalink / raw)
  To: linux-usb; +Cc: Sanath.S, Basavaraj.Natikar, superm1, westeri

Hello Sanath, Mario, Mika and Basavaraj,

I would like to add test evidence from two Strix Halo hosts to this
USB4 reset discussion. My reset test used a backport of the ORIGINAL
f1de1fc5f632 commit, not the newer deferred-reset implementation.
I have not tested the revisions discussed on September 15.

I also observed a separate 20/40 Gbit/s negotiation issue involving RJ45.
I am keeping the observations separate because I have not established
a common cause.

The OdinLink side of this report is also tracked here:
https://github.com/Geramy/OdinLink-Five/issues/31

Separate physical-link negotiation observation
---------------------------------------------
I connected both hosts to the LAN through their RJ45 Ethernet ports
and directly to each other with one USB4 cable using the rear ports.
With both RJ45 connections active, USB4 reported RX/TX 10 Gbit/s x2
(20 Gbit/s per direction) on both hosts. An initial USB4 reconnection
still resulted in that lower rate.

I then disconnected the RJ45 cable on Host B, leaving Host A's RJ45
connection in place. I subsequently unplugged and reconnected USB4.
Both hosts then reported RX/TX 20 Gbit/s x2 (40 Gbit/s per direction).
The kernel journal records Host B's Ethernet link going down, followed
about 3.5 seconds later by USB4 disconnection and about 13 seconds after
that by peer rediscovery. Neither host rebooted during this sequence.

I then reconnected only Host B's RJ45 cable. Its Ethernet link returned
at 1 Gbit/s full duplex, and USB4 REMAINED at 40 Gbit/s on both hosts.
Thus 40 Gbit/s was recovered after removing one RJ45 connection and
reconnecting USB4; it did not require keeping RJ45 disconnected afterward.
This is an observed association during negotiation, not proof that RJ45
causes the lower rate: both Ethernet link state and USB4 reconnection
changed before the successful negotiation, and no repeated controlled
A/B series isolated them.

A separate boot test with both RJ45 connections active also showed
20 Gbit/s USB4 before OdinLink was loaded; loading the identical OdinLink
driver afterward did not change that rate.

The 40 Gbit/s value describes negotiated physical link speed, not a
successful throughput test. After recovery, Thunderbolt networking still
exhibited a TX stall after approximately 255 packets. These observations
are separate from the reset-patch experiment below; a common cause has
not been established.

My observation with an Intel Thunderbolt peer
-----------------------------------------------------------
I also observed successful 40 Gbit/s link negotiation when connecting an
AMD USB4 host to a third computer with an Intel Thunderbolt controller.
This is a useful comparison with the AMD-to-AMD connection
that negotiated only 20 Gbit/s in the cases above.

I have not yet located a specific captured log for this observation.
I have not established the exact Intel controller ID, which AMD host and
port I used, cable identity, RJ45 state, software versions, or repeat count
for this report. I therefore cannot present it as a controlled A/B
comparison or as evidence that all Intel-peer connections are stable.
It suggests investigating peer-dependent link negotiation, but does not
isolate AMD hardware, firmware, Linux software, or the cable as the cause.

Original reset-patch experiment
------------------------------

I tested the original commit
f1de1fc5f632cdeae1f5c2984572ab710d4dfcaa, backported to Fedora
7.2.4-200.fc44.x86_64. I have not tested the v5 hot-unplug patch or the
subsequent deferred-reset/workqueue proposal discussed on September 10.

Setup
-----
Two AMD Ryzen AI Max+ 395 / Strix Halo hosts connected directly over USB4.
AMD USB4 NHI PCI device IDs present: 1022:158d and 1022:158e.
Both hosts ran the same Fedora kernel version, thunderbolt_net, and the
out-of-tree OdinLink driver (odl_tb5), wkljohn/OdinLink-Five commit
bbcd87c2249935769960362dc891b700ec6ace27.
OdinLink was READY and Thunderbolt networking was reachable before teardown.
I stopped model inference for the transport test.

Backport detail
---------------
The Fedora source used nhi_pci_shutdown where the newer upstream context
used nhi_pci_release_irq. The existing shutdown callback was retained and
the reset_interface = nhi_reset_interface callback was added. No other
changes to the patch were made. The module was built against the running
kernel headers and Module.symvers and loaded temporarily, with Secure Boot
disabled. This is a modified-kernel, out-of-tree-driver configuration.

Comparison
----------
With the stock Fedora thunderbolt module, I completed two series of five
OdinLink unload/reload and bidirectional integrity-test cycles (10/10).
The second series also checked Thunderbolt-network ping each cycle.
Each load test used eight streams per host, 400 messages per stream, with
message sizes alternating between 1048576, 350208, 128, and 8192 bytes.
No new CRC, overflow, fragment-sequence, or payload-integrity errors were
reported in those baseline cycles.

With the backported original reset patch, the first coordinated OdinLink
unload failed. The logs confirm that the host-interface reset ran on both
hosts. Both logged ring_interrupt_active warnings. Host A also logged
tb_ring_free warnings. Host B logged a dangling control request and a
tb_ctl_stop warning before its reset, followed by ring warnings, list_del
corruption, and kernel BUG at lib/list_debug.c:62!.

On Host B, rmmod odl_tb5 remained in uninterruptible sleep. Thunderbolt
network connectivity was lost, and the intended load test could not start:
0/5 patched cycles completed. Software reboot requests did not complete
during the subsequent observation period; hard power cycles were required.
I did not install the test module persistently.

Interpretation and limits
-------------------------
This is additional evidence of the original host-wide reset interacting
badly with concurrent networking and an out-of-tree DMA service. It does
not isolate a mainline-only bug or establish the precise cause of the list
corruption. The failure was observed in one patched teardown attempt; the
stock baseline passed ten cycles. The newer workqueue proposal has not
been tested in this setup.

These results do not establish a cause or fix for a separate 20-versus-40
Gbit/s physical link negotiation issue. They also do not validate a fix
for sporadic CRC errors, which did not reproduce in the baseline series.

Would this concurrent tbnet/OdinLink teardown case be useful for reviewing
the deferred-reset implementation, and which specific trace points would
help distinguish an OdinLink teardown issue from a core NHI reset issue?

Please feel free to reply to this email if you need additional information,
sanitized logs, or specific test steps. I can help with targeted testing
on the affected hardware, subject to system availability. Please include
the exact patch/version and test procedure you would like me to use.

Related discussion:
https://lkml.iu.edu/2609.1/08435.html

Current thread context:
https://lists.openwall.net/linux-kernel/2026/09/15/576


^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: AMD Strix Halo USB4: 20/40 Gbit/s negotiation with RJ45 and original host-reset patch regression
  2026-09-15 11:21 AMD Strix Halo USB4: 20/40 Gbit/s negotiation with RJ45 and original host-reset patch regression Nils Arnold
@ 2026-09-15 12:49 ` Mika Westerberg
  2026-09-15 20:42   ` Mario Limonciello
  0 siblings, 1 reply; 7+ messages in thread
From: Mika Westerberg @ 2026-09-15 12:49 UTC (permalink / raw)
  To: Nils Arnold; +Cc: linux-usb, Sanath.S, Basavaraj.Natikar, superm1, westeri

Hi,

On Tue, Sep 15, 2026 at 01:21:39PM +0200, Nils Arnold wrote:
> Hello Sanath, Mario, Mika and Basavaraj,
> 
> I would like to add test evidence from two Strix Halo hosts to this
> USB4 reset discussion. My reset test used a backport of the ORIGINAL
> f1de1fc5f632 commit, not the newer deferred-reset implementation.
> I have not tested the revisions discussed on September 15.

Okay the last message on the thread I asked if we should revert the commit
for v7.3-rcX and then try to get the "better" fix into v7.4. At least what
I understood it seems to solve the issue without the deadlock and killing
the whole host interface underneath everybody else.

Once we have that patch, you can test and see if that helps in your case
too. Now that we know AMD guys can CC you when they submit it.

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: AMD Strix Halo USB4: 20/40 Gbit/s negotiation with RJ45 and original host-reset patch regression
  2026-09-15 12:49 ` Mika Westerberg
@ 2026-09-15 20:42   ` Mario Limonciello
  2026-09-16  7:38     ` Mika Westerberg
  0 siblings, 1 reply; 7+ messages in thread
From: Mario Limonciello @ 2026-09-15 20:42 UTC (permalink / raw)
  To: Mika Westerberg, Nils Arnold
  Cc: linux-usb, Sanath.S, Basavaraj.Natikar, westeri

On 9/15/26 07:49, Mika Westerberg wrote:
> Hi,
> 
> On Tue, Sep 15, 2026 at 01:21:39PM +0200, Nils Arnold wrote:
>> Hello Sanath, Mario, Mika and Basavaraj,
>>
>> I would like to add test evidence from two Strix Halo hosts to this
>> USB4 reset discussion. My reset test used a backport of the ORIGINAL
>> f1de1fc5f632 commit, not the newer deferred-reset implementation.
>> I have not tested the revisions discussed on September 15.
> 
> Okay the last message on the thread I asked if we should revert the commit
> for v7.3-rcX and then try to get the "better" fix into v7.4. At least what
> I understood it seems to solve the issue without the deadlock and killing
> the whole host interface underneath everybody else.

FWIW to Nils I left a note on that thread and agree with Mika.

> 
> Once we have that patch, you can test and see if that helps in your case
> too. Now that we know AMD guys can CC you when they submit it.

One thing I'd like to note though; I would rather that we keep out of 
tree modules out of the conversation on the kernel mailing list when it 
comes to upstream behavior.

There are multiple ways to do xdomain tunnels in the kernel, and we 
should collectively fixate our time and energy on bugs related to those 
instead of PoC out of tree drivers.

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: AMD Strix Halo USB4: 20/40 Gbit/s negotiation with RJ45 and original host-reset patch regression
  2026-09-15 20:42   ` Mario Limonciello
@ 2026-09-16  7:38     ` Mika Westerberg
  2026-09-16 13:26       ` Mario Limonciello
  0 siblings, 1 reply; 7+ messages in thread
From: Mika Westerberg @ 2026-09-16  7:38 UTC (permalink / raw)
  To: Mario Limonciello
  Cc: Nils Arnold, linux-usb, Sanath.S, Basavaraj.Natikar, westeri

Hi,

On Tue, Sep 15, 2026 at 03:42:43PM -0500, Mario Limonciello wrote:
> On 9/15/26 07:49, Mika Westerberg wrote:
> > Hi,
> > 
> > On Tue, Sep 15, 2026 at 01:21:39PM +0200, Nils Arnold wrote:
> > > Hello Sanath, Mario, Mika and Basavaraj,
> > > 
> > > I would like to add test evidence from two Strix Halo hosts to this
> > > USB4 reset discussion. My reset test used a backport of the ORIGINAL
> > > f1de1fc5f632 commit, not the newer deferred-reset implementation.
> > > I have not tested the revisions discussed on September 15.
> > 
> > Okay the last message on the thread I asked if we should revert the commit
> > for v7.3-rcX and then try to get the "better" fix into v7.4. At least what
> > I understood it seems to solve the issue without the deadlock and killing
> > the whole host interface underneath everybody else.
> 
> FWIW to Nils I left a note on that thread and agree with Mika.
> 
> > 
> > Once we have that patch, you can test and see if that helps in your case
> > too. Now that we know AMD guys can CC you when they submit it.
> 
> One thing I'd like to note though; I would rather that we keep out of tree
> modules out of the conversation on the kernel mailing list when it comes to
> upstream behavior.
> 
> There are multiple ways to do xdomain tunnels in the kernel, and we should
> collectively fixate our time and energy on bugs related to those instead of
> PoC out of tree drivers.

Fully agree. But the same behaviour is with the in-tree Thunderbolt
networking and stream (and maybe DMA test) drivers so those should at least
be working with the AMD controller without wedging.

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: AMD Strix Halo USB4: 20/40 Gbit/s negotiation with RJ45 and original host-reset patch regression
  2026-09-16  7:38     ` Mika Westerberg
@ 2026-09-16 13:26       ` Mario Limonciello
  2026-09-16 13:35         ` Mika Westerberg
  0 siblings, 1 reply; 7+ messages in thread
From: Mario Limonciello @ 2026-09-16 13:26 UTC (permalink / raw)
  To: Mika Westerberg
  Cc: Nils Arnold, linux-usb, Sanath.S, Basavaraj.Natikar, westeri

On 9/16/26 02:38, Mika Westerberg wrote:
> Hi,
> 
> On Tue, Sep 15, 2026 at 03:42:43PM -0500, Mario Limonciello wrote:
>> On 9/15/26 07:49, Mika Westerberg wrote:
>>> Hi,
>>>
>>> On Tue, Sep 15, 2026 at 01:21:39PM +0200, Nils Arnold wrote:
>>>> Hello Sanath, Mario, Mika and Basavaraj,
>>>>
>>>> I would like to add test evidence from two Strix Halo hosts to this
>>>> USB4 reset discussion. My reset test used a backport of the ORIGINAL
>>>> f1de1fc5f632 commit, not the newer deferred-reset implementation.
>>>> I have not tested the revisions discussed on September 15.
>>>
>>> Okay the last message on the thread I asked if we should revert the commit
>>> for v7.3-rcX and then try to get the "better" fix into v7.4. At least what
>>> I understood it seems to solve the issue without the deadlock and killing
>>> the whole host interface underneath everybody else.
>>
>> FWIW to Nils I left a note on that thread and agree with Mika.
>>
>>>
>>> Once we have that patch, you can test and see if that helps in your case
>>> too. Now that we know AMD guys can CC you when they submit it.
>>
>> One thing I'd like to note though; I would rather that we keep out of tree
>> modules out of the conversation on the kernel mailing list when it comes to
>> upstream behavior.
>>
>> There are multiple ways to do xdomain tunnels in the kernel, and we should
>> collectively fixate our time and energy on bugs related to those instead of
>> PoC out of tree drivers.
> 
> Fully agree. But the same behaviour is with the in-tree Thunderbolt
> networking and stream (and maybe DMA test) drivers so those should at least
> be working with the AMD controller without wedging.

Yes and no.  That out of tree driver might be configured as a module, 
but at least I haven't audited it.  It could be causing other problems 
too when it accesses certain symbols.

For our own sanity I suggest we work with those in tree drivers and 
untainted kernel.  If the problem can reproduce with those we can work 
on fixing it.

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: AMD Strix Halo USB4: 20/40 Gbit/s negotiation with RJ45 and original host-reset patch regression
  2026-09-16 13:26       ` Mario Limonciello
@ 2026-09-16 13:35         ` Mika Westerberg
  2026-09-16 13:36           ` Mario Limonciello
  0 siblings, 1 reply; 7+ messages in thread
From: Mika Westerberg @ 2026-09-16 13:35 UTC (permalink / raw)
  To: Mario Limonciello
  Cc: Nils Arnold, linux-usb, Sanath.S, Basavaraj.Natikar, westeri

On Wed, Sep 16, 2026 at 08:26:07AM -0500, Mario Limonciello wrote:
> On 9/16/26 02:38, Mika Westerberg wrote:
> > Hi,
> > 
> > On Tue, Sep 15, 2026 at 03:42:43PM -0500, Mario Limonciello wrote:
> > > On 9/15/26 07:49, Mika Westerberg wrote:
> > > > Hi,
> > > > 
> > > > On Tue, Sep 15, 2026 at 01:21:39PM +0200, Nils Arnold wrote:
> > > > > Hello Sanath, Mario, Mika and Basavaraj,
> > > > > 
> > > > > I would like to add test evidence from two Strix Halo hosts to this
> > > > > USB4 reset discussion. My reset test used a backport of the ORIGINAL
> > > > > f1de1fc5f632 commit, not the newer deferred-reset implementation.
> > > > > I have not tested the revisions discussed on September 15.
> > > > 
> > > > Okay the last message on the thread I asked if we should revert the commit
> > > > for v7.3-rcX and then try to get the "better" fix into v7.4. At least what
> > > > I understood it seems to solve the issue without the deadlock and killing
> > > > the whole host interface underneath everybody else.
> > > 
> > > FWIW to Nils I left a note on that thread and agree with Mika.
> > > 
> > > > 
> > > > Once we have that patch, you can test and see if that helps in your case
> > > > too. Now that we know AMD guys can CC you when they submit it.
> > > 
> > > One thing I'd like to note though; I would rather that we keep out of tree
> > > modules out of the conversation on the kernel mailing list when it comes to
> > > upstream behavior.
> > > 
> > > There are multiple ways to do xdomain tunnels in the kernel, and we should
> > > collectively fixate our time and energy on bugs related to those instead of
> > > PoC out of tree drivers.
> > 
> > Fully agree. But the same behaviour is with the in-tree Thunderbolt
> > networking and stream (and maybe DMA test) drivers so those should at least
> > be working with the AMD controller without wedging.
> 
> Yes and no.  That out of tree driver might be configured as a module, but at
> least I haven't audited it.  It could be causing other problems too when it
> accesses certain symbols.

Yes agree this but I mean if the issue happens with the in-tree drivers we
should try to fix it the best way possible. And my understanding is that it
happens in your own validation too, right? This is the same AMD controller
wedge issue we are talking about where the fix was just reverted by you and
now you (AMD) are working on better one? Or I'm missing something :)

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: AMD Strix Halo USB4: 20/40 Gbit/s negotiation with RJ45 and original host-reset patch regression
  2026-09-16 13:35         ` Mika Westerberg
@ 2026-09-16 13:36           ` Mario Limonciello
  0 siblings, 0 replies; 7+ messages in thread
From: Mario Limonciello @ 2026-09-16 13:36 UTC (permalink / raw)
  To: Mika Westerberg
  Cc: Nils Arnold, linux-usb, Sanath.S, Basavaraj.Natikar, westeri

On 9/16/26 08:35, Mika Westerberg wrote:
> On Wed, Sep 16, 2026 at 08:26:07AM -0500, Mario Limonciello wrote:
>> On 9/16/26 02:38, Mika Westerberg wrote:
>>> Hi,
>>>
>>> On Tue, Sep 15, 2026 at 03:42:43PM -0500, Mario Limonciello wrote:
>>>> On 9/15/26 07:49, Mika Westerberg wrote:
>>>>> Hi,
>>>>>
>>>>> On Tue, Sep 15, 2026 at 01:21:39PM +0200, Nils Arnold wrote:
>>>>>> Hello Sanath, Mario, Mika and Basavaraj,
>>>>>>
>>>>>> I would like to add test evidence from two Strix Halo hosts to this
>>>>>> USB4 reset discussion. My reset test used a backport of the ORIGINAL
>>>>>> f1de1fc5f632 commit, not the newer deferred-reset implementation.
>>>>>> I have not tested the revisions discussed on September 15.
>>>>>
>>>>> Okay the last message on the thread I asked if we should revert the commit
>>>>> for v7.3-rcX and then try to get the "better" fix into v7.4. At least what
>>>>> I understood it seems to solve the issue without the deadlock and killing
>>>>> the whole host interface underneath everybody else.
>>>>
>>>> FWIW to Nils I left a note on that thread and agree with Mika.
>>>>
>>>>>
>>>>> Once we have that patch, you can test and see if that helps in your case
>>>>> too. Now that we know AMD guys can CC you when they submit it.
>>>>
>>>> One thing I'd like to note though; I would rather that we keep out of tree
>>>> modules out of the conversation on the kernel mailing list when it comes to
>>>> upstream behavior.
>>>>
>>>> There are multiple ways to do xdomain tunnels in the kernel, and we should
>>>> collectively fixate our time and energy on bugs related to those instead of
>>>> PoC out of tree drivers.
>>>
>>> Fully agree. But the same behaviour is with the in-tree Thunderbolt
>>> networking and stream (and maybe DMA test) drivers so those should at least
>>> be working with the AMD controller without wedging.
>>
>> Yes and no.  That out of tree driver might be configured as a module, but at
>> least I haven't audited it.  It could be causing other problems too when it
>> accesses certain symbols.
> 
> Yes agree this but I mean if the issue happens with the in-tree drivers we
> should try to fix it the best way possible. And my understanding is that it
> happens in your own validation too, right? This is the same AMD controller
> wedge issue we are talking about where the fix was just reverted by you and
> now you (AMD) are working on better one? Or I'm missing something :)

Yep, totally!

^ permalink raw reply	[flat|nested] 7+ messages in thread

end of thread, other threads:[~2026-09-16 13:36 UTC | newest]

Thread overview: 7+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-15 11:21 AMD Strix Halo USB4: 20/40 Gbit/s negotiation with RJ45 and original host-reset patch regression Nils Arnold
2026-09-15 12:49 ` Mika Westerberg
2026-09-15 20:42   ` Mario Limonciello
2026-09-16  7:38     ` Mika Westerberg
2026-09-16 13:26       ` Mario Limonciello
2026-09-16 13:35         ` Mika Westerberg
2026-09-16 13:36           ` Mario Limonciello

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox