* Re: [REGRESSION] mt7925: MLO connectivity silently stalls with 6GHz link active
@ 2026-08-19 9:03 Jonas Hort
2026-08-22 20:18 ` Jonas Hort
0 siblings, 1 reply; 21+ messages in thread
From: Jonas Hort @ 2026-08-19 9:03 UTC (permalink / raw)
To: Devin Wittmayer
Cc: Thorsten Leemhuis, regressions, linux-wireless, Felix Fietkau,
lorenzo.bianconi83
Thanks for the detailed breakdown.
I'll build v7.0 vanilla and test it, as suggested. Fair warning
though: I'm on vacation this week, so I'll pick this up next week.
I've also never compiled a kernel before, so it'll likely take me a
bit of trial and error the first time around - please bear with me
if it takes a little longer than expected.
Will report back once I have results.
Thanks again,
Jonas
19.08.2026 03:18:34 Devin Wittmayer <lucid_duck@justthetip.ca>:
> Thank you very much, that answers both things.
>
> The ROC tracing is the more useful of the two even though it came back
> negative. Two of the three freezes have no ROC activity in them at all, so a
> link switch that never finished cannot be what starts this. The middle one
> does have rocabort, mloroc and rocwork in it, but one out of three makes
> that look like the exception rather than the pattern. So the area I sent you
> looking at is out, and that is worth knowing before you spend nights on
> builds.
>
> One other thing worth saying first. There is a five patch mt76 series on the
> list at the moment and two of the patches look like they were written for
> exactly this bug. I do not think they were, and it is your own numbers that
> show it. The failure 4/5 fixes stops mt76_txq_schedule_list from servicing
> the queue, and the one 5/5 fixes stops mt76_txq_send_burst once the non-AQL
> count reaches its cap. Both of those keep frames from ever reaching the
> hardware, so if either were your problem head would be sitting still
> alongside tail. Yours does the opposite. Head climbs 390 to 408 while tail
> stays at 260, so the frames are getting into the ring and nothing is
> finishing them, which is the far end of the same path. 2/5 is a use after
> free when an interface goes away, so it does not fit either. I would not
> expect that series to change what you see.
>
> On the bisect I would build v7.0 next. There are 32 mt7925 commits between
> 7.0 and 7.1 and 19 of them are one run of work from Sean Wang, reworking how
> the driver tracks the per link mlink and WCID for an MLO station. That is
> the kind of change that fits a bug only showing up with two links up. If
> v7.0 comes back clean, that series is where I would look. If v7.0 is already
> broken then it is off the hook and 6.19 becomes the next split. The mt76
> core and mac80211 both moved in the same window, so mt7925 is where I would
> look first rather than the only place worth looking.
>
> Devin
>
> Am 17.08.26 um 23:34 schrieb Jonas Hort:
>
>> Quick follow-up: managed to confirm 6.18 as a clean baseline on my
>> own hardware now (not just secondhand from others in the forum
>> thread) - running Linux 6.18.42-1-cachyos-lts with MLO active
>> (5GHz+6GHz, same FritzBox 5690 Pro) for 3 hours straight, no freeze
>> at all.
>
> Am 17.08.26 um 16:25 schrieb Jonas Hort:
>
>> One correction to how I described this earlier: the connection does
>> NOT reliably self-heal on its own. I have manually intervened every
>> single time to restore connectivity [...]
>>
>> Update on the ROC tracing: three real freezes captured now with the
>> kprobes active (all confirmed via the WFDMA0 tail-frozen signature).
>>
>> - Freeze #1 (15:37): no ROC activity in the trace.
>> - Freeze #2 (15:43): [...] The trace shows several ROC events
>> (rocabort, mloroc, rocwork) clustered together.
>> - Freeze #3 (16:15): no ROC activity again.
>>
>> Uploaded all three logs to the bugzilla ticket if useful:
>> https://bugzilla.kernel.org/show_bug.cgi?id=221884
^ permalink raw reply [flat|nested] 21+ messages in thread* Re: [REGRESSION] mt7925: MLO connectivity silently stalls with 6GHz link active 2026-08-19 9:03 [REGRESSION] mt7925: MLO connectivity silently stalls with 6GHz link active Jonas Hort @ 2026-08-22 20:18 ` Jonas Hort 2026-08-23 20:54 ` Jonas Hort 0 siblings, 1 reply; 21+ messages in thread From: Jonas Hort @ 2026-08-22 20:18 UTC (permalink / raw) To: Devin Wittmayer Cc: Thorsten Leemhuis, regressions, linux-wireless, Felix Fietkau, lorenzo.bianconi83 Quick update: v7.0 vanilla has been running clean for over 2 hours now with MLO active (5GHz+6GHz, same FritzBox 5690 Pro), no freeze at all. Am 19.08.26 um 11:03 schrieb Jonas Hort: > Thanks for the detailed breakdown. > > I'll build v7.0 vanilla and test it, as suggested. Fair warning > though: I'm on vacation this week, so I'll pick this up next week. > I've also never compiled a kernel before, so it'll likely take me a > bit of trial and error the first time around - please bear with me > if it takes a little longer than expected. > > Will report back once I have results. > > Thanks again, > Jonas > > 19.08.2026 03:18:34 Devin Wittmayer <lucid_duck@justthetip.ca>: > >> Thank you very much, that answers both things. >> >> The ROC tracing is the more useful of the two even though it came back >> negative. Two of the three freezes have no ROC activity in them at all, so a >> link switch that never finished cannot be what starts this. The middle one >> does have rocabort, mloroc and rocwork in it, but one out of three makes >> that look like the exception rather than the pattern. So the area I sent you >> looking at is out, and that is worth knowing before you spend nights on >> builds. >> >> One other thing worth saying first. There is a five patch mt76 series on the >> list at the moment and two of the patches look like they were written for >> exactly this bug. I do not think they were, and it is your own numbers that >> show it. The failure 4/5 fixes stops mt76_txq_schedule_list from servicing >> the queue, and the one 5/5 fixes stops mt76_txq_send_burst once the non-AQL >> count reaches its cap. Both of those keep frames from ever reaching the >> hardware, so if either were your problem head would be sitting still >> alongside tail. Yours does the opposite. Head climbs 390 to 408 while tail >> stays at 260, so the frames are getting into the ring and nothing is >> finishing them, which is the far end of the same path. 2/5 is a use after >> free when an interface goes away, so it does not fit either. I would not >> expect that series to change what you see. >> >> On the bisect I would build v7.0 next. There are 32 mt7925 commits between >> 7.0 and 7.1 and 19 of them are one run of work from Sean Wang, reworking how >> the driver tracks the per link mlink and WCID for an MLO station. That is >> the kind of change that fits a bug only showing up with two links up. If >> v7.0 comes back clean, that series is where I would look. If v7.0 is already >> broken then it is off the hook and 6.19 becomes the next split. The mt76 >> core and mac80211 both moved in the same window, so mt7925 is where I would >> look first rather than the only place worth looking. >> >> Devin >> >> Am 17.08.26 um 23:34 schrieb Jonas Hort: >> >>> Quick follow-up: managed to confirm 6.18 as a clean baseline on my >>> own hardware now (not just secondhand from others in the forum >>> thread) - running Linux 6.18.42-1-cachyos-lts with MLO active >>> (5GHz+6GHz, same FritzBox 5690 Pro) for 3 hours straight, no freeze >>> at all. >> Am 17.08.26 um 16:25 schrieb Jonas Hort: >> >>> One correction to how I described this earlier: the connection does >>> NOT reliably self-heal on its own. I have manually intervened every >>> single time to restore connectivity [...] >>> >>> Update on the ROC tracing: three real freezes captured now with the >>> kprobes active (all confirmed via the WFDMA0 tail-frozen signature). >>> >>> - Freeze #1 (15:37): no ROC activity in the trace. >>> - Freeze #2 (15:43): [...] The trace shows several ROC events >>> (rocabort, mloroc, rocwork) clustered together. >>> - Freeze #3 (16:15): no ROC activity again. >>> >>> Uploaded all three logs to the bugzilla ticket if useful: >>> https://bugzilla.kernel.org/show_bug.cgi?id=221884 ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] mt7925: MLO connectivity silently stalls with 6GHz link active 2026-08-22 20:18 ` Jonas Hort @ 2026-08-23 20:54 ` Jonas Hort 2026-08-24 4:50 ` Devin Wittmayer 2026-08-24 4:52 ` Thorsten Leemhuis 0 siblings, 2 replies; 21+ messages in thread From: Jonas Hort @ 2026-08-23 20:54 UTC (permalink / raw) To: Devin Wittmayer Cc: Thorsten Leemhuis, regressions, linux-wireless, Felix Fietkau, lorenzo.bianconi83 Bisect is done. First bad commit: ff643b81bc38eaff6c0ab783a62e4ba9e10d2476 wifi: mt76: mt7925: pass mlink and mconf to sta_mld_tlv() Sean Wang, 2026-03-06 https://patch.msgid.link/20260306232238.2039675-4-sean.wang@kernel.org Bisected on vanilla, path-restricted to drivers/net/wireless/mediatek/mt76/mt7925/, using v7.0 as good and v7.1 as bad. Each step was booted and tested with MLO active (5GHz+6GHz, FritzBox 5690 Pro), and every "bad" verdict was confirmed by the WFDMA0 signature (tail frozen while head keeps advancing), not just by the connection dropping. ea757740dd87 pass WCID indices to bss_basic_tlv() good (45 min clean) ff643b81bc38 pass mlink and mconf to sta_mld_tlv() bad (freeze after ~4 min) dc019e3294c7 pass mlink to mcu_sta_update() bad (freeze after ~3 min) 9e4d518a4707 pass mlink to mac_link_sta_remove() bad (freeze after ~4 min) cf9db836b1e0 pass mlink to set_link_key() bad (freeze after ~4 min) git bisect log and the incident logs for each step are available if useful - happy to attach them to the bugzilla ticket or send them here. Let me know if you want anything else tested. Thanks, Jonas Am 22.08.26 um 22:18 schrieb Jonas Hort: > Quick update: v7.0 vanilla has been running clean for over 2 hours now > with MLO active (5GHz+6GHz, same FritzBox 5690 Pro), no freeze at all. > > Am 19.08.26 um 11:03 schrieb Jonas Hort: >> Thanks for the detailed breakdown. >> >> I'll build v7.0 vanilla and test it, as suggested. Fair warning >> though: I'm on vacation this week, so I'll pick this up next week. >> I've also never compiled a kernel before, so it'll likely take me a >> bit of trial and error the first time around - please bear with me >> if it takes a little longer than expected. >> >> Will report back once I have results. >> >> Thanks again, >> Jonas >> >> 19.08.2026 03:18:34 Devin Wittmayer <lucid_duck@justthetip.ca>: >> >>> Thank you very much, that answers both things. >>> >>> The ROC tracing is the more useful of the two even though it came back >>> negative. Two of the three freezes have no ROC activity in them at >>> all, so a >>> link switch that never finished cannot be what starts this. The >>> middle one >>> does have rocabort, mloroc and rocwork in it, but one out of three >>> makes >>> that look like the exception rather than the pattern. So the area I >>> sent you >>> looking at is out, and that is worth knowing before you spend nights on >>> builds. >>> >>> One other thing worth saying first. There is a five patch mt76 >>> series on the >>> list at the moment and two of the patches look like they were >>> written for >>> exactly this bug. I do not think they were, and it is your own >>> numbers that >>> show it. The failure 4/5 fixes stops mt76_txq_schedule_list from >>> servicing >>> the queue, and the one 5/5 fixes stops mt76_txq_send_burst once the >>> non-AQL >>> count reaches its cap. Both of those keep frames from ever reaching the >>> hardware, so if either were your problem head would be sitting still >>> alongside tail. Yours does the opposite. Head climbs 390 to 408 >>> while tail >>> stays at 260, so the frames are getting into the ring and nothing is >>> finishing them, which is the far end of the same path. 2/5 is a use >>> after >>> free when an interface goes away, so it does not fit either. I would >>> not >>> expect that series to change what you see. >>> >>> On the bisect I would build v7.0 next. There are 32 mt7925 commits >>> between >>> 7.0 and 7.1 and 19 of them are one run of work from Sean Wang, >>> reworking how >>> the driver tracks the per link mlink and WCID for an MLO station. >>> That is >>> the kind of change that fits a bug only showing up with two links >>> up. If >>> v7.0 comes back clean, that series is where I would look. If v7.0 is >>> already >>> broken then it is off the hook and 6.19 becomes the next split. The >>> mt76 >>> core and mac80211 both moved in the same window, so mt7925 is where >>> I would >>> look first rather than the only place worth looking. >>> >>> Devin >>> >>> Am 17.08.26 um 23:34 schrieb Jonas Hort: >>> >>>> Quick follow-up: managed to confirm 6.18 as a clean baseline on my >>>> own hardware now (not just secondhand from others in the forum >>>> thread) - running Linux 6.18.42-1-cachyos-lts with MLO active >>>> (5GHz+6GHz, same FritzBox 5690 Pro) for 3 hours straight, no freeze >>>> at all. >>> Am 17.08.26 um 16:25 schrieb Jonas Hort: >>> >>>> One correction to how I described this earlier: the connection does >>>> NOT reliably self-heal on its own. I have manually intervened every >>>> single time to restore connectivity [...] >>>> >>>> Update on the ROC tracing: three real freezes captured now with the >>>> kprobes active (all confirmed via the WFDMA0 tail-frozen signature). >>>> >>>> - Freeze #1 (15:37): no ROC activity in the trace. >>>> - Freeze #2 (15:43): [...] The trace shows several ROC events >>>> (rocabort, mloroc, rocwork) clustered together. >>>> - Freeze #3 (16:15): no ROC activity again. >>>> >>>> Uploaded all three logs to the bugzilla ticket if useful: >>>> https://bugzilla.kernel.org/show_bug.cgi?id=221884 ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] mt7925: MLO connectivity silently stalls with 6GHz link active 2026-08-23 20:54 ` Jonas Hort @ 2026-08-24 4:50 ` Devin Wittmayer 2026-08-24 4:52 ` Thorsten Leemhuis 1 sibling, 0 replies; 21+ messages in thread From: Devin Wittmayer @ 2026-08-24 4:50 UTC (permalink / raw) To: Jonas Hort Cc: Sean Wang, Felix Fietkau, Thorsten Leemhuis, regressions, linux-wireless, lorenzo.bianconi83 On Sun, 2026-08-23 at 20:54 +0000, Jonas Hort wrote: > Bisect is done. First bad commit: > > ff643b81bc38eaff6c0ab783a62e4ba9e10d2476 > wifi: mt76: mt7925: pass mlink and mconf to sta_mld_tlv() Jonas, that is a careful bisect and it holds up here. The commit is Sean Wang's, absent from v7.0, present in v7.1, and ea757740dd87 really is its parent. Adding Sean, who knows this code far better than I do. The part I would look at first is the link count. Before, it came from the number of valid links, capped at two. After, the primary is always written and a second only when the caller passes one, with the count taken from however many got written. So a two link station updated on the primary would tell firmware it has one link where it used to say two. Subtle, and easy to miss in a change that is otherwise a straight simplification. That is from reading the diff rather than reproducing it, so Sean may well see something I have not. Nothing has touched that function since, so 7.1, 7.2 and mainline all behave the same way. Jonas, yes please to the bisect log. If the setup is still up, that commit also added a rate limited warning about a missing primary mconf, so it may be worth grepping your failing runs for MLD_TLV_LINK. Thank you for the work you put into this, it was not a small ask. Devin ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] mt7925: MLO connectivity silently stalls with 6GHz link active 2026-08-23 20:54 ` Jonas Hort 2026-08-24 4:50 ` Devin Wittmayer @ 2026-08-24 4:52 ` Thorsten Leemhuis 2026-08-24 8:18 ` Jonas Hort 1 sibling, 1 reply; 21+ messages in thread From: Thorsten Leemhuis @ 2026-08-24 4:52 UTC (permalink / raw) To: Sean Wang Cc: regressions, linux-wireless, Devin Wittmayer, Felix Fietkau, Jonas Hort, lorenzo.bianconi83, linux-mediatek On 8/23/26 22:54, Jonas Hort wrote: > Bisect is done. First bad commit: > > ff643b81bc38eaff6c0ab783a62e4ba9e10d2476 > wifi: mt76: mt7925: pass mlink and mconf to sta_mld_tlv() > Sean Wang, 2026-03-06 > https://patch.msgid.link/20260306232238.2039675-4-sean.wang@kernel.org Then let's add Sean to the list of recipients. :-D Sean, FWIW, this thread starts here: https://lore.kernel.org/all/066b30cc-a9e6-4aeb-964d-71551e8ea3ef@posteo.de/t/#u To quote the summary: "WiFi 7 MLO (5GHz+6GHz) on MT7925 silently stops passing traffic after a few minutes, while the driver continues to report a fully healthy link. Root cause appears to be a stalled WFDMA0 TX hardware queue." Ciao, Thorsten > Bisected on vanilla, path-restricted to > drivers/net/wireless/mediatek/mt76/mt7925/, using v7.0 as good and > v7.1 as bad. Each step was booted and tested with MLO active > (5GHz+6GHz, FritzBox 5690 Pro), and every "bad" verdict was confirmed > by the WFDMA0 signature (tail frozen while head keeps advancing), > not just by the connection dropping. > > ea757740dd87 pass WCID indices to bss_basic_tlv() good (45 min > clean) > ff643b81bc38 pass mlink and mconf to sta_mld_tlv() bad (freeze > after ~4 min) > dc019e3294c7 pass mlink to mcu_sta_update() bad (freeze > after ~3 min) > 9e4d518a4707 pass mlink to mac_link_sta_remove() bad (freeze > after ~4 min) > cf9db836b1e0 pass mlink to set_link_key() bad (freeze > after ~4 min) > > git bisect log and the incident logs for each step are available if > useful - happy to attach them to the bugzilla ticket or send them > here. > > Let me know if you want anything else tested. > > Thanks, > Jonas > > Am 22.08.26 um 22:18 schrieb Jonas Hort: >> Quick update: v7.0 vanilla has been running clean for over 2 hours now >> with MLO active (5GHz+6GHz, same FritzBox 5690 Pro), no freeze at all. >> >> Am 19.08.26 um 11:03 schrieb Jonas Hort: >>> Thanks for the detailed breakdown. >>> >>> I'll build v7.0 vanilla and test it, as suggested. Fair warning >>> though: I'm on vacation this week, so I'll pick this up next week. >>> I've also never compiled a kernel before, so it'll likely take me a >>> bit of trial and error the first time around - please bear with me >>> if it takes a little longer than expected. >>> >>> Will report back once I have results. >>> >>> Thanks again, >>> Jonas >>> >>> 19.08.2026 03:18:34 Devin Wittmayer <lucid_duck@justthetip.ca>: >>> >>>> Thank you very much, that answers both things. >>>> >>>> The ROC tracing is the more useful of the two even though it came back >>>> negative. Two of the three freezes have no ROC activity in them at >>>> all, so a >>>> link switch that never finished cannot be what starts this. The >>>> middle one >>>> does have rocabort, mloroc and rocwork in it, but one out of three >>>> makes >>>> that look like the exception rather than the pattern. So the area I >>>> sent you >>>> looking at is out, and that is worth knowing before you spend nights on >>>> builds. >>>> >>>> One other thing worth saying first. There is a five patch mt76 >>>> series on the >>>> list at the moment and two of the patches look like they were >>>> written for >>>> exactly this bug. I do not think they were, and it is your own >>>> numbers that >>>> show it. The failure 4/5 fixes stops mt76_txq_schedule_list from >>>> servicing >>>> the queue, and the one 5/5 fixes stops mt76_txq_send_burst once the >>>> non-AQL >>>> count reaches its cap. Both of those keep frames from ever reaching the >>>> hardware, so if either were your problem head would be sitting still >>>> alongside tail. Yours does the opposite. Head climbs 390 to 408 >>>> while tail >>>> stays at 260, so the frames are getting into the ring and nothing is >>>> finishing them, which is the far end of the same path. 2/5 is a use >>>> after >>>> free when an interface goes away, so it does not fit either. I would >>>> not >>>> expect that series to change what you see. >>>> >>>> On the bisect I would build v7.0 next. There are 32 mt7925 commits >>>> between >>>> 7.0 and 7.1 and 19 of them are one run of work from Sean Wang, >>>> reworking how >>>> the driver tracks the per link mlink and WCID for an MLO station. >>>> That is >>>> the kind of change that fits a bug only showing up with two links >>>> up. If >>>> v7.0 comes back clean, that series is where I would look. If v7.0 is >>>> already >>>> broken then it is off the hook and 6.19 becomes the next split. The >>>> mt76 >>>> core and mac80211 both moved in the same window, so mt7925 is where >>>> I would >>>> look first rather than the only place worth looking. >>>> >>>> Devin >>>> >>>> Am 17.08.26 um 23:34 schrieb Jonas Hort: >>>> >>>>> Quick follow-up: managed to confirm 6.18 as a clean baseline on my >>>>> own hardware now (not just secondhand from others in the forum >>>>> thread) - running Linux 6.18.42-1-cachyos-lts with MLO active >>>>> (5GHz+6GHz, same FritzBox 5690 Pro) for 3 hours straight, no freeze >>>>> at all. >>>> Am 17.08.26 um 16:25 schrieb Jonas Hort: >>>> >>>>> One correction to how I described this earlier: the connection does >>>>> NOT reliably self-heal on its own. I have manually intervened every >>>>> single time to restore connectivity [...] >>>>> >>>>> Update on the ROC tracing: three real freezes captured now with the >>>>> kprobes active (all confirmed via the WFDMA0 tail-frozen signature). >>>>> >>>>> - Freeze #1 (15:37): no ROC activity in the trace. >>>>> - Freeze #2 (15:43): [...] The trace shows several ROC events >>>>> (rocabort, mloroc, rocwork) clustered together. >>>>> - Freeze #3 (16:15): no ROC activity again. >>>>> >>>>> Uploaded all three logs to the bugzilla ticket if useful: >>>>> https://bugzilla.kernel.org/show_bug.cgi?id=221884 ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] mt7925: MLO connectivity silently stalls with 6GHz link active 2026-08-24 4:52 ` Thorsten Leemhuis @ 2026-08-24 8:18 ` Jonas Hort 2026-08-29 22:13 ` Devin Wittmayer 0 siblings, 1 reply; 21+ messages in thread From: Jonas Hort @ 2026-08-24 8:18 UTC (permalink / raw) To: Thorsten Leemhuis, Sean Wang Cc: regressions, linux-wireless, Devin Wittmayer, Felix Fietkau, lorenzo.bianconi83, linux-mediatek Thanks Devin, and thanks for pulling Sean in. git biscet log: # bad: [8cd9520d35a6c38db6567e97dd93b1f11f185dc6] Linux 7.1 # good: [028ef9c96e96197026887c0f092424679298aae8] Linux 7.0 git bisect start 'v7.1' 'v7.0' '--' 'drivers/net/wireless/mediatek/mt76/mt7925/' # good: [ea757740dd87c0b00c4844dd3282dff4d83fa3c7] wifi: mt76: mt7925: pass WCID indices to bss_basic_tlv() git bisect good ea757740dd87c0b00c4844dd3282dff4d83fa3c7 # bad: [cf9db836b1e069a3d6a80c72b9bdc12df78b0dd1] wifi: mt76: mt7925: pass mlink to set_link_key() git bisect bad cf9db836b1e069a3d6a80c72b9bdc12df78b0dd1 # bad: [9e4d518a4707175e1154876b760d4f2b39967e9d] wifi: mt76: mt7925: pass mlink to mac_link_sta_remove() git bisect bad 9e4d518a4707175e1154876b760d4f2b39967e9d # bad: [dc019e3294c7e3c6e997bb11732d46ce0d9211e9] wifi: mt76: mt7925: pass mlink to mcu_sta_update() git bisect bad dc019e3294c7e3c6e997bb11732d46ce0d9211e9 # bad: [ff643b81bc38eaff6c0ab783a62e4ba9e10d2476] wifi: mt76: mt7925: pass mlink and mconf to sta_mld_tlv() git bisect bad ff643b81bc38eaff6c0ab783a62e4ba9e10d2476 # first 'bad' commit: [ff643b81bc38eaff6c0ab783a62e4ba9e10d2476] wifi: mt76: mt7925: pass mlink and mconf to sta_mld_tlv() On the warning: no hits. I grepped all incident logs for MLD_TLV_LINK and for anything mentioning primary/mconf, and also checked the current boot directly. The only mt7925 lines present are the usual init ones: mt7925e 0000:09:00.0: ASIC revision: 79250000 mt7925e 0000:09:00.0: HW/SW Version: 0x8a108a10, Build Time: 20260605184651a mt7925e 0000:09:00.0: WM Firmware Version: ____000000, Build Time: 20260605184805 Each incident log captures journalctl -k for the 5 minutes leading up to the freeze plus the last 400 dmesg lines, so if it had fired around the stall it should have been in there. Still on the ff643b81bc38 kernel here, so happy to re-check with a different search term if I grepped for the wrong thing. Setup is still up and I can build and test whatever is useful. Thanks, Jonas Am 24.08.26 um 06:52 schrieb Thorsten Leemhuis: > On 8/23/26 22:54, Jonas Hort wrote: >> Bisect is done. First bad commit: >> >> ff643b81bc38eaff6c0ab783a62e4ba9e10d2476 >> wifi: mt76: mt7925: pass mlink and mconf to sta_mld_tlv() >> Sean Wang, 2026-03-06 >> https://patch.msgid.link/20260306232238.2039675-4-sean.wang@kernel.org > Then let's add Sean to the list of recipients. :-D > > Sean, FWIW, this thread starts here: > > https://lore.kernel.org/all/066b30cc-a9e6-4aeb-964d-71551e8ea3ef@posteo.de/t/#u > > To quote the summary: "WiFi 7 MLO (5GHz+6GHz) on MT7925 silently stops > passing traffic after a few minutes, while the driver continues to > report a fully healthy link. Root cause appears to be a stalled WFDMA0 > TX hardware queue." > > Ciao, Thorsten >> Bisected on vanilla, path-restricted to >> drivers/net/wireless/mediatek/mt76/mt7925/, using v7.0 as good and >> v7.1 as bad. Each step was booted and tested with MLO active >> (5GHz+6GHz, FritzBox 5690 Pro), and every "bad" verdict was confirmed >> by the WFDMA0 signature (tail frozen while head keeps advancing), >> not just by the connection dropping. >> >> ea757740dd87 pass WCID indices to bss_basic_tlv() good (45 min >> clean) >> ff643b81bc38 pass mlink and mconf to sta_mld_tlv() bad (freeze >> after ~4 min) >> dc019e3294c7 pass mlink to mcu_sta_update() bad (freeze >> after ~3 min) >> 9e4d518a4707 pass mlink to mac_link_sta_remove() bad (freeze >> after ~4 min) >> cf9db836b1e0 pass mlink to set_link_key() bad (freeze >> after ~4 min) >> >> git bisect log and the incident logs for each step are available if >> useful - happy to attach them to the bugzilla ticket or send them >> here. >> >> Let me know if you want anything else tested. >> >> Thanks, >> Jonas >> >> Am 22.08.26 um 22:18 schrieb Jonas Hort: >>> Quick update: v7.0 vanilla has been running clean for over 2 hours now >>> with MLO active (5GHz+6GHz, same FritzBox 5690 Pro), no freeze at all. >>> >>> Am 19.08.26 um 11:03 schrieb Jonas Hort: >>>> Thanks for the detailed breakdown. >>>> >>>> I'll build v7.0 vanilla and test it, as suggested. Fair warning >>>> though: I'm on vacation this week, so I'll pick this up next week. >>>> I've also never compiled a kernel before, so it'll likely take me a >>>> bit of trial and error the first time around - please bear with me >>>> if it takes a little longer than expected. >>>> >>>> Will report back once I have results. >>>> >>>> Thanks again, >>>> Jonas >>>> >>>> 19.08.2026 03:18:34 Devin Wittmayer <lucid_duck@justthetip.ca>: >>>> >>>>> Thank you very much, that answers both things. >>>>> >>>>> The ROC tracing is the more useful of the two even though it came back >>>>> negative. Two of the three freezes have no ROC activity in them at >>>>> all, so a >>>>> link switch that never finished cannot be what starts this. The >>>>> middle one >>>>> does have rocabort, mloroc and rocwork in it, but one out of three >>>>> makes >>>>> that look like the exception rather than the pattern. So the area I >>>>> sent you >>>>> looking at is out, and that is worth knowing before you spend nights on >>>>> builds. >>>>> >>>>> One other thing worth saying first. There is a five patch mt76 >>>>> series on the >>>>> list at the moment and two of the patches look like they were >>>>> written for >>>>> exactly this bug. I do not think they were, and it is your own >>>>> numbers that >>>>> show it. The failure 4/5 fixes stops mt76_txq_schedule_list from >>>>> servicing >>>>> the queue, and the one 5/5 fixes stops mt76_txq_send_burst once the >>>>> non-AQL >>>>> count reaches its cap. Both of those keep frames from ever reaching the >>>>> hardware, so if either were your problem head would be sitting still >>>>> alongside tail. Yours does the opposite. Head climbs 390 to 408 >>>>> while tail >>>>> stays at 260, so the frames are getting into the ring and nothing is >>>>> finishing them, which is the far end of the same path. 2/5 is a use >>>>> after >>>>> free when an interface goes away, so it does not fit either. I would >>>>> not >>>>> expect that series to change what you see. >>>>> >>>>> On the bisect I would build v7.0 next. There are 32 mt7925 commits >>>>> between >>>>> 7.0 and 7.1 and 19 of them are one run of work from Sean Wang, >>>>> reworking how >>>>> the driver tracks the per link mlink and WCID for an MLO station. >>>>> That is >>>>> the kind of change that fits a bug only showing up with two links >>>>> up. If >>>>> v7.0 comes back clean, that series is where I would look. If v7.0 is >>>>> already >>>>> broken then it is off the hook and 6.19 becomes the next split. The >>>>> mt76 >>>>> core and mac80211 both moved in the same window, so mt7925 is where >>>>> I would >>>>> look first rather than the only place worth looking. >>>>> >>>>> Devin >>>>> >>>>> Am 17.08.26 um 23:34 schrieb Jonas Hort: >>>>> >>>>>> Quick follow-up: managed to confirm 6.18 as a clean baseline on my >>>>>> own hardware now (not just secondhand from others in the forum >>>>>> thread) - running Linux 6.18.42-1-cachyos-lts with MLO active >>>>>> (5GHz+6GHz, same FritzBox 5690 Pro) for 3 hours straight, no freeze >>>>>> at all. >>>>> Am 17.08.26 um 16:25 schrieb Jonas Hort: >>>>> >>>>>> One correction to how I described this earlier: the connection does >>>>>> NOT reliably self-heal on its own. I have manually intervened every >>>>>> single time to restore connectivity [...] >>>>>> >>>>>> Update on the ROC tracing: three real freezes captured now with the >>>>>> kprobes active (all confirmed via the WFDMA0 tail-frozen signature). >>>>>> >>>>>> - Freeze #1 (15:37): no ROC activity in the trace. >>>>>> - Freeze #2 (15:43): [...] The trace shows several ROC events >>>>>> (rocabort, mloroc, rocwork) clustered together. >>>>>> - Freeze #3 (16:15): no ROC activity again. >>>>>> >>>>>> Uploaded all three logs to the bugzilla ticket if useful: >>>>>> https://bugzilla.kernel.org/show_bug.cgi?id=221884 ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] mt7925: MLO connectivity silently stalls with 6GHz link active 2026-08-24 8:18 ` Jonas Hort @ 2026-08-29 22:13 ` Devin Wittmayer 2026-10-02 17:31 ` Jonas Hort 2026-10-02 18:11 ` Jonas Hort 0 siblings, 2 replies; 21+ messages in thread From: Devin Wittmayer @ 2026-08-29 22:13 UTC (permalink / raw) To: Jonas Hort, Sean Wang, Thorsten Leemhuis Cc: Sean Wang, Felix Fietkau, lorenzo.bianconi83, regressions, linux-wireless, linux-mediatek On Mon, 2026-08-24 at 08:18 +0000, Jonas Hort wrote: > Setup is still up and I can build and test whatever is useful. Your bisect holds here, on a different machine and a different access point, running your June firmware. Three states, each replicated: both commits reverted clean, two runs of six minutes only the later one reverted stalls, two runs neither reverted stalls, three runs Tail welded and head climbing, which is your signature. The tree is 7.2-rc5 with one unrelated ACPI patch of mine, nowhere near this path. db134691924f lands fifteen commits after ff643b81bc38, and reverting it alone still stalls, so the builder rewrite is the one that matters. Reverting the builder on its own is not available either: v7.0's version dereferences a link pointer that the later commit leaves unpublished until after the firmware call, so it would need a check neither version has. A correction to my last message. I pointed you at the link count, since the new code reports one link whenever the update concerns the primary. That undercount is real but it's not the cause. Telling firmware one link always is clean over two six-minute runs. Skipping the undercounting update entirely, so firmware is never told anything wrong, stalls on both runs. Writing the entries in v7.0's order stalls on all three. On the August firmware, stock still stalls, twice, the detector firing 95 and 141 seconds into the run. The band pair matters. Same driver, same August firmware, two runs each: 5 GHz + 6 GHz stalls 2.4 GHz + 5 GHz clean, full duration 2.4 GHz + 6 GHz clean, full duration Not the channel width. Narrowing the 5 GHz link to 40 MHz, so that pair carries the same narrow-plus-wide mix as the clean ones, still stalls on both runs. The sharpest pair there is the two runs where the iperf3 stream never established, leaving only the UDP pressure: 5 plus 6 stalled with the queue peaking at 187, and 2.4 plus 6 stayed clean at 186. The other three clean runs peaked above 1000. Sean, the two things I can't see from out here: what firmware does with a multi-link record naming one link while a second is associated, and why 5 plus 6 should differ from 2.4 plus 6. Jonas, thank you for the bisect. Five steps, every bad verdict confirmed on the queue signature. Devin ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] mt7925: MLO connectivity silently stalls with 6GHz link active 2026-08-29 22:13 ` Devin Wittmayer @ 2026-10-02 17:31 ` Jonas Hort 2026-10-02 18:11 ` Jonas Hort 1 sibling, 0 replies; 21+ messages in thread From: Jonas Hort @ 2026-10-02 17:31 UTC (permalink / raw) To: Devin Wittmayer, Sean Wang, Thorsten Leemhuis Cc: Sean Wang, Felix Fietkau, lorenzo.bianconi83, regressions, linux-wireless, linux-mediatek Hi all, Friendly ping on this one - any news on the mt7925 MLO stall? No pressure, I know everyone's busy. My setup is still in place, so I'm happy to build and test patches or collect further traces if that would help. Nothing has changed here: still reproducible with the 5 GHz + 6 GHz pair. Thanks for all the help so far. Jonas Am 30.08.26 um 00:13 schrieb Devin Wittmayer: > On Mon, 2026-08-24 at 08:18 +0000, Jonas Hort wrote: >> Setup is still up and I can build and test whatever is useful. > Your bisect holds here, on a different machine and a different access > point, running your June firmware. Three states, each replicated: > > both commits reverted clean, two runs of six minutes > only the later one reverted stalls, two runs > neither reverted stalls, three runs > > Tail welded and head climbing, which is your signature. The tree is > 7.2-rc5 with one unrelated ACPI patch of mine, nowhere near this path. > > db134691924f lands fifteen commits after ff643b81bc38, and reverting it > alone still stalls, so the builder rewrite is the one that matters. > Reverting the builder on its own is not available either: v7.0's version > dereferences a link pointer that the later commit leaves unpublished > until after the firmware call, so it would need a check neither version > has. > > A correction to my last message. I pointed you at the link count, since > the new code reports one link whenever the update concerns the primary. > That undercount is real but it's not the cause. Telling firmware one > link always is clean over two six-minute runs. Skipping the > undercounting update entirely, so firmware is never told anything wrong, > stalls on both runs. Writing the entries in v7.0's order stalls on all > three. > > On the August firmware, stock still stalls, twice, the detector firing > 95 and 141 seconds into the run. > > The band pair matters. Same driver, same August firmware, two runs each: > > 5 GHz + 6 GHz stalls > 2.4 GHz + 5 GHz clean, full duration > 2.4 GHz + 6 GHz clean, full duration > > Not the channel width. Narrowing the 5 GHz link to 40 MHz, so that pair > carries the same narrow-plus-wide mix as the clean ones, still stalls on > both runs. > > The sharpest pair there is the two runs where the iperf3 stream never > established, leaving only the UDP pressure: 5 plus 6 stalled with the > queue peaking at 187, and 2.4 plus 6 stayed clean at 186. The other > three clean runs peaked above 1000. > > Sean, the two things I can't see from out here: what firmware does with > a multi-link record naming one link while a second is associated, and > why 5 plus 6 should differ from 2.4 plus 6. > > Jonas, thank you for the bisect. Five steps, every bad verdict confirmed > on the queue signature. > > Devin ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] mt7925: MLO connectivity silently stalls with 6GHz link active 2026-08-29 22:13 ` Devin Wittmayer 2026-10-02 17:31 ` Jonas Hort @ 2026-10-02 18:11 ` Jonas Hort 2026-10-03 18:03 ` Devin Wittmayer 1 sibling, 1 reply; 21+ messages in thread From: Jonas Hort @ 2026-10-02 18:11 UTC (permalink / raw) To: Devin Wittmayer, Sean Wang, Thorsten Leemhuis Cc: Sean Wang, Felix Fietkau, lorenzo.bianconi83, regressions, linux-wireless, linux-mediatek Hi all, I just came across this patch from Andrei Rusu de Castro, posted on 2026-09-02 with "Fixes: ff643b81bc38" (the commit I bisected to): Patch wifi: mt76: mt7925: stabilize STA_REC_MLD link selection https://ratatoskr.run/linux-wireless/2026/09/17497919 Could this be the fix for this regression? Devin, you mentioned the link count undercount is real but likely not the root cause - does this patch change that assessment? I'm happy to test it on my hardware if that helps. Thanks, Jonas Am 30.08.26 um 00:13 schrieb Devin Wittmayer: > On Mon, 2026-08-24 at 08:18 +0000, Jonas Hort wrote: >> Setup is still up and I can build and test whatever is useful. > Your bisect holds here, on a different machine and a different access > point, running your June firmware. Three states, each replicated: > > both commits reverted clean, two runs of six minutes > only the later one reverted stalls, two runs > neither reverted stalls, three runs > > Tail welded and head climbing, which is your signature. The tree is > 7.2-rc5 with one unrelated ACPI patch of mine, nowhere near this path. > > db134691924f lands fifteen commits after ff643b81bc38, and reverting it > alone still stalls, so the builder rewrite is the one that matters. > Reverting the builder on its own is not available either: v7.0's version > dereferences a link pointer that the later commit leaves unpublished > until after the firmware call, so it would need a check neither version > has. > > A correction to my last message. I pointed you at the link count, since > the new code reports one link whenever the update concerns the primary. > That undercount is real but it's not the cause. Telling firmware one > link always is clean over two six-minute runs. Skipping the > undercounting update entirely, so firmware is never told anything wrong, > stalls on both runs. Writing the entries in v7.0's order stalls on all > three. > > On the August firmware, stock still stalls, twice, the detector firing > 95 and 141 seconds into the run. > > The band pair matters. Same driver, same August firmware, two runs each: > > 5 GHz + 6 GHz stalls > 2.4 GHz + 5 GHz clean, full duration > 2.4 GHz + 6 GHz clean, full duration > > Not the channel width. Narrowing the 5 GHz link to 40 MHz, so that pair > carries the same narrow-plus-wide mix as the clean ones, still stalls on > both runs. > > The sharpest pair there is the two runs where the iperf3 stream never > established, leaving only the UDP pressure: 5 plus 6 stalled with the > queue peaking at 187, and 2.4 plus 6 stayed clean at 186. The other > three clean runs peaked above 1000. > > Sean, the two things I can't see from out here: what firmware does with > a multi-link record naming one link while a second is associated, and > why 5 plus 6 should differ from 2.4 plus 6. > > Jonas, thank you for the bisect. Five steps, every bad verdict confirmed > on the queue signature. > > Devin ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] mt7925: MLO connectivity silently stalls with 6GHz link active 2026-10-02 18:11 ` Jonas Hort @ 2026-10-03 18:03 ` Devin Wittmayer 2026-10-04 15:48 ` Andrei Rusu de Castro 0 siblings, 1 reply; 21+ messages in thread From: Devin Wittmayer @ 2026-10-03 18:03 UTC (permalink / raw) To: Jonas Hort, Sean Wang, Thorsten Leemhuis Cc: Sean Wang, Felix Fietkau, lorenzo.bianconi83, Andrei Rusu de Castro, regressions, linux-wireless, linux-mediatek On Fri, 2026-10-02 at 18:11 +0000, Jonas Hort wrote: > I just came across this patch from Andrei Rusu de Castro, posted on > 2026-09-02 with "Fixes: ff643b81bc38" (the commit I bisected to): > > Patch wifi: mt76: mt7925: stabilize STA_REC_MLD link selection > > Could this be the fix for this regression? Devin, you mentioned the > link count undercount is real but likely not the root cause - does > this patch change that assessment? The patch you found does not fix this stall. Andrei never claimed it would and it did no harm here, so this is not an argument against it. bench A 2 Sep MT7925 PCIe, 7.2.0-rc5 + lockdep, morrownr @ 6b0ef22e bench B 3 Oct MT7925U USB, 7.2.6, morrownr @ 8e9309b both firmware 20260813113118, DevLabMLD on 5745 and 6135 MHz, iw scan every 4 s under bulk upload. Neither is in-tree. run bench last worked at 600 s stock A 120 s dead + the patch A 141 s dead stock B 122 s dead + the patch, run 1 B 123 s dead + the patch, run 2 B 123 s dead primary link only B 583 s alive Every two-link run died between 120 and 141 seconds, patched or not, and none recovered. Each stayed associated on both links, reporting a healthy rate while the AP heard nothing from the client. The patch describes both links on every update. In August I got to that same record another way, dropping the undercounting update, and it stalled then too. The arm that stays up describes the primary link only. Not a fix, it hides the second link, but the stall tracks whether that second link is in the record. Thanks for your patience, and for the bisect and for finding the patch yourself. Devin ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] mt7925: MLO connectivity silently stalls with 6GHz link active 2026-10-03 18:03 ` Devin Wittmayer @ 2026-10-04 15:48 ` Andrei Rusu de Castro 2026-10-04 15:49 ` [PATCH v2] wifi: mt76: mt7925: keep MLD membership consistent during link add Andrei Rusu de Castro 0 siblings, 1 reply; 21+ messages in thread From: Andrei Rusu de Castro @ 2026-10-04 15:48 UTC (permalink / raw) To: linux-wireless, Devin Wittmayer, Jonas Hort Cc: Felix Fietkau, Lorenzo Bianconi, Ryder Lee, Shayne Chen, Sean Wang, Sean Wang, Thorsten Leemhuis, regressions, linux-mediatek Hi Devin, Jonas, I reproduced the stall with my September patch on an MT7925 PCIe Z13 and traced the station commands sent to firmware. The missing case was the primary update inside secondary link addition. mt7925_mac_link_sta_add() updates the primary WCID first, then the new secondary WCID. The new secondary is not in msta->link[] or valid_links until both commands succeed. My first patch included an unpublished link only when that link was the subject of the current command. It therefore still sent a one-link MLD description to the primary WCID, followed by a two-link description to the secondary WCID. These are the captured fields during secondary addition, omitting an earlier primary-only association update: command WCID count entries (WCID/BSS) September patch 1 1 1/0 2 2 1/0, 2/1 September patch + early 1 2 1/0, 2/1 link publication 2 2 1/0, 2/1 v2 1 2 1/0, 2/1 2 2 1/0, 2/1 V2 passes the initialized pending secondary explicitly through both station updates, without publishing it early. It keeps the existing publication-after-success and host error cleanup. It also keeps stable membership for subsequent updates. There is no new shared pending state. On Linux 7.3-rc5 with firmware 20260813113118 and an ASUS GT-BE98: upstream MLD path stalled at 157 seconds September patch stalled at 289 seconds; repeated at 150 v2 813 seconds, 48 samples, no failures v2, second clean boot 812 seconds, no sustained stall under 20 Mbit/s traffic in each direction The AP advertises three links. The active client pair was 5+6 GHz (active_links=0x3, valid_links=0x7). Failed runs had roughly 27-second full-band scans; both v2 runs returned to roughly 6.3-7.1 seconds. The second v2 run transferred about 1.9 GB in each direction over 760 seconds. It had one gateway-ping timeout after a neighbor flush, with REACHABLE neighbor state and successful IP/HTTPS probes in the same sample, so it did not pass the strict zero-failure gate. Traffic did not enter the persistent associated-but-dead state. Tests used the in-tree driver plus the selected fix, with the separate scheduled-scan withdrawal and WM2 reset correction held constant. Other local platform patches did not change between these arms. My production variant additionally retains the applicable local safety guards in a separate patch; those guards are not part of this submission. That variant passed the full zero-failure gate on two PCIe Z13 machines (813 and 810 seconds), and both booted it successfully as their normal default. Five ordinary reconnects also passed on the first local variant. The changed objects build with W=1 and -Werror on both the tested rc5 source and the current mt76 integration branch. Source-extracted fixtures exercise the pending-link records and host cleanup at eight failing MCU command positions. Those tests do not establish firmware rollback after a partly successful add. Direct debugfs link switches and chip-reset tests exposed failures or aggregation teardown warnings that also occur with my previous working full-revert kernel. V2 is not claimed to resolve those. USB hardware and your exact FritzBox setup have not been tested here. The v2 posted in reply to this message is based on mt76 commit 0dbc9c9fa9b9. Could you try it on the setups where the September patch still stalled? Thanks, Andrei ^ permalink raw reply [flat|nested] 21+ messages in thread
* [PATCH v2] wifi: mt76: mt7925: keep MLD membership consistent during link add 2026-10-04 15:48 ` Andrei Rusu de Castro @ 2026-10-04 15:49 ` Andrei Rusu de Castro 2026-10-04 18:25 ` Jonas Hort 2026-10-05 2:44 ` Devin Wittmayer 0 siblings, 2 replies; 21+ messages in thread From: Andrei Rusu de Castro @ 2026-10-04 15:49 UTC (permalink / raw) To: linux-wireless, Devin Wittmayer, Jonas Hort Cc: Felix Fietkau, Lorenzo Bianconi, Ryder Lee, Shayne Chen, Sean Wang, Sean Wang, Thorsten Leemhuis, regressions, linux-mediatek mt7925_mac_link_sta_add() sends an ASSOC update for the primary WCID before sending the update for a new secondary WCID. The new link is published in msta->link[] and msta->valid_links only after both commands succeed. Selecting the secondary entry from the subject of each command gives the primary STA_REC_MLD one link and the secondary STA_REC_MLD two links. Enumerating published links alone still omits the pending link from the primary command. Firmware-bound command captures reproduce this n=1/n=2 sequence during a secondary addition. Build each MLD TLV from the station's published links and an explicit pending link. Pass the initialized pending link through both add-time station updates, including the update whose subject is the primary. Keep the primary first, bound entries by the firmware array, and skip links without station and BSS state. Other update callers have no pending link. This leaves msta->link[] publication after successful link setup and preserves the existing add-failure cleanup. The pending pointer is used synchronously to populate the command, not stored in shared state. Fixes: ff643b81bc38 ("wifi: mt76: mt7925: pass mlink and mconf to sta_mld_tlv()") Link: https://lore.kernel.org/linux-wireless/066b30cc-a9e6-4aeb-964d-71551e8ea3ef@posteo.de/ Signed-off-by: Andrei Rusu de Castro <arc@empyreal.works> --- Changes since v1: - Carry an explicit initialized pending link through both add-time updates, not just the update whose subject is the pending secondary. - Preserve msta->link[] publication after successful setup. - Describe consistency between per-WCID commands without assuming a shared firmware record overwritten by the last command. V1: https://lore.kernel.org/all/20260902-mt7925-0cbea623@empyreal.works/ Tested on MT7925 PCIe, Linux 7.3-rc5, firmware 20260813113118, ASUS GT-BE98 advertising three links with a 5+6 GHz active pair. The in-tree MLD path stalled after 157 seconds, v1 after 289 and 150 seconds. V2 passed a 48-sample, 813-second scan/traffic run; full-band scans returned to 6.3-7.1 seconds from about 27 seconds. A second clean boot did not stall through 812 seconds and 760 seconds of bilateral 20 Mbit/s traffic. One gateway ping timed out while same-sample IP/HTTPS passed, so that second run did not pass the strict zero-failure gate. The local deployment variant, with separate retained safety guards, passed the same gate on two PCIe Z13 machines (813 and 810 seconds). Both then passed ordinary default boots. The scheduled-scan withdrawal and independent WM2 reset correction were retained throughout testing. USB hardware and the reporter's FritzBox setup remain untested here. W=1/-Werror builds of changed objects pass on both rc5 and the mt76 base below. Source-extracted fixtures under ASan/UBSan cover pending-link encoding and host cleanup at eight failed BSS/STA command positions; they do not model firmware rollback. Direct partial-link switches and reset aggregation warnings also fail on the old full-revert baseline and are not claimed fixed by this patch. .../net/wireless/mediatek/mt76/mt7925/mac.c | 2 +- .../net/wireless/mediatek/mt76/mt7925/main.c | 16 ++--- .../net/wireless/mediatek/mt76/mt7925/mcu.c | 59 +++++++++++++++---- .../wireless/mediatek/mt76/mt7925/mt7925.h | 3 +- 4 files changed, 60 insertions(+), 20 deletions(-) diff --git a/drivers/net/wireless/mediatek/mt76/mt7925/mac.c b/drivers/net/wireless/mediatek/mt76/mt7925/mac.c index 101f571b027f..cbc18dbcbad9 100644 --- a/drivers/net/wireless/mediatek/mt76/mt7925/mac.c +++ b/drivers/net/wireless/mediatek/mt76/mt7925/mac.c @@ -1502,7 +1502,7 @@ mt7925_vif_connect_iter(void *priv, u8 *mac, true, NULL); mt7925_mcu_sta_update(dev, NULL, vif, &mvif->sta.deflink, true, - MT76_STA_INFO_STATE_NONE); + MT76_STA_INFO_STATE_NONE, NULL); mt7925_mcu_uni_add_beacon_offload(dev, hw, vif, true); } } diff --git a/drivers/net/wireless/mediatek/mt76/mt7925/main.c b/drivers/net/wireless/mediatek/mt76/mt7925/main.c index c882952f5df1..f536fffffa08 100644 --- a/drivers/net/wireless/mediatek/mt76/mt7925/main.c +++ b/drivers/net/wireless/mediatek/mt76/mt7925/main.c @@ -1006,7 +1006,7 @@ static int mt7925_mac_link_sta_add(struct mt76_dev *mdev, link_sta == mlink->pri_link) { ret = mt7925_mcu_sta_update(dev, link_sta, vif, mlink, true, - MT76_STA_INFO_STATE_NONE); + MT76_STA_INFO_STATE_NONE, NULL); if (ret) goto out_pm; } else if (ieee80211_vif_is_mld(vif) && @@ -1028,19 +1028,19 @@ static int mt7925_mac_link_sta_add(struct mt76_dev *mdev, ret = mt7925_mcu_sta_update(dev, mlink->pri_link, vif, pri_mlink, true, - MT76_STA_INFO_STATE_ASSOC); + MT76_STA_INFO_STATE_ASSOC, mlink); if (ret) goto out_pm; ret = mt7925_mcu_sta_update(dev, link_sta, vif, mlink, true, - MT76_STA_INFO_STATE_ASSOC); + MT76_STA_INFO_STATE_ASSOC, mlink); if (ret) goto out_pm; } else { ret = mt7925_mcu_sta_update(dev, link_sta, vif, mlink, true, - MT76_STA_INFO_STATE_NONE); + MT76_STA_INFO_STATE_NONE, NULL); if (ret) goto out_pm; } @@ -1248,7 +1248,7 @@ static void mt7925_mac_link_sta_assoc(struct mt76_dev *mdev, memset(mlink->airtime_ac, 0, sizeof(mlink->airtime_ac)); mt7925_mcu_sta_update(dev, link_sta, vif, mlink, true, - MT76_STA_INFO_STATE_ASSOC); + MT76_STA_INFO_STATE_ASSOC, NULL); mt792x_mutex_release(dev); } @@ -1308,7 +1308,7 @@ static void mt7925_mac_link_sta_remove(struct mt76_dev *mdev, mt76_connac_pm_wake(&dev->mphy, &dev->pm); mt7925_mcu_sta_update(dev, link_sta, vif, mlink, false, - MT76_STA_INFO_STATE_NONE); + MT76_STA_INFO_STATE_NONE, NULL); mt7925_mac_wtbl_update(dev, mlink->wcid.idx, MT_WTBL_UPDATE_ADM_COUNT_CLEAR); @@ -1979,7 +1979,7 @@ mt7925_start_ap(struct ieee80211_hw *hw, struct ieee80211_vif *vif, err = mt7925_mcu_sta_update(dev, NULL, vif, &mvif->sta.deflink, true, - MT76_STA_INFO_STATE_NONE); + MT76_STA_INFO_STATE_NONE, NULL); out: mt792x_mutex_release(dev); @@ -2123,7 +2123,7 @@ static void mt7925_vif_cfg_changed(struct ieee80211_hw *hw, if (changed & BSS_CHANGED_ASSOC) { mt7925_mcu_sta_update(dev, NULL, vif, &mvif->sta.deflink, true, - MT76_STA_INFO_STATE_ASSOC); + MT76_STA_INFO_STATE_ASSOC, NULL); mt7925_mcu_set_beacon_filter(dev, vif, vif->cfg.assoc); if (ieee80211_vif_is_mld(vif)) diff --git a/drivers/net/wireless/mediatek/mt76/mt7925/mcu.c b/drivers/net/wireless/mediatek/mt76/mt7925/mcu.c index 2afd3f5e3266..634062730e65 100644 --- a/drivers/net/wireless/mediatek/mt76/mt7925/mcu.c +++ b/drivers/net/wireless/mediatek/mt76/mt7925/mcu.c @@ -2065,14 +2065,18 @@ mt7925_mcu_sta_mld_tlv(struct sk_buff *skb, struct ieee80211_vif *vif, struct ieee80211_sta *sta, struct mt792x_bss_conf *mconf, - struct mt792x_link_sta *mlink) + struct mt792x_link_sta *mlink, + struct mt792x_link_sta *pending) { struct mt792x_vif *mvif = (struct mt792x_vif *)vif->drv_priv; struct mt792x_sta *msta = (struct mt792x_sta *)sta->drv_priv; struct mt792x_dev *dev = mvif->phy->dev; + unsigned long valid = msta->valid_links; struct mt792x_bss_conf *mconf_pri; struct sta_rec_mld *mld; + unsigned int link_id; struct tlv *tlv; + u8 max_links; u8 cnt = 0; /* Primary link always uses driver's deflink WCID. */ @@ -2101,11 +2105,44 @@ mt7925_mcu_sta_mld_tlv(struct sk_buff *skb, mld->link[cnt].wlan_id = cpu_to_le16(msta->deflink.wcid.idx); mld->link[cnt++].bss_idx = mconf_pri->mt76.idx; - /* Optionally encode the currently-updated secondary link. */ - if (mlink && mlink != &msta->deflink && mconf) { - mld->secondary_id = cpu_to_le16(mlink->wcid.idx); - mld->link[cnt].wlan_id = cpu_to_le16(mlink->wcid.idx); - mld->link[cnt++].bss_idx = mconf->mt76.idx; + /* Describe the same station links in each per-WCID STA_REC_MLD, + * rather than selecting the secondary from the current command. + */ + max_links = ARRAY_SIZE(mld->link); + + /* Adding a secondary link updates both primary and secondary STA + * records before publishing the new link in msta->link[]. Include it + * in both commands without moving that publication before success. + */ + if (pending && pending != &msta->deflink) + valid |= BIT(pending->wcid.link_id); + + for_each_set_bit(link_id, &valid, IEEE80211_MLD_MAX_NUM_LINKS) { + struct mt792x_link_sta *mlink_sec; + struct mt792x_bss_conf *mconf_sec; + + if (cnt == max_links) + break; + + if (link_id == msta->deflink_id) + continue; + + mlink_sec = mt792x_sta_to_link(msta, link_id); + if (!mlink_sec && pending && link_id == pending->wcid.link_id) + mlink_sec = pending; + if (!mlink_sec || mlink_sec == &msta->deflink) + continue; + + mconf_sec = rcu_dereference_protected(mvif->link_conf[link_id], + lockdep_is_held(&dev->mt76.mutex)); + if (!mconf_sec) + continue; + + if (cnt == 1) + mld->secondary_id = cpu_to_le16(mlink_sec->wcid.idx); + + mld->link[cnt].wlan_id = cpu_to_le16(mlink_sec->wcid.idx); + mld->link[cnt++].bss_idx = mconf_sec->mt76.idx; } mld->link_num = cnt; @@ -2124,7 +2161,8 @@ mt7925_mcu_sta_remove_tlv(struct sk_buff *skb) static int mt7925_mcu_sta_cmd(struct mt76_phy *phy, - struct mt76_sta_cmd_info *info) + struct mt76_sta_cmd_info *info, + struct mt792x_link_sta *pending) { struct mt792x_vif *mvif = (struct mt792x_vif *)info->vif->drv_priv; struct mt76_dev *dev = phy->dev; @@ -2165,7 +2203,7 @@ mt7925_mcu_sta_cmd(struct mt76_phy *phy, if (info->state != MT76_STA_INFO_STATE_NONE) { mt7925_mcu_sta_mld_tlv(skb, info->vif, info->link_sta->sta, - mconf, mlink); + mconf, mlink, pending); mt7925_mcu_sta_eht_mld_tlv(skb, info->vif, info->link_sta->sta); } @@ -2190,7 +2228,8 @@ int mt7925_mcu_sta_update(struct mt792x_dev *dev, struct ieee80211_vif *vif, struct mt792x_link_sta *mlink, bool enable, - enum mt76_sta_info_state state) + enum mt76_sta_info_state state, + struct mt792x_link_sta *pending) { struct mt792x_vif *mvif = (struct mt792x_vif *)vif->drv_priv; int rssi = -ewma_rssi_read(&mvif->bss_conf.rssi); @@ -2208,7 +2247,7 @@ int mt7925_mcu_sta_update(struct mt792x_dev *dev, info.wcid = &mlink->wcid; info.newly = state != MT76_STA_INFO_STATE_ASSOC; - return mt7925_mcu_sta_cmd(&dev->mphy, &info); + return mt7925_mcu_sta_cmd(&dev->mphy, &info, pending); } int mt7925_mcu_set_beacon_filter(struct mt792x_dev *dev, diff --git a/drivers/net/wireless/mediatek/mt76/mt7925/mt7925.h b/drivers/net/wireless/mediatek/mt76/mt7925/mt7925.h index 33782d9ba9ed..49fa94e5cfd9 100644 --- a/drivers/net/wireless/mediatek/mt76/mt7925/mt7925.h +++ b/drivers/net/wireless/mediatek/mt76/mt7925/mt7925.h @@ -286,7 +286,8 @@ int mt7925_mcu_sta_update(struct mt792x_dev *dev, struct ieee80211_vif *vif, struct mt792x_link_sta *mlink, bool enable, - enum mt76_sta_info_state state); + enum mt76_sta_info_state state, + struct mt792x_link_sta *pending); int mt7925_mcu_set_chan_info(struct mt792x_phy *phy, u16 tag); int mt7925_mcu_set_tx(struct mt792x_dev *dev, struct ieee80211_bss_conf *bss_conf); int mt7925_mcu_set_eeprom(struct mt792x_dev *dev); base-commit: 0dbc9c9fa9b92767c2d556504f38b544cb57a96a -- 2.54.0 ^ permalink raw reply related [flat|nested] 21+ messages in thread
* Re: [PATCH v2] wifi: mt76: mt7925: keep MLD membership consistent during link add 2026-10-04 15:49 ` [PATCH v2] wifi: mt76: mt7925: keep MLD membership consistent during link add Andrei Rusu de Castro @ 2026-10-04 18:25 ` Jonas Hort 2026-10-05 2:44 ` Devin Wittmayer 1 sibling, 0 replies; 21+ messages in thread From: Jonas Hort @ 2026-10-04 18:25 UTC (permalink / raw) To: Andrei Rusu de Castro, linux-wireless, Devin Wittmayer Cc: Felix Fietkau, Lorenzo Bianconi, Ryder Lee, Shayne Chen, Sean Wang, Sean Wang, Thorsten Leemhuis, regressions, linux-mediatek Hi Andrei, Tested v2 on top of v7.3-rc5 on my setup: hardware MT7925 PCIe (mt7925e) firmware 20260813113118 AP FritzBox 5690 Pro, FRITZ!OS 8.25 links 5GHz + 6GHz MLO (5200 + 6135 MHz) stock 7.1.8 / 7.2-rc7 stalled after 3-5 minutes (WFDMA0 tail frozen) v7.3-rc5 + v2 46 minutes, no stall Both links stayed active throughout, with about 4.2 GB received and 340 MB sent during the run. My watchdog (ping + WFDMA0 queue monitoring) did not trigger once. Tested-by: Jonas Hort <jonas.hort@posteo.de> Thanks for the fix! Jonas ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [PATCH v2] wifi: mt76: mt7925: keep MLD membership consistent during link add 2026-10-04 15:49 ` [PATCH v2] wifi: mt76: mt7925: keep MLD membership consistent during link add Andrei Rusu de Castro 2026-10-04 18:25 ` Jonas Hort @ 2026-10-05 2:44 ` Devin Wittmayer 1 sibling, 0 replies; 21+ messages in thread From: Devin Wittmayer @ 2026-10-05 2:44 UTC (permalink / raw) To: Andrei Rusu de Castro Cc: linux-wireless, Jonas Hort, Felix Fietkau, Lorenzo Bianconi, Ryder Lee, Shayne Chen, Sean Wang, Sean Wang, Thorsten Leemhuis, regressions, linux-mediatek Andrei Rusu de Castro wrote: > USB hardware and the reporter's FritzBox setup remain untested here. Tested on a Netgear A9000 (MT7925U, USB), Linux 7.2.6, firmware 20260813113118, morrownr/mt76 at ce4ee3e, against a two-link mt7996 AP MLD on 5745 and 6135 MHz. Your STA_REC_MLD capture reproduces here. Without v2 the AP got nothing on 5745 in any run I checked. With v2, 5745 carried traffic under the scan load from my earlier report: stock v2 scan load, 600 s, mean per ~10 s 249, 267 MB 801, 837 MB scan load, 600 s, lowest ~10 s 126 B, 33 MB 643, 687 MB scan load, share of bytes on 5745 0 about 60% scanning off, share on 5745 0 0 Stock dipped under that load but did not stall for good this time, so the stall result is Jonas's run, not mine. Tested-by: Devin Wittmayer <lucid_duck@justthetip.ca> Thanks, Devin ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] mt7925: MLO connectivity silently stalls with 6GHz link active
@ 2026-08-19 9:03 Jonas Hort
0 siblings, 0 replies; 21+ messages in thread
From: Jonas Hort @ 2026-08-19 9:03 UTC (permalink / raw)
To: Devin Wittmayer
Cc: Thorsten Leemhuis, regressions, linux-wireless, Felix Fietkau,
lorenzo.bianconi83
Thanks for the detailed breakdown.
I'll build v7.0 vanilla and test it, as suggested. Fair warning
though: I'm on vacation this week, so I'll pick this up next week.
I've also never compiled a kernel before, so it'll likely take me a
bit of trial and error the first time around - please bear with me
if it takes a little longer than expected.
Will report back once I have results.
Thanks again,
Jonas
19.08.2026 03:18:34 Devin Wittmayer <lucid_duck@justthetip.ca>:
> Thank you very much, that answers both things.
>
> The ROC tracing is the more useful of the two even though it came back
> negative. Two of the three freezes have no ROC activity in them at all, so a
> link switch that never finished cannot be what starts this. The middle one
> does have rocabort, mloroc and rocwork in it, but one out of three makes
> that look like the exception rather than the pattern. So the area I sent you
> looking at is out, and that is worth knowing before you spend nights on
> builds.
>
> One other thing worth saying first. There is a five patch mt76 series on the
> list at the moment and two of the patches look like they were written for
> exactly this bug. I do not think they were, and it is your own numbers that
> show it. The failure 4/5 fixes stops mt76_txq_schedule_list from servicing
> the queue, and the one 5/5 fixes stops mt76_txq_send_burst once the non-AQL
> count reaches its cap. Both of those keep frames from ever reaching the
> hardware, so if either were your problem head would be sitting still
> alongside tail. Yours does the opposite. Head climbs 390 to 408 while tail
> stays at 260, so the frames are getting into the ring and nothing is
> finishing them, which is the far end of the same path. 2/5 is a use after
> free when an interface goes away, so it does not fit either. I would not
> expect that series to change what you see.
>
> On the bisect I would build v7.0 next. There are 32 mt7925 commits between
> 7.0 and 7.1 and 19 of them are one run of work from Sean Wang, reworking how
> the driver tracks the per link mlink and WCID for an MLO station. That is
> the kind of change that fits a bug only showing up with two links up. If
> v7.0 comes back clean, that series is where I would look. If v7.0 is already
> broken then it is off the hook and 6.19 becomes the next split. The mt76
> core and mac80211 both moved in the same window, so mt7925 is where I would
> look first rather than the only place worth looking.
>
> Devin
>
> Am 17.08.26 um 23:34 schrieb Jonas Hort:
>
>> Quick follow-up: managed to confirm 6.18 as a clean baseline on my
>> own hardware now (not just secondhand from others in the forum
>> thread) - running Linux 6.18.42-1-cachyos-lts with MLO active
>> (5GHz+6GHz, same FritzBox 5690 Pro) for 3 hours straight, no freeze
>> at all.
>
> Am 17.08.26 um 16:25 schrieb Jonas Hort:
>
>> One correction to how I described this earlier: the connection does
>> NOT reliably self-heal on its own. I have manually intervened every
>> single time to restore connectivity [...]
>>
>> Update on the ROC tracing: three real freezes captured now with the
>> kprobes active (all confirmed via the WFDMA0 tail-frozen signature).
>>
>> - Freeze #1 (15:37): no ROC activity in the trace.
>> - Freeze #2 (15:43): [...] The trace shows several ROC events
>> (rocabort, mloroc, rocwork) clustered together.
>> - Freeze #3 (16:15): no ROC activity again.
>>
>> Uploaded all three logs to the bugzilla ticket if useful:
>> https://bugzilla.kernel.org/show_bug.cgi?id=221884
^ permalink raw reply [flat|nested] 21+ messages in thread* [REGRESSION] mt7925: MLO connectivity silently stalls with 6GHz link active @ 2026-08-14 14:04 Jonas Hort 2026-08-15 5:51 ` Thorsten Leemhuis 0 siblings, 1 reply; 21+ messages in thread From: Jonas Hort @ 2026-08-14 14:04 UTC (permalink / raw) To: regressions; +Cc: linux-wireless, lorenzo.bianconi83 Filed as: https://bugzilla.kernel.org/show_bug.cgi?id=221884 Summary: WiFi 7 MLO (5GHz+6GHz) on MT7925 silently stops passing traffic after a few minutes, while the driver continues to report a fully healthy link. Root cause appears to be a stalled WFDMA0 TX hardware queue. Reproduced on both linux 7.1.8 (stable) and 7.2.0-rc7 (mainline pre-release), not fixed as of rc7. Disabling the 6GHz radio on the access point eliminates the freeze entirely in my testing. This is a regression: reported by multiple users on the CachyOS forum as broken starting with the 7.1.x kernel line; earlier kernels (6.18.x longterm) are reported to work fine. I have not yet performed a commit-level bisection but can if guided on the best approach given the AP-side 6GHz dependency. --- Detailed description --- Symptom: When connected via MLO (Multi-Link Operation, EHT/WiFi 7) with two active links - one on 5GHz and one on 6GHz - network traffic silently stops after a few minutes of normal use (web browsing, etc). The connection continues to show as healthy: - `iw link` / `iw station dump` report the connection as associated, authenticated, good signal (-60 to -65 dBm), high negotiated bitrate (800-1700 Mbit/s), 0 tx retries, 0 tx failed. - Nothing is logged to dmesg/journalctl at the time of the freeze. - Ping to the gateway shows 100% packet loss for several minutes, until the connection eventually recovers on its own. I built a small watchdog script (ping-based failure detection + live kernel log monitoring) to catch the freeze automatically and sample mt76 debugfs state repeatedly across the event. This revealed the actual mechanism: `/sys/kernel/debug/ieee80211/phy0/mt76/xmit-queues` shows the WFDMA0 hardware TX queue's `tail` pointer completely frozen while `head` keeps advancing - i.e. packets keep getting enqueued but the firmware stops draining the queue. RX stops in the same instant (rx byte/packet counters in `iw station dump` freeze completely). No tx retries or tx failures are ever reported by the driver, so it does not appear to notice the stall itself. Example from one incident (kernel 7.2.0-rc7): t+1s: WFDMA0: queued=130 head=390 tail=260 rx bytes=184791894 t+9s: WFDMA0: queued=148 head=408 tail=260 rx bytes=184791894 (unchanged) `tail` never advances during the whole stall window while `head` keeps growing - the queue is being filled but never drained. Reproduced this exact signature four times total (twice on 7.1.8-1, twice on 7.2.0-rc7-1), always with MLO 5GHz+6GHz active, always within 3-5 minutes of normal use. 6GHz correlation: My MLO links are 5GHz (5180 MHz) + 6GHz (varies, e.g. 6135 MHz). Fully disabling the 6GHz radio on the access point (FritzBox 5690 Pro, FRITZ!OS 8.25 - not just client-side band restriction via NetworkManager, which does not reliably suppress the second MLO link) results in 30+ minutes of clean operation with no freeze. Re-enabling 6GHz reproduces the freeze again within minutes. Regression info: Multiple users on the CachyOS forum report this started with the 7.1.x kernel line; the previous longterm kernel (6.18.x) is reported to work fine by another user with different affected hardware (different AP). I have not personally tested 6.18.x myself, and have not performed a commit-level bisection yet - happy to do so if pointed toward the most likely area of the driver, given the AP-side 6GHz dependency makes a fully automated bisection awkward. Steps to reproduce: 1. Connect to an AP advertising WiFi 7 MLO with 5GHz+6GHz link options, so the client negotiates both links. 2. Use the connection normally (web browsing is sufficient). 3. Within roughly 3-5 minutes, new connections start hanging; existing traffic stops. 4. Check `iw link` - link still reports as connected/healthy. 5. Check `/sys/kernel/debug/ieee80211/phy*/mt76/xmit-queues` - WFDMA0 tail pointer frozen while head continues to advance. Environment: - Kernel: linux-cachyos 7.1.8-1 and linux-cachyos-rc 7.2.0-rc7-1 (CachyOS, near-vanilla Arch-based build) - cat /proc/version: Linux version 7.2.0-rc7-1-cachyos-rc (linux-cachyos-rc@cachyos) (clang version 22.1.8, LLD 22.1.8) #1 SMP PREEMPT_DYNAMIC Mon, 10 Aug 2026 20:51:00 +0000 - Distribution: CachyOS - Architecture: x86_64 - Kernel tainted: no (/proc/sys/kernel/tainted = 0) - WiFi hardware: MediaTek MT7925 (mt7925e driver, mt76 core), WM Firmware Build 20260605184805, ASIC revision 79250000 - Access point: AVM FritzBox 5690 Pro, FRITZ!OS 8.25 Related community discussion: https://discuss.cachyos.org/t/mt7925-wifi-7-mlo-breaks-connectivity-on-7-1-x-mt76-regression/31972 (I am the thread starter; several other users report matching symptoms with different hardware/APs) Possibly related (different symptom, same 6GHz/MT7925 area, worth checking for a common root cause): Bug 221627 - MT7925E - System locks up with flashing caps lock (comment 1 there: "caused by switching between AP's with the same 6GHz wifi name") Attachments (will add to this ticket): - Full incident logs for all four reproductions (dmesg + journalctl + mt76 debugfs time series across each freeze) - iw station dump / iw link output during freeze - Kernel .config Note: checked bug 220424 (mlo_sta_cmd/sta_cmd regression) before filing - different symptom (persistent near-zero throughput vs. our intermittent full stall with healthy link stats) and different root cause (broadcast wtbl hdr_trans_tlv handling). That fix landed in 6.16.6/6.17-rc5, well before the kernel versions tested here, so it is not the cause of this issue. ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] mt7925: MLO connectivity silently stalls with 6GHz link active 2026-08-14 14:04 Jonas Hort @ 2026-08-15 5:51 ` Thorsten Leemhuis 2026-08-16 20:48 ` Devin Wittmayer 0 siblings, 1 reply; 21+ messages in thread From: Thorsten Leemhuis @ 2026-08-15 5:51 UTC (permalink / raw) To: Jonas Hort, regressions Cc: linux-wireless, lorenzo.bianconi83, Devin Wittmayer On 8/14/26 16:04, Jonas Hort wrote: > Filed as: https://bugzilla.kernel.org/show_bug.cgi?id=221884 > > Summary: WiFi 7 MLO (5GHz+6GHz) on MT7925 silently stops passing traffic > after a few minutes, while the driver continues to report a fully > healthy link. Root cause appears to be a stalled WFDMA0 TX hardware > queue. Reproduced on both linux 7.1.8 (stable) and 7.2.0-rc7 (mainline > pre-release), not fixed as of rc7. Disabling the 6GHz radio on the > access point eliminates the freeze entirely in my testing. Not my area of expertise, but there was one patch was reverted on Thursday that is somewhat related (https://git.kernel.org/torvalds/c/3aa1dcaa4f6f5ae08936491e08bd456f331f2d40 ), but that was more about shutdown/module unload aiui and likely something else. But there were other reports that sounds somewhat related (but careful, I might send you on the wrong track here; CCing Devin, who wrote two of the following messages): https://lore.kernel.org/linux-wireless/20260629083543.153564-1-jb.tsai@mediatek.com/ https://lore.kernel.org/linux-wireless/20260804185118.19705-1-lucid_duck@justthetip.ca/ https://lore.kernel.org/all/20260812155811.10950-1-lucid_duck@justthetip.ca/ Ciao, Thorsten > This is a regression: reported by multiple users on the CachyOS forum > as broken starting with the 7.1.x kernel line; earlier kernels > (6.18.x longterm) are reported to work fine. I have not yet performed > a commit-level bisection but can if guided on the best approach given > the AP-side 6GHz dependency. > > --- Detailed description --- > > Symptom: > When connected via MLO (Multi-Link Operation, EHT/WiFi 7) with two > active links - one on 5GHz and one on 6GHz - network traffic silently > stops after a few minutes of normal use (web browsing, etc). The > connection continues to show as healthy: > > - `iw link` / `iw station dump` report the connection as associated, > authenticated, good signal (-60 to -65 dBm), high negotiated > bitrate (800-1700 Mbit/s), 0 tx retries, 0 tx failed. > - Nothing is logged to dmesg/journalctl at the time of the freeze. > - Ping to the gateway shows 100% packet loss for several minutes, > until the connection eventually recovers on its own. > > I built a small watchdog script (ping-based failure detection + live > kernel log monitoring) to catch the freeze automatically and sample > mt76 debugfs state repeatedly across the event. This revealed the > actual mechanism: > > `/sys/kernel/debug/ieee80211/phy0/mt76/xmit-queues` shows the WFDMA0 > hardware TX queue's `tail` pointer completely frozen while `head` > keeps advancing - i.e. packets keep getting enqueued but the firmware > stops draining the queue. RX stops in the same instant (rx byte/packet > counters in `iw station dump` freeze completely). No tx retries or tx > failures are ever reported by the driver, so it does not appear to > notice the stall itself. > > Example from one incident (kernel 7.2.0-rc7): > t+1s: WFDMA0: queued=130 head=390 tail=260 rx bytes=184791894 > t+9s: WFDMA0: queued=148 head=408 tail=260 rx bytes=184791894 (unchanged) > > `tail` never advances during the whole stall window while `head` > keeps growing - the queue is being filled but never drained. > > Reproduced this exact signature four times total (twice on 7.1.8-1, > twice on 7.2.0-rc7-1), always with MLO 5GHz+6GHz active, always > within 3-5 minutes of normal use. > > 6GHz correlation: > My MLO links are 5GHz (5180 MHz) + 6GHz (varies, e.g. 6135 MHz). > Fully disabling the 6GHz radio on the access point (FritzBox 5690 > Pro, FRITZ!OS 8.25 - not just client-side band restriction via > NetworkManager, which does not reliably suppress the second MLO > link) results in 30+ minutes of clean operation with no freeze. > Re-enabling 6GHz reproduces the freeze again within minutes. > > Regression info: > Multiple users on the CachyOS forum report this started with the > 7.1.x kernel line; the previous longterm kernel (6.18.x) is reported > to work fine by another user with different affected hardware > (different AP). I have not personally tested 6.18.x myself, and have > not performed a commit-level bisection yet - happy to do so if > pointed toward the most likely area of the driver, given the AP-side > 6GHz dependency makes a fully automated bisection awkward. > > Steps to reproduce: > 1. Connect to an AP advertising WiFi 7 MLO with 5GHz+6GHz link > options, so the client negotiates both links. > 2. Use the connection normally (web browsing is sufficient). > 3. Within roughly 3-5 minutes, new connections start hanging; > existing traffic stops. > 4. Check `iw link` - link still reports as connected/healthy. > 5. Check `/sys/kernel/debug/ieee80211/phy*/mt76/xmit-queues` - > WFDMA0 tail pointer frozen while head continues to advance. > > Environment: > - Kernel: linux-cachyos 7.1.8-1 and linux-cachyos-rc 7.2.0-rc7-1 > (CachyOS, near-vanilla Arch-based build) > - cat /proc/version: > Linux version 7.2.0-rc7-1-cachyos-rc (linux-cachyos-rc@cachyos) > (clang version 22.1.8, LLD 22.1.8) #1 SMP PREEMPT_DYNAMIC > Mon, 10 Aug 2026 20:51:00 +0000 > - Distribution: CachyOS > - Architecture: x86_64 > - Kernel tainted: no (/proc/sys/kernel/tainted = 0) > - WiFi hardware: MediaTek MT7925 (mt7925e driver, mt76 core), > WM Firmware Build 20260605184805, ASIC revision 79250000 > - Access point: AVM FritzBox 5690 Pro, FRITZ!OS 8.25 > > Related community discussion: > https://discuss.cachyos.org/t/mt7925-wifi-7-mlo-breaks-connectivity-on-7-1-x-mt76-regression/31972 > (I am the thread starter; several other users report matching > symptoms with different hardware/APs) > > Possibly related (different symptom, same 6GHz/MT7925 area, worth > checking for a common root cause): > Bug 221627 - MT7925E - System locks up with flashing caps lock > (comment 1 there: "caused by switching between AP's with the same > 6GHz wifi name") > > Attachments (will add to this ticket): > - Full incident logs for all four reproductions (dmesg + journalctl > + mt76 debugfs time series across each freeze) > - iw station dump / iw link output during freeze > - Kernel .config > > Note: checked bug 220424 (mlo_sta_cmd/sta_cmd regression) before filing - > different symptom (persistent near-zero throughput vs. our intermittent > full stall with healthy link stats) and different root cause (broadcast > wtbl hdr_trans_tlv handling). That fix landed in 6.16.6/6.17-rc5, well > before the kernel versions tested here, so it is not the cause of this > issue. > > ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] mt7925: MLO connectivity silently stalls with 6GHz link active 2026-08-15 5:51 ` Thorsten Leemhuis @ 2026-08-16 20:48 ` Devin Wittmayer 2026-08-17 14:25 ` Jonas Hort 0 siblings, 1 reply; 21+ messages in thread From: Devin Wittmayer @ 2026-08-16 20:48 UTC (permalink / raw) To: Jonas Hort, Thorsten Leemhuis, regressions Cc: linux-wireless, lorenzo.bianconi83 On 15/08/2026 07:51, Thorsten Leemhuis wrote: > CCing Devin, who wrote two of the following messages Those are the mt7921 regd deadlock. mt7925 does not have it, there is no equivalent of the mt7921_mac_sta_add() call site that creates it. Worth ruling in or out before you bisect. mt7925 sets IEEE80211_MLD_CAP_OP_MAX_SIMUL_LINKS to 0 in mt7925/main.c. It is an N-1 field, mt7996 sets MT7996_MAX_RADIOS - 1, so 0 means one simultaneous link, and mt7925/mcu.c then requests MT7925_ROC_REQ_MLSR_AG or _AA. MLSR is multi-link single radio: your two links share one radio, and ROC is what moves it between them. If 5 GHz is your deflink, your pair is named on that path in mt7925_mac_set_links(): if (band == NL80211_BAND_2GHZ || (band == NL80211_BAND_5GHZ && secondary_band == NL80211_BAND_6GHZ)) { mt7925_abort_roc(...); mt7925_set_mlo_roc(...); } That one only runs at association. The path that can run mid-session is mt7925_change_vif_links(), which calls mt7925_set_mlo_roc() for each added link. Both reach mt7925_mcu_set_mlo_roc(), and all of this is the same in 7.1 and 7.2-rc7. Under your existing watchdog: cd /sys/kernel/debug/tracing echo 'r:mloroc mt7925_mcu_set_mlo_roc ret=$retval:s32' > kprobe_events echo 'r:rocabort mt7925_mcu_abort_roc ret=$retval:s32' >> kprobe_events echo 'p:rocwork mt7925_roc_work' >> kprobe_events echo 1 > events/kprobes/enable Association will fire mloroc once, so anything later is a mid-session switch. If the WFDMA0 tail freezes while one of those is in flight, it is a link switch that did not finish. If they are silent across a stall, the whole path is ruled out and that is worth as much. On the bisect itself: your 6.18 good point is second-hand, from different hardware and a different AP. Worth confirming on your own box before spending steps against it. Devin ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] mt7925: MLO connectivity silently stalls with 6GHz link active 2026-08-16 20:48 ` Devin Wittmayer @ 2026-08-17 14:25 ` Jonas Hort 2026-08-17 21:34 ` Jonas Hort 0 siblings, 1 reply; 21+ messages in thread From: Jonas Hort @ 2026-08-17 14:25 UTC (permalink / raw) To: Devin Wittmayer, Thorsten Leemhuis, regressions Cc: linux-wireless, lorenzo.bianconi83 First, thanks to everyone helping out with this - really appreciate the time you're all putting in. One correction to how I described this earlier: the connection does NOT reliably self-heal on its own. I have manually intervened every single time to restore connectivity - either by disconnecting and reconnecting the WiFi connection, or by switching to my band-lock workaround (forcing 5GHz-only, which disables MLO). What I can say for certain: the system itself has never needed a reboot - it stays fully responsive throughout, only the WiFi link itself needs manual action to recover. Wanted to correct that record before it causes confusion. Update on the ROC tracing: three real freezes captured now with the kprobes active (all confirmed via the WFDMA0 tail-frozen signature). - Freeze #1 (15:37): no ROC activity in the trace. Fixed by disconnecting/reconnecting the WiFi connection fairly quickly. - Freeze #2 (15:43): this time I deliberately waited longer before intervening. The trace shows several ROC events (rocabort, mloroc, rocwork) clustered together. Fixed by switching to the band-lock workaround (5GHz-only). - Freeze #3 (16:15): no ROC activity again. Fixed by disconnecting/reconnecting the WiFi connection. Uploaded all three logs to the bugzilla ticket if useful: https://bugzilla.kernel.org/show_bug.cgi?id=221884 Still haven't gotten to confirming 6.18 as a clean baseline on my own hardware - that's next on my list. Am 16.08.26 um 22:48 schrieb Devin Wittmayer: > On 15/08/2026 07:51, Thorsten Leemhuis wrote: >> CCing Devin, who wrote two of the following messages > Those are the mt7921 regd deadlock. mt7925 does not have it, there is no > equivalent of the mt7921_mac_sta_add() call site that creates it. > > Worth ruling in or out before you bisect. mt7925 sets > IEEE80211_MLD_CAP_OP_MAX_SIMUL_LINKS to 0 in mt7925/main.c. It is an N-1 > field, mt7996 sets MT7996_MAX_RADIOS - 1, so 0 means one simultaneous link, > and mt7925/mcu.c then requests MT7925_ROC_REQ_MLSR_AG or _AA. MLSR is > multi-link single radio: your two links share one radio, and ROC is what > moves it between them. > > If 5 GHz is your deflink, your pair is named on that path in > mt7925_mac_set_links(): > > if (band == NL80211_BAND_2GHZ || > (band == NL80211_BAND_5GHZ && secondary_band == NL80211_BAND_6GHZ)) { > mt7925_abort_roc(...); > mt7925_set_mlo_roc(...); > } > > That one only runs at association. The path that can run mid-session is > mt7925_change_vif_links(), which calls mt7925_set_mlo_roc() for each added > link. Both reach mt7925_mcu_set_mlo_roc(), and all of this is the same in > 7.1 and 7.2-rc7. > > Under your existing watchdog: > > cd /sys/kernel/debug/tracing > echo 'r:mloroc mt7925_mcu_set_mlo_roc ret=$retval:s32' > kprobe_events > echo 'r:rocabort mt7925_mcu_abort_roc ret=$retval:s32' >> kprobe_events > echo 'p:rocwork mt7925_roc_work' >> kprobe_events > echo 1 > events/kprobes/enable > > Association will fire mloroc once, so anything later is a mid-session > switch. If the WFDMA0 tail freezes while one of those is in flight, it is a > link switch that did not finish. If they are silent across a stall, the > whole path is ruled out and that is worth as much. > > On the bisect itself: your 6.18 good point is second-hand, from different > hardware and a different AP. Worth confirming on your own box before > spending steps against it. > > Devin ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] mt7925: MLO connectivity silently stalls with 6GHz link active 2026-08-17 14:25 ` Jonas Hort @ 2026-08-17 21:34 ` Jonas Hort 2026-08-19 1:18 ` Devin Wittmayer 0 siblings, 1 reply; 21+ messages in thread From: Jonas Hort @ 2026-08-17 21:34 UTC (permalink / raw) To: Devin Wittmayer, Thorsten Leemhuis, regressions Cc: linux-wireless, lorenzo.bianconi83 Quick follow-up: managed to confirm 6.18 as a clean baseline on my own hardware now (not just secondhand from others in the forum thread) - running Linux 6.18.42-1-cachyos-lts with MLO active (5GHz+6GHz, same FritzBox 5690 Pro) for 3 hours straight, no freeze at all. $ uname -r 6.18.42-1-cachyos-lts $ cat /proc/version Linux version 6.18.42-1-cachyos-lts (linux-cachyos-lts@cachyos) (gcc (GCC) 16.1.1 20260728, GNU ld (GNU Binutils) 2.47) #1 SMP PREEMPT_DYNAMIC Mon, 03 Aug 2026 17:38:28 +0000 $ uptime -p up 3 hours, 2 minutes $ iw dev wlan0 link Connected to 96:fc:7d:1a:7f:f0 (on wlan0) SSID: BKA Diensttelefon #52 Link 1 BSSID b6:fc:7d:1a:7f:f0 freq: 5200.0 Link 2 BSSID c6:fc:7d:1a:7f:f0 freq: 5975.0 MLD 96:fc:7d:1a:7f:f0 stats: RX: 448724534 bytes (2158217 packets) TX: 257975286 bytes (1259211 packets) signal: -64 dBm tx bitrate: 1921.5 MBit/s 160MHz EHT-MCS 9 EHT-NSS 2 EHT-GI 0 Regards, Jonas Am 17.08.26 um 16:25 schrieb Jonas Hort: > First, thanks to everyone helping out with this - really appreciate > the time you're all putting in. > > One correction to how I described this earlier: the connection does > NOT reliably self-heal on its own. I have manually intervened every > single time to restore connectivity - either by disconnecting and > reconnecting the WiFi connection, or by switching to my band-lock > workaround (forcing 5GHz-only, which disables MLO). What I can say > for certain: the system itself has never needed a reboot - it stays > fully responsive throughout, only the WiFi link itself needs manual > action to recover. Wanted to correct that record before it causes > confusion. > > Update on the ROC tracing: three real freezes captured now with the > kprobes active (all confirmed via the WFDMA0 tail-frozen signature). > > - Freeze #1 (15:37): no ROC activity in the trace. Fixed by > disconnecting/reconnecting the WiFi connection fairly quickly. > - Freeze #2 (15:43): this time I deliberately waited longer before > intervening. The trace shows several ROC events (rocabort, mloroc, > rocwork) clustered together. Fixed by switching to the band-lock > workaround (5GHz-only). > - Freeze #3 (16:15): no ROC activity again. Fixed by > disconnecting/reconnecting the WiFi connection. > > Uploaded all three logs to the bugzilla ticket if useful: > https://bugzilla.kernel.org/show_bug.cgi?id=221884 > > Still haven't gotten to confirming 6.18 as a clean baseline on my own > hardware - that's next on my list. > > Am 16.08.26 um 22:48 schrieb Devin Wittmayer: >> On 15/08/2026 07:51, Thorsten Leemhuis wrote: >>> CCing Devin, who wrote two of the following messages >> Those are the mt7921 regd deadlock. mt7925 does not have it, there is no >> equivalent of the mt7921_mac_sta_add() call site that creates it. >> >> Worth ruling in or out before you bisect. mt7925 sets >> IEEE80211_MLD_CAP_OP_MAX_SIMUL_LINKS to 0 in mt7925/main.c. It is an N-1 >> field, mt7996 sets MT7996_MAX_RADIOS - 1, so 0 means one simultaneous >> link, >> and mt7925/mcu.c then requests MT7925_ROC_REQ_MLSR_AG or _AA. MLSR is >> multi-link single radio: your two links share one radio, and ROC is what >> moves it between them. >> >> If 5 GHz is your deflink, your pair is named on that path in >> mt7925_mac_set_links(): >> >> if (band == NL80211_BAND_2GHZ || >> (band == NL80211_BAND_5GHZ && secondary_band == >> NL80211_BAND_6GHZ)) { >> mt7925_abort_roc(...); >> mt7925_set_mlo_roc(...); >> } >> >> That one only runs at association. The path that can run mid-session is >> mt7925_change_vif_links(), which calls mt7925_set_mlo_roc() for each >> added >> link. Both reach mt7925_mcu_set_mlo_roc(), and all of this is the >> same in >> 7.1 and 7.2-rc7. >> >> Under your existing watchdog: >> >> cd /sys/kernel/debug/tracing >> echo 'r:mloroc mt7925_mcu_set_mlo_roc ret=$retval:s32' > >> kprobe_events >> echo 'r:rocabort mt7925_mcu_abort_roc ret=$retval:s32' >> >> kprobe_events >> echo 'p:rocwork mt7925_roc_work' >> kprobe_events >> echo 1 > events/kprobes/enable >> >> Association will fire mloroc once, so anything later is a mid-session >> switch. If the WFDMA0 tail freezes while one of those is in flight, >> it is a >> link switch that did not finish. If they are silent across a stall, the >> whole path is ruled out and that is worth as much. >> >> On the bisect itself: your 6.18 good point is second-hand, from >> different >> hardware and a different AP. Worth confirming on your own box before >> spending steps against it. >> >> Devin ^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [REGRESSION] mt7925: MLO connectivity silently stalls with 6GHz link active 2026-08-17 21:34 ` Jonas Hort @ 2026-08-19 1:18 ` Devin Wittmayer 0 siblings, 0 replies; 21+ messages in thread From: Devin Wittmayer @ 2026-08-19 1:18 UTC (permalink / raw) To: Jonas Hort, Thorsten Leemhuis, regressions Cc: linux-wireless, Felix Fietkau, lorenzo.bianconi83 Thank you very much, that answers both things. The ROC tracing is the more useful of the two even though it came back negative. Two of the three freezes have no ROC activity in them at all, so a link switch that never finished cannot be what starts this. The middle one does have rocabort, mloroc and rocwork in it, but one out of three makes that look like the exception rather than the pattern. So the area I sent you looking at is out, and that is worth knowing before you spend nights on builds. One other thing worth saying first. There is a five patch mt76 series on the list at the moment and two of the patches look like they were written for exactly this bug. I do not think they were, and it is your own numbers that show it. The failure 4/5 fixes stops mt76_txq_schedule_list from servicing the queue, and the one 5/5 fixes stops mt76_txq_send_burst once the non-AQL count reaches its cap. Both of those keep frames from ever reaching the hardware, so if either were your problem head would be sitting still alongside tail. Yours does the opposite. Head climbs 390 to 408 while tail stays at 260, so the frames are getting into the ring and nothing is finishing them, which is the far end of the same path. 2/5 is a use after free when an interface goes away, so it does not fit either. I would not expect that series to change what you see. On the bisect I would build v7.0 next. There are 32 mt7925 commits between 7.0 and 7.1 and 19 of them are one run of work from Sean Wang, reworking how the driver tracks the per link mlink and WCID for an MLO station. That is the kind of change that fits a bug only showing up with two links up. If v7.0 comes back clean, that series is where I would look. If v7.0 is already broken then it is off the hook and 6.19 becomes the next split. The mt76 core and mac80211 both moved in the same window, so mt7925 is where I would look first rather than the only place worth looking. Devin Am 17.08.26 um 23:34 schrieb Jonas Hort: > Quick follow-up: managed to confirm 6.18 as a clean baseline on my > own hardware now (not just secondhand from others in the forum > thread) - running Linux 6.18.42-1-cachyos-lts with MLO active > (5GHz+6GHz, same FritzBox 5690 Pro) for 3 hours straight, no freeze > at all. Am 17.08.26 um 16:25 schrieb Jonas Hort: > One correction to how I described this earlier: the connection does > NOT reliably self-heal on its own. I have manually intervened every > single time to restore connectivity [...] > > Update on the ROC tracing: three real freezes captured now with the > kprobes active (all confirmed via the WFDMA0 tail-frozen signature). > > - Freeze #1 (15:37): no ROC activity in the trace. > - Freeze #2 (15:43): [...] The trace shows several ROC events > (rocabort, mloroc, rocwork) clustered together. > - Freeze #3 (16:15): no ROC activity again. > > Uploaded all three logs to the bugzilla ticket if useful: > https://bugzilla.kernel.org/show_bug.cgi?id=221884 ^ permalink raw reply [flat|nested] 21+ messages in thread
end of thread, other threads:[~2026-10-05 2:44 UTC | newest] Thread overview: 21+ messages (download: mbox.gz follow: Atom feed -- links below jump to the message on this page -- 2026-08-19 9:03 [REGRESSION] mt7925: MLO connectivity silently stalls with 6GHz link active Jonas Hort 2026-08-22 20:18 ` Jonas Hort 2026-08-23 20:54 ` Jonas Hort 2026-08-24 4:50 ` Devin Wittmayer 2026-08-24 4:52 ` Thorsten Leemhuis 2026-08-24 8:18 ` Jonas Hort 2026-08-29 22:13 ` Devin Wittmayer 2026-10-02 17:31 ` Jonas Hort 2026-10-02 18:11 ` Jonas Hort 2026-10-03 18:03 ` Devin Wittmayer 2026-10-04 15:48 ` Andrei Rusu de Castro 2026-10-04 15:49 ` [PATCH v2] wifi: mt76: mt7925: keep MLD membership consistent during link add Andrei Rusu de Castro 2026-10-04 18:25 ` Jonas Hort 2026-10-05 2:44 ` Devin Wittmayer -- strict thread matches above, loose matches on Subject: below -- 2026-08-19 9:03 [REGRESSION] mt7925: MLO connectivity silently stalls with 6GHz link active Jonas Hort 2026-08-14 14:04 Jonas Hort 2026-08-15 5:51 ` Thorsten Leemhuis 2026-08-16 20:48 ` Devin Wittmayer 2026-08-17 14:25 ` Jonas Hort 2026-08-17 21:34 ` Jonas Hort 2026-08-19 1:18 ` Devin Wittmayer
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox