* [BUG] rtw88 8821ce: connection wedges (100% loss until reboot) with station power save
@ 2026-08-26 16:25 Abdurrahman Karadag
2026-08-26 18:00 ` Abdurrahman Karadag
2026-08-28 3:48 ` Ping-Ke Shih
0 siblings, 2 replies; 13+ messages in thread
From: Abdurrahman Karadag @ 2026-08-26 16:25 UTC (permalink / raw)
To: linux-wireless; +Cc: pkshih
Hardware:
RTL8821CE [10ec:c821], subsystem AzureWave [1a3b:304a]
PCIe root port Intel [8086:51be] (00:1c.6)
ASUS Vivobook X1504ZA, firmware rtw8821c_fw.bin 24.11.0
(feature word 0x7: SIG|LPS_C2H|LCLK, no TX_WAKE)
Arch Linux, kernel 7.1.9 (also seen on 7.0.12).
The same laptop had none of this under Windows.
Symptom:
At some unpredictable point the connection wedges: traffic goes to 100%
packet loss and STAYS there until a reboot. The interface still shows as
associated (iwd/networkd report it connected, DHCP lease held), but nothing
passes - gateway ARP goes INCOMPLETE and stays that way. It does not
gradually degrade or self-heal; it is a hard stop that only a reboot clears.
It tends to hit soon after boot / when joining a network rather than on a
connection that has already been up for a long time.
The one reliable workaround is disabling station power save:
iw dev wlan0 set power_save off
With power save off I have not hit the wedge; with it on it recurs. This is
the strongest signal I have that the station-PS path is involved.
What I ruled out (each tested on this machine):
- rtw88_pci disable_aspm=y : no effect (the module param only gates the
device DBI 0x719 bit; it does not call
pci_disable_link_state, so host ASPM
L1/L1ss stay on - link/l1* remain 1)
- rtw88_core disable_lps_deep=y : no effect (firmware deep-PS off)
- full cold power-off boot : no effect
- suspend/resume : NOT the trigger - connectivity recovers
fine after resume in my tests
- The correctable PCIe "Physical Layer / RxErr" this card logs is DECOUPLED
from the failure: I measured the link working perfectly both while RxErr is
being logged and while it is absent. RxErr is not the cause.
Reproduction (honest):
I could not reproduce the wedge deterministically. Controlled tests all
passed on the stock driver: steady-state idle (minutes), long idle with an
off-device pinger sending to the sleeping STA, and disconnect/reconnect
loops. So it is intermittent and tied to station PS being active, but I have
not found the exact trigger - which is why the reliable handle is
"power_save off makes it stop".
Questions:
- Is a hard wedge (RX/TX stops until reboot) with station power save a known
failure mode on 8821ce? Does the driver have any watchdog/recovery for a
firmware or RX-DMA stall on this chip, or does it rely on the firmware?
- rtw_enter_lps_core() (ps.c) programs the same PS config for every chip
(rlbm=1, smart_ps=2, awake_interval=1) with no chip-specific override.
Given fw feature word 0x7 (no TX_WAKE) on 8821ce, could smart-PS mode 2
be the problem, and would legacy PS-Poll (smart_ps=0) for
RTW_CHIP_TYPE_8821C be worth trying?
I am happy to test any patch and report back with Tested-by, and to capture
logs, iw/AER dumps, rtw88 debugfs, or a PS/wake trace when it wedges.
Thanks.
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [BUG] rtw88 8821ce: connection wedges (100% loss until reboot) with station power save
2026-08-26 16:25 [BUG] rtw88 8821ce: connection wedges (100% loss until reboot) with station power save Abdurrahman Karadag
@ 2026-08-26 18:00 ` Abdurrahman Karadag
2026-08-28 3:57 ` Ping-Ke Shih
2026-08-28 3:48 ` Ping-Ke Shih
1 sibling, 1 reply; 13+ messages in thread
From: Abdurrahman Karadag @ 2026-08-26 18:00 UTC (permalink / raw)
To: linux-wireless; +Cc: pkshih
Correction and new data.
First, a correction to my report: I have now hit the wedge with station
power save OFF, so power save is not the trigger. My earlier "power_save
off makes it stop" was coincidence on an intermittent bug. Please disregard
the PS/smart_ps angle.
Second, I captured a wedge with Wireshark on wlan0 (802.3 view, i.e. at the
netdev boundary), and the signature is much more specific than "connection
wedges":
- RX is fully intact, including unicast: DHCP OFFER/ACK addressed to my
MAC, an ICMP echo request from the router, TLS data from the router,
and the gateway's ARP requests sent *unicast* to me were all received.
- My DHCP DISCOVER/REQUEST (342/345 bytes, L2 broadcast) reach the AP and
are answered within milliseconds - on two different APs (an Android
hotspot and a MikroTik router).
- 14-30 ms after those successful DHCP exchanges, my ARP requests for the
gateway (42 bytes, L2 broadcast; 89 of them, 1/s) get zero replies on
both networks.
- The gateway ARPs *me* (unicast, 10 times at ~0.77 s intervals, then
falls back to broadcast). I receive every request and reply immediately
(42-byte unicast ARP reply) - yet it keeps asking, so my replies never
reach it.
- My TCP SYNs (78 bytes, unicast to the gateway MAC) get no SYN-ACK.
So the failing set is small STA->AP frames (42-byte ARP, both broadcast
and unicast; 78-byte SYN) and the working set is 342-byte broadcast DHCP,
with RX working throughout and no kernel/driver messages. Since DHCP
succeeds tens of milliseconds before ARP fails, this looks like a
per-frame property (frame size, or possibly ethertype) rather than a
temporal stall. Two unrelated APs show the identical pattern, so it is
not AP-specific. The interface stays associated; only a reboot clears it
(a live driver reload froze the machine once, so I avoid that).
Next time it wedges I will run a size probe (static ARP entry for the
gateway, then ping -s 8/56/200/400/1000) to confirm whether it is
size-dependent, plus station-dump tx-failed/retry deltas and a
neigh-flush -> reconnect -> link down/up ladder to see which layer holds
the wedge. If there is anything specific on the 8821c TX side you would
like me to dump (tx desc, debugfs, registers) while it is wedged, tell me
and I will capture it.
Capture available on request.
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [BUG] rtw88 8821ce: connection wedges (100% loss until reboot) with station power save
2026-08-28 3:57 ` Ping-Ke Shih
@ 2026-08-28 3:37 ` Abdurrahman Karadag
2026-08-31 3:35 ` Ping-Ke Shih
2026-09-23 17:29 ` Bitterblue Smith
0 siblings, 2 replies; 13+ messages in thread
From: Abdurrahman Karadag @ 2026-08-28 3:37 UTC (permalink / raw)
To: linux-wireless; +Cc: pkshih
> Is there are version it works well on your platform?
> I'm thinking bisect is a method to find cause.
Yes - and I have to correct my first mail, which said "also seen on
7.0.12". Going back through the system journal and pacman log:
- 7.0.12 (Arch) ran on this laptop from 2026-06-16 to 2026-08-23. The
14 boots the journal holds from that period show no wedge episode
(defined as: DHCP lease acquired, then the gateway unreachable / DNS
failing continuously until reboot). The few candidates are either
sessions caught in the separate 63 s regulatory reconnect loop, or
sporadic DNS timeouts spread over hours-long otherwise-working
sessions (14-27 timeouts in 0.8-3.5 h), plus one 176 s blip - none is
the 100%-loss-until-reboot pattern.
- On 2026-08-23 18:41 pacman upgraded linux 7.0.12 -> 7.1.9 (in the same
transaction: linux-firmware-realtek 20260519 -> 20260810, whose
rtw8821c_fw.bin is byte-identical by sha256, and systemd 260.2 ->
261.2 a few minutes earlier).
- The first boot on 7.1.9 (2026-08-23 23:37) logged 12 wedge episodes,
and it has recurred on most days since.
So for a bisect: good = 7.0.12, bad = 7.1.9 (with the caveat that
systemd changed in the same upgrade; I don't think networkd can explain
a per-AC L2 TX failure, but I mention it for completeness). I still have
the 7.0.12 package and can confirm by running it again for a few days,
and I'm happy to bisect the rtw88/mac80211 range between the two if you
think that's the right next step. (Under Windows the same laptop has
never shown this.)
> The full cold power-off boot includes above two settings, right?
Yes. /etc/modprobe.d had "options rtw88_pci disable_aspm=y" and
"options rtw88_core disable_lps_deep=y", initramfs rebuilt, full power
off, and after boot both parameters read Y in /sys/module/. The wedge
still happened on that boot.
> How did you measure the link? CAT-C?
Nothing special: 30-second windows of ping to the gateway plus
"ip neigh" state and "iw dev wlan0 station dump" counters, while
recording the delta of /sys/bus/pci/devices/0000:02:00.0/aer_dev_correctable
and counting RxErr lines in the kernel log for the same window. Examples:
one window had 0% loss with 6 new RxErr; another had 0% loss with 0
RxErr; over that boot 329 RxErr accumulated while the link was fine. And
the wedge itself occurred in a window with almost no RxErr. So I could
not find any correlation in either direction.
> Can you setup another WiFi as monitor mode to capture 802.11 packets?
> Use another WiFi monitor to see if RTL8821CE actually transmitted the packets.
Yes, I will. I need a second adapter for that; I'll do it on a 2.4 GHz
20 MHz network so a simple monitor-mode dongle is enough, and filter on
the laptop's TA/RA. I'll report whether the ARP/BE frames appear on the
air at all, and if they do, whether the AP ACKs them.
> I'm not sure why the size can affect the result. Normally large size
> is harder to transmit basically though.
You are right, and I withdraw the size theory. Re-examining the same
capture by 802.11 access category instead of size explains every frame
with no exceptions:
- every laptop TX frame that provably reached the AP carries IP DSCP
0xc0 (CS6 -> UP 6 -> AC_VO): the DHCP DISCOVER/REQUEST frames
(systemd-networkd's DHCP client marks them CS6) and IGMPv3 reports;
- every laptop TX frame that provably never arrived is AC_BE:
89 ARP requests (no IP header -> BE), 13 unicast ARP replies,
83 DNS queries (0 answers), 2 TCP SYNs (no SYN-ACK), mDNS/LLMNR;
- sizes overlap the wrong way for a size theory: BE frames of 42-201 B
all fail, VO frames of 54-345 B all pass.
So the wedge looks like "AC_BE TX dead, AC_VO TX alive", RX intact, and
it persists across disconnect/reconnect and across AP changes (so not
per-association state), cleared only by a reboot. It happened with
power save fully off.
As far as I can tell from reading the driver, BE and VO take different
paths (separate PCIe TX rings per AC, and on 8821C BE/BK on the LOW
TX-FIFO queue vs VO/VI on NORMAL), so a per-AC stall - e.g. the BE ring's
mac80211 queue left stopped, or the LOW queue paused/page-starved - would
match what I see (BE frames silently aged out, VO flowing, no driver
message). Does that sound plausible to you, or is there a more likely
place for a per-AC TX stall on this chip? The monitor-mode capture
should at least show whether BE frames reach the air at all.
When it wedges, besides the air capture, I plan to dump before rebooting:
/sys/kernel/debug/ieee80211/phy0/queues (per-hw-queue stop reasons),
stations/<ap>/aqm (per-TID backlog), and via rtw88 debugfs read_reg:
REG_TXPAUSE (0x522), REG_TXDMA_STATUS (0x210), REG_FIFOPAGE_INFO_2/3
(0x234/0x238) and the BE/VO TXBD indices (0x3A8/0x3A0), compared with a
healthy baseline. If there are better registers or a debugfs page for
the BE ring / LOW queue state on 8821C, please tell me and I'll dump
those instead.
> Please fully turn off power save when you do the tests to reduce one
> factor that can possibly cause TX slowly or stuck.
Will do - power save stays off for all further tests.
Thanks a lot for looking at this.
^ permalink raw reply [flat|nested] 13+ messages in thread
* RE: [BUG] rtw88 8821ce: connection wedges (100% loss until reboot) with station power save
2026-08-26 16:25 [BUG] rtw88 8821ce: connection wedges (100% loss until reboot) with station power save Abdurrahman Karadag
2026-08-26 18:00 ` Abdurrahman Karadag
@ 2026-08-28 3:48 ` Ping-Ke Shih
1 sibling, 0 replies; 13+ messages in thread
From: Ping-Ke Shih @ 2026-08-28 3:48 UTC (permalink / raw)
To: Abdurrahman Karadag, linux-wireless@vger.kernel.org
Abdurrahman Karadag <abdurrahmankaradag19@gmail.com> wrote:
> Hardware:
> RTL8821CE [10ec:c821], subsystem AzureWave [1a3b:304a]
> PCIe root port Intel [8086:51be] (00:1c.6)
> ASUS Vivobook X1504ZA, firmware rtw8821c_fw.bin 24.11.0
> (feature word 0x7: SIG|LPS_C2H|LCLK, no TX_WAKE)
> Arch Linux, kernel 7.1.9 (also seen on 7.0.12).
Is there are version it works well on your platform?
I'm thinking bisect is a method to find cause.
>
> Symptom:
> At some unpredictable point the connection wedges: traffic goes to 100%
> packet loss and STAYS there until a reboot. The interface still shows as
> associated (iwd/networkd report it connected, DHCP lease held), but nothing
> passes - gateway ARP goes INCOMPLETE and stays that way. It does not
> gradually degrade or self-heal; it is a hard stop that only a reboot clears.
> It tends to hit soon after boot / when joining a network rather than on a
> connection that has already been up for a long time.
I'd reply this by the latter mail of yours.
>
> The one reliable workaround is disabling station power save:
> iw dev wlan0 set power_save off
> With power save off I have not hit the wedge; with it on it recurs. This is
> the strongest signal I have that the station-PS path is involved.
(Asked to disregard this.)
>
> What I ruled out (each tested on this machine):
> - rtw88_pci disable_aspm=y : no effect (the module param only gates the
> device DBI 0x719 bit; it does not call
> pci_disable_link_state, so host ASPM
> L1/L1ss stay on - link/l1* remain 1)
> - rtw88_core disable_lps_deep=y : no effect (firmware deep-PS off)
The full cold power-off boot includes above two settings, right?
> - full cold power-off boot : no effect
> - suspend/resume : NOT the trigger - connectivity recovers
> fine after resume in my tests
> - The correctable PCIe "Physical Layer / RxErr" this card logs is DECOUPLED
> from the failure: I measured the link working perfectly both while RxErr is
> being logged and while it is absent. RxErr is not the cause.
How did you measure the link? CAT-C?
>
> Reproduction (honest):
> I could not reproduce the wedge deterministically. Controlled tests all
> passed on the stock driver: steady-state idle (minutes), long idle with an
> off-device pinger sending to the sleeping STA, and disconnect/reconnect
> loops. So it is intermittent and tied to station PS being active, but I have
> not found the exact trigger - which is why the reliable handle is
> "power_save off makes it stop".
>
> Questions:
> - Is a hard wedge (RX/TX stops until reboot) with station power save a known
> failure mode on 8821ce? Does the driver have any watchdog/recovery for a
> firmware or RX-DMA stall on this chip, or does it rely on the firmware?
The latter mail of yours explain RX is fully intact.
For RX path, if beacon gets loss, it will disconnect, so I think RX still works
for your case.
> - rtw_enter_lps_core() (ps.c) programs the same PS config for every chip
> (rlbm=1, smart_ps=2, awake_interval=1) with no chip-specific override.
> Given fw feature word 0x7 (no TX_WAKE) on 8821ce, could smart-PS mode 2
> be the problem, and would legacy PS-Poll (smart_ps=0) for
> RTW_CHIP_TYPE_8821C be worth trying?
It looks like it still happened if you entirely turn off power save, so
ignore this...
^ permalink raw reply [flat|nested] 13+ messages in thread
* RE: [BUG] rtw88 8821ce: connection wedges (100% loss until reboot) with station power save
2026-08-26 18:00 ` Abdurrahman Karadag
@ 2026-08-28 3:57 ` Ping-Ke Shih
2026-08-28 3:37 ` Abdurrahman Karadag
0 siblings, 1 reply; 13+ messages in thread
From: Ping-Ke Shih @ 2026-08-28 3:57 UTC (permalink / raw)
To: Abdurrahman Karadag, linux-wireless@vger.kernel.org
Abdurrahman Karadag <abdurrahmankaradag19@gmail.com> wrote:
> Correction and new data.
>
> First, a correction to my report: I have now hit the wedge with station
> power save OFF, so power save is not the trigger. My earlier "power_save
> off makes it stop" was coincidence on an intermittent bug. Please disregard
> the PS/smart_ps angle.
>
> Second, I captured a wedge with Wireshark on wlan0 (802.3 view, i.e. at the
> netdev boundary), and the signature is much more specific than "connection
> wedges":
>
> - RX is fully intact, including unicast: DHCP OFFER/ACK addressed to my
> MAC, an ICMP echo request from the router, TLS data from the router,
> and the gateway's ARP requests sent *unicast* to me were all received.
Good to know RX is good.
Can you setup another WiFi as monitor mode to capture 802.11 packets?
> - My DHCP DISCOVER/REQUEST (342/345 bytes, L2 broadcast) reach the AP and
> are answered within milliseconds - on two different APs (an Android
> hotspot and a MikroTik router).
> - 14-30 ms after those successful DHCP exchanges, my ARP requests for the
> gateway (42 bytes, L2 broadcast; 89 of them, 1/s) get zero replies on
> both networks.
Use another WiFi monitor to see if RTL8821CE actually transmitted the packets.
> - The gateway ARPs *me* (unicast, 10 times at ~0.77 s intervals, then
> falls back to broadcast). I receive every request and reply immediately
> (42-byte unicast ARP reply) - yet it keeps asking, so my replies never
> reach it.
> - My TCP SYNs (78 bytes, unicast to the gateway MAC) get no SYN-ACK.
>
> So the failing set is small STA->AP frames (42-byte ARP, both broadcast
> and unicast; 78-byte SYN) and the working set is 342-byte broadcast DHCP,
> with RX working throughout and no kernel/driver messages. Since DHCP
> succeeds tens of milliseconds before ARP fails, this looks like a
> per-frame property (frame size, or possibly ethertype) rather than a
> temporal stall. Two unrelated APs show the identical pattern, so it is
> not AP-specific. The interface stays associated; only a reboot clears it
> (a live driver reload froze the machine once, so I avoid that).
>
> Next time it wedges I will run a size probe (static ARP entry for the
> gateway, then ping -s 8/56/200/400/1000) to confirm whether it is
> size-dependent, plus station-dump tx-failed/retry deltas and a
> neigh-flush -> reconnect -> link down/up ladder to see which layer holds
> the wedge. If there is anything specific on the 8821c TX side you would
> like me to dump (tx desc, debugfs, registers) while it is wedged, tell me
> and I will capture it.
I'm not sure why the size can affect the result. Normally large size
is harder to transmit basically though.
Please fully turn off power save when you do the tests to reduce one
factor that can possibly cause TX slowly or stuck.
>
> Capture available on request.
802.11 capture by another WiFi monitor is better.
^ permalink raw reply [flat|nested] 13+ messages in thread
* RE: [BUG] rtw88 8821ce: connection wedges (100% loss until reboot) with station power save
2026-08-28 3:37 ` Abdurrahman Karadag
@ 2026-08-31 3:35 ` Ping-Ke Shih
2026-09-02 9:15 ` Abdurrahman Karadag
2026-09-23 17:29 ` Bitterblue Smith
1 sibling, 1 reply; 13+ messages in thread
From: Ping-Ke Shih @ 2026-08-31 3:35 UTC (permalink / raw)
To: Abdurrahman Karadag, linux-wireless@vger.kernel.org
Abdurrahman Karadag <abdurrahmankaradag19@gmail.com> wrote:
> > Is there are version it works well on your platform?
> > I'm thinking bisect is a method to find cause.
>
> Yes - and I have to correct my first mail, which said "also seen on
> 7.0.12". Going back through the system journal and pacman log:
>
> - 7.0.12 (Arch) ran on this laptop from 2026-06-16 to 2026-08-23. The
> 14 boots the journal holds from that period show no wedge episode
> (defined as: DHCP lease acquired, then the gateway unreachable / DNS
> failing continuously until reboot). The few candidates are either
> sessions caught in the separate 63 s regulatory reconnect loop, or
> sporadic DNS timeouts spread over hours-long otherwise-working
> sessions (14-27 timeouts in 0.8-3.5 h), plus one 176 s blip - none is
> the 100%-loss-until-reboot pattern.
> - On 2026-08-23 18:41 pacman upgraded linux 7.0.12 -> 7.1.9 (in the same
> transaction: linux-firmware-realtek 20260519 -> 20260810, whose
> rtw8821c_fw.bin is byte-identical by sha256, and systemd 260.2 ->
> 261.2 a few minutes earlier).
> - The first boot on 7.1.9 (2026-08-23 23:37) logged 12 wedge episodes,
> and it has recurred on most days since.
>
> So for a bisect: good = 7.0.12, bad = 7.1.9 (with the caveat that
> systemd changed in the same upgrade; I don't think networkd can explain
> a per-AC L2 TX failure, but I mention it for completeness). I still have
> the 7.0.12 package and can confirm by running it again for a few days,
> and I'm happy to bisect the rtw88/mac80211 range between the two if you
> think that's the right next step. (Under Windows the same laptop has
> never shown this.)
Checking commits between 7.0.12 and 7.1.9, the only related commit might be
c95323ea9dfb ("wifi: rtw88: coex: Solve LE-HID lag & update coex version to 26020420")
You can revert the patch from 7.1.9 to see if it becomes normal.
Another simple way is to use 7.0.12 kernel + 7.1.9 rtw88 driver and
opposite combination to address the cause, like
kernel rtw88 driver result
------ ------------ ------
7.0.12 (built-in) Good
7.0.12 7.1.9
7.1.9 (built-in) NG
7.1.9 7.0.12
This can also bisect if the cause is driver or mac80211.
>
> > I'm not sure why the size can affect the result. Normally large size
> > is harder to transmit basically though.
>
> You are right, and I withdraw the size theory. Re-examining the same
> capture by 802.11 access category instead of size explains every frame
> with no exceptions:
>
> - every laptop TX frame that provably reached the AP carries IP DSCP
> 0xc0 (CS6 -> UP 6 -> AC_VO): the DHCP DISCOVER/REQUEST frames
> (systemd-networkd's DHCP client marks them CS6) and IGMPv3 reports;
> - every laptop TX frame that provably never arrived is AC_BE:
> 89 ARP requests (no IP header -> BE), 13 unicast ARP replies,
> 83 DNS queries (0 answers), 2 TCP SYNs (no SYN-ACK), mDNS/LLMNR;
> - sizes overlap the wrong way for a size theory: BE frames of 42-201 B
> all fail, VO frames of 54-345 B all pass.
>
> So the wedge looks like "AC_BE TX dead, AC_VO TX alive", RX intact, and
> it persists across disconnect/reconnect and across AP changes (so not
> per-association state), cleared only by a reboot. It happened with
> power save fully off.
>
> As far as I can tell from reading the driver, BE and VO take different
> paths (separate PCIe TX rings per AC, and on 8821C BE/BK on the LOW
> TX-FIFO queue vs VO/VI on NORMAL), so a per-AC stall - e.g. the BE ring's
> mac80211 queue left stopped, or the LOW queue paused/page-starved - would
> match what I see (BE frames silently aged out, VO flowing, no driver
> message). Does that sound plausible to you, or is there a more likely
> place for a per-AC TX stall on this chip? The monitor-mode capture
> should at least show whether BE frames reach the air at all.
I think BE and VO should almost the same, but as your perspective
BE and VO go via different paths. It is worth to do an experiment
to let all packets (by driver modification) go via VO queue.
>
> When it wedges, besides the air capture, I plan to dump before rebooting:
> /sys/kernel/debug/ieee80211/phy0/queues (per-hw-queue stop reasons),
> stations/<ap>/aqm (per-TID backlog), and via rtw88 debugfs read_reg:
> REG_TXPAUSE (0x522), REG_TXDMA_STATUS (0x210), REG_FIFOPAGE_INFO_2/3
> (0x234/0x238) and the BE/VO TXBD indices (0x3A8/0x3A0), compared with a
> healthy baseline. If there are better registers or a debugfs page for
> the BE ring / LOW queue state on 8821C, please tell me and I'll dump
> those instead.
Before checking these values, let's bisect the code, check sniffer, and
do experiment first.
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [BUG] rtw88 8821ce: connection wedges (100% loss until reboot) with station power save
2026-08-31 3:35 ` Ping-Ke Shih
@ 2026-09-02 9:15 ` Abdurrahman Karadag
2026-09-06 4:09 ` Ping-Ke Shih
0 siblings, 1 reply; 13+ messages in thread
From: Abdurrahman Karadag @ 2026-09-02 9:15 UTC (permalink / raw)
To: linux-wireless; +Cc: pkshih
> Checking commits between 7.0.12 and 7.1.9, the only related commit might be
> c95323ea9dfb ("wifi: rtw88: coex: Solve LE-HID lag & update coex version
> to 26020420")
> You can revert the patch from 7.1.9 to see if it becomes normal.
I reverted it from 7.1.9 and the wedge still happened, about an hour after
boot, so that commit is not the cause. BT is disabled on this machine
anyway (coex_info reports "BT disabled", BT status non-conn).
Rather than continue bisecting, I set up a watchdog that dumps
queues/aqm/TXBD indices at the moment of the wedge, before anything
touches the interface. I have five such dumps now (I discount one of
them, taken on a captive portal network where the ping probes are
unreliable), plus healthy baselines, and they show a consistent
driver-level failure. I have also been able to reproduce that failure
deterministically and to recover from it. Details below; I am
sending a patch as a separate mail.
1. What the wedge looks like from inside the driver
---------------------------------------------------
This is the dump from 2026-08-28 22:05, taken with power save off, on
5 GHz, 91 minutes into the association:
/sys/kernel/debug/ieee80211/phy0/queues
00: 0x00000000/0 (VO)
01: 0x00000000/0 (VI)
02: 0x00000001/0 (BE) <- stopped, reason bit 0 = DRIVER
03: 0x00000000/0 (BK)
stations/<ap>/aqm tid 0 (BE): backlog 20941 -> 26960 bytes,
94 -> 144 packets, flags DIRTY, still growing while sampled
iw station dump, 5 s apart:
tx packets 116017 -> 116017 (frozen)
rx packets 244885 -> 244992
beacon rx 53304 -> 53353
TXBD_IDX_BEQ (0x3A8) = 0x003c003a in all four samples, taken over
about 15 seconds
dmesg, shortly before: "failed to get tx report from firmware"
The BE queue is stopped, and rtw_pci_tx_write() only does that when the
ring has no room left:
if (avail_desc(ring->r.wp, ring->r.rp, ring->r.len) < 2) {
ieee80211_stop_queue(rtwdev->hw, skb_get_queue_mapping(skb));
ring->queue_stopped = true;
}
The only ieee80211_wake_queue() for that ring is inside the completion
loop of rtw_pci_tx_isr():
count = cur_rp - ring->r.rp;
while (count--) {
...
if (ring->queue_stopped && avail_desc(...) > 4)
ieee80211_wake_queue(hw, q_map);
...
}
so it can only run if the hardware read index has advanced. In the dump
that index is identical in every sample, the ring stays full, and the
queue is therefore never woken again. Nothing is logged, no counter
moves, and there is no watchdog for this state, which is why the
interface stays associated with RX working while TX is silently dead.
A second dump (2026-08-30 09:12, this one with power save on) shows the
same thing: BE stopped with reason DRIVER, backlog growing, tx packets
frozen at 115902 while rx packets and beacon counters keep increasing,
TXBD_IDX_BEQ = 0x001c001a identical at +2 s, +7 s and +2 min.
Two other dumps caught an earlier stage of the same failure: the read
index was already frozen while the ring still had room, so the queue was
not stopped yet. In the 2026-08-28 20:32 dump the read index stays at
0x45 across four samples while the write index moves 0x19 -> 0x1b ->
0x20 -> 0x22, i.e. frames were still being written into a ring that the
hardware had stopped draining.
2. Reproducing it deterministically
------------------------------------
Since I could not wait for another occurrence, I broke the TX path on
purpose: a temporary module parameter that skips the doorbell
write in rtw_pci_tx_kick_off_queue() for the BE queue, simulating a lost
kick-off. With BE traffic running, this reproduces the failure in about
12 seconds:
BE ring 0x000e000e q02 0x00000001 100% loss
that is, the same observable state as the dumps above: BE queue stopped
with reason DRIVER, ring full, read index not moving, TX dead, RX
unaffected. A Wireshark capture on wlan0 during the artificial stall
looks like the ones taken during real wedges: our ARP requests leave the
netdev and no replies arrive, while RX keeps working (artificial: 40 TX
/ 51 RX, 9 ARP requests, 0 replies; real: 193 / 300, 186 requests, 0
replies).
3. Recovery
-----------
I then added a TX stall check to rtw_watch_dog_work(): for each TX ring,
if there is outstanding work (queue stopped, or cur_rp != ring->r.wp)
and the hardware read index has not moved for three consecutive rounds
(about 6 s), run the kick-off for that queue again. Controlled run, with
the recovery selectable at runtime so both arms use the same build:
13:01:16.293 TX queue 1 stalled (rp 14 wp 170), recovery DISABLED
13:01:22.309 TX queue 1 stalled (rp 14 wp 5), recovery DISABLED
13:01:28.325 TX queue 1 stalled (rp 14 wp 12), recovery DISABLED
13:01:34.341 TX queue 1 stalled (rp 14 wp 12), recovery DISABLED
13:01:40.293 TX queue 1 stalled (rp 14 wp 12), recovery DISABLED
13:01:46.309 TX queue 1 stalled (rp 14 wp 12, stopped), kicking
With recovery disabled the stall persisted for 24 s across five
detection rounds and did not clear by itself. One kick-off was enough:
the ring went from 0x000e000e to 0x00470047, the queue stop flag
cleared, and ping returned to 0% loss within a second of the kick. The
wake up then happens through the existing path, because once the
hardware starts consuming again rtw_pci_tx_isr() runs with a non-zero
count.
4. What this does and does not prove
-------------------------------------
The artificial stall simulates a lost doorbell, so I cannot claim the
root cause in the field is the same. I looked for a register-level
difference between the two and did not find one. I had suspected the
real dumps were special because the hardware read index sits a couple of
slots ahead of the last announced write index, but after instrumenting
the driver I see the same relationship during perfectly healthy
operation, when the ring fills up under load and the queue is stopped
and woken again within milliseconds. The only thing that distinguishes
the failure is that the read index stops advancing at all.
What the experiment does show is that once a TX ring stops draining, the
driver has no way back, and that a re-kick recovers at least the
lost-doorbell case.
I still do not know why the hardware stops consuming descriptors in the
field, and it is not easy to catch: since the dumps above I have gone
about 74 hours without a single occurrence, across five driver builds
(including the stock one), five networks and both bands, with the
association living well past the ages at which it used to happen. It
seems to come in bursts - three occurrences in four hours on one
evening, one two days later, then nothing.
The 802.11 monitor capture you asked for is still on my list (the second
adapter is on the way), and my watchdog now also re-writes the doorbell
during a real wedge so that I can tell whether a re-kick is enough there
too. I will report that separately.
5. Corrections to my earlier mails
-----------------------------------
- "persists across disconnect/reconnect ... cleared only by a reboot":
only true once the ring is full. Two of the wedges were cleared by a
reconnect, and the 2026-08-30 one by ip link set wlan0 down/up,
which reinitialises the rings.
- "AC_BE TX dead, AC_VO TX alive": that was the early stage. In the
two dumps where the ring is full, ping -Q 0xc0 fails as well, so by
then TX is dead on VO too.
- "it tends to hit soon after boot": wrong. The association ages at
the moment of the four wedges were 95, 91, 36 and 61 minutes.
- Power save: three of the five dumps were taken with PS on, because
my udev rule only ran on ACTION=="add" and the setting reverts after
a link down/up. It is now forced off every 20 s. The two dumps I
rely on most, 2026-08-28 20:32 and 22:05, were both taken with PS
off, and they show the same state as the PS-on ones, so PS does not
look like a factor.
I will send the patch separately so that it lands in patchwork on its
own. It only adds the stall detection and the re-kick; it does not touch
the normal TX path, it skips rings with nothing in flight so an idle
device is never poked, and it warns once per stall rather than on every
retry.
One question, in case a re-kick turns out not to be enough in the field.
The out-of-tree rtw88 driver has a pci_old.c for the older PCIe
generation, where the PCIe DMA is reset after a TRX hang using the two
status bits at REG_DBI_CTRL + 3 (bit 0 TX, bit 1 RX), with bit 2
enabling the detection. Is that status bit valid on 8821CE as well? If
it is, I can dump it during the next wedge and, if it confirms a DMA
hang, a reset could be added as a second stage.
Full dumps (five wedges, healthy baselines, the controlled run) and the
captures are available on request.
^ permalink raw reply [flat|nested] 13+ messages in thread
* RE: [BUG] rtw88 8821ce: connection wedges (100% loss until reboot) with station power save
2026-09-02 9:15 ` Abdurrahman Karadag
@ 2026-09-06 4:09 ` Ping-Ke Shih
2026-09-23 15:00 ` abkarada
0 siblings, 1 reply; 13+ messages in thread
From: Ping-Ke Shih @ 2026-09-06 4:09 UTC (permalink / raw)
To: Abdurrahman Karadag, linux-wireless@vger.kernel.org
Abdurrahman Karadag <abdurrahmankaradag19@gmail.com> wrote:
> > Checking commits between 7.0.12 and 7.1.9, the only related commit might be
> > c95323ea9dfb ("wifi: rtw88: coex: Solve LE-HID lag & update coex version
> > to 26020420")
> > You can revert the patch from 7.1.9 to see if it becomes normal.
>
> I reverted it from 7.1.9 and the wedge still happened, about an hour after
> boot, so that commit is not the cause. BT is disabled on this machine
> anyway (coex_info reports "BT disabled", BT status non-conn).
I wonder even you use driver of 7.0.12 (on 7.1.9). It will still happen.
Not sure if host side change something?
Can you 100% ensure 7.0.12 is fine?
>
> Rather than continue bisecting, I set up a watchdog that dumps
> queues/aqm/TXBD indices at the moment of the wedge, before anything
> touches the interface. I have five such dumps now (I discount one of
> them, taken on a captive portal network where the ping probes are
> unreliable), plus healthy baselines, and they show a consistent
> driver-level failure. I have also been able to reproduce that failure
> deterministically and to recover from it. Details below; I am
> sending a patch as a separate mail.
>
> 1. What the wedge looks like from inside the driver
> ---------------------------------------------------
>
> This is the dump from 2026-08-28 22:05, taken with power save off, on
> 5 GHz, 91 minutes into the association:
>
> /sys/kernel/debug/ieee80211/phy0/queues
> 00: 0x00000000/0 (VO)
> 01: 0x00000000/0 (VI)
> 02: 0x00000001/0 (BE) <- stopped, reason bit 0 = DRIVER
> 03: 0x00000000/0 (BK)
>
> stations/<ap>/aqm tid 0 (BE): backlog 20941 -> 26960 bytes,
> 94 -> 144 packets, flags DIRTY, still growing while sampled
>
> iw station dump, 5 s apart:
> tx packets 116017 -> 116017 (frozen)
> rx packets 244885 -> 244992
> beacon rx 53304 -> 53353
>
> TXBD_IDX_BEQ (0x3A8) = 0x003c003a in all four samples, taken over
> about 15 seconds
Hardware read index is 0x3c, and host write index is 0x3a.
So, it reaches the limit of stop queue.
[...]
>
> I will send the patch separately so that it lands in patchwork on its
> own. It only adds the stall detection and the re-kick; it does not touch
> the normal TX path, it skips rings with nothing in flight so an idle
> device is never poked, and it warns once per stall rather than on every
> retry.
As your experiments, the cause is hardware never reads TX buffer and gets
stuck, right?
I quickly check the patch. It looks like the way to unlock this state is
to call rtw_pci_tx_kick_off_queue() again?
Which means not a driver side bug (stop queue but not restart queue properly),
right?
>
> One question, in case a re-kick turns out not to be enough in the field.
> The out-of-tree rtw88 driver has a pci_old.c for the older PCIe
> generation, where the PCIe DMA is reset after a TRX hang using the two
> status bits at REG_DBI_CTRL + 3 (bit 0 TX, bit 1 RX), with bit 2
> enabling the detection. Is that status bit valid on 8821CE as well? If
> it is, I can dump it during the next wedge and, if it confirms a DMA
> hang, a reset could be added as a second stage.
I checked vendor driver. It only does this at initial step, not to recover
it at runtime.
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [BUG] rtw88 8821ce: connection wedges (100% loss until reboot) with station power save
2026-09-06 4:09 ` Ping-Ke Shih
@ 2026-09-23 15:00 ` abkarada
2026-09-24 2:40 ` Ping-Ke Shih
0 siblings, 1 reply; 13+ messages in thread
From: abkarada @ 2026-09-23 15:00 UTC (permalink / raw)
To: linux-wireless; +Cc: pkshih
Sorry for the long silence. I spent it on the monitor capture you asked
for. I set up a second machine in monitor mode and checked it against a
healthy link first: 531 QoS-data frames from my station, all TID 0, and
157 ACKs back from the AP, so the capture side works and would show me
whether the chip transmitted.
Then I waited two and a half weeks and the failure did not happen once.
I cannot tell you whether something masked it or whether it is simply
that rare. The rig stays in place, and that capture is the first thing
you will get if it happens again.
In the meantime I went back over the data I already had, and I think I
had been treating two separate moments as one.
> I wonder even you use driver of 7.0.12 (on 7.1.9). It will still happen.
> Can you 100% ensure 7.0.12 is fine?
I cannot, and it is not. My watchdog caught two stalls on 2026-09-14,
17:48 and 19:09, in one boot, both on the stock 7.0.12 kernel and its
own rtw88. I also withdraw the "12 episodes on the first 7.1.9 boot"
count from my earlier mail: it came from a log heuristic that counted
"DHCP lease acquired, then connectivity lost", and that boot logged 51
"deauthenticating by local choice (Reason: 3)" events, each followed by
a reassociation about a minute later. That is a separate regulatory
disconnect loop and it happens on 7.0.12 boots too. There is no
regression here. My good/bad version claim was never measured properly.
> As your experiments, the cause is hardware never reads TX buffer and
> gets stuck, right?
Yes, and I could not get any further than that. The BE ring's hardware
read index stops moving while the host write index keeps advancing. To
rule out a missed doorbell I taught the watchdog to rewrite the TXBD
host index during a live stall and re-read it. Three independent
occurrences, each pair taken immediately before and immediately after
the rewrite:
BE read index BE write index
2026-09-06 03:17 0x00b7 -> 0x00b7 0x00a2 -> 0x00a3
2026-09-14 17:48 0x00c9 -> 0x00c9 0x006c -> 0x006d
2026-09-14 19:09 0x009e -> 0x009e 0x0050 -> 0x0051
The read index did not move in any of them, and it was not a momentary
sample either: in the 19:09 capture it held 0x9e from the first register
read to the last, minutes apart, while the write index went 0x2b -> 0x4c.
Traffic stayed at 100% loss throughout.
So rewriting the TXBD host index is not sufficient to recover the ring.
I have asked for that patch to be dropped.
> I checked vendor driver. It only does this at initial step, not to
> recover it at runtime.
Thank you for checking the vendor driver. That settles it: those bits
are an init-time thing, not a runtime recovery, so I have dropped the
idea.
> Which means not a driver side bug (stop queue but not restart queue
> properly), right?
Partly. I think we have been looking at two different moments of the
same failure, and there is a gap at each of them: the driver never
notices that a ring has stopped advancing, and it has no way back unless
the hardware resumes by itself. Neither gap is the cause. Both are
things the driver could do something about.
> Hardware read index is 0x3c, and host write index is 0x3a.
> So, it reaches the limit of stop queue.
Agreed, and that dump is the late moment. avail_desc() is 1 there, so
ieee80211_stop_queue() is exactly right and the driver is doing what it
should.
But both stalls I caught in September are from an earlier moment, before
the ring fills:
17:48 0x3a8: 0x00c9004c -> 0x00c9004e -> 0x00c90053
read index 0xc9 frozen, write index 0x4c -> 0x53, avail 124
19:09 0x3a8: 0x009e002b -> 0x009e002d -> 0x009e0034
read index 0x9e frozen, write index 0x2b -> 0x34, avail 115
The ring still has 124 and 115 free descriptors, the queue is not
stopped, nothing has hit any limit - and traffic is already 100% lost,
because the hardware has stopped consuming descriptors while the driver
keeps writing them and ringing the doorbell. Nothing in the driver
notices this. The watchdog runs, the link stays associated, RX keeps
working, and there is no counter or log line anywhere that says the TX
ring has not advanced.
So the first gap is detection, and it happens before your dump.
The second gap is what happens after it. Once the ring does fill and the
queue is stopped, everything that could undo that sits behind the same
frozen index. In rtw_pci_tx_isr():
if (cur_rp >= ring->r.rp)
count = cur_rp - ring->r.rp;
else
count = ring->r.len - (ring->r.rp - cur_rp);
while (count--) {
...
if (ring->queue_stopped &&
avail_desc(ring->r.wp, rp_idx, ring->r.len) > 4) {
q_map = skb_get_queue_mapping(skb);
ieee80211_wake_queue(hw, q_map);
ring->queue_stopped = false;
}
...
}
ring->r.rp = cur_rp;
The only ieee80211_wake_queue() in pci.c is inside that loop. As long as
cur_rp remains equal to ring->r.rp, count stays zero, so the body is
skipped and neither ring->r.rp nor the stopped queue is ever
reconsidered. Meanwhile rtw_pci_tx_write() keeps consuming descriptors
until avail_desc() < 2 and calls ieee80211_stop_queue().
To be fair to the design: this is not a deadlock. The descriptors are
still sitting in the ring, so the normal completion path can recover the
queue - but only if the hardware resumes consuming them. When it does
not, and in the three occurrences above it did not, there is nothing
else: no timeout, no ring reset, no device reset. The interface stays
dead until userspace does something about it.
Since the stall itself is rare and I cannot produce it on demand, I
built a way to put the driver into the stopped-queue state instead, so
that a recovery path could at least be tested. rtw_pci now takes a
debug-only module parameter that makes the TX completion handler behave
as if the read index never advanced, for a chosen ring - the same idea
as the existing fw_crash debugfs knob, which crashes the firmware
deliberately to exercise the recovery path. On 7.0.12, BE ring, station
otherwise idle:
14:29:58 injection on. The driver's own counters, logged from the
ISR, one second apart:
r.rp=174 r.wp=175 avail=254 stopped=0
r.rp=174 r.wp=172 avail=1 stopped=1
mac80211 BE queue: 0x1, IEEE80211_QUEUE_STOP_REASON_DRIVER
14:30:00 ping to the gateway: 100% loss
14:30:30 injection turned OFF
14:32:30 r.rp=174 r.wp=172 avail=1 stopped=1, queue still 0x1,
120 s later, with no recovery
The important observation is the final state: the injection is gone and
completions are no longer being suppressed, but nothing in the driver
revisits the stopped queue. By then the hardware may already have
consumed all descriptors that were submitted during the injection, so
there is no further TX completion to generate an interrupt and make the
driver observe the current hardware read index. With the mac80211 queue
stopped, no new descriptor is submitted either.
I should be precise about the limits of this. It is not your failure
reproduced. The field stall leaves real descriptors pending in the ring,
so if that hardware resumed, the driver would recover on its own; with
the injection, the hardware can consume the submitted descriptors while
the software read pointer is deliberately held back, leaving no later
completion event to make the driver revisit the stopped queue. What it
does give me is a deterministic way to reach that stopped-queue state,
which is what I would need to test any recovery patch - I did not want
to propose one I had no way of exercising.
In an earlier run of the same injection I also captured on wlan0 while
it was wedged. Once the ring is full, frames are either dropped in
rtw_pci_tx_write_data() with -ENOSPC or held in the qdisc; either way
they do not reach the air, but the ones that reach the driver are still
visible to tcpdump. So from userspace it looks like this: ARP requests
to the gateway repeat once a second and are never answered, while the
station still receives ARP requests from another host on the same
subnet. That is exactly what I see in the field.
So I am not asking you to treat the stall itself as a driver bug. I am
asking about the two gaps around it: the driver currently has no
mechanism to detect that a ring has stopped advancing, and no way back
if the hardware does not resume on its own. Would you consider a patch
that addresses those - per-ring detection of "read index has not moved
for a few watchdog rounds while descriptors are pending", triggering the
recovery the driver already has for a firmware crash,
rtw_fw_recovery() -> fw_recovery_work -> ieee80211_restart_hw()?
In the field a full interface restart is what recovers the link in both
occurrences where I got that far; on 2026-09-14 19:09 it came back 3 s
after a down/up and reassociated without a reboot. __fw_recovery_work()
does firmware-crash-specific work, so this would need a lighter variant,
and I would rather hear your view on the direction than send something
you do not want.
If you would rather not add a recovery path at all, the detection half
is still worth something on its own: a log line at the moment a ring
stops advancing, instead of a user discovering that the network died. It
would also have saved me most of the last two months, and it would make
the next report of this kind arrive with the register state already in
it.
Full dumps, the injection patch and the test script are available on
request.
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [BUG] rtw88 8821ce: connection wedges (100% loss until reboot) with station power save
2026-08-28 3:37 ` Abdurrahman Karadag
2026-08-31 3:35 ` Ping-Ke Shih
@ 2026-09-23 17:29 ` Bitterblue Smith
1 sibling, 0 replies; 13+ messages in thread
From: Bitterblue Smith @ 2026-09-23 17:29 UTC (permalink / raw)
To: Abdurrahman Karadag, linux-wireless; +Cc: pkshih
On 28/08/2026 06:37, Abdurrahman Karadag wrote:
>> Is there are version it works well on your platform?
>> I'm thinking bisect is a method to find cause.
>
> Yes - and I have to correct my first mail, which said "also seen on
> 7.0.12". Going back through the system journal and pacman log:
>
> - 7.0.12 (Arch) ran on this laptop from 2026-06-16 to 2026-08-23. The
> 14 boots the journal holds from that period show no wedge episode
> (defined as: DHCP lease acquired, then the gateway unreachable / DNS
> failing continuously until reboot). The few candidates are either
> sessions caught in the separate 63 s regulatory reconnect loop, or
> sporadic DNS timeouts spread over hours-long otherwise-working
> sessions (14-27 timeouts in 0.8-3.5 h), plus one 176 s blip - none is
> the 100%-loss-until-reboot pattern.
> - On 2026-08-23 18:41 pacman upgraded linux 7.0.12 -> 7.1.9 (in the same
> transaction: linux-firmware-realtek 20260519 -> 20260810, whose
> rtw8821c_fw.bin is byte-identical by sha256, and systemd 260.2 ->
> 261.2 a few minutes earlier).
> - The first boot on 7.1.9 (2026-08-23 23:37) logged 12 wedge episodes,
> and it has recurred on most days since.
>
> So for a bisect: good = 7.0.12, bad = 7.1.9 (with the caveat that
> systemd changed in the same upgrade; I don't think networkd can explain
> a per-AC L2 TX failure, but I mention it for completeness). I still have
> the 7.0.12 package and can confirm by running it again for a few days,
> and I'm happy to bisect the rtw88/mac80211 range between the two if you
> think that's the right next step. (Under Windows the same laptop has
> never shown this.)
>
>> The full cold power-off boot includes above two settings, right?
>
> Yes. /etc/modprobe.d had "options rtw88_pci disable_aspm=y" and
> "options rtw88_core disable_lps_deep=y", initramfs rebuilt, full power
> off, and after boot both parameters read Y in /sys/module/. The wedge
> still happened on that boot.
>
>> How did you measure the link? CAT-C?
>
> Nothing special: 30-second windows of ping to the gateway plus
> "ip neigh" state and "iw dev wlan0 station dump" counters, while
> recording the delta of /sys/bus/pci/devices/0000:02:00.0/aer_dev_correctable
> and counting RxErr lines in the kernel log for the same window. Examples:
> one window had 0% loss with 6 new RxErr; another had 0% loss with 0
> RxErr; over that boot 329 RxErr accumulated while the link was fine. And
> the wedge itself occurred in a window with almost no RxErr. So I could
> not find any correlation in either direction.
>
>> Can you setup another WiFi as monitor mode to capture 802.11 packets?
>> Use another WiFi monitor to see if RTL8821CE actually transmitted the packets.
>
> Yes, I will. I need a second adapter for that; I'll do it on a 2.4 GHz
> 20 MHz network so a simple monitor-mode dongle is enough, and filter on
> the laptop's TA/RA. I'll report whether the ARP/BE frames appear on the
> air at all, and if they do, whether the AP ACKs them.
>
>> I'm not sure why the size can affect the result. Normally large size
>> is harder to transmit basically though.
>
> You are right, and I withdraw the size theory. Re-examining the same
> capture by 802.11 access category instead of size explains every frame
> with no exceptions:
>
> - every laptop TX frame that provably reached the AP carries IP DSCP
> 0xc0 (CS6 -> UP 6 -> AC_VO): the DHCP DISCOVER/REQUEST frames
> (systemd-networkd's DHCP client marks them CS6) and IGMPv3 reports;
> - every laptop TX frame that provably never arrived is AC_BE:
> 89 ARP requests (no IP header -> BE), 13 unicast ARP replies,
> 83 DNS queries (0 answers), 2 TCP SYNs (no SYN-ACK), mDNS/LLMNR;
> - sizes overlap the wrong way for a size theory: BE frames of 42-201 B
> all fail, VO frames of 54-345 B all pass.
>
> So the wedge looks like "AC_BE TX dead, AC_VO TX alive", RX intact, and
> it persists across disconnect/reconnect and across AP changes (so not
> per-association state), cleared only by a reboot. It happened with
> power save fully off.
>
> As far as I can tell from reading the driver, BE and VO take different
> paths (separate PCIe TX rings per AC, and on 8821C BE/BK on the LOW
> TX-FIFO queue vs VO/VI on NORMAL), so a per-AC stall - e.g. the BE ring's
> mac80211 queue left stopped, or the LOW queue paused/page-starved - would
> match what I see (BE frames silently aged out, VO flowing, no driver
> message). Does that sound plausible to you, or is there a more likely
> place for a per-AC TX stall on this chip? The monitor-mode capture
> should at least show whether BE frames reach the air at all.
>
> When it wedges, besides the air capture, I plan to dump before rebooting:
> /sys/kernel/debug/ieee80211/phy0/queues (per-hw-queue stop reasons),
> stations/<ap>/aqm (per-TID backlog), and via rtw88 debugfs read_reg:
> REG_TXPAUSE (0x522), REG_TXDMA_STATUS (0x210), REG_FIFOPAGE_INFO_2/3
> (0x234/0x238) and the BE/VO TXBD indices (0x3A8/0x3A0), compared with a
> healthy baseline. If there are better registers or a debugfs page for
> the BE ring / LOW queue state on 8821C, please tell me and I'll dump
> those instead.
>
Have a look at RTK_PCI_HISR0 and RTK_PCI_HISR1 too.
>> Please fully turn off power save when you do the tests to reduce one
>> factor that can possibly cause TX slowly or stuck.
>
> Will do - power save stays off for all further tests.
>
> Thanks a lot for looking at this.
^ permalink raw reply [flat|nested] 13+ messages in thread
* RE: [BUG] rtw88 8821ce: connection wedges (100% loss until reboot) with station power save
2026-09-23 15:00 ` abkarada
@ 2026-09-24 2:40 ` Ping-Ke Shih
2026-10-02 23:11 ` Abdurrahman Karadag
0 siblings, 1 reply; 13+ messages in thread
From: Ping-Ke Shih @ 2026-09-24 2:40 UTC (permalink / raw)
To: abkarada, linux-wireless@vger.kernel.org
abkarada <abdurrahmankaradag19@gmail.com> wrote:
> Yes, and I could not get any further than that. The BE ring's hardware
> read index stops moving while the host write index keeps advancing. To
> rule out a missed doorbell I taught the watchdog to rewrite the TXBD
> host index during a live stall and re-read it. Three independent
> occurrences, each pair taken immediately before and immediately after
> the rewrite:
>
> BE read index BE write index
> 2026-09-06 03:17 0x00b7 -> 0x00b7 0x00a2 -> 0x00a3
> 2026-09-14 17:48 0x00c9 -> 0x00c9 0x006c -> 0x006d
> 2026-09-14 19:09 0x009e -> 0x009e 0x0050 -> 0x0051
>
> The read index did not move in any of them, and it was not a momentary
> sample either: in the 19:09 capture it held 0x9e from the first register
> read to the last, minutes apart, while the write index went 0x2b -> 0x4c.
> Traffic stayed at 100% loss throughout.
>
> So rewriting the TXBD host index is not sufficient to recover the ring.
> I have asked for that patch to be dropped.
Can I say that touching host write index doesn't affect hw read index?
>
> > Hardware read index is 0x3c, and host write index is 0x3a.
> > So, it reaches the limit of stop queue.
>
> Agreed, and that dump is the late moment. avail_desc() is 1 there, so
> ieee80211_stop_queue() is exactly right and the driver is doing what it
> should.
>
> But both stalls I caught in September are from an earlier moment, before
> the ring fills:
>
> 17:48 0x3a8: 0x00c9004c -> 0x00c9004e -> 0x00c90053
> read index 0xc9 frozen, write index 0x4c -> 0x53, avail 124
> 19:09 0x3a8: 0x009e002b -> 0x009e002d -> 0x009e0034
> read index 0x9e frozen, write index 0x2b -> 0x34, avail 115
>
> The ring still has 124 and 115 free descriptors, the queue is not
> stopped, nothing has hit any limit - and traffic is already 100% lost,
> because the hardware has stopped consuming descriptors while the driver
> keeps writing them and ringing the doorbell. Nothing in the driver
> notices this. The watchdog runs, the link stays associated, RX keeps
> working, and there is no counter or log line anywhere that says the TX
> ring has not advanced.
>
> So the first gap is detection, and it happens before your dump.
>
> The second gap is what happens after it. Once the ring does fill and the
> queue is stopped, everything that could undo that sits behind the same
> frozen index. In rtw_pci_tx_isr():
>
> if (cur_rp >= ring->r.rp)
> count = cur_rp - ring->r.rp;
> else
> count = ring->r.len - (ring->r.rp - cur_rp);
>
> while (count--) {
> ...
> if (ring->queue_stopped &&
> avail_desc(ring->r.wp, rp_idx, ring->r.len) > 4) {
> q_map = skb_get_queue_mapping(skb);
> ieee80211_wake_queue(hw, q_map);
> ring->queue_stopped = false;
> }
> ...
> }
>
> ring->r.rp = cur_rp;
>
> The only ieee80211_wake_queue() in pci.c is inside that loop. As long as
> cur_rp remains equal to ring->r.rp, count stays zero, so the body is
> skipped and neither ring->r.rp nor the stopped queue is ever
> reconsidered. Meanwhile rtw_pci_tx_write() keeps consuming descriptors
> until avail_desc() < 2 and calls ieee80211_stop_queue().
>
> To be fair to the design: this is not a deadlock. The descriptors are
> still sitting in the ring, so the normal completion path can recover the
> queue - but only if the hardware resumes consuming them. When it does
> not, and in the three occurrences above it did not, there is nothing
> else: no timeout, no ring reset, no device reset. The interface stays
> dead until userspace does something about it.
As this is hardware get stuck, we might skip to discuss
ieee80211_stop_queue()/ieee80211_wake_queue() for this moment.
> 14:29:58 injection on. The driver's own counters, logged from the
> ISR, one second apart:
> r.rp=174 r.wp=175 avail=254 stopped=0
> r.rp=174 r.wp=172 avail=1 stopped=1
> mac80211 BE queue: 0x1, IEEE80211_QUEUE_STOP_REASON_DRIVER
> 14:30:00 ping to the gateway: 100% loss
> 14:30:30 injection turned OFF
> 14:32:30 r.rp=174 r.wp=172 avail=1 stopped=1, queue still 0x1,
> 120 s later, with no recovery
>
[snip... Since I don't quit understand this test after I read twice.
Maybe I can read it again when I have free time]
> So I am not asking you to treat the stall itself as a driver bug. I am
> asking about the two gaps around it: the driver currently has no
> mechanism to detect that a ring has stopped advancing, and no way back
> if the hardware does not resume on its own. Would you consider a patch
> that addresses those - per-ring detection of "read index has not moved
> for a few watchdog rounds while descriptors are pending", triggering the
> recovery the driver already has for a firmware crash,
> rtw_fw_recovery() -> fw_recovery_work -> ieee80211_restart_hw()?
>
> In the field a full interface restart is what recovers the link in both
> occurrences where I got that far; on 2026-09-14 19:09 it came back 3 s
> after a down/up and reassociated without a reboot. __fw_recovery_work()
> does firmware-crash-specific work, so this would need a lighter variant,
> and I would rather hear your view on the direction than send something
> you do not want.
As your subject "100% loss until reboot", I can't say if this can help.
But here you mentioned "it came back 3 s after a down/up...".
Can you trigger the recovery to see if it can resolve the stuck?
>
> If you would rather not add a recovery path at all, the detection half
> is still worth something on its own: a log line at the moment a ring
> stops advancing, instead of a user discovering that the network died. It
> would also have saved me most of the last two months, and it would make
> the next report of this kind arrive with the register state already in
> it.
Did you mean detection stuck + recovery is the new proposal you want to
do? I think this can be a candidate solution.
Before that, can you summarize the methods that can resolve the stuck?
1. if up/down?
2. recovery?
3. (X) write host write index
4. ... (more)
By the way, recently people want to disable deep LPS and ASMP for this
chip, because they encountered hard system freezes [1], which they did
disable_lps_deep=y and disable_aspm=y before.
Can you also try the settings on your platform?
[1] https://lore.kernel.org/linux-wireless/20260918232801.119348-1-eexto@aol.com/
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [BUG] rtw88 8821ce: connection wedges (100% loss until reboot) with station power save
2026-09-24 2:40 ` Ping-Ke Shih
@ 2026-10-02 23:11 ` Abdurrahman Karadag
2026-10-05 1:42 ` Ping-Ke Shih
0 siblings, 1 reply; 13+ messages in thread
From: Abdurrahman Karadag @ 2026-10-02 23:11 UTC (permalink / raw)
To: pkshih, linux-wireless; +Cc: rtl8821cerfe2
> Can you trigger the recovery to see if it can resolve the stuck?
Not with the stock driver. The restart itself completes - the chip is
reinitialised and the station reassociates - but a queue that was
already stopped keeps its stop reason across the ring reset, so TX
never resumes. That is a driver bug and it is fixed in the first of the
two patches I have just sent; thank you for asking the question that
found it.
I could produce a stalled ring after all. My own 08-28 capture had it
and I had dismissed it: with REG_TXPAUSE at 0xff the BE read index sat
at 0x9a while the write index ran 0x12 -> 0x1b. Pausing TX in hardware
freezes the read index while the driver keeps submitting, which is the
same shape the chip shows when it wedges on its own. Not the same cause
- in the field TXPAUSE was 0x00 - but the same ring state, which is
what your question needed.
From that state, with the stock driver:
REG_TXPAUSE <- 0x0f, push traffic
BE 0x3a8: 0x0081007f, avail_desc() 1, BE stop reason 0x1, in 1 s
rewrite the host write index
0x0081007f -> 0x0081007f, no effect
call rtw_fw_recovery()
"firmware crash, start reset and recover"
"ieee80211 phy1: Hardware restart was requested"
REG_TXPAUSE back to 0x00, BE 0x3a8: 0x00000000, ring empty
wlan0: associated (twice, within the 90 s)
BE stop reason still 0x1, 100% loss, no recovery in 90 s
So the link came back and the ring was reset, and the queue stayed
stopped. ieee80211_wake_queue() is reached only from the completion
loop in rtw_pci_tx_isr(), and ring->queue_stopped is cleared only
there. The reset drops the pending descriptors and frees their skbs
through rtw_pci_free_tx_ring_skbs(), so that loop never runs for them;
the flag and the mac80211 stop reason then sit on an empty ring that
will never complete anything again.
Patch 1 records the queue mappings stopped by the ring-full path and
releases them, both from the normal completion and when the reset
empties the ring. Same script, same procedure, only the patch differs:
without patch with patch
BE queue stopped after 1 s 1 s
doorbell rewrite no effect no effect
rtw_fw_recovery() ran ran
BE stop reason after 0x1 0x0
traffic none in 90 s back within 2 s
> As this is hardware get stuck, we might skip to discuss
> ieee80211_stop_queue()/ieee80211_wake_queue() for this moment.
I did drop it, and then the test you asked for landed in exactly that
code. I want to be clear about the scope though: this is not why the
hardware stops, and it is not what killed the five events below, four
of which never reached the stop threshold at all. It only means that
once a queue has been stopped, no restart can bring it back. So it
matters for any recovery, not for the stall itself.
> Can I say that touching host write index doesn't affect hw read index?
In the stalled state, yes. On a healthy ring the write index is exactly
what makes the read index advance, so not in general. What I rewrote
was the value already there, through the same register
rtw_pci_tx_kick_off_queue() uses. It had no effect in the three field
events where I tried it, and none in the controlled stall above either.
> Before that, can you summarize the methods that can resolve the stuck?
First a correction: there are five events, not three, across three
boots, all with power save off and REG_TXPAUSE at 0x00.
2026-08-28 20:32 BE 0x45 frozen, wp 0x19 -> 0x22 queue open
2026-08-28 22:05 BE 0x3c frozen, wp 0x3a queue stopped
2026-09-06 03:17 BE 0xb7 frozen, wp 0x84 -> 0x9e queue open
2026-09-14 17:48 BE 0xc9 frozen, wp 0x4c -> 0x68 queue open
2026-09-14 19:09 BE 0x9e frozen, wp 0x2b -> 0x4c queue open
The second line is your dump. "Hardware read index is 0x3c, and host
write index is 0x3a" came from my own 08-28 capture, 93 minutes after
the 20:32 one in the same boot, and avail_desc() is 1 there so the
queue was stopped exactly as it should be. I replied as if your dump
were a different case and said all my events were from before the ring
fills. That was wrong, and it is the state patch 1 is about.
The ladder stops at the first step that works, so the denominators are
not comparable:
1. Rewrite host write index 0/3 no effect
2. ip neigh flush 0/4 no effect
3. iwd reconnect (new PTK) 2/4 fixed 08-28 20:32, 09-06 03:17
4. ip link set wlan0 down/up 1/2 fixed 09-14 17:48
5. Reboot always
6. rtw_fw_recovery() 0/1 measured above, before patch 1
I previously told you an interface restart is what recovers this. The
data does not support that: a reconnect cleared it twice and a down/up
once, and the ordering gave the reconnect the earlier attempts. And on
09-14 19:09 my ladder printed a timeout at step 4 only because it
waited 25 s; the journal shows the link up 3 s after the down/up and
reassociated 90 s later, no reboot. The subject line's "until reboot"
describes what a user sees, but three of these recovered without one.
The debugfs fw_crash knob, by the way, does nothing on this chip. The
write only pokes REG_HRCV_MSG; recovery is entered later, and only if
the firmware raises the C2H interrupt and rtw_fw_c2h_cmd_isr() sees
REG_MCU_TST_CFG == VAL_FW_TRIGGER. Here that never happens, so I called
rtw_fw_recovery() directly from a debug build instead.
> By the way, recently people want to disable deep LPS and ASMP [...]
> Can you also try the settings on your platform?
Both have been in /etc/modprobe.d since 2026-08-24, before every event
above. I should be careful though: my captures do not record module
parameters, so I cannot prove from them what was in effect. What they
do record is power save off, "IPS/ Low Power/ PS mode = 0/ 0/ 0", and
REG_TXPAUSE = 0x00 in every sample. And on the one occasion I dumped
sysfs, disable_aspm=Y had not changed the link's L0s/L1 state, so I
would not claim ASPM was actually off.
So: not prevented by disable_lps_deep=y, and nothing was pausing TX in
hardware. I cannot say ASPM was off.
> Have a look at RTK_PCI_HISR0 and RTK_PCI_HISR1 too.
Added, 0x0b4 and 0x0bc, twice two seconds apart next to the TXBD
indices. If the TX-done bit is asserted and stays asserted while the
read index does not move, that says something quite different about
where this sits. Thank you - I would not have thought to look.
Patch 2 is the detection half of the patch I withdrew, without the
doorbell rewrite. It is only a log line; I am not proposing a recovery
for the stall itself, because the table above does not single one out.
^ permalink raw reply [flat|nested] 13+ messages in thread
* RE: [BUG] rtw88 8821ce: connection wedges (100% loss until reboot) with station power save
2026-10-02 23:11 ` Abdurrahman Karadag
@ 2026-10-05 1:42 ` Ping-Ke Shih
0 siblings, 0 replies; 13+ messages in thread
From: Ping-Ke Shih @ 2026-10-05 1:42 UTC (permalink / raw)
To: Abdurrahman Karadag, linux-wireless@vger.kernel.org
Cc: rtl8821cerfe2@gmail.com
Abdurrahman Karadag <abdurrahmankaradag19@gmail.com> wrote:
> > Can you trigger the recovery to see if it can resolve the stuck?
>
> Not with the stock driver. The restart itself completes - the chip is
> reinitialised and the station reassociates - but a queue that was
> already stopped keeps its stop reason across the ring reset, so TX
> never resumes. That is a driver bug and it is fixed in the first of the
> two patches I have just sent; thank you for asking the question that
> found it.
>
> I could produce a stalled ring after all. My own 08-28 capture had it
> and I had dismissed it: with REG_TXPAUSE at 0xff the BE read index sat
> at 0x9a while the write index ran 0x12 -> 0x1b. Pausing TX in hardware
> freezes the read index while the driver keeps submitting, which is the
> same shape the chip shows when it wedges on its own. Not the same cause
> - in the field TXPAUSE was 0x00 - but the same ring state, which is
> what your question needed.
Can you dig why REG_TXPAUSE becomes 0xff? I wonder this is the cause
PCI TX gets stuck?
>
> From that state, with the stock driver:
>
> REG_TXPAUSE <- 0x0f, push traffic
> BE 0x3a8: 0x0081007f, avail_desc() 1, BE stop reason 0x1, in 1 s
> rewrite the host write index
> 0x0081007f -> 0x0081007f, no effect
> call rtw_fw_recovery()
> "firmware crash, start reset and recover"
> "ieee80211 phy1: Hardware restart was requested"
> REG_TXPAUSE back to 0x00, BE 0x3a8: 0x00000000, ring empty
> wlan0: associated (twice, within the 90 s)
> BE stop reason still 0x1, 100% loss, no recovery in 90 s
>
> So the link came back and the ring was reset, and the queue stayed
> stopped. ieee80211_wake_queue() is reached only from the completion
> loop in rtw_pci_tx_isr(), and ring->queue_stopped is cleared only
> there. The reset drops the pending descriptors and frees their skbs
> through rtw_pci_free_tx_ring_skbs(), so that loop never runs for them;
> the flag and the mac80211 stop reason then sit on an empty ring that
> will never complete anything again.
>
> Patch 1 records the queue mappings stopped by the ring-full path and
> releases them, both from the normal completion and when the reset
> empties the ring. Same script, same procedure, only the patch differs:
>
> without patch with patch
> BE queue stopped after 1 s 1 s
> doorbell rewrite no effect no effect
> rtw_fw_recovery() ran ran
> BE stop reason after 0x1 0x0
> traffic none in 90 s back within 2 s
I feel patch 1 can only do ieee80211_wake_queues() instead of iterative
each queue by ieee80211_wake_queue().
By the way, the detail is good, but could you please give summarize the
causes you found and the solutions you adopted? This will be easier
for me to understand this quickly.
>
> > As this is hardware get stuck, we might skip to discuss
> > ieee80211_stop_queue()/ieee80211_wake_queue() for this moment.
>
> I did drop it, and then the test you asked for landed in exactly that
> code. I want to be clear about the scope though: this is not why the
> hardware stops, and it is not what killed the five events below, four
> of which never reached the stop threshold at all. It only means that
> once a queue has been stopped, no restart can bring it back. So it
> matters for any recovery, not for the stall itself.
Before rtw_fw_recovery() isn't a solution, queues will not be relevant,
right? It looks like the recovery missed to consider doing wake up all
queues if any queue stopped.
> > Before that, can you summarize the methods that can resolve the stuck?
>
> First a correction: there are five events, not three, across three
> boots, all with power save off and REG_TXPAUSE at 0x00.
I'm confused. REG_TXPAUSE is 0x00 for below cases? Or it becomes 0xff?
>
> 2026-08-28 20:32 BE 0x45 frozen, wp 0x19 -> 0x22 queue open
> 2026-08-28 22:05 BE 0x3c frozen, wp 0x3a queue stopped
> 2026-09-06 03:17 BE 0xb7 frozen, wp 0x84 -> 0x9e queue open
> 2026-09-14 17:48 BE 0xc9 frozen, wp 0x4c -> 0x68 queue open
> 2026-09-14 19:09 BE 0x9e frozen, wp 0x2b -> 0x4c queue open
>
> The second line is your dump. "Hardware read index is 0x3c, and host
> write index is 0x3a" came from my own 08-28 capture, 93 minutes after
> the 20:32 one in the same boot, and avail_desc() is 1 there so the
> queue was stopped exactly as it should be. I replied as if your dump
> were a different case and said all my events were from before the ring
> fills. That was wrong, and it is the state patch 1 is about.
>
> The ladder stops at the first step that works, so the denominators are
> not comparable:
>
> 1. Rewrite host write index 0/3 no effect
> 2. ip neigh flush 0/4 no effect
> 3. iwd reconnect (new PTK) 2/4 fixed 08-28 20:32, 09-06 03:17
> 4. ip link set wlan0 down/up 1/2 fixed 09-14 17:48
> 5. Reboot always
> 6. rtw_fw_recovery() 0/1 measured above, before patch 1
>
> I previously told you an interface restart is what recovers this. The
> data does not support that: a reconnect cleared it twice and a down/up
> once, and the ordering gave the reconnect the earlier attempts. And on
> 09-14 19:09 my ladder printed a timeout at step 4 only because it
> waited 25 s; the journal shows the link up 3 s after the down/up and
> reassociated 90 s later, no reboot. The subject line's "until reboot"
> describes what a user sees, but three of these recovered without one.
>
> The debugfs fw_crash knob, by the way, does nothing on this chip. The
> write only pokes REG_HRCV_MSG; recovery is entered later, and only if
> the firmware raises the C2H interrupt and rtw_fw_c2h_cmd_isr() sees
> REG_MCU_TST_CFG == VAL_FW_TRIGGER. Here that never happens, so I called
> rtw_fw_recovery() directly from a debug build instead.
The debufs fw_crash only can take effect on RTL8822C.
>
> > By the way, recently people want to disable deep LPS and ASMP [...]
> > Can you also try the settings on your platform?
>
> Both have been in /etc/modprobe.d since 2026-08-24, before every event
> above. I should be careful though: my captures do not record module
> parameters, so I cannot prove from them what was in effect. What they
> do record is power save off, "IPS/ Low Power/ PS mode = 0/ 0/ 0", and
> REG_TXPAUSE = 0x00 in every sample. And on the one occasion I dumped
> sysfs, disable_aspm=Y had not changed the link's L0s/L1 state, so I
> would not claim ASPM was actually off.
>
> So: not prevented by disable_lps_deep=y, and nothing was pausing TX in
> hardware. I cannot say ASPM was off.
It is enough to make sure you have disable_aspm=y and disable_lps_deep=y,
and do cold reboot before testing.
^ permalink raw reply [flat|nested] 13+ messages in thread
end of thread, other threads:[~2026-10-05 1:42 UTC | newest]
Thread overview: 13+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-26 16:25 [BUG] rtw88 8821ce: connection wedges (100% loss until reboot) with station power save Abdurrahman Karadag
2026-08-26 18:00 ` Abdurrahman Karadag
2026-08-28 3:57 ` Ping-Ke Shih
2026-08-28 3:37 ` Abdurrahman Karadag
2026-08-31 3:35 ` Ping-Ke Shih
2026-09-02 9:15 ` Abdurrahman Karadag
2026-09-06 4:09 ` Ping-Ke Shih
2026-09-23 15:00 ` abkarada
2026-09-24 2:40 ` Ping-Ke Shih
2026-10-02 23:11 ` Abdurrahman Karadag
2026-10-05 1:42 ` Ping-Ke Shih
2026-09-23 17:29 ` Bitterblue Smith
2026-08-28 3:48 ` Ping-Ke Shih
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.