From: Abdurrahman Karadag <abdurrahmankaradag19@gmail.com>
To: pkshih@realtek.com, linux-wireless@vger.kernel.org
Cc: rtl8821cerfe2@gmail.com
Subject: Re: [BUG] rtw88 8821ce: connection wedges (100% loss until reboot) with station power save
Date: Sat, 3 Oct 2026 02:11:09 +0300 [thread overview]
Message-ID: <20261002231109.12770-1-abdurrahmankaradag19@gmail.com> (raw)
In-Reply-To: <fd4a810147fb4daeb42e5b8cae981b27@realtek.com>
> Can you trigger the recovery to see if it can resolve the stuck?
Not with the stock driver. The restart itself completes - the chip is
reinitialised and the station reassociates - but a queue that was
already stopped keeps its stop reason across the ring reset, so TX
never resumes. That is a driver bug and it is fixed in the first of the
two patches I have just sent; thank you for asking the question that
found it.
I could produce a stalled ring after all. My own 08-28 capture had it
and I had dismissed it: with REG_TXPAUSE at 0xff the BE read index sat
at 0x9a while the write index ran 0x12 -> 0x1b. Pausing TX in hardware
freezes the read index while the driver keeps submitting, which is the
same shape the chip shows when it wedges on its own. Not the same cause
- in the field TXPAUSE was 0x00 - but the same ring state, which is
what your question needed.
From that state, with the stock driver:
REG_TXPAUSE <- 0x0f, push traffic
BE 0x3a8: 0x0081007f, avail_desc() 1, BE stop reason 0x1, in 1 s
rewrite the host write index
0x0081007f -> 0x0081007f, no effect
call rtw_fw_recovery()
"firmware crash, start reset and recover"
"ieee80211 phy1: Hardware restart was requested"
REG_TXPAUSE back to 0x00, BE 0x3a8: 0x00000000, ring empty
wlan0: associated (twice, within the 90 s)
BE stop reason still 0x1, 100% loss, no recovery in 90 s
So the link came back and the ring was reset, and the queue stayed
stopped. ieee80211_wake_queue() is reached only from the completion
loop in rtw_pci_tx_isr(), and ring->queue_stopped is cleared only
there. The reset drops the pending descriptors and frees their skbs
through rtw_pci_free_tx_ring_skbs(), so that loop never runs for them;
the flag and the mac80211 stop reason then sit on an empty ring that
will never complete anything again.
Patch 1 records the queue mappings stopped by the ring-full path and
releases them, both from the normal completion and when the reset
empties the ring. Same script, same procedure, only the patch differs:
without patch with patch
BE queue stopped after 1 s 1 s
doorbell rewrite no effect no effect
rtw_fw_recovery() ran ran
BE stop reason after 0x1 0x0
traffic none in 90 s back within 2 s
> As this is hardware get stuck, we might skip to discuss
> ieee80211_stop_queue()/ieee80211_wake_queue() for this moment.
I did drop it, and then the test you asked for landed in exactly that
code. I want to be clear about the scope though: this is not why the
hardware stops, and it is not what killed the five events below, four
of which never reached the stop threshold at all. It only means that
once a queue has been stopped, no restart can bring it back. So it
matters for any recovery, not for the stall itself.
> Can I say that touching host write index doesn't affect hw read index?
In the stalled state, yes. On a healthy ring the write index is exactly
what makes the read index advance, so not in general. What I rewrote
was the value already there, through the same register
rtw_pci_tx_kick_off_queue() uses. It had no effect in the three field
events where I tried it, and none in the controlled stall above either.
> Before that, can you summarize the methods that can resolve the stuck?
First a correction: there are five events, not three, across three
boots, all with power save off and REG_TXPAUSE at 0x00.
2026-08-28 20:32 BE 0x45 frozen, wp 0x19 -> 0x22 queue open
2026-08-28 22:05 BE 0x3c frozen, wp 0x3a queue stopped
2026-09-06 03:17 BE 0xb7 frozen, wp 0x84 -> 0x9e queue open
2026-09-14 17:48 BE 0xc9 frozen, wp 0x4c -> 0x68 queue open
2026-09-14 19:09 BE 0x9e frozen, wp 0x2b -> 0x4c queue open
The second line is your dump. "Hardware read index is 0x3c, and host
write index is 0x3a" came from my own 08-28 capture, 93 minutes after
the 20:32 one in the same boot, and avail_desc() is 1 there so the
queue was stopped exactly as it should be. I replied as if your dump
were a different case and said all my events were from before the ring
fills. That was wrong, and it is the state patch 1 is about.
The ladder stops at the first step that works, so the denominators are
not comparable:
1. Rewrite host write index 0/3 no effect
2. ip neigh flush 0/4 no effect
3. iwd reconnect (new PTK) 2/4 fixed 08-28 20:32, 09-06 03:17
4. ip link set wlan0 down/up 1/2 fixed 09-14 17:48
5. Reboot always
6. rtw_fw_recovery() 0/1 measured above, before patch 1
I previously told you an interface restart is what recovers this. The
data does not support that: a reconnect cleared it twice and a down/up
once, and the ordering gave the reconnect the earlier attempts. And on
09-14 19:09 my ladder printed a timeout at step 4 only because it
waited 25 s; the journal shows the link up 3 s after the down/up and
reassociated 90 s later, no reboot. The subject line's "until reboot"
describes what a user sees, but three of these recovered without one.
The debugfs fw_crash knob, by the way, does nothing on this chip. The
write only pokes REG_HRCV_MSG; recovery is entered later, and only if
the firmware raises the C2H interrupt and rtw_fw_c2h_cmd_isr() sees
REG_MCU_TST_CFG == VAL_FW_TRIGGER. Here that never happens, so I called
rtw_fw_recovery() directly from a debug build instead.
> By the way, recently people want to disable deep LPS and ASMP [...]
> Can you also try the settings on your platform?
Both have been in /etc/modprobe.d since 2026-08-24, before every event
above. I should be careful though: my captures do not record module
parameters, so I cannot prove from them what was in effect. What they
do record is power save off, "IPS/ Low Power/ PS mode = 0/ 0/ 0", and
REG_TXPAUSE = 0x00 in every sample. And on the one occasion I dumped
sysfs, disable_aspm=Y had not changed the link's L0s/L1 state, so I
would not claim ASPM was actually off.
So: not prevented by disable_lps_deep=y, and nothing was pausing TX in
hardware. I cannot say ASPM was off.
> Have a look at RTK_PCI_HISR0 and RTK_PCI_HISR1 too.
Added, 0x0b4 and 0x0bc, twice two seconds apart next to the TXBD
indices. If the TX-done bit is asserted and stays asserted while the
read index does not move, that says something quite different about
where this sits. Thank you - I would not have thought to look.
Patch 2 is the detection half of the patch I withdrew, without the
doorbell rewrite. It is only a log line; I am not proposing a recovery
for the stall itself, because the table above does not single one out.
next prev parent reply other threads:[~2026-10-02 23:11 UTC|newest]
Thread overview: 13+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-26 16:25 [BUG] rtw88 8821ce: connection wedges (100% loss until reboot) with station power save Abdurrahman Karadag
2026-08-26 18:00 ` Abdurrahman Karadag
2026-08-28 3:57 ` Ping-Ke Shih
2026-08-28 3:37 ` Abdurrahman Karadag
2026-08-31 3:35 ` Ping-Ke Shih
2026-09-02 9:15 ` Abdurrahman Karadag
2026-09-06 4:09 ` Ping-Ke Shih
2026-09-23 15:00 ` abkarada
2026-09-24 2:40 ` Ping-Ke Shih
2026-10-02 23:11 ` Abdurrahman Karadag [this message]
2026-10-05 1:42 ` Ping-Ke Shih
2026-09-23 17:29 ` Bitterblue Smith
2026-08-28 3:48 ` Ping-Ke Shih
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261002231109.12770-1-abdurrahmankaradag19@gmail.com \
--to=abdurrahmankaradag19@gmail.com \
--cc=linux-wireless@vger.kernel.org \
--cc=pkshih@realtek.com \
--cc=rtl8821cerfe2@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox