From: Ping-Ke Shih <pkshih@realtek.com>
To: abkarada <abdurrahmankaradag19@gmail.com>,
"linux-wireless@vger.kernel.org" <linux-wireless@vger.kernel.org>
Subject: RE: [BUG] rtw88 8821ce: connection wedges (100% loss until reboot) with station power save
Date: Thu, 24 Sep 2026 02:40:41 +0000 [thread overview]
Message-ID: <fd4a810147fb4daeb42e5b8cae981b27@realtek.com> (raw)
In-Reply-To: <20260923150017.181310-1-abdurrahmankaradag19@gmail.com>
abkarada <abdurrahmankaradag19@gmail.com> wrote:
> Yes, and I could not get any further than that. The BE ring's hardware
> read index stops moving while the host write index keeps advancing. To
> rule out a missed doorbell I taught the watchdog to rewrite the TXBD
> host index during a live stall and re-read it. Three independent
> occurrences, each pair taken immediately before and immediately after
> the rewrite:
>
> BE read index BE write index
> 2026-09-06 03:17 0x00b7 -> 0x00b7 0x00a2 -> 0x00a3
> 2026-09-14 17:48 0x00c9 -> 0x00c9 0x006c -> 0x006d
> 2026-09-14 19:09 0x009e -> 0x009e 0x0050 -> 0x0051
>
> The read index did not move in any of them, and it was not a momentary
> sample either: in the 19:09 capture it held 0x9e from the first register
> read to the last, minutes apart, while the write index went 0x2b -> 0x4c.
> Traffic stayed at 100% loss throughout.
>
> So rewriting the TXBD host index is not sufficient to recover the ring.
> I have asked for that patch to be dropped.
Can I say that touching host write index doesn't affect hw read index?
>
> > Hardware read index is 0x3c, and host write index is 0x3a.
> > So, it reaches the limit of stop queue.
>
> Agreed, and that dump is the late moment. avail_desc() is 1 there, so
> ieee80211_stop_queue() is exactly right and the driver is doing what it
> should.
>
> But both stalls I caught in September are from an earlier moment, before
> the ring fills:
>
> 17:48 0x3a8: 0x00c9004c -> 0x00c9004e -> 0x00c90053
> read index 0xc9 frozen, write index 0x4c -> 0x53, avail 124
> 19:09 0x3a8: 0x009e002b -> 0x009e002d -> 0x009e0034
> read index 0x9e frozen, write index 0x2b -> 0x34, avail 115
>
> The ring still has 124 and 115 free descriptors, the queue is not
> stopped, nothing has hit any limit - and traffic is already 100% lost,
> because the hardware has stopped consuming descriptors while the driver
> keeps writing them and ringing the doorbell. Nothing in the driver
> notices this. The watchdog runs, the link stays associated, RX keeps
> working, and there is no counter or log line anywhere that says the TX
> ring has not advanced.
>
> So the first gap is detection, and it happens before your dump.
>
> The second gap is what happens after it. Once the ring does fill and the
> queue is stopped, everything that could undo that sits behind the same
> frozen index. In rtw_pci_tx_isr():
>
> if (cur_rp >= ring->r.rp)
> count = cur_rp - ring->r.rp;
> else
> count = ring->r.len - (ring->r.rp - cur_rp);
>
> while (count--) {
> ...
> if (ring->queue_stopped &&
> avail_desc(ring->r.wp, rp_idx, ring->r.len) > 4) {
> q_map = skb_get_queue_mapping(skb);
> ieee80211_wake_queue(hw, q_map);
> ring->queue_stopped = false;
> }
> ...
> }
>
> ring->r.rp = cur_rp;
>
> The only ieee80211_wake_queue() in pci.c is inside that loop. As long as
> cur_rp remains equal to ring->r.rp, count stays zero, so the body is
> skipped and neither ring->r.rp nor the stopped queue is ever
> reconsidered. Meanwhile rtw_pci_tx_write() keeps consuming descriptors
> until avail_desc() < 2 and calls ieee80211_stop_queue().
>
> To be fair to the design: this is not a deadlock. The descriptors are
> still sitting in the ring, so the normal completion path can recover the
> queue - but only if the hardware resumes consuming them. When it does
> not, and in the three occurrences above it did not, there is nothing
> else: no timeout, no ring reset, no device reset. The interface stays
> dead until userspace does something about it.
As this is hardware get stuck, we might skip to discuss
ieee80211_stop_queue()/ieee80211_wake_queue() for this moment.
> 14:29:58 injection on. The driver's own counters, logged from the
> ISR, one second apart:
> r.rp=174 r.wp=175 avail=254 stopped=0
> r.rp=174 r.wp=172 avail=1 stopped=1
> mac80211 BE queue: 0x1, IEEE80211_QUEUE_STOP_REASON_DRIVER
> 14:30:00 ping to the gateway: 100% loss
> 14:30:30 injection turned OFF
> 14:32:30 r.rp=174 r.wp=172 avail=1 stopped=1, queue still 0x1,
> 120 s later, with no recovery
>
[snip... Since I don't quit understand this test after I read twice.
Maybe I can read it again when I have free time]
> So I am not asking you to treat the stall itself as a driver bug. I am
> asking about the two gaps around it: the driver currently has no
> mechanism to detect that a ring has stopped advancing, and no way back
> if the hardware does not resume on its own. Would you consider a patch
> that addresses those - per-ring detection of "read index has not moved
> for a few watchdog rounds while descriptors are pending", triggering the
> recovery the driver already has for a firmware crash,
> rtw_fw_recovery() -> fw_recovery_work -> ieee80211_restart_hw()?
>
> In the field a full interface restart is what recovers the link in both
> occurrences where I got that far; on 2026-09-14 19:09 it came back 3 s
> after a down/up and reassociated without a reboot. __fw_recovery_work()
> does firmware-crash-specific work, so this would need a lighter variant,
> and I would rather hear your view on the direction than send something
> you do not want.
As your subject "100% loss until reboot", I can't say if this can help.
But here you mentioned "it came back 3 s after a down/up...".
Can you trigger the recovery to see if it can resolve the stuck?
>
> If you would rather not add a recovery path at all, the detection half
> is still worth something on its own: a log line at the moment a ring
> stops advancing, instead of a user discovering that the network died. It
> would also have saved me most of the last two months, and it would make
> the next report of this kind arrive with the register state already in
> it.
Did you mean detection stuck + recovery is the new proposal you want to
do? I think this can be a candidate solution.
Before that, can you summarize the methods that can resolve the stuck?
1. if up/down?
2. recovery?
3. (X) write host write index
4. ... (more)
By the way, recently people want to disable deep LPS and ASMP for this
chip, because they encountered hard system freezes [1], which they did
disable_lps_deep=y and disable_aspm=y before.
Can you also try the settings on your platform?
[1] https://lore.kernel.org/linux-wireless/20260918232801.119348-1-eexto@aol.com/
next prev parent reply other threads:[~2026-09-24 2:40 UTC|newest]
Thread overview: 13+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-26 16:25 [BUG] rtw88 8821ce: connection wedges (100% loss until reboot) with station power save Abdurrahman Karadag
2026-08-26 18:00 ` Abdurrahman Karadag
2026-08-28 3:57 ` Ping-Ke Shih
2026-08-28 3:37 ` Abdurrahman Karadag
2026-08-31 3:35 ` Ping-Ke Shih
2026-09-02 9:15 ` Abdurrahman Karadag
2026-09-06 4:09 ` Ping-Ke Shih
2026-09-23 15:00 ` abkarada
2026-09-24 2:40 ` Ping-Ke Shih [this message]
2026-10-02 23:11 ` Abdurrahman Karadag
2026-10-05 1:42 ` Ping-Ke Shih
2026-09-23 17:29 ` Bitterblue Smith
2026-08-28 3:48 ` Ping-Ke Shih
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=fd4a810147fb4daeb42e5b8cae981b27@realtek.com \
--to=pkshih@realtek.com \
--cc=abdurrahmankaradag19@gmail.com \
--cc=linux-wireless@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox