From: abkarada <abdurrahmankaradag19@gmail.com>
To: linux-wireless@vger.kernel.org
Cc: pkshih@realtek.com
Subject: Re: [BUG] rtw88 8821ce: connection wedges (100% loss until reboot) with station power save
Date: Wed, 23 Sep 2026 18:00:17 +0300 [thread overview]
Message-ID: <20260923150017.181310-1-abdurrahmankaradag19@gmail.com> (raw)
In-Reply-To: <9d64e207c4434411981bc133ef820d7e@realtek.com>
Sorry for the long silence. I spent it on the monitor capture you asked
for. I set up a second machine in monitor mode and checked it against a
healthy link first: 531 QoS-data frames from my station, all TID 0, and
157 ACKs back from the AP, so the capture side works and would show me
whether the chip transmitted.
Then I waited two and a half weeks and the failure did not happen once.
I cannot tell you whether something masked it or whether it is simply
that rare. The rig stays in place, and that capture is the first thing
you will get if it happens again.
In the meantime I went back over the data I already had, and I think I
had been treating two separate moments as one.
> I wonder even you use driver of 7.0.12 (on 7.1.9). It will still happen.
> Can you 100% ensure 7.0.12 is fine?
I cannot, and it is not. My watchdog caught two stalls on 2026-09-14,
17:48 and 19:09, in one boot, both on the stock 7.0.12 kernel and its
own rtw88. I also withdraw the "12 episodes on the first 7.1.9 boot"
count from my earlier mail: it came from a log heuristic that counted
"DHCP lease acquired, then connectivity lost", and that boot logged 51
"deauthenticating by local choice (Reason: 3)" events, each followed by
a reassociation about a minute later. That is a separate regulatory
disconnect loop and it happens on 7.0.12 boots too. There is no
regression here. My good/bad version claim was never measured properly.
> As your experiments, the cause is hardware never reads TX buffer and
> gets stuck, right?
Yes, and I could not get any further than that. The BE ring's hardware
read index stops moving while the host write index keeps advancing. To
rule out a missed doorbell I taught the watchdog to rewrite the TXBD
host index during a live stall and re-read it. Three independent
occurrences, each pair taken immediately before and immediately after
the rewrite:
BE read index BE write index
2026-09-06 03:17 0x00b7 -> 0x00b7 0x00a2 -> 0x00a3
2026-09-14 17:48 0x00c9 -> 0x00c9 0x006c -> 0x006d
2026-09-14 19:09 0x009e -> 0x009e 0x0050 -> 0x0051
The read index did not move in any of them, and it was not a momentary
sample either: in the 19:09 capture it held 0x9e from the first register
read to the last, minutes apart, while the write index went 0x2b -> 0x4c.
Traffic stayed at 100% loss throughout.
So rewriting the TXBD host index is not sufficient to recover the ring.
I have asked for that patch to be dropped.
> I checked vendor driver. It only does this at initial step, not to
> recover it at runtime.
Thank you for checking the vendor driver. That settles it: those bits
are an init-time thing, not a runtime recovery, so I have dropped the
idea.
> Which means not a driver side bug (stop queue but not restart queue
> properly), right?
Partly. I think we have been looking at two different moments of the
same failure, and there is a gap at each of them: the driver never
notices that a ring has stopped advancing, and it has no way back unless
the hardware resumes by itself. Neither gap is the cause. Both are
things the driver could do something about.
> Hardware read index is 0x3c, and host write index is 0x3a.
> So, it reaches the limit of stop queue.
Agreed, and that dump is the late moment. avail_desc() is 1 there, so
ieee80211_stop_queue() is exactly right and the driver is doing what it
should.
But both stalls I caught in September are from an earlier moment, before
the ring fills:
17:48 0x3a8: 0x00c9004c -> 0x00c9004e -> 0x00c90053
read index 0xc9 frozen, write index 0x4c -> 0x53, avail 124
19:09 0x3a8: 0x009e002b -> 0x009e002d -> 0x009e0034
read index 0x9e frozen, write index 0x2b -> 0x34, avail 115
The ring still has 124 and 115 free descriptors, the queue is not
stopped, nothing has hit any limit - and traffic is already 100% lost,
because the hardware has stopped consuming descriptors while the driver
keeps writing them and ringing the doorbell. Nothing in the driver
notices this. The watchdog runs, the link stays associated, RX keeps
working, and there is no counter or log line anywhere that says the TX
ring has not advanced.
So the first gap is detection, and it happens before your dump.
The second gap is what happens after it. Once the ring does fill and the
queue is stopped, everything that could undo that sits behind the same
frozen index. In rtw_pci_tx_isr():
if (cur_rp >= ring->r.rp)
count = cur_rp - ring->r.rp;
else
count = ring->r.len - (ring->r.rp - cur_rp);
while (count--) {
...
if (ring->queue_stopped &&
avail_desc(ring->r.wp, rp_idx, ring->r.len) > 4) {
q_map = skb_get_queue_mapping(skb);
ieee80211_wake_queue(hw, q_map);
ring->queue_stopped = false;
}
...
}
ring->r.rp = cur_rp;
The only ieee80211_wake_queue() in pci.c is inside that loop. As long as
cur_rp remains equal to ring->r.rp, count stays zero, so the body is
skipped and neither ring->r.rp nor the stopped queue is ever
reconsidered. Meanwhile rtw_pci_tx_write() keeps consuming descriptors
until avail_desc() < 2 and calls ieee80211_stop_queue().
To be fair to the design: this is not a deadlock. The descriptors are
still sitting in the ring, so the normal completion path can recover the
queue - but only if the hardware resumes consuming them. When it does
not, and in the three occurrences above it did not, there is nothing
else: no timeout, no ring reset, no device reset. The interface stays
dead until userspace does something about it.
Since the stall itself is rare and I cannot produce it on demand, I
built a way to put the driver into the stopped-queue state instead, so
that a recovery path could at least be tested. rtw_pci now takes a
debug-only module parameter that makes the TX completion handler behave
as if the read index never advanced, for a chosen ring - the same idea
as the existing fw_crash debugfs knob, which crashes the firmware
deliberately to exercise the recovery path. On 7.0.12, BE ring, station
otherwise idle:
14:29:58 injection on. The driver's own counters, logged from the
ISR, one second apart:
r.rp=174 r.wp=175 avail=254 stopped=0
r.rp=174 r.wp=172 avail=1 stopped=1
mac80211 BE queue: 0x1, IEEE80211_QUEUE_STOP_REASON_DRIVER
14:30:00 ping to the gateway: 100% loss
14:30:30 injection turned OFF
14:32:30 r.rp=174 r.wp=172 avail=1 stopped=1, queue still 0x1,
120 s later, with no recovery
The important observation is the final state: the injection is gone and
completions are no longer being suppressed, but nothing in the driver
revisits the stopped queue. By then the hardware may already have
consumed all descriptors that were submitted during the injection, so
there is no further TX completion to generate an interrupt and make the
driver observe the current hardware read index. With the mac80211 queue
stopped, no new descriptor is submitted either.
I should be precise about the limits of this. It is not your failure
reproduced. The field stall leaves real descriptors pending in the ring,
so if that hardware resumed, the driver would recover on its own; with
the injection, the hardware can consume the submitted descriptors while
the software read pointer is deliberately held back, leaving no later
completion event to make the driver revisit the stopped queue. What it
does give me is a deterministic way to reach that stopped-queue state,
which is what I would need to test any recovery patch - I did not want
to propose one I had no way of exercising.
In an earlier run of the same injection I also captured on wlan0 while
it was wedged. Once the ring is full, frames are either dropped in
rtw_pci_tx_write_data() with -ENOSPC or held in the qdisc; either way
they do not reach the air, but the ones that reach the driver are still
visible to tcpdump. So from userspace it looks like this: ARP requests
to the gateway repeat once a second and are never answered, while the
station still receives ARP requests from another host on the same
subnet. That is exactly what I see in the field.
So I am not asking you to treat the stall itself as a driver bug. I am
asking about the two gaps around it: the driver currently has no
mechanism to detect that a ring has stopped advancing, and no way back
if the hardware does not resume on its own. Would you consider a patch
that addresses those - per-ring detection of "read index has not moved
for a few watchdog rounds while descriptors are pending", triggering the
recovery the driver already has for a firmware crash,
rtw_fw_recovery() -> fw_recovery_work -> ieee80211_restart_hw()?
In the field a full interface restart is what recovers the link in both
occurrences where I got that far; on 2026-09-14 19:09 it came back 3 s
after a down/up and reassociated without a reboot. __fw_recovery_work()
does firmware-crash-specific work, so this would need a lighter variant,
and I would rather hear your view on the direction than send something
you do not want.
If you would rather not add a recovery path at all, the detection half
is still worth something on its own: a log line at the moment a ring
stops advancing, instead of a user discovering that the network died. It
would also have saved me most of the last two months, and it would make
the next report of this kind arrive with the register state already in
it.
Full dumps, the injection patch and the test script are available on
request.
next prev parent reply other threads:[~2026-09-23 15:00 UTC|newest]
Thread overview: 13+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-26 16:25 [BUG] rtw88 8821ce: connection wedges (100% loss until reboot) with station power save Abdurrahman Karadag
2026-08-26 18:00 ` Abdurrahman Karadag
2026-08-28 3:57 ` Ping-Ke Shih
2026-08-28 3:37 ` Abdurrahman Karadag
2026-08-31 3:35 ` Ping-Ke Shih
2026-09-02 9:15 ` Abdurrahman Karadag
2026-09-06 4:09 ` Ping-Ke Shih
2026-09-23 15:00 ` abkarada [this message]
2026-09-24 2:40 ` Ping-Ke Shih
2026-10-02 23:11 ` Abdurrahman Karadag
2026-10-05 1:42 ` Ping-Ke Shih
2026-09-23 17:29 ` Bitterblue Smith
2026-08-28 3:48 ` Ping-Ke Shih
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260923150017.181310-1-abdurrahmankaradag19@gmail.com \
--to=abdurrahmankaradag19@gmail.com \
--cc=linux-wireless@vger.kernel.org \
--cc=pkshih@realtek.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox