All of lore.kernel.org
 help / color / mirror / Atom feed
From: Ping-Ke Shih <pkshih@realtek.com>
To: abkarada <abdurrahmankaradag19@gmail.com>,
	"linux-wireless@vger.kernel.org" <linux-wireless@vger.kernel.org>
Subject: RE: [BUG] rtw88 8821ce: connection wedges (100% loss until reboot) with station power save
Date: Thu, 24 Sep 2026 02:40:41 +0000	[thread overview]
Message-ID: <fd4a810147fb4daeb42e5b8cae981b27@realtek.com> (raw)
In-Reply-To: <20260923150017.181310-1-abdurrahmankaradag19@gmail.com>

abkarada <abdurrahmankaradag19@gmail.com> wrote:
> Yes, and I could not get any further than that. The BE ring's hardware
> read index stops moving while the host write index keeps advancing. To
> rule out a missed doorbell I taught the watchdog to rewrite the TXBD
> host index during a live stall and re-read it. Three independent
> occurrences, each pair taken immediately before and immediately after
> the rewrite:
> 
>                      BE read index      BE write index
>   2026-09-06 03:17   0x00b7 -> 0x00b7   0x00a2 -> 0x00a3
>   2026-09-14 17:48   0x00c9 -> 0x00c9   0x006c -> 0x006d
>   2026-09-14 19:09   0x009e -> 0x009e   0x0050 -> 0x0051
> 
> The read index did not move in any of them, and it was not a momentary
> sample either: in the 19:09 capture it held 0x9e from the first register
> read to the last, minutes apart, while the write index went 0x2b -> 0x4c.
> Traffic stayed at 100% loss throughout.
> 
> So rewriting the TXBD host index is not sufficient to recover the ring.
> I have asked for that patch to be dropped.

Can I say that touching host write index doesn't affect hw read index?

> 
> > Hardware read index is 0x3c, and host write index is 0x3a.
> > So, it reaches the limit of stop queue.
> 
> Agreed, and that dump is the late moment. avail_desc() is 1 there, so
> ieee80211_stop_queue() is exactly right and the driver is doing what it
> should.
> 
> But both stalls I caught in September are from an earlier moment, before
> the ring fills:
> 
>   17:48  0x3a8: 0x00c9004c -> 0x00c9004e -> 0x00c90053
>          read index 0xc9 frozen, write index 0x4c -> 0x53, avail 124
>   19:09  0x3a8: 0x009e002b -> 0x009e002d -> 0x009e0034
>          read index 0x9e frozen, write index 0x2b -> 0x34, avail 115
> 
> The ring still has 124 and 115 free descriptors, the queue is not
> stopped, nothing has hit any limit - and traffic is already 100% lost,
> because the hardware has stopped consuming descriptors while the driver
> keeps writing them and ringing the doorbell. Nothing in the driver
> notices this. The watchdog runs, the link stays associated, RX keeps
> working, and there is no counter or log line anywhere that says the TX
> ring has not advanced.
> 
> So the first gap is detection, and it happens before your dump.
> 
> The second gap is what happens after it. Once the ring does fill and the
> queue is stopped, everything that could undo that sits behind the same
> frozen index. In rtw_pci_tx_isr():
> 
>         if (cur_rp >= ring->r.rp)
>                 count = cur_rp - ring->r.rp;
>         else
>                 count = ring->r.len - (ring->r.rp - cur_rp);
> 
>         while (count--) {
>                 ...
>                 if (ring->queue_stopped &&
>                     avail_desc(ring->r.wp, rp_idx, ring->r.len) > 4) {
>                         q_map = skb_get_queue_mapping(skb);
>                         ieee80211_wake_queue(hw, q_map);
>                         ring->queue_stopped = false;
>                 }
>                 ...
>         }
> 
>         ring->r.rp = cur_rp;
> 
> The only ieee80211_wake_queue() in pci.c is inside that loop. As long as
> cur_rp remains equal to ring->r.rp, count stays zero, so the body is
> skipped and neither ring->r.rp nor the stopped queue is ever
> reconsidered. Meanwhile rtw_pci_tx_write() keeps consuming descriptors
> until avail_desc() < 2 and calls ieee80211_stop_queue().
> 
> To be fair to the design: this is not a deadlock. The descriptors are
> still sitting in the ring, so the normal completion path can recover the
> queue - but only if the hardware resumes consuming them. When it does
> not, and in the three occurrences above it did not, there is nothing
> else: no timeout, no ring reset, no device reset. The interface stays
> dead until userspace does something about it.

As this is hardware get stuck, we might skip to discuss 
ieee80211_stop_queue()/ieee80211_wake_queue() for this moment.

>   14:29:58  injection on. The driver's own counters, logged from the
>             ISR, one second apart:
>               r.rp=174 r.wp=175 avail=254 stopped=0
>               r.rp=174 r.wp=172 avail=1   stopped=1
>             mac80211 BE queue: 0x1, IEEE80211_QUEUE_STOP_REASON_DRIVER
>   14:30:00  ping to the gateway: 100% loss
>   14:30:30  injection turned OFF
>   14:32:30  r.rp=174 r.wp=172 avail=1 stopped=1, queue still 0x1,
>             120 s later, with no recovery
> 

[snip... Since I don't quit understand this test after I read twice. 
Maybe I can read it again when I have free time]


> So I am not asking you to treat the stall itself as a driver bug. I am
> asking about the two gaps around it: the driver currently has no
> mechanism to detect that a ring has stopped advancing, and no way back
> if the hardware does not resume on its own. Would you consider a patch
> that addresses those - per-ring detection of "read index has not moved
> for a few watchdog rounds while descriptors are pending", triggering the
> recovery the driver already has for a firmware crash,
> rtw_fw_recovery() -> fw_recovery_work -> ieee80211_restart_hw()?
> 
> In the field a full interface restart is what recovers the link in both
> occurrences where I got that far; on 2026-09-14 19:09 it came back 3 s
> after a down/up and reassociated without a reboot. __fw_recovery_work()
> does firmware-crash-specific work, so this would need a lighter variant,
> and I would rather hear your view on the direction than send something
> you do not want.

As your subject "100% loss until reboot", I can't say if this can help.
But here you mentioned "it came back 3 s after a down/up...".

Can you trigger the recovery to see if it can resolve the stuck?

> 
> If you would rather not add a recovery path at all, the detection half
> is still worth something on its own: a log line at the moment a ring
> stops advancing, instead of a user discovering that the network died. It
> would also have saved me most of the last two months, and it would make
> the next report of this kind arrive with the register state already in
> it.

Did you mean detection stuck + recovery is the new proposal you want to
do? I think this can be a candidate solution. 

Before that, can you summarize the methods that can resolve the stuck?
1. if up/down?
2. recovery?
3. (X) write host write index
4. ... (more)

By the way, recently people want to disable deep LPS and ASMP for this
chip, because they encountered hard system freezes [1], which they did
disable_lps_deep=y and disable_aspm=y before.

Can you also try the settings on your platform?

[1] https://lore.kernel.org/linux-wireless/20260918232801.119348-1-eexto@aol.com/




  reply	other threads:[~2026-09-24  2:40 UTC|newest]

Thread overview: 13+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-26 16:25 [BUG] rtw88 8821ce: connection wedges (100% loss until reboot) with station power save Abdurrahman Karadag
2026-08-26 18:00 ` Abdurrahman Karadag
2026-08-28  3:57   ` Ping-Ke Shih
2026-08-28  3:37     ` Abdurrahman Karadag
2026-08-31  3:35       ` Ping-Ke Shih
2026-09-02  9:15         ` Abdurrahman Karadag
2026-09-06  4:09           ` Ping-Ke Shih
2026-09-23 15:00             ` abkarada
2026-09-24  2:40               ` Ping-Ke Shih [this message]
2026-10-02 23:11                 ` Abdurrahman Karadag
2026-10-05  1:42                   ` Ping-Ke Shih
2026-09-23 17:29       ` Bitterblue Smith
2026-08-28  3:48 ` Ping-Ke Shih

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=fd4a810147fb4daeb42e5b8cae981b27@realtek.com \
    --to=pkshih@realtek.com \
    --cc=abdurrahmankaradag19@gmail.com \
    --cc=linux-wireless@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.