From: Vladislav Rysin <vrysin@gmail.com>
To: lucid_duck@justthetip.ca
Cc: linux-wireless@vger.kernel.org, nbd@nbd.name
Subject: Re: [BUG] mt7925e/MT7927: RX page_pool buffers stamped every 16 bytes, corrupting skb_shared_info->frag_list
Date: Thu, 10 Sep 2026 14:00:55 -0400 [thread overview]
Message-ID: <20260910180055.100698-1-vrysin@gmail.com> (raw)
In-Reply-To: <20260910134748.75684-1-vrysin@gmail.com>
Devin,
The rxdmad_c loop is the most useful thing anyone has put in this thread.
I could not have found it -- the DKMS source went with the package and I
have no tree on this box any more.
dma.c:184-185 for (i = 0; i < q->ndesc; i++)
dmad[i].data3 = cpu_to_le32(data3);
A bounded loop writing only +12 across ndesc 16-byte records is precisely
the footprint, and it changes what I think I am looking at. Two things
follow.
First, it explains something I could not: the partial pages. 219/256 and
204/256 do not fit continuous hardware DMA, which has no reason to stop
mid-page. They fit a loop with a bounded count landing on memory that is
not the ring it was written for.
Second, the timing. The fourth panic was taken inside a reset:
mt7925e: Message 00020016 (seq 4) timeout
Workqueue: mt76 mt7925_mac_reset_work [mt7925_common]
mt7925e_mac_reset -> __local_bh_enable_ip -> do_softirq
-> net_rx_action -> skb_defer_free_flush -> kfree_skb_list_reason
and the three earlier chip re-inits in my logs each followed an MCU
timeout too. So the question I would ask the source, if I still had it,
is whether any queue-init or fill loop on the reset path can run with
q->desc or q->ndesc describing something other than the coherent ring it
was set up for -- a queue re-initialised while its descriptor pointer
still refers to, or has been reassigned to, page_pool memory. The
corruption itself is already there before the flush; the reset is only
what walks into it.
The constant does not match, granted: 0x0080a624 against 0xf0000000, and
the flag gated to mt7996. So either a different writer with the same
shape, or the same shape with data3 coming from somewhere else. Worth
checking whether anything else in the tree writes a +12 word across a
descriptor array, including on paths shared with 7925.
On your card -- that is the part of this I cannot do any more. Mine is out
of the machine. If you can reproduce it, the dump I never managed to get
is one taken with
core_collector makedumpfile -l --message-level 7 -d 1
Mine were -d 31, which excluded the sk_buff slab page, so I could only
ever inspect the data page. With -d 1 the skb itself survives, and so do
the module pages, which would let you walk the queue state -- q->desc,
q->ndesc, the page_pool -- at the moment of the fault rather than
inferring it from a data page as I had to.
What reproduced it here, for what it is worth:
traffic-dependent, not timed: 22h34m, 27h55m, 4h53m, 8h50m of uptime
plain VHT80, 5745 MHz, 866 Mbit/s, NSS 2 -- never EHT or 320 MHz
multi-AP mesh, frequent roams and Reason 2 deauths
Wi-Fi power save made no difference, on or off
reproduced on 7.1.10 and 7.1.12 with mediatek-mt7927-dkms 2.14, and on
7.2.3 in-tree with the kernel Not tainted
Everything is here:
https://drive.google.com/drive/folders/1-NWRQVhtiIBTeNFLe4QukUyrVCM23YDt?usp=sharing
127.0.0.1-2026-09-03-20:42:25/ 1.5 GB 7.1.12, DKMS 2.14
127.0.0.1-2026-09-04-11:09:39/ 1.4 GB 7.2.3, in-tree, Not tainted
127.0.0.1-2026-09-05-02:16:37/ 1.3 GB 7.2.3, in-tree, the reset one
analyze.py drgn script
Each directory has vmcore, vmcore-dmesg.txt and kexec-dmesg.log. The
script prints the stack, the frag_list bytes, the stamped-record count and
the page map from any of the three; it wants kernel-debuginfo matching the
dump's kernel, and the install line for the 7.2.3 vanilla build is in its
docstring. crash 9.0.1 cannot open these at all -- it dies with "invalid
structure member offset: kmem_cache_s_num" on the 7.2 slab layout, which
is why the script is drgn.
All three are -d 31, so the sk_buff page is missing from every one of them.
Thanks,
Vladislav Rysin
prev parent reply other threads:[~2026-09-10 18:02 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-06 16:27 [BUG] mt7925e/MT7927: RX page_pool buffers stamped every 16 bytes, corrupting skb_shared_info->frag_list -> #SS in kfree_skb_list_reason Vladislav Rysin
2026-09-10 4:03 ` [BUG] mt7925e/MT7927: RX page_pool buffers stamped every 16 bytes, corrupting skb_shared_info->frag_list Devin Wittmayer
2026-09-10 13:47 ` Vladislav Rysin
2026-09-10 16:24 ` Devin Wittmayer
2026-09-10 18:00 ` Vladislav Rysin [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260910180055.100698-1-vrysin@gmail.com \
--to=vrysin@gmail.com \
--cc=linux-wireless@vger.kernel.org \
--cc=lucid_duck@justthetip.ca \
--cc=nbd@nbd.name \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.