* [BUG] mt7925e/MT7927: RX page_pool buffers stamped every 16 bytes, corrupting skb_shared_info->frag_list -> #SS in kfree_skb_list_reason
@ 2026-09-06 16:27 Vladislav Rysin
2026-09-10 4:03 ` [BUG] mt7925e/MT7927: RX page_pool buffers stamped every 16 bytes, corrupting skb_shared_info->frag_list Devin Wittmayer
2026-09-10 13:47 ` Vladislav Rysin
0 siblings, 2 replies; 5+ messages in thread
From: Vladislav Rysin @ 2026-09-06 16:27 UTC (permalink / raw)
To: linux-wireless; +Cc: nbd
Hi,
I am hitting a reproducible kernel panic with an MT7927 (Filogic 380) card on
mt7925e. Four panics over five days, all with byte-identical registers. I have
analysed two vmcores and believe I have the mechanism, though not the culprit.
Hardware / software
-------------------
Dell Precision 5820 Tower, BIOS 2.48.0, 250 GB RAM
MEDIATEK MT7927 802.11be 2x2 [14c3:7927], subsystem 105b:e104
behind a dedicated root port (b2:00.0 -> b3:00.0), PCIe 8GT/s x1
(LnkSta matches LnkCap; UESta clean, no AER events)
driver: mt7925e, MT7927 raw CHIPID=0x0000, forcing chip=0x7927
firmware: linux-firmware-20260810
HW/SW Version 0x8a108a10, Build Time 20260414133750a
WM Firmware Version ____000000, Build Time 20260414134255
link at time of crashes: VHT80, 5745 MHz, 866.7 Mbit/s, VHT-NSS 2 (not EHT)
Reproduces on:
7.2.3 vanilla (in-tree mt7925e, kernel Not tainted) <- crashes 3 and 4
7.1.12 and 7.1.10 with the mediatek-mt7927-dkms 2.14 backport <- crashes 1, 2
Uptime at panic: 22h34m, 27h55m, 4h53m, 8h50m. Traffic-dependent, not timed.
The panic
---------
Oops: stack segment: 0000 [#1] SMP NOPTI
CPU: 6 UID: 0 PID: 808 Comm: napi/phy0-0 Kdump: loaded Not tainted
7.2.3-300.vanilla.fc44.x86_64
RIP: 0010:kfree_skb_list_reason+0xce/0x260
RBX: 0080a62400000000 RDI: 0080a62400000000 RBP: 0080a62400000000
Call Trace:
skb_release_data+0x184/0x220
napi_consume_skb+0x7a/0x160
skb_defer_free_flush+0x81/0xc0
napi_threaded_poll_loop+0x14e/0x290
napi_threaded_poll+0x42/0xb0
A second panic took the same fault from softirq context instead, during the
driver's own recovery, right after an MCU timeout:
mt7925e 0000:b3:00.0: Message 00020016 (seq 4) timeout
Workqueue: mt76 mt7925_mac_reset_work [mt7925_common]
mt7925e_mac_reset+0x20e/0x400 -> __local_bh_enable_ip -> do_softirq
-> net_rx_action -> skb_defer_free_flush -> kfree_skb_list_reason
So the flushing context is incidental; the skb is already corrupt on the
per-CPU defer list. The corruption happens at RX time.
What the vmcore shows
---------------------
Analysed with drgn 0.2.0 (crash 9.0.1 cannot open a 7.2 dump:
"invalid structure member offset: kmem_cache_s_num").
The RX data page carrying the skb is stamped: every 16 bytes, the dword at
offset +12 is overwritten with 24 a6 80 00 (LE 0x0080a624). 256 of 256
16-byte records in the page, and several other page_pool pages are fully
stamped as well (scattered, not contiguous).
The unstamped bytes are genuine payload -- local MAC, AP MAC and the local
IPv4 address are all plainly visible in the same page:
+0e80 00 00 66 ac f7 29 1b 46 f8 1a 2b 1a 24 a6 80 00
+0ea0 c0 a8 56 22 01 bb e3 fc 94 66 fb f5 24 a6 80 00
Why the register value is always identical:
- skb_shared_info sits at the buffer tail, page offset 0xec0 (fixed
q->buf_size); identical in both dumps
- frag_list is at shinfo+8
- the 16-byte-stride stamp lands on shinfo+12, i.e. the *upper* four bytes
of the frag_list pointer
- the lower half stays 0, so frag_list becomes 0x0080a624_00000000
- non-NULL, so skb_release_data() calls kfree_skb_list_reason() on it, which
does mov 0x0(%rbp),%rbp on a non-canonical address -> #SS
That is why four panics on two kernel versions, two driver builds and two
different call paths produced byte-identical registers. It is not random
memory corruption; it is one field, stamped by one writer, every time.
Interpretation (this part is inference)
---------------------------------------
A 16-byte stride touching only DW3 matches
struct mt76_desc { __le32 buf0, ctrl, buf1, info; } -- info is DW3 at
offset 12. That would mean descriptor-shaped writes are landing on
page_pool RX data pages. The stride, the offset and the constant are
measured; the field identity is a guess and I would not want it taken as
more than that.
I could not identify what 0x0080a624 decodes to as a descriptor word.
Ruled out
---------
- Out-of-tree driver: reproduces on in-tree 7.2.3 with the kernel Not tainted
- Hardware / PCIe link: link trains at full capability, UESta clean, no AER,
EDAC CE/UE both zero. A marginal link does not reproduce a fixed value at a
fixed offset four times.
- Wi-Fi power save: disabled at runtime and persistently; panicked anyway
- ASPM: already disabled by the driver at probe (LnkCtl: ASPM Disabled)
- GRO fraglist: rx-gro-list is off, so nothing should be building a frag_list
at all -- consistent with shinfo being overwritten rather than populated
I have two vmcores (1.4 GB and 1.5 GB, 7.2.3, makedumpfile -l -d 31) and can
run any drgn query against them, or provide dumps if useful. Happy to test
patches -- the box reproduces this within a day of normal use.
Note the dumps were taken with -d 31, which excluded the sk_buff slab page
itself; only the data page survived. I am lowering the filter for future
dumps.
Thanks,
Vladislav Rysin
^ permalink raw reply [flat|nested] 5+ messages in thread* Re: [BUG] mt7925e/MT7927: RX page_pool buffers stamped every 16 bytes, corrupting skb_shared_info->frag_list
2026-09-06 16:27 [BUG] mt7925e/MT7927: RX page_pool buffers stamped every 16 bytes, corrupting skb_shared_info->frag_list -> #SS in kfree_skb_list_reason Vladislav Rysin
@ 2026-09-10 4:03 ` Devin Wittmayer
2026-09-10 13:47 ` Vladislav Rysin
1 sibling, 0 replies; 5+ messages in thread
From: Devin Wittmayer @ 2026-09-10 4:03 UTC (permalink / raw)
To: Vladislav Rysin; +Cc: linux-wireless, Felix Fietkau
Vladislav,
Thank you for the two vmcores, and for pinning it to a single field.
Your arithmetic checks out against 7.2.3:
buffer size 2048
shinfo at buf + 1728
buffer at page + 2048 -> 0xec0
frag_list at shinfo + 8
rx ring 1536 entries, 24 KB, six pages
That last line is the scattered pages you found.
The shape you picked is the only candidate in the driver. Nothing else
is sixteen bytes with a word at offset twelve.
The constant is not worth more of your time. Those info fields only get
read by the 7915 and 7996 parts, and this driver never looks at that
word.
Devin
^ permalink raw reply [flat|nested] 5+ messages in thread* Re: [BUG] mt7925e/MT7927: RX page_pool buffers stamped every 16 bytes, corrupting skb_shared_info->frag_list
2026-09-06 16:27 [BUG] mt7925e/MT7927: RX page_pool buffers stamped every 16 bytes, corrupting skb_shared_info->frag_list -> #SS in kfree_skb_list_reason Vladislav Rysin
2026-09-10 4:03 ` [BUG] mt7925e/MT7927: RX page_pool buffers stamped every 16 bytes, corrupting skb_shared_info->frag_list Devin Wittmayer
@ 2026-09-10 13:47 ` Vladislav Rysin
2026-09-10 16:24 ` Devin Wittmayer
2026-09-10 18:00 ` Vladislav Rysin
1 sibling, 2 replies; 5+ messages in thread
From: Vladislav Rysin @ 2026-09-10 13:47 UTC (permalink / raw)
To: lucid_duck; +Cc: linux-wireless, nbd
Devin,
One correction first: you have not got the vmcores. I offered them in the
first mail but attached nothing, so anything below is from my analysis, not
from dumps you have looked at. Say the word and I will get them to you.
Your buffer arithmetic agrees with what I measured, and buf_size 2048 turns
out to be load-bearing in a way that helps. Separating what I read out of
the dumps from what is derived:
measured shinfo at page offset 0xec0 in both dumps
corrupt bytes at shinfo+8 (frag_list), upper half only
derived sizeof(skb_shared_info) = 48 + 17*16 = 320, already
64-byte aligned so SKB_DATA_ALIGN is a no-op, 2048-320
= 1728
Where I think the reading goes wrong is the ring.
page_pool hands out 2048-byte fragments, so a 4 KB page holds *two* RX
buffers. Their shinfo lands at page+1728 (0x6c0) and page+3776 (0xec0),
and the frag_list upper half at 0x6cc and 0xecc. Both are 12 mod 16. A
16-byte-stride write at offset 12 therefore lands on the frag_list high
half of *every* buffer in the page, not just occasionally. That is why
four panics produced byte-identical registers rather than intermittent
garbage, and it only makes sense if these pages are page_pool buffers.
It also confirms your 2048 from my side.
The ring cannot be what I measured, for four reasons.
1. Different allocators, and only one of them is contiguous. In the
in-tree mt76.ko on 7.2.3 the undefined symbols are:
U dmam_alloc_attrs <- descriptors, coherent, contiguous
U page_pool_create
U page_pool_alloc_frag <- RX data, scattered fragments
U page_pool_put_unrefed_netmem
That is the only coherent allocation path in the module; there is no
vmalloc allocation (is_vmalloc_addr is only a test). So a ring of any
size is consecutive pages in the kernel mapping. Mine are not
consecutive -- stamped 16-byte records per 4 KB page, across the
17-page window I sampled around the offending buffer:
8053f 256 80540 256 80541 0 80542 219 80543 204
80544 0 80545 0 80546 0 80547 256* 80548 0
80549 0 8054a 256 8054b 256 8054c 256 8054d 256
8054e 0 8054f 0
(* the page holding the shinfo)
Seven fully stamped, two partial, interleaved with clean pages. The
window was my sample, not the extent, so there may be more. I have not
verified your 1536 entries / 24 KB / six pages, but I do not need to:
contiguity rules the ring out at any size.
2. Two pages are only partially stamped, 219/256 and 204/256. Every slot
in a filled ring holds a descriptor.
3. The page holding the shinfo carries live payload in the bytes the
stamp does not cover -- local MAC, AP MAC and local IPv4 all legible:
+0e80 00 00 66 ac f7 29 1b 46 f8 1a 2b 1a 24 a6 80 00
+0ea0 c0 a8 56 22 01 bb e3 fc 94 66 fb f5 24 a6 80 00
A descriptor ring does not carry payload.
4. If these were ring pages there would be no bug to report. The driver
writing its own ring is normal, and skb_shared_info would never be in
range of it.
I also went looking for a benign writer and could not find one. There is
no WED/RRO offload in play: zero wed/rro/airoha/ppe undefined symbols in
mt76, mt7925e, mt7925-common, mt792x-lib and mt76-connac-lib on this
kernel. And the hardware does write an RX descriptor into the buffer --
but at the buffer start, once, where the driver parses it. Neither
produces a repeating 16-byte stride across a whole page and over shinfo.
Which is why I read your last paragraph the opposite way. If this driver
never reads that word, and it is only meaningful to 7915/7996, then the
"driver scribbling its own ring" explanation is gone and the question
narrows instead of closing: something writes DW3-shaped words into
page_pool data buffers, at a 16-byte stride, with the same constant across
four panics, three kernel versions (7.1.10, 7.1.12, 7.2.3) and two driver
builds (the mediatek-mt7927-dkms 2.14 backport and in-tree, the latter
with the kernel Not tainted). Whatever performs that write is the defect.
I agree the constant's *meaning* is a dead end -- I am treating it purely
as a fingerprint for identifying the writer.
Your claim that mt76_desc is the only 16-byte structure in the driver with
a word at offset 12 I cannot check; the DKMS source tree went with the
package. I am content for the shape to be wrong. The stride, the offset,
the constant, the page map and the payload-in-page are measurements and do
not depend on it.
Two limits on what I can still do. The dumps were captured with
makedumpfile -d 31, which excluded the sk_buff slab page, so the skb
cannot be inspected in them -- only the data page survived. And the card
is out of the machine, replaced by an MT7925 USB adapter (mt7925u), which
has not reproduced this in four days. So I cannot capture a cleaner dump
or test a patch against this hardware any more. What I have is the two
dumps and a drgn script that reproduces the above from either of them;
both are yours on request.
Thanks,
Vladislav Rysin
^ permalink raw reply [flat|nested] 5+ messages in thread* Re: [BUG] mt7925e/MT7927: RX page_pool buffers stamped every 16 bytes, corrupting skb_shared_info->frag_list
2026-09-10 13:47 ` Vladislav Rysin
@ 2026-09-10 16:24 ` Devin Wittmayer
2026-09-10 18:00 ` Vladislav Rysin
1 sibling, 0 replies; 5+ messages in thread
From: Devin Wittmayer @ 2026-09-10 16:24 UTC (permalink / raw)
To: Vladislav Rysin; +Cc: linux-wireless, Felix Fietkau
Vladislav,
My mistake on the vmcores. Sorry about that.
You are right about the ring. I had what I needed to see that myself and
had already noted the descriptors come from a coherent allocation. Your
first mail said the stamped pages were scattered. Those two do not sit
together and I never put them side by side.
I got the shape claim wrong too, and it is the one you could not check.
There is a second 16-byte structure, and a loop that writes only its word
at +12:
dma.h:136 struct mt76_rro_rxdmad_c, four __le32
dma.c:184-185 for (i = 0; i < q->ndesc; i++)
dmad[i].data3 = cpu_to_le32(data3);
Your stride and your offset exactly. Not your writer though: it puts
0xf0000000 there, and the queue flag it needs is set only by mt7996. So
the pattern is one the driver itself makes, which seems worth knowing
even though this instance is not it.
Your point about both buffers landing at 12 mod 16 explains the identical
registers better than anything I sent.
Same card here, still in a machine:
aa:00.0 MT7927 [14c3:7927] mt7925e
Send the dumps and the drgn script whenever suits, I have drgn 0.2.0 to
match yours.
Devin
^ permalink raw reply [flat|nested] 5+ messages in thread* Re: [BUG] mt7925e/MT7927: RX page_pool buffers stamped every 16 bytes, corrupting skb_shared_info->frag_list
2026-09-10 13:47 ` Vladislav Rysin
2026-09-10 16:24 ` Devin Wittmayer
@ 2026-09-10 18:00 ` Vladislav Rysin
1 sibling, 0 replies; 5+ messages in thread
From: Vladislav Rysin @ 2026-09-10 18:00 UTC (permalink / raw)
To: lucid_duck; +Cc: linux-wireless, nbd
Devin,
The rxdmad_c loop is the most useful thing anyone has put in this thread.
I could not have found it -- the DKMS source went with the package and I
have no tree on this box any more.
dma.c:184-185 for (i = 0; i < q->ndesc; i++)
dmad[i].data3 = cpu_to_le32(data3);
A bounded loop writing only +12 across ndesc 16-byte records is precisely
the footprint, and it changes what I think I am looking at. Two things
follow.
First, it explains something I could not: the partial pages. 219/256 and
204/256 do not fit continuous hardware DMA, which has no reason to stop
mid-page. They fit a loop with a bounded count landing on memory that is
not the ring it was written for.
Second, the timing. The fourth panic was taken inside a reset:
mt7925e: Message 00020016 (seq 4) timeout
Workqueue: mt76 mt7925_mac_reset_work [mt7925_common]
mt7925e_mac_reset -> __local_bh_enable_ip -> do_softirq
-> net_rx_action -> skb_defer_free_flush -> kfree_skb_list_reason
and the three earlier chip re-inits in my logs each followed an MCU
timeout too. So the question I would ask the source, if I still had it,
is whether any queue-init or fill loop on the reset path can run with
q->desc or q->ndesc describing something other than the coherent ring it
was set up for -- a queue re-initialised while its descriptor pointer
still refers to, or has been reassigned to, page_pool memory. The
corruption itself is already there before the flush; the reset is only
what walks into it.
The constant does not match, granted: 0x0080a624 against 0xf0000000, and
the flag gated to mt7996. So either a different writer with the same
shape, or the same shape with data3 coming from somewhere else. Worth
checking whether anything else in the tree writes a +12 word across a
descriptor array, including on paths shared with 7925.
On your card -- that is the part of this I cannot do any more. Mine is out
of the machine. If you can reproduce it, the dump I never managed to get
is one taken with
core_collector makedumpfile -l --message-level 7 -d 1
Mine were -d 31, which excluded the sk_buff slab page, so I could only
ever inspect the data page. With -d 1 the skb itself survives, and so do
the module pages, which would let you walk the queue state -- q->desc,
q->ndesc, the page_pool -- at the moment of the fault rather than
inferring it from a data page as I had to.
What reproduced it here, for what it is worth:
traffic-dependent, not timed: 22h34m, 27h55m, 4h53m, 8h50m of uptime
plain VHT80, 5745 MHz, 866 Mbit/s, NSS 2 -- never EHT or 320 MHz
multi-AP mesh, frequent roams and Reason 2 deauths
Wi-Fi power save made no difference, on or off
reproduced on 7.1.10 and 7.1.12 with mediatek-mt7927-dkms 2.14, and on
7.2.3 in-tree with the kernel Not tainted
Everything is here:
https://drive.google.com/drive/folders/1-NWRQVhtiIBTeNFLe4QukUyrVCM23YDt?usp=sharing
127.0.0.1-2026-09-03-20:42:25/ 1.5 GB 7.1.12, DKMS 2.14
127.0.0.1-2026-09-04-11:09:39/ 1.4 GB 7.2.3, in-tree, Not tainted
127.0.0.1-2026-09-05-02:16:37/ 1.3 GB 7.2.3, in-tree, the reset one
analyze.py drgn script
Each directory has vmcore, vmcore-dmesg.txt and kexec-dmesg.log. The
script prints the stack, the frag_list bytes, the stamped-record count and
the page map from any of the three; it wants kernel-debuginfo matching the
dump's kernel, and the install line for the 7.2.3 vanilla build is in its
docstring. crash 9.0.1 cannot open these at all -- it dies with "invalid
structure member offset: kmem_cache_s_num" on the 7.2 slab layout, which
is why the script is drgn.
All three are -d 31, so the sk_buff page is missing from every one of them.
Thanks,
Vladislav Rysin
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-09-10 18:02 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-06 16:27 [BUG] mt7925e/MT7927: RX page_pool buffers stamped every 16 bytes, corrupting skb_shared_info->frag_list -> #SS in kfree_skb_list_reason Vladislav Rysin
2026-09-10 4:03 ` [BUG] mt7925e/MT7927: RX page_pool buffers stamped every 16 bytes, corrupting skb_shared_info->frag_list Devin Wittmayer
2026-09-10 13:47 ` Vladislav Rysin
2026-09-10 16:24 ` Devin Wittmayer
2026-09-10 18:00 ` Vladislav Rysin
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.