From: Vladislav Rysin <vrysin@gmail.com>
To: lucid_duck@justthetip.ca
Cc: linux-wireless@vger.kernel.org, nbd@nbd.name
Subject: Re: [BUG] mt7925e/MT7927: RX page_pool buffers stamped every 16 bytes, corrupting skb_shared_info->frag_list
Date: Thu, 10 Sep 2026 09:47:48 -0400 [thread overview]
Message-ID: <20260910134748.75684-1-vrysin@gmail.com> (raw)
In-Reply-To: <20260906162742.15333-1-vrysin@gmail.com>
Devin,
One correction first: you have not got the vmcores. I offered them in the
first mail but attached nothing, so anything below is from my analysis, not
from dumps you have looked at. Say the word and I will get them to you.
Your buffer arithmetic agrees with what I measured, and buf_size 2048 turns
out to be load-bearing in a way that helps. Separating what I read out of
the dumps from what is derived:
measured shinfo at page offset 0xec0 in both dumps
corrupt bytes at shinfo+8 (frag_list), upper half only
derived sizeof(skb_shared_info) = 48 + 17*16 = 320, already
64-byte aligned so SKB_DATA_ALIGN is a no-op, 2048-320
= 1728
Where I think the reading goes wrong is the ring.
page_pool hands out 2048-byte fragments, so a 4 KB page holds *two* RX
buffers. Their shinfo lands at page+1728 (0x6c0) and page+3776 (0xec0),
and the frag_list upper half at 0x6cc and 0xecc. Both are 12 mod 16. A
16-byte-stride write at offset 12 therefore lands on the frag_list high
half of *every* buffer in the page, not just occasionally. That is why
four panics produced byte-identical registers rather than intermittent
garbage, and it only makes sense if these pages are page_pool buffers.
It also confirms your 2048 from my side.
The ring cannot be what I measured, for four reasons.
1. Different allocators, and only one of them is contiguous. In the
in-tree mt76.ko on 7.2.3 the undefined symbols are:
U dmam_alloc_attrs <- descriptors, coherent, contiguous
U page_pool_create
U page_pool_alloc_frag <- RX data, scattered fragments
U page_pool_put_unrefed_netmem
That is the only coherent allocation path in the module; there is no
vmalloc allocation (is_vmalloc_addr is only a test). So a ring of any
size is consecutive pages in the kernel mapping. Mine are not
consecutive -- stamped 16-byte records per 4 KB page, across the
17-page window I sampled around the offending buffer:
8053f 256 80540 256 80541 0 80542 219 80543 204
80544 0 80545 0 80546 0 80547 256* 80548 0
80549 0 8054a 256 8054b 256 8054c 256 8054d 256
8054e 0 8054f 0
(* the page holding the shinfo)
Seven fully stamped, two partial, interleaved with clean pages. The
window was my sample, not the extent, so there may be more. I have not
verified your 1536 entries / 24 KB / six pages, but I do not need to:
contiguity rules the ring out at any size.
2. Two pages are only partially stamped, 219/256 and 204/256. Every slot
in a filled ring holds a descriptor.
3. The page holding the shinfo carries live payload in the bytes the
stamp does not cover -- local MAC, AP MAC and local IPv4 all legible:
+0e80 00 00 66 ac f7 29 1b 46 f8 1a 2b 1a 24 a6 80 00
+0ea0 c0 a8 56 22 01 bb e3 fc 94 66 fb f5 24 a6 80 00
A descriptor ring does not carry payload.
4. If these were ring pages there would be no bug to report. The driver
writing its own ring is normal, and skb_shared_info would never be in
range of it.
I also went looking for a benign writer and could not find one. There is
no WED/RRO offload in play: zero wed/rro/airoha/ppe undefined symbols in
mt76, mt7925e, mt7925-common, mt792x-lib and mt76-connac-lib on this
kernel. And the hardware does write an RX descriptor into the buffer --
but at the buffer start, once, where the driver parses it. Neither
produces a repeating 16-byte stride across a whole page and over shinfo.
Which is why I read your last paragraph the opposite way. If this driver
never reads that word, and it is only meaningful to 7915/7996, then the
"driver scribbling its own ring" explanation is gone and the question
narrows instead of closing: something writes DW3-shaped words into
page_pool data buffers, at a 16-byte stride, with the same constant across
four panics, three kernel versions (7.1.10, 7.1.12, 7.2.3) and two driver
builds (the mediatek-mt7927-dkms 2.14 backport and in-tree, the latter
with the kernel Not tainted). Whatever performs that write is the defect.
I agree the constant's *meaning* is a dead end -- I am treating it purely
as a fingerprint for identifying the writer.
Your claim that mt76_desc is the only 16-byte structure in the driver with
a word at offset 12 I cannot check; the DKMS source tree went with the
package. I am content for the shape to be wrong. The stride, the offset,
the constant, the page map and the payload-in-page are measurements and do
not depend on it.
Two limits on what I can still do. The dumps were captured with
makedumpfile -d 31, which excluded the sk_buff slab page, so the skb
cannot be inspected in them -- only the data page survived. And the card
is out of the machine, replaced by an MT7925 USB adapter (mt7925u), which
has not reproduced this in four days. So I cannot capture a cleaner dump
or test a patch against this hardware any more. What I have is the two
dumps and a drgn script that reproduces the above from either of them;
both are yours on request.
Thanks,
Vladislav Rysin
next prev parent reply other threads:[~2026-09-10 13:48 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-06 16:27 [BUG] mt7925e/MT7927: RX page_pool buffers stamped every 16 bytes, corrupting skb_shared_info->frag_list -> #SS in kfree_skb_list_reason Vladislav Rysin
2026-09-10 4:03 ` [BUG] mt7925e/MT7927: RX page_pool buffers stamped every 16 bytes, corrupting skb_shared_info->frag_list Devin Wittmayer
2026-09-10 13:47 ` Vladislav Rysin [this message]
2026-09-10 16:24 ` Devin Wittmayer
2026-09-10 18:00 ` Vladislav Rysin
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260910134748.75684-1-vrysin@gmail.com \
--to=vrysin@gmail.com \
--cc=linux-wireless@vger.kernel.org \
--cc=lucid_duck@justthetip.ca \
--cc=nbd@nbd.name \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.