All of lore.kernel.org
 help / color / mirror / Atom feed
From: Vladislav Rysin <vrysin@gmail.com>
To: linux-wireless@vger.kernel.org
Cc: nbd@nbd.name
Subject: [BUG] mt7925e/MT7927: RX page_pool buffers stamped every 16 bytes, corrupting skb_shared_info->frag_list -> #SS in kfree_skb_list_reason
Date: Sun,  6 Sep 2026 12:27:42 -0400	[thread overview]
Message-ID: <20260906162742.15333-1-vrysin@gmail.com> (raw)

Hi,

I am hitting a reproducible kernel panic with an MT7927 (Filogic 380) card on
mt7925e. Four panics over five days, all with byte-identical registers. I have
analysed two vmcores and believe I have the mechanism, though not the culprit.

Hardware / software
-------------------
Dell Precision 5820 Tower, BIOS 2.48.0, 250 GB RAM
MEDIATEK MT7927 802.11be 2x2 [14c3:7927], subsystem 105b:e104
  behind a dedicated root port (b2:00.0 -> b3:00.0), PCIe 8GT/s x1
  (LnkSta matches LnkCap; UESta clean, no AER events)
driver: mt7925e, MT7927 raw CHIPID=0x0000, forcing chip=0x7927
firmware: linux-firmware-20260810
  HW/SW Version 0x8a108a10, Build Time 20260414133750a
  WM Firmware Version ____000000, Build Time 20260414134255
link at time of crashes: VHT80, 5745 MHz, 866.7 Mbit/s, VHT-NSS 2 (not EHT)

Reproduces on:
  7.2.3 vanilla (in-tree mt7925e, kernel Not tainted)   <- crashes 3 and 4
  7.1.12 and 7.1.10 with the mediatek-mt7927-dkms 2.14 backport <- crashes 1, 2

Uptime at panic: 22h34m, 27h55m, 4h53m, 8h50m. Traffic-dependent, not timed.

The panic
---------
Oops: stack segment: 0000 [#1] SMP NOPTI
CPU: 6 UID: 0 PID: 808 Comm: napi/phy0-0 Kdump: loaded Not tainted
     7.2.3-300.vanilla.fc44.x86_64
RIP: 0010:kfree_skb_list_reason+0xce/0x260
RBX: 0080a62400000000 RDI: 0080a62400000000 RBP: 0080a62400000000
Call Trace:
 skb_release_data+0x184/0x220
 napi_consume_skb+0x7a/0x160
 skb_defer_free_flush+0x81/0xc0
 napi_threaded_poll_loop+0x14e/0x290
 napi_threaded_poll+0x42/0xb0

A second panic took the same fault from softirq context instead, during the
driver's own recovery, right after an MCU timeout:

 mt7925e 0000:b3:00.0: Message 00020016 (seq 4) timeout
 Workqueue: mt76 mt7925_mac_reset_work [mt7925_common]
 mt7925e_mac_reset+0x20e/0x400 -> __local_bh_enable_ip -> do_softirq
 -> net_rx_action -> skb_defer_free_flush -> kfree_skb_list_reason

So the flushing context is incidental; the skb is already corrupt on the
per-CPU defer list. The corruption happens at RX time.

What the vmcore shows
---------------------
Analysed with drgn 0.2.0 (crash 9.0.1 cannot open a 7.2 dump:
"invalid structure member offset: kmem_cache_s_num").

The RX data page carrying the skb is stamped: every 16 bytes, the dword at
offset +12 is overwritten with 24 a6 80 00 (LE 0x0080a624). 256 of 256
16-byte records in the page, and several other page_pool pages are fully
stamped as well (scattered, not contiguous).

The unstamped bytes are genuine payload -- local MAC, AP MAC and the local
IPv4 address are all plainly visible in the same page:

  +0e80  00 00 66 ac f7 29 1b 46 f8 1a 2b 1a  24 a6 80 00
  +0ea0  c0 a8 56 22 01 bb e3 fc 94 66 fb f5  24 a6 80 00

Why the register value is always identical:

  - skb_shared_info sits at the buffer tail, page offset 0xec0 (fixed
    q->buf_size); identical in both dumps
  - frag_list is at shinfo+8
  - the 16-byte-stride stamp lands on shinfo+12, i.e. the *upper* four bytes
    of the frag_list pointer
  - the lower half stays 0, so frag_list becomes 0x0080a624_00000000
  - non-NULL, so skb_release_data() calls kfree_skb_list_reason() on it, which
    does mov 0x0(%rbp),%rbp on a non-canonical address -> #SS

That is why four panics on two kernel versions, two driver builds and two
different call paths produced byte-identical registers. It is not random
memory corruption; it is one field, stamped by one writer, every time.

Interpretation (this part is inference)
---------------------------------------
A 16-byte stride touching only DW3 matches
struct mt76_desc { __le32 buf0, ctrl, buf1, info; } -- info is DW3 at
offset 12. That would mean descriptor-shaped writes are landing on
page_pool RX data pages. The stride, the offset and the constant are
measured; the field identity is a guess and I would not want it taken as
more than that.

I could not identify what 0x0080a624 decodes to as a descriptor word.

Ruled out
---------
 - Out-of-tree driver: reproduces on in-tree 7.2.3 with the kernel Not tainted
 - Hardware / PCIe link: link trains at full capability, UESta clean, no AER,
   EDAC CE/UE both zero. A marginal link does not reproduce a fixed value at a
   fixed offset four times.
 - Wi-Fi power save: disabled at runtime and persistently; panicked anyway
 - ASPM: already disabled by the driver at probe (LnkCtl: ASPM Disabled)
 - GRO fraglist: rx-gro-list is off, so nothing should be building a frag_list
   at all -- consistent with shinfo being overwritten rather than populated

I have two vmcores (1.4 GB and 1.5 GB, 7.2.3, makedumpfile -l -d 31) and can
run any drgn query against them, or provide dumps if useful. Happy to test
patches -- the box reproduces this within a day of normal use.

Note the dumps were taken with -d 31, which excluded the sk_buff slab page
itself; only the data page survived. I am lowering the filter for future
dumps.

Thanks,
Vladislav Rysin

             reply	other threads:[~2026-09-06 16:27 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-06 16:27 Vladislav Rysin [this message]
2026-09-10  4:03 ` [BUG] mt7925e/MT7927: RX page_pool buffers stamped every 16 bytes, corrupting skb_shared_info->frag_list Devin Wittmayer
2026-09-10 13:47 ` Vladislav Rysin
2026-09-10 16:24   ` Devin Wittmayer
2026-09-10 18:00   ` Vladislav Rysin

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260906162742.15333-1-vrysin@gmail.com \
    --to=vrysin@gmail.com \
    --cc=linux-wireless@vger.kernel.org \
    --cc=nbd@nbd.name \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.