Netdev List
 help / color / mirror / Atom feed
* [PATCH net] page_pool: keep fragments cache-line aligned
@ 2026-09-03 19:17 Fikret Pirim
  0 siblings, 0 replies; only message in thread
From: Fikret Pirim @ 2026-09-03 19:17 UTC (permalink / raw)
  To: fpirim, netdev

page_pool_alloc_frag_netmem() rounds the requested fragment size up to
dma_get_cache_alignment(), which is 1 on architectures that do not define
ARCH_DMA_MINALIGN -- x86 among them, where the helper falls through to
"return 1" (include/linux/dma-mapping.h, include/linux/cache.h).

Two callers on the shared per-CPU system page_pool ask for arbitrary byte
lengths.  skb_pp_cow_data() in its fragment loop:

	size = min_t(u32, len, PAGE_SIZE);
	truesize = size;
	page = page_pool_dev_alloc(pool, &page_off, &truesize);

and xdp_copy_frags_from_zc() likewise (u32 truesize = len).  Once such a
fragment is handed out, pool->frag_offset sits at an arbitrary byte offset,
and the next page_pool_alloc_va() from the same pool returns a misaligned
buffer.  Both callers turn that buffer into an skb with napi_build_skb(),
and __build_skb_around() stores it verbatim:

	skb->head = data;
	skb_set_end_offset(skb, size);

so skb_shinfo(skb) == skb->head + skb->end inherits the misalignment.
atomic_set(&shinfo->dataref, 1) is a plain store and stays silent, but the
first skb_clone() does atomic_inc() on it.  dataref sits at offset 32 in
struct skb_shared_info, so whenever (head + end) % 64 lands in 61..63 the
locked dword straddles a cache line and x86 raises #AC.  On a kernel with
split-lock detection fatal for kernel code that is a panic:

	Oops: Split lock detected
	RIP: 0010:skb_clone+0x159/0x1e0
	 raw_local_deliver+0x1dd/0x290		/* or __udp4_lib_rcv+0x703 */
	 ip_protocol_deliver_rcu+0x51/0x1a0
	 ip_local_deliver_finish+0x85/0x100
	 __netif_receive_skb_one_core+0x85/0xa0
	 process_backlog+0x98/0x150
	 ...
	 __local_bh_enable_ip+0x62/0x70
	 __dev_queue_xmit+0x3cf/0x11a0
	 udp_sendmsg+0xb79/0xfe0

I hit this three times on an i9-13900K, on 6.18.48 and on 7.1.11, while a
generic-mode XDP program with frags support was attached to lo by a NetBird
VPN agent.  Locally generated traffic on lo always has headroom <
XDP_PACKET_HEADROOM, so netif_skb_check_for_xdp() -> skb_cow_data_for_xdp()
-> skb_pp_cow_data() runs for every packet, and any nonlinear one poisons
the pool for the packets that follow.  The three dumps agree with the
arithmetic: skb_shinfo % 64 was 30, 29 and 30, putting the 4-byte dataref
at 62..65, 61..64 and 62..65.  Cloning came from multicast fan-out and from
raw sockets.  On the same machine, with no special load, a read-only
bpftrace kretprobe on __napi_build_skb() counted 40 misaligned heads in 85
seconds against 2033 aligned ones.

Where split-lock detection is not available the misplaced dataref corrupts
whatever it lands on instead; I suspect this is also behind the syzbot
report "KASAN: use-after-free Read in skb_clone" (31 Aug 2026,
extid=22c4f9a7026c86bcc3b8), where the head is freed early.

Round fragments up to at least SMP_CACHE_BYTES.  Architectures with
ARCH_DMA_MINALIGN already get this or stricter rounding from the same line,
so this only restores on x86 the alignment that page_pool_alloc_va() users
implicitly rely on.

Fixes: FIXME_HASH ("page_pool: introduce page_pool_alloc() API")
Signed-off-by: Fikret Pirim <fpirim@gmail.com>
---
Tested on 6.18.48 in an 8 vCPU VM: an XDP_PASS program declared
SEC("xdp.frags") attached to lo in generic mode, 8000-byte UDP datagrams
streamed over lo to drive the fragment loop, and 86-byte multicast to eight
listeners to clone the next skb.  Misalignment is visible without
split-lock hardware by adding

	WARN(!IS_ALIGNED((unsigned long)data, 8), "unaligned head %px frag_size %u\n", ...)

at the top of __build_skb_around() and booting with panic_on_warn=1:

	RIP: 0010:__build_skb_around+0x28c/0x370
	R12: ffff88812a58bba6			/* head, 38 mod 64 */
	R13: 000000000000045a			/* frag_size */
	 __napi_build_skb+0x39/0x50
	 skb_pp_cow_data+0x233/0xfc0
	 do_xdp_generic+0x377/0xd90
	 __netif_receive_skb_core.constprop.0+0x533/0x3f10
	 __netif_receive_skb_one_core+0xe5/0x520
	 process_backlog+0x196/0x5a0
	 __napi_poll+0xb2/0x450
	 net_rx_action+0x418/0xdc0
	 handle_softirqs+0x19e/0x4b0
	 </IRQ>
	 __local_bh_enable_ip+0x62/0x70
	 __dev_queue_xmit+0xcaf/0x3510
	 ip_finish_output2+0x990/0x1f20
	 udp_send_skb+0x60f/0x1310
	 udp_sendmsg+0x1586/0x2230

  unpatched                            trap after ~25 s (twice)
  this patch                           375 s clean, 97 M + 146 M deliveries
  align the caller instead             378 s clean, 91 M + 138 M deliveries
  unpatched, XDP program without frags 272 s clean (pskb_expand_head path)
  unpatched, no XDP program            200 s clean

I also have the caller-side alternative that rounds up the length
skb_pp_cow_data() asks for; I went with this one because
xdp_copy_frags_from_zc() has the same pattern and any future
page_pool_alloc_va() user would share the exposure.  Happy to send that
version instead if you prefer to leave page_pool's rounding alone.  Either
way a DEBUG_NET_WARN_ON_ONCE() on the head alignment in
__build_skb_around() would have caught this immediately.

 net/core/page_pool.c | 11 ++++++++++-
 1 file changed, 10 insertions(+), 1 deletion(-)

--- a/net/core/page_pool.c
+++ b/net/core/page_pool.c
@@ -1061,7 +1061,16 @@
 	if (WARN_ON(size > max_size))
 		return 0;
 
-	size = ALIGN(size, dma_get_cache_alignment());
+	/* Keep fragments cache-line aligned. dma_get_cache_alignment() is 1 on
+	 * architectures without ARCH_DMA_MINALIGN (x86 among them), and callers
+	 * such as skb_pp_cow_data() and xdp_copy_frags_from_zc() request
+	 * arbitrary byte lengths; an odd-sized fragment would leave frag_offset
+	 * misaligned and the next page_pool_alloc_va() buffer would be built into
+	 * an skb whose skb_shared_info is misaligned (split-lock #AC in the first
+	 * skb_clone() on x86, refcount corruption elsewhere).
+	 */
+	size = ALIGN(size, max_t(unsigned int, dma_get_cache_alignment(),
+				 SMP_CACHE_BYTES));
 	*offset = pool->frag_offset;
 
 	if (netmem && *offset + size > max_size) {

^ permalink raw reply	[flat|nested] only message in thread

only message in thread, other threads:[~2026-09-03 19:19 UTC | newest]

Thread overview: (only message) (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-03 19:17 [PATCH net] page_pool: keep fragments cache-line aligned Fikret Pirim

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox