From: Jiayuan Chen <jiayuan.chen@linux.dev>
To: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Cc: bpf@vger.kernel.org, netdev@vger.kernel.org,
syzbot+237bbeed8dfe0699b7f5@syzkaller.appspotmail.com,
"Andrew Lunn" <andrew+netdev@lunn.ch>,
"David S. Miller" <davem@davemloft.net>,
"Eric Dumazet" <edumazet@google.com>,
"Jakub Kicinski" <kuba@kernel.org>,
"Paolo Abeni" <pabeni@redhat.com>,
"Alexei Starovoitov" <ast@kernel.org>,
"Daniel Borkmann" <daniel@iogearbox.net>,
"Jesper Dangaard Brouer" <hawk@kernel.org>,
"John Fastabend" <john.fastabend@gmail.com>,
"Stanislav Fomichev" <sdf@fomichev.me>,
"Simon Horman" <horms@kernel.org>,
"Martin KaFai Lau" <martin.lau@linux.dev>,
"Andrii Nakryiko" <andrii@kernel.org>,
"Eduard Zingerman" <eddyz87@gmail.com>,
"Kumar Kartikeya Dwivedi" <memxor@gmail.com>,
"Song Liu" <song@kernel.org>,
"Yonghong Song" <yonghong.song@linux.dev>,
"Jiri Olsa" <jolsa@kernel.org>,
"Emil Tsalapatis" <emil@etsalapatis.com>,
"Ihor Solodrai" <ihor.solodrai@linux.dev>,
"Shuah Khan" <shuah@kernel.org>,
"Kuniyuki Iwashima" <kuniyu@google.com>,
"Hangbin Liu" <liuhangbin@gmail.com>,
"Krishna Kumar" <krikku@gmail.com>,
"Martin Karsten" <mkarsten@uwaterloo.ca>,
"Toke Høiland-Jørgensen" <toke@redhat.com>,
linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org
Subject: Re: [PATCH bpf v2 1/2] bpf, veth: xdp: fix page_pool page leak on skb-backed XDP
Date: Mon, 24 Aug 2026 20:19:44 +0800 [thread overview]
Message-ID: <01fd251d-1134-4251-9c11-79943e367ffe@linux.dev> (raw)
In-Reply-To: <aowdem3KQVfvocE7@lore-desk>
On 8/24/26 6:31 PM, Lorenzo Bianconi wrote:
>> bpf_xdp_shrink_data() frees a released frag via __xdp_return() using
>> xdp->rxq->mem.type, but that type is wrong for skb-backed XDP: the skb is
>> cow'd into page_pool memory while the rxq still says MEM_TYPE_PAGE_SHARED,
>> so the page_pool page is freed with page_frag_free() and we hit
>> "Bad page state ... page_pool leak".
>>
>> Both generic XDP and veth are affected. A non-linear skb is cow'd into
>> page_pool memory (skb_cow_data_for_xdp() -> skb_pp_cow_data() for generic
>> XDP, veth_convert_skb_to_xdp_buff() for veth), so its frags become
>> page_pool pages while the rxq keeps MEM_TYPE_PAGE_SHARED.
>>
>> We can't just fix rxq->mem.type in place:
>> - generic XDP: xdp->rxq is dev->_rx[queue].xdp_rxq (see
>> bpf_prog_run_generic_xdp()), a shared rxq that other CPUs may access in
>> parallel, so we must not write to it.
>> - veth: rq->xdp_rxq.mem is shared per-queue state that veth resets on XDP
>> teardown, and with GRO that reset runs without stopping in-flight NAPI,
>> so a type stashed there can be clobbered under a packet still in flight.
>>
>> Adding a check in __xdp_return() or bpf_xdp_shrink_data() itself is not an
>> option either: without recording it somewhere, both can only guess the
>> frag's memory type, which quickly gets confusing.
>>
>> So record it in the xdp_buff. Add a XDP_FLAGS_FRAGS_PAGE_POOL flag; the two
>> skb-cow sites set it, and bpf_xdp_shrink_data() frees the frag to the
>> page_pool when it is set, otherwise it keeps falling back to
>> xdp->rxq->mem.type unchanged. No other path changes behaviour.
>>
>> Fixes: e6d5dbdd20aa ("xdp: add multi-buff support for xdp running in generic mode")
>> Fixes: 0ebab78cbcbf ("net: veth: add page_pool for page recycling")
> Hi Jiayuan Chen,
>
> thx for fixing it. Can we do something like the patch below instead?
>
> Regards,
> Lorenzo
>
> diff --git a/net/core/xdp.c b/net/core/xdp.c
> index 1d679e8fd649..4ed659b58141 100644
> --- a/net/core/xdp.c
> +++ b/net/core/xdp.c
> @@ -433,16 +433,16 @@ EXPORT_SYMBOL_GPL(xdp_rxq_info_attach_page_pool);
> void __xdp_return(netmem_ref netmem, enum xdp_mem_type mem_type,
> bool napi_direct, struct xdp_buff *xdp)
> {
> + netmem_ref head_netmem = netmem_compound_head(netmem);
> + if (netmem_is_pp(head_netmem))
> + mem_type = MEM_TYPE_PAGE_POOL;
> +
> switch (mem_type) {
> case MEM_TYPE_PAGE_POOL:
> - netmem = netmem_compound_head(netmem);
> if (napi_direct && xdp_return_frame_no_direct())
> napi_direct = false;
> - /* No need to check netmem_is_pp() as mem->type knows this a
> - * page_pool page
> - */
> - page_pool_put_full_netmem(netmem_get_pp(netmem), netmem,
> - napi_direct);
> + page_pool_put_full_netmem(netmem_get_pp(head_netmem),
> + head_netmem, napi_direct);
> break;
> case MEM_TYPE_PAGE_SHARED:
> page_frag_free(__netmem_address(netmem));
Hi Lorenzo,
I tried this, but it regresses the bpf selftest with a page_pool ref
underflow (0 warns on master, 45 with the patch):
WARNING: include/net/page_pool/helpers.h:297 at
page_pool_alloc_frag_netmem
skb_pp_cow_data
veth_xdp_rcv_skb
On XDP_TX/XDP_REDIRECT veth has to take plain page refs via get_page()
(veth_xdp_get()) and then
consume_skb(): the skb itself must be freed while the data pages stay
alive for the frame. consume_skb()
already returns the skb's page_pool ref, so what the frame holds
afterwards is a plain page ref, to be
dropped with page_frag_free().
netmem_is_pp() can't see that: it only says the page still belongs to a
pool (other users may still hold pool refs on the same page),
not what kind of ref we're dropping. So __xdp_return() turns those
plain-ref drops into a second pool
put and pp_ref goes negative.
That's why I kept the type in the xdp_buff and only override it in the
shrink path, where we know the
frag ref is the cow'd page_pool one.
Regards,
Jiayuan
next prev parent reply other threads:[~2026-08-24 12:19 UTC|newest]
Thread overview: 14+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-24 3:02 [PATCH bpf v2 0/2] net: xdp: fix bpf_xdp_shrink_data() page handling on generic XDP and veth Jiayuan Chen
2026-08-24 3:06 ` [PATCH bpf v2 1/2] bpf, veth: xdp: fix page_pool page leak on skb-backed XDP Jiayuan Chen
2026-08-24 3:06 ` [PATCH bpf v2 2/2] selftests/bpf: add xdp_shrink_frags Jiayuan Chen
2026-08-24 3:58 ` bot+bpf-ci
2026-08-24 4:53 ` Jiayuan Chen
2026-08-27 19:19 ` Jakub Kicinski
2026-08-24 4:11 ` [PATCH bpf v2 1/2] bpf, veth: xdp: fix page_pool page leak on skb-backed XDP bot+bpf-ci
2026-08-24 10:31 ` Lorenzo Bianconi
2026-08-24 12:19 ` Jiayuan Chen [this message]
2026-08-24 14:50 ` Lorenzo Bianconi
2026-08-25 12:06 ` Jiayuan Chen
2026-08-27 19:19 ` Jakub Kicinski
2026-08-27 19:20 ` Jakub Kicinski
2026-09-04 23:16 ` Emil Tsalapatis
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=01fd251d-1134-4251-9c11-79943e367ffe@linux.dev \
--to=jiayuan.chen@linux.dev \
--cc=andrew+netdev@lunn.ch \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=bpf@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=davem@davemloft.net \
--cc=eddyz87@gmail.com \
--cc=edumazet@google.com \
--cc=emil@etsalapatis.com \
--cc=hawk@kernel.org \
--cc=horms@kernel.org \
--cc=ihor.solodrai@linux.dev \
--cc=john.fastabend@gmail.com \
--cc=jolsa@kernel.org \
--cc=krikku@gmail.com \
--cc=kuba@kernel.org \
--cc=kuniyu@google.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=liuhangbin@gmail.com \
--cc=lorenzo.bianconi@oss.qualcomm.com \
--cc=martin.lau@linux.dev \
--cc=memxor@gmail.com \
--cc=mkarsten@uwaterloo.ca \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=sdf@fomichev.me \
--cc=shuah@kernel.org \
--cc=song@kernel.org \
--cc=syzbot+237bbeed8dfe0699b7f5@syzkaller.appspotmail.com \
--cc=toke@redhat.com \
--cc=yonghong.song@linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.