From: "Emil Tsalapatis" <emil@etsalapatis.com>
To: "Junseo Lim" <zirajs7@gmail.com>,
"John Fastabend" <john.fastabend@gmail.com>,
"Jakub Sitnicki" <jakub@cloudflare.com>,
"Jiayuan Chen" <jiayuan.chen@linux.dev>
Cc: "David S. Miller" <davem@davemloft.net>,
"Eric Dumazet" <edumazet@google.com>,
"Jakub Kicinski" <kuba@kernel.org>,
"Paolo Abeni" <pabeni@redhat.com>,
"Simon Horman" <horms@kernel.org>,
"Martin KaFai Lau" <martin.lau@kernel.org>,
<linux-kernel@vger.kernel.org>, <bpf@vger.kernel.org>,
<netdev@vger.kernel.org>, "Sechang Lim" <rhkrqnwk98@gmail.com>
Subject: Re: [PATCH bpf] bpf, sockmap: fix page_counter underflow in strparser SK_PASS
Date: Wed, 29 Jul 2026 02:08:27 -0400 [thread overview]
Message-ID: <DKATWVU2EXPJ.1QNREP4E46RY3@etsalapatis.com> (raw)
In-Reply-To: <20260723065244.186916-1-zirajs7@gmail.com>
On Thu Jul 23, 2026 at 2:52 AM EDT, Junseo Lim wrote:
> tcp_bpf_strp_read_sock() delays cleanup of SK_PASS bytes by
> subtracting psock->ingress_bytes from the amount passed to
> __tcp_cleanup_rbuf(). But when sk_psock_verdict_apply() queues the skb
> directly through sk_psock_skb_ingress_self(), skb_set_owner_r() is called
> unconditionally and charges the skb again.
Can you specify in the commit that this is about sk_forward_alloc (AFAICT)?
The eventual page_counter underflow is a second-order effect even if it's
what eventually causes the crash.
>
> The duplicated charge is later released independently and can trigger a
> page_counter underflow.
>
> Add a charge_skb argument to sk_psock_skb_ingress_self() and skip
> skb_set_owner_r() only for the direct strparser SK_PASS path. Keep existing
> accounting for the other self-ingress caller and for non-strparser SK_PASS.
>
> Fixes: 36b62df5683c ("bpf: Fix wrong copied_seq calculation")
> Signed-off-by: Junseo Lim <zirajs7@gmail.com>
> ---
> Crash reproduced on Linux tree 94515f3a7d4256a5062176b7d6ed0471938cd51a
> with KASAN, MEMCG, panic_on_warn=1, and oops=panic.
Can you make a selftest out of the reproducer?
>
> Reproducer/log/config: https://gist.github.com/ZirAjs/16c95c89972ace73910d9b5ac78a5807
>
> The reproducer drives the strparser SK_PASS path until teardown reports:
>
> page_counter underflow
> Workqueue: events sk_psock_destroy
> Kernel panic - not syncing: kernel: panic_on_warn set ...
>
> net/core/skmsg.c | 29 +++++++++++++++++------------
> 1 file changed, 17 insertions(+), 12 deletions(-)
>
> diff --git a/net/core/skmsg.c b/net/core/skmsg.c
> index 2521b643fa05..17260681479a 100644
> --- a/net/core/skmsg.c
> +++ b/net/core/skmsg.c
> @@ -586,7 +586,8 @@ static int sk_psock_skb_ingress_enqueue(struct sk_buff *skb,
> }
>
> static int sk_psock_skb_ingress_self(struct sk_psock *psock, struct sk_buff *skb,
> - u32 off, u32 len, bool take_ref);
> + u32 off, u32 len, bool take_ref,
> + bool charge_skb);
>
> static int sk_psock_skb_ingress(struct sk_psock *psock, struct sk_buff *skb,
> u32 off, u32 len)
> @@ -595,12 +596,9 @@ static int sk_psock_skb_ingress(struct sk_psock *psock, struct sk_buff *skb,
> struct sk_msg *msg;
> int err;
>
> - /* If we are receiving on the same sock skb->sk is already assigned,
> - * skip memory accounting and owner transition seeing it already set
> - * correctly.
> - */
> if (unlikely(skb->sk == sk))
> - return sk_psock_skb_ingress_self(psock, skb, off, len, true);
> + return sk_psock_skb_ingress_self(psock, skb, off, len, true,
> + true);
> msg = sk_psock_create_ingress_msg(sk, skb);
> if (!msg)
> return -EAGAIN;
> @@ -618,12 +616,14 @@ static int sk_psock_skb_ingress(struct sk_psock *psock, struct sk_buff *skb,
> return err;
> }
>
> -/* Puts an skb on the ingress queue of the socket already assigned to the
> - * skb. In this case we do not need to check memory limits or skb_set_owner_r
> - * because the skb is already accounted for here.
> +/* Puts an skb on the ingress queue for psock->sk.
> + *
> + * When charge_skb is false, the direct strparser SK_PASS path keeps the TCP
> + * receive queue accounting in place and must not call skb_set_owner_r().
> */
> static int sk_psock_skb_ingress_self(struct sk_psock *psock, struct sk_buff *skb,
> - u32 off, u32 len, bool take_ref)
> + u32 off, u32 len, bool take_ref,
> + bool charge_skb)
> {
> struct sk_msg *msg = alloc_sk_msg(GFP_ATOMIC);
> struct sock *sk = psock->sk;
> @@ -631,7 +631,8 @@ static int sk_psock_skb_ingress_self(struct sk_psock *psock, struct sk_buff *skb
>
> if (unlikely(!msg))
> return -EAGAIN;
> - skb_set_owner_r(skb, sk);
> + if (charge_skb)
> + skb_set_owner_r(skb, sk);
Sashiko's right that skipping this leaves the skb half-configured. Since
the point is to not double-count sk_forward_alloc but we still have
to account truesize, can you try something like:
if (!skb_charge)
sk_rmem_schedule(sk, skb, 0);
skb_set_owner_r(skb, sk);
to just ajust the sk_forward_alloc?
pw-bot: cr
>
> /* This is used in tcp_bpf_recvmsg_parser() to determine whether the
> * data originates from the socket's own protocol stack. No need to
> @@ -1017,6 +1018,8 @@ static int sk_psock_verdict_apply(struct sk_psock *psock, struct sk_buff *skb,
> * retrying later from workqueue.
> */
> if (skb_queue_empty(&psock->ingress_skb)) {
> + bool charge_skb = true;
> +
> len = skb->len;
> off = 0;
> if (skb_bpf_strparser(skb)) {
> @@ -1024,8 +1027,10 @@ static int sk_psock_verdict_apply(struct sk_psock *psock, struct sk_buff *skb,
>
> off = stm->offset;
> len = stm->full_len;
> + charge_skb = false;
> }
> - err = sk_psock_skb_ingress_self(psock, skb, off, len, false);
> + err = sk_psock_skb_ingress_self(psock, skb, off, len,
> + false, charge_skb);
> }
> if (err < 0) {
> spin_lock_bh(&psock->ingress_lock);
prev parent reply other threads:[~2026-07-29 6:08 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-23 6:52 [PATCH bpf] bpf, sockmap: fix page_counter underflow in strparser SK_PASS Junseo Lim
2026-07-23 7:20 ` sashiko-bot
2026-07-29 6:08 ` Emil Tsalapatis [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=DKATWVU2EXPJ.1QNREP4E46RY3@etsalapatis.com \
--to=emil@etsalapatis.com \
--cc=bpf@vger.kernel.org \
--cc=davem@davemloft.net \
--cc=edumazet@google.com \
--cc=horms@kernel.org \
--cc=jakub@cloudflare.com \
--cc=jiayuan.chen@linux.dev \
--cc=john.fastabend@gmail.com \
--cc=kuba@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=martin.lau@kernel.org \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=rhkrqnwk98@gmail.com \
--cc=zirajs7@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.