Netdev List
 help / color / mirror / Atom feed
* [PATCH bpf 0/3] bpf: Fix socket leaks around connect(AF_UNSPEC)+listen()
@ 2026-07-23 11:33 Michal Luczaj
  2026-07-23 11:33 ` [PATCH bpf 1/3] bpf, sockmap: Use sock_hold() instead of refcount_inc_not_zero() in lookup Michal Luczaj
                   ` (2 more replies)
  0 siblings, 3 replies; 6+ messages in thread
From: Michal Luczaj @ 2026-07-23 11:33 UTC (permalink / raw)
  To: Eric Dumazet, Kuniyuki Iwashima, Paolo Abeni, Willem de Bruijn,
	John Fastabend, Jakub Sitnicki, Jiayuan Chen, David S. Miller,
	Jakub Kicinski, Simon Horman, Daniel Borkmann, Stanislav Fomichev,
	Martin KaFai Lau, Alexei Starovoitov, Andrii Nakryiko,
	Eduard Zingerman, Kumar Kartikeya Dwivedi, Song Liu,
	Yonghong Song, Jiri Olsa, Emil Tsalapatis, Joe Stringer
  Cc: netdev, bpf, linux-kernel, Michal Luczaj, Sashiko

This is a follow-up to Sashiko's report[1].

Several BPF socket helpers acquire a socket reference only when
sk_is_refcounted() == true, and release it, independently, by
re-evaluating sk_is_refcounted() again at the time the release runs. TCP
connect(AF_UNSPEC)+listen() sets SOCK_RCU_FREE on an established socket.
If that happens while a reference is outstanding, the release side sees
sk_is_refcounted() == false and skips the put; the socket is leaked.

unreferenced object 0xffff88811617ce00 (size 3200):
  comm "softirq", pid 0, jiffies 4294848512
  hex dump (first 32 bytes):
    7f 00 00 01 7f 00 00 01 4d 43 02 f6 00 00 00 00  ........MC......
    02 00 07 41 00 00 00 00 00 00 00 00 00 00 00 00  ...A............
  backtrace (crc fb5bd4c8):
    kmem_cache_alloc_noprof+0x53e/0x640
    sk_prot_alloc+0x69/0x240
    sk_clone+0x79/0x1230
    inet_csk_clone_lock+0x30/0x760
    tcp_create_openreq_child+0x34/0x2750
    tcp_v4_syn_recv_sock+0x12e/0x1080
    tcp_check_req+0x447/0x2310
    tcp_v4_rcv+0x1026/0x3c90
    ip_protocol_deliver_rcu+0x93/0x340
    ip_local_deliver_finish+0x356/0x5c0
    ip_local_deliver+0x184/0x4a0
    ip_rcv+0x4f4/0x5b0
    __netif_receive_skb_one_core+0x153/0x1b0
    process_backlog+0x28d/0x1190
    __napi_poll+0xab/0x520
    net_rx_action+0x3f0/0xca0

The approach taken is to make acquire and release unconditional and
symmetric.

Note: TC bpf_sk_assign() has the same issue; it takes a reference only when
sk_is_refcounted() is true at assign time, but sock_pfree() (the skb
destructor it installs) re-checks sk_is_refcounted() independently at
release time. The same connect(AF_UNSPEC)+listen() transition leaks the
socket here too. I'd welcome suggestions on the right way to handle this.

[1]: https://lore.kernel.org/bpf/20260701235552.2B0AA1F00A3F@smtp.kernel.org/

Signed-off-by: Michal Luczaj <mhal@rbox.co>
---
Michal Luczaj (3):
      bpf, sockmap: Use sock_hold() instead of refcount_inc_not_zero() in lookup
      bpf: Extract shared reqsk-to-listener upgrade
      bpf: Unconditionally take socket references in lookup helpers

 net/core/filter.c   | 80 ++++++++++++++++++++++++++---------------------------
 net/core/sock_map.c | 12 +++-----
 2 files changed, 44 insertions(+), 48 deletions(-)
---
base-commit: 94515f3a7d4256a5062176b7d6ed0471938cd51a
change-id: 20260628-sockmap-lookup-tcp-leak-bdaba3e083c5

Best regards,
--  
Michal Luczaj <mhal@rbox.co>


^ permalink raw reply	[flat|nested] 6+ messages in thread

end of thread, other threads:[~2026-07-25  0:00 UTC | newest]

Thread overview: 6+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-23 11:33 [PATCH bpf 0/3] bpf: Fix socket leaks around connect(AF_UNSPEC)+listen() Michal Luczaj
2026-07-23 11:33 ` [PATCH bpf 1/3] bpf, sockmap: Use sock_hold() instead of refcount_inc_not_zero() in lookup Michal Luczaj
2026-07-25  0:00   ` John Fastabend
2026-07-23 11:33 ` [PATCH bpf 2/3] bpf: Extract shared reqsk-to-listener upgrade Michal Luczaj
2026-07-23 11:33 ` [PATCH bpf 3/3] bpf: Unconditionally take socket references in lookup helpers Michal Luczaj
2026-07-23 15:25   ` bot+bpf-ci

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox