From: Jiayuan Chen <jiayuan.chen@linux.dev>
To: Kuniyuki Iwashima <kuniyu@google.com>, bot+bpf-ci@kernel.org
Cc: bpf@vger.kernel.org, daniel@iogearbox.net,
john.fastabend@gmail.com, sdf@fomichev.me, martin.lau@linux.dev,
ast@kernel.org, andrii@kernel.org, eddyz87@gmail.com,
memxor@gmail.com, song@kernel.org, yonghong.song@linux.dev,
jolsa@kernel.org, emil@etsalapatis.com, ihor.solodrai@linux.dev,
davem@davemloft.net, edumazet@google.com, kuba@kernel.org,
pabeni@redhat.com, horms@kernel.org, ncardwell@google.com,
shuah@kernel.org, aditi.ghag@isovalent.com,
netdev@vger.kernel.org, linux-kernel@vger.kernel.org,
linux-kselftest@vger.kernel.org, martin.lau@kernel.org,
mason@kernel.org
Subject: Re: [PATCH bpf v2 2/3] tcp: Skip cond_resched() in inet_csk_listen_stop() under BPF context
Date: Tue, 8 Sep 2026 20:22:02 +0800 [thread overview]
Message-ID: <c6c4fd69-3074-4d6d-804b-cf8ff62359da@linux.dev> (raw)
In-Reply-To: <a44c1e4d-e9a8-4df3-aa25-20875cd4d30e@linux.dev>
On 9/8/26 4:07 PM, Jiayuan Chen wrote:
>
> On 9/8/26 7:34 AM, Kuniyuki Iwashima wrote:
>> On Sun, Sep 6, 2026 at 1:23 AM <bot+bpf-ci@kernel.org> wrote:
>>>> tcp: Skip cond_resched() in inet_csk_listen_stop() under BPF context
>>>>
>>>> bpf_sock_destroy() runs from the tcp iterator, under
>>>> rcu_read_lock(). If
>>>> the sock is a listener that still has children in its accept queue,
>>>> tcp_abort() ends up in inet_csk_listen_stop() and the cond_resched()
>>>> there trips the debug check:
>>>>
>>>> BUG: sleeping function called from invalid context at
>>>> net/ipv4/inet_connection_sock.c:1523
>>>> in_atomic(): 0, irqs_disabled(): 0, non_block: 0, pid: 628, name:
>>>> test_progs
>>>> preempt_count: 0, expected: 0
>>>> RCU nest depth: 1, expected: 0
>>>> locks held by test_progs/628: 3, last CPU#3:
>>>> #0: ffff8881158cee18 (&p->lock){+.+.}-{4:4}, at:
>>>> bpf_seq_read+0x56/0x1210
>>>> #1: ffff8881106bb858 (sk_lock-AF_INET6){+.+.}-{0:0}, at:
>>>> bpf_iter_tcp_seq_show+0x32b/0x4b0
>>>> #2: ffffffffb435af20 (rcu_read_lock){....}-{1:3}, at:
>>>> bpf_iter_run_prog+0x46b/0xde0
>>>> CPU: 3 UID: 0 PID: 628 Comm: test_progs Tainted: G W
>>>> 7.2.0+ #65 PREEMPT
>>>> Tainted: [W]=WARN
>>>> Call Trace:
>>>> <TASK>
>>>> dump_stack_lvl+0xc1/0xf0
>>>> dump_stack+0x10/0x20
>>>> __might_resched+0x3d2/0x610
>>>> inet_csk_listen_stop+0x7b/0xbf0
>>>> tcp_abort+0x23b/0x3b0
>>>> bpf_sock_destroy+0xfc/0x140
>>>> bpf_prog_448133d24601754f_iter_tcp6_server+0x81/0x8a
>>>> bpf_iter_run_prog+0x538/0xde0
>>>> bpf_iter_tcp_seq_show+0x26b/0x4b0
>>>> bpf_seq_read+0x424/0x1210
>>>> vfs_read+0x197/0xe40
>>>> ksys_read+0x119/0x240
>>>> __x64_sys_read+0x72/0xc0
>>>> x64_sys_call+0x647/0x27e0
>>>> do_syscall_64+0xe5/0x610
>>>> entry_SYSCALL_64_after_hwframe+0x76/0x7e
>>>> RIP: 0033:0x7fad39b28aca
>>>> RSP: 002b:00007ffc381c61c0 EFLAGS: 00000246 ORIG_RAX: 0000000000000000
>>>> RAX: ffffffffffffffda RBX: 00007ffc381c6a88 RCX: 00007fad39b28aca
>>>> RDX: 0000000000000032 RSI: 00007ffc381c6250 RDI: 0000000000000014
>>>> RBP: 00007ffc381c61e0 R08: 0000000000000000 R09: 0000000000000000
>>>> R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000003
>>>> R13: 0000000000000000 R14: 000055f077c1bbb0 R15: 00007fad3a0f3000
>>>> </TASK>
>>>>
>>>> The commit that added the kfunc already guards lock_sock() in
>>>> tcp_abort()
>>>> and udp_abort() with has_current_bpf_ctx(), but missed the listener
>>>> path.
>>>> Do the same for the cond_resched(), it can't reschedule there anyway.
>>>>
>>>> Fixes: 4ddbcb886268 ("bpf: Add bpf_sock_destroy kfunc")
>>>> Signed-off-by: Jiayuan Chen <jiayuan.chen@linux.dev>
>>> Is the justification "it can't reschedule there anyway" accurate?
>>>
>>> With CONFIG_PREEMPT_DYNAMIC=y booted with preempt=none or
>>> preempt=voluntary,
>>> cond_resched() expands to __cond_resched() which can actually
>>> reschedule.
>>> The splat in the commit message confirms preempt_count is 0 while
>>> RCU nest
>>> depth is 1. With preempt_count==0, should_resched(0) can be true and
>>> __cond_resched() will call preempt_schedule_common() for a real
>>> reschedule.
>>>
>>> Additionally, in configurations with CONFIG_PREEMPT_RCU=n where
>>> rcu_read_lock() is preempt_disable(), __cond_resched() falls through to
>>> rcu_all_qs() which calls rcu_qs() to report a quiescent state from
>>> inside
>>> an RCU read-side critical section. That's a correctness problem
>>> beyond just
>>> the debug check.
>>>
>>> So the call can either reschedule (PREEMPT_DYNAMIC none/voluntary)
>>> or report
>>> a bogus quiescent state (non-preemptible RCU). Could the
>>> justification be
>>> reworded to explain that the loop runs inside the iterator's RCU
>>> read-side
>>> critical section and must not reschedule or report a quiescent state
>>> there?
>>> The code change itself is correct and matches the existing pattern in
>>> tcp_abort() and udp_abort().
>> or maybe simply remove cond_resched(), hoping 7dadeaa6e851 would
>> resolve the scheduling issue.
>>
>
> Good suggestion. cond_resched has become old practice under
> CONFIG_PREEMPT_LAZY
After reconsideration, I think it's not a good idea to remove it if we
treat it as a fix and the fix will be backported to LTS.
Or we just drop both cond_resched and Fixes tag.
next prev parent reply other threads:[~2026-09-08 12:22 UTC|newest]
Thread overview: 11+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-06 7:41 [PATCH bpf v2 0/3] bpf,tcp: Fix bpf_sock_destroy() on TIME_WAIT and listener socks Jiayuan Chen
2026-09-06 7:41 ` [PATCH bpf v2 1/3] bpf: Fix out-of-bounds read of sk_protocol in bpf_sock_destroy() Jiayuan Chen
2026-09-07 23:23 ` Kuniyuki Iwashima
2026-09-06 7:41 ` [PATCH bpf v2 2/3] tcp: Skip cond_resched() in inet_csk_listen_stop() under BPF context Jiayuan Chen
2026-09-06 8:01 ` sashiko-bot
2026-09-06 8:23 ` bot+bpf-ci
2026-09-07 23:34 ` Kuniyuki Iwashima
2026-09-08 8:07 ` Jiayuan Chen
2026-09-08 12:22 ` Jiayuan Chen [this message]
2026-09-06 7:41 ` [PATCH bpf v2 3/3] selftests/bpf: Test bpf_sock_destroy() on TIME_WAIT and listener socks Jiayuan Chen
2026-09-06 7:50 ` sashiko-bot
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=c6c4fd69-3074-4d6d-804b-cf8ff62359da@linux.dev \
--to=jiayuan.chen@linux.dev \
--cc=aditi.ghag@isovalent.com \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=bot+bpf-ci@kernel.org \
--cc=bpf@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=davem@davemloft.net \
--cc=eddyz87@gmail.com \
--cc=edumazet@google.com \
--cc=emil@etsalapatis.com \
--cc=horms@kernel.org \
--cc=ihor.solodrai@linux.dev \
--cc=john.fastabend@gmail.com \
--cc=jolsa@kernel.org \
--cc=kuba@kernel.org \
--cc=kuniyu@google.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=martin.lau@kernel.org \
--cc=martin.lau@linux.dev \
--cc=mason@kernel.org \
--cc=memxor@gmail.com \
--cc=ncardwell@google.com \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=sdf@fomichev.me \
--cc=shuah@kernel.org \
--cc=song@kernel.org \
--cc=yonghong.song@linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.