From: "Ömer Mete Kaya" <omermetekaya0@gmail.com>
To: Ben Greear <greearb@candelatech.com>, linux-wireless@vger.kernel.org
Subject: Re: [PATCH] wifi: cfg80211: avoid holding rtnl_mutex across all cfg80211_leave() calls
Date: Wed, 9 Sep 2026 01:32:52 +0300 [thread overview]
Message-ID: <35a68705-0bcf-4b65-9c0b-050c0344f985@gmail.com> (raw)
In-Reply-To: <bbb382ee-9ada-f060-9128-c2cb6e5f524c@candelatech.com>
On 9/8/26 23:09, Ben Greear wrote:
> On 9/8/26 12:52, Ömer Mete Kaya wrote:
>>
>>
>> On 9/8/26 19:50, Ben Greear wrote:
>>> On 9/8/26 01:58, Ömer Mete Kaya wrote:
>>>>
>>>>
>>>> On 9/7/26 19:38, Ben Greear wrote:
>>>>> On 9/5/26 5:25 PM, Ömer Mete Kaya wrote:
>>>>>> reg_check_chans_work() holds rtnl_mutex for the entire duration of
>>>>>> iterating over all registered devices and calling cfg80211_leave() on
>>>>>> each invalid wdev. cfg80211_leave() can be slow (disconnect, stop AP,
>>>>>> leave mesh), causing rtnl_mutex starvation when many wireless
>>>>>> interfaces
>>>>>> are present. This results in tasks waiting for rtnl_mutex for longer
>>>>>> than hung_task_timeout_secs:
>>>>>>
>>>>>> INFO: task hung in inet_rtm_newaddr
>>>>>> INFO: task hung in inet6_rtm_newaddr
>>>>>> INFO: task hung in nsim_destroy
>>>>>> INFO: task hung in tun_chr_close
>>>>>> INFO: task hung in switchdev_deferred_process_work
>>>>>
>>>>> Hello Omer,
>>>>>
>>>>> Considering that maybe something has mis-diagnosed the problem, could
>>>>> you share details of the stack traces of
>>>>> the hung processes and lockdep output to see if the hang is actually
>>>>> elsewhere? What kernel version are you
>>>>> testing?
>>>>>
>>>>> Thanks,
>>>>> Ben
>>>>>
>>>>
>>>>
>>>> Hi Ben,
>>>>
>>>> Here is the evidence with stack traces and lockdep output. My kernel
>>>> version is 7.2.0-02677-g544d85de4dc2 (net/main HEAD) and both unpatched
>>>> and patched kernels were tested on the same setup(52 mac80211_hwsim
>>>> radios (mac80211_hwsim.radios=51)).
>>>>
>>>> Unpatched kernel:
>>>>
>>>> The lockdep output shows reg_check_chans_work as the rtnl_mutex holder
>>>> and ip as the waiter:
>>>>
>>>> locks held by kworker/0:2/11408: 3, on CPU#0:
>>>> #1: (reg_check_chans).work
>>>> #2: ffffffff91118180 (rtnl_mutex){+.+.}-{4:4},
>>>> at: reg_check_chans_work+0xad/0x1330
>>>>
>>>> locks held by ip/14180: 1, on CPU#0:
>>>> #0: ffffffff91118180 (rtnl_mutex){+.+.}-{4:4},
>>>> at: rtnl_getlink+0xbfb/0x13b0
>>>>
>>>> Hung task call trace:
>>>>
>>>> INFO: task ip:14180 blocked for more than 5 seconds.
>>>> Call Trace:
>>>> __schedule+0x1cba/0x6a40
>>>> schedule+0xe2/0x2e0
>>>> schedule_preempt_disabled+0x13/0x30
>>>> __mutex_lock+0x871/0x1cc0
>>>> rtnl_getlink+0xbfb/0x13b0
>>>> rtnetlink_rcv_msg+0x9a1/0xee0
>>>> netlink_rcv_skb+0x186/0x450
>>>> netlink_sendmsg+0x8e5/0xde0
>>>>
>>>> Kernel panic, not syncing: hung_task: blocked tasks
>>>>
>>>> I injected msleep(200) per wdev in reg_leave_invalid_chans() to model a
>>>> slow cfg80211_leave() operation, to measure the hold duration. With 52
>>>> interfaces this produced approximately 10.5 seconds of
>>>> continuous rtnl_mutex hold:
>>>>
>>>> cfg80211: reg_check_chans_work: rtnl held for 10499 ms
>>>>
>>>> The structural problem is clear regardless of the exact duration:
>>>> rtnl_mutex is held across all per-interface cleanup for every
>>>> registered device in a single acquisition.
>>>>
>>>> Patched kernel:
>>>>
>>>> rtnl_mutex is acquired per-device, each hold is brief:
>>>>
>>>> cfg80211: reg_check_chans_work: device 0 rtnl held 0 ms
>>>> cfg80211: reg_check_chans_work: device 1 rtnl held 0 ms
>>>> ...
>>>> cfg80211: reg_check_chans_work: device 51 rtnl held 0 ms
>>>>
>>>> No hung tasks were observed.
>>>
>>> At least my kernel tends to only warn about hung tasks that are blocked
>>> for minutes. Did you adjust your kernel to make this timeout smaller?
>>
>> Yes, I reduced hung_task_timeout_secs to 5 seconds for the test.
>> With the default 140-second timeout my reproduction would not have
>> triggered a hang, though syzbot does with that timeout. Actually, I've
>> started to have doubts about the extent of its real-world impact, it
>> depends heavily on callback latency and number of active interfaces.
>>
>>> Do you actually have 52 real devices that cause this problem? If so,
>>> what
>>> hardware is this? Or you can only reproduce this with modified
>>> kernel that
>>> injects a 200ms sleep?
>>
>> I used mac80211_hwsim virtual radios,
>> and the hang could only be reproduced with artificial delays.
>> The 200ms value was chosen as a rough estimate of a slow
>> driver callback, not based on measured real hardware latency.
>
> So you created a fake problem to make it fit what something
> thinks is the root cause. I don't think that is helpful.
Agreed, but I already pointed out "my analysis was primarily code-based".
> Possibly root cause below is that something else holds wiphy.mtx, so
> reg_check_chans_work is blocked
> and cannot make progress.
Actually a possibility, I wont be oppose since I cannot prove what
absolutely causes that 143 second timeout.
Do you have full dmesg dump that includes all
> blocked processes
> and full lockdep dump? Are you actually able to reproduce the problem
> w/out adding fake sleeps?
I cant produce from my own machine without fake sleeps because I lack
the actual hardware or a fuzzing ability as syzbot, its impossible with
a laptop and virtual radios. But according to syzbot's full crash report:
INFO: task syz-executor:5585 blocked for more than 143 seconds.
Not tainted syzkaller #0
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
task:syz-executor state:D stack:25304 pid:5585 tgid:5585 ppid:1
task_flags:0x400140 flags:0x00080002
Call Trace:
<TASK>
context_switch kernel/sched/core.c:5510 [inline]
__schedule+0x17d9/0x56c0 kernel/sched/core.c:7234
__schedule_loop kernel/sched/core.c:7311 [inline]
schedule+0x164/0x2b0 kernel/sched/core.c:7326
schedule_preempt_disabled+0x13/0x30 kernel/sched/core.c:7383
__mutex_lock_common kernel/locking/mutex.c:726 [inline]
__mutex_lock+0x7bf/0x1550 kernel/locking/mutex.c:821
rtnl_net_lock include/linux/rtnetlink.h:130 [inline]
inet_rtm_newaddr+0x464/0x19e0 net/ipv4/devinet.c:978
rtnetlink_rcv_msg+0x802/0xc00 net/core/rtnetlink.c:7076
netlink_rcv_skb+0x226/0x4a0 net/netlink/af_netlink.c:2556
netlink_unicast_kernel net/netlink/af_netlink.c:1319 [inline]
netlink_unicast+0x7bb/0x940 net/netlink/af_netlink.c:1345
netlink_sendmsg+0x813/0xb40 net/netlink/af_netlink.c:1900
sock_sendmsg_nosec+0x13a/0x180 net/socket.c:775
__sock_sendmsg net/socket.c:790 [inline]
__sys_sendto+0x408/0x5a0 net/socket.c:2252
__do_sys_sendto net/socket.c:2259 [inline]
__se_sys_sendto net/socket.c:2255 [inline]
__x64_sys_sendto+0xde/0x100 net/socket.c:2255
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0x174/0x580 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
RIP: 0033:0x7f0ee475d68e
RSP: 002b:00007ffe75fcc6c8 EFLAGS: 00000246 ORIG_RAX: 000000000000002c
RAX: ffffffffffffffda RBX: 000055559003e500 RCX: 00007f0ee475d68e
RDX: 0000000000000028 RSI: 00007f0ee5544670 RDI: 0000000000000003
RBP: 0000000000000001 R08: 00007ffe75fcc744 R09: 000000000000000c
R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000003
R13: 0000000000000000 R14: 00007f0ee5544670 R15: 0000000000000000
</TASK>
Showing all locks held in the system:
4 locks held by kworker/0:1/10:
#0: ffff88801aca5540
((wq_completion)events_power_efficient){+.+.}-{0:0}, at:
process_one_work kernel/workqueue.c:3297 [inline]
#0: ffff88801aca5540
((wq_completion)events_power_efficient){+.+.}-{0:0}, at:
process_scheduled_works+0xa20/0x14e0 kernel/workqueue.c:3405
#1: ffffc9000023fc40 ((reg_check_chans).work){+.+.}-{0:0}, at:
process_one_work kernel/workqueue.c:3297 [inline]
#1: ffffc9000023fc40 ((reg_check_chans).work){+.+.}-{0:0}, at:
process_scheduled_works+0xa20/0x14e0 kernel/workqueue.c:3405
#2: ffffffff8fded300 (rtnl_mutex){+.+.}-{4:4}, at:
reg_check_chans_work+0xac/0x1110 net/wireless/reg.c:2469
#3: ffff8880129707a0 (&rdev->wiphy.mtx){+.+.}-{4:4}, at:
class_wiphy_constructor include/net/cfg80211.h:6884 [inline]
#3: ffff8880129707a0 (&rdev->wiphy.mtx){+.+.}-{4:4}, at:
reg_leave_invalid_chans net/wireless/reg.c:2457 [inline]
#3: ffff8880129707a0 (&rdev->wiphy.mtx){+.+.}-{4:4}, at:
reg_check_chans_work+0x1a1/0x1110 net/wireless/reg.c:2472
1 lock held by khungtaskd/25:
#0: ffffffff8e959c20 (rcu_read_lock){....}-{1:3}, at: rcu_lock_acquire
include/linux/rcupdate.h:300 [inline]
#0: ffffffff8e959c20 (rcu_read_lock){....}-{1:3}, at: rcu_read_lock
include/linux/rcupdate.h:840 [inline]
#0: ffffffff8e959c20 (rcu_read_lock){....}-{1:3}, at:
debug_show_all_locks+0x2e/0x180 kernel/locking/lockdep.c:6775
2 locks held by kworker/u4:6/160:
1 lock held by klogd/4693:
1 lock held by dhcpcd/5000:
#0: ffffffff8fded300 (rtnl_mutex){+.+.}-{4:4}, at: rtnl_net_lock
include/linux/rtnetlink.h:130 [inline]
#0: ffffffff8fded300 (rtnl_mutex){+.+.}-{4:4}, at:
inet6_rtm_newaddr+0x65e/0xdf0 net/ipv6/addrconf.c:5062
2 locks held by getty/5093:
#0: ffff88803fac20a0 (&tty->ldisc_sem){++++}-{0:0}, at:
tty_ldisc_ref_wait+0x25/0x70 drivers/tty/tty_ldisc.c:243
#1: ffffc90000ad82e8 (&ldata->atomic_read_lock){+.+.}-{4:4}, at:
n_tty_read+0x45a/0x1360 drivers/tty/n_tty.c:2211
3 locks held by kworker/u4:0/5354:
6 locks held by kworker/u4:1/5428:
#0: ffff88801beb4940 ((wq_completion)netns){+.+.}-{0:0}, at:
process_one_work kernel/workqueue.c:3297 [inline]
#0: ffff88801beb4940 ((wq_completion)netns){+.+.}-{0:0}, at:
process_scheduled_works+0xa20/0x14e0 kernel/workqueue.c:3405
#1: ffffc90002467c40 (net_cleanup_work){+.+.}-{0:0}, at:
process_one_work kernel/workqueue.c:3297 [inline]
#1: ffffc90002467c40 (net_cleanup_work){+.+.}-{0:0}, at:
process_scheduled_works+0xa20/0x14e0 kernel/workqueue.c:3405
#2: ffffffff8fdde748 (pernet_ops_rwsem){++++}-{4:4}, at:
cleanup_net+0xf5/0x810 net/core/net_namespace.c:673
#3: ffff8880122f4128 (&dev->mutex){....}-{4:4}, at: device_lock
include/linux/device.h:1102 [inline]
#3: ffff8880122f4128 (&dev->mutex){....}-{4:4}, at: devl_dev_lock
net/devlink/devl_internal.h:124 [inline]
#3: ffff8880122f4128 (&dev->mutex){....}-{4:4}, at:
devlink_pernet_pre_exit+0x129/0x420 net/devlink/core.c:557
Especially this:
#2: ffffffff8fded300 (rtnl_mutex){+.+.}-{4:4}, at:
reg_check_chans_work+0xac/0x1110 net/wireless/reg.c:2469
#3: ffff8880129707a0 (&rdev->wiphy.mtx){+.+.}-{4:4}, at:
reg_leave_invalid_chans net/wireless/reg.c:2457 [inline]
#3: ffff8880129707a0 (&rdev->wiphy.mtx){+.+.}-{4:4}, at:
reg_check_chans_work+0x1a1/0x1110 net/wireless/reg.c:2472
>
>> actual syzbot crash report :
>>
>> locks held by kworker running reg_check_chans_work:
>> #2: rtnl_mutex, at: reg_check_chans_work+0xac net/wireless/
>> reg.c:2469
>> #3: wiphy.mtx, at: reg_check_chans_work+0x1a1 net/wireless/
>> reg.c:2472
>>
>> locks held by syz-executor (inet6_rtm_newaddr):
>> #0: rtnl_mutex, at: inet6_rtm_newaddr+0x65e
>>
>> INFO: task syz-executor blocked for more than 143 seconds.
>>
>> This directly shows reg_check_chans_work holding rtnl_mutex while
>> inet6_rtm_newaddr waits — 143 seconds, with the default timeout.
>
> You or your LLM is overly confident of root cause in my opinion.
I am not confident, I can stop defending it immediately if I see a
reason to the contrary, I just couldnt find another cause yet. If you
did, please let me know. You said you made a different fix before,
perhaps showing it might open my eyes.
> I will be happy to talk with a human about this since I am interested in
> weird wifi problems at scale, but please do not send more LLM responses.
Was never the case, probably a misunderstanding.
> Thanks,
> Ben
>
Thanks,
Ömer Mete
next prev parent reply other threads:[~2026-09-08 22:32 UTC|newest]
Thread overview: 17+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-03 11:13 [PATCH] wifi: cfg80211: avoid holding rtnl_mutex across all cfg80211_leave() calls Ömer Mete Kaya
2026-09-03 15:13 ` Ömer Mete Kaya
2026-09-03 15:13 ` [PATCH v2] " Ömer Mete Kaya
2026-09-05 18:45 ` Simon Horman
2026-09-06 0:25 ` Ömer Mete Kaya
2026-09-06 0:25 ` [PATCH] " Ömer Mete Kaya
2026-09-06 11:57 ` Johannes Berg
2026-09-07 16:38 ` Ben Greear
2026-09-08 8:58 ` Ömer Mete Kaya
2026-09-08 16:50 ` Ben Greear
2026-09-08 19:52 ` Ömer Mete Kaya
2026-09-08 20:09 ` Ben Greear
2026-09-08 22:32 ` Ömer Mete Kaya [this message]
2026-09-08 23:40 ` Ben Greear
2026-09-09 17:54 ` Ömer Mete Kaya
2026-09-09 15:26 ` netdev-bot+sashiko
-- strict thread matches above, loose matches on Subject: below --
2026-09-03 15:05 Ömer Mete Kaya
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=35a68705-0bcf-4b65-9c0b-050c0344f985@gmail.com \
--to=omermetekaya0@gmail.com \
--cc=greearb@candelatech.com \
--cc=linux-wireless@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.