* [PATCH v2] can: j1939: cancel pending address claim timers on rx release
@ 2026-09-29 10:25 Tetsuo Handa
2026-09-29 10:40 ` sashiko-bot
0 siblings, 1 reply; 2+ messages in thread
From: Tetsuo Handa @ 2026-09-29 10:25 UTC (permalink / raw)
To: linux-can, Marc Kleine-Budde, Oleksij Rempel,
Robin van der Gracht, kernel, Oliver Hartkopp
syzbot is reporting "struct j1939_ecu" refcount leak, for
j1939_ecu_get(ecu);
priv->ents[ecu->addr] = ecu;
in j1939_ecu_map_locked() from j1939_ecu_timer_handler() can succeed
even after
priv->ents[ecu->addr] = NULL;
j1939_ecu_put(ecu);
in j1939_ecu_unmap_locked() from j1939_ecu_unmap_all() from
__j1939_rx_release() from j1939_netdev_stop() has completed.
unregister_netdevice: waiting for vxcan1 to become free. Usage count = 3
ref_tracker: netdev@ffff8880710f0700 has 1/2 users at
__netdev_tracker_alloc include/linux/netdevice.h:4496 [inline]
netdev_hold include/linux/netdevice.h:4525 [inline]
j1939_ecu_create_locked+0x1c9/0x400 net/can/j1939/bus.c:159
j1939_local_ecu_get+0xeb/0x220 net/can/j1939/bus.c:293
j1939_sk_bind+0x70a/0xc60 net/can/j1939/socket.c:529
__sys_bind_socket net/socket.c:1920 [inline]
__sys_bind+0x2e3/0x410 net/socket.c:1951
__do_sys_bind net/socket.c:1956 [inline]
__se_sys_bind net/socket.c:1954 [inline]
__x64_sys_bind+0x7a/0x90 net/socket.c:1954
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0x174/0x580 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
ref_tracker: netdev@ffff8880710f0700 has 1/2 users at
__netdev_tracker_alloc include/linux/netdevice.h:4496 [inline]
netdev_hold include/linux/netdevice.h:4525 [inline]
j1939_priv_create net/can/j1939/main.c:140 [inline]
j1939_netdev_start+0x387/0xb20 net/can/j1939/main.c:268
j1939_sk_bind+0x946/0xc60 net/can/j1939/socket.c:506
__sys_bind_socket net/socket.c:1920 [inline]
__sys_bind+0x2e3/0x410 net/socket.c:1951
__do_sys_bind net/socket.c:1956 [inline]
__se_sys_bind net/socket.c:1954 [inline]
__x64_sys_bind+0x7a/0x90 net/socket.c:1954
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0x174/0x580 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
Fix this race condition by canceling address claim timers before calling
j1939_ecu_unmap_all() from __j1939_rx_release() from j1939_netdev_stop().
Reported-by: syzbot+e2af46126e0644cbebdd@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=e2af46126e0644cbebdd
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928193312.553632-1-mkl%40pengutronix.de # [PATCH net 06/22]
Assisted-by: Gemini-Pro gpt-6-astra opus-5-5
Signed-off-by: Tetsuo Handa <penguin-kernel@I-love.SAKURA.ne.jp>
---
Changes in v2:
- The "can: j1939: j1939_sk_bind(): fix j1939_ecu leak when re-bind failed"
patch has passed a review by sashiko@sashiko.dev . But another review by
sashiko@netdev-ai.bots.linux.dev mentioned that this leak is caused by
not "re-bind failure" but "j1939_ecu_timer_handler() race".
net/can/j1939/main.c | 30 ++++++++++++++++++++++++++++++
1 file changed, 30 insertions(+)
diff --git a/net/can/j1939/main.c b/net/can/j1939/main.c
index 5e5e6c228f22..f9222056ac9a 100644
--- a/net/can/j1939/main.c
+++ b/net/can/j1939/main.c
@@ -212,8 +212,38 @@ static void __j1939_rx_release(struct kref *kref)
{
struct j1939_priv *priv = container_of(kref, struct j1939_priv,
rx_kref);
+ struct j1939_ecu *ecu;
j1939_can_rx_unregister(priv);
+
+ /* can_rx_unregister() uses call_rcu() internally and is asynchronous.
+ * We must wait for an RCU grace period to ensure that any in-flight
+ * j1939_can_recv() instances on other CPUs have fully completed.
+ */
+ synchronize_rcu();
+
+ /* Cancel all pending address claim timers before unmapping the ECUs.
+ * This prevents an orphaned timer from re-mapping an ECU after the
+ * rx path has been completely torn down.
+ */
+rescan_timers:
+ write_lock_bh(&priv->lock);
+ list_for_each_entry(ecu, &priv->ecus, list) {
+ int ret = hrtimer_try_to_cancel(&ecu->ac_timer);
+
+ if (ret == 0) /* Was not running. */
+ continue;
+ kref_get(&ecu->kref);
+ write_unlock_bh(&priv->lock);
+ if (ret == -1) /* Was running. Wait without lock. */
+ j1939_ecu_timer_cancel(ecu);
+ else /* Was canceled. Drop ref from j1939_ecu_timer_start(). */
+ j1939_ecu_put(ecu);
+ j1939_ecu_put(ecu);
+ goto rescan_timers;
+ }
+ write_unlock_bh(&priv->lock);
+
j1939_ecu_unmap_all(priv);
j1939_priv_set(priv->ndev, NULL);
mutex_unlock(&j1939_netdev_lock);
--
2.52.0
^ permalink raw reply related [flat|nested] 2+ messages in thread
* Re: [PATCH v2] can: j1939: cancel pending address claim timers on rx release
2026-09-29 10:25 [PATCH v2] can: j1939: cancel pending address claim timers on rx release Tetsuo Handa
@ 2026-09-29 10:40 ` sashiko-bot
0 siblings, 0 replies; 2+ messages in thread
From: sashiko-bot @ 2026-09-29 10:40 UTC (permalink / raw)
To: Tetsuo Handa
Cc: Oleksij Rempel, Oliver Hartkopp, linux-can, Vincent Mailhol,
Marc Kleine-Budde
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] can: j1939: list corruption via lockless j1939_ecu_put()
--
commit 0dc29971aa78a56ec257a9c34f0e743a7e1b0390
Author: Tetsuo Handa <penguin-kernel@I-love.SAKURA.ne.jp>
can: j1939: cancel pending address claim timers on rx release
This commit attempts to fix a struct j1939_ecu refcount leak by cancelling
address claim timers before calling j1939_ecu_unmap_all() during device
teardown. It adds a rescan loop in __j1939_rx_release() to safely cancel
timers while avoiding deadlocks.
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928193312.553632-1-mkl%40pengutronix.de
> diff --git a/net/can/j1939/main.c b/net/can/j1939/main.c
> index 5e5e6c228f225..3a4de839cf9a2 100644
> --- a/net/can/j1939/main.c
> +++ b/net/can/j1939/main.c
[ ... ]
> @@ -212,8 +212,38 @@ static void __j1939_rx_release(struct kref *kref)
> {
> struct j1939_priv *priv = container_of(kref, struct j1939_priv,
> rx_kref);
> + struct j1939_ecu *ecu;
>
> j1939_can_rx_unregister(priv);
> +
> + /* can_rx_unregister() uses call_rcu() internally and is asynchronous.
> + * We must wait for an RCU grace period to ensure that any in-flight
> + * j1939_can_recv() instances on other CPUs have fully completed.
> + */
> + synchronize_rcu();
> +
> + /* Cancel all pending address claim timers before unmapping the ECUs.
> + * This prevents an orphaned timer from re-mapping an ECU after the
> + * rx path has been completely torn down.
> + */
> +rescan_timers:
> + write_lock_bh(&priv->lock);
> + list_for_each_entry(ecu, &priv->ecus, list) {
> + int ret = hrtimer_try_to_cancel(&ecu->ac_timer);
> +
> + if (ret == 0) /* Was not running. */
> + continue;
> + kref_get(&ecu->kref);
> + write_unlock_bh(&priv->lock);
> + if (ret == -1) /* Was running. Wait without lock. */
> + j1939_ecu_timer_cancel(ecu);
> + else /* Was canceled. Drop ref from j1939_ecu_timer_start(). */
> + j1939_ecu_put(ecu);
> + j1939_ecu_put(ecu);
[Severity: High]
Could these lockless j1939_ecu_put() calls trigger list corruption?
When __j1939_rx_release() cancels the timer, it drops priv->lock before
dropping the references. If this ecu's reference count drops to zero here
(for example, if the timer was the last reference), __j1939_ecu_release()
will run locklessly.
Looking at __j1939_ecu_release() in net/can/j1939/bus.c, it modifies the
list without acquiring priv->lock:
net/can/j1939/bus.c:__j1939_ecu_release() {
...
list_del(&ecu->list);
...
}
Does this deterministic path lead to a kernel panic when other threads traverse
or modify priv->ecus during device teardown?
As an aside, looking at j1939_ac_recv() in net/can/j1939/address-claim.c, it
has a pre-existing example of dropping a reference locklessly. Is it possible
that the underlying list deletion vulnerability exists elsewhere as well?
> + goto rescan_timers;
> + }
> + write_unlock_bh(&priv->lock);
> +
> j1939_ecu_unmap_all(priv);
--
Sashiko AI review · https://sashiko.dev/#/patchset/853e770e-4eee-4da8-8cf0-08731f9651ab@I-love.SAKURA.ne.jp?part=1
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-09-29 10:40 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-29 10:25 [PATCH v2] can: j1939: cancel pending address claim timers on rx release Tetsuo Handa
2026-09-29 10:40 ` sashiko-bot
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox