Linux CAN drivers development
 help / color / mirror / Atom feed
* [PATCH v2] can: j1939: cancel pending address claim timers on rx release
@ 2026-09-29 10:25 Tetsuo Handa
  2026-09-29 10:40 ` sashiko-bot
  0 siblings, 1 reply; 2+ messages in thread
From: Tetsuo Handa @ 2026-09-29 10:25 UTC (permalink / raw)
  To: linux-can, Marc Kleine-Budde, Oleksij Rempel,
	Robin van der Gracht, kernel, Oliver Hartkopp

syzbot is reporting "struct j1939_ecu" refcount leak, for

  j1939_ecu_get(ecu);
  priv->ents[ecu->addr] = ecu;

in j1939_ecu_map_locked() from j1939_ecu_timer_handler() can succeed
even after

  priv->ents[ecu->addr] = NULL;
  j1939_ecu_put(ecu);

in j1939_ecu_unmap_locked() from j1939_ecu_unmap_all() from
__j1939_rx_release() from j1939_netdev_stop() has completed.

  unregister_netdevice: waiting for vxcan1 to become free. Usage count = 3
  ref_tracker: netdev@ffff8880710f0700 has 1/2 users at
       __netdev_tracker_alloc include/linux/netdevice.h:4496 [inline]
       netdev_hold include/linux/netdevice.h:4525 [inline]
       j1939_ecu_create_locked+0x1c9/0x400 net/can/j1939/bus.c:159
       j1939_local_ecu_get+0xeb/0x220 net/can/j1939/bus.c:293
       j1939_sk_bind+0x70a/0xc60 net/can/j1939/socket.c:529
       __sys_bind_socket net/socket.c:1920 [inline]
       __sys_bind+0x2e3/0x410 net/socket.c:1951
       __do_sys_bind net/socket.c:1956 [inline]
       __se_sys_bind net/socket.c:1954 [inline]
       __x64_sys_bind+0x7a/0x90 net/socket.c:1954
       do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
       do_syscall_64+0x174/0x580 arch/x86/entry/syscall_64.c:94
       entry_SYSCALL_64_after_hwframe+0x77/0x7f

  ref_tracker: netdev@ffff8880710f0700 has 1/2 users at
       __netdev_tracker_alloc include/linux/netdevice.h:4496 [inline]
       netdev_hold include/linux/netdevice.h:4525 [inline]
       j1939_priv_create net/can/j1939/main.c:140 [inline]
       j1939_netdev_start+0x387/0xb20 net/can/j1939/main.c:268
       j1939_sk_bind+0x946/0xc60 net/can/j1939/socket.c:506
       __sys_bind_socket net/socket.c:1920 [inline]
       __sys_bind+0x2e3/0x410 net/socket.c:1951
       __do_sys_bind net/socket.c:1956 [inline]
       __se_sys_bind net/socket.c:1954 [inline]
       __x64_sys_bind+0x7a/0x90 net/socket.c:1954
       do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
       do_syscall_64+0x174/0x580 arch/x86/entry/syscall_64.c:94
       entry_SYSCALL_64_after_hwframe+0x77/0x7f

Fix this race condition by canceling address claim timers before calling
j1939_ecu_unmap_all() from __j1939_rx_release() from j1939_netdev_stop().

Reported-by: syzbot+e2af46126e0644cbebdd@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=e2af46126e0644cbebdd
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928193312.553632-1-mkl%40pengutronix.de # [PATCH net 06/22]
Assisted-by: Gemini-Pro gpt-6-astra opus-5-5
Signed-off-by: Tetsuo Handa <penguin-kernel@I-love.SAKURA.ne.jp>
---
Changes in v2:
  - The "can: j1939: j1939_sk_bind(): fix j1939_ecu leak when re-bind failed"
    patch has passed a review by sashiko@sashiko.dev . But another review by
    sashiko@netdev-ai.bots.linux.dev mentioned that this leak is caused by
    not "re-bind failure" but "j1939_ecu_timer_handler() race".

 net/can/j1939/main.c | 30 ++++++++++++++++++++++++++++++
 1 file changed, 30 insertions(+)

diff --git a/net/can/j1939/main.c b/net/can/j1939/main.c
index 5e5e6c228f22..f9222056ac9a 100644
--- a/net/can/j1939/main.c
+++ b/net/can/j1939/main.c
@@ -212,8 +212,38 @@ static void __j1939_rx_release(struct kref *kref)
 {
 	struct j1939_priv *priv = container_of(kref, struct j1939_priv,
 					       rx_kref);
+	struct j1939_ecu *ecu;
 
 	j1939_can_rx_unregister(priv);
+
+	/* can_rx_unregister() uses call_rcu() internally and is asynchronous.
+	 * We must wait for an RCU grace period to ensure that any in-flight
+	 * j1939_can_recv() instances on other CPUs have fully completed.
+	 */
+	synchronize_rcu();
+
+	/* Cancel all pending address claim timers before unmapping the ECUs.
+	 * This prevents an orphaned timer from re-mapping an ECU after the
+	 * rx path has been completely torn down.
+	 */
+rescan_timers:
+	write_lock_bh(&priv->lock);
+	list_for_each_entry(ecu, &priv->ecus, list) {
+		int ret = hrtimer_try_to_cancel(&ecu->ac_timer);
+
+		if (ret == 0) /* Was not running. */
+			continue;
+		kref_get(&ecu->kref);
+		write_unlock_bh(&priv->lock);
+		if (ret == -1) /* Was running. Wait without lock. */
+			j1939_ecu_timer_cancel(ecu);
+		else /* Was canceled. Drop ref from j1939_ecu_timer_start(). */
+			j1939_ecu_put(ecu);
+		j1939_ecu_put(ecu);
+		goto rescan_timers;
+	}
+	write_unlock_bh(&priv->lock);
+
 	j1939_ecu_unmap_all(priv);
 	j1939_priv_set(priv->ndev, NULL);
 	mutex_unlock(&j1939_netdev_lock);
-- 
2.52.0

^ permalink raw reply related	[flat|nested] 2+ messages in thread

* Re: [PATCH v2] can: j1939: cancel pending address claim timers on rx release
  2026-09-29 10:25 [PATCH v2] can: j1939: cancel pending address claim timers on rx release Tetsuo Handa
@ 2026-09-29 10:40 ` sashiko-bot
  0 siblings, 0 replies; 2+ messages in thread
From: sashiko-bot @ 2026-09-29 10:40 UTC (permalink / raw)
  To: Tetsuo Handa
  Cc: Oleksij Rempel, Oliver Hartkopp, linux-can, Vincent Mailhol,
	Marc Kleine-Budde

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] can: j1939: list corruption via lockless j1939_ecu_put()
--

commit 0dc29971aa78a56ec257a9c34f0e743a7e1b0390
Author: Tetsuo Handa <penguin-kernel@I-love.SAKURA.ne.jp>

can: j1939: cancel pending address claim timers on rx release

This commit attempts to fix a struct j1939_ecu refcount leak by cancelling
address claim timers before calling j1939_ecu_unmap_all() during device
teardown. It adds a rescan loop in __j1939_rx_release() to safely cancel
timers while avoiding deadlocks.

Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928193312.553632-1-mkl%40pengutronix.de

> diff --git a/net/can/j1939/main.c b/net/can/j1939/main.c
> index 5e5e6c228f225..3a4de839cf9a2 100644
> --- a/net/can/j1939/main.c
> +++ b/net/can/j1939/main.c

[ ... ]

> @@ -212,8 +212,38 @@ static void __j1939_rx_release(struct kref *kref)
>  {
>  	struct j1939_priv *priv = container_of(kref, struct j1939_priv,
>  					       rx_kref);
> +	struct j1939_ecu *ecu;
>  
>  	j1939_can_rx_unregister(priv);
> +
> +	/* can_rx_unregister() uses call_rcu() internally and is asynchronous.
> +	 * We must wait for an RCU grace period to ensure that any in-flight
> +	 * j1939_can_recv() instances on other CPUs have fully completed.
> +	 */
> +	synchronize_rcu();
> +
> +	/* Cancel all pending address claim timers before unmapping the ECUs.
> +	 * This prevents an orphaned timer from re-mapping an ECU after the
> +	 * rx path has been completely torn down.
> +	 */
> +rescan_timers:
> +	write_lock_bh(&priv->lock);
> +	list_for_each_entry(ecu, &priv->ecus, list) {
> +		int ret = hrtimer_try_to_cancel(&ecu->ac_timer);
> +
> +		if (ret == 0) /* Was not running. */
> +			continue;
> +		kref_get(&ecu->kref);
> +		write_unlock_bh(&priv->lock);
> +		if (ret == -1) /* Was running. Wait without lock. */
> +			j1939_ecu_timer_cancel(ecu);
> +		else /* Was canceled. Drop ref from j1939_ecu_timer_start(). */
> +			j1939_ecu_put(ecu);
> +		j1939_ecu_put(ecu);

[Severity: High]
Could these lockless j1939_ecu_put() calls trigger list corruption?

When __j1939_rx_release() cancels the timer, it drops priv->lock before
dropping the references. If this ecu's reference count drops to zero here
(for example, if the timer was the last reference), __j1939_ecu_release() 
will run locklessly.

Looking at __j1939_ecu_release() in net/can/j1939/bus.c, it modifies the 
list without acquiring priv->lock:

net/can/j1939/bus.c:__j1939_ecu_release() {
    ...
    list_del(&ecu->list);
    ...
}

Does this deterministic path lead to a kernel panic when other threads traverse
or modify priv->ecus during device teardown?

As an aside, looking at j1939_ac_recv() in net/can/j1939/address-claim.c, it
has a pre-existing example of dropping a reference locklessly. Is it possible 
that the underlying list deletion vulnerability exists elsewhere as well?

> +		goto rescan_timers;
> +	}
> +	write_unlock_bh(&priv->lock);
> +
>  	j1939_ecu_unmap_all(priv);

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/853e770e-4eee-4da8-8cf0-08731f9651ab@I-love.SAKURA.ne.jp?part=1

^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-09-29 10:40 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-29 10:25 [PATCH v2] can: j1939: cancel pending address claim timers on rx release Tetsuo Handa
2026-09-29 10:40 ` sashiko-bot

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox