* [PATCH] can: j1939: fix memory leaks caused by pending address claim timer
@ 2026-09-01 7:05 syzbot
2026-09-01 7:19 ` sashiko-bot
0 siblings, 1 reply; 2+ messages in thread
From: syzbot @ 2026-09-01 7:05 UTC (permalink / raw)
To: syzkaller-bugs, Slawomir Stepien, linux-can, Marc Kleine-Budde,
Oleksij Rempel, Robin van der Gracht, Oliver Hartkopp
Cc: kernel, linux-kernel, syzbot
From: Slawomir Stepien <sst@poczta.fm>
A memory leak of struct j1939_ecu and struct j1939_priv objects occurs when
a CAN netdev is stopped while an Address Claim timer is pending:
BUG: memory leak
unreferenced object 0xffff888198faa000 (size 8192):
backtrace (crc 42a920ae):
j1939_priv_create net/can/j1939/main.c:131 [inline]
j1939_netdev_start+0x11b/0x5c0 net/can/j1939/main.c:268
j1939_sk_bind+0x42c/0x4a0 net/can/j1939/socket.c:506
__sys_bind+0x1fa/0x2c0 net/socket.c:1951
__x64_sys_bind+0x1c/0x30 net/socket.c:1954
do_syscall_64+0x14f/0x3c0 arch/x86/entry/syscall_64.c:94
BUG: memory leak
unreferenced object 0xffff8881978f36c0 (size 192):
backtrace (crc f1bed932):
j1939_ecu_create_locked+0x4e/0x1c0 net/can/j1939/bus.c:155
j1939_local_ecu_get+0xe6/0x1b0 net/can/j1939/bus.c:293
j1939_sk_bind+0x300/0x4a0 net/can/j1939/socket.c:529
__sys_bind+0x1fa/0x2c0 net/socket.c:1951
__x64_sys_bind+0x1c/0x30 net/socket.c:1954
do_syscall_64+0x14f/0x3c0 arch/x86/entry/syscall_64.c:94
When an Address Claim message is processed, j1939_ecu_timer_start() starts
a 250ms timer (ecu->ac_timer) and acquires a reference to the ECU. At this
stage, the ECU is linked to priv->ecus but not yet mapped into priv->ents.
If the socket is closed or the netdev is stopped before the timer expires,
j1939_netdev_stop() calls __j1939_rx_release(), which invokes
j1939_ecu_unmap_all(). However, j1939_ecu_unmap_all() only unmapped entries
in priv->ents, leaving the pending ECU timer running. When ecu->ac_timer
expires after netdev teardown, j1939_ecu_timer_handler() unconditionally
maps the ECU into priv->ents of the stopped priv instance, where it will
never be unmapped or freed. Since the leaked ECU retains a reference to
priv, both the ECU and priv structures are leaked.
Furthermore, calling hrtimer_cancel() while holding priv->lock can lead to
a deadlock if j1939_ecu_timer_handler() runs concurrently on another CPU
waiting to acquire priv->lock.
Fix the memory leak and deadlock by:
- Iterating over priv->ecus in j1939_ecu_unmap_all() to cancel any pending
ECU timers with j1939_ecu_timer_cancel().
- Checking kref_read(&priv->rx_kref) > 0 in j1939_ecu_timer_handler()
before calling j1939_ecu_map_locked(ecu) so that unmapped ECUs are not
mapped into a stopped netdev.
- Using hrtimer_try_to_cancel() in j1939_ecu_timer_cancel() to avoid
spinning with priv->lock held.
- Taking temporary references with j1939_priv_get() and j1939_priv_put()
around j1939_ecu_unmap_all() and j1939_ecu_timer_handler() to ensure priv
memory remains valid while held.
Fixes: 9d71dd0c7009 ("can: add support of SAE J1939 protocol")
Assisted-by: Gemini:gemini-3.7-flash Gemini:gemini-3.1-pro-preview syzbot
Reported-by: syzbot+489e907b2a026a6f5fa0@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=489e907b2a026a6f5fa0
Link: https://syzkaller.appspot.com/ai_job?id=c118979e-3d39-430b-b07e-608d26aeaa70
Signed-off-by: Slawomir Stepien <sst@poczta.fm>
---
diff --git a/net/can/j1939/bus.c b/net/can/j1939/bus.c
index cdc3c0a71..71e52f25a 100644
--- a/net/can/j1939/bus.c
+++ b/net/can/j1939/bus.c
@@ -98,13 +98,18 @@ void j1939_ecu_unmap(struct j1939_ecu *ecu)
void j1939_ecu_unmap_all(struct j1939_priv *priv)
{
+ struct j1939_ecu *ecu, *tmp;
int i;
+ j1939_priv_get(priv);
write_lock_bh(&priv->lock);
for (i = 0; i < ARRAY_SIZE(priv->ents); i++)
if (priv->ents[i].ecu)
j1939_ecu_unmap_locked(priv->ents[i].ecu);
+ list_for_each_entry_safe(ecu, tmp, &priv->ecus, list)
+ j1939_ecu_timer_cancel(ecu);
write_unlock_bh(&priv->lock);
+ j1939_priv_put(priv);
}
void j1939_ecu_timer_start(struct j1939_ecu *ecu)
@@ -121,7 +126,7 @@ void j1939_ecu_timer_start(struct j1939_ecu *ecu)
void j1939_ecu_timer_cancel(struct j1939_ecu *ecu)
{
- if (hrtimer_cancel(&ecu->ac_timer))
+ if (hrtimer_try_to_cancel(&ecu->ac_timer) == 1)
j1939_ecu_put(ecu);
}
@@ -131,17 +136,18 @@ static enum hrtimer_restart j1939_ecu_timer_handler(struct hrtimer *hrtimer)
container_of(hrtimer, struct j1939_ecu, ac_timer);
struct j1939_priv *priv = ecu->priv;
+ j1939_priv_get(priv);
write_lock_bh(&priv->lock);
- /* TODO: can we test if ecu->addr is unicast before starting
- * the timer?
- */
- j1939_ecu_map_locked(ecu);
+ /* Only map the ECU if the netdev is still active */
+ if (kref_read(&priv->rx_kref) > 0)
+ j1939_ecu_map_locked(ecu);
/* The corresponding j1939_ecu_get() is in
* j1939_ecu_timer_start().
*/
j1939_ecu_put(ecu);
write_unlock_bh(&priv->lock);
+ j1939_priv_put(priv);
return HRTIMER_NORESTART;
}
base-commit: 8d3ae59288f1e7d58d76558a6ee96d533bc5019f
--
See https://goo.gle/syzbot-ai-patches for information about AI-generated patches.
The person who has signed off on the patch is responsible for
addressing comments.
syzbot engineers can be reached at syzkaller@googlegroups.com.
^ permalink raw reply related [flat|nested] 2+ messages in thread* Re: [PATCH] can: j1939: fix memory leaks caused by pending address claim timer
2026-09-01 7:05 [PATCH] can: j1939: fix memory leaks caused by pending address claim timer syzbot
@ 2026-09-01 7:19 ` sashiko-bot
0 siblings, 0 replies; 2+ messages in thread
From: sashiko-bot @ 2026-09-01 7:19 UTC (permalink / raw)
To: syzbot
Cc: Vincent Mailhol, Oliver Hartkopp, Marc Kleine-Budde,
Oleksij Rempel, linux-can
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
New issues:
- [High] Ignoring hrtimer_try_to_cancel() failure introduces race conditions leading to inappropriate ECU mapping (NETDEV_DOWN, IDLE claims, and delay bypass).
Pre-existing issues:
- [Critical] Lockless list_del in __j1939_ecu_release leads to linked list corruption and use-after-free.
--
commit 2b6762bc65c5c7e152f9be9bd0a6d303989290db
Author: Slawomir Stepien <sst@poczta.fm>
can: j1939: fix memory leaks caused by pending address claim timer
Fixes a memory leak and deadlock involving pending Address Claim timers during
netdev teardown. It cancels timers during j1939_ecu_unmap_all() and adds
kref_read(&priv->rx_kref) checks to prevent remapping ECUs on stopped
interfaces.
Link: https://syzkaller.appspot.com/ai_job?id=c118979e-3d39-430b-b07e-608d26aeaa70
> diff --git a/net/can/j1939/bus.c b/net/can/j1939/bus.c
> index cdc3c0a719379..71e52f25aa604 100644
> --- a/net/can/j1939/bus.c
> +++ b/net/can/j1939/bus.c
[ ... ]
> @@ -121,7 +126,7 @@ void j1939_ecu_timer_start(struct j1939_ecu *ecu)
>
> void j1939_ecu_timer_cancel(struct j1939_ecu *ecu)
> {
> - if (hrtimer_cancel(&ecu->ac_timer))
> + if (hrtimer_try_to_cancel(&ecu->ac_timer) == 1)
> j1939_ecu_put(ecu);
> }
[Severity: High]
Does ignoring the -1 return value from hrtimer_try_to_cancel() introduce race
conditions that could lead to inappropriate ECU mapping?
The patch replaces a deadlocking hrtimer_cancel() with hrtimer_try_to_cancel()
but ignores failures. If the timer callback is running concurrently on another
CPU, it will block on write_lock_bh(&priv->lock).
If the current CPU executes j1939_ecu_unmap_all() or processes an IDLE claim
in j1939_ac_process(), it holds priv->lock, calls hrtimer_try_to_cancel()
(which fails and returns -1), modifies the state, and unlocks.
The other CPU running the timer callback then acquires the lock and could
execute j1939_ecu_map_locked() based on old state. Could this cause the ECU to
be mapped back onto a DOWN interface, ignore IDLE claims by re-mapping the old
address, or bypass the J1939 250ms delay by immediately mapping the new
address?
[ ... ]
> @@ -131,17 +136,18 @@ static enum hrtimer_restart j1939_ecu_timer_handler(struct hrtimer *hrtimer)
> container_of(hrtimer, struct j1939_ecu, ac_timer);
> struct j1939_priv *priv = ecu->priv;
>
> + j1939_priv_get(priv);
> write_lock_bh(&priv->lock);
> - /* TODO: can we test if ecu->addr is unicast before starting
> - * the timer?
> - */
> - j1939_ecu_map_locked(ecu);
> + /* Only map the ECU if the netdev is still active */
> + if (kref_read(&priv->rx_kref) > 0)
> + j1939_ecu_map_locked(ecu);
>
> /* The corresponding j1939_ecu_get() is in
> * j1939_ecu_timer_start().
> */
> j1939_ecu_put(ecu);
[Severity: Critical]
This is a pre-existing issue, but could dropping the last ECU reference
outside of priv->lock lead to list corruption?
In __j1939_ecu_release(), there is a lockless call to list_del(&ecu->list).
If CPU 1 receives an address claim, j1939_ac_recv() takes a temporary
reference via j1939_ecu_get_by_addr(). If CPU 2 concurrently executes
j1939_ecu_unmap_all(), unmapping the ECU and dropping its mapping reference,
then when CPU 1 finishes processing it calls j1939_ecu_put(), dropping the
final reference.
Because paths like j1939_ac_recv() and j1939_ac_fixup() call
j1939_ecu_put(ecu) without holding priv->lock, does this cause
__j1939_ecu_release() to execute list_del(&ecu->list) locklessly, resulting
in list corruption and use-after-free crashes on subsequent traversals?
> write_unlock_bh(&priv->lock);
> + j1939_priv_put(priv);
>
> return HRTIMER_NORESTART;
> }
--
Sashiko AI review · https://sashiko.dev/#/patchset/1ca51efc-8929-4df1-ab11-72407b6d496f@mail.kernel.org?part=1
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-09-01 7:19 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-01 7:05 [PATCH] can: j1939: fix memory leaks caused by pending address claim timer syzbot
2026-09-01 7:19 ` sashiko-bot
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox