Netdev List
 help / color / mirror / Atom feed
From: netdev-bot+sashiko@kernel.org
To: mkl@pengutronix.de
Cc: netdev@vger.kernel.org, davem@davemloft.net, kuba@kernel.org,
	linux-can@vger.kernel.org, kernel@pengutronix.de,
	hemanth.selam@gmail.com
Subject: Re: [PATCH net-next 26/37] can: ems_usb: use usb_kill_urb() to stop the intr URB
Date: Sat, 10 Oct 2026 14:03:02 +0000	[thread overview]
Message-ID: <179164098298.434549.8955529744448846100@kernel.org> (raw)
In-Reply-To: <20261009134323.64064-27-mkl@pengutronix.de>

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] The commit message describes the disconnect race and leaves out
  the more easily reached ifdown/reopen failure that this change also…

Pre-existing issues:
- [Critical] Double free of the DMA-coherent RX buffers in
  drivers/net/can/usb/ems_usb.c.
- [Medium] ems_usb_start() returns errors after URBs are already in
  flight and does not stop them, and ems_usb_open() does not clean up…

--

can: ems_usb: use usb_kill_urb() to stop the intr URB

This commit replaces usb_unlink_urb() with usb_kill_urb() for the
interrupt URB in unlink_all_urbs(). Teardown now waits for
ems_usb_read_interrupt_callback() to return before ems_usb_disconnect()
frees intr_in_buffer and the candev.

Link: https://patch.msgid.link/20260923030522.409344-1-fanwu01@zju.edu.cn

> The handler can therefore still be running while ems_usb_disconnect()
> frees its data: it reads the transfer buffer dev->intr_in_buffer, which
> is kfree()d there, and dereferences the private context, which is
> released via free_candev() together with the network device.

[Severity: Low]
This disconnect window does exist, but it is fairly narrow.
ems_usb_driver does not set soft_unbind, so usb_unbind_interface() does
this before calling ->disconnect:

drivers/usb/core/driver.c:usb_unbind_interface() {
    ...
	if (!driver->soft_unbind || udev->state == USB_STATE_NOTATTACHED)
		usb_disable_interface(udev, intf, false);

	driver->disconnect(intf);
    ...
}

usb_disable_interface() reaches usb_hcd_flush_endpoint(), which kills
the URBs still queued on the endpoint. The only case left is a URB that
the HCD has already dequeued but whose BH giveback has not run yet.

The ifdown path seems easier to hit. In ems_usb_close(), the old
usb_unlink_urb() call did not raise urb->reject. So if the intr URB
completed with status 0 at that moment, the handler resubmitted it:

ems_usb_read_interrupt_callback() {
    ...
	case 0:
		dev->free_slots = dev->intr_in_buffer[1];
    ...
	err = usb_submit_urb(urb, GFP_ATOMIC);
    ...
}

The URB then kept polling after ifdown. The next ems_usb_open() failed
with -EBUSY, and usb_submit_urb() printed the "URB submitted while
active" WARN. This patch fixes that case too.

Should the commit message mention it, since stable backporters may use
it to judge impact?

> The handler treats the -ENOENT completion of a killed URB as terminal and
> does not take RTNL or any sleeping lock, so the resubmit loop is cut and
> no RTNL deadlock occurs.

If the URB completes with status 0 just as it is being killed, the
handler does not see -ENOENT. In that case the loop is stopped by
urb->reject in usb_kill_urb(): the resubmit fails and "failed
resubmitting intr urb" is logged. Could the wording reflect both cases?

> diff --git a/drivers/net/can/usb/ems_usb.c b/drivers/net/can/usb/ems_usb.c
> index 24cf8f651f8f1..2d31f0b86859c 100644
> --- a/drivers/net/can/usb/ems_usb.c
> +++ b/drivers/net/can/usb/ems_usb.c
> @@ -748,7 +748,7 @@ static void unlink_all_urbs(struct ems_usb *dev)
>  {
>  	int i;
>  
> -	usb_unlink_urb(dev->intr_urb);
> +	usb_kill_urb(dev->intr_urb);
>  
>  	usb_kill_anchored_urbs(&dev->rx_submitted);
>  

[Severity: Critical]
This is a pre-existing issue and was not introduced by this patch. It is
in the same function and on the same disconnect path, though. Can the
RX coherent buffers be freed twice here?

The loop right after this frees each buffer but never clears the saved
pointers:

	for (i = 0; i < MAX_RX_URBS; ++i)
		usb_free_coherent(dev->udev, RX_BUFFER_SIZE,
				  dev->rxbuf[i], dev->rxbuf_dma[i]);

If the interface has been opened at least once, the first free happens
in ems_usb_close(). That runs on a normal ifdown, or from
unregister_netdev() when the device is unplugged while up.
ems_usb_disconnect() then calls unlink_all_urbs() again:

ems_usb_disconnect()
    unregister_netdev(dev->netdev);   -> ems_usb_close()->unlink_all_urbs()
    unlink_all_urbs(dev);             -> usb_free_coherent() on same rxbuf[]

hcd_buffer_free() only skips NULL addresses. So the stale pointers go on
to dma_pool_free(), kfree() or gen_pool_free() a second time.

A similar case: after an open/close cycle, a later ems_usb_start() can
leave its RX loop early on an alloc or submit failure. The remaining
rxbuf[] slots then still point at buffers that were already freed.

The free loop came from commit 9969e3c5f40c ("can: ems_usb:
ems_usb_start(): fix possible memory leak"). Would setting
dev->rxbuf[i] = NULL after the free fix this? Another option would be to
avoid the second unlink_all_urbs() call in ems_usb_disconnect().

[Severity: Medium]
This is also pre-existing and was not introduced by this patch. Can
ems_usb_start() return an error while URBs are still in flight?

If the intr URB submit fails, the function returns with the RX URBs
still anchored on dev->rx_submitted and running:

	err = usb_submit_urb(dev->intr_urb, GFP_KERNEL);
	if (err) {
		netdev_warn(netdev, "intr URB submit failed: %d\n", err);

		return err;
	}

If a later ems_usb_control_cmd() or ems_usb_write_mode() fails, the
failed label only prints a warning and returns. The intr URB and the RX
URBs keep running.

ems_usb_open() handles this with only close_candev(). ndo_stop is not
called after a failed ndo_open, so unlink_all_urbs() never runs. Both
completion handlers keep resubmitting while the interface is down.

On the next open, usb_submit_urb() on the still-active intr_urb hits:

	if (urb->hcpriv) {
		WARN_ONCE(1, "URB %p submitted while active\n", urb);
		return -EBUSY;
	}

The open keeps failing until the device is replugged.

Each retry also overwrites dev->rxbuf[] and dev->rxbuf_dma[]. That leaks
the coherent buffers from the earlier attempt: disconnect kills their
URBs through the anchor but never frees the buffers.

With this patch, disconnect at least kills the leftover intr URB
synchronously.

Should the ems_usb_start() error paths kill the intr URB and the
anchored RX URBs before returning?

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261009134323.64064-2-mkl%40pengutronix.de

  reply	other threads:[~2026-10-10 14:03 UTC|newest]

Thread overview: 57+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-09 13:27 [PATCH net-next 0/37] pull-request: can-next 2026-10-09 Marc Kleine-Budde
2026-10-09 13:27 ` [PATCH net-next 01/37] can: dev: can_dropped_invalid_skb: drop CAN XL frames on non-CAN XL devices Marc Kleine-Budde
2026-10-09 13:27 ` [PATCH net-next 02/37] can: raw: remove redundant NULL check before netdev_hold() Marc Kleine-Budde
2026-10-09 13:27 ` [PATCH net-next 03/37] can: convert unreliable ARPHRD_CAN type checks to robust can_get_ml_priv() Marc Kleine-Budde
2026-10-10 14:02   ` netdev-bot+sashiko
2026-10-09 13:27 ` [PATCH net-next 04/37] can: proc: reset pkg_stats atomics individually Marc Kleine-Budde
2026-10-09 13:27 ` [PATCH net-next 05/37] can: proc: remove pointers from CAN specific proc output Marc Kleine-Budde
2026-10-10 14:02   ` netdev-bot+sashiko
2026-10-09 13:27 ` [PATCH net-next 06/37] can: j1939: cancel pending address claim timers from j1939_ecu_unmap_all() Marc Kleine-Budde
2026-10-10 14:02   ` netdev-bot+sashiko
2026-10-09 13:27 ` [PATCH net-next 07/37] can: isotp: check the frame type, not just the length Marc Kleine-Budde
2026-10-09 13:27 ` [PATCH net-next 08/37] dt-bindings: can: renesas,rcar-canfd: Document RZ/G3S SoC Marc Kleine-Budde
2026-10-09 13:27 ` [PATCH net-next 09/37] can: rcar_canfd: Fix typos in macro names Marc Kleine-Budde
2026-10-09 13:27 ` [PATCH net-next 10/37] can: skb: make echo skb freeing safe in any IRQ context Marc Kleine-Budde
2026-10-10 14:02   ` netdev-bot+sashiko
2026-10-09 13:27 ` [PATCH net-next 11/37] can: rcar_canfd: Allow the CAN FD clock to be sourced from fck Marc Kleine-Budde
2026-10-10 14:02   ` netdev-bot+sashiko
2026-10-09 13:27 ` [PATCH net-next 12/37] can: skb: make CAN skb allocation failure paths IRQ-safe Marc Kleine-Budde
2026-10-09 13:27 ` [PATCH net-next 13/37] can: rcar_canfd: Do not set registers selecting the CAN mode Marc Kleine-Budde
2026-10-10 14:02   ` netdev-bot+sashiko
2026-10-09 13:27 ` [PATCH net-next 14/37] can: dev: can_put_echo_skb(): free skb on invalid echo index Marc Kleine-Budde
2026-10-09 13:27 ` [PATCH net-next 15/37] can: rcar_canfd: Add support for Renesas RZ/G3S Marc Kleine-Budde
2026-10-10 14:02   ` netdev-bot+sashiko
2026-10-09 13:27 ` [PATCH net-next 16/37] dt-bindings: can: renesas,rcar-canfd: Document RZ/G3L SoC Marc Kleine-Budde
2026-10-10 14:02   ` netdev-bot+sashiko
2026-10-09 13:27 ` [PATCH net-next 17/37] can: rcar_canfd: Derive max_channels from the device tree Marc Kleine-Budde
2026-10-10 14:02   ` netdev-bot+sashiko
2026-10-09 13:27 ` [PATCH net-next 18/37] dt-bindings: net: can: convert grcan to DT schema Marc Kleine-Budde
2026-10-10 14:02   ` netdev-bot+sashiko
2026-10-09 13:27 ` [PATCH net-next 19/37] can: rcar_canfd: Add support for Renesas RZ/G3L Marc Kleine-Budde
2026-10-10 14:02   ` netdev-bot+sashiko
2026-10-09 13:27 ` [PATCH net-next 20/37] dt-bindings: can: renesas,rcar-canfd: Restrict resets in top-level Marc Kleine-Budde
2026-10-10 14:03   ` netdev-bot+sashiko
2026-10-09 13:27 ` [PATCH net-next 21/37] can: grcan: update the binding file reference in the driver comment Marc Kleine-Budde
2026-10-09 13:27 ` [PATCH net-next 22/37] can: remove Softing CANcard driver Marc Kleine-Budde
2026-10-09 13:27 ` [PATCH net-next 23/37] can: Convert to DEFINE_SIMPLE_DEV_PM_OPS() Marc Kleine-Budde
2026-10-09 13:27 ` [PATCH net-next 24/37] can: cc770: don't discard the IRQ lookup error in probe Marc Kleine-Budde
2026-10-10 14:03   ` netdev-bot+sashiko
2026-10-09 13:27 ` [PATCH net-next 25/37] can: cc770: fix the clock divider check on the platform bus Marc Kleine-Budde
2026-10-09 13:27 ` [PATCH net-next 26/37] can: ems_usb: use usb_kill_urb() to stop the intr URB Marc Kleine-Budde
2026-10-10 14:03   ` netdev-bot+sashiko [this message]
2026-10-09 13:27 ` [PATCH net-next 27/37] can: esd: acc_start_xmit(): do not touch skb after can_put_echo_skb() Marc Kleine-Budde
2026-10-10 14:03   ` netdev-bot+sashiko
2026-10-09 13:28 ` [PATCH net-next 28/37] can: flexcan: flexcan_setup_stop_mode_gpr: fix OF node reference leak Marc Kleine-Budde
2026-10-09 13:28 ` [PATCH net-next 29/37] can: f81604: f81604_close(): fix use-after-free on disconnect Marc Kleine-Budde
2026-10-10 14:03   ` netdev-bot+sashiko
2026-10-09 13:28 ` [PATCH net-next 30/37] can: hi311x: drop hi3110_lock before free_irq() on open failure Marc Kleine-Budde
2026-10-10 14:03   ` netdev-bot+sashiko
2026-10-09 13:28 ` [PATCH net-next 31/37] can: kvaser_usb: refactor endpoint lookup Marc Kleine-Budde
2026-10-09 13:28 ` [PATCH net-next 32/37] can: kvaser_usb: validate command format before parsing in hydra receive path Marc Kleine-Budde
2026-10-09 13:28 ` [PATCH net-next 33/37] can: kvaser_pciefd: fix use-after-free in bec poll timer Marc Kleine-Budde
2026-10-09 13:28 ` [PATCH net-next 34/37] can: mcp251xfd: mcp251xfd_probe(): reject devices without match data Marc Kleine-Budde
2026-10-09 13:28 ` [PATCH net-next 35/37] can: sun4i_can: sun4ican_probe(): fix clk leak Marc Kleine-Budde
2026-10-10 14:03   ` netdev-bot+sashiko
2026-10-09 13:28 ` [PATCH net-next 36/37] can: ucan: fix repeated word 'is' in comment Marc Kleine-Budde
2026-10-09 13:28 ` [PATCH net-next 37/37] can: xilinx_can: set CAN FD flags on received frames Marc Kleine-Budde
2026-10-09 13:57 ` [PATCH net-next 0/37] pull-request: can-next 2026-10-09 Marc Kleine-Budde

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=179164098298.434549.8955529744448846100@kernel.org \
    --to=netdev-bot+sashiko@kernel.org \
    --cc=davem@davemloft.net \
    --cc=hemanth.selam@gmail.com \
    --cc=kernel@pengutronix.de \
    --cc=kuba@kernel.org \
    --cc=linux-can@vger.kernel.org \
    --cc=mkl@pengutronix.de \
    --cc=netdev@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox