Linux kernel and device drivers for NXP i.MX platforms
 help / color / mirror / Atom feed
* [PATCH] can: flexcan: fix synchronous external abort in flexcan_get_berr_counter()
@ 2026-09-15  6:35 chunyuzhiqiang
  2026-09-21  9:29 ` Bough Chen
  0 siblings, 1 reply; 2+ messages in thread
From: chunyuzhiqiang @ 2026-09-15  6:35 UTC (permalink / raw)
  To: imx; +Cc: shawnguo, mkl, ChunYuZhiQiang

From: ChunYuZhiQiang <chunyuzhiqiang@yinhe.ht>

When the FlexCAN interface is down, can_fill_info() still calls
flexcan_get_berr_counter() to fill the error counters for a netlink
dump. This function calls pm_runtime_resume_and_get(), which only
enables the clocks via flexcan_runtime_resume(), but does not clear
the MCR[MDIS] bit. Since the interface has never been opened, the
FlexCAN module is still disabled (MDIS=1) from register_flexcandev().
Accessing the ECR register then triggers a synchronous external abort.

This can be reproduced on an i.MX8QXP board by simply running
`ip link show` without ever bringing the CAN interface up:

    Internal error: synchronous external abort: 0000000096000210 [#1] PREEMPT SMP
    pc : flexcan_read_le+0x0/0x18
    lr : flexcan_get_berr_counter+0x4c/0x8c
    Call trace:
     flexcan_read_le+0x0/0x18
     can_fill_info+0x1f8/0x434
     rtnl_fill_ifinfo+0x8fc/0xbb4
     rtnl_dump_ifinfo+0x364/0x448
     ...

ftrace shows the exact path:

  can_fill_info() {
    flexcan_get_berr_counter() {
      __pm_runtime_resume() {
        rpm_resume()
          rpm_callback()
            __rpm_callback()
              pm_generic_runtime_resume()
                flexcan_runtime_resume()
                  flexcan_clks_enable() {
                    clk_prepare(); clk_enable();
                    ...
                  }
      }
      do_mem_abort() {
        do_sea()
        ...
      }
    }
  }

Fix this by returning early if the interface is not running.

Tested on i.MX8QXP: after the patch, `ip link show` no longer triggers
the abort.

Cc: stable@vger.kernel.org
Signed-off-by: ChunYuZhiQiang <chunyuzhiqiang@yinhe.ht>
---
 drivers/net/can/flexcan/flexcan-core.c | 3 +++
 1 file changed, 3 insertions(+)

diff --git a/drivers/net/can/flexcan/flexcan-core.c b/drivers/net/can/flexcan/flexcan-core.c
index 06d5d35fc..b08779ee2 100644
--- a/drivers/net/can/flexcan/flexcan-core.c
+++ b/drivers/net/can/flexcan/flexcan-core.c
@@ -765,6 +765,9 @@ static int flexcan_get_berr_counter(const struct net_device *dev,
 	const struct flexcan_priv *priv = netdev_priv(dev);
 	int err;
 
+	if (!netif_running(dev))
+		return 0;
+
 	err = pm_runtime_resume_and_get(priv->dev);
 	if (err < 0)
 		return err;
-- 
2.47.3


^ permalink raw reply related	[flat|nested] 2+ messages in thread

* Re: [PATCH] can: flexcan: fix synchronous external abort in flexcan_get_berr_counter()
  2026-09-15  6:35 [PATCH] can: flexcan: fix synchronous external abort in flexcan_get_berr_counter() chunyuzhiqiang
@ 2026-09-21  9:29 ` Bough Chen
  0 siblings, 0 replies; 2+ messages in thread
From: Bough Chen @ 2026-09-21  9:29 UTC (permalink / raw)
  To: chunyuzhiqiang; +Cc: imx, shawnguo, mkl

On Tue, Sep 15, 2026 at 02:35:22PM +0800, chunyuzhiqiang wrote:
> From: ChunYuZhiQiang <chunyuzhiqiang@yinhe.ht>
> 
> When the FlexCAN interface is down, can_fill_info() still calls
> flexcan_get_berr_counter() to fill the error counters for a netlink
> dump. This function calls pm_runtime_resume_and_get(), which only
> enables the clocks via flexcan_runtime_resume(), but does not clear
> the MCR[MDIS] bit. Since the interface has never been opened, the
> FlexCAN module is still disabled (MDIS=1) from register_flexcandev().
> Accessing the ECR register then triggers a synchronous external abort.
> 

Hi ChunYuZhiQiang,

Thanks for the patch. The fix direction is correct, but I think the
root-cause description needs to be corrected before this goes upstream.

The MDIS explanation doesn't hold. Setting MCR[MDIS]=1 only puts the CAN
protocol engine into disable/low-power state; it does *not* make the
register slave interface stop responding to bus accesses. Reading ECR
with MDIS=1 on a powered, clocked FlexCAN module does not raise a
synchronous external abort. In fact register_flexcandev() reads and
writes MCR (and toggles MDIS via flexcan_chip_disable()/
flexcan_chip_enable()) throughout probe without ever aborting. So "the
module is still disabled (MDIS=1), therefore accessing ECR triggers a
SEA" is not the actual mechanism.

The real trigger is platform-specific: the register target is not truly
clocked/powered when the access is issued. On i.MX8QXP the CAN clocks
are gated through the SCU LPCG and the whole subsystem shares the
IMX_SC_R_CAN_0 power domain. flexcan_runtime_resume() only calls
flexcan_clks_enable(); combined with the LPCG e10858 synchronization
erratum, there is a window where pm_runtime_resume_and_get() has
returned but the clock has not actually reached the module yet. The
immediate readl(&regs->ecr) in that window hits an un-clocked target and
faults with the external abort (0x96000210, abort on load).

This is why the abort is not reproducible on i.MX95 / i.MX952. Those
SoCs use SCMI-based centralized clocking (IMX95_CLK_CANx /
IMX952_CLK_CANx with BUSWAKEUP/BUSAON), have no LPCG e10858 erratum, and
do not share one power domain across CAN instances. There, after
pm_runtime_resume_and_get() returns the registers are genuinely
accessible, so reading ECR is safe even with MDIS=1. This actually
disproves the MDIS theory: if MDIS were the cause, i.MX95/i.MX952 would
abort too, since MDIS=1 there as well.

So the underlying issue is a logic flaw common to all platforms -
can_fill_info() unconditionally calls do_get_berr_counter() regardless
of interface state - but the SEA only manifests on topologies like
i.MX8QXP. This is the same class of problem that was already fixed for
m_can:

  commit 91a55c72a821d ("can: m_can: m_can_get_berr_counter(): don't
  wake up controller if interface is down")

Suggestions for v2:

1. Reword the commit message: drop the MDIS reasoning and instead state
   that when the interface is down the controller may be unpowered / its
   clock not yet settled (e.g. the i.MX8QXP SCU+LPCG topology), so
   resuming and immediately accessing ECR can trigger an external abort.
   Reference the m_can precedent above.

2. Add a Fixes: tag (the runtime-PM introduction, or ec56acfef2af1 which
   first added the clock enable in flexcan_get_berr_counter()).


Thanks,
Bough

> This can be reproduced on an i.MX8QXP board by simply running
> `ip link show` without ever bringing the CAN interface up:
> 
>     Internal error: synchronous external abort: 0000000096000210 [#1] PREEMPT SMP
>     pc : flexcan_read_le+0x0/0x18
>     lr : flexcan_get_berr_counter+0x4c/0x8c
>     Call trace:
>      flexcan_read_le+0x0/0x18
>      can_fill_info+0x1f8/0x434
>      rtnl_fill_ifinfo+0x8fc/0xbb4
>      rtnl_dump_ifinfo+0x364/0x448
>      ...
> 
> ftrace shows the exact path:
> 
>   can_fill_info() {
>     flexcan_get_berr_counter() {
>       __pm_runtime_resume() {
>         rpm_resume()
>           rpm_callback()
>             __rpm_callback()
>               pm_generic_runtime_resume()
>                 flexcan_runtime_resume()
>                   flexcan_clks_enable() {
>                     clk_prepare(); clk_enable();
>                     ...
>                   }
>       }
>       do_mem_abort() {
>         do_sea()
>         ...
>       }
>     }
>   }
> 
> Fix this by returning early if the interface is not running.
> 
> Tested on i.MX8QXP: after the patch, `ip link show` no longer triggers
> the abort.
> 
> Cc: stable@vger.kernel.org
> Signed-off-by: ChunYuZhiQiang <chunyuzhiqiang@yinhe.ht>
> ---
>  drivers/net/can/flexcan/flexcan-core.c | 3 +++
>  1 file changed, 3 insertions(+)
> 
> diff --git a/drivers/net/can/flexcan/flexcan-core.c b/drivers/net/can/flexcan/flexcan-core.c
> index 06d5d35fc..b08779ee2 100644
> --- a/drivers/net/can/flexcan/flexcan-core.c
> +++ b/drivers/net/can/flexcan/flexcan-core.c
> @@ -765,6 +765,9 @@ static int flexcan_get_berr_counter(const struct net_device *dev,
>  	const struct flexcan_priv *priv = netdev_priv(dev);
>  	int err;
>  
> +	if (!netif_running(dev))
> +		return 0;
> +
>  	err = pm_runtime_resume_and_get(priv->dev);
>  	if (err < 0)
>  		return err;
> -- 
> 2.47.3
> 

^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-09-21  9:25 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-15  6:35 [PATCH] can: flexcan: fix synchronous external abort in flexcan_get_berr_counter() chunyuzhiqiang
2026-09-21  9:29 ` Bough Chen

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox