Linux kernel and device drivers for NXP i.MX platforms
 help / color / mirror / Atom feed
From: Bough Chen <haibo.chen@oss.nxp.com>
To: chunyuzhiqiang <chunyuzhiqiang@yinhe.ht>
Cc: imx@lists.linux.dev, shawnguo@kernel.org, mkl@pengutronix.de
Subject: Re: [PATCH] can: flexcan: fix synchronous external abort in flexcan_get_berr_counter()
Date: Mon, 21 Sep 2026 17:29:23 +0800	[thread overview]
Message-ID: <20260921092923.jwyluujse2uet3js@shlinux89> (raw)
In-Reply-To: <20260915063522.3010750-1-chunyuzhiqiang@yinhe.ht>

On Tue, Sep 15, 2026 at 02:35:22PM +0800, chunyuzhiqiang wrote:
> From: ChunYuZhiQiang <chunyuzhiqiang@yinhe.ht>
> 
> When the FlexCAN interface is down, can_fill_info() still calls
> flexcan_get_berr_counter() to fill the error counters for a netlink
> dump. This function calls pm_runtime_resume_and_get(), which only
> enables the clocks via flexcan_runtime_resume(), but does not clear
> the MCR[MDIS] bit. Since the interface has never been opened, the
> FlexCAN module is still disabled (MDIS=1) from register_flexcandev().
> Accessing the ECR register then triggers a synchronous external abort.
> 

Hi ChunYuZhiQiang,

Thanks for the patch. The fix direction is correct, but I think the
root-cause description needs to be corrected before this goes upstream.

The MDIS explanation doesn't hold. Setting MCR[MDIS]=1 only puts the CAN
protocol engine into disable/low-power state; it does *not* make the
register slave interface stop responding to bus accesses. Reading ECR
with MDIS=1 on a powered, clocked FlexCAN module does not raise a
synchronous external abort. In fact register_flexcandev() reads and
writes MCR (and toggles MDIS via flexcan_chip_disable()/
flexcan_chip_enable()) throughout probe without ever aborting. So "the
module is still disabled (MDIS=1), therefore accessing ECR triggers a
SEA" is not the actual mechanism.

The real trigger is platform-specific: the register target is not truly
clocked/powered when the access is issued. On i.MX8QXP the CAN clocks
are gated through the SCU LPCG and the whole subsystem shares the
IMX_SC_R_CAN_0 power domain. flexcan_runtime_resume() only calls
flexcan_clks_enable(); combined with the LPCG e10858 synchronization
erratum, there is a window where pm_runtime_resume_and_get() has
returned but the clock has not actually reached the module yet. The
immediate readl(&regs->ecr) in that window hits an un-clocked target and
faults with the external abort (0x96000210, abort on load).

This is why the abort is not reproducible on i.MX95 / i.MX952. Those
SoCs use SCMI-based centralized clocking (IMX95_CLK_CANx /
IMX952_CLK_CANx with BUSWAKEUP/BUSAON), have no LPCG e10858 erratum, and
do not share one power domain across CAN instances. There, after
pm_runtime_resume_and_get() returns the registers are genuinely
accessible, so reading ECR is safe even with MDIS=1. This actually
disproves the MDIS theory: if MDIS were the cause, i.MX95/i.MX952 would
abort too, since MDIS=1 there as well.

So the underlying issue is a logic flaw common to all platforms -
can_fill_info() unconditionally calls do_get_berr_counter() regardless
of interface state - but the SEA only manifests on topologies like
i.MX8QXP. This is the same class of problem that was already fixed for
m_can:

  commit 91a55c72a821d ("can: m_can: m_can_get_berr_counter(): don't
  wake up controller if interface is down")

Suggestions for v2:

1. Reword the commit message: drop the MDIS reasoning and instead state
   that when the interface is down the controller may be unpowered / its
   clock not yet settled (e.g. the i.MX8QXP SCU+LPCG topology), so
   resuming and immediately accessing ECR can trigger an external abort.
   Reference the m_can precedent above.

2. Add a Fixes: tag (the runtime-PM introduction, or ec56acfef2af1 which
   first added the clock enable in flexcan_get_berr_counter()).


Thanks,
Bough

> This can be reproduced on an i.MX8QXP board by simply running
> `ip link show` without ever bringing the CAN interface up:
> 
>     Internal error: synchronous external abort: 0000000096000210 [#1] PREEMPT SMP
>     pc : flexcan_read_le+0x0/0x18
>     lr : flexcan_get_berr_counter+0x4c/0x8c
>     Call trace:
>      flexcan_read_le+0x0/0x18
>      can_fill_info+0x1f8/0x434
>      rtnl_fill_ifinfo+0x8fc/0xbb4
>      rtnl_dump_ifinfo+0x364/0x448
>      ...
> 
> ftrace shows the exact path:
> 
>   can_fill_info() {
>     flexcan_get_berr_counter() {
>       __pm_runtime_resume() {
>         rpm_resume()
>           rpm_callback()
>             __rpm_callback()
>               pm_generic_runtime_resume()
>                 flexcan_runtime_resume()
>                   flexcan_clks_enable() {
>                     clk_prepare(); clk_enable();
>                     ...
>                   }
>       }
>       do_mem_abort() {
>         do_sea()
>         ...
>       }
>     }
>   }
> 
> Fix this by returning early if the interface is not running.
> 
> Tested on i.MX8QXP: after the patch, `ip link show` no longer triggers
> the abort.
> 
> Cc: stable@vger.kernel.org
> Signed-off-by: ChunYuZhiQiang <chunyuzhiqiang@yinhe.ht>
> ---
>  drivers/net/can/flexcan/flexcan-core.c | 3 +++
>  1 file changed, 3 insertions(+)
> 
> diff --git a/drivers/net/can/flexcan/flexcan-core.c b/drivers/net/can/flexcan/flexcan-core.c
> index 06d5d35fc..b08779ee2 100644
> --- a/drivers/net/can/flexcan/flexcan-core.c
> +++ b/drivers/net/can/flexcan/flexcan-core.c
> @@ -765,6 +765,9 @@ static int flexcan_get_berr_counter(const struct net_device *dev,
>  	const struct flexcan_priv *priv = netdev_priv(dev);
>  	int err;
>  
> +	if (!netif_running(dev))
> +		return 0;
> +
>  	err = pm_runtime_resume_and_get(priv->dev);
>  	if (err < 0)
>  		return err;
> -- 
> 2.47.3
> 

      reply	other threads:[~2026-09-21  9:25 UTC|newest]

Thread overview: 2+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-15  6:35 [PATCH] can: flexcan: fix synchronous external abort in flexcan_get_berr_counter() chunyuzhiqiang
2026-09-21  9:29 ` Bough Chen [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260921092923.jwyluujse2uet3js@shlinux89 \
    --to=haibo.chen@oss.nxp.com \
    --cc=chunyuzhiqiang@yinhe.ht \
    --cc=imx@lists.linux.dev \
    --cc=mkl@pengutronix.de \
    --cc=shawnguo@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox