* [PATCH] can: flexcan: fix synchronous external abort in flexcan_get_berr_counter()
@ 2026-09-15 6:35 chunyuzhiqiang
2026-09-21 9:29 ` Bough Chen
0 siblings, 1 reply; 2+ messages in thread
From: chunyuzhiqiang @ 2026-09-15 6:35 UTC (permalink / raw)
To: imx; +Cc: shawnguo, mkl, ChunYuZhiQiang
From: ChunYuZhiQiang <chunyuzhiqiang@yinhe.ht>
When the FlexCAN interface is down, can_fill_info() still calls
flexcan_get_berr_counter() to fill the error counters for a netlink
dump. This function calls pm_runtime_resume_and_get(), which only
enables the clocks via flexcan_runtime_resume(), but does not clear
the MCR[MDIS] bit. Since the interface has never been opened, the
FlexCAN module is still disabled (MDIS=1) from register_flexcandev().
Accessing the ECR register then triggers a synchronous external abort.
This can be reproduced on an i.MX8QXP board by simply running
`ip link show` without ever bringing the CAN interface up:
Internal error: synchronous external abort: 0000000096000210 [#1] PREEMPT SMP
pc : flexcan_read_le+0x0/0x18
lr : flexcan_get_berr_counter+0x4c/0x8c
Call trace:
flexcan_read_le+0x0/0x18
can_fill_info+0x1f8/0x434
rtnl_fill_ifinfo+0x8fc/0xbb4
rtnl_dump_ifinfo+0x364/0x448
...
ftrace shows the exact path:
can_fill_info() {
flexcan_get_berr_counter() {
__pm_runtime_resume() {
rpm_resume()
rpm_callback()
__rpm_callback()
pm_generic_runtime_resume()
flexcan_runtime_resume()
flexcan_clks_enable() {
clk_prepare(); clk_enable();
...
}
}
do_mem_abort() {
do_sea()
...
}
}
}
Fix this by returning early if the interface is not running.
Tested on i.MX8QXP: after the patch, `ip link show` no longer triggers
the abort.
Cc: stable@vger.kernel.org
Signed-off-by: ChunYuZhiQiang <chunyuzhiqiang@yinhe.ht>
---
drivers/net/can/flexcan/flexcan-core.c | 3 +++
1 file changed, 3 insertions(+)
diff --git a/drivers/net/can/flexcan/flexcan-core.c b/drivers/net/can/flexcan/flexcan-core.c
index 06d5d35fc..b08779ee2 100644
--- a/drivers/net/can/flexcan/flexcan-core.c
+++ b/drivers/net/can/flexcan/flexcan-core.c
@@ -765,6 +765,9 @@ static int flexcan_get_berr_counter(const struct net_device *dev,
const struct flexcan_priv *priv = netdev_priv(dev);
int err;
+ if (!netif_running(dev))
+ return 0;
+
err = pm_runtime_resume_and_get(priv->dev);
if (err < 0)
return err;
--
2.47.3
^ permalink raw reply related [flat|nested] 2+ messages in thread
* Re: [PATCH] can: flexcan: fix synchronous external abort in flexcan_get_berr_counter()
2026-09-15 6:35 [PATCH] can: flexcan: fix synchronous external abort in flexcan_get_berr_counter() chunyuzhiqiang
@ 2026-09-21 9:29 ` Bough Chen
0 siblings, 0 replies; 2+ messages in thread
From: Bough Chen @ 2026-09-21 9:29 UTC (permalink / raw)
To: chunyuzhiqiang; +Cc: imx, shawnguo, mkl
On Tue, Sep 15, 2026 at 02:35:22PM +0800, chunyuzhiqiang wrote:
> From: ChunYuZhiQiang <chunyuzhiqiang@yinhe.ht>
>
> When the FlexCAN interface is down, can_fill_info() still calls
> flexcan_get_berr_counter() to fill the error counters for a netlink
> dump. This function calls pm_runtime_resume_and_get(), which only
> enables the clocks via flexcan_runtime_resume(), but does not clear
> the MCR[MDIS] bit. Since the interface has never been opened, the
> FlexCAN module is still disabled (MDIS=1) from register_flexcandev().
> Accessing the ECR register then triggers a synchronous external abort.
>
Hi ChunYuZhiQiang,
Thanks for the patch. The fix direction is correct, but I think the
root-cause description needs to be corrected before this goes upstream.
The MDIS explanation doesn't hold. Setting MCR[MDIS]=1 only puts the CAN
protocol engine into disable/low-power state; it does *not* make the
register slave interface stop responding to bus accesses. Reading ECR
with MDIS=1 on a powered, clocked FlexCAN module does not raise a
synchronous external abort. In fact register_flexcandev() reads and
writes MCR (and toggles MDIS via flexcan_chip_disable()/
flexcan_chip_enable()) throughout probe without ever aborting. So "the
module is still disabled (MDIS=1), therefore accessing ECR triggers a
SEA" is not the actual mechanism.
The real trigger is platform-specific: the register target is not truly
clocked/powered when the access is issued. On i.MX8QXP the CAN clocks
are gated through the SCU LPCG and the whole subsystem shares the
IMX_SC_R_CAN_0 power domain. flexcan_runtime_resume() only calls
flexcan_clks_enable(); combined with the LPCG e10858 synchronization
erratum, there is a window where pm_runtime_resume_and_get() has
returned but the clock has not actually reached the module yet. The
immediate readl(®s->ecr) in that window hits an un-clocked target and
faults with the external abort (0x96000210, abort on load).
This is why the abort is not reproducible on i.MX95 / i.MX952. Those
SoCs use SCMI-based centralized clocking (IMX95_CLK_CANx /
IMX952_CLK_CANx with BUSWAKEUP/BUSAON), have no LPCG e10858 erratum, and
do not share one power domain across CAN instances. There, after
pm_runtime_resume_and_get() returns the registers are genuinely
accessible, so reading ECR is safe even with MDIS=1. This actually
disproves the MDIS theory: if MDIS were the cause, i.MX95/i.MX952 would
abort too, since MDIS=1 there as well.
So the underlying issue is a logic flaw common to all platforms -
can_fill_info() unconditionally calls do_get_berr_counter() regardless
of interface state - but the SEA only manifests on topologies like
i.MX8QXP. This is the same class of problem that was already fixed for
m_can:
commit 91a55c72a821d ("can: m_can: m_can_get_berr_counter(): don't
wake up controller if interface is down")
Suggestions for v2:
1. Reword the commit message: drop the MDIS reasoning and instead state
that when the interface is down the controller may be unpowered / its
clock not yet settled (e.g. the i.MX8QXP SCU+LPCG topology), so
resuming and immediately accessing ECR can trigger an external abort.
Reference the m_can precedent above.
2. Add a Fixes: tag (the runtime-PM introduction, or ec56acfef2af1 which
first added the clock enable in flexcan_get_berr_counter()).
Thanks,
Bough
> This can be reproduced on an i.MX8QXP board by simply running
> `ip link show` without ever bringing the CAN interface up:
>
> Internal error: synchronous external abort: 0000000096000210 [#1] PREEMPT SMP
> pc : flexcan_read_le+0x0/0x18
> lr : flexcan_get_berr_counter+0x4c/0x8c
> Call trace:
> flexcan_read_le+0x0/0x18
> can_fill_info+0x1f8/0x434
> rtnl_fill_ifinfo+0x8fc/0xbb4
> rtnl_dump_ifinfo+0x364/0x448
> ...
>
> ftrace shows the exact path:
>
> can_fill_info() {
> flexcan_get_berr_counter() {
> __pm_runtime_resume() {
> rpm_resume()
> rpm_callback()
> __rpm_callback()
> pm_generic_runtime_resume()
> flexcan_runtime_resume()
> flexcan_clks_enable() {
> clk_prepare(); clk_enable();
> ...
> }
> }
> do_mem_abort() {
> do_sea()
> ...
> }
> }
> }
>
> Fix this by returning early if the interface is not running.
>
> Tested on i.MX8QXP: after the patch, `ip link show` no longer triggers
> the abort.
>
> Cc: stable@vger.kernel.org
> Signed-off-by: ChunYuZhiQiang <chunyuzhiqiang@yinhe.ht>
> ---
> drivers/net/can/flexcan/flexcan-core.c | 3 +++
> 1 file changed, 3 insertions(+)
>
> diff --git a/drivers/net/can/flexcan/flexcan-core.c b/drivers/net/can/flexcan/flexcan-core.c
> index 06d5d35fc..b08779ee2 100644
> --- a/drivers/net/can/flexcan/flexcan-core.c
> +++ b/drivers/net/can/flexcan/flexcan-core.c
> @@ -765,6 +765,9 @@ static int flexcan_get_berr_counter(const struct net_device *dev,
> const struct flexcan_priv *priv = netdev_priv(dev);
> int err;
>
> + if (!netif_running(dev))
> + return 0;
> +
> err = pm_runtime_resume_and_get(priv->dev);
> if (err < 0)
> return err;
> --
> 2.47.3
>
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-09-21 9:25 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-15 6:35 [PATCH] can: flexcan: fix synchronous external abort in flexcan_get_berr_counter() chunyuzhiqiang
2026-09-21 9:29 ` Bough Chen
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox