From: Kohei Enju <enju.kohei@fujitsu.com>
To: Marc Zyngier <maz@kernel.org>
Cc: Tomohiro Misono <misono.tomohiro@fujitsu.com>,
Catalin Marinas <catalin.marinas@arm.com>,
Will Deacon <will@kernel.org>,
Mark Rutland <mark.rutland@arm.com>,
Jonathan Corbet <corbet@lwn.net>,
Shuah Khan <skhan@linuxfoundation.org>,
Randy Dunlap <rdunlap@infradead.org>,
Thomas Gleixner <tglx@kernel.org>, Radu Rendec <radu@rendec.net>,
linux-arm-kernel@lists.infradead.org,
linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org,
linux-perf-users@vger.kernel.org
Subject: Re: [PATCH 4/4] irqchip/gicv3: Add workaround for FUJITSU-MONAKA erratum E#030003
Date: Tue, 6 Oct 2026 18:08:19 +0900 [thread overview]
Message-ID: <asS5r0WNCQpMIo4p@FCCLS0092175.localdomain> (raw)
In-Reply-To: <86tsn42sqr.wl-maz@kernel.org>
On 10/02 13:26, Marc Zyngier wrote:
> On Fri, 02 Oct 2026 11:26:52 +0100,
> Tomohiro Misono <misono.tomohiro@fujitsu.com> wrote:
> >
> > From: Kohei Enju <enju.kohei@fujitsu.com>
> >
> > On affected FUJITSU-MONAKA CPUs, an SGI generated by writing to
> > ICC_SGI0R_EL1, ICC_SGI1R_EL1, or ICC_ASGI1R_EL1 may be lost if the
> > operation races with CPU interface processing triggered by the arrival
> > of a higher-priority interrupt, an update to a pending interrupt, or a
> > transition of the PE to the Sleep state. When this occurs, the system
> > register write does not complete, causing the issuing core to hang.
> >
>
> Is the SGI lost? Or is the sender core hanging?
Both: the SGI is lost, and the sysreg write does not complete, causing
the sending core to hang.
>
> I expect that virtualised accesses to these registers are not
> affected, since they trap, but it'd be good to document this.
You're right. I'll document it in v2.
>
> > Work around the erratum by writing 1 to the bit corresponding to the
> > target SGI INTID in GICR_ISPENDR0 of each target PE's GIC Redistributor,
> > instead of writing to the affected system registers.
> >
> > The affinity topology of affected systems has Aff0 == 0 for every PE,
> > with PEs distinguished by higher affinity levels. Consequently, an
> > ICC_SGI1R_EL1 write can target only one PE, so the workaround does not
> > increase the number of writes required to send an SGI to multiple PEs.
> >
> > Accessing a target PE's GICR_ISPENDR0 requires its Redistributor base to
> > have been discovered. Affected platforms conform to SBBR, which requires
> > PSCI for secondary CPU boot. They therefore do not use the ACPI Parking
> > protocol, which sends an IPI before the secondary CPU has initialized
> > its Redistributor. With PSCI, a secondary CPU discovers its
> > Redistributor before becoming an IPI target.
> >
> > Signed-off-by: Kohei Enju <enju.kohei@fujitsu.com>
> > ---
> > Documentation/arch/arm64/silicon-errata.rst | 2 ++
> > drivers/irqchip/irq-gic-v3.c | 48 +++++++++++++++++++++++++++++
> > 2 files changed, 50 insertions(+)
> >
> > diff --git a/Documentation/arch/arm64/silicon-errata.rst b/Documentation/arch/arm64/silicon-errata.rst
> > index 429bdf9a3d9d..13328a30df59 100644
> > --- a/Documentation/arch/arm64/silicon-errata.rst
> > +++ b/Documentation/arch/arm64/silicon-errata.rst
> > @@ -377,6 +377,8 @@ stable kernels.
> > +----------------+-----------------+-----------------+-----------------------------+
> > | Fujitsu | MONAKA | E#030002 | FUJITSU_ERRATUM_030002 |
> > +----------------+-----------------+-----------------+-----------------------------+
> > +| Fujitsu | MONAKA GICv3/v4 | E#030003 | N/A |
> > ++----------------+-----------------+-----------------+-----------------------------+
> > +----------------+-----------------+-----------------+-----------------------------+
> > | ASR | ASR8601 | #8601001 | N/A |
> > +----------------+-----------------+-----------------+-----------------------------+
> > diff --git a/drivers/irqchip/irq-gic-v3.c b/drivers/irqchip/irq-gic-v3.c
> > index 6e1fa5b247fc..e73fea8a0e27 100644
> > --- a/drivers/irqchip/irq-gic-v3.c
> > +++ b/drivers/irqchip/irq-gic-v3.c
> > @@ -80,6 +80,8 @@ static DEFINE_STATIC_KEY_FALSE(gic_nvidia_t241_erratum);
> >
> > static DEFINE_STATIC_KEY_FALSE(gic_arm64_2941627_erratum);
> >
> > +static DEFINE_STATIC_KEY_FALSE(gic_fujitsu_030003_erratum);
> > +
> > static struct gic_chip_data gic_data __read_mostly;
> > static DEFINE_STATIC_KEY_TRUE(supports_deactivate_key);
> >
> > @@ -238,6 +240,9 @@ static DEFINE_PER_CPU(bool, has_rss);
> > #define gic_data_rdist() (this_cpu_ptr(gic_data.rdists.rdist))
> > #define gic_data_rdist_rd_base() (gic_data_rdist()->rd_base)
> > #define gic_data_rdist_sgi_base() (gic_data_rdist_rd_base() + SZ_64K)
> > +#define gic_data_rdist_cpu(cpu) (per_cpu_ptr(gic_data.rdists.rdist, cpu))
> > +#define gic_data_rdist_rd_base_cpu(cpu) (gic_data_rdist_cpu(cpu)->rd_base)
> > +#define gic_data_rdist_sgi_base_cpu(cpu) (gic_data_rdist_rd_base_cpu(cpu) + SZ_64K)
> >
> > /* Our default, arbitrary priority value. Linux only uses one anyway. */
> > #define DEFAULT_PMR_VALUE 0xf0
> > @@ -1373,6 +1378,13 @@ static void gic_send_sgi(u64 cluster_id, u16 tlist, unsigned int irq)
> > gic_write_sgi1r(val);
> > }
> >
> > +static void gic_send_sgi_via_rdist(int cpu, unsigned int irq)
> > +{
> > + void __iomem *base = gic_data_rdist_sgi_base_cpu(cpu);
> > +
> > + writel_relaxed(BIT(irq), base + GICR_ISPENDR0);
> > +}
> > +
>
> I don't think there is any need for a helper, given that there is a
> single caller.
Ack.
>
> > static void gic_ipi_send_mask(struct irq_data *d, const struct cpumask *mask)
> > {
> > int cpu;
> > @@ -1386,6 +1398,22 @@ static void gic_ipi_send_mask(struct irq_data *d, const struct cpumask *mask)
> > */
> > dsb(ishst);
> >
> > + if (static_branch_unlikely(&gic_fujitsu_030003_erratum)) {
> > + /*
> > + * The affinity topology of affected systems has Aff0 == 0 for
> > + * every PE; PEs are distinguished by higher affinity levels.
> > + * ICC_SGI1R_EL1 therefore targets only one PE per write, so
> > + * using GICR_ISPENDR0 does not increase the number of writes
> > + * required to send an SGI to multiple PEs.
>
> This is not about the number of writes, but the cost of individual
> writes. A sysreg access is almost free (at least it is on decent
> implementations), while an MMIO access is probably one of the worst
> offenders. Therefore implying that there is no extra overhead is
> likely to be misleading. The commit message has the same problem.
>
> In any case, I don't think you need to justify anything here, as the
> choice between costly SGIs and a dead CPU is pretty moot.
Ack. I'll remove the comment and the corresponding explanation from the
commit message.
For context, on this implementation, we do not expect a significant
latency difference between the MMIO and system register accesses, as
both ultimately reach the same hardware block.
>
> > + */
> > + for_each_cpu(cpu, mask)
> > + gic_send_sgi_via_rdist(cpu, d->hwirq);
> > +
> > + /* Force the above writes to GICR_ISPENDR0 to be executed */
> > + dsb(st);
>
> This doesn't force things to be executed. This is about completion of
> the access, and with an nGnRE mapping, it doesn't enforce that the
> stores actually reach the RDs, only an arbitrary point in the memory
> subsystem. The only way to guarantee this is to perform a read-back.
Understood.
I hadn't considered this when writing v1, but on further reflection, I
don't think gic_ipi_send_mask() needs to wait for the target CPUs to
handle the IPIs. Given that the GICv2 driver does not wait for MMIO
write completion either, I don't think we need to ensure completion of
these writes before returning.
If my understanding is correct, I'll remove the dsb(st) in v2.
>
> If this is relying on some additional properties that are
> implementation specific, then this requires to be documented (for
> example, if the implementation treats nGnRE as nGnRnE).
>
> > + return;
> > + }
> > +
> > for_each_cpu(cpu, mask) {
> > u64 cluster_id = MPIDR_TO_SGI_CLUSTER_ID(gic_cpu_to_affinity(cpu));
> > u16 tlist;
> > @@ -1871,6 +1899,20 @@ static bool gic_enable_quirk_rk3399(void *data)
> > return false;
> > }
> >
> > +#define SMCCC_SOC_ID_FUJITSU_MONAKA 0x00040003
> > +
> > +static bool gic_enable_quirk_fujitsu_030003(void *data)
> > +{
> > + s32 soc_id = arm_smccc_get_soc_id_version();
> > +
> > + /* Check JEP106 code for FUJITSU-MONAKA chip (0004:0003) */
> > + if (soc_id != SMCCC_SOC_ID_FUJITSU_MONAKA)
> > + return false;
>
> Why is this keyed on some firmware interface, while it is the CPU
> interface that is at fault? I'd expect that looking at the MIDR would
> be more reliable.
Ack.
>
> > +
> > + static_branch_enable(&gic_fujitsu_030003_erratum);
> > + return true;
> > +}
> > +
> > static bool rd_set_non_coherent(void *data)
> > {
> > struct gic_chip_data *d = data;
> > @@ -1951,6 +1993,12 @@ static const struct gic_quirk gic_quirks[] = {
> > .mask = 0xff000fff,
> > .init = gic_enable_quirk_rk3399,
> > },
> > + {
> > + .desc = "GICv3: Fujitsu erratum 030003",
> > + .iidr = 0x0403043b,
> > + .mask = 0xffffffff,
> > + .init = gic_enable_quirk_fujitsu_030003,
>
> Same problem. This is looking that the distributor instead of the CPU.
> It's OK to use it as a proxy for further filtering, but the final
> decision should probably rest on the MIDR.
Ack. I'll use the MIDR instead of the SoC ID.
Thanks,
Kohei
>
> Thanks,
>
> M.
>
> --
> Without deviation from the norm, progress is not possible.
next prev parent reply other threads:[~2026-10-06 9:08 UTC|newest]
Thread overview: 20+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-02 10:26 [PATCH 0/4] arm64: Add workarounds for FUJITSU-MONAKA CPU Tomohiro Misono
2026-10-02 10:26 ` [PATCH 1/4] arm64: cputype: Add FUJITSU MONAKA definition Tomohiro Misono
2026-10-02 10:31 ` sashiko-bot
2026-10-02 10:26 ` [PATCH 2/4] arm64: tlb: Add tlbi workaround for FUJITSU-MONAKA Erratum E#030001 Tomohiro Misono
2026-10-02 10:41 ` sashiko-bot
2026-10-02 12:43 ` Will Deacon
2026-10-06 9:48 ` Tomohiro Misono (Fujitsu)
2026-10-02 13:01 ` Mark Rutland
2026-10-06 9:45 ` Tomohiro Misono (Fujitsu)
2026-10-02 10:26 ` [PATCH 3/4] perf: arm_pmu: Add workaround for FUJITSU-MONAKA Erratum E#030002 Tomohiro Misono
2026-10-02 10:35 ` sashiko-bot
2026-10-02 12:44 ` Will Deacon
2026-10-06 9:52 ` Tomohiro Misono (Fujitsu)
2026-10-02 10:26 ` [PATCH 4/4] irqchip/gicv3: Add workaround for FUJITSU-MONAKA erratum E#030003 Tomohiro Misono
2026-10-02 10:41 ` sashiko-bot
2026-10-07 11:37 ` Kohei Enju
2026-10-02 12:26 ` Marc Zyngier
2026-10-06 9:08 ` Kohei Enju [this message]
2026-10-06 9:42 ` Marc Zyngier
2026-10-07 8:17 ` Kohei Enju
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=asS5r0WNCQpMIo4p@FCCLS0092175.localdomain \
--to=enju.kohei@fujitsu.com \
--cc=catalin.marinas@arm.com \
--cc=corbet@lwn.net \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-perf-users@vger.kernel.org \
--cc=mark.rutland@arm.com \
--cc=maz@kernel.org \
--cc=misono.tomohiro@fujitsu.com \
--cc=radu@rendec.net \
--cc=rdunlap@infradead.org \
--cc=skhan@linuxfoundation.org \
--cc=tglx@kernel.org \
--cc=will@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox