* [PATCH Dovetail 1/2] genirq: irq_pipeline: Remove irq_is_oob()
2026-09-09 14:32 [PATCH Dovetail 0/2] Obsolete marking tick IRQs with IRQF_TIMER Florian Bezdeka
@ 2026-09-09 14:32 ` Florian Bezdeka
2026-09-09 14:32 ` [PATCH Dovetail 2/2] irq_pipeline: Introduce IRQ_TICK to obsolete IRQF_TIMER on tick IRQs Florian Bezdeka
` (2 subsequent siblings)
3 siblings, 0 replies; 13+ messages in thread
From: Florian Bezdeka @ 2026-09-09 14:32 UTC (permalink / raw)
To: Xenomai; +Cc: Philippe Gerum, Gerte Hoogewerf, Andrew MacPherson,
Florian Bezdeka
There are no users, so let's remove the dead code.
No functional change.
Signed-off-by: Florian Bezdeka <florian.bezdeka@siemens.com>
---
include/linux/irqdesc.h | 5 -----
1 file changed, 5 deletions(-)
diff --git a/include/linux/irqdesc.h b/include/linux/irqdesc.h
index 260ffd288bd8110b810540e2998966083cb9bc25..b8347550f87f75e1d3f590502fbcc6a6612508e3 100644
--- a/include/linux/irqdesc.h
+++ b/include/linux/irqdesc.h
@@ -269,11 +269,6 @@ static inline bool irq_is_percpu_devid(unsigned int irq)
return irq_check_status_bit(irq, IRQ_PER_CPU_DEVID);
}
-static inline int irq_is_oob(unsigned int irq)
-{
- return irq_check_status_bit(irq, IRQ_OOB);
-}
-
void __irq_set_lockdep_class(unsigned int irq, struct lock_class_key *lock_class,
struct lock_class_key *request_class);
static inline void
--
2.55.0
^ permalink raw reply related [flat|nested] 13+ messages in thread* [PATCH Dovetail 2/2] irq_pipeline: Introduce IRQ_TICK to obsolete IRQF_TIMER on tick IRQs
2026-09-09 14:32 [PATCH Dovetail 0/2] Obsolete marking tick IRQs with IRQF_TIMER Florian Bezdeka
2026-09-09 14:32 ` [PATCH Dovetail 1/2] genirq: irq_pipeline: Remove irq_is_oob() Florian Bezdeka
@ 2026-09-09 14:32 ` Florian Bezdeka
2026-09-09 15:43 ` [PATCH Dovetail 0/2] Obsolete marking tick IRQs with IRQF_TIMER Gerte Hoogewerf
2026-09-10 12:49 ` Andrew MacPherson
3 siblings, 0 replies; 13+ messages in thread
From: Florian Bezdeka @ 2026-09-09 14:32 UTC (permalink / raw)
To: Xenomai; +Cc: Philippe Gerum, Gerte Hoogewerf, Andrew MacPherson,
Florian Bezdeka
Currently, all clock tick IRQs need to be marked with the IRQF_TIMER
flag, so that the registers could be saved when handling a pipeline
entry for timer IRQs / ticks.
It turned out that it is quite easy to miss that flag when enabling
the drivers for the pipelining case. In addition, IRQF_TIMER also
includes setting IRQF_NO_SUSPEND and IRQF_NO_THREAD, which is
misleading.
We hook into the clock event device registration now and check for
the CLOCK_EVT_FEAT_PIPELINE flag on the event device. If this flag
is set, we mark the underlying IRQ as IRQ_TICK.
The IRQ_TICK flag is only used for that pipeline internal purpose.
Suggested-by: Philippe Gerum <rpm@xenomai.org>
Signed-off-by: Florian Bezdeka <florian.bezdeka@siemens.com>
---
include/linux/irq.h | 11 +++++++++--
kernel/irq/debug.h | 1 +
kernel/irq/manage.c | 11 +++++++++++
kernel/irq/pipeline.c | 3 ++-
kernel/irq/settings.h | 17 +++++++++++++++++
kernel/time/clockevents.c | 3 +++
kernel/time/tick-common.c | 5 +++++
7 files changed, 48 insertions(+), 3 deletions(-)
diff --git a/include/linux/irq.h b/include/linux/irq.h
index b13e4e90ab18f027e43b4f892ca37ef482a6cf0a..5f7a2c78b3ca7a7a23894f68e16503c785eb5686 100644
--- a/include/linux/irq.h
+++ b/include/linux/irq.h
@@ -81,6 +81,7 @@ enum irqchip_irq_state;
* when pipelining is enabled (CONFIG_IRQ_PIPELINE),
* regardless of the (virtualized) interrupt state
* maintained by local_irq_save/disable().
+ * IRQ_TICK - Interrupt is a timer tick source.
*/
enum {
IRQ_TYPE_NONE = 0x00000000,
@@ -109,14 +110,15 @@ enum {
IRQ_HIDDEN = (1 << 20),
IRQ_NO_DEBUG = (1 << 21),
IRQ_OOB = (1 << 22),
- IRQ_RESERVED = (1 << 23),
+ IRQ_TICK = (1 << 23),
+ IRQ_RESERVED = (1 << 24),
};
#define IRQF_MODIFY_MASK \
(IRQ_TYPE_SENSE_MASK | IRQ_NOPROBE | IRQ_NOREQUEST | \
IRQ_NOAUTOEN | IRQ_LEVEL | IRQ_NO_BALANCING | \
IRQ_PER_CPU | IRQ_NESTED_THREAD | IRQ_NOTHREAD | IRQ_PER_CPU_DEVID | \
- IRQ_IS_POLLED | IRQ_DISABLE_UNLAZY | IRQ_HIDDEN | IRQ_OOB)
+ IRQ_IS_POLLED | IRQ_DISABLE_UNLAZY | IRQ_HIDDEN | IRQ_OOB | IRQ_TICK)
#define IRQ_NO_BALANCING_MASK (IRQ_PER_CPU | IRQ_NO_BALANCING)
@@ -1258,6 +1260,7 @@ static inline struct irq_chip_type *irq_data_get_chip_type(struct irq_data *d)
#ifdef CONFIG_IRQ_PIPELINE
int irq_switch_oob(unsigned int irq, bool on);
+void irq_switch_tick(unsigned int irq, bool on);
void irq_clear_deferral(struct irq_desc *desc);
void irq_clear_forward(struct irq_desc *desc);
#else
@@ -1266,6 +1269,10 @@ static inline int irq_switch_oob(unsigned int irq, bool on)
return 0;
}
+static inline void irq_switch_tick(unsigned int irq, bool on)
+{
+}
+
static inline void irq_clear_deferral(struct irq_desc *desc) { }
static inline void irq_clear_forward(struct irq_desc *desc) { }
#endif /* !CONFIG_IRQ_PIPELINE */
diff --git a/kernel/irq/debug.h b/kernel/irq/debug.h
index 4eafd04a629628553c11d639b4958d5007bbc012..9813c8ec66a447e4acde5ac2ca2fa1f8ab70021b 100644
--- a/kernel/irq/debug.h
+++ b/kernel/irq/debug.h
@@ -34,6 +34,7 @@ static inline void print_irq_desc(unsigned int irq, struct irq_desc *desc)
___P(IRQ_NOTHREAD);
___P(IRQ_NOAUTOEN);
___P(IRQ_OOB);
+ ___P(IRQ_TICK);
___PS(IRQS_AUTODETECT);
___PS(IRQS_REPLAY);
diff --git a/kernel/irq/manage.c b/kernel/irq/manage.c
index 1c1920f5fbc083ec299433dcddd4fae53e0d6ce8..a10d1081bea8c8887c61cb3ef5a87b274e1c9e08 100644
--- a/kernel/irq/manage.c
+++ b/kernel/irq/manage.c
@@ -949,6 +949,17 @@ int irq_switch_oob(unsigned int irq, bool on)
}
EXPORT_SYMBOL_GPL(irq_switch_oob);
+void irq_switch_tick(unsigned int irq, bool on)
+{
+ scoped_irqdesc_get_and_lock(irq, 0) {
+ if (on)
+ irq_settings_set_tick(scoped_irqdesc);
+ else
+ irq_settings_clr_tick(scoped_irqdesc);
+ }
+}
+EXPORT_SYMBOL_GPL(irq_switch_tick);
+
#endif /* CONFIG_IRQ_PIPELINE */
/*
diff --git a/kernel/irq/pipeline.c b/kernel/irq/pipeline.c
index 85ec0cbf5fb1e1fa38dee98d1f79a3ee4ae18774..88911489aeb632a00aab9784407f09ff835f7861 100644
--- a/kernel/irq/pipeline.c
+++ b/kernel/irq/pipeline.c
@@ -1064,8 +1064,9 @@ void copy_timer_regs(struct irq_desc *desc, struct pt_regs *regs)
{
struct irq_pipeline_data *p;
- if (desc->action == NULL || !(desc->action->flags & __IRQF_TIMER))
+ if (!irq_settings_is_tick(desc))
return;
+
/*
* Given our deferred dispatching model for regular IRQs, we
* record the preempted context registers only for the latest
diff --git a/kernel/irq/settings.h b/kernel/irq/settings.h
index 27a37d992f237f6429f58b0b51eb9db39a62e8f7..c4b48d22d271a24e49951c4289784e44090f02dd 100644
--- a/kernel/irq/settings.h
+++ b/kernel/irq/settings.h
@@ -19,6 +19,7 @@ enum {
_IRQ_HIDDEN = IRQ_HIDDEN,
_IRQ_NO_DEBUG = IRQ_NO_DEBUG,
_IRQ_OOB = IRQ_OOB,
+ _IRQ_TICK = IRQ_TICK,
_IRQ_PROC_VALID = IRQ_RESERVED,
_IRQF_MODIFY_MASK = IRQF_MODIFY_MASK,
};
@@ -37,6 +38,7 @@ enum {
#define IRQ_HIDDEN GOT_YOU_MORON
#define IRQ_NO_DEBUG GOT_YOU_MORON
#define IRQ_OOB GOT_YOU_MORON
+#define IRQ_TICK GOT_YOU_MORON
#define IRQ_RESERVED GOT_YOU_MORON
#undef IRQF_MODIFY_MASK
#define IRQF_MODIFY_MASK GOT_YOU_MORON
@@ -210,3 +212,18 @@ static inline void irq_settings_set_oob(struct irq_desc *desc)
{
desc->status_use_accessors |= _IRQ_OOB;
}
+
+static inline bool irq_settings_is_tick(struct irq_desc *desc)
+{
+ return desc->status_use_accessors & _IRQ_TICK;
+}
+
+static inline void irq_settings_clr_tick(struct irq_desc *desc)
+{
+ desc->status_use_accessors &= ~_IRQ_TICK;
+}
+
+static inline void irq_settings_set_tick(struct irq_desc *desc)
+{
+ desc->status_use_accessors |= _IRQ_TICK;
+}
diff --git a/kernel/time/clockevents.c b/kernel/time/clockevents.c
index 0ed4122d40986708bd57f32ebf825aef60756f7d..d8d2dd43baaa0053a73bb18764ece562e8b49024 100644
--- a/kernel/time/clockevents.c
+++ b/kernel/time/clockevents.c
@@ -12,6 +12,7 @@
#include <linux/init.h>
#include <linux/module.h>
#include <linux/smp.h>
+#include <linux/irq.h>
#include <linux/device.h>
#include "tick-internal.h"
@@ -177,6 +178,8 @@ void clockevents_shutdown(struct clock_event_device *dev)
clockevents_switch_state(dev, CLOCK_EVT_STATE_SHUTDOWN);
dev->next_event = KTIME_MAX;
dev->next_event_forced = 0;
+ if (dev->features & CLOCK_EVT_FEAT_PIPELINE)
+ irq_switch_tick(dev->irq, false);
}
/**
diff --git a/kernel/time/tick-common.c b/kernel/time/tick-common.c
index 90fae659e4ea68578a8ab5a739eccc1184ed5ff9..97e3ff40405dfbe670fe45aebdff015163cb6f31 100644
--- a/kernel/time/tick-common.c
+++ b/kernel/time/tick-common.c
@@ -12,7 +12,9 @@
#include <linux/err.h>
#include <linux/hrtimer.h>
#include <linux/interrupt.h>
+#include <linux/irq.h>
#include <linux/nmi.h>
+#include <linux/irq.h>
#include <linux/percpu.h>
#include <linux/profile.h>
#include <linux/sched.h>
@@ -244,6 +246,9 @@ static void tick_setup_device(struct tick_device *td,
if (!cpumask_equal(newdev->cpumask, cpumask))
irq_set_affinity(newdev->irq, cpumask);
+ if (newdev->features & CLOCK_EVT_FEAT_PIPELINE)
+ irq_switch_tick(newdev->irq, true);
+
/*
* When global broadcasting is active, check if the current
* device is registered as a placeholder for broadcast mode.
--
2.55.0
^ permalink raw reply related [flat|nested] 13+ messages in thread* Re: [PATCH Dovetail 0/2] Obsolete marking tick IRQs with IRQF_TIMER
2026-09-09 14:32 [PATCH Dovetail 0/2] Obsolete marking tick IRQs with IRQF_TIMER Florian Bezdeka
2026-09-09 14:32 ` [PATCH Dovetail 1/2] genirq: irq_pipeline: Remove irq_is_oob() Florian Bezdeka
2026-09-09 14:32 ` [PATCH Dovetail 2/2] irq_pipeline: Introduce IRQ_TICK to obsolete IRQF_TIMER on tick IRQs Florian Bezdeka
@ 2026-09-09 15:43 ` Gerte Hoogewerf
2026-09-10 5:53 ` Gerte Hoogewerf
2026-09-10 12:49 ` Andrew MacPherson
3 siblings, 1 reply; 13+ messages in thread
From: Gerte Hoogewerf @ 2026-09-09 15:43 UTC (permalink / raw)
To: Florian Bezdeka; +Cc: Xenomai, Philippe Gerum, Andrew MacPherson
[-- Attachment #1: Type: text/plain, Size: 999 bytes --]
Hi Florian,
> The following is the try to fix the problems reported by Gerte and Andrew.
This seems to work for me. I will test this further in the coming days.
Caveat: our US+ board is not yet on version 7.2. I had to back-port it to 6.18.
Attached is the code I tested with (base-commit: f01f7033b0 -- tag:
v6.18.29-dovetail3-rebase).
I'm not asking you to do anything with this. I'm just illustrating
that I'm currently testing a slightly different variant (there were
minor merge conflicts).
Thank you so much.
Cheers,
--
This
email and any attachment(s) it may contain is confidential and is
intended
solely for the use of the individual(s) to whom it is addressed.
If you are not
the intended recipient of this email, you must not take
action based on the
contents, nor distribute, nor expose any part of the
content(s) to entities or
person(s) beyond the original distribution list.
Please contact the sender and
delete the email if you have received it in
error. Thank you.
[-- Attachment #2: 0001-irq_pipeline-Introduce-IRQ_TICK-to-obsolete-IRQF_TIM.patch --]
[-- Type: text/plain, Size: 7310 bytes --]
From 7a4e4c1f781192588b1e6b45a7215fc5059ebe17 Mon Sep 17 00:00:00 2001
From: Gerte Hoogewerf <ghoogewerf@lmi3d.com>
Date: Wed, 9 Sep 2026 17:04:59 +0200
Subject: [PATCH] irq_pipeline: Introduce IRQ_TICK to obsolete IRQF_TIMER on
tick IRQs
Currently, all clock tick IRQs need to be marked with the IRQF_TIMER
flag, so that the registers could be saved when handling a pipeline
entry for timer IRQs / ticks.
It turned out that it is quite easy to miss that flag when enabling
the drivers for the pipelining case. In addition, IRQF_TIMER also
includes setting IRQF_NO_SUSPEND and IRQF_NO_THREAD, which is
misleading.
We hook into the clock event device registration now and check for
the CLOCK_EVT_FEAT_PIPELINE flag on the event device. If this flag
is set, we mark the underlying IRQ as IRQ_TICK.
The IRQ_TICK flag is only used for that pipeline internal purpose.
---
include/linux/irq.h | 9 ++++++++-
kernel/irq/debug.h | 1 +
kernel/irq/manage.c | 11 +++++++++++
kernel/irq/pipeline.c | 10 ++++------
kernel/irq/settings.h | 17 +++++++++++++++++
kernel/time/clockevents.c | 3 +++
kernel/time/tick-common.c | 4 ++++
7 files changed, 48 insertions(+), 7 deletions(-)
diff --git a/include/linux/irq.h b/include/linux/irq.h
index 450fbca73ca2..956a7fac57f2 100644
--- a/include/linux/irq.h
+++ b/include/linux/irq.h
@@ -78,6 +78,7 @@ enum irqchip_irq_state;
* regardless of the (virtualized) interrupt state
* maintained by local_irq_save/disable().
* IRQ_CHAINED - Interrupt is chained.
+ * IRQ_TICK - Interrupt is a timer tick source.
*/
enum {
IRQ_TYPE_NONE = 0x00000000,
@@ -107,13 +108,14 @@ enum {
IRQ_NO_DEBUG = (1 << 21),
IRQ_OOB = (1 << 22),
IRQ_CHAINED = (1 << 23),
+ IRQ_TICK = (1 << 24),
};
#define IRQF_MODIFY_MASK \
(IRQ_TYPE_SENSE_MASK | IRQ_NOPROBE | IRQ_NOREQUEST | \
IRQ_NOAUTOEN | IRQ_LEVEL | IRQ_NO_BALANCING | \
IRQ_PER_CPU | IRQ_NESTED_THREAD | IRQ_NOTHREAD | IRQ_PER_CPU_DEVID | \
- IRQ_IS_POLLED | IRQ_DISABLE_UNLAZY | IRQ_HIDDEN | IRQ_OOB)
+ IRQ_IS_POLLED | IRQ_DISABLE_UNLAZY | IRQ_HIDDEN | IRQ_OOB | IRQ_TICK)
#define IRQ_NO_BALANCING_MASK (IRQ_PER_CPU | IRQ_NO_BALANCING)
@@ -1253,6 +1255,7 @@ static inline struct irq_chip_type *irq_data_get_chip_type(struct irq_data *d)
#ifdef CONFIG_IRQ_PIPELINE
int irq_switch_oob(unsigned int irq, bool on);
+void irq_switch_tick(unsigned int irq, bool on);
void irq_clear_deferral(struct irq_desc *desc);
void irq_clear_forward(struct irq_desc *desc);
#else
@@ -1261,6 +1264,10 @@ static inline int irq_switch_oob(unsigned int irq, bool on)
return 0;
}
+static inline void irq_switch_tick(unsigned int irq, bool on)
+{
+}
+
static inline void irq_clear_deferral(struct irq_desc *desc) { }
static inline void irq_clear_forward(struct irq_desc *desc) { }
#endif /* !CONFIG_IRQ_PIPELINE */
diff --git a/kernel/irq/debug.h b/kernel/irq/debug.h
index 40f726845748..9e1844570296 100644
--- a/kernel/irq/debug.h
+++ b/kernel/irq/debug.h
@@ -35,6 +35,7 @@ static inline void print_irq_desc(unsigned int irq, struct irq_desc *desc)
___P(IRQ_NOAUTOEN);
___P(IRQ_OOB);
___P(IRQ_CHAINED);
+ ___P(IRQ_TICK);
___PS(IRQS_AUTODETECT);
___PS(IRQS_REPLAY);
diff --git a/kernel/irq/manage.c b/kernel/irq/manage.c
index b7899ac21634..0f8740231dea 100644
--- a/kernel/irq/manage.c
+++ b/kernel/irq/manage.c
@@ -938,6 +938,17 @@ int irq_switch_oob(unsigned int irq, bool on)
}
EXPORT_SYMBOL_GPL(irq_switch_oob);
+void irq_switch_tick(unsigned int irq, bool on)
+{
+ scoped_irqdesc_get_and_lock(irq, 0) {
+ if (on)
+ irq_settings_set_tick(scoped_irqdesc);
+ else
+ irq_settings_clr_tick(scoped_irqdesc);
+ }
+}
+EXPORT_SYMBOL_GPL(irq_switch_tick);
+
#endif /* CONFIG_IRQ_PIPELINE */
/*
diff --git a/kernel/irq/pipeline.c b/kernel/irq/pipeline.c
index d2bbe2263e01..c64ebe62bd7c 100644
--- a/kernel/irq/pipeline.c
+++ b/kernel/irq/pipeline.c
@@ -1063,10 +1063,6 @@ bool handle_oob_irq(struct irq_desc *desc)
static inline
void copy_timer_regs(struct irq_desc *desc, struct pt_regs *regs)
{
- struct irq_pipeline_data *p;
-
- if (desc->action == NULL || !(desc->action->flags & __IRQF_TIMER))
- return;
/*
* Given our deferred dispatching model for regular IRQs, we
* record the preempted context registers only for the latest
@@ -1074,8 +1070,10 @@ void copy_timer_regs(struct irq_desc *desc, struct pt_regs *regs)
* CPU times properly. It is assumed that no other interrupt
* handler cares for such information.
*/
- p = raw_cpu_ptr(&irq_pipeline);
- arch_save_timer_regs(&p->tick_regs, regs);
+ if (irq_settings_is_tick(desc)) {
+ struct irq_pipeline_data *p = raw_cpu_ptr(&irq_pipeline);
+ arch_save_timer_regs(&p->tick_regs, regs);
+ }
}
static __always_inline
diff --git a/kernel/irq/settings.h b/kernel/irq/settings.h
index 1f5c49545a88..70fde1fa0982 100644
--- a/kernel/irq/settings.h
+++ b/kernel/irq/settings.h
@@ -20,6 +20,7 @@ enum {
_IRQ_NO_DEBUG = IRQ_NO_DEBUG,
_IRQ_OOB = IRQ_OOB,
_IRQ_CHAINED = IRQ_CHAINED,
+ _IRQ_TICK = IRQ_TICK,
_IRQF_MODIFY_MASK = IRQF_MODIFY_MASK,
};
@@ -38,6 +39,7 @@ enum {
#define IRQ_NO_DEBUG GOT_YOU_MORON
#define IRQ_OOB GOT_YOU_MORON
#define IRQ_CHAINED GOT_YOU_MORON
+#define IRQ_TICK GOT_YOU_MORON
#undef IRQF_MODIFY_MASK
#define IRQF_MODIFY_MASK GOT_YOU_MORON
@@ -214,3 +216,18 @@ static inline void irq_settings_clr_chained(struct irq_desc *desc)
{
desc->status_use_accessors &= ~_IRQ_CHAINED;
}
+
+static inline bool irq_settings_is_tick(struct irq_desc *desc)
+{
+ return desc->status_use_accessors & _IRQ_TICK;
+}
+
+static inline void irq_settings_clr_tick(struct irq_desc *desc)
+{
+ desc->status_use_accessors &= ~_IRQ_TICK;
+}
+
+static inline void irq_settings_set_tick(struct irq_desc *desc)
+{
+ desc->status_use_accessors |= _IRQ_TICK;
+}
diff --git a/kernel/time/clockevents.c b/kernel/time/clockevents.c
index dc21840a6080..3c9ece0ada86 100644
--- a/kernel/time/clockevents.c
+++ b/kernel/time/clockevents.c
@@ -12,6 +12,7 @@
#include <linux/init.h>
#include <linux/module.h>
#include <linux/smp.h>
+#include <linux/irq.h>
#include <linux/device.h>
#include "tick-internal.h"
@@ -173,6 +174,8 @@ void clockevents_shutdown(struct clock_event_device *dev)
{
clockevents_switch_state(dev, CLOCK_EVT_STATE_SHUTDOWN);
dev->next_event = KTIME_MAX;
+ if (dev->features & CLOCK_EVT_FEAT_PIPELINE)
+ irq_switch_tick(dev->irq, false);
}
/**
diff --git a/kernel/time/tick-common.c b/kernel/time/tick-common.c
index f53324c89c9f..a0db59614133 100644
--- a/kernel/time/tick-common.c
+++ b/kernel/time/tick-common.c
@@ -13,6 +13,7 @@
#include <linux/hrtimer.h>
#include <linux/interrupt.h>
#include <linux/nmi.h>
+#include <linux/irq.h>
#include <linux/percpu.h>
#include <linux/profile.h>
#include <linux/sched.h>
@@ -243,6 +244,9 @@ static void tick_setup_device(struct tick_device *td,
if (!cpumask_equal(newdev->cpumask, cpumask))
irq_set_affinity(newdev->irq, cpumask);
+ if (newdev->features & CLOCK_EVT_FEAT_PIPELINE)
+ irq_switch_tick(newdev->irq, true);
+
/*
* When global broadcasting is active, check if the current
* device is registered as a placeholder for broadcast mode.
--
2.55.0
^ permalink raw reply related [flat|nested] 13+ messages in thread* Re: [PATCH Dovetail 0/2] Obsolete marking tick IRQs with IRQF_TIMER
2026-09-09 15:43 ` [PATCH Dovetail 0/2] Obsolete marking tick IRQs with IRQF_TIMER Gerte Hoogewerf
@ 2026-09-10 5:53 ` Gerte Hoogewerf
0 siblings, 0 replies; 13+ messages in thread
From: Gerte Hoogewerf @ 2026-09-10 5:53 UTC (permalink / raw)
To: Florian Bezdeka; +Cc: Xenomai, Philippe Gerum, Andrew MacPherson
Hi Florian,
I ported our board support (US+ custom board) to Linux 7.2 (base
commit 58e8e0f74) and started testing your patch. The base commit
reproduced the issue. The patch solved the issue.
Awesome.
Thanks,
--
This
email and any attachment(s) it may contain is confidential and is
intended
solely for the use of the individual(s) to whom it is addressed.
If you are not
the intended recipient of this email, you must not take
action based on the
contents, nor distribute, nor expose any part of the
content(s) to entities or
person(s) beyond the original distribution list.
Please contact the sender and
delete the email if you have received it in
error. Thank you.
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [PATCH Dovetail 0/2] Obsolete marking tick IRQs with IRQF_TIMER
2026-09-09 14:32 [PATCH Dovetail 0/2] Obsolete marking tick IRQs with IRQF_TIMER Florian Bezdeka
` (2 preceding siblings ...)
2026-09-09 15:43 ` [PATCH Dovetail 0/2] Obsolete marking tick IRQs with IRQF_TIMER Gerte Hoogewerf
@ 2026-09-10 12:49 ` Andrew MacPherson
2026-09-10 12:57 ` Florian Bezdeka
3 siblings, 1 reply; 13+ messages in thread
From: Andrew MacPherson @ 2026-09-10 12:49 UTC (permalink / raw)
To: Florian Bezdeka; +Cc: Xenomai, Philippe Gerum, Gerte Hoogewerf
Hi Florian,
On Wed, 9 Sept 2026 at 16:33, Florian Bezdeka
<florian.bezdeka@siemens.com> wrote:
>
> Hi all,
>
> The following is the try to fix the problems reported by Gerte and Andrew.
> We forgot to add the IRQF_TIMER to some of the arm/arm64 clocksource
> drivers.
>
> While the patch provided by Andrew [3] is correct, I tried to find a way
> to get rid of this Dovetail specific IRQF_TIMER marking, as it could
> easily be forgotten.
>
> Review comments and intensive testing is highly appreciated.
> Tests were done on qemu for arm, arm64 and x86.
>
> If accepted / merged we should clean up the dovetail IRQF_TIMER markers
> afterwards.
>
> Best regards,
> Florian
>
> [1] https://lore.kernel.org/xenomai/CAKndYJF00RPbfD7U8YCAn_vzxDFdtwz5m9O=JVqhH5fdXHxY9w@mail.gmail.com/T/
> [2] https://lore.kernel.org/xenomai/CACiVkw=Vs1KUgfZ=Mq-rXU=4gPFDtmuoNoMv=rzy_kvV7KgpKg@mail.gmail.com/
> [3] https://lore.kernel.org/xenomai/20260907083708.2566-1-andrew@elk.audio/T/#u
>
> To: Xenomai <xenomai@lists.linux.dev>
> Cc: Philippe Gerum <rpm@xenomai.org>
> Cc: Gerte Hoogewerf <ghoogewerf@lmi3d.com>
> Cc: Andrew MacPherson <andrew@elk.audio>
>
> Suggested-by: Philippe Gerum <rpm@xenomai.org>
> Signed-off-by: Florian Bezdeka <florian.bezdeka@siemens.com>
> ---
> Florian Bezdeka (2):
> genirq: irq_pipeline: Remove irq_is_oob()
> irq_pipeline: Introduce IRQ_TICK to obsolete IRQF_TIMER on tick IRQs
>
> include/linux/irq.h | 11 +++++++++--
> include/linux/irqdesc.h | 5 -----
> kernel/irq/debug.h | 1 +
> kernel/irq/manage.c | 11 +++++++++++
> kernel/irq/pipeline.c | 3 ++-
> kernel/irq/settings.h | 17 +++++++++++++++++
> kernel/time/clockevents.c | 3 +++
> kernel/time/tick-common.c | 5 +++++
> 8 files changed, 48 insertions(+), 8 deletions(-)
> ---
> base-commit: 58e8e0f74c2b20f6228b7d327ec2e955ffd5b491
> change-id: 20260907-wip-flo-v7-2-irq-cleanup-65fb73c34f9c
>
> Best regards,
> --
> Florian Bezdeka <florian.bezdeka@siemens.com>
>
Thanks for looking into this! I backported the patch to the older
kernel (6.6.48) we're using since it would be difficult to quickly try
on 7.2 and realized that my earlier testing didn't go far enough.
What I really want to run is "perf record -g --call-graph dwarf"
against a running process, however in testing I was just using "perf
top" as a quick proxy. With both your patch and the one I submitted
earlier "perf top" starts returning valid results (not all 0s as
before). However, both patches result in a kernel panic when trying to
use them to get a call graph.
A snapshot of the panic looks like this:
Unable to handle kernel NULL pointer dereference at virtual address
00000000 when read
Internal error: Oops: 5 [#1] PREEMPT SMP ARM
CPU: 0 PID: 762 Comm: XXXXXX Tainted: G O 6.6.48 #2
PC is at unwind_exec_insn+0x2dc/0x47c
LR is at unwind_frame+0x18c/0x424
unwind_exec_insn from unwind_frame+0x18c/0x424
unwind_frame from walk_stackframe+0x30/0x3c
walk_stackframe from perf_callchain_kernel+0x60/0x84
perf_callchain_kernel from get_perf_callchain+0xa8/0x220
get_perf_callchain from perf_callchain+0x84/0x98
perf_callchain from perf_prepare_sample+0x4ec/0x650
perf_prepare_sample from perf_event_output_forward+0x58/0xd0
perf_event_output_forward from __perf_event_overflow+0x98/0x25c
__perf_event_overflow from armv7pmu_handle_irq+0x11c/0x150
armv7pmu_handle_irq from armpmu_dispatch_irq+0x28/0x80
armpmu_dispatch_irq from __handle_irq_event_percpu+0x68/0x224
__handle_irq_event_percpu from handle_irq_event+0x58/0xf8
handle_irq_event from handle_fasteoi_irq+0x154/0x2f4
handle_fasteoi_irq from handle_irq_desc+0x20/0x30
handle_irq_desc from arch_do_IRQ_pipelined+0x58/0x98
arch_do_IRQ_pipelined from sync_current_irq_stage+0x90/0xc0
sync_current_irq_stage from handle_irq_pipelined_finish+0xa0/0x198
handle_irq_pipelined_finish from __irq_usr+0x60/0x80
Kernel panic - not syncing: Fatal exception in interrupt
As before this could just be due to the older kernel and maybe this is
fixed in the 7.2 branch. I also tried hacking around this issue with
the help of an LLM, basically to skip unwinding frames that would
trigger it, which lead to another panic:
kernel BUG at kernel/irq_work.c:245!
Internal error: Oops - BUG: 0 [#1] PREEMPT SMP ARM
CPU: 0 PID: 520 Comm: XXXXXX Tainted: G O 6.6.48 #4
PC is at irq_work_run_list+0x14/0x64
LR is at irq_work_run_list+0xc/0x64
irq_work_run_list from irq_work_run+0x28/0x3c
irq_work_run from armv7pmu_handle_irq+0x148/0x150
armv7pmu_handle_irq from armpmu_dispatch_irq+0x28/0x80
armpmu_dispatch_irq from __handle_irq_event_percpu+0x68/0x224
__handle_irq_event_percpu from handle_irq_event+0x58/0xf8
handle_irq_event from handle_fasteoi_irq+0x154/0x2f4
handle_fasteoi_irq from handle_irq_desc+0x20/0x30
handle_irq_desc from arch_do_IRQ_pipelined+0x58/0x90
arch_do_IRQ_pipelined from sync_current_irq_stage+0xac/0xe0
sync_current_irq_stage from handle_irq_pipelined_finish+0xa0/0x198
handle_irq_pipelined_finish from __irq_usr+0x60/0x80
Kernel panic - not syncing: Fatal exception in interrupt
Hacking also around this one I do get expected call graph results from
perf, ignoring some bogus sample frames caused by the hacks.
Let me know if I can provide any more detail here, I didn't want to
post patches that are obviously incorrect but if they're useful for
debugging in any way then I can.
Thanks again,
Andrew
^ permalink raw reply [flat|nested] 13+ messages in thread* Re: [PATCH Dovetail 0/2] Obsolete marking tick IRQs with IRQF_TIMER
2026-09-10 12:49 ` Andrew MacPherson
@ 2026-09-10 12:57 ` Florian Bezdeka
2026-09-10 14:38 ` Andrew MacPherson
0 siblings, 1 reply; 13+ messages in thread
From: Florian Bezdeka @ 2026-09-10 12:57 UTC (permalink / raw)
To: Andrew MacPherson; +Cc: Xenomai, Philippe Gerum, Gerte Hoogewerf
On Thu, 2026-09-10 at 14:49 +0200, Andrew MacPherson wrote:
> Hi Florian,
>
> On Wed, 9 Sept 2026 at 16:33, Florian Bezdeka
> <florian.bezdeka@siemens.com> wrote:
> >
> > Hi all,
> >
> > The following is the try to fix the problems reported by Gerte and Andrew.
> > We forgot to add the IRQF_TIMER to some of the arm/arm64 clocksource
> > drivers.
> >
> > While the patch provided by Andrew [3] is correct, I tried to find a way
> > to get rid of this Dovetail specific IRQF_TIMER marking, as it could
> > easily be forgotten.
> >
> > Review comments and intensive testing is highly appreciated.
> > Tests were done on qemu for arm, arm64 and x86.
> >
> > If accepted / merged we should clean up the dovetail IRQF_TIMER markers
> > afterwards.
> >
> > Best regards,
> > Florian
> >
> > [1] https://lore.kernel.org/xenomai/CAKndYJF00RPbfD7U8YCAn_vzxDFdtwz5m9O=JVqhH5fdXHxY9w@mail.gmail.com/T/
> > [2] https://lore.kernel.org/xenomai/CACiVkw=Vs1KUgfZ=Mq-rXU=4gPFDtmuoNoMv=rzy_kvV7KgpKg@mail.gmail.com/
> > [3] https://lore.kernel.org/xenomai/20260907083708.2566-1-andrew@elk.audio/T/#u
> >
> > To: Xenomai <xenomai@lists.linux.dev>
> > Cc: Philippe Gerum <rpm@xenomai.org>
> > Cc: Gerte Hoogewerf <ghoogewerf@lmi3d.com>
> > Cc: Andrew MacPherson <andrew@elk.audio>
> >
> > Suggested-by: Philippe Gerum <rpm@xenomai.org>
> > Signed-off-by: Florian Bezdeka <florian.bezdeka@siemens.com>
> > ---
> > Florian Bezdeka (2):
> > genirq: irq_pipeline: Remove irq_is_oob()
> > irq_pipeline: Introduce IRQ_TICK to obsolete IRQF_TIMER on tick IRQs
> >
> > include/linux/irq.h | 11 +++++++++--
> > include/linux/irqdesc.h | 5 -----
> > kernel/irq/debug.h | 1 +
> > kernel/irq/manage.c | 11 +++++++++++
> > kernel/irq/pipeline.c | 3 ++-
> > kernel/irq/settings.h | 17 +++++++++++++++++
> > kernel/time/clockevents.c | 3 +++
> > kernel/time/tick-common.c | 5 +++++
> > 8 files changed, 48 insertions(+), 8 deletions(-)
> > ---
> > base-commit: 58e8e0f74c2b20f6228b7d327ec2e955ffd5b491
> > change-id: 20260907-wip-flo-v7-2-irq-cleanup-65fb73c34f9c
> >
> > Best regards,
> > --
> > Florian Bezdeka <florian.bezdeka@siemens.com>
> >
>
> Thanks for looking into this! I backported the patch to the older
> kernel (6.6.48) we're using since it would be difficult to quickly try
> on 7.2 and realized that my earlier testing didn't go far enough.
>
> What I really want to run is "perf record -g --call-graph dwarf"
> against a running process, however in testing I was just using "perf
> top" as a quick proxy. With both your patch and the one I submitted
> earlier "perf top" starts returning valid results (not all 0s as
> before). However, both patches result in a kernel panic when trying to
> use them to get a call graph.
>
> A snapshot of the panic looks like this:
>
> Unable to handle kernel NULL pointer dereference at virtual address
> 00000000 when read
> Internal error: Oops: 5 [#1] PREEMPT SMP ARM
> CPU: 0 PID: 762 Comm: XXXXXX Tainted: G O 6.6.48 #2
> PC is at unwind_exec_insn+0x2dc/0x47c
> LR is at unwind_frame+0x18c/0x424
>
> unwind_exec_insn from unwind_frame+0x18c/0x424
> unwind_frame from walk_stackframe+0x30/0x3c
> walk_stackframe from perf_callchain_kernel+0x60/0x84
> perf_callchain_kernel from get_perf_callchain+0xa8/0x220
> get_perf_callchain from perf_callchain+0x84/0x98
> perf_callchain from perf_prepare_sample+0x4ec/0x650
> perf_prepare_sample from perf_event_output_forward+0x58/0xd0
> perf_event_output_forward from __perf_event_overflow+0x98/0x25c
> __perf_event_overflow from armv7pmu_handle_irq+0x11c/0x150
> armv7pmu_handle_irq from armpmu_dispatch_irq+0x28/0x80
> armpmu_dispatch_irq from __handle_irq_event_percpu+0x68/0x224
> __handle_irq_event_percpu from handle_irq_event+0x58/0xf8
> handle_irq_event from handle_fasteoi_irq+0x154/0x2f4
> handle_fasteoi_irq from handle_irq_desc+0x20/0x30
> handle_irq_desc from arch_do_IRQ_pipelined+0x58/0x98
> arch_do_IRQ_pipelined from sync_current_irq_stage+0x90/0xc0
> sync_current_irq_stage from handle_irq_pipelined_finish+0xa0/0x198
> handle_irq_pipelined_finish from __irq_usr+0x60/0x80
>
> Kernel panic - not syncing: Fatal exception in interrupt
>
> As before this could just be due to the older kernel and maybe this is
> fixed in the 7.2 branch. I also tried hacking around this issue with
> the help of an LLM, basically to skip unwinding frames that would
> trigger it, which lead to another panic:
>
> kernel BUG at kernel/irq_work.c:245!
> Internal error: Oops - BUG: 0 [#1] PREEMPT SMP ARM
> CPU: 0 PID: 520 Comm: XXXXXX Tainted: G O 6.6.48 #4
> PC is at irq_work_run_list+0x14/0x64
> LR is at irq_work_run_list+0xc/0x64
>
> irq_work_run_list from irq_work_run+0x28/0x3c
> irq_work_run from armv7pmu_handle_irq+0x148/0x150
> armv7pmu_handle_irq from armpmu_dispatch_irq+0x28/0x80
> armpmu_dispatch_irq from __handle_irq_event_percpu+0x68/0x224
> __handle_irq_event_percpu from handle_irq_event+0x58/0xf8
> handle_irq_event from handle_fasteoi_irq+0x154/0x2f4
> handle_fasteoi_irq from handle_irq_desc+0x20/0x30
> handle_irq_desc from arch_do_IRQ_pipelined+0x58/0x90
> arch_do_IRQ_pipelined from sync_current_irq_stage+0xac/0xe0
> sync_current_irq_stage from handle_irq_pipelined_finish+0xa0/0x198
> handle_irq_pipelined_finish from __irq_usr+0x60/0x80
>
> Kernel panic - not syncing: Fatal exception in interrupt
>
> Hacking also around this one I do get expected call graph results from
> perf, ignoring some bogus sample frames caused by the hacks.
>
> Let me know if I can provide any more detail here, I didn't want to
> post patches that are obviously incorrect but if they're useful for
> debugging in any way then I can.
>
Thanks for letting me know. Could you please try backporting
58e8e0f74c2b ("ARM: irq_pipeline: save registers used for walking tick
frames")
as well and try again.
That was fixed in 7.2 recently and it seems backporting did not start
yet. Your branch is EOL, so... do not expect it to land there.
HTH,
Florian
^ permalink raw reply [flat|nested] 13+ messages in thread* Re: [PATCH Dovetail 0/2] Obsolete marking tick IRQs with IRQF_TIMER
2026-09-10 12:57 ` Florian Bezdeka
@ 2026-09-10 14:38 ` Andrew MacPherson
2026-09-11 8:48 ` Florian Bezdeka
0 siblings, 1 reply; 13+ messages in thread
From: Andrew MacPherson @ 2026-09-10 14:38 UTC (permalink / raw)
To: Florian Bezdeka; +Cc: Xenomai, Philippe Gerum, Gerte Hoogewerf
On Thu, 10 Sept 2026 at 14:57, Florian Bezdeka
<florian.bezdeka@siemens.com> wrote:
>
> On Thu, 2026-09-10 at 14:49 +0200, Andrew MacPherson wrote:
> > Hi Florian,
> >
> > On Wed, 9 Sept 2026 at 16:33, Florian Bezdeka
> > <florian.bezdeka@siemens.com> wrote:
> > >
> > > Hi all,
> > >
> > > The following is the try to fix the problems reported by Gerte and Andrew.
> > > We forgot to add the IRQF_TIMER to some of the arm/arm64 clocksource
> > > drivers.
> > >
> > > While the patch provided by Andrew [3] is correct, I tried to find a way
> > > to get rid of this Dovetail specific IRQF_TIMER marking, as it could
> > > easily be forgotten.
> > >
> > > Review comments and intensive testing is highly appreciated.
> > > Tests were done on qemu for arm, arm64 and x86.
> > >
> > > If accepted / merged we should clean up the dovetail IRQF_TIMER markers
> > > afterwards.
> > >
> > > Best regards,
> > > Florian
> > >
> > > [1] https://lore.kernel.org/xenomai/CAKndYJF00RPbfD7U8YCAn_vzxDFdtwz5m9O=JVqhH5fdXHxY9w@mail.gmail.com/T/
> > > [2] https://lore.kernel.org/xenomai/CACiVkw=Vs1KUgfZ=Mq-rXU=4gPFDtmuoNoMv=rzy_kvV7KgpKg@mail.gmail.com/
> > > [3] https://lore.kernel.org/xenomai/20260907083708.2566-1-andrew@elk.audio/T/#u
> > >
> > > To: Xenomai <xenomai@lists.linux.dev>
> > > Cc: Philippe Gerum <rpm@xenomai.org>
> > > Cc: Gerte Hoogewerf <ghoogewerf@lmi3d.com>
> > > Cc: Andrew MacPherson <andrew@elk.audio>
> > >
> > > Suggested-by: Philippe Gerum <rpm@xenomai.org>
> > > Signed-off-by: Florian Bezdeka <florian.bezdeka@siemens.com>
> > > ---
> > > Florian Bezdeka (2):
> > > genirq: irq_pipeline: Remove irq_is_oob()
> > > irq_pipeline: Introduce IRQ_TICK to obsolete IRQF_TIMER on tick IRQs
> > >
> > > include/linux/irq.h | 11 +++++++++--
> > > include/linux/irqdesc.h | 5 -----
> > > kernel/irq/debug.h | 1 +
> > > kernel/irq/manage.c | 11 +++++++++++
> > > kernel/irq/pipeline.c | 3 ++-
> > > kernel/irq/settings.h | 17 +++++++++++++++++
> > > kernel/time/clockevents.c | 3 +++
> > > kernel/time/tick-common.c | 5 +++++
> > > 8 files changed, 48 insertions(+), 8 deletions(-)
> > > ---
> > > base-commit: 58e8e0f74c2b20f6228b7d327ec2e955ffd5b491
> > > change-id: 20260907-wip-flo-v7-2-irq-cleanup-65fb73c34f9c
> > >
> > > Best regards,
> > > --
> > > Florian Bezdeka <florian.bezdeka@siemens.com>
> > >
> >
> > Thanks for looking into this! I backported the patch to the older
> > kernel (6.6.48) we're using since it would be difficult to quickly try
> > on 7.2 and realized that my earlier testing didn't go far enough.
> >
> > What I really want to run is "perf record -g --call-graph dwarf"
> > against a running process, however in testing I was just using "perf
> > top" as a quick proxy. With both your patch and the one I submitted
> > earlier "perf top" starts returning valid results (not all 0s as
> > before). However, both patches result in a kernel panic when trying to
> > use them to get a call graph.
> >
> > A snapshot of the panic looks like this:
> >
> > Unable to handle kernel NULL pointer dereference at virtual address
> > 00000000 when read
> > Internal error: Oops: 5 [#1] PREEMPT SMP ARM
> > CPU: 0 PID: 762 Comm: XXXXXX Tainted: G O 6.6.48 #2
> > PC is at unwind_exec_insn+0x2dc/0x47c
> > LR is at unwind_frame+0x18c/0x424
> >
> > unwind_exec_insn from unwind_frame+0x18c/0x424
> > unwind_frame from walk_stackframe+0x30/0x3c
> > walk_stackframe from perf_callchain_kernel+0x60/0x84
> > perf_callchain_kernel from get_perf_callchain+0xa8/0x220
> > get_perf_callchain from perf_callchain+0x84/0x98
> > perf_callchain from perf_prepare_sample+0x4ec/0x650
> > perf_prepare_sample from perf_event_output_forward+0x58/0xd0
> > perf_event_output_forward from __perf_event_overflow+0x98/0x25c
> > __perf_event_overflow from armv7pmu_handle_irq+0x11c/0x150
> > armv7pmu_handle_irq from armpmu_dispatch_irq+0x28/0x80
> > armpmu_dispatch_irq from __handle_irq_event_percpu+0x68/0x224
> > __handle_irq_event_percpu from handle_irq_event+0x58/0xf8
> > handle_irq_event from handle_fasteoi_irq+0x154/0x2f4
> > handle_fasteoi_irq from handle_irq_desc+0x20/0x30
> > handle_irq_desc from arch_do_IRQ_pipelined+0x58/0x98
> > arch_do_IRQ_pipelined from sync_current_irq_stage+0x90/0xc0
> > sync_current_irq_stage from handle_irq_pipelined_finish+0xa0/0x198
> > handle_irq_pipelined_finish from __irq_usr+0x60/0x80
> >
> > Kernel panic - not syncing: Fatal exception in interrupt
> >
> > As before this could just be due to the older kernel and maybe this is
> > fixed in the 7.2 branch. I also tried hacking around this issue with
> > the help of an LLM, basically to skip unwinding frames that would
> > trigger it, which lead to another panic:
> >
> > kernel BUG at kernel/irq_work.c:245!
> > Internal error: Oops - BUG: 0 [#1] PREEMPT SMP ARM
> > CPU: 0 PID: 520 Comm: XXXXXX Tainted: G O 6.6.48 #4
> > PC is at irq_work_run_list+0x14/0x64
> > LR is at irq_work_run_list+0xc/0x64
> >
> > irq_work_run_list from irq_work_run+0x28/0x3c
> > irq_work_run from armv7pmu_handle_irq+0x148/0x150
> > armv7pmu_handle_irq from armpmu_dispatch_irq+0x28/0x80
> > armpmu_dispatch_irq from __handle_irq_event_percpu+0x68/0x224
> > __handle_irq_event_percpu from handle_irq_event+0x58/0xf8
> > handle_irq_event from handle_fasteoi_irq+0x154/0x2f4
> > handle_fasteoi_irq from handle_irq_desc+0x20/0x30
> > handle_irq_desc from arch_do_IRQ_pipelined+0x58/0x90
> > arch_do_IRQ_pipelined from sync_current_irq_stage+0xac/0xe0
> > sync_current_irq_stage from handle_irq_pipelined_finish+0xa0/0x198
> > handle_irq_pipelined_finish from __irq_usr+0x60/0x80
> >
> > Kernel panic - not syncing: Fatal exception in interrupt
> >
> > Hacking also around this one I do get expected call graph results from
> > perf, ignoring some bogus sample frames caused by the hacks.
> >
> > Let me know if I can provide any more detail here, I didn't want to
> > post patches that are obviously incorrect but if they're useful for
> > debugging in any way then I can.
> >
>
> Thanks for letting me know. Could you please try backporting
>
> 58e8e0f74c2b ("ARM: irq_pipeline: save registers used for walking tick
> frames")
>
> as well and try again.
>
> That was fixed in 7.2 recently and it seems backporting did not start
> yet. Your branch is EOL, so... do not expect it to land there.
>
> HTH,
> Florian
>
That patch resolved the first crash (in unwind_exec_insn) but the
second one still remains (in irq_work_run_list).
And apologies for the older kernel, we're aware that we're on a
deprecated version and are in the process of upgrading, just wanted to
report what we're seeing here in case it's relevant in later versions.
Thanks!
Andrew
^ permalink raw reply [flat|nested] 13+ messages in thread* Re: [PATCH Dovetail 0/2] Obsolete marking tick IRQs with IRQF_TIMER
2026-09-10 14:38 ` Andrew MacPherson
@ 2026-09-11 8:48 ` Florian Bezdeka
2026-09-11 12:53 ` Andrew MacPherson
0 siblings, 1 reply; 13+ messages in thread
From: Florian Bezdeka @ 2026-09-11 8:48 UTC (permalink / raw)
To: Andrew MacPherson; +Cc: Xenomai, Philippe Gerum, Gerte Hoogewerf
On Thu, 2026-09-10 at 16:38 +0200, Andrew MacPherson wrote:
>
> > kernel BUG at kernel/irq_work.c:245!
> > Internal error: Oops - BUG: 0 [#1] PREEMPT SMP ARM
> > CPU: 0 PID: 520 Comm: XXXXXX Tainted: G O 6.6.48 #4
> > PC is at irq_work_run_list+0x14/0x64
> > LR is at irq_work_run_list+0xc/0x64
> > >
> > > irq_work_run_list from irq_work_run+0x28/0x3c
> > > irq_work_run from armv7pmu_handle_irq+0x148/0x150
> > > armv7pmu_handle_irq from armpmu_dispatch_irq+0x28/0x80
> > > armpmu_dispatch_irq from __handle_irq_event_percpu+0x68/0x224
> > > __handle_irq_event_percpu from handle_irq_event+0x58/0xf8
> > > handle_irq_event from handle_fasteoi_irq+0x154/0x2f4
> > > handle_fasteoi_irq from handle_irq_desc+0x20/0x30
> > > handle_irq_desc from arch_do_IRQ_pipelined+0x58/0x90
> > > arch_do_IRQ_pipelined from sync_current_irq_stage+0xac/0xe0
> > > sync_current_irq_stage from handle_irq_pipelined_finish+0xa0/0x198
> > > handle_irq_pipelined_finish from __irq_usr+0x60/0x80
> > >
> > > Kernel panic - not syncing: Fatal exception in interrupt
> > >
> > >
I can't reproduce that one here. Not on 7.2 nor on 6.6.49.
Might be that this is depending on
- your workload (irq work)
- kernel configuration
- xenomai version (which one do you use?)
- qemu vs. real hw
I'm quite sure that this is a different issue as we hit a BUG() in
irq_work_run_list():
BUG_ON(!irqs_disabled() && !IS_ENABLED(CONFIG_PREEMPT_RT));
On first glance that looks like a corruption of the virtual interrupt
state. An irq_work event was waiting in the IRQ log and got applied with
a wrong inband IRQ state.
Last time we found issues on arm was probably [1]. Seems that this
series was not backported (yet). [1] got merged into 7.1. Maybe you can
give it a try. I don't have futher ideas at the moment.
Best regards,
Florian
[1] https://lore.kernel.org/xenomai/87ik6qvqck.fsf@xenomai.org/T/#m31701ab94206574c96cc2d2b24d77d9793359892
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [PATCH Dovetail 0/2] Obsolete marking tick IRQs with IRQF_TIMER
2026-09-11 8:48 ` Florian Bezdeka
@ 2026-09-11 12:53 ` Andrew MacPherson
2026-09-11 14:01 ` Florian Bezdeka
0 siblings, 1 reply; 13+ messages in thread
From: Andrew MacPherson @ 2026-09-11 12:53 UTC (permalink / raw)
To: Florian Bezdeka; +Cc: Xenomai, Philippe Gerum, Gerte Hoogewerf
On Fri, 11 Sept 2026 at 10:48, Florian Bezdeka
<florian.bezdeka@siemens.com> wrote:
>
> On Thu, 2026-09-10 at 16:38 +0200, Andrew MacPherson wrote:
> >
> > > kernel BUG at kernel/irq_work.c:245!
> > > Internal error: Oops - BUG: 0 [#1] PREEMPT SMP ARM
> > > CPU: 0 PID: 520 Comm: XXXXXX Tainted: G O 6.6.48 #4
> > > PC is at irq_work_run_list+0x14/0x64
> > > LR is at irq_work_run_list+0xc/0x64
> > > >
> > > > irq_work_run_list from irq_work_run+0x28/0x3c
> > > > irq_work_run from armv7pmu_handle_irq+0x148/0x150
> > > > armv7pmu_handle_irq from armpmu_dispatch_irq+0x28/0x80
> > > > armpmu_dispatch_irq from __handle_irq_event_percpu+0x68/0x224
> > > > __handle_irq_event_percpu from handle_irq_event+0x58/0xf8
> > > > handle_irq_event from handle_fasteoi_irq+0x154/0x2f4
> > > > handle_fasteoi_irq from handle_irq_desc+0x20/0x30
> > > > handle_irq_desc from arch_do_IRQ_pipelined+0x58/0x90
> > > > arch_do_IRQ_pipelined from sync_current_irq_stage+0xac/0xe0
> > > > sync_current_irq_stage from handle_irq_pipelined_finish+0xa0/0x198
> > > > handle_irq_pipelined_finish from __irq_usr+0x60/0x80
> > > >
> > > > Kernel panic - not syncing: Fatal exception in interrupt
> > > >
> > > >
>
> I can't reproduce that one here. Not on 7.2 nor on 6.6.49.
> Might be that this is depending on
> - your workload (irq work)
> - kernel configuration
> - xenomai version (which one do you use?)
> - qemu vs. real hw
>
> I'm quite sure that this is a different issue as we hit a BUG() in
> irq_work_run_list():
>
> BUG_ON(!irqs_disabled() && !IS_ENABLED(CONFIG_PREEMPT_RT));
>
> On first glance that looks like a corruption of the virtual interrupt
> state. An irq_work event was waiting in the IRQ log and got applied with
> a wrong inband IRQ state.
>
> Last time we found issues on arm was probably [1]. Seems that this
> series was not backported (yet). [1] got merged into 7.1. Maybe you can
> give it a try. I don't have futher ideas at the moment.
>
> Best regards,
> Florian
>
> [1] https://lore.kernel.org/xenomai/87ik6qvqck.fsf@xenomai.org/T/#m31701ab94206574c96cc2d2b24d77d9793359892
I tried backporting the patches from [1] and they do in fact resolve
the crash! Narrowing it down somewhat it seems that specifically
patches 5+6 together are enough to do it, if I only apply those and
skip the rest the crash is still resolved. With these patches plus the
earlier one I now get clean results from a perf callgraph run.
An LLM's static analysis of the issue is: do_page_fault() checked the
raw hardware IRQ flag instead of the correct in-band-stall state
before re-enabling interrupts - and page faults are frequent enough
(routine memory access) that this wrong check fired constantly. Since
real hardware IRQs are deliberately left on during Dovetail's in-band
IRQ replay, and the stall bit that guards against reentrancy is a
single un-counted flag, that wrong early local_irq_enable() call
opened a window for a genuinely nested interrupt to corrupt the
stall/hardirq bookkeeping that irq_work_run_list()'s assertion depends
on.
Thanks again for the help tracking this down!
Cheers,
Andrew
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [PATCH Dovetail 0/2] Obsolete marking tick IRQs with IRQF_TIMER
2026-09-11 12:53 ` Andrew MacPherson
@ 2026-09-11 14:01 ` Florian Bezdeka
2026-09-11 14:11 ` Bezdeka, Florian
0 siblings, 1 reply; 13+ messages in thread
From: Florian Bezdeka @ 2026-09-11 14:01 UTC (permalink / raw)
To: Andrew MacPherson, Philippe Gerum, Jan Kiszka; +Cc: Xenomai, Gerte Hoogewerf
On Fri, 2026-09-11 at 14:53 +0200, Andrew MacPherson wrote:
> On Fri, 11 Sept 2026 at 10:48, Florian Bezdeka
> <florian.bezdeka@siemens.com> wrote:
> >
> > On Thu, 2026-09-10 at 16:38 +0200, Andrew MacPherson wrote:
> > >
> > > > kernel BUG at kernel/irq_work.c:245!
> > > > Internal error: Oops - BUG: 0 [#1] PREEMPT SMP ARM
> > > > CPU: 0 PID: 520 Comm: XXXXXX Tainted: G O 6.6.48 #4
> > > > PC is at irq_work_run_list+0x14/0x64
> > > > LR is at irq_work_run_list+0xc/0x64
> > > > >
> > > > > irq_work_run_list from irq_work_run+0x28/0x3c
> > > > > irq_work_run from armv7pmu_handle_irq+0x148/0x150
> > > > > armv7pmu_handle_irq from armpmu_dispatch_irq+0x28/0x80
> > > > > armpmu_dispatch_irq from __handle_irq_event_percpu+0x68/0x224
> > > > > __handle_irq_event_percpu from handle_irq_event+0x58/0xf8
> > > > > handle_irq_event from handle_fasteoi_irq+0x154/0x2f4
> > > > > handle_fasteoi_irq from handle_irq_desc+0x20/0x30
> > > > > handle_irq_desc from arch_do_IRQ_pipelined+0x58/0x90
> > > > > arch_do_IRQ_pipelined from sync_current_irq_stage+0xac/0xe0
> > > > > sync_current_irq_stage from handle_irq_pipelined_finish+0xa0/0x198
> > > > > handle_irq_pipelined_finish from __irq_usr+0x60/0x80
> > > > >
> > > > > Kernel panic - not syncing: Fatal exception in interrupt
> > > > >
> > > > >
> >
> > I can't reproduce that one here. Not on 7.2 nor on 6.6.49.
> > Might be that this is depending on
> > - your workload (irq work)
> > - kernel configuration
> > - xenomai version (which one do you use?)
> > - qemu vs. real hw
> >
> > I'm quite sure that this is a different issue as we hit a BUG() in
> > irq_work_run_list():
> >
> > BUG_ON(!irqs_disabled() && !IS_ENABLED(CONFIG_PREEMPT_RT));
> >
> > On first glance that looks like a corruption of the virtual interrupt
> > state. An irq_work event was waiting in the IRQ log and got applied with
> > a wrong inband IRQ state.
> >
> > Last time we found issues on arm was probably [1]. Seems that this
> > series was not backported (yet). [1] got merged into 7.1. Maybe you can
> > give it a try. I don't have futher ideas at the moment.
> >
> > Best regards,
> > Florian
> >
> > [1] https://lore.kernel.org/xenomai/87ik6qvqck.fsf@xenomai.org/T/#m31701ab94206574c96cc2d2b24d77d9793359892
>
> I tried backporting the patches from [1] and they do in fact resolve
> the crash! Narrowing it down somewhat it seems that specifically
> patches 5+6 together are enough to do it, if I only apply those and
> skip the rest the crash is still resolved. With these patches plus the
> earlier one I now get clean results from a perf callgraph run.
>
> An LLM's static analysis of the issue is: do_page_fault() checked the
> raw hardware IRQ flag instead of the correct in-band-stall state
> before re-enabling interrupts - and page faults are frequent enough
> (routine memory access) that this wrong check fired constantly. Since
> real hardware IRQs are deliberately left on during Dovetail's in-band
> IRQ replay, and the stall bit that guards against reentrancy is a
> single un-counted flag, that wrong early local_irq_enable() call
> opened a window for a genuinely nested interrupt to corrupt the
> stall/hardirq bookkeeping that irq_work_run_list()'s assertion depends
> on.
>
> Thanks again for the help tracking this down!
Thanks for testing and reporting back. Highly appreciated!
So let me inform "stable" maintainers that [1] fixes a real problem. I
was just reviewing parts of the arm pipeline implementation back then
and realized that there are potential problems. Now they got real ;-)
@Jan, Philippe:
Could you please take care of [1] being applied into older, but still
maintained branches. Thanks!
This series would also be needed to fix some perf problems on arm/arm64.
But let's wait for some more feedback first.
Thanks all,
Florian
[1] https://lore.kernel.org/xenomai/87ik6qvqck.fsf@xenomai.org/T/#m31701ab94206574c96cc2d2b24d77d9793359892
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [PATCH Dovetail 0/2] Obsolete marking tick IRQs with IRQF_TIMER
2026-09-11 14:01 ` Florian Bezdeka
@ 2026-09-11 14:11 ` Bezdeka, Florian
2026-09-11 14:31 ` Philippe Gerum
0 siblings, 1 reply; 13+ messages in thread
From: Bezdeka, Florian @ 2026-09-11 14:11 UTC (permalink / raw)
To: andrew@elk.audio, rpm@xenomai.org, Kiszka, Jan
Cc: ghoogewerf@lmi3d.com, xenomai@lists.linux.dev
On Fri, 2026-09-11 at 16:01 +0200, Florian Bezdeka wrote:
> On Fri, 2026-09-11 at 14:53 +0200, Andrew MacPherson wrote:
> > On Fri, 11 Sept 2026 at 10:48, Florian Bezdeka
> > <florian.bezdeka@siemens.com> wrote:
> > >
> > > On Thu, 2026-09-10 at 16:38 +0200, Andrew MacPherson wrote:
> > > >
> > > > > kernel BUG at kernel/irq_work.c:245!
> > > > > Internal error: Oops - BUG: 0 [#1] PREEMPT SMP ARM
> > > > > CPU: 0 PID: 520 Comm: XXXXXX Tainted: G O 6.6.48 #4
> > > > > PC is at irq_work_run_list+0x14/0x64
> > > > > LR is at irq_work_run_list+0xc/0x64
> > > > > >
> > > > > > irq_work_run_list from irq_work_run+0x28/0x3c
> > > > > > irq_work_run from armv7pmu_handle_irq+0x148/0x150
> > > > > > armv7pmu_handle_irq from armpmu_dispatch_irq+0x28/0x80
> > > > > > armpmu_dispatch_irq from __handle_irq_event_percpu+0x68/0x224
> > > > > > __handle_irq_event_percpu from handle_irq_event+0x58/0xf8
> > > > > > handle_irq_event from handle_fasteoi_irq+0x154/0x2f4
> > > > > > handle_fasteoi_irq from handle_irq_desc+0x20/0x30
> > > > > > handle_irq_desc from arch_do_IRQ_pipelined+0x58/0x90
> > > > > > arch_do_IRQ_pipelined from sync_current_irq_stage+0xac/0xe0
> > > > > > sync_current_irq_stage from handle_irq_pipelined_finish+0xa0/0x198
> > > > > > handle_irq_pipelined_finish from __irq_usr+0x60/0x80
> > > > > >
> > > > > > Kernel panic - not syncing: Fatal exception in interrupt
> > > > > >
> > > > > >
> > >
> > > I can't reproduce that one here. Not on 7.2 nor on 6.6.49.
> > > Might be that this is depending on
> > > - your workload (irq work)
> > > - kernel configuration
> > > - xenomai version (which one do you use?)
> > > - qemu vs. real hw
> > >
> > > I'm quite sure that this is a different issue as we hit a BUG() in
> > > irq_work_run_list():
> > >
> > > BUG_ON(!irqs_disabled() && !IS_ENABLED(CONFIG_PREEMPT_RT));
> > >
> > > On first glance that looks like a corruption of the virtual interrupt
> > > state. An irq_work event was waiting in the IRQ log and got applied with
> > > a wrong inband IRQ state.
> > >
> > > Last time we found issues on arm was probably [1]. Seems that this
> > > series was not backported (yet). [1] got merged into 7.1. Maybe you can
> > > give it a try. I don't have futher ideas at the moment.
> > >
> > > Best regards,
> > > Florian
> > >
> > > [1] https://lore.kernel.org/xenomai/87ik6qvqck.fsf@xenomai.org/T/#m31701ab94206574c96cc2d2b24d77d9793359892
> >
> > I tried backporting the patches from [1] and they do in fact resolve
> > the crash! Narrowing it down somewhat it seems that specifically
> > patches 5+6 together are enough to do it, if I only apply those and
> > skip the rest the crash is still resolved. With these patches plus the
> > earlier one I now get clean results from a perf callgraph run.
> >
> > An LLM's static analysis of the issue is: do_page_fault() checked the
> > raw hardware IRQ flag instead of the correct in-band-stall state
> > before re-enabling interrupts - and page faults are frequent enough
> > (routine memory access) that this wrong check fired constantly. Since
> > real hardware IRQs are deliberately left on during Dovetail's in-band
> > IRQ replay, and the stall bit that guards against reentrancy is a
> > single un-counted flag, that wrong early local_irq_enable() call
> > opened a window for a genuinely nested interrupt to corrupt the
> > stall/hardirq bookkeeping that irq_work_run_list()'s assertion depends
> > on.
> >
> > Thanks again for the help tracking this down!
>
> Thanks for testing and reporting back. Highly appreciated!
>
> So let me inform "stable" maintainers that [1] fixes a real problem. I
> was just reviewing parts of the arm pipeline implementation back then
> and realized that there are potential problems. Now they got real ;-)
>
> @Jan, Philippe:
> Could you please take care of [1] being applied into older, but still
> maintained branches. Thanks!
>
> This series would also be needed to fix some perf problems on arm/arm64.
> But let's wait for some more feedback first.
Forgot one end: Additionally we need
58e8e0f74c2b ("ARM: irq_pipeline: save registers used for walking tick frames")
to fix the perf issues. Couldn't find it on the list. Seems it sneaked
in silently ;-)
>
> [1] https://lore.kernel.org/xenomai/87ik6qvqck.fsf@xenomai.org/T/#m31701ab94206574c96cc2d2b24d77d9793359892
^ permalink raw reply [flat|nested] 13+ messages in thread* Re: [PATCH Dovetail 0/2] Obsolete marking tick IRQs with IRQF_TIMER
2026-09-11 14:11 ` Bezdeka, Florian
@ 2026-09-11 14:31 ` Philippe Gerum
0 siblings, 0 replies; 13+ messages in thread
From: Philippe Gerum @ 2026-09-11 14:31 UTC (permalink / raw)
To: Bezdeka, Florian
Cc: andrew@elk.audio, Kiszka, Jan, ghoogewerf@lmi3d.com,
xenomai@lists.linux.dev
"Bezdeka, Florian" <florian.bezdeka@siemens.com> writes:
> On Fri, 2026-09-11 at 16:01 +0200, Florian Bezdeka wrote:
>> On Fri, 2026-09-11 at 14:53 +0200, Andrew MacPherson wrote:
>> > On Fri, 11 Sept 2026 at 10:48, Florian Bezdeka
>> > <florian.bezdeka@siemens.com> wrote:
>> > >
>> > > On Thu, 2026-09-10 at 16:38 +0200, Andrew MacPherson wrote:
>> > > >
>> > > > > kernel BUG at kernel/irq_work.c:245!
>> > > > > Internal error: Oops - BUG: 0 [#1] PREEMPT SMP ARM
>> > > > > CPU: 0 PID: 520 Comm: XXXXXX Tainted: G O 6.6.48 #4
>> > > > > PC is at irq_work_run_list+0x14/0x64
>> > > > > LR is at irq_work_run_list+0xc/0x64
>> > > > > >
>> > > > > > irq_work_run_list from irq_work_run+0x28/0x3c
>> > > > > > irq_work_run from armv7pmu_handle_irq+0x148/0x150
>> > > > > > armv7pmu_handle_irq from armpmu_dispatch_irq+0x28/0x80
>> > > > > > armpmu_dispatch_irq from __handle_irq_event_percpu+0x68/0x224
>> > > > > > __handle_irq_event_percpu from handle_irq_event+0x58/0xf8
>> > > > > > handle_irq_event from handle_fasteoi_irq+0x154/0x2f4
>> > > > > > handle_fasteoi_irq from handle_irq_desc+0x20/0x30
>> > > > > > handle_irq_desc from arch_do_IRQ_pipelined+0x58/0x90
>> > > > > > arch_do_IRQ_pipelined from sync_current_irq_stage+0xac/0xe0
>> > > > > > sync_current_irq_stage from handle_irq_pipelined_finish+0xa0/0x198
>> > > > > > handle_irq_pipelined_finish from __irq_usr+0x60/0x80
>> > > > > >
>> > > > > > Kernel panic - not syncing: Fatal exception in interrupt
>> > > > > >
>> > > > > >
>> > >
>> > > I can't reproduce that one here. Not on 7.2 nor on 6.6.49.
>> > > Might be that this is depending on
>> > > - your workload (irq work)
>> > > - kernel configuration
>> > > - xenomai version (which one do you use?)
>> > > - qemu vs. real hw
>> > >
>> > > I'm quite sure that this is a different issue as we hit a BUG() in
>> > > irq_work_run_list():
>> > >
>> > > BUG_ON(!irqs_disabled() && !IS_ENABLED(CONFIG_PREEMPT_RT));
>> > >
>> > > On first glance that looks like a corruption of the virtual interrupt
>> > > state. An irq_work event was waiting in the IRQ log and got applied with
>> > > a wrong inband IRQ state.
>> > >
>> > > Last time we found issues on arm was probably [1]. Seems that this
>> > > series was not backported (yet). [1] got merged into 7.1. Maybe you can
>> > > give it a try. I don't have futher ideas at the moment.
>> > >
>> > > Best regards,
>> > > Florian
>> > >
>> > > [1] https://lore.kernel.org/xenomai/87ik6qvqck.fsf@xenomai.org/T/#m31701ab94206574c96cc2d2b24d77d9793359892
>> >
>> > I tried backporting the patches from [1] and they do in fact resolve
>> > the crash! Narrowing it down somewhat it seems that specifically
>> > patches 5+6 together are enough to do it, if I only apply those and
>> > skip the rest the crash is still resolved. With these patches plus the
>> > earlier one I now get clean results from a perf callgraph run.
>> >
>> > An LLM's static analysis of the issue is: do_page_fault() checked the
>> > raw hardware IRQ flag instead of the correct in-band-stall state
>> > before re-enabling interrupts - and page faults are frequent enough
>> > (routine memory access) that this wrong check fired constantly. Since
>> > real hardware IRQs are deliberately left on during Dovetail's in-band
>> > IRQ replay, and the stall bit that guards against reentrancy is a
>> > single un-counted flag, that wrong early local_irq_enable() call
>> > opened a window for a genuinely nested interrupt to corrupt the
>> > stall/hardirq bookkeeping that irq_work_run_list()'s assertion depends
>> > on.
>> >
>> > Thanks again for the help tracking this down!
>>
>> Thanks for testing and reporting back. Highly appreciated!
>>
>> So let me inform "stable" maintainers that [1] fixes a real problem. I
>> was just reviewing parts of the arm pipeline implementation back then
>> and realized that there are potential problems. Now they got real ;-)
>>
>> @Jan, Philippe:
>> Could you please take care of [1] being applied into older, but still
>> maintained branches. Thanks!
>>
>> This series would also be needed to fix some perf problems on arm/arm64.
>> But let's wait for some more feedback first.
>
> Forgot one end: Additionally we need
>
> 58e8e0f74c2b ("ARM: irq_pipeline: save registers used for walking tick frames")
>
> to fix the perf issues. Couldn't find it on the list. Seems it sneaked
> in silently ;-)
>
https://lore.kernel.org/xenomai/CAKndYJHo--bOwSh4Sg7wyuPuWnFcHkPrg_bBRLNYWBm3VQ-VnQ@mail.gmail.com/T/#mf36158864c1a3225ef5305f1557052c955260ae2
--
Philippe.
^ permalink raw reply [flat|nested] 13+ messages in thread