Linux real-time development
 help / color / mirror / Atom feed
* [RFC PATCH] sched: add CONFIG_SCHED_EXT_RT_EXPERIMENTAL to move SCHED_EXT above SCHED_RR/SCHED_FIFO
@ 2026-04-14 12:59 Fredrik Boettger Bratberg
  2026-04-14 17:23 ` Tejun Heo
  0 siblings, 1 reply; 7+ messages in thread
From: Fredrik Boettger Bratberg @ 2026-04-14 12:59 UTC (permalink / raw)
  To: fredrik.b.bratberg
  Cc: Henrik Austad, Peter Ziljstra, Tejun Heo, Steven Rostedt,
	Thomas Gleixner, linux-rt-devel, sched-ext

From: fredrikBB <fredrik.b.bratberg@gmail.com>

sched_ext is a great tool to quickly develop and test custom scheduling
policies. However, a current limitation of sched_ext is that it is fixed
between the idle and fair classes in the scheduling class hierarchy. This
makes it problematic to implement new real-time policies with sched_ext.
For research and experimental purposes, it is useful to have the option to
increase sched_ext priority relative to the other classes.

One particular use case is the implementation of mixed-criticality
schedulers such as EDF-VD, where real-time capabilities are required. In
such cases, a priority equal to or higher than that of SCHED_RR and
SCHED_FIFO is desired in order to limit interference by other tasks.
Framed differently, this option enables experimentation with custom
real-time scheduling policies in sched_ext.

A recent paper published in the 2025 Real-Time Systems Symposium (RTSS)
"Enabling Flexible Scheduling in ROS 2", would presumably also benefint
from this feature. The paper aims to address and improve the real-time
capabilities of Robot Operating System 2 (ROS2). As part of their paper
they implement an EDF using sched_ext. They do not explicitly address the
issue of their EDF SCHED_EXT tasks effectively having a lower priority than
SCHED_NORMAL, but it is a known limitation of the current sched_ext class
placement

This change conditionally reorders sched_class selection such that, when
enabled:

  STOP > DL > EXT > RT > FAIR > IDLE

Normally, it is possible to schedule all fair-class policies under
sched_ext. CONFIG_SCHED_EXT_RT_EXPERIMENTAL removes this possibility in
order to prevent fair-class policies from effectively increasing their
precedence in the scheduling class hierarchy. This is done by forcing
scx_switched_all() and scx_switching_all to false.

This patch has been tested on a BeagleV-Fire development board with
PREEMPT_RT enabled on the 6.18.6 kernel, and it made SCHED_EXT tasks be
preempted by SCHED_RR / SCHED_FIFO tasks to a much lesser extent as
desired. However, it still occasionally happen that SCHED_EXT tasks gets
preempted by other tasks which do not run in an hard IRQ context. Comments
regarding why this can be the case are also greatly appreciated.

Reviewed-by: Henrik Austad <henrik@austad.us>
Signed-off-by: Fredrik Boettger Bratberg <fredrik.b.bratberg@gmail.com>
Cc: Peter Ziljstra <peterz@infradead.org> 
Cc: Tejun Heo <tj@kernel.org>
Cc: Steven Rostedt <rostedt@goodmis.org>
Cc: Thomas Gleixner <tglx@kernel.org>
Cc: Henrik Austad <henrik@austad.us>
Cc: linux-rt-devel@lists.linux.dev
Cc: sched-ext@lists.linux.dev
---
 include/asm-generic/vmlinux.lds.h | 13 +++++++++++++
 kernel/Kconfig.preempt            | 16 ++++++++++++++++
 kernel/sched/core.c               |  6 ++++++
 kernel/sched/ext.c                | 11 +++++++++++
 kernel/sched/sched.h              |  8 ++++++++
 5 files changed, 54 insertions(+)

diff --git a/include/asm-generic/vmlinux.lds.h b/include/asm-generic/vmlinux.lds.h
index e04d56a53..3fe9504c4 100644
--- a/include/asm-generic/vmlinux.lds.h
+++ b/include/asm-generic/vmlinux.lds.h
@@ -133,6 +133,18 @@ defined(CONFIG_AUTOFDO_CLANG) || defined(CONFIG_PROPELLER_CLANG)
  * used to determine the order of the priority of each sched class in
  * relation to each other.
  */
+#ifdef CONFIG_SCHED_EXT_RT_EXPERIMENTAL
+#define SCHED_DATA				\
+	STRUCT_ALIGN();				\
+	__sched_class_highest = .;		\
+	*(__stop_sched_class)			\
+	*(__dl_sched_class)			\
+	*(__ext_sched_class)			\
+	*(__rt_sched_class)			\
+	*(__fair_sched_class)			\
+	*(__idle_sched_class)			\
+	__sched_class_lowest = .;
+#else
 #define SCHED_DATA				\
 	STRUCT_ALIGN();				\
 	__sched_class_highest = .;		\
@@ -143,6 +155,7 @@ defined(CONFIG_AUTOFDO_CLANG) || defined(CONFIG_PROPELLER_CLANG)
 	*(__ext_sched_class)			\
 	*(__idle_sched_class)			\
 	__sched_class_lowest = .;
+#endif
 
 /* The actual configuration determine if the init/exit sections
  * are handled as text/data or they can be discarded (which
diff --git a/kernel/Kconfig.preempt b/kernel/Kconfig.preempt
index da326800c..553d22605 100644
--- a/kernel/Kconfig.preempt
+++ b/kernel/Kconfig.preempt
@@ -189,3 +189,19 @@ config SCHED_CLASS_EXT
 	  For more information:
 	    Documentation/scheduler/sched-ext.rst
 	    https://github.com/sched-ext/scx
+
+config SCHED_EXT_RT_EXPERIMENTAL
+	bool "[EXPERIMENTAL] Increase sched_ext priority above SCHED_CLASS_RT"
+	depends on SCHED_CLASS_EXT
+	help
+	  This option increases sched_ext's class precedence above SCHED_RR and
+	  SCHED_FIFO which improves the real-time capabilities of sched_ext.
+	  Normally the sched_ext scheduler is placed in between the fair
+	  and idle scheduling class, which can cause exessive preemption of
+	  SCHED_EXT tasks by both real-time and fair class tasks.
+
+	  Note that this option will force SCX_OPS_SWITCH_PARTIAL, meaning only tasks
+	  with the SCHED_EXT policy can be scheduled by sched_ext. This
+	  is in contrast to the default behaviour where all SCHED_NORMAL,
+	  SCHED_BATCH, SCHED_IDLE, and SCHED_EXT tasks are scheduled by
+	  sched_ext when the BPF scheduler is loaded and running.
diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index eb47d294e..93d16f665 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -8654,8 +8654,14 @@ void __init sched_init(void)
 	BUG_ON(!sched_class_above(&rt_sched_class, &fair_sched_class));
 	BUG_ON(!sched_class_above(&fair_sched_class, &idle_sched_class));
 #ifdef CONFIG_SCHED_CLASS_EXT
+	/* Reorder BUG_ON if sched_ext is moved in the hierarchy */
+	#ifdef CONFIG_SCHED_EXT_RT_EXPERIMENTAL
+	BUG_ON(!sched_class_above(&dl_sched_class, &ext_sched_class));
+	BUG_ON(!sched_class_above(&ext_sched_class, &rt_sched_class));
+	#else
 	BUG_ON(!sched_class_above(&fair_sched_class, &ext_sched_class));
 	BUG_ON(!sched_class_above(&ext_sched_class, &idle_sched_class));
+	#endif
 #endif
 
 	wait_bit_init();
diff --git a/kernel/sched/ext.c b/kernel/sched/ext.c
index 31eda2a56..dea807971 100644
--- a/kernel/sched/ext.c
+++ b/kernel/sched/ext.c
@@ -4776,7 +4776,15 @@ static int scx_enable(struct sched_ext_ops *ops, struct bpf_link *link)
 	 * All tasks are READY. It's safe to turn on scx_enabled() and switch
 	 * all eligible tasks.
 	 */
+/*
+ * If sched_ext is elevated in the scheduling hierarchy, we cannot allow
+ * sched_ext to schedule tasks part of the fair classes.
+ */
+#ifdef CONFIG_SCHED_EXT_RT_EXPERIMENTAL
+	WRITE_ONCE(scx_switching_all, false);
+#else
 	WRITE_ONCE(scx_switching_all, !(ops->flags & SCX_OPS_SWITCH_PARTIAL));
+#endif
 	static_branch_enable(&__scx_enabled);
 
 	/*
@@ -4819,8 +4827,11 @@ static int scx_enable(struct sched_ext_ops *ops, struct bpf_link *link)
 		goto err_disable;
 	}
 
+/* When running with SCHED_EXT_RT, only allow switch_partial */
+#ifndef CONFIG_SCHED_EXT_RT_EXPERIMENTAL
 	if (!(ops->flags & SCX_OPS_SWITCH_PARTIAL))
 		static_branch_enable(&__scx_switched_all);
+#endif
 
 	pr_info("sched_ext: BPF scheduler \"%s\" enabled%s\n",
 		sch->ops.name, scx_switched_all() ? "" : " (partial)");
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index 2f8b06b12..6e9c1cae7 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -1762,7 +1762,15 @@ DECLARE_STATIC_KEY_FALSE(__scx_enabled);	/* SCX BPF scheduler loaded */
 DECLARE_STATIC_KEY_FALSE(__scx_switched_all);	/* all fair class tasks on SCX */
 
 #define scx_enabled()		static_branch_unlikely(&__scx_enabled)
+/*
+ * If sched_ext is elevated in the scheduling hierarchy, we cannot allow
+ * sched_ext to schedule tasks part of the fair classes.
+ */
+#ifdef CONFIG_SCHED_EXT_RT_EXPERIMENTAL
+#define scx_switched_all()	false
+#else
 #define scx_switched_all()	static_branch_unlikely(&__scx_switched_all)
+#endif
 
 static inline void scx_rq_clock_update(struct rq *rq, u64 clock)
 {
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 7+ messages in thread

end of thread, other threads:[~2026-09-28 15:39 UTC | newest]

Thread overview: 7+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-04-14 12:59 [RFC PATCH] sched: add CONFIG_SCHED_EXT_RT_EXPERIMENTAL to move SCHED_EXT above SCHED_RR/SCHED_FIFO Fredrik Boettger Bratberg
2026-04-14 17:23 ` Tejun Heo
2026-04-16  8:11   ` Fredrik Bratberg
2026-04-16 17:59     ` Tejun Heo
2026-04-17  6:30       ` Fredrik Bratberg
2026-09-21  7:47   ` Tengfei Fan
2026-09-28 15:39     ` Christian Loehle

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox