* [PATCH v14 01/13] sched/cputime: Add kcpustat_field_total helper
2026-09-28 5:37 [PATCH v14 00/13] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff Shrikanth Hegde
@ 2026-09-28 5:37 ` Shrikanth Hegde
2026-09-28 5:54 ` sashiko-bot
2026-09-28 5:37 ` [PATCH v14 02/13] cpumask: Introduce cpumask_intersects_and Shrikanth Hegde
` (11 subsequent siblings)
12 siblings, 1 reply; 32+ messages in thread
From: Shrikanth Hegde @ 2026-09-28 5:37 UTC (permalink / raw)
To: linux-kernel, mingo, peterz, juri.lelli, vincent.guittot,
yury.norov, kprateek.nayak, iii, corbet, meted, ynorov
Cc: sshegde, tglx, gregkh, pbonzini, seanjc, vschneid, huschle,
rostedt, dietmar.eggemann, maddy, srikar, hdanton, chleroy,
vineeth, frederic, arighi, pauld, christian.loehle, tj,
tommaso.cucinotta, maz, rafael, rdunlap, kernellwp, linux-doc,
jgross, virtualization, sunlightlinux
Provide a new helper function which sums up a given type of cpustat
over a specified cpumask.
This allows the caller's code to be simpler and avoids duplication.
For example, subsequent patch in the steal governor use this exact
same pattern when calculating steal time.
Number of cpus can be derived from cpumask_weight() where necessary.
Suggested-by: Yury Norov <yury.norov@gmail.com>
Acked-by: Frederic Weisbecker <frederic@kernel.org>
Reviewed-by: Yury Norov <ynorov@nvidia.com>
Reviewed-by: Mete Durlu <meted@linux.ibm.com>
Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
---
arch/s390/kernel/hiperdispatch.c | 10 +++-------
fs/proc/uptime.c | 6 +-----
include/linux/kernel_stat.h | 11 +++++++++++
3 files changed, 15 insertions(+), 12 deletions(-)
diff --git a/arch/s390/kernel/hiperdispatch.c b/arch/s390/kernel/hiperdispatch.c
index 217206522266..c21496f0a141 100644
--- a/arch/s390/kernel/hiperdispatch.c
+++ b/arch/s390/kernel/hiperdispatch.c
@@ -207,16 +207,12 @@ static unsigned long hd_calculate_steal_percentage(void)
{
unsigned long time_delta, steal_delta, steal, percentage;
static ktime_t prev;
- int cpus, cpu;
+ int cpus;
ktime_t now;
- cpus = 0;
- steal = 0;
percentage = 0;
- for_each_cpu(cpu, &hd_vmvl_cpumask) {
- steal += kcpustat_cpu(cpu).cpustat[CPUTIME_STEAL];
- cpus++;
- }
+ steal = kcpustat_field_total(CPUTIME_STEAL, &hd_vmvl_cpumask);
+ cpus = cpumask_weight(&hd_vmvl_cpumask);
/*
* If there is no vertical medium and low CPUs steal time
* is 0 as vertical high CPUs shouldn't experience steal time.
diff --git a/fs/proc/uptime.c b/fs/proc/uptime.c
index 433aa947cd57..53143c66cbe1 100644
--- a/fs/proc/uptime.c
+++ b/fs/proc/uptime.c
@@ -15,12 +15,8 @@ static int uptime_proc_show(struct seq_file *m, void *v)
struct timespec64 idle;
u64 idle_nsec;
u32 rem;
- int i;
-
- idle_nsec = 0;
- for_each_possible_cpu(i)
- idle_nsec += kcpustat_field(CPUTIME_IDLE, i);
+ idle_nsec = kcpustat_field_total(CPUTIME_IDLE, cpu_possible_mask);
ktime_get_boottime_ts64(&uptime);
timens_add_boottime(&uptime);
diff --git a/include/linux/kernel_stat.h b/include/linux/kernel_stat.h
index 9ca6c2259dfe..c1e85550bf12 100644
--- a/include/linux/kernel_stat.h
+++ b/include/linux/kernel_stat.h
@@ -196,6 +196,17 @@ static inline void kcpustat_cpu_fetch(struct kernel_cpustat *dst, int cpu)
}
#endif /* !CONFIG_VIRT_CPU_ACCOUNTING_GEN */
+static inline u64 kcpustat_field_total(enum cpu_usage_stat usage, const struct cpumask *cpus)
+{
+ u64 total = 0;
+ int cpu;
+
+ for_each_cpu(cpu, cpus)
+ total += kcpustat_field(usage, cpu);
+
+ return total;
+}
+
extern void account_user_time(struct task_struct *, u64);
extern void account_guest_time(struct task_struct *, u64);
extern void account_system_time(struct task_struct *, int, u64);
--
2.52.0
^ permalink raw reply related [flat|nested] 32+ messages in thread* Re: [PATCH v14 01/13] sched/cputime: Add kcpustat_field_total helper
2026-09-28 5:37 ` [PATCH v14 01/13] sched/cputime: Add kcpustat_field_total helper Shrikanth Hegde
@ 2026-09-28 5:54 ` sashiko-bot
0 siblings, 0 replies; 32+ messages in thread
From: sashiko-bot @ 2026-09-28 5:54 UTC (permalink / raw)
To: Shrikanth Hegde; +Cc: virtualization, Michael S. Tsirkin, Eugenio Perez
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] s390/hiperdispatch: TOCTOU race in hd_calculate_steal_percentage() causes artificial steal percentage spikes
--
commit 543d41ec158ad38d026cd1a0e58754cc9393d347
Author: Shrikanth Hegde <sshegde@linux.ibm.com>
sched/cputime: Add kcpustat_field_total helper
This patch introduces a new helper function to sum up a given cpustat type
over a specified cpumask. It replaces the open-coded steal calculation in
s390 hiperdispatch and idle time calculation in proc uptime to simplify
the callers.
> diff --git a/arch/s390/kernel/hiperdispatch.c b/arch/s390/kernel/hiperdispatch.c
> index 2172065222661..c21496f0a141a 100644
> --- a/arch/s390/kernel/hiperdispatch.c
> +++ b/arch/s390/kernel/hiperdispatch.c
[ ... ]
> @@ -207,16 +207,12 @@ static unsigned long hd_calculate_steal_percentage(void)
> {
> unsigned long time_delta, steal_delta, steal, percentage;
> static ktime_t prev;
> - int cpus, cpu;
> + int cpus;
> ktime_t now;
>
> - cpus = 0;
> - steal = 0;
> percentage = 0;
> - for_each_cpu(cpu, &hd_vmvl_cpumask) {
> - steal += kcpustat_cpu(cpu).cpustat[CPUTIME_STEAL];
> - cpus++;
> - }
> + steal = kcpustat_field_total(CPUTIME_STEAL, &hd_vmvl_cpumask);
> + cpus = cpumask_weight(&hd_vmvl_cpumask);
[Severity: High]
Could separating the summation from the divisor count open a race window
against concurrent mask updates?
In the previous single-loop calculation, the total steal time and CPU count
were evaluated together on a stable view of the mask. With these changes,
if a concurrent CPU hotplug event calls update_cpu_masks() while
hd_calculate_steal_percentage() is running, hd_vmvl_cpumask could be cleared
and repopulated asynchronously.
If the mask changes from multiple CPUs to 1 CPU between the
kcpustat_field_total calculation and the cpumask_weight call, steal will
reflect the delta of the larger number of CPUs, but cpus will evaluate to 1.
Because the percentage is computed as steal_delta / cpus, dividing the
larger steal_delta by 1 would artificially inflate the steal percentage.
Can this erroneous spike trip the hd_steal_threshold, causing the scheduler
topology to wrongly disable Hiperdispatch and downgrade high-capacity cores?
[ ... ]
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260928053728.797539-1-sshegde@linux.ibm.com?part=1
^ permalink raw reply [flat|nested] 32+ messages in thread
* [PATCH v14 02/13] cpumask: Introduce cpumask_intersects_and
2026-09-28 5:37 [PATCH v14 00/13] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff Shrikanth Hegde
2026-09-28 5:37 ` [PATCH v14 01/13] sched/cputime: Add kcpustat_field_total helper Shrikanth Hegde
@ 2026-09-28 5:37 ` Shrikanth Hegde
2026-09-28 5:44 ` sashiko-bot
2026-09-28 5:37 ` [PATCH v14 03/13] sched/docs: Document cpu_preferred_mask and Preferred CPU concept Shrikanth Hegde
` (10 subsequent siblings)
12 siblings, 1 reply; 32+ messages in thread
From: Shrikanth Hegde @ 2026-09-28 5:37 UTC (permalink / raw)
To: linux-kernel, mingo, peterz, juri.lelli, vincent.guittot,
yury.norov, kprateek.nayak, iii, corbet, meted, ynorov
Cc: sshegde, tglx, gregkh, pbonzini, seanjc, vschneid, huschle,
rostedt, dietmar.eggemann, maddy, srikar, hdanton, chleroy,
vineeth, frederic, arighi, pauld, christian.loehle, tj,
tommaso.cucinotta, maz, rafael, rdunlap, kernellwp, linux-doc,
jgross, virtualization, sunlightlinux
Introduce bitmap_intersects_and() to determine whether the intersection
of three bitmaps is non-empty. Unlike cpumask_first_and_and(), this
returns immediately when an intersecting word is found and does not
calculate the first matching bit.
Add cpumask_intersects_and() as the corresponding cpumask wrapper.
A subsequent patch uses the helper to determine whether a task
has a CPU that is present in its affinity mask, the preferred CPU mask,
and task possible CPU mask.
Suggested-by: Yury Norov <yury.norov@gmail.com>
Reviewed-by: Yury Norov <yury.norov@gmail.com>
Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
---
include/linux/bitmap.h | 14 ++++++++++++++
include/linux/cpumask.h | 18 ++++++++++++++++++
lib/bitmap.c | 17 +++++++++++++++++
3 files changed, 49 insertions(+)
diff --git a/include/linux/bitmap.h b/include/linux/bitmap.h
index 7df1573a409c..adafbcf2016b 100644
--- a/include/linux/bitmap.h
+++ b/include/linux/bitmap.h
@@ -52,6 +52,7 @@ struct device;
* bitmap_complement(dst, src, nbits) *dst = ~(*src)
* bitmap_equal(src1, src2, nbits) Are *src1 and *src2 equal?
* bitmap_intersects(src1, src2, nbits) Do *src1 and *src2 overlap?
+ * bitmap_intersects_and(src1, src2, src3, nbits) Do *src1, *src2 and *src3 overlap?
* bitmap_subset(src1, src2, nbits) Is *src1 a subset of *src2?
* bitmap_empty(src, nbits) Are all bits zero in *src?
* bitmap_full(src, nbits) Are all bits set in *src?
@@ -181,6 +182,9 @@ void __bitmap_replace(unsigned long *dst,
const unsigned long *mask, unsigned int nbits);
bool __bitmap_intersects(const unsigned long *bitmap1,
const unsigned long *bitmap2, unsigned int nbits);
+bool __bitmap_intersects_and(const unsigned long *bitmap1,
+ const unsigned long *bitmap2,
+ const unsigned long *bitmap3, unsigned int nbits);
bool __bitmap_subset(const unsigned long *bitmap1,
const unsigned long *bitmap2, unsigned int nbits);
unsigned int __bitmap_weight(const unsigned long *bitmap, unsigned int nbits);
@@ -445,6 +449,16 @@ bool bitmap_intersects(const unsigned long *src1, const unsigned long *src2, uns
return __bitmap_intersects(src1, src2, nbits);
}
+static __always_inline
+bool bitmap_intersects_and(const unsigned long *src1, const unsigned long *src2,
+ const unsigned long *src3, unsigned int nbits)
+{
+ if (small_const_nbits(nbits))
+ return ((*src1 & *src2 & *src3) & BITMAP_LAST_WORD_MASK(nbits)) != 0;
+ else
+ return __bitmap_intersects_and(src1, src2, src3, nbits);
+}
+
static __always_inline
bool bitmap_subset(const unsigned long *src1, const unsigned long *src2, unsigned int nbits)
{
diff --git a/include/linux/cpumask.h b/include/linux/cpumask.h
index 4c8bb6953107..7c8f16797f94 100644
--- a/include/linux/cpumask.h
+++ b/include/linux/cpumask.h
@@ -824,6 +824,24 @@ bool cpumask_intersects(const struct cpumask *src1p, const struct cpumask *src2p
small_cpumask_bits);
}
+/**
+ * cpumask_intersects_and - (*src1p & *src2p & *src3p) != 0
+ * @src1p: the first input
+ * @src2p: the second input
+ * @src3p: the third input
+ *
+ * Return: true if AND of the three cpumasks is non-empty,
+ * otherwise false
+ */
+static __always_inline
+bool cpumask_intersects_and(const struct cpumask *src1p,
+ const struct cpumask *src2p,
+ const struct cpumask *src3p)
+{
+ return bitmap_intersects_and(cpumask_bits(src1p), cpumask_bits(src2p),
+ cpumask_bits(src3p), small_cpumask_bits);
+}
+
/**
* cpumask_subset - (*src1p & ~*src2p) == 0
* @src1p: the first input
diff --git a/lib/bitmap.c b/lib/bitmap.c
index ed685127a107..d1cb8a507c60 100644
--- a/lib/bitmap.c
+++ b/lib/bitmap.c
@@ -308,6 +308,23 @@ bool __bitmap_intersects(const unsigned long *bitmap1,
}
EXPORT_SYMBOL(__bitmap_intersects);
+bool __bitmap_intersects_and(const unsigned long *bitmap1,
+ const unsigned long *bitmap2,
+ const unsigned long *bitmap3, unsigned int bits)
+{
+ unsigned int k, lim = bits / BITS_PER_LONG;
+
+ for (k = 0; k < lim; ++k)
+ if (bitmap1[k] & bitmap2[k] & bitmap3[k])
+ return true;
+
+ if (bits % BITS_PER_LONG)
+ if ((bitmap1[k] & bitmap2[k] & bitmap3[k]) & BITMAP_LAST_WORD_MASK(bits))
+ return true;
+ return false;
+}
+EXPORT_SYMBOL(__bitmap_intersects_and);
+
bool __bitmap_subset(const unsigned long *bitmap1,
const unsigned long *bitmap2, unsigned int bits)
{
--
2.52.0
^ permalink raw reply related [flat|nested] 32+ messages in thread* Re: [PATCH v14 02/13] cpumask: Introduce cpumask_intersects_and
2026-09-28 5:37 ` [PATCH v14 02/13] cpumask: Introduce cpumask_intersects_and Shrikanth Hegde
@ 2026-09-28 5:44 ` sashiko-bot
0 siblings, 0 replies; 32+ messages in thread
From: sashiko-bot @ 2026-09-28 5:44 UTC (permalink / raw)
To: Shrikanth Hegde; +Cc: virtualization, Eugenio Perez, Michael S. Tsirkin
> Introduce bitmap_intersects_and() to determine whether the intersection
> of three bitmaps is non-empty. Unlike cpumask_first_and_and(), this
> returns immediately when an intersecting word is found and does not
> calculate the first matching bit.
>
> Add cpumask_intersects_and() as the corresponding cpumask wrapper.
>
> A subsequent patch uses the helper to determine whether a task
> has a CPU that is present in its affinity mask, the preferred CPU mask,
> and task possible CPU mask.
>
> Suggested-by: Yury Norov <yury.norov@gmail.com>
> Reviewed-by: Yury Norov <yury.norov@gmail.com>
> Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
Sashiko has reviewed this patch and found no issues. It looks great!
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260928053728.797539-1-sshegde@linux.ibm.com?part=2
^ permalink raw reply [flat|nested] 32+ messages in thread
* [PATCH v14 03/13] sched/docs: Document cpu_preferred_mask and Preferred CPU concept
2026-09-28 5:37 [PATCH v14 00/13] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff Shrikanth Hegde
2026-09-28 5:37 ` [PATCH v14 01/13] sched/cputime: Add kcpustat_field_total helper Shrikanth Hegde
2026-09-28 5:37 ` [PATCH v14 02/13] cpumask: Introduce cpumask_intersects_and Shrikanth Hegde
@ 2026-09-28 5:37 ` Shrikanth Hegde
2026-09-28 5:42 ` sashiko-bot
2026-09-28 5:37 ` [PATCH v14 04/13] cpumask: Introduce cpu_preferred_mask Shrikanth Hegde
` (9 subsequent siblings)
12 siblings, 1 reply; 32+ messages in thread
From: Shrikanth Hegde @ 2026-09-28 5:37 UTC (permalink / raw)
To: linux-kernel, mingo, peterz, juri.lelli, vincent.guittot,
yury.norov, kprateek.nayak, iii, corbet, meted, ynorov
Cc: sshegde, tglx, gregkh, pbonzini, seanjc, vschneid, huschle,
rostedt, dietmar.eggemann, maddy, srikar, hdanton, chleroy,
vineeth, frederic, arighi, pauld, christian.loehle, tj,
tommaso.cucinotta, maz, rafael, rdunlap, kernellwp, linux-doc,
jgross, virtualization, sunlightlinux
Add documentation for new CPU state called preferred CPU state and
corresponding cpumask called cpu_preferred_mask.
This could help users in understanding what it is and how to use it.
Document the role of scheduler and driver for this feature to work.
Details regarding the driver documentation will be added in
later patches under Documentation/driver-api/steal-governor.rst.
Newly added file could be used for other paravirt usecase documentation.
Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
---
Documentation/scheduler/index.rst | 1 +
Documentation/scheduler/sched-paravirt.rst | 67 ++++++++++++++++++++++
2 files changed, 68 insertions(+)
create mode 100644 Documentation/scheduler/sched-paravirt.rst
diff --git a/Documentation/scheduler/index.rst b/Documentation/scheduler/index.rst
index 17ce8d76befc..a43647b9706d 100644
--- a/Documentation/scheduler/index.rst
+++ b/Documentation/scheduler/index.rst
@@ -23,5 +23,6 @@ Scheduler
sched-stats
sched-ext
sched-debug
+ sched-paravirt
text_files
diff --git a/Documentation/scheduler/sched-paravirt.rst b/Documentation/scheduler/sched-paravirt.rst
new file mode 100644
index 000000000000..311cdbf2722b
--- /dev/null
+++ b/Documentation/scheduler/sched-paravirt.rst
@@ -0,0 +1,67 @@
+.. SPDX-License-Identifier: GPL-2.0
+.. _sched-paravirt:
+
+Preferred CPUs
+==============
+
+In paravirtualized environments CPU overcommit is a common scenario.
+i.e. the sum of virtual CPUs (vCPUs) of all VMs is greater than number of
+physical CPUs (pCPUs). Under such conditions when all or many VMs have
+high utilization, hypervisor won't be able to satisfy the CPU requirement
+and has to context switch within or across VMs. The hypervisor needs to
+preempt one vCPU to run another. This is called vCPU preemption.
+This is more expensive compared to task context switch within a vCPU, since
+hypervisor lacks vCPU context and could preempt a critical section which
+slows forward progress.
+
+In such cases it is better that combined vCPU demand from all VMs is reduced
+by not using some of the vCPUs in each VM. vCPUs where workload can be safely
+scheduled which won't increase any contention for pCPU are called
+"Preferred CPUs".
+
+One of the main design constructs is that preferred CPUs are always
+a subset of active CPUs. In most cases preferred CPUs will be same as
+active CPUs. When there is pCPU contention, Preferred CPUs will reduce
+based on the steal time. When the pCPU contention goes away as indicated
+by steal time, Preferred CPUs could become same as active CPUs again.
+The policy decisions are to be taken by driver.
+For example, steal_governor. Look at its documentation for more
+details. (``drivers/virt/steal_governor.c``)
+
+Scheduling decisions such as wakeup, pushing the task etc, need this
+CPU state info. This is maintained in ``cpu_preferred_mask``.
+vCPUs which are not in ``cpu_preferred_mask`` should be treated as vCPUs which
+should not be used at this moment provided it doesn't break user affinity.
+
+This is achieved by:
+
+1. Selecting a preferred CPU at wakeup using fallback mechanism.
+2. Pushing the task away from non-preferred CPU at tick.
+3. Selecting only preferred CPUs for load balance.
+
+``/sys/devices/system/cpu/preferred`` prints the current ``cpu_preferred_mask``
+in cpulist format.
+
+Notes:
+
+1. This feature is available under ``CONFIG_PREFERRED_CPU``. Driver which
+ makes decisions should enable it. For example, steal_governor driver
+ (``CONFIG_STEAL_GOVERNOR``). On enabling the driver, CPU preferred state
+ can change based on steal time. Without the driver, preferred CPUs is
+ same as active CPUs.
+
+2. This feature works for the FAIR class only.
+
+3. A pinned task, which can't be moved to preferred CPUs will continue
+ to run based on its affinity. But no load balancing happens if it is affined
+ only on non-preferred CPUs.
+
+4. Decision to change the preferred CPU state is driven by the kernel.
+ Hence it shouldn't break user affinities. One of the main reasons why
+ CPU hotplug or Isolated cpuset partitions was not a solution.
+
+5. This feature works best only when all the Guest VMs enable the feature as
+ it is a co-operative scheme. If a specific VM doesn't enable this feature
+ it may end up with more CPUs than others, still should lead to better
+ performance when seen from system view.
+ Users who enable this driver must ensure it is enabled in all Guest VMs.
--
2.52.0
^ permalink raw reply related [flat|nested] 32+ messages in thread* [PATCH v14 04/13] cpumask: Introduce cpu_preferred_mask
2026-09-28 5:37 [PATCH v14 00/13] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff Shrikanth Hegde
` (2 preceding siblings ...)
2026-09-28 5:37 ` [PATCH v14 03/13] sched/docs: Document cpu_preferred_mask and Preferred CPU concept Shrikanth Hegde
@ 2026-09-28 5:37 ` Shrikanth Hegde
2026-09-28 5:46 ` sashiko-bot
2026-09-28 5:37 ` [PATCH v14 05/13] sysfs: Add preferred CPU file Shrikanth Hegde
` (8 subsequent siblings)
12 siblings, 1 reply; 32+ messages in thread
From: Shrikanth Hegde @ 2026-09-28 5:37 UTC (permalink / raw)
To: linux-kernel, mingo, peterz, juri.lelli, vincent.guittot,
yury.norov, kprateek.nayak, iii, corbet, meted, ynorov
Cc: sshegde, tglx, gregkh, pbonzini, seanjc, vschneid, huschle,
rostedt, dietmar.eggemann, maddy, srikar, hdanton, chleroy,
vineeth, frederic, arighi, pauld, christian.loehle, tj,
tommaso.cucinotta, maz, rafael, rdunlap, kernellwp, linux-doc,
jgross, virtualization, sunlightlinux
Provide the preferred CPU infrastructure. Define get/set macros
which could be used to get/set CPU state as preferred.
CONFIG_PREFERRED_CPU will be selected by the driver which handles
steal time values. It is going to set/clear preferred CPU state.
This driver will be called steal_governor and it is introduced in
subsequent patches. It periodically computes the steal ratio and
decides on preferred CPU state.
A CPU is set to preferred when it becomes active. Later it may be
marked as non-preferred depending on steal ratio by the steal_governor.
Always maintain design construct of preferred is subset of active.
i.e. preferred ⊆ active ⊆ online ⊆ present ⊆ possible
With CONFIG_PREFERRED_CPU=n, ensure set_cpu_preferred is a nop and get
method returns the active state in that case.
Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
---
include/linux/cpumask.h | 24 ++++++++++++++++++++++++
kernel/Kconfig.preempt | 4 ++++
kernel/cpu.c | 6 ++++++
kernel/sched/core.c | 5 +++++
4 files changed, 39 insertions(+)
diff --git a/include/linux/cpumask.h b/include/linux/cpumask.h
index 7c8f16797f94..bf89bb3f30f6 100644
--- a/include/linux/cpumask.h
+++ b/include/linux/cpumask.h
@@ -121,12 +121,20 @@ extern struct cpumask __cpu_enabled_mask;
extern struct cpumask __cpu_present_mask;
extern struct cpumask __cpu_active_mask;
extern struct cpumask __cpu_dying_mask;
+
+#ifdef CONFIG_PREFERRED_CPU
+extern struct cpumask __cpu_preferred_mask;
+#else
+#define __cpu_preferred_mask __cpu_active_mask
+#endif
+
#define cpu_possible_mask ((const struct cpumask *)&__cpu_possible_mask)
#define cpu_online_mask ((const struct cpumask *)&__cpu_online_mask)
#define cpu_enabled_mask ((const struct cpumask *)&__cpu_enabled_mask)
#define cpu_present_mask ((const struct cpumask *)&__cpu_present_mask)
#define cpu_active_mask ((const struct cpumask *)&__cpu_active_mask)
#define cpu_dying_mask ((const struct cpumask *)&__cpu_dying_mask)
+#define cpu_preferred_mask ((const struct cpumask *)&__cpu_preferred_mask)
extern atomic_t __num_online_cpus;
extern unsigned int __num_possible_cpus;
@@ -1181,6 +1189,12 @@ void init_cpu_possible(const struct cpumask *src);
#define set_cpu_active(cpu, active) assign_cpu((cpu), &__cpu_active_mask, (active))
#define set_cpu_dying(cpu, dying) assign_cpu((cpu), &__cpu_dying_mask, (dying))
+#ifdef CONFIG_PREFERRED_CPU
+#define set_cpu_preferred(cpu, preferred) assign_cpu((cpu), &__cpu_preferred_mask, (preferred))
+#else
+#define set_cpu_preferred(cpu, preferred) do { } while (0)
+#endif
+
void set_cpu_online(unsigned int cpu, bool online);
void set_cpu_possible(unsigned int cpu, bool possible);
@@ -1275,6 +1289,11 @@ static __always_inline bool cpu_dying(unsigned int cpu)
return cpumask_test_cpu(cpu, cpu_dying_mask);
}
+static __always_inline bool cpu_preferred(unsigned int cpu)
+{
+ return cpumask_test_cpu(cpu, cpu_preferred_mask);
+}
+
#else
#define num_online_cpus() 1U
@@ -1313,6 +1332,11 @@ static __always_inline bool cpu_dying(unsigned int cpu)
return false;
}
+static __always_inline bool cpu_preferred(unsigned int cpu)
+{
+ return cpu == 0;
+}
+
#endif /* NR_CPUS > 1 */
#define cpu_is_offline(cpu) unlikely(!cpu_online(cpu))
diff --git a/kernel/Kconfig.preempt b/kernel/Kconfig.preempt
index 985aea617cfe..edc067a0c422 100644
--- a/kernel/Kconfig.preempt
+++ b/kernel/Kconfig.preempt
@@ -194,3 +194,7 @@ config SCHED_CLASS_EXT
For more information:
Documentation/scheduler/sched-ext.rst
https://github.com/sched-ext/scx
+
+config PREFERRED_CPU
+ bool
+ depends on SMP && PARAVIRT
diff --git a/kernel/cpu.c b/kernel/cpu.c
index b3c8553d7bd6..376d297a6292 100644
--- a/kernel/cpu.c
+++ b/kernel/cpu.c
@@ -3103,6 +3103,11 @@ EXPORT_SYMBOL(__cpu_dying_mask);
atomic_t __num_online_cpus __read_mostly;
EXPORT_SYMBOL(__num_online_cpus);
+#ifdef CONFIG_PREFERRED_CPU
+struct cpumask __cpu_preferred_mask __read_mostly;
+EXPORT_SYMBOL_GPL(__cpu_preferred_mask);
+#endif
+
void init_cpu_present(const struct cpumask *src)
{
cpumask_copy(&__cpu_present_mask, src);
@@ -3160,6 +3165,7 @@ void __init boot_cpu_init(void)
/* Mark the boot cpu "present", "online" etc for SMP and UP case */
set_cpu_online(cpu, true);
set_cpu_active(cpu, true);
+ set_cpu_preferred(cpu, true);
set_cpu_present(cpu, true);
set_cpu_possible(cpu, true);
diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index ee9b443f760d..0183c75d7d51 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -8580,6 +8580,9 @@ int sched_cpu_activate(unsigned int cpu)
*/
sched_set_rq_online(rq, cpu);
+ /* preferred is subset of active and follows its state */
+ set_cpu_preferred(cpu, true);
+
return 0;
}
@@ -8593,6 +8596,8 @@ int sched_cpu_deactivate(unsigned int cpu)
if (ret)
return ret;
+ set_cpu_preferred(cpu, false);
+
/*
* Remove CPU from nohz.idle_cpus_mask to prevent participating in
* load balancing when not active
--
2.52.0
^ permalink raw reply related [flat|nested] 32+ messages in thread* Re: [PATCH v14 04/13] cpumask: Introduce cpu_preferred_mask
2026-09-28 5:37 ` [PATCH v14 04/13] cpumask: Introduce cpu_preferred_mask Shrikanth Hegde
@ 2026-09-28 5:46 ` sashiko-bot
0 siblings, 0 replies; 32+ messages in thread
From: sashiko-bot @ 2026-09-28 5:46 UTC (permalink / raw)
To: Shrikanth Hegde; +Cc: Michael S. Tsirkin, Eugenio Perez, virtualization
> Provide the preferred CPU infrastructure. Define get/set macros
> which could be used to get/set CPU state as preferred.
>
> CONFIG_PREFERRED_CPU will be selected by the driver which handles
> steal time values. It is going to set/clear preferred CPU state.
> This driver will be called steal_governor and it is introduced in
> subsequent patches. It periodically computes the steal ratio and
> decides on preferred CPU state.
>
> A CPU is set to preferred when it becomes active. Later it may be
> marked as non-preferred depending on steal ratio by the steal_governor.
>
> Always maintain design construct of preferred is subset of active.
> i.e. preferred ⊆ active ⊆ online ⊆ present ⊆ possible
>
> With CONFIG_PREFERRED_CPU=n, ensure set_cpu_preferred is a nop and get
> method returns the active state in that case.
>
> Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
Sashiko has reviewed this patch and found no issues. It looks great!
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260928053728.797539-1-sshegde@linux.ibm.com?part=4
^ permalink raw reply [flat|nested] 32+ messages in thread
* [PATCH v14 05/13] sysfs: Add preferred CPU file
2026-09-28 5:37 [PATCH v14 00/13] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff Shrikanth Hegde
` (3 preceding siblings ...)
2026-09-28 5:37 ` [PATCH v14 04/13] cpumask: Introduce cpu_preferred_mask Shrikanth Hegde
@ 2026-09-28 5:37 ` Shrikanth Hegde
2026-09-28 5:47 ` sashiko-bot
2026-09-28 5:37 ` [PATCH v14 06/13] sched/core: Try to use a preferred CPU in is_cpu_allowed Shrikanth Hegde
` (7 subsequent siblings)
12 siblings, 1 reply; 32+ messages in thread
From: Shrikanth Hegde @ 2026-09-28 5:37 UTC (permalink / raw)
To: linux-kernel, mingo, peterz, juri.lelli, vincent.guittot,
yury.norov, kprateek.nayak, iii, corbet, meted, ynorov
Cc: sshegde, tglx, gregkh, pbonzini, seanjc, vschneid, huschle,
rostedt, dietmar.eggemann, maddy, srikar, hdanton, chleroy,
vineeth, frederic, arighi, pauld, christian.loehle, tj,
tommaso.cucinotta, maz, rafael, rdunlap, kernellwp, linux-doc,
jgross, virtualization, sunlightlinux
Add a "preferred" file in /sys/devices/system/cpu/ when kernel is
built with CONFIG_PREFERRED_CPU=y.
This would help
- Users to quickly check which CPUs are marked as preferred.
- Userspace daemons such as irqbalance to use this mask to
route irqs into preferred CPUs.
For example:
cat /sys/devices/system/cpu/online
0-719
cat /sys/devices/system/cpu/preferred
0-599 <<< Implies 0-599 are preferred for workloads and 600-719
should be avoided at this moment.
cat /sys/devices/system/cpu/preferred
0-719 <<< All CPUs are usable. There is no preference.
Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
---
Documentation/ABI/testing/sysfs-devices-system-cpu | 14 ++++++++++++++
drivers/base/cpu.c | 12 ++++++++++++
2 files changed, 26 insertions(+)
diff --git a/Documentation/ABI/testing/sysfs-devices-system-cpu b/Documentation/ABI/testing/sysfs-devices-system-cpu
index 82d10d556cc8..8f250b8693b8 100644
--- a/Documentation/ABI/testing/sysfs-devices-system-cpu
+++ b/Documentation/ABI/testing/sysfs-devices-system-cpu
@@ -806,3 +806,17 @@ Date: Nov 2022
Contact: Linux kernel mailing list <linux-kernel@vger.kernel.org>
Description:
(RO) the list of CPUs that can be brought online.
+
+What: /sys/devices/system/cpu/preferred
+Date: Sep 2026
+Contact: Linux kernel mailing list <linux-kernel@vger.kernel.org>
+Description:
+ (RO) the list of preferred CPUs applicable in
+ paravirtualized environments.
+
+ The steal governor driver dynamically adjusts this mask
+ based on observed steal time. Scheduling tasks on
+ CPUs outside of this list may lead to performance
+ degradations due to underlying physical CPU contention.
+
+ See Documentation/scheduler/sched-paravirt.rst for more details.
diff --git a/drivers/base/cpu.c b/drivers/base/cpu.c
index 69e52fed4241..747915ff974f 100644
--- a/drivers/base/cpu.c
+++ b/drivers/base/cpu.c
@@ -391,6 +391,15 @@ static int cpu_uevent(const struct device *dev, struct kobj_uevent_env *env)
}
#endif
+#ifdef CONFIG_PREFERRED_CPU
+static ssize_t preferred_show(struct device *dev,
+ struct device_attribute *attr, char *buf)
+{
+ return sysfs_emit(buf, "%*pbl\n", cpumask_pr_args(cpu_preferred_mask));
+}
+static DEVICE_ATTR_RO(preferred);
+#endif
+
const struct bus_type cpu_subsys = {
.name = "cpu",
.dev_name = "cpu",
@@ -531,6 +540,9 @@ static struct attribute *cpu_root_attrs[] = {
#endif
#ifdef CONFIG_GENERIC_CPU_AUTOPROBE
&dev_attr_modalias.attr,
+#endif
+#ifdef CONFIG_PREFERRED_CPU
+ &dev_attr_preferred.attr,
#endif
NULL
};
--
2.52.0
^ permalink raw reply related [flat|nested] 32+ messages in thread* Re: [PATCH v14 05/13] sysfs: Add preferred CPU file
2026-09-28 5:37 ` [PATCH v14 05/13] sysfs: Add preferred CPU file Shrikanth Hegde
@ 2026-09-28 5:47 ` sashiko-bot
0 siblings, 0 replies; 32+ messages in thread
From: sashiko-bot @ 2026-09-28 5:47 UTC (permalink / raw)
To: Shrikanth Hegde; +Cc: Michael S. Tsirkin, Eugenio Perez, virtualization
> Add a "preferred" file in /sys/devices/system/cpu/ when kernel is
> built with CONFIG_PREFERRED_CPU=y.
>
> This would help
> - Users to quickly check which CPUs are marked as preferred.
> - Userspace daemons such as irqbalance to use this mask to
> route irqs into preferred CPUs.
>
> For example:
> cat /sys/devices/system/cpu/online
> 0-719
> cat /sys/devices/system/cpu/preferred
> 0-599 <<< Implies 0-599 are preferred for workloads and 600-719
> should be avoided at this moment.
>
> cat /sys/devices/system/cpu/preferred
> 0-719 <<< All CPUs are usable. There is no preference.
>
> Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
Sashiko has reviewed this patch and found no issues. It looks great!
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260928053728.797539-1-sshegde@linux.ibm.com?part=5
^ permalink raw reply [flat|nested] 32+ messages in thread
* [PATCH v14 06/13] sched/core: Try to use a preferred CPU in is_cpu_allowed
2026-09-28 5:37 [PATCH v14 00/13] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff Shrikanth Hegde
` (4 preceding siblings ...)
2026-09-28 5:37 ` [PATCH v14 05/13] sysfs: Add preferred CPU file Shrikanth Hegde
@ 2026-09-28 5:37 ` Shrikanth Hegde
2026-09-28 6:00 ` sashiko-bot
2026-09-28 5:37 ` [PATCH v14 07/13] sched/fair: Load balance only among preferred CPUs Shrikanth Hegde
` (6 subsequent siblings)
12 siblings, 1 reply; 32+ messages in thread
From: Shrikanth Hegde @ 2026-09-28 5:37 UTC (permalink / raw)
To: linux-kernel, mingo, peterz, juri.lelli, vincent.guittot,
yury.norov, kprateek.nayak, iii, corbet, meted, ynorov
Cc: sshegde, tglx, gregkh, pbonzini, seanjc, vschneid, huschle,
rostedt, dietmar.eggemann, maddy, srikar, hdanton, chleroy,
vineeth, frederic, arighi, pauld, christian.loehle, tj,
tommaso.cucinotta, maz, rafael, rdunlap, kernellwp, linux-doc,
jgross, virtualization, sunlightlinux
When possible, try to choose a preferred CPU.
This is essential to maintain user affinities when preferred
CPUs change. A task pinned on a non-preferred CPU should continue
to run there, since this is a non-user triggered event.
If a CPU is non-preferred and the task can run on other CPUs which are
currently preferred, then choose a preferred CPU instead.
This is decided by checking if cpus_ptr, cpu_preferred_mask and
task possible CPUs intersect or not. If yes, then the task has
other preferred CPUs. This takes care of tasks with architecture-specific
CPU masks (e.g., 32-bit tasks on arm64).
The push task mechanism uses a stopper thread which calls
select_fallback_rq() and uses this mechanism to pick a preferred CPU.
This takes care of the wakeup path for FAIR tasks too.
is_cpu_allowed() is called to ensure wakeups happen on preferred CPUs.
With that, additional checks in available_idle_cpu() are not necessary.
Ignore the preferred CPU state if a task's affinity is changing and
its new mask no longer includes the CPU it is currently running on.
This ensures migration_cpu_stop() does not abort, preventing the task
from being stranded outside its allowed affinity.
For the majority of cases, this would still keep select_fallback_rq()
as O(N). cpumask_intersects_and(), which is O(N), is called only if
!cpu_preferred. The task running there is expected to move out.
Subsequently, it should run on a preferred CPU. This becomes O(N**2)
only for tasks pinned solely to non-preferred CPUs. That is a rare case.
Overhead is minimal when the CPU is preferred.
Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
---
kernel/sched/core.c | 30 ++++++++++++++++++++++++++++--
1 file changed, 28 insertions(+), 2 deletions(-)
diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index 0183c75d7d51..04400f934cc7 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -2504,6 +2504,24 @@ static inline bool rq_has_pinned_tasks(struct rq *rq)
return rq->nr_pinned;
}
+static inline bool task_can_migrate_to_preferred(struct task_struct *p, int cpu)
+{
+ /* No need to migrate from a preferred CPU */
+ if (cpu_preferred(cpu))
+ return false;
+
+ /* Only FAIR tasks honor preferred CPU state */
+ if (unlikely(p->sched_class != &fair_sched_class))
+ return false;
+
+ /* Ignore preferred state if task affinity is changing */
+ if (unlikely(!cpumask_test_cpu(task_cpu(p), p->cpus_ptr)))
+ return false;
+
+ return cpumask_intersects_and(p->cpus_ptr, cpu_preferred_mask,
+ task_cpu_possible_mask(p));
+}
+
/*
* Per-CPU kthreads are allowed to run on !active && online CPUs, see
* __set_cpus_allowed_ptr() and select_fallback_rq().
@@ -2519,8 +2537,12 @@ static inline bool is_cpu_allowed(struct task_struct *p, int cpu)
return cpu_online(cpu);
/* Non kernel threads are not allowed during either online or offline. */
- if (!(p->flags & PF_KTHREAD))
+ if (!(p->flags & PF_KTHREAD)) {
+ /* Try to use preferred CPU if task's affinity allows */
+ if (task_can_migrate_to_preferred(p, cpu))
+ return false;
return cpu_active(cpu);
+ }
/* KTHREAD_IS_PER_CPU is always allowed. */
if (kthread_is_per_cpu(p))
@@ -2530,7 +2552,11 @@ static inline bool is_cpu_allowed(struct task_struct *p, int cpu)
if (cpu_dying(cpu))
return false;
- /* But are allowed during online. */
+ /* Try to keep unbound kthreads on a preferred CPU if possible. */
+ if (task_can_migrate_to_preferred(p, cpu))
+ return false;
+
+ /* Otherwise, they are allowed to run on online CPU. */
return cpu_online(cpu);
}
--
2.52.0
^ permalink raw reply related [flat|nested] 32+ messages in thread* Re: [PATCH v14 06/13] sched/core: Try to use a preferred CPU in is_cpu_allowed
2026-09-28 5:37 ` [PATCH v14 06/13] sched/core: Try to use a preferred CPU in is_cpu_allowed Shrikanth Hegde
@ 2026-09-28 6:00 ` sashiko-bot
2026-09-28 6:48 ` Shrikanth Hegde
0 siblings, 1 reply; 32+ messages in thread
From: sashiko-bot @ 2026-09-28 6:00 UTC (permalink / raw)
To: Shrikanth Hegde; +Cc: Eugenio Perez, virtualization, Michael S. Tsirkin
Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
- [High] sched/core: O(N^2) wakeup latency and task clumping in fallback path
- [High] sched/fair: Active load balancing to non-preferred CPUs silently fails
- [High] sched/core: TOCTOU race with CPU hotplug breaks task affinity
--
commit dfa0f159ec643e02807bc453bbfcc49c9f1e7915
Author: Shrikanth Hegde <sshegde@linux.ibm.com>
sched/core: Try to use a preferred CPU in is_cpu_allowed
This commit modifies the scheduler to preferentially select preferred CPUs
when checking if a task is allowed to run on a given CPU. It introduces a
helper function to check if a task can migrate to a preferred CPU based on
its affinity mask.
> For the majority of cases, this would still keep select_fallback_rq()
> as O(N). cpumask_intersects_and(), which is O(N), is called only if
> !cpu_preferred. The task running there is expected to move out.
> Subsequently, it should run on a preferred CPU. This becomes O(N**2)
> only for tasks pinned solely to non-preferred CPUs. That is a rare case.
Is this description accurate?
Because the fallback loop iterates over the affinity mask sequentially, if
the allowed preferred CPUs happen to have higher indices than the
non-preferred CPUs, the loop will perform the intersection operation for
every non-preferred CPU it checks. This means any task with preferred CPUs
at higher indices could experience the quadratic latency overhead, rather
than only tasks pinned solely to non-preferred CPUs.
> diff --git a/kernel/sched/core.c b/kernel/sched/core.c
> index 0183c75d7d518..04400f934cc72 100644
> --- a/kernel/sched/core.c
> +++ b/kernel/sched/core.c
[ ... ]
> @@ -2504,6 +2504,24 @@ static inline bool rq_has_pinned_tasks(struct rq *rq)
> return rq->nr_pinned;
> }
>
> +static inline bool task_can_migrate_to_preferred(struct task_struct *p, int cpu)
> +{
> + /* No need to migrate from a preferred CPU */
> + if (cpu_preferred(cpu))
> + return false;
> +
> + /* Only FAIR tasks honor preferred CPU state */
> + if (unlikely(p->sched_class != &fair_sched_class))
> + return false;
> +
> + /* Ignore preferred state if task affinity is changing */
> + if (unlikely(!cpumask_test_cpu(task_cpu(p), p->cpus_ptr)))
> + return false;
> +
> + return cpumask_intersects_and(p->cpus_ptr, cpu_preferred_mask,
> + task_cpu_possible_mask(p));
> +}
[Severity: High]
Does this intersection operation lead to quadratic latency during wakeups
when combined with the fallback loop?
When select_task_rq() in kernel/sched/core.c picks a non-preferred CPU
because it is idle, is_cpu_allowed() will reject it, forcing the task into
select_fallback_rq(). The fallback loop iterates sequentially over all CPUs:
kernel/sched/core.c:select_fallback_rq() {
...
for_each_cpu(dest_cpu, p->cpus_ptr) {
if (!is_cpu_allowed(p, dest_cpu))
continue;
...
goto out;
}
...
}
If the allowed preferred CPUs are at the end of the cpumask, wouldn't this
perform an intersection for every non-preferred CPU it checks, leading to
severe scheduling latency? Additionally, wouldn't this deterministic loop
always dump the task on the first preferred CPU it finds, destroying load
distribution by clumping tasks on the lowest-ID preferred CPU?
This performance degradation from wakeups on non-preferred CPUs does not
appear to be resolved in the final state of the patch series.
[ ... ]
> @@ -2519,8 +2537,12 @@ static inline bool is_cpu_allowed(struct task_struct *p, int cpu)
> return cpu_online(cpu);
>
> /* Non kernel threads are not allowed during either online or offline. */
> - if (!(p->flags & PF_KTHREAD))
> + if (!(p->flags & PF_KTHREAD)) {
> + /* Try to use preferred CPU if task's affinity allows */
> + if (task_can_migrate_to_preferred(p, cpu))
> + return false;
[Severity: High]
Could returning false for non-preferred CPUs here cause active load
balancing to silently fail?
During routine background load balancing, if the system attempts to pull
load from a preferred CPU to an idle non-preferred CPU, the stopper thread
will run __migrate_task().
kernel/sched/core.c:__migrate_task() {
...
if (!is_cpu_allowed(p, dest_cpu))
return rq;
...
}
Since the destination is non-preferred, won't is_cpu_allowed() reject it,
causing the stopper thread to silently abort the migration and break load
distribution?
While follow-up patches in the series restrict the busiest CPU to preferred
CPUs, they do not appear to prevent the destination from being non-preferred,
leaving this active load balancing path broken.
[Severity: High]
Can this rejection introduce a race condition with CPU hotplug that breaks
a task's user-configured affinity?
In select_fallback_rq() in kernel/sched/core.c, the loop might skip an
active, non-preferred CPU because task_can_migrate_to_preferred() indicates
another preferred CPU is available in the affinity mask. If a concurrent
CPU hotplug operation removes that preferred CPU from the active mask before
the loop reaches it, the loop will skip it too:
kernel/sched/core.c:select_fallback_rq() {
...
for_each_cpu(dest_cpu, p->cpus_ptr) {
if (!is_cpu_allowed(p, dest_cpu))
continue;
...
}
Could the loop then terminate without finding any allowed CPU, permanently
forcing the task outside of its allowed affinity?
> return cpu_active(cpu);
> + }
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260928053728.797539-1-sshegde@linux.ibm.com?part=6
^ permalink raw reply [flat|nested] 32+ messages in thread* Re: [PATCH v14 06/13] sched/core: Try to use a preferred CPU in is_cpu_allowed
2026-09-28 6:00 ` sashiko-bot
@ 2026-09-28 6:48 ` Shrikanth Hegde
0 siblings, 0 replies; 32+ messages in thread
From: Shrikanth Hegde @ 2026-09-28 6:48 UTC (permalink / raw)
To: sashiko-reviews; +Cc: Eugenio Perez, virtualization, Michael S. Tsirkin
I believe most of the comments have been addressed in the previous rounds.
On 9/28/26 11:30 AM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
> - [High] sched/core: O(N^2) wakeup latency and task clumping in fallback path
> - [High] sched/fair: Active load balancing to non-preferred CPUs silently fails
> - [High] sched/core: TOCTOU race with CPU hotplug breaks task affinity
> --
>
> commit dfa0f159ec643e02807bc453bbfcc49c9f1e7915
> Author: Shrikanth Hegde <sshegde@linux.ibm.com>
>
> sched/core: Try to use a preferred CPU in is_cpu_allowed
>
> This commit modifies the scheduler to preferentially select preferred CPUs
> when checking if a task is allowed to run on a given CPU. It introduces a
> helper function to check if a task can migrate to a preferred CPU based on
> its affinity mask.
>
>> For the majority of cases, this would still keep select_fallback_rq()
>> as O(N). cpumask_intersects_and(), which is O(N), is called only if
>> !cpu_preferred. The task running there is expected to move out.
>> Subsequently, it should run on a preferred CPU. This becomes O(N**2)
>> only for tasks pinned solely to non-preferred CPUs. That is a rare case.
>
> Is this description accurate?
>
> Because the fallback loop iterates over the affinity mask sequentially, if
> the allowed preferred CPUs happen to have higher indices than the
> non-preferred CPUs, the loop will perform the intersection operation for
> every non-preferred CPU it checks. This means any task with preferred CPUs
> at higher indices could experience the quadratic latency overhead, rather
> than only tasks pinned solely to non-preferred CPUs.
>
>> diff --git a/kernel/sched/core.c b/kernel/sched/core.c
>> index 0183c75d7d518..04400f934cc72 100644
>> --- a/kernel/sched/core.c
>> +++ b/kernel/sched/core.c
> [ ... ]
>> @@ -2504,6 +2504,24 @@ static inline bool rq_has_pinned_tasks(struct rq *rq)
>> return rq->nr_pinned;
>> }
>>
>> +static inline bool task_can_migrate_to_preferred(struct task_struct *p, int cpu)
>> +{
>> + /* No need to migrate from a preferred CPU */
>> + if (cpu_preferred(cpu))
>> + return false;
>> +
>> + /* Only FAIR tasks honor preferred CPU state */
>> + if (unlikely(p->sched_class != &fair_sched_class))
>> + return false;
>> +
>> + /* Ignore preferred state if task affinity is changing */
>> + if (unlikely(!cpumask_test_cpu(task_cpu(p), p->cpus_ptr)))
>> + return false;
>> +
>> + return cpumask_intersects_and(p->cpus_ptr, cpu_preferred_mask,
>> + task_cpu_possible_mask(p));
>> +}
>
> [Severity: High]
> Does this intersection operation lead to quadratic latency during wakeups
> when combined with the fallback loop?
>
> When select_task_rq() in kernel/sched/core.c picks a non-preferred CPU
> because it is idle, is_cpu_allowed() will reject it, forcing the task into
> select_fallback_rq(). The fallback loop iterates sequentially over all CPUs:
>
> kernel/sched/core.c:select_fallback_rq() {
> ...
> for_each_cpu(dest_cpu, p->cpus_ptr) {
> if (!is_cpu_allowed(p, dest_cpu))
> continue;
> ...
> goto out;
> }
> ...
> }
>
> If the allowed preferred CPUs are at the end of the cpumask, wouldn't this
> perform an intersection for every non-preferred CPU it checks, leading to
> severe scheduling latency? Additionally, wouldn't this deterministic loop
> always dump the task on the first preferred CPU it finds, destroying load
> distribution by clumping tasks on the lowest-ID preferred CPU?
>
> This performance degradation from wakeups on non-preferred CPUs does not
> appear to be resolved in the final state of the patch series.
It is mentioned in the changelog and this is true only when wakeup happens on
a non-preferred CPUs. Even if it runs there is meant to pushed out. So take a
bit more cycles and do the right thing.
>
> [ ... ]
>> @@ -2519,8 +2537,12 @@ static inline bool is_cpu_allowed(struct task_struct *p, int cpu)
>> return cpu_online(cpu);
>>
>> /* Non kernel threads are not allowed during either online or offline. */
>> - if (!(p->flags & PF_KTHREAD))
>> + if (!(p->flags & PF_KTHREAD)) {
>> + /* Try to use preferred CPU if task's affinity allows */
>> + if (task_can_migrate_to_preferred(p, cpu))
>> + return false;
>
> [Severity: High]
> Could returning false for non-preferred CPUs here cause active load
> balancing to silently fail?
Spreading load to non-preferred CPUs is by design and expected.
>
> During routine background load balancing, if the system attempts to pull
> load from a preferred CPU to an idle non-preferred CPU, the stopper thread
> will run __migrate_task().
>
> kernel/sched/core.c:__migrate_task() {
> ...
> if (!is_cpu_allowed(p, dest_cpu))
> return rq;
> ...
> }
>
> Since the destination is non-preferred, won't is_cpu_allowed() reject it,
> causing the stopper thread to silently abort the migration and break load
> distribution?
>
> While follow-up patches in the series restrict the busiest CPU to preferred
> CPUs, they do not appear to prevent the destination from being non-preferred,
> leaving this active load balancing path broken.
>
> [Severity: High]
> Can this rejection introduce a race condition with CPU hotplug that breaks
> a task's user-configured affinity?
>
> In select_fallback_rq() in kernel/sched/core.c, the loop might skip an
> active, non-preferred CPU because task_can_migrate_to_preferred() indicates
> another preferred CPU is available in the affinity mask. If a concurrent
> CPU hotplug operation removes that preferred CPU from the active mask before
> the loop reaches it, the loop will skip it too:
There is two iterations and steal governor's minimal interval is 100ms.
So it is likely safe.
>
> kernel/sched/core.c:select_fallback_rq() {
> ...
> for_each_cpu(dest_cpu, p->cpus_ptr) {
> if (!is_cpu_allowed(p, dest_cpu))
> continue;
> ...
> }
>
> Could the loop then terminate without finding any allowed CPU, permanently
> forcing the task outside of its allowed affinity?
>
>> return cpu_active(cpu);
>> + }
>
^ permalink raw reply [flat|nested] 32+ messages in thread
* [PATCH v14 07/13] sched/fair: Load balance only among preferred CPUs
2026-09-28 5:37 [PATCH v14 00/13] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff Shrikanth Hegde
` (5 preceding siblings ...)
2026-09-28 5:37 ` [PATCH v14 06/13] sched/core: Try to use a preferred CPU in is_cpu_allowed Shrikanth Hegde
@ 2026-09-28 5:37 ` Shrikanth Hegde
2026-09-28 5:55 ` sashiko-bot
2026-09-28 5:37 ` [PATCH v14 08/13] sched/core: Push current task from non preferred CPU Shrikanth Hegde
` (5 subsequent siblings)
12 siblings, 1 reply; 32+ messages in thread
From: Shrikanth Hegde @ 2026-09-28 5:37 UTC (permalink / raw)
To: linux-kernel, mingo, peterz, juri.lelli, vincent.guittot,
yury.norov, kprateek.nayak, iii, corbet, meted, ynorov
Cc: sshegde, tglx, gregkh, pbonzini, seanjc, vschneid, huschle,
rostedt, dietmar.eggemann, maddy, srikar, hdanton, chleroy,
vineeth, frederic, arighi, pauld, christian.loehle, tj,
tommaso.cucinotta, maz, rafael, rdunlap, kernellwp, linux-doc,
jgross, virtualization, sunlightlinux
When a CPU is marked as non-preferred, any load pulled towards it is
pointless since the task will be pushed out again in the next tick.
So, consider only preferred CPUs for load balancing.
This ensures load balancing does not fight against the push task mechanism
which happens at the tick. Also, this stops active balancing from happening
on a non-preferred CPU pulling the load.
This also means there is no load balancing if a task is pinned only to
non-preferred CPUs. They will continue to run where they were previously
running before the CPUs were marked as non-preferred.
Bail out early for NEWIDLE balancing, as load balancing is done only on
preferred CPUs. Note that idle balancing is allowed to go through, since
that naturally updates nohz.next_balance when all the idle CPUs are
non-preferred.
Also, optimization in find_new_ilb() is skipped. The steal governor driver,
which is introduced in later patches, updates the preferred CPUs state in
descending order. find_new_ilb() checks for idle CPUs in ascending order.
Hence, in most common scenarios, the idle CPU found by find_new_ilb() will
already be a preferred CPU. When all idle CPUs are non-preferred, the first
idle CPU has to be chosen anyway. All of this is naturally handled in
find_new_ilb() currently. Adding additional complexity to it for rare
edge cases is not necessary.
Reviewed-by: Yury Norov <ynorov@nvidia.com>
Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
---
kernel/sched/fair.c | 8 +++-----
1 file changed, 3 insertions(+), 5 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 03206e15e6fe..1c687c3c70f1 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -13467,7 +13467,7 @@ static int sched_balance_rq(int this_cpu, struct rq *this_rq,
};
bool need_unlock = false;
- cpumask_and(cpus, sched_domain_span(sd), cpu_active_mask);
+ cpumask_and(cpus, sched_domain_span(sd), cpu_preferred_mask);
schedstat_inc(sd->lb_count[idle]);
@@ -14582,10 +14582,8 @@ static int sched_balance_newidle(struct rq *this_rq, struct rq_flags *rf)
*/
this_rq->idle_stamp = rq_clock(this_rq);
- /*
- * Do not pull tasks towards !active CPUs...
- */
- if (!cpu_active(this_cpu))
+ /* Do not pull tasks towards !preferred CPUs */
+ if (!cpu_preferred(this_cpu))
return 0;
/*
--
2.52.0
^ permalink raw reply related [flat|nested] 32+ messages in thread* Re: [PATCH v14 07/13] sched/fair: Load balance only among preferred CPUs
2026-09-28 5:37 ` [PATCH v14 07/13] sched/fair: Load balance only among preferred CPUs Shrikanth Hegde
@ 2026-09-28 5:55 ` sashiko-bot
0 siblings, 0 replies; 32+ messages in thread
From: sashiko-bot @ 2026-09-28 5:55 UTC (permalink / raw)
To: Shrikanth Hegde; +Cc: Michael S. Tsirkin, Eugenio Perez, virtualization
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] sched: capacity mismatch in load balance due to restricted env->cpus
--
commit 614fa1366931a4c3a90955d4b7c85dd2e0e3bb49
Author: Shrikanth Hegde <sshegde@linux.ibm.com>
Subject: sched/fair: Load balance only among preferred CPUs
This commit restricts load balancing to preferred CPUs to avoid conflicting
with the tick's push task mechanism. It replaces cpu_active_mask with
cpu_preferred_mask during load balance CPU mask initialization and bails
out early in newidle balancing if the CPU is not preferred.
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 03206e15e6fe4..1c687c3c70f11 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -13467,7 +13467,7 @@ static int sched_balance_rq(int this_cpu, struct rq *this_rq,
> };
> bool need_unlock = false;
>
> - cpumask_and(cpus, sched_domain_span(sd), cpu_active_mask);
> + cpumask_and(cpus, sched_domain_span(sd), cpu_preferred_mask);
[Severity: High]
Will this change create a capacity mismatch in the load balancer?
When env->cpus is restricted to cpu_preferred_mask here, update_sg_lb_stats()
will only aggregate group load for those preferred CPUs:
kernel/sched/fair.c:update_sg_lb_stats() {
...
for_each_cpu_and(i, sched_group_span(group), env->cpus) {
struct rq *rq = cpu_rq(i);
unsigned long load = cpu_load(rq);
sgs->group_load += load;
...
}
However, the group's capacity is set using the full capacity of all active
CPUs:
kernel/sched/fair.c:update_sg_lb_stats() {
...
sgs->group_capacity = group->sgc->capacity;
...
}
Could the mismatch between the deflated group_load and the full group_capacity
artificially lower sgs->avg_load?
This phantom spare capacity might cause the load balancer to aggressively pull
tasks from other groups into the preferred CPUs of the destination group,
overloading them.
Can this also prevent tasks on non-preferred CPUs from ever being load
balanced?
kernel/sched/fair.c:sched_balance_find_src_rq() {
...
for_each_cpu_and(i, sched_group_span(group), env->cpus) {
...
}
Because this loop uses env->cpus, the load balancer ignores non-preferred
CPUs entirely, even if they are overloaded.
[ ... ]
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260928053728.797539-1-sshegde@linux.ibm.com?part=7
^ permalink raw reply [flat|nested] 32+ messages in thread
* [PATCH v14 08/13] sched/core: Push current task from non preferred CPU
2026-09-28 5:37 [PATCH v14 00/13] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff Shrikanth Hegde
` (6 preceding siblings ...)
2026-09-28 5:37 ` [PATCH v14 07/13] sched/fair: Load balance only among preferred CPUs Shrikanth Hegde
@ 2026-09-28 5:37 ` Shrikanth Hegde
2026-09-28 5:56 ` sashiko-bot
2026-09-28 5:37 ` [PATCH v14 09/13] sched/debug: Add migration stats due to non preferred CPUs Shrikanth Hegde
` (4 subsequent siblings)
12 siblings, 1 reply; 32+ messages in thread
From: Shrikanth Hegde @ 2026-09-28 5:37 UTC (permalink / raw)
To: linux-kernel, mingo, peterz, juri.lelli, vincent.guittot,
yury.norov, kprateek.nayak, iii, corbet, meted, ynorov
Cc: sshegde, tglx, gregkh, pbonzini, seanjc, vschneid, huschle,
rostedt, dietmar.eggemann, maddy, srikar, hdanton, chleroy,
vineeth, frederic, arighi, pauld, christian.loehle, tj,
tommaso.cucinotta, maz, rafael, rdunlap, kernellwp, linux-doc,
jgross, virtualization, sunlightlinux
Actively push out the current running task on a non-preferred CPU (NPC).
Since the task is currently running, a stopper thread must be queued
to push the task out. However, if the task is pinned only to
non-preferred CPUs, it will continue running there.
This helps to maintain userspace affinities, unlike CPU hotplug
or isolated cpusets.
The implementation follows the structure of __balance_push_cpu_stop(),
but is kept separate to handle the preferred-CPU specific conditions and
pending-work state under CONFIG_PREFERRED_CPU.
Add the npc_push_work_pending flag to protect the work buffer.
For now, only the currently running task is pushed out. This keeps the code
simpler. In the future, an optimization may be added to move all queued
tasks on the runqueue.
This works only for the FAIR scheduling class.
Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
---
kernel/sched/core.c | 83 ++++++++++++++++++++++++++++++++++++++++++++
kernel/sched/sched.h | 9 +++++
2 files changed, 92 insertions(+)
diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index 04400f934cc7..5049eff58fb7 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -5809,6 +5809,9 @@ void sched_tick(void)
unsigned long hw_pressure;
u64 resched_latency;
+ if (!cpu_preferred(cpu))
+ sched_push_current_non_preferred_cpu(rq);
+
if (housekeeping_cpu(cpu, HK_TYPE_KERNEL_NOISE))
arch_scale_freq_tick();
@@ -11201,3 +11204,83 @@ void sched_change_end(struct sched_change_ctx *ctx)
p->sched_class->prio_changed(rq, p, ctx->prio);
}
}
+
+#ifdef CONFIG_PREFERRED_CPU
+static DEFINE_PER_CPU(struct cpu_stop_work, npc_push_task_work);
+
+static int sched_non_preferred_cpu_push_stop(void *arg)
+{
+ struct task_struct *p = arg;
+ struct rq *rq = this_rq();
+ struct rq_flags rf;
+ int cpu;
+
+ if (cpu_preferred(rq->cpu)) {
+ scoped_guard(rq_lock_irqsave, rq)
+ rq->npc_push_work_pending = false;
+ put_task_struct(p);
+ return 0;
+ }
+
+ scoped_guard (raw_spinlock_irq, &p->pi_lock) {
+ /*
+ * select_fallback_rq() may acquire the rq lock in case of
+ * fallback. So call it before grabbing rq lock. If the task
+ * migrates to another CPU before the rq lock is acquired,
+ * subsequent validation of task's current rq will help to
+ * safely bail out.
+ */
+ cpu = select_fallback_rq(rq->cpu, p);
+ rq_lock(rq, &rf);
+ rq->npc_push_work_pending = false;
+ update_rq_clock(rq);
+ context_unsafe_alias(rq);
+
+ if (task_rq(p) == rq && task_on_rq_queued(p))
+ rq = __migrate_task(rq, &rf, p, cpu);
+ rq_unlock(rq, &rf);
+ }
+
+ put_task_struct(p);
+ return 0;
+}
+
+/*
+ * Push the current task running on non-preferred CPU(npc).
+ * Using this non preferred CPU will lead to more contention
+ * in the host. So it is better not to use this CPU.
+ *
+ * Since task is running, call a stopper to push the task out. This is
+ * similar to how task moves during hotplug. In select_fallback_rq() a
+ * preferred CPU will be chosen and henceforth task shouldn't come back to
+ * this CPU again.
+ *
+ * Works for FAIR class only.
+ *
+ * If task is affined only on non-preferred CPUs, no point in moving it out.
+ */
+void sched_push_current_non_preferred_cpu(struct rq *rq)
+{
+ struct task_struct *push_task = rq->curr;
+
+ scoped_guard(rq_lock, rq) {
+ /* Push the task if its explicit affinity allows */
+ if (!task_can_migrate_to_preferred(push_task, rq->cpu))
+ return;
+
+ /* There is already a stopper thread. Don't race with it. */
+ if (rq->npc_push_work_pending)
+ return;
+
+ if (is_migration_disabled(push_task))
+ return;
+
+ rq->npc_push_work_pending = true;
+ }
+
+ /* sched_tick runs with interrupts disabled. */
+ get_task_struct(push_task);
+ stop_one_cpu_nowait(rq->cpu, sched_non_preferred_cpu_push_stop,
+ push_task, this_cpu_ptr(&npc_push_task_work));
+}
+#endif
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index b98084e1f5b0..ee482bb12a66 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -1326,6 +1326,9 @@ struct rq {
#ifdef CONFIG_PARAVIRT_TIME_ACCOUNTING
u64 prev_steal_time_rq;
#endif
+#ifdef CONFIG_PREFERRED_CPU
+ bool npc_push_work_pending;
+#endif
/* calc_load related fields */
unsigned long calc_load_update;
@@ -4292,4 +4295,10 @@ DEFINE_CLASS_IS_UNCONDITIONAL(sched_change)
#include "ext/ext.h"
+#ifdef CONFIG_PREFERRED_CPU
+void sched_push_current_non_preferred_cpu(struct rq *rq);
+#else /* !CONFIG_PREFERRED_CPU */
+static inline void sched_push_current_non_preferred_cpu(struct rq *rq) { }
+#endif
+
#endif /* _KERNEL_SCHED_SCHED_H */
--
2.52.0
^ permalink raw reply related [flat|nested] 32+ messages in thread* Re: [PATCH v14 08/13] sched/core: Push current task from non preferred CPU
2026-09-28 5:37 ` [PATCH v14 08/13] sched/core: Push current task from non preferred CPU Shrikanth Hegde
@ 2026-09-28 5:56 ` sashiko-bot
0 siblings, 0 replies; 32+ messages in thread
From: sashiko-bot @ 2026-09-28 5:56 UTC (permalink / raw)
To: Shrikanth Hegde; +Cc: Eugenio Perez, virtualization, Michael S. Tsirkin
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] sched/core: FAIR kthreads cause stopper thread livelock on non-preferred CPUs
--
commit a6718dc77694942ae8b63a9eb25ace9c01b3c5ef
Author: Shrikanth Hegde <sshegde@linux.ibm.com>
sched/core: Push current task from non preferred CPU
This patch actively pushes the currently running task off a non-preferred
CPU using a stopper thread, specifically targeting tasks in the FAIR
scheduling class. It adds the npc_push_work_pending flag to avoid queueing
multiple stopper threads.
> diff --git a/kernel/sched/core.c b/kernel/sched/core.c
> --- a/kernel/sched/core.c
> +++ b/kernel/sched/core.c
[ ... ]
> +static int sched_non_preferred_cpu_push_stop(void *arg)
> +{
> + struct task_struct *p = arg;
> + struct rq *rq = this_rq();
> + struct rq_flags rf;
> + int cpu;
> +
> + if (cpu_preferred(rq->cpu)) {
> + scoped_guard(rq_lock_irqsave, rq)
> + rq->npc_push_work_pending = false;
> + put_task_struct(p);
> + return 0;
> + }
> +
> + scoped_guard (raw_spinlock_irq, &p->pi_lock) {
> + /*
> + * select_fallback_rq() may acquire the rq lock in case of
> + * fallback. So call it before grabbing rq lock. If the task
> + * migrates to another CPU before the rq lock is acquired,
> + * subsequent validation of task's current rq will help to
> + * safely bail out.
> + */
> + cpu = select_fallback_rq(rq->cpu, p);
[ ... ]
> +void sched_push_current_non_preferred_cpu(struct rq *rq)
> +{
> + struct task_struct *push_task = rq->curr;
> +
> + scoped_guard(rq_lock, rq) {
> + /* Push the task if its explicit affinity allows */
> + if (!task_can_migrate_to_preferred(push_task, rq->cpu))
> + return;
[Severity: High]
Can this result in a livelock with endless stopper thread preemptions when
the task being pushed is a FAIR kthread (such as an unbound kworker)?
When task_can_migrate_to_preferred() returns true for a FAIR kthread,
this function successfully queues a stopper thread. However, when the
stopper thread runs sched_non_preferred_cpu_push_stop() and calls
select_fallback_rq(), it ultimately evaluates is_cpu_allowed(), which
skips the preferred-CPU check for kthreads:
kernel/sched/core.c:is_cpu_allowed() {
...
/* Non kernel threads are not allowed during either online or offline. */
if (!(p->flags & PF_KTHREAD)) {
/* Try to use preferred CPU if task's affinity allows */
if (task_can_migrate_to_preferred(p, cpu))
return false;
return cpu_active(cpu);
}
...
}
Because the kthread bypasses the preferred-CPU enforcement,
select_fallback_rq() can pick the first online CPU in the local node,
which might be the exact same non-preferred CPU we are trying to push it
away from.
If that happens, the migration becomes a no-op, npc_push_work_pending gets
cleared, and sched_tick() will just restart the same sequence on the next
tick, wasting CPU cycles in an endless loop.
> +
> + /* There is already a stopper thread. Don't race with it. */
> + if (rq->npc_push_work_pending)
> + return;
[ ... ]
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260928053728.797539-1-sshegde@linux.ibm.com?part=8
^ permalink raw reply [flat|nested] 32+ messages in thread
* [PATCH v14 09/13] sched/debug: Add migration stats due to non preferred CPUs
2026-09-28 5:37 [PATCH v14 00/13] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff Shrikanth Hegde
` (7 preceding siblings ...)
2026-09-28 5:37 ` [PATCH v14 08/13] sched/core: Push current task from non preferred CPU Shrikanth Hegde
@ 2026-09-28 5:37 ` Shrikanth Hegde
2026-09-28 5:48 ` sashiko-bot
2026-09-29 12:18 ` Nathan Chancellor
2026-09-28 5:37 ` [PATCH v14 10/13] virt: Introduce steal governor driver Shrikanth Hegde
` (3 subsequent siblings)
12 siblings, 2 replies; 32+ messages in thread
From: Shrikanth Hegde @ 2026-09-28 5:37 UTC (permalink / raw)
To: linux-kernel, mingo, peterz, juri.lelli, vincent.guittot,
yury.norov, kprateek.nayak, iii, corbet, meted, ynorov
Cc: sshegde, tglx, gregkh, pbonzini, seanjc, vschneid, huschle,
rostedt, dietmar.eggemann, maddy, srikar, hdanton, chleroy,
vineeth, frederic, arighi, pauld, christian.loehle, tj,
tommaso.cucinotta, maz, rafael, rdunlap, kernellwp, linux-doc,
jgross, virtualization, sunlightlinux
Add a new per-task stat,
- nr_migrations_cpu_non_preferred: number of push migrations while the
CPU is non-preferred.
Since this new stat is per-task, it changes only /proc/<pid>/sched.
It doesn't update /proc/schedstat. Hence increasing the schedstat version
is not necessary.
Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
---
include/linux/sched.h | 1 +
kernel/sched/core.c | 9 +++++++--
kernel/sched/debug.c | 1 +
3 files changed, 9 insertions(+), 2 deletions(-)
diff --git a/include/linux/sched.h b/include/linux/sched.h
index d7cc77181ef9..48639b0c75d0 100644
--- a/include/linux/sched.h
+++ b/include/linux/sched.h
@@ -554,6 +554,7 @@ struct sched_statistics {
u64 nr_failed_migrations_running;
u64 nr_failed_migrations_hot;
u64 nr_forced_migrations;
+ u64 nr_migrations_cpu_non_preferred;
u64 nr_wakeups;
u64 nr_wakeups_sync;
diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index 5049eff58fb7..7465c983e6f6 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -11236,8 +11236,13 @@ static int sched_non_preferred_cpu_push_stop(void *arg)
update_rq_clock(rq);
context_unsafe_alias(rq);
- if (task_rq(p) == rq && task_on_rq_queued(p))
- rq = __migrate_task(rq, &rf, p, cpu);
+ if (task_rq(p) == rq && task_on_rq_queued(p)) {
+ struct rq *dest_rq = __migrate_task(rq, &rf, p, cpu);
+
+ if (rq != dest_rq)
+ schedstat_inc(p->stats.nr_migrations_cpu_non_preferred);
+ rq = dest_rq;
+ }
rq_unlock(rq, &rf);
}
diff --git a/kernel/sched/debug.c b/kernel/sched/debug.c
index fda10b3fa249..e6a3b516c703 100644
--- a/kernel/sched/debug.c
+++ b/kernel/sched/debug.c
@@ -1440,6 +1440,7 @@ void proc_sched_show_task(struct task_struct *p, struct pid_namespace *ns,
P_SCHEDSTAT(nr_failed_migrations_running);
P_SCHEDSTAT(nr_failed_migrations_hot);
P_SCHEDSTAT(nr_forced_migrations);
+ P_SCHEDSTAT(nr_migrations_cpu_non_preferred);
P_SCHEDSTAT(nr_wakeups);
P_SCHEDSTAT(nr_wakeups_sync);
P_SCHEDSTAT(nr_wakeups_migrate);
--
2.52.0
^ permalink raw reply related [flat|nested] 32+ messages in thread* Re: [PATCH v14 09/13] sched/debug: Add migration stats due to non preferred CPUs
2026-09-28 5:37 ` [PATCH v14 09/13] sched/debug: Add migration stats due to non preferred CPUs Shrikanth Hegde
@ 2026-09-28 5:48 ` sashiko-bot
2026-09-29 12:18 ` Nathan Chancellor
1 sibling, 0 replies; 32+ messages in thread
From: sashiko-bot @ 2026-09-28 5:48 UTC (permalink / raw)
To: Shrikanth Hegde; +Cc: Michael S. Tsirkin, Eugenio Perez, virtualization
> Add a new per-task stat,
> - nr_migrations_cpu_non_preferred: number of push migrations while the
> CPU is non-preferred.
>
> Since this new stat is per-task, it changes only /proc/<pid>/sched.
> It doesn't update /proc/schedstat. Hence increasing the schedstat version
> is not necessary.
>
> Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
Sashiko has reviewed this patch and found no issues. It looks great!
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260928053728.797539-1-sshegde@linux.ibm.com?part=9
^ permalink raw reply [flat|nested] 32+ messages in thread
* Re: [PATCH v14 09/13] sched/debug: Add migration stats due to non preferred CPUs
2026-09-28 5:37 ` [PATCH v14 09/13] sched/debug: Add migration stats due to non preferred CPUs Shrikanth Hegde
2026-09-28 5:48 ` sashiko-bot
@ 2026-09-29 12:18 ` Nathan Chancellor
2026-09-29 12:43 ` Shrikanth Hegde
1 sibling, 1 reply; 32+ messages in thread
From: Nathan Chancellor @ 2026-09-29 12:18 UTC (permalink / raw)
To: Shrikanth Hegde
Cc: linux-kernel, mingo, peterz, juri.lelli, vincent.guittot,
yury.norov, kprateek.nayak, iii, corbet, meted, ynorov, tglx,
gregkh, pbonzini, seanjc, vschneid, huschle, rostedt,
dietmar.eggemann, maddy, srikar, hdanton, chleroy, vineeth,
frederic, arighi, pauld, christian.loehle, tj, tommaso.cucinotta,
maz, rafael, rdunlap, kernellwp, linux-doc, jgross,
virtualization, sunlightlinux, Marco Elver, llvm
On Mon, Sep 28, 2026 at 11:07:24AM +0530, Shrikanth Hegde wrote:
> Add a new per-task stat,
> - nr_migrations_cpu_non_preferred: number of push migrations while the
> CPU is non-preferred.
>
> Since this new stat is per-task, it changes only /proc/<pid>/sched.
> It doesn't update /proc/schedstat. Hence increasing the schedstat version
> is not necessary.
>
> Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
...
> diff --git a/kernel/sched/core.c b/kernel/sched/core.c
> index 5049eff58fb7..7465c983e6f6 100644
> --- a/kernel/sched/core.c
> +++ b/kernel/sched/core.c
> @@ -11236,8 +11236,13 @@ static int sched_non_preferred_cpu_push_stop(void *arg)
> update_rq_clock(rq);
> context_unsafe_alias(rq);
>
> - if (task_rq(p) == rq && task_on_rq_queued(p))
> - rq = __migrate_task(rq, &rf, p, cpu);
> + if (task_rq(p) == rq && task_on_rq_queued(p)) {
> + struct rq *dest_rq = __migrate_task(rq, &rf, p, cpu);
> +
> + if (rq != dest_rq)
> + schedstat_inc(p->stats.nr_migrations_cpu_non_preferred);
> + rq = dest_rq;
> + }
> rq_unlock(rq, &rf);
> }
This breaks the build for me with clang-23+ (which have context analysis
enabled by default), although it bisects to the final patch of the
series since this is under CONFIG_PREFERRED_CPU and it is not selected
until then.
kernel/sched/core.c:11292:25: error: calling function '__migrate_task' requires holding raw_spinlock 'rq_lockp(rq)' exclusively [-Werror,-Wthread-safety-analysis]
11292 | struct rq *dest_rq = __migrate_task(rq, &rf, p, cpu);
| ^
kernel/sched/core.c:11298:3: error: releasing raw_spinlock 'rq_lockp(rq)' that was not held [-Werror,-Wthread-safety-analysis]
11298 | rq_unlock(rq, &rf);
| ^
kernel/sched/core.c:11303:1: error: raw_spinlock 'rq_lockp(__this_rq())' is not held on every path through here [-Werror,-Wthread-safety-analysis]
11303 | }
| ^
kernel/sched/core.c:11286:3: note: raw_spinlock acquired here
11286 | rq_lock(rq, &rf);
| ^
3 errors generated.
Not sure what the proper fix for this is, maybe another
context_unsafe_alias()? cc Marco just in case
--
Cheers,
Nathan
^ permalink raw reply [flat|nested] 32+ messages in thread* Re: [PATCH v14 09/13] sched/debug: Add migration stats due to non preferred CPUs
2026-09-29 12:18 ` Nathan Chancellor
@ 2026-09-29 12:43 ` Shrikanth Hegde
2026-09-29 14:57 ` Shrikanth Hegde
0 siblings, 1 reply; 32+ messages in thread
From: Shrikanth Hegde @ 2026-09-29 12:43 UTC (permalink / raw)
To: Nathan Chancellor
Cc: linux-kernel, mingo, peterz, juri.lelli, vincent.guittot,
yury.norov, kprateek.nayak, iii, corbet, meted, ynorov, tglx,
gregkh, pbonzini, seanjc, vschneid, huschle, rostedt,
dietmar.eggemann, maddy, srikar, hdanton, chleroy, vineeth,
frederic, arighi, pauld, christian.loehle, tj, tommaso.cucinotta,
maz, rafael, rdunlap, kernellwp, linux-doc, jgross,
virtualization, sunlightlinux, Marco Elver, llvm
Hi Nathan. Thanks for report.
On 9/29/26 5:48 PM, Nathan Chancellor wrote:
> On Mon, Sep 28, 2026 at 11:07:24AM +0530, Shrikanth Hegde wrote:
>> Add a new per-task stat,
>> - nr_migrations_cpu_non_preferred: number of push migrations while the
>> CPU is non-preferred.
>>
>> Since this new stat is per-task, it changes only /proc/<pid>/sched.
>> It doesn't update /proc/schedstat. Hence increasing the schedstat version
>> is not necessary.
>>
>> Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
> ...
>> diff --git a/kernel/sched/core.c b/kernel/sched/core.c
>> index 5049eff58fb7..7465c983e6f6 100644
>> --- a/kernel/sched/core.c
>> +++ b/kernel/sched/core.c
>> @@ -11236,8 +11236,13 @@ static int sched_non_preferred_cpu_push_stop(void *arg)
>> update_rq_clock(rq);
>> context_unsafe_alias(rq);
>>
>> - if (task_rq(p) == rq && task_on_rq_queued(p))
>> - rq = __migrate_task(rq, &rf, p, cpu);
>> + if (task_rq(p) == rq && task_on_rq_queued(p)) {
>> + struct rq *dest_rq = __migrate_task(rq, &rf, p, cpu);
>> +
>> + if (rq != dest_rq)
>> + schedstat_inc(p->stats.nr_migrations_cpu_non_preferred);
>> + rq = dest_rq;
>> + }
>> rq_unlock(rq, &rf);
>> }
>
> This breaks the build for me with clang-23+ (which have context analysis
> enabled by default), although it bisects to the final patch of the
> series since this is under CONFIG_PREFERRED_CPU and it is not selected
> until then.
>
> kernel/sched/core.c:11292:25: error: calling function '__migrate_task' requires holding raw_spinlock 'rq_lockp(rq)' exclusively [-Werror,-Wthread-safety-analysis]
> 11292 | struct rq *dest_rq = __migrate_task(rq, &rf, p, cpu);
> | ^
> kernel/sched/core.c:11298:3: error: releasing raw_spinlock 'rq_lockp(rq)' that was not held [-Werror,-Wthread-safety-analysis]
> 11298 | rq_unlock(rq, &rf);
> | ^
> kernel/sched/core.c:11303:1: error: raw_spinlock 'rq_lockp(__this_rq())' is not held on every path through here [-Werror,-Wthread-safety-analysis]
> 11303 | }
> | ^
> kernel/sched/core.c:11286:3: note: raw_spinlock acquired here
> 11286 | rq_lock(rq, &rf);
> | ^
> 3 errors generated.
>
> Not sure what the proper fix for this is, maybe another
> context_unsafe_alias()? cc Marco just in case
>
I suspect it is due to using of dest_rq = rq and rq is changing context.
let me try locally and see the fix.
One fix is use rq->cpu instead of rq comparison to see if migration happened.
^ permalink raw reply [flat|nested] 32+ messages in thread* Re: [PATCH v14 09/13] sched/debug: Add migration stats due to non preferred CPUs
2026-09-29 12:43 ` Shrikanth Hegde
@ 2026-09-29 14:57 ` Shrikanth Hegde
0 siblings, 0 replies; 32+ messages in thread
From: Shrikanth Hegde @ 2026-09-29 14:57 UTC (permalink / raw)
To: Nathan Chancellor, peterz
Cc: linux-kernel, mingo, juri.lelli, vincent.guittot, yury.norov,
kprateek.nayak, iii, corbet, meted, ynorov, tglx, gregkh,
pbonzini, seanjc, vschneid, huschle, rostedt, dietmar.eggemann,
maddy, srikar, hdanton, chleroy, vineeth, frederic, arighi, pauld,
christian.loehle, tj, tommaso.cucinotta, maz, rafael, rdunlap,
kernellwp, linux-doc, jgross, virtualization, sunlightlinux,
Marco Elver, llvm
Hi Nathan.
On 9/29/26 6:13 PM, Shrikanth Hegde wrote:
> Hi Nathan. Thanks for report.
>
> On 9/29/26 5:48 PM, Nathan Chancellor wrote:
>> On Mon, Sep 28, 2026 at 11:07:24AM +0530, Shrikanth Hegde wrote:
>>> Add a new per-task stat,
>>> - nr_migrations_cpu_non_preferred: number of push migrations while the
>>> CPU is non-preferred.
>>>
>>> Since this new stat is per-task, it changes only /proc/<pid>/sched.
>>> It doesn't update /proc/schedstat. Hence increasing the schedstat version
>>> is not necessary.
>>>
>>> Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
>> ...
>>> diff --git a/kernel/sched/core.c b/kernel/sched/core.c
>>> index 5049eff58fb7..7465c983e6f6 100644
>>> --- a/kernel/sched/core.c
>>> +++ b/kernel/sched/core.c
>>> @@ -11236,8 +11236,13 @@ static int sched_non_preferred_cpu_push_stop(void *arg)
>>> update_rq_clock(rq);
>>> context_unsafe_alias(rq);
>>> - if (task_rq(p) == rq && task_on_rq_queued(p))
>>> - rq = __migrate_task(rq, &rf, p, cpu);
>>> + if (task_rq(p) == rq && task_on_rq_queued(p)) {
>>> + struct rq *dest_rq = __migrate_task(rq, &rf, p, cpu);
>>> +
>>> + if (rq != dest_rq)
>>> + schedstat_inc(p->stats.nr_migrations_cpu_non_preferred);
>>> + rq = dest_rq;
>>> + }
>>> rq_unlock(rq, &rf);
>>> }
>>
>> This breaks the build for me with clang-23+ (which have context analysis
>> enabled by default), although it bisects to the final patch of the
>> series since this is under CONFIG_PREFERRED_CPU and it is not selected
>> until then.
>>
>> kernel/sched/core.c:11292:25: error: calling function '__migrate_task' requires holding raw_spinlock 'rq_lockp(rq)' exclusively [-Werror,-Wthread-safety-analysis]
>> 11292 | struct rq *dest_rq = __migrate_task(rq, &rf, p, cpu);
>> | ^
>> kernel/sched/core.c:11298:3: error: releasing raw_spinlock 'rq_lockp(rq)' that was not held [-Werror,-Wthread-safety-analysis]
>> 11298 | rq_unlock(rq, &rf);
>> | ^
>> kernel/sched/core.c:11303:1: error: raw_spinlock 'rq_lockp(__this_rq())' is not held on every path through here [-Werror,-Wthread-safety-analysis]
>> 11303 | }
>> | ^
>> kernel/sched/core.c:11286:3: note: raw_spinlock acquired here
>> 11286 | rq_lock(rq, &rf);
>> | ^
>> 3 errors generated.
>>
>> Not sure what the proper fix for this is, maybe another
>> context_unsafe_alias()? cc Marco just in case
>>
>
> I suspect it is due to using of dest_rq = rq and rq is changing context.
>
> let me try locally and see the fix.
> One fix is use rq->cpu instead of rq comparison to see if migration happened.
Well, it was rather due to Patch 8/13 which had rq marked as context_unsafe_alias
after acquiring it.
I have tried below and that helps to fix the warnings.
I will write a changelog and send it across soon.
Let me know if it works for you.
---
diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index 0bb86a43a592..23677d76f9d2 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -11283,10 +11283,10 @@ static int sched_non_preferred_cpu_push_stop(void *arg)
* safely bail out.
*/
cpu = select_fallback_rq(rq->cpu, p);
+ context_unsafe_alias(rq);
rq_lock(rq, &rf);
rq->npc_push_work_pending = false;
update_rq_clock(rq);
- context_unsafe_alias(rq);
if (task_rq(p) == rq && task_on_rq_queued(p)) {
struct rq *dest_rq = __migrate_task(rq, &rf, p, cpu);
^ permalink raw reply related [flat|nested] 32+ messages in thread
* [PATCH v14 10/13] virt: Introduce steal governor driver
2026-09-28 5:37 [PATCH v14 00/13] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff Shrikanth Hegde
` (8 preceding siblings ...)
2026-09-28 5:37 ` [PATCH v14 09/13] sched/debug: Add migration stats due to non preferred CPUs Shrikanth Hegde
@ 2026-09-28 5:37 ` Shrikanth Hegde
2026-09-28 5:47 ` sashiko-bot
2026-09-28 5:37 ` [PATCH v14 11/13] virt/steal_governor: Add control knobs for handling steal values Shrikanth Hegde
` (2 subsequent siblings)
12 siblings, 1 reply; 32+ messages in thread
From: Shrikanth Hegde @ 2026-09-28 5:37 UTC (permalink / raw)
To: linux-kernel, mingo, peterz, juri.lelli, vincent.guittot,
yury.norov, kprateek.nayak, iii, corbet, meted, ynorov
Cc: sshegde, tglx, gregkh, pbonzini, seanjc, vschneid, huschle,
rostedt, dietmar.eggemann, maddy, srikar, hdanton, chleroy,
vineeth, frederic, arighi, pauld, christian.loehle, tj,
tommaso.cucinotta, maz, rafael, rdunlap, kernellwp, linux-doc,
jgross, virtualization, sunlightlinux
Introduce a new driver in virt named steal_governor. This driver
will compute the steal time and drive the policy decisions regarding the
preferred CPU state.
Note that this driver is strictly intended for actual guests. Hence
block it on Xen dom0.
More details can be found in Documentation/driver-api/steal-governor.rst.
A new kconfig called STEAL_GOVERNOR is introduced in subsequent patches,
which enables this driver. This driver will select CONFIG_PREFERRED_CPU.
This makes configs driven by user preference/configuration.
When the driver is disabled, preferred CPUs remain the same as active CPUs.
The file layout of the driver is kept simple for now. The code is in
drivers/virt/steal_governor.c, and the configs are part of
drivers/virt/Kconfig.
The main structure of the steal governor contains:
- work, delay: Deferred periodic work function variables.
- steal, time: Used to calculate deltas during periodic work.
- interval_ms, high_threshold, low_threshold: Tuning knobs for the
steal governor.
While there, add MAINTAINERS entry for this new driver.
Suggested-by: Yury Norov <yury.norov@gmail.com>
Suggested-by: K Prateek Nayak <kprateek.nayak@amd.com>
Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
---
Documentation/driver-api/index.rst | 1 +
Documentation/driver-api/steal-governor.rst | 151 ++++++++++++++++++++
MAINTAINERS | 9 ++
drivers/virt/steal_governor.c | 78 ++++++++++
4 files changed, 239 insertions(+)
create mode 100644 Documentation/driver-api/steal-governor.rst
create mode 100644 drivers/virt/steal_governor.c
diff --git a/Documentation/driver-api/index.rst b/Documentation/driver-api/index.rst
index 6601a258690f..26b7638a327d 100644
--- a/Documentation/driver-api/index.rst
+++ b/Documentation/driver-api/index.rst
@@ -139,6 +139,7 @@ Subsystem-specific APIs
sm501
soundwire/index
spi
+ steal-governor
surface_aggregator/index
switchtec
sync_file
diff --git a/Documentation/driver-api/steal-governor.rst b/Documentation/driver-api/steal-governor.rst
new file mode 100644
index 000000000000..3817eedb38d7
--- /dev/null
+++ b/Documentation/driver-api/steal-governor.rst
@@ -0,0 +1,151 @@
+.. SPDX-License-Identifier: GPL-2.0
+
+Steal Governor
+==============
+
+:Author: Shrikanth Hegde <sshegde@linux.ibm.com>
+
+Introduction
+============
+
+The steal governor is aimed at mitigating the Noisy Neighbour problem
+which occurs in paravirtualized environments with CPU overcommit.
+The performance of a workload running in one VM gets degraded by
+the activity of other VMs on the same host. As a result, all VMs
+collectively make slower forward progress.
+
+In such systems, high utilization in all VMs causes the hypervisor to
+frequently preempt vCPUs. This vCPU preemption is expensive.
+To mitigate this, the kernel aims to restrict workloads to a subset of
+Preferred CPUs to reduce physical CPU contention.
+A detailed explanation of Preferred CPUs is available in
+``Documentation/scheduler/sched-paravirt.rst``.
+
+The steal governor selects ``CONFIG_PREFERRED_CPU=y`` which enables the
+scheduler core infrastructure to move the tasks to Preferred CPUs where
+possible. The driver controls the policy decisions regarding the state of
+preferred CPUs. That is, this driver decides which CPUs are preferred
+and which CPUs are non-preferred.
+
+The driver code is available at ``drivers/virt/steal_governor.c``.
+
+Core idea
+=========
+
+steal time is an indication available today in Guest which shows contention
+for underlying physical CPU. Use it as a hint in the guest to fold the
+workload to a reduced set of vCPUs. When there is contention, steal time
+will show up in all the guests. When each guest honors the hint and folds
+the workload to a smaller set of vCPUs (Preferred CPUs), it reduces the
+contention and thereby reduces vCPU preemption.
+This is achieved without any cross-guest communication.
+
+Steal governor driver effectively does:
+
+1. Periodically computes the steal ratio using accumulated steal time
+ across possible CPUs, normalized by the number of active CPUs.
+
+2. If steal ratio is greater than high threshold, reduce the number of
+ preferred CPUs by 1 core. Ensure at least one core is left always.
+ Skip changing the state of offline CPUs in that core.
+
+3. If steal ratio is less than or equal to low threshold, increase the
+ number of preferred CPUs by 1 core. If preferred is same as active,
+ nothing to be done. Skip changing the state of offline CPUs.
+ This helps to handle cases where few CPUs are offline in a core and
+ those offline CPUs will not be marked as preferred.
+
+4. Ensure preferred CPUs is always subset of active CPUs.
+ On feature disable it is same as active CPUs.
+
+This feature works best only when all the VMs enable the feature as
+it is a co-operative scheme. If a specific VM doesn't enable this feature
+it may end up with more CPUs than others, still should lead to better
+performance when seen from system view. Those who enable this driver must
+ensure it is enabled in all VMs.
+
+Note that this driver is strictly intended for actual guests; for example,
+loading this module in a privileged VM like Xen Dom0 is blocked.
+
+Workload considerations
+=======================
+
+The steal governor is useful for workloads where vCPU preemption has
+costs beyond the lost CPU time, such as lock-holder preemption, critical
+sections, communicating threads, and cache or TLB disruption.
+
+Pure CPU-time workloads with independent workers may not benefit and
+could see a small regression due to additional guest scheduling overhead.
+
+Module Parameters
+=================
+
+interval_ms
+-----------
+
+How often steal governor checks for steal time.
+Default: 1000 i.e. 1 second. Value should be in between 100ms to 100sec.
+
+This controls how fast steal governor driver reacts to changes to the
+contention of physical CPUs. Since it does a fair amount of work, setting
+too low may have overhead. Setting it too high might render it ineffective.
+
+low_threshold
+-------------
+
+lower threshold value in percentage * 100.
+Default: 200, i.e. 2% steal is considered as low threshold.
+Can't be higher than high_threshold.
+
+This determines what values should be considered as nil/no steal values.
+When steal governor sees steal ratio is less than or equal to this value,
+it will increase the preferred CPUs by 1 core.
+Using zero might cause oscillations.
+
+high_threshold
+--------------
+
+higher threshold value in percentage * 100
+Default: 500, i.e. 5% steal is considered as high threshold.
+Can't be lower than low_threshold. Must be less than 10000.
+
+This determines what values should be considered as high steal values.
+When steal governor sees steal ratio is higher than this value, it will
+reduce the preferred CPUs by 1 core.
+
+Limitations of default values
+-----------------------------
+
+Because of the vast diversity in VM configurations and different
+architectures, the default thresholds may not be optimal for all systems.
+Users may need to tune these parameters based on the system under
+test to achieve the best results.
+
+The governor sums the steal time across all possible CPUs, which ensures
+the accumulated steal time remains a monotonically increasing value.
+However, to calculate the effective steal ratio, it divides this sum
+by the number of active CPUs. Because only active CPUs contribute to
+the steal time delta, this prevents threshold dilution on sparsely
+populated systems.
+
+The driver reduces/increases preferred CPUs by core-level. This could provide
+faster convergence for hypervisors such as powerVM. But on KVM and Xen
+convergence could be slower depending on the configuration.
+Using a smaller interval_ms could help one to expedite it.
+
+Reasons for CONFIG_STEAL_GOVERNOR=m
+===================================
+
+Selecting this driver makes CONFIG_PREFERRED_CPU=y. That makes configs
+driven by user preference. Though one can have CONFIG_STEAL_GOVERNOR=y,
+It is recommended to build CONFIG_STEAL_GOVERNOR=m due to below reasons:
+
+1. Doing periodic work has additional overheads. Enabling this driver
+ in systems where steal time cannot happen is of no use. There is no
+ benefit with additional overheads in such systems.
+
+2. This works well when all VMs work in a co-operative manner. When an
+ administrative user enables it in one VM, he/she will likely enable
+ it in all VMs.
+
+3. User can tweak the module parameters by reloading the module.
diff --git a/MAINTAINERS b/MAINTAINERS
index 3a19da74d00c..28003d36b3f7 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -26223,6 +26223,15 @@ F: rust/helpers/jump_label.c
F: rust/kernel/generated_arch_static_branch_asm.rs.S
F: rust/kernel/jump_label.rs
+STEAL GOVERNOR DRIVER
+M: Shrikanth Hegde <sshegde@linux.ibm.com>
+R: Yury Norov <yury.norov@gmail.com>
+L: linux-kernel@vger.kernel.org
+S: Maintained
+T: git git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git sched/core
+F: Documentation/driver-api/steal-governor.rst
+F: drivers/virt/steal_governor.c
+
STI AUDIO (ASoC) DRIVERS
M: Arnaud Pouliquen <arnaud.pouliquen@foss.st.com>
L: linux-sound@vger.kernel.org
diff --git a/drivers/virt/steal_governor.c b/drivers/virt/steal_governor.c
new file mode 100644
index 000000000000..2320cbfa4b47
--- /dev/null
+++ b/drivers/virt/steal_governor.c
@@ -0,0 +1,78 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/*
+ * Steal time governor driver periodically computes steal time.
+ * Based on the thresholds it either reduce/increase the preferred
+ * CPUs which can be used by the workload to avoid vCPU preemption
+ * to an extent possible in paravirtualized environment.
+ *
+ * Available with CONFIG_STEAL_GOVERNOR
+ *
+ * Copyright (C) 2026 IBM
+ * Author: Shrikanth Hegde <sshegde@linux.ibm.com>
+ */
+
+#define pr_fmt(fmt) KBUILD_MODNAME ": " fmt
+
+#include <linux/cpuhplock.h>
+#include <linux/cpumask.h>
+#include <linux/init.h>
+#include <linux/kernel.h>
+#include <linux/kconfig.h>
+#include <linux/ktime.h>
+#include <linux/module.h>
+#include <linux/types.h>
+#include <linux/workqueue.h>
+#ifdef CONFIG_XEN
+#include <xen/xen.h>
+#endif
+
+#if !IS_ENABLED(CONFIG_PREFERRED_CPU)
+#error "Steal Governor requires CONFIG_PREFERRED_CPU"
+#endif
+
+struct steal_governor {
+ ktime_t time;
+ u64 steal;
+ unsigned long delay;
+ unsigned int interval_ms;
+ unsigned int high_threshold;
+ unsigned int low_threshold;
+ struct delayed_work work;
+};
+
+static struct steal_governor sg_ctx;
+
+static void restore_preferred_to_active(void)
+{
+ int cpu;
+
+ guard(cpus_read_lock)();
+ for_each_cpu(cpu, cpu_active_mask)
+ set_cpu_preferred(cpu, true);
+}
+
+static int __init steal_governor_init(void)
+{
+#ifdef CONFIG_XEN
+ if (xen_initial_domain()) {
+ pr_err("Cannot load in Xen Dom0 (Host OS). Driver is for guests only.\n");
+ return -ENODEV;
+ }
+#endif
+
+ pr_info("enabled\n");
+ return 0;
+}
+
+static void __exit steal_governor_exit(void)
+{
+ restore_preferred_to_active();
+ pr_info("disabled\n");
+}
+
+module_init(steal_governor_init);
+module_exit(steal_governor_exit);
+
+MODULE_LICENSE("GPL");
+MODULE_AUTHOR("IBM Corporation");
+MODULE_DESCRIPTION("Virtualization Steal Time Governor");
--
2.52.0
^ permalink raw reply related [flat|nested] 32+ messages in thread* Re: [PATCH v14 10/13] virt: Introduce steal governor driver
2026-09-28 5:37 ` [PATCH v14 10/13] virt: Introduce steal governor driver Shrikanth Hegde
@ 2026-09-28 5:47 ` sashiko-bot
0 siblings, 0 replies; 32+ messages in thread
From: sashiko-bot @ 2026-09-28 5:47 UTC (permalink / raw)
To: Shrikanth Hegde; +Cc: Eugenio Perez, virtualization, Michael S. Tsirkin
> Introduce a new driver in virt named steal_governor. This driver
> will compute the steal time and drive the policy decisions regarding the
> preferred CPU state.
>
> Note that this driver is strictly intended for actual guests. Hence
> block it on Xen dom0.
> More details can be found in Documentation/driver-api/steal-governor.rst.
>
> A new kconfig called STEAL_GOVERNOR is introduced in subsequent patches,
> which enables this driver. This driver will select CONFIG_PREFERRED_CPU.
> This makes configs driven by user preference/configuration.
> When the driver is disabled, preferred CPUs remain the same as active CPUs.
>
> The file layout of the driver is kept simple for now. The code is in
> drivers/virt/steal_governor.c, and the configs are part of
> drivers/virt/Kconfig.
>
> The main structure of the steal governor contains:
> - work, delay: Deferred periodic work function variables.
> - steal, time: Used to calculate deltas during periodic work.
> - interval_ms, high_threshold, low_threshold: Tuning knobs for the
> steal governor.
>
> While there, add MAINTAINERS entry for this new driver.
>
> Suggested-by: Yury Norov <yury.norov@gmail.com>
> Suggested-by: K Prateek Nayak <kprateek.nayak@amd.com>
> Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
Sashiko has reviewed this patch and found no issues. It looks great!
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260928053728.797539-1-sshegde@linux.ibm.com?part=10
^ permalink raw reply [flat|nested] 32+ messages in thread
* [PATCH v14 11/13] virt/steal_governor: Add control knobs for handling steal values
2026-09-28 5:37 [PATCH v14 00/13] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff Shrikanth Hegde
` (9 preceding siblings ...)
2026-09-28 5:37 ` [PATCH v14 10/13] virt: Introduce steal governor driver Shrikanth Hegde
@ 2026-09-28 5:37 ` Shrikanth Hegde
2026-09-28 5:46 ` sashiko-bot
2026-09-28 5:37 ` [PATCH v14 12/13] virt/steal_governor: Implement steal_governor policy loop Shrikanth Hegde
2026-09-28 5:37 ` [PATCH v14 13/13] virt/steal_governor: Enable the driver Shrikanth Hegde
12 siblings, 1 reply; 32+ messages in thread
From: Shrikanth Hegde @ 2026-09-28 5:37 UTC (permalink / raw)
To: linux-kernel, mingo, peterz, juri.lelli, vincent.guittot,
yury.norov, kprateek.nayak, iii, corbet, meted, ynorov
Cc: sshegde, tglx, gregkh, pbonzini, seanjc, vschneid, huschle,
rostedt, dietmar.eggemann, maddy, srikar, hdanton, chleroy,
vineeth, frederic, arighi, pauld, christian.loehle, tj,
tommaso.cucinotta, maz, rafael, rdunlap, kernellwp, linux-doc,
jgross, virtualization, sunlightlinux
These are the knobs to control the steal_governor.
interval_ms:
How often steal governor checks for steal time.
(Default: 1000 i.e 1 second)
This controls how fast steal governor driver reacts to changes to
the contention of physical CPUs.
Can be set between 100 to 100000. i.e. 100ms to 100seconds.
100ms is kept as minimum to ensure few meaningful steal values
accumulate even with HZ=100.
low_threshold:
lower threshold value in percentage * 100.
(Default: 200, i.e 2% steal is considered as low threshold)
This determines what values should be considered as nil/no steal values.
When steal governor see steal ratio is below or equal to this value, it
will increase the preferred CPUs by 1 core. Having value as zero
might cause oscillations
high_threshold:
higher threshold value in percentage * 100
(Default: 500, i.e 5% steal is considered as high threshold)
This determines what values should be considered as high steal values.
When steal governor sees steal ratio is higher than this value, it will
reduce the preferred CPUs by 1 core.
module_param_cb methods are used to do the validation checks.
This helps to ensure one configures sane values.
Since low and high are dependent, that check is done at module init.
Notes:
- Parameters values can't be changed at runtime. One has to unload
the module and change it. Hence recommended to build it as module.
- Default values may not work well for all configurations. Tune it
according to the system under test.
Documentation is available at: Documentation/driver-api/steal-governor.rst
Suggested-by: Yury Norov <yury.norov@gmail.com>
Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
---
drivers/virt/steal_governor.c | 73 ++++++++++++++++++++++++++++++++++-
1 file changed, 71 insertions(+), 2 deletions(-)
diff --git a/drivers/virt/steal_governor.c b/drivers/virt/steal_governor.c
index 2320cbfa4b47..27f53ea16498 100644
--- a/drivers/virt/steal_governor.c
+++ b/drivers/virt/steal_governor.c
@@ -40,7 +40,11 @@ struct steal_governor {
struct delayed_work work;
};
-static struct steal_governor sg_ctx;
+static struct steal_governor sg_ctx = {
+ .interval_ms = 1000, /* 1 second */
+ .high_threshold = 500, /* 5% */
+ .low_threshold = 200, /* 2% */
+};
static void restore_preferred_to_active(void)
{
@@ -51,6 +55,62 @@ static void restore_preferred_to_active(void)
set_cpu_preferred(cpu, true);
}
+static int param_set_interval_ms(const char *val, const struct kernel_param *kp)
+{
+ unsigned int interval;
+ int ret;
+
+ ret = kstrtouint(val, 0, &interval);
+ if (ret)
+ return ret;
+
+ if (interval < 100 || interval > 100000) {
+ pr_err("interval_ms must be between 100 and 100000\n");
+ return -EINVAL;
+ }
+
+ return param_set_uint(val, kp);
+}
+
+static const struct kernel_param_ops interval_ms_ops = {
+ .set = param_set_interval_ms,
+ .get = param_get_uint,
+};
+
+module_param_cb(interval_ms, &interval_ms_ops, &sg_ctx.interval_ms, 0444);
+MODULE_PARM_DESC(interval_ms,
+ "Sampling frequency in milliseconds. default: 1000");
+
+static int param_set_high_threshold(const char *val, const struct kernel_param *kp)
+{
+ unsigned int threshold;
+ int ret;
+
+ ret = kstrtouint(val, 0, &threshold);
+ if (ret)
+ return ret;
+
+ if (threshold >= 100 * 100) {
+ pr_err("high_threshold (%u) can't be more than 99.99%%\n", threshold);
+ return -EINVAL;
+ }
+
+ return param_set_uint(val, kp);
+}
+
+static const struct kernel_param_ops high_threshold_ops = {
+ .set = param_set_high_threshold,
+ .get = param_get_uint,
+};
+
+module_param_cb(high_threshold, &high_threshold_ops, &sg_ctx.high_threshold, 0444);
+MODULE_PARM_DESC(high_threshold,
+ "High steal threshold. default: 500 i.e 5%. Must be > low_threshold");
+
+module_param_named(low_threshold, sg_ctx.low_threshold, uint, 0444);
+MODULE_PARM_DESC(low_threshold,
+ "Low steal threshold. default: 200 i.e 2%. Must be < high_threshold");
+
static int __init steal_governor_init(void)
{
#ifdef CONFIG_XEN
@@ -60,7 +120,16 @@ static int __init steal_governor_init(void)
}
#endif
- pr_info("enabled\n");
+ if (sg_ctx.low_threshold >= sg_ctx.high_threshold) {
+ pr_err("low_threshold (%u) must be less than high_threshold (%u)\n",
+ sg_ctx.low_threshold, sg_ctx.high_threshold);
+ return -EINVAL;
+ }
+
+ sg_ctx.delay = msecs_to_jiffies(sg_ctx.interval_ms);
+ pr_info("enabled. interval: %ums, high_threshold: %u, low_threshold: %u\n",
+ sg_ctx.interval_ms, sg_ctx.high_threshold, sg_ctx.low_threshold);
+
return 0;
}
--
2.52.0
^ permalink raw reply related [flat|nested] 32+ messages in thread* Re: [PATCH v14 11/13] virt/steal_governor: Add control knobs for handling steal values
2026-09-28 5:37 ` [PATCH v14 11/13] virt/steal_governor: Add control knobs for handling steal values Shrikanth Hegde
@ 2026-09-28 5:46 ` sashiko-bot
0 siblings, 0 replies; 32+ messages in thread
From: sashiko-bot @ 2026-09-28 5:46 UTC (permalink / raw)
To: Shrikanth Hegde; +Cc: Michael S. Tsirkin, Eugenio Perez, virtualization
> These are the knobs to control the steal_governor.
>
> interval_ms:
> How often steal governor checks for steal time.
> (Default: 1000 i.e 1 second)
> This controls how fast steal governor driver reacts to changes to
> the contention of physical CPUs.
> Can be set between 100 to 100000. i.e. 100ms to 100seconds.
> 100ms is kept as minimum to ensure few meaningful steal values
> accumulate even with HZ=100.
>
> low_threshold:
> lower threshold value in percentage * 100.
> (Default: 200, i.e 2% steal is considered as low threshold)
> This determines what values should be considered as nil/no steal values.
> [ ... ]
>
> Documentation is available at: Documentation/driver-api/steal-governor.rst
>
> Suggested-by: Yury Norov <yury.norov@gmail.com>
> Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
Sashiko has reviewed this patch and found no issues. It looks great!
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260928053728.797539-1-sshegde@linux.ibm.com?part=11
^ permalink raw reply [flat|nested] 32+ messages in thread
* [PATCH v14 12/13] virt/steal_governor: Implement steal_governor policy loop
2026-09-28 5:37 [PATCH v14 00/13] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff Shrikanth Hegde
` (10 preceding siblings ...)
2026-09-28 5:37 ` [PATCH v14 11/13] virt/steal_governor: Add control knobs for handling steal values Shrikanth Hegde
@ 2026-09-28 5:37 ` Shrikanth Hegde
2026-09-28 5:51 ` sashiko-bot
2026-09-28 5:37 ` [PATCH v14 13/13] virt/steal_governor: Enable the driver Shrikanth Hegde
12 siblings, 1 reply; 32+ messages in thread
From: Shrikanth Hegde @ 2026-09-28 5:37 UTC (permalink / raw)
To: linux-kernel, mingo, peterz, juri.lelli, vincent.guittot,
yury.norov, kprateek.nayak, iii, corbet, meted, ynorov
Cc: sshegde, tglx, gregkh, pbonzini, seanjc, vschneid, huschle,
rostedt, dietmar.eggemann, maddy, srikar, hdanton, chleroy,
vineeth, frederic, arighi, pauld, christian.loehle, tj,
tommaso.cucinotta, maz, rafael, rdunlap, kernellwp, linux-doc,
jgross, virtualization, sunlightlinux
Schedule work at regular intervals to implement the steal_governor
policy loop, which monitors steal time and takes action on the state of
preferred CPUs. The interval is determined by the interval_ms parameter.
schedule_delayed_work() is used since interval_ms is on the order of
milliseconds and the work does not need to happen instantly.
Periodic policy loop essentially does:
- Gets the total/delta steal values and cpus to use steal_ratio.
- Calculate the steal_ratio as below.
steal_ratio = (delta_steal * 100*100)/(delta_ns * num_cpus())
It is calculated this way to consider the fractional values of steal
time. I.e 10 means 0.1% steal time. A few tricks such as
divide by 10,000 are used to avoid possible overflow.
- If steal ratio is higher than high threshold, call the method to reduce
the preferred CPUs.
- If steal ratio is lower or equal to low threshold, call the method to
increase the preferred CPUs.
- If the steal ratio falls in between, no action is taken.
- Ensures design constraints always met.
1. At least one core/CPU must be there in preferred mask.
2. preferred CPUs is subset of active CPUs.
If not met, then restore preferred CPUs to active and stop
requeue of the work. Driver is effectively non-functional after that.
Note that design checks are always performed. This helps avoid placing
driver-specific design constraints inside the core CPU hotplug mechanism.
User may offline specific set of CPUs that could leave the preferred
mask as empty. With the design check performed always, driver gracefully
shuts down upon detecting that edge case.
In order to help the above loop, a few helper functions have been added.
1. get_system_steal_time()
- steal governor takes global view of steal time instead of individual
vCPU. Collect the steal values across the vCPUs of interest.
- Sum up steal time values across possible CPUs. This helps to keep it
a monotonically increasing number and avoids spikes due to CPU
hotplug.
2. decrease_preferred_cpus()
- Called when there is high steal time. It needs to decide which CPUs to
mark as non-preferred.
- Get first housekeeping CPU and its core mask. Mark it as
protected core. This helps to keep at least one core as preferred.
(kernel ensures at least one housekeeping CPU stays active.)
- Find the last CPU outside of this protected core mask. i.e target CPU
- Based on that target CPU, get its sibling and mark them as
non-preferred.
3. increase_preferred_cpus()
- Called when there is low steal time. It needs to decide which CPUs to
mark as preferred and set that state.
- Get the first active non-preferred CPUs. This likely is the last
set of CPUs being marked as non-preferred.
- get the siblings of that CPU and mark them as preferred.
4. get_system_cpus()
- informs how many CPUs needs to be considered for steal_ratio
calculations.
- Since only active CPUs effectively contribute to steal time delta,
returns number of active CPUs. This also helps to avoid dilution of
thresholds in sparsely populated systems.
Notes:
1. Using core instead of individual CPUs performs better as SMT is
quite common and some hypervisor such as powerVM does core scheduling.
2. This doesn't do any NUMA splicing to keep the code simpler and
minimal overhead. Current code expects CPUs spread uniformly
across NUMA nodes.
Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
---
drivers/virt/steal_governor.c | 149 ++++++++++++++++++++++++++++++++++
1 file changed, 149 insertions(+)
diff --git a/drivers/virt/steal_governor.c b/drivers/virt/steal_governor.c
index 27f53ea16498..6e31f9923dea 100644
--- a/drivers/virt/steal_governor.c
+++ b/drivers/virt/steal_governor.c
@@ -13,13 +13,18 @@
#define pr_fmt(fmt) KBUILD_MODNAME ": " fmt
+#include <linux/cleanup.h>
#include <linux/cpuhplock.h>
#include <linux/cpumask.h>
#include <linux/init.h>
#include <linux/kernel.h>
+#include <linux/kernel_stat.h>
#include <linux/kconfig.h>
#include <linux/ktime.h>
+#include <linux/math64.h>
#include <linux/module.h>
+#include <linux/sched/isolation.h>
+#include <linux/topology.h>
#include <linux/types.h>
#include <linux/workqueue.h>
#ifdef CONFIG_XEN
@@ -111,6 +116,145 @@ module_param_named(low_threshold, sg_ctx.low_threshold, uint, 0444);
MODULE_PARM_DESC(low_threshold,
"Low steal threshold. default: 200 i.e 2%. Must be < high_threshold");
+/* Return collective steal time across system. */
+static u64 get_system_steal_time(void)
+{
+ return kcpustat_field_total(CPUTIME_STEAL, cpu_possible_mask);
+}
+
+/* Return number of CPUs to consider for steal ratio. */
+static unsigned int get_system_cpus(void)
+{
+ return num_active_cpus();
+}
+
+/*
+ * Called when the steal governor detects high physical CPU contention.
+ * It finds the last active core in the preferred mask and mark those
+ * CPUs as non-preferred.
+ *
+ * Must ensure:
+ * - at least one core is always kept as preferred
+ * - preferred is always subset of active.
+ */
+static void decrease_preferred_cpus(void)
+{
+ const struct cpumask *first_hk_core;
+ int target_cpu = nr_cpu_ids;
+ int cpu;
+
+ guard(cpus_read_lock)();
+ cpu = cpumask_first_and(housekeeping_cpumask(HK_TYPE_KERNEL_NOISE),
+ cpu_preferred_mask);
+ if (cpu >= nr_cpu_ids)
+ return;
+
+ /* Always leave first housekeeping core as preferred. */
+ first_hk_core = topology_sibling_cpumask(cpu);
+ cpu = cpumask_last(cpu_preferred_mask);
+ if (cpu >= nr_cpu_ids)
+ return;
+
+ /* Find the last CPU which doesn't belong to that first hk_core. */
+ if (!cpumask_test_cpu(cpu, first_hk_core)) {
+ target_cpu = cpu;
+ } else {
+ for_each_cpu_andnot(cpu, cpu_preferred_mask, first_hk_core)
+ target_cpu = cpu;
+ }
+
+ /* Only the first housekeeping core remains */
+ if (target_cpu >= nr_cpu_ids)
+ return;
+
+ for_each_cpu_and(cpu, topology_sibling_cpumask(target_cpu),
+ cpu_preferred_mask)
+ set_cpu_preferred(cpu, false);
+}
+
+/*
+ * Called when the steal governor detects no/low physical CPU contention.
+ * It finds the first active core outside of preferred mask and mark
+ * those CPUs as preferred.
+ *
+ * Must ensure preferred is subset of active.
+ */
+static void increase_preferred_cpus(void)
+{
+ int first_cpu, cpu;
+
+ guard(cpus_read_lock)();
+ first_cpu = cpumask_first_andnot(cpu_active_mask, cpu_preferred_mask);
+
+ /* All CPUs are preferred. Nothing to increase further */
+ if (first_cpu >= nr_cpu_ids)
+ return;
+
+ for_each_cpu_and(cpu, topology_sibling_cpumask(first_cpu),
+ cpu_active_mask)
+ set_cpu_preferred(cpu, true);
+}
+
+static bool preferred_cpus_valid(void)
+{
+ if (cpumask_empty(cpu_preferred_mask)) {
+ pr_err("empty preferred mask. stopping\n");
+ return false;
+ }
+
+ if (!cpumask_subset(cpu_preferred_mask, cpu_active_mask)) {
+ pr_err("preferred: %*pbl is not subset of active: %*pbl, stopping\n",
+ cpumask_pr_args(cpu_preferred_mask),
+ cpumask_pr_args(cpu_active_mask));
+ return false;
+ }
+
+ return true;
+}
+
+static void steal_governor_loop(struct work_struct *work)
+{
+ u64 curr_steal, delta_steal, delta_ns, steal_ratio;
+ ktime_t now;
+
+ now = ktime_get();
+ delta_ns = ktime_to_ns(ktime_sub(now, sg_ctx.time));
+
+ if (unlikely(delta_ns < NSEC_PER_MSEC)) {
+ pr_err_ratelimited("work scheduled too soon delta_ns: %llu\n", delta_ns);
+ goto requeue_work;
+ }
+
+ curr_steal = get_system_steal_time();
+ delta_steal = curr_steal > sg_ctx.steal ? curr_steal - sg_ctx.steal : 0;
+ sg_ctx.steal = curr_steal;
+ sg_ctx.time = now;
+
+ /*
+ * steal_ratio = (delta_steal * 100*100)/(delta_ns * num_cpus())
+ * To avoid possible overflow, divide the denominator early.
+ * Note minimum interval is 100ms.
+ */
+ delta_ns = max_t(u64, div_u64(delta_ns * get_system_cpus(), 10000), 1);
+ steal_ratio = div64_u64(delta_steal, delta_ns);
+
+ if (steal_ratio > sg_ctx.high_threshold)
+ decrease_preferred_cpus();
+ else if (steal_ratio <= sg_ctx.low_threshold)
+ increase_preferred_cpus();
+ /*
+ * else: steal ratio is within bounds. Still do design checks so that
+ * module restores to active if CPU hotplug breaks those assumptions.
+ */
+ if (!preferred_cpus_valid()) {
+ restore_preferred_to_active();
+ return;
+ }
+
+requeue_work:
+ schedule_delayed_work(&sg_ctx.work, sg_ctx.delay);
+}
+
static int __init steal_governor_init(void)
{
#ifdef CONFIG_XEN
@@ -127,6 +271,10 @@ static int __init steal_governor_init(void)
}
sg_ctx.delay = msecs_to_jiffies(sg_ctx.interval_ms);
+ INIT_DELAYED_WORK(&sg_ctx.work, steal_governor_loop);
+ sg_ctx.steal = get_system_steal_time();
+ sg_ctx.time = ktime_get();
+ schedule_delayed_work(&sg_ctx.work, sg_ctx.delay);
pr_info("enabled. interval: %ums, high_threshold: %u, low_threshold: %u\n",
sg_ctx.interval_ms, sg_ctx.high_threshold, sg_ctx.low_threshold);
@@ -135,6 +283,7 @@ static int __init steal_governor_init(void)
static void __exit steal_governor_exit(void)
{
+ disable_delayed_work_sync(&sg_ctx.work);
restore_preferred_to_active();
pr_info("disabled\n");
}
--
2.52.0
^ permalink raw reply related [flat|nested] 32+ messages in thread* Re: [PATCH v14 12/13] virt/steal_governor: Implement steal_governor policy loop
2026-09-28 5:37 ` [PATCH v14 12/13] virt/steal_governor: Implement steal_governor policy loop Shrikanth Hegde
@ 2026-09-28 5:51 ` sashiko-bot
2026-09-28 6:52 ` Shrikanth Hegde
0 siblings, 1 reply; 32+ messages in thread
From: sashiko-bot @ 2026-09-28 5:51 UTC (permalink / raw)
To: Shrikanth Hegde; +Cc: Eugenio Perez, virtualization, Michael S. Tsirkin
Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
- [High] Missing bounds check on interval_ms allows infinite workqueue spinloop
- [High] TOCTOU data race in preferred_cpus_valid() causes erroneous driver shutdown
- [Medium] Integer overflow in steal_ratio denominator when VM is paused
--
commit 5f38399332407b64e0ce2da43681c27f1e1287d5
Author: Shrikanth Hegde <sshegde@linux.ibm.com>
virt/steal_governor: Implement steal_governor policy loop
Schedule work at regular intervals to implement the steal_governor
policy loop, which monitors steal time and takes action on the state of
preferred CPUs. The interval is determined by the interval_ms parameter.
> diff --git a/drivers/virt/steal_governor.c b/drivers/virt/steal_governor.c
> index 27f53ea16498b..6e31f9923dea4 100644
> --- a/drivers/virt/steal_governor.c
> +++ b/drivers/virt/steal_governor.c
[ ... ]
> +static bool preferred_cpus_valid(void)
> +{
> + if (cpumask_empty(cpu_preferred_mask)) {
> + pr_err("empty preferred mask. stopping\n");
> + return false;
> + }
> +
> + if (!cpumask_subset(cpu_preferred_mask, cpu_active_mask)) {
[Severity: High]
Does this lockless read of cpu_preferred_mask and cpu_active_mask race
with CPU hotplug?
Since CPU hotplug clears the preferred mask bit before the active mask bit,
it seems cpumask_subset() can read a torn state where a word is read from
the preferred mask before the bit is cleared, and the corresponding word
is read from the active mask after the bit is cleared, returning false.
> + pr_err("preferred: %*pbl is not subset of active: %*pbl, stopping\n",
> + cpumask_pr_args(cpu_preferred_mask),
> + cpumask_pr_args(cpu_active_mask));
> + return false;
> + }
> +
> + return true;
> +}
> +
> +static void steal_governor_loop(struct work_struct *work)
> +{
> + u64 curr_steal, delta_steal, delta_ns, steal_ratio;
> + ktime_t now;
> +
> + now = ktime_get();
> + delta_ns = ktime_to_ns(ktime_sub(now, sg_ctx.time));
> +
> + if (unlikely(delta_ns < NSEC_PER_MSEC)) {
> + pr_err_ratelimited("work scheduled too soon delta_ns: %llu\n", delta_ns);
> + goto requeue_work;
[Severity: High]
Can this create an infinite spinloop if interval_ms is set to 0?
If interval_ms is 0, the delay is 0 jiffies. When the work executes, it will
take this branch because delta_ns < 1ms. It jumps directly to requeue_work
without updating sg_ctx.time, which schedules it again with 0 delay. This
will continuously spin and consume an entire CPU core.
> + }
> +
> + curr_steal = get_system_steal_time();
> + delta_steal = curr_steal > sg_ctx.steal ? curr_steal - sg_ctx.steal : 0;
> + sg_ctx.steal = curr_steal;
> + sg_ctx.time = now;
> +
> + /*
> + * steal_ratio = (delta_steal * 100*100)/(delta_ns * num_cpus())
> + * To avoid possible overflow, divide the denominator early.
> + * Note minimum interval is 100ms.
> + */
> + delta_ns = max_t(u64, div_u64(delta_ns * get_system_cpus(), 10000), 1);
[Severity: Medium]
Will this silently overflow the 64-bit multiplication if the VM is paused
for an extended period?
Since the multiplication of delta_ns * get_system_cpus() happens before
the division, a large time delta (e.g., a large VM suspended for a long time)
can overflow U64_MAX, leading to a wildly incorrect steal_ratio calculation.
> + steal_ratio = div64_u64(delta_steal, delta_ns);
> +
> + if (steal_ratio > sg_ctx.high_threshold)
> + decrease_preferred_cpus();
> + else if (steal_ratio <= sg_ctx.low_threshold)
> + increase_preferred_cpus();
> + /*
> + * else: steal ratio is within bounds. Still do design checks so that
> + * module restores to active if CPU hotplug breaks those assumptions.
> + */
> + if (!preferred_cpus_valid()) {
[Severity: High]
Does calling preferred_cpus_valid() here without cpus_read_lock() lead
to a false teardown?
If a hotplug operation is happening concurrently, the torn state read in
preferred_cpus_valid() will cause the governor to erroneously shut down
the policy loop permanently.
> + restore_preferred_to_active();
> + return;
> + }
> +
> +requeue_work:
> + schedule_delayed_work(&sg_ctx.work, sg_ctx.delay);
[Severity: High]
Does this immediately spin if sg_ctx.delay is 0, as mentioned above?
> +}
> +
> static int __init steal_governor_init(void)
> {
[ ... ]
> sg_ctx.delay = msecs_to_jiffies(sg_ctx.interval_ms);
[Severity: High]
Is it possible for sg_ctx.interval_ms to be 0 here?
The module parameter interval_ms doesn't appear to be bounds-checked
before this point, meaning it could be set to 0 and trigger the infinite
workqueue spinloop.
> + INIT_DELAYED_WORK(&sg_ctx.work, steal_governor_loop);
> + sg_ctx.steal = get_system_steal_time();
> + sg_ctx.time = ktime_get();
> + schedule_delayed_work(&sg_ctx.work, sg_ctx.delay);
> pr_info("enabled. interval: %ums, high_threshold: %u, low_threshold: %u\n",
> sg_ctx.interval_ms, sg_ctx.high_threshold, sg_ctx.low_threshold);
>
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260928053728.797539-1-sshegde@linux.ibm.com?part=12
^ permalink raw reply [flat|nested] 32+ messages in thread* Re: [PATCH v14 12/13] virt/steal_governor: Implement steal_governor policy loop
2026-09-28 5:51 ` sashiko-bot
@ 2026-09-28 6:52 ` Shrikanth Hegde
0 siblings, 0 replies; 32+ messages in thread
From: Shrikanth Hegde @ 2026-09-28 6:52 UTC (permalink / raw)
To: sashiko-reviews; +Cc: Eugenio Perez, virtualization, Michael S. Tsirkin
I don't see any real concerns so far. Some optimizations yes, but they
can be deferred post the merge if they are necessary.
On 9/28/26 11:21 AM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
> - [High] Missing bounds check on interval_ms allows infinite workqueue spinloop
> - [High] TOCTOU data race in preferred_cpus_valid() causes erroneous driver shutdown
> - [Medium] Integer overflow in steal_ratio denominator when VM is paused
> --
>
> commit 5f38399332407b64e0ce2da43681c27f1e1287d5
> Author: Shrikanth Hegde <sshegde@linux.ibm.com>
>
> virt/steal_governor: Implement steal_governor policy loop
>
> Schedule work at regular intervals to implement the steal_governor
> policy loop, which monitors steal time and takes action on the state of
> preferred CPUs. The interval is determined by the interval_ms parameter.
>
>> diff --git a/drivers/virt/steal_governor.c b/drivers/virt/steal_governor.c
>> index 27f53ea16498b..6e31f9923dea4 100644
>> --- a/drivers/virt/steal_governor.c
>> +++ b/drivers/virt/steal_governor.c
> [ ... ]
>> +static bool preferred_cpus_valid(void)
>> +{
>> + if (cpumask_empty(cpu_preferred_mask)) {
>> + pr_err("empty preferred mask. stopping\n");
>> + return false;
>> + }
>> +
>> + if (!cpumask_subset(cpu_preferred_mask, cpu_active_mask)) {
>
> [Severity: High]
> Does this lockless read of cpu_preferred_mask and cpu_active_mask race
> with CPU hotplug?
>
> Since CPU hotplug clears the preferred mask bit before the active mask bit,
> it seems cpumask_subset() can read a torn state where a word is read from
> the preferred mask before the bit is cleared, and the corresponding word
> is read from the active mask after the bit is cleared, returning false.
Update ordering is taken care. Not an issue is clear/set bit involves atomics.
>
>> + pr_err("preferred: %*pbl is not subset of active: %*pbl, stopping\n",
>> + cpumask_pr_args(cpu_preferred_mask),
>> + cpumask_pr_args(cpu_active_mask));
>> + return false;
>> + }
>> +
>> + return true;
>> +}
>> +
>> +static void steal_governor_loop(struct work_struct *work)
>> +{
>> + u64 curr_steal, delta_steal, delta_ns, steal_ratio;
>> + ktime_t now;
>> +
>> + now = ktime_get();
>> + delta_ns = ktime_to_ns(ktime_sub(now, sg_ctx.time));
>> +
>> + if (unlikely(delta_ns < NSEC_PER_MSEC)) {
>> + pr_err_ratelimited("work scheduled too soon delta_ns: %llu\n", delta_ns);
>> + goto requeue_work;
>
> [Severity: High]
> Can this create an infinite spinloop if interval_ms is set to 0?
>
> If interval_ms is 0, the delay is 0 jiffies. When the work executes, it will
> take this branch because delta_ns < 1ms. It jumps directly to requeue_work
> without updating sg_ctx.time, which schedules it again with 0 delay. This
> will continuously spin and consume an entire CPU core.
>
No. Interval can't be set to 0. Minimum is 100ms.
>> + }
>> +
>> + curr_steal = get_system_steal_time();
>> + delta_steal = curr_steal > sg_ctx.steal ? curr_steal - sg_ctx.steal : 0;
>> + sg_ctx.steal = curr_steal;
>> + sg_ctx.time = now;
>> +
>> + /*
>> + * steal_ratio = (delta_steal * 100*100)/(delta_ns * num_cpus())
>> + * To avoid possible overflow, divide the denominator early.
>> + * Note minimum interval is 100ms.
>> + */
>> + delta_ns = max_t(u64, div_u64(delta_ns * get_system_cpus(), 10000), 1);
>
> [Severity: Medium]
> Will this silently overflow the 64-bit multiplication if the VM is paused
> for an extended period?
>
> Since the multiplication of delta_ns * get_system_cpus() happens before
> the division, a large time delta (e.g., a large VM suspended for a long time)
> can overflow U64_MAX, leading to a wildly incorrect steal_ratio calculation.
>
Not a reason to fret about. It is way too long duration.
>> + steal_ratio = div64_u64(delta_steal, delta_ns);
>> +
>> + if (steal_ratio > sg_ctx.high_threshold)
>> + decrease_preferred_cpus();
>> + else if (steal_ratio <= sg_ctx.low_threshold)
>> + increase_preferred_cpus();
>> + /*
>> + * else: steal ratio is within bounds. Still do design checks so that
>> + * module restores to active if CPU hotplug breaks those assumptions.
>> + */
>> + if (!preferred_cpus_valid()) {
>
> [Severity: High]
> Does calling preferred_cpus_valid() here without cpus_read_lock() lead
> to a false teardown?
>
> If a hotplug operation is happening concurrently, the torn state read in
> preferred_cpus_valid() will cause the governor to erroneously shut down
> the policy loop permanently.
>
>> + restore_preferred_to_active();
>> + return;
>> + }
>> +
>> +requeue_work:
>> + schedule_delayed_work(&sg_ctx.work, sg_ctx.delay);
>
> [Severity: High]
> Does this immediately spin if sg_ctx.delay is 0, as mentioned above?
As said above, interval cannot be less than 100ms.
>
>> +}
>> +
>> static int __init steal_governor_init(void)
>> {
> [ ... ]
>> sg_ctx.delay = msecs_to_jiffies(sg_ctx.interval_ms);
>
> [Severity: High]
> Is it possible for sg_ctx.interval_ms to be 0 here?
>
> The module parameter interval_ms doesn't appear to be bounds-checked
> before this point, meaning it could be set to 0 and trigger the infinite
> workqueue spinloop.
As said above, interval cannot be less than 100ms.
>
>> + INIT_DELAYED_WORK(&sg_ctx.work, steal_governor_loop);
>> + sg_ctx.steal = get_system_steal_time();
>> + sg_ctx.time = ktime_get();
>> + schedule_delayed_work(&sg_ctx.work, sg_ctx.delay);
>> pr_info("enabled. interval: %ums, high_threshold: %u, low_threshold: %u\n",
>> sg_ctx.interval_ms, sg_ctx.high_threshold, sg_ctx.low_threshold);
>>
>
^ permalink raw reply [flat|nested] 32+ messages in thread
* [PATCH v14 13/13] virt/steal_governor: Enable the driver
2026-09-28 5:37 [PATCH v14 00/13] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff Shrikanth Hegde
` (11 preceding siblings ...)
2026-09-28 5:37 ` [PATCH v14 12/13] virt/steal_governor: Implement steal_governor policy loop Shrikanth Hegde
@ 2026-09-28 5:37 ` Shrikanth Hegde
2026-09-28 5:48 ` sashiko-bot
12 siblings, 1 reply; 32+ messages in thread
From: Shrikanth Hegde @ 2026-09-28 5:37 UTC (permalink / raw)
To: linux-kernel, mingo, peterz, juri.lelli, vincent.guittot,
yury.norov, kprateek.nayak, iii, corbet, meted, ynorov
Cc: sshegde, tglx, gregkh, pbonzini, seanjc, vschneid, huschle,
rostedt, dietmar.eggemann, maddy, srikar, hdanton, chleroy,
vineeth, frederic, arighi, pauld, christian.loehle, tj,
tommaso.cucinotta, maz, rafael, rdunlap, kernellwp, linux-doc,
jgross, virtualization, sunlightlinux
Provide a config option to enable the steal_governor driver.
Since the feature targets paravirtualized environments and requires SMP,
enforce those dependencies. The driver selects CONFIG_PREFERRED_CPU
for the core scheduler mechanisms to work.
It is recommended to build the driver as a module (m) instead of
built-in (y) due to the following reasons:
- Module parameters are read-only after initialization. Building as a
module allows updating these parameters by simply reloading the module.
Default module parameters cannot work in all configurations.
- The driver can be completely disabled by unloading the module.
- This feature works best when all VMs operate in a cooperative manner.
Requiring an explicit module load ensures intentional deployment
across all VMs by the system administrator.
Suggested-by: Yury Norov <yury.norov@gmail.com>
Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
---
drivers/virt/Kconfig | 17 +++++++++++++++++
drivers/virt/Makefile | 1 +
2 files changed, 18 insertions(+)
diff --git a/drivers/virt/Kconfig b/drivers/virt/Kconfig
index 52eb7e4ba71f..eeb84e578ddf 100644
--- a/drivers/virt/Kconfig
+++ b/drivers/virt/Kconfig
@@ -41,6 +41,23 @@ config FSL_HV_MANAGER
4) A kernel interface for receiving callbacks when a managed
partition shuts down.
+config STEAL_GOVERNOR
+ tristate "Dynamic vCPU management based on steal time"
+ depends on PARAVIRT && SMP
+ select PREFERRED_CPU
+ default m
+ help
+ This driver helps to reduce the steal time in paravirtualized
+ environments, thereby reducing vCPU preemption costs.
+
+ By default preferred CPUs will be same as active CPUs. Depending
+ on the steal time when steal_governor driver is enabled,
+ preferred CPUs could become subset of active CPUs.
+ More details are at: Documentation/driver-api/steal-governor.rst
+
+ It is recommended to build it as module and load the module
+ to enable it.
+
source "drivers/virt/vboxguest/Kconfig"
source "drivers/virt/nitro_enclaves/Kconfig"
diff --git a/drivers/virt/Makefile b/drivers/virt/Makefile
index f29901bd7820..05fb075ef5b8 100644
--- a/drivers/virt/Makefile
+++ b/drivers/virt/Makefile
@@ -5,6 +5,7 @@
obj-$(CONFIG_FSL_HV_MANAGER) += fsl_hypervisor.o
obj-$(CONFIG_VMGENID) += vmgenid.o
+obj-$(CONFIG_STEAL_GOVERNOR) += steal_governor.o
obj-y += vboxguest/
obj-$(CONFIG_NITRO_ENCLAVES) += nitro_enclaves/
--
2.52.0
^ permalink raw reply related [flat|nested] 32+ messages in thread* Re: [PATCH v14 13/13] virt/steal_governor: Enable the driver
2026-09-28 5:37 ` [PATCH v14 13/13] virt/steal_governor: Enable the driver Shrikanth Hegde
@ 2026-09-28 5:48 ` sashiko-bot
0 siblings, 0 replies; 32+ messages in thread
From: sashiko-bot @ 2026-09-28 5:48 UTC (permalink / raw)
To: Shrikanth Hegde; +Cc: virtualization, Michael S. Tsirkin, Eugenio Perez
> Provide a config option to enable the steal_governor driver.
> Since the feature targets paravirtualized environments and requires SMP,
> enforce those dependencies. The driver selects CONFIG_PREFERRED_CPU
> for the core scheduler mechanisms to work.
>
> It is recommended to build the driver as a module (m) instead of
> built-in (y) due to the following reasons:
>
> - Module parameters are read-only after initialization. Building as a
> module allows updating these parameters by simply reloading the module.
> Default module parameters cannot work in all configurations.
>
> - The driver can be completely disabled by unloading the module.
>
> - This feature works best when all VMs operate in a cooperative manner.
> Requiring an explicit module load ensures intentional deployment
> across all VMs by the system administrator.
>
> Suggested-by: Yury Norov <yury.norov@gmail.com>
> Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
Sashiko has reviewed this patch and found no issues. It looks great!
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260928053728.797539-1-sshegde@linux.ibm.com?part=13
^ permalink raw reply [flat|nested] 32+ messages in thread