* [PATCH RFC 0/1] cpuidle: add per-CPU latency_limit_ns sysfs attribute
@ 2026-09-15 13:13 Anthony Harivel
2026-09-15 13:13 ` [PATCH 1/1] " Anthony Harivel
0 siblings, 1 reply; 3+ messages in thread
From: Anthony Harivel @ 2026-09-15 13:13 UTC (permalink / raw)
To: linux-pm; +Cc: rafael, daniel.lezcano, seanjc, pbonzini, kvm, Anthony Harivel
This is an RFC for a new per-CPU sysfs attribute that lets privileged
userspace set a governor-respected upper bound on idle state exit
latency.
This follows the discussion on KVM_CAP_CSTATE_POLICY (RFC v3,
Message-ID: 20260914143702.915401-1-aharivel@redhat.com) where Sean
and Paolo concluded that per-CPU cpuidle controls are the right
abstraction rather than a KVM-level interface [1][2].
== Problem ==
Cloud operators running mixed NFV workloads want to reduce energy
consumption by disabling halt-polling (halt_poll_ns=0). When vCPUs
enter HLT, the kernel cpuidle governor picks deep C-states (C6,
~133us wakeup) by default — good for power savings, bad for
latency-sensitive VMs.
Existing per-CPU controls (stateN/disable) work but require
knowledge of the C-state table for each CPU microarchitecture.
There is no latency-based per-CPU ceiling that works portably
across Intel/AMD/ARM.
== Solution ==
New sysfs attribute:
/sys/devices/system/cpu/cpuN/cpuidle/latency_limit_ns
When set to a non-zero value, cpuidle_governor_latency_req() returns
the minimum of the existing PM QoS constraints and latency_limit_ns.
All governors (menu, TEO, haltpoll) automatically respect it — no
per-governor modifications needed.
# Cap CPU 4 to ~C1 wakeup latency
echo 2000 > /sys/devices/system/cpu/cpu4/cpuidle/latency_limit_ns
# Remove limit
echo 0 > /sys/devices/system/cpu/cpu4/cpuidle/latency_limit_ns
The interface is latency-based (nanoseconds) rather than
C-state-index-based, making it portable across microarchitectures
without per-uarch tuning — as Sean suggested [1].
== Integration ==
For the KVM/NFV use case: userspace (OpenStack Nova, libvirt, or a
simple script) pins vCPUs to pCPUs and writes latency_limit_ns on
those CPUs. No KVM or QEMU changes needed. This also works for
non-KVM use cases (DPDK, bare-metal NFV).
== Test results ==
Tested on Dell R640 (Intel Xeon Gold 5118, intel_idle driver,
states: POLL/C1/C1E/C6).
Feature selftest (7/7 pass):
ok 1 sysfs attribute exists
ok 2 default value is 0
ok 3 write/readback
ok 4 reset to 0
ok 5 attribute on all 48 CPUs
ok 6 per-CPU isolation
ok 7 functional enforcement (deep state entered 1 time with limit)
Multi-VM demo (2 VMs, 60s, stock QEMU, same host):
VM-A: CPUs 2,4 with latency_limit_ns=2000
VM-B: CPUs 6,8 with no limit
VM-A (limit=2000ns) VM-B (no limit)
C1 usage delta: +24031 / +22369 +7668 / +3044
C1E usage delta: +0 / +0 +6888 / +7784
C6 usage delta: +1 / +1 +6990 / +14843
VM-A stays in C1 (C1E and C6 completely blocked). VM-B freely
enters deep C-states. Same host, same moment, stock QEMU.
== Design notes ==
- latency_limit_ns defaults to 0 (no limit, existing behavior).
- Requires CAP_SYS_ADMIN to write (same as stateN/disable).
- Integrates at cpuidle_governor_latency_req() level, so it
composes with existing PM QoS constraints (takes the minimum).
- Does NOT reuse forced_idle_latency_limit_ns — that field bypasses
the governor entirely (used by play_idle_precise() for idle
injection). latency_limit_ns is a governor ceiling, not a bypass.
Looking for feedback on the approach. Happy to add a selftest or
documentation patch in a follow-up.
[1] https://lore.kernel.org/kvm/aqgRj7mfDhCUqWw4@google.com/
[2] https://lore.kernel.org/kvm/ (Paolo's reply in same thread)
Anthony Harivel (1):
cpuidle: add per-CPU latency_limit_ns sysfs attribute
drivers/cpuidle/governor.c | 10 +++++++++-
drivers/cpuidle/sysfs.c | 36 ++++++++++++++++++++++++++++++++++++
include/linux/cpuidle.h | 1 +
3 files changed, 46 insertions(+), 1 deletion(-)
--
2.55.0
^ permalink raw reply [flat|nested] 3+ messages in thread* [PATCH 1/1] cpuidle: add per-CPU latency_limit_ns sysfs attribute
2026-09-15 13:13 [PATCH RFC 0/1] cpuidle: add per-CPU latency_limit_ns sysfs attribute Anthony Harivel
@ 2026-09-15 13:13 ` Anthony Harivel
2026-09-25 17:36 ` Rafael J. Wysocki (Intel)
0 siblings, 1 reply; 3+ messages in thread
From: Anthony Harivel @ 2026-09-15 13:13 UTC (permalink / raw)
To: linux-pm; +Cc: rafael, daniel.lezcano, seanjc, pbonzini, kvm, Anthony Harivel
Add a per-CPU sysfs attribute at:
/sys/devices/system/cpu/cpuN/cpuidle/latency_limit_ns
This allows privileged userspace to set a governor-respected upper
bound on the exit latency for idle state selection on a given CPU.
When set (non-zero), cpuidle_governor_latency_req() returns the
minimum of the existing PM QoS constraints and latency_limit_ns,
so all governors (menu, TEO, haltpoll) automatically respect it
without per-governor modifications.
Use case: cloud operators running mixed workloads can cap idle
depth on CPUs pinned to latency-sensitive VMs while allowing other
CPUs to enter deep C-states for energy savings. The interface is
latency-based (nanoseconds) rather than C-state-index-based, making
it portable across Intel/AMD/ARM without uarch-specific tuning.
Signed-off-by: Anthony Harivel <aharivel@redhat.com>
---
drivers/cpuidle/governor.c | 10 +++++++++-
drivers/cpuidle/sysfs.c | 36 ++++++++++++++++++++++++++++++++++++
include/linux/cpuidle.h | 1 +
3 files changed, 46 insertions(+), 1 deletion(-)
diff --git a/drivers/cpuidle/governor.c b/drivers/cpuidle/governor.c
index 5d0e7f78c6c5..4c6f77ce2028 100644
--- a/drivers/cpuidle/governor.c
+++ b/drivers/cpuidle/governor.c
@@ -112,6 +112,8 @@ s64 cpuidle_governor_latency_req(unsigned int cpu)
int device_req = dev_pm_qos_raw_resume_latency(device);
int global_req = cpu_latency_qos_limit();
int global_wake_req = cpu_wakeup_latency_qos_limit();
+ struct cpuidle_device *dev;
+ s64 result;
if (global_req > global_wake_req)
global_req = global_wake_req;
@@ -119,5 +121,11 @@ s64 cpuidle_governor_latency_req(unsigned int cpu)
if (device_req > global_req)
device_req = global_req;
- return (s64)device_req * NSEC_PER_USEC;
+ result = (s64)device_req * NSEC_PER_USEC;
+
+ dev = per_cpu(cpuidle_devices, cpu);
+ if (dev && dev->latency_limit_ns && dev->latency_limit_ns < result)
+ result = dev->latency_limit_ns;
+
+ return result;
}
diff --git a/drivers/cpuidle/sysfs.c b/drivers/cpuidle/sysfs.c
index b81d22479234..d060a4b7facc 100644
--- a/drivers/cpuidle/sysfs.c
+++ b/drivers/cpuidle/sysfs.c
@@ -207,8 +207,44 @@ static void cpuidle_sysfs_release(struct kobject *kobj)
complete(&kdev->kobj_unregister);
}
+static ssize_t show_latency_limit_ns(struct cpuidle_device *dev, char *buf)
+{
+ return sysfs_emit(buf, "%llu\n", dev->latency_limit_ns);
+}
+
+static ssize_t store_latency_limit_ns(struct cpuidle_device *dev,
+ const char *buf, size_t count)
+{
+ u64 value;
+ int err;
+
+ if (!capable(CAP_SYS_ADMIN))
+ return -EPERM;
+
+ err = kstrtou64(buf, 0, &value);
+ if (err)
+ return err;
+
+ dev->latency_limit_ns = value;
+
+ return count;
+}
+
+static struct cpuidle_attr attr_latency_limit_ns = {
+ .attr = { .name = "latency_limit_ns", .mode = 0644 },
+ .show = show_latency_limit_ns,
+ .store = store_latency_limit_ns,
+};
+
+static struct attribute *cpuidle_device_default_attrs[] = {
+ &attr_latency_limit_ns.attr,
+ NULL,
+};
+ATTRIBUTE_GROUPS(cpuidle_device_default);
+
static const struct kobj_type ktype_cpuidle = {
.sysfs_ops = &cpuidle_sysfs_ops,
+ .default_groups = cpuidle_device_default_groups,
.release = cpuidle_sysfs_release,
};
diff --git a/include/linux/cpuidle.h b/include/linux/cpuidle.h
index a2485348def3..f70d6288e13d 100644
--- a/include/linux/cpuidle.h
+++ b/include/linux/cpuidle.h
@@ -101,6 +101,7 @@ struct cpuidle_device {
u64 last_residency_ns;
u64 poll_limit_ns;
u64 forced_idle_latency_limit_ns;
+ u64 latency_limit_ns;
struct cpuidle_state_usage states_usage[CPUIDLE_STATE_MAX];
struct cpuidle_state_kobj *kobjs[CPUIDLE_STATE_MAX];
struct cpuidle_driver_kobj *kobj_driver;
--
2.55.0
^ permalink raw reply related [flat|nested] 3+ messages in thread* Re: [PATCH 1/1] cpuidle: add per-CPU latency_limit_ns sysfs attribute
2026-09-15 13:13 ` [PATCH 1/1] " Anthony Harivel
@ 2026-09-25 17:36 ` Rafael J. Wysocki (Intel)
0 siblings, 0 replies; 3+ messages in thread
From: Rafael J. Wysocki (Intel) @ 2026-09-25 17:36 UTC (permalink / raw)
To: Anthony Harivel; +Cc: linux-pm, rafael, daniel.lezcano, seanjc, pbonzini, kvm
On Tue, Sep 15, 2026 at 3:13 PM Anthony Harivel <aharivel@redhat.com> wrote:
>
> Add a per-CPU sysfs attribute at:
> /sys/devices/system/cpu/cpuN/cpuidle/latency_limit_ns
>
> This allows privileged userspace to set a governor-respected upper
> bound on the exit latency for idle state selection on a given CPU.
Which is already present in the form of
/sys/devices/system/cpu/cpuN/power/pm_qos_resume_latency_us
See Documentation/admin-guide/pm/cpuidle.rst for more information.
> When set (non-zero), cpuidle_governor_latency_req() returns the
> minimum of the existing PM QoS constraints and latency_limit_ns,
> so all governors (menu, TEO, haltpoll) automatically respect it
> without per-governor modifications.
>
> Use case: cloud operators running mixed workloads can cap idle
> depth on CPUs pinned to latency-sensitive VMs while allowing other
> CPUs to enter deep C-states for energy savings. The interface is
> latency-based (nanoseconds) rather than C-state-index-based, making
> it portable across Intel/AMD/ARM without uarch-specific tuning.
>
> Signed-off-by: Anthony Harivel <aharivel@redhat.com>
> ---
> drivers/cpuidle/governor.c | 10 +++++++++-
> drivers/cpuidle/sysfs.c | 36 ++++++++++++++++++++++++++++++++++++
> include/linux/cpuidle.h | 1 +
> 3 files changed, 46 insertions(+), 1 deletion(-)
>
> diff --git a/drivers/cpuidle/governor.c b/drivers/cpuidle/governor.c
> index 5d0e7f78c6c5..4c6f77ce2028 100644
> --- a/drivers/cpuidle/governor.c
> +++ b/drivers/cpuidle/governor.c
> @@ -112,6 +112,8 @@ s64 cpuidle_governor_latency_req(unsigned int cpu)
> int device_req = dev_pm_qos_raw_resume_latency(device);
> int global_req = cpu_latency_qos_limit();
> int global_wake_req = cpu_wakeup_latency_qos_limit();
> + struct cpuidle_device *dev;
> + s64 result;
>
> if (global_req > global_wake_req)
> global_req = global_wake_req;
> @@ -119,5 +121,11 @@ s64 cpuidle_governor_latency_req(unsigned int cpu)
> if (device_req > global_req)
> device_req = global_req;
>
> - return (s64)device_req * NSEC_PER_USEC;
> + result = (s64)device_req * NSEC_PER_USEC;
> +
> + dev = per_cpu(cpuidle_devices, cpu);
> + if (dev && dev->latency_limit_ns && dev->latency_limit_ns < result)
> + result = dev->latency_limit_ns;
> +
> + return result;
> }
> diff --git a/drivers/cpuidle/sysfs.c b/drivers/cpuidle/sysfs.c
> index b81d22479234..d060a4b7facc 100644
> --- a/drivers/cpuidle/sysfs.c
> +++ b/drivers/cpuidle/sysfs.c
> @@ -207,8 +207,44 @@ static void cpuidle_sysfs_release(struct kobject *kobj)
> complete(&kdev->kobj_unregister);
> }
>
> +static ssize_t show_latency_limit_ns(struct cpuidle_device *dev, char *buf)
> +{
> + return sysfs_emit(buf, "%llu\n", dev->latency_limit_ns);
> +}
> +
> +static ssize_t store_latency_limit_ns(struct cpuidle_device *dev,
> + const char *buf, size_t count)
> +{
> + u64 value;
> + int err;
> +
> + if (!capable(CAP_SYS_ADMIN))
> + return -EPERM;
> +
> + err = kstrtou64(buf, 0, &value);
> + if (err)
> + return err;
> +
> + dev->latency_limit_ns = value;
> +
> + return count;
> +}
> +
> +static struct cpuidle_attr attr_latency_limit_ns = {
> + .attr = { .name = "latency_limit_ns", .mode = 0644 },
> + .show = show_latency_limit_ns,
> + .store = store_latency_limit_ns,
> +};
> +
> +static struct attribute *cpuidle_device_default_attrs[] = {
> + &attr_latency_limit_ns.attr,
> + NULL,
> +};
> +ATTRIBUTE_GROUPS(cpuidle_device_default);
> +
> static const struct kobj_type ktype_cpuidle = {
> .sysfs_ops = &cpuidle_sysfs_ops,
> + .default_groups = cpuidle_device_default_groups,
> .release = cpuidle_sysfs_release,
> };
>
> diff --git a/include/linux/cpuidle.h b/include/linux/cpuidle.h
> index a2485348def3..f70d6288e13d 100644
> --- a/include/linux/cpuidle.h
> +++ b/include/linux/cpuidle.h
> @@ -101,6 +101,7 @@ struct cpuidle_device {
> u64 last_residency_ns;
> u64 poll_limit_ns;
> u64 forced_idle_latency_limit_ns;
> + u64 latency_limit_ns;
> struct cpuidle_state_usage states_usage[CPUIDLE_STATE_MAX];
> struct cpuidle_state_kobj *kobjs[CPUIDLE_STATE_MAX];
> struct cpuidle_driver_kobj *kobj_driver;
> --
> 2.55.0
>
>
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-09-25 17:37 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-15 13:13 [PATCH RFC 0/1] cpuidle: add per-CPU latency_limit_ns sysfs attribute Anthony Harivel
2026-09-15 13:13 ` [PATCH 1/1] " Anthony Harivel
2026-09-25 17:36 ` Rafael J. Wysocki (Intel)
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox