Linux Power Management development
 help / color / mirror / Atom feed
* [PATCH RFC 0/1] cpuidle: add per-CPU latency_limit_ns sysfs attribute
@ 2026-09-15 13:13 Anthony Harivel
  2026-09-15 13:13 ` [PATCH 1/1] " Anthony Harivel
  0 siblings, 1 reply; 3+ messages in thread
From: Anthony Harivel @ 2026-09-15 13:13 UTC (permalink / raw)
  To: linux-pm; +Cc: rafael, daniel.lezcano, seanjc, pbonzini, kvm, Anthony Harivel

This is an RFC for a new per-CPU sysfs attribute that lets privileged
userspace set a governor-respected upper bound on idle state exit
latency.

This follows the discussion on KVM_CAP_CSTATE_POLICY (RFC v3,
Message-ID: 20260914143702.915401-1-aharivel@redhat.com) where Sean
and Paolo concluded that per-CPU cpuidle controls are the right
abstraction rather than a KVM-level interface [1][2].

== Problem ==

Cloud operators running mixed NFV workloads want to reduce energy
consumption by disabling halt-polling (halt_poll_ns=0). When vCPUs
enter HLT, the kernel cpuidle governor picks deep C-states (C6,
~133us wakeup) by default — good for power savings, bad for
latency-sensitive VMs.

Existing per-CPU controls (stateN/disable) work but require
knowledge of the C-state table for each CPU microarchitecture.
There is no latency-based per-CPU ceiling that works portably
across Intel/AMD/ARM.

== Solution ==

New sysfs attribute:
  /sys/devices/system/cpu/cpuN/cpuidle/latency_limit_ns

When set to a non-zero value, cpuidle_governor_latency_req() returns
the minimum of the existing PM QoS constraints and latency_limit_ns.
All governors (menu, TEO, haltpoll) automatically respect it — no
per-governor modifications needed.

  # Cap CPU 4 to ~C1 wakeup latency
  echo 2000 > /sys/devices/system/cpu/cpu4/cpuidle/latency_limit_ns

  # Remove limit
  echo 0 > /sys/devices/system/cpu/cpu4/cpuidle/latency_limit_ns

The interface is latency-based (nanoseconds) rather than
C-state-index-based, making it portable across microarchitectures
without per-uarch tuning — as Sean suggested [1].

== Integration ==

For the KVM/NFV use case: userspace (OpenStack Nova, libvirt, or a
simple script) pins vCPUs to pCPUs and writes latency_limit_ns on
those CPUs. No KVM or QEMU changes needed. This also works for
non-KVM use cases (DPDK, bare-metal NFV).

== Test results ==

Tested on Dell R640 (Intel Xeon Gold 5118, intel_idle driver,
states: POLL/C1/C1E/C6).

Feature selftest (7/7 pass):

  ok 1 sysfs attribute exists
  ok 2 default value is 0
  ok 3 write/readback
  ok 4 reset to 0
  ok 5 attribute on all 48 CPUs
  ok 6 per-CPU isolation
  ok 7 functional enforcement (deep state entered 1 time with limit)

Multi-VM demo (2 VMs, 60s, stock QEMU, same host):

  VM-A: CPUs 2,4 with latency_limit_ns=2000
  VM-B: CPUs 6,8 with no limit

                    VM-A (limit=2000ns)    VM-B (no limit)
  C1 usage delta:   +24031 / +22369        +7668 / +3044
  C1E usage delta:  +0 / +0               +6888 / +7784
  C6 usage delta:   +1 / +1               +6990 / +14843

VM-A stays in C1 (C1E and C6 completely blocked). VM-B freely
enters deep C-states. Same host, same moment, stock QEMU.

== Design notes ==

- latency_limit_ns defaults to 0 (no limit, existing behavior).
- Requires CAP_SYS_ADMIN to write (same as stateN/disable).
- Integrates at cpuidle_governor_latency_req() level, so it
  composes with existing PM QoS constraints (takes the minimum).
- Does NOT reuse forced_idle_latency_limit_ns — that field bypasses
  the governor entirely (used by play_idle_precise() for idle
  injection). latency_limit_ns is a governor ceiling, not a bypass.

Looking for feedback on the approach. Happy to add a selftest or
documentation patch in a follow-up.

[1] https://lore.kernel.org/kvm/aqgRj7mfDhCUqWw4@google.com/
[2] https://lore.kernel.org/kvm/ (Paolo's reply in same thread)

Anthony Harivel (1):
  cpuidle: add per-CPU latency_limit_ns sysfs attribute

 drivers/cpuidle/governor.c | 10 +++++++++-
 drivers/cpuidle/sysfs.c    | 36 ++++++++++++++++++++++++++++++++++++
 include/linux/cpuidle.h    |  1 +
 3 files changed, 46 insertions(+), 1 deletion(-)

-- 
2.55.0


^ permalink raw reply	[flat|nested] 3+ messages in thread

* [PATCH 1/1] cpuidle: add per-CPU latency_limit_ns sysfs attribute
  2026-09-15 13:13 [PATCH RFC 0/1] cpuidle: add per-CPU latency_limit_ns sysfs attribute Anthony Harivel
@ 2026-09-15 13:13 ` Anthony Harivel
  2026-09-25 17:36   ` Rafael J. Wysocki (Intel)
  0 siblings, 1 reply; 3+ messages in thread
From: Anthony Harivel @ 2026-09-15 13:13 UTC (permalink / raw)
  To: linux-pm; +Cc: rafael, daniel.lezcano, seanjc, pbonzini, kvm, Anthony Harivel

Add a per-CPU sysfs attribute at:
  /sys/devices/system/cpu/cpuN/cpuidle/latency_limit_ns

This allows privileged userspace to set a governor-respected upper
bound on the exit latency for idle state selection on a given CPU.
When set (non-zero), cpuidle_governor_latency_req() returns the
minimum of the existing PM QoS constraints and latency_limit_ns,
so all governors (menu, TEO, haltpoll) automatically respect it
without per-governor modifications.

Use case: cloud operators running mixed workloads can cap idle
depth on CPUs pinned to latency-sensitive VMs while allowing other
CPUs to enter deep C-states for energy savings. The interface is
latency-based (nanoseconds) rather than C-state-index-based, making
it portable across Intel/AMD/ARM without uarch-specific tuning.

Signed-off-by: Anthony Harivel <aharivel@redhat.com>
---
 drivers/cpuidle/governor.c | 10 +++++++++-
 drivers/cpuidle/sysfs.c    | 36 ++++++++++++++++++++++++++++++++++++
 include/linux/cpuidle.h    |  1 +
 3 files changed, 46 insertions(+), 1 deletion(-)

diff --git a/drivers/cpuidle/governor.c b/drivers/cpuidle/governor.c
index 5d0e7f78c6c5..4c6f77ce2028 100644
--- a/drivers/cpuidle/governor.c
+++ b/drivers/cpuidle/governor.c
@@ -112,6 +112,8 @@ s64 cpuidle_governor_latency_req(unsigned int cpu)
 	int device_req = dev_pm_qos_raw_resume_latency(device);
 	int global_req = cpu_latency_qos_limit();
 	int global_wake_req = cpu_wakeup_latency_qos_limit();
+	struct cpuidle_device *dev;
+	s64 result;
 
 	if (global_req > global_wake_req)
 		global_req = global_wake_req;
@@ -119,5 +121,11 @@ s64 cpuidle_governor_latency_req(unsigned int cpu)
 	if (device_req > global_req)
 		device_req = global_req;
 
-	return (s64)device_req * NSEC_PER_USEC;
+	result = (s64)device_req * NSEC_PER_USEC;
+
+	dev = per_cpu(cpuidle_devices, cpu);
+	if (dev && dev->latency_limit_ns && dev->latency_limit_ns < result)
+		result = dev->latency_limit_ns;
+
+	return result;
 }
diff --git a/drivers/cpuidle/sysfs.c b/drivers/cpuidle/sysfs.c
index b81d22479234..d060a4b7facc 100644
--- a/drivers/cpuidle/sysfs.c
+++ b/drivers/cpuidle/sysfs.c
@@ -207,8 +207,44 @@ static void cpuidle_sysfs_release(struct kobject *kobj)
 	complete(&kdev->kobj_unregister);
 }
 
+static ssize_t show_latency_limit_ns(struct cpuidle_device *dev, char *buf)
+{
+	return sysfs_emit(buf, "%llu\n", dev->latency_limit_ns);
+}
+
+static ssize_t store_latency_limit_ns(struct cpuidle_device *dev,
+				       const char *buf, size_t count)
+{
+	u64 value;
+	int err;
+
+	if (!capable(CAP_SYS_ADMIN))
+		return -EPERM;
+
+	err = kstrtou64(buf, 0, &value);
+	if (err)
+		return err;
+
+	dev->latency_limit_ns = value;
+
+	return count;
+}
+
+static struct cpuidle_attr attr_latency_limit_ns = {
+	.attr = { .name = "latency_limit_ns", .mode = 0644 },
+	.show = show_latency_limit_ns,
+	.store = store_latency_limit_ns,
+};
+
+static struct attribute *cpuidle_device_default_attrs[] = {
+	&attr_latency_limit_ns.attr,
+	NULL,
+};
+ATTRIBUTE_GROUPS(cpuidle_device_default);
+
 static const struct kobj_type ktype_cpuidle = {
 	.sysfs_ops = &cpuidle_sysfs_ops,
+	.default_groups = cpuidle_device_default_groups,
 	.release = cpuidle_sysfs_release,
 };
 
diff --git a/include/linux/cpuidle.h b/include/linux/cpuidle.h
index a2485348def3..f70d6288e13d 100644
--- a/include/linux/cpuidle.h
+++ b/include/linux/cpuidle.h
@@ -101,6 +101,7 @@ struct cpuidle_device {
 	u64			last_residency_ns;
 	u64			poll_limit_ns;
 	u64			forced_idle_latency_limit_ns;
+	u64			latency_limit_ns;
 	struct cpuidle_state_usage	states_usage[CPUIDLE_STATE_MAX];
 	struct cpuidle_state_kobj *kobjs[CPUIDLE_STATE_MAX];
 	struct cpuidle_driver_kobj *kobj_driver;
-- 
2.55.0


^ permalink raw reply related	[flat|nested] 3+ messages in thread

* Re: [PATCH 1/1] cpuidle: add per-CPU latency_limit_ns sysfs attribute
  2026-09-15 13:13 ` [PATCH 1/1] " Anthony Harivel
@ 2026-09-25 17:36   ` Rafael J. Wysocki (Intel)
  0 siblings, 0 replies; 3+ messages in thread
From: Rafael J. Wysocki (Intel) @ 2026-09-25 17:36 UTC (permalink / raw)
  To: Anthony Harivel; +Cc: linux-pm, rafael, daniel.lezcano, seanjc, pbonzini, kvm

On Tue, Sep 15, 2026 at 3:13 PM Anthony Harivel <aharivel@redhat.com> wrote:
>
> Add a per-CPU sysfs attribute at:
>   /sys/devices/system/cpu/cpuN/cpuidle/latency_limit_ns
>
> This allows privileged userspace to set a governor-respected upper
> bound on the exit latency for idle state selection on a given CPU.

Which is already present in the form of

/sys/devices/system/cpu/cpuN/power/pm_qos_resume_latency_us

See Documentation/admin-guide/pm/cpuidle.rst for more information.

> When set (non-zero), cpuidle_governor_latency_req() returns the
> minimum of the existing PM QoS constraints and latency_limit_ns,
> so all governors (menu, TEO, haltpoll) automatically respect it
> without per-governor modifications.
>
> Use case: cloud operators running mixed workloads can cap idle
> depth on CPUs pinned to latency-sensitive VMs while allowing other
> CPUs to enter deep C-states for energy savings. The interface is
> latency-based (nanoseconds) rather than C-state-index-based, making
> it portable across Intel/AMD/ARM without uarch-specific tuning.
>
> Signed-off-by: Anthony Harivel <aharivel@redhat.com>
> ---
>  drivers/cpuidle/governor.c | 10 +++++++++-
>  drivers/cpuidle/sysfs.c    | 36 ++++++++++++++++++++++++++++++++++++
>  include/linux/cpuidle.h    |  1 +
>  3 files changed, 46 insertions(+), 1 deletion(-)
>
> diff --git a/drivers/cpuidle/governor.c b/drivers/cpuidle/governor.c
> index 5d0e7f78c6c5..4c6f77ce2028 100644
> --- a/drivers/cpuidle/governor.c
> +++ b/drivers/cpuidle/governor.c
> @@ -112,6 +112,8 @@ s64 cpuidle_governor_latency_req(unsigned int cpu)
>         int device_req = dev_pm_qos_raw_resume_latency(device);
>         int global_req = cpu_latency_qos_limit();
>         int global_wake_req = cpu_wakeup_latency_qos_limit();
> +       struct cpuidle_device *dev;
> +       s64 result;
>
>         if (global_req > global_wake_req)
>                 global_req = global_wake_req;
> @@ -119,5 +121,11 @@ s64 cpuidle_governor_latency_req(unsigned int cpu)
>         if (device_req > global_req)
>                 device_req = global_req;
>
> -       return (s64)device_req * NSEC_PER_USEC;
> +       result = (s64)device_req * NSEC_PER_USEC;
> +
> +       dev = per_cpu(cpuidle_devices, cpu);
> +       if (dev && dev->latency_limit_ns && dev->latency_limit_ns < result)
> +               result = dev->latency_limit_ns;
> +
> +       return result;
>  }
> diff --git a/drivers/cpuidle/sysfs.c b/drivers/cpuidle/sysfs.c
> index b81d22479234..d060a4b7facc 100644
> --- a/drivers/cpuidle/sysfs.c
> +++ b/drivers/cpuidle/sysfs.c
> @@ -207,8 +207,44 @@ static void cpuidle_sysfs_release(struct kobject *kobj)
>         complete(&kdev->kobj_unregister);
>  }
>
> +static ssize_t show_latency_limit_ns(struct cpuidle_device *dev, char *buf)
> +{
> +       return sysfs_emit(buf, "%llu\n", dev->latency_limit_ns);
> +}
> +
> +static ssize_t store_latency_limit_ns(struct cpuidle_device *dev,
> +                                      const char *buf, size_t count)
> +{
> +       u64 value;
> +       int err;
> +
> +       if (!capable(CAP_SYS_ADMIN))
> +               return -EPERM;
> +
> +       err = kstrtou64(buf, 0, &value);
> +       if (err)
> +               return err;
> +
> +       dev->latency_limit_ns = value;
> +
> +       return count;
> +}
> +
> +static struct cpuidle_attr attr_latency_limit_ns = {
> +       .attr = { .name = "latency_limit_ns", .mode = 0644 },
> +       .show = show_latency_limit_ns,
> +       .store = store_latency_limit_ns,
> +};
> +
> +static struct attribute *cpuidle_device_default_attrs[] = {
> +       &attr_latency_limit_ns.attr,
> +       NULL,
> +};
> +ATTRIBUTE_GROUPS(cpuidle_device_default);
> +
>  static const struct kobj_type ktype_cpuidle = {
>         .sysfs_ops = &cpuidle_sysfs_ops,
> +       .default_groups = cpuidle_device_default_groups,
>         .release = cpuidle_sysfs_release,
>  };
>
> diff --git a/include/linux/cpuidle.h b/include/linux/cpuidle.h
> index a2485348def3..f70d6288e13d 100644
> --- a/include/linux/cpuidle.h
> +++ b/include/linux/cpuidle.h
> @@ -101,6 +101,7 @@ struct cpuidle_device {
>         u64                     last_residency_ns;
>         u64                     poll_limit_ns;
>         u64                     forced_idle_latency_limit_ns;
> +       u64                     latency_limit_ns;
>         struct cpuidle_state_usage      states_usage[CPUIDLE_STATE_MAX];
>         struct cpuidle_state_kobj *kobjs[CPUIDLE_STATE_MAX];
>         struct cpuidle_driver_kobj *kobj_driver;
> --
> 2.55.0
>
>

^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-09-25 17:37 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-15 13:13 [PATCH RFC 0/1] cpuidle: add per-CPU latency_limit_ns sysfs attribute Anthony Harivel
2026-09-15 13:13 ` [PATCH 1/1] " Anthony Harivel
2026-09-25 17:36   ` Rafael J. Wysocki (Intel)

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox