From: Anthony Harivel <aharivel@redhat.com>
To: linux-pm@vger.kernel.org
Cc: rafael@kernel.org, daniel.lezcano@linaro.org, seanjc@google.com,
pbonzini@redhat.com, kvm@vger.kernel.org,
Anthony Harivel <aharivel@redhat.com>
Subject: [PATCH RFC 0/1] cpuidle: add per-CPU latency_limit_ns sysfs attribute
Date: Tue, 15 Sep 2026 15:13:06 +0200 [thread overview]
Message-ID: <20260915131310.1053834-1-aharivel@redhat.com> (raw)
This is an RFC for a new per-CPU sysfs attribute that lets privileged
userspace set a governor-respected upper bound on idle state exit
latency.
This follows the discussion on KVM_CAP_CSTATE_POLICY (RFC v3,
Message-ID: 20260914143702.915401-1-aharivel@redhat.com) where Sean
and Paolo concluded that per-CPU cpuidle controls are the right
abstraction rather than a KVM-level interface [1][2].
== Problem ==
Cloud operators running mixed NFV workloads want to reduce energy
consumption by disabling halt-polling (halt_poll_ns=0). When vCPUs
enter HLT, the kernel cpuidle governor picks deep C-states (C6,
~133us wakeup) by default — good for power savings, bad for
latency-sensitive VMs.
Existing per-CPU controls (stateN/disable) work but require
knowledge of the C-state table for each CPU microarchitecture.
There is no latency-based per-CPU ceiling that works portably
across Intel/AMD/ARM.
== Solution ==
New sysfs attribute:
/sys/devices/system/cpu/cpuN/cpuidle/latency_limit_ns
When set to a non-zero value, cpuidle_governor_latency_req() returns
the minimum of the existing PM QoS constraints and latency_limit_ns.
All governors (menu, TEO, haltpoll) automatically respect it — no
per-governor modifications needed.
# Cap CPU 4 to ~C1 wakeup latency
echo 2000 > /sys/devices/system/cpu/cpu4/cpuidle/latency_limit_ns
# Remove limit
echo 0 > /sys/devices/system/cpu/cpu4/cpuidle/latency_limit_ns
The interface is latency-based (nanoseconds) rather than
C-state-index-based, making it portable across microarchitectures
without per-uarch tuning — as Sean suggested [1].
== Integration ==
For the KVM/NFV use case: userspace (OpenStack Nova, libvirt, or a
simple script) pins vCPUs to pCPUs and writes latency_limit_ns on
those CPUs. No KVM or QEMU changes needed. This also works for
non-KVM use cases (DPDK, bare-metal NFV).
== Test results ==
Tested on Dell R640 (Intel Xeon Gold 5118, intel_idle driver,
states: POLL/C1/C1E/C6).
Feature selftest (7/7 pass):
ok 1 sysfs attribute exists
ok 2 default value is 0
ok 3 write/readback
ok 4 reset to 0
ok 5 attribute on all 48 CPUs
ok 6 per-CPU isolation
ok 7 functional enforcement (deep state entered 1 time with limit)
Multi-VM demo (2 VMs, 60s, stock QEMU, same host):
VM-A: CPUs 2,4 with latency_limit_ns=2000
VM-B: CPUs 6,8 with no limit
VM-A (limit=2000ns) VM-B (no limit)
C1 usage delta: +24031 / +22369 +7668 / +3044
C1E usage delta: +0 / +0 +6888 / +7784
C6 usage delta: +1 / +1 +6990 / +14843
VM-A stays in C1 (C1E and C6 completely blocked). VM-B freely
enters deep C-states. Same host, same moment, stock QEMU.
== Design notes ==
- latency_limit_ns defaults to 0 (no limit, existing behavior).
- Requires CAP_SYS_ADMIN to write (same as stateN/disable).
- Integrates at cpuidle_governor_latency_req() level, so it
composes with existing PM QoS constraints (takes the minimum).
- Does NOT reuse forced_idle_latency_limit_ns — that field bypasses
the governor entirely (used by play_idle_precise() for idle
injection). latency_limit_ns is a governor ceiling, not a bypass.
Looking for feedback on the approach. Happy to add a selftest or
documentation patch in a follow-up.
[1] https://lore.kernel.org/kvm/aqgRj7mfDhCUqWw4@google.com/
[2] https://lore.kernel.org/kvm/ (Paolo's reply in same thread)
Anthony Harivel (1):
cpuidle: add per-CPU latency_limit_ns sysfs attribute
drivers/cpuidle/governor.c | 10 +++++++++-
drivers/cpuidle/sysfs.c | 36 ++++++++++++++++++++++++++++++++++++
include/linux/cpuidle.h | 1 +
3 files changed, 46 insertions(+), 1 deletion(-)
--
2.55.0
next reply other threads:[~2026-09-15 13:13 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-15 13:13 Anthony Harivel [this message]
2026-09-15 13:13 ` [PATCH 1/1] cpuidle: add per-CPU latency_limit_ns sysfs attribute Anthony Harivel
2026-09-25 17:36 ` Rafael J. Wysocki (Intel)
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260915131310.1053834-1-aharivel@redhat.com \
--to=aharivel@redhat.com \
--cc=daniel.lezcano@linaro.org \
--cc=kvm@vger.kernel.org \
--cc=linux-pm@vger.kernel.org \
--cc=pbonzini@redhat.com \
--cc=rafael@kernel.org \
--cc=seanjc@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox