Linux Power Management development
 help / color / mirror / Atom feed
From: Anthony Harivel <aharivel@redhat.com>
To: linux-pm@vger.kernel.org
Cc: rafael@kernel.org, daniel.lezcano@linaro.org, seanjc@google.com,
	pbonzini@redhat.com, kvm@vger.kernel.org,
	Anthony Harivel <aharivel@redhat.com>
Subject: [PATCH RFC 0/1] cpuidle: add per-CPU latency_limit_ns sysfs attribute
Date: Tue, 15 Sep 2026 15:13:06 +0200	[thread overview]
Message-ID: <20260915131310.1053834-1-aharivel@redhat.com> (raw)

This is an RFC for a new per-CPU sysfs attribute that lets privileged
userspace set a governor-respected upper bound on idle state exit
latency.

This follows the discussion on KVM_CAP_CSTATE_POLICY (RFC v3,
Message-ID: 20260914143702.915401-1-aharivel@redhat.com) where Sean
and Paolo concluded that per-CPU cpuidle controls are the right
abstraction rather than a KVM-level interface [1][2].

== Problem ==

Cloud operators running mixed NFV workloads want to reduce energy
consumption by disabling halt-polling (halt_poll_ns=0). When vCPUs
enter HLT, the kernel cpuidle governor picks deep C-states (C6,
~133us wakeup) by default — good for power savings, bad for
latency-sensitive VMs.

Existing per-CPU controls (stateN/disable) work but require
knowledge of the C-state table for each CPU microarchitecture.
There is no latency-based per-CPU ceiling that works portably
across Intel/AMD/ARM.

== Solution ==

New sysfs attribute:
  /sys/devices/system/cpu/cpuN/cpuidle/latency_limit_ns

When set to a non-zero value, cpuidle_governor_latency_req() returns
the minimum of the existing PM QoS constraints and latency_limit_ns.
All governors (menu, TEO, haltpoll) automatically respect it — no
per-governor modifications needed.

  # Cap CPU 4 to ~C1 wakeup latency
  echo 2000 > /sys/devices/system/cpu/cpu4/cpuidle/latency_limit_ns

  # Remove limit
  echo 0 > /sys/devices/system/cpu/cpu4/cpuidle/latency_limit_ns

The interface is latency-based (nanoseconds) rather than
C-state-index-based, making it portable across microarchitectures
without per-uarch tuning — as Sean suggested [1].

== Integration ==

For the KVM/NFV use case: userspace (OpenStack Nova, libvirt, or a
simple script) pins vCPUs to pCPUs and writes latency_limit_ns on
those CPUs. No KVM or QEMU changes needed. This also works for
non-KVM use cases (DPDK, bare-metal NFV).

== Test results ==

Tested on Dell R640 (Intel Xeon Gold 5118, intel_idle driver,
states: POLL/C1/C1E/C6).

Feature selftest (7/7 pass):

  ok 1 sysfs attribute exists
  ok 2 default value is 0
  ok 3 write/readback
  ok 4 reset to 0
  ok 5 attribute on all 48 CPUs
  ok 6 per-CPU isolation
  ok 7 functional enforcement (deep state entered 1 time with limit)

Multi-VM demo (2 VMs, 60s, stock QEMU, same host):

  VM-A: CPUs 2,4 with latency_limit_ns=2000
  VM-B: CPUs 6,8 with no limit

                    VM-A (limit=2000ns)    VM-B (no limit)
  C1 usage delta:   +24031 / +22369        +7668 / +3044
  C1E usage delta:  +0 / +0               +6888 / +7784
  C6 usage delta:   +1 / +1               +6990 / +14843

VM-A stays in C1 (C1E and C6 completely blocked). VM-B freely
enters deep C-states. Same host, same moment, stock QEMU.

== Design notes ==

- latency_limit_ns defaults to 0 (no limit, existing behavior).
- Requires CAP_SYS_ADMIN to write (same as stateN/disable).
- Integrates at cpuidle_governor_latency_req() level, so it
  composes with existing PM QoS constraints (takes the minimum).
- Does NOT reuse forced_idle_latency_limit_ns — that field bypasses
  the governor entirely (used by play_idle_precise() for idle
  injection). latency_limit_ns is a governor ceiling, not a bypass.

Looking for feedback on the approach. Happy to add a selftest or
documentation patch in a follow-up.

[1] https://lore.kernel.org/kvm/aqgRj7mfDhCUqWw4@google.com/
[2] https://lore.kernel.org/kvm/ (Paolo's reply in same thread)

Anthony Harivel (1):
  cpuidle: add per-CPU latency_limit_ns sysfs attribute

 drivers/cpuidle/governor.c | 10 +++++++++-
 drivers/cpuidle/sysfs.c    | 36 ++++++++++++++++++++++++++++++++++++
 include/linux/cpuidle.h    |  1 +
 3 files changed, 46 insertions(+), 1 deletion(-)

-- 
2.55.0


             reply	other threads:[~2026-09-15 13:13 UTC|newest]

Thread overview: 3+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-15 13:13 Anthony Harivel [this message]
2026-09-15 13:13 ` [PATCH 1/1] cpuidle: add per-CPU latency_limit_ns sysfs attribute Anthony Harivel
2026-09-25 17:36   ` Rafael J. Wysocki (Intel)

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260915131310.1053834-1-aharivel@redhat.com \
    --to=aharivel@redhat.com \
    --cc=daniel.lezcano@linaro.org \
    --cc=kvm@vger.kernel.org \
    --cc=linux-pm@vger.kernel.org \
    --cc=pbonzini@redhat.com \
    --cc=rafael@kernel.org \
    --cc=seanjc@google.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox