From: Anthony Harivel <aharivel@redhat.com>
To: kvm@vger.kernel.org
Cc: pbonzini@redhat.com, seanjc@google.com, aharivel@redhat.com
Subject: [RFC PATCH] KVM: x86: RFC for per-VM C-state policy enforcement (KVM_CAP_CSTATE_POLICY)
Date: Wed, 15 Jul 2026 14:42:11 +0200 [thread overview]
Message-ID: <20260715124235.1444895-1-aharivel@redhat.com> (raw)
KVM currently provides two mutually exclusive modes for guest idle
management: full interception (cpu-pm=off, host controls C-states but
pays VM exit overhead) or full delegation (cpu-pm=on, guest MWAIT
executes directly with zero overhead but host loses all control).
There is no middle ground. With cpu-pm=on, the host cannot override
guest C-state decisions - not via cpuidle sysfs, not via MSR 0xE2
(BIOS-locked), and not via any KVM interface.
This RFC proposes KVM_CAP_CSTATE_POLICY, a new VM-scoped capability
that lets the host set a per-VM maximum C-state ceiling. When active,
KVM intercepts MWAIT, inspects the requested C-state, and enforces
the ceiling before executing the idle instruction on behalf of the
guest.
The capability follows the KVM_CAP_HALT_POLL pattern: a per-VM ioctl,
changeable at runtime without VM restart.
Semantics:
args[0] = max_cstate
-1 = unrestricted (current cpu-pm=on, no interception)
0 = force C0 (no idle)
1 = cap at C1 (kill switch for deep C-states)
6 = cap at C6 (monitor but don't restrict)
Measured VM exit overhead on Haswell-EP (ftrace, x86-tsc clock):
Median: 867 ns (0.87 us)
p99: 3,740 ns (3.74 us)
Relative to C-state exit latencies this is negligible: +0.65% for C6
(133 us exit latency), +2.6% for C3 (33 us).
Use case: NFV deployments where the VNF vendor owns C-state selection
via guest cpuidle, but the cloud operator needs a host-side kill
switch to cap deep C-states per-VM for debugging or SLA enforcement,
without guest cooperation.
Design questions in rfc/KVM_CAP_CSTATE_POLICY.txt:
(a) VMCS "MWAIT exiting" runtime toggle pattern
(b) MWAIT hint to logical C-state mapping across CPU families
(c) HLT handling (no C-state hint in instruction)
(d) Interaction with halt_poll_ns
(e) Per-vCPU vs per-VM granularity
I am looking for feedback on the overall approach. If there is
consensus that this capability belongs in KVM, I am willing to
implement the full stack - starting with the KVM ioctl, followed
by QEMU and libvirt integration.
Signed-off-by: Anthony Harivel <aharivel@redhat.com>
---
rfc/KVM_CAP_CSTATE_POLICY.txt | 179 ++++++++++++++++++++++++++++++++++
1 file changed, 179 insertions(+)
create mode 100644 rfc/KVM_CAP_CSTATE_POLICY.txt
diff --git a/rfc/KVM_CAP_CSTATE_POLICY.txt b/rfc/KVM_CAP_CSTATE_POLICY.txt
new file mode 100644
index 000000000000..15439e6e7ebb
--- /dev/null
+++ b/rfc/KVM_CAP_CSTATE_POLICY.txt
@@ -0,0 +1,179 @@
+KVM: x86: Proposal for per-VM C-state policy enforcement
+========================================================
+
+1. Problem statement
+--------------------
+
+KVM currently offers two mutually exclusive modes for guest idle
+management on x86:
+
+ (a) Default (cpu-pm=off): guest HLT/MWAIT causes a VM exit. KVM
+ intercepts and decides the C-state via the host cpuidle governor.
+ The host has full control but pays exit overhead on every idle.
+
+ (b) Delegated (cpu-pm=on, QEMU -overcommit cpu-pm=on): guest MWAIT
+ executes directly on hardware via VMCS secondary execution control
+ "MWAIT exiting" = 0. Zero VM exit overhead, but the host has no
+ visibility or control over which C-states the guest enters.
+
+There is no middle ground. With delegation active, the host cannot cap
+the deepest C-state a guest may enter — not via host cpuidle sysfs, not
+via MSR 0xE2 (PKG_CST_CONFIG_CONTROL, typically BIOS-locked), and not
+via any KVM interface.
+
+This creates an operational gap for NFV and latency-sensitive cloud
+deployments where:
+
+ - The guest (VNF) should own C-state selection because it knows its
+ own latency requirements.
+ - The host operator needs a "kill switch" to cap deep C-states per-VM
+ for debugging or SLA enforcement, without VM restart and without
+ requiring guest cooperation (SSH/agent may be unavailable).
+
+2. Proposed capability: KVM_CAP_CSTATE_POLICY
+----------------------------------------------
+
+A new KVM VM-scoped capability that lets the host set a maximum C-state
+ceiling per VM. When a policy is active, KVM intercepts MWAIT/HLT,
+inspects the requested C-state, and enforces the ceiling before
+executing the idle instruction on behalf of the guest.
+
+ KVM_CAP_CSTATE_POLICY
+ Architecture: x86
+ Target: VM (struct kvm)
+ Parameters:
+ args[0] = max_cstate (-1 to 6)
+ args[1] = flags (reserved, must be 0)
+
+ Semantics:
+ max_cstate = -1 No interception; current cpu-pm=on behavior.
+ MWAIT executes directly (VMCS "MWAIT exiting" = 0).
+ max_cstate = 0 Force C0. Guest halts return immediately (no idle).
+ max_cstate = 1 Cap at C1. Guest requests for C3/C6 are downgraded.
+ max_cstate = 6 Cap at C6 (effectively unrestricted on most
+ platforms, but interception is active for
+ monitoring).
+
+The capability follows the same pattern as KVM_CAP_HALT_POLL: a per-VM
+ioctl that overrides system-wide behavior for a specific VM, changeable
+at runtime without VM restart.
+
+3. Internal flow
+----------------
+
+When max_cstate >= 0 (policy active):
+
+ 1. VMCS "MWAIT exiting" = 1 (intercept MWAIT).
+ 2. On VM exit for MWAIT:
+ a. Extract the C-state hint from the MWAIT operand.
+ b. Map the hint to a logical C-state level (MWAIT sub-states
+ are CPU-family-specific; a translation table is needed).
+ c. If requested_cstate <= max_cstate: execute native MWAIT with
+ the guest's original hint.
+ d. If requested_cstate > max_cstate: execute MWAIT with the
+ capped hint, or return to guest immediately (C0).
+ 3. For HLT exits: KVM selects a C-state up to max_cstate via the
+ host cpuidle governor (existing behavior, but now capped).
+
+When max_cstate = -1 (unrestricted):
+
+ 1. VMCS "MWAIT exiting" = 0 (direct execution, no VM exit).
+ 2. Equivalent to current cpu-pm=on behavior.
+
+Transitioning between -1 and >= 0 requires toggling the VMCS secondary
+execution control at runtime. This is the same mechanism used for other
+VMCS control toggles and is well-precedented in KVM.
+
+4. Measured VM exit overhead
+----------------------------
+
+We benchmarked the exit-entry round-trip cost on Intel Xeon E5-2630 v3
+(Haswell-EP, 2x8C/16T) using ftrace with x86-tsc clock (nanosecond
+precision). MSR_WRITE exits were used as a proxy for pure exit mechanism
+cost — fast synchronous round-trips with no scheduling or blocking.
+
+ Samples: 154
+ Min: 433 ns
+ Median: 867 ns (0.87 us)
+ Mean: 1,523 ns (1.52 us)
+ p95: 3,037 ns (3.04 us)
+ p99: 3,740 ns (3.74 us)
+ Max: 13,963 ns (13.96 us)
+
+Relative to C-state exit latencies:
+
+ C-state Exit latency VM exit adds Relative overhead
+ C1E 10 us +0.87 us +8.7%
+ C3 33 us +0.87 us +2.6%
+ C6 133 us +0.87 us +0.65%
+
+The interception overhead is sub-microsecond at median and negligible
+relative to the C-state transitions it governs. For the deepest C-states
+where power savings matter most (C3/C6), the overhead is under 3%.
+
+5. Design questions for discussion
+-----------------------------------
+
+(a) VMCS control toggling: switching "MWAIT exiting" at runtime when
+ the policy changes between -1 and >= 0. This requires a VMCS
+ update on each vCPU. Safe to do via kvm_vcpu_kick() + request
+ flag, or is there a better pattern?
+
+(b) MWAIT hint to C-state mapping: MWAIT sub-states (EAX[7:4] for
+ C-state, EAX[3:0] for sub-state) vary across CPU families. Should
+ KVM maintain a per-model translation table, or should the policy
+ operate on raw MWAIT hints directly?
+
+(c) HLT handling: when a guest uses HLT (no C-state hint), KVM
+ currently enters kvm_vcpu_halt() which may poll (halt_poll_ns)
+ then block. With a max_cstate policy, should KVM cap the C-state
+ the host cpuidle governor selects for the blocked vCPU thread?
+ This would require hooking into cpuidle or using
+ MONITOR/MWAIT-based idle with a capped hint instead of schedule().
+
+(d) Interaction with halt_poll_ns: when policy is active (max_cstate
+ >= 0), interception is on, so halt polling applies again. The two
+ knobs become orthogonal:
+ halt_poll_ns = spin duration before any C-state
+ max_cstate = deepest C-state after polling fails
+ Is this the right interaction model?
+
+(e) Per-vCPU vs per-VM: starting with per-VM is simpler. Per-vCPU
+ would allow mixed policies (dataplane vCPUs in C0, control-plane
+ vCPUs in C6). Worth the complexity for v1?
+
+6. Userspace integration path
+-----------------------------
+
+ Kernel: KVM ioctl (this proposal)
+ QEMU: New -overcommit sub-option: cpu-pm-max-cstate=N
+ libvirt: New XML element in <features><kvm>
+
+Runtime changes via the ioctl would allow virsh/QEMU monitor commands
+to adjust the policy on a live VM without restart.
+
+7. Prior art
+------------
+
+ - KVM_CAP_HALT_POLL: per-VM halt polling override. Same design
+ pattern — VM capability, ioctl, per-VM value. Merged in 2020.
+
+ - Xen ACPI idle driver: Xen has per-domain C-state policy via its
+ ACPI integration. KVM has no equivalent.
+
+ - Host kernel max_cstate: the kernel parameter
+ intel_idle.max_cstate= caps C-states system-wide, but has no
+ per-VM or per-CPU granularity.
+
+ - QEMU cpu-pm=on: delegates MWAIT/HLT to guest but provides no
+ policy mechanism. libvirt support for cpu-pm was never merged
+ upstream despite being documented.
+
+8. Next steps
+-------------
+
+I am looking for feedback on the overall approach and the design
+questions above. If there is consensus that this capability belongs
+in KVM, I am willing to implement the full stack — starting with the
+KVM ioctl and MWAIT/HLT exit handler changes, followed by the QEMU
+and libvirt integration.
--
2.53.0
next reply other threads:[~2026-07-15 12:42 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-15 12:42 Anthony Harivel [this message]
2026-07-15 12:48 ` [RFC PATCH] KVM: x86: RFC for per-VM C-state policy enforcement (KVM_CAP_CSTATE_POLICY) sashiko-bot
[not found] ` <CACUOEi9wRzr2yJJiu1fAk1xYFuDu8NDbpKKzbhQ-9bZhxaKRpQ@mail.gmail.com>
2026-07-15 13:52 ` Anthony Harivel
2026-07-16 10:05 ` Anthony Harivel
2026-07-21 12:22 ` [RFC PATCH v2] KVM: x86: RFC v2 " Anthony Harivel
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260715124235.1444895-1-aharivel@redhat.com \
--to=aharivel@redhat.com \
--cc=kvm@vger.kernel.org \
--cc=pbonzini@redhat.com \
--cc=seanjc@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox