* [RFC PATCH] KVM: x86: RFC for per-VM C-state policy enforcement (KVM_CAP_CSTATE_POLICY)
@ 2026-07-15 12:42 Anthony Harivel
2026-07-15 12:48 ` sashiko-bot
` (2 more replies)
0 siblings, 3 replies; 5+ messages in thread
From: Anthony Harivel @ 2026-07-15 12:42 UTC (permalink / raw)
To: kvm; +Cc: pbonzini, seanjc, aharivel
KVM currently provides two mutually exclusive modes for guest idle
management: full interception (cpu-pm=off, host controls C-states but
pays VM exit overhead) or full delegation (cpu-pm=on, guest MWAIT
executes directly with zero overhead but host loses all control).
There is no middle ground. With cpu-pm=on, the host cannot override
guest C-state decisions - not via cpuidle sysfs, not via MSR 0xE2
(BIOS-locked), and not via any KVM interface.
This RFC proposes KVM_CAP_CSTATE_POLICY, a new VM-scoped capability
that lets the host set a per-VM maximum C-state ceiling. When active,
KVM intercepts MWAIT, inspects the requested C-state, and enforces
the ceiling before executing the idle instruction on behalf of the
guest.
The capability follows the KVM_CAP_HALT_POLL pattern: a per-VM ioctl,
changeable at runtime without VM restart.
Semantics:
args[0] = max_cstate
-1 = unrestricted (current cpu-pm=on, no interception)
0 = force C0 (no idle)
1 = cap at C1 (kill switch for deep C-states)
6 = cap at C6 (monitor but don't restrict)
Measured VM exit overhead on Haswell-EP (ftrace, x86-tsc clock):
Median: 867 ns (0.87 us)
p99: 3,740 ns (3.74 us)
Relative to C-state exit latencies this is negligible: +0.65% for C6
(133 us exit latency), +2.6% for C3 (33 us).
Use case: NFV deployments where the VNF vendor owns C-state selection
via guest cpuidle, but the cloud operator needs a host-side kill
switch to cap deep C-states per-VM for debugging or SLA enforcement,
without guest cooperation.
Design questions in rfc/KVM_CAP_CSTATE_POLICY.txt:
(a) VMCS "MWAIT exiting" runtime toggle pattern
(b) MWAIT hint to logical C-state mapping across CPU families
(c) HLT handling (no C-state hint in instruction)
(d) Interaction with halt_poll_ns
(e) Per-vCPU vs per-VM granularity
I am looking for feedback on the overall approach. If there is
consensus that this capability belongs in KVM, I am willing to
implement the full stack - starting with the KVM ioctl, followed
by QEMU and libvirt integration.
Signed-off-by: Anthony Harivel <aharivel@redhat.com>
---
rfc/KVM_CAP_CSTATE_POLICY.txt | 179 ++++++++++++++++++++++++++++++++++
1 file changed, 179 insertions(+)
create mode 100644 rfc/KVM_CAP_CSTATE_POLICY.txt
diff --git a/rfc/KVM_CAP_CSTATE_POLICY.txt b/rfc/KVM_CAP_CSTATE_POLICY.txt
new file mode 100644
index 000000000000..15439e6e7ebb
--- /dev/null
+++ b/rfc/KVM_CAP_CSTATE_POLICY.txt
@@ -0,0 +1,179 @@
+KVM: x86: Proposal for per-VM C-state policy enforcement
+========================================================
+
+1. Problem statement
+--------------------
+
+KVM currently offers two mutually exclusive modes for guest idle
+management on x86:
+
+ (a) Default (cpu-pm=off): guest HLT/MWAIT causes a VM exit. KVM
+ intercepts and decides the C-state via the host cpuidle governor.
+ The host has full control but pays exit overhead on every idle.
+
+ (b) Delegated (cpu-pm=on, QEMU -overcommit cpu-pm=on): guest MWAIT
+ executes directly on hardware via VMCS secondary execution control
+ "MWAIT exiting" = 0. Zero VM exit overhead, but the host has no
+ visibility or control over which C-states the guest enters.
+
+There is no middle ground. With delegation active, the host cannot cap
+the deepest C-state a guest may enter — not via host cpuidle sysfs, not
+via MSR 0xE2 (PKG_CST_CONFIG_CONTROL, typically BIOS-locked), and not
+via any KVM interface.
+
+This creates an operational gap for NFV and latency-sensitive cloud
+deployments where:
+
+ - The guest (VNF) should own C-state selection because it knows its
+ own latency requirements.
+ - The host operator needs a "kill switch" to cap deep C-states per-VM
+ for debugging or SLA enforcement, without VM restart and without
+ requiring guest cooperation (SSH/agent may be unavailable).
+
+2. Proposed capability: KVM_CAP_CSTATE_POLICY
+----------------------------------------------
+
+A new KVM VM-scoped capability that lets the host set a maximum C-state
+ceiling per VM. When a policy is active, KVM intercepts MWAIT/HLT,
+inspects the requested C-state, and enforces the ceiling before
+executing the idle instruction on behalf of the guest.
+
+ KVM_CAP_CSTATE_POLICY
+ Architecture: x86
+ Target: VM (struct kvm)
+ Parameters:
+ args[0] = max_cstate (-1 to 6)
+ args[1] = flags (reserved, must be 0)
+
+ Semantics:
+ max_cstate = -1 No interception; current cpu-pm=on behavior.
+ MWAIT executes directly (VMCS "MWAIT exiting" = 0).
+ max_cstate = 0 Force C0. Guest halts return immediately (no idle).
+ max_cstate = 1 Cap at C1. Guest requests for C3/C6 are downgraded.
+ max_cstate = 6 Cap at C6 (effectively unrestricted on most
+ platforms, but interception is active for
+ monitoring).
+
+The capability follows the same pattern as KVM_CAP_HALT_POLL: a per-VM
+ioctl that overrides system-wide behavior for a specific VM, changeable
+at runtime without VM restart.
+
+3. Internal flow
+----------------
+
+When max_cstate >= 0 (policy active):
+
+ 1. VMCS "MWAIT exiting" = 1 (intercept MWAIT).
+ 2. On VM exit for MWAIT:
+ a. Extract the C-state hint from the MWAIT operand.
+ b. Map the hint to a logical C-state level (MWAIT sub-states
+ are CPU-family-specific; a translation table is needed).
+ c. If requested_cstate <= max_cstate: execute native MWAIT with
+ the guest's original hint.
+ d. If requested_cstate > max_cstate: execute MWAIT with the
+ capped hint, or return to guest immediately (C0).
+ 3. For HLT exits: KVM selects a C-state up to max_cstate via the
+ host cpuidle governor (existing behavior, but now capped).
+
+When max_cstate = -1 (unrestricted):
+
+ 1. VMCS "MWAIT exiting" = 0 (direct execution, no VM exit).
+ 2. Equivalent to current cpu-pm=on behavior.
+
+Transitioning between -1 and >= 0 requires toggling the VMCS secondary
+execution control at runtime. This is the same mechanism used for other
+VMCS control toggles and is well-precedented in KVM.
+
+4. Measured VM exit overhead
+----------------------------
+
+We benchmarked the exit-entry round-trip cost on Intel Xeon E5-2630 v3
+(Haswell-EP, 2x8C/16T) using ftrace with x86-tsc clock (nanosecond
+precision). MSR_WRITE exits were used as a proxy for pure exit mechanism
+cost — fast synchronous round-trips with no scheduling or blocking.
+
+ Samples: 154
+ Min: 433 ns
+ Median: 867 ns (0.87 us)
+ Mean: 1,523 ns (1.52 us)
+ p95: 3,037 ns (3.04 us)
+ p99: 3,740 ns (3.74 us)
+ Max: 13,963 ns (13.96 us)
+
+Relative to C-state exit latencies:
+
+ C-state Exit latency VM exit adds Relative overhead
+ C1E 10 us +0.87 us +8.7%
+ C3 33 us +0.87 us +2.6%
+ C6 133 us +0.87 us +0.65%
+
+The interception overhead is sub-microsecond at median and negligible
+relative to the C-state transitions it governs. For the deepest C-states
+where power savings matter most (C3/C6), the overhead is under 3%.
+
+5. Design questions for discussion
+-----------------------------------
+
+(a) VMCS control toggling: switching "MWAIT exiting" at runtime when
+ the policy changes between -1 and >= 0. This requires a VMCS
+ update on each vCPU. Safe to do via kvm_vcpu_kick() + request
+ flag, or is there a better pattern?
+
+(b) MWAIT hint to C-state mapping: MWAIT sub-states (EAX[7:4] for
+ C-state, EAX[3:0] for sub-state) vary across CPU families. Should
+ KVM maintain a per-model translation table, or should the policy
+ operate on raw MWAIT hints directly?
+
+(c) HLT handling: when a guest uses HLT (no C-state hint), KVM
+ currently enters kvm_vcpu_halt() which may poll (halt_poll_ns)
+ then block. With a max_cstate policy, should KVM cap the C-state
+ the host cpuidle governor selects for the blocked vCPU thread?
+ This would require hooking into cpuidle or using
+ MONITOR/MWAIT-based idle with a capped hint instead of schedule().
+
+(d) Interaction with halt_poll_ns: when policy is active (max_cstate
+ >= 0), interception is on, so halt polling applies again. The two
+ knobs become orthogonal:
+ halt_poll_ns = spin duration before any C-state
+ max_cstate = deepest C-state after polling fails
+ Is this the right interaction model?
+
+(e) Per-vCPU vs per-VM: starting with per-VM is simpler. Per-vCPU
+ would allow mixed policies (dataplane vCPUs in C0, control-plane
+ vCPUs in C6). Worth the complexity for v1?
+
+6. Userspace integration path
+-----------------------------
+
+ Kernel: KVM ioctl (this proposal)
+ QEMU: New -overcommit sub-option: cpu-pm-max-cstate=N
+ libvirt: New XML element in <features><kvm>
+
+Runtime changes via the ioctl would allow virsh/QEMU monitor commands
+to adjust the policy on a live VM without restart.
+
+7. Prior art
+------------
+
+ - KVM_CAP_HALT_POLL: per-VM halt polling override. Same design
+ pattern — VM capability, ioctl, per-VM value. Merged in 2020.
+
+ - Xen ACPI idle driver: Xen has per-domain C-state policy via its
+ ACPI integration. KVM has no equivalent.
+
+ - Host kernel max_cstate: the kernel parameter
+ intel_idle.max_cstate= caps C-states system-wide, but has no
+ per-VM or per-CPU granularity.
+
+ - QEMU cpu-pm=on: delegates MWAIT/HLT to guest but provides no
+ policy mechanism. libvirt support for cpu-pm was never merged
+ upstream despite being documented.
+
+8. Next steps
+-------------
+
+I am looking for feedback on the overall approach and the design
+questions above. If there is consensus that this capability belongs
+in KVM, I am willing to implement the full stack — starting with the
+KVM ioctl and MWAIT/HLT exit handler changes, followed by the QEMU
+and libvirt integration.
--
2.53.0
^ permalink raw reply related [flat|nested] 5+ messages in thread* Re: [RFC PATCH] KVM: x86: RFC for per-VM C-state policy enforcement (KVM_CAP_CSTATE_POLICY)
2026-07-15 12:42 [RFC PATCH] KVM: x86: RFC for per-VM C-state policy enforcement (KVM_CAP_CSTATE_POLICY) Anthony Harivel
@ 2026-07-15 12:48 ` sashiko-bot
[not found] ` <CACUOEi9wRzr2yJJiu1fAk1xYFuDu8NDbpKKzbhQ-9bZhxaKRpQ@mail.gmail.com>
2026-07-16 10:05 ` Anthony Harivel
2026-07-21 12:22 ` [RFC PATCH v2] KVM: x86: RFC v2 " Anthony Harivel
2 siblings, 1 reply; 5+ messages in thread
From: sashiko-bot @ 2026-07-15 12:48 UTC (permalink / raw)
To: Anthony Harivel; +Cc: kvm
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] VM-Exits clear the CPU's monitor event pending flag, causing host-side MWAIT to immediately fall through (NOP).
- [High] Executing MWAIT synchronously within the KVM exit handler violates Linux scheduler and RCU invariants.
--
--- Patch [1]: [RFC PATCH] KVM: x86: RFC for per-VM C-state policy enforcement (KVM_CAP_CSTATE_POLICY) ---
commit ba053ea08a3b49a2e72662168226ef280171ed4d
Author: Anthony Harivel <aharivel@redhat.com>
KVM: x86: RFC for per-VM C-state policy enforcement (KVM_CAP_CSTATE_POLICY)
This RFC proposes KVM_CAP_CSTATE_POLICY, a new VM-scoped capability
that lets the host set a per-VM maximum C-state ceiling. When active,
KVM intercepts MWAIT, inspects the requested C-state, and enforces
the ceiling by executing a capped MWAIT on behalf of the guest.
> diff --git a/rfc/KVM_CAP_CSTATE_POLICY.txt b/rfc/KVM_CAP_CSTATE_POLICY.txt
> new file mode 100644
> index 0000000000000..15439e6e7ebbf
> --- /dev/null
> +++ b/rfc/KVM_CAP_CSTATE_POLICY.txt
[ ... ]
> @@ -60,18 +60,18 @@
> 3. Internal flow
> ----------------
>
> When max_cstate >= 0 (policy active):
>
> 1. VMCS "MWAIT exiting" = 1 (intercept MWAIT).
> 2. On VM exit for MWAIT:
> a. Extract the C-state hint from the MWAIT operand.
> b. Map the hint to a logical C-state level (MWAIT sub-states
> are CPU-family-specific; a translation table is needed).
> c. If requested_cstate <= max_cstate: execute native MWAIT with
> the guest's original hint.
> d. If requested_cstate > max_cstate: execute MWAIT with the
> capped hint, or return to guest immediately (C0).
[Severity: High]
Does executing native MWAIT on the host side work as intended here?
According to the Intel SDM (Vol 3, 27.5.5 Updates to Non-Register State),
a VM-Exit unconditionally clears the state of the monitor hardware.
Since the state established by the guest's preceding MONITOR instruction
is wiped out by the VM-Exit, won't the subsequent host-side MWAIT find no
armed monitor and immediately fall through as a NOP, resulting in a busy
wait loop instead of entering an idle power state?
[Severity: High]
If we execute MWAIT synchronously within the KVM exit handler, how does
this interact with the Linux scheduler and RCU?
If the CPU successfully enters a deep sleep state via an inline host
MWAIT, it seems the KVM vCPU thread would remain in the TASK_RUNNING
state without yielding via schedule() or notifying RCU via
rcu_idle_enter().
Could this hide the CPU's true idle status from the scheduler, break load
balancing, and stall RCU grace periods by bypassing the host cpuidle
framework?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260715124235.1444895-1-aharivel@redhat.com?part=1
^ permalink raw reply [flat|nested] 5+ messages in thread* Re: [RFC PATCH] KVM: x86: RFC for per-VM C-state policy enforcement (KVM_CAP_CSTATE_POLICY)
2026-07-15 12:42 [RFC PATCH] KVM: x86: RFC for per-VM C-state policy enforcement (KVM_CAP_CSTATE_POLICY) Anthony Harivel
2026-07-15 12:48 ` sashiko-bot
@ 2026-07-16 10:05 ` Anthony Harivel
2026-07-21 12:22 ` [RFC PATCH v2] KVM: x86: RFC v2 " Anthony Harivel
2 siblings, 0 replies; 5+ messages in thread
From: Anthony Harivel @ 2026-07-16 10:05 UTC (permalink / raw)
To: aharivel; +Cc: kvm, pbonzini, seanjc
Hi Paolo, Sean,
I sent an RFC to kvm@ for a per-VM C-state policy capability
(KVM_CAP_CSTATE_POLICY). The idea is to let the host set a max C-state ceiling
per VM, filling the gap between full interception (cpu-pm=off) and full
delegation (cpu-pm=on).
https://lore.kernel.org/kvm/20260715124235.1444895-1-aharivel@redhat.com/
I've been tracing the implementation path and I've hit a design question I'd
value your input on before going further.
The enforcement point is kvm_vcpu_halt() → kvm_vcpu_block() → schedule()
→ cpuidle. The problem is that cpuidle has no per-task C-state constraint
— states_usage[].disable is per-CPU, and forced_idle_latency_limit_ns is also
per-CPU.
So three options:
(a) Temporarily toggle states_usage[i].disable on the pinned pCPU
before/after kvm_vcpu_block(). Simple, works with existing API,
but only correct with dedicated pinning (overcommit with mixed
policies would race on the disable flags).
(b) Use forced_idle_latency_limit_ns on the pCPU — same per-CPU
limitation, and latency-based rather than state-index-based.
(c) Propose a new cpuidle API for per-task idle constraints (e.g.
a per-task_struct annotation checked by the governor). Correct
for all cases, but bigger scope and needs cpuidle maintainer
buy-in (Rafael Wysocki).
My use case is NFV with dedicated pinned cores, so option (a) would work.
But I'm wondering:
1. Is this use case too niche for upstream KVM? The target is cloud
operators running VNFs on dedicated cores who need a host-side
kill switch for deep C-states without guest cooperation.
2. If it is worth pursuing — would you accept option (a) scoped to
the dedicated-pinning case for v1, or would you want the cleaner
option (c) from the start?
3. Is there a simpler approach I'm missing entirely? I looked at
whether KVM could cap things before reaching schedule() (e.g.
custom idle with MONITOR+MWAIT), but that breaks scheduler/RCU
invariants as the sashiko review pointed out.
Happy to prototype whichever direction you think has the best chance of landing.
Anthony
^ permalink raw reply [flat|nested] 5+ messages in thread* [RFC PATCH v2] KVM: x86: RFC v2 for per-VM C-state policy enforcement (KVM_CAP_CSTATE_POLICY)
2026-07-15 12:42 [RFC PATCH] KVM: x86: RFC for per-VM C-state policy enforcement (KVM_CAP_CSTATE_POLICY) Anthony Harivel
2026-07-15 12:48 ` sashiko-bot
2026-07-16 10:05 ` Anthony Harivel
@ 2026-07-21 12:22 ` Anthony Harivel
2 siblings, 0 replies; 5+ messages in thread
From: Anthony Harivel @ 2026-07-21 12:22 UTC (permalink / raw)
To: kvm; +Cc: pbonzini, seanjc
v2: Correct internal flow to use kvm_vcpu_halt() + cpuidle constraint
instead of inline MWAIT. Expand enforcement mechanism design question
with three concrete options (A/B/C) and cpuidle call chain analysis.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---
rfc/KVM_CAP_CSTATE_POLICY.txt | 57 +++++++++++++++++++++++++++++------
1 file changed, 48 insertions(+), 9 deletions(-)
diff --git a/rfc/KVM_CAP_CSTATE_POLICY.txt b/rfc/KVM_CAP_CSTATE_POLICY.txt
index f376adf77af1..dec369e328e2 100644
--- a/rfc/KVM_CAP_CSTATE_POLICY.txt
+++ b/rfc/KVM_CAP_CSTATE_POLICY.txt
@@ -1,6 +1,16 @@
KVM: x86: Proposal for per-VM C-state policy enforcement
========================================================
+Changes since v1:
+ - Corrected internal flow (section 3): enforcement goes through
+ kvm_vcpu_halt() + cpuidle constraint, NOT inline MWAIT execution.
+ v1 was ambiguous on this point.
+ - Added rationale for why inline MWAIT in the exit handler is wrong:
+ monitor state cleared on VM exit (SDM 27.5.5), scheduler/RCU
+ invariant violations.
+ - Expanded design question (a) with concrete enforcement options
+ (A/B/C) and cpuidle call chain analysis.
+
1. Problem statement
--------------------
@@ -136,23 +146,52 @@ where power savings matter most (C3/C6), the overhead is under 3%.
5. Design questions for discussion
-----------------------------------
-(a) VMCS control toggling: switching "MWAIT exiting" at runtime when
+(a) Enforcement mechanism — cpuidle constraint
+
+ The enforcement point is kvm_vcpu_halt() → kvm_vcpu_block() →
+ schedule() → cpuidle. The problem: cpuidle has no per-task
+ C-state constraint. states_usage[].disable is per-CPU, and
+ forced_idle_latency_limit_ns is also per-CPU.
+
+ Three options I see:
+
+ Option A: Temporarily toggle states_usage[i].disable on the
+ pinned pCPU before/after kvm_vcpu_block(). Set
+ CPUIDLE_STATE_DISABLED_BY_DRIVER for states > max_cstate,
+ restore after wakeup. Simple, works with existing API, but
+ only correct with dedicated pinning — overcommit with mixed
+ policies would race on the disable flags.
+
+ Option B: Use forced_idle_latency_limit_ns on the pCPU.
+ Same per-CPU limitation, and latency-based rather than
+ state-index-based — less precise.
+
+ Option C: Propose a new cpuidle API for per-task idle
+ constraints (e.g. a per-task_struct annotation checked by
+ the governor during select()). Correct for all cases, but
+ bigger scope and needs cpuidle maintainer buy-in.
+
+ My use case is NFV with dedicated pinned cores, so option A
+ would work for v1 scoped to that case. But I'd rather not
+ invest in prototyping the wrong approach.
+
+ Is this use case too niche for upstream KVM? The target is
+ cloud operators running VNFs on dedicated cores who need a
+ host-side kill switch for deep C-states without guest
+ cooperation. If it is worth pursuing — would you accept
+ option A scoped to dedicated pinning for v1, or would you
+ want option C from the start?
+
+(b) VMCS control toggling: switching "MWAIT exiting" at runtime when
the policy changes between -1 and >= 0. This requires a VMCS
update on each vCPU. Safe to do via kvm_vcpu_kick() + request
flag, or is there a better pattern?
-(b) MWAIT hint to C-state mapping: MWAIT sub-states (EAX[7:4] for
+(c) MWAIT hint to C-state mapping: MWAIT sub-states (EAX[7:4] for
C-state, EAX[3:0] for sub-state) vary across CPU families. Should
KVM maintain a per-model translation table, or should the policy
operate on raw MWAIT hints directly?
-(c) cpuidle constraint mechanism: the proposed flow caps C-states
- within kvm_vcpu_halt() via the host cpuidle framework. The
- cleanest approach seems to be temporarily disabling states via
- dev->states_usage[].disable before blocking. Is there a better
- interface, or should we propose a new cpuidle API for transient
- per-CPU state constraints?
-
(d) Interaction with halt_poll_ns: when policy is active (max_cstate
>= 0), interception is on, so halt polling applies again. The two
knobs become orthogonal:
--
2.53.0
^ permalink raw reply related [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-07-21 12:22 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-15 12:42 [RFC PATCH] KVM: x86: RFC for per-VM C-state policy enforcement (KVM_CAP_CSTATE_POLICY) Anthony Harivel
2026-07-15 12:48 ` sashiko-bot
[not found] ` <CACUOEi9wRzr2yJJiu1fAk1xYFuDu8NDbpKKzbhQ-9bZhxaKRpQ@mail.gmail.com>
2026-07-15 13:52 ` Anthony Harivel
2026-07-16 10:05 ` Anthony Harivel
2026-07-21 12:22 ` [RFC PATCH v2] KVM: x86: RFC v2 " Anthony Harivel
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox