From: Anthony Harivel <aharivel@redhat.com>
To: kvm@vger.kernel.org
Cc: pbonzini@redhat.com, seanjc@google.com
Subject: [RFC PATCH v2] KVM: x86: RFC v2 for per-VM C-state policy enforcement (KVM_CAP_CSTATE_POLICY)
Date: Tue, 21 Jul 2026 14:22:15 +0200 [thread overview]
Message-ID: <20260721122215.1963145-1-aharivel@redhat.com> (raw)
In-Reply-To: <20260715124235.1444895-1-aharivel@redhat.com>
v2: Correct internal flow to use kvm_vcpu_halt() + cpuidle constraint
instead of inline MWAIT. Expand enforcement mechanism design question
with three concrete options (A/B/C) and cpuidle call chain analysis.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---
rfc/KVM_CAP_CSTATE_POLICY.txt | 57 +++++++++++++++++++++++++++++------
1 file changed, 48 insertions(+), 9 deletions(-)
diff --git a/rfc/KVM_CAP_CSTATE_POLICY.txt b/rfc/KVM_CAP_CSTATE_POLICY.txt
index f376adf77af1..dec369e328e2 100644
--- a/rfc/KVM_CAP_CSTATE_POLICY.txt
+++ b/rfc/KVM_CAP_CSTATE_POLICY.txt
@@ -1,6 +1,16 @@
KVM: x86: Proposal for per-VM C-state policy enforcement
========================================================
+Changes since v1:
+ - Corrected internal flow (section 3): enforcement goes through
+ kvm_vcpu_halt() + cpuidle constraint, NOT inline MWAIT execution.
+ v1 was ambiguous on this point.
+ - Added rationale for why inline MWAIT in the exit handler is wrong:
+ monitor state cleared on VM exit (SDM 27.5.5), scheduler/RCU
+ invariant violations.
+ - Expanded design question (a) with concrete enforcement options
+ (A/B/C) and cpuidle call chain analysis.
+
1. Problem statement
--------------------
@@ -136,23 +146,52 @@ where power savings matter most (C3/C6), the overhead is under 3%.
5. Design questions for discussion
-----------------------------------
-(a) VMCS control toggling: switching "MWAIT exiting" at runtime when
+(a) Enforcement mechanism — cpuidle constraint
+
+ The enforcement point is kvm_vcpu_halt() → kvm_vcpu_block() →
+ schedule() → cpuidle. The problem: cpuidle has no per-task
+ C-state constraint. states_usage[].disable is per-CPU, and
+ forced_idle_latency_limit_ns is also per-CPU.
+
+ Three options I see:
+
+ Option A: Temporarily toggle states_usage[i].disable on the
+ pinned pCPU before/after kvm_vcpu_block(). Set
+ CPUIDLE_STATE_DISABLED_BY_DRIVER for states > max_cstate,
+ restore after wakeup. Simple, works with existing API, but
+ only correct with dedicated pinning — overcommit with mixed
+ policies would race on the disable flags.
+
+ Option B: Use forced_idle_latency_limit_ns on the pCPU.
+ Same per-CPU limitation, and latency-based rather than
+ state-index-based — less precise.
+
+ Option C: Propose a new cpuidle API for per-task idle
+ constraints (e.g. a per-task_struct annotation checked by
+ the governor during select()). Correct for all cases, but
+ bigger scope and needs cpuidle maintainer buy-in.
+
+ My use case is NFV with dedicated pinned cores, so option A
+ would work for v1 scoped to that case. But I'd rather not
+ invest in prototyping the wrong approach.
+
+ Is this use case too niche for upstream KVM? The target is
+ cloud operators running VNFs on dedicated cores who need a
+ host-side kill switch for deep C-states without guest
+ cooperation. If it is worth pursuing — would you accept
+ option A scoped to dedicated pinning for v1, or would you
+ want option C from the start?
+
+(b) VMCS control toggling: switching "MWAIT exiting" at runtime when
the policy changes between -1 and >= 0. This requires a VMCS
update on each vCPU. Safe to do via kvm_vcpu_kick() + request
flag, or is there a better pattern?
-(b) MWAIT hint to C-state mapping: MWAIT sub-states (EAX[7:4] for
+(c) MWAIT hint to C-state mapping: MWAIT sub-states (EAX[7:4] for
C-state, EAX[3:0] for sub-state) vary across CPU families. Should
KVM maintain a per-model translation table, or should the policy
operate on raw MWAIT hints directly?
-(c) cpuidle constraint mechanism: the proposed flow caps C-states
- within kvm_vcpu_halt() via the host cpuidle framework. The
- cleanest approach seems to be temporarily disabling states via
- dev->states_usage[].disable before blocking. Is there a better
- interface, or should we propose a new cpuidle API for transient
- per-CPU state constraints?
-
(d) Interaction with halt_poll_ns: when policy is active (max_cstate
>= 0), interception is on, so halt polling applies again. The two
knobs become orthogonal:
--
2.53.0
prev parent reply other threads:[~2026-07-21 12:22 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-15 12:42 [RFC PATCH] KVM: x86: RFC for per-VM C-state policy enforcement (KVM_CAP_CSTATE_POLICY) Anthony Harivel
2026-07-15 12:48 ` sashiko-bot
[not found] ` <CACUOEi9wRzr2yJJiu1fAk1xYFuDu8NDbpKKzbhQ-9bZhxaKRpQ@mail.gmail.com>
2026-07-15 13:52 ` Anthony Harivel
2026-07-16 10:05 ` Anthony Harivel
2026-07-21 12:22 ` Anthony Harivel [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260721122215.1963145-1-aharivel@redhat.com \
--to=aharivel@redhat.com \
--cc=kvm@vger.kernel.org \
--cc=pbonzini@redhat.com \
--cc=seanjc@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox