Kernel KVM virtualization development
 help / color / mirror / Atom feed
From: Anthony Harivel <aharivel@redhat.com>
To: kvm@vger.kernel.org
Cc: pbonzini@redhat.com, seanjc@google.com
Subject: [RFC PATCH v2] KVM: x86: RFC v2 for per-VM C-state policy enforcement (KVM_CAP_CSTATE_POLICY)
Date: Tue, 21 Jul 2026 14:22:15 +0200	[thread overview]
Message-ID: <20260721122215.1963145-1-aharivel@redhat.com> (raw)
In-Reply-To: <20260715124235.1444895-1-aharivel@redhat.com>

v2: Correct internal flow to use kvm_vcpu_halt() + cpuidle constraint
instead of inline MWAIT. Expand enforcement mechanism design question
with three concrete options (A/B/C) and cpuidle call chain analysis.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---
 rfc/KVM_CAP_CSTATE_POLICY.txt | 57 +++++++++++++++++++++++++++++------
 1 file changed, 48 insertions(+), 9 deletions(-)

diff --git a/rfc/KVM_CAP_CSTATE_POLICY.txt b/rfc/KVM_CAP_CSTATE_POLICY.txt
index f376adf77af1..dec369e328e2 100644
--- a/rfc/KVM_CAP_CSTATE_POLICY.txt
+++ b/rfc/KVM_CAP_CSTATE_POLICY.txt
@@ -1,6 +1,16 @@
 KVM: x86: Proposal for per-VM C-state policy enforcement
 ========================================================
 
+Changes since v1:
+  - Corrected internal flow (section 3): enforcement goes through
+    kvm_vcpu_halt() + cpuidle constraint, NOT inline MWAIT execution.
+    v1 was ambiguous on this point.
+  - Added rationale for why inline MWAIT in the exit handler is wrong:
+    monitor state cleared on VM exit (SDM 27.5.5), scheduler/RCU
+    invariant violations.
+  - Expanded design question (a) with concrete enforcement options
+    (A/B/C) and cpuidle call chain analysis.
+
 1. Problem statement
 --------------------
 
@@ -136,23 +146,52 @@ where power savings matter most (C3/C6), the overhead is under 3%.
 5. Design questions for discussion
 -----------------------------------
 
-(a) VMCS control toggling: switching "MWAIT exiting" at runtime when
+(a) Enforcement mechanism — cpuidle constraint
+
+    The enforcement point is kvm_vcpu_halt() → kvm_vcpu_block() →
+    schedule() → cpuidle. The problem: cpuidle has no per-task
+    C-state constraint. states_usage[].disable is per-CPU, and
+    forced_idle_latency_limit_ns is also per-CPU.
+
+    Three options I see:
+
+      Option A: Temporarily toggle states_usage[i].disable on the
+      pinned pCPU before/after kvm_vcpu_block(). Set
+      CPUIDLE_STATE_DISABLED_BY_DRIVER for states > max_cstate,
+      restore after wakeup. Simple, works with existing API, but
+      only correct with dedicated pinning — overcommit with mixed
+      policies would race on the disable flags.
+
+      Option B: Use forced_idle_latency_limit_ns on the pCPU.
+      Same per-CPU limitation, and latency-based rather than
+      state-index-based — less precise.
+
+      Option C: Propose a new cpuidle API for per-task idle
+      constraints (e.g. a per-task_struct annotation checked by
+      the governor during select()). Correct for all cases, but
+      bigger scope and needs cpuidle maintainer buy-in.
+
+    My use case is NFV with dedicated pinned cores, so option A
+    would work for v1 scoped to that case. But I'd rather not
+    invest in prototyping the wrong approach.
+
+    Is this use case too niche for upstream KVM? The target is
+    cloud operators running VNFs on dedicated cores who need a
+    host-side kill switch for deep C-states without guest
+    cooperation. If it is worth pursuing — would you accept
+    option A scoped to dedicated pinning for v1, or would you
+    want option C from the start?
+
+(b) VMCS control toggling: switching "MWAIT exiting" at runtime when
     the policy changes between -1 and >= 0. This requires a VMCS
     update on each vCPU. Safe to do via kvm_vcpu_kick() + request
     flag, or is there a better pattern?
 
-(b) MWAIT hint to C-state mapping: MWAIT sub-states (EAX[7:4] for
+(c) MWAIT hint to C-state mapping: MWAIT sub-states (EAX[7:4] for
     C-state, EAX[3:0] for sub-state) vary across CPU families. Should
     KVM maintain a per-model translation table, or should the policy
     operate on raw MWAIT hints directly?
 
-(c) cpuidle constraint mechanism: the proposed flow caps C-states
-    within kvm_vcpu_halt() via the host cpuidle framework. The
-    cleanest approach seems to be temporarily disabling states via
-    dev->states_usage[].disable before blocking. Is there a better
-    interface, or should we propose a new cpuidle API for transient
-    per-CPU state constraints?
-
 (d) Interaction with halt_poll_ns: when policy is active (max_cstate
     >= 0), interception is on, so halt polling applies again. The two
     knobs become orthogonal:
-- 
2.53.0


      parent reply	other threads:[~2026-07-21 12:22 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-15 12:42 [RFC PATCH] KVM: x86: RFC for per-VM C-state policy enforcement (KVM_CAP_CSTATE_POLICY) Anthony Harivel
2026-07-15 12:48 ` sashiko-bot
     [not found]   ` <CACUOEi9wRzr2yJJiu1fAk1xYFuDu8NDbpKKzbhQ-9bZhxaKRpQ@mail.gmail.com>
2026-07-15 13:52     ` Anthony Harivel
2026-07-16 10:05 ` Anthony Harivel
2026-07-21 12:22 ` Anthony Harivel [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260721122215.1963145-1-aharivel@redhat.com \
    --to=aharivel@redhat.com \
    --cc=kvm@vger.kernel.org \
    --cc=pbonzini@redhat.com \
    --cc=seanjc@google.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox