Kernel KVM virtualization development
 help / color / mirror / Atom feed
From: Yosry Ahmed <yosry@kernel.org>
To: Sean Christopherson <seanjc@google.com>
Cc: sashiko-reviews@lists.linux.dev, kvm@vger.kernel.org
Subject: Re: [RFC PATCH v2 17/25] KVM: nSVM: Service local TLB flushes before nested transitions
Date: Fri, 24 Jul 2026 22:27:23 +0000	[thread overview]
Message-ID: <amPmtfdm69hw_CM3@google.com> (raw)
In-Reply-To: <amIYWOgQw9z3zedA@google.com>

On Thu, Jul 23, 2026 at 06:34:16AM -0700, Sean Christopherson wrote:
> On Thu, Jul 23, 2026, Yosry Ahmed wrote:
> > On Wed, Jul 22, 2026 at 05:46:31PM -0700, Sean Christopherson wrote:
> > > On Tue, Jun 16, 2026, Yosry Ahmed wrote:
> > > > On Tue, Jun 16, 2026 at 11:06 AM Yosry Ahmed <yosry@kernel.org> wrote:
> > > > > > > diff --git a/arch/x86/kvm/svm/nested.c b/arch/x86/kvm/svm/nested.c
> > > > > > > index 234724d8b4c54..7b19191e0e43f 100644
> > > > > > > --- a/arch/x86/kvm/svm/nested.c
> > > > > > > +++ b/arch/x86/kvm/svm/nested.c
> > > > > > > @@ -717,9 +717,16 @@ static void svm_switch_vmcb(struct vcpu_svm *svm, struct kvm_vmcb_info *target_v
> > > > > > >  {
> > > > > > >       struct kvm_vcpu *vcpu = &svm->vcpu;
> > > > > > >
> > > > > > > +     /*
> > > > > > > +      * TLB flushes are applied to the VMCB, so apply any pending TLB flushes
> > > > > > > +      * on the current VMCB before switching to a new one..
> > > > > > > +      */
> > > > > > > +     kvm_service_local_tlb_flush_requests(vcpu);
> > > > > >
> > > > > > [Severity: High]
> > > > > > Does this clear global TLB flush requests (like KVM_REQ_TLB_FLUSH_GUEST)
> > > > > > while applying them only to the outgoing VMCB?
> > > > > >
> > > > > > If a host-initiated flush occurs (e.g., from page reclaim), the request
> > > > > > bit will be consumed and applied to the outgoing VMCB, leaving the incoming
> > > > > > VMCB to miss the flush and run with a stale TLB.
> > > > >
> > > > > Yeah it is intentional to apply the flushes to the
> > > > > current/old/outgoing context. If there are flushes needed in the new
> > > > > context, they are queued by nested_svm_{entry/exit}_tlb_flush().
> > > > > That's similar to what VMX does.
> > > > >
> > > > > >
> > > > > > Also, is there a context mismatch here during nested VM-Exit?
> > > > > >
> > > > > > In nested_svm_vmexit(), leave_guest_mode(vcpu) is called before
> > > > > > svm_switch_vmcb(svm, &svm->vmcb01).
> > > > > >
> > > > > > Because of this, kvm_service_local_tlb_flush_requests() will see
> > > > > > is_guest_mode(vcpu) as false. If Hyper-V is enabled, this means
> > > > > > kvm_hv_purge_tlb_flush_fifo() will incorrectly target L1's FIFO while the
> > > > > > hardware flushes are actually being applied to L2's vmcb02.
> > > > >
> > > > > Ugh.. yes. This is annoying. kvm_service_local_tlb_flush_requests()
> > > > > needs to be called on both the current/old/outgoing VMCB *and* guest
> > > > > mode. So we'll need to open-code the call in a bunch of places before
> > > > > svm_switch_vmcb() and {enter/leave}_guest_mode(). I really liked
> > > > > putting it in svm_switch_vmcb() together with
> > > > > nested_svm_{entry/exit}_tlb_flush() so that all the TLB flushing logic
> > > > > for nested transitions live in one place and the ordering needs to be
> > > > > handled in one place.
> > > > 
> > > > Maybe we can just re-order the code to always call svm_switch_vmcb()
> > > > before {enter/leave}_guest_mode(). We already do that on the entry
> > > > side, and seems to be straightforward on the exit side.
> > > 
> > > As stated earlier, I'd prefer to explicitly do flushing stuff where it fits from
> > > an architectural perspective.
> > 
> > I did it this way because (as you also stated) it's more robust, and it
> > also documents the ordering requirements:
> > - We need to service local flushes before switching to the new VMCB.
> 
> No, that's not the requirement.  The requirement is that KVM faithlyfully emulates
> the SVM architecture.  For KVM's implementation, that _mostly_ aligns with
> switching between vmcb01 and vmcb02, but that's not a hard guarantee.  Unlike VMX,
> SVM doesn't force KVM doesn't need to switch the active VMCB in order to make
> changes to a VMCB, so nSVM may never end up with as many switches as nVMX that
> don't need to trigger a TLB flush, but conceptually it's still inaccurate.
> 
> And regarding robustness, the flaw Sashiko pointed out regarding servicing pending
> flushes after leave_guest_mode(vcpu) highlights that burying architectural behaviors
> in what are effectively utility functions can be dangerous.  The counter-argument is
> that we ended up with a similar bug in nVMX where KVM straight up forgot to service
> the flushes, but my point is that handling this in svm_switch_vmcb() isn't a silver
> bullet.  And I really don't like adding an arbitrary constraint that the "new" VMCB
> needs to "match" the current L1 vs. L2 mode.
> 
> > - We need to queue new flushes after servicing local flushes.
> > 
> > I like that it's all in one place. All that being said, I did go
> > back-and-forth on this so I am not opposed to open-coding it. Let me know if
> > the above argument swayed you or if you still prefer open-coding.
> 
> I still prefer open-coding the calls to perform/request/serivce flushes.

For the record, I had locally added a warning to svm_switch_vmcb() to
make sure that guest_mode and the active VMCB are consistent when
kvm_service_local_tlb_flush_requests() is called. Given that we are
going with open coding the calls, I will add a wrapper for the
warn+call:

static void nested_svm_service_local_tlb_flushes(struct vcpu *vcpu)
{
        struct vcpu_svm *svm = to_svm(vcpu);

        /*
         * TLB flushes are applied to the VMCB, so apply any pending TLB flushes
         * on the outgoing VMCB before switching to a new one. A TLB flush could
         * purge the relevant HV TLB flush FIFO depending on guest_mode, so make
         * sure the VMCB and guest_mode context is consistent.
         */
        WARN_ON_ONCE(is_guest_mode(vcpu) != (svm->vmcb == svm->nested.vmcb02.ptr));
        kvm_service_local_tlb_flush_requests(vcpu);
}

  reply	other threads:[~2026-07-24 22:27 UTC|newest]

Thread overview: 92+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-06-16  0:41 [RFC PATCH v2 00/25] Optimize nSVM TLB flushes Yosry Ahmed
2026-06-16  0:41 ` [RFC PATCH v2 01/25] KVM: nSVM: Flush the TLB after forcefully leaving nested Yosry Ahmed
2026-06-16  0:41 ` [RFC PATCH v2 02/25] KVM: SVM: Passthrough the number of supported ASIDs Yosry Ahmed
2026-07-14 13:20   ` Yosry Ahmed
2026-07-14 21:28     ` Jim Mattson
2026-07-14 21:43       ` Yosry Ahmed
2026-07-14 23:41         ` Sean Christopherson
2026-07-15  6:32           ` Jim Mattson
2026-07-15 17:45             ` Yosry Ahmed
2026-07-15 18:20               ` Jim Mattson
2026-07-15 19:12                 ` Yosry Ahmed
2026-07-16 23:09                   ` Jim Mattson
2026-07-22 22:09             ` Sean Christopherson
2026-07-22 22:39               ` Jim Mattson
2026-07-23  0:27                 ` Sean Christopherson
2026-07-23  2:46                   ` Jim Mattson
     [not found]                     ` <CALMp9eQi=604LVn=ZMnzoUy55LVb1qorKzPKm5HVK_ACsPuP_A@mail.gmail.com>
2026-07-23 14:07                       ` Sean Christopherson
2026-07-23 15:57                         ` Jim Mattson
2026-07-23 16:32                           ` Yosry Ahmed
2026-07-23 16:54                             ` Jim Mattson
2026-07-23 16:57                               ` Yosry Ahmed
2026-07-23 17:07                                 ` Jim Mattson
2026-07-23 17:26                                   ` Sean Christopherson
2026-07-23 17:34                                     ` Yosry Ahmed
2026-07-15 17:41           ` Yosry Ahmed
2026-06-16  0:41 ` [RFC PATCH v2 03/25] KVM: VMX: Generalize VPID allocation to be vendor-neutral Yosry Ahmed
2026-07-22 22:26   ` Sean Christopherson
2026-07-22 22:36     ` Yosry Ahmed
2026-07-23 13:26       ` Sean Christopherson
2026-06-16  0:41 ` [RFC PATCH v2 04/25] KVM: x86/mmu: Support specifying a minimum TLB tag Yosry Ahmed
2026-07-23  0:28   ` Sean Christopherson
2026-07-23  5:00     ` Yosry Ahmed
2026-07-23 14:15       ` Sean Christopherson
2026-06-16  0:41 ` [RFC PATCH v2 05/25] KVM: SVM: Add helpers to set/clear ASID flush in VMCB Yosry Ahmed
2026-06-16  0:41 ` [RFC PATCH v2 06/25] KVM: SVM: Fallback to flush everything if FLUSHBYASID is not available Yosry Ahmed
2026-07-23  0:29   ` Sean Christopherson
2026-06-16  0:41 ` [RFC PATCH v2 07/25] KVM: SVM: Duplicate pre-run ASID check for SEV and non-SEV guests Yosry Ahmed
2026-06-16  0:41 ` [RFC PATCH v2 08/25] KVM: SEV: Stop using per-vCPU ASID for SEV VMs Yosry Ahmed
2026-06-16  1:06   ` sashiko-bot
2026-06-16 17:50     ` Yosry Ahmed
2026-07-07 21:30       ` Yosry Ahmed
2026-07-23  0:33         ` Sean Christopherson
2026-07-23  5:02           ` Yosry Ahmed
2026-06-16  0:41 ` [RFC PATCH v2 09/25] KVM: SVM: Use a static ASID per vCPU Yosry Ahmed
2026-06-16  1:08   ` sashiko-bot
2026-06-16 17:58     ` Yosry Ahmed
2026-07-07 21:28       ` Yosry Ahmed
2026-06-16  0:41 ` [RFC PATCH v2 10/25] KVM: nSVM: Add a placeholder ASID for L2 Yosry Ahmed
2026-06-16  0:41 ` [RFC PATCH v2 11/25] KVM: x86: hyper-v: Rename kvm_hv_vcpu_purge_flush_tlb() Yosry Ahmed
2026-06-16  0:41 ` [RFC PATCH v2 12/25] KVM: x86: hyper-v: Allow puring all TLB flush FIFOs Yosry Ahmed
2026-06-16  0:41 ` [RFC PATCH v2 13/25] KVM: nSVM: Flush both L1 and L2 ASIDs on KVM_REQ_TLB_FLUSH Yosry Ahmed
2026-06-16  1:05   ` sashiko-bot
2026-06-16 18:00     ` Yosry Ahmed
2026-07-23  0:37       ` Sean Christopherson
2026-06-16  0:41 ` [RFC PATCH v2 14/25] KVM: nSVM: Move svm_switch_vmcb() to nested.c Yosry Ahmed
2026-06-16  0:41 ` [RFC PATCH v2 15/25] KVM: nSVM: Call nested_svm_transition_tlb_flush() on every VMCB switch Yosry Ahmed
2026-07-23  0:42   ` Sean Christopherson
2026-07-23  5:07     ` Yosry Ahmed
2026-06-16  0:41 ` [RFC PATCH v2 16/25] KVM: nSVM: Split nested_svm_transition_tlb_flush() into entry/exit fns Yosry Ahmed
2026-06-16  0:41 ` [RFC PATCH v2 17/25] KVM: nSVM: Service local TLB flushes before nested transitions Yosry Ahmed
2026-06-16  1:20   ` sashiko-bot
2026-06-16 18:06     ` Yosry Ahmed
2026-06-17  0:21       ` Yosry Ahmed
2026-07-23  0:46         ` Sean Christopherson
2026-07-23  5:06           ` Yosry Ahmed
2026-07-23 13:34             ` Sean Christopherson
2026-07-24 22:27               ` Yosry Ahmed [this message]
2026-06-16  0:41 ` [RFC PATCH v2 18/25] KVM: nSVM: Handle nested TLB flush requests through TLB_CONTROL Yosry Ahmed
2026-06-16  0:41 ` [RFC PATCH v2 19/25] KVM: nSVM: Flush the TLB if L1 changes L2's ASID in vmcb12 Yosry Ahmed
2026-06-16  0:41 ` [RFC PATCH v2 20/25] KVM: nSVM: Do not reset TLB_CONTROL in vmcb02 on nested VM-Enter Yosry Ahmed
2026-06-16  0:41 ` [RFC PATCH v2 21/25] KVM: x86/mmu: rename __kvm_mmu_invalidate_addr() Yosry Ahmed
2026-06-16  0:41 ` [RFC PATCH v2 22/25] KVM: x86/mmu: Refactor kvm_mmu_invlpg() to allow skipping the gva flush Yosry Ahmed
2026-07-23  0:53   ` Sean Christopherson
2026-07-23  0:56     ` Sean Christopherson
2026-07-23  5:11       ` Yosry Ahmed
2026-07-23 15:23         ` Sean Christopherson
2026-07-23 21:48           ` Yosry Ahmed
2026-07-23 22:03             ` Sean Christopherson
2026-07-23 22:10               ` Yosry Ahmed
2026-07-24  1:12                 ` Sean Christopherson
2026-07-24 16:45           ` Yosry Ahmed
2026-06-16  0:41 ` [RFC PATCH v2 23/25] KVM: nSVM: Flush L2's ASID when emulating INVLPGA Yosry Ahmed
2026-06-16  0:41 ` [RFC PATCH v2 24/25] KVM: nSVM: Use different ASIDs for L1 and L2 Yosry Ahmed
2026-06-16  1:30   ` sashiko-bot
2026-06-16 18:14     ` Yosry Ahmed
2026-06-16 18:16       ` Yosry Ahmed
2026-06-16 18:28       ` Yosry Ahmed
2026-06-16 19:54         ` Jim Mattson
2026-06-16 19:56           ` Yosry Ahmed
2026-06-16 21:49             ` Yosry Ahmed
2026-06-16  0:41 ` [RFC PATCH v2 25/25] DO NOT MERGE: Add nested_tlb_force_flush Yosry Ahmed
2026-06-16  1:21   ` sashiko-bot

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=amPmtfdm69hw_CM3@google.com \
    --to=yosry@kernel.org \
    --cc=kvm@vger.kernel.org \
    --cc=sashiko-reviews@lists.linux.dev \
    --cc=seanjc@google.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox