kvm.vger.kernel.org archive mirror
 help / color / mirror / Atom feed
From: Hao Zhang <hao_zhang_kdev@163.com>
To: Sean Christopherson <seanjc@google.com>
Cc: Hao Zhang <hao_zhang_kdev@163.com>,
	Paolo Bonzini <pbonzini@redhat.com>,
	kvm@vger.kernel.org
Subject: Re: [PATCH v3] KVM: x86: Mark nested RSM-created VM-Enter as untrusted
Date: Wed, 22 Jul 2026 09:26:17 +0800	[thread overview]
Message-ID: <amAcObySFW2LOtXH@192.168.1.215> (raw)
In-Reply-To: <al-SKXoha-dqz047@google.com>

On Tue, Jul 21, 2026, Sean Christopherson wrote:
> On Tue, Jul 21, 2026, Hao Zhang wrote:
> > From: Hao Zhang <zhanghao1@kylinos.cn>
> > 
> > RSM can restore state that indicates the vCPU was running L2 when SMI
> > arrived.  KVM then re-enters nested guest mode and marks nested_run_pending
> > so that the next KVM_RUN completes the pending nested transition.
> > 
> > Unlike a normal nested VMLAUNCH/VMRESUME, the state restored by RSM comes
> > from SMRAM.  Userspace can modify SMRAM before RSM.  The resulting L2 state
> > may require invalid-guest-state emulation.  If KVM trusts that pending run,
> > VMX can synthesize EXIT_REASON_INVALID_STATE before hardware VM-Entry
> > completes the pending run and trip the nested_run_pending invariant.
> > 
> > Do not treat an RSM-created pending nested run as KVM-trusted state.  Mark
> > the run as KVM_NESTED_RUN_PENDING_UNTRUSTED for both VMX and SVM, matching
> > the existing userspace-modified-state handling in kvm_x86_vcpu_pre_run().
> > 
> > Keep the VMX BUG check for trusted pending nested VM-Enters, but allow
> > untrusted pending state to reach the existing invalid-L2-state handling.
> > If that handling synthesizes a nested VM-Exit back to L1, retain the
> > pending state while synchronizing VMCS02 to VMCS12, and consume it before
> > restoring L1 state so that the flag does not leak into L1 and block event
> > injection.
> > 
> > Fixes: 2bb8cafea80b ("KVM: vVMX: signal failure for nested VMEntry if emulation_required")
> > Suggested-by: Sean Christopherson <seanjc@google.com>
> > Signed-off-by: Hao Zhang <zhanghao1@kylinos.cn>
> > ---
> > diff --git a/arch/x86/kvm/vmx/vmx.c b/arch/x86/kvm/vmx/vmx.c
> > index cc75feec05da..cf5c3f78ebfc 100644
> > --- a/arch/x86/kvm/vmx/vmx.c
> > +++ b/arch/x86/kvm/vmx/vmx.c
> > @@ -6715,12 +6715,14 @@ static int __vmx_handle_exit(struct kvm_vcpu *vcpu, fastpath_t exit_fastpath)
> >  		return 0;
> >  
> >  	/*
> > -	 * KVM should never reach this point with a pending nested VM-Enter.
> > -	 * More specifically, short-circuiting VM-Entry to emulate L2 due to
> > -	 * invalid guest state should never happen as that means KVM knowingly
> > -	 * allowed a nested VM-Enter with an invalid vmcs12.  More below.
> > -	 */
> > -	if (KVM_BUG_ON(vcpu->arch.nested_run_pending, vcpu->kvm))
> > +	 * KVM should never reach this point with a trusted pending nested
> > +	 * VM-Enter.  Short-circuiting VM-Entry to emulate L2 due to invalid
> > +	 * guest state after a normal nested VM-Enter would mean KVM knowingly
> > +	 * allowed an invalid vmcs12.  Pending VM-Enters restored from untrusted
> > +	 * state are handled below.
> > +	 */
> > +	if (KVM_BUG_ON(vcpu->arch.nested_run_pending == KVM_NESTED_RUN_PENDING,
> > +		       vcpu->kvm))
> 
> Ugh, I missed that my suggestion would require adjusting this KVM_BUG_ON().
> That's not the end of the world, but it's not great either because it further
> degrades the sanity checks provided by nested_run_pending.
> 
> And while Sashiko was hallucinating wildly with respect to a "mailicious L2
> modifying SMRAM", Sashiko did have a point about leaving nested_run_pending set
> beyond the emulated RSM.  KVM really should check for the impossible state after
> loading vCPU state from SMRAM in emulator_leave_smm(), so that forcibly leaving
> nested mode to synthesize TRIPLE_FAULT into L1 naturally Just Works.
> 
> The only reason I didn't suggest that is because I didn't want to add another
> SMM-specific kvm_x86_ops hook, but I think we can avoid that and (arguably)
> improve other paths at the same time.
> 
> Rather than provide a generic .vcpu_pre_run() hook, which is somewhat confusing
> because SVM has it's own pre_svm_run(), what if we provide separate hooks for
> unhandleable_emulation_required() and vcpu_needs_initialization()?  That bleeds
> some vendor specific details into common x86, but in many ways, that's a _good_
> thing in this case because I think it's helpful to highlight the need for
> additional initialization for CoCo vCPUs.
> 
> Do you happen to have a reproducer I can test?  Either way, I'll try to add a
> testcase to tools/testing/selftests/kvm/x86/vmx_invalid_nested_guest_state.c to
> cover the SMRAM => RSM path.
> 

I don't have a standalone C reproducer, but I do have the syzkaller program
that hit the original WARNING.  It uses syz_kvm_setup_cpu$x86 and includes the
RSM path; I used it via syz-repro against the original crash bucket.

The repro seed is below.  It is not yet converted into a KVM selftest, but it
should be useful as a starting point for the SMRAM => RSM coverage in
vmx_invalid_nested_guest_state.c.

r0 = openat$kvm(0xffffffffffffff9c, &(0x7f0000000000), 0x2, 0x0)
r1 = ioctl$KVM_CREATE_VM(r0, 0xae01, 0x0)
r2 = ioctl$KVM_CREATE_VCPU(r1, 0xae41, 0x1)
syz_kvm_setup_cpu$x86(r1, r2, &(0x7f0000000000/0x18000)=nil, &(0x7f0000000080)=[@text64={0x40, 0x0}], 0x1, 0x54, 0x0, 0x0)
madvise(&(0x7f000010c000/0xe000)=nil, 0xe000, 0x66)
ioctl$KVM_RUN(r2, 0xae80, 0x0)
r3 = openat$kvm(0xffffffffffffff9c, &(0x7f0000000000), 0x2, 0x0)
r4 = ioctl$KVM_CREATE_VM(r3, 0xae01, 0x0)
r5 = ioctl$KVM_CREATE_VCPU(r4, 0xae41, 0x0)
syz_kvm_setup_cpu$x86(r4, r5, &(0x7f0000000000/0x18000)=nil, &(0x7f0000000080)=[@text32={0x20, &(0x7f0000000180)="6526f30f09660faade7e07b805000000b9ce8400000f01c1dd5d140f00d4ea000000008600440f20c03504000000440f22c0360f01d10f07", 0x38}], 0x1, 0x20, 0x0, 0x0)
syz_kvm_setup_cpu$x86(r1, r2, &(0x7f0000d9e000/0x18000)=nil, &(0x7f00000001c0)=[@text32={0x20, &(0x7f0000000240)="66ba4000edb8cd0000000f23d00f21f835100000020f23f8b99f0b00000f32dd1f66ba6100ec670f00d3b9800000c00f3235000100000f30c4e2e19274b900b9800000c00f3235008000000f3026650fc76c5589", 0x54}], 0x1, 0x8, &(0x7f0000000200)=[@cstype0={0x4, 0x8}], 0x1)
r6 = openat$kvm(0xffffffffffffff9c, &(0x7f0000000080), 0x0, 0x0)
r7 = ioctl$KVM_CREATE_VM(r6, 0xae01, 0x0)
r8 = ioctl$KVM_CREATE_VCPU(r7, 0xae41, 0x0)
syz_kvm_setup_cpu$x86(r7, r8, &(0x7f0000fe1000/0x18000)=nil, &(0x7f0000000040)=[@text64={0x40, &(0x7f00000000c0)="0f6f2848b802000000000000000f23d80f21f835c00000b00f23f8470f220240a6dae9643644a5410f01c80f0fb500800000aeb9cb020000b851320000ba000000000f300f006b00", 0x48}], 0x1, 0x51, 0x0, 0x0)
ioctl$KVM_RUN(r8, 0xae80, 0x0)
syz_kvm_setup_cpu$x86(r1, r2, &(0x7f0000fe5000/0x18000)=nil, &(0x7f00000002c0)=[@textreal={0x8, 0x0}], 0x1, 0x22, 0x0, 0x0)
ioctl$KVM_RUN(r2, 0xae80, 0x0)

--
Thanks,
Hao


  reply	other threads:[~2026-07-22  1:26 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-21  8:40 [PATCH v3] KVM: x86: Mark nested RSM-created VM-Enter as untrusted Hao Zhang
2026-07-21 15:37 ` Sean Christopherson
2026-07-22  1:26   ` Hao Zhang [this message]
2026-07-27  7:53   ` Hao Zhang
2026-07-28  0:26     ` Sean Christopherson

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=amAcObySFW2LOtXH@192.168.1.215 \
    --to=hao_zhang_kdev@163.com \
    --cc=kvm@vger.kernel.org \
    --cc=pbonzini@redhat.com \
    --cc=seanjc@google.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).