From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pl1-f182.google.com (mail-pl1-f182.google.com [209.85.214.182]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A9F3C21FF2A for ; Sat, 1 Aug 2026 09:51:53 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.182 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785577916; cv=none; b=a7i36EvG9QAlraLVjlFH+IyNbrKzznrD+4ij0ZyoLLDuTpyMz+wjN3kGv9HCZXk7lSRDID0RxtIprQcdGZXxsngv1R21wZyAcRYJ6owhuGuV2vUJsi1rlCYhBkW3vY/kMnWBaQFSYzJWKHTehBtJxX6nR8RPTBPXObRe0070fIY= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785577916; c=relaxed/simple; bh=yKljj4SYUug21tLfqZ9l0H49iuoeZFS5V5gXTGXAKVg=; h=Date:From:To:Cc:Subject:Message-ID:MIME-Version:Content-Type: Content-Disposition; b=TtA2OXj+GStDN3LUzdQbz8/RVx2YIZXQJXd92ceZff+TkWzc0uLLpejHsWD23Hyj+ESOM6rbmBb63U3Jvorent2jmJq4RMXu3C/6cly755Khmm4v5d2XCJPJGFhQI5XCt1OzvtlvUXjgUE62oK1l1qzL2QAj/C7gfsbAQvQZSMQ= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=Oc60jGEB; arc=none smtp.client-ip=209.85.214.182 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="Oc60jGEB" Received: by mail-pl1-f182.google.com with SMTP id d9443c01a7336-2ce98cb8165so19121235ad.1 for ; Sat, 01 Aug 2026 02:51:53 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1785577913; x=1786182713; darn=vger.kernel.org; h=content-disposition:content-type:mime-version:message-id:subject:cc :to:from:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=IvDuzw2HHT6P+745ujwE5DH7nkh13wNIypUn+wAvqck=; b=Oc60jGEB/HbSdPqiyig0v+oQPdyM4CUN53LYfgXVfrxdj3yt4xl3JNBbnDrouD+my9 PLrzt+7pt6dXRIv2HegduoFpGrv1j4ExUZfgpZey0cdf1pMbEclt8oR+eKj/4iBf5HJs 9zxDBZqwUeeoTV9MbSMx3/LeXnEb5UZlsBaASGVyZ88qw4r5BpxWGTEKnxjpi1PbgHgd c21mVJHjC70FXCuqKj8yE5lSrLhJxTwTL8C3egK34tE7cGI9zrkDDZkJqHYI8tKxZ3w5 mWeQolRD2PLnEWNYU+G0GjloWfppQaTitbSJfRDLObnS1A0YEMrl7Pon1KhHBniGpMEN +HYA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785577913; x=1786182713; h=content-disposition:content-type:mime-version:message-id:subject:cc :to:from:date:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=IvDuzw2HHT6P+745ujwE5DH7nkh13wNIypUn+wAvqck=; b=MqBiJMhwpzy5c9kv3oGs+hGITHbxQANl+D7focq5vTNHsG0MofOMd/8aK7hLya3cK7 NFMGa4D6nWlAiSeK/g5+26T2+an9hclxW5hBl7/wvU0RJczzxtMbR/y91vWxWxMb2HuT mN+7sa31AbakZb7fA8CDrfzQluuq2YuuKeq/7UkQNdkODSMILU+Mcl4+s2LAK32i0VP+ C6kH8Kdb87sItqR7WuaAfj8hYBIUN3vozAHk0JSheKD2YQ49pu9gz35vIfW49EkoTrMW rqWPlIwZ3sEvWcQ/Qx6xbjU9NZm3nFy7fLp3ZvXonTQqrHrpl1bPPxMxGraAtBf1iGqC rP5g== X-Forwarded-Encrypted: i=1; AHgh+RoLOMMwBD+4q0YNWyutlDW3hQrdCCnbquSjGyYDWpADms2rv5qg4ROoRPcabSiTQNqDhzA9wly9N2UMdw8=@vger.kernel.org X-Gm-Message-State: AOJu0YxNi9Y3diEDMft6Dn2F+0oY2hW1CFD5I10QPgkSqRh3P7PeOBPf 6dmYFPQY+oVJBMmcDEu0uRrnTX7k1dd/ahnz0kuvev/XsbTCWCDPcVj9 X-Gm-Gg: AR+sD13F+4tyziKllx1MxHI9LYInVG2KF3K9RwWUYKgfteGOWtHYGr8TlMTUxogSIRM Be2t6iAzuo/bMI7Nt6K368F1PpE/Ng//HWDtz91NR2twspDK0SwIFJpnjftYzHAIkFCSkiT6Cre gTdLA2OHNsrdSDVqvfuANtxIpFoxksR4/3yvhJPZGCy2OHWs5z9zNhRZaTJcc7rBjxgOcIOu15N DstkDCOpsX0ew4ThnHh1BuYtK/kydF/Cne117GA/H+0UDFjEorEgY+HXBKCdxeEgYpcjOZ+AkeN QWTqZT/z0jUNjHvQC5kYghcAcKcGLvQywjQeBCKB3cRh1fo3mbIEYySiNyI1s8yI6MCNywPIuAa NFOu17JOSDcGQywrvXyGsdxVsS8wvUgy++oL07t0hDG1W6FFkrzCaIlwwbLe4wGwWOI7Qaz2XRW vlZOGOWpOI2fJMlxICv8zLk5Eo1B4v+um/M+T7IiaOEcmW7d3pQCiZ+wT2CRw5Klx/uQrk28905 LP8MXL18XJBVhFSWVPv X-Received: by 2002:a17:903:3905:b0:2cc:80d6:b72 with SMTP id d9443c01a7336-2d0536caf12mr24874305ad.20.1785577912888; Sat, 01 Aug 2026 02:51:52 -0700 (PDT) Received: from v4bel ([58.123.110.97]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2d04b0eb598sm16202195ad.51.2026.08.01.02.51.49 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 01 Aug 2026 02:51:52 -0700 (PDT) Date: Sat, 1 Aug 2026 18:51:47 +0900 From: Hyunwoo Kim To: seanjc@google.com, pbonzini@redhat.com, tglx@kernel.org, mingo@redhat.com, bp@alien8.de, dave.hansen@linux.intel.com, x86@kernel.org, dwmw2@infradead.org Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org, imv4bel@gmail.com Subject: [PATCH] KVM: nVMX: Don't load L1's host state when freeing a vCPU Message-ID: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline Don't load L1's host state when kicking a vCPU out of nested guest mode as part of freeing the vCPU, as loading host state processes vmcs12's VM-Exit MSR load list, i.e. reads an (index, value) pair out of guest memory and feeds it to kvm_emulate_msr_write() with host_initiated=false. Letting the guest emulate WRMSR against VM-scope state that KVM is actively tearing down goes sideways in at least two ways. Writing HV_X64_MSR_ICR sends an IPI, which for a non-shorthand, non-broadcast destination walks kvm->arch.apic_map to dereference the target's local APIC. kvm_free_lapic() neither rebuilds nor dirties the map, and the map is freed only after all vCPUs are destroyed, i.e. the map still points at the already-freed local APIC of a previously destroyed vCPU. This requires userspace to expose Hyper-V's CPUID to the guest. Writing MSR_KVM_SYSTEM_TIME_NEW activates the kvmclock gfn=>pfn cache, which leaves the cache's list entry, resident in the about-to-be-freed vCPU, linked into kvm->gpc_list; the next vCPU to manipulate the list writes through that entry. The vCPU will never run again, so nothing can observe the loaded host state. Simply restore KVM's MMU pointers so that they aren't left pointing at the nested MMU, and bail. nSVM does the same, i.e. doesn't emulate a VM-Exit when forcibly leaving nested mode, and performs only the equivalent MMU cleanup. Bail just before the branch that splits the success and VM-Fail paths, as leaving guest mode, canceling the VMX-preemption timer, and switching back to vmcs01 are all needed by the free path. Canceling the timer is in fact the only reason the free path goes through an emulated VM-Exit, see commit b4b65b5642d6 ("KVM: x86: cleanup freeing of nested state"). Fixes: b4b65b5642d6 ("KVM: x86: cleanup freeing of nested state") Cc: stable@vger.kernel.org Signed-off-by: Hyunwoo Kim --- arch/x86/kvm/vmx/nested.c | 12 ++++++++++++ arch/x86/kvm/vmx/vmx.h | 3 +++ 2 files changed, 15 insertions(+) diff --git a/arch/x86/kvm/vmx/nested.c b/arch/x86/kvm/vmx/nested.c index ddf6df7bee93b2..8d58547acb1887 100644 --- a/arch/x86/kvm/vmx/nested.c +++ b/arch/x86/kvm/vmx/nested.c @@ -384,6 +384,8 @@ static void free_nested(struct kvm_vcpu *vcpu) */ void nested_vmx_free_vcpu(struct kvm_vcpu *vcpu) { + to_vmx(vcpu)->nested.vcpu_is_dying = true; + vcpu_load(vcpu); vmx_leave_nested(vcpu); vcpu_put(vcpu); @@ -5173,6 +5175,16 @@ void __nested_vmx_vmexit(struct kvm_vcpu *vcpu, u32 vm_exit_reason, /* in case we halted in L2 */ kvm_set_mp_state(vcpu, KVM_MP_STATE_RUNNABLE); + /* + * Don't emulate guest-controlled state, e.g. vmcs12's VM-Exit MSR load + * list, when freeing the vCPU. Bail only after leaving guest mode, + * canceling the preemption timer, and switching back to vmcs01. + */ + if (vmx->nested.vcpu_is_dying) { + nested_ept_uninit_mmu_context(vcpu); + return; + } + if (likely(!vmx->fail)) { if (vm_exit_reason != -1) trace_kvm_nested_vmexit_inject(vmcs12->vm_exit_reason, diff --git a/arch/x86/kvm/vmx/vmx.h b/arch/x86/kvm/vmx/vmx.h index dc8517f15bc463..2bacd3fe4c7ded 100644 --- a/arch/x86/kvm/vmx/vmx.h +++ b/arch/x86/kvm/vmx/vmx.h @@ -76,6 +76,9 @@ struct nested_vmx { gpa_t vmxon_ptr; bool pml_full; + /* Set when freeing the vCPU, to suppress emulation of guest state. */ + bool vcpu_is_dying; + /* The guest-physical address of the current VMCS L1 keeps for L2 */ gpa_t current_vmptr; /* -- 2.43.0