Kernel KVM virtualization development
 help / color / mirror / Atom feed
From: Kyle Zeng <kylebot@openai.com>
To: Sean Christopherson <seanjc@google.com>
Cc: sashiko-reviews@lists.linux.dev, kvm@vger.kernel.org
Subject: Re: [PATCH v3] KVM: x86: Restrict saved GPA writes to hardware write faults
Date: Thu, 17 Sep 2026 20:33:44 -0700	[thread overview]
Message-ID: <aqyxGC1kwlFO8Z4x@com-75606> (raw)
In-Reply-To: <aqrqb-lz2f0gCwuF@google.com>

On Wed, Sep 16, 2026 at 12:13:51PM -0700, Sean Christopherson wrote:
> On Tue, Sep 15, 2026, Kyle Zeng wrote:
> > On Tue, Sep 15, 2026 at 10:41:03PM +0000, sashiko-bot@kernel.org wrote:
> > > > @@ -6454,7 +6456,18 @@ int x86_emulate_instruction(struct kvm_vcpu *vcpu, gpa_t cr2_or_gpa,
> > > >  
> > > >  		/* With shadow page tables, cr2 contains a GVA or nGPA. */
> > > >  		if (vcpu->arch.mmu->root_role.direct) {
> > > > -			ctxt->gpa_available = true;
> > > > +			ctxt->gpa_access = ACC_READ_MASK;
> > > > +			/*
> > > > +			 * Always allow writes for SEV guests, as the guest's
> > > > +			 * page tables are encrypted, i.e. KVM can't walk the
> > > > +			 * guest's page tables and so must always use the GPA
> > > > +			 * from the initial fault.  Restricting use of the GPA
> > > > +			 * to the access type that faulted would prevent KVM
> > > > +			 * from emulating RMW operations for SEV guests.
> > > > +			 */
> > > > +			if ((emulation_type & EMULTYPE_PF_WRITE) ||
> > > > +			    (cpu_feature_enabled(X86_FEATURE_SVM) && is_sev_guest(vcpu)))
> > > 
> > > [Severity: High]
> > > Could this lead to an out-of-bounds read and type confusion?
>  
> ...
> 
> > Is there a configuration that gets past the initialization checks and
> > demonstrates an incorrect access? I'd like to establish that before
> > treating this as a functional issue.
> 
> Eh, doesn't matter, my suggestion was terrible.  I forgot that is_sev_guest()
> needs to check embedded data; I was thinking of TDX and SNP VMs, which have
> dedicated VM types.
> 
> I think the right way to handle this is to track if a VM has protected page
> tables.  That'd also help communicate/document why KVM has this weird behavior.
> 
> Untested...
> 
> diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h
> index 30ffaa65f589..90c950c12900 100644
> --- a/arch/x86/include/asm/kvm_host.h
> +++ b/arch/x86/include/asm/kvm_host.h
> @@ -1166,6 +1166,7 @@ struct kvm_arch {
>  	u8 mmu_valid_gen;
>  	u8 vm_type;
>  	bool has_private_mem;
> +	bool has_protected_page_tables;
>  	bool has_protected_state;
>  	bool has_protected_eoi;
>  	bool has_protected_pmu;
> diff --git a/arch/x86/kvm/svm/sev.c b/arch/x86/kvm/svm/sev.c
> index b56c39c53c98..7f544658b1cc 100644
> --- a/arch/x86/kvm/svm/sev.c
> +++ b/arch/x86/kvm/svm/sev.c
> @@ -2955,6 +2955,7 @@ void sev_vm_init(struct kvm *kvm)
>  		kvm->arch.has_protected_state = true;
>  		fallthrough;
>  	case KVM_X86_SEV_VM:
> +		kvm->arch.has_protected_page_tables = true;
>  		kvm->arch.pre_fault_allowed = !kvm->arch.has_private_mem;
>  		to_kvm_sev_info(kvm)->need_init = true;
>  		break;
> diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c
> index 014545c1839b..35022f308e06 100644
> --- a/arch/x86/kvm/vmx/tdx.c
> +++ b/arch/x86/kvm/vmx/tdx.c
> @@ -616,6 +616,7 @@ int tdx_vm_init(struct kvm *kvm)
>  {
>  	struct kvm_tdx *kvm_tdx = to_kvm_tdx(kvm);
>  
> +	kvm->arch.has_protected_page_tables = true;
>  	kvm->arch.has_protected_state = true;
>  	/*
>  	 * TDX Module doesn't allow the hypervisor to modify the EOI-bitmap,
> diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c
> index f36b952624c8..7e5e2e587086 100644
> --- a/arch/x86/kvm/x86.c
> +++ b/arch/x86/kvm/x86.c
> @@ -33,7 +33,6 @@
>  #include "lapic.h"
>  #include "xen.h"
>  #include "smm.h"
> -#include "svm/svm.h"
>  
>  #include <linux/clocksource.h>
>  #include <linux/timekeeping.h>
> @@ -6412,15 +6411,14 @@ int x86_emulate_instruction(struct kvm_vcpu *vcpu, gpa_t cr2_or_gpa,
>  		if (vcpu->arch.mmu->root_role.direct) {
>  			ctxt->gpa_access = ACC_READ_MASK;
>  			/*
> -			 * Always allow writes for SEV guests, as the guest's
> -			 * page tables are encrypted, i.e. KVM can't walk the
> -			 * guest's page tables and so must always use the GPA
> -			 * from the initial fault.  Restricting use of the GPA
> -			 * to the access type that faulted would prevent KVM
> -			 * from emulating RMW operations for SEV guests.
> +			 * Always allow writes for guests with protected page
> +			 * tables, as KVM can't walk the guest's page tables,
> +			 * i.e. KVM can't get the RMW protections for a given
> +			 * GVA to see if the write side of a RMW operation
> +			 * should be allowed.
>  			 */
>  			if ((emulation_type & EMULTYPE_PF_WRITE) ||
> -			    (cpu_feature_enabled(X86_FEATURE_SVM) && is_sev_guest(vcpu)))
> +			    vcpu->kvm->arch.has_protected_page_tables)
>  				ctxt->gpa_access |= ACC_WRITE_MASK;
>  			ctxt->gpa_val = cr2_or_gpa;
>  		}
> 

Hi Sean,

I think the proposed patch will cause regression in legacy-SEV.

The new flag records whether KVM can read a guest´s page tables. The problem is that the patch sets it only when the VM is created, but older SEV APIs enable encryption later.

A modern SEV VM follows this sequence:
  Create VM with type KVM_X86_SEV_VM
  -> sev_vm_init() sets has_protected_page_tables = true
  -> initialize SEV

A legacy SEV VM follows a different sequence:
  Create VM with type KVM_X86_DEFAULT_VM
  -> flag stays false
  -> KVM_SEV_INIT enables SEV
  -> page tables are encrypted, but the flag is still false

The `is_sev_guest()` check recognizes the second case because it reads SEV´s actual active state. The proposed replacement would miss it.
That matters for an MMIO read-modify-write instruction. If hardware faults on the initial read, KVM grants the saved GPA only read permission. When emulation reaches the write, the incorrect flag prevents the SEV exception from applying. KVM then tries to check write permission by walking the guest´s encrypted page tables, which can fail and break the guest.
An ordinary hardware write fault still works because EMULTYPE_PF_WRITE grants write access independently.
The fix is to set the new flag when SEV initialization succeeds. Copying or migrating an encryption context must also set it on the destination VM. Setting it for every DEFAULT_VM would be wrong, because ordinary guests would then regain the vulnerable write shortcut.

V4 will be sent separately to address this issue.

Best,
Kyle

  reply	other threads:[~2026-09-18  3:33 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-15 22:26 [PATCH v3] KVM: x86: Restrict saved GPA writes to hardware write faults Kyle Zeng
2026-09-15 22:41 ` sashiko-bot
2026-09-15 22:56   ` Kyle Zeng
2026-09-16 19:13     ` Sean Christopherson
2026-09-18  3:33       ` Kyle Zeng [this message]
2026-09-16 19:16 ` Sean Christopherson

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=aqyxGC1kwlFO8Z4x@com-75606 \
    --to=kylebot@openai.com \
    --cc=kvm@vger.kernel.org \
    --cc=sashiko-reviews@lists.linux.dev \
    --cc=seanjc@google.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox