All of lore.kernel.org
 help / color / mirror / Atom feed
From: Kyle Zeng <kylebot@openai.com>
To: kvm@vger.kernel.org
Cc: Sean Christopherson <seanjc@google.com>,
	Paolo Bonzini <pbonzini@redhat.com>,
	Thomas Gleixner <tglx@kernel.org>, Ingo Molnar <mingo@redhat.com>,
	Borislav Petkov <bp@alien8.de>,
	Dave Hansen <dave.hansen@linux.intel.com>,
	x86@kernel.org, Chris Ayoub <cayoub@openai.com>,
	Kyle Zeng <kylebot@openai.com>, Oleg Boiko <oboiko@openai.com>
Subject: [PATCH] KVM: x86: Restrict saved GPA writes to hardware write faults
Date: Fri, 28 Aug 2026 12:30:55 -0700	[thread overview]
Message-ID: <20260828193055.59623-1-kylebot@openai.com> (raw)

A GPA supplied by a hardware page fault describes the access that
faulted, not an arbitrary instruction decoded afterwards.  A guest can
change an MMIO read into a store before KVM fetches the instruction.  The
saved-GPA shortcut then skips the write-aware guest page-table walk and
can issue a write through a guest-read-only mapping.

Carry final-write fault information into the emulator and retain it
alongside the saved GPA.  Only reuse that GPA for a write when hardware
reported a final data write.  Reads, faults with unknown access direction,
and implicit guest page-table writes do not authorize an emulated write;
fall back to the existing permission-aware translation in those cases.

Keep the authorization across MMIO completion without decoding a new
instruction, and reset it when initializing a fresh emulator context.
This also covers writes reached through the cmpxchg fallback.

Preserve the shortcut for hardware-authorized writes.  In particular,
legacy SEV guests can have encrypted page tables that KVM cannot walk,
so forcing every emulated write through a software walk is not suitable.

Fixes: 0f89b207b04a ("kvm: svm: Use the hardware provided GPA instead of page walk")
Reported-by: Oleg Boiko <oboiko@openai.com>
Assisted-by: Codex:gpt-5.6-sol
Signed-off-by: Kyle Zeng <kylebot@openai.com>
---
 arch/x86/kvm/kvm_emulate.h |  3 ++-
 arch/x86/kvm/mmu/mmu.c     | 10 ++++++++++
 arch/x86/kvm/x86.c         |  9 ++++++++-
 arch/x86/kvm/x86.h         |  4 ++++
 4 files changed, 24 insertions(+), 2 deletions(-)

diff --git a/arch/x86/kvm/kvm_emulate.h b/arch/x86/kvm/kvm_emulate.h
index 3e375af15c03..aceca9d3a16d 100644
--- a/arch/x86/kvm/kvm_emulate.h
+++ b/arch/x86/kvm/kvm_emulate.h
@@ -354,8 +354,9 @@ struct x86_emulate_ctxt {
 	bool have_exception;
 	struct x86_exception exception;
 
-	/* GPA available */
+	/* GPA and write access supplied by a hardware page fault. */
 	bool gpa_available;
+	bool gpa_write;
 	gpa_t gpa_val;
 
 	/*
diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c
index 064ecc33b926..223698cbb9be 100644
--- a/arch/x86/kvm/mmu/mmu.c
+++ b/arch/x86/kvm/mmu/mmu.c
@@ -6632,6 +6632,16 @@ int noinline kvm_mmu_page_fault(struct kvm_vcpu *vcpu, gpa_t cr2_or_gpa, u64 err
 		return r;
 
 emulate:
+	/*
+	 * A write during a guest page walk does not authorize the instruction
+	 * itself to write to the faulting GPA.  Require a final write access
+	 * before allowing the emulator to reuse the GPA for writes.
+	 */
+	if (direct && (error_code & PFERR_WRITE_MASK) &&
+	    (error_code & PFERR_GUEST_FINAL_MASK) &&
+	    !(error_code & PFERR_GUEST_PAGE_MASK))
+		emulation_type |= EMULTYPE_PF_WRITE;
+
 	return x86_emulate_instruction(vcpu, cr2_or_gpa, emulation_type, insn,
 				       insn_len);
 }
diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c
index 79468ddfe473..8c46aa3468c2 100644
--- a/arch/x86/kvm/x86.c
+++ b/arch/x86/kvm/x86.c
@@ -5088,8 +5088,13 @@ static int emulator_read_write_onepage(unsigned long addr, void *val,
 	 * Note, this cannot be used on string operations since string
 	 * operation using rep will only have the initial GPA from the NPF
 	 * occurred.
+	 *
+	 * The guest may have changed the instruction since the fault.  Reuse
+	 * the GPA for writes only if hardware reported a final write access;
+	 * otherwise, recheck permissions for the decoded access.
 	 */
-	if (ctxt->gpa_available && emulator_can_use_gpa(ctxt) &&
+	if (ctxt->gpa_available && (!write || ctxt->gpa_write) &&
+	    emulator_can_use_gpa(ctxt) &&
 	    (addr & ~PAGE_MASK) == (ctxt->gpa_val & ~PAGE_MASK)) {
 		gpa = ctxt->gpa_val;
 		ret = vcpu_is_mmio_gpa(vcpu, addr, gpa, write);
@@ -5939,6 +5944,7 @@ static void init_emulate_ctxt(struct kvm_vcpu *vcpu)
 	kvm_x86_call(get_cs_db_l_bits)(vcpu, &cs_db, &cs_l);
 
 	ctxt->gpa_available = false;
+	ctxt->gpa_write = false;
 	ctxt->eflags = kvm_get_rflags(vcpu);
 	ctxt->tf = (ctxt->eflags & X86_EFLAGS_TF) != 0;
 
@@ -6455,6 +6461,7 @@ int x86_emulate_instruction(struct kvm_vcpu *vcpu, gpa_t cr2_or_gpa,
 		/* With shadow page tables, cr2 contains a GVA or nGPA. */
 		if (vcpu->arch.mmu->root_role.direct) {
 			ctxt->gpa_available = true;
+			ctxt->gpa_write = emulation_type & EMULTYPE_PF_WRITE;
 			ctxt->gpa_val = cr2_or_gpa;
 		}
 	} else {
diff --git a/arch/x86/kvm/x86.h b/arch/x86/kvm/x86.h
index 0f5919b092e4..2dd456ad3e2d 100644
--- a/arch/x86/kvm/x86.h
+++ b/arch/x86/kvm/x86.h
@@ -407,6 +407,9 @@ int x86_emulate_instruction(struct kvm_vcpu *vcpu, gpa_t cr2_or_gpa,
  * EMULTYPE_PF - Set when an intercepted #PF triggers the emulation, in which case
  *		 the CR2/GPA value pass on the stack is valid.
  *
+ * EMULTYPE_PF_WRITE - Set with EMULTYPE_PF when hardware reports a write to the
+ *		     final GPA, not an implicit guest page-table access.
+ *
  * EMULTYPE_COMPLETE_USER_EXIT - Set when the emulator should update interruptibility
  *				 state and inject single-step #DBs after skipping
  *				 an instruction (after completing userspace I/O).
@@ -445,6 +448,7 @@ int x86_emulate_instruction(struct kvm_vcpu *vcpu, gpa_t cr2_or_gpa,
 #define EMULTYPE_COMPLETE_USER_EXIT (1 << 7)
 #define EMULTYPE_WRITE_PF_TO_SP	    (1 << 8)
 #define EMULTYPE_SKIP_SOFT_INT	    (1 << 9)
+#define EMULTYPE_PF_WRITE	    BIT(10)
 
 #define EMULTYPE_SET_SOFT_INT_VECTOR(v)	((u32)((v) & 0xff) << 16)
 #define EMULTYPE_GET_SOFT_INT_VECTOR(e)	(((e) >> 16) & 0xff)

base-commit: 73e3f0710014fe6d4ed98cfc02292f6121db7558

             reply	other threads:[~2026-08-28 19:31 UTC|newest]

Thread overview: 4+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-28 19:30 Kyle Zeng [this message]
2026-08-28 19:49 ` [PATCH] KVM: x86: Restrict saved GPA writes to hardware write faults sashiko-bot
2026-08-28 21:02   ` Kyle Zeng
2026-08-28 23:12 ` Sean Christopherson

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260828193055.59623-1-kylebot@openai.com \
    --to=kylebot@openai.com \
    --cc=bp@alien8.de \
    --cc=cayoub@openai.com \
    --cc=dave.hansen@linux.intel.com \
    --cc=kvm@vger.kernel.org \
    --cc=mingo@redhat.com \
    --cc=oboiko@openai.com \
    --cc=pbonzini@redhat.com \
    --cc=seanjc@google.com \
    --cc=tglx@kernel.org \
    --cc=x86@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.