From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pj1-f70.google.com (mail-pj1-f70.google.com [209.85.216.70]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5A5F53CB8FC for ; Tue, 25 Aug 2026 21:04:44 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.70 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787691887; cv=none; b=mY46mPGQWlTbe0Q/kpai4ZXtZayAonBFNLfF/srbA6xkYmBDrZiMycUGgIREBYleyVYvoOV9lCJVwM1htVs63KsL2nNNOtiMDz9hb1rv8mpd2H49HFTn/Wx0WXeovatHAYyhcgP2RtW9I53TXBKwwiEzT6j/DrMtiRvfCjy5moI= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787691887; c=relaxed/simple; bh=y3JaxVV/ItmvBl2jv06UEW18IWm8frWqmaUiPg1L5gs=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=hTQ0z7/h141qG+do9jCkzK9ZT3hKhHJHbtrUmxwSWp1hi2edZCwi7EntoSsywqT6dbx90CIXoLP5xoK+vHnbFQlRF7RYy1/GMOk6QfOF5sNUNo2swPTd3+QR65yhyc8cN1/9H0Rs5EhDIigHdrKXdzl5CNjxjVUvPzzpthVBm2Q= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=FJvIl9YH; arc=none smtp.client-ip=209.85.216.70 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="FJvIl9YH" Received: by mail-pj1-f70.google.com with SMTP id 98e67ed59e1d1-38e22137fb3so564045a91.0 for ; Tue, 25 Aug 2026 14:04:44 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1787691884; x=1788296684; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=kdCwpbXCm4diJMO0qCIhObrfaL6xSD7Kcdv0HWKzDdk=; b=FJvIl9YHrrDqcf6WazXNV9GBcqnltCDX++DO6e3rTb6Nl1jmof8rTVaYgtwn8A/9/W 8Xbny2Eu7EVhVI67CMUWR0WTIQklvvoi6UyfuTYTV2WTj/kUAQsrkgwgSyRLlUzRaUZL MQPzFC4gjnbvqYRa2sVbdhKXlVLs3J15QDD0+aCS4XV7Acdo1EamFDCQLMZNmKzNtbTS lfXo7BU7jJOgN+x12hhC9dBr+qMK7t0XKb1PCihdujl1ycTykNG/hq+AtsfW5FBG8qkq i7DT7WbVANUgRqOJDtadHEtrTaN6WaIcNVIakXsN/iSVk2iuvhEklY3n6gvf4KDVijRT 4Rig== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787691884; x=1788296684; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=kdCwpbXCm4diJMO0qCIhObrfaL6xSD7Kcdv0HWKzDdk=; b=Hf9FntopRTeP0HernOJyv2fkRJREVM8gdV7M86GyLDI1AZ+CKvl7d3s8Gip9bYckT4 1Y+TW5hRexdmPnVHHiXgt97bzi8mI5Aetc5ZI4QakML1vqAAwMoosCAz9u1yzgBg/UH8 vx7bEcBtPpwkhyqzWe7Zu4xGksvjkdJQWNt5bD1VOTIhglBCwKwu+SLPXHqdkFkAtraf AHgdvBdZzWH7VHvBUWUXeFWkjr920E6+KOwf9uUOsuVbMBPlDh3fshUH2c95H4zdGIuq ERkCLuctXIk8bSohSmu4buo9V6BezrZPMQlwbrbmsmHABWxPGICW7kkvwVl0mzlTwiGR PQlw== X-Forwarded-Encrypted: i=1; AHgh+Ro3v0PBsgK4LznzKB1+1oeN02CtInpN7e/e/II+sGJzLpnDja40dNLu6d2/l3ruBEmvOjJyn9f5qnRLDhM=@vger.kernel.org X-Gm-Message-State: AFuF++m57XORBUiMje1f/V9r/2t8xWzWOdf6h+gldHtKsXkAPV35BpDx 1hRBSj5eGlEVGRpv7HxWNncgOMrX9rILDk6rmx9upjKvpwIDp8l0owJu361gzS1Jg7v/7GrfLRV uE81Z/g== X-Received: from pgjt20.prod.google.com ([2002:a63:f354:0:b0:cc1:2989:9d19]) (user=seanjc job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a20:cfa8:b0:3cc:8f53:26c0 with SMTP id adf61e73a8af0-3cf86dd6053mr2583459637.17.1787691883538; Tue, 25 Aug 2026 14:04:43 -0700 (PDT) Date: Tue, 25 Aug 2026 14:04:42 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260825140159.70997-1-mamarang@amazon.com> Message-ID: Subject: Re: [RFC] KVM: x86/mmu: Prefetch forward run of pages on TDP page faults From: Sean Christopherson To: James Houghton Cc: Marco Marangoni , Paolo Bonzini , Thomas Gleixner , Ingo Molnar , Borislav Petkov , Dave Hansen , "x86@kernel.org" , "H. Peter Anvin" , "kvm@vger.kernel.org" , "linux-kernel@vger.kernel.org" , "kernel-patches@amazon.com" , Riccardo Mancini , Michael Zoumboulakis , Marco Marangoni Content-Type: text/plain; charset="us-ascii" On Tue, Aug 25, 2026, Sean Christopherson wrote: > Somewhat off the cuff and *very* lightly tested, but this seems to do what I want. > If it provides comparable performance, I'll write a changelog (or two? e.g. to > have direct MMUs switch in a separate patch), and let Sashiko and other bots rip > apart my idea. > > Note! This has a hard dependency on in-flight prefaulting fixes[*]. Without > those, prefaulting will hang the vCPU if the root is invalidated. > [*] https://lore.kernel.org/all/20260806214050.78058-1-seanjc@google.com > > Note #2! The below deliberately ignores A/D-disabled MMUs. I can't think of > any reason why it matters whether or not KVM can precisely detect accessed SPTEs, > all of the aging stuff is already extremely fuzzy. And of course I posted an untested version (I ripped out the direct MMU prefetching as an afterthough, and dropped a printk). This version should actually compile. diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c index 79c450d677b4..88b6aa1f840f 100644 --- a/arch/x86/kvm/mmu/mmu.c +++ b/arch/x86/kvm/mmu/mmu.c @@ -118,6 +118,9 @@ EXPORT_SYMBOL_FOR_KVM_INTERNAL(tdp_mmu_enabled); bool __read_mostly eager_page_split = true; module_param(eager_page_split, bool, 0644); +unsigned int __read_mostly auto_prefault_nr_pages = KVM_PAGES_PER_HPAGE(PG_LEVEL_2M); +module_param(auto_prefault_nr_pages, uint, 0644); + static int max_huge_page_level __read_mostly; static int tdp_root_level __read_mostly; static int max_tdp_level __read_mostly; @@ -3205,69 +3208,6 @@ static bool kvm_mmu_prefetch_sptes(struct kvm_vcpu *vcpu, gfn_t gfn, u64 *sptep, return true; } -static bool direct_pte_prefetch_many(struct kvm_vcpu *vcpu, - struct kvm_mmu_page *sp, - u64 *start, u64 *end) -{ - gfn_t gfn = kvm_mmu_page_get_gfn(sp, spte_index(start)); - unsigned int access = sp->role.access; - - return kvm_mmu_prefetch_sptes(vcpu, gfn, start, end - start, access); -} - -static void __direct_pte_prefetch(struct kvm_vcpu *vcpu, - struct kvm_mmu_page *sp, u64 *sptep) -{ - u64 *spte, *start = NULL; - int i; - - WARN_ON_ONCE(!sp->role.direct); - - i = spte_index(sptep) & ~(PTE_PREFETCH_NUM - 1); - spte = sp->spt + i; - - for (i = 0; i < PTE_PREFETCH_NUM; i++, spte++) { - if (is_shadow_present_pte(*spte) || spte == sptep) { - if (!start) - continue; - if (!direct_pte_prefetch_many(vcpu, sp, start, spte)) - return; - - start = NULL; - } else if (!start) - start = spte; - } - if (start) - direct_pte_prefetch_many(vcpu, sp, start, spte); -} - -static void direct_pte_prefetch(struct kvm_vcpu *vcpu, u64 *sptep) -{ - struct kvm_mmu_page *sp; - - sp = sptep_to_sp(sptep); - - /* - * Without accessed bits, there's no way to distinguish between - * actually accessed translations and prefetched, so disable pte - * prefetch if accessed bits aren't available. - */ - if (sp_ad_disabled(sp)) - return; - - if (sp->role.level > PG_LEVEL_4K) - return; - - /* - * If addresses are being invalidated, skip prefetching to avoid - * accidentally prefetching those addresses. - */ - if (unlikely(vcpu->kvm->mmu_invalidate_in_progress)) - return; - - __direct_pte_prefetch(vcpu, sp, sptep); -} - /* * Lookup the mapping level for @gfn in the current mm. * @@ -3502,8 +3442,8 @@ static int direct_map(struct kvm_vcpu *vcpu, struct kvm_page_fault *fault) { struct kvm_shadow_walk_iterator it; struct kvm_mmu_page *sp; - int ret, access; gfn_t base_gfn = fault->gfn; + int access; kvm_mmu_hugepage_adjust(vcpu, fault); @@ -3534,13 +3474,8 @@ static int direct_map(struct kvm_vcpu *vcpu, struct kvm_page_fault *fault) if (WARN_ON_ONCE(it.level != fault->goal_level)) return -EFAULT; - ret = mmu_set_spte(vcpu, fault->slot, it.sptep, access, - base_gfn, fault->pfn, fault); - if (ret == RET_PF_SPURIOUS) - return ret; - - direct_pte_prefetch(vcpu, it.sptep); - return ret; + return mmu_set_spte(vcpu, fault->slot, it.sptep, access, base_gfn, + fault->pfn, fault); } static void kvm_send_hwpoison_signal(struct kvm_memory_slot *slot, gfn_t gfn) @@ -6580,11 +6515,38 @@ static int kvm_mmu_write_protect_fault(struct kvm_vcpu *vcpu, gpa_t cr2_or_gpa, return RET_PF_EMULATE; } +static void kvm_mmu_auto_prefault(struct kvm_vcpu *vcpu, gpa_t start, + u64 error_code, u8 level) +{ + gfn_t nr_pages = READ_ONCE(auto_prefault_nr_pages); + int nr_pages_msb; + gfn_t i; + + if (unlikely(error_code & PFERR_RSVD_MASK)) + return; + + nr_pages = min(nr_pages, KVM_PAGES_PER_HPAGE(PG_LEVEL_1G)); + nr_pages_msb = find_last_bit((unsigned long *)&nr_pages, sizeof(nr_pages)); + + start = ALIGN_DOWN(start, gfn_to_gpa(BIT_ULL(nr_pages_msb))); + + for (i = KVM_PAGES_PER_HPAGE(level); i < nr_pages; i += KVM_PAGES_PER_HPAGE(level)) { + gpa_t gpa = start + gfn_to_gpa(i); + + if (gpa < start || gpa_to_gfn(gpa) > kvm_mmu_max_gfn()) + return; + + if (kvm_tdp_page_prefault(vcpu, gpa, error_code, &level)) + return; + } +} + int noinline kvm_mmu_page_fault(struct kvm_vcpu *vcpu, gpa_t cr2_or_gpa, u64 error_code, void *insn, int insn_len) { int r, emulation_type = EMULTYPE_PF; bool direct = vcpu->arch.mmu->root_role.direct; + u8 level; if (WARN_ON_ONCE(!VALID_PAGE(vcpu->arch.mmu->root.hpa))) return RET_PF_RETRY; @@ -6617,7 +6579,7 @@ int noinline kvm_mmu_page_fault(struct kvm_vcpu *vcpu, gpa_t cr2_or_gpa, u64 err vcpu->stat.pf_taken++; r = kvm_mmu_do_page_fault(vcpu, cr2_or_gpa, error_code, false, - &emulation_type, NULL); + &emulation_type, &level); if (KVM_BUG_ON(r == RET_PF_INVALID, vcpu->kvm)) return -EIO; } @@ -6628,6 +6590,8 @@ int noinline kvm_mmu_page_fault(struct kvm_vcpu *vcpu, gpa_t cr2_or_gpa, u64 err if (r == RET_PF_WRITE_PROTECTED) r = kvm_mmu_write_protect_fault(vcpu, cr2_or_gpa, error_code, &emulation_type); + else if (r == RET_PF_FIXED) + kvm_mmu_auto_prefault(vcpu, cr2_or_gpa, error_code, level); if (r == RET_PF_FIXED) vcpu->stat.pf_fixed++; diff --git a/arch/x86/kvm/mmu/paging_tmpl.h b/arch/x86/kvm/mmu/paging_tmpl.h index 27427e7f22fa..b41b4b78cc81 100644 --- a/arch/x86/kvm/mmu/paging_tmpl.h +++ b/arch/x86/kvm/mmu/paging_tmpl.h @@ -617,7 +617,7 @@ static void FNAME(pte_prefetch)(struct kvm_vcpu *vcpu, struct guest_walker *gw, sp = sptep_to_sp(sptep); - if (sp->role.level > PG_LEVEL_4K) + if (sp->role.level > PG_LEVEL_4K || sp->role.direct) return; /* @@ -627,9 +627,6 @@ static void FNAME(pte_prefetch)(struct kvm_vcpu *vcpu, struct guest_walker *gw, if (unlikely(vcpu->kvm->mmu_invalidate_in_progress)) return; - if (sp->role.direct) - return __direct_pte_prefetch(vcpu, sp, sptep); - i = spte_index(sptep) & ~(PTE_PREFETCH_NUM - 1); spte = sp->spt + i;