From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 556D330BF70; Mon, 31 Aug 2026 00:25:21 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788135921; cv=none; b=SqykyjpRzsJN5E13hFLH1mDW31S1o6agrRGCgrHeMIkBsaQlbtT9WJIN9NnQPc8pVZH4DFVOFrsfQYtM8mV8DGl3nnj98iANQ7blgofAYKPfrG13qUne0IsVqKUa7GOpmxZ7GlWLbCfSUTkDKQClQet0Esi35W6GXD8Dh1wBjlM= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788135921; c=relaxed/simple; bh=z59hdgGozA63z31zC23BZ4GXqCsjybzsO4sM4OB80Dg=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=dBb3Ff5Lc+E8xr7Itzeqs1cqq8beHBgJBe0P2Xls//hmOaMqDto7gjvQ7eWgJoJYEHISFTxIoIQVZXseT8QuW83s4Di5A+2whfnMuLXuPvhqSpiq0xmffBq4StZRlhB8owhF0qf5Vbv04I92xdIfTbHG2S3+8xvSn13lLA3aoL4= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=DpR8VEQ0; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="DpR8VEQ0" Received: by smtp.kernel.org (Postfix) with ESMTPS id 24482C4DDE7; Mon, 31 Aug 2026 00:25:21 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1788135921; bh=z59hdgGozA63z31zC23BZ4GXqCsjybzsO4sM4OB80Dg=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=DpR8VEQ0UA58HaY31wEnwLZOCzp290pfHU61d3OSuj5SE4bgeUIjQGmPtrArEmfmT Bn6I0PLu8PiCZkzeqe5O9527n28QKyFJAW3rRC7f4xktgzT3I9G9qxFQKL078NBFBc iNqlPguHzpCvgruRgwpGAnMeY+flKzeXe/eftNDB6rsi0q3DM1uchmcCTOg99x+EG6 rq3WEyw0zzW6rsUFR+rFNNFuB3H2QQsmo6DqJXDA4RTp7TcQUeAMBXjHMxuBZLVhAu OxJfMImjWtWAzXOuiiJs9ewO9CND9s0UCeeXDRV5tROKHGVvFWiBuA51K5eRaGaYQ6 aQoICo1yoGMAw== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 09CD0C61DF1; Mon, 31 Aug 2026 00:25:21 +0000 (UTC) From: Ackerley Tng via B4 Relay Date: Sun, 30 Aug 2026 17:25:17 -0700 Subject: [PATCH v12 16/45] KVM: guest_memfd: Call arch make_shared callback for to-shared conversion Precedence: bulk X-Mailing-List: linux-doc@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: 7bit Message-Id: <20260830-gmem-inplace-conversion-v12-16-85e5fd25252a@google.com> References: <20260830-gmem-inplace-conversion-v12-0-85e5fd25252a@google.com> In-Reply-To: <20260830-gmem-inplace-conversion-v12-0-85e5fd25252a@google.com> To: aik@amd.com, andrew.jones@linux.dev, binbin.wu@linux.intel.com, brauner@kernel.org, chao.p.peng@linux.intel.com, david@kernel.org, jmattson@google.com, jthoughton@google.com, michael.roth@amd.com, oupton@kernel.org, pankaj.gupta@amd.com, qperret@google.com, rick.p.edgecombe@intel.com, rientjes@google.com, shivankg@amd.com, steven.price@arm.com, willy@infradead.org, wyihan@google.com, yan.y.zhao@intel.com, forkloop@google.com, pratyush@kernel.org, suzuki.poulose@arm.com, aneesh.kumar@kernel.org, liam@infradead.org, Paolo Bonzini , Sean Christopherson , Thomas Gleixner , Ingo Molnar , Borislav Petkov , Dave Hansen , x86@kernel.org, "H. Peter Anvin" , Steven Rostedt , Masami Hiramatsu , Mathieu Desnoyers , Jonathan Corbet , Shuah Khan , Shuah Khan , Vishal Annapurve , Andrew Morton , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Youngjun Park , Qi Zheng , Shakeel Butt , Kiryl Shutsemau , Baoquan He , Jason Gunthorpe , John Hubbard , Peter Xu , tarunsahu@google.com, Randy Dunlap , Lorenzo Stoakes , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jason Gunthorpe , Fuad Tabba , Vlastimil Babka , Baoquan He Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-mm@kvack.org, linux-coco@lists.linux.dev, Ackerley Tng , Fuad Tabba X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=ed25519-sha256; t=1788135915; l=5997; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=cwEhZDpHTNx7ou7qscSIiavupcBz2iS+igqE5zUMd8Q=; b=t6DynLmrFZFZfNeyDVzLpduvArXHO0s1vSOrUsIWpNPwO/0vT3Y8DyHNEr/kDVnMFXQfuCFtv SR6kCWPCwtjCJlL15dvtAK08r074OdOuuf9mOk9x9z7bVFOMfIzN+/a X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Endpoint-Received: by B4 Relay for ackerleytng@google.com/20260225 with auth_id=649 X-Original-From: Ackerley Tng Reply-To: ackerleytng@google.com From: Ackerley Tng When doing in-place conversion from PRIVATE to SHARED, immediately inform arch code of the conversion for all allocated pages/folios, e.g. so that arch code can put hardware metadata tables in the correct state. Eagerly updating the table for to SHARED conversions avoids having to implement on-demand updates, e.g. when faulting in host userspace mappings. Skip the entire flow if the arch doesn't implement conversion callbacks, as getting folios from the filemap is noticeably expensive, especially when converting large chunks of memory. Deliberately don't eagerly update the metadata table on conversions from SHARED to PRIVATE, because assigning a page to a VM (versus "returning" it to the host) requires the exact GFN associated with the page, i.e would require walking the memslot bindings. And because KVM *must* do on-demand metadata updates when getting a PFN for KVM-internal usage, as that's the only time a relevant memslot binding is guaranteed to exist. Note! Inform arch code of the conversion within the protection of the invalidation sequence, to ensure that any existing mappings are dropped before hardware is updated, and to ensure that new mappings can't be established until after the conversion is complete. Reviewed-by: Fuad Tabba Signed-off-by: Ackerley Tng --- arch/x86/include/asm/kvm-x86-ops.h | 2 +- arch/x86/include/asm/kvm_host.h | 2 +- arch/x86/kvm/x86.c | 5 +++++ include/linux/kvm_host.h | 1 + virt/kvm/guest_memfd.c | 42 ++++++++++++++++++++++++++++++++++++++ 5 files changed, 50 insertions(+), 2 deletions(-) diff --git a/arch/x86/include/asm/kvm-x86-ops.h b/arch/x86/include/asm/kvm-x86-ops.h index e213c9ae3e301..67b43c167045b 100644 --- a/arch/x86/include/asm/kvm-x86-ops.h +++ b/arch/x86/include/asm/kvm-x86-ops.h @@ -150,7 +150,7 @@ KVM_X86_OP_OPTIONAL(alloc_apic_backing_page) #ifdef CONFIG_HAVE_KVM_ARCH_GMEM_CONVERT KVM_X86_OP_OPTIONAL_RET0(gmem_make_private) #endif -#ifdef CONFIG_HAVE_KVM_ARCH_GMEM_RECLAIM +#if defined(CONFIG_HAVE_KVM_ARCH_GMEM_CONVERT) || defined(CONFIG_HAVE_KVM_ARCH_GMEM_RECLAIM) KVM_X86_OP_OPTIONAL(gmem_make_shared) #endif #ifdef CONFIG_HAVE_KVM_ARCH_GMEM_INVALIDATE diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h index 744c1f6ff03ed..83e26ce45fb79 100644 --- a/arch/x86/include/asm/kvm_host.h +++ b/arch/x86/include/asm/kvm_host.h @@ -1732,7 +1732,7 @@ struct kvm_x86_ops { int (*gmem_make_private)(struct kvm *kvm, gfn_t gfn, kvm_pfn_t pfn, kvm_pfn_t nr_pages); #endif -#ifdef CONFIG_HAVE_KVM_ARCH_GMEM_RECLAIM +#if defined(CONFIG_HAVE_KVM_ARCH_GMEM_CONVERT) || defined(CONFIG_HAVE_KVM_ARCH_GMEM_RECLAIM) void (*gmem_make_shared)(kvm_pfn_t pfn, kvm_pfn_t nr_pages); #endif #ifdef CONFIG_HAVE_KVM_ARCH_GMEM_INVALIDATE diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c index 2292249570314..75e03a2f79db2 100644 --- a/arch/x86/kvm/x86.c +++ b/arch/x86/kvm/x86.c @@ -10653,6 +10653,11 @@ int kvm_arch_gmem_make_private(struct kvm *kvm, gfn_t gfn, kvm_pfn_t pfn, { return kvm_x86_call(gmem_make_private)(kvm, gfn, pfn, nr_pages); } + +void kvm_arch_gmem_make_shared(kvm_pfn_t pfn, kvm_pfn_t nr_pages) +{ + kvm_x86_call(gmem_make_shared)(pfn, nr_pages); +} #endif #ifdef CONFIG_HAVE_KVM_ARCH_GMEM_RECLAIM diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h index ab87effdd221f..485f18454eb45 100644 --- a/include/linux/kvm_host.h +++ b/include/linux/kvm_host.h @@ -2610,6 +2610,7 @@ static inline int kvm_gmem_get_pfn(struct kvm *kvm, int kvm_arch_gmem_make_private(struct kvm *kvm, gfn_t gfn, kvm_pfn_t pfn, kvm_pfn_t nr_pages); +void kvm_arch_gmem_make_shared(kvm_pfn_t pfn, kvm_pfn_t nr_pages); #ifndef CONFIG_HAVE_KVM_ARCH_GMEM_CONVERT #define kvm_arch_has_gmem_convert() false #endif diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c index fe02c47c85fb5..d14a7024bdc7b 100644 --- a/virt/kvm/guest_memfd.c +++ b/virt/kvm/guest_memfd.c @@ -575,6 +575,43 @@ static bool kvm_gmem_has_outstanding_references(struct inode *inode, return has_outstanding; } +#ifdef CONFIG_HAVE_KVM_ARCH_GMEM_CONVERT +static void kvm_gmem_make_shared(struct inode *inode, pgoff_t start, pgoff_t end) +{ + struct folio_batch fbatch; + pgoff_t next = start; + int i; + + folio_batch_init(&fbatch); + while (filemap_get_folios(inode->i_mapping, &next, end - 1, &fbatch)) { + for (i = 0; i < folio_batch_count(&fbatch); ++i) { + struct folio *folio = fbatch.folios[i]; + pgoff_t start_index, end_index; + kvm_pfn_t start_pfn; + kvm_pfn_t nr_pages; + + start_index = max(start, folio->index); + end_index = min(end, folio_next_index(folio)); + /* + * end_index is either in folio or points to + * the first page of the next folio. Hence, + * all pages in range [start_index, end_index) + * are contiguous. + */ + start_pfn = folio_file_pfn(folio, start_index); + nr_pages = end_index - start_index; + + kvm_arch_gmem_make_shared(start_pfn, nr_pages); + } + + folio_batch_release(&fbatch); + cond_resched(); + } +} +#else +static void kvm_gmem_make_shared(struct inode *inode, pgoff_t start, pgoff_t end) {} +#endif + static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start, size_t nr_pages, uint64_t attrs, pgoff_t *err_index) @@ -624,7 +661,12 @@ static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start, filter = to_private ? KVM_FILTER_SHARED : KVM_FILTER_PRIVATE; kvm_gmem_invalidate_start(inode, start, end, filter); + + if (!to_private && kvm_arch_has_gmem_convert()) + kvm_gmem_make_shared(inode, start, end); + mas_store_prealloc(&mas, xa_mk_value(attrs)); + kvm_gmem_invalidate_end(inode, start, end); out: filemap_invalidate_unlock(mapping); -- 2.55.0.897.gb25b4bd76c-goog