From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D95633BAD95; Thu, 10 Sep 2026 23:55:38 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789084539; cv=none; b=rZG6uS7eWMgHxUJ5RnpsY3L9s8oNnvKsEdM5xNsDZSojRpOEHvO7DPjQK0SzQ7XG7Qi+OpqfVqASX3l991PF/5A1w/l7QBJ7K8XNYQhaXa8wp7OjJLBccYpb7P/6l+Scm5v4/RH+79tVApFURM1tYPQblHMjezvgCHDfi5N7WF8= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789084539; c=relaxed/simple; bh=gGM6cYb630gdWoJEJn4fvEBtwAcYmnpAVbuU3m8hixg=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=tfB0Y3Pl7T9P8h09VI7whXTmu1c2jWu3dQwXvXE49AOg6t6BUtBwQSGNiWGZXP2HkYhSUrV7Uwxk941Q07+qeE4xXVgBVmTwJQCzJCcNUJa0FKZjhVDJv55jf1DYkZjJTQXqrV+UI/MxSVVBLKZyQ+9e6Izz8QddIuwenQupPNw= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=rCyKGBzL; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="rCyKGBzL" Received: by smtp.kernel.org (Postfix) with ESMTPS id 5F5A9C2BCC7; Thu, 10 Sep 2026 23:55:38 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1789084538; bh=gGM6cYb630gdWoJEJn4fvEBtwAcYmnpAVbuU3m8hixg=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=rCyKGBzLHZWHjByQEvp/3J+bhaghA38hro0nxowD1vP4RKw0B1LAa3zU7l9MJmkje FYMTzDXJQu99FujV0FHz4M56GC0TI5f73xy3qYXDSMiX2hm9wLISM6lwzMDuf2cHAE TbSNIGlHEvf3e/acZnayLgtk9UzjeYt7b3EwNQLdF7VIV/0S5J+AW3djIrmIU7t/YW w/3zJO4V4rlcYidR7ob3cFGPDwCRqGa8dKlMJR3D/Y0kmzd3vXcA0kKTgbBv9xRmuG DebQPv3jLe+zobRMgPuiHWwaSncHCEgeu44Uy3wkH3oP5Bxurtk4PMZcV8qLELcqwk xOcJ0BaBQQpsQ== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 46592C79FBB; Thu, 10 Sep 2026 23:55:38 +0000 (UTC) From: Ackerley Tng via B4 Relay Date: Thu, 10 Sep 2026 16:55:42 -0700 Subject: [PATCH v13 16/44] KVM: guest_memfd: Call arch make_shared callback for to-shared conversion Precedence: bulk X-Mailing-List: kvm@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: 7bit Message-Id: <20260910-gmem-inplace-conversion-v13-16-dd6fbf94f4e1@google.com> References: <20260910-gmem-inplace-conversion-v13-0-dd6fbf94f4e1@google.com> In-Reply-To: <20260910-gmem-inplace-conversion-v13-0-dd6fbf94f4e1@google.com> To: aik@amd.com, andrew.jones@linux.dev, binbin.wu@linux.intel.com, brauner@kernel.org, chao.p.peng@linux.intel.com, david@kernel.org, jmattson@google.com, jthoughton@google.com, michael.roth@amd.com, oupton@kernel.org, pankaj.gupta@amd.com, qperret@google.com, rick.p.edgecombe@intel.com, rientjes@google.com, shivankg@amd.com, steven.price@arm.com, willy@infradead.org, wyihan@google.com, yan.y.zhao@intel.com, forkloop@google.com, pratyush@kernel.org, suzuki.poulose@arm.com, aneesh.kumar@kernel.org, liam@infradead.org, Paolo Bonzini , Sean Christopherson , Thomas Gleixner , Ingo Molnar , Borislav Petkov , Dave Hansen , x86@kernel.org, "H. Peter Anvin" , Steven Rostedt , Masami Hiramatsu , Mathieu Desnoyers , Jonathan Corbet , Shuah Khan , Shuah Khan , Vishal Annapurve , Andrew Morton , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Youngjun Park , Qi Zheng , Shakeel Butt , Kiryl Shutsemau , Baoquan He , Jason Gunthorpe , John Hubbard , Peter Xu , tarunsahu@google.com, Randy Dunlap , Lorenzo Stoakes , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jason Gunthorpe , Fuad Tabba , Vlastimil Babka , Baoquan He Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-mm@kvack.org, linux-coco@lists.linux.dev, Ackerley Tng , Fuad Tabba X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=ed25519-sha256; t=1789084533; l=5988; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=pnjcuYiGgOrTACJiDuCpEUUXJmfVPZdc0nxxctxlq1I=; b=8n/zqT12bJsDEGglxVJpV0qF7heq6Y/6E8z8+2TPGozDYy2KIXFbkJruoR9CqjTm5xYcISqXB uNl2w215XnvChWobjBOo60K/Wh3YvU13WDJU5+Z2PMtVpViNfPkT54t X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Endpoint-Received: by B4 Relay for ackerleytng@google.com/20260225 with auth_id=649 X-Original-From: Ackerley Tng Reply-To: ackerleytng@google.com From: Ackerley Tng When doing in-place conversion from PRIVATE to SHARED, immediately inform arch code of the conversion for all allocated pages/folios, e.g. so that arch code can put hardware metadata tables in the correct state. Eagerly updating the table for to SHARED conversions avoids having to implement on-demand updates, e.g. when faulting in host userspace mappings. Skip the entire flow if the arch doesn't implement conversion callbacks, as getting folios from the filemap is noticeably expensive, especially when converting large chunks of memory. Deliberately don't eagerly update the metadata table on conversions from SHARED to PRIVATE, because assigning a page to a VM (versus "returning" it to the host) requires the exact GFN associated with the page, i.e would require walking the memslot bindings. And because KVM *must* do on-demand metadata updates when getting a PFN for KVM-internal usage, as that's the only time a relevant memslot binding is guaranteed to exist. Note! Inform arch code of the conversion within the protection of the invalidation sequence, to ensure that any existing mappings are dropped before hardware is updated, and to ensure that new mappings can't be established until after the conversion is complete. Signed-off-by: Ackerley Tng Reviewed-by: Fuad Tabba --- arch/x86/include/asm/kvm-x86-ops.h | 2 +- arch/x86/include/asm/kvm_host.h | 2 +- arch/x86/kvm/x86.c | 5 +++++ include/linux/kvm_host.h | 1 + virt/kvm/guest_memfd.c | 42 ++++++++++++++++++++++++++++++++++++++ 5 files changed, 50 insertions(+), 2 deletions(-) diff --git a/arch/x86/include/asm/kvm-x86-ops.h b/arch/x86/include/asm/kvm-x86-ops.h index e213c9ae3e301..67b43c167045b 100644 --- a/arch/x86/include/asm/kvm-x86-ops.h +++ b/arch/x86/include/asm/kvm-x86-ops.h @@ -150,7 +150,7 @@ KVM_X86_OP_OPTIONAL(alloc_apic_backing_page) #ifdef CONFIG_HAVE_KVM_ARCH_GMEM_CONVERT KVM_X86_OP_OPTIONAL_RET0(gmem_make_private) #endif -#ifdef CONFIG_HAVE_KVM_ARCH_GMEM_RECLAIM +#if defined(CONFIG_HAVE_KVM_ARCH_GMEM_CONVERT) || defined(CONFIG_HAVE_KVM_ARCH_GMEM_RECLAIM) KVM_X86_OP_OPTIONAL(gmem_make_shared) #endif #ifdef CONFIG_HAVE_KVM_ARCH_GMEM_INVALIDATE diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h index 744c1f6ff03ed..83e26ce45fb79 100644 --- a/arch/x86/include/asm/kvm_host.h +++ b/arch/x86/include/asm/kvm_host.h @@ -1732,7 +1732,7 @@ struct kvm_x86_ops { int (*gmem_make_private)(struct kvm *kvm, gfn_t gfn, kvm_pfn_t pfn, kvm_pfn_t nr_pages); #endif -#ifdef CONFIG_HAVE_KVM_ARCH_GMEM_RECLAIM +#if defined(CONFIG_HAVE_KVM_ARCH_GMEM_CONVERT) || defined(CONFIG_HAVE_KVM_ARCH_GMEM_RECLAIM) void (*gmem_make_shared)(kvm_pfn_t pfn, kvm_pfn_t nr_pages); #endif #ifdef CONFIG_HAVE_KVM_ARCH_GMEM_INVALIDATE diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c index 2292249570314..75e03a2f79db2 100644 --- a/arch/x86/kvm/x86.c +++ b/arch/x86/kvm/x86.c @@ -10653,6 +10653,11 @@ int kvm_arch_gmem_make_private(struct kvm *kvm, gfn_t gfn, kvm_pfn_t pfn, { return kvm_x86_call(gmem_make_private)(kvm, gfn, pfn, nr_pages); } + +void kvm_arch_gmem_make_shared(kvm_pfn_t pfn, kvm_pfn_t nr_pages) +{ + kvm_x86_call(gmem_make_shared)(pfn, nr_pages); +} #endif #ifdef CONFIG_HAVE_KVM_ARCH_GMEM_RECLAIM diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h index ab87effdd221f..485f18454eb45 100644 --- a/include/linux/kvm_host.h +++ b/include/linux/kvm_host.h @@ -2610,6 +2610,7 @@ static inline int kvm_gmem_get_pfn(struct kvm *kvm, int kvm_arch_gmem_make_private(struct kvm *kvm, gfn_t gfn, kvm_pfn_t pfn, kvm_pfn_t nr_pages); +void kvm_arch_gmem_make_shared(kvm_pfn_t pfn, kvm_pfn_t nr_pages); #ifndef CONFIG_HAVE_KVM_ARCH_GMEM_CONVERT #define kvm_arch_has_gmem_convert() false #endif diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c index 2de39ce8aec13..85446fe720603 100644 --- a/virt/kvm/guest_memfd.c +++ b/virt/kvm/guest_memfd.c @@ -588,6 +588,43 @@ static bool kvm_gmem_has_outstanding_references(struct inode *inode, return false; } +#ifdef CONFIG_HAVE_KVM_ARCH_GMEM_CONVERT +static void kvm_gmem_make_shared(struct inode *inode, pgoff_t start, pgoff_t end) +{ + struct folio_batch fbatch; + pgoff_t next = start; + int i; + + folio_batch_init(&fbatch); + while (filemap_get_folios(inode->i_mapping, &next, end - 1, &fbatch)) { + for (i = 0; i < folio_batch_count(&fbatch); ++i) { + struct folio *folio = fbatch.folios[i]; + pgoff_t start_index, end_index; + kvm_pfn_t start_pfn; + kvm_pfn_t nr_pages; + + start_index = max(start, folio->index); + end_index = min(end, folio_next_index(folio)); + /* + * end_index is either in folio or points to + * the first page of the next folio. Hence, + * all pages in range [start_index, end_index) + * are contiguous. + */ + start_pfn = folio_file_pfn(folio, start_index); + nr_pages = end_index - start_index; + + kvm_arch_gmem_make_shared(start_pfn, nr_pages); + } + + folio_batch_release(&fbatch); + cond_resched(); + } +} +#else +static void kvm_gmem_make_shared(struct inode *inode, pgoff_t start, pgoff_t end) {} +#endif + static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start, size_t nr_pages, uint64_t attrs, pgoff_t *err_index) @@ -637,7 +674,12 @@ static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start, filter = to_private ? KVM_FILTER_SHARED : KVM_FILTER_PRIVATE; kvm_gmem_invalidate_start(inode, start, end, filter); + + if (!to_private && kvm_arch_has_gmem_convert()) + kvm_gmem_make_shared(inode, start, end); + mas_store_prealloc(&mas, xa_mk_value(attrs)); + kvm_gmem_invalidate_end(inode, start, end); out: filemap_invalidate_unlock(mapping); -- 2.55.0.1007.g17ff1f9808-goog