From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EA1372F361E; Mon, 31 Aug 2026 00:25:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788135921; cv=none; b=sAKUYuVRaKycotpEIEzICzgrXyBje36wz+4HR9gqelfZ+xqCMagCic2UaVBP1tgjIk94+iPNcbfSqyrlQAtR7yba/9+ZcUFdqdEp7hprdus5SEdVKvLXXwXrcfAbj7lgpT8IGZPxcY7hOdWyHieAKEBNNCUnqf+uAgzhCrxfDDk= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788135921; c=relaxed/simple; bh=pgbQeqwV6IRb5CiPCfNQynStgiWTbolys+BVEEK4H2w=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=WEyl4bKrUR/zLO3dQV2z1PobC0zi48IxP2nlxgX5+EjdcorXrpzFMoUVM7DFz7YKYSjSXNsdgfojr1ooZLFuU38h5X1WuOGyJNwVIJnBgkCOJXu1G9YEXsSvQuUnFOqVbttQtLILvFNW8g9W3tKEV0Y6cnFli7OO/HQ9Pz6ya/k= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=mZSTa6ni; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="mZSTa6ni" Received: by smtp.kernel.org (Postfix) with ESMTPS id B8328C4AF54; Mon, 31 Aug 2026 00:25:20 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1788135920; bh=pgbQeqwV6IRb5CiPCfNQynStgiWTbolys+BVEEK4H2w=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=mZSTa6ni/SkaGctJC95sGPqw7Ik2tnPRymeXDDfMOLYPN72DoTj7xjt+asjczUfcc bppappQ9c9LUoY9EIPrTqJttmF6zojfgB2nzktvtf8rApIZv0SvZfihCbDJ7S2JB+U d2/XrY+rsOxIGTJBaYDtZG5LP6cIBB2pJ2qRa9UKxdruov3qKzP8mQef1N6WDnGvJq iquSBeuh+UxlIfCqyP9i9gNWfg2YJ2L3PtnNIH5ztsr+EKIfJsKR1SAlgX64MCzz1y Lj7ej/FBYV/ydqd5HKGBSe1XV3FH0KVEHGF0Fta3/DEHiovgEKvHNC77PH9HmfRubV e6ronKXVy4p8g== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 971A0C61DFF; Mon, 31 Aug 2026 00:25:20 +0000 (UTC) From: Ackerley Tng via B4 Relay Date: Sun, 30 Aug 2026 17:25:13 -0700 Subject: [PATCH v12 12/45] KVM: guest_memfd: Always fault from guest_memfd if in-place conversion is enabled Precedence: bulk X-Mailing-List: linux-trace-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: 7bit Message-Id: <20260830-gmem-inplace-conversion-v12-12-85e5fd25252a@google.com> References: <20260830-gmem-inplace-conversion-v12-0-85e5fd25252a@google.com> In-Reply-To: <20260830-gmem-inplace-conversion-v12-0-85e5fd25252a@google.com> To: aik@amd.com, andrew.jones@linux.dev, binbin.wu@linux.intel.com, brauner@kernel.org, chao.p.peng@linux.intel.com, david@kernel.org, jmattson@google.com, jthoughton@google.com, michael.roth@amd.com, oupton@kernel.org, pankaj.gupta@amd.com, qperret@google.com, rick.p.edgecombe@intel.com, rientjes@google.com, shivankg@amd.com, steven.price@arm.com, willy@infradead.org, wyihan@google.com, yan.y.zhao@intel.com, forkloop@google.com, pratyush@kernel.org, suzuki.poulose@arm.com, aneesh.kumar@kernel.org, liam@infradead.org, Paolo Bonzini , Sean Christopherson , Thomas Gleixner , Ingo Molnar , Borislav Petkov , Dave Hansen , x86@kernel.org, "H. Peter Anvin" , Steven Rostedt , Masami Hiramatsu , Mathieu Desnoyers , Jonathan Corbet , Shuah Khan , Shuah Khan , Vishal Annapurve , Andrew Morton , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Youngjun Park , Qi Zheng , Shakeel Butt , Kiryl Shutsemau , Baoquan He , Jason Gunthorpe , John Hubbard , Peter Xu , tarunsahu@google.com, Randy Dunlap , Lorenzo Stoakes , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jason Gunthorpe , Fuad Tabba , Vlastimil Babka , Baoquan He Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-mm@kvack.org, linux-coco@lists.linux.dev, Ackerley Tng X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=ed25519-sha256; t=1788135915; l=4450; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=X0oHuYGMuEMc+mIS2D25NifHLJhyfzoCyNVUkFKAoCE=; b=EtQxw+QEVWJoZlS70xZzAwu5ZPYmJaQfBkoC/h5utIHwFISWqanhlhfA/RfJ33NDQYkw7Y4kT x0KHakJQKe3CW/InZ0fVH/Fv/TtLlWGt56gk9mRZgklUQUp8zSa/3yP X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Endpoint-Received: by B4 Relay for ackerleytng@google.com/20260225 with auth_id=649 X-Original-From: Ackerley Tng Reply-To: ackerleytng@google.com From: Ackerley Tng If a guest_memfd memslot is created but the guest_memfd does not have the GUEST_MEMFD_FLAG_MMAP set, KVM still fulfils guest faults by looking up the memslot's userspace_addr. Set KVM_MEMSLOT_GMEM_ONLY if in-place conversion is enabled so that the guest_memfd's memory will be used for both shared and private memory. With in-place conversion, guest_memfd will be the only backing memory for the memslot. No validation is performed to require userspace_addr to be a mapping from the associated guest_memfd because even after validation, userspace is free to remap something else at the provided userspace_addr. userspace_addr will still be used by functions like kvm_read_guest(), and if userspace_addr does not match up with the corresponding memory in the memslot's guest_memfd (whether userspace_addr points to the wrong offset or some non-guest_memfd memory, etc), that is a user error. Requiring both shared and private memory to come from the only associated guest_memfd simplifies invalidation in stage 2 page tables. On a PUNCH_HOLE operation on a guest_memfd, the invalidation is now guaranteed to be invalidating only memory mapped from the given guest_memfd. Suggested-by: Sean Christopherson Signed-off-by: Ackerley Tng --- Documentation/virt/kvm/api.rst | 22 ++++++++++++++-------- virt/kvm/guest_memfd.c | 2 +- 2 files changed, 15 insertions(+), 9 deletions(-) diff --git a/Documentation/virt/kvm/api.rst b/Documentation/virt/kvm/api.rst index 90a29424c54c8..668886f50024d 100644 --- a/Documentation/virt/kvm/api.rst +++ b/Documentation/virt/kvm/api.rst @@ -6381,10 +6381,16 @@ mapping for userspace_addr is not required to be valid/populated at the time of KVM_SET_USER_MEMORY_REGION2, e.g. shared memory can be lazily mapped/allocated on-demand. -When mapping a gfn into the guest, KVM selects shared vs. private, i.e consumes -userspace_addr vs. guest_memfd, based on the state in guest_memfd, which is the -sole authority on private vs. shared memory. See :ref:`KVM_CREATE_GUEST_MEMFD` -to find out more about the creation-time shared/private status. +When mapping a gfn into the guest, guest faults are always serviced from +guest_memfd regardless of whether memory is shared or private. KVM determines +shared vs. private based on the state in guest_memfd, which is the sole +authority on private vs. shared memory. See :ref:`KVM_CREATE_GUEST_MEMFD` to +find out more about the creation-time shared/private status. + +userspace_addr is expected to be the mmap()-ed address corresponding to the +right offset within the guest_memfd. Any mismatch between userspace_addr and +guest_memfd is not validated and is a user error. userspace_addr is only used +for host-side guest accesses such as kvm_read_guest(). If in-place conversion is disabled, KVM selects shared vs. private, i.e consumes userspace_addr vs. guest_memfd, based on the gfn's KVM_MEMORY_ATTRIBUTE_PRIVATE @@ -6490,10 +6496,10 @@ specified via KVM_CREATE_GUEST_MEMFD. Currently defined flags: page tables. Private memory cannot. ============================ ================================================ -When the KVM MMU performs a PFN lookup to service a guest fault and the backing -guest_memfd has the GUEST_MEMFD_FLAG_MMAP set, then the fault will always be -consumed from guest_memfd, regardless of whether it is a shared or a private -fault. +When the KVM MMU performs a PFN lookup to service a guest fault, the fault will +always be consumed from guest_memfd, regardless of whether it is a shared or a +private fault (unless in-place conversion is disabled and the backing +guest_memfd does not have the GUEST_MEMFD_FLAG_MMAP flag set). See KVM_SET_USER_MEMORY_REGION2 for additional details. diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c index 0afe1468d2d9d..e41802944756b 100644 --- a/virt/kvm/guest_memfd.c +++ b/virt/kvm/guest_memfd.c @@ -746,7 +746,7 @@ int kvm_gmem_bind(struct kvm *kvm, struct kvm_memory_slot *slot, */ WRITE_ONCE(slot->gmem.file, file); slot->gmem.pgoff = start; - if (kvm_gmem_supports_mmap(inode)) + if (gmem_in_place_conversion || kvm_gmem_supports_mmap(inode)) slot->flags |= KVM_MEMSLOT_GMEM_ONLY; xa_store_range(&f->bindings, start, end - 1, slot, GFP_KERNEL); -- 2.55.0.897.gb25b4bd76c-goog