From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pl1-f200.google.com (mail-pl1-f200.google.com [209.85.214.200]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 96A6246F4BA for ; Fri, 28 Aug 2026 14:57:03 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.200 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787929025; cv=none; b=RRa72hhKZgEJEnafF9H3WPp+nd9Lqkuxe5/JmjIzg1oX8u/KuhdkWdRt/dYo/F2soClbT7+1GxU4+yhKo8gQaGs5GPD1QEWdPGIZFpNjX3y81SIvocTgPW4lNpkh/l1RkG41uru3OxnNS6qiqa6Vr1Er+TwXXAhW3VHZe6Ac5zM= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787929025; c=relaxed/simple; bh=LZpPvS21SRGUjRRQo9YMu2uphCB+Yb/O1WCRXm1EFjc=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=b+4Q34gfA+TG5LU2oh0NALWQPhA+iA+Oa4nqfQ0LIuyRcG/D44QcxpV2EINAbDXuezqwT3lFuqwqqhlvDF0BHgwWOVDqVD8fZI4qn82z8XqloV5kiqw6ZnnCpmRuMC5gcLqeTrOlSgWdAse/HdRwi+p0nhVuJyYyaQGxZl7m2e8= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=HKi8VdUO; arc=none smtp.client-ip=209.85.214.200 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="HKi8VdUO" Received: by mail-pl1-f200.google.com with SMTP id d9443c01a7336-2d7293f13c6so14643425ad.2 for ; Fri, 28 Aug 2026 07:57:03 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1787929023; x=1788533823; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=rKyTv2j02zbnlF+PSfGKNUB7r5AsnRVQJaAhKBT0UgU=; b=HKi8VdUOMQ+VhGwRtStYGF2qLx+TKzH7citcBlKEg4hjaizk6X5JWUqFLqtF7AizQ3 ayxD6h+lwClPnw6h1ap2n6vfGqEVYLJ2LjsafVfVzZyMekM1kM8ZYNK2Z1azvsXDDPVs y1hqvhtzpK/60tJN1XmbkQM0BDmxgUaznhlQ5bLTWAxRWWIJeDWN+IEHSfH0qvUFR1FL fj7bz6dko7atEYi1As+Sb18pSNahp6AmydieQp7TV82IiGh5UIQC4u4sPrudKgfvUAlO K5AZ2FO9HITDBDFoVv0NDYV5z/n37h+Mgs7BgeuVTMvGTJJA+Ro4thpoa0pAf9tPq9b3 CXtA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787929023; x=1788533823; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=rKyTv2j02zbnlF+PSfGKNUB7r5AsnRVQJaAhKBT0UgU=; b=SAT7qspCG/m70x/U3zXakV6iTrdfkjfHUIXhKKLRB49B1x8NuvdhsWGp7UVy58xbGi qX1kedxxv8cIqVU6nuu8BMb3EQ2x1UWetkpzGjlFuNN+NaUluuIZQf2ikzs3oqLekuQk lE1p1xLEQOURKiwnAF7Tb9sSnBvBRnzsHf+baboXp4r/ZYDGZFLAohdUpRAui0OYeyjT /VDMbxFb0LnHFfR4gfNcowZAu2VAFjZmwKa9l5YLFlRA+J8kVDcgU6oGr0WbQ+3w2vt7 ifED1E1oZ66ZpqzyY9GUjrmKJi08wqqW/DMchrfH91jFFbU5UhEyOmVoKrNijoNNk8qj SyQQ== X-Forwarded-Encrypted: i=1; AHgh+Ro5KpQLRpiLoSpRJN56+5sgbcT06skddf9y9MJjJWSJDV5GDfkacAU11crEh7y4crdUoDSI7iZrC8M=@vger.kernel.org X-Gm-Message-State: AFuF++lWPXpXz+rupF6HaQfKAKxoHpv7WgraBil+vL4GSHtu86Yo49RM pO4G1lprpg2Px8jD/XKKBYO1At6ESYBmvngg8WKC5GANfqZilMdrpf2JwwALOW2Z396tWDKqsMB +0/h6Kw== X-Received: from plbkz13.prod.google.com ([2002:a17:902:f9cd:b0:2ca:ed29:ea82]) (user=seanjc job=prod-delivery.src-stubby-dispatcher) by 2002:a17:903:1984:b0:2d8:d4d3:da4f with SMTP id d9443c01a7336-2d8d4d3db7dmr6783485ad.19.1787929022578; Fri, 28 Aug 2026 07:57:02 -0700 (PDT) Date: Fri, 28 Aug 2026 07:57:01 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-doc@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260826-gmem-inplace-conversion-v11-0-0a15d8a799aa@google.com> <20260826-gmem-inplace-conversion-v11-15-0a15d8a799aa@google.com> Message-ID: Subject: Re: [PATCH v11 15/46] KVM: guest_memfd: Call arch make_shared callback for to-shared conversion From: Sean Christopherson To: Ackerley Tng Cc: aik@amd.com, andrew.jones@linux.dev, binbin.wu@linux.intel.com, brauner@kernel.org, chao.p.peng@linux.intel.com, david@kernel.org, jmattson@google.com, jthoughton@google.com, michael.roth@amd.com, oupton@kernel.org, pankaj.gupta@amd.com, qperret@google.com, rick.p.edgecombe@intel.com, rientjes@google.com, shivankg@amd.com, steven.price@arm.com, willy@infradead.org, wyihan@google.com, yan.y.zhao@intel.com, forkloop@google.com, pratyush@kernel.org, suzuki.poulose@arm.com, aneesh.kumar@kernel.org, liam@infradead.org, Paolo Bonzini , Thomas Gleixner , Ingo Molnar , Borislav Petkov , Dave Hansen , x86@kernel.org, "H. Peter Anvin" , Steven Rostedt , Masami Hiramatsu , Mathieu Desnoyers , Jonathan Corbet , Shuah Khan , Shuah Khan , Vishal Annapurve , Andrew Morton , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Youngjun Park , Qi Zheng , Shakeel Butt , Kiryl Shutsemau , Baoquan He , Jason Gunthorpe , John Hubbard , Peter Xu , tarunsahu@google.com, Fuad Tabba , Vlastimil Babka , kvm@vger.kernel.org, linux-kernel@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-mm@kvack.org, linux-coco@lists.linux.dev Content-Type: text/plain; charset="us-ascii" On Thu, Aug 27, 2026, Ackerley Tng wrote: > Sean Christopherson writes: > > > On Wed, Aug 26, 2026, Ackerley Tng wrote: > >> Omit support for calling the arch hook to make private, since SNP, the only > >> implementer of the arch make-private hook today, would actually prefer > >> making private only just before faulting memory into the NPTs. > >> > >> Calling the make-private arch hook would require iterating both bindings > >> and the filemap to find the intersection of bindings and allocated > >> folios. > > > > Why would KVM need to iterate over the bindings? Only the RMP needs to be updated, > > whether or not the RMP is currently reachable is irrelevant, no? > > > > Subsequent calls to kvm_arch_gmem_make_private() from kvm_gmem_get_pfn() would be > > superfluous, but that's already possible, e.g. if an NPT mappings is removed for > > whatever reason. > > > > make_private takes a gfn and pfn, so __kvm_gmem_set_attributes() would > need to iterate bindings to get gfns and filemap to find folios to get > pfns. Oooh, right, unassigned a page in the RMP only needs the PFN, but assigned a page needs the ASID and GFN. > > I don't care terribly about how SNP handles this, but I do want accurate reasoning > > and justification so that if/when we revisit any of this in the future, we can make > > informed decisions. Because unless I'm missing something, this is an optimization > > choice (eager vs. lazy to-private conversions), not a complexity tradeoff, and it's > > not clear to me how we decided the lazy approach would provide better performance. > > > > How's this, to replace the entire commit message? I hope it captures > points from this discussion: > > When memory in guest_memfd is converted from private to shared, the > platform-specific state associated with the guest-private pages must > be invalidated or cleaned up. > > Iterate over the folios in the affected range and call the > kvm_arch_gmem_make_shared() hook for each PFN range. This allows > architectures to update hardware metadata or encryption states to > transition pages to the shared state, instead of leaving hardware > state as private while guest_memfd tracks it as shared. Transitioning > hardware state ensures that guest_memfd upholds the guarantee that > userspace only maps shared memory. Nit, don't reference functions by name when it's easy-ish to avoid doing so. And don't give a play-by-play: the patch makes it pretty obvious the code is iterating over folios, what isn't obvious is *why* the code does that. > Invoke this helper after indicating to KVM's mmu code that an > invalidation is in progress to stop in-flight page faults from > succeeding. Calling the invalidation helper also calls the arch > invalidate hook. For SNP, this kicks any vCPU with a registered VMSA > within the range being converted out of the guest. This ensures that > make_shared never fails due to the VMSA page being in-use and is > important because if make_shared fails, the RMP table would track the > page as private while guest_memfd is unaware and tracks the page as > shared. Exaclty what SNP does isn't relevant. Or rather, it's but on example of how this needs to work. I.e. the conversion needs to happen within the invalidtion sequence because thems the rules for KVM. > Omit support for calling the arch hook to make private during > to-private conversions. Making private lazily at fault time aligns > with how it works on other platforms like TDX. To me, this isn't a valid argument. We've fully committed to relying on vendor specific behavior, and SNP can't truly work like TDX because the underlying implementations are so different. > Furthermore, making pages shared only requires PFNs, which are > obtained by iterating folios in the filemap. In contrast, making pages > private in the RMP also requires the GFN, which would require > iterating bindings to get GFNs and the filemap to get PFNs from > allocated folios. Deferring the transition to fault time avoids this > additional complexity. It's not just complexity, it's that the bindings might not even exist. I.e. for all intents and purposes, doing on-demand updates is mandatory, because that's the only time a relevant memslot binding is guaranteed to exist. All in all, this? When doing in-place conversion from PRIVATE to SHARED, immediately inform arch code of the conversion for all allocated pages/folios, e.g. so that arch code can put hardware metadata tables in the correct state. Eagerly updating the table for to SHARED conversions avoids having to implement on-demand updates, e.g. when faulting in host userspace mappings. Skip the entire flow if the arch doesn't implement conversion callbacks, as getting folios from the filemap is noticeably expensive, especially when converting large chunks of memory. Deliberately don't eagerly update the metadata table on conversions from SHARED to PRIVATE, because assigning a page to a VM (versus "returning" it to the host) requires the exact GFN associated with the page, i.e would require walking the memslot bindings. And because KVM *must* do on-demand metadata updates when getting a PFN for KVM-internal usage, as that's the only time a relevant memslot binding is guaranteed to exist. Note! Inform arch code of the conversion within the protection of the invalidation sequence, to ensure that any existing mappings are dropped before hardware is updated, and to ensure that new mappings can't be established until after the conversion is complete.