Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Ackerley Tng <ackerleytng@google.com>
To: aik@amd.com, andrew.jones@linux.dev, binbin.wu@linux.intel.com,
	 brauner@kernel.org, chao.p.peng@linux.intel.com,
	david@kernel.org,  jmattson@google.com, jthoughton@google.com,
	michael.roth@amd.com,  oupton@kernel.org, pankaj.gupta@amd.com,
	qperret@google.com,  rick.p.edgecombe@intel.com,
	rientjes@google.com, shivankg@amd.com,  steven.price@arm.com,
	willy@infradead.org, wyihan@google.com,  yan.y.zhao@intel.com,
	forkloop@google.com, pratyush@kernel.org,
	 suzuki.poulose@arm.com, aneesh.kumar@kernel.org,
	liam@infradead.org,  Paolo Bonzini <pbonzini@redhat.com>,
	Sean Christopherson <seanjc@google.com>,
	 Thomas Gleixner <tglx@kernel.org>,
	Ingo Molnar <mingo@redhat.com>, Borislav Petkov <bp@alien8.de>,
	 Dave Hansen <dave.hansen@linux.intel.com>,
	x86@kernel.org,  "H. Peter Anvin" <hpa@zytor.com>,
	Steven Rostedt <rostedt@goodmis.org>,
	 Masami Hiramatsu <mhiramat@kernel.org>,
	Mathieu Desnoyers <mathieu.desnoyers@efficios.com>,
	 Jonathan Corbet <corbet@lwn.net>,
	Shuah Khan <skhan@linuxfoundation.org>,
	 Shuah Khan <shuah@kernel.org>,
	Vishal Annapurve <vannapurve@google.com>,
	 Andrew Morton <akpm@linux-foundation.org>,
	Chris Li <chrisl@kernel.org>,  Kairui Song <kasong@tencent.com>,
	Kemeng Shi <shikemeng@huaweicloud.com>,
	 Nhat Pham <nphamcs@gmail.com>, Barry Song <baohua@kernel.org>,
	 Axel Rasmussen <axelrasmussen@google.com>,
	Yuanchu Xie <yuanchu@google.com>,  Wei Xu <weixugc@google.com>,
	Youngjun Park <youngjun.park@lge.com>,
	 Qi Zheng <qi.zheng@linux.dev>,
	Shakeel Butt <shakeel.butt@linux.dev>,
	 Kiryl Shutsemau <kas@kernel.org>,
	Baoquan He <baoquan.he@linux.dev>, Jason Gunthorpe <jgg@ziepe.ca>,
	 John Hubbard <jhubbard@nvidia.com>, Peter Xu <peterx@redhat.com>,
	tarunsahu@google.com,  Fuad Tabba <fuad.tabba@linux.dev>,
	Vlastimil Babka <vbabka@kernel.org>
Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org,
	 linux-trace-kernel@vger.kernel.org, linux-doc@vger.kernel.org,
	 linux-kselftest@vger.kernel.org, linux-mm@kvack.org,
	 linux-coco@lists.linux.dev,
	Ackerley Tng <ackerleytng@google.com>,
	 Xiaoyao Li <xiaoyao.li@intel.com>
Subject: [PATCH v11 23/46] KVM: Let userspace disable per-VM mem attributes, enable per-gmem attributes
Date: Wed, 26 Aug 2026 09:18:21 +0000	[thread overview]
Message-ID: <20260826-gmem-inplace-conversion-v11-23-0a15d8a799aa@google.com> (raw)
In-Reply-To: <20260826-gmem-inplace-conversion-v11-0-0a15d8a799aa@google.com>

From: Sean Christopherson <seanjc@google.com>

Allow the user to disable KVM_VM_MEMORY_ATTRIBUTES even when KVM supports
PRIVATE and SHARED attributes, and expose gmem_in_place_conversion as a
module parameter when per-VM attributes are supported.  I.e. let userspace
enable in-place PRIVATE<=>SHARED conversion of guest_memfd pages.

Provide both a Kconfig option and a (conditional) module param so that
deployments that use a custom kernel can fully disable per-VM tracking,
while not forcing distros to ship two separate kernels in order to provide
backwards compatibility for downstream users.

Don't allow running VMs with mixed tracking for a given instance of KVM,
i.e. disallow toggling the module param after KVM is loaded, as the extra
complexity needed to handle per-VM behavior far outweighs any potential
benefit.  E.g. neither TDX nor SNP supports live migration, so in effect
the requirement is that existing deployments that want to support both the
old and the new models would need to tell their VMM which flavor of
tracking to use.

Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Fuad Tabba <tabba@google.com>
Tested-by: Shivank Garg <shivankg@amd.com>
[Define module_param only if CONFIG_KVM_VM_MEMORY_ATTRIBUTES is enabled]
Suggested-by: Xiaoyao Li <xiaoyao.li@intel.com>
Reviewed-by: Xiaoyao Li <xiaoyao.li@intel.com>
Co-developed-by: Ackerley Tng <ackerleytng@google.com>
Signed-off-by: Ackerley Tng <ackerleytng@google.com>
---
 Documentation/admin-guide/kernel-parameters.txt | 25 +++++++++++++++++++++++++
 Documentation/virt/kvm/api.rst                  |  7 +++++--
 arch/x86/include/asm/kvm_host.h                 |  4 +++-
 arch/x86/kvm/Kconfig                            | 14 ++++++++++----
 virt/kvm/kvm_main.c                             |  5 ++++-
 5 files changed, 47 insertions(+), 8 deletions(-)

diff --git a/Documentation/admin-guide/kernel-parameters.txt b/Documentation/admin-guide/kernel-parameters.txt
index 1069806616b94..0ac44e1bccd28 100644
--- a/Documentation/admin-guide/kernel-parameters.txt
+++ b/Documentation/admin-guide/kernel-parameters.txt
@@ -3114,6 +3114,31 @@ Kernel parameters
 	kvm.enable_vmware_backdoor=[KVM] Support VMware backdoor PV interface.
 				   Default is false (don't support).
 
+	kvm.gmem_in_place_conversion=
+			[KVM] Controls whether KVM enables in-place conversion
+			support for guest_memfd and tracks the private/shared
+			state of memory per guest_memfd instead of per VM.
+
+			If enabled, KVM enables the KVM_SET_MEMORY_ATTRIBUTES2
+			ioctl on guest_memfd file descriptors and disables the
+			legacy VM-scoped KVM_SET_MEMORY_ATTRIBUTES ioctl for
+			private memory state tracking. Only the
+			KVM_MEMORY_ATTRIBUTE_PRIVATE attribute moves to
+			per-guest_memfd tracking; other attributes remain
+			per-VM.
+
+			This parameter toggles KVM's in-place conversion
+			capability support. Whether a VMM uses separate backends
+			or out-of-place memory management is determined by
+			userspace VMM design.
+
+			Note, this parameter is only available when
+			CONFIG_KVM_VM_MEMORY_ATTRIBUTES=y. When
+			CONFIG_KVM_VM_MEMORY_ATTRIBUTES is not set, in-place
+			conversion is unconditionally enabled.
+
+			Default is Y (on).
+
 	kvm.nx_huge_pages=
 			[KVM] Controls the software workaround for the
 			X86_BUG_ITLB_MULTIHIT bug.
diff --git a/Documentation/virt/kvm/api.rst b/Documentation/virt/kvm/api.rst
index f88d65b78c504..d976f8ec2e2dc 100644
--- a/Documentation/virt/kvm/api.rst
+++ b/Documentation/virt/kvm/api.rst
@@ -6383,9 +6383,12 @@ on-demand.
 
 When mapping a gfn into the guest, KVM selects shared vs. private, i.e consumes
 userspace_addr vs. guest_memfd, based on the gfn's KVM_MEMORY_ATTRIBUTE_PRIVATE
-state.  At VM creation time, all memory is shared, i.e. the PRIVATE attribute
-is '0' for all gfns.  Userspace can control whether memory is shared/private by
+state.  If in-place conversion is disabled, i.e. PRIVATE is tracked per-VM,
+then at VM creation time, all memory is shared, i.e. the PRIVATE attribute is
+'0' for all gfns.  Userspace can control whether memory is shared/private by
 toggling KVM_MEMORY_ATTRIBUTE_PRIVATE via KVM_SET_MEMORY_ATTRIBUTES as needed.
+If in-place conversion is enabled, then the starting PRIVATE vs. SHARED state
+of a gfn is determined by the relevant guest_memfd instance.
 
 S390:
 ^^^^^
diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h
index 83e26ce45fb79..e840418427a1d 100644
--- a/arch/x86/include/asm/kvm_host.h
+++ b/arch/x86/include/asm/kvm_host.h
@@ -1851,7 +1851,9 @@ enum kvm_intr_type {
 	((vcpu) && (vcpu)->arch.handling_intr_from_guest && \
 	 (!!in_nmi() == ((vcpu)->arch.handling_intr_from_guest == KVM_HANDLING_NMI)))
 
-#ifdef CONFIG_KVM_VM_MEMORY_ATTRIBUTES
+#if defined(CONFIG_KVM_SW_PROTECTED_VM) ||	\
+    defined(CONFIG_KVM_INTEL_TDX) ||		\
+    defined(CONFIG_KVM_AMD_SEV)
 #define kvm_arch_has_private_mem(kvm) ((kvm)->arch.has_private_mem)
 #endif
 #ifdef CONFIG_HAVE_KVM_ARCH_GMEM_CONVERT
diff --git a/arch/x86/kvm/Kconfig b/arch/x86/kvm/Kconfig
index abb108886733a..2c3c22aeafa54 100644
--- a/arch/x86/kvm/Kconfig
+++ b/arch/x86/kvm/Kconfig
@@ -81,13 +81,21 @@ config KVM_WERROR
 	  If in doubt, say "N".
 
 config KVM_VM_MEMORY_ATTRIBUTES
-	bool
+	bool "Enable per-VM PRIVATE vs. SHARED attributes (for CoCo VMs)"
+	depends on KVM_SW_PROTECTED_VM || KVM_INTEL_TDX || KVM_AMD_SEV
+	help
+	  Enable support for tracking PRIVATE vs. SHARED memory using per-VM
+	  memory attributes.  Using per-VM attributes is deprecated in favor of
+	  tracking PRIVATE state in guest_memfd.  Select this if you need to run
+	  CoCo VMs using a VMM that doesn't support guest_memfd memory
+	  attributes.
+
+	  If unsure, say N.
 
 config KVM_SW_PROTECTED_VM
 	bool "Enable support for KVM software-protected VMs"
 	depends on EXPERT
 	depends on KVM_X86 && X86_64
-	select KVM_VM_MEMORY_ATTRIBUTES
 	help
 	  Enable support for KVM software-protected VMs.  Currently, software-
 	  protected VMs are purely a development and testing vehicle for
@@ -138,7 +146,6 @@ config KVM_INTEL_TDX
 	bool "Intel Trust Domain Extensions (TDX) support"
 	default y
 	depends on INTEL_TDX_HOST
-	select KVM_VM_MEMORY_ATTRIBUTES
 	select HAVE_KVM_ARCH_GMEM_POPULATE
 	help
 	  Provides support for launching Intel Trust Domain Extensions (TDX)
@@ -162,7 +169,6 @@ config KVM_AMD_SEV
 	depends on KVM_AMD && X86_64
 	depends on CRYPTO_DEV_SP_PSP && !(KVM_AMD=y && CRYPTO_DEV_CCP_DD=m)
 	select ARCH_HAS_CC_PLATFORM
-	select KVM_VM_MEMORY_ATTRIBUTES
 	select HAVE_KVM_ARCH_GMEM_CONVERT
 	select HAVE_KVM_ARCH_GMEM_RECLAIM
 	select HAVE_KVM_ARCH_GMEM_INVALIDATE
diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c
index 05c518c9b8078..929fd3e1a01e6 100644
--- a/virt/kvm/kvm_main.c
+++ b/virt/kvm/kvm_main.c
@@ -103,7 +103,10 @@ static bool __ro_after_init allow_unsafe_mappings;
 module_param(allow_unsafe_mappings, bool, 0444);
 
 #ifdef kvm_arch_has_private_mem
-bool __ro_after_init gmem_in_place_conversion = false;
+bool __ro_after_init gmem_in_place_conversion = !IS_ENABLED(CONFIG_KVM_VM_MEMORY_ATTRIBUTES);
+#ifdef CONFIG_KVM_VM_MEMORY_ATTRIBUTES
+module_param(gmem_in_place_conversion, bool, 0444);
+#endif
 EXPORT_SYMBOL_FOR_KVM_INTERNAL(gmem_in_place_conversion);
 #endif
 

-- 
2.55.0.887.g758fc8c411-goog



  parent reply	other threads:[~2026-08-26  9:19 UTC|newest]

Thread overview: 51+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-26  9:17 [PATCH v11 00/46] guest_memfd: In-place conversion support Ackerley Tng
2026-08-26  9:17 ` [PATCH v11 01/46] KVM: guest_memfd: Optimize away conversion overheads via dead-code elimination Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 02/46] KVM: guest_memfd: Use kvm_mem_is_private() when populating guest_memfd memory Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 03/46] KVM: guest_memfd: Introduce per-gmem attributes, use to guard user mappings Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 04/46] KVM: Rename KVM_GENERIC_MEMORY_ATTRIBUTES to KVM_VM_MEMORY_ATTRIBUTES Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 05/46] KVM: Enumerate support for PRIVATE memory iff kvm_arch_has_private_mem is defined Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 06/46] KVM: Rename memory attribute APIs to prepare for in-place gmem conversion Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 07/46] KVM: Rename kvm_mem_is_private() to kvm_is_private_gfn() Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 08/46] KVM: Provide generic interface for checking memory private/shared status Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 09/46] KVM: guest_memfd: Stub in ability to enable in-place shared<=>private conversion Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 10/46] KVM: Consolidate private memory and guest_memfd ifdeffery in kvm_host.h Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 11/46] KVM: guest_memfd: Invalidate both SHARED and PRIVATE mappings for in-place conversions Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 12/46] KVM: guest_memfd: Pass mapping type filter to invalidation helper Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 13/46] KVM: guest_memfd: Add base support for KVM_SET_MEMORY_ATTRIBUTES2 Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 14/46] KVM: guest_memfd: Ensure pages are not in use before conversion Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 15/46] KVM: guest_memfd: Call arch make_shared callback for to-shared conversion Ackerley Tng
2026-08-26 19:44   ` Sean Christopherson
2026-08-26 22:17     ` Michael Roth
2026-08-26 22:33       ` Sean Christopherson
2026-08-26 23:47         ` Michael Roth
2026-08-26  9:18 ` [PATCH v11 16/46] KVM: guest_memfd: Return early if range already has requested attributes Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 17/46] mm/gup: factor out LRU cache draining for folio into lru_cache_drain_for_folio() Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 18/46] KVM: guest_memfd: Handle lru_add fbatch refcounts during conversion safety check Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 19/46] KVM: guest_memfd: Zero page while getting pfn Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 20/46] KVM: SEV: Make 'uaddr' parameter optional for KVM_SEV_SNP_LAUNCH_UPDATE Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 21/46] KVM: TDX: Make source page optional for KVM_TDX_INIT_MEM_REGION Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 22/46] KVM: Move KVM_VM_MEMORY_ATTRIBUTES config definition to x86 Ackerley Tng
2026-08-26  9:18 ` Ackerley Tng [this message]
2026-08-26  9:18 ` [PATCH v11 24/46] KVM: guest_memfd: Enable INIT_SHARED on guest_memfd for x86 Coco VMs Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 25/46] KVM: selftests: Create gmem fd before "regular" fd when adding memslot Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 26/46] KVM: selftests: Rename guest_memfd{,_offset} to gmem_{fd,offset} Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 27/46] KVM: selftests: Add support for mmap() on guest_memfd in core library Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 28/46] KVM: selftests: Add selftests global for guest memory attributes capability Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 29/46] KVM: selftests: Add helpers for calling ioctls on guest_memfd Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 30/46] KVM: selftests: Test basic single-page conversion flow Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 31/46] KVM: selftests: Test conversion flow when INIT_SHARED Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 32/46] KVM: selftests: Test conversion precision in guest_memfd Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 33/46] KVM: selftests: Test conversion before allocation Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 34/46] KVM: selftests: Convert with allocated folios in different layouts Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 35/46] KVM: selftests: Test that truncation does not change shared/private status Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 36/46] KVM: selftests: Test that shared/private status is consistent across processes Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 37/46] KVM: selftests: Add helpers to pin pages with CONFIG_GUP_TEST Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 38/46] KVM: selftests: Test conversion with elevated page refcount Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 39/46] KVM: selftests: Reset shared memory after hole-punching Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 40/46] KVM: selftests: Provide function to look up guest_memfd details from gpa Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 41/46] KVM: selftests: Provide common function to set memory attributes Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 42/46] KVM: selftests: Make TEST_EXPECT_SIGBUS thread-safe Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 43/46] KVM: selftests: Support guest_memfd attributes in private_mem_conversions_test Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 44/46] KVM: selftests: Set up page size and alignment independently for guest_memfd Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 45/46] KVM: selftests: Test in-place conversions in private_mem_conversions_test Ackerley Tng
2026-08-26  9:18 ` [PATCH v11 46/46] KVM: selftests: Update private memory exits test to work with per-gmem attributes Ackerley Tng

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260826-gmem-inplace-conversion-v11-23-0a15d8a799aa@google.com \
    --to=ackerleytng@google.com \
    --cc=aik@amd.com \
    --cc=akpm@linux-foundation.org \
    --cc=andrew.jones@linux.dev \
    --cc=aneesh.kumar@kernel.org \
    --cc=axelrasmussen@google.com \
    --cc=baohua@kernel.org \
    --cc=baoquan.he@linux.dev \
    --cc=binbin.wu@linux.intel.com \
    --cc=bp@alien8.de \
    --cc=brauner@kernel.org \
    --cc=chao.p.peng@linux.intel.com \
    --cc=chrisl@kernel.org \
    --cc=corbet@lwn.net \
    --cc=dave.hansen@linux.intel.com \
    --cc=david@kernel.org \
    --cc=forkloop@google.com \
    --cc=fuad.tabba@linux.dev \
    --cc=hpa@zytor.com \
    --cc=jgg@ziepe.ca \
    --cc=jhubbard@nvidia.com \
    --cc=jmattson@google.com \
    --cc=jthoughton@google.com \
    --cc=kas@kernel.org \
    --cc=kasong@tencent.com \
    --cc=kvm@vger.kernel.org \
    --cc=liam@infradead.org \
    --cc=linux-coco@lists.linux.dev \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=linux-trace-kernel@vger.kernel.org \
    --cc=mathieu.desnoyers@efficios.com \
    --cc=mhiramat@kernel.org \
    --cc=michael.roth@amd.com \
    --cc=mingo@redhat.com \
    --cc=nphamcs@gmail.com \
    --cc=oupton@kernel.org \
    --cc=pankaj.gupta@amd.com \
    --cc=pbonzini@redhat.com \
    --cc=peterx@redhat.com \
    --cc=pratyush@kernel.org \
    --cc=qi.zheng@linux.dev \
    --cc=qperret@google.com \
    --cc=rick.p.edgecombe@intel.com \
    --cc=rientjes@google.com \
    --cc=rostedt@goodmis.org \
    --cc=seanjc@google.com \
    --cc=shakeel.butt@linux.dev \
    --cc=shikemeng@huaweicloud.com \
    --cc=shivankg@amd.com \
    --cc=shuah@kernel.org \
    --cc=skhan@linuxfoundation.org \
    --cc=steven.price@arm.com \
    --cc=suzuki.poulose@arm.com \
    --cc=tarunsahu@google.com \
    --cc=tglx@kernel.org \
    --cc=vannapurve@google.com \
    --cc=vbabka@kernel.org \
    --cc=weixugc@google.com \
    --cc=willy@infradead.org \
    --cc=wyihan@google.com \
    --cc=x86@kernel.org \
    --cc=xiaoyao.li@intel.com \
    --cc=yan.y.zhao@intel.com \
    --cc=youngjun.park@lge.com \
    --cc=yuanchu@google.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox