All of lore.kernel.org
 help / color / mirror / Atom feed
From: Yan Zhao <yan.y.zhao@intel.com>
To: <ackerleytng@google.com>, <aik@amd.com>, <andrew.jones@linux.dev>,
	<binbin.wu@linux.intel.com>, <brauner@kernel.org>,
	<chao.p.peng@linux.intel.com>, <david@kernel.org>,
	<jmattson@google.com>, <jthoughton@google.com>,
	<michael.roth@amd.com>, <oupton@kernel.org>,
	<pankaj.gupta@amd.com>, <qperret@google.com>,
	<rick.p.edgecombe@intel.com>, <rientjes@google.com>,
	<shivankg@amd.com>, <steven.price@arm.com>, <tabba@google.com>,
	<willy@infradead.org>, <wyihan@google.com>, <forkloop@google.com>,
	<pratyush@kernel.org>, <suzuki.poulose@arm.com>,
	<aneesh.kumar@kernel.org>, <liam@infradead.org>,
	Paolo Bonzini <pbonzini@redhat.com>,
	Sean Christopherson <seanjc@google.com>,
	"Thomas Gleixner" <tglx@kernel.org>,
	Ingo Molnar <mingo@redhat.com>, Borislav Petkov <bp@alien8.de>,
	Dave Hansen <dave.hansen@linux.intel.com>, <x86@kernel.org>,
	"H. Peter Anvin" <hpa@zytor.com>,
	Steven Rostedt <rostedt@goodmis.org>,
	Masami Hiramatsu <mhiramat@kernel.org>,
	Mathieu Desnoyers <mathieu.desnoyers@efficios.com>,
	Jonathan Corbet <corbet@lwn.net>,
	"Shuah Khan" <skhan@linuxfoundation.org>,
	Shuah Khan <shuah@kernel.org>,
	"Vishal Annapurve" <vannapurve@google.com>,
	Andrew Morton <akpm@linux-foundation.org>,
	Chris Li <chrisl@kernel.org>, Kairui Song <kasong@tencent.com>,
	Kemeng Shi <shikemeng@huaweicloud.com>,
	Nhat Pham <nphamcs@gmail.com>, Barry Song <baohua@kernel.org>,
	Axel Rasmussen <axelrasmussen@google.com>,
	Yuanchu Xie <yuanchu@google.com>, Wei Xu <weixugc@google.com>,
	Youngjun Park <youngjun.park@lge.com>,
	Qi Zheng <qi.zheng@linux.dev>,
	Shakeel Butt <shakeel.butt@linux.dev>,
	Kiryl Shutsemau <kas@kernel.org>,
	Baoquan He <baoquan.he@linux.dev>, Jason Gunthorpe <jgg@ziepe.ca>,
	John Hubbard <jhubbard@nvidia.com>, Peter Xu <peterx@redhat.com>,
	<tarunsahu@google.com>, Vlastimil Babka <vbabka@kernel.org>,
	<kvm@vger.kernel.org>, <linux-kernel@vger.kernel.org>,
	<linux-trace-kernel@vger.kernel.org>, <linux-doc@vger.kernel.org>,
	<linux-kselftest@vger.kernel.org>, <linux-mm@kvack.org>,
	<linux-coco@lists.linux.dev>
Subject: Re: [PATCH v10 11/41] KVM: guest_memfd: Ensure pages are not in use before conversion
Date: Mon, 10 Aug 2026 05:51:40 +0800	[thread overview]
Message-ID: <anj2bPrGqzZgZWMF@yzhao56-desk.sh.intel.com> (raw)
In-Reply-To: <anZ4W9o5pTWIEgMY@yzhao56-desk.sh.intel.com>

On Sat, Aug 08, 2026 at 08:29:15AM +0800, Yan Zhao wrote:
> On Fri, Aug 07, 2026 at 02:52:50PM -0700, Ackerley Tng via B4 Relay wrote:
> > +static bool kvm_gmem_is_safe_for_conversion(struct inode *inode, pgoff_t start,
> > +					    size_t nr_pages, pgoff_t *err_index)
> > +{
> > +	struct address_space *mapping = inode->i_mapping;
> > +	const int filemap_get_folios_refcount = 1;
> > +	pgoff_t last = start + nr_pages - 1;
> > +	struct folio_batch fbatch;
> > +	bool safe = true;
> > +	pgoff_t next;
> > +	int i;
> > +
> > +	folio_batch_init(&fbatch);
> > +
> > +	next = start;
> > +	while (safe && filemap_get_folios(mapping, &next, last, &fbatch)) {
> > +		for (i = 0; i < folio_batch_count(&fbatch); ++i) {
> > +			struct folio *folio = fbatch.folios[i];
> > +
> > +			if (folio_ref_count(folio) !=
> > +			    folio_nr_pages(folio) + filemap_get_folios_refcount) {
> > +				safe = false;
> > +				*err_index = max(start, folio->index);
> > +				break;
> > +			}
> > +		}
> > +
> > +		folio_batch_release(&fbatch);
> > +		cond_resched();
> > +	}
> > +
> > +	return safe;
> > +}
> > +
> >  static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start,
> > -				     size_t nr_pages, uint64_t attrs)
> > +				     size_t nr_pages, uint64_t attrs,
> > +				     pgoff_t *err_index)
> >  {
> >  	bool to_private = attrs & KVM_MEMORY_ATTRIBUTE_PRIVATE;
> >  	struct address_space *mapping = inode->i_mapping;
> > @@ -542,8 +576,21 @@ static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start,
> >  
> >  	mas_init(&mas, mt, start);
> >  	r = kvm_gmem_mas_preallocate(&mas, attrs, start, nr_pages);
> > -	if (r)
> > +	if (r) {
> > +		*err_index = start;
> >  		goto out;
> > +	}
> > +
> > +	if (to_private) {
> > +		unmap_mapping_pages(mapping, start, nr_pages, false);
> > +
> > +		if (!kvm_gmem_is_safe_for_conversion(inode, start, nr_pages,
> > +						     err_index)) {
> Note: conversion failures could occur if another vCPU is attempting to map a GFN
> within this range.
> 
> CPU 0 (setting attributes)          CPU 1 (attempting to map)
> --------------------------          --------------------
>                                  A: mmu_invalidate_retry_gfn_unsafe
>                                     filemap_invalidate_lock_shared
>                                     __kvm_gmem_get_pfn ==> folio refcount++
>                                     filemap_invalidate_unlock_shared
> 
> filemap_invalidate_lock
> filemap_get_folios
> check folio_ref_count(folio) ==> Not match !!
> filemap_invalidate_unlock
> 
>                                  B: read_lock(&vcpu->kvm->mmu_lock);
>                                     is_page_fault_stale
>                                     kvm_mmu_finish_page_fault ==>folio recount--
> 				    read_unlock(&vcpu->kvm->mmu_lock);
> 
> 
> Retrying in kvm_gmem_is_safe_for_conversion() or moving the invocation of
> kvm_mmu_invalidate_start() + kvm_mmu_invalidate_range_add() to an earlier
> position does not help as long as CPU 1 stays at stage A.
> 
> So, should we avoid this failure?
> e.g., by moving filemap_invalidate_unlock_shared() from stage A to after
> stage B?

Or what about having KVM always treat gmem page as non-refcounted, and have
kvm_gmem_get_pfn() put folio refcount before releasing the filemap invalidate
lock?
Below patch is applied and tested at the end of this series.

From 8c2f29bc15bceb6a8fa103cf2585ec11354fd74e Mon Sep 17 00:00:00 2001
From: Yan Zhao <yan.y.zhao@intel.com>
Date: Mon, 10 Aug 2026 06:24:52 +0800
Subject: [PATCH] KVM: guest_memfd: Return gmem page as non-refcounted

Have kvm_gmem_get_pfn() put gmem page refcount before releasing filemap
invalidate lock and return the gmem page as non-refcounted. This avoids
gmem memory attribute conversion failure caused by temporarily holding gmem
page after faulting and before completing mapping.

guest_memfd always holds gmem page in filemap cache. TDX does not increment
gmem page refcount when having gmem pages mapped in S-EPT. Additionally,
as gmem pages are not swappable, setting dirty or accessed bit is not
necessary. Therefore, there's no need to treat gmem pages as refcounted
pages.

Signed-off-by: Yan Zhao <yan.y.zhao@intel.com>
---
 virt/kvm/guest_memfd.c | 5 ++---
 1 file changed, 2 insertions(+), 3 deletions(-)

diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c
index 2115e73e455a..e357b4ffa777 100644
--- a/virt/kvm/guest_memfd.c
+++ b/virt/kvm/guest_memfd.c
@@ -1332,11 +1332,10 @@ int kvm_gmem_get_pfn(struct kvm *kvm, struct kvm_memory_slot *slot,
 #endif
 
 	folio_unlock(folio);
+	folio_put(folio);
 
 	if (!r)
-		*page = folio_file_page(folio, index);
-	else
-		folio_put(folio);
+		*page = NULL;
 
 out:
 	filemap_invalidate_unlock_shared(file_inode(file)->i_mapping);
-- 
2.43.2

  reply	other threads:[~2026-08-09 22:32 UTC|newest]

Thread overview: 109+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-07 21:52 [PATCH v10 00/41] guest_memfd: In-place conversion support Ackerley Tng via B4 Relay
2026-08-07 21:52 ` Ackerley Tng
2026-08-07 21:52 ` [PATCH v10 01/41] KVM: guest_memfd: Use kvm_mem_is_private() when populating guest_memfd memory Ackerley Tng via B4 Relay
2026-08-07 21:52   ` Ackerley Tng
2026-08-07 21:52 ` [PATCH v10 02/41] KVM: guest_memfd: Introduce per-gmem attributes, use to guard user mappings Ackerley Tng via B4 Relay
2026-08-07 21:52   ` Ackerley Tng
2026-08-10 15:12   ` Sean Christopherson
2026-08-07 21:52 ` [PATCH v10 03/41] KVM: Rename KVM_GENERIC_MEMORY_ATTRIBUTES to KVM_VM_MEMORY_ATTRIBUTES Ackerley Tng via B4 Relay
2026-08-07 21:52   ` Ackerley Tng
2026-08-07 21:52 ` [PATCH v10 04/41] KVM: Enumerate support for PRIVATE memory iff kvm_arch_has_private_mem is defined Ackerley Tng via B4 Relay
2026-08-07 21:52   ` Ackerley Tng
2026-08-07 21:52 ` [PATCH v10 05/41] KVM: Rename memory attribute APIs to prepare for in-place gmem conversion Ackerley Tng via B4 Relay
2026-08-07 21:52   ` Ackerley Tng
2026-08-07 21:52 ` [PATCH v10 06/41] KVM: Provide generic interface for checking memory private/shared status Ackerley Tng via B4 Relay
2026-08-07 21:52   ` Ackerley Tng
2026-08-07 21:52 ` [PATCH v10 07/41] KVM: guest_memfd: Stub in ability to enable in-place shared<=>private conversion Ackerley Tng via B4 Relay
2026-08-07 21:52   ` Ackerley Tng
2026-08-10  8:49   ` David Hildenbrand (Arm)
2026-08-10 15:01     ` Sean Christopherson
2026-08-07 21:52 ` [PATCH v10 08/41] KVM: Consolidate private memory and guest_memfd ifdeffery in kvm_host.h Ackerley Tng via B4 Relay
2026-08-07 21:52   ` Ackerley Tng
2026-08-10  8:50   ` David Hildenbrand (Arm)
2026-08-07 21:52 ` [PATCH v10 09/41] KVM: guest_memfd: Filter both shared and private when invalidating Ackerley Tng via B4 Relay
2026-08-07 21:52   ` Ackerley Tng
2026-08-10  9:06   ` David Hildenbrand (Arm)
2026-08-10  9:27     ` Suzuki K Poulose
2026-08-10 19:11       ` Ackerley Tng
2026-08-07 21:52 ` [PATCH v10 10/41] KVM: guest_memfd: Add base support for KVM_SET_MEMORY_ATTRIBUTES2 Ackerley Tng via B4 Relay
2026-08-07 21:52   ` Ackerley Tng
2026-08-07 21:52 ` [PATCH v10 11/41] KVM: guest_memfd: Ensure pages are not in use before conversion Ackerley Tng via B4 Relay
2026-08-07 21:52   ` Ackerley Tng
2026-08-08  0:29   ` Yan Zhao
2026-08-09 21:51     ` Yan Zhao [this message]
2026-08-10 21:06       ` Ackerley Tng
2026-08-11  1:04         ` Yan Zhao
2026-08-11  2:17           ` Ackerley Tng
2026-08-11  4:50             ` Yan Zhao
2026-08-10  9:19   ` David Hildenbrand (Arm)
2026-08-10 21:41     ` Ackerley Tng
2026-08-10 22:26       ` Sean Christopherson
2026-08-07 21:52 ` [PATCH v10 12/41] KVM: guest_memfd: Call arch make_shared callback for to-shared conversion Ackerley Tng via B4 Relay
2026-08-07 21:52   ` Ackerley Tng
2026-08-07 21:52 ` [PATCH v10 13/41] KVM: guest_memfd: Return early if range already has requested attributes Ackerley Tng via B4 Relay
2026-08-07 21:52   ` Ackerley Tng
2026-08-07 21:52 ` [PATCH v10 14/41] mm/gup: factor out LRU cache draining for folio into lru_cache_drain_for_folio() Ackerley Tng via B4 Relay
2026-08-07 21:52   ` Ackerley Tng
2026-08-10  9:22   ` David Hildenbrand (Arm)
2026-08-07 21:52 ` [PATCH v10 15/41] KVM: guest_memfd: Handle lru_add fbatch refcounts during conversion safety check Ackerley Tng via B4 Relay
2026-08-07 21:52   ` Ackerley Tng
2026-08-10  9:25   ` David Hildenbrand (Arm)
2026-08-10 21:29     ` Ackerley Tng
2026-08-07 21:52 ` [PATCH v10 16/41] KVM: guest_memfd: Zero page while getting pfn Ackerley Tng via B4 Relay
2026-08-07 21:52   ` Ackerley Tng
2026-08-10  9:30   ` David Hildenbrand (Arm)
2026-08-10 23:41     ` Ackerley Tng
2026-08-11  0:18       ` Sean Christopherson
2026-08-07 21:52 ` [PATCH v10 17/41] KVM: SEV: Make 'uaddr' parameter optional for KVM_SEV_SNP_LAUNCH_UPDATE Ackerley Tng via B4 Relay
2026-08-07 21:52   ` Ackerley Tng
2026-08-07 21:52 ` [PATCH v10 18/41] KVM: TDX: Make source page optional for KVM_TDX_INIT_MEM_REGION Ackerley Tng via B4 Relay
2026-08-07 21:52   ` Ackerley Tng
2026-08-07 21:52 ` [PATCH v10 19/41] KVM: Move KVM_VM_MEMORY_ATTRIBUTES config definition to x86 Ackerley Tng via B4 Relay
2026-08-07 21:52   ` Ackerley Tng
2026-08-10  9:32   ` David Hildenbrand (Arm)
2026-08-07 21:52 ` [PATCH v10 20/41] KVM: Let userspace disable per-VM mem attributes, enable per-gmem attributes Ackerley Tng via B4 Relay
2026-08-07 21:52   ` Ackerley Tng
2026-08-10  9:36   ` David Hildenbrand (Arm)
2026-08-07 21:53 ` [PATCH v10 21/41] KVM: guest_memfd: Enable INIT_SHARED on guest_memfd for x86 Coco VMs Ackerley Tng via B4 Relay
2026-08-07 21:53   ` Ackerley Tng
2026-08-07 21:53 ` [PATCH v10 22/41] KVM: selftests: Create gmem fd before "regular" fd when adding memslot Ackerley Tng via B4 Relay
2026-08-07 21:53   ` Ackerley Tng
2026-08-07 21:53 ` [PATCH v10 23/41] KVM: selftests: Rename guest_memfd{,_offset} to gmem_{fd,offset} Ackerley Tng via B4 Relay
2026-08-07 21:53   ` Ackerley Tng
2026-08-07 21:53 ` [PATCH v10 24/41] KVM: selftests: Add support for mmap() on guest_memfd in core library Ackerley Tng via B4 Relay
2026-08-07 21:53   ` Ackerley Tng
2026-08-07 21:53 ` [PATCH v10 25/41] KVM: selftests: Add selftests global for guest memory attributes capability Ackerley Tng via B4 Relay
2026-08-07 21:53   ` Ackerley Tng
2026-08-07 21:53 ` [PATCH v10 26/41] KVM: selftests: Add helpers for calling ioctls on guest_memfd Ackerley Tng via B4 Relay
2026-08-07 21:53   ` Ackerley Tng
2026-08-07 21:53 ` [PATCH v10 27/41] KVM: selftests: Test basic single-page conversion flow Ackerley Tng via B4 Relay
2026-08-07 21:53   ` Ackerley Tng
2026-08-07 21:53 ` [PATCH v10 28/41] KVM: selftests: Test conversion flow when INIT_SHARED Ackerley Tng via B4 Relay
2026-08-07 21:53   ` Ackerley Tng
2026-08-07 21:53 ` [PATCH v10 29/41] KVM: selftests: Test conversion precision in guest_memfd Ackerley Tng via B4 Relay
2026-08-07 21:53   ` Ackerley Tng
2026-08-07 21:53 ` [PATCH v10 30/41] KVM: selftests: Test conversion before allocation Ackerley Tng via B4 Relay
2026-08-07 21:53   ` Ackerley Tng
2026-08-07 21:53 ` [PATCH v10 31/41] KVM: selftests: Convert with allocated folios in different layouts Ackerley Tng via B4 Relay
2026-08-07 21:53   ` Ackerley Tng
2026-08-07 21:53 ` [PATCH v10 32/41] KVM: selftests: Test that truncation does not change shared/private status Ackerley Tng via B4 Relay
2026-08-07 21:53   ` Ackerley Tng
2026-08-07 21:53 ` [PATCH v10 33/41] KVM: selftests: Test that shared/private status is consistent across processes Ackerley Tng via B4 Relay
2026-08-07 21:53   ` Ackerley Tng
2026-08-07 21:53 ` [PATCH v10 34/41] KVM: selftests: Add helpers to pin pages with CONFIG_GUP_TEST Ackerley Tng via B4 Relay
2026-08-07 21:53   ` Ackerley Tng
2026-08-07 21:53 ` [PATCH v10 35/41] KVM: selftests: Test conversion with elevated page refcount Ackerley Tng via B4 Relay
2026-08-07 21:53   ` Ackerley Tng
2026-08-07 21:53 ` [PATCH v10 36/41] KVM: selftests: Reset shared memory after hole-punching Ackerley Tng via B4 Relay
2026-08-07 21:53   ` Ackerley Tng
2026-08-07 21:53 ` [PATCH v10 37/41] KVM: selftests: Provide function to look up guest_memfd details from gpa Ackerley Tng via B4 Relay
2026-08-07 21:53   ` Ackerley Tng
2026-08-07 21:53 ` [PATCH v10 38/41] KVM: selftests: Provide common function to set memory attributes Ackerley Tng via B4 Relay
2026-08-07 21:53   ` Ackerley Tng
2026-08-07 21:53 ` [PATCH v10 39/41] KVM: selftests: Make TEST_EXPECT_SIGBUS thread-safe Ackerley Tng via B4 Relay
2026-08-07 21:53   ` Ackerley Tng
2026-08-07 21:53 ` [PATCH v10 40/41] KVM: selftests: Update private_mem_conversions_test to mmap() guest_memfd Ackerley Tng via B4 Relay
2026-08-07 21:53   ` Ackerley Tng
2026-08-07 21:53 ` [PATCH v10 41/41] KVM: selftests: Update private memory exits test to work with per-gmem attributes Ackerley Tng via B4 Relay
2026-08-07 21:53   ` Ackerley Tng
2026-08-10  8:43 ` [PATCH v10 00/41] guest_memfd: In-place conversion support David Hildenbrand (Arm)

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=anj2bPrGqzZgZWMF@yzhao56-desk.sh.intel.com \
    --to=yan.y.zhao@intel.com \
    --cc=ackerleytng@google.com \
    --cc=aik@amd.com \
    --cc=akpm@linux-foundation.org \
    --cc=andrew.jones@linux.dev \
    --cc=aneesh.kumar@kernel.org \
    --cc=axelrasmussen@google.com \
    --cc=baohua@kernel.org \
    --cc=baoquan.he@linux.dev \
    --cc=binbin.wu@linux.intel.com \
    --cc=bp@alien8.de \
    --cc=brauner@kernel.org \
    --cc=chao.p.peng@linux.intel.com \
    --cc=chrisl@kernel.org \
    --cc=corbet@lwn.net \
    --cc=dave.hansen@linux.intel.com \
    --cc=david@kernel.org \
    --cc=forkloop@google.com \
    --cc=hpa@zytor.com \
    --cc=jgg@ziepe.ca \
    --cc=jhubbard@nvidia.com \
    --cc=jmattson@google.com \
    --cc=jthoughton@google.com \
    --cc=kas@kernel.org \
    --cc=kasong@tencent.com \
    --cc=kvm@vger.kernel.org \
    --cc=liam@infradead.org \
    --cc=linux-coco@lists.linux.dev \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=linux-trace-kernel@vger.kernel.org \
    --cc=mathieu.desnoyers@efficios.com \
    --cc=mhiramat@kernel.org \
    --cc=michael.roth@amd.com \
    --cc=mingo@redhat.com \
    --cc=nphamcs@gmail.com \
    --cc=oupton@kernel.org \
    --cc=pankaj.gupta@amd.com \
    --cc=pbonzini@redhat.com \
    --cc=peterx@redhat.com \
    --cc=pratyush@kernel.org \
    --cc=qi.zheng@linux.dev \
    --cc=qperret@google.com \
    --cc=rick.p.edgecombe@intel.com \
    --cc=rientjes@google.com \
    --cc=rostedt@goodmis.org \
    --cc=seanjc@google.com \
    --cc=shakeel.butt@linux.dev \
    --cc=shikemeng@huaweicloud.com \
    --cc=shivankg@amd.com \
    --cc=shuah@kernel.org \
    --cc=skhan@linuxfoundation.org \
    --cc=steven.price@arm.com \
    --cc=suzuki.poulose@arm.com \
    --cc=tabba@google.com \
    --cc=tarunsahu@google.com \
    --cc=tglx@kernel.org \
    --cc=vannapurve@google.com \
    --cc=vbabka@kernel.org \
    --cc=weixugc@google.com \
    --cc=willy@infradead.org \
    --cc=wyihan@google.com \
    --cc=x86@kernel.org \
    --cc=youngjun.park@lge.com \
    --cc=yuanchu@google.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.