From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pg1-f198.google.com (mail-pg1-f198.google.com [209.85.215.198]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 44D683BBFAB for ; Wed, 26 Aug 2026 09:18:41 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.198 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787735926; cv=none; b=VHU7v3xysVwuAIdL2lHyu+ZYbRomrPj+8vYNdVnWhepB3YR+l2t3ckHkTb1Pg07H6MgcFlGBoXkIyJDzsU84jUG7WxHIEKiEWw/FGW8hlUbBKNEPWjZR67tBy2tsAJunuO3rbGb/liBIEuqsbd8VFS3QQPA3MD+b95dHR8caPuI= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787735926; c=relaxed/simple; bh=ZTwn4UyMShrGZfEU90otJVup2JEojFMgiToKCnnSVVE=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=dOSfNGliHsllhdit428dXLcTPcx5CwroOIInVKa1BRjEKIpq9V1sx5BDboSWrqOR62fcsLhGQq1tHXuZnifPVGdLz+jW5aLyQZxs9Fw3s0vLwnLqDRZVnjbgynPqrjEQ7KpzzjlToLANNHtAHyKiyYLJ97Cfo+IOp9Hx3tQUyoc= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--ackerleytng.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=cTfr9TBD; arc=none smtp.client-ip=209.85.215.198 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--ackerleytng.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="cTfr9TBD" Received: by mail-pg1-f198.google.com with SMTP id 41be03b00d2f7-cbb467e56aaso608917a12.1 for ; Wed, 26 Aug 2026 02:18:40 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1787735919; x=1788340719; darn=lists.linux.dev; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=S/nf3e4536rn08VcIwbRdla3FrzM9D86FJQTH62XB18=; b=cTfr9TBDdfKLNt/2juYqIFpkHTDGxA8kmZ7zkZPHDtrEYbxn4Ugot6rH9U55lXnCZM wlVbG5VSGlklNCB4GxoWVyl581McemmnbUgvsP6ABLiMTDYVj3fhtuxNmSz5OOia13y1 79F7Ab1HCGfuTH6+ru5hLjGdUz9e4JLxvlLHp4Vrj6J5vgjHMjU1TeDsU4WplVb+zbU0 3d73gop3Z6lcM2pAft3p+PAxBTvt8KossX60JNvKtjV1wKU5J2oHcdvG6frYRcmWs1Rm KokcbAVwKxDsJPZ9htucxOHiPGT7qU8Ft0Ha7LoSONQfK1cUq51uAqUD8hH/5CHq0ctD bikA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787735919; x=1788340719; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=S/nf3e4536rn08VcIwbRdla3FrzM9D86FJQTH62XB18=; b=hAK7UI0IJRLyDnlP1oPks8GuOtdSDoaglDX7NXLUa69QNxqPMV82Zr9voQM/TVUCIx pUmPS/3s2GSW1/Rqf7ou50ylRb+5R5NLri5/0IJKG3MBbRYpPeBftfVLKhXX0Undf0xs FMGAWAYYMZWM/AIJds9RIaNyjOdZj5fH1NiOH3Q5vtIG9tg1f0dGKJHkQjNLKSgJKiKz qz7DsCRMn1+4Y+7/VQhIXcVYnbJtXMoli81G8wHLTVyt6HfSLxQ3EazR5mu9rf8kH28c t6rkRanlMDU7upQFWiV2udelgYli15ikfqPztkjaUWVrct1uuYxo4pVvXLAzgmzr0v10 fGxA== X-Forwarded-Encrypted: i=1; AHgh+RrKpnd12HHagdWYrV0yLR5mV9Vm9I1U0/b2kM/E1sPJx+3Xw8lWMeAtpF2pic0llNuvb0WUoI+qW5lk@lists.linux.dev X-Gm-Message-State: AFuF++kbjPaV5zWTrHIxOGoLfUux2G+5XNhevi+6/rayoCEbOX0kpxbV yRwNqZZ6h76QAUl1P0mHGB5wfhk/BlsA3QXQZKbcb8/aO15DUBrbi4BdOJWaU6XIiYrcoRKTIzB Wnd/o/BHS5jPLPcXLjlMTO//nQQ== X-Received: from pgig7.prod.google.com ([2002:a63:f407:0:b0:cc1:c7d3:7c7b]) (user=ackerleytng job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a21:a383:b0:3cd:8d15:895e with SMTP id adf61e73a8af0-3cf84b53689mr11572460637.12.1787735918168; Wed, 26 Aug 2026 02:18:38 -0700 (PDT) Date: Wed, 26 Aug 2026 09:18:16 +0000 In-Reply-To: <20260826-gmem-inplace-conversion-v11-0-0a15d8a799aa@google.com> Precedence: bulk X-Mailing-List: linux-coco@lists.linux.dev List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260826-gmem-inplace-conversion-v11-0-0a15d8a799aa@google.com> X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Developer-Signature: v=1; a=ed25519-sha256; t=1787735885; l=4708; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=ZTwn4UyMShrGZfEU90otJVup2JEojFMgiToKCnnSVVE=; b=dDBRMblwUg6UwX2I1hz2Tpxt+EI3WovENt6hqBknKpraB0X285GZYRtS86E6IMk5svSgepv73 EfxqI68rjF/CyC9ce+QdXDFmPsPxg8E1UYpFu6gAUyL8f5QVEBtho9j X-Mailer: b4 0.16.0 Message-ID: <20260826-gmem-inplace-conversion-v11-18-0a15d8a799aa@google.com> Subject: [PATCH v11 18/46] KVM: guest_memfd: Handle lru_add fbatch refcounts during conversion safety check From: Ackerley Tng To: aik@amd.com, andrew.jones@linux.dev, binbin.wu@linux.intel.com, brauner@kernel.org, chao.p.peng@linux.intel.com, david@kernel.org, jmattson@google.com, jthoughton@google.com, michael.roth@amd.com, oupton@kernel.org, pankaj.gupta@amd.com, qperret@google.com, rick.p.edgecombe@intel.com, rientjes@google.com, shivankg@amd.com, steven.price@arm.com, willy@infradead.org, wyihan@google.com, yan.y.zhao@intel.com, forkloop@google.com, pratyush@kernel.org, suzuki.poulose@arm.com, aneesh.kumar@kernel.org, liam@infradead.org, Paolo Bonzini , Sean Christopherson , Thomas Gleixner , Ingo Molnar , Borislav Petkov , Dave Hansen , x86@kernel.org, "H. Peter Anvin" , Steven Rostedt , Masami Hiramatsu , Mathieu Desnoyers , Jonathan Corbet , Shuah Khan , Shuah Khan , Vishal Annapurve , Andrew Morton , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Youngjun Park , Qi Zheng , Shakeel Butt , Kiryl Shutsemau , Baoquan He , Jason Gunthorpe , John Hubbard , Peter Xu , tarunsahu@google.com, Fuad Tabba , Vlastimil Babka Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-mm@kvack.org, linux-coco@lists.linux.dev, Ackerley Tng , "Vlastimil Babka (SUSE)" , Fuad Tabba Content-Type: text/plain; charset="utf-8" A guest_memfd folio has no outstanding references if guest_memfd holds the only references on it. Any other references on the folio may indicate another user, and guest_memfd cannot convert it to private if there may be an existing host user. A folio will have outstanding references if it is present in a per-CPU lru_add fbatch. guest_memfd does not actually participate in LRU, but freshly-allocated folios are still added to the lru_add fbatch for batch LRU statistics processing. A folio may also have extra refcounts if it is on the mlock fbatch. These two known "usages" of the folio are handled by calling lru_cache_drain_for_folio, which drains both the lru_add and mlock fbatches. After draining, if the refcount is still elevated, then there are truly outstanding references. If the page may be dma pinned, DMA is using it and hence there are outstanding references. folio_maybe_dma_pinned() can have false positives, but that's only with a significant number of refcounts, at which point draining LRU is not going to move the needle - it can still be concluded that the folio has outstanding references. If the page is still mapped after guest_memfd tried to unmap it earlier in the conversion process, it also has outstanding references. Return true and exit early to avoid unnecessary draining in these 2 cases. Provide a drain status to only drain once ever while processing a batch of folios. Acked-by: Vlastimil Babka (SUSE) Suggested-by: David Hildenbrand Reviewed-by: Fuad Tabba Reviewed-by: Binbin Wu Signed-off-by: Ackerley Tng --- mm/swap.c | 2 ++ virt/kvm/guest_memfd.c | 30 ++++++++++++++++++++++-------- 2 files changed, 24 insertions(+), 8 deletions(-) diff --git a/mm/swap.c b/mm/swap.c index 8e965c8ce9aa9..9f511b97ab110 100644 --- a/mm/swap.c +++ b/mm/swap.c @@ -37,6 +37,7 @@ #include #include #include +#include #include "internal.h" @@ -995,6 +996,7 @@ void lru_cache_drain_for_folio(const struct folio *folio, *drained = LRU_CACHE_DRAINED_ALL; } } +EXPORT_SYMBOL_FOR_KVM(lru_cache_drain_for_folio); atomic_t lru_disable_count = ATOMIC_INIT(0); diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c index 6dc199be0eb87..4912f90567fe8 100644 --- a/virt/kvm/guest_memfd.c +++ b/virt/kvm/guest_memfd.c @@ -8,6 +8,7 @@ #include #include #include +#include #include "kvm_mm.h" #include "guest_memfd.h" @@ -556,10 +557,28 @@ static int kvm_gmem_mas_preallocate(struct ma_state *mas, u64 attributes, return mas_preallocate(mas, xa_mk_value(attributes), GFP_KERNEL); } +static bool __folio_has_outstanding_references(struct folio *folio, + enum lru_cache_drained *drained) +{ + if (folio_maybe_dma_pinned(folio) || folio_mapped(folio)) + return true; + + /* 1 reference held by filemap_get_folios() in the folio batch. */ + lru_cache_drain_for_folio(folio, 1, drained); + + /* + * Outstanding references are anything other than those from the page + * cache, plus 1 temporary reference held by filemap_get_folios() in the + * folio batch. + */ + return folio_ref_count(folio) != folio_nr_pages(folio) + 1; +} + static bool kvm_gmem_has_outstanding_references(struct inode *inode, pgoff_t start, size_t nr_pages, pgoff_t *err_index) { + enum lru_cache_drained drained = LRU_CACHE_NOT_DRAINED; struct address_space *mapping = inode->i_mapping; pgoff_t last = start + nr_pages - 1; bool has_outstanding = false; @@ -570,17 +589,12 @@ static bool kvm_gmem_has_outstanding_references(struct inode *inode, folio_batch_init(&fbatch); next = start; - while (has_outstanding && filemap_get_folios(mapping, &next, last, &fbatch)) { + while (!has_outstanding && filemap_get_folios(mapping, &next, last, &fbatch)) { for (i = 0; i < folio_batch_count(&fbatch); ++i) { struct folio *folio = fbatch.folios[i]; - /* - * Outstanding references are anything other than those - * from the page cache, plus 1 temporary reference held - * by filemap_get_folios() in the folio batch. - */ - if (folio_ref_count(folio) != folio_nr_pages(folio) + 1) { - has_outstanding = true; + has_outstanding = __folio_has_outstanding_references(folio, &drained); + if (has_outstanding) { *err_index = max(start, folio->index); break; } -- 2.55.0.887.g758fc8c411-goog