From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7C3C02DCF55; Wed, 29 Jul 2026 00:35:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785285309; cv=none; b=qo3155QI4X09pXdA8HRRmz4mez7OEenZsxlaq7zVoXSHz2prgyWA3CqyMWEWTpNXDmL3OV57/lQmSEJJU3zy9KgU9Lwo8Dkk6+oXasVX4u2Hhg0v0RxhuSUyDg5gEgjK6uIaHZFVEmETJ8AjmrcRsyvAV8HKhhHZ2dIwMQdOfu4= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785285309; c=relaxed/simple; bh=y9s8lGEG7811uVLgjbEHjBppzW1/vDu3LGD/tpzmX0c=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=E54PNZBVugyTGS2lGJ2cwpHyL6oPpIcq9mEEdwJt3KiBlmBTDz64lV8O9XodNJMI6rc9v8Gv1rZbuBPqF+/8a5JQ84fwWQRXR5OHuTSPO5YgXk85nnEagpqRotDF/HX4rSCyMS3vcVQywzhlNAsggAVPqbRFFvXSorrhLmITykc= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Ol292dno; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Ol292dno" Received: by smtp.kernel.org (Postfix) with ESMTPS id 129C0C4DE1F; Wed, 29 Jul 2026 00:35:09 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1785285309; bh=y9s8lGEG7811uVLgjbEHjBppzW1/vDu3LGD/tpzmX0c=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=Ol292dnoGjvwiuG6ZBKhP1JoEp7flx3SnTt9QThs7Oab8vfYH+oB/tIkzB5Z3jusr YoUETk6buG13nHngxL2vYZ546byUno/xHbpxCUHQ/IqTr4X9fe3vrrjLmMuOU7sGS8 XGMpXjZL6+/XO4kjILtfUL9cMy8uO1pNC8kbxaxiTPVQE00TxrJxFwWnbY+FPBZLxu xFWZCsP31Ncvy92MNWSr4MnvqJhUKtKEXJoYCZk2OKgYnD1CmJdWScDbUaccQWfU7X T2+Fnwdcds++F7ubbIBj+ClZJ7ECDsHpTbVqljyxec5Y3ZrEy0Z4bI3stcWVQKwnek B65Cyw5n0xgqA== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id E893CC54F52; Wed, 29 Jul 2026 00:35:08 +0000 (UTC) From: Ackerley Tng via B4 Relay Date: Tue, 28 Jul 2026 17:35:14 -0700 Subject: [PATCH v9 15/41] KVM: guest_memfd: Handle lru_add fbatch refcounts during conversion safety check Precedence: bulk X-Mailing-List: linux-doc@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: 7bit Message-Id: <20260728-gmem-inplace-conversion-v9-15-35f9aec2aed2@google.com> References: <20260728-gmem-inplace-conversion-v9-0-35f9aec2aed2@google.com> In-Reply-To: <20260728-gmem-inplace-conversion-v9-0-35f9aec2aed2@google.com> To: aik@amd.com, andrew.jones@linux.dev, binbin.wu@linux.intel.com, brauner@kernel.org, chao.p.peng@linux.intel.com, david@kernel.org, jmattson@google.com, jthoughton@google.com, michael.roth@amd.com, oupton@kernel.org, pankaj.gupta@amd.com, qperret@google.com, rick.p.edgecombe@intel.com, rientjes@google.com, shivankg@amd.com, steven.price@arm.com, tabba@google.com, willy@infradead.org, wyihan@google.com, yan.y.zhao@intel.com, forkloop@google.com, pratyush@kernel.org, suzuki.poulose@arm.com, aneesh.kumar@kernel.org, liam@infradead.org, Paolo Bonzini , Sean Christopherson , Thomas Gleixner , Ingo Molnar , Borislav Petkov , Dave Hansen , x86@kernel.org, "H. Peter Anvin" , Steven Rostedt , Masami Hiramatsu , Mathieu Desnoyers , Jonathan Corbet , Shuah Khan , Shuah Khan , Vishal Annapurve , Andrew Morton , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Youngjun Park , Qi Zheng , Shakeel Butt , Kiryl Shutsemau , Baoquan He , Jason Gunthorpe , John Hubbard , Peter Xu , Jason Gunthorpe , Vlastimil Babka , Baoquan He Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-mm@kvack.org, linux-coco@lists.linux.dev, Ackerley Tng X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=ed25519-sha256; t=1785285305; l=2739; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=yzGmlEDFy/cjLSw+iMJeM2ur0Qa630sqxH0mSnKpbLc=; b=bHbpb3CAou9uPxE9qLS+0Ete0qwRLtl5oyfQAGGSUyU/rOr+nAFuNfXSrcqFW31J95vZyoagb J95vuSBHVIqDOJiMFt4O3/MaVCuQAloHhhFBIU/6XCZMXUwmIiUmtAk X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Endpoint-Received: by B4 Relay for ackerleytng@google.com/20260225 with auth_id=649 X-Original-From: Ackerley Tng Reply-To: ackerleytng@google.com From: Ackerley Tng When checking if a guest_memfd folio is safe for conversion, its refcount is examined. A folio may be present in a per-CPU lru_add fbatch, which temporarily increases its refcount. This can lead to a false positive, incorrectly indicating that the folio is in use and preventing the conversion, even if it is otherwise safe. The conversion process might not be on the same CPU that holds the folio in its fbatch. Hence, use lru_add_drain_progressive() to progressively drain lru_add fbatches. guest_memfd folios are unevictable, so they can only reside in the lru_add fbatch. If the folio's refcount is still unsafe after draining, then the conversion is truly unsafe and has to be aborted. Signed-off-by: Ackerley Tng --- mm/swap.c | 2 ++ virt/kvm/guest_memfd.c | 12 ++++++++++-- 2 files changed, 12 insertions(+), 2 deletions(-) diff --git a/mm/swap.c b/mm/swap.c index 0f9465d31fe52..4427d76c88d6d 100644 --- a/mm/swap.c +++ b/mm/swap.c @@ -37,6 +37,7 @@ #include #include #include +#include #include "internal.h" @@ -964,6 +965,7 @@ bool lru_add_drain_progressive(int *drain_state) } return false; } +EXPORT_SYMBOL_FOR_KVM(lru_add_drain_progressive); atomic_t lru_disable_count = ATOMIC_INIT(0); diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c index 928fc7f4a01e4..505bb6747620b 100644 --- a/virt/kvm/guest_memfd.c +++ b/virt/kvm/guest_memfd.c @@ -8,6 +8,7 @@ #include #include #include +#include #include "kvm_mm.h" @@ -554,6 +555,7 @@ static bool kvm_gmem_is_safe_for_conversion(struct inode *inode, pgoff_t start, const int filemap_get_folios_refcount = 1; pgoff_t last = start + nr_pages - 1; struct folio_batch fbatch; + int drain_state = 0; bool safe = true; pgoff_t next; int i; @@ -565,9 +567,15 @@ static bool kvm_gmem_is_safe_for_conversion(struct inode *inode, pgoff_t start, for (i = 0; i < folio_batch_count(&fbatch); ++i) { struct folio *folio = fbatch.folios[i]; + int expected_refcount = folio_nr_pages(folio) + + filemap_get_folios_refcount; - if (folio_ref_count(folio) != - folio_nr_pages(folio) + filemap_get_folios_refcount) { + while (folio_may_be_lru_cached(folio) && + folio_ref_count(folio) != expected_refcount && + lru_add_drain_progressive(&drain_state)) + ; + + if (folio_ref_count(folio) != expected_refcount) { safe = false; *err_index = max(start, folio->index); break; -- 2.55.0.508.g3f0d502094-goog