From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-yw1-f181.google.com (mail-yw1-f181.google.com [209.85.128.181]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 01BDE364043 for ; Wed, 2 Sep 2026 02:21:24 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.181 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788315686; cv=none; b=OWRF8Gda/xmzjzOyDZl82J/fO3D0ivf6UM9Yys8X1S4uPDWe7fbspY1cLrc1yWQN7yjV7z68IicazPJVxGVHeqd1sTzS3LGiqcS/R9dXg+XjfLcT+gKho9h4rdtdx8OrSAFU5u2ZCKM92W46LDwM1FMOUIgeSqxKz/VaGUICX4w= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788315686; c=relaxed/simple; bh=dMv4nlHu8Y249mLMfP00bq7DYL6IzUHigBCjVTtONZg=; h=Date:From:To:cc:Subject:In-Reply-To:Message-ID:References: MIME-Version:Content-Type; b=PSpXQf5ck6E+y22mdu7P8sEqpO5Pn0Yuq2rDGIAYENsZooNeGgDbw6xe3nrM8jgJgUbEvqzvvgW9J5t/tZeLMQcwpVF32sagw0Vrd+XUhnx/psAY79G20f9Qn9af0mnjYBXvuiKts4SFFm97uLkRlmLmbMRLYqQC0riq3E9te9o= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=meLSm2rQ; arc=none smtp.client-ip=209.85.128.181 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="meLSm2rQ" Received: by mail-yw1-f181.google.com with SMTP id 00721157ae682-85aeb506b78so10838087b3.3 for ; Tue, 01 Sep 2026 19:21:24 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1788315684; x=1788920484; darn=vger.kernel.org; h=content-type:mime-version:references:message-id:in-reply-to:subject :cc:to:from:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=GP2nihoXTh44eq0onDqBGjACptMCLq276d/ZwwaKLtA=; b=meLSm2rQNh+3UVY1uUdEjjMfsuClBtztj7u2+Tyeq5hpPRLDsRnLTzMjAFEXNA22Cp OlRM/CrYljMqUZKdlY22Ltu4/786tcDa9kBNMUmHjafmJKGCJMW0FMaXbz7/nia6010A Y00BkzuWN/1i302PQLhgTqjysmRQEEZ6ZSxmDvE7PSJq5i7ughq/fgwIzihqdiBvxlho wTNerakvDwMdKpY0ReezEeR6O2lwFCaxhVkst9NV7Ev3uFskpWIBUvM+9lSELqeI+Dsm zvrum3dxAWwBGqyWXgtNUQd+TSiS8kmT7jWa3WIZs7Ys5r5tXN6i9Ds6yK3CxOKCG/BF yJAw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788315684; x=1788920484; h=content-type:mime-version:references:message-id:in-reply-to:subject :cc:to:from:date:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=GP2nihoXTh44eq0onDqBGjACptMCLq276d/ZwwaKLtA=; b=GTHEpeq2eUULGRGQRuifpZvmK+4F2piHKpIz65+a0OA8xRU9P5jA+j0RsxgtPFSdB0 6gez8TTYk7lnHkD71rQolKhLR/B0Iv7lrvrSiA/zkpn0osX5QCJjbK5y/HuQZteawDuS g/oCMP6T9w9r36OQpmpg6tAATn6ofF10vzoo/DB/BK+E28lVAjtKel28QHvmefBs2f1G rSEC9z559POJifY/soBkpJ+TWDQctVHyjG26cFMacmHLyGg5dsBMGncFpAAGL0bW0mfA 9kbsqWWvbLKYMgD1I+B0WOk9zrxVKzTqQ6NGDXVr7bZFxkbe1ROsJvWzFKDBuxssU6Cf vxXg== X-Forwarded-Encrypted: i=1; AKwUvBxg/G1n5U/+UxhmGSDs1d1aDokI+o3n5o+EmMqOOxLB+2ahJNVsM603p+B+Glw5esduV6rbJW+bGzQ=@vger.kernel.org X-Gm-Message-State: AFuF++mGJZf78oro6XddbxU/Z3UQdcrtDqziTpGMYeWZve/Af03sZELA 1ome5tYoe0hyeGurqabaUvKH/NZ72N9BHhyfmFLjlaCpeT6t63QXkp89cA0rZma8ng== X-Gm-Gg: AYBFou3knhnIuwiijBPT30DpYLr1ILYbTUcMGmmJz+4YENMmT/58Fh+aO/1Hf8LtKnB 5DEM4cxRnoXjElI4a9drH66fNnK8G76m5dCbchJYCLaFCavaxpQR7IRAZVbUrPGqSiUoDBwO0rd AQnZAACMxcqOj16M6XXmOphyVUkk5iI4ZFY2jJX/CXfbU7I5Nf/x3JFh8V/1EEqlnKGvFpRDCeM VpLK67Hgo7mgoouDYAn45ZGVTa7IpV6SLUXqxJuuAgQuSBjRb7AXj6Ed8M0HLRKkW8pjxIlEr2e aTZp90E7xpEmWsuBMyeYC8uaJv1RXylmHYjfsOJ7GtEmLuQJR8fY1FyuK5E23n+B8IIrHvaOeRa PhinmsLIzX0aAzNWlgXuzItGYyIRLULB26xIYo4vnT8oYPRVaxjGWuLfcc/ZhrMVnljmcBRzzoO pzBpVFifj0YV3bvA1QpP8pwY9bIBz4lxl2DWnn3A1QHEbOVo1AKHJr6+9xNL6cHxsyeeGRqjtAd Q3u+60tSpoK9KiXMFPjpf5XivQCTkAArYJ2rWsPwL6J1l5k X-Received: by 2002:a05:690e:4419:b0:66f:7d51:b5ef with SMTP id 956f58d0204a3-66f9b90a9d6mr365005d50.3.1788315681539; Tue, 01 Sep 2026 19:21:21 -0700 (PDT) Received: from darker.attlocal.net (172-10-233-147.lightspeed.sntcca.sbcglobal.net. [172.10.233.147]) by smtp.gmail.com with ESMTPSA id 956f58d0204a3-66f986b612esm893663d50.21.2026.09.01.19.21.15 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 01 Sep 2026 19:21:20 -0700 (PDT) Date: Tue, 1 Sep 2026 19:21:05 -0700 (PDT) From: Hugh Dickins To: Ackerley Tng , Andrew Morton cc: aik@amd.com, andrew.jones@linux.dev, binbin.wu@linux.intel.com, brauner@kernel.org, chao.p.peng@linux.intel.com, david@kernel.org, jmattson@google.com, jthoughton@google.com, michael.roth@amd.com, oupton@kernel.org, pankaj.gupta@amd.com, qperret@google.com, rick.p.edgecombe@intel.com, rientjes@google.com, shivankg@amd.com, steven.price@arm.com, willy@infradead.org, wyihan@google.com, yan.y.zhao@intel.com, forkloop@google.com, pratyush@kernel.org, suzuki.poulose@arm.com, aneesh.kumar@kernel.org, liam@infradead.org, Paolo Bonzini , Sean Christopherson , Thomas Gleixner , Ingo Molnar , Borislav Petkov , Dave Hansen , x86@kernel.org, "H. Peter Anvin" , Steven Rostedt , Masami Hiramatsu , Mathieu Desnoyers , Jonathan Corbet , Shuah Khan , Shuah Khan , Vishal Annapurve , Andrew Morton , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Youngjun Park , Qi Zheng , Shakeel Butt , Kiryl Shutsemau , Baoquan He , Jason Gunthorpe , John Hubbard , Peter Xu , tarunsahu@google.com, Randy Dunlap , Lorenzo Stoakes , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Fuad Tabba , kvm@vger.kernel.org, linux-kernel@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-mm@kvack.org, linux-coco@lists.linux.dev Subject: Re: [PATCH v12 18/45] KVM: guest_memfd: Handle lru_add fbatch refcounts during conversion safety check In-Reply-To: <20260830-gmem-inplace-conversion-v12-18-85e5fd25252a@google.com> Message-ID: References: <20260830-gmem-inplace-conversion-v12-0-85e5fd25252a@google.com> <20260830-gmem-inplace-conversion-v12-18-85e5fd25252a@google.com> Precedence: bulk X-Mailing-List: linux-doc@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=US-ASCII On Sun, 30 Aug 2026, Ackerley Tng via B4 Relay wrote: > From: Ackerley Tng > > A guest_memfd folio has no outstanding references if guest_memfd holds the > only references on it. Any other references on the folio may indicate > another user, and guest_memfd cannot convert it to private if there may be > an existing host user. > > A folio will have outstanding references if it is present in a per-CPU > lru_add fbatch. guest_memfd does not actually participate in LRU, but > freshly-allocated folios are still added to the lru_add fbatch for batch > LRU statistics processing. > > A folio may also have extra refcounts if it is on the mlock fbatch. > > These two known "usages" of the folio are handled by calling > lru_cache_drain_for_folio, which drains both the lru_add and mlock > fbatches. After draining, if the refcount is still elevated, then there are > truly outstanding references. > > If the page may be dma pinned, DMA is using it and hence there are > outstanding references. folio_maybe_dma_pinned() can have false positives, > but that's only with a significant number of refcounts, at which point > draining LRU is not going to move the needle - it can still be concluded > that the folio has outstanding references. > > If the page is still mapped after guest_memfd tried to unmap it earlier in > the conversion process, it also has outstanding references. > > Return true and exit early to avoid unnecessary draining in these 2 cases. > > Provide a drain status to only drain once ever while processing a batch of > folios. > > Acked-by: Vlastimil Babka (SUSE) > Suggested-by: David Hildenbrand > Reviewed-by: Fuad Tabba > Reviewed-by: Binbin Wu > Signed-off-by: Ackerley Tng > --- > mm/folio.c | 2 ++ > virt/kvm/guest_memfd.c | 30 ++++++++++++++++++++++-------- > 2 files changed, 24 insertions(+), 8 deletions(-) > > diff --git a/mm/folio.c b/mm/folio.c > index c02dcea9c03c2..50a6dbe55998e 100644 > --- a/mm/folio.c > +++ b/mm/folio.c > @@ -33,6 +33,7 @@ > #include > #include > #include > +#include > > #include "internal.h" > #include "page_alloc.h" > @@ -926,6 +927,7 @@ void lru_cache_drain_for_folio(const struct folio *folio, > *drained = LRU_CACHE_DRAINED_ALL; > } > } > +EXPORT_SYMBOL_FOR_KVM(lru_cache_drain_for_folio); > > atomic_t lru_disable_count = ATOMIC_INIT(0); > I don't mind about the virt/kvm/guest_memfd.c part of it, but I'm finding a KVM patchset modifying mm/folio.c there hard to deal with: and notice Sean also suggesting to separate this part out. As you know, I've worked up a patchset "mm/fbatch: drain lru_add_drain() and _all()" which finally removes the problem lru_cache_drain_for_folio() works around. In the initial version posted a week ago, there was no lru_cache_drain_for_folio() in the tree. Now 7.3-rc1 has it, so I intended a replacement 13/25 in my series, giving you just an empty inline lru_cache_drain_for_folio() stub (and enum lru_cache_drained) in linux/swap.h. But that won't work for you, if you're adding an EXPORT_SYMBOL_FOR_KVM() in mm/folio.c, and of course conflicts with my removals (in context both above and below your EXPORT line). It's easy for me to remove what's in mm/gup.c and mm/folio.c, but I cannot remove what is not yet there. I've wasted hours on this, hoping not to trouble either of you; but seeing now that I shall have to rebase anyway (an unrelated mlock fix), I'm electing to take the only clean way out: I'm going to submit this mm/folio.c part of your patch to Andrew tonight (with a shorter Cc list!), in the hope that it can be accelerated into 7.3-rc2 (or at least get an mm-stable stable base-commit id) which we can both work off independently. Whether that's acceptable to Ackerley and to Andrew, I don't know (just as we don't know when either of our patchsets will go further), but let me try. Thanks, Hugh