From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-yx2-f12.google.com (mail-yx2-f12.google.com [74.125.224.140]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 217AA3AEB2C for ; Wed, 9 Sep 2026 09:49:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.224.140 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788947365; cv=none; b=GNEDSYeRqhq8htGEYMCH+85WzGNzKYiVJBXbbvCzgiaRtKvGKhnJLn54j5g5HH031DIx58EZcg8HbFQEm0Ul0Et4yUj6B0GxabmT5Bki86w8GMPWS10Sb4KC5skmxPVUl10Ie0AkjWPvGsgOUHZZBdz1Csbe5D71X4hQw2b8EEk= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788947365; c=relaxed/simple; bh=bmiYMnfxqxqXTndNn9OHugj8rDer4W9a5xt6AtmTM94=; h=Date:From:To:cc:Subject:In-Reply-To:Message-ID:References: MIME-Version:Content-Type; b=fnqDtnXUCglnaPj06vXPasjzrE7VP5ObSp1hE1dXbJHwvdxxULOfVxA2sV7XnS2QqvadGz/9JpKK9C7NSlLxjs3ZX5y+eidE5Xmxhiop0GUwzLKmAxIrT8eSkhykW9ePHn2ADECh1hZwTT/sAUaFdIDUYJNIakNHMLi+Gr55t/Q= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=Xy1CfrsA; arc=none smtp.client-ip=74.125.224.140 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="Xy1CfrsA" Received: by mail-yx2-f12.google.com with SMTP id 00721157ae682-85d46e4cdcdso9307217b3.2 for ; Wed, 09 Sep 2026 02:49:23 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1788947362; x=1789552162; darn=vger.kernel.org; h=content-type:mime-version:references:message-id:in-reply-to:subject :cc:to:from:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=UufK0vMlC6Z+KADxCs5ZoL9LcMsl7R0DiwRuRItb70g=; b=Xy1CfrsA3ykgtEBzxr7he0QYrrmY3qjtfiMjhWVfHWvD82Lmk495w9h/A7UoEC7EzO eXtlFRoWGRO+0ipdKtnQhjzpxrzCObInhbqjlZX0+cuIUkWzJ+btylZsFb17ElAEFbik +ubFA4B6r3+ze2+ybMtgE1NZJdL7C6dlgkpDkXOfSWqSw/p0uJ4OeRFX7fPPlX0PUegI G2NVTzmc+E7kIyPHIw4GIAqIIEH524CGYYb0yO8ZP8ggJQKyznkO1d9hFGkUkYyULa3N w5CidnpQp/TmCwcvnk50txFpDrhOeYS49TZYOhF2wDUDLiwtFUl7lI1N6ez09IQRfpRu 7bFA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788947362; x=1789552162; h=content-type:mime-version:references:message-id:in-reply-to:subject :cc:to:from:date:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=UufK0vMlC6Z+KADxCs5ZoL9LcMsl7R0DiwRuRItb70g=; b=C8mJA/jiLvGUolndvoizBw8q7Sf0zV9w6vAudBxqCwhhXCFUoedEmb9WvzsaXGVXaj uNInCvF7C7ge8fv1YmpLoC97MQ3O6zyWlhknXzZ3VaZJ4+VA9+Vg3cDW1UrRa0bpQ6BC s/jYFSkd4G7/99cQjMU1hkmMDVfLfqVgBVstyn4MYP0n+efTLU/bms7BXzFsrkRXVpRy yuIEJWIL3RDOasG9eYbAgszoJ0H7X019Af9utzTjQzkSvZRICvyTpVJSkprLFxK7i7/9 j9C25rYB8H0CqJ/rJYs58nATrFCvwqpYXRTXnWTaeAaBiUkKzj4k2WK1PO2nW4PsWGg4 ncNg== X-Forwarded-Encrypted: i=1; AKwUvBzFxzbtCy3Il3gKpEyDGyKh7D12QUllRP38Ic/0MUiDBwe0Dyevhn1wmJLmmFlivg/YCGT8irXraaIOcPkP@vger.kernel.org X-Gm-Message-State: AFuF++k+WVDKosZsr7Od0xWsaVr0PFyb2+8ea+fCa2KhwzYjHOMXaDla vgaTdLlq47fm6+sgHo35NFCk3t03ki1fnBeYHEPMFSWbJoxhoI+TFFWo9UrABrDEDA== X-Gm-Gg: AYBFou2q+WuV0p+tdMfIjy/0vvqT9QrlRZerMU4dx8pgL/FOT0SUiWvIJ6o76OM4o42 OOu8+nk+yRy4kpBnXwun3c1FAF9CcJf9a261FbEwAhmNuV6k3UJgVQ1U0w27rcUTDqd4E+nQ/2I /g8aHzpKO6sN4VzKOxfxcJn0TUOAWxO5NDGOOETPG9RhQw2VpaYINDetHp8fSTb2CE4SfnoqXIV 3hbgjWvWYuMw5oA1EDVG5d3oSbms27xOut9TtvjRTEdhLZdhTMLXa8VSZL8hQjPFLkvkEGmU7GG Y0iHevxt2Q7zIMFm/GG2ikeHUUTbWBWdNEfTHI9PYkzlyJqYf5AV84rOvatI/kRiT1e/wGm6Ky/ eTQ70zHCHoAQyTmV1mDjvPKV6IJUoNRpPGX6DBYAy0HeRJBoWCfykfaxU0LG3Ez5wiRad8nU0Fn Ga4x0dQae4t3UHEF5BTLnmxPslNTFDcXam23RNvRKmzw8x1f+0cBt0+KKKb3JfgfMd/dKvZyo88 qDhVYzVJVshgDdVaEUOkQqbUC5hGZLnu2+pFljKrkSkJLd/Z9bORw+bqtg= X-Received: by 2002:a05:690c:e3f2:b0:87c:35a5:4aea with SMTP id 00721157ae682-87f22b7eaf5mr23815837b3.6.1788947361454; Wed, 09 Sep 2026 02:49:21 -0700 (PDT) Received: from darker.attlocal.net (172-10-233-147.lightspeed.sntcca.sbcglobal.net. [172.10.233.147]) by smtp.gmail.com with ESMTPSA id 00721157ae682-87149600f48sm107580907b3.13.2026.09.09.02.49.16 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 09 Sep 2026 02:49:20 -0700 (PDT) Date: Wed, 9 Sep 2026 02:49:15 -0700 (PDT) From: Hugh Dickins To: Andrew Morton cc: Ackerley Tng , Alexander Viro , Alexandre Ghiti , Baolin Wang , Barry Song , Binbin Wu , Christian Brauner , Christoph Hellwig , Christoph Lameter , Claudio Imbrenda , David Hildenbrand , JP Kobryn , Jan Kara , Jens Axboe , Johannes Weiner , Kairui Song , Kiryl Shutsemau , Lance Yang , Leonardo Bras , Lorenzo Stoakes , Marcelo Tosatti , Matthew Wilcox , Mel Gorman , Miaohe Lin , Michal Hocko , Minchan Kim , Muchun Song , Oscar Salvador , Peter Zijlstra , Qi Zheng , Rik van Riel , Sebastian Andrzej Siewior , Shakeel Butt , Suren Baghdasaryan , Vlastimil Babka , Yang Shi , Yu Zhao , Zach O'Keefe , Zi Yan , linux-block@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org Subject: [PATCH v2 04/26] mm/fbatch: lru bit set, no extra ref, while folio on per-cpu fbatch In-Reply-To: Message-ID: <61e15506-940f-3532-5fc9-4086f7612c10@google.com> References: Precedence: bulk X-Mailing-List: linux-fsdevel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=US-ASCII Treat folios on a per-cpu fbatch as if they were already on the lruvec: with PG_lru set, without holding an extra reference. This will enable the removal of most lru_add_drain() and lru_add_drain_all() calls soon. Recognize such a folio by 0x02 set in the folio->lru.next pointer by folio_add_lru(). Then lruvec_del_folio() (aided by "lru_add_del_folio") can pretend to unlink it, and folio_batch_move_lru()'s lru_add case can check whether one of the others has already moved it to lruvec. (That bit is also used in a transient way by set_page_pfmemalloc(), to inform interested callers whether page_is_pfmemalloc(): but those callers are in networking, not putting folios on LRU; and accept that any use of the page->lru field already erases page_is_pfmemalloc() information.) Let folio->lru.next point to the lru_add fbatch entry, but this is now just for debugging: it seemed to be important for folio_batch_move_lru() to distinguish fresh from stale entries, but then it turned out that it has to processs them identically. Activate, deactivates and move_tail, holding no reference on the folio, might come to act on a stale folio when the fbatch is drained: but it's acquired by try_get and test_clear_lru, so safe even when suboptimal. Reclaim is not an exact science, and there have been no complaints of missed actions since 5.11 commit fc574c23558c ("mm/swap.c: serialize memcg changes in pagevec_lru_move_fn") introduced the TestClearPageLRU protocol: so don't expect complaints of a few surprisingly taken actions. Signed-off-by: Hugh Dickins --- include/linux/mm_inline.h | 25 ++++++++ include/linux/mm_types.h | 6 +- mm/folio.c | 119 +++++++++++++------------------------- mm/huge_memory.c | 6 +- 4 files changed, 74 insertions(+), 82 deletions(-) diff --git a/include/linux/mm_inline.h b/include/linux/mm_inline.h index 621c8653d8f7..8420b1276535 100644 --- a/include/linux/mm_inline.h +++ b/include/linux/mm_inline.h @@ -343,6 +343,29 @@ static inline void folio_migrate_refs(struct folio *new, const struct folio *old } #endif /* CONFIG_LRU_GEN */ +enum { + LRU_NEXT_NEVER_TAIL = 0, /* Used by a tail's compound_head */ + LRU_NEXT_BATCHED = 1, /* Not used by any aligned pointer */ + NR_LRU_NEXT_FLAGS +}; + +static __always_inline +bool lru_add_del_folio(struct folio *folio) +{ + unsigned long lru_next = READ_ONCE(folio->lru_next); + + /* BUG_ON(folio_test_lru(folio) && folio_ref_count(folio)); */ + if (!(lru_next & BIT(LRU_NEXT_BATCHED))) + return false; + + WRITE_ONCE(folio->lru.next, LIST_POISON1); + /* BUG_ON(folio->lru_next & BIT(LRU_NEXT_BATCHED)); */ + + /* Ensure folio->lru_next visible when folio_set_lru() called later */ + smp_mb__before_atomic(); + return true; +} + static __always_inline void lruvec_add_folio(struct lruvec *lruvec, struct folio *folio) { @@ -384,6 +407,8 @@ void lruvec_del_folio(struct lruvec *lruvec, struct folio *folio) if (lru_gen_del_folio(lruvec, folio, false)) return; + if (lru_add_del_folio(folio)) + return; if (lru != LRU_UNEVICTABLE) list_del(&folio->lru); diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h index 6d815f6440c9..fe6220b97cf3 100644 --- a/include/linux/mm_types.h +++ b/include/linux/mm_types.h @@ -85,6 +85,8 @@ struct page { * WARNING: bit 0 of the first word is used for PageTail(). That * means the other users of this union MUST NOT use the bit to * avoid collision and false-positive PageTail(). + * Bit 1 of the first word is used by page_is_pfmemalloc(). + * Bit 1 of the first word (lru_next) is also used by folio_add_lru(). */ union { struct { /* Page cache and anonymous pages */ @@ -410,10 +412,8 @@ struct folio { union { struct list_head lru; /* private: avoid cluttering the output */ - /* For the Unevictable "LRU list" slot */ struct { - /* Avoid compound_info */ - void *__filler; + unsigned long lru_next; /* public: */ unsigned int mlock_count; /* private: */ diff --git a/mm/folio.c b/mm/folio.c index b9dc4f5e10a6..e743cd539b9e 100644 --- a/mm/folio.c +++ b/mm/folio.c @@ -152,57 +152,34 @@ static void folio_batch_move_lru(struct folio_batch *fbatch, move_fn_t move_fn) int i; struct lruvec *lruvec = NULL; unsigned long flags = 0; - struct folio_batch free_fbatch; - bool is_lru_add = (move_fn == lru_add); - - /* - * If we're adding to the LRU, preemptively filter dead folios. Use - * this dedicated folio batch for temp storage and deferred cleanup. - */ - if (is_lru_add) - folio_batch_init(&free_fbatch); for (i = 0; i < folio_batch_count(fbatch); i++) { struct folio *folio = fbatch->folios[i]; - /* block memcg migration while the folio moves between lru */ - if (!is_lru_add && !folio_test_clear_lru(folio)) - continue; - - /* - * Filter dead folios by moving them from the add batch to the temp - * batch for freeing after this loop. - * - * We're bypassing normal cleanup. Clear flags that are not - * applicable to dead folios. - * - * Since the folio may be part of a huge page, unqueue from - * deferred split list to avoid a dangling list entry. - */ - if (is_lru_add && folio_ref_freeze(folio, 1)) { - __folio_clear_active(folio); - __folio_clear_unevictable(folio); - folio_unqueue_deferred_split(folio); + if (!folio_try_get(folio)) { fbatch->folios[i] = NULL; - folio_batch_add(&free_fbatch, folio); continue; } + if (!folio_test_clear_lru(folio)) + continue; + + /* Do not add to LRU if it has already been added */ + if (move_fn == lru_add && !lru_add_del_folio(folio)) + goto restore_lru; + folio_lruvec_relock_irqsave(folio, &lruvec, &flags); move_fn(lruvec, folio); + /* Do add to LRU if not already there (move_fn skipped) */ + if (lru_add_del_folio(folio)) + lruvec_add_folio(lruvec, folio); +restore_lru: folio_set_lru(folio); } if (lruvec) lruvec_unlock_irqrestore(lruvec, flags); - - /* Cleanup filtered dead folios. */ - if (is_lru_add) { - mem_cgroup_uncharge_folios(&free_fbatch); - free_unref_folios(&free_fbatch); - } - folios_put(fbatch); } @@ -211,8 +188,6 @@ static void __folio_batch_add_and_move(struct folio_batch __percpu *fbatch, { unsigned long flags; - folio_get(folio); - if (disable_irq) local_lock_irqsave(&cpu_fbatches.lock_irq, flags); else @@ -273,7 +248,6 @@ static void lru_activate(struct lruvec *lruvec, struct folio *folio) if (folio_test_active(folio) || folio_test_unevictable(folio)) return; - lruvec_del_folio(lruvec, folio); folio_set_active(folio); lruvec_add_folio(lruvec, folio); @@ -289,37 +263,12 @@ void folio_activate(struct folio *folio) !folio_test_lru(folio)) return; - folio_batch_add_and_move(folio, lru_activate); -} - -static void __lru_cache_activate_folio(struct folio *folio) -{ - struct folio_batch *fbatch; - int i; - - local_lock(&cpu_fbatches.lock); - fbatch = this_cpu_ptr(&cpu_fbatches.lru_add); - /* - * Search backwards on the optimistic assumption that the folio being - * activated has just been added to this batch. Note that only - * the local batch is examined as a !LRU folio could be in the - * process of being released, reclaimed, migrated or on a remote - * batch that is currently being drained. Furthermore, marking - * a remote batch's folio active potentially hits a race where - * a folio is marked active just after it is added to the inactive - * list causing accounting errors and BUG_ON checks to trigger. + * XXX: It is curiously difficult to recreate safely the old + * __lru_cache_activate_folio() optimization (folio_set_active() + * directly if it's on the local lru_add fbatch): revisit later. */ - for (i = folio_batch_count(fbatch) - 1; i >= 0; i--) { - struct folio *batch_folio = fbatch->folios[i]; - - if (batch_folio == folio) { - folio_set_active(folio); - break; - } - } - - local_unlock(&cpu_fbatches.lock); + folio_batch_add_and_move(folio, lru_activate); } #ifdef CONFIG_LRU_GEN @@ -410,16 +359,7 @@ void folio_mark_accessed(struct folio *folio) * unevictable page accessed has no effect. */ } else if (!folio_test_active(folio)) { - /* - * If the folio is on the LRU, queue it for activation via - * cpu_fbatches.lru_activate. Otherwise, assume the folio is in a - * folio_batch, mark it active and it'll be moved to the active - * LRU on the next drain. - */ - if (folio_test_lru(folio)) - folio_activate(folio); - else - __lru_cache_activate_folio(folio); + folio_activate(folio); folio_clear_referenced(folio); workingset_activation(folio); } @@ -439,6 +379,10 @@ EXPORT_SYMBOL(folio_mark_accessed); */ void folio_add_lru(struct folio *folio) { + struct folio_batch *fbatch; + unsigned long lru_next; + bool full; + VM_BUG_ON_FOLIO(folio_test_active(folio) && folio_test_unevictable(folio), folio); VM_BUG_ON_FOLIO(folio_test_lru(folio), folio); @@ -458,7 +402,26 @@ void folio_add_lru(struct folio *folio) folio_mark_accessed(folio); } - folio_batch_add_and_move(folio, lru_add); + local_lock(&cpu_fbatches.lock); + fbatch = this_cpu_ptr(&cpu_fbatches.lru_add); + + /* Storing this address is only for debugging */ + lru_next = (unsigned long)&fbatch->folios[fbatch->nr]; + /* This mask will do nothing on 64-bit */ + lru_next &= ~(BIT(NR_LRU_NEXT_FLAGS) - 1); + lru_next |= BIT(LRU_NEXT_BATCHED); + folio->lru_next = lru_next; + + full = !folio_batch_add(fbatch, folio); + + /* Ensure folio->lru_next visible to folio_test_clear_lru() callers */ + smp_mb__before_atomic(); + folio_set_lru(folio); + + if (full || !folio_may_be_lru_cached(folio) || lru_cache_disabled()) + folio_batch_move_lru(fbatch, lru_add); + + local_unlock(&cpu_fbatches.lock); } EXPORT_SYMBOL(folio_add_lru); diff --git a/mm/huge_memory.c b/mm/huge_memory.c index afbb5974bd22..c7510d875433 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -3995,8 +3995,12 @@ static int __folio_freeze_and_split_unmapped(struct folio *folio, unsigned int n } /* lock lru list/PageCompound, ref frozen by page_ref_freeze */ - if (do_lru) + if (do_lru) { lruvec = folio_lruvec_lock(folio); + /* Move from fbatch to lruvec before lru_add_split_folio()s */ + if (lru_add_del_folio(folio)) + lruvec_add_folio(lruvec, folio); + } ret = __split_unmapped_folio(folio, new_order, split_at, xas, mapping, split_type); -- 2.51.0