* + mm-vmscan-avoid-anon-scanning-for-gfp_noio-with-low-swapcache.patch added to mm-new branch
@ 2026-09-09 1:28 Andrew Morton
0 siblings, 0 replies; only message in thread
From: Andrew Morton @ 2026-09-09 1:28 UTC (permalink / raw)
To: mm-commits, zhangbo56, shakeel.butt, mhocko, ljs, kasong, hannes,
david, baohua, zhangbo0325, akpm
The patch titled
Subject: mm: vmscan: avoid anon scanning for GFP_NOIO with low swapcache
has been added to the -mm mm-new branch. Its filename is
mm-vmscan-avoid-anon-scanning-for-gfp_noio-with-low-swapcache.patch
This patch will shortly appear at
https://git.kernel.org/pub/scm/linux/kernel/git/akpm/25-new.git/tree/patches/mm-vmscan-avoid-anon-scanning-for-gfp_noio-with-low-swapcache.patch
This patch will later appear in the mm-new branch at
git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm
Note, mm-new is a provisional staging ground for work-in-progress
patches, and acceptance into mm-new is a notification for others take
notice and to finish up reviews. Please do not hesitate to respond to
review feedback and post updated versions to replace or incrementally
fixup patches in mm-new.
The mm-new branch of mm.git is not included in linux-next
If a few days of testing in mm-new is successful, the patch will me moved
into mm.git's mm-unstable branch, which is included in linux-next
Before you just go and hit "reply", please:
a) Consider who else should be cc'ed
b) Prefer to cc a suitable mailing list as well
c) Ideally: find the original patch on the mailing list and do a
reply-to-all to that, adding suitable additional cc's
*** Remember to use Documentation/process/submit-checklist.rst when testing your code ***
The -mm tree is included into linux-next via various
branches at git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm
and is updated there most days
------------------------------------------------------
From: Bo Zhang <zhangbo0325@gmail.com>
Subject: mm: vmscan: avoid anon scanning for GFP_NOIO with low swapcache
Date: Tue, 8 Sep 2026 14:26:49 +0800
We have observed some cases where memory is allocated with GFP_NOIO, so we
cannot reclaim any anon folios unless they are in swapcache. We can end
up spending more than 150 ms looping in `shrink_folio_list()` scanning
non-swapcache folios without reclaiming a single folio. This is pure
overhead.
This is particularly true on systems using zRAM, where swapcache is
relatively rare. So let's check whether anon reclaim is allowed by GFP_IO
and whether there is enough swapcache to make it worthwhile. If the
swapcache is extremely low, we're essentially searching for a needle in a
haystack, so let's avoid scanning anon in the first place.
On Android this is triggered by dm-verity hash-block reads through
dm-bufio, which legitimately use GFP_NOIO because they run underneath the
IO path:
verity_verify_io -> verity_hash_for_block -> verity_verify_level
-> dm_bufio_read_with_ioprio -> new_read -> __bufio_new
-> alloc_buffer
gfp: GFP_NOIO | __GFP_NORETRY | __GFP_NOMEMALLOC | __GFP_NOWARN
Such a reclaimer can land on a memcg with a large, unswapped anon LRU and
a tiny file LRU (e.g. inactive_anon ~335 MB vs inactive_file ~4 MB, with
negligible swapcache). shrink_lruvec() then keeps feeding that huge anon
list into shrink_folio_list() - ~2400 shrink_folio_list() calls, ~93,000
anon folios scanned - where every folio is kept because it needs IO. The
150+ ms above is one such single shrink_lruvec() pass (not accumulated
across a reclaim cycle), and it reclaims nothing; the actual progress
comes entirely from the file side.
Aging anon alongside file does have some value for a later __GFP_IO
reclaimer, so it is not strictly pure overhead. But that aging is only
deferred, not lost: kswapd and other __GFP_IO reclaimers still walk and
age anon. Spending ~168 ms aging memory that this context cannot reclaim
is not a worthwhile trade-off in a latency-sensitive path.
To stay conservative, this only skips anon when the swapcache is really
tiny - below 1/64 of the anon LRU - i.e. when essentially no anon on the
list can be reclaimed without IO. Whenever there is a meaningful amount
of swapcached anon, the normal path is used and anon is scanned and aged
as before.
Note this only addresses the traditional active/inactive LRU. MGLRU
selects anon vs file scanning in its own path and is not covered here;
fixing the MGLRU case is left as a TODO.
Link: https://lore.kernel.org/20260908062649.1045883-1-zhangbo56@xiaomi.com
Signed-off-by: Bo Zhang <zhangbo56@xiaomi.com>
Reviewed-by: Barry Song <baohua@kernel.org>
Cc: David Hildenbrand <david@kernel.org>
Cc: Johannes Weiner <hannes@cmpxchg.org>
Cc: Kairui Song <kasong@tencent.com>
Cc: Lorenzo Stoakes <ljs@kernel.org>
Cc: Michal Hocko <mhocko@kernel.org>
Cc: Shakeel Butt <shakeel.butt@linux.dev>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
---
mm/vmscan.c | 62 +++++++++++++++++++++++++++++++++++++++++++++-----
1 file changed, 57 insertions(+), 5 deletions(-)
--- a/mm/vmscan.c~mm-vmscan-avoid-anon-scanning-for-gfp_noio-with-low-swapcache
+++ a/mm/vmscan.c
@@ -362,20 +362,72 @@ static bool can_demote(int nid, struct s
return !nodes_empty(allowed_mask);
}
+#ifdef CONFIG_SWAP
+static inline bool reclaimable_anon_is_low(struct mem_cgroup *memcg,
+ int nid, struct scan_control *sc)
+{
+ pg_data_t *pgdat = NODE_DATA(nid);
+ unsigned long anon_pages, swapcache;
+
+ /*
+ * A GFP_NOIO reclaimer can only reclaim anon that is already in the
+ * swapcache (adding anon to the swapcache needs IO). When swapcache is
+ * far below the anon LRU, scanning anon reclaims nothing and only burns
+ * CPU. The 1/64 threshold keeps this to the case where anon is
+ * effectively unreclaimable.
+ */
+ if (!sc || (sc->gfp_mask & __GFP_IO))
+ return false;
+
+ /*
+ * FIXME: MGLRU doesn't fully respect can_reclaim_anon_pages() for the
+ * scanning type, so only apply this to the traditional LRU for now.
+ */
+ if (lru_gen_enabled())
+ return false;
+
+ if (memcg) {
+ struct lruvec *lruvec = mem_cgroup_lruvec(memcg, pgdat);
+
+ anon_pages = lruvec_page_state(lruvec, NR_INACTIVE_ANON) +
+ lruvec_page_state(lruvec, NR_ACTIVE_ANON);
+ swapcache = lruvec_page_state(lruvec, NR_SWAPCACHE);
+ } else {
+ anon_pages = node_page_state(pgdat, NR_INACTIVE_ANON) +
+ node_page_state(pgdat, NR_ACTIVE_ANON);
+ swapcache = node_page_state(pgdat, NR_SWAPCACHE);
+ }
+
+ return swapcache < (anon_pages >> 6);
+}
+#else
+static inline bool reclaimable_anon_is_low(struct mem_cgroup *memcg,
+ int nid, struct scan_control *sc)
+{
+ return true;
+}
+#endif /* CONFIG_SWAP */
+
static inline bool can_reclaim_anon_pages(struct mem_cgroup *memcg,
int nid,
struct scan_control *sc)
{
if (memcg == NULL) {
/*
- * For non-memcg reclaim, is there
- * space in any swap device?
+ * For non-memcg reclaim, is there space in any swap device?
+ * And under GFP_NOIO, is there enough swapcached anon to make
+ * scanning anon worthwhile?
*/
- if (get_nr_swap_pages() > 0)
+ if (get_nr_swap_pages() > 0 &&
+ !reclaimable_anon_is_low(memcg, nid, sc))
return true;
} else {
- /* Is the memcg below its swap limit? */
- if (mem_cgroup_get_nr_swap_pages(memcg) > 0)
+ /*
+ * Is the memcg below its swap limit, and under GFP_NOIO does
+ * it have enough swapcached anon to make scanning worthwhile?
+ */
+ if (mem_cgroup_get_nr_swap_pages(memcg) > 0 &&
+ !reclaimable_anon_is_low(memcg, nid, sc))
return true;
}
_
Patches currently in -mm which might be from zhangbo0325@gmail.com are
mm-vmscan-avoid-anon-scanning-for-gfp_noio-with-low-swapcache.patch
^ permalink raw reply [flat|nested] only message in thread
only message in thread, other threads:[~2026-09-09 1:28 UTC | newest]
Thread overview: (only message) (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-09 1:28 + mm-vmscan-avoid-anon-scanning-for-gfp_noio-with-low-swapcache.patch added to mm-new branch Andrew Morton
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.