From: Usama Arif <usama.arif@linux.dev>
To: Usama Arif <usama.arif@linux.dev>
Cc: Andrew Morton <akpm@linux-foundation.org>,
david@kernel.org, ljs@kernel.org, liam@infradead.org,
vbabka@kernel.org, rppt@kernel.org, surenb@google.com,
mhocko@suse.com, kasong@tencent.com, qi.zheng@linux.dev,
shakeel.butt@linux.dev, axelrasmussen@google.com,
yuanchu@google.com, weixugc@google.com, chrisl@kernel.org,
nphamcs@gmail.com, baoquan.he@linux.dev, youngjun.park@lge.com,
hannes@cmpxchg.org, roman.gushchin@linux.dev,
muchun.song@linux.dev, linux-mm@kvack.org,
linux-kernel@vger.kernel.org, cgroups@vger.kernel.org,
rientjes@google.com, kernel-team@meta.com
Subject: Re: [PATCH v5 3/3] mm/vmscan: reduce lru_lock contention via vmstat-derived scan-balance cost
Date: Wed, 29 Jul 2026 06:02:39 -0700 [thread overview]
Message-ID: <20260729130241.3327679-1-usama.arif@linux.dev> (raw)
In-Reply-To: <20260727162550.2032-4-usama.arif@linux.dev>
On Mon, 27 Jul 2026 09:23:25 -0700 Usama Arif <usama.arif@linux.dev> wrote:
> The anon/file scan balance in get_scan_count() is driven by two scalars
> in struct lruvec, anon_cost and file_cost, accumulated by every reclaim
> producer under lruvec->lru_lock. The acquisition sites for cost work
> specifically are:
>
> - shrink_inactive_list() re-takes lru_lock at function exit purely
> to call lru_note_cost_unlock_irq() with (nr_pageout, nr_scanned -
> nr_reclaimed). One acquisition per inactive shrink.
> - shrink_active_list() does the same with (0, nr_rotated). One
> acquisition per active shrink.
> - workingset_refault() takes the lock via folio_lruvec_lock_irq()
> purely to record the refault cost. One acquisition per refault.
> - prepare_scan_control() takes lru_lock just to snapshot the two
> scalars into sc->{anon,file}_cost.
> - lru_note_cost_unlock_irq() itself walks parent_lruvec and
> re-acquires lru_lock on each ancestor to propagate the update,
> adding O(memcg-depth) acquisitions per producer call.
>
> This hurts because lru_lock is already a heavy contention point on
> memory-heavy workloads: every isolate_lru_folios(), move_folios_to_lru()
> and folio_add_lru() takes it. The cost work itself is trivial (two
> scalar bumps and one comparison), but it contends with and causes
> contention for actual LRU manipulation. The parent_lruvec() walk also
> multiplies cost-update overhead by memcg hierarchy depth.
>
> The balance formula for anon and file, respectively, is this:
>
> cost = nr_io * SWAP_CLUSTER_MAX + nr_rotated
>
> Instead of recording cost and running averaging logic directly when
> these events occur, snapshot running vmstat counters once per reclaim
> cycle and derive the balance from event deltas since the last run.
>
> Use PGROTATE_* from the preceding patch for the rotation input.
> WORKINGSET_RESTORE_* and NR_VMSCAN_WRITE provide the remaining event
> counters. Charge NR_VMSCAN_WRITE through lruvec stats so all inputs can
> be sampled per lruvec and aggregated through the memcg hierarchy. This
> is overall cheaper and has fewer lock acquisition sites.
>
> Moving accumulation and decay to the reclaim side also improves the cost
> model across reclaim gaps. With producer-side decay, events that happen
> while reclaim is idle still age each other before reclaim ever samples
> the costs. If a workload refaults a large anon set and then a smaller
> file set before reclaim runs again, the later file activity can age the
> earlier anon activity out of the cost model. The new scheme observes the
> whole between-reclaim delta and decays anon and file proportionally, so
> the scan-balance history better represents what happened since the last
> reclaim pass.
>
> A dedicated per-lruvec spinlock, cost_lock, serialises the delta
> extraction, the cost->count update and the halving loop against
> concurrent reclaimers in the same memcg+node.
>
> NR_VMSCAN_WRITE is accounted at writeout(), so reclaim_stat.nr_pageout is
> no longer needed and is removed.
>
> memcg-v1's memory.stat anon_cost/file_cost is now sourced from
> cost[].count instead of the removed lruvec anon_cost/file_cost fields.
> The reported values only refresh when prepare_scan_control() runs and
> are bounded at ~lrusize/4 by the halving loop; the scan-balance signal
> they express is unchanged.
>
> Under pure MGLRU the scan-balance signal itself is not consumed (both
> prepare_scan_control() and get_scan_count() are short-circuited on the
> MGLRU paths, and MGLRU's own type/tier selection comes from read_ctrl_pos()
> on lrugen->{avg_refaulted,avg_total,refaulted,evicted}, not from
> anon_cost/file_cost). NR_VMSCAN_WRITE naturally covers writeout from
> either reclaim implementation. The preceding patch also bumps
> PGROTATE_{ANON,FILE} from evict_folios(), so rotation-driven reclaim
> work is accounted consistently across both implementations.
>
> Acked-by: Shakeel Butt <shakeel.butt@linux.dev>
> Acked-by: Johannes Weiner <hannes@cmpxchg.org>
> Signed-off-by: Usama Arif <usama.arif@linux.dev>
> ---
> include/linux/mmzone.h | 13 +++++--
> include/linux/swap.h | 3 --
> include/linux/vmstat.h | 1 -
> mm/memcontrol-v1.c | 4 +--
> mm/memcontrol.c | 1 +
> mm/mmzone.c | 1 +
> mm/swap.c | 69 ------------------------------------
> mm/vmscan.c | 79 +++++++++++++++++++++++++++++++++++-------
> mm/workingset.c | 5 ---
> 9 files changed, 81 insertions(+), 95 deletions(-)
>
> diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h
> index aab06fb6c6d5..85303c5867c8 100644
> --- a/include/linux/mmzone.h
> +++ b/include/linux/mmzone.h
> @@ -757,6 +757,12 @@ void lru_gen_reparent_memcg(struct mem_cgroup *memcg, struct mem_cgroup *parent,
>
> #endif /* CONFIG_LRU_GEN */
>
> +struct lru_cost {
> + unsigned long count;
> + unsigned long last_rotated;
> + unsigned long last_io;
> +};
> +
> struct lruvec {
> struct list_head lists[NR_LRU_LISTS];
> /* per lruvec lru_lock for memcg */
> @@ -765,9 +771,12 @@ struct lruvec {
> * These track the cost of reclaiming one LRU - file or anon -
> * over the other. As the observed cost of reclaiming one LRU
> * increases, the reclaim scan balance tips toward the other.
> + * Updated and decayed at prepare_scan_control() time; cost_lock
> + * serialises that update.
> */
> - unsigned long anon_cost;
> - unsigned long file_cost;
> + struct lru_cost cost[ANON_AND_FILE];
> + /* Protects cost[]. */
> + spinlock_t cost_lock;
> /* Non-resident age, driven by LRU movement */
> atomic_long_t nonresident_age;
> /* Refaults at the time of last reclaim cycle */
> diff --git a/include/linux/swap.h b/include/linux/swap.h
> index 6d72778e6cc3..d35a4761ebd7 100644
> --- a/include/linux/swap.h
> +++ b/include/linux/swap.h
> @@ -309,9 +309,6 @@ extern unsigned long totalreserve_pages;
>
>
> /* linux/mm/swap.c */
> -void lru_note_cost_unlock_irq(struct lruvec *lruvec, bool file,
> - unsigned int nr_io, unsigned int nr_rotated);
> -void lru_note_cost_refault(struct folio *);
> void folio_add_lru(struct folio *);
> void folio_add_lru_vma(struct folio *, struct vm_area_struct *);
> void mark_page_accessed(struct page *);
> diff --git a/include/linux/vmstat.h b/include/linux/vmstat.h
> index fb8c76289e02..5b31d8e7ae40 100644
> --- a/include/linux/vmstat.h
> +++ b/include/linux/vmstat.h
> @@ -20,7 +20,6 @@ struct reclaim_stat {
> unsigned nr_congested;
> unsigned nr_writeback;
> unsigned nr_immediate;
> - unsigned nr_pageout;
> unsigned nr_activate[ANON_AND_FILE];
> unsigned nr_ref_keep;
> unsigned nr_unmap_fail;
> diff --git a/mm/memcontrol-v1.c b/mm/memcontrol-v1.c
> index 765069211567..091bc9ffee44 100644
> --- a/mm/memcontrol-v1.c
> +++ b/mm/memcontrol-v1.c
> @@ -1988,8 +1988,8 @@ void memcg1_stat_format(struct mem_cgroup *memcg, struct seq_buf *s)
> for_each_online_pgdat(pgdat) {
> mz = memcg->nodeinfo[pgdat->node_id];
>
> - anon_cost += mz->lruvec.anon_cost;
> - file_cost += mz->lruvec.file_cost;
> + anon_cost += mz->lruvec.cost[WORKINGSET_ANON].count;
> + file_cost += mz->lruvec.cost[WORKINGSET_FILE].count;
> }
> seq_buf_printf(s, "anon_cost %lu\n", anon_cost);
> seq_buf_printf(s, "file_cost %lu\n", file_cost);
> diff --git a/mm/memcontrol.c b/mm/memcontrol.c
> index 23adb698dadd..42f4351c69cc 100644
> --- a/mm/memcontrol.c
> +++ b/mm/memcontrol.c
> @@ -393,6 +393,7 @@ static const unsigned int memcg_node_stat_items[] = {
> NR_SHMEM_THPS,
> NR_FILE_THPS,
> NR_ANON_THPS,
> + NR_VMSCAN_WRITE,
> NR_VMALLOC,
> NR_KERNEL_STACK_KB,
> NR_PAGETABLE,
> diff --git a/mm/mmzone.c b/mm/mmzone.c
> index 0c8f181d9d50..17139db4d291 100644
> --- a/mm/mmzone.c
> +++ b/mm/mmzone.c
> @@ -78,6 +78,7 @@ void lruvec_init(struct lruvec *lruvec)
>
> memset(lruvec, 0, sizeof(struct lruvec));
> spin_lock_init(&lruvec->lru_lock);
> + spin_lock_init(&lruvec->cost_lock);
> zswap_lruvec_state_init(lruvec);
>
> for_each_lru(lru)
> diff --git a/mm/swap.c b/mm/swap.c
> index 588f50d8f1a8..74b281778cbc 100644
> --- a/mm/swap.c
> +++ b/mm/swap.c
> @@ -272,73 +272,6 @@ void folio_rotate_reclaimable(struct folio *folio)
> folio_batch_add_and_move(folio, lru_move_tail);
> }
>
> -void lru_note_cost_unlock_irq(struct lruvec *lruvec, bool file,
> - unsigned int nr_io, unsigned int nr_rotated)
> - __releases(lruvec->lru_lock)
> - __releases(rcu)
> -{
> - unsigned long cost;
> -
> - /*
> - * Reflect the relative cost of incurring IO and spending CPU
> - * time on rotations. This doesn't attempt to make a precise
> - * comparison, it just says: if reloads are about comparable
> - * between the LRU lists, or rotations are overwhelmingly
> - * different between them, adjust scan balance for CPU work.
> - */
> - cost = nr_io * SWAP_CLUSTER_MAX + nr_rotated;
> - if (!cost) {
> - spin_unlock_irq(&lruvec->lru_lock);
> - rcu_read_unlock();
> - return;
> - }
> -
> - for (;;) {
> - unsigned long lrusize;
> -
> - /* Record cost event */
> - if (file)
> - lruvec->file_cost += cost;
> - else
> - lruvec->anon_cost += cost;
> -
> - /*
> - * Decay previous events
> - *
> - * Because workloads change over time (and to avoid
> - * overflow) we keep these statistics as a floating
> - * average, which ends up weighing recent refaults
> - * more than old ones.
> - */
> - lrusize = lruvec_page_state(lruvec, NR_INACTIVE_ANON) +
> - lruvec_page_state(lruvec, NR_ACTIVE_ANON) +
> - lruvec_page_state(lruvec, NR_INACTIVE_FILE) +
> - lruvec_page_state(lruvec, NR_ACTIVE_FILE);
> -
> - if (lruvec->file_cost + lruvec->anon_cost > lrusize / 4) {
> - lruvec->file_cost /= 2;
> - lruvec->anon_cost /= 2;
> - }
> -
> - spin_unlock_irq(&lruvec->lru_lock);
> - lruvec = parent_lruvec(lruvec);
> - if (!lruvec) {
> - rcu_read_unlock();
> - break;
> - }
> - spin_lock_irq(&lruvec->lru_lock);
> - }
> -}
> -
> -void lru_note_cost_refault(struct folio *folio)
> -{
> - struct lruvec *lruvec;
> -
> - lruvec = folio_lruvec_lock_irq(folio);
> - lru_note_cost_unlock_irq(lruvec, folio_is_file_lru(folio),
> - folio_nr_pages(folio), 0);
> -}
> -
> static void lru_activate(struct lruvec *lruvec, struct folio *folio)
> {
> long nr_pages = folio_nr_pages(folio);
> @@ -1164,8 +1097,6 @@ void lru_reparent_memcg(struct mem_cgroup *memcg, struct mem_cgroup *parent, int
>
> child_lruvec = mem_cgroup_lruvec(memcg, NODE_DATA(nid));
> parent_lruvec = mem_cgroup_lruvec(parent, NODE_DATA(nid));
> - parent_lruvec->anon_cost += child_lruvec->anon_cost;
> - parent_lruvec->file_cost += child_lruvec->file_cost;
>
> for_each_lru(lru)
> lruvec_reparent_lru(child_lruvec, parent_lruvec, lru, nid);
> diff --git a/mm/vmscan.c b/mm/vmscan.c
> index 053f41584989..0f6334005610 100644
> --- a/mm/vmscan.c
> +++ b/mm/vmscan.c
> @@ -641,7 +641,7 @@ static pageout_t writeout(struct folio *folio, struct address_space *mapping,
> folio_clear_reclaim(folio);
>
> trace_mm_vmscan_write_folio(folio);
> - node_stat_add_folio(folio, NR_VMSCAN_WRITE);
> + lruvec_stat_mod_folio(folio, NR_VMSCAN_WRITE, folio_nr_pages(folio));
> return PAGE_SUCCESS;
> }
>
> @@ -1418,8 +1418,6 @@ static unsigned int shrink_folio_list(struct list_head *folio_list,
> sc->nr_scanned -= (nr_pages - 1);
> nr_pages = 1;
> }
> - stat->nr_pageout += nr_pages;
> -
> if (folio_test_writeback(folio))
> goto keep;
> if (folio_test_dirty(folio))
> @@ -2047,9 +2045,6 @@ static unsigned long shrink_inactive_list(unsigned long nr_to_scan,
> mod_lruvec_state(lruvec, PGROTATE_ANON + file,
> nr_scanned - nr_reclaimed);
>
> - lruvec_lock_irq(lruvec);
> - lru_note_cost_unlock_irq(lruvec, file, stat.nr_pageout,
> - nr_scanned - nr_reclaimed);
> handle_reclaim_writeback(nr_taken, pgdat, sc, &stat);
> trace_mm_vmscan_lru_shrink_inactive(pgdat->node_id,
> nr_scanned, nr_reclaimed, &stat, sc->priority, file);
> @@ -2158,8 +2153,6 @@ static void shrink_active_list(unsigned long nr_to_scan,
> if (nr_rotated)
> mod_lruvec_state(lruvec, PGROTATE_ANON + file, nr_rotated);
>
> - lruvec_lock_irq(lruvec);
> - lru_note_cost_unlock_irq(lruvec, file, 0, nr_rotated);
> trace_mm_vmscan_lru_shrink_active(pgdat->node_id, nr_taken, nr_activate,
> nr_deactivate, nr_rotated, sc->priority, file);
> }
> @@ -2292,8 +2285,10 @@ enum scan_balance {
>
> static void prepare_scan_control(pg_data_t *pgdat, struct scan_control *sc)
> {
> - unsigned long file;
> + struct lru_cost *anon_cost, *file_cost;
> struct lruvec *target_lruvec;
> + unsigned long lrusize;
> + unsigned long file;
>
> if (lru_gen_enabled() && !lru_gen_switching())
> return;
> @@ -2309,11 +2304,69 @@ static void prepare_scan_control(pg_data_t *pgdat, struct scan_control *sc)
>
> /*
> * Determine the scan balance between anon and file LRUs.
> + *
> + * The cost model is based on rotations, refaults and
> + * reclaim-driven writes (anon only) on each side.
> + *
> + * These event counters are monotonic, so each reclaim cycle
> + * the delta since the last scan is extracted and incorporated
> + * into a decaying average. This ensures currency, as workloads
> + * change over time, and avoids overflow in the calculations.
> + *
> + * Use lruvec_page_state_monotonic() so unsigned subtraction
> + * yields the correct delta across a signed-long wraparound of
> + * the underlying counter (a real hazard on 32-bit that the
> + * clamp in lruvec_page_state() would otherwise turn into a huge
> + * spurious delta).
> */
> - spin_lock_irq(&target_lruvec->lru_lock);
> - sc->anon_cost = target_lruvec->anon_cost;
> - sc->file_cost = target_lruvec->file_cost;
> - spin_unlock_irq(&target_lruvec->lru_lock);
> + spin_lock(&target_lruvec->cost_lock);
> +
> + for (int f = 0; f <= 1; f++) {
> + struct lru_cost *cost = &target_lruvec->cost[f];
> + unsigned long rotated, io, nr_rotated, nr_io;
> +
> + rotated = lruvec_page_state_monotonic(target_lruvec,
> + PGROTATE_ANON + f);
> + io = lruvec_page_state_monotonic(target_lruvec,
> + WORKINGSET_RESTORE_BASE + f);
> + if (f == WORKINGSET_ANON)
> + io += lruvec_page_state_monotonic(target_lruvec,
> + NR_VMSCAN_WRITE);
> +
> + nr_rotated = rotated - cost->last_rotated;
> + nr_io = io - cost->last_io;
> +
> + /*
> + * Reflect the relative cost of incurring IO and spending
> + * CPU time on rotations. This doesn't attempt to make a
> + * precise comparison, it just says: if reloads are about
> + * comparable between the LRU lists, or rotations are
> + * overwhelmingly different between them, adjust scan
> + * balance for CPU work.
> + */
> + cost->count += nr_io * SWAP_CLUSTER_MAX + nr_rotated;
Sashiko review:
---
Can this calculation overflow on 32-bit architectures?
...
When global reclaim finally runs, nr_io can be a massive delta. Multiplyingi
this by SWAP_CLUSTER_MAX (32) could silently overflow the 32-bit unsigned long,
producing a random garbage cost.
---
I dont think we need to worry about this.
ULONG_MAX / SWAP_CLUSTER_MAX = 2^32 / 32 = 134,217,728
nr_io is a per-lruvec, per-cycle delta of (WORKINGSET_RESTORE_{ANON,FILE} on the
file side; those two plus NR_VMSCAN_WRITE on the anon side), in
pages. So the multiplication wraps only when this lruvec has
observed ~134 M new page-events since its own last
prepare_scan_control() call.
On 32-bit, addressable RAM tops out at ~4 GiB (~1 M pages), or ~64 GiB with
PAE (~16 M pages) in the extreme. 134 M events therefore requires
~130 full swap-and-refault turnovers of a 4 GiB LRU, or ~8
turnovers of a full 64 GiB PAE LRU, without a single intervening
prepare_scan_control() call on that lruvec.
> +
> + cost->last_rotated = rotated;
> + cost->last_io = io;
> + }
> +
> + anon_cost = &target_lruvec->cost[WORKINGSET_ANON];
> + file_cost = &target_lruvec->cost[WORKINGSET_FILE];
> +
> + lrusize = lruvec_page_state(target_lruvec, NR_INACTIVE_ANON) +
> + lruvec_page_state(target_lruvec, NR_ACTIVE_ANON) +
> + lruvec_page_state(target_lruvec, NR_INACTIVE_FILE) +
> + lruvec_page_state(target_lruvec, NR_ACTIVE_FILE);
> +
> + while (anon_cost->count + file_cost->count > lrusize / 4) {
Sashiko review:
---
If the newly added garbage cost caused anon_cost->count + file_cost->count to
wrap around and overflow to a small value (<= lrusize / 4), would this skip
the decay loop entirely?
---
The same 32-bit / ~134M events with no intervening scan reasoning from the
previous finding gates when this can happen at all.
Same as above, IMO safe to ignore.
> + anon_cost->count /= 2;
> + file_cost->count /= 2;
> + }
> +
> + sc->anon_cost = anon_cost->count;
> + sc->file_cost = file_cost->count;
> +
> + spin_unlock(&target_lruvec->cost_lock);
>
> /*
> * Target desirable inactive:active list ratios for the anon
> diff --git a/mm/workingset.c b/mm/workingset.c
> index f351798e723a..7ac2b88c80ae 100644
> --- a/mm/workingset.c
> +++ b/mm/workingset.c
> @@ -584,11 +584,6 @@ void workingset_refault(struct folio *folio, void *shadow)
> /* Folio was active prior to eviction */
> if (workingset) {
> folio_set_workingset(folio);
> - /*
> - * XXX: Move to folio_add_lru() when it supports new vs
> - * putback
> - */
> - lru_note_cost_refault(folio);
> mod_lruvec_state(lruvec, WORKINGSET_RESTORE_BASE + file, nr);
> }
> out:
> --
> 2.53.0-Meta
>
>
prev parent reply other threads:[~2026-07-29 13:02 UTC|newest]
Thread overview: 9+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-27 16:23 [PATCH v5 0/3] mm/vmscan: reduce lru_lock contention via vmstat-derived scan-balance cost Usama Arif
2026-07-27 16:23 ` [PATCH v5 1/3] mm/vmstat, mm/memcontrol: add _monotonic vmstat readers Usama Arif
2026-07-27 16:23 ` [PATCH v5 2/3] mm/vmscan: add pgrotate_anon and pgrotate_file vmstat counters Usama Arif
2026-07-27 16:33 ` Shakeel Butt
2026-07-27 17:23 ` Johannes Weiner
2026-07-29 13:09 ` Usama Arif
2026-07-27 16:23 ` [PATCH v5 3/3] mm/vmscan: reduce lru_lock contention via vmstat-derived scan-balance cost Usama Arif
2026-07-27 16:35 ` Shakeel Butt
2026-07-29 13:02 ` Usama Arif [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260729130241.3327679-1-usama.arif@linux.dev \
--to=usama.arif@linux.dev \
--cc=akpm@linux-foundation.org \
--cc=axelrasmussen@google.com \
--cc=baoquan.he@linux.dev \
--cc=cgroups@vger.kernel.org \
--cc=chrisl@kernel.org \
--cc=david@kernel.org \
--cc=hannes@cmpxchg.org \
--cc=kasong@tencent.com \
--cc=kernel-team@meta.com \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@suse.com \
--cc=muchun.song@linux.dev \
--cc=nphamcs@gmail.com \
--cc=qi.zheng@linux.dev \
--cc=rientjes@google.com \
--cc=roman.gushchin@linux.dev \
--cc=rppt@kernel.org \
--cc=shakeel.butt@linux.dev \
--cc=surenb@google.com \
--cc=vbabka@kernel.org \
--cc=weixugc@google.com \
--cc=youngjun.park@lge.com \
--cc=yuanchu@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox