From: Tao Cui <cui.tao@linux.dev>
To: Andrew Morton <akpm@linux-foundation.org>
Cc: linux-mm@kvack.org, Suren Baghdasaryan <surenb@google.com>,
Michal Hocko <mhocko@suse.com>,
David Hildenbrand <david@kernel.org>,
Lorenzo Stoakes <ljs@kernel.org>,
"Liam R. Howlett" <liam@infradead.org>,
Vlastimil Babka <vbabka@kernel.org>,
Mike Rapoport <rppt@kernel.org>,
linux-kernel@vger.kernel.org, Tao Cui <cuitao@kylinos.cn>,
cui.tao@linux.dev
Subject: [PATCH] mm/vmpressure: scale the vmpressure window with machine size
Date: Fri, 24 Jul 2026 13:43:05 +0800 [thread overview]
Message-ID: <20260724054305.516126-1-cui.tao@linux.dev> (raw)
From: Tao Cui <cuitao@kylinos.cn>
vmpressure_win -- the number of pages the reclaimer must scan before
socket pressure is re-evaluated -- has been a fixed 512 pages (2 MB) since
vmpressure was introduced. The reclaimer scans more pages on a larger
machine, so the window is reached far more often there and socket pressure
is re-armed every handful of scanned pages. The comment at the definition
has asked for the window to scale with machine size "as we do for vmstat
thresholds" for over a decade.
Scale it the same way calculate_normal_threshold() does: logarithmically
with memory (fls of memory in 128 MB units), computed once in a
subsys_initcall once totalram_pages() is known. A machine under 128 MB
keeps the historical 512 pages; the window then grows by SWAP_CLUSTER_MAX
* 16 per doubling of memory.
Why this matters: the scanned/reclaimed ratio that drives socket pressure
is averaged over the window, and the window rate-limits the evaluation.
With a fixed 2 MB window the evaluation runs the same number of times
regardless of machine size, which is disproportionately many on a large
machine. Measured by cold-booting one VM at each size and running the
same cgroup-bound reclaim workload (so the page count is identical across
sizes):
config scaled_win pages scanned 512-win evals scaled evals
4 GB 3072 11.9 M 23267 3877
8 GB 3584 11.9 M 23306 3329
16 GB 4096 11.9 M 23281 2910
32 GB 4608 11.9 M 23268 2585
64 GB 5120 11.9 M 23281 2328
For the same reclaim work the fixed window evaluates ~23k times at every
machine size; the scaled window evaluates fewer times the larger the
machine -- a 6x reduction at 4 GB growing to 10x at 64 GB (and the
logarithmic growth continues: ~11x projected at 128 GB). Each evaluation
takes the per-memcg sr_lock and may write the socket_pressure seqlock, so
on larger machines with more memcgs under pressure this is real overhead
the fixed window pays needlessly.
The default stays 512 until the initcall runs, so early-boot reclaim is
unchanged.
Signed-off-by: Tao Cui <cuitao@kylinos.cn>
---
include/linux/vmpressure.h | 2 +-
mm/vmpressure.c | 20 +++++++++++++++++---
2 files changed, 18 insertions(+), 4 deletions(-)
diff --git a/include/linux/vmpressure.h b/include/linux/vmpressure.h
index b4d13457bc2a..09111f5bdc88 100644
--- a/include/linux/vmpressure.h
+++ b/include/linux/vmpressure.h
@@ -51,7 +51,7 @@ extern struct vmpressure *memcg_to_vmpressure(struct mem_cgroup *memcg);
extern struct mem_cgroup *vmpressure_to_memcg(struct vmpressure *vmpr);
/* Shared with the v1 vmpressure block in mm/memcontrol-v1.c. */
-extern const unsigned long vmpressure_win;
+extern unsigned long vmpressure_win;
extern enum vmpressure_levels vmpressure_calc_level(unsigned long scanned,
unsigned long reclaimed);
diff --git a/mm/vmpressure.c b/mm/vmpressure.c
index 9629240d77ad..4c8671273c79 100644
--- a/mm/vmpressure.c
+++ b/mm/vmpressure.c
@@ -31,10 +31,24 @@
* As the vmscan reclaimer logic works with chunks which are multiple of
* SWAP_CLUSTER_MAX, it makes sense to use it for the window size as well.
*
- * TODO: Make the window size depend on machine size, as we do for vmstat
- * thresholds. Currently we set it to 512 pages (2MB for 4KB pages).
+ * Scale the window with machine size, the way vmstat thresholds do: on a
+ * larger machine the reclaimer scans more pages, so a fixed window would
+ * re-evaluate pressure every handful of pages. The scaling is logarithmic
+ * (fls, like calculate_normal_threshold()), keeping the growth moderate.
+ * The default is the historical 512 pages; an early initcall applies the
+ * scaling once totalram_pages() is known.
*/
-const unsigned long vmpressure_win = SWAP_CLUSTER_MAX * 16;
+unsigned long vmpressure_win __read_mostly = SWAP_CLUSTER_MAX * 16;
+
+static int __init vmpressure_init_window(void)
+{
+ unsigned long mem128m; /* machine memory in 128MB units */
+
+ mem128m = totalram_pages() >> (27 - PAGE_SHIFT);
+ vmpressure_win = SWAP_CLUSTER_MAX * 16 * (1 + fls(mem128m));
+ return 0;
+}
+subsys_initcall(vmpressure_init_window);
/*
* These thresholds are used when we account memory pressure through
--
2.43.0
next reply other threads:[~2026-07-24 5:43 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-24 5:43 Tao Cui [this message]
2026-07-24 7:30 ` [PATCH] mm/vmpressure: scale the vmpressure window with machine size Lorenzo Stoakes (ARM)
2026-07-24 8:56 ` Tao Cui
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260724054305.516126-1-cui.tao@linux.dev \
--to=cui.tao@linux.dev \
--cc=akpm@linux-foundation.org \
--cc=cuitao@kylinos.cn \
--cc=david@kernel.org \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@suse.com \
--cc=rppt@kernel.org \
--cc=surenb@google.com \
--cc=vbabka@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.