From: Baoquan He <baoquan.he@linux.dev>
To: Baoquan He <hebaoquan@kylinos.cn>
Cc: linux-mm@kvack.org, akpm@linux-foundation.org, david@kernel.org,
rostedt@goodmis.org, mhiramat@kernel.org, kasong@tencent.com,
baohua@kernel.org, qi.zheng@linux.dev, shakeel.butt@linux.dev,
axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com,
baolin.wang@linux.alibaba.com, hannes@cmpxchg.org
Subject: Re: [PATCH 0/9] mm/mglru: suppress empty page table walks during aging
Date: Mon, 24 Aug 2026 16:11:51 +0800 [thread overview]
Message-ID: <aov8x5b4w-cnv8sI@fedora> (raw)
In-Reply-To: <20260824073806.629593-1-hebaoquan@kylinos.cn>
On 08/24/26 at 03:37pm, Baoquan He wrote:
> In the current MGLRU, lru_gen_use_mm() will mark an mm for all nodes at
> each context switch, so on each node aging will walk into each mm's page
> table independently. On a multi-NUMA node system, one process launched
> on one or a subset of nodes, its mm is walked by all other nodes's aging
> while finds on pages for their lruvec. This is pure waste (100% walks
^ s/on/no/, typo
> are empty on those other nodes)
>
> This patch series suppresses these empty page table walks with two
> complementary mechanisms at different levels:
>
> - empty_map (cross-node specifc). At mm granularity, one process's mm
> whose walk found no pages for one node's lruvec is skipped on that
> node between re-scan passes.
>
> - PUD-level Bloom filter (general). One level up from the existing PMD
> filters, it skips any 1GB PUD that had no young entries last
> generation - namely whose 512 PMDs all failed PMD test. Cross-node
> empty walks can prove its effect the best, but it also suppresses
> purely code PUD inside local mms. So it reduces unnecessary walking
> in any workload with cold areas.
>
> The two complement each other: empty_map reduces the *number* of
> cross-node walks, the PUD-level filter reduces the *cost* of the walks
> that remain.
>
> Measurement test (3-socket Xeon 6766E, 32 x 256MB workers pinned to node 0):
>
> node 1 (foreign mm walks):
> empty_map off (PUD filter on): 3300 walks, all empty, ~1238 PTE scans/walk
> both on: 900 walks (-73%), ~0 PTE scans/walk
>
> - empty_map cuts node 1's walk count by ~73% (every 4th pass re-scans
> to close migration/fault windows).
> - the PUD-level filter makes the remaining walks nearly free: the 1GB
> subtrees are skipped instead of iterating their 512 PMDs
> (leaf_total drops from ~1238 to ~0 per walk).
> - worker RSS is unchanged in both modes - no premature reclaim. A
> force_scan pass confirms the local pages are found young and eligible
> (leaf_eligible == young == 21320 on node 0) when the filters are
> bypassed, i.e. the pages are hot and the aging walker can find them.
>
> Performance regression test (make -j4 in a 3G memory cgroup, 4 vCPU / 8GB
> 2-NUMA VM, median of 3 runs):
>
> baseline patch diff
> build time 10m01s 9m41s -3.4%
> pgpgin 136868 138080 +0.9%
> pgmajfault 1832 1862 +1.6%
>
> Build time, page-in and major-fault counts are within run-to-run
> variance of a no-patch baseline - no measurable regression. The build
> time is if anything slightly lower, consistent with the suppression
> reducing reclaim overhead (fewer cross-node empty walks) during a real
> memory-pressure workload.
>
> The test codes/scripts (numa_workload + run_test_v4.sh) are available at:
> https://github.com/baoquan-he/mglru-empty-walk-test
>
> RFC-v1:
> - Bloom filter helpers are named symmetrically -
> test/update/reset_pmd_bloom_filter() alongside the PUD-level ones, and
> the struct field filters -> pmd_filters (Barry).
>
> - An mm is marked empty only when its page tables were actually walked
> (a failed mmap_read_trylock() or a stale seq is not empty), and the skip
> is invalidated on the major page-fault and migration paths. (Sashiko)
>
> - Reworked the re-scan: the old shared counter across nodes got stuck at
> the slowest node, so the skip never really engaged. Now each node re-scans
> its empty mms every N passes on its own clock. N is the skip_empty knob
> (default 4, proper read/write with input validation). (Sashiko)
>
> - Added measurement counters/tracepoint, and a kernel-build regression
> test (no measurable impact).
>
>
> Baoquan He (9):
> mm/mglru: add MM_WALK_EMPTY stats and tracepoint
> mm/mglru: suppress cross-node empty page table walks
> mm/mglru: add debugfs knob for the empty-walk skip threshold
> mm/mglru: invalidate empty-walk skip on page fault and migration
> mm/mglru: add PUD-level Bloom filter state
> mm/mglru: refactor Bloom filter helpers for two filter levels
> mm/mglru: skip PUD subtrees during aging
> mm/mglru: report hot PUDs from the rmap feedback path
> mm/mglru: count PUD subtrees skipped by the PUD-level filter
>
> include/linux/mm_types.h | 22 ++++
> include/linux/mmzone.h | 13 ++-
> include/trace/events/vmscan.h | 30 +++++
> mm/huge_memory.c | 3 +
> mm/memory.c | 15 +++
> mm/migrate.c | 4 +
> mm/vmscan.c | 205 +++++++++++++++++++++++++++++-----
> 7 files changed, 264 insertions(+), 28 deletions(-)
>
>
> base-commit: efecab401cb15fd3bb9bc05990609acb6b267ff2
> --
> 2.54.0
>
>
next prev parent reply other threads:[~2026-08-24 8:12 UTC|newest]
Thread overview: 18+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-24 7:37 [PATCH 0/9] mm/mglru: suppress empty page table walks during aging Baoquan He
2026-08-24 7:37 ` [PATCH 1/9] mm/mglru: add MM_WALK_EMPTY stats and tracepoint Baoquan He
2026-08-28 5:37 ` Barry Song
2026-08-28 7:46 ` Baoquan He
2026-08-24 7:37 ` [PATCH 2/9] mm/mglru: suppress cross-node empty page table walks Baoquan He
2026-08-28 6:35 ` Barry Song
2026-08-28 7:26 ` Baoquan He
2026-08-24 7:38 ` [PATCH 3/9] mm/mglru: add debugfs knob for the empty-walk skip threshold Baoquan He
2026-08-24 7:38 ` [PATCH 4/9] mm/mglru: invalidate empty-walk skip on page fault and migration Baoquan He
2026-08-24 7:38 ` [PATCH 5/9] mm/mglru: add PUD-level Bloom filter state Baoquan He
2026-08-24 7:38 ` [PATCH 6/9] mm/mglru: refactor Bloom filter helpers for two filter levels Baoquan He
2026-08-24 7:38 ` [PATCH 7/9] mm/mglru: skip PUD subtrees during aging Baoquan He
2026-08-24 7:38 ` [PATCH 8/9] mm/mglru: report hot PUDs from the rmap feedback path Baoquan He
2026-08-24 7:38 ` [PATCH 9/9] mm/mglru: count PUD subtrees skipped by the PUD-level filter Baoquan He
2026-08-24 8:11 ` Baoquan He [this message]
2026-08-24 8:42 ` [PATCH 0/9] mm/mglru: suppress empty page table walks during aging Baoquan He
2026-08-28 6:11 ` Barry Song
2026-08-28 7:42 ` Baoquan He
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aov8x5b4w-cnv8sI@fedora \
--to=baoquan.he@linux.dev \
--cc=akpm@linux-foundation.org \
--cc=axelrasmussen@google.com \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=david@kernel.org \
--cc=hannes@cmpxchg.org \
--cc=hebaoquan@kylinos.cn \
--cc=kasong@tencent.com \
--cc=linux-mm@kvack.org \
--cc=mhiramat@kernel.org \
--cc=qi.zheng@linux.dev \
--cc=rostedt@goodmis.org \
--cc=shakeel.butt@linux.dev \
--cc=weixugc@google.com \
--cc=yuanchu@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox