From: Davidlohr Bueso <dave@stgolabs.net>
To: Bharata B Rao <bharata@amd.com>
Cc: linux-kernel@vger.kernel.org, linux-mm@kvack.org,
jic23@kernel.org, dave.hansen@intel.com, gourry@gourry.net,
mgorman@techsingularity.net, mingo@redhat.com,
peterz@infradead.org, raghavendra.kt@amd.com, riel@surriel.com,
rientjes@google.com, sj@kernel.org, weixugc@google.com,
willy@infradead.org, ying.huang@linux.alibaba.com,
ziy@nvidia.com, nifan.cxl@gmail.com, xuezhengchu@huawei.com,
yiannis@zptcorp.com, akpm@linux-foundation.org, david@kernel.org,
byungchul@sk.com, kinseyho@google.com, joshua.hahnjy@gmail.com,
yuanchu@google.com, balbirs@nvidia.com, alok.rathore@samsung.com,
shivankg@amd.com, donettom@linux.ibm.com
Subject: Re: [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure
Date: Tue, 29 Sep 2026 12:56:46 -0700 [thread overview]
Message-ID: <20260929195646.4ttbl22clnrrktg7@offworld> (raw)
In-Reply-To: <31442591-d020-47b5-8f1c-b87fb632c226@amd.com>
On Sun, 27 Sep 2026, Bharata B Rao wrote:
>Hotness promotion engine
>------------------------
>This is not something that was written for pghot from scratch, but instead it is
>the same hot page promotion engine that is part of NUMAB2 which is now
>generalized and moved to pghot. So this is not the complexity that pghot
>introduces afresh.
>
>So considering all these, I see pghot as a light-weight and low-overhead
>mechanism to track per-PFN hotness and do async batch migration. Initial
>versions had fancy double data structures; a large hash a small binary tree of
>promotion-ready records and associated synchronization mechanism, but that is
>all past now.
I think we are all in agreement that the async batch migration is wanted.
>
>pghot interface for sampling
>============================
>pghot_record_access() interface was designed keeping the existing NUMAB2 source
>in mind. It fits that and it fits other sources like IBS Memory Profiler. So
>sampling sources report an access and the shared promotion engine acts upon it.
>
>But for sources like CHMU, from what you describe, I gather that a bulk
>reporting interface plus an indication to bypass the engine to treat the PFNs as
>migrate-ready, is what is required. Should those migrate-ready PFNs go through
>the regular pghot tracking (getting into section hotmaps to be picked up by
>kmigrated) or even that should be bypassed?
>
>However, in the context of PTE A bit based source, I have often thought about
>extending the interface for
>
>- bulk reporting where more than one PFN gets reported.
>- indicating the bypass options (frequency check bypass, recency check bypass etc)
>
>NUMAB2 has to perform better in pghot
>=====================================
>pghot is about a sub-system that makes it possible to have multiple sources to
>coexist with reuse of common hot page promotion engine.
>
>NUMAB2 source resides within the scheduler and the promotion engine is also part
>of the scheduler. Through pghot, I am separating the source (NUMA hint faults)
>from the engine and moving that existing engine into pghot, to a common place
>where it gets reused for other sources as well.
>
>It is the same NUMA hint faults and more or less the same engine and hence my
>main objective is to ensure that there is no regression during this move.
>Additional performance optimizations can be done to the engine itself separately
>but that shouldn't be the baseline expectation from pghot.
>
>Is moving hot page promotion out of scheduler into a dedicated system, a good
>thing in general? I believe so as scheduler isn't the right place for it to
>reside. However I would like to hear from scheduler folks on this.
So two of the autonuma balancing og authors are scheduler experts - and iirc
*the* reason back then was locality. And Peter has already nacked the IBS
stuff in the past. But indeed the batch async part would be good to get nack/ack;
albeit the cgroup charging situation.
>
>Do we even need a centralized hot page promotion engine?
>========================================================
>NUMAB2 is good and serves as a good baseline for any new source that comes up.
>But with different kinds of sources becoming available, do we want all of them
>to duplicate the hot page heuristics and promote hot pages on their own? I
>thought that may not be preferable and hence started this effort.
What are these sources that will become available?
>CXL HMU may not need the promotion engine, but IBS Memory Profiler needs. It
>needs a promotion engine with full recency and frequency considerations before
>promoting. I don't think an arch driver like IBS Memory Profiler should be doing
>hotness heuristics within itself but instead be using the existing engine. In
>fact in my early posts, the driver based on primary IBS instance was feeding the
>"access sample" as "NUMA hint fault" to NUMAB1/B2 so that rest of NUMA
>Balancing/Hot page promotion engine just worked. But I think pghot is a better
>approach than that.
I agree that the IBS driver should not be doing hotness heuristics, it's the hw
that should.
>
>Why full pghot? Isn't async batch migration enough?
>===================================================
>Some of the above reasons apply but during the course of iterations, I have had
>implementations of just the migrator (kmigrated [1]).
>
>If every sub-system/source has intelligence of its own and just wants to
>handover a list of pages to async migrator thread, that's not much of an effort
>as this implementation showed.
>
>But then if some source needs rate-limiting and another source needs only
>hottest pages to be promoted, then again we overlap with the existing NUMAB2 engine.
>
>Then if we want to be slightly generic and want two sources to complement each
>other or the promoter to differentiate between lukewarm vs hot pages/regions
>then we may have to maintain hotness records and may soon end up with something
>similar to pghot's hotness tracking and reporting mechanism.
>
>Hardware sources have to out-perform NUMAB2
>===========================================
>Different sources will have different characteristics and capabilities and will
>help different workloads differently. So it is the choice that one could
>provide. Sometimes sources can complement each other as well.
>
>What IBS Memory Profiler has shown is that it can match and/or exceed (for
>Graph500, ptr-chase, llama) NUMAB2 with no hint faults overhead [2]. Also please
>check the initial XSBench numbers in my inline reply.
It's all about the numbers, and the ones you have just don't really sell - which is
one of the reasons this has been going on for years. If numa balancing didn't
exist, then maybe adding all this would make sense. For the XSBench I don't think
making decisions based on the non-overcommitted case is worthwhile.
>
>Cost of an unused source
>========================
>Not all the sources are required for every situation. Sources can be disabled at
>compile time or not enabled at run time with no cost or effect on other sources.
>But hotmap allocations would remain as a static cost even when no source is
>enabled (built out at compile time, the map is gone; sources off at runtime, the
>map remains)
>
>Tracking granularity: per-PFN vs region
>=======================================
>Often times this question comes up when pghot is compared with DAMON.
>
>Firstly, the baseline that we have (NUMAB2), tracks hotness at page granularity.
>To support this, pghot started with per-PFN granularity. Naturally two concerns
>come up:
>
>1. Memory overhead: I have shown the numbers above. It is lower-tier only and
>not much IMHO.
>
>2. Scan/CPU overhead: pghot started with scanning all PFNs, that was expensive.
>Then it introduced hotness bit per memory section so that only those sections
>which are marked hot are scanned. Now I have added (yet to be posted) a
>sub-section level hotness tracking where a hotness bit is maintained for each
>fixed 2M region within a section. This has considerably reduced the CPU overhead
>for kmigrated thread. While the ptr-chase numbers that I shared with the
>separated out IBS RFC v0 post of IBS Memory Profiler [2] does show the kmigrated
>utilization numbers with sub-section tracking, I plan to have some more numbers
>ready for LPC.
I look forward to seeing any new numbers you have. Region granularity is certainly
more aligned with willy as well as chmu.
>I am beginning to feel that this may be a good middle ground between real
>region-level tracking (where all pages of region are promoted irrespective of
>their real hotness) vs scanning at sub-section granularity but performing
>per-PFN precise promotion.
[...]
>I would be interested to understand more about how you ran XSBench, the
>parameters used, the promotion stats, if demotion stats etc. Do share when you
>get time.
The actual command is 'XSBench -g 90424' so only set the gridpoints, everything
else is default, so total ~44Gb footprint. I don't have the vmstats currently
(I do not run these benchmarks) but will share them once I get them - I can
affirm that demotion is in fact enabled, so full TPP up and down.
For the chmu: 32GB device (1:1 dram and cxl), this is with a 4k unit size, 1s
epoch, reporting mode is always on, threshold value is 1024.
The numa balancing mode was set to 3, the rest used the default values:
pghot_freq_threshold=2, pghot_promote_freq_window_ms=3000,
pghot_promote_rage_limit_MBps=65536, kmigrated_sleep_ms=100, kmigrated_batch_nr=512.
I also have numbers for two more benchmarks with a real chmu, with basically
the same numab parameters:
(i) TaoBench almost 2x throughput, going from ~250 qps to ~480 qps, vs NUMAB3.
(dram:cxl is 16:32Gb with a 32Gb memsize, num_clients=2, clients_per_thread=75)
(ii) Graph500 only shows a smaller ~20% improvement vs NUMAB3 (bfs mean time drops
from 5.0 to 4.15 secs). This was for a memory ratio dram:cxl as 64:32Gb. The
algorithm is BFS, edge factor 512, 16 mpi processes.
(mpiexec.openmpi -n 16 ./graph500_reference_bfs 22 512)
>
>I did a round of testing to compare base-NUMAB2, pghot-hintfaults(NUMAB2) and
>pghot-hwhints (IBS Memory Profiler). I am still experimenting with options,
>placement etc but here are my initial numbers:
>
>Non-Overcommitted case: XSBench working set fits fully within toptier but starts
>on lower tier before the measurement phase
Those are nice numbers, but I don't think this is the methodology to use...
It is more representative for the workload's working set to be > total dram and
therefore spill into slower tier(s), instead of artificially starting in the slow
memory and moving up. Do you have data for the over committed case?
Thanks,
Davidlohr
>(XSBench -s XL -g 238847 -t 64 -p 60000000 -l 34, footprint ~116.56 GiB)
>
> Runtime (s) lookups/s Promotions (pages)
>base-NUMAB0 498.6 4.09M 0
>base-NUMAB2 422.4 4.83M 30.5M
>pghot-hintfaults 329.6 6.20M 30.5M
>pghot-hwhints 109.2 18.69M 2.7M
>
>Same amount of pages promoted by both base-NUMAB2 and pghot-hintfaults but
>runtime is better with pghot. This could be the benefit of async batched
>migration showing and no adverse effect of losing cache locality.
>
>pghot-hwhints shows good results. As I said this is just a first peek to the
>experimental numbers, I should have more concrete numbers and conclusion in LPC.
>
>[1] Kmigrated -
>https://lore.kernel.org/linux-mm/20250616133931.206626-1-bharata@amd.com/#t
>[2] IBS Memory Profiler RFC v0 -
>https://lore.kernel.org/linux-mm/20260924062206.319314-1-bharata@amd.com/
>
>Regards,
>Bharata.
prev parent reply other threads:[~2026-09-29 19:57 UTC|newest]
Thread overview: 56+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-28 5:43 [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure Bharata B Rao
2026-07-28 5:43 ` [PATCH v8 1/8] mm: migrate: Allow misplaced migration without VMA Bharata B Rao
2026-09-29 22:06 ` Davidlohr Bueso
2026-07-28 5:43 ` [PATCH v8 2/8] mm: migrate: Add promote_misplaced_memcg_folios() Bharata B Rao
2026-07-30 6:34 ` Bharata B Rao
2026-09-29 22:08 ` Davidlohr Bueso
2026-07-28 5:43 ` [PATCH v8 3/8] mm: Hot page tracking and promotion - pghot Bharata B Rao
2026-07-31 16:14 ` Bharata B Rao
2026-09-27 23:25 ` Davidlohr Bueso
2026-09-28 4:14 ` Bharata B Rao
2026-07-28 5:43 ` [PATCH v8 4/8] mm: pghot: Precision mode for pghot Bharata B Rao
2026-07-31 16:27 ` Bharata B Rao
2026-07-28 5:43 ` [PATCH v8 5/8] mm: sched: move NUMA balancing tiering promotion to pghot Bharata B Rao
2026-08-03 8:23 ` Bharata B Rao
2026-07-28 5:43 ` [PATCH v8 6/8] x86/ibs: Move IBS caps definitions into its own header Bharata B Rao
2026-07-28 5:43 ` [PATCH v8 7/8] x86/mm/ibs: In-kernel driver for AMD IBS Memory Profiler Bharata B Rao
2026-08-04 5:00 ` Bharata B Rao
2026-07-28 5:43 ` [PATCH v8 8/8] x86/mm/ibs: Add runtime controls for IBS memprofiler Bharata B Rao
2026-08-04 5:20 ` Bharata B Rao
2026-07-28 5:55 ` [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure - microbenchmark numbers Bharata B Rao
2026-07-28 5:59 ` [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure - NAS BT Bharata B Rao
2026-07-28 6:02 ` [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure - Graph500 Bharata B Rao
2026-07-28 6:05 ` [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure - redis-memtier Bharata B Rao
2026-07-28 6:17 ` [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure - llama-bench Bharata B Rao
2026-07-28 18:14 ` [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure Andrew Morton
2026-07-28 18:24 ` Matthew Wilcox
2026-07-28 18:57 ` Gregory Price
2026-07-28 19:20 ` David Hildenbrand (Arm)
2026-07-28 19:59 ` Gregory Price
2026-07-29 11:45 ` Bharata B Rao
2026-08-10 3:38 ` Yongting Lin
2026-08-10 4:16 ` Matthew Wilcox
2026-08-10 5:35 ` Bharata B Rao
2026-08-11 7:15 ` Yongting Lin
2026-08-13 2:21 ` Gregory Price
2026-08-10 14:37 ` SJ Park
2026-08-11 6:37 ` Yongting Lin
2026-07-29 9:35 ` Bharata B Rao
2026-07-29 13:54 ` SJ Park
2026-08-04 1:23 ` SJ Park
2026-08-06 5:49 ` Bharata B Rao
2026-08-06 13:44 ` SJ Park
2026-08-10 4:46 ` Bharata B Rao
2026-08-10 14:25 ` SJ Park
2026-09-11 21:08 ` Joshua Hahn
2026-09-16 3:08 ` Bharata B Rao
2026-09-16 20:52 ` Joshua Hahn
2026-09-17 5:23 ` Bharata B Rao
2026-09-27 23:21 ` Davidlohr Bueso
2026-09-28 4:13 ` Bharata B Rao
2026-09-25 2:12 ` Davidlohr Bueso
2026-09-25 9:50 ` SJ Park
2026-09-25 16:30 ` Davidlohr Bueso
2026-09-25 16:57 ` Gregory Price
2026-09-27 15:51 ` Bharata B Rao
2026-09-29 19:56 ` Davidlohr Bueso [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260929195646.4ttbl22clnrrktg7@offworld \
--to=dave@stgolabs.net \
--cc=akpm@linux-foundation.org \
--cc=alok.rathore@samsung.com \
--cc=balbirs@nvidia.com \
--cc=bharata@amd.com \
--cc=byungchul@sk.com \
--cc=dave.hansen@intel.com \
--cc=david@kernel.org \
--cc=donettom@linux.ibm.com \
--cc=gourry@gourry.net \
--cc=jic23@kernel.org \
--cc=joshua.hahnjy@gmail.com \
--cc=kinseyho@google.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=mgorman@techsingularity.net \
--cc=mingo@redhat.com \
--cc=nifan.cxl@gmail.com \
--cc=peterz@infradead.org \
--cc=raghavendra.kt@amd.com \
--cc=riel@surriel.com \
--cc=rientjes@google.com \
--cc=shivankg@amd.com \
--cc=sj@kernel.org \
--cc=weixugc@google.com \
--cc=willy@infradead.org \
--cc=xuezhengchu@huawei.com \
--cc=yiannis@zptcorp.com \
--cc=ying.huang@linux.alibaba.com \
--cc=yuanchu@google.com \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.