linux-mm.kvack.org archive mirror
 help / color / mirror / Atom feed
From: David Rientjes <rientjes@google.com>
To: Davidlohr Bueso <dave@stgolabs.net>, Fan Ni <nifan.cxl@gmail.com>,
	 Frank van der Linden <fvdl@google.com>,
	Gregory Price <gourry@gourry.net>,
	 Jonathan Cameron <jic23@kernel.org>,
	Joshua Hahn <joshua.hahnjy@gmail.com>,
	 Raghavendra K T <rkodsara@amd.com>,
	 "Rao, Bharata Bhasker" <bharata@amd.com>,
	SeongJae Park <sj@kernel.org>,  Wei Xu <weixugc@google.com>,
	Xuezheng Chu <xuezhengchu@huawei.com>,
	 Yiannis Nikolakopoulos <yiannis@zptcorp.com>,
	Zi Yan <ziy@nvidia.com>
Cc: linux-mm@kvack.org
Subject: [Linux Memory Hotness and Promotion] Notes from September 24, 2026
Date: Sat, 26 Sep 2026 14:43:23 -0700 (PDT)	[thread overview]
Message-ID: <0c8a7227-c7eb-bef1-7326-c7b04cde78a8@google.com> (raw)

Hi everybody,

Here are the notes from the last Linux Memory Hotness and Promotion call
that happened on Thursday, September 24.  Thanks to everybody who was 
involved!

These notes are intended to bring people up to speed who could not attend 
the call as well as keep the conversation going in between meetings.

----->o-----
Bharata updated on the status of his work.  He posted pghot-hwhints for 
the IBS Memory Profiler as a separate but dependent patch series as it is 
x86 specific and the pghot patch series is becoming too big with multiple 
sources of hotness information posted together.

Bharata noted that the benchmark numbers for pghot-hwhints so far looks 
promising and that he is looking forward to discussing this topic at LPC 
coming up.

----->o-----
Huan updated on the status of his work with Yiannis for non-temporal 
stores in migrate_pages(), he prepared a patch series to send for feedback 
to the group.  He looked into Shivank's feedback and work being done for 
batching with 2MB pages.  Today, we need to do a loop for each 4KB page, 
Shivank's work speeds this up but Huan has not yet integrated this since 
Shivank's work is not yet upstream.  There are also mixed results posted 
upstream.  Thus, for now this is being considered orthogonally until 
Shivank's changes land upstream.

I asked if Meta was also going to be looking into these optimizations; 
Gregory noted that they'll be waiting for it to land upstream.  They've 
been looking at basic correctness of NUMA Balancing rather than 
optimizations for now.  Per his discussion upstream, it turns out that 
NUMA Balancing, at least for tiering, has been broken for the last four 
years.

Shivank updated on his patches[1] for non-temporal stores without 
overloading migrate_mode.  Feedback on this would be useful.  He will also 
be presenting at LPC including additional experiments that have been run.

----->o-----
Davidlohr expressed concern about the direction of pghot overall.  He 
noted the upstream discussion has focused on not having any regressions, 
but there isn't an obvious gain either.  It looks like Joshua had the same 
conclusion although PSI may have improved.  Davidlohr has been working on 
integrating CHMU into pghot and found the experience to be complicated so 
he ended up bypassing most of it.  Whenever he got an interrupt from the 
device, he would try to create as many contiguous HPA ranges as he could 
to feed to pghot; once he got that, he decided to promote instead of doing 
decay and got a 3x improvement over NUMA Balancing that he was expecting.  
He suggested we may want to go back to the drawing board with regard to 
the complexity and interfaces of pghot.

Davidlohr suggested we may want to rely more on user-driven proactive 
reclaim and can then tune CHMU as desired so there is some control over 
freeing and promotion.  Another option would be to integrate this into 
DAMON as well.  One concern is that if this is integrated into pghot then 
we lose locality information.

Gregory said that pghot provides value in doing this asynchronously, which 
NUMA Balancing does not do, and that avoids lengthy stalled faults.  In 
this case, Davidlohr suggested we may just punt it to a kthread that would 
recheck if it's hot and promote.  Gregory thinks we're proposing 
significant complexity in pghot that doesn't yet carry all its weight but 
if we were to focus on the async part then this could be a start and we 
could build incrementally on top.  Davidlohr said we're still losing 
locality information; Gregory suggested this is a general limitation of 
tiering in general.  Davidlohr said we could feed the locality information 
to the kthread although this would include overhead for tracking.

Davidlohr agreed with the value of the asynchronous promotion handling, 
including for NUMA Balancing, but not with catering to IBS necessarily.  
Gregory is a proponent of splitting out kpromoted from the rest of the 
system.

Wei agreed with Gregory's assessment although the data structures are up 
for debate.  We could iterate on the data structures for this.  DAMON has 
a ton of additional support and we need to discuss what all comes along 
with it if this is a general tiering solution.  Gregory suggested the 
asynchronous promotion work could also be used by DAMON.

Joshua broke down the components of pghot: device drivers, kscand, 
kmigrated, and pghot interface.  If they get separated, it would also be 
useful for code reviews.  He was concerned that pghot may be heading in 
the same direction as DAMON.  Davidlohr agreed with this.

Bharata, have you thought about the possibility of only introducing the 
asynchronous promotion support first, getting alignment from stakeholders 
on that, and then extending pghot with additional enhancements later?

----->o-----
NOTE!!!  The next meeting will be canceled due to LPC 2026.

Next meeting will be on Thursday, October 22 at 8:30am PDT (UTC-7),
everybody is welcome: https://meet.google.com/jak-ytdx-hnm

Topics for the next meeting:

 - debrief discussions at LPC
 - directional discussion for pghot and how to build it incrementally:
   focus on async promotion mechanisms first and then add enhancements on
   top of this separately
 - update on combined patch series for supporting non-temporal stores in
   migrate_pages() with memory error handling (series from Yiannis +
   Huan) for 4KB, extensions for 2MB later from Shivank
 - v6 of Shivank's series for enlightening migrate_pages() for hardware
   assists and his rmap batch series as well as non-temporal stores 
   without overloading migrate_mode
 - Teja's update on SDXI page migration based on AMD patches and hardware
   issues being encountered that do not result in page migration
 - update on tier-aware memcg limits status and production testing based
   on the latest major overhaul
 - first class support for virtualization based memory tier support, how
   to leverge memory tiers in the guest
 - discuss generalized subsystem for providing bandwidth information
   independent of the underlying platform, ideally through resctrl
   + preferably this bandwidth monitoring is not per NUMA node but rather
     slow and fast

Please let me know if you'd like to propose additional topics for
discussion, thank you!

[1]
https://lore.kernel.org/linux-mm/20260902-migrate-refactor-shivank-v1-9-9dcca87669c4@amd.com/ 


             reply	other threads:[~2026-09-26 21:43 UTC|newest]

Thread overview: 2+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-26 21:43 David Rientjes [this message]
2026-09-27 15:56 ` [Linux Memory Hotness and Promotion] Notes from September 24, 2026 Bharata B Rao

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=0c8a7227-c7eb-bef1-7326-c7b04cde78a8@google.com \
    --to=rientjes@google.com \
    --cc=bharata@amd.com \
    --cc=dave@stgolabs.net \
    --cc=fvdl@google.com \
    --cc=gourry@gourry.net \
    --cc=jic23@kernel.org \
    --cc=joshua.hahnjy@gmail.com \
    --cc=linux-mm@kvack.org \
    --cc=nifan.cxl@gmail.com \
    --cc=rkodsara@amd.com \
    --cc=sj@kernel.org \
    --cc=weixugc@google.com \
    --cc=xuezhengchu@huawei.com \
    --cc=yiannis@zptcorp.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).