Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [Linux Memory Hotness and Promotion] Notes from September 10, 2026
@ 2026-09-12 23:15 David Rientjes
  0 siblings, 0 replies; only message in thread
From: David Rientjes @ 2026-09-12 23:15 UTC (permalink / raw)
  To: Davidlohr Bueso, Fan Ni, Frank van der Linden, Gregory Price,
	Jonathan Cameron, Joshua Hahn, Raghavendra K T,
	Rao, Bharata Bhasker, SeongJae Park, Wei Xu, Xuezheng Chu,
	Yiannis Nikolakopoulos, Zi Yan
  Cc: linux-mm

Hi everybody,

Here are the notes from the last Linux Memory Hotness and Promotion call
that happened on Thursday, September 10.  Thanks to everybody who was 
involved!

These notes are intended to bring people up to speed who could not attend 
the call as well as keep the conversation going in between meetings.

----->o-----
Gregory shared his progress on end-to-end compressed RAM and performance 
numbers obtained through qemu experimentation.  He will also be talking 
about this topic at LPC coming up next month.

The theme focused on a first principles approach for a memory architecture 
that includes devices that necessarily lie about their capacities.  For 
example, devices can tell Linux that they have 2GB of memory when they 
actually only have 1GB of memory because compression is enabled.  Some 
problems with this include releasing capacity back to the devices and 
including poisoning of memory completely (not just best effort 
prevention).

There are two ways to access this memory: files and user mappings.  Files 
go through the kernel direct map whereas user mappings go through the user 
page tables.  If you made a poor assumption about the compression ratio 
that you have, then accessible pages don't actually have a real backing.  
When you put file mapped memory on compressed RAM, you lose control over 
who can access it.

Writing memory causes physical memory to be loaded into a cache that gets 
invalidated and stored in a different area of memory in compressed form.  
Compression ratios can change out from under you, thus we're expecting 
Linux to operate on faith that we won't end up poisoning memory -- there 
is not great precident for subsystem components in Linux that operate 
based entirely on faith.  This is particularly concerning when dealing 
with memory.

How much capacity do we actually expose?  Zram and zswap have had this 
debate; in this case, we make the decision at boot.  With software 
compression, you intercept every access through page tables; with 
hardware, you don't have this and the result can be catastrophic.  When 
writing to a file, for example, the page allocator can continually hand 
out pages but this can result in poison of that memory.

Accounting is another problem: zram and zswap can charge the compressed 
memory back to the cgroup; this is important for oom killing.  We 
completely lose this accounting with compressed RAM since it's just a 
memcpy().  In the extreme, consider task A storing 1GB of incompressible 
memory that has a real usage of 1GB whereas task B storing 2GB of zero 
pages has a real usage of 0KB.  In this case, the latter would actually be 
selected for oom kill.

For memory poisoning of memory, this is detected at read.  The poison 
occurs on write, but it is detected on read.  In the future, interrupts 
may provide insight into this.

If you have the ability to cut off future writes and allocations, it is 
possible to avoid poison entirely; Gregory stated that we shouldn't try to 
migrate our way out of this problem.  How we do this for file backed 
memory and user mapped memory is entirely different.

----->o-----
If the entire device is saturated and the compression ratio drops over 
time, there may be an interrupt to inform us that we may need to start 
reclaiming.  If we have uncontended writes, we need to reclaim faster than 
the rate of writes; this is likely to be a losing proposition.  There's 
also the opposite problem where the device is largely empty but a user is 
writing /dev/random to the device -- the page allocator will continue 
giving out pages but they won't eventually be backed by anything.

Gregory has implemented a number of solutions for these problems and has 
approached it from first principles for isolation.  He suggests we need 
normal pages that are mappable and can be reclaimed; demotion should "just 
work" like zram or zswap does.  There should be no ZONE_DEVICE requirement 
to support this.  Additionally, Gregory suggested the memory should be 
mappable to read-only, we should support dynamic sizing (ballooning), and 
allocation control.  The kernel has to be able to control who can access 
the memory especially for writes.

For page cache, cleancache would have been preferable but is no longer 
available in the kernel.

Gregory provided a link to the github where this is implemented[1] using 
btrfs.

----->o-----
Gregory tested this work by setting up a system with lots of qemu 
instances on hardware with actual CXL memory.  He sought to prove two 
things: that his work does not impose an additional overhead on the kernel 
and that the page cache implementation is safe.  He primarily used NVMe 
based swap-in and zswap swap-in as the comparisons.  His implementation 
was much faster than both, closer to actual memory speeds.  He suggested 
that we could not compare this to standard DRAM that supports uncontended 
writes.

Gregory noted that for read-heavy workloads there was little to no 
performance implications and the additional capacity resulted in better 
throughput.  Performance dropped with a lot of writing that required 
promotions, which is intuitive.

Frank van der Linden noted that he'd been running experiments with similar 
devices and that he had largely come to the same conclusions, so he was 
very supportive of this work.  He asked about promotion on write when DRAM 
was under pressure.  Gregory suggested that we just need to reclaim in 
this case -- it's an allocation just like anything else, promotion is not 
special here.

Wei Xu noted that some users have the ability to kill jobs to free memory 
capacity, which may not be possible for all users.  Gregory noted that 
bandwidth on these devices are ~45GB/s, we might get 20-30GB/s if it's all 
writes.  If we have 1TB of compressed RAM, it's difficult to reason about 
being able to keep up with the pace of writes happening to the device and 
be able to make strong guarantees that we will not start poisoning memory.

----->o-----
Next meeting will be on Thursday, September 24 at 8:30am PDT (UTC-7),
everybody is welcome: https://meet.google.com/jak-ytdx-hnm

Topics for the next meeting:

 - update on combined patch series for supporting non-temporal stores in
   migrate_pages() with memory error handling (series from Yiannis +
   Huan)
 - v9 of pghot and the PTE A bit based source (kscand) for inclusion in
   the upstream kernel
 - v6 of Shivank's series for enlightening migrate_pages() for hardware
   assists and his rmap batch series and LPC discussion
 - Teja's update on SDXI page migration based on AMD patches and hardware
   issues being encountered that do not result in page migration
 - update on tier-aware memcg limits status and production testing based
   on the latest major overhaul
 - first class support for virtualization based memory tier support, how
   to leverge memory tiers in the guest
 - discuss generalized subsystem for providing bandwidth information
   independent of the underlying platform, ideally through resctrl,
   otherwise utilizing bandwidth information will be challenging
   + preferably this bandwidth monitoring is not per NUMA node but rather
     slow and fast

Please let me know if you'd like to propose additional topics for
discussion, thank you!

[1]
https://github.com/gourryinverse/linux/commits/scratch/gourry/cramtest/cram_72/


^ permalink raw reply	[flat|nested] only message in thread

only message in thread, other threads:[~2026-09-12 23:16 UTC | newest]

Thread overview: (only message) (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-12 23:15 [Linux Memory Hotness and Promotion] Notes from September 10, 2026 David Rientjes

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox