From: David Rientjes <rientjes@google.com>
To: Davidlohr Bueso <dave@stgolabs.net>, Fan Ni <nifan.cxl@gmail.com>,
Frank van der Linden <fvdl@google.com>,
Gregory Price <gourry@gourry.net>,
Jonathan Cameron <jic23@kernel.org>,
Joshua Hahn <joshua.hahnjy@gmail.com>,
Raghavendra K T <rkodsara@amd.com>,
"Rao, Bharata Bhasker" <bharata@amd.com>,
SeongJae Park <sj@kernel.org>, Wei Xu <weixugc@google.com>,
Xuezheng Chu <xuezhengchu@huawei.com>,
Yiannis Nikolakopoulos <yiannis@zptcorp.com>,
Zi Yan <ziy@nvidia.com>
Cc: linux-mm@kvack.org
Subject: [Linux Memory Hotness and Promotion] Notes from September 10, 2026
Date: Sat, 12 Sep 2026 16:15:52 -0700 (PDT) [thread overview]
Message-ID: <f5a9467b-1cdd-732c-a70b-625be9fc8dd5@google.com> (raw)
Hi everybody,
Here are the notes from the last Linux Memory Hotness and Promotion call
that happened on Thursday, September 10. Thanks to everybody who was
involved!
These notes are intended to bring people up to speed who could not attend
the call as well as keep the conversation going in between meetings.
----->o-----
Gregory shared his progress on end-to-end compressed RAM and performance
numbers obtained through qemu experimentation. He will also be talking
about this topic at LPC coming up next month.
The theme focused on a first principles approach for a memory architecture
that includes devices that necessarily lie about their capacities. For
example, devices can tell Linux that they have 2GB of memory when they
actually only have 1GB of memory because compression is enabled. Some
problems with this include releasing capacity back to the devices and
including poisoning of memory completely (not just best effort
prevention).
There are two ways to access this memory: files and user mappings. Files
go through the kernel direct map whereas user mappings go through the user
page tables. If you made a poor assumption about the compression ratio
that you have, then accessible pages don't actually have a real backing.
When you put file mapped memory on compressed RAM, you lose control over
who can access it.
Writing memory causes physical memory to be loaded into a cache that gets
invalidated and stored in a different area of memory in compressed form.
Compression ratios can change out from under you, thus we're expecting
Linux to operate on faith that we won't end up poisoning memory -- there
is not great precident for subsystem components in Linux that operate
based entirely on faith. This is particularly concerning when dealing
with memory.
How much capacity do we actually expose? Zram and zswap have had this
debate; in this case, we make the decision at boot. With software
compression, you intercept every access through page tables; with
hardware, you don't have this and the result can be catastrophic. When
writing to a file, for example, the page allocator can continually hand
out pages but this can result in poison of that memory.
Accounting is another problem: zram and zswap can charge the compressed
memory back to the cgroup; this is important for oom killing. We
completely lose this accounting with compressed RAM since it's just a
memcpy(). In the extreme, consider task A storing 1GB of incompressible
memory that has a real usage of 1GB whereas task B storing 2GB of zero
pages has a real usage of 0KB. In this case, the latter would actually be
selected for oom kill.
For memory poisoning of memory, this is detected at read. The poison
occurs on write, but it is detected on read. In the future, interrupts
may provide insight into this.
If you have the ability to cut off future writes and allocations, it is
possible to avoid poison entirely; Gregory stated that we shouldn't try to
migrate our way out of this problem. How we do this for file backed
memory and user mapped memory is entirely different.
----->o-----
If the entire device is saturated and the compression ratio drops over
time, there may be an interrupt to inform us that we may need to start
reclaiming. If we have uncontended writes, we need to reclaim faster than
the rate of writes; this is likely to be a losing proposition. There's
also the opposite problem where the device is largely empty but a user is
writing /dev/random to the device -- the page allocator will continue
giving out pages but they won't eventually be backed by anything.
Gregory has implemented a number of solutions for these problems and has
approached it from first principles for isolation. He suggests we need
normal pages that are mappable and can be reclaimed; demotion should "just
work" like zram or zswap does. There should be no ZONE_DEVICE requirement
to support this. Additionally, Gregory suggested the memory should be
mappable to read-only, we should support dynamic sizing (ballooning), and
allocation control. The kernel has to be able to control who can access
the memory especially for writes.
For page cache, cleancache would have been preferable but is no longer
available in the kernel.
Gregory provided a link to the github where this is implemented[1] using
btrfs.
----->o-----
Gregory tested this work by setting up a system with lots of qemu
instances on hardware with actual CXL memory. He sought to prove two
things: that his work does not impose an additional overhead on the kernel
and that the page cache implementation is safe. He primarily used NVMe
based swap-in and zswap swap-in as the comparisons. His implementation
was much faster than both, closer to actual memory speeds. He suggested
that we could not compare this to standard DRAM that supports uncontended
writes.
Gregory noted that for read-heavy workloads there was little to no
performance implications and the additional capacity resulted in better
throughput. Performance dropped with a lot of writing that required
promotions, which is intuitive.
Frank van der Linden noted that he'd been running experiments with similar
devices and that he had largely come to the same conclusions, so he was
very supportive of this work. He asked about promotion on write when DRAM
was under pressure. Gregory suggested that we just need to reclaim in
this case -- it's an allocation just like anything else, promotion is not
special here.
Wei Xu noted that some users have the ability to kill jobs to free memory
capacity, which may not be possible for all users. Gregory noted that
bandwidth on these devices are ~45GB/s, we might get 20-30GB/s if it's all
writes. If we have 1TB of compressed RAM, it's difficult to reason about
being able to keep up with the pace of writes happening to the device and
be able to make strong guarantees that we will not start poisoning memory.
----->o-----
Next meeting will be on Thursday, September 24 at 8:30am PDT (UTC-7),
everybody is welcome: https://meet.google.com/jak-ytdx-hnm
Topics for the next meeting:
- update on combined patch series for supporting non-temporal stores in
migrate_pages() with memory error handling (series from Yiannis +
Huan)
- v9 of pghot and the PTE A bit based source (kscand) for inclusion in
the upstream kernel
- v6 of Shivank's series for enlightening migrate_pages() for hardware
assists and his rmap batch series and LPC discussion
- Teja's update on SDXI page migration based on AMD patches and hardware
issues being encountered that do not result in page migration
- update on tier-aware memcg limits status and production testing based
on the latest major overhaul
- first class support for virtualization based memory tier support, how
to leverge memory tiers in the guest
- discuss generalized subsystem for providing bandwidth information
independent of the underlying platform, ideally through resctrl,
otherwise utilizing bandwidth information will be challenging
+ preferably this bandwidth monitoring is not per NUMA node but rather
slow and fast
Please let me know if you'd like to propose additional topics for
discussion, thank you!
[1]
https://github.com/gourryinverse/linux/commits/scratch/gourry/cramtest/cram_72/
reply other threads:[~2026-09-12 23:16 UTC|newest]
Thread overview: [no followups] expand[flat|nested] mbox.gz Atom feed
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=f5a9467b-1cdd-732c-a70b-625be9fc8dd5@google.com \
--to=rientjes@google.com \
--cc=bharata@amd.com \
--cc=dave@stgolabs.net \
--cc=fvdl@google.com \
--cc=gourry@gourry.net \
--cc=jic23@kernel.org \
--cc=joshua.hahnjy@gmail.com \
--cc=linux-mm@kvack.org \
--cc=nifan.cxl@gmail.com \
--cc=rkodsara@amd.com \
--cc=sj@kernel.org \
--cc=weixugc@google.com \
--cc=xuezhengchu@huawei.com \
--cc=yiannis@zptcorp.com \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox