Linux filesystem development
 help / color / mirror / Atom feed
From: Joanne Koong <joannelkoong@gmail.com>
To: akpm@linux-foundation.org, hannes@cmpxchg.org,
	shakeel.butt@linux.dev, roman.gushchin@linux.dev,
	willy@infradead.org, jack@suse.cz
Cc: mhocko@suse.com, muchun.song@linux.dev, david@kernel.org,
	ljs@kernel.org, vbabka@kernel.org, liam@infradead.org,
	rppt@kernel.org, surenb@google.com, riel@surriel.com,
	linux-mm@kvack.org, cgroups@vger.kernel.org,
	linux-fsdevel@vger.kernel.org
Subject: [PATCH v1 0/3] mm/readahead: avoid per-folio memcg reclaim
Date: Fri,  2 Oct 2026 17:15:52 -0700	[thread overview]
Message-ID: <20261003001555.3498357-1-joannelkoong@gmail.com> (raw)

Readahead adds the folios in its window to the page cache one at a time. Each
folio is charged separately. When the memcg is at its limit, each one of those
charges triggers reclaim. With many tasks faulting in the same cgroup, the
margin one reclaim pass frees gets consumed by the others, so the faulting
tasks keep reclaiming, all to make room for speculative folios, and in the
worst case, reclaim livelocks.

On Meta's fleet, this per-folio reclaim is a large cost. From fleet-wide CPU
profiles at Meta:

 - Page cache insertion (filemap_add_folio()) triggers about half of all
   memcg limit reclaim CPU, more than anonymous faults, swap-in and
   memory.high combined.

 - 93% of that comes from mmap fault readahead: filemap_fault() ->
   do_sync_mmap_readahead() -> page_cache_ra_unbounded() ->
   filemap_add_folio() -> mem_cgroup_charge() -> try_charge_memcg() ->
   try_to_free_mem_cgroup_pages().

 - In filemap_add_folio(), 83% of the CPU is memcg reclaim while the page
   cache insertion itself is 6%.

Most of it comes from services whose worker cgroups run at their limit while
many threads fault in mmapped files, mostly on btrfs, which our hosts mount
with compress-force=zstd:3.

This series adds readahead folios to the page cache without direct
reclaim. When the memcg is at its limit, readahead reclaims for the rest of
its window all at once instead of once per folio, and tries a second time if
the window runs out of room again. Please note that this only applies to
speculative readahead folios. For the folio a fault or read actually needs to
read in, it is still charged with the mapping's normal gfp mask (with the
reclaim flag set), like before.

On a 26-core/52-thread machine with btrfs, running 26, 52 or 104 processes
(one per core, one per thread, and 2x oversubscribed) that each mmap their own
file in one memcg with a 1G memory.max (before and after measured in the same
boot), reclaim passes per major fault drop by 94-96% for both random and
sequential reads. With data compressing ~3:1 under compress-force=zstd:3, the
runs that livelocked in reclaim without this series (4 of 18) no longer do.
Throughput with incompressible data, where the disk is the bottleneck, is
within 5% of before in either direction and the drops are within run-to-run
variation. With the compressed data, it is 9% higher at 104 processes, where
reclaim contention is.

More details on the results seen (medians of 3 runs of 10 secs, 5 for seq, rand)
are as follows:

Reclaim passes per major fault:
                       26 procs      52 procs         104 procs
cold, rand          23.2 -> 0.88    23.2 -> 0.88    23.1 -> 0.89
cold, z3            21.7 -> 0.95    20.0 -> 0.91    17.9 -> 0.85
seq,  rand          78.3 -> 3.66    45.8 -> 2.37    26.8 -> 1.24
seq,  z3            42.0 -> 2.59    37.2 -> 2.01    22.9 -> 1.24

Throughput (pages/s), after vs before:
                       26 procs      52 procs         104 procs
cold, rand            -5%             +5%             -5%  (disk-bound)
cold, z3              -3%             -1%             +9%
seq,  rand            -1%              0%             +4%

(read_ahead_kb=128, "cold" = read 64 pages from random offsets, "seq" = read
4096 sequential pages, "rand" data doesn't compress , "z3" data compresses
~3:1, which moves the bottleneck from the disk to the CPU).

Cold readers get the same readahead as before (pages per major fault within
2%). Sequential readers get up to 18% fewer pages per major fault, but their 
throughput and read bandwidth are unchanged or slightly higher, so readahead
pretty much does the same IO but in more and smaller pieces.

There were a few alternative approaches considered and tested in experiments:
 - don't reclaim in readahead at all:
   This cut reclaim to 0.02 passes per fault, but effectively disabled readahead
   in cgroups that live at their limit, and led to ~100x as many major faults.
   Each read request sent to storage was around 5KB instead of ~120KB.

 - charging the whole window up front:
   This leads to overcharging, as it charges for folios that may turn out to
   already be cached. With every other 16 page jjchunk of the file cached it
   triggered ~15x more reclaim passes per fault than this series.

 - One reclaim attempt per request:
   This led to the fewest reclaim passes (0.75 per cold fault), but at 52 procs
   10x as many requests end early and sequential readers lose 29-35% of their
   pages per fault. This is because with many tasks in one memcg, the room one
   reclaim attempt makes is often used up by other tasks before the window has
   completed, and with no second attempt, readahead stops early.

 - Two attempts per request without the immediate drain-and-retry:
   3x as many requests end early as this series and sequential pages per fault
   drop 16-22%.

 - Up to four attempts per request:
   At 52 procs on z3 data it gave 14% more pages per fault for 8% more reclaim
   than two attempts, but this is within the noise of the two attempts'
   run-to-run variation (13%)

Swap readahead charges the folios in its window one at a time in the same way,
so it can hit the same per-folio reclaim at the memcg limit. This will be
further investigated and followed-up on in a separate independent series.

This series was run through an LLM for sanity-checking / reviewing,
structuring the cover letter and improving the wording of commit messages,
helping generate some test scripts, and for bouncing ideas for different
alternative designs that could work better.

Joanne Koong (3):
  mm: memcontrol: factor reclaim logic out of try_charge_memcg()
  mm: memcontrol: add mem_cgroup_reclaim_for_batch()
  mm/readahead: avoid per-folio memcg reclaim

 include/linux/pagemap.h |   1 +
 mm/internal.h           |  10 +++
 mm/memcontrol.c         | 131 ++++++++++++++++++++++++++++++++--------
 mm/readahead.c          |  51 ++++++++++++++--
 4 files changed, 163 insertions(+), 30 deletions(-)

-- 
2.52.0


             reply	other threads:[~2026-10-03  0:18 UTC|newest]

Thread overview: 8+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-03  0:15 Joanne Koong [this message]
2026-10-03  0:15 ` [PATCH v1 1/3] mm: memcontrol: factor reclaim logic out of try_charge_memcg() Joanne Koong
2026-10-05 18:49   ` Rik van Riel
2026-10-03  0:15 ` [PATCH v1 2/3] mm: memcontrol: add mem_cgroup_reclaim_for_batch() Joanne Koong
2026-10-03  0:15 ` [PATCH v1 3/3] mm/readahead: avoid per-folio memcg reclaim Joanne Koong
2026-10-05 12:26 ` [PATCH v1 0/3] " Jan Kara
2026-10-06 14:01   ` Joanne Koong
2026-10-07 15:30     ` Jan Kara

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261003001555.3498357-1-joannelkoong@gmail.com \
    --to=joannelkoong@gmail.com \
    --cc=akpm@linux-foundation.org \
    --cc=cgroups@vger.kernel.org \
    --cc=david@kernel.org \
    --cc=hannes@cmpxchg.org \
    --cc=jack@suse.cz \
    --cc=liam@infradead.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=mhocko@suse.com \
    --cc=muchun.song@linux.dev \
    --cc=riel@surriel.com \
    --cc=roman.gushchin@linux.dev \
    --cc=rppt@kernel.org \
    --cc=shakeel.butt@linux.dev \
    --cc=surenb@google.com \
    --cc=vbabka@kernel.org \
    --cc=willy@infradead.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox