* [PATCH v1 0/3] mm/readahead: avoid per-folio memcg reclaim
@ 2026-10-03 0:15 Joanne Koong
2026-10-03 0:15 ` [PATCH v1 1/3] mm: memcontrol: factor reclaim logic out of try_charge_memcg() Joanne Koong
` (3 more replies)
0 siblings, 4 replies; 8+ messages in thread
From: Joanne Koong @ 2026-10-03 0:15 UTC (permalink / raw)
To: akpm, hannes, shakeel.butt, roman.gushchin, willy, jack
Cc: mhocko, muchun.song, david, ljs, vbabka, liam, rppt, surenb, riel,
linux-mm, cgroups, linux-fsdevel
Readahead adds the folios in its window to the page cache one at a time. Each
folio is charged separately. When the memcg is at its limit, each one of those
charges triggers reclaim. With many tasks faulting in the same cgroup, the
margin one reclaim pass frees gets consumed by the others, so the faulting
tasks keep reclaiming, all to make room for speculative folios, and in the
worst case, reclaim livelocks.
On Meta's fleet, this per-folio reclaim is a large cost. From fleet-wide CPU
profiles at Meta:
- Page cache insertion (filemap_add_folio()) triggers about half of all
memcg limit reclaim CPU, more than anonymous faults, swap-in and
memory.high combined.
- 93% of that comes from mmap fault readahead: filemap_fault() ->
do_sync_mmap_readahead() -> page_cache_ra_unbounded() ->
filemap_add_folio() -> mem_cgroup_charge() -> try_charge_memcg() ->
try_to_free_mem_cgroup_pages().
- In filemap_add_folio(), 83% of the CPU is memcg reclaim while the page
cache insertion itself is 6%.
Most of it comes from services whose worker cgroups run at their limit while
many threads fault in mmapped files, mostly on btrfs, which our hosts mount
with compress-force=zstd:3.
This series adds readahead folios to the page cache without direct
reclaim. When the memcg is at its limit, readahead reclaims for the rest of
its window all at once instead of once per folio, and tries a second time if
the window runs out of room again. Please note that this only applies to
speculative readahead folios. For the folio a fault or read actually needs to
read in, it is still charged with the mapping's normal gfp mask (with the
reclaim flag set), like before.
On a 26-core/52-thread machine with btrfs, running 26, 52 or 104 processes
(one per core, one per thread, and 2x oversubscribed) that each mmap their own
file in one memcg with a 1G memory.max (before and after measured in the same
boot), reclaim passes per major fault drop by 94-96% for both random and
sequential reads. With data compressing ~3:1 under compress-force=zstd:3, the
runs that livelocked in reclaim without this series (4 of 18) no longer do.
Throughput with incompressible data, where the disk is the bottleneck, is
within 5% of before in either direction and the drops are within run-to-run
variation. With the compressed data, it is 9% higher at 104 processes, where
reclaim contention is.
More details on the results seen (medians of 3 runs of 10 secs, 5 for seq, rand)
are as follows:
Reclaim passes per major fault:
26 procs 52 procs 104 procs
cold, rand 23.2 -> 0.88 23.2 -> 0.88 23.1 -> 0.89
cold, z3 21.7 -> 0.95 20.0 -> 0.91 17.9 -> 0.85
seq, rand 78.3 -> 3.66 45.8 -> 2.37 26.8 -> 1.24
seq, z3 42.0 -> 2.59 37.2 -> 2.01 22.9 -> 1.24
Throughput (pages/s), after vs before:
26 procs 52 procs 104 procs
cold, rand -5% +5% -5% (disk-bound)
cold, z3 -3% -1% +9%
seq, rand -1% 0% +4%
(read_ahead_kb=128, "cold" = read 64 pages from random offsets, "seq" = read
4096 sequential pages, "rand" data doesn't compress , "z3" data compresses
~3:1, which moves the bottleneck from the disk to the CPU).
Cold readers get the same readahead as before (pages per major fault within
2%). Sequential readers get up to 18% fewer pages per major fault, but their
throughput and read bandwidth are unchanged or slightly higher, so readahead
pretty much does the same IO but in more and smaller pieces.
There were a few alternative approaches considered and tested in experiments:
- don't reclaim in readahead at all:
This cut reclaim to 0.02 passes per fault, but effectively disabled readahead
in cgroups that live at their limit, and led to ~100x as many major faults.
Each read request sent to storage was around 5KB instead of ~120KB.
- charging the whole window up front:
This leads to overcharging, as it charges for folios that may turn out to
already be cached. With every other 16 page jjchunk of the file cached it
triggered ~15x more reclaim passes per fault than this series.
- One reclaim attempt per request:
This led to the fewest reclaim passes (0.75 per cold fault), but at 52 procs
10x as many requests end early and sequential readers lose 29-35% of their
pages per fault. This is because with many tasks in one memcg, the room one
reclaim attempt makes is often used up by other tasks before the window has
completed, and with no second attempt, readahead stops early.
- Two attempts per request without the immediate drain-and-retry:
3x as many requests end early as this series and sequential pages per fault
drop 16-22%.
- Up to four attempts per request:
At 52 procs on z3 data it gave 14% more pages per fault for 8% more reclaim
than two attempts, but this is within the noise of the two attempts'
run-to-run variation (13%)
Swap readahead charges the folios in its window one at a time in the same way,
so it can hit the same per-folio reclaim at the memcg limit. This will be
further investigated and followed-up on in a separate independent series.
This series was run through an LLM for sanity-checking / reviewing,
structuring the cover letter and improving the wording of commit messages,
helping generate some test scripts, and for bouncing ideas for different
alternative designs that could work better.
Joanne Koong (3):
mm: memcontrol: factor reclaim logic out of try_charge_memcg()
mm: memcontrol: add mem_cgroup_reclaim_for_batch()
mm/readahead: avoid per-folio memcg reclaim
include/linux/pagemap.h | 1 +
mm/internal.h | 10 +++
mm/memcontrol.c | 131 ++++++++++++++++++++++++++++++++--------
mm/readahead.c | 51 ++++++++++++++--
4 files changed, 163 insertions(+), 30 deletions(-)
--
2.52.0
^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH v1 1/3] mm: memcontrol: factor reclaim logic out of try_charge_memcg()
2026-10-03 0:15 [PATCH v1 0/3] mm/readahead: avoid per-folio memcg reclaim Joanne Koong
@ 2026-10-03 0:15 ` Joanne Koong
2026-10-05 18:49 ` Rik van Riel
2026-10-03 0:15 ` [PATCH v1 2/3] mm: memcontrol: add mem_cgroup_reclaim_for_batch() Joanne Koong
` (2 subsequent siblings)
3 siblings, 1 reply; 8+ messages in thread
From: Joanne Koong @ 2026-10-03 0:15 UTC (permalink / raw)
To: akpm, hannes, shakeel.butt, roman.gushchin, willy, jack
Cc: mhocko, muchun.song, david, ljs, vbabka, liam, rppt, surenb, riel,
linux-mm, cgroups, linux-fsdevel
Move reclaim logic in try_charge_memcg() into helper functions so that
the next patch can reuse them.
In memcg_charge_reclaim(), we can call memcg_memory_event() directly
instead of passing in allow_spinning. allow_spinning must always be true
here since memcg_charge_may_reclaim() has already checked that the gfp
mask allows blocking, which implies spinning is also allowed.
No functional changes.
Signed-off-by: Joanne Koong <joannelkoong@gmail.com>
---
mm/memcontrol.c | 67 +++++++++++++++++++++++++++++++------------------
1 file changed, 43 insertions(+), 24 deletions(-)
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index aad0498a7bd6..1d4b603085a5 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -2696,6 +2696,46 @@ void __mem_cgroup_handle_over_high(gfp_t gfp_mask)
css_put(&memcg->css);
}
+static bool memcg_charge_may_reclaim(gfp_t gfp_mask)
+{
+ if (unlikely(task_in_memcg_oom(current)))
+ return false;
+
+ if (!gfpflags_allow_blocking(gfp_mask))
+ return false;
+
+ /*
+ * OOM victim still needs to charge memory to exit. OOM reaper should
+ * help but it might fail on mmap_lock contention. If the victim is a
+ * large thread group then all exiting threads might compete on oom_lock
+ * just to learn that there is nothing really killable anymore. Bail
+ * out early and fail the charge to expedite their exit. They are
+ * considered fully reclaimed by the oom reaper and they shouldn't
+ * contribute further charges.
+ */
+ if (tsk_is_oom_victim(current) &&
+ mm_flags_test(MMF_OOM_SKIP, current->signal->oom_mm))
+ return false;
+
+ return true;
+}
+
+static unsigned long memcg_charge_reclaim(struct mem_cgroup *memcg,
+ unsigned long nr_pages,
+ gfp_t gfp_mask,
+ unsigned int reclaim_options)
+{
+ unsigned long nr_reclaimed, pflags;
+
+ memcg_memory_event(memcg, MEMCG_MAX);
+ psi_memstall_enter(&pflags);
+ nr_reclaimed = try_to_free_mem_cgroup_pages(memcg, nr_pages, gfp_mask,
+ reclaim_options, NULL);
+ psi_memstall_leave(&pflags);
+
+ return nr_reclaimed;
+}
+
static int try_charge_memcg(struct mem_cgroup *memcg, gfp_t gfp_mask,
unsigned int nr_pages)
{
@@ -2708,7 +2748,6 @@ static int try_charge_memcg(struct mem_cgroup *memcg, gfp_t gfp_mask,
unsigned int reclaim_options;
bool drained = false;
bool raised_max_event = false;
- unsigned long pflags;
bool allow_spinning = gfpflags_allow_spinning(gfp_mask);
int ret = 0;
@@ -2747,33 +2786,13 @@ static int try_charge_memcg(struct mem_cgroup *memcg, gfp_t gfp_mask,
if (unlikely(current->flags & PF_MEMALLOC))
goto force;
- if (unlikely(task_in_memcg_oom(current)))
- goto nomem;
-
- if (!gfpflags_allow_blocking(gfp_mask))
- goto nomem;
-
- /*
- * OOM victim still needs to charge memory to exit. OOM reaper should
- * help but it might fail on mmap_lock contention. If the victim is a
- * large thread group then all exiting threads might compete on oom_lock
- * just to learn that there is nothing really killable anymore. Bail
- * out early and fail the charge to expedite their exit. They are
- * considered fully reclaimed by the oom reaper and they shouldn't
- * contribute further charges.
- */
- if (tsk_is_oom_victim(current) &&
- mm_flags_test(MMF_OOM_SKIP, current->signal->oom_mm))
+ if (!memcg_charge_may_reclaim(gfp_mask))
goto nomem;
- __memcg_memory_event(mem_over_limit, MEMCG_MAX, allow_spinning);
+ nr_reclaimed = memcg_charge_reclaim(mem_over_limit, nr_pages, gfp_mask,
+ reclaim_options);
raised_max_event = true;
- psi_memstall_enter(&pflags);
- nr_reclaimed = try_to_free_mem_cgroup_pages(mem_over_limit, nr_pages,
- gfp_mask, reclaim_options, NULL);
- psi_memstall_leave(&pflags);
-
if (mem_cgroup_margin(mem_over_limit) >= nr_pages)
goto retry;
--
2.52.0
^ permalink raw reply related [flat|nested] 8+ messages in thread
* [PATCH v1 2/3] mm: memcontrol: add mem_cgroup_reclaim_for_batch()
2026-10-03 0:15 [PATCH v1 0/3] mm/readahead: avoid per-folio memcg reclaim Joanne Koong
2026-10-03 0:15 ` [PATCH v1 1/3] mm: memcontrol: factor reclaim logic out of try_charge_memcg() Joanne Koong
@ 2026-10-03 0:15 ` Joanne Koong
2026-10-03 0:15 ` [PATCH v1 3/3] mm/readahead: avoid per-folio memcg reclaim Joanne Koong
2026-10-05 12:26 ` [PATCH v1 0/3] " Jan Kara
3 siblings, 0 replies; 8+ messages in thread
From: Joanne Koong @ 2026-10-03 0:15 UTC (permalink / raw)
To: akpm, hannes, shakeel.butt, roman.gushchin, willy, jack
Cc: mhocko, muchun.song, david, ljs, vbabka, liam, rppt, surenb, riel,
linux-mm, cgroups, linux-fsdevel
Some callers charge a batch of folios one at a time, where each charge
may reclaim. At the memcg limit, each one of those charges then enters
reclaim separately. Readahead is the main example. This was observed on
Meta's fleet, where a single fault's readahead window can trigger more
than a dozen reclaim passes.
Add mem_cgroup_reclaim_for_batch() so that such callers can reclaim for
a whole batch at once instead of once per folio. Callers can then charge
without reclaim and if a charge fails, call
mem_cgroup_reclaim_for_batch() to make room for the rest of the batch.
Unlike try_charge_memcg(), it retries reclaim only once rather than up
to MAX_RECLAIM_RETRIES since failing is acceptable for the callers'
charges.
The next patch uses it for readahead.
Signed-off-by: Joanne Koong <joannelkoong@gmail.com>
---
mm/internal.h | 10 ++++++++
mm/memcontrol.c | 64 +++++++++++++++++++++++++++++++++++++++++++++++++
2 files changed, 74 insertions(+)
diff --git a/mm/internal.h b/mm/internal.h
index ff28b940e0cd..6d914f6a6138 100644
--- a/mm/internal.h
+++ b/mm/internal.h
@@ -1652,4 +1652,14 @@ static inline bool can_spin_trylock(void)
/* char-mem.c */
bool file_is_dev_zero(const struct file *file);
+#ifdef CONFIG_MEMCG
+bool mem_cgroup_reclaim_for_batch(unsigned long nr_pages, gfp_t gfp);
+#else
+static inline bool mem_cgroup_reclaim_for_batch(unsigned long nr_pages,
+ gfp_t gfp)
+{
+ return false;
+}
+#endif
+
#endif /* __MM_INTERNAL_H */
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index 1d4b603085a5..9aea684ea755 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -2736,6 +2736,70 @@ static unsigned long memcg_charge_reclaim(struct mem_cgroup *memcg,
return nr_reclaimed;
}
+/**
+ * mem_cgroup_reclaim_for_batch - make room for a batch of upcoming charges
+ * @nr_pages: number of pages about to be charged
+ * @gfp: reclaim context
+ *
+ * Reclaim is needed when a memcg's margin (limit minus usage) is less than
+ * @nr_pages. Starting from the memcg that the current task's charges go to,
+ * reclaim @nr_pages from the closest ancestor whose margin is too small, with
+ * at most one retry. This lets callers that charge a batch of folios one at a
+ * time (like readahead) reclaim @nr_pages at once instead of once per folio.
+ * Callers must be ok with the charges failing.
+ *
+ * Returns true if reclaim was needed and done or false if no reclaim was
+ * needed or if the current task may not reclaim.
+ */
+bool mem_cgroup_reclaim_for_batch(unsigned long nr_pages, gfp_t gfp)
+{
+ unsigned int reclaim_options = MEMCG_RECLAIM_MAY_SWAP;
+ struct mem_cgroup *orig, *memcg;
+
+ if (mem_cgroup_disabled() || !memcg_charge_may_reclaim(gfp))
+ return false;
+
+ orig = get_mem_cgroup_from_mm(NULL);
+ for (memcg = orig; !mem_cgroup_is_root(memcg);
+ memcg = parent_mem_cgroup(memcg))
+ if (mem_cgroup_margin(memcg) < nr_pages)
+ break;
+
+ /* No reclaim needed. No memcg up to the root lacks the margin */
+ if (mem_cgroup_is_root(memcg)) {
+ css_put(&orig->css);
+ return false;
+ }
+
+ /*
+ * cgroup v1 can limit memory+swap together (memsw). If memsw is the
+ * limit being hit, swapping a page out doesn't help, so reclaim without
+ * swapping.
+ */
+ if (do_memsw_account() &&
+ page_counter_read(&memcg->memsw) + nr_pages >
+ READ_ONCE(memcg->memsw.max))
+ reclaim_options &= ~MEMCG_RECLAIM_MAY_SWAP;
+
+ memcg_charge_reclaim(memcg, nr_pages, gfp, reclaim_options);
+ /*
+ * Like try_charge_memcg(), if reclaiming didn't make enough room, drain
+ * the per-cpu stocks and if that doesn't help enough either, try
+ * reclaiming again. Other tasks charging the same memcg may have used
+ * up what was just reclaimed. Unlike try_charge_memcg(), only retry
+ * once rather than up to MAX_RECLAIM_RETRIES times, since failing is
+ * acceptable for the caller's charges.
+ */
+ if (mem_cgroup_margin(memcg) < nr_pages) {
+ drain_all_stock(memcg);
+ if (mem_cgroup_margin(memcg) < nr_pages)
+ memcg_charge_reclaim(memcg, nr_pages, gfp,
+ reclaim_options);
+ }
+ css_put(&orig->css);
+ return true;
+}
+
static int try_charge_memcg(struct mem_cgroup *memcg, gfp_t gfp_mask,
unsigned int nr_pages)
{
--
2.52.0
^ permalink raw reply related [flat|nested] 8+ messages in thread
* [PATCH v1 3/3] mm/readahead: avoid per-folio memcg reclaim
2026-10-03 0:15 [PATCH v1 0/3] mm/readahead: avoid per-folio memcg reclaim Joanne Koong
2026-10-03 0:15 ` [PATCH v1 1/3] mm: memcontrol: factor reclaim logic out of try_charge_memcg() Joanne Koong
2026-10-03 0:15 ` [PATCH v1 2/3] mm: memcontrol: add mem_cgroup_reclaim_for_batch() Joanne Koong
@ 2026-10-03 0:15 ` Joanne Koong
2026-10-05 12:26 ` [PATCH v1 0/3] " Jan Kara
3 siblings, 0 replies; 8+ messages in thread
From: Joanne Koong @ 2026-10-03 0:15 UTC (permalink / raw)
To: akpm, hannes, shakeel.butt, roman.gushchin, willy, jack
Cc: mhocko, muchun.song, david, ljs, vbabka, liam, rppt, surenb, riel,
linux-mm, cgroups, linux-fsdevel
Readahead adds the folios in its window to the page cache one at a time.
Each folio is charged separately. When the memcg is at its limit, each
one of those charges triggers reclaim (readahead's __GFP_NORETRY only
stops try_charge_memcg() after a reclaim pass). With many tasks faulting
in the same cgroup, the margin one reclaim pass frees gets consumed by
the others, so the faulting tasks keep reclaiming, all to make room for
speculative folios, and in the worst case, reclaim livelocks.
On Meta's fleet, this per-folio reclaim is a large cost. Page cache
insertion triggers around half of all memcg limit reclaim CPU, 93% of it
from mmap fault readahead, in services whose cgroups run at their limit
while many threads fault in mmapped files. Of the CPU time spent in
filemap_add_folio(), 83% is memcg reclaim while inserting the folio into
the page cache itself is 6%.
To avoid this, add readahead folios to the page cache without direct
reclaim, and when the memcg is at its limit, reclaim for the rest of the
window at once with mem_cgroup_reclaim_for_batch() instead of once per
folio. If the memcg hits its limit again before the whole window has
been added, try one more time, since other tasks charging the same memcg
may have used up the room the first attempt made. If there's still no
room, stop readahead early rather than reclaim harder, since readahead
folios are speculative. The count of reclaim attempts lives in
readahead_control so that it covers every path that adds folios for the
request.
This only affects speculative readahead folios. There's no change in
behavior for the folio a fault or read actually needs. If readahead
didn't bring that folio in, filemap_fault() and filemap_read() allocate
and charge it with the mapping's normal gfp mask, like before.
There are two other differences from the current behavior. The first is
that readahead charges that push a cgroup past memory.high leave the
high reclaim to the return to user space, as other non-blocking charges
do. The second is that under global memory pressure, readahead's page
cache xarray node allocations no longer enter direct reclaim (they still
wake kswapd). If one fails, readahead stops early, as it does when a
folio allocation fails.
Tested on a 26-core/52-thread machine with btrfs, running 26, 52 or 104
processes (one per core, one per thread, and 2x oversubscribed) that
each mmap their own file in one memcg with a 1G memory.max (before and
after measured in the same boot), reclaim passes per major fault drop by
94-96% for both random and sequential reads. With data compressing ~3:1
under compress-force=zstd:3, the runs that livelocked in reclaim without
the patch (4 of 18) no longer do. Throughput with incompressible data,
where the disk is the bottleneck, is within 5% of before in either
direction, and the drops are within run-to-run variation. With the
compressed data, it is 9% higher at 104 processes, where reclaim
contention is. Sequential readers get up to 18% fewer pages per major
fault, since readahead now stops early when there is still no room, but
their throughput and read bandwidth are unchanged or slightly higher.
Reported-by: Rik van Riel <riel@surriel.com>
Signed-off-by: Joanne Koong <joannelkoong@gmail.com>
---
include/linux/pagemap.h | 1 +
mm/readahead.c | 51 ++++++++++++++++++++++++++++++++++++-----
2 files changed, 46 insertions(+), 6 deletions(-)
diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h
index 73af18a37367..973a836076bd 100644
--- a/include/linux/pagemap.h
+++ b/include/linux/pagemap.h
@@ -1409,6 +1409,7 @@ struct readahead_control {
pgoff_t _index;
unsigned int _nr_pages;
unsigned int _batch_count;
+ unsigned int _nr_memcg_reclaims;
bool dropbehind;
bool _workingset;
unsigned long _pflags;
diff --git a/mm/readahead.c b/mm/readahead.c
index 6e5563290287..e06375559e9b 100644
--- a/mm/readahead.c
+++ b/mm/readahead.c
@@ -204,6 +204,41 @@ static struct folio *ractl_alloc_folio(struct readahead_control *ractl,
return folio;
}
+/*
+ * Max number of memcg reclaim attempts per readahead request. The first
+ * makes room for the rest of the readahead window. If the window runs out
+ * of room again, try one more time since other readers in the same memcg
+ * might have used that room up.
+ */
+#define READAHEAD_MAX_MEMCG_RECLAIMS 2
+
+/*
+ * Add a readahead folio to the page cache without direct reclaim. This prevents
+ * a memcg at its limit from running reclaim for every folio in the readahead
+ * window (folios in the window are charged one at a time). If adding the folio
+ * to the page cache returns -ENOMEM, reclaim enough room for the rest of the
+ * window and retry adding it.
+ */
+static int readahead_add_folio(struct readahead_control *ractl,
+ struct folio *folio, pgoff_t index,
+ unsigned long nr_pages_left, gfp_t gfp)
+{
+ struct address_space *mapping = ractl->mapping;
+ int ret;
+
+ ret = filemap_add_folio(mapping, folio, index,
+ gfp & ~__GFP_DIRECT_RECLAIM);
+ if (ret == -ENOMEM &&
+ ractl->_nr_memcg_reclaims < READAHEAD_MAX_MEMCG_RECLAIMS) {
+ ractl->_nr_memcg_reclaims++;
+ nr_pages_left = max(nr_pages_left, folio_nr_pages(folio));
+ if (mem_cgroup_reclaim_for_batch(nr_pages_left, gfp))
+ ret = filemap_add_folio(mapping, folio, index,
+ gfp & ~__GFP_DIRECT_RECLAIM);
+ }
+ return ret;
+}
+
/**
* page_cache_ra_unbounded - Start unchecked readahead.
* @ractl: Readahead control.
@@ -290,7 +325,8 @@ void page_cache_ra_unbounded(struct readahead_control *ractl,
if (!folio)
break;
- ret = filemap_add_folio(mapping, folio, index + i, gfp_mask);
+ ret = readahead_add_folio(ractl, folio, index + i,
+ nr_to_read - i, gfp_mask);
if (ret < 0) {
folio_put(folio);
if (ret == -ENOMEM)
@@ -457,7 +493,7 @@ static unsigned long get_next_ra_size(struct file_ra_state *ra,
*/
static inline int ra_alloc_folio(struct readahead_control *ractl, pgoff_t index,
- pgoff_t mark, unsigned int order, gfp_t gfp)
+ pgoff_t mark, pgoff_t limit, unsigned int order, gfp_t gfp)
{
int err;
struct folio *folio = ractl_alloc_folio(ractl, gfp, order);
@@ -467,7 +503,7 @@ static inline int ra_alloc_folio(struct readahead_control *ractl, pgoff_t index,
mark = round_down(mark, 1UL << order);
if (index == mark)
folio_set_readahead(folio);
- err = filemap_add_folio(ractl->mapping, folio, index, gfp);
+ err = readahead_add_folio(ractl, folio, index, limit - index + 1, gfp);
if (err) {
folio_put(folio);
return err;
@@ -532,7 +568,7 @@ void page_cache_ra_order(struct readahead_control *ractl,
/* Don't allocate pages past EOF */
while (order > min_order && index + (1UL << order) - 1 > limit)
order--;
- err = ra_alloc_folio(ractl, index, mark, order, gfp);
+ err = ra_alloc_folio(ractl, index, mark, limit, order, gfp);
if (err)
break;
index += 1UL << order;
@@ -813,7 +849,8 @@ void readahead_expand(struct readahead_control *ractl,
return;
index = mapping_align_index(mapping, index);
- if (filemap_add_folio(mapping, folio, index, gfp_mask) < 0) {
+ if (readahead_add_folio(ractl, folio, index,
+ ractl->_index - new_index, gfp_mask) < 0) {
folio_put(folio);
return;
}
@@ -842,7 +879,9 @@ void readahead_expand(struct readahead_control *ractl,
return;
index = mapping_align_index(mapping, index);
- if (filemap_add_folio(mapping, folio, index, gfp_mask) < 0) {
+ if (readahead_add_folio(ractl, folio, index,
+ new_nr_pages - ractl->_nr_pages,
+ gfp_mask) < 0) {
folio_put(folio);
return;
}
--
2.52.0
^ permalink raw reply related [flat|nested] 8+ messages in thread
* Re: [PATCH v1 0/3] mm/readahead: avoid per-folio memcg reclaim
2026-10-03 0:15 [PATCH v1 0/3] mm/readahead: avoid per-folio memcg reclaim Joanne Koong
` (2 preceding siblings ...)
2026-10-03 0:15 ` [PATCH v1 3/3] mm/readahead: avoid per-folio memcg reclaim Joanne Koong
@ 2026-10-05 12:26 ` Jan Kara
2026-10-06 14:01 ` Joanne Koong
3 siblings, 1 reply; 8+ messages in thread
From: Jan Kara @ 2026-10-05 12:26 UTC (permalink / raw)
To: Joanne Koong
Cc: akpm, hannes, shakeel.butt, roman.gushchin, willy, jack, mhocko,
muchun.song, david, ljs, vbabka, liam, rppt, surenb, riel,
linux-mm, cgroups, linux-fsdevel
Hello Joanne!
On Fri 02-10-26 17:15:52, Joanne Koong wrote:
> Readahead adds the folios in its window to the page cache one at a time. Each
> folio is charged separately. When the memcg is at its limit, each one of those
> charges triggers reclaim. With many tasks faulting in the same cgroup, the
> margin one reclaim pass frees gets consumed by the others, so the faulting
> tasks keep reclaiming, all to make room for speculative folios, and in the
> worst case, reclaim livelocks.
>
> On Meta's fleet, this per-folio reclaim is a large cost. From fleet-wide CPU
> profiles at Meta:
>
> - Page cache insertion (filemap_add_folio()) triggers about half of all
> memcg limit reclaim CPU, more than anonymous faults, swap-in and
> memory.high combined.
>
> - 93% of that comes from mmap fault readahead: filemap_fault() ->
> do_sync_mmap_readahead() -> page_cache_ra_unbounded() ->
> filemap_add_folio() -> mem_cgroup_charge() -> try_charge_memcg() ->
> try_to_free_mem_cgroup_pages().
>
> - In filemap_add_folio(), 83% of the CPU is memcg reclaim while the page
> cache insertion itself is 6%.
>
> Most of it comes from services whose worker cgroups run at their limit while
> many threads fault in mmapped files, mostly on btrfs, which our hosts mount
> with compress-force=zstd:3.
>
> This series adds readahead folios to the page cache without direct
> reclaim. When the memcg is at its limit, readahead reclaims for the rest of
> its window all at once instead of once per folio, and tries a second time if
> the window runs out of room again. Please note that this only applies to
> speculative readahead folios. For the folio a fault or read actually needs to
> read in, it is still charged with the mapping's normal gfp mask (with the
> reclaim flag set), like before.
I cannot really comment on how much sense this makes from memcg POV (I'll
defer to people understanding that area better). When I'm reading your
problem description I'm wondering about one thing: Are the folios we are
faulting around really used in your workloads before they get reclaimed?
Because what I'd expect to happen in cases under heavy memory pressure is
that ra->mmap_miss would quickly grow over the limit (100) and we'd just
stop doing mmap readahead for the files. If the read-around pages aren't
getting used before being evicted, then the first natural thing would be to
actually stop reading them in the first place :) The mmap_miss logic is
actually rather dumb so I wouldn't be surprised if it doesn't work in your
case...
The experiments you have conducted below indicate that if you effectively
disable readahead you get more page faults & smaller IO requests (expected)
but for your realistic loads under memory pressure is that really an
overall performance loss?
Honza
>
> On a 26-core/52-thread machine with btrfs, running 26, 52 or 104 processes
> (one per core, one per thread, and 2x oversubscribed) that each mmap their own
> file in one memcg with a 1G memory.max (before and after measured in the same
> boot), reclaim passes per major fault drop by 94-96% for both random and
> sequential reads. With data compressing ~3:1 under compress-force=zstd:3, the
> runs that livelocked in reclaim without this series (4 of 18) no longer do.
> Throughput with incompressible data, where the disk is the bottleneck, is
> within 5% of before in either direction and the drops are within run-to-run
> variation. With the compressed data, it is 9% higher at 104 processes, where
> reclaim contention is.
>
> More details on the results seen (medians of 3 runs of 10 secs, 5 for seq, rand)
> are as follows:
>
> Reclaim passes per major fault:
> 26 procs 52 procs 104 procs
> cold, rand 23.2 -> 0.88 23.2 -> 0.88 23.1 -> 0.89
> cold, z3 21.7 -> 0.95 20.0 -> 0.91 17.9 -> 0.85
> seq, rand 78.3 -> 3.66 45.8 -> 2.37 26.8 -> 1.24
> seq, z3 42.0 -> 2.59 37.2 -> 2.01 22.9 -> 1.24
>
> Throughput (pages/s), after vs before:
> 26 procs 52 procs 104 procs
> cold, rand -5% +5% -5% (disk-bound)
> cold, z3 -3% -1% +9%
> seq, rand -1% 0% +4%
>
> (read_ahead_kb=128, "cold" = read 64 pages from random offsets, "seq" = read
> 4096 sequential pages, "rand" data doesn't compress , "z3" data compresses
> ~3:1, which moves the bottleneck from the disk to the CPU).
>
> Cold readers get the same readahead as before (pages per major fault within
> 2%). Sequential readers get up to 18% fewer pages per major fault, but their
> throughput and read bandwidth are unchanged or slightly higher, so readahead
> pretty much does the same IO but in more and smaller pieces.
>
> There were a few alternative approaches considered and tested in experiments:
> - don't reclaim in readahead at all:
> This cut reclaim to 0.02 passes per fault, but effectively disabled readahead
> in cgroups that live at their limit, and led to ~100x as many major faults.
> Each read request sent to storage was around 5KB instead of ~120KB.
>
> - charging the whole window up front:
> This leads to overcharging, as it charges for folios that may turn out to
> already be cached. With every other 16 page jjchunk of the file cached it
> triggered ~15x more reclaim passes per fault than this series.
>
> - One reclaim attempt per request:
> This led to the fewest reclaim passes (0.75 per cold fault), but at 52 procs
> 10x as many requests end early and sequential readers lose 29-35% of their
> pages per fault. This is because with many tasks in one memcg, the room one
> reclaim attempt makes is often used up by other tasks before the window has
> completed, and with no second attempt, readahead stops early.
>
> - Two attempts per request without the immediate drain-and-retry:
> 3x as many requests end early as this series and sequential pages per fault
> drop 16-22%.
>
> - Up to four attempts per request:
> At 52 procs on z3 data it gave 14% more pages per fault for 8% more reclaim
> than two attempts, but this is within the noise of the two attempts'
> run-to-run variation (13%)
>
> Swap readahead charges the folios in its window one at a time in the same way,
> so it can hit the same per-folio reclaim at the memcg limit. This will be
> further investigated and followed-up on in a separate independent series.
>
> This series was run through an LLM for sanity-checking / reviewing,
> structuring the cover letter and improving the wording of commit messages,
> helping generate some test scripts, and for bouncing ideas for different
> alternative designs that could work better.
>
> Joanne Koong (3):
> mm: memcontrol: factor reclaim logic out of try_charge_memcg()
> mm: memcontrol: add mem_cgroup_reclaim_for_batch()
> mm/readahead: avoid per-folio memcg reclaim
>
> include/linux/pagemap.h | 1 +
> mm/internal.h | 10 +++
> mm/memcontrol.c | 131 ++++++++++++++++++++++++++++++++--------
> mm/readahead.c | 51 ++++++++++++++--
> 4 files changed, 163 insertions(+), 30 deletions(-)
>
> --
> 2.52.0
>
--
Jan Kara <jack@suse.com>
SUSE Labs, CR
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH v1 1/3] mm: memcontrol: factor reclaim logic out of try_charge_memcg()
2026-10-03 0:15 ` [PATCH v1 1/3] mm: memcontrol: factor reclaim logic out of try_charge_memcg() Joanne Koong
@ 2026-10-05 18:49 ` Rik van Riel
0 siblings, 0 replies; 8+ messages in thread
From: Rik van Riel @ 2026-10-05 18:49 UTC (permalink / raw)
To: Joanne Koong, akpm, hannes, shakeel.butt, roman.gushchin, willy,
jack
Cc: mhocko, muchun.song, david, ljs, vbabka, liam, rppt, surenb,
linux-mm, cgroups, linux-fsdevel
On Fri, 2026-10-02 at 17:15 -0700, Joanne Koong wrote:
> Move reclaim logic in try_charge_memcg() into helper functions so
> that
> the next patch can reuse them.
>
> In memcg_charge_reclaim(), we can call memcg_memory_event() directly
> instead of passing in allow_spinning. allow_spinning must always be
> true
> here since memcg_charge_may_reclaim() has already checked that the
> gfp
> mask allows blocking, which implies spinning is also allowed.
>
> No functional changes.
>
> Signed-off-by: Joanne Koong <joannelkoong@gmail.com>
>
Nice cleanup.
Reviewed-by: Rik van Riel <riel@surriel.com>
--
All Rights Reversed.
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH v1 0/3] mm/readahead: avoid per-folio memcg reclaim
2026-10-05 12:26 ` [PATCH v1 0/3] " Jan Kara
@ 2026-10-06 14:01 ` Joanne Koong
2026-10-07 15:30 ` Jan Kara
0 siblings, 1 reply; 8+ messages in thread
From: Joanne Koong @ 2026-10-06 14:01 UTC (permalink / raw)
To: Jan Kara
Cc: akpm, hannes, shakeel.butt, roman.gushchin, willy, mhocko,
muchun.song, david, ljs, vbabka, liam, rppt, surenb, riel,
linux-mm, cgroups, linux-fsdevel, open list:BTRFS FILE SYSTEM
Hi Jan!
Thanks for taking a look at this.
On Mon, Oct 5, 2026 at 2:27 PM Jan Kara <jack@suse.cz> wrote:
>
> Hello Joanne!
>
> On Fri 02-10-26 17:15:52, Joanne Koong wrote:
> > Readahead adds the folios in its window to the page cache one at a time. Each
> > folio is charged separately. When the memcg is at its limit, each one of those
> > charges triggers reclaim. With many tasks faulting in the same cgroup, the
> > margin one reclaim pass frees gets consumed by the others, so the faulting
> > tasks keep reclaiming, all to make room for speculative folios, and in the
> > worst case, reclaim livelocks.
> >
> > On Meta's fleet, this per-folio reclaim is a large cost. From fleet-wide CPU
> > profiles at Meta:
> >
> > - Page cache insertion (filemap_add_folio()) triggers about half of all
> > memcg limit reclaim CPU, more than anonymous faults, swap-in and
> > memory.high combined.
> >
> > - 93% of that comes from mmap fault readahead: filemap_fault() ->
> > do_sync_mmap_readahead() -> page_cache_ra_unbounded() ->
> > filemap_add_folio() -> mem_cgroup_charge() -> try_charge_memcg() ->
> > try_to_free_mem_cgroup_pages().
> >
> > - In filemap_add_folio(), 83% of the CPU is memcg reclaim while the page
> > cache insertion itself is 6%.
> >
> > Most of it comes from services whose worker cgroups run at their limit while
> > many threads fault in mmapped files, mostly on btrfs, which our hosts mount
> > with compress-force=zstd:3.
> >
> > This series adds readahead folios to the page cache without direct
> > reclaim. When the memcg is at its limit, readahead reclaims for the rest of
> > its window all at once instead of once per folio, and tries a second time if
> > the window runs out of room again. Please note that this only applies to
> > speculative readahead folios. For the folio a fault or read actually needs to
> > read in, it is still charged with the mapping's normal gfp mask (with the
> > reclaim flag set), like before.
>
> I cannot really comment on how much sense this makes from memcg POV (I'll
> defer to people understanding that area better). When I'm reading your
> problem description I'm wondering about one thing: Are the folios we are
> faulting around really used in your workloads before they get reclaimed?
I ran some traces on two production hosts and am roughly seeing:
* of the readahead folios that reclaim dropped during runs, ~66-78%
had never been mapped
* running a 60 second sample: mmap read-around accounts for ~86% of
readahead on both hosts. 77% of the folios it adds are never mapped.
Around 10-15% are faulted on directly and the rest are mapped by
fault-around. mmap_miss skipped read-around for only ~2-4% of major
faults. For faults that triggered read-around, mmap_miss was below 10
for 74%-96% of them.
This looks specific to mmap-readaround though. On these same hosts,
I'm seeing much better utilization by buffered read readaheads (only
~30% never accessed).
I think a big part of what's causing this is that btrfs by default
sets ra_pages to at least 4M (in open_ctree() in fs/btrfs/disk-io.c).
In do_sync_mmap_readahead(), I see "ra->size = ra->ra_pages;". So it
looks like mmap uses this same value directly for how many pages to
read-around (with 4M ra_pages, this is 1024 pages), even if the faults
are random accesses. I think this also accounts for why mmap_miss is
so low too, since with the readaround windows that large, most faults
land on pages that are already in the page cache, and each of those
minor faults count as up to 16 hits, while each major fault only adds
one miss and happens a lot less frequently. With many processes
mapping the same file, I think it makes the problem worse (eg each
process has minor faults to map it to its page tables, but there was
only 1 major fault to load it into the cache).
(Adding the btrfs list to cc. Apologies for not doing this from the
beginning. The full thread is at
https://lore.kernel.org/linux-fsdevel/20261003001555.3498357-1-joannelkoong@gmail.com/).
Instead of having the mmap readaround size coupled 1:1 with ra_pages,
do you think it makes sense to compute it differently? Like either
having some sort of fixed cap (readaround never bigger than
VM_READAHEAD_PAGES) or dynamically adapting the window where it starts
small and grows if/when readaround folios get used and shrinking when
they don't? I don't think this would degrade sequential mmap access
(eg async readahead would ramp the window back up to ra_pages in a few
steps once the fault hits the PG_readahead marker) and I think it
could be beneficial in general too, as less pages need to get read in
speculatively that wouldn't be used, which would also lead to less
reclaim.
> Because what I'd expect to happen in cases under heavy memory pressure is
> that ra->mmap_miss would quickly grow over the limit (100) and we'd just
> stop doing mmap readahead for the files. If the read-around pages aren't
> getting used before being evicted, then the first natural thing would be to
> actually stop reading them in the first place :) The mmap_miss logic is
> actually rather dumb so I wouldn't be surprised if it doesn't work in your
> case...
Do you think it'd be better to have mmap_miss count pages read vs
pages used instead of counting misses and hits seen at fault time? I
think the problem we're running into here at Meta which mmap_miss is
not detecting is that there's lots of wasted pages that never reach a
fault. It also seems to overcount hits, eg a single fault can subtract
mmap_miss by up to 16 if there are nearby cached pages (which might
have been cached by buffered reads, not mmap), even if those cached
pages never actually get touched. And the same cached page can be
counted as a hit several times, once for each process that minor
faults it in. Or maybe there's some simple way to get the info from
reclaim instead? Where if reclaim is evicting a readahead page that
was never accessed, that's counted as a miss? imo that seems a lot
more accurate as a signal though that also seems a lot more invasive,
like needing a new page flag.
As I understand it, both these improvements (detangling mmap
readaround size from ra_pages and improving mmap_miss heuristics)
would be complementary to this series rather than a replacement for
the per-folio reclaim fix. Does this match your view?
>
> The experiments you have conducted below indicate that if you effectively
> disable readahead you get more page faults & smaller IO requests (expected)
> but for your realistic loads under memory pressure is that really an
> overall performance loss?
I haven't measured end-to-end performance for that on production loads
but my suspicion is that losing read-around at the memcg limit
wouldn't affect random-access workloads like these much, while it
would for sequential readers.
Thanks,
Joanne
>
> Honza
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH v1 0/3] mm/readahead: avoid per-folio memcg reclaim
2026-10-06 14:01 ` Joanne Koong
@ 2026-10-07 15:30 ` Jan Kara
0 siblings, 0 replies; 8+ messages in thread
From: Jan Kara @ 2026-10-07 15:30 UTC (permalink / raw)
To: Joanne Koong
Cc: Jan Kara, akpm, hannes, shakeel.butt, roman.gushchin, willy,
mhocko, muchun.song, david, ljs, vbabka, liam, rppt, surenb, riel,
linux-mm, cgroups, linux-fsdevel, open list:BTRFS FILE SYSTEM
Hi Joanne!
On Tue 06-10-26 16:01:28, Joanne Koong wrote:
> On Mon, Oct 5, 2026 at 2:27 PM Jan Kara <jack@suse.cz> wrote:
> > On Fri 02-10-26 17:15:52, Joanne Koong wrote:
> > > Readahead adds the folios in its window to the page cache one at a time. Each
> > > folio is charged separately. When the memcg is at its limit, each one of those
> > > charges triggers reclaim. With many tasks faulting in the same cgroup, the
> > > margin one reclaim pass frees gets consumed by the others, so the faulting
> > > tasks keep reclaiming, all to make room for speculative folios, and in the
> > > worst case, reclaim livelocks.
> > >
> > > On Meta's fleet, this per-folio reclaim is a large cost. From fleet-wide CPU
> > > profiles at Meta:
> > >
> > > - Page cache insertion (filemap_add_folio()) triggers about half of all
> > > memcg limit reclaim CPU, more than anonymous faults, swap-in and
> > > memory.high combined.
> > >
> > > - 93% of that comes from mmap fault readahead: filemap_fault() ->
> > > do_sync_mmap_readahead() -> page_cache_ra_unbounded() ->
> > > filemap_add_folio() -> mem_cgroup_charge() -> try_charge_memcg() ->
> > > try_to_free_mem_cgroup_pages().
> > >
> > > - In filemap_add_folio(), 83% of the CPU is memcg reclaim while the page
> > > cache insertion itself is 6%.
> > >
> > > Most of it comes from services whose worker cgroups run at their limit while
> > > many threads fault in mmapped files, mostly on btrfs, which our hosts mount
> > > with compress-force=zstd:3.
> > >
> > > This series adds readahead folios to the page cache without direct
> > > reclaim. When the memcg is at its limit, readahead reclaims for the rest of
> > > its window all at once instead of once per folio, and tries a second time if
> > > the window runs out of room again. Please note that this only applies to
> > > speculative readahead folios. For the folio a fault or read actually needs to
> > > read in, it is still charged with the mapping's normal gfp mask (with the
> > > reclaim flag set), like before.
> >
> > I cannot really comment on how much sense this makes from memcg POV (I'll
> > defer to people understanding that area better). When I'm reading your
> > problem description I'm wondering about one thing: Are the folios we are
> > faulting around really used in your workloads before they get reclaimed?
>
> I ran some traces on two production hosts and am roughly seeing:
> * of the readahead folios that reclaim dropped during runs, ~66-78%
> had never been mapped
> * running a 60 second sample: mmap read-around accounts for ~86% of
> readahead on both hosts. 77% of the folios it adds are never mapped.
> Around 10-15% are faulted on directly and the rest are mapped by
> fault-around. mmap_miss skipped read-around for only ~2-4% of major
> faults. For faults that triggered read-around, mmap_miss was below 10
> for 74%-96% of them.
>
> This looks specific to mmap-readaround though. On these same hosts,
> I'm seeing much better utilization by buffered read readaheads (only
> ~30% never accessed).
>
> I think a big part of what's causing this is that btrfs by default
> sets ra_pages to at least 4M (in open_ctree() in fs/btrfs/disk-io.c).
> In do_sync_mmap_readahead(), I see "ra->size = ra->ra_pages;". So it
> looks like mmap uses this same value directly for how many pages to
> read-around (with 4M ra_pages, this is 1024 pages), even if the faults
> are random accesses. I think this also accounts for why mmap_miss is
> so low too, since with the readaround windows that large, most faults
> land on pages that are already in the page cache, and each of those
> minor faults count as up to 16 hits, while each major fault only adds
> one miss and happens a lot less frequently. With many processes
> mapping the same file, I think it makes the problem worse (eg each
> process has minor faults to map it to its page tables, but there was
> only 1 major fault to load it into the cache).
> (Adding the btrfs list to cc. Apologies for not doing this from the
> beginning. The full thread is at
> https://lore.kernel.org/linux-fsdevel/20261003001555.3498357-1-joannelkoong@gmail.com/).
Thanks for looking into this quickly! So basically your tracing seems to
confirm my suspicion our read-around logic for mmap is quite bad.
> Instead of having the mmap readaround size coupled 1:1 with ra_pages,
> do you think it makes sense to compute it differently? Like either
> having some sort of fixed cap (readaround never bigger than
> VM_READAHEAD_PAGES) or dynamically adapting the window where it starts
> small and grows if/when readaround folios get used and shrinking when
> they don't? I don't think this would degrade sequential mmap access
> (eg async readahead would ramp the window back up to ra_pages in a few
> steps once the fault hits the PG_readahead marker) and I think it
> could be beneficial in general too, as less pages need to get read in
> speculatively that wouldn't be used, which would also lead to less
> reclaim.
So I think we definitely need to fix the detection whether the fault-around
pages are getting used or not because since we've added filemap_map_pages()
the mmap_miss accounting was broken. And once that signal is fixed, we can
come up with some logic to size the readaround sensibly.
> > Because what I'd expect to happen in cases under heavy memory pressure is
> > that ra->mmap_miss would quickly grow over the limit (100) and we'd just
> > stop doing mmap readahead for the files. If the read-around pages aren't
> > getting used before being evicted, then the first natural thing would be to
> > actually stop reading them in the first place :) The mmap_miss logic is
> > actually rather dumb so I wouldn't be surprised if it doesn't work in your
> > case...
>
> Do you think it'd be better to have mmap_miss count pages read vs
> pages used instead of counting misses and hits seen at fault time? I
> think the problem we're running into here at Meta which mmap_miss is
> not detecting is that there's lots of wasted pages that never reach a
> fault. It also seems to overcount hits, eg a single fault can subtract
> mmap_miss by up to 16 if there are nearby cached pages (which might
> have been cached by buffered reads, not mmap), even if those cached
> pages never actually get touched. And the same cached page can be
> counted as a hit several times, once for each process that minor
> faults it in.
Yes, these are mostly the problems I was hinting at above :).
> Or maybe there's some simple way to get the info from
> reclaim instead? Where if reclaim is evicting a readahead page that
> was never accessed, that's counted as a miss? imo that seems a lot
> more accurate as a signal though that also seems a lot more invasive,
> like needing a new page flag.
Yes, something like that would be a good indication our readahead isn't
working as it should. But the problem he is how to track this information
with the folio... Maybe as an xarray tag? Liam was mentioning to me a few
days ago that with maple tree we have more xarray tags available ;).
> As I understand it, both these improvements (detangling mmap
> readaround size from ra_pages and improving mmap_miss heuristics)
> would be complementary to this series rather than a replacement for
> the per-folio reclaim fix. Does this match your view?
Yes, they are related but mostly independent from the issue you're
addressing in this series.
Honza
--
Jan Kara <jack@suse.com>
SUSE Labs, CR
^ permalink raw reply [flat|nested] 8+ messages in thread
end of thread, other threads:[~2026-10-07 15:31 UTC | newest]
Thread overview: 8+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-03 0:15 [PATCH v1 0/3] mm/readahead: avoid per-folio memcg reclaim Joanne Koong
2026-10-03 0:15 ` [PATCH v1 1/3] mm: memcontrol: factor reclaim logic out of try_charge_memcg() Joanne Koong
2026-10-05 18:49 ` Rik van Riel
2026-10-03 0:15 ` [PATCH v1 2/3] mm: memcontrol: add mem_cgroup_reclaim_for_batch() Joanne Koong
2026-10-03 0:15 ` [PATCH v1 3/3] mm/readahead: avoid per-folio memcg reclaim Joanne Koong
2026-10-05 12:26 ` [PATCH v1 0/3] " Jan Kara
2026-10-06 14:01 ` Joanne Koong
2026-10-07 15:30 ` Jan Kara
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox