From: Joanne Koong <joannelkoong@gmail.com>
To: akpm@linux-foundation.org, hannes@cmpxchg.org,
shakeel.butt@linux.dev, roman.gushchin@linux.dev,
willy@infradead.org, jack@suse.cz
Cc: mhocko@suse.com, muchun.song@linux.dev, david@kernel.org,
ljs@kernel.org, vbabka@kernel.org, liam@infradead.org,
rppt@kernel.org, surenb@google.com, riel@surriel.com,
linux-mm@kvack.org, cgroups@vger.kernel.org,
linux-fsdevel@vger.kernel.org
Subject: [PATCH v1 2/3] mm: memcontrol: add mem_cgroup_reclaim_for_batch()
Date: Fri, 2 Oct 2026 17:15:54 -0700 [thread overview]
Message-ID: <20261003001555.3498357-3-joannelkoong@gmail.com> (raw)
In-Reply-To: <20261003001555.3498357-1-joannelkoong@gmail.com>
Some callers charge a batch of folios one at a time, where each charge
may reclaim. At the memcg limit, each one of those charges then enters
reclaim separately. Readahead is the main example. This was observed on
Meta's fleet, where a single fault's readahead window can trigger more
than a dozen reclaim passes.
Add mem_cgroup_reclaim_for_batch() so that such callers can reclaim for
a whole batch at once instead of once per folio. Callers can then charge
without reclaim and if a charge fails, call
mem_cgroup_reclaim_for_batch() to make room for the rest of the batch.
Unlike try_charge_memcg(), it retries reclaim only once rather than up
to MAX_RECLAIM_RETRIES since failing is acceptable for the callers'
charges.
The next patch uses it for readahead.
Signed-off-by: Joanne Koong <joannelkoong@gmail.com>
---
mm/internal.h | 10 ++++++++
mm/memcontrol.c | 64 +++++++++++++++++++++++++++++++++++++++++++++++++
2 files changed, 74 insertions(+)
diff --git a/mm/internal.h b/mm/internal.h
index ff28b940e0cd..6d914f6a6138 100644
--- a/mm/internal.h
+++ b/mm/internal.h
@@ -1652,4 +1652,14 @@ static inline bool can_spin_trylock(void)
/* char-mem.c */
bool file_is_dev_zero(const struct file *file);
+#ifdef CONFIG_MEMCG
+bool mem_cgroup_reclaim_for_batch(unsigned long nr_pages, gfp_t gfp);
+#else
+static inline bool mem_cgroup_reclaim_for_batch(unsigned long nr_pages,
+ gfp_t gfp)
+{
+ return false;
+}
+#endif
+
#endif /* __MM_INTERNAL_H */
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index 1d4b603085a5..9aea684ea755 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -2736,6 +2736,70 @@ static unsigned long memcg_charge_reclaim(struct mem_cgroup *memcg,
return nr_reclaimed;
}
+/**
+ * mem_cgroup_reclaim_for_batch - make room for a batch of upcoming charges
+ * @nr_pages: number of pages about to be charged
+ * @gfp: reclaim context
+ *
+ * Reclaim is needed when a memcg's margin (limit minus usage) is less than
+ * @nr_pages. Starting from the memcg that the current task's charges go to,
+ * reclaim @nr_pages from the closest ancestor whose margin is too small, with
+ * at most one retry. This lets callers that charge a batch of folios one at a
+ * time (like readahead) reclaim @nr_pages at once instead of once per folio.
+ * Callers must be ok with the charges failing.
+ *
+ * Returns true if reclaim was needed and done or false if no reclaim was
+ * needed or if the current task may not reclaim.
+ */
+bool mem_cgroup_reclaim_for_batch(unsigned long nr_pages, gfp_t gfp)
+{
+ unsigned int reclaim_options = MEMCG_RECLAIM_MAY_SWAP;
+ struct mem_cgroup *orig, *memcg;
+
+ if (mem_cgroup_disabled() || !memcg_charge_may_reclaim(gfp))
+ return false;
+
+ orig = get_mem_cgroup_from_mm(NULL);
+ for (memcg = orig; !mem_cgroup_is_root(memcg);
+ memcg = parent_mem_cgroup(memcg))
+ if (mem_cgroup_margin(memcg) < nr_pages)
+ break;
+
+ /* No reclaim needed. No memcg up to the root lacks the margin */
+ if (mem_cgroup_is_root(memcg)) {
+ css_put(&orig->css);
+ return false;
+ }
+
+ /*
+ * cgroup v1 can limit memory+swap together (memsw). If memsw is the
+ * limit being hit, swapping a page out doesn't help, so reclaim without
+ * swapping.
+ */
+ if (do_memsw_account() &&
+ page_counter_read(&memcg->memsw) + nr_pages >
+ READ_ONCE(memcg->memsw.max))
+ reclaim_options &= ~MEMCG_RECLAIM_MAY_SWAP;
+
+ memcg_charge_reclaim(memcg, nr_pages, gfp, reclaim_options);
+ /*
+ * Like try_charge_memcg(), if reclaiming didn't make enough room, drain
+ * the per-cpu stocks and if that doesn't help enough either, try
+ * reclaiming again. Other tasks charging the same memcg may have used
+ * up what was just reclaimed. Unlike try_charge_memcg(), only retry
+ * once rather than up to MAX_RECLAIM_RETRIES times, since failing is
+ * acceptable for the caller's charges.
+ */
+ if (mem_cgroup_margin(memcg) < nr_pages) {
+ drain_all_stock(memcg);
+ if (mem_cgroup_margin(memcg) < nr_pages)
+ memcg_charge_reclaim(memcg, nr_pages, gfp,
+ reclaim_options);
+ }
+ css_put(&orig->css);
+ return true;
+}
+
static int try_charge_memcg(struct mem_cgroup *memcg, gfp_t gfp_mask,
unsigned int nr_pages)
{
--
2.52.0
next prev parent reply other threads:[~2026-10-03 0:19 UTC|newest]
Thread overview: 8+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-03 0:15 [PATCH v1 0/3] mm/readahead: avoid per-folio memcg reclaim Joanne Koong
2026-10-03 0:15 ` [PATCH v1 1/3] mm: memcontrol: factor reclaim logic out of try_charge_memcg() Joanne Koong
2026-10-05 18:49 ` Rik van Riel
2026-10-03 0:15 ` Joanne Koong [this message]
2026-10-03 0:15 ` [PATCH v1 3/3] mm/readahead: avoid per-folio memcg reclaim Joanne Koong
2026-10-05 12:26 ` [PATCH v1 0/3] " Jan Kara
2026-10-06 14:01 ` Joanne Koong
2026-10-07 15:30 ` Jan Kara
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261003001555.3498357-3-joannelkoong@gmail.com \
--to=joannelkoong@gmail.com \
--cc=akpm@linux-foundation.org \
--cc=cgroups@vger.kernel.org \
--cc=david@kernel.org \
--cc=hannes@cmpxchg.org \
--cc=jack@suse.cz \
--cc=liam@infradead.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@suse.com \
--cc=muchun.song@linux.dev \
--cc=riel@surriel.com \
--cc=roman.gushchin@linux.dev \
--cc=rppt@kernel.org \
--cc=shakeel.butt@linux.dev \
--cc=surenb@google.com \
--cc=vbabka@kernel.org \
--cc=willy@infradead.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox