Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Joanne Koong <joannelkoong@gmail.com>
To: akpm@linux-foundation.org, hannes@cmpxchg.org,
	shakeel.butt@linux.dev, roman.gushchin@linux.dev,
	willy@infradead.org, jack@suse.cz
Cc: mhocko@suse.com, muchun.song@linux.dev, david@kernel.org,
	ljs@kernel.org, vbabka@kernel.org, liam@infradead.org,
	rppt@kernel.org, surenb@google.com, riel@surriel.com,
	linux-mm@kvack.org, cgroups@vger.kernel.org,
	linux-fsdevel@vger.kernel.org
Subject: [PATCH v1 2/3] mm: memcontrol: add mem_cgroup_reclaim_for_batch()
Date: Fri,  2 Oct 2026 17:15:54 -0700	[thread overview]
Message-ID: <20261003001555.3498357-3-joannelkoong@gmail.com> (raw)
In-Reply-To: <20261003001555.3498357-1-joannelkoong@gmail.com>

Some callers charge a batch of folios one at a time, where each charge
may reclaim. At the memcg limit, each one of those charges then enters
reclaim separately. Readahead is the main example. This was observed on
Meta's fleet, where a single fault's readahead window can trigger more
than a dozen reclaim passes.

Add mem_cgroup_reclaim_for_batch() so that such callers can reclaim for
a whole batch at once instead of once per folio. Callers can then charge
without reclaim and if a charge fails, call
mem_cgroup_reclaim_for_batch() to make room for the rest of the batch.
Unlike try_charge_memcg(), it retries reclaim only once rather than up
to MAX_RECLAIM_RETRIES since failing is acceptable for the callers'
charges.

The next patch uses it for readahead.

Signed-off-by: Joanne Koong <joannelkoong@gmail.com>
---
 mm/internal.h   | 10 ++++++++
 mm/memcontrol.c | 64 +++++++++++++++++++++++++++++++++++++++++++++++++
 2 files changed, 74 insertions(+)

diff --git a/mm/internal.h b/mm/internal.h
index ff28b940e0cd..6d914f6a6138 100644
--- a/mm/internal.h
+++ b/mm/internal.h
@@ -1652,4 +1652,14 @@ static inline bool can_spin_trylock(void)
 /* char-mem.c */
 bool file_is_dev_zero(const struct file *file);
 
+#ifdef CONFIG_MEMCG
+bool mem_cgroup_reclaim_for_batch(unsigned long nr_pages, gfp_t gfp);
+#else
+static inline bool mem_cgroup_reclaim_for_batch(unsigned long nr_pages,
+						gfp_t gfp)
+{
+	return false;
+}
+#endif
+
 #endif	/* __MM_INTERNAL_H */
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index 1d4b603085a5..9aea684ea755 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -2736,6 +2736,70 @@ static unsigned long memcg_charge_reclaim(struct mem_cgroup *memcg,
 	return nr_reclaimed;
 }
 
+/**
+ * mem_cgroup_reclaim_for_batch - make room for a batch of upcoming charges
+ * @nr_pages: number of pages about to be charged
+ * @gfp: reclaim context
+ *
+ * Reclaim is needed when a memcg's margin (limit minus usage) is less than
+ * @nr_pages. Starting from the memcg that the current task's charges go to,
+ * reclaim @nr_pages from the closest ancestor whose margin is too small, with
+ * at most one retry. This lets callers that charge a batch of folios one at a
+ * time (like readahead) reclaim @nr_pages at once instead of once per folio.
+ * Callers must be ok with the charges failing.
+ *
+ * Returns true if reclaim was needed and done or false if no reclaim was
+ * needed or if the current task may not reclaim.
+ */
+bool mem_cgroup_reclaim_for_batch(unsigned long nr_pages, gfp_t gfp)
+{
+	unsigned int reclaim_options = MEMCG_RECLAIM_MAY_SWAP;
+	struct mem_cgroup *orig, *memcg;
+
+	if (mem_cgroup_disabled() || !memcg_charge_may_reclaim(gfp))
+		return false;
+
+	orig = get_mem_cgroup_from_mm(NULL);
+	for (memcg = orig; !mem_cgroup_is_root(memcg);
+	     memcg = parent_mem_cgroup(memcg))
+		if (mem_cgroup_margin(memcg) < nr_pages)
+			break;
+
+	/* No reclaim needed. No memcg up to the root lacks the margin */
+	if (mem_cgroup_is_root(memcg)) {
+		css_put(&orig->css);
+		return false;
+	}
+
+	/*
+	 * cgroup v1 can limit memory+swap together (memsw). If memsw is the
+	 * limit being hit, swapping a page out doesn't help, so reclaim without
+	 * swapping.
+	 */
+	if (do_memsw_account() &&
+	    page_counter_read(&memcg->memsw) + nr_pages >
+	    READ_ONCE(memcg->memsw.max))
+		reclaim_options &= ~MEMCG_RECLAIM_MAY_SWAP;
+
+	memcg_charge_reclaim(memcg, nr_pages, gfp, reclaim_options);
+	/*
+	 * Like try_charge_memcg(), if reclaiming didn't make enough room, drain
+	 * the per-cpu stocks and if that doesn't help enough either, try
+	 * reclaiming again. Other tasks charging the same memcg may have used
+	 * up what was just reclaimed. Unlike try_charge_memcg(), only retry
+	 * once rather than up to MAX_RECLAIM_RETRIES times, since failing is
+	 * acceptable for the caller's charges.
+	 */
+	if (mem_cgroup_margin(memcg) < nr_pages) {
+		drain_all_stock(memcg);
+		if (mem_cgroup_margin(memcg) < nr_pages)
+			memcg_charge_reclaim(memcg, nr_pages, gfp,
+					     reclaim_options);
+	}
+	css_put(&orig->css);
+	return true;
+}
+
 static int try_charge_memcg(struct mem_cgroup *memcg, gfp_t gfp_mask,
 			    unsigned int nr_pages)
 {
-- 
2.52.0



  parent reply	other threads:[~2026-10-03  0:19 UTC|newest]

Thread overview: 8+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-03  0:15 [PATCH v1 0/3] mm/readahead: avoid per-folio memcg reclaim Joanne Koong
2026-10-03  0:15 ` [PATCH v1 1/3] mm: memcontrol: factor reclaim logic out of try_charge_memcg() Joanne Koong
2026-10-05 18:49   ` Rik van Riel
2026-10-03  0:15 ` Joanne Koong [this message]
2026-10-03  0:15 ` [PATCH v1 3/3] mm/readahead: avoid per-folio memcg reclaim Joanne Koong
2026-10-05 12:26 ` [PATCH v1 0/3] " Jan Kara
2026-10-06 14:01   ` Joanne Koong
2026-10-07 15:30     ` Jan Kara

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261003001555.3498357-3-joannelkoong@gmail.com \
    --to=joannelkoong@gmail.com \
    --cc=akpm@linux-foundation.org \
    --cc=cgroups@vger.kernel.org \
    --cc=david@kernel.org \
    --cc=hannes@cmpxchg.org \
    --cc=jack@suse.cz \
    --cc=liam@infradead.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=mhocko@suse.com \
    --cc=muchun.song@linux.dev \
    --cc=riel@surriel.com \
    --cc=roman.gushchin@linux.dev \
    --cc=rppt@kernel.org \
    --cc=shakeel.butt@linux.dev \
    --cc=surenb@google.com \
    --cc=vbabka@kernel.org \
    --cc=willy@infradead.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox