Linux filesystem development
 help / color / mirror / Atom feed
From: Joanne Koong <joannelkoong@gmail.com>
To: akpm@linux-foundation.org, hannes@cmpxchg.org,
	shakeel.butt@linux.dev, roman.gushchin@linux.dev,
	willy@infradead.org, jack@suse.cz
Cc: mhocko@suse.com, muchun.song@linux.dev, david@kernel.org,
	ljs@kernel.org, vbabka@kernel.org, liam@infradead.org,
	rppt@kernel.org, surenb@google.com, riel@surriel.com,
	linux-mm@kvack.org, cgroups@vger.kernel.org,
	linux-fsdevel@vger.kernel.org
Subject: [PATCH v1 3/3] mm/readahead: avoid per-folio memcg reclaim
Date: Fri,  2 Oct 2026 17:15:55 -0700	[thread overview]
Message-ID: <20261003001555.3498357-4-joannelkoong@gmail.com> (raw)
In-Reply-To: <20261003001555.3498357-1-joannelkoong@gmail.com>

Readahead adds the folios in its window to the page cache one at a time.
Each folio is charged separately. When the memcg is at its limit, each
one of those charges triggers reclaim (readahead's __GFP_NORETRY only
stops try_charge_memcg() after a reclaim pass). With many tasks faulting
in the same cgroup, the margin one reclaim pass frees gets consumed by
the others, so the faulting tasks keep reclaiming, all to make room for
speculative folios, and in the worst case, reclaim livelocks.

On Meta's fleet, this per-folio reclaim is a large cost. Page cache
insertion triggers around half of all memcg limit reclaim CPU, 93% of it
from mmap fault readahead, in services whose cgroups run at their limit
while many threads fault in mmapped files. Of the CPU time spent in
filemap_add_folio(), 83% is memcg reclaim while inserting the folio into
the page cache itself is 6%.

To avoid this, add readahead folios to the page cache without direct
reclaim, and when the memcg is at its limit, reclaim for the rest of the
window at once with mem_cgroup_reclaim_for_batch() instead of once per
folio. If the memcg hits its limit again before the whole window has
been added, try one more time, since other tasks charging the same memcg
may have used up the room the first attempt made. If there's still no
room, stop readahead early rather than reclaim harder, since readahead
folios are speculative. The count of reclaim attempts lives in
readahead_control so that it covers every path that adds folios for the
request.

This only affects speculative readahead folios. There's no change in
behavior for the folio a fault or read actually needs. If readahead
didn't bring that folio in, filemap_fault() and filemap_read() allocate
and charge it with the mapping's normal gfp mask, like before.

There are two other differences from the current behavior. The first is
that readahead charges that push a cgroup past memory.high leave the
high reclaim to the return to user space, as other non-blocking charges
do. The second is that under global memory pressure, readahead's page
cache xarray node allocations no longer enter direct reclaim (they still
wake kswapd). If one fails, readahead stops early, as it does when a
folio allocation fails.

Tested on a 26-core/52-thread machine with btrfs, running 26, 52 or 104
processes (one per core, one per thread, and 2x oversubscribed) that
each mmap their own file in one memcg with a 1G memory.max (before and
after measured in the same boot), reclaim passes per major fault drop by
94-96% for both random and sequential reads. With data compressing ~3:1
under compress-force=zstd:3, the runs that livelocked in reclaim without
the patch (4 of 18) no longer do. Throughput with incompressible data,
where the disk is the bottleneck, is within 5% of before in either
direction, and the drops are within run-to-run variation. With the
compressed data, it is 9% higher at 104 processes, where reclaim
contention is. Sequential readers get up to 18% fewer pages per major
fault, since readahead now stops early when there is still no room, but
their throughput and read bandwidth are unchanged or slightly higher.

Reported-by: Rik van Riel <riel@surriel.com>
Signed-off-by: Joanne Koong <joannelkoong@gmail.com>
---
 include/linux/pagemap.h |  1 +
 mm/readahead.c          | 51 ++++++++++++++++++++++++++++++++++++-----
 2 files changed, 46 insertions(+), 6 deletions(-)

diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h
index 73af18a37367..973a836076bd 100644
--- a/include/linux/pagemap.h
+++ b/include/linux/pagemap.h
@@ -1409,6 +1409,7 @@ struct readahead_control {
 	pgoff_t _index;
 	unsigned int _nr_pages;
 	unsigned int _batch_count;
+	unsigned int _nr_memcg_reclaims;
 	bool dropbehind;
 	bool _workingset;
 	unsigned long _pflags;
diff --git a/mm/readahead.c b/mm/readahead.c
index 6e5563290287..e06375559e9b 100644
--- a/mm/readahead.c
+++ b/mm/readahead.c
@@ -204,6 +204,41 @@ static struct folio *ractl_alloc_folio(struct readahead_control *ractl,
 	return folio;
 }
 
+/*
+ * Max number of memcg reclaim attempts per readahead request. The first
+ * makes room for the rest of the readahead window. If the window runs out
+ * of room again, try one more time since other readers in the same memcg
+ * might have used that room up.
+ */
+#define READAHEAD_MAX_MEMCG_RECLAIMS	2
+
+/*
+ * Add a readahead folio to the page cache without direct reclaim. This prevents
+ * a memcg at its limit from running reclaim for every folio in the readahead
+ * window (folios in the window are charged one at a time). If adding the folio
+ * to the page cache returns -ENOMEM, reclaim enough room for the rest of the
+ * window and retry adding it.
+ */
+static int readahead_add_folio(struct readahead_control *ractl,
+			       struct folio *folio, pgoff_t index,
+			       unsigned long nr_pages_left, gfp_t gfp)
+{
+	struct address_space *mapping = ractl->mapping;
+	int ret;
+
+	ret = filemap_add_folio(mapping, folio, index,
+				gfp & ~__GFP_DIRECT_RECLAIM);
+	if (ret == -ENOMEM &&
+	    ractl->_nr_memcg_reclaims < READAHEAD_MAX_MEMCG_RECLAIMS) {
+		ractl->_nr_memcg_reclaims++;
+		nr_pages_left = max(nr_pages_left, folio_nr_pages(folio));
+		if (mem_cgroup_reclaim_for_batch(nr_pages_left, gfp))
+			ret = filemap_add_folio(mapping, folio, index,
+						gfp & ~__GFP_DIRECT_RECLAIM);
+	}
+	return ret;
+}
+
 /**
  * page_cache_ra_unbounded - Start unchecked readahead.
  * @ractl: Readahead control.
@@ -290,7 +325,8 @@ void page_cache_ra_unbounded(struct readahead_control *ractl,
 		if (!folio)
 			break;
 
-		ret = filemap_add_folio(mapping, folio, index + i, gfp_mask);
+		ret = readahead_add_folio(ractl, folio, index + i,
+					  nr_to_read - i, gfp_mask);
 		if (ret < 0) {
 			folio_put(folio);
 			if (ret == -ENOMEM)
@@ -457,7 +493,7 @@ static unsigned long get_next_ra_size(struct file_ra_state *ra,
  */
 
 static inline int ra_alloc_folio(struct readahead_control *ractl, pgoff_t index,
-		pgoff_t mark, unsigned int order, gfp_t gfp)
+		pgoff_t mark, pgoff_t limit, unsigned int order, gfp_t gfp)
 {
 	int err;
 	struct folio *folio = ractl_alloc_folio(ractl, gfp, order);
@@ -467,7 +503,7 @@ static inline int ra_alloc_folio(struct readahead_control *ractl, pgoff_t index,
 	mark = round_down(mark, 1UL << order);
 	if (index == mark)
 		folio_set_readahead(folio);
-	err = filemap_add_folio(ractl->mapping, folio, index, gfp);
+	err = readahead_add_folio(ractl, folio, index, limit - index + 1, gfp);
 	if (err) {
 		folio_put(folio);
 		return err;
@@ -532,7 +568,7 @@ void page_cache_ra_order(struct readahead_control *ractl,
 		/* Don't allocate pages past EOF */
 		while (order > min_order && index + (1UL << order) - 1 > limit)
 			order--;
-		err = ra_alloc_folio(ractl, index, mark, order, gfp);
+		err = ra_alloc_folio(ractl, index, mark, limit, order, gfp);
 		if (err)
 			break;
 		index += 1UL << order;
@@ -813,7 +849,8 @@ void readahead_expand(struct readahead_control *ractl,
 			return;
 
 		index = mapping_align_index(mapping, index);
-		if (filemap_add_folio(mapping, folio, index, gfp_mask) < 0) {
+		if (readahead_add_folio(ractl, folio, index,
+					ractl->_index - new_index, gfp_mask) < 0) {
 			folio_put(folio);
 			return;
 		}
@@ -842,7 +879,9 @@ void readahead_expand(struct readahead_control *ractl,
 			return;
 
 		index = mapping_align_index(mapping, index);
-		if (filemap_add_folio(mapping, folio, index, gfp_mask) < 0) {
+		if (readahead_add_folio(ractl, folio, index,
+					new_nr_pages - ractl->_nr_pages,
+					gfp_mask) < 0) {
 			folio_put(folio);
 			return;
 		}
-- 
2.52.0


  parent reply	other threads:[~2026-10-03  0:19 UTC|newest]

Thread overview: 8+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-03  0:15 [PATCH v1 0/3] mm/readahead: avoid per-folio memcg reclaim Joanne Koong
2026-10-03  0:15 ` [PATCH v1 1/3] mm: memcontrol: factor reclaim logic out of try_charge_memcg() Joanne Koong
2026-10-05 18:49   ` Rik van Riel
2026-10-03  0:15 ` [PATCH v1 2/3] mm: memcontrol: add mem_cgroup_reclaim_for_batch() Joanne Koong
2026-10-03  0:15 ` Joanne Koong [this message]
2026-10-05 12:26 ` [PATCH v1 0/3] mm/readahead: avoid per-folio memcg reclaim Jan Kara
2026-10-06 14:01   ` Joanne Koong
2026-10-07 15:30     ` Jan Kara

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261003001555.3498357-4-joannelkoong@gmail.com \
    --to=joannelkoong@gmail.com \
    --cc=akpm@linux-foundation.org \
    --cc=cgroups@vger.kernel.org \
    --cc=david@kernel.org \
    --cc=hannes@cmpxchg.org \
    --cc=jack@suse.cz \
    --cc=liam@infradead.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=mhocko@suse.com \
    --cc=muchun.song@linux.dev \
    --cc=riel@surriel.com \
    --cc=roman.gushchin@linux.dev \
    --cc=rppt@kernel.org \
    --cc=shakeel.butt@linux.dev \
    --cc=surenb@google.com \
    --cc=vbabka@kernel.org \
    --cc=willy@infradead.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox