Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Qinyun Tan <qinyuntan@linux.alibaba.com>
To: Andrew Morton <akpm@linux-foundation.org>
Cc: Johannes Weiner <hannes@cmpxchg.org>,
	Michal Hocko <mhocko@suse.com>,
	Roman Gushchin <roman.gushchin@linux.dev>,
	Shakeel Butt <shakeel.butt@linux.dev>,
	Muchun Song <muchun.song@linux.dev>,
	David Hildenbrand <david@kernel.org>,
	Lorenzo Stoakes <ljs@kernel.org>, Zi Yan <ziy@nvidia.com>,
	Baolin Wang <baolin.wang@linux.alibaba.com>,
	Xunlei Pang <xlpang@linux.alibaba.com>,
	"Liam R . Howlett" <liam@infradead.org>,
	Nico Pache <nico.pache@linux.dev>,
	Ryan Roberts <ryan.roberts@arm.com>, Dev Jain <dev.jain@arm.com>,
	Barry Song <baohua@kernel.org>, Lance Yang <lance.yang@linux.dev>,
	Usama Arif <usama.arif@linux.dev>,
	Vlastimil Babka <vbabka@kernel.org>,
	Mike Rapoport <rppt@kernel.org>,
	Suren Baghdasaryan <surenb@google.com>,
	Chris Down <chris@chrisdown.name>,
	Chuanhua Han <hanchuanhua@oppo.com>,
	Kairui Song <kasong@tencent.com>,
	linux-mm@kvack.org, cgroups@vger.kernel.org,
	linux-kernel@vger.kernel.org,
	Qinyun Tan <qinyuntan@linux.alibaba.com>,
	stable@vger.kernel.org
Subject: [PATCH 1/2] mm: memcg: settle memory.high debt after THP faults with non-blocking gfp
Date: Fri,  4 Sep 2026 11:54:06 +0800	[thread overview]
Message-ID: <20260904035407.4098627-2-qinyuntan@linux.alibaba.com> (raw)
In-Reply-To: <20260904035407.4098627-1-qinyuntan@linux.alibaba.com>

Anonymous THP faults happening in a kernel loop that does not return
to userspace -- the populate loop of a single mlock() call, or any
GUP-driven population -- can drive a memcg's usage from memory.high
all the way up to memory.max with zero reclaim and zero penalty
sleep.

This defeats the containment memory.high is supposed to provide:
above high, the documented promise is that "the processes of the
cgroup are throttled and put under heavy reclaim pressure", and
userspace OOM handlers (oomd, Kubernetes) rely on the high..max
buffer as their reaction window.  Only after hitting memory.max does
the non-blocking charge fail, THP fall back to 4K, and
folio_prealloc()'s GFP_KERNEL charge finally restore throttling --
by which point the entire buffer has been consumed.

memory.high is enforced at two points after a charge succeeds:

  1. from resume_user_mode_work() on return to userspace, requested
     via set_notify_resume();
  2. synchronously in try_charge_memcg() for large overcharges, added
     by commit c9afe31ec443 ("memcg: synchronously enforce memory.high
     for large overcharges"), gated on gfpflags_allow_blocking().

A populate loop does not return to userspace between faults, so gate
1 never runs.  Gate 2 is defeated by the charge gfp: since
commit 3b3636924dfe ("mm, memcg: sync allocation and memcg charge
gfp flags for THP"), the THP fault path passes the allocation gfp
from vma_thp_gfp_mask() to mem_cgroup_charge().  With defrag=defer
that gfp is GFP_TRANSHUGE_LIGHT | __GFP_KSWAPD_RECLAIM; with the
default defrag=madvise and no MADV_HUGEPAGE it is plain
GFP_TRANSHUGE_LIGHT.  Neither allows blocking.  This is the right
policy for the physical allocation -- a THP is not worth direct
compaction, fall back to 4K instead -- but try_charge_memcg() also
interprets it as "this context cannot sleep" and skips the
synchronous enforcement, even
though fault context sleeps just fine (it holds the mmap or per-VMA
read lock).

Fix this in the fault paths, which know their context can sleep:
after a successful THP/mTHP charge, settle any accrued over-high
debt via mem_cgroup_handle_over_high(GFP_KERNEL).  This reuses the
existing throttling machinery (reclaim + calculate_high_delay()
penalty sleep) and is a no-op read of
current->memcg_nr_pages_over_high when there is no debt.

Deliberately not changed:

 - The charge gfp itself is kept coupled to the allocation gfp, so the
   fail-fast behaviour at memory.max (charge fails -> fall back to 4K
   instead of reclaiming or OOMing for a THP) that the coupling was
   introduced for is fully preserved.

 - try_charge_memcg() is not touched: gfpflags_allow_blocking() is the
   only signal it has, and it must stay conservative for callers that
   genuinely cannot sleep.

The pre-existing selftest test_memcg_high_sync, added alongside the
synchronous enforcement by commit 6323ec54b450 ("selftests: memcg:
test high limit for single entry allocation"), readily reproduces
this: it mlocks 200M against memory.high=30M and memory.max=140M
with swap disabled, and expects high events with no max events.  On
systems with transparent_hugepage/enabled=always it fails without
this patch -- the population bursts through to memory.max -- and
passes with it.

Fixes: c9afe31ec443 ("memcg: synchronously enforce memory.high for large overcharges")
Cc: <stable@vger.kernel.org>
Signed-off-by: Qinyun Tan <qinyuntan@linux.alibaba.com>
---

Note for stable backports: mem_cgroup_handle_over_high() only gained
its gfp_mask argument in v6.6, from commit 9ea9cb00a82b ("mm:
memcontrol: fix GFP_NOFS recursion in memory.high enforcement"); on
older kernels the call sites take no argument.

 mm/huge_memory.c | 8 ++++++++
 mm/memory.c      | 2 ++
 2 files changed, 10 insertions(+)

diff --git a/mm/huge_memory.c b/mm/huge_memory.c
index ced400f72d43..543ba4a74dc3 100644
--- a/mm/huge_memory.c
+++ b/mm/huge_memory.c
@@ -1329,6 +1329,14 @@ static struct folio *vma_alloc_anon_folio_pmd(struct vm_area_struct *vma,
 		return NULL;
 	}
 
+	/*
+	 * The charge gfp encodes THP allocation policy and may not allow
+	 * blocking, which makes try_charge skip its synchronous memory.high
+	 * throttling. Fault context can sleep, so settle any over-high debt
+	 * here instead of letting usage grow unthrottled up to memory.max.
+	 */
+	mem_cgroup_handle_over_high(GFP_KERNEL);
+
 	if (folio_memcg_alloc_deferred(folio)) {
 		folio_put(folio);
 		count_vm_event(THP_FAULT_FALLBACK);
diff --git a/mm/memory.c b/mm/memory.c
index 8b0c2c735d3d..24cbf2a26905 100644
--- a/mm/memory.c
+++ b/mm/memory.c
@@ -5362,6 +5362,8 @@ static struct folio *alloc_anon_folio(struct vm_fault *vmf)
 			folio_put(folio);
 			goto next;
 		}
+		/* Same reasoning as in vma_alloc_anon_folio_pmd(). */
+		mem_cgroup_handle_over_high(GFP_KERNEL);
 		if (order > 1 && folio_memcg_alloc_deferred(folio)) {
 			folio_put(folio);
 			goto fallback;
-- 
2.55.0



  reply	other threads:[~2026-09-04  3:54 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-04  3:54 [PATCH 0/2] mm: memcg: settle memory.high debt for non-blocking THP charges Qinyun Tan
2026-09-04  3:54 ` Qinyun Tan [this message]
2026-09-04  6:50   ` [PATCH 1/2] mm: memcg: settle memory.high debt after THP faults with non-blocking gfp Baolin Wang
2026-09-04  9:05     ` Qinyun Tan
2026-09-04 15:10   ` Zi Yan
2026-09-04  3:54 ` [PATCH 2/2] mm: memcg: settle memory.high debt after large folio swapin Qinyun Tan

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260904035407.4098627-2-qinyuntan@linux.alibaba.com \
    --to=qinyuntan@linux.alibaba.com \
    --cc=akpm@linux-foundation.org \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=cgroups@vger.kernel.org \
    --cc=chris@chrisdown.name \
    --cc=david@kernel.org \
    --cc=dev.jain@arm.com \
    --cc=hanchuanhua@oppo.com \
    --cc=hannes@cmpxchg.org \
    --cc=kasong@tencent.com \
    --cc=lance.yang@linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=mhocko@suse.com \
    --cc=muchun.song@linux.dev \
    --cc=nico.pache@linux.dev \
    --cc=roman.gushchin@linux.dev \
    --cc=rppt@kernel.org \
    --cc=ryan.roberts@arm.com \
    --cc=shakeel.butt@linux.dev \
    --cc=stable@vger.kernel.org \
    --cc=surenb@google.com \
    --cc=usama.arif@linux.dev \
    --cc=vbabka@kernel.org \
    --cc=xlpang@linux.alibaba.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox