From: Michal Hocko <mhocko@suse.com>
To: Shakeel Butt <shakeel.butt@linux.dev>
Cc: Andrew Morton <akpm@linux-foundation.org>,
David Rientjes <rientjes@google.com>,
Johannes Weiner <hannes@cmpxchg.org>,
Roman Gushchin <roman.gushchin@linux.dev>,
Muchun Song <muchun.song@linux.dev>,
Suren Baghdasaryan <surenb@google.com>,
Usama Arif <usama.arif@linux.dev>,
Rik van Riel <riel@surriel.com>, Nhat Pham <nphamcs@gmail.com>,
Meta kernel team <kernel-team@meta.com>,
linux-mm@kvack.org, cgroups@vger.kernel.org,
linux-kernel@vger.kernel.org
Subject: Re: [PATCH] memcg: bypass the reclaim and oom killer for dying tasks once oom_reaper is done
Date: Thu, 30 Jul 2026 08:57:44 +0200 [thread overview]
Message-ID: <amr16BmkkidgRolL@tiehlicka> (raw)
In-Reply-To: <20260729024612.3369005-1-shakeel.butt@linux.dev>
On Tue 28-07-26 19:46:12, Shakeel Butt wrote:
> At Meta, we are seeing instances where an OOM killed job is stuck in the
> exit path for several hours. In one particular case, the job was stuck
> for more than 8 hours and I had to manually remove the memory.max limits
> to allow the process to exit.
>
> The job was a single process job and had ~55 GiB memory.max and zswap
> enabled. It had almost 0 anon in memory and ~111 GiB in zswap compressed
> to ~51 GiB zswap pool (i.e. almost all of memory.current was zswap).
> Nothing was left on the LRUs to reclaim.
>
> On further inspection, I observed ~20k threads of that process stuck
> with the following stack:
>
> [<0>] mem_cgroup_out_of_memory+0x4e/0xa0
> [<0>] charge_memcg+0x8bf/0x990
> [<0>] mem_cgroup_swapin_charge_folio+0x4e/0x80
> [<0>] __read_swap_cache_async+0x10c/0x260
> [<0>] swapin_readahead+0x116/0x3f0
> [<0>] do_swap_page+0x13c/0x1ce0
> [<0>] handle_mm_fault+0x61d/0x11f0
> [<0>] do_user_addr_fault+0x3e7/0x6d0
> [<0>] exc_page_fault+0x8f/0x110
> [<0>] asm_exc_page_fault+0x22/0x30
> [<0>] __get_user_8+0x14/0x20
> [<0>] futex_cleanup+0x27/0x1c0
> [<0>] futex_exit_release+0x47/0x60
> [<0>] do_exit+0x107/0x940
> [<0>] do_group_exit+0x81/0xa0
> [<0>] get_signal+0x2b1/0x6e0
> [<0>] arch_do_signal_or_restart+0x1a/0x1c0
> [<0>] exit_to_user_mode_loop+0xa8/0x1c0
> [<0>] do_syscall_64+0x152/0x250
> [<0>] entry_SYSCALL_64_after_hwframe+0x4b/0x53
>
> In addition the dmesg was filled with "Out of memory and no killable
> processes..." messages.
>
> I have no idea why oom reaper was not able to reap/unmap the process. My
> guess is that since oom reaper tries to acquire mmap_lock in read mode
> limited number of times and then gives up, there might be a thread of
> that process which had mmap_lock in write mode at that time.
>
> My initial suspicion was the futex_cleanup and kernel page fault causing
> infinite fault and charge retries but that was put to rest in previous
> discussions happened on similar problem [1].
>
> My current theory is that it is just a simple slow serialization behind
> the oom_lock. Unlike page allocator, memcg charge code takes the
> oom_lock without the "try". Though memcg oom code uses
> mutex_lock_killable(), note that in the call stack get_signal() consumes
> SIGKILL (or sigdelset(SIGKILL)) before calling do_group_exit(). So this
> mutex_lock_killable() is just a mutex_lock() here. Therefore 10s of
> thousands of threads are waiting on oom_lock and one by one they get
> -EFAULT from get_user() in the futex cleanup code and bails out.
>
> Discussion from [1] lead to the commit a75ffa26122b ("memcg, oom: do not
> bypass oom killer for dying tasks") which routes dying tasks into the OOM
> path precisely so the oom_reaper can reap their mm and free the memory
> asynchronously. But the reaper is best-effort and one-shot: if it cannot
> take mmap_lock for read (e.g. a sibling thread holds it for write) it
> sets MMF_OOM_SKIP and never retries, leaving only the glacial
> oom_lock-serialized synchronous drain.
>
> Once MMF_OOM_SKIP is set there is no more asynchronous reclaim coming for
> the mm, so a dying task charging against it has nothing left to wait for:
> it frees its memory only once it finishes exiting. Running reclaim and the
> (no-victim) OOM killer for it is then pointless, and doing it for 10s of
> thousands of exiting threads is what serializes them behind oom_lock. So
> before reclaim, if current is an OOM victim whose reaper is done, fail the
> charge.
>
> Reproduced with 20k threads, each parking a robust futex head on
> its own zswapped page, OOM-group-killed while a sibling holds mmap_lock
> for write so the reaper gives up and sets MMF_OOM_SKIP. Tested on
> next-20260728 and baseline show ~90 seconds exit time while with the
> patch the exit time reduced to ~3 seconds.
>
> Link: https://lore.kernel.org/7a4e5591f45df455e6a485fc5400989569d3d22d.camel@surriel.com/ [1]
> Signed-off-by: Shakeel Butt <shakeel.butt@linux.dev>
Acked-by: Michal Hocko <mhocko@suse.com>
Thanks
> ---
> mm/memcontrol.c | 13 +++++++++++++
> 1 file changed, 13 insertions(+)
>
> diff --git a/mm/memcontrol.c b/mm/memcontrol.c
> index 8319ad8c5c23..f7a5f8a6cfee 100644
> --- a/mm/memcontrol.c
> +++ b/mm/memcontrol.c
> @@ -2653,6 +2653,19 @@ static int try_charge_memcg(struct mem_cgroup *memcg, gfp_t gfp_mask,
> if (!gfpflags_allow_blocking(gfp_mask))
> goto nomem;
>
> + /*
> + * OOM victim still needs to charge memory to exit. OOM reaper should
> + * help but it might fail on mmap_lock contention. If the victim is a
> + * large thread group then all exiting threads might compete on oom_lock
> + * just to learn that there is nothing really killable anymore. Bail
> + * out early and fail the charge to expedite their exit. They are
> + * considered fully reclaimed by the oom reaper and they shouldn't
> + * contribute further charges.
> + */
> + if (tsk_is_oom_victim(current) &&
> + mm_flags_test(MMF_OOM_SKIP, current->signal->oom_mm))
> + goto nomem;
> +
> __memcg_memory_event(mem_over_limit, MEMCG_MAX, allow_spinning);
> raised_max_event = true;
>
> --
> 2.53.0-Meta
>
--
Michal Hocko
SUSE Labs
prev parent reply other threads:[~2026-07-30 6:57 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-29 2:46 [PATCH] memcg: bypass the reclaim and oom killer for dying tasks once oom_reaper is done Shakeel Butt
2026-07-29 18:51 ` Johannes Weiner
2026-07-30 1:13 ` Andrew Morton
2026-07-30 4:37 ` Shakeel Butt
2026-07-30 6:57 ` Michal Hocko [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=amr16BmkkidgRolL@tiehlicka \
--to=mhocko@suse.com \
--cc=akpm@linux-foundation.org \
--cc=cgroups@vger.kernel.org \
--cc=hannes@cmpxchg.org \
--cc=kernel-team@meta.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=muchun.song@linux.dev \
--cc=nphamcs@gmail.com \
--cc=riel@surriel.com \
--cc=rientjes@google.com \
--cc=roman.gushchin@linux.dev \
--cc=shakeel.butt@linux.dev \
--cc=surenb@google.com \
--cc=usama.arif@linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.