From: Shakeel Butt <shakeel.butt@linux.dev>
To: Michal Hocko <mhocko@suse.com>
Cc: Audra Mitchell <audra@redhat.com>,
david@kernel.org, jocolema@redhat.com, raquini@redhat.com,
Johannes Weiner <hannes@cmpxchg.org>,
Roman Gushchin <roman.gushchin@linux.dev>,
Muchun Song <muchun.song@linux.dev>,
Andrew Morton <akpm@linux-foundation.org>,
cgroups@vger.kernel.org, linux-mm@kvack.org,
linux-kernel@vger.kernel.org
Subject: Re: [PATCH] Fix unbounded loop within try_charge_memcg
Date: Fri, 7 Aug 2026 12:10:20 -0700 [thread overview]
Message-ID: <anYr6UIi1nevGJkA@linux.dev> (raw)
In-Reply-To: <anYoLs1B0MwUNR2n@tiehlicka>
On Fri, Aug 07, 2026 at 08:47:10PM +0200, Michal Hocko wrote:
> On Fri 07-08-26 11:40:06, Audra Mitchell wrote:
> > On Fri, Aug 07, 2026 at 10:06:53AM +0200, Michal Hocko wrote:
> > > On Thu 06-08-26 11:10:03, Audra Mitchell wrote:
> > > > Originally nr_retries was actually nr_oom_retries and we used it to track (and
> > > > limit) the number of times we entered the mem_cgroup_oom path and then attempted
> > > > a retry. The purpose of nr_retries counter changed with the introduction of
> > > > 9b1306192d33 ("mm: memcontrol: retry reclaim for oom-disabled and __GFP_NOFAIL
> > > > charges") so that the oom-disabled and __GFP_NOFAIL charges would also continue
> > > > to retry within the desired nr_retries threshold. Later d977aa939fca
> > > > ("mm, memcg: unify reclaim retry limits with page allocator") changed the
> > > > nr_retries counter from 5 to 16.
> > > >
> > > > As the function has evolved we now have multiple paths that have a goto retry
> > > > path and we have lost the original purpose of the nr_retries counter, allowing
> > > > us to take a goto retry path an unbounded number of times.
> > > >
> > > > Fix the unbounded retries by nesting the code in a loop and decrementing the
> > > > nr_retries counter correctly.
> > >
> > > Are you trying to fix a theoretical problem spotted by the code review
> > > or is there any actual problem that you are trying to fix?
> >
> > We have had some customer complaints that performance has slowed to a crawl when
> > the cgroup memory limit has come close to the maximum limit. In those cases, we
> > have noticed that each process is spending a large amount of time in the direct
> > reclaim path acquiring just enough memory for their specific allocation, thus
> > by-passing the oom condition yet degrading the system's overall performance.
>
> Yes, this is entirely possible scenario.
>
> > In the global case, direct reclaim is bounded by DEF_PRIORITY,
>
> Well, both global and memcg reclaim share the reclaim logic. Both of
> them try to exercise all reclaim priorities (i.e. check whole eligible
> LRU lists) and they fall back to OOM killer only if there is no other
> option left. For the global case should_reclaim_retry is the gate keeper
> around direct reclaim retries while for the memcg we have more or less
> fixed number of retries.
>
> > however, a cgroup
> > will go through the try_charge_memcg path which will call
> > try_to_free_mem_cgroup_pages->do_try_to_free_pages each time it does a retry (16
> > times).
>
> > If we bound the loop in try_charge_memcg, the worst case is 16*12 passes
> > attempting to reclaim. This patch is meant to address the unbound case, limiting
> > the loops to 16 attempts at following the direct reclaim path. An argument could
> > be made to reduce nr_retries as well, but given that the nr_retries has been set
> > to 16 for sometime, it seemed unlikely such a change would be considered.
>
> As Shakeel said in other reply, this is a deliberate implementation
> decision. The OOM killer is the very last resort and we are giving
> chance to userspace to handle close to OOM situation much more
> gracefully and also workload aware. Keep in mind that what might be seen
> as a slow progress for one workload might be acceptable for others where
> OOM killer could mean a lot of work being lost.
>
> From what you are describing above those users might be hitting reclaim trashing.
> I.e. last small portion of a reclaimable memory is bounced back and
> forth for the workload to make tiny but steady forward progress. While
> OOM killer might help to stop the suffering and restart the workload
> sooner I would generally recommend revisiting limits set for the
> particular workload.
+1 to this. Beside memory.pressure, we also have refault and reclaim metrics
in memory.stat which can further help in debugging if the workload is thrashing
due to workingset larger than the limits.
> Especially if restarting it might lead to the same
> state sooner or later. Watching PSI metric would be a good start to see
> how the workload behaves wrt memory stalling. User space oom handlers
> might be a proper measure as well but that will always be safeguard
> rather than a solution.
> --
> Michal Hocko
> SUSE Labs
next prev parent reply other threads:[~2026-08-07 19:10 UTC|newest]
Thread overview: 8+ messages / expand[flat|nested] mbox.gz Atom feed top
[not found] <20260806151004.1825320-1-audra@redhat.com>
[not found] ` <anWSHc04JJBJ1VbS@tiehlicka>
2026-08-07 15:40 ` [PATCH] Fix unbounded loop within try_charge_memcg Audra Mitchell
2026-08-07 17:28 ` Shakeel Butt
2026-08-07 18:47 ` Michal Hocko
2026-08-07 19:10 ` Shakeel Butt [this message]
2026-08-07 19:13 ` Audra Mitchell
2026-08-07 20:07 ` Michal Hocko
2026-08-06 18:24 Audra Mitchell
2026-08-07 9:17 ` Michal Hocko
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=anYr6UIi1nevGJkA@linux.dev \
--to=shakeel.butt@linux.dev \
--cc=akpm@linux-foundation.org \
--cc=audra@redhat.com \
--cc=cgroups@vger.kernel.org \
--cc=david@kernel.org \
--cc=hannes@cmpxchg.org \
--cc=jocolema@redhat.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=mhocko@suse.com \
--cc=muchun.song@linux.dev \
--cc=raquini@redhat.com \
--cc=roman.gushchin@linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox