From: Shakeel Butt <shakeel.butt@linux.dev>
To: Kairui Song <ryncsn@gmail.com>
Cc: linux-mm@kvack.org, Andrew Morton <akpm@linux-foundation.org>,
Barry Song <baohua@kernel.org>,
Axel Rasmussen <axelrasmussen@google.com>,
Yuanchu Xie <yuanchu@google.com>, Wei Xu <weixugc@google.com>,
Baoquan He <baoquan.he@linux.dev>,
Johannes Weiner <hannes@cmpxchg.org>,
Michal Hocko <mhocko@kernel.org>,
Roman Gushchin <roman.gushchin@linux.dev>,
Muchun Song <muchun.song@linux.dev>,
Chris Li <chrisl@kernel.org>,
Baolin Wang <baolin.wang@linux.alibaba.com>,
David Hildenbrand <david@kernel.org>,
Lorenzo Stoakes <ljs@kernel.org>,
"Liam R. Howlett" <liam@infradead.org>,
Vlastimil Babka <vbabka@kernel.org>,
Ridong Chen <ridong.chen@linux.dev>,
Lian Wang <lianux.mm@gmail.com>, Yu Zhao <yuzhao@google.com>,
Zi Yan <ziy@nvidia.com>, Qi Zheng <qi.zheng@linux.dev>,
cgroups@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH v4 1/6] mm/memcontrol: make lru_zone_size atomic and simplify sanity check
Date: Tue, 1 Sep 2026 10:37:14 -0700 [thread overview]
Message-ID: <apcJgcFLHKZfI40m@linux.dev> (raw)
In-Reply-To: <CAMgjq7D6V4eTZJ1Q1Xf2J6yTCq=JjvzT-W-HHX6EnMfzm3BK1g@mail.gmail.com>
On Tue, Sep 01, 2026 at 01:20:33PM +0800, Kairui Song wrote:
> On Tue, Sep 1, 2026 at 12:38 PM Shakeel Butt <shakeel.butt@linux.dev> wrote:
> >
> > On Mon, Aug 31, 2026 at 02:43:31AM +0800, Kairui Song via B4 Relay wrote:
> > > From: Kairui Song <kasong@tencent.com>
> > >
> > > commit ca707239e8a7 ("mm: update_lru_size warn and reset bad lru_size")
> > > introduced a sanity check to catch memcg counter underflow, which was
> > > more of a workaround for another bug: lru_zone_size is unsigned, so
> > > underflow wraps it around and returns an enormously large number, then
> > > the memcg shrinker loops almost forever as the calculated number of
> > > folios to shrink is huge. That commit also checked if a zero value
> > > matches the empty LRU list, so we have to hold the LRU lock, and
> > > handle the positive and negative deltas separately.
> > >
> > > But later commit b4536f0c829c ("mm, memcg: fix the active list aging
> > > for lowmem requests when memcg is enabled") already removed the LRU
> > > emptiness check, so handling the deltas separately is no longer
> > > needed. And if we just turn it into an atomic long, underflow isn't a
> > > big issue either,
> >
> > Why atomic long and not just long?
>
> Long is fine too; I was just trying to be defensive. The original
> check was a defensive design for leaks or races, and atomic long is
> more immune to race conditions. See below for the other conditions.
Unless we really need it, let's stay with "long" for now. As you mentioned below
this is busy/hot path, no need to add overhead of atomic here.
>
> > > So let's turn the counter into an atomic long and check at the reader
> > > side instead, which has a smaller overhead. The underflow correction
> > > is removed: a massive leak of the LRU size counter would indicate
> > > that something else has gone very wrong, and one should fix that
> > > leaking site instead. Besides, the updater-side sanity check is
> > > unlikely to catch the leaking site anyway: if a folio was removed
> > > without updating the counter while other folios remain on the LRU,
> > > the WARN only triggers much later, from a likely innocent callsite.
> >
> > Do you have any data to support your claim that updater-side sanity check is not
> > that useful? Also can you explain the motivation to move the check from the
> > update side to reader side?
>
> I think I described that in the commit message. Commit ca707239e8a7
> introduced the original emptiness check. The goal was to fix an
> over-busy reclaim scenario: upon underflow, an unsigned long returns
> an extremely large value. The reclaimer calculates the page number to
> scan based on that value, and thus enters a busy loop trying to
> reclaim much more than the actual usage. It's really a defensive
> design or workaround.
>
> And later commit b4536f0c829c just removed that emptiness check,
> leaving only the underflow check. So at this point we are just
> defending against underflow. Currently, I don't believe there are any
> known leaks or underflow issues in mainline, still purely a defensive
> design.
>
> Reading these counters (mem_cgroup_get_zone_lru_size()) is much rare
> than updating them: the only readers are reclaimer - per loop, each
> loop will scan a large chunk of folios. Or during reparenting - per
> child.
>
> Writers, however, operate on every LRU folio.
>
> So checking it in the reader side is much lighter and cleaner code-wise.
So, one motivation is cleaner code?
>
> Checking in the reader side also tolerates deferred accounting false
> positives: if a folio is added to / removed from LRU but have the
> accounting happens later, that's fine, but might cause a temporary
> underflow (there shouldn't be such usage right now, but cleanly the
> oldest commit I mentioned is trying to catch something). It should
> just proceed without performing extra operations on the counter or
> raising a warning on the updater side.
>
> The underflow won't affect anything besides the two readers mentioned
> above (reclaimer and reparenting).
>
> Besides, I don't think warning on the writer side will help catch the
> real leak either, it just adds overhead to a busy path. The folio
> triggering the underflow is very unlikely to be the one leaking the
> count. And if that leak is actually caused by a race, the atomic long
> change should cure it.
I think you are going and looking too much into the history of this warning. I
think you just need to explain that the current warning in the update side
already can't always catch the actual leak and is costly. Moving to read side
make the code cleaner and warnings will be at the location where the consumers
we care about are going to consume it. Also add couple of sentences on behavior
difference it will have i.e. warning on flipping into negative vs we will
continue to operate even in negative but only warn on negative consumption.
Another question: why is this patch part of this series? Just some orthogonal
cleanups or are they related?
next prev parent reply other threads:[~2026-09-01 17:37 UTC|newest]
Thread overview: 17+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-30 18:43 [PATCH v4 0/6] mm/mglru: clean up folio counters and flag usage Kairui Song via B4 Relay
2026-08-30 18:43 ` [PATCH v4 1/6] mm/memcontrol: make lru_zone_size atomic and simplify sanity check Kairui Song via B4 Relay
2026-09-01 4:38 ` Shakeel Butt
2026-09-01 5:20 ` Kairui Song
2026-09-01 17:37 ` Shakeel Butt [this message]
2026-09-01 18:11 ` Kairui Song
2026-09-01 18:34 ` Shakeel Butt
2026-08-30 18:43 ` [PATCH v4 2/6] mm/mglru: introduce helpers for manipulating gen and refs flags Kairui Song via B4 Relay
2026-09-01 2:05 ` Baolin Wang
2026-09-01 10:32 ` Barry Song
2026-08-30 18:43 ` [PATCH v4 3/6] mm/migrate: copy all referenced state via folio_migrate_lru_refs Kairui Song via B4 Relay
2026-08-30 18:43 ` [PATCH v4 4/6] mm/mglru: move max_seq read into walk_update_folio Kairui Song via B4 Relay
2026-08-30 18:43 ` [PATCH v4 5/6] mm/mglru: use explicit tier range in read_ctrl_pos() Kairui Song via B4 Relay
2026-08-30 18:43 ` [PATCH v4 6/6] mm/mglru: fix potential generation folio number leak Kairui Song via B4 Relay
2026-09-01 10:56 ` Barry Song
2026-09-01 11:13 ` Kairui Song
2026-09-01 11:33 ` Barry Song
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=apcJgcFLHKZfI40m@linux.dev \
--to=shakeel.butt@linux.dev \
--cc=akpm@linux-foundation.org \
--cc=axelrasmussen@google.com \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=baoquan.he@linux.dev \
--cc=cgroups@vger.kernel.org \
--cc=chrisl@kernel.org \
--cc=david@kernel.org \
--cc=hannes@cmpxchg.org \
--cc=liam@infradead.org \
--cc=lianux.mm@gmail.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@kernel.org \
--cc=muchun.song@linux.dev \
--cc=qi.zheng@linux.dev \
--cc=ridong.chen@linux.dev \
--cc=roman.gushchin@linux.dev \
--cc=ryncsn@gmail.com \
--cc=vbabka@kernel.org \
--cc=weixugc@google.com \
--cc=yuanchu@google.com \
--cc=yuzhao@google.com \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox