From: Baolin Wang <baolin.wang@linux.alibaba.com>
To: Barry Song <baohua@kernel.org>
Cc: Hui Zhu <hui.zhu@linux.dev>,
Andrew Morton <akpm@linux-foundation.org>,
Johannes Weiner <hannes@cmpxchg.org>,
David Hildenbrand <david@kernel.org>,
Michal Hocko <mhocko@kernel.org>, Qi Zheng <qi.zheng@linux.dev>,
Shakeel Butt <shakeel.butt@linux.dev>,
Lorenzo Stoakes <ljs@kernel.org>,
Kairui Song <kasong@tencent.com>,
Axel Rasmussen <axelrasmussen@google.com>,
Yuanchu Xie <yuanchu@google.com>, Wei Xu <weixugc@google.com>,
linux-mm@kvack.org, linux-kernel@vger.kernel.org,
Hui Zhu <zhuhui@kylinos.cn>
Subject: Re: [PATCH] mm/mglru: Fix young counter undercount for large folios
Date: Thu, 13 Aug 2026 09:31:53 +0800 [thread overview]
Message-ID: <d3471389-98b9-485a-9426-1796327b69c8@linux.alibaba.com> (raw)
In-Reply-To: <CAGsJ_4z-nfbeL7VxZyiWgJ1Ut3D5P+AfMXfqqmndwqMxmmsRJg@mail.gmail.com>
On 8/13/26 9:20 AM, Barry Song wrote:
> On Thu, Aug 13, 2026 at 9:09 AM Baolin Wang
> <baolin.wang@linux.alibaba.com> wrote:
>>
>>
>>
>> On 8/13/26 8:53 AM, Barry Song wrote:
>>> On Wed, Aug 12, 2026 at 6:17 PM Baolin Wang
>>> <baolin.wang@linux.alibaba.com> wrote:
>>>>
>>>>
>>>>
>>>> On 8/12/26 2:59 PM, Hui Zhu wrote:
>>>>> From: Hui Zhu <zhuhui@kylinos.cn>
>>>>>
>>>>> In lru_gen_look_around(), the young counter tracks the number of young
>>>>> PTEs. The original folio's contribution is represented by the initial
>>>>> value of young: test_and_clear_young_ptes_notify() is called on it at
>>>>> function entry, and the function returns early if it is not young. In
>>>>> the subsequent loop, the original folio is skipped (its accessed bits
>>>>> were already cleared), so it is not double-counted.
>>>>>
>>>>> However, young is initialized to 1 regardless of the folio size. When
>>>>> the original folio is a large folio with nr PTEs, its young count is
>>>>> underestimated by nr - 1. This inconsistency can cause
>>>>> suitable_to_scan() to return false, preventing the PMD from being added
>>>>> to the bloom filter and reducing aging accuracy for mTHP workloads.
>>>>>
>>>>> Initialize young to nr so the original folio is accounted the same way
>>>>> as other young folios in the loop (young += nr).
>>>>>
>>>>> Signed-off-by: Hui Zhu <zhuhui@kylinos.cn>
>>>>> ---
>>>>
>>>> Good catch. Please also add the Fixes tag:
>>>>
>>>> Fixes: 56e5b60b2114 ("mm: support batched checking of the young flag for
>>>> MGLRU")
>>>>
>>>> With that,
>>>> Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
>>>
>>> Hi Baolin, Hui,
>>>
>>> I am not convinced this is the correct patch. test_and_clear_young_ptes_notify()
>>> only indicates that there is at least one young PTE among the nr PTEs;
>>> it does not mean that all of the PTEs are young.
>>>
>>> Am I missing something?
>>
>> You are right. But I explained why this is done in my original commit
>> 56e5b60b2114:
>>
>> "
>> Note that we also update the 'young' counter and
>> 'mm_stats[MM_LEAF_YOUNG]' counter with the batched count in the
>> lru_gen_look_around() and walk_pte_range(). However, the batched
>> operations may inflate these two counters, because in a large folio not
>> all PTEs may have been accessed. (Additionally, tracking how many PTEs
>> have been accessed within a large folio is not very meaningful, since
>> the mm core actually tracks access/dirty on a per-folio basis, not per
>> page). The impact analysis is as follows:
>>
>> 1. The 'mm_stats[MM_LEAF_YOUNG]' counter has no functional impact and is
>> mainly for debugging.
>>
>> 2. The 'young' counter is used to decide whether to place the current
>> PMD entry into the bloom filters by suitable_to_scan() (so that next
>> time we can check whether it has been accessed again), which may set the
>> hash bit in the bloom filters for a PMD entry that hasn't seen much
>> access. However, bloom filters inherently allow some error, so this
>> effect appears negligible.
>> "
>>
>> Based on this, I think changing it to 'nr' is reasonable. For an
>> accessed large folio, it's better to have the bloom filter rescan the
>> PMD and keep it in memory instead of reclaiming it incorrectly.
>>
>
> I am not sure if this is the best policy, but we don't seem to have
> a practical way to get the exact number of accessed PTEs, so this may
> be acceptable.
As I mentioned earlier, it seems unnecessary to implement this, since
core-mm tracks access flag at per-folio granularity. Moreover, bloom
filter itself allows for some error.
However, could we at least update the changelog to
> clarify that this is intentional?
>
> " However, young is initialized to 1 regardless of the folio size. When
> the original folio is a large folio with nr PTEs, its young count is
> underestimated by nr - 1. This inconsistency can cause
> suitable_to_scan() to return false, preventing the PMD from being added"
Agree. Looks better.
> Its young count is not underestimated; we are intentionally
> overestimating it. Also, nr does not necessarily equal
> folio_nr_pages(), does it?
Right.
prev parent reply other threads:[~2026-08-13 1:32 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-12 6:59 [PATCH] mm/mglru: Fix young counter undercount for large folios Hui Zhu
2026-08-12 10:17 ` Baolin Wang
2026-08-13 0:53 ` Barry Song
2026-08-13 1:09 ` Baolin Wang
2026-08-13 1:20 ` Barry Song
2026-08-13 1:31 ` Baolin Wang [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=d3471389-98b9-485a-9426-1796327b69c8@linux.alibaba.com \
--to=baolin.wang@linux.alibaba.com \
--cc=akpm@linux-foundation.org \
--cc=axelrasmussen@google.com \
--cc=baohua@kernel.org \
--cc=david@kernel.org \
--cc=hannes@cmpxchg.org \
--cc=hui.zhu@linux.dev \
--cc=kasong@tencent.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@kernel.org \
--cc=qi.zheng@linux.dev \
--cc=shakeel.butt@linux.dev \
--cc=weixugc@google.com \
--cc=yuanchu@google.com \
--cc=zhuhui@kylinos.cn \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox