Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Baolin Wang <baolin.wang@linux.alibaba.com>
To: Barry Song <baohua@kernel.org>
Cc: Hui Zhu <hui.zhu@linux.dev>,
	Andrew Morton <akpm@linux-foundation.org>,
	Johannes Weiner <hannes@cmpxchg.org>,
	David Hildenbrand <david@kernel.org>,
	Michal Hocko <mhocko@kernel.org>, Qi Zheng <qi.zheng@linux.dev>,
	Shakeel Butt <shakeel.butt@linux.dev>,
	Lorenzo Stoakes <ljs@kernel.org>,
	Kairui Song <kasong@tencent.com>,
	Axel Rasmussen <axelrasmussen@google.com>,
	Yuanchu Xie <yuanchu@google.com>, Wei Xu <weixugc@google.com>,
	linux-mm@kvack.org, linux-kernel@vger.kernel.org,
	Hui Zhu <zhuhui@kylinos.cn>
Subject: Re: [PATCH] mm/mglru: Fix young counter undercount for large folios
Date: Thu, 13 Aug 2026 09:31:53 +0800	[thread overview]
Message-ID: <d3471389-98b9-485a-9426-1796327b69c8@linux.alibaba.com> (raw)
In-Reply-To: <CAGsJ_4z-nfbeL7VxZyiWgJ1Ut3D5P+AfMXfqqmndwqMxmmsRJg@mail.gmail.com>



On 8/13/26 9:20 AM, Barry Song wrote:
> On Thu, Aug 13, 2026 at 9:09 AM Baolin Wang
> <baolin.wang@linux.alibaba.com> wrote:
>>
>>
>>
>> On 8/13/26 8:53 AM, Barry Song wrote:
>>> On Wed, Aug 12, 2026 at 6:17 PM Baolin Wang
>>> <baolin.wang@linux.alibaba.com> wrote:
>>>>
>>>>
>>>>
>>>> On 8/12/26 2:59 PM, Hui Zhu wrote:
>>>>> From: Hui Zhu <zhuhui@kylinos.cn>
>>>>>
>>>>> In lru_gen_look_around(), the young counter tracks the number of young
>>>>> PTEs.  The original folio's contribution is represented by the initial
>>>>> value of young: test_and_clear_young_ptes_notify() is called on it at
>>>>> function entry, and the function returns early if it is not young.  In
>>>>> the subsequent loop, the original folio is skipped (its accessed bits
>>>>> were already cleared), so it is not double-counted.
>>>>>
>>>>> However, young is initialized to 1 regardless of the folio size.  When
>>>>> the original folio is a large folio with nr PTEs, its young count is
>>>>> underestimated by nr - 1.  This inconsistency can cause
>>>>> suitable_to_scan() to return false, preventing the PMD from being added
>>>>> to the bloom filter and reducing aging accuracy for mTHP workloads.
>>>>>
>>>>> Initialize young to nr so the original folio is accounted the same way
>>>>> as other young folios in the loop (young += nr).
>>>>>
>>>>> Signed-off-by: Hui Zhu <zhuhui@kylinos.cn>
>>>>> ---
>>>>
>>>> Good catch. Please also add the Fixes tag:
>>>>
>>>> Fixes: 56e5b60b2114 ("mm: support batched checking of the young flag for
>>>> MGLRU")
>>>>
>>>> With that,
>>>> Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
>>>
>>> Hi Baolin, Hui,
>>>
>>> I am not convinced this is the correct patch. test_and_clear_young_ptes_notify()
>>> only indicates that there is at least one young PTE among the nr PTEs;
>>> it does not mean that all of the PTEs are young.
>>>
>>> Am I missing something?
>>
>> You are right. But I explained why this is done in my original commit
>> 56e5b60b2114:
>>
>> "
>> Note that we also update the 'young' counter and
>> 'mm_stats[MM_LEAF_YOUNG]' counter with the batched count in the
>> lru_gen_look_around() and walk_pte_range().  However, the batched
>> operations may inflate these two counters, because in a large folio not
>> all PTEs may have been accessed. (Additionally, tracking how many PTEs
>> have been accessed within a large folio is not very meaningful, since
>> the mm core actually tracks access/dirty on a per-folio basis, not per
>> page).  The impact analysis is as follows:
>>
>> 1. The 'mm_stats[MM_LEAF_YOUNG]' counter has no functional impact and is
>> mainly for debugging.
>>
>> 2. The 'young' counter is used to decide whether to place the current
>> PMD entry into the bloom filters by suitable_to_scan() (so that next
>> time we can check whether it has been accessed again), which may set the
>> hash bit in the bloom filters for a PMD entry that hasn't seen much
>> access.  However, bloom filters inherently allow some error, so this
>> effect appears negligible.
>> "
>>
>> Based on this, I think changing it to 'nr' is reasonable. For an
>> accessed large folio, it's better to have the bloom filter rescan the
>> PMD and keep it in memory instead of reclaiming it incorrectly.
>>
> 
> I am not sure if this is the best policy, but we don't seem to have
> a practical way to get the exact number of accessed PTEs, so this may
> be acceptable. 

As I mentioned earlier, it seems unnecessary to implement this, since 
core-mm tracks access flag at per-folio granularity. Moreover, bloom 
filter itself allows for some error.

However, could we at least update the changelog to
> clarify that this is intentional?
> 
> " However, young is initialized to 1 regardless of the folio size.  When
>   the original folio is a large folio with nr PTEs, its young count is
>   underestimated by nr - 1.  This inconsistency can cause
>   suitable_to_scan() to return false, preventing the PMD from being added"

Agree. Looks better.

> Its young count is not underestimated; we are intentionally
> overestimating it. Also, nr does not necessarily equal
> folio_nr_pages(), does it?

Right.


      reply	other threads:[~2026-08-13  1:32 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-12  6:59 [PATCH] mm/mglru: Fix young counter undercount for large folios Hui Zhu
2026-08-12 10:17 ` Baolin Wang
2026-08-13  0:53   ` Barry Song
2026-08-13  1:09     ` Baolin Wang
2026-08-13  1:20       ` Barry Song
2026-08-13  1:31         ` Baolin Wang [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=d3471389-98b9-485a-9426-1796327b69c8@linux.alibaba.com \
    --to=baolin.wang@linux.alibaba.com \
    --cc=akpm@linux-foundation.org \
    --cc=axelrasmussen@google.com \
    --cc=baohua@kernel.org \
    --cc=david@kernel.org \
    --cc=hannes@cmpxchg.org \
    --cc=hui.zhu@linux.dev \
    --cc=kasong@tencent.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=mhocko@kernel.org \
    --cc=qi.zheng@linux.dev \
    --cc=shakeel.butt@linux.dev \
    --cc=weixugc@google.com \
    --cc=yuanchu@google.com \
    --cc=zhuhui@kylinos.cn \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox