From: Qi Zheng <qi.zheng@linux.dev>
To: Andrew Morton <akpm@linux-foundation.org>
Cc: hughd@google.com, baolin.wang@linux.alibaba.com,
usama.arif@linux.dev, brauner@kernel.org, david@kernel.org,
linux-mm@kvack.org, linux-kernel@vger.kernel.org,
Qi Zheng <zhengqi.arch@bytedance.com>
Subject: Re: [PATCH v4 0/4] make unused huge shrinker memcg aware
Date: Fri, 28 Aug 2026 10:48:38 +0800 [thread overview]
Message-ID: <51f48626-c38b-4cd6-a826-92d107ae35bb@linux.dev> (raw)
In-Reply-To: <20260827155020.530e28b0afa3b43fede9aa9c@linux-foundation.org>
Hi Andrew,
On 8/28/26 6:50 AM, Andrew Morton wrote:
> On Mon, 17 Aug 2026 17:03:24 +0800 Qi Zheng <qi.zheng@linux.dev> wrote:
>
>> Changes in v4:
>> Changes in v3:
>> Changes in v2:
>
> Thanks for the diligent versioning info. fyi, it is conventional to
> maintain this below the --- separator. It's not really the most
> important part of the [0/N]!
Got it.
>
>>
>> The shmem unused huge shrinker maintains a per-superblock list of inodes
>> whose tail huge folio extends beyond i_size. Because this list is not
>> memcg aware, reclaim triggered by memcg A can scan inodes across the
>> entire superblock and split huge folios charged to unrelated memcg B,
>> causing unexpected impact on it.
>>
>> In the worst case, memcg A has no reclaimable shmem at all, making the
>> reclaim entirely useless and incurring unnecessary latency. We observed
>> this in production, where page lock contention during split caused
>> multi-hundred-millisecond stalls:
>
> Ugh. That's the most important part!
>
>> tid 11340 comm scanner locked a page for 182264 us! kstack:
>> unlock_page+1
>> split_huge_page_to_list+3135
>> shmem_unused_huge_shrink+767
>> super_cache_scan+329
>> do_shrink_slab+291
>> shrink_slab+533
>> shrink_node+400
>> do_try_to_free_pages+206
>> try_to_free_mem_cgroup_pages+262
>> try_charge_memcg+591
>> mem_cgroup_charge+136
>> __handle_mm_fault+2431
>> handle_mm_fault+194
>> do_user_addr_fault+462
>> __do_page_fault+176
>> do_page_fault+48
>> page_fault+62
>>
>> Usama's recent patch [1] prevents the shmem unused shrinker from being
>> invoked during memcg-level reclaim altogether, but this is overly
>> conservative: we can do better by reclaiming only the shmem charged to
>> the reclaiming memcg.
>>
>> This series converts the shrinker list to a memcg-aware list_lru, so
>> that non-root memcg reclaim walks only candidates charged to the
>> reclaiming memcg. Global reclaim, root memcg reclaim and shmem quota
>> reclaim retain their existing global semantics.
>>
>> To avoid pinning a dying memcg through a long-lived CSS reference, each
>> inode stores an obj_cgroup reference instead of a mem_cgroup reference.
>> The list_lru add/delete paths resolve the current memcg from the objcg
>> under RCU, staying consistent with list_lru's own memcg migration on
>> offline.
>
> Sashiko said a few things and they look disturbing-if-true:
>
> https://sashiko.dev/#/patchset/cover.1786955972.git.zhengqi.arch@bytedance.com
>
> (Apologies if this has already been considered - we don't have ways of
> tracking all this (yet, I hope) apart from personal memory and personal
> memorys are quite fried at present)
As both Usama and I have pointed out [1][2], [PATCH v4 1/4] is the part
that got dropped during the merge. The complete patch [3] was actually
reviewed a while ago.
[1].https://lore.kernel.org/all/20260810101954.822260-1-usama.arif@linux.dev/
[2].
https://lore.kernel.org/all/9c7efd5f-f8d3-4926-acb4-34c326ffb1c3@linux.dev/
[3].
https://lore.kernel.org/all/20260715103516.2410175-1-usama.arif@linux.dev/
Thanks,
Qi
>
prev parent reply other threads:[~2026-08-28 2:49 UTC|newest]
Thread overview: 12+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-17 9:03 [PATCH v4 0/4] make unused huge shrinker memcg aware Qi Zheng
2026-08-17 9:03 ` [PATCH v4 1/4] fs: fix missed removal of super_fs_objects_eligible() Qi Zheng
2026-08-28 18:41 ` Andrew Morton
2026-08-29 1:49 ` Qi Zheng
2026-08-17 9:03 ` [PATCH v4 2/4] mm: memcontrol: make obj_cgroup_memcg() handle NULL objcg Qi Zheng
2026-08-17 11:29 ` Qi Zheng
2026-08-17 16:05 ` Shakeel Butt
2026-08-17 9:03 ` [PATCH v4 3/4] mm: shmem: move unused huge shrinklist queuing past the truncation check Qi Zheng
2026-08-17 9:03 ` [PATCH v4 4/4] mm: shmem: make unused huge shrinker memcg aware Qi Zheng
2026-08-18 3:54 ` Baolin Wang
2026-08-27 22:50 ` [PATCH v4 0/4] " Andrew Morton
2026-08-28 2:48 ` Qi Zheng [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=51f48626-c38b-4cd6-a826-92d107ae35bb@linux.dev \
--to=qi.zheng@linux.dev \
--cc=akpm@linux-foundation.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=brauner@kernel.org \
--cc=david@kernel.org \
--cc=hughd@google.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=usama.arif@linux.dev \
--cc=zhengqi.arch@bytedance.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox