From: Qi Zheng <qi.zheng@linux.dev>
To: Andrew Morton <akpm@linux-foundation.org>
Cc: hughd@google.com, baolin.wang@linux.alibaba.com,
usama.arif@linux.dev, brauner@kernel.org, david@kernel.org,
linux-mm@kvack.org, linux-kernel@vger.kernel.org,
Qi Zheng <zhengqi.arch@bytedance.com>
Subject: Re: [PATCH v4 0/4] make unused huge shrinker memcg aware
Date: Fri, 28 Aug 2026 10:48:38 +0800 [thread overview]
Message-ID: <51f48626-c38b-4cd6-a826-92d107ae35bb@linux.dev> (raw)
In-Reply-To: <20260827155020.530e28b0afa3b43fede9aa9c@linux-foundation.org>
Hi Andrew,
On 8/28/26 6:50 AM, Andrew Morton wrote:
> On Mon, 17 Aug 2026 17:03:24 +0800 Qi Zheng <qi.zheng@linux.dev> wrote:
>
>> Changes in v4:
>> Changes in v3:
>> Changes in v2:
>
> Thanks for the diligent versioning info. fyi, it is conventional to
> maintain this below the --- separator. It's not really the most
> important part of the [0/N]!
Got it.
>
>>
>> The shmem unused huge shrinker maintains a per-superblock list of inodes
>> whose tail huge folio extends beyond i_size. Because this list is not
>> memcg aware, reclaim triggered by memcg A can scan inodes across the
>> entire superblock and split huge folios charged to unrelated memcg B,
>> causing unexpected impact on it.
>>
>> In the worst case, memcg A has no reclaimable shmem at all, making the
>> reclaim entirely useless and incurring unnecessary latency. We observed
>> this in production, where page lock contention during split caused
>> multi-hundred-millisecond stalls:
>
> Ugh. That's the most important part!
>
>> tid 11340 comm scanner locked a page for 182264 us! kstack:
>> unlock_page+1
>> split_huge_page_to_list+3135
>> shmem_unused_huge_shrink+767
>> super_cache_scan+329
>> do_shrink_slab+291
>> shrink_slab+533
>> shrink_node+400
>> do_try_to_free_pages+206
>> try_to_free_mem_cgroup_pages+262
>> try_charge_memcg+591
>> mem_cgroup_charge+136
>> __handle_mm_fault+2431
>> handle_mm_fault+194
>> do_user_addr_fault+462
>> __do_page_fault+176
>> do_page_fault+48
>> page_fault+62
>>
>> Usama's recent patch [1] prevents the shmem unused shrinker from being
>> invoked during memcg-level reclaim altogether, but this is overly
>> conservative: we can do better by reclaiming only the shmem charged to
>> the reclaiming memcg.
>>
>> This series converts the shrinker list to a memcg-aware list_lru, so
>> that non-root memcg reclaim walks only candidates charged to the
>> reclaiming memcg. Global reclaim, root memcg reclaim and shmem quota
>> reclaim retain their existing global semantics.
>>
>> To avoid pinning a dying memcg through a long-lived CSS reference, each
>> inode stores an obj_cgroup reference instead of a mem_cgroup reference.
>> The list_lru add/delete paths resolve the current memcg from the objcg
>> under RCU, staying consistent with list_lru's own memcg migration on
>> offline.
>
> Sashiko said a few things and they look disturbing-if-true:
>
> https://sashiko.dev/#/patchset/cover.1786955972.git.zhengqi.arch@bytedance.com
>
> (Apologies if this has already been considered - we don't have ways of
> tracking all this (yet, I hope) apart from personal memory and personal
> memorys are quite fried at present)
As both Usama and I have pointed out [1][2], [PATCH v4 1/4] is the part
that got dropped during the merge. The complete patch [3] was actually
reviewed a while ago.
[1].https://lore.kernel.org/all/20260810101954.822260-1-usama.arif@linux.dev/
[2].
https://lore.kernel.org/all/9c7efd5f-f8d3-4926-acb4-34c326ffb1c3@linux.dev/
[3].
https://lore.kernel.org/all/20260715103516.2410175-1-usama.arif@linux.dev/
Thanks,
Qi
>
prev parent reply other threads:[~2026-08-28 2:49 UTC|newest]
Thread overview: 17+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-17 9:03 [PATCH v4 0/4] make unused huge shrinker memcg aware Qi Zheng
2026-08-17 9:03 ` [PATCH v4 1/4] fs: fix missed removal of super_fs_objects_eligible() Qi Zheng
2026-08-28 18:41 ` Andrew Morton
2026-08-29 1:49 ` Qi Zheng
2026-08-29 23:23 ` Andrew Morton
2026-08-31 2:28 ` Qi Zheng
2026-09-01 2:21 ` Andrew Morton
2026-09-01 2:25 ` Qi Zheng
2026-09-01 9:51 ` Usama Arif
2026-08-17 9:03 ` [PATCH v4 2/4] mm: memcontrol: make obj_cgroup_memcg() handle NULL objcg Qi Zheng
2026-08-17 11:29 ` Qi Zheng
2026-08-17 16:05 ` Shakeel Butt
2026-08-17 9:03 ` [PATCH v4 3/4] mm: shmem: move unused huge shrinklist queuing past the truncation check Qi Zheng
2026-08-17 9:03 ` [PATCH v4 4/4] mm: shmem: make unused huge shrinker memcg aware Qi Zheng
2026-08-18 3:54 ` Baolin Wang
2026-08-27 22:50 ` [PATCH v4 0/4] " Andrew Morton
2026-08-28 2:48 ` Qi Zheng [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=51f48626-c38b-4cd6-a826-92d107ae35bb@linux.dev \
--to=qi.zheng@linux.dev \
--cc=akpm@linux-foundation.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=brauner@kernel.org \
--cc=david@kernel.org \
--cc=hughd@google.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=usama.arif@linux.dev \
--cc=zhengqi.arch@bytedance.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.