From: Andrew Morton <akpm@linux-foundation.org>
To: Richard Chang <richardycc@google.com>
Cc: Kairui Song <kasong@tencent.com>, Qi Zheng <qi.zheng@linux.dev>,
Shakeel Butt <shakeel.butt@linux.dev>,
Barry Song <baohua@kernel.org>,
Axel Rasmussen <axelrasmussen@google.com>,
Yuanchu Xie <yuanchu@google.com>, Wei Xu <weixugc@google.com>,
Johannes Weiner <hannes@cmpxchg.org>,
David Hildenbrand <david@kernel.org>,
Michal Hocko <mhocko@kernel.org>,
Lorenzo Stoakes <ljs@kernel.org>, Oleg Nesterov <oleg@redhat.com>,
Suren Baghdasaryan <surenb@google.com>,
"T . J . Mercier" <tjmercier@google.com>,
Martin Liu <liumartin@google.com>,
Minchan Kim <minchan@kernel.org>,
linux-mm@kvack.org, linux-kernel@vger.kernel.org,
Michal Hocko <mhocko@suse.com>
Subject: Re: [PATCH v4] mm: vmscan: abort proactive reclaim early when freezing for suspend
Date: Sun, 19 Jul 2026 22:23:06 -0700 [thread overview]
Message-ID: <20260719222306.540829de219e12e21396d260@linux-foundation.org> (raw)
In-Reply-To: <20260720044103.905191-1-richardycc@google.com>
On Mon, 20 Jul 2026 04:41:03 +0000 Richard Chang <richardycc@google.com> wrote:
> Proactive reclaim (triggered via memory.reclaim or node sysfs) checks
> for pending signals in its outer loop in user_proactive_reclaim().
> However, the inner reclaim loops—specifically scanning cgroups in
> shrink_many() and evicting/aging folios in try_to_shrink_lruvec()—can
> run for a long time before returning to the outer loop, especially on
> systems with many cgroups or large memory sizes.
>
> During system suspend, the PM freezer attempts to freeze all tasks by
> sending fake signals (setting TIF_SIGPENDING). Because the inner loops
> do not check for pending signals, the proactive reclaim task can remain
> stuck in kernel space for seconds, failing to enter the refrigerator in
> a timely manner. This leads to suspend failures due to freeze timeouts,
> a behavior observed on Android devices.
>
> This latency issue is specific to proactive reclaim because of its
> large, user-defined reclaim targets (could be gigabytes). Since commit
> 287d5fedb377 ("mm: memcg: use larger batches for proactive reclaim"),
> proactive reclaim uses larger decaying batch sizes (starting at 1/4 of
> the remaining target) to maintain throughput. This keeps the task in
> the inner reclaim loop for extended periods. In contrast, reactive
> reclaim (global/memcg) uses small targets (SWAP_CLUSTER_MAX, typically
> 32 pages), allowing it to return to the outer loop and check signals
> frequently.
So 287d5fedb377 led to suspend failures on MGLRU-using kernels.
That's a regression which justifies a Fixes: and a cc:stable, don't
people agree?
AI review asked a couple of serious-sounding questions:
https://sashiko.dev/#/patchset/20260720044103.905191-1-richardycc@google.com
> To fix this, add a signal_pending() check to should_abort_scan() for
> proactive reclaim paths. Since should_abort_scan() is called within
> the inner scanning and eviction loops, this allows proactive reclaim to
> abort early and return to the outer loop in user_proactive_reclaim().
>
> Additionally, return -ERESTARTSYS instead of -EINTR in
> user_proactive_reclaim(). When interrupted by system suspend, returning
> -ERESTARTSYS allows the task to enter the refrigerator and automatically
> restart the syscall upon resume, making the freezer transparent to
> userspace. For real signals, the signal layer will either restart the
> syscall (if SA_RESTART is set) or return -EINTR to userspace.
>
> This fix specifically targets Multi-Gen LRU (MGLRU). Classic LRU's scan
> targets per iteration are strictly bounded by get_scan_count(), which
> ensures it returns to the outer loop more frequently.
>
> The check in should_abort_scan() is limited to proactive reclaim
> (sc->proactive) to avoid inadvertently affecting reactive reclaim paths,
> and is wrapped in unlikely() as it is a slow path.
>
next prev parent reply other threads:[~2026-07-20 5:23 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
[not found] <CALC_0q8ntqjCwC6CBN_0h_z9-_+pDp8-pqQgZ39O2s8ke==9cQ@mail.gmail.com>
2026-07-17 18:12 ` [PATCH v3] mm: vmscan: abort proactive reclaim early when freezing for suspend Richard Chang
2026-07-17 18:47 ` Oleg Nesterov
2026-07-20 4:37 ` Richard Chang
2026-07-20 4:41 ` [PATCH v4] " Richard Chang
2026-07-20 5:23 ` Andrew Morton [this message]
2026-07-20 12:04 ` Richard Chang
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260719222306.540829de219e12e21396d260@linux-foundation.org \
--to=akpm@linux-foundation.org \
--cc=axelrasmussen@google.com \
--cc=baohua@kernel.org \
--cc=david@kernel.org \
--cc=hannes@cmpxchg.org \
--cc=kasong@tencent.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=liumartin@google.com \
--cc=ljs@kernel.org \
--cc=mhocko@kernel.org \
--cc=mhocko@suse.com \
--cc=minchan@kernel.org \
--cc=oleg@redhat.com \
--cc=qi.zheng@linux.dev \
--cc=richardycc@google.com \
--cc=shakeel.butt@linux.dev \
--cc=surenb@google.com \
--cc=tjmercier@google.com \
--cc=weixugc@google.com \
--cc=yuanchu@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox