From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wm1-f41.google.com (mail-wm1-f41.google.com [209.85.128.41]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 567ED1A275 for ; Tue, 21 Jul 2026 12:30:41 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.41 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784637043; cv=none; b=fpBVxWmITf6Y/vg59SKZsdhOJiPKHyNEC0OLLhNZQ+vnv0GMK3fm/3YWbOEX+JNygpeVoiPu3RJ1B0S6CMHuf0thuQMX3wkfgxDuCxpitXPu5lGFaYg7OzvGOCyJny/lQgEmqPW/Mw/MBKBB/CAQ52Oodg41t6F/S//S6NhihpY= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784637043; c=relaxed/simple; bh=Or/jQdZg73vOsMYQP/hRQxpMOgegn1tRQN3R9cniKEM=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=bzWQFG59A3brwzexODVr35t1h3HeLqCw5KtkOvr9fvQxxQzuauRKZ9vZcaWeXCIVseXPglBjGynMPWNF3xB42/gVP6ev+3ufqEizitQx4o4jfwzb9m6kFEzu/O+96Fqm1OCviir1QKz4xlYU7qEh9ssmvNGl8ch7WbOpUQi9GQ8= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=suse.com; spf=pass smtp.mailfrom=suse.com; dkim=pass (2048-bit key) header.d=suse.com header.i=@suse.com header.b=NRBUdwN/; arc=none smtp.client-ip=209.85.128.41 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=suse.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=suse.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=suse.com header.i=@suse.com header.b="NRBUdwN/" Received: by mail-wm1-f41.google.com with SMTP id 5b1f17b1804b1-4953e04ef16so53834705e9.2 for ; Tue, 21 Jul 2026 05:30:41 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=suse.com; s=google; t=1784637039; x=1785241839; darn=vger.kernel.org; h=in-reply-to:content-transfer-encoding:content-disposition :content-type:mime-version:references:message-id:subject:cc:to:from :date:from:to:cc:subject:date:message-id:reply-to:content-type; bh=Vw+R7y8hfT7A4z5+X72bXa8COEHFqKBWa6qMTipwJ+w=; b=NRBUdwN/9LglRjnbeS/8ruMA17WNyyP7uBoufO601dvkclUo0qMfeE9VRbKQRwWzuo hDEgmbzZ0cyyP80sjetULEc68k6njC3aE1faZ+zvvqjNrTkqFOly9rodAv9JIweHxLev jlDTFQ1mhC5FtjY8zMb1pQ4nxsSg21G5FX6Ocu2QodaKBU5sgDxFx4SeAXNFQxnwOSTq JngW71hFcJQPLDvyOZ72hT+nV9LkhWLvC4fEBFyPtPwj2hK9/ppGgHYA69pTK/G9lTxE Ai1Mt5w6SmO31vKES0OLDdyBiItI56sfv77GSxp/RJQxxZlF+HYBfA/jsid9n/8E0Dlj iMAA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784637039; x=1785241839; h=in-reply-to:content-transfer-encoding:content-disposition :content-type:mime-version:references:message-id:subject:cc:to:from :date:x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to:content-type; bh=Vw+R7y8hfT7A4z5+X72bXa8COEHFqKBWa6qMTipwJ+w=; b=cASUCOccZPlS3MNVJxy7uZnqIwKRHhpwRzxBeymdXdZw7MuJIWE1+/nMpPQcTagAx5 CJOf12GX9y4BMajZYsWZVWb2YgifVNcg3+jeMhgliHPeVYJDF/fIYDTDbKrZn2C6YTDK w8iHpY+NcwmJjI9xZ0CUoJmYNXIFg8c6numPf2tNfJOTBOQBgijkLqgDI4I+yIXZMWQb lPjFxzXeGqBH3ya8/JObeP3DJiiD7vA1vl6dcE4j4yHhomuUOPoJ9qxfYZ+3jWImniww pnupdpi9e6hIT5Z41XU5OpT5Z/Pq8aVbSLGr0cyWLxKI75EAB9NFPhwVMzMPJI0vIDL6 1UVw== X-Forwarded-Encrypted: i=1; AHgh+RrwtiKZNXcgKQmwQQC40h3uAnFQhWfwxZ+EKh+1z+GFGfs7tfTY0hS0Q9AQlqKv5h/I7STz5JEuMdZnLOs=@vger.kernel.org X-Gm-Message-State: AOJu0YzocJ+EJNhE9aKeJkLcidxYK4ktYLUcS3wA1h07gdhkYAXVZ4ZI 5n8J4vUP76r88hx79ailqYa4FHa1x7Kf9VnI4AYygjUjwG1LytAGCoTmD0A4DzbedvM= X-Gm-Gg: AfdE7ckO1xDI94dk+uBhRgbsT1qB+DNhHC4Usq7pX23nCvucREvb3GTUt62wq7wp0dG seS7wO9i2ez3Z0vK90CsWPtXDOVMAKP/QPipsFkqdRx5voe9T23vYW6X6mNB9fzqJ8d5aaMhd5M 3dFgNRHxM6Wp/5QlnsZA8PhlCZKgIAmL2Xkcn2t7qrZXJgf7z7HvB1dY8yc/H4HxJ4+vxYarBvp 3rNzlCT4xxwp34tkw88Gfl5so1+0+SnPbdbjtn9teoATexhnBkJkeG5LLAcXXtUn7QQAbx5Ixem X7SBqE+Wodk+YD8kzurcZe5b6jYbvhwvqtEZoQLue14KlDvWYe1DvqrrWzsFrBCBYKDuXS5DgD1 5I8rG63oMLO4ZluThLXr/gK3cmeXpQCB4Y44OmqMTNccFVxAs4i68H3WGPkW2HggiEpuMMgbwcD 7NFNipIIkr5Bae X-Received: by 2002:a05:600c:6dc1:b0:493:b967:178d with SMTP id 5b1f17b1804b1-4954a402dd9mr129364835e9.19.1784637039238; Tue, 21 Jul 2026 05:30:39 -0700 (PDT) Received: from localhost (109-81-80-79.rct.o2.cz. [109.81.80.79]) by smtp.gmail.com with ESMTPSA id ffacd0b85a97d-47f63ed1911sm41461110f8f.22.2026.07.21.05.30.38 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 21 Jul 2026 05:30:38 -0700 (PDT) Date: Tue, 21 Jul 2026 14:30:37 +0200 From: Michal Hocko To: Richard Chang Cc: Andrew Morton , Kairui Song , Qi Zheng , Shakeel Butt , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Johannes Weiner , David Hildenbrand , Lorenzo Stoakes , Oleg Nesterov , Suren Baghdasaryan , "T . J . Mercier" , Martin Liu , Minchan Kim , linux-mm@kvack.org, linux-kernel@vger.kernel.org Subject: Re: [PATCH v4] mm: vmscan: abort proactive reclaim early when freezing for suspend Message-ID: References: <20260720044103.905191-1-richardycc@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Disposition: inline Content-Transfer-Encoding: 8bit In-Reply-To: <20260720044103.905191-1-richardycc@google.com> On Mon 20-07-26 04:41:03, Richard Chang wrote: > Proactive reclaim (triggered via memory.reclaim or node sysfs) checks > for pending signals in its outer loop in user_proactive_reclaim(). > However, the inner reclaim loops—specifically scanning cgroups in > shrink_many() and evicting/aging folios in try_to_shrink_lruvec()—can > run for a long time before returning to the outer loop, especially on > systems with many cgroups or large memory sizes. > > During system suspend, the PM freezer attempts to freeze all tasks by > sending fake signals (setting TIF_SIGPENDING). Because the inner loops > do not check for pending signals, the proactive reclaim task can remain > stuck in kernel space for seconds, failing to enter the refrigerator in > a timely manner. This leads to suspend failures due to freeze timeouts, > a behavior observed on Android devices. > > This latency issue is specific to proactive reclaim because of its > large, user-defined reclaim targets (could be gigabytes). Since commit > 287d5fedb377 ("mm: memcg: use larger batches for proactive reclaim"), > proactive reclaim uses larger decaying batch sizes (starting at 1/4 of > the remaining target) to maintain throughput. This keeps the task in > the inner reclaim loop for extended periods. In contrast, reactive > reclaim (global/memcg) uses small targets (SWAP_CLUSTER_MAX, typically > 32 pages), allowing it to return to the outer loop and check signals > frequently. > > To fix this, add a signal_pending() check to should_abort_scan() for > proactive reclaim paths. Since should_abort_scan() is called within > the inner scanning and eviction loops, this allows proactive reclaim to > abort early and return to the outer loop in user_proactive_reclaim(). > > Additionally, return -ERESTARTSYS instead of -EINTR in > user_proactive_reclaim(). When interrupted by system suspend, returning > -ERESTARTSYS allows the task to enter the refrigerator and automatically > restart the syscall upon resume, making the freezer transparent to > userspace. For real signals, the signal layer will either restart the > syscall (if SA_RESTART is set) or return -EINTR to userspace. > > This fix specifically targets Multi-Gen LRU (MGLRU). Classic LRU's scan > targets per iteration are strictly bounded by get_scan_count(), which > ensures it returns to the outer loop more frequently. > > The check in should_abort_scan() is limited to proactive reclaim > (sc->proactive) to avoid inadvertently affecting reactive reclaim paths, > and is wrapped in unlikely() as it is a slow path. > > Suggested-by: Michal Hocko > Suggested-by: Oleg Nesterov > Signed-off-by: Richard Chang Acked-by: Michal Hocko Thanks > --- > v2: Update the commit message > v3: Return -ERESTARTSYS instead of -EINTR in user_proactive_reclaim > v4: Clarify -ERESTARTSYS vs SA_RESTART behavior > > mm/vmscan.c | 12 +++++++++++- > 1 file changed, 11 insertions(+), 1 deletion(-) > > diff --git a/mm/vmscan.c b/mm/vmscan.c > index 35c3bb15ae96..5aa4becacb7f 100644 > --- a/mm/vmscan.c > +++ b/mm/vmscan.c > @@ -4929,6 +4929,9 @@ static bool should_abort_scan(struct lruvec *lruvec, struct scan_control *sc) > int i; > enum zone_watermarks mark; > > + if (unlikely(sc->proactive && signal_pending(current))) > + return true; > + > if (sc->nr_reclaimed >= max(sc->nr_to_reclaim, compact_gap(sc->order))) > return true; > > @@ -7909,8 +7912,15 @@ int user_proactive_reclaim(char *buf, > unsigned long batch_size = (nr_to_reclaim - nr_reclaimed) / 4; > unsigned long reclaimed; > > + /* > + * Return -ERESTARTSYS to allow the freezer to interrupt the > + * task. The syscall will be transparently restarted upon > + * resume. For real signals, it either restarts the syscall > + * (if SA_RESTART is set) or is converted to -EINTR by the > + * signal layer. > + */ > if (signal_pending(current)) > - return -EINTR; > + return -ERESTARTSYS; > > /* > * This is the final attempt, drain percpu lru caches in the > -- > 2.55.0.229.g6434b31f56-goog -- Michal Hocko SUSE Labs