From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wm1-f52.google.com (mail-wm1-f52.google.com [209.85.128.52]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 87B69411A1B for ; Fri, 7 Aug 2026 20:07:34 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.52 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786133256; cv=none; b=ey/aW9qTTwIYPCEdgA3M4F7V7ZFhzhl/p4lviHpLGa9yiEGi+kiKKZZzW8kLiPr6LJCL8bEbJ22MeXaOzYLaIItkDaaqWH/5PjijHkq3dva1tCsMKxcj79EKC6KFIqyGFaK8P/bBU63JqaNTy1XJkkvdAUZH4umvlPxXdLTt7Cc= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786133256; c=relaxed/simple; bh=h+x6kpisysewhLc9kX8A2S5kwjryOVqBoH0OmWTWEFM=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=ZlBE7Up7mv3p4/J0FN9458jv0jg1SJ9Dk4TuyN2UWUc5ywkGYHIl1kGyxgdw+/i8rqeCaYnESyZrK+Q4NvdyrQinafpdIm2AXV7C+ebksBE2KatMkDX8IPW6xWL9NQWKE6uPNGXSaBbRI+K/BzHq+oGE+pPaIUNcwTsyj5Ct6po= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=suse.com; spf=pass smtp.mailfrom=suse.com; dkim=pass (2048-bit key) header.d=suse.com header.i=@suse.com header.b=eNO6GHWm; arc=none smtp.client-ip=209.85.128.52 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=suse.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=suse.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=suse.com header.i=@suse.com header.b="eNO6GHWm" Received: by mail-wm1-f52.google.com with SMTP id 5b1f17b1804b1-4954f5e8020so20207725e9.2 for ; Fri, 07 Aug 2026 13:07:34 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=suse.com; s=google; t=1786133253; x=1786738053; darn=vger.kernel.org; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:from:to:cc:subject :date:message-id:reply-to:content-type; bh=d/MJr6Oe7Mwi0NIRhRgjfbHGJr7/OP8f2gsXk/Z1AZs=; b=eNO6GHWmXCg/rfx/iGUK3ncE1Oi3XhNPMzlonDEFEK/5uVPe90/92vkprAQj///2Qc Rwe+LAuUlWmXf212EuFot5YDxybaN1gSwrd739Sp8z3Se9aYYs87hiB3x7OUUkO6gCDg XV/FenGEnrecvx/jFFOqZcjhkZ1pTbw8Blzp76q1JLmBON2f5chI8azTXxKYJindaVpA zm0c/3+boPjz/+a+EnKbZyFBI2Pq4f0dziDIvKgvrYPEQCTsEli3bNDBnP3ZWLpuEhDi 47ow6m5gvuUg7/FXROkrorAT0i30vurlThteVD5oYHh6bsb2YeQ39RoihACwYFSfEYDO pXDg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786133253; x=1786738053; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=d/MJr6Oe7Mwi0NIRhRgjfbHGJr7/OP8f2gsXk/Z1AZs=; b=J9yzzUcU86v8NSKIzWVw2KgEvcol893LADmFXMl9HkSubOOFD8bF541A+IG1HtAA3B MqOlji1R6T3d8ibfjoqC++9KnPceo4+UcgT98c/MhZ09X2j6FEFdZ4YjaliquNAoCQo9 uL+vPYLwo9GUckV9CrYberXT9Qrrm0kNA78sl5BkffC44Ch5ZNZJEUnxtTKuWpu+TSlw mEK22MGk5dpISfTq2klIMbT6w/KzkSOOg0qw0ZfrfBRaa4ZD+YczghqqNajBrj2KNQeX j6/d/D+bJuadVpkzGnpcAq2IifvLIt2SUSuqY9TJ+H/bwTdEveyJVbRSV8oUKLxEWgu3 95Tg== X-Forwarded-Encrypted: i=1; AHgh+RpncspLN+j7FfYqs+Bo7+IGqmEk+73HPFr5AXiBvFA2sH3//7hIMZ2NwBLXGHlz+up+DgeMMX9ex5svb2s=@vger.kernel.org X-Gm-Message-State: AOJu0Yw5KsrhQG5afGcf088nIA8uTxwbNB1ZuVsRtEGoFDcFLGjjV+Pb HXjZEyNkV0OEWouUelvgMsdMaJ1lnW61pMN43bCU3MEdP7XB+MbPb+lUP5Dis6h0xDw= X-Gm-Gg: AR+sD12zRba2WcdXpz+QQPjzlzPPCTWs3mYN8QICtIBSOWGfn+YOEJN2wMroY3jBDLc FIdqJgOjXmCaprbaN5ILk2C6XiekdeYega+d9T5O7FlreOeBU4p9b304tzJG87M2xcyF5qTySfI jwwZzJboff2GPGgZX56G2gW45YvI4vVQi0ZMVZgokHUlvMsoCWxFOqNfSPOMZYMX5BvVC8Lq214 UfFF9K1EljTvHFoqfgZJbIlthih0tlh+dAFqWEVfu33ezNkoodHX8HlbXtm2laAyD+VyzqButgr drDkNaEAOJODXiSR41VrvQyE2oppDNOSaUD0O4ohtM48c1Lox+5iTwTw5U5NzPJ37eQjhLgA0qM FoKN5hvIkQY7cGo9jaoNQT8eBs2HDHJaSMkdb5JWsuE/E/pLjzKcE9/Qqg0ciAl3Yz+fzM5g/PM Pd+hlULuEugC25/lJygyj5PvFc92dToCJahLJQdGICuMHCu+txWA0e1dGNKmxhnY/iIz1pdRzuv MhITXGgy9XI X-Received: by 2002:a05:600c:b85:b0:496:c93d:e2f with SMTP id 5b1f17b1804b1-4995e0dff0amr84684845e9.15.1786133252699; Fri, 07 Aug 2026 13:07:32 -0700 (PDT) Received: from localhost (109-81-83-166.rct.o2.cz. [109.81.83.166]) by smtp.gmail.com with ESMTPSA id ffacd0b85a97d-480021fd392sm8251138f8f.31.2026.08.07.13.07.31 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 07 Aug 2026 13:07:32 -0700 (PDT) Date: Fri, 7 Aug 2026 22:07:30 +0200 From: Michal Hocko To: Audra Mitchell Cc: david@kernel.org, jocolema@redhat.com, raquini@redhat.com, Johannes Weiner , Roman Gushchin , Shakeel Butt , Muchun Song , Andrew Morton , cgroups@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org Subject: Re: [PATCH] Fix unbounded loop within try_charge_memcg Message-ID: References: <20260806151004.1825320-1-audra@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: On Fri 07-08-26 15:13:22, Audra Mitchell wrote: > On Fri, Aug 07, 2026 at 08:47:10PM +0200, Michal Hocko wrote: > > On Fri 07-08-26 11:40:06, Audra Mitchell wrote: > > > On Fri, Aug 07, 2026 at 10:06:53AM +0200, Michal Hocko wrote: > > > > On Thu 06-08-26 11:10:03, Audra Mitchell wrote: > > > > Well, both global and memcg reclaim share the reclaim logic. Both of > > them try to exercise all reclaim priorities (i.e. check whole eligible > > LRU lists) and they fall back to OOM killer only if there is no other > > option left. For the global case should_reclaim_retry is the gate keeper > > around direct reclaim retries while for the memcg we have more or less > > fixed number of retries. > > > From what you are describing above those users might be hitting reclaim trashing. > > I.e. last small portion of a reclaimable memory is bounced back and > > forth for the workload to make tiny but steady forward progress. While > > OOM killer might help to stop the suffering and restart the workload > > sooner I would generally recommend revisiting limits set for the > > particular workload. Especially if restarting it might lead to the same > > state sooner or later. Watching PSI metric would be a good start to see > > how the workload behaves wrt memory stalling. User space oom handlers > > might be a proper measure as well but that will always be safeguard > > rather than a solution. > > Both of these recommendations were made to the end customer, with a > significant emphasis on reviewing the workload and the limits in place. > > However, while reviewing the code, I was truly surprised to find retry > pathways in try_charge_memcg that do not decrement the counter, such as this: > > nr_reclaimed = try_to_free_mem_cgroup_pages(mem_over_limit, nr_pages, > gfp_mask, reclaim_options, NULL); > psi_memstall_leave(&pflags); > > if (mem_cgroup_margin(mem_over_limit) >= nr_pages) > goto retry; > > If the intent is to have a counter of nr_retries, I am genuinely surprised > we are comfortable not adhering to said counter. The said counter is conuting retries without any real progress. Similarly to the global reclaim. If we get some margin here we are making progress so we are not really OOM yet. > The patch is not meant to > change any heuristic, but only enforce a counter that is already in place. It is very much chaning the heuristic ;) With your proposed change we would be hitting OOM much more eaiser and often. Not something a lot of user would appreciate. -- Michal Hocko SUSE Labs