From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wm1-f47.google.com (mail-wm1-f47.google.com [209.85.128.47]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E0AAC2E7167 for ; Tue, 6 Jan 2026 14:00:41 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.47 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1767708044; cv=none; b=VCzEBcyNhWujJVT6OVrzd+90tusBCxdbgdGFde7amVR+89JFFsmYrd8oXoKRH4pSPL8bNuxxa+2Ud9s9JjMmxcylgEJsnMDXU86GGOzeTpu1HIscFSShUyFrGXj1H6esShNDjFt+7SKsctnN+AxWGTiQ5f2GRcxCvbpXMfNbkFc= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1767708044; c=relaxed/simple; bh=VXmH6a10VB6SS3cIM+CcfdU7cSBwS61iUr1Hse6wjTQ=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=KOIPS5dyf8j+AFdPUlbDs8kSTt+4n0IQ7C4vXRYhAsQX9pUJpPHCyTTthueE42dgeS6Ojda/aWMvRA5vK7UVTf+FyWzUAliwCHFyjn7wlcnLHz518A6AMNsQB1DBQ93cohQCZQjQy2op9e2agTlKmuN+0DRGV7pW+0mgzUqp0Mo= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=suse.com; spf=pass smtp.mailfrom=suse.com; dkim=pass (2048-bit key) header.d=suse.com header.i=@suse.com header.b=G3FVvxPv; arc=none smtp.client-ip=209.85.128.47 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=suse.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=suse.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=suse.com header.i=@suse.com header.b="G3FVvxPv" Received: by mail-wm1-f47.google.com with SMTP id 5b1f17b1804b1-4775ae77516so10923965e9.1 for ; Tue, 06 Jan 2026 06:00:41 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=suse.com; s=google; t=1767708040; x=1768312840; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=qD8hUJoWpmRPh3qLUw4WPXMyrpMQhfKMlK8mtIKaFIY=; b=G3FVvxPvUAFvFpioczYHADJR5jVfkzjZ1PgkUU27pBAhtZGtZWQJBVhwBNcIaYbwM+ fhJ2Cmt8mzGJDBA1k34oMwNhPRi7OSdbiNoxOmeGplHns6ncty/Xo1geNeyRYxuSVhdx pKnFGwNCSnDccjwRpCUKNlEfWLvoF8ZqYuYgJYOTM/BWDBFJ/O/ExMQxgHcbd0s1ijnj qEvBOIAabD7rkKWQQ5gtQ/mc55tFXkBhtaVtL5YLlxfslSFPZSt6JC5PSD8sk/f+JVAW Uo27ceR2/XWrsJV5bmsAfPGcsuZ1eXHFMoij0Hlk+/4EALaD0XMXvZprgB/sAqjqy2Rl 90VQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1767708040; x=1768312840; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-gg:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=qD8hUJoWpmRPh3qLUw4WPXMyrpMQhfKMlK8mtIKaFIY=; b=MMiaTHziDtkpGzuxen2yFIfwmUN+bHcw0DJDbz4OlgKDQKbTB2yqGv/ICngYrheoC8 D6KGlE28Qn3kv8GSd1hhlDmxO3miN8PWt/kmUcO0Dh3yUoyQEJaeHsCuP6QEVkExFRap Inf6yu1u/XkLFsnPyym/5KbaP/ftIzGOj9DWuGE8rloEFY8NAPwasgrpngAXAEu35hY4 TSGku6FS49nzDCCM2LZpIHAmSIMw0SKc6/69GXvmugxfVwAPz5txUR5q1nAT0v8v69W8 zsc/pgkp6kgJ9+5QAwcYK4o+uRtxpNmjEFlUhhnrdGxnHYK5B0JGgXaRAtpPtoEtOwoN NnpA== X-Forwarded-Encrypted: i=1; AJvYcCUZlbY74oq8y8skc5VNW1tazezMrHHkSUxRSSzGrfol7llSp9n46gd44xXisKv6mYi+l8PEDkfMU/PFM8o=@vger.kernel.org X-Gm-Message-State: AOJu0YxO0KPyEFXOP0MeAjyuwK/Gi1CzhQV1jvPdBrlY0C+IwOUzCu2E ZhyvuJKGQB6HFj1Z94Is4MvLV/5rasLStMDAKH283PcLkRWM1UuBb/Ajc1w8ss2SpoE= X-Gm-Gg: AY/fxX7VLLRvSUKP9BQkNM1fWky/lB9HlpOVXKbJmDcQXhQB3CNrHLjpBKy0GonTwNr 5GBb3OUXYCrtGJX95FssgUQ46aTbszptTmjtDPUUVqCKkSZ/wdadfBaF1kzDRz9t/FTyVGfROMj YCcnXt0eSeh2K5q8Y4wZuTpXmCZmB2qOF5qr5zTVBHhhg4uqaGQbkAz3vL3oJuVv5yvHIeBgmUQ 5cNcdfARUJcU5cR5SHfitvGj9kd29CBVif62sMp6LG0eu5qM3KSAW1b+e4WubKkR3zuB6KEhJ8z 6Vj2uzo9TJEhirCoWQPT0ie+meNmF6wKNK/OV+1WoJzZFpxuH1NJQi0rA9shY5kkXAG3JlaJC1R RUxjch4n05wp9rr2FmmgYPJAiH/Ij1UlXlXhKMxEcPn9Hi0jFVs8N3U7WTAb7aEdvVW+1FUYOhK CXqwWIgyXIhlN6YWUE8ftZ2i5N X-Google-Smtp-Source: AGHT+IGHQFZt2QvFYPr9z2ERUKepowjPPfzDfhbLzQI8A0YPz1euQTjI1RCGsUbf75Xmp/4xK6snFw== X-Received: by 2002:a05:600c:4e93:b0:479:3a86:dc1b with SMTP id 5b1f17b1804b1-47d7f0a9556mr30131465e9.37.1767708040017; Tue, 06 Jan 2026 06:00:40 -0800 (PST) Received: from localhost (109-81-90-116.rct.o2.cz. [109.81.90.116]) by smtp.gmail.com with ESMTPSA id ffacd0b85a97d-432bd5ee5e3sm4534181f8f.35.2026.01.06.06.00.39 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 06 Jan 2026 06:00:39 -0800 (PST) Date: Tue, 6 Jan 2026 15:00:38 +0100 From: Michal Hocko To: Vlastimil Babka Cc: Andrew Morton , Suren Baghdasaryan , Brendan Jackman , Johannes Weiner , Zi Yan , David Rientjes , David Hildenbrand , Lorenzo Stoakes , "Liam R. Howlett" , Mike Rapoport , Joshua Hahn , Pedro Falcato , linux-mm@kvack.org, linux-kernel@vger.kernel.org Subject: Re: [PATCH mm-unstable v3 3/3] mm/page_alloc: simplify __alloc_pages_slowpath() flow Message-ID: References: <20260106-thp-thisnode-tweak-v3-0-f5d67c21a193@suse.cz> <20260106-thp-thisnode-tweak-v3-3-f5d67c21a193@suse.cz> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20260106-thp-thisnode-tweak-v3-3-f5d67c21a193@suse.cz> On Tue 06-01-26 12:52:38, Vlastimil Babka wrote: > The actions done before entering the main retry loop include waking up > kswapds and an allocation attempt with the precise alloc_flags. > Then in the loop we keep waking up kswapds, and we retry the allocation > with flags potentially further adjusted by being allowed to use reserves > (due to e.g. becoming an OOM killer victim). > > We can adjust the retry loop to keep only one instance of waking up > kswapds and allocation attempt. Introduce the can_retry_reserves > variable for retrying once when we become eligible for reserves. It is > still useful not to evaluate reserve_flags immediately for the first > allocation attempt, because it's better to first try succeed in a > non-preferred zone above the min watermark before allocating immediately > from the preferred zone below min watermark. > > Additionally move the cpuset update checks introduced by e05741fb10c3 > ("mm/page_alloc.c: avoid infinite retries caused by cpuset race") > further down the retry loop. It's enough to do the checks only before > reaching any potentially infinite 'goto retry;' loop. > > There should be no meaningful functional changes. The change of exact > moments the retry for reserves and cpuset updates are checked should not > result in different outomes modulo races with concurrent allocator > activity. > > Signed-off-by: Vlastimil Babka LGTM Acked-by: Michal Hocko > --- > mm/page_alloc.c | 41 +++++++++++++++++++++++------------------ > 1 file changed, 23 insertions(+), 18 deletions(-) > > diff --git a/mm/page_alloc.c b/mm/page_alloc.c > index 3b2579c5716f..c02564042618 100644 > --- a/mm/page_alloc.c > +++ b/mm/page_alloc.c > @@ -4716,6 +4716,7 @@ __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order, > unsigned int zonelist_iter_cookie; > int reserve_flags; > bool compact_first = false; > + bool can_retry_reserves = true; > > if (unlikely(nofail)) { > /* > @@ -4783,6 +4784,8 @@ __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order, > goto nopage; > } > > +retry: > + /* Ensure kswapd doesn't accidentally go to sleep as long as we loop */ > if (alloc_flags & ALLOC_KSWAPD) > wake_all_kswapds(order, gfp_mask, ac); > > @@ -4794,19 +4797,6 @@ __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order, > if (page) > goto got_pg; > > -retry: > - /* > - * Deal with possible cpuset update races or zonelist updates to avoid > - * infinite retries. > - */ > - if (check_retry_cpuset(cpuset_mems_cookie, ac) || > - check_retry_zonelist(zonelist_iter_cookie)) > - goto restart; > - > - /* Ensure kswapd doesn't accidentally go to sleep as long as we loop */ > - if (alloc_flags & ALLOC_KSWAPD) > - wake_all_kswapds(order, gfp_mask, ac); > - > reserve_flags = __gfp_pfmemalloc_flags(gfp_mask); > if (reserve_flags) > alloc_flags = gfp_to_alloc_flags_cma(gfp_mask, reserve_flags) | > @@ -4821,12 +4811,18 @@ __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order, > ac->nodemask = NULL; > ac->preferred_zoneref = first_zones_zonelist(ac->zonelist, > ac->highest_zoneidx, ac->nodemask); > - } > > - /* Attempt with potentially adjusted zonelist and alloc_flags */ > - page = get_page_from_freelist(gfp_mask, order, alloc_flags, ac); > - if (page) > - goto got_pg; > + /* > + * The first time we adjust anything due to being allowed to > + * ignore memory policies or watermarks, retry immediately. This > + * allows us to keep the first allocation attempt optimistic so > + * it can succeed in a zone that is still above watermarks. > + */ > + if (can_retry_reserves) { > + can_retry_reserves = false; > + goto retry; > + } > + } > > /* Caller is not willing to reclaim, we can't balance anything */ > if (!can_direct_reclaim) > @@ -4889,6 +4885,15 @@ __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order, > !(gfp_mask & __GFP_RETRY_MAYFAIL))) > goto nopage; > > + /* > + * Deal with possible cpuset update races or zonelist updates to avoid > + * infinite retries. No "goto retry;" can be placed above this check > + * unless it can execute just once. > + */ > + if (check_retry_cpuset(cpuset_mems_cookie, ac) || > + check_retry_zonelist(zonelist_iter_cookie)) > + goto restart; > + > if (should_reclaim_retry(gfp_mask, order, ac, alloc_flags, > did_some_progress > 0, &no_progress_loops)) > goto retry; > > -- > 2.52.0 -- Michal Hocko SUSE Labs