From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 479D41A38F9 for ; Mon, 23 Mar 2026 13:36:32 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1774272992; cv=none; b=EI9u93/XNTwIT0OFUH/iDLJrr5jKNyLtwY/3SlhmMdVGFZcrkFh/m6CXVs/YOp4vkEDjPHYoaXjn9ByKLsXQNdZa0oZfApcZIRWw4i9/Y1alEzj4/IAoqwEeDKh+e3LA6gAWA5TPm4/gGH83TLuExxbzNNU0P2iaq63CZWD52+g= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1774272992; c=relaxed/simple; bh=CPpDOfhpdW9UBLtwsezPEAAxhyR2+9b1dCeOhUI5JGQ=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=H5KbFoaWl3uFG4ErdCKU7RMT4aHAy+XStjJo5l2NhKUC56eHlMVGbtKWJbIeX6ohb2CHuE5cr6pWaSdxaL/w3pnnv412ZVK7/2szvBwt1bGH3tqfAOVtu6XAJ9vr35exwm/7rQJtbpoDmIy0HHcLTC//W4hE4cRgkBmBXkMFZtQ= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=BO9uapdY; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="BO9uapdY" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 7C8D5C2BC9E; Mon, 23 Mar 2026 13:36:30 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1774272991; bh=CPpDOfhpdW9UBLtwsezPEAAxhyR2+9b1dCeOhUI5JGQ=; h=Date:Subject:To:Cc:References:From:In-Reply-To:From; b=BO9uapdYjE0Kvpst573nixVmOGKL3AhixRvJB0IGJRs6pLvII6DVOdZd5tv0R2Pit UaPMiUqUbaZcRHHQh5SDOQaFT0NmR/wCthC3xqMQ2E/MWmEIUGN7+8fzZXO3hPH4pw K0lmplr7nfH49zdAHS9cm+zqIwrKCJ/bUlBtQBIBf6bDiZ3lxKXgXI62aFR9+g6v+W MCeB5N8aECpgkXhMyKOE4UnxRN+cCAg0KSQ+/4UkHs4gQyXIrhDvWi25HaXO0gLRlW JKcs9KnFCM62E7op1Xu9SkZvz+PWlu63FoYXOxh3LKO1xHzGH+4qbF6hbtVaIopZhs 1MNLV8L8o7CAg== Message-ID: <44aadc9c-28d2-497d-ba4e-659517e8ca47@kernel.org> Date: Mon, 23 Mar 2026 14:36:28 +0100 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH] mm/page_alloc: don't increase highatomic reserve after pcp alloc To: Frank van der Linden , akpm@linux-foundation.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org Cc: Michal Hocko , Johannes Weiner , Zhiguo Jiang References: <20260320173426.1831267-1-fvdl@google.com> From: "Vlastimil Babka (SUSE)" Content-Language: en-US In-Reply-To: <20260320173426.1831267-1-fvdl@google.com> Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 7bit On 3/20/26 6:34 PM, Frank van der Linden wrote: > Higher order GFP_ATOMIC allocations can be served through a > PCP list with ALLOC_HIGHATOMIC set. Such an allocation can > e.g. happen if a zone is between the low and min watermarks, > and get_page_from_freelist is retried after the alloc_flags > are relaxed. > > The call to reserve_highatomic_pageblock() after such a PCP > allocation will result in an increase every single time: > the page from the (unmovable) PCP list will never have > migrate type MIGRATE_HIGHATOMIC, since MIGRATE_HIGHATOMIC > pages do not appear on the unmovable PCP list. So a new > pageblock is converted to MIGRATE_HIGHATOMIC. > > Eventually that leads to the maximum of 1% of the zone being > used up by (often mostly free) MIGRATE_HIGHATOMIC pageblocks, > for no good reason. Since this space is not available for > normal allocations, this wastes memory and will push things > in to reclaim too soon. > > This was observed on a system that ran a test with bursts of > memory activity, pared with GFP_ATOMIC SLUB activity. These > would lead to a new slab being allocated with GFP_ATOMIC, > sometimes hitting the get_page_from_freelist retry path by > being below the low watermark. While the frequency of those > allocations was low, it kept adding up over time, and the > number of MIGRATE_ATOMIC pageblocks kept increasing. > > If a higher order atomic allocation can be served by > the unmovable PCP list, there is probably no need yet to > extend the reserves. So, move the check and possible extension > of the highatomic reserves to the buddy case only, and > do not refill the PCP list for ALLOC_HIGHATOMIC if it's > empty. This way, the PCP list is tried for ALLOC_HIGHATOMIC > for a fast atomic allocation. But it will immediately fall > back to rmqueue_buddy() if it's empty. In rmqueue_buddy(), > the MIGRATE_HIGHATOMIC buddy lists are tried first (as before), > and the reserves are extended only if that fails. > > With this change, the test was stable. Highatomic reserves > were built up, but to a normal level. No highatomic failures > were seen. > > This is similar to the patch proposed in [1] by Zhiguo Jiang, > but re-arranged a bit. > > Signed-off-by: Zhiguo Jiang > Signed-off-by: Frank van der Linden > Link: https://lore.kernel.org/all/20231122013925.1507-1-justinjiang@vivo.com/ [1] > Fixes: 44042b4498728 ("mm/page_alloc: allow high-order pages to be stored on the per-cpu lists") Makes sense to me and looks ok. Thanks. Reviewed-by: Vlastimil Babka (SUSE) > --- > mm/page_alloc.c | 30 +++++++++++++++++++++++------- > 1 file changed, 23 insertions(+), 7 deletions(-) > > diff --git a/mm/page_alloc.c b/mm/page_alloc.c > index 2d4b6f1a554e..57e17a15dae5 100644 > --- a/mm/page_alloc.c > +++ b/mm/page_alloc.c > @@ -243,6 +243,8 @@ unsigned int pageblock_order __read_mostly; > > static void __free_pages_ok(struct page *page, unsigned int order, > fpi_t fpi_flags); > +static void reserve_highatomic_pageblock(struct page *page, int order, > + struct zone *zone); > > /* > * results with 256, 32 in the lowmem_reserve sysctl: > @@ -3275,6 +3277,13 @@ struct page *rmqueue_buddy(struct zone *preferred_zone, struct zone *zone, > spin_unlock_irqrestore(&zone->lock, flags); > } while (check_new_pages(page, order)); > > + /* > + * If this is a high-order atomic allocation then check > + * if the pageblock should be reserved for the future > + */ > + if (unlikely(alloc_flags & ALLOC_HIGHATOMIC)) > + reserve_highatomic_pageblock(page, order, zone); > + > __count_zid_vm_events(PGALLOC, page_zonenum(page), 1 << order); > zone_statistics(preferred_zone, zone, 1); > > @@ -3346,6 +3355,20 @@ struct page *__rmqueue_pcplist(struct zone *zone, unsigned int order, > int batch = nr_pcp_alloc(pcp, zone, order); > int alloced; > > + /* > + * Don't refill the list for a higher order atomic > + * allocation under memory pressure, as this would > + * not build up any HIGHATOMIC reserves, which > + * might be needed soon. > + * > + * Instead, direct it towards the reserves by > + * returning NULL, which will make the caller fall > + * back to rmqueue_buddy. This will try to use the > + * reserves first and grow them if needed. > + */ > + if (alloc_flags & ALLOC_HIGHATOMIC) > + return NULL; > + > alloced = rmqueue_bulk(zone, order, > batch, list, > migratetype, alloc_flags); > @@ -3961,13 +3984,6 @@ get_page_from_freelist(gfp_t gfp_mask, unsigned int order, int alloc_flags, > if (page) { > prep_new_page(page, order, gfp_mask, alloc_flags); > > - /* > - * If this is a high-order atomic allocation then check > - * if the pageblock should be reserved for the future > - */ > - if (unlikely(alloc_flags & ALLOC_HIGHATOMIC)) > - reserve_highatomic_pageblock(page, order, zone); > - > return page; > } else { > if (cond_accept_memory(zone, order, alloc_flags))