From: Salvatore Dipietro <dipiets@amazon.it>
To: <vbabka@kernel.org>, <hannes@cmpxchg.org>, <willy@infradead.org>
Cc: <abuehaze@amazon.com>, <akpm@linux-foundation.org>,
<alisaidi@amazon.com>, <blakgeof@amazon.com>,
<rvvandan@amazon.com>, <brauner@kernel.org>, <dgc@kernel.org>,
<dipietro.salvatore@gmail.com>, <dipiets@amazon.it>,
<djwong@kernel.org>, <hch@infradead.org>, <jackmanb@google.com>,
<linux-fsdevel@vger.kernel.org>, <linux-kernel@vger.kernel.org>,
<linux-mm@kvack.org>, <linux-xfs@vger.kernel.org>,
<mhocko@suse.com>, <ritesh.list@gmail.com>,
<stable@vger.kernel.org>, <surenb@google.com>, <ziy@nvidia.com>
Subject: Re: [PATCH v3] mm/page_alloc: avoid direct compaction for costly __GFP_NORETRY allocations
Date: Thu, 27 Aug 2026 09:21:00 +0000 [thread overview]
Message-ID: <20260827092100.3988386-1-dipiets@amazon.it> (raw)
In-Reply-To: <7ddb704c-76a9-4071-8969-06fa1403eb07@kernel.org>
On Wed, Jul 22, 2026 at 02:35:00PM +0200, Vlastimil Babka (SUSE) wrote:
> AFAICS both your v3 and the implemented Johannes' suggestion effectively
> make all the compact_first logic dead code - at least for the intended THP
> use case based on costly_order. It will only continue to do something for
> the non-movable order>0 case. So that's a significant change which will not
> show up in your results, but might affect other workloads (I'd expect mainly
> by having fewer THPs).
Right - the change above modifies THP allocation requests as well. So instead
of applying it to every costly __GFP_NORETRY allocation, I restrict it to
callers that do not ask for __GFP_THISNODE. That leaves the local-node THP
attempt untouched: the same bit pair is already tested further down the function
to bail out of the compact_first block (page_alloc.c:4899), and alloc_pages_mpol()
sets both flags on the local-node-first attempt (mempolicy.c:2476-2478).
I tested this with my PostgreSQL workload on kernel v7.2-rc7 and results look good:
Baseline: 70,735 tps
This version: 144,670 tps (145,280 / 144,037 / 144,694) +104.5%
diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index ee902a468c2f..1e4b0c1e7296 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -4725,10 +4725,10 @@ static inline struct page *
__alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order,
struct alloc_context *ac)
{
- bool can_direct_reclaim = gfp_mask & __GFP_DIRECT_RECLAIM;
- bool can_compact = can_direct_reclaim && gfp_compaction_allowed(gfp_mask);
- bool nofail = gfp_mask & __GFP_NOFAIL;
const bool costly_order = order > PAGE_ALLOC_COSTLY_ORDER;
+ bool can_direct_reclaim;
+ bool can_compact;
+ bool nofail;
struct page *page = NULL;
unsigned int alloc_flags;
unsigned long did_some_progress;
@@ -4743,6 +4743,18 @@ __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order,
bool can_retry_reserves = true;
unsigned long alloc_start_time = jiffies;
+ /*
+ * Costly __GFP_NORETRY callers have a cheap fallback, so don't stall
+ * them in reclaim or compaction. __GFP_THISNODE callers are exempt.
+ */
+ if (costly_order && (gfp_mask & __GFP_NORETRY) &&
+ !(gfp_mask & __GFP_THISNODE))
+ gfp_mask &= ~__GFP_DIRECT_RECLAIM;
+
+ can_direct_reclaim = gfp_mask & __GFP_DIRECT_RECLAIM;
+ can_compact = can_direct_reclaim && gfp_compaction_allowed(gfp_mask);
+ nofail = gfp_mask & __GFP_NOFAIL;
+
if (unlikely(nofail)) {
/*
* Also we don't support __GFP_NOFAIL without __GFP_DIRECT_RECLAIM,
AMAZON DEVELOPMENT CENTER ITALY SRL, viale Monte Grappa 3/5, 20124 Milano, Italia, Registro delle Imprese di Milano Monza Brianza Lodi REA n. 2504859, Capitale Sociale: 10.000 EUR i.v., Cod. Fisc. e P.IVA 10100050961, Societa con Socio Unico
next prev parent reply other threads:[~2026-08-27 9:21 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-10 14:34 [PATCH v3] mm/page_alloc: avoid direct compaction for costly __GFP_NORETRY allocations Salvatore Dipietro
2026-07-10 15:22 ` Johannes Weiner
2026-07-10 18:03 ` Matthew Wilcox
2026-07-14 12:02 ` Salvatore Dipietro
2026-07-22 12:35 ` Vlastimil Babka (SUSE)
2026-08-27 9:21 ` Salvatore Dipietro [this message]
2026-07-22 12:17 ` Vlastimil Babka (SUSE)
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260827092100.3988386-1-dipiets@amazon.it \
--to=dipiets@amazon.it \
--cc=abuehaze@amazon.com \
--cc=akpm@linux-foundation.org \
--cc=alisaidi@amazon.com \
--cc=blakgeof@amazon.com \
--cc=brauner@kernel.org \
--cc=dgc@kernel.org \
--cc=dipietro.salvatore@gmail.com \
--cc=djwong@kernel.org \
--cc=hannes@cmpxchg.org \
--cc=hch@infradead.org \
--cc=jackmanb@google.com \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux-xfs@vger.kernel.org \
--cc=mhocko@suse.com \
--cc=ritesh.list@gmail.com \
--cc=rvvandan@amazon.com \
--cc=stable@vger.kernel.org \
--cc=surenb@google.com \
--cc=vbabka@kernel.org \
--cc=willy@infradead.org \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox