All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH v4] mm/page_alloc: avoid direct compaction for costly __GFP_NORETRY allocations
@ 2026-09-04 11:56 Salvatore Dipietro
  2026-09-04 14:11 ` Vlastimil Babka (SUSE)
                   ` (3 more replies)
  0 siblings, 4 replies; 34+ messages in thread
From: Salvatore Dipietro @ 2026-09-04 11:56 UTC (permalink / raw)
  To: linux-kernel, hch
  Cc: abuehaze, akpm, alisaidi, blakgeof, brauner, dgc,
	dipietro.salvatore, dipiets, djwong, hannes, jackmanb,
	linux-fsdevel, linux-mm, linux-xfs, mhocko, ritesh.list, rvvandan,
	stable, surenb, vbabka, willy, ziy, Vlastimil Babka,
	David Hildenbrand, Christoph Hellwig, Brendan Jackman

Commit 5d8edfb900d5 ("iomap: Copy larger chunks from userspace")
introduced high-order folio allocations in the iomap buffered write
path.  When memory is fragmented, each failed costly-order allocation
enters __alloc_pages_slowpath() which runs direct compaction and
drain_all_pages(), causing a 0.38x throughput drop on PostgreSQL
pgbench (simple-update) with 1024 clients on a 96-vCPU arm64 system.

The root issue is that direct compaction is too expensive for hot
allocation paths that have fallbacks to smaller allocations.
__filemap_get_folio_mpol() already marks higher-order allocations with
__GFP_NORETRY | __GFP_NOWARN, signalling that the caller can handle
failure.  However, the page allocator still attempts full direct
compaction for costly orders with __GFP_NORETRY, which is unnecessarily
aggressive when the caller will simply retry at a lower order.

For costly-order allocations with __GFP_NORETRY, clear
__GFP_DIRECT_RECLAIM at the very start of the slowpath, before
can_direct_reclaim, can_compact and the nofail checks are evaluated.
This makes the entire slowpath treat the request as non-blocking: no
direct reclaim, no direct compaction and no drain_all_pages() IPI
across every CPU. kswapd (and in turn kcompactd) is still woken further
down for background defragmentation, so compaction keeps working for
long-term system health while being removed from the latency-critical
direct allocation path.

Allocations that also request __GFP_THISNODE are exempted.  That flag
pairing identifies the local-node-first THP attempt issued by
alloc_pages_mpol() (mempolicy.c), which relies on direct compaction to
form transparent huge pages.

Test environment:
  Hardware:  AWS EC2 m8g.24xlarge (96 vCPU, arm64)
             12x 1TB IO2 32000 IOPS RAID0 XFS
  OS:        AL2023
  Kernel:    v7.3-rc1
  Database:  PostgreSQL 18.4
  Workload:  pgbench simple-update, 1024 clients, 96 threads, 1200s

Results (average of 3 runs, TPS):

  Config                   Avg TPS      % vs Baseline
  baseline (no patch)       59,408       -
  With this patch          155,409      +161.6%

Link: https://lore.kernel.org/all/20260403193535.9970-1-dipiets@amazon.it/T/#t [v1]
Link: https://lore.kernel.org/linux-mm/20260420161404.642-1-dipiets@amazon.it/T/#u [v2]
Link: https://lore.kernel.org/all/20260710143437.12379-1-dipiets@amazon.it/T/#u [v3]
Fixes: 5d8edfb900d5 ("iomap: Copy larger chunks from userspace")
Cc: stable@vger.kernel.org
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: Vlastimil Babka <vbabka@suse.cz>
Cc: David Hildenbrand <david@redhat.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Johannes Weiner <hannes@cmpxchg.org>
Cc: Matthew Wilcox <willy@infradead.org>
Cc: Christoph Hellwig <hch@lst.de>
Cc: Dave Chinner <dgc@kernel.org>
Cc: Ritesh Harjani <ritesh.list@gmail.com>
Cc: linux-mm@kvack.org
Cc: linux-fsdevel@vger.kernel.org
Cc: linux-xfs@vger.kernel.org
Signed-off-by: Salvatore Dipietro <dipiets@amazon.it>
---
v4: Clear __GFP_DIRECT_RECLAIM early in the slowpath and exempt
    __GFP_THISNODE so THP attempt keeps using direct compaction
v3: Move to mm/page_alloc.c, wake kcompactd instead of avoiding it
v2: Move from fs/iomap/buffered-io.c to mm/filemap.c
v1: Avoid compaction in iomap folio allocation

 mm/page_alloc.c | 18 +++++++++++++++---
 1 file changed, 15 insertions(+), 3 deletions(-)

diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index 12fac9084c48..542c2ec31061 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -4784,10 +4784,10 @@ static inline struct page *
 __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order,
 						struct alloc_context *ac)
 {
-	bool can_direct_reclaim = gfp_mask & __GFP_DIRECT_RECLAIM;
-	bool can_compact = can_direct_reclaim && gfp_compaction_allowed(gfp_mask);
-	bool nofail = gfp_mask & __GFP_NOFAIL;
 	const bool costly_order = order > PAGE_ALLOC_COSTLY_ORDER;
+	bool can_direct_reclaim;
+	bool can_compact;
+	bool nofail;
 	struct page *page = NULL;
 	unsigned int alloc_flags;
 	unsigned long did_some_progress;
@@ -4802,6 +4802,18 @@ __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order,
 	bool can_retry_reserves = true;
 	unsigned long alloc_start_time = jiffies;
 
+	/*
+	 * Costly __GFP_NORETRY callers have a cheap fallback, so don't stall
+	 * them in reclaim or compaction. __GFP_THISNODE callers are exempt.
+	 */
+	if (costly_order && (gfp_mask & __GFP_NORETRY) &&
+	    !(gfp_mask & __GFP_THISNODE))
+		gfp_mask &= ~__GFP_DIRECT_RECLAIM;
+
+	can_direct_reclaim = gfp_mask & __GFP_DIRECT_RECLAIM;
+	can_compact = can_direct_reclaim && gfp_compaction_allowed(gfp_mask);
+	nofail = gfp_mask & __GFP_NOFAIL;
+
 	if (unlikely(nofail)) {
 		/*
 		 * Also we don't support __GFP_NOFAIL without __GFP_DIRECT_RECLAIM,
-- 
2.50.1




AMAZON DEVELOPMENT CENTER ITALY SRL, viale Monte Grappa 3/5, 20124 Milano, Italia, Registro delle Imprese di Milano Monza Brianza Lodi REA n. 2504859, Capitale Sociale: 10.000 EUR i.v., Cod. Fisc. e P.IVA 10100050961, Societa con Socio Unico




^ permalink raw reply related	[flat|nested] 34+ messages in thread

end of thread, other threads:[~2026-09-24 14:34 UTC | newest]

Thread overview: 34+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-04 11:56 [PATCH v4] mm/page_alloc: avoid direct compaction for costly __GFP_NORETRY allocations Salvatore Dipietro
2026-09-04 14:11 ` Vlastimil Babka (SUSE)
2026-09-04 15:08   ` Zi Yan
2026-09-07  7:30     ` Vlastimil Babka (SUSE)
2026-09-09  2:33       ` Zi Yan
2026-09-09  8:51         ` Vlastimil Babka (SUSE)
2026-09-04 16:10 ` Johannes Weiner
2026-09-06  0:42 ` Andrew Morton
2026-09-06 23:05   ` Dave Chinner
2026-09-10 11:46   ` Salvatore Dipietro
2026-09-10 22:00     ` Andrew Morton
2026-09-11 14:30       ` Salvatore Dipietro
2026-09-11 15:59     ` Johannes Weiner
2026-09-16 11:24       ` Vlastimil Babka (SUSE)
2026-09-16 15:58         ` Johannes Weiner
2026-09-16 22:34           ` Andrew Morton
2026-09-16 22:35             ` Andrew Morton
2026-09-19  4:13               ` Matthew Wilcox
2026-09-21  9:35                 ` Vlastimil Babka (SUSE)
2026-09-23  9:25                   ` Salvatore Dipietro
2026-09-18  7:05           ` Vlastimil Babka (SUSE)
2026-09-18 21:28             ` Andrew Morton
2026-09-22  8:44               ` Vlastimil Babka (SUSE)
2026-09-21 14:37             ` Johannes Weiner
2026-09-21 14:38               ` [PATCH 1/2] mm: page_alloc: do not give all non-blocking requests reserve access Johannes Weiner
2026-09-21 14:54                 ` Matthew Wilcox
2026-09-21 15:58                   ` Johannes Weiner
2026-09-22 11:52                     ` Vlastimil Babka (SUSE)
2026-09-22 13:56                       ` Johannes Weiner
2026-09-24 14:34                         ` Vlastimil Babka (SUSE)
2026-09-21 14:39               ` [PATCH 2/2] mm: page_alloc: remove ALLOC_NON_BLOCK from ALLOC_RESERVES Johannes Weiner
2026-09-22 12:06                 ` Vlastimil Babka (SUSE)
2026-09-22 14:00                   ` Johannes Weiner
2026-09-07  5:54 ` [PATCH v4] mm/page_alloc: avoid direct compaction for costly __GFP_NORETRY allocations Christoph Hellwig

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.