All of lore.kernel.org
 help / color / mirror / Atom feed
From: Ferran Duarri <ferran.duarri@me.com>
To: Andrew Morton <akpm@linux-foundation.org>,
	David Hildenbrand <david@kernel.org>,
	Lorenzo Stoakes <ljs@kernel.org>
Cc: Zi Yan <ziy@nvidia.com>,
	Baolin Wang <baolin.wang@linux.alibaba.com>,
	Ryan Roberts <ryan.roberts@arm.com>,
	Barry Song <baohua@kernel.org>,
	linux-mm@kvack.org, linux-kernel@vger.kernel.org,
	Ferran Duarri <ferran.duarri@me.com>
Subject: [PATCH] mm: thp: default defrag mode to defer+madvise
Date: Thu, 20 Aug 2026 21:08:23 +0200	[thread overview]
Message-ID: <20260820190825.221308-1-ferran.duarri@me.com> (raw)

transparent_hugepage_flags unconditionally sets
TRANSPARENT_HUGEPAGE_DEFRAG_REQ_MADV_FLAG, so the built-in defrag mode
is always "madvise" regardless of which of CONFIG_TRANSPARENT_HUGEPAGE_
ALWAYS or _MADVISE was selected. Those Kconfig options choose whether
THP is applied by default; they do not express a preference about how
hard the allocator should work to produce a huge page, yet the defrag
default is pinned as if they did.

In "madvise" mode a fault on a MADV_HUGEPAGE region performs direct
compaction in the fault path. Paired with transparent_hugepage=madvise
that is a bounded cost, since only regions that asked for it are
affected. Paired with transparent_hugepage=always every anonymous fault
becomes eligible, and under memory pressure the faulting thread can
stall in compaction.

"defer+madvise" keeps the same set of regions eligible for huge pages
and the same allocation attempt, but on failure it wakes kswapd and
khugepaged to compact in the background instead of doing it inline. The
fault proceeds with small pages and the region is collapsed later. This
is the configuration long recommended to users running THP=always with
large anonymous working sets, and it is a strictly weaker stall
guarantee to make the default.

This changes only the compiled-in default;
/sys/kernel/mm/transparent_hugepage/defrag continues to accept every
existing mode, including the current "madvise" behaviour.

No stall measurement is offered with this patch. On the machine that
prompted it the fault path never reaches direct compaction at all:
thp_fault_alloc is 60682 against thp_fault_fallback 0, and compact_stall
is 0, because memory has stayed abundant enough that no huge-page
allocation has had to fall back. That makes the box unable to testify
either way, and a number gathered under those conditions would measure
nothing. The argument above is a correctness one about which knob the
Kconfig choice is entitled to set, and it should be judged on that.

Signed-off-by: Ferran Duarri <ferran.duarri@me.com>
---
 mm/huge_memory.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/mm/huge_memory.c b/mm/huge_memory.c
index b118bcd392cb..749f669cca56 100644
--- a/mm/huge_memory.c
+++ b/mm/huge_memory.c
@@ -63,7 +63,7 @@ unsigned long transparent_hugepage_flags __read_mostly =
 #ifdef CONFIG_TRANSPARENT_HUGEPAGE_MADVISE
 	(1<<TRANSPARENT_HUGEPAGE_REQ_MADV_FLAG)|
 #endif
-	(1<<TRANSPARENT_HUGEPAGE_DEFRAG_REQ_MADV_FLAG)|
+	(1<<TRANSPARENT_HUGEPAGE_DEFRAG_KSWAPD_OR_MADV_FLAG)|
 	(1<<TRANSPARENT_HUGEPAGE_DEFRAG_KHUGEPAGED_FLAG)|
 	(1<<TRANSPARENT_HUGEPAGE_USE_ZERO_PAGE_FLAG);
 


             reply	other threads:[~2026-08-20 19:08 UTC|newest]

Thread overview: 8+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-20 19:08 Ferran Duarri [this message]
2026-08-20 20:15 ` [PATCH] mm: thp: default defrag mode to defer+madvise Zi Yan
2026-08-21 16:41   ` Lorenzo Stoakes (ARM)
2026-08-21 14:18 ` David Hildenbrand (Arm)
2026-08-21 17:10   ` Ferran Duarri
2026-08-21 16:53 ` Lorenzo Stoakes (ARM)
2026-08-21 17:26   ` Ferran Duarri
2026-08-21 17:47   ` Ferran

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260820190825.221308-1-ferran.duarri@me.com \
    --to=ferran.duarri@me.com \
    --cc=akpm@linux-foundation.org \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=david@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=ryan.roberts@arm.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.