All of lore.kernel.org
 help / color / mirror / Atom feed
From: "Zi Yan" <ziy@nvidia.com>
To: "Ferran Duarri" <ferran.duarri@me.com>,
	"Andrew Morton" <akpm@linux-foundation.org>,
	"David Hildenbrand" <david@kernel.org>,
	"Lorenzo Stoakes" <ljs@kernel.org>,
	"Johannes Weiner" <hannes@cmpxchg.org>
Cc: "Baolin Wang" <baolin.wang@linux.alibaba.com>,
	"Ryan Roberts" <ryan.roberts@arm.com>,
	"Barry Song" <baohua@kernel.org>, <linux-mm@kvack.org>,
	<linux-kernel@vger.kernel.org>
Subject: Re: [PATCH] mm: thp: default defrag mode to defer+madvise
Date: Thu, 20 Aug 2026 16:15:30 -0400	[thread overview]
Message-ID: <DKU1PF6DM6NN.YOF6NN3OWQUF@nvidia.com> (raw)
In-Reply-To: <20260820190825.221308-1-ferran.duarri@me.com>

On Thu Aug 20, 2026 at 3:08 PM EDT, Ferran Duarri wrote:
> transparent_hugepage_flags unconditionally sets
> TRANSPARENT_HUGEPAGE_DEFRAG_REQ_MADV_FLAG, so the built-in defrag mode
> is always "madvise" regardless of which of CONFIG_TRANSPARENT_HUGEPAGE_
> ALWAYS or _MADVISE was selected. Those Kconfig options choose whether
> THP is applied by default; they do not express a preference about how
> hard the allocator should work to produce a huge page, yet the defrag
> default is pinned as if they did.
>
> In "madvise" mode a fault on a MADV_HUGEPAGE region performs direct
> compaction in the fault path. Paired with transparent_hugepage=madvise
> that is a bounded cost, since only regions that asked for it are
> affected. Paired with transparent_hugepage=always every anonymous fault
> becomes eligible, and under memory pressure the faulting thread can
> stall in compaction.
>
> "defer+madvise" keeps the same set of regions eligible for huge pages
> and the same allocation attempt, but on failure it wakes kswapd and
> khugepaged to compact in the background instead of doing it inline. The
> fault proceeds with small pages and the region is collapsed later. This
> is the configuration long recommended to users running THP=always with
> large anonymous working sets, and it is a strictly weaker stall
> guarantee to make the default.
>
> This changes only the compiled-in default;
> /sys/kernel/mm/transparent_hugepage/defrag continues to accept every
> existing mode, including the current "madvise" behaviour.
>
> No stall measurement is offered with this patch. On the machine that
> prompted it the fault path never reaches direct compaction at all:
> thp_fault_alloc is 60682 against thp_fault_fallback 0, and compact_stall
> is 0, because memory has stayed abundant enough that no huge-page
> allocation has had to fall back. That makes the box unable to testify
> either way, and a number gathered under those conditions would measure
> nothing. The argument above is a correctness one about which knob the
> Kconfig choice is entitled to set, and it should be judged on that.
>
> Signed-off-by: Ferran Duarri <ferran.duarri@me.com>
> ---
>  mm/huge_memory.c | 2 +-
>  1 file changed, 1 insertion(+), 1 deletion(-)
>
> diff --git a/mm/huge_memory.c b/mm/huge_memory.c
> index b118bcd392cb..749f669cca56 100644
> --- a/mm/huge_memory.c
> +++ b/mm/huge_memory.c
> @@ -63,7 +63,7 @@ unsigned long transparent_hugepage_flags __read_mostly =
>  #ifdef CONFIG_TRANSPARENT_HUGEPAGE_MADVISE
>  	(1<<TRANSPARENT_HUGEPAGE_REQ_MADV_FLAG)|
>  #endif
> -	(1<<TRANSPARENT_HUGEPAGE_DEFRAG_REQ_MADV_FLAG)|
> +	(1<<TRANSPARENT_HUGEPAGE_DEFRAG_KSWAPD_OR_MADV_FLAG)|
>  	(1<<TRANSPARENT_HUGEPAGE_DEFRAG_KHUGEPAGED_FLAG)|
>  	(1<<TRANSPARENT_HUGEPAGE_USE_ZERO_PAGE_FLAG);
>  

+Johannes, since he said Meta heavily rely khugepaged to make good use
of THP, he might give more insight here.

The patch adds additional kswapd/kcompactd work when a system fails to
allocate THP for non madvised THP-eligible memory regions. If the
system's CPU utilization is high, kswapd/kcompactd can take CPU
resrouces from others, right? Performance numbers are needed to justify
this change.

-- 
Best Regards,
Yan, Zi


      reply	other threads:[~2026-08-20 20:15 UTC|newest]

Thread overview: 2+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-20 19:08 [PATCH] mm: thp: default defrag mode to defer+madvise Ferran Duarri
2026-08-20 20:15 ` Zi Yan [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=DKU1PF6DM6NN.YOF6NN3OWQUF@nvidia.com \
    --to=ziy@nvidia.com \
    --cc=akpm@linux-foundation.org \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=david@kernel.org \
    --cc=ferran.duarri@me.com \
    --cc=hannes@cmpxchg.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=ryan.roberts@arm.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.