Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: "Lorenzo Stoakes (ARM)" <ljs@kernel.org>
To: Nimrod Oren <noren@nvidia.com>
Cc: Andrew Morton <akpm@linux-foundation.org>,
	 David Hildenbrand <david@kernel.org>,
	Jonathan Corbet <corbet@lwn.net>,
	 Vlastimil Babka <vbabka@kernel.org>,
	"Liam R. Howlett" <liam@infradead.org>,
	 Mike Rapoport <rppt@kernel.org>,
	Suren Baghdasaryan <surenb@google.com>,
	 Michal Hocko <mhocko@suse.com>,
	Shuah Khan <skhan@linuxfoundation.org>,
	 Randy Dunlap <rdunlap@infradead.org>, Zi Yan <ziy@nvidia.com>,
	 Baolin Wang <baolin.wang@linux.alibaba.com>,
	Nico Pache <nico.pache@linux.dev>,
	 Ryan Roberts <ryan.roberts@arm.com>, Dev Jain <dev.jain@arm.com>,
	Barry Song <baohua@kernel.org>,
	 Lance Yang <lance.yang@linux.dev>,
	Usama Arif <usama.arif@linux.dev>,
	 Kiryl Shutsemau <kas@kernel.org>,
	Nirmoy Das <nirmoyd@nvidia.com>,
	 Dragos Tatulea <dtatulea@nvidia.com>,
	linux-mm@kvack.org, linux-doc@vger.kernel.org,
	 linux-kernel@vger.kernel.org
Subject: Re: [PATCH v2] mm/khugepaged: cap min_free_kbytes recommendation at 1 GiB
Date: Mon, 31 Aug 2026 10:00:07 +0100	[thread overview]
Message-ID: <apVCAeICeuxaL5Ed@gremlin> (raw)
In-Reply-To: <20260831075635.2244437-1-noren@nvidia.com>

On Mon, Aug 31, 2026 at 10:56:35AM +0300, Nimrod Oren wrote:
> When THP is enabled, set_recommended_min_free_kbytes() may raise
> min_free_kbytes using a heuristic that scales with pageblock_nr_pages.
> With MIGRATE_PCPTYPES equal to 3, the formula accounts for 11
> pageblocks for every eligible populated zone before capping the result
> at 5% of low memory.
>
> This is reasonable when a pageblock is 2 MiB, as on common 4 KiB page
> configurations, but scales poorly with larger base page sizes. With
> the default arm64 pageblock sizes, the contribution per eligible zone
> before the 5% cap is:
>
>   4 KiB pages:   2 MiB pageblock,   22 MiB per zone
>  16 KiB pages:  32 MiB pageblock,  352 MiB per zone
>  64 KiB pages: 512 MiB pageblock,  5.5 GiB per zone
>
> Consequently, min_free_kbytes can reach excessive and unwanted levels.
>
> Add an absolute 1 GiB cap to the recommendation, in addition to the
> existing percentage cap. This bounds the automatic recommendation to a
> sane value on systems with large pageblocks while preserving existing
> behavior for typical systems with 2 MiB pageblocks.
>
> Link: https://lore.kernel.org/r/20260716173504.760369-1-noren@nvidia.com/
> Suggested-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
> Signed-off-by: Nimrod Oren <noren@nvidia.com>

I do think this is a practical way of preventing runaway increase of reserves.

In the long run we need to rework the pageblock model I think, but this is
a reasonable way forward in the short-medium term.

Obviously I defer to David and the community on this if they think this is
horrible, then let's not.

But I do worry that we'll carry on with unworkably huge reserves on 64 KiB
page size while rejecting every way out until pageblocks are _entirely_
reworked to account for this (probably the long term solution).

So I think this the least-worst approach of those proposed so far for now
:) Therefore:

Acked-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>

> ---
> v2:
> * Use a named constant for the cap.
> * Drop explicit linux/sizes.h include.
> * Update min_free_kbytes documentation.
>
> v1:
> * Cap the final recommendation at 1 GiB as suggested by Lorenzo.
> https://lore.kernel.org/r/20260728202014.2517142-1-noren@nvidia.com/
>
> RFC v1:
> https://lore.kernel.org/r/20260716173504.760369-1-noren@nvidia.com/
> ---
>  Documentation/admin-guide/sysctl/vm.rst | 8 ++++++++
>  mm/khugepaged.c                         | 7 ++++++-
>  2 files changed, 14 insertions(+), 1 deletion(-)
>
> diff --git a/Documentation/admin-guide/sysctl/vm.rst b/Documentation/admin-guide/sysctl/vm.rst
> index 5b318d17aa4b..c0c925e6f504 100644
> --- a/Documentation/admin-guide/sysctl/vm.rst
> +++ b/Documentation/admin-guide/sysctl/vm.rst
> @@ -539,6 +539,14 @@ watermark[WMARK_MIN] value for each lowmem zone in the system.
>  Each lowmem zone gets a number of reserved free pages based
>  proportionally on its size.
>
> +When Transparent Hugepage (THP) support is enabled through global or
> +per-size controls, the kernel may raise this value automatically to
> +help keep pageblocks free and reduce fragmentation for THP allocations.
> +The automatic recommendation cannot exceed 5% of low memory or 1 GiB,
> +whichever is lower.  This limit applies only to the automatic
> +recommendation and does not constrain a higher value written by
> +userspace.
> +
>  Some minimal amount of memory is needed to satisfy PF_MEMALLOC
>  allocations; if you set this to lower than 1024KB, your system will
>  become subtly broken, and prone to deadlock under high loads.
> diff --git a/mm/khugepaged.c b/mm/khugepaged.c
> index f49a6710933b..a8fc061d2488 100644
> --- a/mm/khugepaged.c
> +++ b/mm/khugepaged.c
> @@ -101,6 +101,7 @@ static DEFINE_READ_MOSTLY_HASHTABLE(mm_slots_hash, MM_SLOTS_HASH_BITS);
>
>  static struct kmem_cache *mm_slot_cache __ro_after_init;
>
> +#define THP_MIN_FREE_MAX_PAGES		(SZ_1G / PAGE_SIZE)
>  #define KHUGEPAGED_MIN_MTHP_ORDER	2
>
>  struct collapse_control {
> @@ -3120,9 +3121,13 @@ void set_recommended_min_free_kbytes(void)
>  	recommended_min += pageblock_nr_pages * nr_zones *
>  			   MIGRATE_PCPTYPES * MIGRATE_PCPTYPES;
>
> -	/* don't ever allow to reserve more than 5% of the lowmem */
> +	/*
> +	 * Don't allow the THP recommendation to exceed 5% of lowmem or an
> +	 * absolute cap, whichever is smaller.
> +	 */
>  	recommended_min = min(recommended_min,
>  			      (unsigned long) nr_free_buffer_pages() / 20);
> +	recommended_min = min(recommended_min, THP_MIN_FREE_MAX_PAGES);
>  	recommended_min <<= (PAGE_SHIFT-10);
>
>  	if (recommended_min > min_free_kbytes) {
> --
> 2.45.0
>

--
Cheers, Lorenzo


  parent reply	other threads:[~2026-08-31  9:00 UTC|newest]

Thread overview: 11+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-31  7:56 [PATCH v2] mm/khugepaged: cap min_free_kbytes recommendation at 1 GiB Nimrod Oren
2026-08-31  8:47 ` Kiryl Shutsemau
2026-08-31  8:57   ` Lorenzo Stoakes (ARM)
2026-08-31  9:00 ` Lorenzo Stoakes (ARM) [this message]
2026-08-31  9:00 ` Lance Yang
2026-08-31 15:01   ` Zi Yan
2026-08-31  9:20 ` Michal Hocko
2026-08-31  9:34   ` Lorenzo Stoakes (ARM)
2026-08-31 15:00     ` Zi Yan
2026-08-31 15:59       ` Lorenzo Stoakes (ARM)
2026-08-31 17:08         ` Nimrod Oren

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=apVCAeICeuxaL5Ed@gremlin \
    --to=ljs@kernel.org \
    --cc=akpm@linux-foundation.org \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=corbet@lwn.net \
    --cc=david@kernel.org \
    --cc=dev.jain@arm.com \
    --cc=dtatulea@nvidia.com \
    --cc=kas@kernel.org \
    --cc=lance.yang@linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=mhocko@suse.com \
    --cc=nico.pache@linux.dev \
    --cc=nirmoyd@nvidia.com \
    --cc=noren@nvidia.com \
    --cc=rdunlap@infradead.org \
    --cc=rppt@kernel.org \
    --cc=ryan.roberts@arm.com \
    --cc=skhan@linuxfoundation.org \
    --cc=surenb@google.com \
    --cc=usama.arif@linux.dev \
    --cc=vbabka@kernel.org \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox