* [PATCH] mm/khugepaged: cap min_free_kbytes recommendation at 1 GiB
@ 2026-07-28 20:20 Nimrod Oren
2026-07-29 0:17 ` Andrew Morton
0 siblings, 1 reply; 2+ messages in thread
From: Nimrod Oren @ 2026-07-28 20:20 UTC (permalink / raw)
To: Andrew Morton, David Hildenbrand, Lorenzo Stoakes
Cc: Zi Yan, Baolin Wang, Liam R. Howlett, Nico Pache, Ryan Roberts,
Dev Jain, Barry Song, Lance Yang, Usama Arif, linux-mm,
linux-kernel, Nimrod Oren
When THP is enabled, set_recommended_min_free_kbytes() may raise
min_free_kbytes using a heuristic that scales with pageblock_nr_pages.
With MIGRATE_PCPTYPES equal to 3, the formula accounts for 11
pageblocks for every eligible populated zone before capping the result
at 5% of low memory.
This is reasonable when a pageblock is 2 MiB, as on common 4 KiB page
configurations, but scales poorly with larger base page sizes. With
the default arm64 pageblock sizes, the contribution per eligible zone
before the 5% cap is:
4 KiB pages: 2 MiB pageblock, 22 MiB per zone
16 KiB pages: 32 MiB pageblock, 352 MiB per zone
64 KiB pages: 512 MiB pageblock, 5.5 GiB per zone
Consequently, min_free_kbytes can reach excessive and unwanted levels.
Add an absolute 1 GiB cap to the recommendation, in addition to the
existing percentage cap. This bounds the automatic recommendation to a
sane value on systems with large pageblocks while preserving existing
behavior for typical systems with 2 MiB pageblocks.
Link: https://lore.kernel.org/r/20260716173504.760369-1-noren@nvidia.com/
Suggested-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
Signed-off-by: Nimrod Oren <noren@nvidia.com>
---
v1:
* Cap the final recommendation at 1 GiB as suggested by Lorenzo.
RFC v1:
https://lore.kernel.org/r/20260716173504.760369-1-noren@nvidia.com/
---
mm/khugepaged.c | 7 ++++++-
1 file changed, 6 insertions(+), 1 deletion(-)
diff --git a/mm/khugepaged.c b/mm/khugepaged.c
index 27e8f3077e80..d9caf10e05b5 100644
--- a/mm/khugepaged.c
+++ b/mm/khugepaged.c
@@ -23,6 +23,7 @@
#include <linux/ksm.h>
#include <linux/pgalloc.h>
#include <linux/backing-dev.h>
+#include <linux/sizes.h>
#include <asm/tlb.h>
#include "internal.h"
@@ -3096,9 +3097,13 @@ void set_recommended_min_free_kbytes(void)
recommended_min += pageblock_nr_pages * nr_zones *
MIGRATE_PCPTYPES * MIGRATE_PCPTYPES;
- /* don't ever allow to reserve more than 5% of the lowmem */
+ /*
+ * Don't ever allow to reserve more than 5% of lowmem or 1 GiB,
+ * whichever is smaller.
+ */
recommended_min = min(recommended_min,
(unsigned long) nr_free_buffer_pages() / 20);
+ recommended_min = min(recommended_min, SZ_1G / PAGE_SIZE);
recommended_min <<= (PAGE_SHIFT-10);
if (recommended_min > min_free_kbytes) {
--
2.45.0
^ permalink raw reply related [flat|nested] 2+ messages in thread
* Re: [PATCH] mm/khugepaged: cap min_free_kbytes recommendation at 1 GiB
2026-07-28 20:20 [PATCH] mm/khugepaged: cap min_free_kbytes recommendation at 1 GiB Nimrod Oren
@ 2026-07-29 0:17 ` Andrew Morton
0 siblings, 0 replies; 2+ messages in thread
From: Andrew Morton @ 2026-07-29 0:17 UTC (permalink / raw)
To: Nimrod Oren
Cc: David Hildenbrand, Lorenzo Stoakes, Zi Yan, Baolin Wang,
Liam R. Howlett, Nico Pache, Ryan Roberts, Dev Jain, Barry Song,
Lance Yang, Usama Arif, linux-mm, linux-kernel, Vlastimil Babka
On Tue, 28 Jul 2026 23:20:14 +0300 Nimrod Oren <noren@nvidia.com> wrote:
> When THP is enabled, set_recommended_min_free_kbytes() may raise
> min_free_kbytes using a heuristic that scales with pageblock_nr_pages.
> With MIGRATE_PCPTYPES equal to 3, the formula accounts for 11
> pageblocks for every eligible populated zone before capping the result
> at 5% of low memory.
>
> This is reasonable when a pageblock is 2 MiB, as on common 4 KiB page
> configurations, but scales poorly with larger base page sizes. With
> the default arm64 pageblock sizes, the contribution per eligible zone
> before the 5% cap is:
>
> 4 KiB pages: 2 MiB pageblock, 22 MiB per zone
> 16 KiB pages: 32 MiB pageblock, 352 MiB per zone
> 64 KiB pages: 512 MiB pageblock, 5.5 GiB per zone
>
> Consequently, min_free_kbytes can reach excessive and unwanted levels.
>
> Add an absolute 1 GiB cap to the recommendation, in addition to the
> existing percentage cap. This bounds the automatic recommendation to a
> sane value on systems with large pageblocks while preserving existing
> behavior for typical systems with 2 MiB pageblocks.
Another hard-coded number isn't pretty :(
> mm/khugepaged.c | 7 ++++++-
Now, why is a page-allocator function sitting in khugepaged.c?
Perhaps to avoid a #ifdef CONFIG_TRANSPARENT_HUGEPAGE.
Or perhaps so that page-allocator maintainers don't get cc'ed on
page-allocator patches ;)
>
> diff --git a/mm/khugepaged.c b/mm/khugepaged.c
> index 27e8f3077e80..d9caf10e05b5 100644
> --- a/mm/khugepaged.c
> +++ b/mm/khugepaged.c
> @@ -23,6 +23,7 @@
> #include <linux/ksm.h>
> #include <linux/pgalloc.h>
> #include <linux/backing-dev.h>
> +#include <linux/sizes.h>
Why this?
> #include <asm/tlb.h>
> #include "internal.h"
> @@ -3096,9 +3097,13 @@ void set_recommended_min_free_kbytes(void)
> recommended_min += pageblock_nr_pages * nr_zones *
> MIGRATE_PCPTYPES * MIGRATE_PCPTYPES;
>
> - /* don't ever allow to reserve more than 5% of the lowmem */
> + /*
> + * Don't ever allow to reserve more than 5% of lowmem or 1 GiB,
> + * whichever is smaller.
> + */
> recommended_min = min(recommended_min,
> (unsigned long) nr_free_buffer_pages() / 20);
> + recommended_min = min(recommended_min, SZ_1G / PAGE_SIZE);
> recommended_min <<= (PAGE_SHIFT-10);
>
> if (recommended_min > min_free_kbytes) {
Is a min_free_kbytes documentation update needed?
Documentation/admin-guide/sysctl/vm.rst.
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-07-29 0:17 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-28 20:20 [PATCH] mm/khugepaged: cap min_free_kbytes recommendation at 1 GiB Nimrod Oren
2026-07-29 0:17 ` Andrew Morton
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.