From: Yafang Shao <laoar.shao@gmail.com>
To: akpm@linux-foundation.org, david@redhat.com, ziy@nvidia.com,
baolin.wang@linux.alibaba.com, lorenzo.stoakes@oracle.com,
Liam.Howlett@oracle.com, npache@redhat.com, ryan.roberts@arm.com,
dev.jain@arm.com, hannes@cmpxchg.org, usamaarif642@gmail.com,
gutierrez.asier@huawei-partners.com, willy@infradead.org,
ast@kernel.org, daniel@iogearbox.net, andrii@kernel.org,
ameryhung@gmail.com, rientjes@google.com
Cc: bpf@vger.kernel.org, linux-mm@kvack.org,
Yafang Shao <laoar.shao@gmail.com>
Subject: [RFC PATCH v5 mm-new 0/5] mm, bpf: BPF based THP order selection
Date: Mon, 18 Aug 2025 13:55:05 +0800 [thread overview]
Message-ID: <20250818055510.968-1-laoar.shao@gmail.com> (raw)
Background
----------
Our production servers consistently configure THP to "never" due to
historical incidents caused by its behavior. Key issues include:
- Increased Memory Consumption
THP significantly raises overall memory usage, reducing available memory
for workloads.
- Latency Spikes
Random latency spikes occur due to frequent memory compaction triggered
by THP.
- Lack of Fine-Grained Control
THP tuning is globally configured, making it unsuitable for containerized
environments. When multiple workloads share a host, enabling THP without
per-workload control leads to unpredictable behavior.
Due to these issues, administrators avoid switching to madvise or always
modes—unless per-workload THP control is implemented.
To address this, we propose BPF-based THP policy for flexible adjustment.
Additionally, as David mentioned [0], this mechanism can also serve as a
policy prototyping tool (test policies via BPF before upstreaming them).
Proposed Solution
-----------------
As suggested by David [0], we introduce a new BPF interface:
/**
* @get_suggested_order: Get the suggested THP orders for allocation
* @mm: mm_struct associated with the THP allocation
* @vma__nullable: vm_area_struct associated with the THP allocation (may be NULL)
* When NULL, the decision should be based on @mm (i.e., when
* triggered from an mm-scope hook rather than a VMA-specific
* context).
* Must belong to @mm (guaranteed by the caller).
* @vma_flags: use these vm_flags instead of @vma->vm_flags (0 if @vma is NULL)
* @tva_flags: TVA flags for current @vma (-1 if @vma is NULL)
* @orders: Bitmask of requested THP orders for this allocation
* - PMD-mapped allocation if PMD_ORDER is set
* - mTHP allocation otherwise
*
* Rerurn: Bitmask of suggested THP orders for allocation. The highest
* suggested order will not exceed the highest requested order
* in @orders.
*/
int (*get_suggested_order)(struct mm_struct *mm, struct vm_area_struct *vma__nullable,
u64 vma_flags, enum tva_type tva_flags, int orders) __rcu;
This interface:
- Supports both use cases (per-workload tuning + policy prototyping).
- Can be extended with BPF helpers (e.g., for memory pressure awareness).
This is an experimental feature. To use it, you must enable
CONFIG_EXPERIMENTAL_BPF_ORDER_SELECTION.
Warning:
- The interface may change
- Behavior may differ in future kernel versions
- We might remove it in the future
A simple test case is included in Patch #4.
Future work:
- Extend it to File THP
Changes:
RFC v4->v5:
- Add support for vma (David)
- Add mTHP support in khugepaged (Zi)
- Use bitmask of all allowed orders instead (Zi)
- Retrieve the page size and PMD order rather than hardcoding them (Zi)
RFC v3->v4: https://lwn.net/Articles/1031829/
- Use a new interface get_suggested_order() (David)
- Mark it as experimental (David, Lorenzo)
- Code improvement in THP (Usama)
- Code improvement in BPF struct ops (Amery)
RFC v2->v3: https://lwn.net/Articles/1024545/
- Finer-graind tuning based on madvise or always mode (David, Lorenzo)
- Use BPF to write more advanced policies logic (David, Lorenzo)
RFC v1->v2: https://lwn.net/Articles/1021783/
The main changes are as follows,
- Use struct_ops instead of fmod_ret (Alexei)
- Introduce a new THP mode (Johannes)
- Introduce new helpers for BPF hook (Zi)
- Refine the commit log
RFC v1: https://lwn.net/Articles/1019290/
Yafang Shao (5):
mm: thp: add support for BPF based THP order selection
mm: thp: add a new kfunc bpf_mm_get_mem_cgroup()
mm: thp: add a new kfunc bpf_mm_get_task()
bpf: mark vma->vm_mm as trusted
selftest/bpf: add selftest for BPF based THP order seletection
include/linux/huge_mm.h | 15 +
include/linux/khugepaged.h | 12 +-
kernel/bpf/verifier.c | 5 +
mm/Kconfig | 12 +
mm/Makefile | 1 +
mm/bpf_thp.c | 269 ++++++++++++++++++
mm/huge_memory.c | 10 +
mm/khugepaged.c | 26 +-
mm/memory.c | 18 +-
tools/testing/selftests/bpf/config | 3 +
.../selftests/bpf/prog_tests/thp_adjust.c | 224 +++++++++++++++
.../selftests/bpf/progs/test_thp_adjust.c | 76 +++++
.../bpf/progs/test_thp_adjust_failure.c | 25 ++
13 files changed, 689 insertions(+), 7 deletions(-)
create mode 100644 mm/bpf_thp.c
create mode 100644 tools/testing/selftests/bpf/prog_tests/thp_adjust.c
create mode 100644 tools/testing/selftests/bpf/progs/test_thp_adjust.c
create mode 100644 tools/testing/selftests/bpf/progs/test_thp_adjust_failure.c
--
2.47.3
next reply other threads:[~2025-08-18 5:55 UTC|newest]
Thread overview: 17+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-08-18 5:55 Yafang Shao [this message]
2025-08-18 5:55 ` [RFC PATCH v5 mm-new 1/5] mm: thp: add support for BPF based THP order selection Yafang Shao
2025-08-18 13:17 ` Usama Arif
2025-08-19 3:08 ` Yafang Shao
2025-08-19 10:11 ` Usama Arif
2025-08-19 11:10 ` Gutierrez Asier
2025-08-19 11:43 ` Yafang Shao
2025-08-18 5:55 ` [RFC PATCH v5 mm-new 2/5] mm: thp: add a new kfunc bpf_mm_get_mem_cgroup() Yafang Shao
2025-08-18 5:55 ` [RFC PATCH v5 mm-new 3/5] mm: thp: add a new kfunc bpf_mm_get_task() Yafang Shao
2025-08-18 5:55 ` [RFC PATCH v5 mm-new 4/5] bpf: mark vma->vm_mm as trusted Yafang Shao
2025-08-18 5:55 ` [RFC PATCH v5 mm-new 5/5] selftest/bpf: add selftest for BPF based THP order seletection Yafang Shao
2025-08-18 14:00 ` Usama Arif
2025-08-19 3:09 ` Yafang Shao
2025-08-18 14:35 ` [RFC PATCH v5 mm-new 0/5] mm, bpf: BPF based THP order selection Usama Arif
2025-08-19 2:41 ` Yafang Shao
2025-08-19 10:44 ` Usama Arif
2025-08-19 11:33 ` Yafang Shao
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20250818055510.968-1-laoar.shao@gmail.com \
--to=laoar.shao@gmail.com \
--cc=Liam.Howlett@oracle.com \
--cc=akpm@linux-foundation.org \
--cc=ameryhung@gmail.com \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=bpf@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=david@redhat.com \
--cc=dev.jain@arm.com \
--cc=gutierrez.asier@huawei-partners.com \
--cc=hannes@cmpxchg.org \
--cc=linux-mm@kvack.org \
--cc=lorenzo.stoakes@oracle.com \
--cc=npache@redhat.com \
--cc=rientjes@google.com \
--cc=ryan.roberts@arm.com \
--cc=usamaarif642@gmail.com \
--cc=willy@infradead.org \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.