* + mm-memfd-fix-hugetlb-reservation-accounting-in-error-paths.patch added to mm-new branch
@ 2026-09-03 20:22 Andrew Morton
0 siblings, 0 replies; only message in thread
From: Andrew Morton @ 2026-09-03 20:22 UTC (permalink / raw)
To: mm-commits, vivek.kasireddy, stable, osalvador, muchun.song,
hughd, david, baolin.wang, lihongfu, akpm
The patch titled
Subject: mm/memfd: fix hugetlb reservation accounting in error paths
has been added to the -mm mm-new branch. Its filename is
mm-memfd-fix-hugetlb-reservation-accounting-in-error-paths.patch
This patch will shortly appear at
https://git.kernel.org/pub/scm/linux/kernel/git/akpm/25-new.git/tree/patches/mm-memfd-fix-hugetlb-reservation-accounting-in-error-paths.patch
This patch will later appear in the mm-new branch at
git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm
Note, mm-new is a provisional staging ground for work-in-progress
patches, and acceptance into mm-new is a notification for others take
notice and to finish up reviews. Please do not hesitate to respond to
review feedback and post updated versions to replace or incrementally
fixup patches in mm-new.
The mm-new branch of mm.git is not included in linux-next
If a few days of testing in mm-new is successful, the patch will me moved
into mm.git's mm-unstable branch, which is included in linux-next
Before you just go and hit "reply", please:
a) Consider who else should be cc'ed
b) Prefer to cc a suitable mailing list as well
c) Ideally: find the original patch on the mailing list and do a
reply-to-all to that, adding suitable additional cc's
*** Remember to use Documentation/process/submit-checklist.rst when testing your code ***
The -mm tree is included into linux-next via various
branches at git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm
and is updated there most days
------------------------------------------------------
From: Hongfu Li <lihongfu@kylinos.cn>
Subject: mm/memfd: fix hugetlb reservation accounting in error paths
Date: Thu, 3 Sep 2026 11:01:34 +0800
If hugetlb_add_to_page_cache() in memfd_alloc_folio() fails with -EEXIST,
a concurrent fault has already instantiated the folio in the page cache,
and the reservation now belongs to that folio. Calling
hugetlb_unreserve_pages() in that case incorrectly removes the region
backing the cached folio. A later truncate or inode eviction then passes
a negative (chg - freed) into hugepage_subpool_put_pages(), corrupting
subpool and resv_huge_pages accounting.
Over time, these corrupted counters would leak huge page reservations.
Applications using hugetlb memfds would eventually find themselves unable
to allocate huge pages, receiving unexpected ENOMEM errors even though
system memory and pool capacities appeared free and healthy.
The corrupted accounting caused hugepage_subpool_put_pages() to receive a
negative value during a later file truncation or inode eviction.
While this typically manifests as kernel logs (WARN traces or badness
flags regarding subpool page counts), it could cause misbehaved resource
tracking that impacts subsequent system operations, unmounts, or process
teardowns interacting with that hugetlb file descriptor.
So hold the hugetlb fault mutex from hugetlb_reserve_pages() until the
error-path unreserve completes to make the reserve, allocate and
instantiate steps atomic against concurrent faults. With the mutex held
from the start, a concurrent fault can no longer consume the reservation
between reserve and allocate/instantiate. If a fault completed before the
mutex was taken, it has already added the region for that index, so
hugetlb_reserve_pages() returns 0 and the error path leaves the region in
place.
Link: https://lore.kernel.org/20260903030134.7407-1-hongfu.li@linux.dev
Fixes: 717cf9357325 ("mm/memfd: reserve hugetlb folios before allocation")
Signed-off-by: Hongfu Li <lihongfu@kylinos.cn>
Cc: Baolin Wang <baolin.wang@linux.alibaba.com>
Cc: David Hildenbrand <david@kernel.org>
Cc: Hugh Dickins <hughd@google.com>
Cc: Muchun Song <muchun.song@linux.dev>
Cc: Oscar Salvador <osalvador@suse.de>
Cc: Vivek Kasireddy <vivek.kasireddy@intel.com>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
---
mm/memfd.c | 31 ++++++++++++++++---------------
1 file changed, 16 insertions(+), 15 deletions(-)
--- a/mm/memfd.c~mm-memfd-fix-hugetlb-reservation-accounting-in-error-paths
+++ a/mm/memfd.c
@@ -82,22 +82,31 @@ struct folio *memfd_alloc_folio(struct f
struct hstate *h = hstate_file(memfd);
int err = -ENOMEM;
long nr_resv;
+ u32 hash;
gfp_mask = htlb_alloc_mask(h);
gfp_mask &= ~(__GFP_HIGHMEM | __GFP_MOVABLE);
idx >>= huge_page_order(h);
+ /*
+ * Serialize hugepage allocation and instantiation to prevent
+ * races with concurrent allocations, as required by all other
+ * callers of hugetlb_add_to_page_cache().
+ */
+ hash = hugetlb_fault_mutex_hash(memfd->f_mapping, idx);
+ mutex_lock(&hugetlb_fault_mutex_table[hash]);
+
nr_resv = hugetlb_reserve_pages(inode, idx, idx + 1, NULL, EMPTY_VMA_FLAGS);
- if (nr_resv < 0)
- return ERR_PTR(nr_resv);
+ if (nr_resv < 0) {
+ err = nr_resv;
+ goto out_unlock;
+ }
folio = alloc_hugetlb_folio_reserve(h,
numa_node_id(),
NULL,
gfp_mask);
if (folio) {
- u32 hash;
-
/*
* Zero the folio to prevent information leaks to userspace.
* Use folio_zero_user() which is optimized for huge/gigantic
@@ -112,20 +121,9 @@ struct folio *memfd_alloc_folio(struct f
*/
__folio_mark_uptodate(folio);
- /*
- * Serialize hugepage allocation and instantiation to prevent
- * races with concurrent allocations, as required by all other
- * callers of hugetlb_add_to_page_cache().
- */
- hash = hugetlb_fault_mutex_hash(memfd->f_mapping, idx);
- mutex_lock(&hugetlb_fault_mutex_table[hash]);
-
err = hugetlb_add_to_page_cache(folio,
memfd->f_mapping,
idx);
-
- mutex_unlock(&hugetlb_fault_mutex_table[hash]);
-
if (err) {
folio_put(folio);
goto err_unresv;
@@ -133,11 +131,14 @@ struct folio *memfd_alloc_folio(struct f
hugetlb_set_folio_subpool(folio, subpool_inode(inode));
folio_unlock(folio);
+ mutex_unlock(&hugetlb_fault_mutex_table[hash]);
return folio;
}
err_unresv:
if (nr_resv > 0)
hugetlb_unreserve_pages(inode, idx, idx + 1, 0);
+out_unlock:
+ mutex_unlock(&hugetlb_fault_mutex_table[hash]);
return ERR_PTR(err);
}
#endif
_
Patches currently in -mm which might be from lihongfu@kylinos.cn are
mm-use-a-folio-in-the-softleaf_is_device_private-path.patch
mm-hugetlb-fix-resv_huge_pages-double-decrement-in-memfd-error-path.patch
mm-memcontrol-remove-unused-memcg-parameter-in-calculate_high_delay.patch
selftests-mm-khugepaged-consolidate-error-exits-via-kselftest-helpers.patch
mm-page_owner-preserve-original-free_pid-free_tgid-during-folio-migration.patch
mm-memfd-fix-hugetlb-reservation-accounting-in-error-paths.patch
^ permalink raw reply [flat|nested] only message in thread
only message in thread, other threads:[~2026-09-03 20:22 UTC | newest]
Thread overview: (only message) (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-03 20:22 + mm-memfd-fix-hugetlb-reservation-accounting-in-error-paths.patch added to mm-new branch Andrew Morton
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.