From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A3DC8402BB3; Tue, 21 Jul 2026 02:53:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784602389; cv=none; b=dQXXzC1kfij/TgaBC0beXNH12z2NDUTfnr2cEYq4dcTPifx1njVj6ZOjkCYxxSNOjRqwM6I1B4i1v9CAPmQOMGHbCnmE9m+xtQVkesHFvejiueWo+z40ZddguG8ccVmSO60ujH1J8BRmODT1Bk3LnEgISWwD9j0nmurUIoSS0tk= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784602389; c=relaxed/simple; bh=kttlEIrRapL9va7fSljR3xrwSzOf2bb6snXuzZxqZ30=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=iMXzVgoVnb9MXWaGLBjf/mXvWD81s7NA63jg4L6F8mF000macsCL4V/86J53XhuJP693hzZanBcGlHOqvviAO6P+PHbI7MKxvqoXIsyWbV4fTqmevIWSr0Ub8oWxEKw8aNV2rBcbwLQIVnmZkPqI/KHIIYn5/FbGbOWX6Ae2OqU= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=ZOhggKih; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="ZOhggKih" Received: by smtp.kernel.org (Postfix) with ESMTPS id 881F8C2BCF7; Tue, 21 Jul 2026 02:53:09 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1784602389; bh=kttlEIrRapL9va7fSljR3xrwSzOf2bb6snXuzZxqZ30=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=ZOhggKih7xOk/MyAuJWaa130N+z1Q9x1F+XEfuajwBORDjzg1QKcUta3uHiqj+QLu vXPnHaoq/uE8J/O1kfEKI1trcY97vfO1cjJWnRJ5bVIInHaZKKPS7o/onuZMvZHEl9 LAwiKs1u5gGTP7z4Ryki0KHNtoZjge6VdqF30jE+gdXuNxEbmj1DVXW5r0C0RYOudc PD9DBKg7uiMcqKXwVe7X6R2Ql3/vjd9aXtc55UfZJYuLdnIEiRN7YRFPzHzeWx+1SX tNqIrk8XWy6dodqh62rbSfIA/XoCoKUGkTpjdK2cYvI0HZilpYvioLpoO72CBstkUT wYy9P+MUM4i7w== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 74A0BC44524; Tue, 21 Jul 2026 02:53:09 +0000 (UTC) From: Ackerley Tng via B4 Relay Date: Mon, 20 Jul 2026 17:25:14 -0700 Subject: [PATCH v3 11/13] WIP: Reproducer for subpool usage leak Precedence: bulk X-Mailing-List: cgroups@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: 7bit Message-Id: <20260720-hugetlb-alloc-failure-fixes-v3-11-7d2a169aa9ee@google.com> References: <20260720-hugetlb-alloc-failure-fixes-v3-0-7d2a169aa9ee@google.com> In-Reply-To: <20260720-hugetlb-alloc-failure-fixes-v3-0-7d2a169aa9ee@google.com> To: Muchun Song , Oscar Salvador , David Hildenbrand , Joshua Hahn , Shakeel Butt , Nhat Pham , Andrew Morton , Peter Xu , Wupeng Ma , fvdl@google.com, rientjes@google.com, jthoughton@google.com, Mike Kravetz , Johannes Weiner , Michal Hocko , Roman Gushchin Cc: vannapurve@google.com, erdemaktas@google.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, Ackerley Tng X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784602387; l=4607; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=inX7G11YaaxbwEU3Aj5PB0crlLMuo9T1V2I+32/3MYg=; b=2snhi0Fk67JboBtC3TcQ7c/BnoASl8jauKmvWjHDI4B5pTaAhEL4c/GXuRyTy283zI1c7EcHa Em1FnEeYordAxOBRStKhF9bHrxmCLw+F4rYPMhqRBYEn6LJ0/c0I9V8 X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Endpoint-Received: by B4 Relay for ackerleytng@google.com/20260225 with auth_id=649 X-Original-From: Ackerley Tng Reply-To: ackerleytng@google.com From: Ackerley Tng (This reproducer was hacked up and not meant to be merged.) The kernel leaks subpool usage and the subpool structure itself if a HugeTLBfs mount specifying size (which sets max_hpages on the subpool) is created. subpool_leak_max_size.sh reproduces this with the following steps: 1. Create mount, specifying size=2M (1 page). This sets max_hpages = 1 on the subpool, but does not reserve any pages. 2. Set nr_hugepages = 0 and nr_overcommit_hugepages = 0 so that physical allocations will fail. 3. Run fallocate -l 2M on a file in the mount. + This calls hugetlbfs_fallocate, which attempts to allocate a page by calling alloc_hugetlb_folio. + alloc_hugetlb_folio calls hugepage_subpool_get_pages to track the allocation against the subpool limit. This increments used_hpages to 1. + Physical allocation fails because nr_hugepages is 0. + Before patch (Buggy): + The error path in alloc_hugetlb_folio sees gbl_chg is 1 (indicating we tried to allocate a global page) and incorrectly skips calling hugepage_subpool_put_pages. + fallocate fails and returns to userspace, but the subpool used_hpages counter remains leaked at 1. + After patch: + The error path always calls hugepage_subpool_put_pages if map_chg is true, restoring used_hpages to 0. 4. Unmount the filesystem. + During unmount, the kernel calls unlock_or_release_subpool to clean up the subpool. + It checks if the subpool is free using subpool_is_free, which returns whether used_hpages is 0. + Before patch (Buggy): + Since used_hpages leaked and is 1, subpool_is_free returns false. + The kernel skips freeing the subpool structure, leaking the hugepage_subpool structure in kernel memory. + After patch: + Since used_hpages is 0, subpool_is_free returns true, and the subpool structure is correctly freed. Signed-off-by: Ackerley Tng --- subpool_leak_max_size.sh | 71 ++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 71 insertions(+) diff --git a/subpool_leak_max_size.sh b/subpool_leak_max_size.sh new file mode 100755 index 0000000000000..bfafa1ba074ea --- /dev/null +++ b/subpool_leak_max_size.sh @@ -0,0 +1,71 @@ +#!/bin/bash + +if [ "$EUID" -ne 0 ]; then + echo "Please run as root" + exit 1 +fi + +MNT_PATH="/tmp/mnt_hugetlb" +FILE_PATH="$MNT_PATH/test_file" + +# Save original values +orig_nr=$(cat /sys/kernel/mm/hugepages/hugepages-2048kB/nr_hugepages) +orig_overcommit=$(cat /sys/kernel/mm/hugepages/hugepages-2048kB/nr_overcommit_hugepages) + +cleanup() { + echo "Cleaning up..." + rm -f "$FILE_PATH" + umount "$MNT_PATH" 2>/dev/null + rmdir "$MNT_PATH" 2>/dev/null + echo "$orig_nr" > /sys/kernel/mm/hugepages/hugepages-2048kB/nr_hugepages + echo "$orig_overcommit" > /sys/kernel/mm/hugepages/hugepages-2048kB/nr_overcommit_hugepages + echo "Cleanup done." +} +trap cleanup EXIT + +# 1. Mount hugetlbfs with size=2M (1 page) +mkdir -p "$MNT_PATH" +if ! mount -t hugetlbfs -o size=2M none "$MNT_PATH"; then + echo "Failed to mount hugetlbfs" + exit 1 +fi + +# 2. Set nr_hugepages to 0, overcommit to 0 +echo 0 > /sys/kernel/mm/hugepages/hugepages-2048kB/nr_hugepages +echo 0 > /sys/kernel/mm/hugepages/hugepages-2048kB/nr_overcommit_hugepages + +# Check subpool usage before running +read total free < <(stat -f -c "%b %f" "$MNT_PATH") +used_before=$((total - free)) +echo "Before test - Subpool total blocks: $total" +echo "Before test - Subpool free blocks: $free" +echo "Before test - Subpool used blocks: $used_before" +if [ "$used_before" -ne 0 ]; then + echo "ERROR: Subpool is not clean before test starts!" + exit 1 +fi + +# Run fallocate (expecting failure) +echo "Running fallocate (expecting failure)..." +if fallocate -l 2M "$FILE_PATH" 2>/dev/null; then + echo "ERROR: fallocate succeeded but should have failed (nr_hugepages is 0)" + exit 1 +fi + +# Check subpool usage via statfs +# %b: Total blocks +# %f: Free blocks +read total free < <(stat -f -c "%b %f" "$MNT_PATH") +used=$((total - free)) + +echo "Subpool total blocks: $total" +echo "Subpool free blocks: $free" +echo "Subpool used blocks (leaked if > 0): $used" + +if [ "$used" -gt 0 ]; then + echo "RESULT: LEAK DETECTED (FAIL)" + exit 1 +else + echo "RESULT: NO LEAK (PASS)" + exit 0 +fi -- 2.55.0.229.g6434b31f56-goog