From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pj1-f70.google.com (mail-pj1-f70.google.com [209.85.216.70]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DF62610F0 for ; Sun, 9 Aug 2026 04:32:52 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.70 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786249974; cv=none; b=fm/57xMCdIdClbLqW9/cPrci8wUJq6FAkT6YkiLJ8Y42+r87GPkmFkKqs+XaD91jWtzUXs3GvD0CwgW/I0BxLd7RxEwAxbgWVfd9fr8fWI2fkvm50TWowrrMkfW5xz4SL61kkBl0OWKnFCr9EOauMAAhIIJOzBsmMXQAwiQzGdU= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786249974; c=relaxed/simple; bh=o0szSX1F0T5quii2vwewPKVDiY4KxycI1vKY28GTroo=; h=Date:Mime-Version:Message-ID:Subject:From:To:Cc:Content-Type; b=mKjKyEUU0n67gnzfL7Cedsf1fYT5nVACblCg4+FG45e+MlY8a9W/PKa2EFTjhMkAQxW3V1QutUW5wEyJ6UH3ygn2TGUReNTnLWOp7d2iyvXOeZLJDXjQ6LjbhEFd+tEcgiHZMO347r2Ezl1LYRZ0Vl0Sxl6HnNyyC76lFFuoKcg= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--souravpanda.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=cD922Zb+; arc=none smtp.client-ip=209.85.216.70 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--souravpanda.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="cD922Zb+" Received: by mail-pj1-f70.google.com with SMTP id 98e67ed59e1d1-38ecc48b3c2so1495454a91.1 for ; Sat, 08 Aug 2026 21:32:52 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1786249972; x=1786854772; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:mime-version:date:from :to:cc:subject:date:message-id:reply-to:content-type; bh=SFcfVMlaONyojBIsJ/DA4x9Y5kX85WT5BYOtPk8UXQs=; b=cD922Zb+xkKrL+NfdQd7DUidkk2tH/GfBKqmJBN1pzIVHAo2btXLeczOe+ZD+5tmH4 fuyiRYXCBEDVlm4q0fqH7R8qjj9BYUK+FJazy5FRr/PU7PJNO2bRgpFk0DYX2vk6FqQO ZN5gGQGH/X1q/lmpuND4I5s5HA79yEuFOexETIk/b01nsfOs6mhggdAAuAoEEsbXPrng 4sAlEUiKhRFgyL2vxn0uKhR4QRCAjt9h2LU79DcrQ3JQXnILmJvSti3xA+8uoREhp7GL N3OD9eSKzljMrSMjsJoWrMbDGkTbYP1LCMTMbGrSfk9r1ltt3d55y3wG36AiBPlzy2Uf KQfQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786249972; x=1786854772; h=content-type:cc:to:from:subject:message-id:mime-version:date :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=SFcfVMlaONyojBIsJ/DA4x9Y5kX85WT5BYOtPk8UXQs=; b=TIu7rcw0PHQjiobfeMdLCNZCMllnGhv2C5rAyGJacNGixZmiKgQHa/6+yW4ajzA//7 /O74+p9Lptdvc+VJyc9ZNUfNEXlpJBrjRg796ypt6SaqVrfMJl2O/z2jrL+fvv2EU38t t7x8pkwCqKT5+ql4MC/cE7Cq6AEiOeRe6czcpojagWZqJzzI0Lx16e1RMp+p3Q40nrFI BKdDswcpcOGeRqvnBFP3YzqCAzpFMs4sOLy/cLPh7X9gEigr7z3io43eejhYPLRIlDui qEhStQKtJUZkjkUQMGcnSHz6aLlNk9XdBUikfbxMWZB2vf7DXVRD+w8gKlwIxQxZ08sj JvBg== X-Forwarded-Encrypted: i=1; AHgh+RoXZqR8ydGS5whkTTHEZheHTe/U0BKZOz0qDeF/3fkmWdNtutIY/Xi3G/vvSQ9GjOKfCZd0K8GuBeF6JfU=@vger.kernel.org X-Gm-Message-State: AOJu0Yy1rD3LA4rBm/sJ3ciunmydTAh3rzBHsXKpiE6CHiDvrc22Jpnp CjlCYfv1waTWzlwQBevkcrs74H7T8mUcI7BqzO7+bb5poNjqjurfdMW5plfDPo0+7/rDIJ01BrD 39ymFyfiSgAfgF9BPp+ucmoRkwg== X-Received: from pjwo24.prod.google.com ([2002:a17:90a:d258:b0:384:f6e1:ff87]) (user=souravpanda job=prod-delivery.src-stubby-dispatcher) by 2002:a17:90b:2249:b0:36b:bec8:94c5 with SMTP id 98e67ed59e1d1-3928242ffc4mr9674877a91.10.1786249972002; Sat, 08 Aug 2026 21:32:52 -0700 (PDT) Date: Sun, 9 Aug 2026 04:32:50 +0000 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 X-Mailer: git-send-email 2.55.0.679.g6767b8d81c-goog Message-ID: <20260809043250.2917406-1-souravpanda@google.com> Subject: [PATCH v5] mm/hugetlb_cma: Fix null nodemask dereference in hugetlb_cma_alloc_frozen_folio From: Sourav Panda To: muchun.song@linux.dev, osalvador@suse.de, akpm@linux-foundation.org Cc: usama.arif@linux.dev, shakeel.butt@linux.dev, wangkefeng.wang@huawei.com, anshuman.khandual@arm.com, david@kernel.org, surenb@google.com, fvdl@google.com, gthelen@google.com, hannes@cmpxchg.org, riel@surriel.com, sj@kernel.org, vbabka@suse.cz, mhocko@suse.com, bjackman@google.com, zi.yan@sent.com, souravpanda@google.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org Content-Type: text/plain; charset="UTF-8" alloc_buddy_hugetlb_folio_with_mpol() can pass a NULL nodemask to alloc_fresh_hugetlb_folio() as a fallback to allocate from all nodes. If order is gigantic, alloc_fresh_hugetlb_folio() propagates the NULL nodemask down to hugetlb_cma_alloc_frozen_folio() via alloc_gigantic_frozen_folio(). Additionally, hugetlb_cma_alloc_frozen_folio() previously attempted allocation on hugetlb_cma[nid] without verifying if nid is included in the caller's nodemask. Adding a node_isset(nid, *nodemask) check ensures the initial preferred node allocation honors the memory policy / nodemask. However, hugetlb_cma_alloc_frozen_folio() dereferences the nodemask in node_isset(nid, *nodemask) and for_each_node_mask(node, *nodemask), leading to a null pointer dereference kernel panic when nodemask is NULL. Fix this by checking if nodemask is NULL in hugetlb_cma_alloc_frozen_folio() and defaulting it to cpuset_current_mems_allowed (safely read using a seqcount retry loop). This ensures that the initial node check and fallback loop safely honor the task's cpuset without violating cpuset constraints or causing NULL pointer dereferences. >From a userspace perspective, this bug allows an unprivileged user to crash the kernel (trigger a panic) by requesting a gigantic hugepage allocation with MPOL_PREFERRED_MANY on a system where CMA is only configured on a subset of NUMA nodes. This can be reproduced by booting a VM with two NUMA nodes, restricting CMA to Node 1 (e.g., hugetlb_cma=1:1G default_hugepagesz=1G hugepagesz=1G hugepages=0), and running a program that allocates a 1GB hugepage area without reserving, restricts allocation to Node 0 using mbind() with MPOL_PREFERRED_MANY, and triggers a page fault: void *ptr = mmap(NULL, 1UL << 30, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS | MAP_HUGETLB | MAP_HUGE_1GB | MAP_NORESERVE, -1, 0); unsigned long nodemask = 1; /* Node 0 */ mbind(ptr, 1UL << 30, MPOL_PREFERRED_MANY, &nodemask, sizeof(nodemask) * 8, 0); memset(ptr, 0, 1UL << 30); /* Trigger fault */ This results in a NULL pointer dereference: BUG: kernel NULL pointer dereference, address: 0000000000000000 #PF: supervisor read access in kernel mode #PF: error_code(0x0000) - not-present page Oops: Oops: 0000 [#1] SMP NOPTI RIP: 0010:hugetlb_cma_alloc_frozen_folio+0x75/0x120 Call Trace: only_alloc_fresh_hugetlb_folio.isra.0+0x2c/0x160 alloc_surplus_hugetlb_folio+0x6d/0x100 alloc_hugetlb_folio+0x3c5/0x660 hugetlb_no_page+0x3d9/0x650 Fixes: eb02f14c4a2b ("mm/hugetlb: allow overcommitting gigantic hugepages") Cc: stable@vger.kernel.org Signed-off-by: Sourav Panda --- Changes in v5: - Replaced defaulting nodemask to &node_states[N_MEMORY] with safely reading cpuset_current_mems_allowed using a seqcount retry loop, ensuring fallback allocations comply with task hardwall cpusets as suggested by Usama Arif. - v4: https://lore.kernel.org/linux-mm/20260726072935.3513996-1-souravpanda@google.com/ - v3: https://lore.kernel.org/linux-mm/20260705175119.440599-1-souravpanda@google.com/ - v2: https://lore.kernel.org/linux-mm/20260704174930.2885785-1-souravpanda@google.com/ - v1: https://lore.kernel.org/linux-mm/20260702215713.627941-1-souravpanda@google.com/ mm/hugetlb_cma.c | 15 ++++++++++++++- 1 file changed, 14 insertions(+), 1 deletion(-) diff --git a/mm/hugetlb_cma.c b/mm/hugetlb_cma.c index 39344d6c78d8..3ae9347078e9 100644 --- a/mm/hugetlb_cma.c +++ b/mm/hugetlb_cma.c @@ -3,6 +3,7 @@ #include #include #include +#include #include #include @@ -30,11 +31,23 @@ struct folio *hugetlb_cma_alloc_frozen_folio(int order, gfp_t gfp_mask, int node; struct folio *folio; struct page *page = NULL; + nodemask_t local_node_mask; if (!hugetlb_cma_size) return NULL; - if (hugetlb_cma[nid]) + if (!nodemask) { + unsigned int cpuset_mems_cookie; + + do { + cpuset_mems_cookie = read_mems_allowed_begin(); + local_node_mask = cpuset_current_mems_allowed; + } while (read_mems_allowed_retry(cpuset_mems_cookie)); + + nodemask = &local_node_mask; + } + + if (hugetlb_cma[nid] && node_isset(nid, *nodemask)) page = cma_alloc_frozen_compound(hugetlb_cma[nid], order); if (!page && !(gfp_mask & __GFP_THISNODE)) { -- 2.55.0