From: Usama Arif <usama.arif@linux.dev>
To: Sourav Panda <souravpanda@google.com>,
Anshuman Khandual <anshuman.khandual@arm.com>
Cc: muchun.song@linux.dev, osalvador@suse.de,
akpm@linux-foundation.org, david@kernel.org, surenb@google.com,
fvdl@google.com, gthelen@google.com, linux-mm@kvack.org,
linux-kernel@vger.kernel.org
Subject: Re: [PATCH v4] mm/hugetlb_cma: Fix null nodemask dereference in hugetlb_cma_alloc_frozen_folio
Date: Thu, 6 Aug 2026 12:07:36 +0100 [thread overview]
Message-ID: <007cdb8b-01e3-4d73-9ac9-1a52ab5798dd@linux.dev> (raw)
In-Reply-To: <CANruzcT9PEXYRu1t_ge9ZNE8hLBpt6rWP4CwsBUkoZQTtpPKdA@mail.gmail.com>
On 06/08/2026 06:33, Sourav Panda wrote:
> On Wed, Aug 5, 2026 at 7:47 PM Anshuman Khandual
> <anshuman.khandual@arm.com> wrote:
>>
>> On Wed, Aug 05, 2026 at 01:24:42PM -0700, Usama Arif wrote:
>>> On Sun, 26 Jul 2026 07:29:34 +0000 Sourav Panda <souravpanda@google.com> wrote:
>>>
>>>> alloc_buddy_hugetlb_folio_with_mpol() can pass a NULL nodemask to
>>>> alloc_fresh_hugetlb_folio() as a fallback to allocate from all
>>>> nodes. If order is gigantic, alloc_fresh_hugetlb_folio() propagates
>>>> the NULL nodemask down to hugetlb_cma_alloc_frozen_folio() via
>>>> alloc_gigantic_frozen_folio().
>>>>
>>>> hugetlb_cma_alloc_frozen_folio() blindly dereferences the nodemask in
>>>> node_isset(nid, *nodemask) and for_each_node_mask(node, *nodemask),
>>>> leading to a null pointer dereference kernel panic.
>>>>
>>>> Fix this by checking if nodemask is NULL in
>>>> hugetlb_cma_alloc_frozen_folio() and defaulting it to
>>>> node_states[N_MEMORY]. This allows hugetlb_cma allocations to fall
>>>> back to any node with memory, keeping behavior consistent with
>>>> alloc_contig_frozen_pages() and alloc_buddy_frozen_folio().
>>>>
>>>> >From a userspace perspective, this bug allows an unprivileged user to
>>>> crash the kernel (trigger a panic) by requesting a gigantic hugepage
>>>> allocation with MPOL_PREFERRED_MANY on a system where CMA is only
>>>> configured on a subset of NUMA nodes.
>>>>
>>>> This can be reproduced by booting a VM with two NUMA nodes, restricting
>>>> CMA to Node 1 (e.g., hugetlb_cma=1:1G default_hugepagesz=1G
>>>> hugepagesz=1G hugepages=0), and running a program that allocates a
>>>> 1GB hugepage area without reserving, restricts allocation to Node 0
>>>> using mbind() with MPOL_PREFERRED_MANY, and triggers a page fault:
>>>>
>>>> void *ptr = mmap(NULL, 1UL << 30, PROT_READ | PROT_WRITE,
>>>> MAP_PRIVATE | MAP_ANONYMOUS | MAP_HUGETLB |
>>>> MAP_HUGE_1GB | MAP_NORESERVE, -1, 0);
>>>> unsigned long nodemask = 1; /* Node 0 */
>>>> mbind(ptr, 1UL << 30, MPOL_PREFERRED_MANY, &nodemask,
>>>> sizeof(nodemask) * 8, 0);
>>>> memset(ptr, 0, 1UL << 30); /* Trigger fault */
>>>>
>>>> This results in a NULL pointer dereference:
>>>>
>>>> BUG: kernel NULL pointer dereference, address: 0000000000000000
>>>> #PF: supervisor read access in kernel mode
>>>> #PF: error_code(0x0000) - not-present page
>>>> Oops: Oops: 0000 [#1] SMP NOPTI
>>>> RIP: 0010:hugetlb_cma_alloc_frozen_folio+0x75/0x120
>>>> Call Trace:
>>>> <TASK>
>>>> only_alloc_fresh_hugetlb_folio.isra.0+0x2c/0x160
>>>> alloc_surplus_hugetlb_folio+0x6d/0x100
>>>> alloc_hugetlb_folio+0x3c5/0x660
>>>> hugetlb_no_page+0x3d9/0x650
>>>>
>>>> Fixes: eb02f14c4a2b ("mm/hugetlb: allow overcommitting gigantic hugepages")
>>>> Cc: stable@vger.kernel.org
>>>> Signed-off-by: Sourav Panda <souravpanda@google.com>
>>>> ---
>>>> Changes in v4:
>>>> - Reverted the alloc_fresh_hugetlb_folio() cpuset snapshot approach from v3.
>>>> As Muchun Song pointed out, snapshotting cpuset_current_mems_allowed does
>>>> not prevent false-positive allocation failures without complex retry loops,
>>>> and alloc_contig_frozen_pages() / alloc_buddy_frozen_folio() already handle
>>>> NULL nodemasks safely internally.
>>>> - Handled NULL nodemask directly inside hugetlb_cma_alloc_frozen_folio()
>>>> by defaulting nodemask to node_states[N_MEMORY] (Option 2), keeping
>>>> HugeTLB allocators clean and consistent.
>>>> - v3: https://lore.kernel.org/linux-mm/20260705175119.440599-1-souravpanda@google.com/
>>>> - v2: https://lore.kernel.org/linux-mm/20260704174930.2885785-1-souravpanda@google.com/
>>>> - v1: https://lore.kernel.org/linux-mm/20260702215713.627941-1-souravpanda@google.com/
>>>>
>>>> mm/hugetlb_cma.c | 5 ++++-
>>>> 1 file changed, 4 insertions(+), 1 deletion(-)
>>>>
>>>> diff --git a/mm/hugetlb_cma.c b/mm/hugetlb_cma.c
>>>> index 39344d6c78d8..5744de0ceeb7 100644
>>>> --- a/mm/hugetlb_cma.c
>>>> +++ b/mm/hugetlb_cma.c
>>>> @@ -34,7 +34,10 @@ struct folio *hugetlb_cma_alloc_frozen_folio(int order, gfp_t gfp_mask,
>>>> if (!hugetlb_cma_size)
>>>> return NULL;
>>>>
>>>> - if (hugetlb_cma[nid])
>>>> + if (!nodemask)
>>>> + nodemask = &node_states[N_MEMORY];
>>>
>>> Hi Sourav,
>>>
>>> hmm what if cpuset.mems only allows allocation from node 0, and hugetlb CMA
>>> is only available on node 1. This will no now allow allocating a gigantic page
>>> on node 1 when its explicitly not allowed?
>>
>> That's fair point but should not the user be also responsible in provding a right
>> nodemask containing CMA memory if it prefers avoiding node_states[N_MEMORY] based
>> fallback mechanism in kernel ?
>>
I don't think the task can be made responsible for providing a CMA-aware fallback mask
here. Userspace supplies the MPOL_PREFERRED_MANY preference, but the kernel itself
changes the second attempt to a NULL nodemask. That permits fallback outside the
preferred policy nodes; it does not permit fallback outside the task's hardwall cpuset.
>
> We have precedence for three options, which makes selecting one difficult :)
>
> 1) nodemask = &cpuset_current_mems_allowed is currently being used in:
> only_alloc_fresh_hugetlb_folio --> alloc_buddy_frozen_folio -->
> __alloc_frozen_pages_noprof --> prepare_alloc_pages.
Here I think using &cpuset_current_mems_allowed is part of the page allocator's wider cpuset handling.
> 2) Some places use cookies with cpuset_current_mems_allowed to prevent
> torn writes. But this adds complexity.
> An example would be dequeue_hugetlb_folio_nodemask()
> 3) node_states[N_MEMORY] is sprinkled all over the kernel (simplest).
>
I would prefer option 2. A NULL mempolicy mask should fall back to
cpuset_current_mems_allowed, not all N_MEMORY nodes.
>>>
>>> Would it not be better to check cpuset as well?
>>>
>>> Thanks,
>>> Usama
>>>
>>>> +
>>>> + if (hugetlb_cma[nid] && node_isset(nid, *nodemask))
>>>> page = cma_alloc_frozen_compound(hugetlb_cma[nid], order);
>>>>
>>>> if (!page && !(gfp_mask & __GFP_THISNODE)) {
>>>> --
>>>> 2.55.0
>>>>
prev parent reply other threads:[~2026-08-06 11:07 UTC|newest]
Thread overview: 11+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-26 7:29 [PATCH v4] mm/hugetlb_cma: Fix null nodemask dereference in hugetlb_cma_alloc_frozen_folio Sourav Panda
2026-07-28 0:57 ` Andrew Morton
2026-08-03 6:37 ` Sourav Panda
2026-08-03 9:48 ` Anshuman Khandual
2026-08-05 4:42 ` Sourav Panda
2026-08-06 2:29 ` Anshuman Khandual
2026-08-05 20:24 ` Usama Arif
2026-08-05 23:05 ` Sourav Panda
2026-08-06 2:47 ` Anshuman Khandual
2026-08-06 5:33 ` Sourav Panda
2026-08-06 11:07 ` Usama Arif [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=007cdb8b-01e3-4d73-9ac9-1a52ab5798dd@linux.dev \
--to=usama.arif@linux.dev \
--cc=akpm@linux-foundation.org \
--cc=anshuman.khandual@arm.com \
--cc=david@kernel.org \
--cc=fvdl@google.com \
--cc=gthelen@google.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=muchun.song@linux.dev \
--cc=osalvador@suse.de \
--cc=souravpanda@google.com \
--cc=surenb@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox