All of lore.kernel.org
 help / color / mirror / Atom feed
From: Usama Arif <usama.arif@linux.dev>
To: Sourav Panda <souravpanda@google.com>,
	Anshuman Khandual <anshuman.khandual@arm.com>
Cc: muchun.song@linux.dev, osalvador@suse.de,
	akpm@linux-foundation.org, david@kernel.org, surenb@google.com,
	fvdl@google.com, gthelen@google.com, linux-mm@kvack.org,
	linux-kernel@vger.kernel.org
Subject: Re: [PATCH v4] mm/hugetlb_cma: Fix null nodemask dereference in hugetlb_cma_alloc_frozen_folio
Date: Thu, 6 Aug 2026 12:07:36 +0100	[thread overview]
Message-ID: <007cdb8b-01e3-4d73-9ac9-1a52ab5798dd@linux.dev> (raw)
In-Reply-To: <CANruzcT9PEXYRu1t_ge9ZNE8hLBpt6rWP4CwsBUkoZQTtpPKdA@mail.gmail.com>



On 06/08/2026 06:33, Sourav Panda wrote:
> On Wed, Aug 5, 2026 at 7:47 PM Anshuman Khandual
> <anshuman.khandual@arm.com> wrote:
>>
>> On Wed, Aug 05, 2026 at 01:24:42PM -0700, Usama Arif wrote:
>>> On Sun, 26 Jul 2026 07:29:34 +0000 Sourav Panda <souravpanda@google.com> wrote:
>>>
>>>> alloc_buddy_hugetlb_folio_with_mpol() can pass a NULL nodemask to
>>>> alloc_fresh_hugetlb_folio() as a fallback to allocate from all
>>>> nodes. If order is gigantic, alloc_fresh_hugetlb_folio() propagates
>>>> the NULL nodemask down to hugetlb_cma_alloc_frozen_folio() via
>>>> alloc_gigantic_frozen_folio().
>>>>
>>>> hugetlb_cma_alloc_frozen_folio() blindly dereferences the nodemask in
>>>> node_isset(nid, *nodemask) and for_each_node_mask(node, *nodemask),
>>>> leading to a null pointer dereference kernel panic.
>>>>
>>>> Fix this by checking if nodemask is NULL in
>>>> hugetlb_cma_alloc_frozen_folio() and defaulting it to
>>>> node_states[N_MEMORY]. This allows hugetlb_cma allocations to fall
>>>> back to any node with memory, keeping behavior consistent with
>>>> alloc_contig_frozen_pages() and alloc_buddy_frozen_folio().
>>>>
>>>> >From a userspace perspective, this bug allows an unprivileged user to
>>>> crash the kernel (trigger a panic) by requesting a gigantic hugepage
>>>> allocation with MPOL_PREFERRED_MANY on a system where CMA is only
>>>> configured on a subset of NUMA nodes.
>>>>
>>>> This can be reproduced by booting a VM with two NUMA nodes, restricting
>>>> CMA to Node 1 (e.g., hugetlb_cma=1:1G default_hugepagesz=1G
>>>> hugepagesz=1G hugepages=0), and running a program that allocates a
>>>> 1GB hugepage area without reserving, restricts allocation to Node 0
>>>> using mbind() with MPOL_PREFERRED_MANY, and triggers a page fault:
>>>>
>>>>   void *ptr = mmap(NULL, 1UL << 30, PROT_READ | PROT_WRITE,
>>>>                    MAP_PRIVATE | MAP_ANONYMOUS | MAP_HUGETLB |
>>>>                    MAP_HUGE_1GB | MAP_NORESERVE, -1, 0);
>>>>   unsigned long nodemask = 1; /* Node 0 */
>>>>   mbind(ptr, 1UL << 30, MPOL_PREFERRED_MANY, &nodemask,
>>>>         sizeof(nodemask) * 8, 0);
>>>>   memset(ptr, 0, 1UL << 30); /* Trigger fault */
>>>>
>>>> This results in a NULL pointer dereference:
>>>>
>>>>   BUG: kernel NULL pointer dereference, address: 0000000000000000
>>>>   #PF: supervisor read access in kernel mode
>>>>   #PF: error_code(0x0000) - not-present page
>>>>   Oops: Oops: 0000 [#1] SMP NOPTI
>>>>   RIP: 0010:hugetlb_cma_alloc_frozen_folio+0x75/0x120
>>>>   Call Trace:
>>>>    <TASK>
>>>>    only_alloc_fresh_hugetlb_folio.isra.0+0x2c/0x160
>>>>    alloc_surplus_hugetlb_folio+0x6d/0x100
>>>>    alloc_hugetlb_folio+0x3c5/0x660
>>>>    hugetlb_no_page+0x3d9/0x650
>>>>
>>>> Fixes: eb02f14c4a2b ("mm/hugetlb: allow overcommitting gigantic hugepages")
>>>> Cc: stable@vger.kernel.org
>>>> Signed-off-by: Sourav Panda <souravpanda@google.com>
>>>> ---
>>>> Changes in v4:
>>>> - Reverted the alloc_fresh_hugetlb_folio() cpuset snapshot approach from v3.
>>>>   As Muchun Song pointed out, snapshotting cpuset_current_mems_allowed does
>>>>   not prevent false-positive allocation failures without complex retry loops,
>>>>   and alloc_contig_frozen_pages() / alloc_buddy_frozen_folio() already handle
>>>>   NULL nodemasks safely internally.
>>>> - Handled NULL nodemask directly inside hugetlb_cma_alloc_frozen_folio()
>>>>   by defaulting nodemask to node_states[N_MEMORY] (Option 2), keeping
>>>>   HugeTLB allocators clean and consistent.
>>>> - v3: https://lore.kernel.org/linux-mm/20260705175119.440599-1-souravpanda@google.com/
>>>> - v2: https://lore.kernel.org/linux-mm/20260704174930.2885785-1-souravpanda@google.com/
>>>> - v1: https://lore.kernel.org/linux-mm/20260702215713.627941-1-souravpanda@google.com/
>>>>
>>>>  mm/hugetlb_cma.c | 5 ++++-
>>>>  1 file changed, 4 insertions(+), 1 deletion(-)
>>>>
>>>> diff --git a/mm/hugetlb_cma.c b/mm/hugetlb_cma.c
>>>> index 39344d6c78d8..5744de0ceeb7 100644
>>>> --- a/mm/hugetlb_cma.c
>>>> +++ b/mm/hugetlb_cma.c
>>>> @@ -34,7 +34,10 @@ struct folio *hugetlb_cma_alloc_frozen_folio(int order, gfp_t gfp_mask,
>>>>     if (!hugetlb_cma_size)
>>>>             return NULL;
>>>>
>>>> -   if (hugetlb_cma[nid])
>>>> +   if (!nodemask)
>>>> +           nodemask = &node_states[N_MEMORY];
>>>
>>> Hi Sourav,
>>>
>>> hmm what if cpuset.mems only allows allocation from node 0, and hugetlb CMA
>>> is only available on node 1. This will no now allow allocating a gigantic page
>>> on node 1 when its explicitly not allowed?
>>
>> That's fair point but should not the user be also responsible in provding a right
>> nodemask containing CMA memory if it prefers avoiding node_states[N_MEMORY] based
>> fallback mechanism in kernel ?
>>

I don't think the task can be made responsible for providing a CMA-aware fallback mask
here.  Userspace supplies the MPOL_PREFERRED_MANY preference, but the kernel itself
changes the second attempt to a NULL nodemask.  That permits fallback outside the
preferred policy nodes; it does not permit fallback outside the task's hardwall cpuset.

> 
> We have precedence for three options, which makes selecting one difficult :)
> 
> 1) nodemask = &cpuset_current_mems_allowed is currently being used in:
>        only_alloc_fresh_hugetlb_folio --> alloc_buddy_frozen_folio -->
> __alloc_frozen_pages_noprof --> prepare_alloc_pages.

Here I think using &cpuset_current_mems_allowed is part of the page allocator's wider cpuset handling.

> 2) Some places use cookies with cpuset_current_mems_allowed to prevent
> torn writes. But this adds complexity.
>        An example would be dequeue_hugetlb_folio_nodemask()
> 3) node_states[N_MEMORY] is sprinkled all over the kernel (simplest).
> 

I would prefer option 2.  A NULL mempolicy mask should fall back to
cpuset_current_mems_allowed, not all N_MEMORY nodes. 


>>>
>>> Would it not be better to check cpuset as well?
>>>
>>> Thanks,
>>> Usama
>>>
>>>> +
>>>> +   if (hugetlb_cma[nid] && node_isset(nid, *nodemask))
>>>>             page = cma_alloc_frozen_compound(hugetlb_cma[nid], order);
>>>>
>>>>     if (!page && !(gfp_mask & __GFP_THISNODE)) {
>>>> --
>>>> 2.55.0
>>>>



  reply	other threads:[~2026-08-06 11:07 UTC|newest]

Thread overview: 14+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-26  7:29 [PATCH v4] mm/hugetlb_cma: Fix null nodemask dereference in hugetlb_cma_alloc_frozen_folio Sourav Panda
2026-07-28  0:57 ` Andrew Morton
2026-08-03  6:37   ` Sourav Panda
2026-08-03  9:48 ` Anshuman Khandual
2026-08-05  4:42   ` Sourav Panda
2026-08-06  2:29     ` Anshuman Khandual
2026-08-05 20:24 ` Usama Arif
2026-08-05 23:05   ` Sourav Panda
2026-08-06  2:47   ` Anshuman Khandual
2026-08-06  5:33     ` Sourav Panda
2026-08-06 11:07       ` Usama Arif [this message]
2026-08-07  4:11         ` Sourav Panda
2026-08-07  5:28           ` Anshuman Khandual
2026-08-07  5:54             ` Sourav Panda

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=007cdb8b-01e3-4d73-9ac9-1a52ab5798dd@linux.dev \
    --to=usama.arif@linux.dev \
    --cc=akpm@linux-foundation.org \
    --cc=anshuman.khandual@arm.com \
    --cc=david@kernel.org \
    --cc=fvdl@google.com \
    --cc=gthelen@google.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=muchun.song@linux.dev \
    --cc=osalvador@suse.de \
    --cc=souravpanda@google.com \
    --cc=surenb@google.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.