* Re: [PATCH v4] mm/hugetlb_cma: Fix null nodemask dereference in hugetlb_cma_alloc_frozen_folio
2026-07-26 7:29 [PATCH v4] mm/hugetlb_cma: Fix null nodemask dereference in hugetlb_cma_alloc_frozen_folio Sourav Panda
@ 2026-07-28 0:57 ` Andrew Morton
2026-08-03 6:37 ` Sourav Panda
2026-08-03 9:48 ` Anshuman Khandual
2026-08-05 20:24 ` Usama Arif
2 siblings, 1 reply; 11+ messages in thread
From: Andrew Morton @ 2026-07-28 0:57 UTC (permalink / raw)
To: Sourav Panda
Cc: muchun.song, osalvador, david, surenb, fvdl, gthelen, linux-mm,
linux-kernel, Usama Arif, Shakeel Butt, Kefeng Wang
On Sun, 26 Jul 2026 07:29:34 +0000 Sourav Panda <souravpanda@google.com> wrote:
> alloc_buddy_hugetlb_folio_with_mpol() can pass a NULL nodemask to
> alloc_fresh_hugetlb_folio() as a fallback to allocate from all
> nodes. If order is gigantic, alloc_fresh_hugetlb_folio() propagates
> the NULL nodemask down to hugetlb_cma_alloc_frozen_folio() via
> alloc_gigantic_frozen_folio().
>
> hugetlb_cma_alloc_frozen_folio() blindly dereferences the nodemask in
> node_isset(nid, *nodemask) and for_each_node_mask(node, *nodemask),
> leading to a null pointer dereference kernel panic.
>
> Fix this by checking if nodemask is NULL in
> hugetlb_cma_alloc_frozen_folio() and defaulting it to
> node_states[N_MEMORY]. This allows hugetlb_cma allocations to fall
> back to any node with memory, keeping behavior consistent with
> alloc_contig_frozen_pages() and alloc_buddy_frozen_folio().
>
> >From a userspace perspective, this bug allows an unprivileged user to
> crash the kernel (trigger a panic) by requesting a gigantic hugepage
> allocation with MPOL_PREFERRED_MANY on a system where CMA is only
> configured on a subset of NUMA nodes.
>
> This can be reproduced by booting a VM with two NUMA nodes, restricting
> CMA to Node 1 (e.g., hugetlb_cma=1:1G default_hugepagesz=1G
> hugepagesz=1G hugepages=0), and running a program that allocates a
> 1GB hugepage area without reserving, restricts allocation to Node 0
> using mbind() with MPOL_PREFERRED_MANY, and triggers a page fault:
>
> void *ptr = mmap(NULL, 1UL << 30, PROT_READ | PROT_WRITE,
> MAP_PRIVATE | MAP_ANONYMOUS | MAP_HUGETLB |
> MAP_HUGE_1GB | MAP_NORESERVE, -1, 0);
> unsigned long nodemask = 1; /* Node 0 */
> mbind(ptr, 1UL << 30, MPOL_PREFERRED_MANY, &nodemask,
> sizeof(nodemask) * 8, 0);
> memset(ptr, 0, 1UL << 30); /* Trigger fault */
>
> This results in a NULL pointer dereference:
>
> BUG: kernel NULL pointer dereference, address: 0000000000000000
> #PF: supervisor read access in kernel mode
> #PF: error_code(0x0000) - not-present page
> Oops: Oops: 0000 [#1] SMP NOPTI
> RIP: 0010:hugetlb_cma_alloc_frozen_folio+0x75/0x120
> Call Trace:
> <TASK>
> only_alloc_fresh_hugetlb_folio.isra.0+0x2c/0x160
> alloc_surplus_hugetlb_folio+0x6d/0x100
> alloc_hugetlb_folio+0x3c5/0x660
> hugetlb_no_page+0x3d9/0x650
Thanks. I'll queue this as a hotfix while we await review.
> Fixes: eb02f14c4a2b ("mm/hugetlb: allow overcommitting gigantic hugepages")
It's a good idea to cc the people who were involved in the to-be-fixed
patch - Author and Reviewers.
>
> ...
>
> --- a/mm/hugetlb_cma.c
> +++ b/mm/hugetlb_cma.c
> @@ -34,7 +34,10 @@ struct folio *hugetlb_cma_alloc_frozen_folio(int order, gfp_t gfp_mask,
> if (!hugetlb_cma_size)
> return NULL;
>
> - if (hugetlb_cma[nid])
> + if (!nodemask)
> + nodemask = &node_states[N_MEMORY];
> +
> + if (hugetlb_cma[nid] && node_isset(nid, *nodemask))
> page = cma_alloc_frozen_compound(hugetlb_cma[nid], order);
>
> if (!page && !(gfp_mask & __GFP_THISNODE)) {
> --
> 2.55.0
^ permalink raw reply [flat|nested] 11+ messages in thread* Re: [PATCH v4] mm/hugetlb_cma: Fix null nodemask dereference in hugetlb_cma_alloc_frozen_folio
2026-07-28 0:57 ` Andrew Morton
@ 2026-08-03 6:37 ` Sourav Panda
0 siblings, 0 replies; 11+ messages in thread
From: Sourav Panda @ 2026-08-03 6:37 UTC (permalink / raw)
To: Andrew Morton
Cc: muchun.song, osalvador, david, surenb, fvdl, gthelen, linux-mm,
linux-kernel, Usama Arif, Shakeel Butt, Kefeng Wang,
Vlastimil Babka, Michal Hocko, Brendan Jackman, Johannes Weiner,
Zi Yan
On Mon, Jul 27, 2026 at 5:57 PM Andrew Morton <akpm@linux-foundation.org> wrote:
>
> On Sun, 26 Jul 2026 07:29:34 +0000 Sourav Panda <souravpanda@google.com> wrote:
>
> > alloc_buddy_hugetlb_folio_with_mpol() can pass a NULL nodemask to
> > alloc_fresh_hugetlb_folio() as a fallback to allocate from all
> > nodes. If order is gigantic, alloc_fresh_hugetlb_folio() propagates
> > the NULL nodemask down to hugetlb_cma_alloc_frozen_folio() via
> > alloc_gigantic_frozen_folio().
> >
> > hugetlb_cma_alloc_frozen_folio() blindly dereferences the nodemask in
> > node_isset(nid, *nodemask) and for_each_node_mask(node, *nodemask),
> > leading to a null pointer dereference kernel panic.
> >
> > Fix this by checking if nodemask is NULL in
> > hugetlb_cma_alloc_frozen_folio() and defaulting it to
> > node_states[N_MEMORY]. This allows hugetlb_cma allocations to fall
> > back to any node with memory, keeping behavior consistent with
> > alloc_contig_frozen_pages() and alloc_buddy_frozen_folio().
> >
> > >From a userspace perspective, this bug allows an unprivileged user to
> > crash the kernel (trigger a panic) by requesting a gigantic hugepage
> > allocation with MPOL_PREFERRED_MANY on a system where CMA is only
> > configured on a subset of NUMA nodes.
> >
> > This can be reproduced by booting a VM with two NUMA nodes, restricting
> > CMA to Node 1 (e.g., hugetlb_cma=1:1G default_hugepagesz=1G
> > hugepagesz=1G hugepages=0), and running a program that allocates a
> > 1GB hugepage area without reserving, restricts allocation to Node 0
> > using mbind() with MPOL_PREFERRED_MANY, and triggers a page fault:
> >
> > void *ptr = mmap(NULL, 1UL << 30, PROT_READ | PROT_WRITE,
> > MAP_PRIVATE | MAP_ANONYMOUS | MAP_HUGETLB |
> > MAP_HUGE_1GB | MAP_NORESERVE, -1, 0);
> > unsigned long nodemask = 1; /* Node 0 */
> > mbind(ptr, 1UL << 30, MPOL_PREFERRED_MANY, &nodemask,
> > sizeof(nodemask) * 8, 0);
> > memset(ptr, 0, 1UL << 30); /* Trigger fault */
> >
> > This results in a NULL pointer dereference:
> >
> > BUG: kernel NULL pointer dereference, address: 0000000000000000
> > #PF: supervisor read access in kernel mode
> > #PF: error_code(0x0000) - not-present page
> > Oops: Oops: 0000 [#1] SMP NOPTI
> > RIP: 0010:hugetlb_cma_alloc_frozen_folio+0x75/0x120
> > Call Trace:
> > <TASK>
> > only_alloc_fresh_hugetlb_folio.isra.0+0x2c/0x160
> > alloc_surplus_hugetlb_folio+0x6d/0x100
> > alloc_hugetlb_folio+0x3c5/0x660
> > hugetlb_no_page+0x3d9/0x650
>
> Thanks. I'll queue this as a hotfix while we await review.
>
> > Fixes: eb02f14c4a2b ("mm/hugetlb: allow overcommitting gigantic hugepages")
>
> It's a good idea to cc the people who were involved in the to-be-fixed
> patch - Author and Reviewers.
>
Done!
> >
> > ...
> >
> > --- a/mm/hugetlb_cma.c
> > +++ b/mm/hugetlb_cma.c
> > @@ -34,7 +34,10 @@ struct folio *hugetlb_cma_alloc_frozen_folio(int order, gfp_t gfp_mask,
> > if (!hugetlb_cma_size)
> > return NULL;
> >
> > - if (hugetlb_cma[nid])
> > + if (!nodemask)
> > + nodemask = &node_states[N_MEMORY];
> > +
> > + if (hugetlb_cma[nid] && node_isset(nid, *nodemask))
> > page = cma_alloc_frozen_compound(hugetlb_cma[nid], order);
> >
> > if (!page && !(gfp_mask & __GFP_THISNODE)) {
> > --
> > 2.55.0
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH v4] mm/hugetlb_cma: Fix null nodemask dereference in hugetlb_cma_alloc_frozen_folio
2026-07-26 7:29 [PATCH v4] mm/hugetlb_cma: Fix null nodemask dereference in hugetlb_cma_alloc_frozen_folio Sourav Panda
2026-07-28 0:57 ` Andrew Morton
@ 2026-08-03 9:48 ` Anshuman Khandual
2026-08-05 4:42 ` Sourav Panda
2026-08-05 20:24 ` Usama Arif
2 siblings, 1 reply; 11+ messages in thread
From: Anshuman Khandual @ 2026-08-03 9:48 UTC (permalink / raw)
To: Sourav Panda
Cc: muchun.song, osalvador, akpm, david, surenb, fvdl, gthelen,
linux-mm, linux-kernel
On Sun, Jul 26, 2026 at 07:29:34AM +0000, Sourav Panda wrote:
> alloc_buddy_hugetlb_folio_with_mpol() can pass a NULL nodemask to
> alloc_fresh_hugetlb_folio() as a fallback to allocate from all
> nodes. If order is gigantic, alloc_fresh_hugetlb_folio() propagates
> the NULL nodemask down to hugetlb_cma_alloc_frozen_folio() via
> alloc_gigantic_frozen_folio().
>
> hugetlb_cma_alloc_frozen_folio() blindly dereferences the nodemask in
> node_isset(nid, *nodemask) and for_each_node_mask(node, *nodemask),
> leading to a null pointer dereference kernel panic.
node_isset() is being addded here in this change right ? OR was there
another path via node_isset() which could have been called on a NULL
valued nodemask ? Although subsequent for_each_node_mask() could do
the required dereference as mentioned earlier.
>
> Fix this by checking if nodemask is NULL in
> hugetlb_cma_alloc_frozen_folio() and defaulting it to
> node_states[N_MEMORY]. This allows hugetlb_cma allocations to fall
> back to any node with memory, keeping behavior consistent with
> alloc_contig_frozen_pages() and alloc_buddy_frozen_folio().
>
> From a userspace perspective, this bug allows an unprivileged user to
> crash the kernel (trigger a panic) by requesting a gigantic hugepage
> allocation with MPOL_PREFERRED_MANY on a system where CMA is only
> configured on a subset of NUMA nodes.
>
> This can be reproduced by booting a VM with two NUMA nodes, restricting
> CMA to Node 1 (e.g., hugetlb_cma=1:1G default_hugepagesz=1G
> hugepagesz=1G hugepages=0), and running a program that allocates a
> 1GB hugepage area without reserving, restricts allocation to Node 0
> using mbind() with MPOL_PREFERRED_MANY, and triggers a page fault:
>
> void *ptr = mmap(NULL, 1UL << 30, PROT_READ | PROT_WRITE,
> MAP_PRIVATE | MAP_ANONYMOUS | MAP_HUGETLB |
> MAP_HUGE_1GB | MAP_NORESERVE, -1, 0);
> unsigned long nodemask = 1; /* Node 0 */
> mbind(ptr, 1UL << 30, MPOL_PREFERRED_MANY, &nodemask,
> sizeof(nodemask) * 8, 0);
> memset(ptr, 0, 1UL << 30); /* Trigger fault */
>
> This results in a NULL pointer dereference:
>
> BUG: kernel NULL pointer dereference, address: 0000000000000000
> #PF: supervisor read access in kernel mode
> #PF: error_code(0x0000) - not-present page
> Oops: Oops: 0000 [#1] SMP NOPTI
> RIP: 0010:hugetlb_cma_alloc_frozen_folio+0x75/0x120
> Call Trace:
> <TASK>
> only_alloc_fresh_hugetlb_folio.isra.0+0x2c/0x160
> alloc_surplus_hugetlb_folio+0x6d/0x100
> alloc_hugetlb_folio+0x3c5/0x660
> hugetlb_no_page+0x3d9/0x650
>
> Fixes: eb02f14c4a2b ("mm/hugetlb: allow overcommitting gigantic hugepages")
> Cc: stable@vger.kernel.org
> Signed-off-by: Sourav Panda <souravpanda@google.com>
> ---
> Changes in v4:
> - Reverted the alloc_fresh_hugetlb_folio() cpuset snapshot approach from v3.
> As Muchun Song pointed out, snapshotting cpuset_current_mems_allowed does
> not prevent false-positive allocation failures without complex retry loops,
> and alloc_contig_frozen_pages() / alloc_buddy_frozen_folio() already handle
> NULL nodemasks safely internally.
> - Handled NULL nodemask directly inside hugetlb_cma_alloc_frozen_folio()
> by defaulting nodemask to node_states[N_MEMORY] (Option 2), keeping
> HugeTLB allocators clean and consistent.
> - v3: https://lore.kernel.org/linux-mm/20260705175119.440599-1-souravpanda@google.com/
> - v2: https://lore.kernel.org/linux-mm/20260704174930.2885785-1-souravpanda@google.com/
> - v1: https://lore.kernel.org/linux-mm/20260702215713.627941-1-souravpanda@google.com/
>
> mm/hugetlb_cma.c | 5 ++++-
> 1 file changed, 4 insertions(+), 1 deletion(-)
>
> diff --git a/mm/hugetlb_cma.c b/mm/hugetlb_cma.c
> index 39344d6c78d8..5744de0ceeb7 100644
> --- a/mm/hugetlb_cma.c
> +++ b/mm/hugetlb_cma.c
> @@ -34,7 +34,10 @@ struct folio *hugetlb_cma_alloc_frozen_folio(int order, gfp_t gfp_mask,
> if (!hugetlb_cma_size)
> return NULL;
>
> - if (hugetlb_cma[nid])
> + if (!nodemask)
> + nodemask = &node_states[N_MEMORY];
> +
> + if (hugetlb_cma[nid] && node_isset(nid, *nodemask))
> page = cma_alloc_frozen_compound(hugetlb_cma[nid], order);
>
> if (!page && !(gfp_mask & __GFP_THISNODE)) {
> --
> 2.55.0
^ permalink raw reply [flat|nested] 11+ messages in thread* Re: [PATCH v4] mm/hugetlb_cma: Fix null nodemask dereference in hugetlb_cma_alloc_frozen_folio
2026-08-03 9:48 ` Anshuman Khandual
@ 2026-08-05 4:42 ` Sourav Panda
2026-08-06 2:29 ` Anshuman Khandual
0 siblings, 1 reply; 11+ messages in thread
From: Sourav Panda @ 2026-08-05 4:42 UTC (permalink / raw)
To: Anshuman Khandual
Cc: muchun.song, osalvador, akpm, david, surenb, fvdl, gthelen,
linux-mm, linux-kernel
On Mon, Aug 3, 2026 at 2:48 AM Anshuman Khandual
<anshuman.khandual@arm.com> wrote:
>
> On Sun, Jul 26, 2026 at 07:29:34AM +0000, Sourav Panda wrote:
> > alloc_buddy_hugetlb_folio_with_mpol() can pass a NULL nodemask to
> > alloc_fresh_hugetlb_folio() as a fallback to allocate from all
> > nodes. If order is gigantic, alloc_fresh_hugetlb_folio() propagates
> > the NULL nodemask down to hugetlb_cma_alloc_frozen_folio() via
> > alloc_gigantic_frozen_folio().
> >
> > hugetlb_cma_alloc_frozen_folio() blindly dereferences the nodemask in
> > node_isset(nid, *nodemask) and for_each_node_mask(node, *nodemask),
> > leading to a null pointer dereference kernel panic.
>
> node_isset() is being addded here in this change right ? OR was there
> another path via node_isset() which could have been called on a NULL
> valued nodemask ? Although subsequent for_each_node_mask() could do
> the required dereference as mentioned earlier.
>
Thanks for the review, Anshuman!
Sashiko [1] suggested adding node_isset(nid, *nodemask) in v1 because
without it, hugetlb_cma_alloc_frozen_folio() could allocate from a node
without checking whether it is actually allowed by the caller's memory
policy / nodemask.
Once we added that check, both node_isset(nid, *nodemask) and
for_each_node_mask(node, *nodemask) became potential NULL pointer
dereferences when nodemask is NULL.
[1] https://sashiko.dev/#/patchset/20260702215713.627941-1-souravpanda@google.com
Thanks!
Sourav
> >
> > Fix this by checking if nodemask is NULL in
> > hugetlb_cma_alloc_frozen_folio() and defaulting it to
> > node_states[N_MEMORY]. This allows hugetlb_cma allocations to fall
> > back to any node with memory, keeping behavior consistent with
> > alloc_contig_frozen_pages() and alloc_buddy_frozen_folio().
> >
> > From a userspace perspective, this bug allows an unprivileged user to
> > crash the kernel (trigger a panic) by requesting a gigantic hugepage
> > allocation with MPOL_PREFERRED_MANY on a system where CMA is only
> > configured on a subset of NUMA nodes.
> >
> > This can be reproduced by booting a VM with two NUMA nodes, restricting
> > CMA to Node 1 (e.g., hugetlb_cma=1:1G default_hugepagesz=1G
> > hugepagesz=1G hugepages=0), and running a program that allocates a
> > 1GB hugepage area without reserving, restricts allocation to Node 0
> > using mbind() with MPOL_PREFERRED_MANY, and triggers a page fault:
> >
> > void *ptr = mmap(NULL, 1UL << 30, PROT_READ | PROT_WRITE,
> > MAP_PRIVATE | MAP_ANONYMOUS | MAP_HUGETLB |
> > MAP_HUGE_1GB | MAP_NORESERVE, -1, 0);
> > unsigned long nodemask = 1; /* Node 0 */
> > mbind(ptr, 1UL << 30, MPOL_PREFERRED_MANY, &nodemask,
> > sizeof(nodemask) * 8, 0);
> > memset(ptr, 0, 1UL << 30); /* Trigger fault */
> >
> > This results in a NULL pointer dereference:
> >
> > BUG: kernel NULL pointer dereference, address: 0000000000000000
> > #PF: supervisor read access in kernel mode
> > #PF: error_code(0x0000) - not-present page
> > Oops: Oops: 0000 [#1] SMP NOPTI
> > RIP: 0010:hugetlb_cma_alloc_frozen_folio+0x75/0x120
> > Call Trace:
> > <TASK>
> > only_alloc_fresh_hugetlb_folio.isra.0+0x2c/0x160
> > alloc_surplus_hugetlb_folio+0x6d/0x100
> > alloc_hugetlb_folio+0x3c5/0x660
> > hugetlb_no_page+0x3d9/0x650
> >
> > Fixes: eb02f14c4a2b ("mm/hugetlb: allow overcommitting gigantic hugepages")
> > Cc: stable@vger.kernel.org
> > Signed-off-by: Sourav Panda <souravpanda@google.com>
> > ---
> > Changes in v4:
> > - Reverted the alloc_fresh_hugetlb_folio() cpuset snapshot approach from v3.
> > As Muchun Song pointed out, snapshotting cpuset_current_mems_allowed does
> > not prevent false-positive allocation failures without complex retry loops,
> > and alloc_contig_frozen_pages() / alloc_buddy_frozen_folio() already handle
> > NULL nodemasks safely internally.
> > - Handled NULL nodemask directly inside hugetlb_cma_alloc_frozen_folio()
> > by defaulting nodemask to node_states[N_MEMORY] (Option 2), keeping
> > HugeTLB allocators clean and consistent.
> > - v3: https://lore.kernel.org/linux-mm/20260705175119.440599-1-souravpanda@google.com/
> > - v2: https://lore.kernel.org/linux-mm/20260704174930.2885785-1-souravpanda@google.com/
> > - v1: https://lore.kernel.org/linux-mm/20260702215713.627941-1-souravpanda@google.com/
> >
> > mm/hugetlb_cma.c | 5 ++++-
> > 1 file changed, 4 insertions(+), 1 deletion(-)
> >
> > diff --git a/mm/hugetlb_cma.c b/mm/hugetlb_cma.c
> > index 39344d6c78d8..5744de0ceeb7 100644
> > --- a/mm/hugetlb_cma.c
> > +++ b/mm/hugetlb_cma.c
> > @@ -34,7 +34,10 @@ struct folio *hugetlb_cma_alloc_frozen_folio(int order, gfp_t gfp_mask,
> > if (!hugetlb_cma_size)
> > return NULL;
> >
> > - if (hugetlb_cma[nid])
> > + if (!nodemask)
> > + nodemask = &node_states[N_MEMORY];
> > +
> > + if (hugetlb_cma[nid] && node_isset(nid, *nodemask))
> > page = cma_alloc_frozen_compound(hugetlb_cma[nid], order);
> >
> > if (!page && !(gfp_mask & __GFP_THISNODE)) {
> > --
> > 2.55.0
^ permalink raw reply [flat|nested] 11+ messages in thread* Re: [PATCH v4] mm/hugetlb_cma: Fix null nodemask dereference in hugetlb_cma_alloc_frozen_folio
2026-08-05 4:42 ` Sourav Panda
@ 2026-08-06 2:29 ` Anshuman Khandual
0 siblings, 0 replies; 11+ messages in thread
From: Anshuman Khandual @ 2026-08-06 2:29 UTC (permalink / raw)
To: Sourav Panda
Cc: muchun.song, osalvador, akpm, david, surenb, fvdl, gthelen,
linux-mm, linux-kernel
On Tue, Aug 04, 2026 at 09:42:45PM -0700, Sourav Panda wrote:
> On Mon, Aug 3, 2026 at 2:48 AM Anshuman Khandual
> <anshuman.khandual@arm.com> wrote:
> >
> > On Sun, Jul 26, 2026 at 07:29:34AM +0000, Sourav Panda wrote:
> > > alloc_buddy_hugetlb_folio_with_mpol() can pass a NULL nodemask to
> > > alloc_fresh_hugetlb_folio() as a fallback to allocate from all
> > > nodes. If order is gigantic, alloc_fresh_hugetlb_folio() propagates
> > > the NULL nodemask down to hugetlb_cma_alloc_frozen_folio() via
> > > alloc_gigantic_frozen_folio().
> > >
> > > hugetlb_cma_alloc_frozen_folio() blindly dereferences the nodemask in
> > > node_isset(nid, *nodemask) and for_each_node_mask(node, *nodemask),
> > > leading to a null pointer dereference kernel panic.
> >
> > node_isset() is being addded here in this change right ? OR was there
> > another path via node_isset() which could have been called on a NULL
> > valued nodemask ? Although subsequent for_each_node_mask() could do
> > the required dereference as mentioned earlier.
> >
>
> Thanks for the review, Anshuman!
>
> Sashiko [1] suggested adding node_isset(nid, *nodemask) in v1 because
> without it, hugetlb_cma_alloc_frozen_folio() could allocate from a node
> without checking whether it is actually allowed by the caller's memory
> policy / nodemask.
Fair enough.
>
> Once we added that check, both node_isset(nid, *nodemask) and
> for_each_node_mask(node, *nodemask) became potential NULL pointer
> dereferences when nodemask is NULL.
Got it, thanks !
>
> [1] https://sashiko.dev/#/patchset/20260702215713.627941-1-souravpanda@google.com
>
> Thanks!
> Sourav
>
>
> > >
> > > Fix this by checking if nodemask is NULL in
> > > hugetlb_cma_alloc_frozen_folio() and defaulting it to
> > > node_states[N_MEMORY]. This allows hugetlb_cma allocations to fall
> > > back to any node with memory, keeping behavior consistent with
> > > alloc_contig_frozen_pages() and alloc_buddy_frozen_folio().
> > >
> > > From a userspace perspective, this bug allows an unprivileged user to
> > > crash the kernel (trigger a panic) by requesting a gigantic hugepage
> > > allocation with MPOL_PREFERRED_MANY on a system where CMA is only
> > > configured on a subset of NUMA nodes.
> > >
> > > This can be reproduced by booting a VM with two NUMA nodes, restricting
> > > CMA to Node 1 (e.g., hugetlb_cma=1:1G default_hugepagesz=1G
> > > hugepagesz=1G hugepages=0), and running a program that allocates a
> > > 1GB hugepage area without reserving, restricts allocation to Node 0
> > > using mbind() with MPOL_PREFERRED_MANY, and triggers a page fault:
> > >
> > > void *ptr = mmap(NULL, 1UL << 30, PROT_READ | PROT_WRITE,
> > > MAP_PRIVATE | MAP_ANONYMOUS | MAP_HUGETLB |
> > > MAP_HUGE_1GB | MAP_NORESERVE, -1, 0);
> > > unsigned long nodemask = 1; /* Node 0 */
> > > mbind(ptr, 1UL << 30, MPOL_PREFERRED_MANY, &nodemask,
> > > sizeof(nodemask) * 8, 0);
> > > memset(ptr, 0, 1UL << 30); /* Trigger fault */
> > >
> > > This results in a NULL pointer dereference:
> > >
> > > BUG: kernel NULL pointer dereference, address: 0000000000000000
> > > #PF: supervisor read access in kernel mode
> > > #PF: error_code(0x0000) - not-present page
> > > Oops: Oops: 0000 [#1] SMP NOPTI
> > > RIP: 0010:hugetlb_cma_alloc_frozen_folio+0x75/0x120
> > > Call Trace:
> > > <TASK>
> > > only_alloc_fresh_hugetlb_folio.isra.0+0x2c/0x160
> > > alloc_surplus_hugetlb_folio+0x6d/0x100
> > > alloc_hugetlb_folio+0x3c5/0x660
> > > hugetlb_no_page+0x3d9/0x650
> > >
> > > Fixes: eb02f14c4a2b ("mm/hugetlb: allow overcommitting gigantic hugepages")
> > > Cc: stable@vger.kernel.org
> > > Signed-off-by: Sourav Panda <souravpanda@google.com>
> > > ---
> > > Changes in v4:
> > > - Reverted the alloc_fresh_hugetlb_folio() cpuset snapshot approach from v3.
> > > As Muchun Song pointed out, snapshotting cpuset_current_mems_allowed does
> > > not prevent false-positive allocation failures without complex retry loops,
> > > and alloc_contig_frozen_pages() / alloc_buddy_frozen_folio() already handle
> > > NULL nodemasks safely internally.
> > > - Handled NULL nodemask directly inside hugetlb_cma_alloc_frozen_folio()
> > > by defaulting nodemask to node_states[N_MEMORY] (Option 2), keeping
> > > HugeTLB allocators clean and consistent.
> > > - v3: https://lore.kernel.org/linux-mm/20260705175119.440599-1-souravpanda@google.com/
> > > - v2: https://lore.kernel.org/linux-mm/20260704174930.2885785-1-souravpanda@google.com/
> > > - v1: https://lore.kernel.org/linux-mm/20260702215713.627941-1-souravpanda@google.com/
> > >
> > > mm/hugetlb_cma.c | 5 ++++-
> > > 1 file changed, 4 insertions(+), 1 deletion(-)
> > >
> > > diff --git a/mm/hugetlb_cma.c b/mm/hugetlb_cma.c
> > > index 39344d6c78d8..5744de0ceeb7 100644
> > > --- a/mm/hugetlb_cma.c
> > > +++ b/mm/hugetlb_cma.c
> > > @@ -34,7 +34,10 @@ struct folio *hugetlb_cma_alloc_frozen_folio(int order, gfp_t gfp_mask,
> > > if (!hugetlb_cma_size)
> > > return NULL;
> > >
> > > - if (hugetlb_cma[nid])
> > > + if (!nodemask)
> > > + nodemask = &node_states[N_MEMORY];
> > > +
> > > + if (hugetlb_cma[nid] && node_isset(nid, *nodemask))
> > > page = cma_alloc_frozen_compound(hugetlb_cma[nid], order);
> > >
> > > if (!page && !(gfp_mask & __GFP_THISNODE)) {
> > > --
> > > 2.55.0
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH v4] mm/hugetlb_cma: Fix null nodemask dereference in hugetlb_cma_alloc_frozen_folio
2026-07-26 7:29 [PATCH v4] mm/hugetlb_cma: Fix null nodemask dereference in hugetlb_cma_alloc_frozen_folio Sourav Panda
2026-07-28 0:57 ` Andrew Morton
2026-08-03 9:48 ` Anshuman Khandual
@ 2026-08-05 20:24 ` Usama Arif
2026-08-05 23:05 ` Sourav Panda
2026-08-06 2:47 ` Anshuman Khandual
2 siblings, 2 replies; 11+ messages in thread
From: Usama Arif @ 2026-08-05 20:24 UTC (permalink / raw)
To: Sourav Panda
Cc: Usama Arif, muchun.song, osalvador, akpm, david, surenb, fvdl,
gthelen, linux-mm, linux-kernel
On Sun, 26 Jul 2026 07:29:34 +0000 Sourav Panda <souravpanda@google.com> wrote:
> alloc_buddy_hugetlb_folio_with_mpol() can pass a NULL nodemask to
> alloc_fresh_hugetlb_folio() as a fallback to allocate from all
> nodes. If order is gigantic, alloc_fresh_hugetlb_folio() propagates
> the NULL nodemask down to hugetlb_cma_alloc_frozen_folio() via
> alloc_gigantic_frozen_folio().
>
> hugetlb_cma_alloc_frozen_folio() blindly dereferences the nodemask in
> node_isset(nid, *nodemask) and for_each_node_mask(node, *nodemask),
> leading to a null pointer dereference kernel panic.
>
> Fix this by checking if nodemask is NULL in
> hugetlb_cma_alloc_frozen_folio() and defaulting it to
> node_states[N_MEMORY]. This allows hugetlb_cma allocations to fall
> back to any node with memory, keeping behavior consistent with
> alloc_contig_frozen_pages() and alloc_buddy_frozen_folio().
>
> >From a userspace perspective, this bug allows an unprivileged user to
> crash the kernel (trigger a panic) by requesting a gigantic hugepage
> allocation with MPOL_PREFERRED_MANY on a system where CMA is only
> configured on a subset of NUMA nodes.
>
> This can be reproduced by booting a VM with two NUMA nodes, restricting
> CMA to Node 1 (e.g., hugetlb_cma=1:1G default_hugepagesz=1G
> hugepagesz=1G hugepages=0), and running a program that allocates a
> 1GB hugepage area without reserving, restricts allocation to Node 0
> using mbind() with MPOL_PREFERRED_MANY, and triggers a page fault:
>
> void *ptr = mmap(NULL, 1UL << 30, PROT_READ | PROT_WRITE,
> MAP_PRIVATE | MAP_ANONYMOUS | MAP_HUGETLB |
> MAP_HUGE_1GB | MAP_NORESERVE, -1, 0);
> unsigned long nodemask = 1; /* Node 0 */
> mbind(ptr, 1UL << 30, MPOL_PREFERRED_MANY, &nodemask,
> sizeof(nodemask) * 8, 0);
> memset(ptr, 0, 1UL << 30); /* Trigger fault */
>
> This results in a NULL pointer dereference:
>
> BUG: kernel NULL pointer dereference, address: 0000000000000000
> #PF: supervisor read access in kernel mode
> #PF: error_code(0x0000) - not-present page
> Oops: Oops: 0000 [#1] SMP NOPTI
> RIP: 0010:hugetlb_cma_alloc_frozen_folio+0x75/0x120
> Call Trace:
> <TASK>
> only_alloc_fresh_hugetlb_folio.isra.0+0x2c/0x160
> alloc_surplus_hugetlb_folio+0x6d/0x100
> alloc_hugetlb_folio+0x3c5/0x660
> hugetlb_no_page+0x3d9/0x650
>
> Fixes: eb02f14c4a2b ("mm/hugetlb: allow overcommitting gigantic hugepages")
> Cc: stable@vger.kernel.org
> Signed-off-by: Sourav Panda <souravpanda@google.com>
> ---
> Changes in v4:
> - Reverted the alloc_fresh_hugetlb_folio() cpuset snapshot approach from v3.
> As Muchun Song pointed out, snapshotting cpuset_current_mems_allowed does
> not prevent false-positive allocation failures without complex retry loops,
> and alloc_contig_frozen_pages() / alloc_buddy_frozen_folio() already handle
> NULL nodemasks safely internally.
> - Handled NULL nodemask directly inside hugetlb_cma_alloc_frozen_folio()
> by defaulting nodemask to node_states[N_MEMORY] (Option 2), keeping
> HugeTLB allocators clean and consistent.
> - v3: https://lore.kernel.org/linux-mm/20260705175119.440599-1-souravpanda@google.com/
> - v2: https://lore.kernel.org/linux-mm/20260704174930.2885785-1-souravpanda@google.com/
> - v1: https://lore.kernel.org/linux-mm/20260702215713.627941-1-souravpanda@google.com/
>
> mm/hugetlb_cma.c | 5 ++++-
> 1 file changed, 4 insertions(+), 1 deletion(-)
>
> diff --git a/mm/hugetlb_cma.c b/mm/hugetlb_cma.c
> index 39344d6c78d8..5744de0ceeb7 100644
> --- a/mm/hugetlb_cma.c
> +++ b/mm/hugetlb_cma.c
> @@ -34,7 +34,10 @@ struct folio *hugetlb_cma_alloc_frozen_folio(int order, gfp_t gfp_mask,
> if (!hugetlb_cma_size)
> return NULL;
>
> - if (hugetlb_cma[nid])
> + if (!nodemask)
> + nodemask = &node_states[N_MEMORY];
Hi Sourav,
hmm what if cpuset.mems only allows allocation from node 0, and hugetlb CMA
is only available on node 1. This will no now allow allocating a gigantic page
on node 1 when its explicitly not allowed?
Would it not be better to check cpuset as well?
Thanks,
Usama
> +
> + if (hugetlb_cma[nid] && node_isset(nid, *nodemask))
> page = cma_alloc_frozen_compound(hugetlb_cma[nid], order);
>
> if (!page && !(gfp_mask & __GFP_THISNODE)) {
> --
> 2.55.0
>
^ permalink raw reply [flat|nested] 11+ messages in thread* Re: [PATCH v4] mm/hugetlb_cma: Fix null nodemask dereference in hugetlb_cma_alloc_frozen_folio
2026-08-05 20:24 ` Usama Arif
@ 2026-08-05 23:05 ` Sourav Panda
2026-08-06 2:47 ` Anshuman Khandual
1 sibling, 0 replies; 11+ messages in thread
From: Sourav Panda @ 2026-08-05 23:05 UTC (permalink / raw)
To: Usama Arif
Cc: muchun.song, osalvador, akpm, david, surenb, fvdl, gthelen,
linux-mm, linux-kernel
On Wed, Aug 5, 2026 at 1:24 PM Usama Arif <usama.arif@linux.dev> wrote:
>
> On Sun, 26 Jul 2026 07:29:34 +0000 Sourav Panda <souravpanda@google.com> wrote:
>
> > alloc_buddy_hugetlb_folio_with_mpol() can pass a NULL nodemask to
> > alloc_fresh_hugetlb_folio() as a fallback to allocate from all
> > nodes. If order is gigantic, alloc_fresh_hugetlb_folio() propagates
> > the NULL nodemask down to hugetlb_cma_alloc_frozen_folio() via
> > alloc_gigantic_frozen_folio().
> >
> > hugetlb_cma_alloc_frozen_folio() blindly dereferences the nodemask in
> > node_isset(nid, *nodemask) and for_each_node_mask(node, *nodemask),
> > leading to a null pointer dereference kernel panic.
> >
> > Fix this by checking if nodemask is NULL in
> > hugetlb_cma_alloc_frozen_folio() and defaulting it to
> > node_states[N_MEMORY]. This allows hugetlb_cma allocations to fall
> > back to any node with memory, keeping behavior consistent with
> > alloc_contig_frozen_pages() and alloc_buddy_frozen_folio().
> >
> > >From a userspace perspective, this bug allows an unprivileged user to
> > crash the kernel (trigger a panic) by requesting a gigantic hugepage
> > allocation with MPOL_PREFERRED_MANY on a system where CMA is only
> > configured on a subset of NUMA nodes.
> >
> > This can be reproduced by booting a VM with two NUMA nodes, restricting
> > CMA to Node 1 (e.g., hugetlb_cma=1:1G default_hugepagesz=1G
> > hugepagesz=1G hugepages=0), and running a program that allocates a
> > 1GB hugepage area without reserving, restricts allocation to Node 0
> > using mbind() with MPOL_PREFERRED_MANY, and triggers a page fault:
> >
> > void *ptr = mmap(NULL, 1UL << 30, PROT_READ | PROT_WRITE,
> > MAP_PRIVATE | MAP_ANONYMOUS | MAP_HUGETLB |
> > MAP_HUGE_1GB | MAP_NORESERVE, -1, 0);
> > unsigned long nodemask = 1; /* Node 0 */
> > mbind(ptr, 1UL << 30, MPOL_PREFERRED_MANY, &nodemask,
> > sizeof(nodemask) * 8, 0);
> > memset(ptr, 0, 1UL << 30); /* Trigger fault */
> >
> > This results in a NULL pointer dereference:
> >
> > BUG: kernel NULL pointer dereference, address: 0000000000000000
> > #PF: supervisor read access in kernel mode
> > #PF: error_code(0x0000) - not-present page
> > Oops: Oops: 0000 [#1] SMP NOPTI
> > RIP: 0010:hugetlb_cma_alloc_frozen_folio+0x75/0x120
> > Call Trace:
> > <TASK>
> > only_alloc_fresh_hugetlb_folio.isra.0+0x2c/0x160
> > alloc_surplus_hugetlb_folio+0x6d/0x100
> > alloc_hugetlb_folio+0x3c5/0x660
> > hugetlb_no_page+0x3d9/0x650
> >
> > Fixes: eb02f14c4a2b ("mm/hugetlb: allow overcommitting gigantic hugepages")
> > Cc: stable@vger.kernel.org
> > Signed-off-by: Sourav Panda <souravpanda@google.com>
> > ---
> > Changes in v4:
> > - Reverted the alloc_fresh_hugetlb_folio() cpuset snapshot approach from v3.
> > As Muchun Song pointed out, snapshotting cpuset_current_mems_allowed does
> > not prevent false-positive allocation failures without complex retry loops,
> > and alloc_contig_frozen_pages() / alloc_buddy_frozen_folio() already handle
> > NULL nodemasks safely internally.
> > - Handled NULL nodemask directly inside hugetlb_cma_alloc_frozen_folio()
> > by defaulting nodemask to node_states[N_MEMORY] (Option 2), keeping
> > HugeTLB allocators clean and consistent.
> > - v3: https://lore.kernel.org/linux-mm/20260705175119.440599-1-souravpanda@google.com/
> > - v2: https://lore.kernel.org/linux-mm/20260704174930.2885785-1-souravpanda@google.com/
> > - v1: https://lore.kernel.org/linux-mm/20260702215713.627941-1-souravpanda@google.com/
> >
> > mm/hugetlb_cma.c | 5 ++++-
> > 1 file changed, 4 insertions(+), 1 deletion(-)
> >
> > diff --git a/mm/hugetlb_cma.c b/mm/hugetlb_cma.c
> > index 39344d6c78d8..5744de0ceeb7 100644
> > --- a/mm/hugetlb_cma.c
> > +++ b/mm/hugetlb_cma.c
> > @@ -34,7 +34,10 @@ struct folio *hugetlb_cma_alloc_frozen_folio(int order, gfp_t gfp_mask,
> > if (!hugetlb_cma_size)
> > return NULL;
> >
> > - if (hugetlb_cma[nid])
> > + if (!nodemask)
> > + nodemask = &node_states[N_MEMORY];
>
> Hi Sourav,
>
> hmm what if cpuset.mems only allows allocation from node 0, and hugetlb CMA
> is only available on node 1. This will no now allow allocating a gigantic page
> on node 1 when its explicitly not allowed?
>
> Would it not be better to check cpuset as well?
>
> Thanks,
> Usama
>
Hi Usama,
In v2 and v3 of this patch series, we actually used
&cpuset_current_mems_allowed (adding a seqcount retry loop in v3 to
safely read it without torn reads). However, during the v3 review,
concerns about using cpuset_current_mems_allowed, noting potential
false-positive allocation issues.
- v3: https://lore.kernel.org/linux-mm/20260705175119.440599-1-souravpanda@google.com/
- v2: https://lore.kernel.org/linux-mm/20260704174930.2885785-1-souravpanda@google.com/
- v1: https://lore.kernel.org/linux-mm/20260702215713.627941-1-souravpanda@google.com/
Let me know if you prefer I add it back!
Thanks!
Sourav
> > +
> > + if (hugetlb_cma[nid] && node_isset(nid, *nodemask))
> > page = cma_alloc_frozen_compound(hugetlb_cma[nid], order);
> >
> > if (!page && !(gfp_mask & __GFP_THISNODE)) {
> > --
> > 2.55.0
> >
^ permalink raw reply [flat|nested] 11+ messages in thread* Re: [PATCH v4] mm/hugetlb_cma: Fix null nodemask dereference in hugetlb_cma_alloc_frozen_folio
2026-08-05 20:24 ` Usama Arif
2026-08-05 23:05 ` Sourav Panda
@ 2026-08-06 2:47 ` Anshuman Khandual
2026-08-06 5:33 ` Sourav Panda
1 sibling, 1 reply; 11+ messages in thread
From: Anshuman Khandual @ 2026-08-06 2:47 UTC (permalink / raw)
To: Usama Arif
Cc: Sourav Panda, muchun.song, osalvador, akpm, david, surenb, fvdl,
gthelen, linux-mm, linux-kernel
On Wed, Aug 05, 2026 at 01:24:42PM -0700, Usama Arif wrote:
> On Sun, 26 Jul 2026 07:29:34 +0000 Sourav Panda <souravpanda@google.com> wrote:
>
> > alloc_buddy_hugetlb_folio_with_mpol() can pass a NULL nodemask to
> > alloc_fresh_hugetlb_folio() as a fallback to allocate from all
> > nodes. If order is gigantic, alloc_fresh_hugetlb_folio() propagates
> > the NULL nodemask down to hugetlb_cma_alloc_frozen_folio() via
> > alloc_gigantic_frozen_folio().
> >
> > hugetlb_cma_alloc_frozen_folio() blindly dereferences the nodemask in
> > node_isset(nid, *nodemask) and for_each_node_mask(node, *nodemask),
> > leading to a null pointer dereference kernel panic.
> >
> > Fix this by checking if nodemask is NULL in
> > hugetlb_cma_alloc_frozen_folio() and defaulting it to
> > node_states[N_MEMORY]. This allows hugetlb_cma allocations to fall
> > back to any node with memory, keeping behavior consistent with
> > alloc_contig_frozen_pages() and alloc_buddy_frozen_folio().
> >
> > >From a userspace perspective, this bug allows an unprivileged user to
> > crash the kernel (trigger a panic) by requesting a gigantic hugepage
> > allocation with MPOL_PREFERRED_MANY on a system where CMA is only
> > configured on a subset of NUMA nodes.
> >
> > This can be reproduced by booting a VM with two NUMA nodes, restricting
> > CMA to Node 1 (e.g., hugetlb_cma=1:1G default_hugepagesz=1G
> > hugepagesz=1G hugepages=0), and running a program that allocates a
> > 1GB hugepage area without reserving, restricts allocation to Node 0
> > using mbind() with MPOL_PREFERRED_MANY, and triggers a page fault:
> >
> > void *ptr = mmap(NULL, 1UL << 30, PROT_READ | PROT_WRITE,
> > MAP_PRIVATE | MAP_ANONYMOUS | MAP_HUGETLB |
> > MAP_HUGE_1GB | MAP_NORESERVE, -1, 0);
> > unsigned long nodemask = 1; /* Node 0 */
> > mbind(ptr, 1UL << 30, MPOL_PREFERRED_MANY, &nodemask,
> > sizeof(nodemask) * 8, 0);
> > memset(ptr, 0, 1UL << 30); /* Trigger fault */
> >
> > This results in a NULL pointer dereference:
> >
> > BUG: kernel NULL pointer dereference, address: 0000000000000000
> > #PF: supervisor read access in kernel mode
> > #PF: error_code(0x0000) - not-present page
> > Oops: Oops: 0000 [#1] SMP NOPTI
> > RIP: 0010:hugetlb_cma_alloc_frozen_folio+0x75/0x120
> > Call Trace:
> > <TASK>
> > only_alloc_fresh_hugetlb_folio.isra.0+0x2c/0x160
> > alloc_surplus_hugetlb_folio+0x6d/0x100
> > alloc_hugetlb_folio+0x3c5/0x660
> > hugetlb_no_page+0x3d9/0x650
> >
> > Fixes: eb02f14c4a2b ("mm/hugetlb: allow overcommitting gigantic hugepages")
> > Cc: stable@vger.kernel.org
> > Signed-off-by: Sourav Panda <souravpanda@google.com>
> > ---
> > Changes in v4:
> > - Reverted the alloc_fresh_hugetlb_folio() cpuset snapshot approach from v3.
> > As Muchun Song pointed out, snapshotting cpuset_current_mems_allowed does
> > not prevent false-positive allocation failures without complex retry loops,
> > and alloc_contig_frozen_pages() / alloc_buddy_frozen_folio() already handle
> > NULL nodemasks safely internally.
> > - Handled NULL nodemask directly inside hugetlb_cma_alloc_frozen_folio()
> > by defaulting nodemask to node_states[N_MEMORY] (Option 2), keeping
> > HugeTLB allocators clean and consistent.
> > - v3: https://lore.kernel.org/linux-mm/20260705175119.440599-1-souravpanda@google.com/
> > - v2: https://lore.kernel.org/linux-mm/20260704174930.2885785-1-souravpanda@google.com/
> > - v1: https://lore.kernel.org/linux-mm/20260702215713.627941-1-souravpanda@google.com/
> >
> > mm/hugetlb_cma.c | 5 ++++-
> > 1 file changed, 4 insertions(+), 1 deletion(-)
> >
> > diff --git a/mm/hugetlb_cma.c b/mm/hugetlb_cma.c
> > index 39344d6c78d8..5744de0ceeb7 100644
> > --- a/mm/hugetlb_cma.c
> > +++ b/mm/hugetlb_cma.c
> > @@ -34,7 +34,10 @@ struct folio *hugetlb_cma_alloc_frozen_folio(int order, gfp_t gfp_mask,
> > if (!hugetlb_cma_size)
> > return NULL;
> >
> > - if (hugetlb_cma[nid])
> > + if (!nodemask)
> > + nodemask = &node_states[N_MEMORY];
>
> Hi Sourav,
>
> hmm what if cpuset.mems only allows allocation from node 0, and hugetlb CMA
> is only available on node 1. This will no now allow allocating a gigantic page
> on node 1 when its explicitly not allowed?
That's fair point but should not the user be also responsible in provding a right
nodemask containing CMA memory if it prefers avoiding node_states[N_MEMORY] based
fallback mechanism in kernel ?
>
> Would it not be better to check cpuset as well?
>
> Thanks,
> Usama
>
> > +
> > + if (hugetlb_cma[nid] && node_isset(nid, *nodemask))
> > page = cma_alloc_frozen_compound(hugetlb_cma[nid], order);
> >
> > if (!page && !(gfp_mask & __GFP_THISNODE)) {
> > --
> > 2.55.0
> >
^ permalink raw reply [flat|nested] 11+ messages in thread* Re: [PATCH v4] mm/hugetlb_cma: Fix null nodemask dereference in hugetlb_cma_alloc_frozen_folio
2026-08-06 2:47 ` Anshuman Khandual
@ 2026-08-06 5:33 ` Sourav Panda
2026-08-06 11:07 ` Usama Arif
0 siblings, 1 reply; 11+ messages in thread
From: Sourav Panda @ 2026-08-06 5:33 UTC (permalink / raw)
To: Anshuman Khandual
Cc: Usama Arif, muchun.song, osalvador, akpm, david, surenb, fvdl,
gthelen, linux-mm, linux-kernel
On Wed, Aug 5, 2026 at 7:47 PM Anshuman Khandual
<anshuman.khandual@arm.com> wrote:
>
> On Wed, Aug 05, 2026 at 01:24:42PM -0700, Usama Arif wrote:
> > On Sun, 26 Jul 2026 07:29:34 +0000 Sourav Panda <souravpanda@google.com> wrote:
> >
> > > alloc_buddy_hugetlb_folio_with_mpol() can pass a NULL nodemask to
> > > alloc_fresh_hugetlb_folio() as a fallback to allocate from all
> > > nodes. If order is gigantic, alloc_fresh_hugetlb_folio() propagates
> > > the NULL nodemask down to hugetlb_cma_alloc_frozen_folio() via
> > > alloc_gigantic_frozen_folio().
> > >
> > > hugetlb_cma_alloc_frozen_folio() blindly dereferences the nodemask in
> > > node_isset(nid, *nodemask) and for_each_node_mask(node, *nodemask),
> > > leading to a null pointer dereference kernel panic.
> > >
> > > Fix this by checking if nodemask is NULL in
> > > hugetlb_cma_alloc_frozen_folio() and defaulting it to
> > > node_states[N_MEMORY]. This allows hugetlb_cma allocations to fall
> > > back to any node with memory, keeping behavior consistent with
> > > alloc_contig_frozen_pages() and alloc_buddy_frozen_folio().
> > >
> > > >From a userspace perspective, this bug allows an unprivileged user to
> > > crash the kernel (trigger a panic) by requesting a gigantic hugepage
> > > allocation with MPOL_PREFERRED_MANY on a system where CMA is only
> > > configured on a subset of NUMA nodes.
> > >
> > > This can be reproduced by booting a VM with two NUMA nodes, restricting
> > > CMA to Node 1 (e.g., hugetlb_cma=1:1G default_hugepagesz=1G
> > > hugepagesz=1G hugepages=0), and running a program that allocates a
> > > 1GB hugepage area without reserving, restricts allocation to Node 0
> > > using mbind() with MPOL_PREFERRED_MANY, and triggers a page fault:
> > >
> > > void *ptr = mmap(NULL, 1UL << 30, PROT_READ | PROT_WRITE,
> > > MAP_PRIVATE | MAP_ANONYMOUS | MAP_HUGETLB |
> > > MAP_HUGE_1GB | MAP_NORESERVE, -1, 0);
> > > unsigned long nodemask = 1; /* Node 0 */
> > > mbind(ptr, 1UL << 30, MPOL_PREFERRED_MANY, &nodemask,
> > > sizeof(nodemask) * 8, 0);
> > > memset(ptr, 0, 1UL << 30); /* Trigger fault */
> > >
> > > This results in a NULL pointer dereference:
> > >
> > > BUG: kernel NULL pointer dereference, address: 0000000000000000
> > > #PF: supervisor read access in kernel mode
> > > #PF: error_code(0x0000) - not-present page
> > > Oops: Oops: 0000 [#1] SMP NOPTI
> > > RIP: 0010:hugetlb_cma_alloc_frozen_folio+0x75/0x120
> > > Call Trace:
> > > <TASK>
> > > only_alloc_fresh_hugetlb_folio.isra.0+0x2c/0x160
> > > alloc_surplus_hugetlb_folio+0x6d/0x100
> > > alloc_hugetlb_folio+0x3c5/0x660
> > > hugetlb_no_page+0x3d9/0x650
> > >
> > > Fixes: eb02f14c4a2b ("mm/hugetlb: allow overcommitting gigantic hugepages")
> > > Cc: stable@vger.kernel.org
> > > Signed-off-by: Sourav Panda <souravpanda@google.com>
> > > ---
> > > Changes in v4:
> > > - Reverted the alloc_fresh_hugetlb_folio() cpuset snapshot approach from v3.
> > > As Muchun Song pointed out, snapshotting cpuset_current_mems_allowed does
> > > not prevent false-positive allocation failures without complex retry loops,
> > > and alloc_contig_frozen_pages() / alloc_buddy_frozen_folio() already handle
> > > NULL nodemasks safely internally.
> > > - Handled NULL nodemask directly inside hugetlb_cma_alloc_frozen_folio()
> > > by defaulting nodemask to node_states[N_MEMORY] (Option 2), keeping
> > > HugeTLB allocators clean and consistent.
> > > - v3: https://lore.kernel.org/linux-mm/20260705175119.440599-1-souravpanda@google.com/
> > > - v2: https://lore.kernel.org/linux-mm/20260704174930.2885785-1-souravpanda@google.com/
> > > - v1: https://lore.kernel.org/linux-mm/20260702215713.627941-1-souravpanda@google.com/
> > >
> > > mm/hugetlb_cma.c | 5 ++++-
> > > 1 file changed, 4 insertions(+), 1 deletion(-)
> > >
> > > diff --git a/mm/hugetlb_cma.c b/mm/hugetlb_cma.c
> > > index 39344d6c78d8..5744de0ceeb7 100644
> > > --- a/mm/hugetlb_cma.c
> > > +++ b/mm/hugetlb_cma.c
> > > @@ -34,7 +34,10 @@ struct folio *hugetlb_cma_alloc_frozen_folio(int order, gfp_t gfp_mask,
> > > if (!hugetlb_cma_size)
> > > return NULL;
> > >
> > > - if (hugetlb_cma[nid])
> > > + if (!nodemask)
> > > + nodemask = &node_states[N_MEMORY];
> >
> > Hi Sourav,
> >
> > hmm what if cpuset.mems only allows allocation from node 0, and hugetlb CMA
> > is only available on node 1. This will no now allow allocating a gigantic page
> > on node 1 when its explicitly not allowed?
>
> That's fair point but should not the user be also responsible in provding a right
> nodemask containing CMA memory if it prefers avoiding node_states[N_MEMORY] based
> fallback mechanism in kernel ?
>
We have precedence for three options, which makes selecting one difficult :)
1) nodemask = &cpuset_current_mems_allowed is currently being used in:
only_alloc_fresh_hugetlb_folio --> alloc_buddy_frozen_folio -->
__alloc_frozen_pages_noprof --> prepare_alloc_pages.
2) Some places use cookies with cpuset_current_mems_allowed to prevent
torn writes. But this adds complexity.
An example would be dequeue_hugetlb_folio_nodemask()
3) node_states[N_MEMORY] is sprinkled all over the kernel (simplest).
> >
> > Would it not be better to check cpuset as well?
> >
> > Thanks,
> > Usama
> >
> > > +
> > > + if (hugetlb_cma[nid] && node_isset(nid, *nodemask))
> > > page = cma_alloc_frozen_compound(hugetlb_cma[nid], order);
> > >
> > > if (!page && !(gfp_mask & __GFP_THISNODE)) {
> > > --
> > > 2.55.0
> > >
^ permalink raw reply [flat|nested] 11+ messages in thread* Re: [PATCH v4] mm/hugetlb_cma: Fix null nodemask dereference in hugetlb_cma_alloc_frozen_folio
2026-08-06 5:33 ` Sourav Panda
@ 2026-08-06 11:07 ` Usama Arif
0 siblings, 0 replies; 11+ messages in thread
From: Usama Arif @ 2026-08-06 11:07 UTC (permalink / raw)
To: Sourav Panda, Anshuman Khandual
Cc: muchun.song, osalvador, akpm, david, surenb, fvdl, gthelen,
linux-mm, linux-kernel
On 06/08/2026 06:33, Sourav Panda wrote:
> On Wed, Aug 5, 2026 at 7:47 PM Anshuman Khandual
> <anshuman.khandual@arm.com> wrote:
>>
>> On Wed, Aug 05, 2026 at 01:24:42PM -0700, Usama Arif wrote:
>>> On Sun, 26 Jul 2026 07:29:34 +0000 Sourav Panda <souravpanda@google.com> wrote:
>>>
>>>> alloc_buddy_hugetlb_folio_with_mpol() can pass a NULL nodemask to
>>>> alloc_fresh_hugetlb_folio() as a fallback to allocate from all
>>>> nodes. If order is gigantic, alloc_fresh_hugetlb_folio() propagates
>>>> the NULL nodemask down to hugetlb_cma_alloc_frozen_folio() via
>>>> alloc_gigantic_frozen_folio().
>>>>
>>>> hugetlb_cma_alloc_frozen_folio() blindly dereferences the nodemask in
>>>> node_isset(nid, *nodemask) and for_each_node_mask(node, *nodemask),
>>>> leading to a null pointer dereference kernel panic.
>>>>
>>>> Fix this by checking if nodemask is NULL in
>>>> hugetlb_cma_alloc_frozen_folio() and defaulting it to
>>>> node_states[N_MEMORY]. This allows hugetlb_cma allocations to fall
>>>> back to any node with memory, keeping behavior consistent with
>>>> alloc_contig_frozen_pages() and alloc_buddy_frozen_folio().
>>>>
>>>> >From a userspace perspective, this bug allows an unprivileged user to
>>>> crash the kernel (trigger a panic) by requesting a gigantic hugepage
>>>> allocation with MPOL_PREFERRED_MANY on a system where CMA is only
>>>> configured on a subset of NUMA nodes.
>>>>
>>>> This can be reproduced by booting a VM with two NUMA nodes, restricting
>>>> CMA to Node 1 (e.g., hugetlb_cma=1:1G default_hugepagesz=1G
>>>> hugepagesz=1G hugepages=0), and running a program that allocates a
>>>> 1GB hugepage area without reserving, restricts allocation to Node 0
>>>> using mbind() with MPOL_PREFERRED_MANY, and triggers a page fault:
>>>>
>>>> void *ptr = mmap(NULL, 1UL << 30, PROT_READ | PROT_WRITE,
>>>> MAP_PRIVATE | MAP_ANONYMOUS | MAP_HUGETLB |
>>>> MAP_HUGE_1GB | MAP_NORESERVE, -1, 0);
>>>> unsigned long nodemask = 1; /* Node 0 */
>>>> mbind(ptr, 1UL << 30, MPOL_PREFERRED_MANY, &nodemask,
>>>> sizeof(nodemask) * 8, 0);
>>>> memset(ptr, 0, 1UL << 30); /* Trigger fault */
>>>>
>>>> This results in a NULL pointer dereference:
>>>>
>>>> BUG: kernel NULL pointer dereference, address: 0000000000000000
>>>> #PF: supervisor read access in kernel mode
>>>> #PF: error_code(0x0000) - not-present page
>>>> Oops: Oops: 0000 [#1] SMP NOPTI
>>>> RIP: 0010:hugetlb_cma_alloc_frozen_folio+0x75/0x120
>>>> Call Trace:
>>>> <TASK>
>>>> only_alloc_fresh_hugetlb_folio.isra.0+0x2c/0x160
>>>> alloc_surplus_hugetlb_folio+0x6d/0x100
>>>> alloc_hugetlb_folio+0x3c5/0x660
>>>> hugetlb_no_page+0x3d9/0x650
>>>>
>>>> Fixes: eb02f14c4a2b ("mm/hugetlb: allow overcommitting gigantic hugepages")
>>>> Cc: stable@vger.kernel.org
>>>> Signed-off-by: Sourav Panda <souravpanda@google.com>
>>>> ---
>>>> Changes in v4:
>>>> - Reverted the alloc_fresh_hugetlb_folio() cpuset snapshot approach from v3.
>>>> As Muchun Song pointed out, snapshotting cpuset_current_mems_allowed does
>>>> not prevent false-positive allocation failures without complex retry loops,
>>>> and alloc_contig_frozen_pages() / alloc_buddy_frozen_folio() already handle
>>>> NULL nodemasks safely internally.
>>>> - Handled NULL nodemask directly inside hugetlb_cma_alloc_frozen_folio()
>>>> by defaulting nodemask to node_states[N_MEMORY] (Option 2), keeping
>>>> HugeTLB allocators clean and consistent.
>>>> - v3: https://lore.kernel.org/linux-mm/20260705175119.440599-1-souravpanda@google.com/
>>>> - v2: https://lore.kernel.org/linux-mm/20260704174930.2885785-1-souravpanda@google.com/
>>>> - v1: https://lore.kernel.org/linux-mm/20260702215713.627941-1-souravpanda@google.com/
>>>>
>>>> mm/hugetlb_cma.c | 5 ++++-
>>>> 1 file changed, 4 insertions(+), 1 deletion(-)
>>>>
>>>> diff --git a/mm/hugetlb_cma.c b/mm/hugetlb_cma.c
>>>> index 39344d6c78d8..5744de0ceeb7 100644
>>>> --- a/mm/hugetlb_cma.c
>>>> +++ b/mm/hugetlb_cma.c
>>>> @@ -34,7 +34,10 @@ struct folio *hugetlb_cma_alloc_frozen_folio(int order, gfp_t gfp_mask,
>>>> if (!hugetlb_cma_size)
>>>> return NULL;
>>>>
>>>> - if (hugetlb_cma[nid])
>>>> + if (!nodemask)
>>>> + nodemask = &node_states[N_MEMORY];
>>>
>>> Hi Sourav,
>>>
>>> hmm what if cpuset.mems only allows allocation from node 0, and hugetlb CMA
>>> is only available on node 1. This will no now allow allocating a gigantic page
>>> on node 1 when its explicitly not allowed?
>>
>> That's fair point but should not the user be also responsible in provding a right
>> nodemask containing CMA memory if it prefers avoiding node_states[N_MEMORY] based
>> fallback mechanism in kernel ?
>>
I don't think the task can be made responsible for providing a CMA-aware fallback mask
here. Userspace supplies the MPOL_PREFERRED_MANY preference, but the kernel itself
changes the second attempt to a NULL nodemask. That permits fallback outside the
preferred policy nodes; it does not permit fallback outside the task's hardwall cpuset.
>
> We have precedence for three options, which makes selecting one difficult :)
>
> 1) nodemask = &cpuset_current_mems_allowed is currently being used in:
> only_alloc_fresh_hugetlb_folio --> alloc_buddy_frozen_folio -->
> __alloc_frozen_pages_noprof --> prepare_alloc_pages.
Here I think using &cpuset_current_mems_allowed is part of the page allocator's wider cpuset handling.
> 2) Some places use cookies with cpuset_current_mems_allowed to prevent
> torn writes. But this adds complexity.
> An example would be dequeue_hugetlb_folio_nodemask()
> 3) node_states[N_MEMORY] is sprinkled all over the kernel (simplest).
>
I would prefer option 2. A NULL mempolicy mask should fall back to
cpuset_current_mems_allowed, not all N_MEMORY nodes.
>>>
>>> Would it not be better to check cpuset as well?
>>>
>>> Thanks,
>>> Usama
>>>
>>>> +
>>>> + if (hugetlb_cma[nid] && node_isset(nid, *nodemask))
>>>> page = cma_alloc_frozen_compound(hugetlb_cma[nid], order);
>>>>
>>>> if (!page && !(gfp_mask & __GFP_THISNODE)) {
>>>> --
>>>> 2.55.0
>>>>
^ permalink raw reply [flat|nested] 11+ messages in thread