From: Yosry Ahmed <yosry@kernel.org>
To: Brendan Jackman <jackmanb@google.com>
Cc: Borislav Petkov <bp@alien8.de>,
Dave Hansen <dave.hansen@linux.intel.com>,
Peter Zijlstra <peterz@infradead.org>,
Andrew Morton <akpm@linux-foundation.org>,
David Hildenbrand <david@kernel.org>,
Vlastimil Babka <vbabka@kernel.org>,
Mike Rapoport <rppt@kernel.org>, Wei Xu <weixugc@google.com>,
Johannes Weiner <hannes@cmpxchg.org>, Zi Yan <ziy@nvidia.com>,
Lorenzo Stoakes <ljs@kernel.org>,
linux-mm@kvack.org, linux-kernel@vger.kernel.org,
x86@kernel.org, Sumit Garg <sumit.garg@oss.qualcomm.com>,
Will Deacon <will@kernel.org>,
rientjes@google.com, "Kalyazin, Nikita" <kalyazin@amazon.co.uk>,
patrick.roy@linux.dev, "Itazuri, Takahiro" <itazur@amazon.co.uk>,
Andy Lutomirski <luto@kernel.org>,
David Kaplan <david.kaplan@amd.com>,
Thomas Gleixner <tglx@kernel.org>,
Patrick Bellasi <derkling@google.com>,
Reiji Watanabe <reijiw@google.com>,
Sean Christopherson <seanjc@google.com>
Subject: Re: [PATCH v3 21/26] mm/page_alloc: implement FREETYPE_UNMAPPED allocations
Date: Tue, 4 Aug 2026 23:41:36 +0000 [thread overview]
Message-ID: <anJylMTl6GeM8OHx@google.com> (raw)
In-Reply-To: <20260726-page_alloc-unmapped-v3-21-6f5729aa9832@google.com>
On Sun, Jul 26, 2026 at 10:22:54PM +0000, Brendan Jackman wrote:
> Currently FREETYPE_UNMAPPED allocs will always fail because, although the
> lists exist to hold them, there is no way to actually create an unmapped
> page block. This commit adds one, and also the logic to map it back
> again when that's needed.
>
> Doing this at pageblock granularity ensures that the pageblock flags can
> be used to infer which freetype a page belongs to. It also provides nice
> batching of TLB flushes, and also avoids creating too much unnecessary
> TLB fragmentation in the physmap.
>
> There are some functional requirements for flipping a block:
>
> - Unmapping requires a TLB shootdown, meaning IRQs must be enabled.
>
> - Updating the pagetables might require allocating a pagetable to break
> down a huge page. This would deadlock if the zone lock was held.
We also need to zero unmapped/sensitive pages before mapping them again,
but seems like the current approach is to differ this to the caller,
which makes sense. The only annoying part is that
want_init_on_{free/alloc}() now silently skip the zeroing for those
pages.
We should document somewhere that the users of ALLOC_UNMAPPED are
responsible for zeroing memory before freeing it?
The current users currently always zero the pages on allocation as well,
I am not sure if this should also be a general requirement, or perhaps
only if want_init_on_alloc() is set?
>
> This makes allocations that need to change sensitivity _somewhat_
s/sensitivity/direct mapping status (or sth)?
> similar to those that need to fallback to a different migratetype. But,
> the locking requirements mean that this can't just be squashed into the
> existing "fallback" allocator logic, instead a new allocator path just
> for this purpose is needed.
>
> The new path is assumed to be much cheaper than the really heavyweight
> stuff like compaction and reclaim. But at present it is treated as less
> desirable than the mobility-related "fallback" and "stealing" logic.
> This might turn out to need revision (in particular, maybe it's a
> problem that __rmqueue_steal(), which causes fragmentation, happens
> before __rmqueue_direct_map()), but that should be treated as a subsequent
> optimisation project.
>
> Adding alloc_flags to gfp_freetype() requires moving it to
> mm/page_alloc.h so it can refer to ALLOC_UNMAPPED. It was already only
> used in internal mm code.
>
> Now that unmapped pageblocks actually exist, exclude them from
> migration. Migrating unmapped pages via the mermap should be possible
> but that's something to be added later when needed.
>
> Signed-off-by: Brendan Jackman <jackmanb@google.com>
[..]
> @@ -3400,6 +3426,127 @@ static inline void zone_statistics(struct zone *preferred_zone, struct zone *z,
> #endif
> }
>
> +#ifdef CONFIG_PAGE_ALLOC_UNMAPPED
> +/* Try to allocate a page by mapping/unmapping a block from the direct map. */
> +static inline struct page *
> +__rmqueue_direct_map(struct zone *zone, unsigned int request_order,
> + unsigned int alloc_flags, freetype_t freetype)
> +{
> + unsigned int ft_flags_other = freetype_flags(freetype) ^ FREETYPE_UNMAPPED;
> + freetype_t ft_other = migrate_to_freetype(free_to_migratetype(freetype),
> + ft_flags_other);
> + bool want_mapped = !(freetype_flags(freetype) & FREETYPE_UNMAPPED);
> + enum rmqueue_mode rmqm = RMQUEUE_NORMAL;
> + unsigned long irq_flags;
> + int nr_pageblocks, nr_freed;
> + struct page *page;
> + int alloc_order;
> + int err;
> +
> + if (freetype_idx(ft_other) < 0)
> + return NULL;
> +
> + /*
> + * Might need a TLB shootdown. Even if IRQs are on this isn't
> + * safe if the caller holds a lock (in case the other CPUs need that
> + * lock to handle the shootdown IPI).
> + */
> + if (alloc_flags & ALLOC_NOBLOCK)
> + return NULL;
Should we only check this if !want_mapped? IIUC we only need a TLB
shootdown when unmapping.
> +
> + if (!can_set_direct_map() || alloc_flags & ALLOC_NOLOCK)
> + return NULL;
> +
> + lockdep_assert(!irqs_disabled() || unlikely(early_boot_irqs_disabled));
> +
> + /*
> + * Need to [un]map a whole pageblock (otherwise it might require
> + * allocating pagetables). First allocate it.
> + */
> + alloc_order = max(request_order, pageblock_order);
> + nr_pageblocks = 1 << (alloc_order - pageblock_order);
> + spin_lock_irqsave(&zone->lock, irq_flags);
> + /* First try a block that already has the right migratetype. */
> + page = __rmqueue(zone, alloc_order, ft_other, alloc_flags, &rmqm);
IIUC, this is called after __rmqueue() will have already failed in the
caller with request_order (a potentially smaller order), so why are we
trying this again here?
> + if (!page) {
> + /* Fallback to changing a block's migratetype. */
> + rmqm = RMQUEUE_CLAIM;
> + page = __rmqueue(zone, alloc_order, ft_other, alloc_flags, &rmqm);
> + }
> + spin_unlock_irqrestore(&zone->lock, irq_flags);
> + if (!page)
> + return NULL;
> +
> + /*
> + * Now that IRQs are on it's safe to do a TLB shootdown, and now that we
> + * released the zone lock it's possible to allocate a pagetable if
> + * needed to split up a huge page.
> + *
> + * Note that modifying the direct map may need to allocate pagetables.
> + * What about unbounded recursion? Here are the assumptions that make it
> + * safe:
> + *
> + * - The direct map starts out fully mapped at boot. (This is not really
> + * an "assumption" as it's in direct control of page_alloc.c).
> + *
> + * - Once pages in the direct map are broken down, they are not
> + * re-aggregated into larger pages again.
> + *
> + * - Pagetables are never allocated with ALLOC_UNMAPPED.
> + *
> + * Under these assumptions, a pagetable might need to be allocated while
> + * _unmapping_ stuff from the direct map during an ALLOC_UNMAPPED
> + * allocation. But, the allocation of that pagetable never requires
> + * allocating a further pagetable.
> + */
> + err = set_direct_map_valid_noflush(page,
> + nr_pageblocks << pageblock_order, want_mapped);
> + if (err == -ENOMEM || WARN_ONCE(err, "err=%d\n", err)) {
> + set_direct_map_valid_noflush(page,
> + nr_pageblocks << pageblock_order, !want_mapped);
> + spin_lock_irqsave(&zone->lock, irq_flags);
> + /* Important: free using _old_ freetype. */
> + __free_one_page(page, page_to_pfn(page), zone,
> + alloc_order, ft_other, FPI_SKIP_REPORT_NOTIFY);
> + spin_unlock_irqrestore(&zone->lock, irq_flags);
> + return NULL;
> + }
> +
> + if (want_mapped) {
> + /* Exposing formerly-protected data; scrub it. */
> + clear_highpages_kasan_tagged(page, nr_pageblocks << pageblock_order);
Shouldn't all unmapped memory be zeroed on free? If we solidify this
assumption we can probably drop this here (and maybe replace it with an
assertion)?
> + } else {
> + unsigned long start = (unsigned long)page_address(page);
> + unsigned long end = start + (nr_pageblocks << (pageblock_order + PAGE_SHIFT));
> +
> + flush_tlb_kernel_range(start, end);
> + }
> +
> + for (int i = 0; i < nr_pageblocks; i++) {
> + struct page *block_page = page + (pageblock_nr_pages * i);
> +
> + set_pageblock_freetype_flags(block_page, freetype_flags(freetype));
> + }
> +
> + if (request_order >= alloc_order)
> + return page;
> +
> + /* Free any remaining pages in the block. */
> + spin_lock_irqsave(&zone->lock, irq_flags);
> + nr_freed = expand(zone, page, request_order, alloc_order, freetype);
> + account_freepages(zone, nr_freed, free_to_migratetype(freetype));
> + spin_unlock_irqrestore(&zone->lock, irq_flags);
> +
> + return page;
> +}
> +#else /* CONFIG_PAGE_ALLOC_UNMAPPED */
> +static inline struct page *__rmqueue_direct_map(struct zone *zone, unsigned int request_order,
> + unsigned int alloc_flags, freetype_t freetype)
> +{
> + return NULL;
> +}
> +#endif /* CONFIG_PAGE_ALLOC_UNMAPPED */
> +
> static __always_inline
> struct page *rmqueue_buddy(struct zone *preferred_zone, struct zone *zone,
> unsigned int order, unsigned int alloc_flags,
> @@ -3433,13 +3580,15 @@ struct page *rmqueue_buddy(struct zone *preferred_zone, struct zone *zone,
> */
> if (!page && (alloc_flags & (ALLOC_OOM|ALLOC_HARDER)))
> page = __rmqueue_smallest(zone, order, ft_high);
> -
> - if (!page) {
> - spin_unlock_irqrestore(&zone->lock, flags);
> - return NULL;
> - }
> }
> spin_unlock_irqrestore(&zone->lock, flags);
> +
> + /* Try changing direct map, now we've released the zone lock */
> + if (!page)
> + page = __rmqueue_direct_map(zone, order, alloc_flags, freetype);
> + if (!page)
> + return NULL;
> +
> } while (check_new_pages(page, order));
>
> /*
> @@ -3660,6 +3809,8 @@ static void reserve_highatomic_pageblock(struct page *page, int order,
> return;
>
> ft_high = freetype_with_migrate(ft, MIGRATE_HIGHATOMIC);
> + if (freetype_idx(ft_high) < 0)
> + return;
Does this belong in "mm/page_alloc: add support for freetypes with no
freelist"?
> if (order < pageblock_order) {
> if (move_freepages_block(zone, page, ft, ft_high) == -1)
> return;
> @@ -3975,13 +4126,15 @@ alloc_flags_nofragment(struct zone *zone, gfp_t gfp_mask)
> }
>
> /* Must be called after current_gfp_context() which can change gfp_mask */
> -static inline unsigned int alloc_flags_cma(gfp_t gfp_mask)
> +static inline unsigned int alloc_flags_cma(gfp_t gfp_mask, unsigned int alloc_flags)
> {
> #ifdef CONFIG_CMA
> - if (free_to_migratetype(gfp_freetype(gfp_mask)) == MIGRATE_MOVABLE)
> - return ALLOC_CMA;
> + if (free_to_migratetype(gfp_freetype(gfp_mask, alloc_flags)) == MIGRATE_MOVABLE)
> + alloc_flags |= ALLOC_CMA;
> #endif
> - return ALLOC_DEFAULT;
> + alloc_flags |= ALLOC_DEFAULT;
> +
> + return alloc_flags;
> }
>
> /*
> @@ -4770,7 +4923,7 @@ alloc_flags_slowpath(gfp_t gfp_mask, unsigned int order)
> } else if (unlikely(rt_or_dl_task(current)) && in_task())
> alloc_flags |= ALLOC_MIN_RESERVE;
>
> - alloc_flags |= alloc_flags_cma(gfp_mask);
> + alloc_flags = alloc_flags_cma(gfp_mask, alloc_flags);
>
> if (defrag_mode)
> alloc_flags |= ALLOC_NOFRAGMENT;
> @@ -5085,7 +5238,7 @@ __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order,
>
> reserve_flags = __gfp_pfmemalloc_flags(gfp_mask);
> if (reserve_flags)
> - alloc_flags = alloc_flags_cma(gfp_mask) | reserve_flags |
> + alloc_flags = alloc_flags_cma(gfp_mask, alloc_flags) | reserve_flags |
> ac->alloc_flags | (alloc_flags & ALLOC_KSWAPD);
Should we pass in ac->alloc_flags here to maintain equivalent
functionality, maybe this:
alloc_flags = alloc_flags_cma(gfp_mask, ac->alloc_flags) |
reserve_flags | (alloc_flags & ALLOC_KSWAPD);
>
> /*
> @@ -5307,7 +5460,11 @@ static inline bool prepare_alloc_pages(gfp_t gfp_mask, unsigned int order,
> ac->highest_zoneidx = gfp_zone(gfp_mask);
> ac->zonelist = node_zonelist(preferred_nid, gfp_mask);
> ac->nodemask = nodemask;
> - ac->freetype = gfp_freetype(gfp_mask);
> + ac->freetype = gfp_freetype(gfp_mask, *alloc_flags);
> +
> + /* Not implemented yet. */
> + if (freetype_flags(ac->freetype) & FREETYPE_UNMAPPED && gfp_mask & __GFP_ZERO)
Nit: Add more parentheses for readability?
if ((freetype_flags(ac->freetype) & FREETYPE_UNMAPPED) && (gfp_mask & __GFP_ZERO))
> + return false;
>
> if (cpusets_enabled()) {
> *alloc_gfp |= __GFP_HARDWALL;
next prev parent reply other threads:[~2026-08-04 23:41 UTC|newest]
Thread overview: 61+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-26 22:22 [PATCH v3 00/26] mm: Add ALLOC_UNMAPPED and AS_NO_DIRECT_MAP Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 01/26] set_memory: add folio_{zap,restore}_direct_map helpers Brendan Jackman
2026-07-27 10:33 ` Mike Rapoport
2026-07-29 11:42 ` Brendan Jackman
2026-07-30 20:34 ` Yosry Ahmed
2026-07-31 5:21 ` Mike Rapoport
2026-07-31 11:57 ` Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 02/26] mm/secretmem: make use of folio_{zap,restore}_direct_map Brendan Jackman
2026-07-27 10:40 ` Mike Rapoport
2026-07-26 22:22 ` [PATCH v3 03/26] mm: introduce AS_NO_DIRECT_MAP Brendan Jackman
2026-07-30 21:06 ` Yosry Ahmed
2026-07-31 12:15 ` Brendan Jackman
2026-07-31 19:28 ` Yosry Ahmed
2026-08-02 16:10 ` Mike Rapoport
2026-07-26 22:22 ` [PATCH v3 04/26] x86/mm: split out preallocate_sub_pgd() Brendan Jackman
2026-07-31 22:10 ` Yosry Ahmed
2026-08-02 16:13 ` Mike Rapoport
2026-07-26 22:22 ` [PATCH v3 05/26] x86: move PAE PMD preallocation defines to header Brendan Jackman
2026-07-31 23:59 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 06/26] x86/tlb: Expose some flush function declarations to modules Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 07/26] x86/mm: introduce mm-local region Brendan Jackman
2026-08-02 16:27 ` Mike Rapoport
2026-08-03 22:29 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 08/26] x86/mm: move LDT remap into " Brendan Jackman
2026-08-03 22:33 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 09/26] mm: Create flags arg for __apply_to_page_range() Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 10/26] mm: Add more flags " Brendan Jackman
2026-08-04 0:08 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 11/26] x86/mm: introduce the mermap Brendan Jackman
2026-08-02 16:40 ` Mike Rapoport
2026-08-04 18:38 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 12/26] mm: KUnit tests for " Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 13/26] mm: introduce freetype_t Brendan Jackman
2026-08-04 22:23 ` Yosry Ahmed
2026-08-04 23:02 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 14/26] mm: move migratetype definitions to freetype.h Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 15/26] mm/page_alloc: add support for freetypes with no freelist Brendan Jackman
2026-07-31 14:13 ` Vlastimil Babka (SUSE)
2026-07-26 22:22 ` [PATCH v3 16/26] mm: add definitions for allocating unmapped pages Brendan Jackman
2026-08-04 19:53 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 17/26] mm: encode freetype flags in pageblock flags Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 18/26] mm/page_alloc: separate pcplists by freetype flags Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 19/26] mm/page_alloc: rename ALLOC_NON_BLOCK back to _HARDER Brendan Jackman
2026-07-31 14:52 ` Vlastimil Babka (SUSE)
2026-08-03 9:20 ` Vlastimil Babka (SUSE)
2026-08-04 21:50 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 20/26] mm/page_alloc: introduce ALLOC_NOBLOCK Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 21/26] mm/page_alloc: implement FREETYPE_UNMAPPED allocations Brendan Jackman
2026-08-03 9:18 ` Vlastimil Babka (SUSE)
2026-08-04 23:41 ` Yosry Ahmed [this message]
2026-08-04 23:53 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 22/26] mm: Minimal KUnit tests for some new page_alloc logic Brendan Jackman
2026-08-03 9:30 ` Vlastimil Babka (SUSE)
2026-07-26 22:22 ` [PATCH v3 23/26] mm: Split out NR_FREE_PAGES_BLOCKS_[UN]MAPPED Brendan Jackman
2026-08-03 9:32 ` Vlastimil Babka (SUSE)
2026-07-26 22:22 ` [PATCH v3 24/26] mm/page_alloc: always direct compact for unmapped allocs Brendan Jackman
2026-08-03 9:44 ` Vlastimil Babka (SUSE)
2026-07-26 22:22 ` [PATCH v3 25/26] mm: plumb alloc flags into some alloc funcs Brendan Jackman
2026-08-03 9:52 ` Vlastimil Babka (SUSE)
2026-07-26 22:22 ` [PATCH v3 26/26] mm: add fast path for AS_NO_DIRECT_MAP Brendan Jackman
2026-07-29 11:52 ` [PATCH v3 00/26] mm: Add ALLOC_UNMAPPED and AS_NO_DIRECT_MAP Brendan Jackman
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=anJylMTl6GeM8OHx@google.com \
--to=yosry@kernel.org \
--cc=akpm@linux-foundation.org \
--cc=bp@alien8.de \
--cc=dave.hansen@linux.intel.com \
--cc=david.kaplan@amd.com \
--cc=david@kernel.org \
--cc=derkling@google.com \
--cc=hannes@cmpxchg.org \
--cc=itazur@amazon.co.uk \
--cc=jackmanb@google.com \
--cc=kalyazin@amazon.co.uk \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=luto@kernel.org \
--cc=patrick.roy@linux.dev \
--cc=peterz@infradead.org \
--cc=reijiw@google.com \
--cc=rientjes@google.com \
--cc=rppt@kernel.org \
--cc=seanjc@google.com \
--cc=sumit.garg@oss.qualcomm.com \
--cc=tglx@kernel.org \
--cc=vbabka@kernel.org \
--cc=weixugc@google.com \
--cc=will@kernel.org \
--cc=x86@kernel.org \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.