From: Yosry Ahmed <yosry@kernel.org>
To: Brendan Jackman <jackmanb@google.com>
Cc: Borislav Petkov <bp@alien8.de>,
Dave Hansen <dave.hansen@linux.intel.com>,
Peter Zijlstra <peterz@infradead.org>,
Andrew Morton <akpm@linux-foundation.org>,
David Hildenbrand <david@kernel.org>,
Vlastimil Babka <vbabka@kernel.org>,
Mike Rapoport <rppt@kernel.org>, Wei Xu <weixugc@google.com>,
Johannes Weiner <hannes@cmpxchg.org>, Zi Yan <ziy@nvidia.com>,
Lorenzo Stoakes <ljs@kernel.org>,
linux-mm@kvack.org, linux-kernel@vger.kernel.org,
x86@kernel.org, Sumit Garg <sumit.garg@oss.qualcomm.com>,
Will Deacon <will@kernel.org>,
rientjes@google.com, "Kalyazin, Nikita" <kalyazin@amazon.co.uk>,
patrick.roy@linux.dev, "Itazuri, Takahiro" <itazur@amazon.co.uk>,
Andy Lutomirski <luto@kernel.org>,
David Kaplan <david.kaplan@amd.com>,
Thomas Gleixner <tglx@kernel.org>,
Patrick Bellasi <derkling@google.com>,
Reiji Watanabe <reijiw@google.com>,
Sean Christopherson <seanjc@google.com>
Subject: Re: [PATCH v3 07/26] x86/mm: introduce mm-local region
Date: Mon, 3 Aug 2026 22:29:52 +0000 [thread overview]
Message-ID: <anES7O7T_PN2bV0v@google.com> (raw)
In-Reply-To: <20260726-page_alloc-unmapped-v3-7-6f5729aa9832@google.com>
On Sun, Jul 26, 2026 at 10:22:40PM +0000, Brendan Jackman wrote:
> Various security features benefit from having process-local address
> mappings within the kernel. Examples include no-direct-map guest_memfd
> [2] and significant optimizations for ASI [1].
>
> With the currently envisaged usecases, there will be many situations
> where almost no processes have any need for the mm-local region.
> Therefore, avoid its overhead (memory cost of pagetables, alloc/free
> overhead during fork/exit) for processes that don't use it by requiring
> its users to explicitly initialize it via the new mm_local_* API.
>
> As pointed out by Andy in [0], x86 already has a PGD entry that is local
> to the mm, which is used for the LDT. In a subsequent patch, the LDT
> remap will be unified with the general mm-local region, but to help
> keep the patch to a manageable size, first just introduce the mm-local
> region.
>
> On 64-bit, give the mm-local region a whole PGD. On 32-bit, just give it
> one PMD. No investigation has been done into whether it's feasible to
> expand the region on 32-bit. Most likely there is no strong usecase for
> that anyway.
>
> In order to combine the need for an on-demand mm initialisation, with
> the desire to transparently handle propagating mappings to userspace
> under KPTI, the user and kernel pagetables are shared at the highest
> level possible. For PAE that means the PTE table is shared and for
> 64-bit the P4D/PUD. This is implemented by pre-allocating the first
> shared table when the mm-local region is first initialised.
>
> [0] https://lore.kernel.org/linux-mm/CALCETrXHbS9VXfZ80kOjiTrreM2EbapYeGp68mvJPbosUtorYA@mail.gmail.com/
> [1] https://linuxasi.dev/
> [2] https://lore.kernel.org/all/20250924151101.2225820-1-patrick.roy@campus.lmu.de
> Signed-off-by: Brendan Jackman <jackmanb@google.com>
> ---
> arch/x86/Kconfig | 2 +
> arch/x86/include/asm/mmu_context.h | 118 +++++++++++++++++++++++++++++++-
> arch/x86/include/asm/pgtable_32_areas.h | 9 ++-
> arch/x86/include/asm/pgtable_64_types.h | 12 +++-
> arch/x86/mm/pgtable.c | 3 +
> include/linux/mm.h | 13 ++++
> include/linux/mm_types.h | 2 +
> kernel/fork.c | 1 +
> mm/Kconfig | 7 ++
> 9 files changed, 161 insertions(+), 6 deletions(-)
>
> diff --git a/arch/x86/Kconfig b/arch/x86/Kconfig
> index fb298e2191792..3efab3524a6cf 100644
> --- a/arch/x86/Kconfig
> +++ b/arch/x86/Kconfig
> @@ -132,6 +132,8 @@ config X86
> select ARCH_SUPPORTS_LTO_CLANG
> select ARCH_SUPPORTS_LTO_CLANG_THIN
> select ARCH_SUPPORTS_RT
> + # LDT remap temporarily clashes with mm-local region, can't have both.
> + select ARCH_SUPPORTS_MM_LOCAL_REGION if X86_64 || X86_PAE && !MODIFY_LDT_SYSCALL
> select ARCH_USE_BUILTIN_BSWAP
> select ARCH_USE_CMPXCHG_LOCKREF
> select ARCH_USE_MEMTEST
> diff --git a/arch/x86/include/asm/mmu_context.h b/arch/x86/include/asm/mmu_context.h
> index ef5b507de34e2..3d4f54673014f 100644
> --- a/arch/x86/include/asm/mmu_context.h
> +++ b/arch/x86/include/asm/mmu_context.h
> @@ -8,8 +8,10 @@
>
> #include <trace/events/tlb.h>
>
> +#include <asm/tlb.h>
> #include <asm/tlbflush.h>
> #include <asm/paravirt.h>
> +#include <asm/pgalloc.h>
> #include <asm/debugreg.h>
> #include <asm/gsseg.h>
> #include <asm/desc.h>
> @@ -223,10 +225,124 @@ static inline int arch_dup_mmap(struct mm_struct *oldmm, struct mm_struct *mm)
> return ldt_dup_context(oldmm, mm);
> }
>
> +#ifdef CONFIG_MM_LOCAL_REGION
> +static inline void mm_local_region_free(struct mm_struct *mm)
> +{
> + if (!mm_local_region_used(mm))
> + return;
> +
> + struct mmu_gather tlb;
> + unsigned long start = MM_LOCAL_BASE_ADDR;
> + unsigned long end = MM_LOCAL_END_ADDR;
These declarations should probably go at the beginning of the function.
> +
> + /*
> + * Although free_pgd_range() is intended for freeing user
> + * page-tables, it also works out for kernel mappings on x86.
> + * Use tlb_gather_mmu_fullmm() to avoid confusing the
> + * range-tracking logic in __tlb_adjust_range().
> + */
> + tlb_gather_mmu_fullmm(&tlb, mm);
> + free_pgd_range(&tlb, start, end, start, end);
> + tlb_finish_mmu(&tlb);
> +
> + mm_flags_clear(MMF_LOCAL_REGION_USED, mm);
> +}
> +
> +#if defined(CONFIG_MITIGATION_PAGE_TABLE_ISOLATION) && defined(CONFIG_X86_PAE)
Would it be clearer to have nested #ifdefs instead?
#ifdef CONFIG_MITIGATION_PAGE_TABLE_ISOLATION
#ifdef CONFIG_X86_PAE
...
#else /* CONFIG_X86_PAE */
...
#endif /* CONFIG_X86_PAE */
#else /* CONFIG_MITIGATION_PAGE_TABLE_ISOLATION */
#endif /* CONFIG_MITIGATION_PAGE_TABLE_ISOLATION */
Maybe not, just thinking out loud.
> +static inline pmd_t *pgd_to_pmd_walk(pgd_t *pgd, unsigned long va)
> +{
> + p4d_t *p4d;
> + pud_t *pud;
> +
> + if (pgd->pgd == 0)
> + return NULL;
> +
> + p4d = p4d_offset(pgd, va);
> + if (p4d_none(*p4d))
> + return NULL;
> +
> + pud = pud_offset(p4d, va);
> + if (pud_none(*pud))
> + return NULL;
> +
> + return pmd_offset(pud, va);
> +}
> +
> +static inline int mm_local_map_to_user(struct mm_struct *mm)
> +{
> + BUILD_BUG_ON(!PREALLOCATED_PMDS);
> + pgd_t *k_pgd = pgd_offset(mm, MM_LOCAL_BASE_ADDR);
> + pgd_t *u_pgd = kernel_to_user_pgdp(k_pgd);
> + pmd_t *k_pmd, *u_pmd;
> + int err;
> +
> + k_pmd = pgd_to_pmd_walk(k_pgd, MM_LOCAL_BASE_ADDR);
> + u_pmd = pgd_to_pmd_walk(u_pgd, MM_LOCAL_BASE_ADDR);
> +
> + BUILD_BUG_ON(MM_LOCAL_END_ADDR - MM_LOCAL_BASE_ADDR > PMD_SIZE);
> +
> + /* Preallocate the PTE table so it can be shared. */
> + err = pte_alloc(mm, k_pmd);
> + if (err)
> + return err;
> +
> + /* Point the userspace PMD at the same PTE as the kernel PMD. */
> + set_pmd(u_pmd, *k_pmd);
> + return 0;
> +}
> +#elif defined(CONFIG_MITIGATION_PAGE_TABLE_ISOLATION)
> +static inline int mm_local_map_to_user(struct mm_struct *mm)
> +{
> + pgd_t *pgd;
> + int err;
> +
> + err = preallocate_sub_pgd(mm, MM_LOCAL_BASE_ADDR);
> + if (err)
> + return err;
> +
> + pgd = pgd_offset(mm, MM_LOCAL_BASE_ADDR);
> + set_pgd(kernel_to_user_pgdp(pgd), *pgd);
> + return 0;
> +}
The code above bears a lot of similarity to the LDT code removed in
patch 8, and reviewing them separately is annoying. I realize that they
were a single patch in the previous version and Dave complained that it
was too large.
What if we go a different way:
1. Move the LDT functions that will be repurposed to mmu_context.h.
2. Rename the functions to the mm_local_* domain where needed.
3. Actually perform the switch for LDT to use mm local region.
Maybe (2) and (3) should be combined, depending on what the git diff
looks like.
I think this will make the diffs much clearer, for example
mm_local_map_to_user() mainly differ from map_ldt_struct_to_user() in
preallocation.
> +#else
> +static inline int mm_local_map_to_user(struct mm_struct *mm)
> +{
> + WARN_ONCE(1, "mm_local_map_to_user() not implemented");
> + return -EINVAL;
> +}
> +#endif
[..]
> diff --git a/arch/x86/include/asm/pgtable_32_areas.h b/arch/x86/include/asm/pgtable_32_areas.h
> index 921148b429676..7fccb887f8b33 100644
> --- a/arch/x86/include/asm/pgtable_32_areas.h
> +++ b/arch/x86/include/asm/pgtable_32_areas.h
> @@ -30,9 +30,14 @@ extern bool __vmalloc_start_set; /* set once high_memory is set */
> #define CPU_ENTRY_AREA_BASE \
> ((FIXADDR_TOT_START - PAGE_SIZE*(CPU_ENTRY_AREA_PAGES+1)) & PMD_MASK)
>
> -#define LDT_BASE_ADDR \
> - ((CPU_ENTRY_AREA_BASE - PAGE_SIZE) & PMD_MASK)
> +/*
> + * On 32-bit the mm-local region is currently completely consumed by the LDT
> + * remap.
> + */
> +#define MM_LOCAL_BASE_ADDR ((CPU_ENTRY_AREA_BASE - PAGE_SIZE) & PMD_MASK)
> +#define MM_LOCAL_END_ADDR (MM_LOCAL_BASE_ADDR + PMD_SIZE)
>
> +#define LDT_BASE_ADDR MM_LOCAL_BASE_ADDR
> #define LDT_END_ADDR (LDT_BASE_ADDR + PMD_SIZE)
>
> #define PKMAP_BASE \
> diff --git a/arch/x86/include/asm/pgtable_64_types.h b/arch/x86/include/asm/pgtable_64_types.h
> index 7eb61ef6a185f..1181565966405 100644
> --- a/arch/x86/include/asm/pgtable_64_types.h
> +++ b/arch/x86/include/asm/pgtable_64_types.h
> @@ -5,8 +5,11 @@
> #include <asm/sparsemem.h>
>
> #ifndef __ASSEMBLER__
> +#include <linux/build_bug.h>
> #include <linux/types.h>
> #include <asm/kaslr.h>
> +#include <asm/page_types.h>
> +#include <uapi/asm/ldt.h>
>
> /*
> * These are used to make use of C type-checking..
> @@ -100,9 +103,12 @@ extern unsigned int ptrs_per_p4d;
> #define GUARD_HOLE_BASE_ADDR (GUARD_HOLE_PGD_ENTRY << PGDIR_SHIFT)
> #define GUARD_HOLE_END_ADDR (GUARD_HOLE_BASE_ADDR + GUARD_HOLE_SIZE)
>
> -#define LDT_PGD_ENTRY -240UL
> -#define LDT_BASE_ADDR (LDT_PGD_ENTRY << PGDIR_SHIFT)
> -#define LDT_END_ADDR (LDT_BASE_ADDR + PGDIR_SIZE)
> +#define MM_LOCAL_PGD_ENTRY -240UL
> +#define MM_LOCAL_BASE_ADDR (MM_LOCAL_PGD_ENTRY << PGDIR_SHIFT)
> +#define MM_LOCAL_END_ADDR ((MM_LOCAL_PGD_ENTRY + 1) << PGDIR_SHIFT)
Any reason not keep the current formula (i.e. MM_LOCAL_BASE_ADDR +
PGDIR_SIZE)?
> +
> +#define LDT_BASE_ADDR MM_LOCAL_BASE_ADDR
> +#define LDT_END_ADDR (LDT_BASE_ADDR + PMD_SIZE)
Looks like the LDT area was silently changed to PMD_SIZE here. I assume
this is to give the rest of the pgd-mapped address space to the mermap,
but maybe we should call this out explicitly, or do it when the mermap
is introduced (or separately)?
>
> #define __VMALLOC_BASE_L4 0xffffc90000000000UL
> #define __VMALLOC_BASE_L5 0xffa0000000000000UL
[..]
next prev parent reply other threads:[~2026-08-03 22:29 UTC|newest]
Thread overview: 61+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-26 22:22 [PATCH v3 00/26] mm: Add ALLOC_UNMAPPED and AS_NO_DIRECT_MAP Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 01/26] set_memory: add folio_{zap,restore}_direct_map helpers Brendan Jackman
2026-07-27 10:33 ` Mike Rapoport
2026-07-29 11:42 ` Brendan Jackman
2026-07-30 20:34 ` Yosry Ahmed
2026-07-31 5:21 ` Mike Rapoport
2026-07-31 11:57 ` Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 02/26] mm/secretmem: make use of folio_{zap,restore}_direct_map Brendan Jackman
2026-07-27 10:40 ` Mike Rapoport
2026-07-26 22:22 ` [PATCH v3 03/26] mm: introduce AS_NO_DIRECT_MAP Brendan Jackman
2026-07-30 21:06 ` Yosry Ahmed
2026-07-31 12:15 ` Brendan Jackman
2026-07-31 19:28 ` Yosry Ahmed
2026-08-02 16:10 ` Mike Rapoport
2026-07-26 22:22 ` [PATCH v3 04/26] x86/mm: split out preallocate_sub_pgd() Brendan Jackman
2026-07-31 22:10 ` Yosry Ahmed
2026-08-02 16:13 ` Mike Rapoport
2026-07-26 22:22 ` [PATCH v3 05/26] x86: move PAE PMD preallocation defines to header Brendan Jackman
2026-07-31 23:59 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 06/26] x86/tlb: Expose some flush function declarations to modules Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 07/26] x86/mm: introduce mm-local region Brendan Jackman
2026-08-02 16:27 ` Mike Rapoport
2026-08-03 22:29 ` Yosry Ahmed [this message]
2026-07-26 22:22 ` [PATCH v3 08/26] x86/mm: move LDT remap into " Brendan Jackman
2026-08-03 22:33 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 09/26] mm: Create flags arg for __apply_to_page_range() Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 10/26] mm: Add more flags " Brendan Jackman
2026-08-04 0:08 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 11/26] x86/mm: introduce the mermap Brendan Jackman
2026-08-02 16:40 ` Mike Rapoport
2026-08-04 18:38 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 12/26] mm: KUnit tests for " Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 13/26] mm: introduce freetype_t Brendan Jackman
2026-08-04 22:23 ` Yosry Ahmed
2026-08-04 23:02 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 14/26] mm: move migratetype definitions to freetype.h Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 15/26] mm/page_alloc: add support for freetypes with no freelist Brendan Jackman
2026-07-31 14:13 ` Vlastimil Babka (SUSE)
2026-07-26 22:22 ` [PATCH v3 16/26] mm: add definitions for allocating unmapped pages Brendan Jackman
2026-08-04 19:53 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 17/26] mm: encode freetype flags in pageblock flags Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 18/26] mm/page_alloc: separate pcplists by freetype flags Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 19/26] mm/page_alloc: rename ALLOC_NON_BLOCK back to _HARDER Brendan Jackman
2026-07-31 14:52 ` Vlastimil Babka (SUSE)
2026-08-03 9:20 ` Vlastimil Babka (SUSE)
2026-08-04 21:50 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 20/26] mm/page_alloc: introduce ALLOC_NOBLOCK Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 21/26] mm/page_alloc: implement FREETYPE_UNMAPPED allocations Brendan Jackman
2026-08-03 9:18 ` Vlastimil Babka (SUSE)
2026-08-04 23:41 ` Yosry Ahmed
2026-08-04 23:53 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 22/26] mm: Minimal KUnit tests for some new page_alloc logic Brendan Jackman
2026-08-03 9:30 ` Vlastimil Babka (SUSE)
2026-07-26 22:22 ` [PATCH v3 23/26] mm: Split out NR_FREE_PAGES_BLOCKS_[UN]MAPPED Brendan Jackman
2026-08-03 9:32 ` Vlastimil Babka (SUSE)
2026-07-26 22:22 ` [PATCH v3 24/26] mm/page_alloc: always direct compact for unmapped allocs Brendan Jackman
2026-08-03 9:44 ` Vlastimil Babka (SUSE)
2026-07-26 22:22 ` [PATCH v3 25/26] mm: plumb alloc flags into some alloc funcs Brendan Jackman
2026-08-03 9:52 ` Vlastimil Babka (SUSE)
2026-07-26 22:22 ` [PATCH v3 26/26] mm: add fast path for AS_NO_DIRECT_MAP Brendan Jackman
2026-07-29 11:52 ` [PATCH v3 00/26] mm: Add ALLOC_UNMAPPED and AS_NO_DIRECT_MAP Brendan Jackman
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=anES7O7T_PN2bV0v@google.com \
--to=yosry@kernel.org \
--cc=akpm@linux-foundation.org \
--cc=bp@alien8.de \
--cc=dave.hansen@linux.intel.com \
--cc=david.kaplan@amd.com \
--cc=david@kernel.org \
--cc=derkling@google.com \
--cc=hannes@cmpxchg.org \
--cc=itazur@amazon.co.uk \
--cc=jackmanb@google.com \
--cc=kalyazin@amazon.co.uk \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=luto@kernel.org \
--cc=patrick.roy@linux.dev \
--cc=peterz@infradead.org \
--cc=reijiw@google.com \
--cc=rientjes@google.com \
--cc=rppt@kernel.org \
--cc=seanjc@google.com \
--cc=sumit.garg@oss.qualcomm.com \
--cc=tglx@kernel.org \
--cc=vbabka@kernel.org \
--cc=weixugc@google.com \
--cc=will@kernel.org \
--cc=x86@kernel.org \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.