Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Yosry Ahmed <yosry@kernel.org>
To: Brendan Jackman <jackmanb@google.com>
Cc: Borislav Petkov <bp@alien8.de>,
	 Dave Hansen <dave.hansen@linux.intel.com>,
	Peter Zijlstra <peterz@infradead.org>,
	 Andrew Morton <akpm@linux-foundation.org>,
	David Hildenbrand <david@kernel.org>,
	 Vlastimil Babka <vbabka@kernel.org>,
	Mike Rapoport <rppt@kernel.org>, Wei Xu <weixugc@google.com>,
	 Johannes Weiner <hannes@cmpxchg.org>, Zi Yan <ziy@nvidia.com>,
	Lorenzo Stoakes <ljs@kernel.org>,
	 linux-mm@kvack.org, linux-kernel@vger.kernel.org,
	x86@kernel.org,  Sumit Garg <sumit.garg@oss.qualcomm.com>,
	Will Deacon <will@kernel.org>,
	rientjes@google.com,  "Kalyazin, Nikita" <kalyazin@amazon.co.uk>,
	patrick.roy@linux.dev, "Itazuri, Takahiro" <itazur@amazon.co.uk>,
	 Andy Lutomirski <luto@kernel.org>,
	David Kaplan <david.kaplan@amd.com>,
	 Thomas Gleixner <tglx@kernel.org>,
	Patrick Bellasi <derkling@google.com>,
	 Reiji Watanabe <reijiw@google.com>,
	Sean Christopherson <seanjc@google.com>
Subject: Re: [PATCH v3 10/26] mm: Add more flags for __apply_to_page_range()
Date: Tue, 4 Aug 2026 00:08:37 +0000	[thread overview]
Message-ID: <anEnO7KbDQg4k5nf@google.com> (raw)
In-Reply-To: <20260726-page_alloc-unmapped-v3-10-6f5729aa9832@google.com>

On Sun, Jul 26, 2026 at 10:22:43PM +0000, Brendan Jackman wrote:
> Add two flags to make this API more generic:
> 
> 1. Separate "create" into two levels - one to allow creating new
>    mappings without allocating pagetables, and one for the current
>    behaviour that allows both of these.
> 
> 2. Create a new flag to report that the caller has taken care of
>    synchronization and no locks are required.
> 
> Both of these will serve to allow calling this API from restricted
> contexts where allocation and pagetable locking are not possible.
> 
> Signed-off-by: Brendan Jackman <jackmanb@google.com>
> ---
>  mm/internal.h | 26 +++++++++++++++++++++++++-
>  mm/memory.c   | 59 ++++++++++++++++++++++++++++++++++-------------------------
>  2 files changed, 59 insertions(+), 26 deletions(-)
> 
> diff --git a/mm/internal.h b/mm/internal.h
> index 395331a12d62d..5a237d9c5fa96 100644
> --- a/mm/internal.h
> +++ b/mm/internal.h
> @@ -1662,9 +1662,33 @@ static inline bool can_spin_trylock(void)
>  
>  /*
>   * Create a mapping if it doesn't exist. (Otherwise, skip regions with no
> - * existing mapping, and return an error for regions with no leaf pagetable).
> + * existing mapping). This doesn't allow allocating, most users will want
> + * PGRANGE_ALLOC.
> + *
> + * Do not test this bit directly as it is implied by PGRANGE_ALLOC, use
> + * pgrange_create() instead.
>   */
>  #define PGRANGE_CREATE		(1 << 0)
> +/*
> + * Allocate a pagetable if one is missing. (Otherwise, return an error for
> + * regions with no leaf pagetable). Also implies PGRANGE_CREATE.
> + *
> + * Note that __apply_to_page_range() assumes that pagetables for the area are
> + * already initialised down to PMD level, so this only affects PTEs in practice.
> + */
> +#define PGRANGE_ALLOC		(1 << 1)
> +/*
> + * Do not take any locks. This means the caller has taken care of
> + * synchronisation. This is incompatible with PGRANGE_ALLOC and also with
> + * mm=&init_mm.
> + */
> +#define PGRANGE_NOLOCK		(1 << 2)

I assume this is used by the mermap as locking is not required because
the mappings are per-CPU and migration is disabled while the mermap is
used?

Also, why is this incompatible with init_mm? It actually seems like
apply_to_pte_range() is always lockless for init_mm (uses
pte_offset_kernel()), probably callers are also synchronizing in their
own way (e.g. exclusive access to a vmap area?).

> +
> +
> +static inline bool pgrange_create(unsigned int flags)
> +{
> +	return flags & (PGRANGE_CREATE | PGRANGE_ALLOC);
> +}
>  
>  int __apply_to_page_range(struct mm_struct *mm, unsigned long addr,
>  			  unsigned long size, pte_fn_t fn,
> diff --git a/mm/memory.c b/mm/memory.c
> index c4defefea1574..d00508db1021e 100644
> --- a/mm/memory.c
> +++ b/mm/memory.c
> @@ -3441,30 +3441,35 @@ static int apply_to_pte_range(struct mm_struct *mm, pmd_t *pmd,
>  				     pte_fn_t fn, void *data, unsigned int flags,
>  				     pgtbl_mod_mask *mask)
>  {
> -	bool create = flags & PGRANGE_CREATE;
>  	pte_t *pte, *mapped_pte;
>  	int err = 0;
>  	spinlock_t *ptl;
>  
> -	if (create) {
> +	if (flags & PGRANGE_ALLOC) {
> +		VM_WARN_ON(flags & PGRANGE_NOLOCK);
> +
>  		mapped_pte = pte = (mm == &init_mm) ?
>  			pte_alloc_kernel_track(pmd, addr, mask) :
>  			pte_alloc_map_lock(mm, pmd, addr, &ptl);
>  		if (!pte)
>  			return -ENOMEM;
>  	} else {
> -		mapped_pte = pte = (mm == &init_mm) ?
> -			pte_offset_kernel(pmd, addr) :
> -			pte_offset_map_lock(mm, pmd, addr, &ptl);
> +		if (mm == &init_mm)
> +			pte = pte_offset_kernel(pmd, addr);
> +		else if (flags & PGRANGE_NOLOCK)
> +			pte = pte_offset_map(pmd, addr);
> +		else
> +			pte = pte_offset_map_lock(mm, pmd, addr, &ptl);
>  		if (!pte)
>  			return -EINVAL;
> +		mapped_pte = pte;
>  	}
>  
>  	lazy_mmu_mode_enable();
>  
>  	if (fn) {
>  		do {
> -			if (create || !pte_none(ptep_get(pte))) {
> +			if (pgrange_create(flags) || !pte_none(ptep_get(pte))) {
>  				err = fn(pte, addr, data);
>  				if (err)
>  					break;
> @@ -3475,8 +3480,13 @@ static int apply_to_pte_range(struct mm_struct *mm, pmd_t *pmd,
>  
>  	lazy_mmu_mode_disable();
>  
> -	if (mm != &init_mm)
> -		pte_unmap_unlock(mapped_pte, ptl);
> +	if (mm != &init_mm) {
> +		if (flags & PGRANGE_NOLOCK)
> +			pte_unmap(mapped_pte);
> +		else
> +			pte_unmap_unlock(mapped_pte, ptl);
> +	}
> +
>  	return err;
>  }
[..]


  reply	other threads:[~2026-08-04  0:08 UTC|newest]

Thread overview: 54+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-26 22:22 [PATCH v3 00/26] mm: Add ALLOC_UNMAPPED and AS_NO_DIRECT_MAP Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 01/26] set_memory: add folio_{zap,restore}_direct_map helpers Brendan Jackman
2026-07-27 10:33   ` Mike Rapoport
2026-07-29 11:42     ` Brendan Jackman
2026-07-30 20:34   ` Yosry Ahmed
2026-07-31  5:21     ` Mike Rapoport
2026-07-31 11:57       ` Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 02/26] mm/secretmem: make use of folio_{zap,restore}_direct_map Brendan Jackman
2026-07-27 10:40   ` Mike Rapoport
2026-07-26 22:22 ` [PATCH v3 03/26] mm: introduce AS_NO_DIRECT_MAP Brendan Jackman
2026-07-30 21:06   ` Yosry Ahmed
2026-07-31 12:15     ` Brendan Jackman
2026-07-31 19:28       ` Yosry Ahmed
2026-08-02 16:10   ` Mike Rapoport
2026-07-26 22:22 ` [PATCH v3 04/26] x86/mm: split out preallocate_sub_pgd() Brendan Jackman
2026-07-31 22:10   ` Yosry Ahmed
2026-08-02 16:13   ` Mike Rapoport
2026-07-26 22:22 ` [PATCH v3 05/26] x86: move PAE PMD preallocation defines to header Brendan Jackman
2026-07-31 23:59   ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 06/26] x86/tlb: Expose some flush function declarations to modules Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 07/26] x86/mm: introduce mm-local region Brendan Jackman
2026-08-02 16:27   ` Mike Rapoport
2026-08-03 22:29   ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 08/26] x86/mm: move LDT remap into " Brendan Jackman
2026-08-03 22:33   ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 09/26] mm: Create flags arg for __apply_to_page_range() Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 10/26] mm: Add more flags " Brendan Jackman
2026-08-04  0:08   ` Yosry Ahmed [this message]
2026-07-26 22:22 ` [PATCH v3 11/26] x86/mm: introduce the mermap Brendan Jackman
2026-08-02 16:40   ` Mike Rapoport
2026-07-26 22:22 ` [PATCH v3 12/26] mm: KUnit tests for " Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 13/26] mm: introduce freetype_t Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 14/26] mm: move migratetype definitions to freetype.h Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 15/26] mm/page_alloc: add support for freetypes with no freelist Brendan Jackman
2026-07-31 14:13   ` Vlastimil Babka (SUSE)
2026-07-26 22:22 ` [PATCH v3 16/26] mm: add definitions for allocating unmapped pages Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 17/26] mm: encode freetype flags in pageblock flags Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 18/26] mm/page_alloc: separate pcplists by freetype flags Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 19/26] mm/page_alloc: rename ALLOC_NON_BLOCK back to _HARDER Brendan Jackman
2026-07-31 14:52   ` Vlastimil Babka (SUSE)
2026-08-03  9:20     ` Vlastimil Babka (SUSE)
2026-07-26 22:22 ` [PATCH v3 20/26] mm/page_alloc: introduce ALLOC_NOBLOCK Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 21/26] mm/page_alloc: implement FREETYPE_UNMAPPED allocations Brendan Jackman
2026-08-03  9:18   ` Vlastimil Babka (SUSE)
2026-07-26 22:22 ` [PATCH v3 22/26] mm: Minimal KUnit tests for some new page_alloc logic Brendan Jackman
2026-08-03  9:30   ` Vlastimil Babka (SUSE)
2026-07-26 22:22 ` [PATCH v3 23/26] mm: Split out NR_FREE_PAGES_BLOCKS_[UN]MAPPED Brendan Jackman
2026-08-03  9:32   ` Vlastimil Babka (SUSE)
2026-07-26 22:22 ` [PATCH v3 24/26] mm/page_alloc: always direct compact for unmapped allocs Brendan Jackman
2026-08-03  9:44   ` Vlastimil Babka (SUSE)
2026-07-26 22:22 ` [PATCH v3 25/26] mm: plumb alloc flags into some alloc funcs Brendan Jackman
2026-08-03  9:52   ` Vlastimil Babka (SUSE)
2026-07-26 22:22 ` [PATCH v3 26/26] mm: add fast path for AS_NO_DIRECT_MAP Brendan Jackman
2026-07-29 11:52 ` [PATCH v3 00/26] mm: Add ALLOC_UNMAPPED and AS_NO_DIRECT_MAP Brendan Jackman

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=anEnO7KbDQg4k5nf@google.com \
    --to=yosry@kernel.org \
    --cc=akpm@linux-foundation.org \
    --cc=bp@alien8.de \
    --cc=dave.hansen@linux.intel.com \
    --cc=david.kaplan@amd.com \
    --cc=david@kernel.org \
    --cc=derkling@google.com \
    --cc=hannes@cmpxchg.org \
    --cc=itazur@amazon.co.uk \
    --cc=jackmanb@google.com \
    --cc=kalyazin@amazon.co.uk \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=luto@kernel.org \
    --cc=patrick.roy@linux.dev \
    --cc=peterz@infradead.org \
    --cc=reijiw@google.com \
    --cc=rientjes@google.com \
    --cc=rppt@kernel.org \
    --cc=seanjc@google.com \
    --cc=sumit.garg@oss.qualcomm.com \
    --cc=tglx@kernel.org \
    --cc=vbabka@kernel.org \
    --cc=weixugc@google.com \
    --cc=will@kernel.org \
    --cc=x86@kernel.org \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox