Linux Trace Kernel
 help / color / mirror / Atom feed
From: "Lorenzo Stoakes (ARM)" <ljs@kernel.org>
To: "Mike Rapoport (Microsoft)" <rppt@kernel.org>
Cc: Andrew Morton <akpm@linux-foundation.org>,
	 David Hildenbrand <david@kernel.org>,
	Baolin Wang <baolin.wang@linux.alibaba.com>,
	 Barry Song <baohua@kernel.org>, Dev Jain <dev.jain@arm.com>,
	Hugh Dickins <hughd@google.com>,  Jann Horn <jannh@google.com>,
	Jason Gunthorpe <jgg@ziepe.ca>,
	 John Hubbard <jhubbard@nvidia.com>,
	Jonathan Corbet <corbet@lwn.net>,
	 Lance Yang <lance.yang@linux.dev>,
	"Liam R. Howlett" <liam@infradead.org>,
	 Masami Hiramatsu <mhiramat@kernel.org>,
	Mathieu Desnoyers <mathieu.desnoyers@efficios.com>,
	 Michal Hocko <mhocko@suse.com>,
	Muchun Song <muchun.song@linux.dev>,
	 Nico Pache <nico.pache@linux.dev>,
	Oscar Salvador <osalvador@suse.de>,
	 Pedro Falcato <pfalcato@suse.de>, Peter Xu <peterx@redhat.com>,
	 Ryan Roberts <ryan.roberts@arm.com>,
	Shakeel Butt <shakeel.butt@linux.dev>,
	 Shuah Khan <skhan@linuxfoundation.org>,
	Steven Rostedt <rostedt@goodmis.org>,
	 Suren Baghdasaryan <surenb@google.com>,
	Usama Arif <usama.arif@linux.dev>,
	 Vlastimil Babka <vbabka@kernel.org>, Zi Yan <ziy@nvidia.com>,
	linux-doc@vger.kernel.org,  linux-fsdevel@vger.kernel.org,
	linux-kernel@vger.kernel.org, linux-mm@kvack.org,
	 linux-trace-kernel@vger.kernel.org
Subject: Re: [PATCH 3/6] userfaultfd: use userfaultfd_*() helpers instead of open coded flag tests
Date: Mon, 24 Aug 2026 16:10:51 +0100	[thread overview]
Message-ID: <aoxdWWMEeQbcG5Pq@gremlin> (raw)
In-Reply-To: <20260823-uffd-vm-flags-v1-v1-3-3086981b33cf@kernel.org>

On Sun, Aug 23, 2026 at 03:17:40PM +0300, Mike Rapoport (Microsoft) wrote:
> Move userfaultfd_{missing,wp,minor,rwp}() and userfaultfd_protected()
> ahead of uffd_disable_huge_pmd_share() and uffd_disable_fault_around()
> and make the latter two use the helpers rather than open coded VMA flag
> masks.
>
> Convert open coded VMA flag test in mfill_get_vma() to userfaultfd_wp()
> as well.

It'd be better to do the moves and the reworks separately. We don't have a limit
on patch count :)

>
> With every user of the per-VMA uffd modes going through the helpers,
> their underlying representation can be changed in the next step.
>
> No functional change.

There is a functional change, or at least seems to be, see below.

>
> Assisted-by: copilot:claude-opus-5
> Signed-off-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
> ---
>  include/linux/userfaultfd_k.h | 70 +++++++++++++++++++++----------------------
>  mm/userfaultfd.c              |  2 +-
>  2 files changed, 35 insertions(+), 37 deletions(-)
>
> diff --git a/include/linux/userfaultfd_k.h b/include/linux/userfaultfd_k.h
> index 3396d270b159..d8262e3dc134 100644
> --- a/include/linux/userfaultfd_k.h
> +++ b/include/linux/userfaultfd_k.h
> @@ -168,42 +168,6 @@ static inline bool is_mergeable_vm_userfaultfd_ctx(struct vm_area_struct *vma,
>  	return vma->vm_userfaultfd_ctx.ctx == vm_ctx.ctx;
>  }
>
> -/*
> - * Never enable huge pmd sharing on some uffd registered vmas:
> - *
> - * - VM_UFFD_WP and VM_UFFD_RWP VMAs, because the write protect / access
> - *   tracking information is per pgtable entry.
> - *
> - * - VM_UFFD_MINOR VMAs, because otherwise we would never get minor faults for
> - *   VMAs which share huge pmds. (If you have two mappings to the same
> - *   underlying pages, and fault in the non-UFFD-registered one with a write,
> - *   with huge pmd sharing this would *also* setup the second UFFD-registered
> - *   mapping, and we'd not get minor faults.)
> - */
> -static inline bool uffd_disable_huge_pmd_share(struct vm_area_struct *vma)
> -{
> -	return vma_test_any_mask(vma,
> -		mk_vma_flags_from_masks(VMA_UFFD_WP, VMA_UFFD_RWP,
> -					VMA_UFFD_MINOR));
> -}
> -
> -/*
> - * Don't do fault around for WP, RWP or MINOR registered uffd range.  For
> - * MINOR registered range, fault around will be a total disaster and ptes can
> - * be installed without notifications; for WP it should mostly be fine as long
> - * as the fault around checks for pte_none() before the installation, however
> - * to be super safe we just forbid it; for RWP, pre-faulted neighbours would
> - * be indistinguishable from accessed pages in PAGEMAP_SCAN (PAGE_IS_ACCESSED)
> - * and pollute the tracked working set, so each page must be populated by its
> - * own fault.
> - */
> -static inline bool uffd_disable_fault_around(struct vm_area_struct *vma)
> -{
> -	return vma_test_any_mask(vma,
> -		mk_vma_flags_from_masks(VMA_UFFD_WP, VMA_UFFD_RWP,
> -					VMA_UFFD_MINOR));
> -}
> -
>  static inline bool userfaultfd_missing(const struct vm_area_struct *vma)
>  {
>  	return vma_test_any_mask(vma, VMA_UFFD_MISSING);
> @@ -235,6 +199,40 @@ static inline bool userfaultfd_protected(const struct vm_area_struct *vma)
>  	return userfaultfd_wp(vma) || userfaultfd_rwp(vma);
>  }
>
> +/*
> + * Never enable huge pmd sharing on some uffd registered vmas:
> + *
> + * - uffd-WP and uffd-RWP VMAs, because the write protect / access tracking
> + *   information is per pgtable entry.
> + *
> + * - uffd-MINOR VMAs, because otherwise we would never get minor faults for
> + *   VMAs which share huge pmds. (If you have two mappings to the same
> + *   underlying pages, and fault in the non-UFFD-registered one with a write,
> + *   with huge pmd sharing this would *also* setup the second UFFD-registered
> + *   mapping, and we'd not get minor faults.)
> + */
> +static inline bool uffd_disable_huge_pmd_share(struct vm_area_struct *vma)
> +{
> +	return userfaultfd_minor(vma) || userfaultfd_wp(vma) ||
> +	       userfaultfd_rwp(vma);
> +}
> +
> +/*
> + * Don't do fault around for WP, RWP or MINOR registered uffd range.  For
> + * MINOR registered range, fault around will be a total disaster and ptes can
> + * be installed without notifications; for WP it should mostly be fine as long
> + * as the fault around checks for pte_none() before the installation, however
> + * to be super safe we just forbid it; for RWP, pre-faulted neighbours would
> + * be indistinguishable from accessed pages in PAGEMAP_SCAN (PAGE_IS_ACCESSED)
> + * and pollute the tracked working set, so each page must be populated by its
> + * own fault.
> + */
> +static inline bool uffd_disable_fault_around(struct vm_area_struct *vma)
> +{
> +	return userfaultfd_minor(vma) || userfaultfd_wp(vma) ||
> +	       userfaultfd_rwp(vma);

This is changing the logic.

Before we were testing only the flags, now we have:

static inline bool userfaultfd_rwp(const struct vm_area_struct *vma)
{
	/*
	 * Callers gate PAGE_NONE usage on this; PAGE_NONE is a BUILD_BUG()
	 * without CONFIG_ARCH_HAS_PTE_PROTNONE, so fold to false.
	 */
	if (!IS_ENABLED(CONFIG_ARCH_HAS_PTE_PROTNONE))
		return false;
	return vma_test_single_mask(vma, VMA_UFFD_RWP);
}

I.e. adding in a CONFIG_ARCH_HAS_PTE_PROTNONE check.

BTW side-note these:

static inline bool userfaultfd_missing(const struct vm_area_struct *vma)
{
	return vma_test_any_mask(vma, VMA_UFFD_MISSING);
}

static inline bool userfaultfd_wp(const struct vm_area_struct *vma)
{
	return vma_test_any_mask(vma, VMA_UFFD_WP);
}

static inline bool userfaultfd_minor(const struct vm_area_struct *vma)
{
	return vma_test_any_mask(vma, VMA_UFFD_MINOR);
}

Should all use vma_test_single_mask() really :)


> +}
> +
>  static inline bool userfaultfd_pte_wp(struct vm_area_struct *vma,
>  				      pte_t pte)
>  {
> diff --git a/mm/userfaultfd.c b/mm/userfaultfd.c
> index 74f04c323c50..32003aa04943 100644
> --- a/mm/userfaultfd.c
> +++ b/mm/userfaultfd.c
> @@ -261,7 +261,7 @@ static int mfill_get_vma(struct mfill_state *state)
>  	 * validate 'mode' now that we know the dst_vma: don't allow
>  	 * a wrprotect copy if the userfaultfd didn't register as WP.
>  	 */
> -	if ((flags & MFILL_ATOMIC_WP) && !(dst_vma->vm_flags & VM_UFFD_WP))
> +	if ((flags & MFILL_ATOMIC_WP) && !userfaultfd_wp(dst_vma))
>  		goto out_unlock;
>
>  	if (is_vm_hugetlb_page(dst_vma))
>
> --
> 2.53.0
>

--
Cheers, Lorenzo

  parent reply	other threads:[~2026-08-24 15:11 UTC|newest]

Thread overview: 25+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-23 12:17 [PATCH 0/6] userfaultfd: decouple uffd mode from VMA flags Mike Rapoport (Microsoft)
2026-08-23 12:17 ` [PATCH 1/6] mm/gup: move gup_can_follow_protnone() to gup.c Mike Rapoport (Microsoft)
2026-08-23 21:03   ` Barry Song
2026-08-24 14:42   ` David Hildenbrand (Arm)
2026-08-24 14:59   ` Lorenzo Stoakes (ARM)
2026-08-25  2:03   ` Zi Yan
2026-08-23 12:17 ` [PATCH 2/6] userfaultfd: constify VMA parameter of userfaultfd_*() helpers Mike Rapoport (Microsoft)
2026-08-23 12:27   ` sashiko-bot
2026-08-23 21:03   ` Barry Song
2026-08-24 15:03   ` Lorenzo Stoakes (ARM)
2026-08-25  2:03   ` Zi Yan
2026-08-23 12:17 ` [PATCH 3/6] userfaultfd: use userfaultfd_*() helpers instead of open coded flag tests Mike Rapoport (Microsoft)
2026-08-23 21:14   ` Barry Song
2026-08-24 15:10   ` Lorenzo Stoakes (ARM) [this message]
2026-08-23 12:17 ` [PATCH 4/6] userfaultfd: rename vm_userfaultfd_ctx to vm_uffd_state Mike Rapoport (Microsoft)
2026-08-24 14:43   ` David Hildenbrand (Arm)
2026-08-24 15:42   ` Lorenzo Stoakes (ARM)
2026-08-23 12:17 ` [PATCH 5/6] userfaultfd: decouple fault reason from VMA flags Mike Rapoport (Microsoft)
2026-08-24  8:12   ` Muchun Song
2026-08-24 14:46   ` David Hildenbrand (Arm)
2026-08-24 16:28   ` Lorenzo Stoakes (ARM)
2026-08-23 12:17 ` [PATCH 6/6] userfaultfd: collapse VM_UFFD_{MISSING,WP,MINOR,RWP} into single VM_UFFD Mike Rapoport (Microsoft)
2026-08-24  7:11   ` Lance Yang
2026-08-24  8:17     ` Mike Rapoport
2026-08-24  8:27       ` Lance Yang

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=aoxdWWMEeQbcG5Pq@gremlin \
    --to=ljs@kernel.org \
    --cc=akpm@linux-foundation.org \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=corbet@lwn.net \
    --cc=david@kernel.org \
    --cc=dev.jain@arm.com \
    --cc=hughd@google.com \
    --cc=jannh@google.com \
    --cc=jgg@ziepe.ca \
    --cc=jhubbard@nvidia.com \
    --cc=lance.yang@linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=linux-trace-kernel@vger.kernel.org \
    --cc=mathieu.desnoyers@efficios.com \
    --cc=mhiramat@kernel.org \
    --cc=mhocko@suse.com \
    --cc=muchun.song@linux.dev \
    --cc=nico.pache@linux.dev \
    --cc=osalvador@suse.de \
    --cc=peterx@redhat.com \
    --cc=pfalcato@suse.de \
    --cc=rostedt@goodmis.org \
    --cc=rppt@kernel.org \
    --cc=ryan.roberts@arm.com \
    --cc=shakeel.butt@linux.dev \
    --cc=skhan@linuxfoundation.org \
    --cc=surenb@google.com \
    --cc=usama.arif@linux.dev \
    --cc=vbabka@kernel.org \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox