Linux Documentation
 help / color / mirror / Atom feed
From: "David Hildenbrand (Arm)" <david@kernel.org>
To: Stanislav Kinsburskii <skinsburskii@gmail.com>,
	airlied@gmail.com, akhilesh@ee.iitb.ac.in,
	akpm@linux-foundation.org, corbet@lwn.net, dakr@kernel.org,
	decui@microsoft.com, haiyangz@microsoft.com, jgg@ziepe.ca,
	kees@kernel.org, kys@microsoft.com, leon@kernel.org,
	liam@infradead.org, lizhi.hou@amd.com, ljs@kernel.org,
	longli@microsoft.com, lyude@redhat.com,
	maarten.lankhorst@linux.intel.com, mamin506@gmail.com,
	mhocko@suse.com, mripard@kernel.org,
	nouveau@lists.freedesktop.org, ogabbay@kernel.org,
	oleg@redhat.com, rppt@kernel.org, shuah@kernel.org,
	simona@ffwll.ch, skhan@linuxfoundation.org, surenb@google.com,
	tzimmermann@suse.de, vbabka@kernel.org, wei.liu@kernel.org
Cc: dri-devel@lists.freedesktop.org, intel-xe@lists.freedesktop.org,
	linux-mm@kvack.org, linux-doc@vger.kernel.org,
	linux-hyperv@vger.kernel.org, linux-kernel@vger.kernel.org,
	linux-kselftest@vger.kernel.org, linux-rdma@vger.kernel.org
Subject: Re: [PATCH v9 1/8] mm/hmm: move page fault handling out of walk callbacks
Date: Tue, 21 Jul 2026 17:19:06 +0200	[thread overview]
Message-ID: <0b9be5b3-93aa-407f-b83d-409bec4b55e1@kernel.org> (raw)
In-Reply-To: <178413935809.1155966.6398279131483632373.stgit@skinsburskii>

On 7/15/26 20:15, Stanislav Kinsburskii wrote:
> hmm_range_fault() currently triggers page faults from inside the page-table
> walk callbacks: hmm_vma_walk_pmd(), hmm_vma_walk_pud(),
> hmm_vma_walk_hugetlb_entry() and the pte-level helper all call
> hmm_vma_fault(), which in turn calls handle_mm_fault() while the walker
> still holds nested locks.  The pte spinlock is dropped explicitly by each
> caller, and the hugetlb path manually drops and retakes
> hugetlb_vma_lock_read around the fault to dodge a deadlock against the walk
> framework's unconditional unlock.
> 
> This layering does not extend cleanly to fault handlers that may release
> mmap_lock (VM_FAULT_RETRY, VM_FAULT_COMPLETED). If the lock is dropped
> while walk_page_range() is mid-traversal, the VMA can be freed before the
> walk framework's matching hugetlb_vma_unlock_read(), turning that unlock
> into a use-after-free.
> 
> Split the responsibilities the way get_user_pages() does. Walk callbacks
> become inspect-only: when they detect a range that needs to be faulted in,
> they record it in struct hmm_vma_walk and return a private sentinel
> (HMM_FAULT_PENDING). The outer loop in hmm_range_fault() then drops out of
> walk_page_range(), invokes a new helper hmm_do_fault() that calls
> handle_mm_fault() with only mmap_lock held, and restarts the walk so the
> now-present entries are collected into hmm_pfns.
> 
> No functional change for existing callers. As a side effect the hugetlb
> callback no longer needs the hugetlb_vma_{un}lock_read dance, and every
> fault-path exit from the callbacks now releases the pte spinlock on a
> single, common path. This refactor is also a precursor for adding an
> unlockable variant of hmm_range_fault() in a follow-up patch.
> 
> Reviewed-by: Jason Gunthorpe <jgg@nvidia.com>
> Signed-off-by: Stanislav Kinsburskii <skinsburskii@gmail.com>
> ---


[...]


>  
>  		pfn = pud_pfn(pud) + ((addr & ~PUD_MASK) >> PAGE_SHIFT);
> @@ -564,21 +561,8 @@ static int hmm_vma_walk_hugetlb_entry(pte_t *pte, unsigned long hmask,
>  	required_fault =
>  		hmm_pte_need_fault(hmm_vma_walk, pfn_req_flags, cpu_flags);
>  	if (required_fault) {
> -		int ret;
> -
>  		spin_unlock(ptl);
> -		hugetlb_vma_unlock_read(vma);
> -		/*
> -		 * Avoid deadlock: drop the vma lock before calling
> -		 * hmm_vma_fault(), which will itself potentially take and
> -		 * drop the vma lock. This is also correct from a
> -		 * protection point of view, because there is no further
> -		 * use here of either pte or ptl after dropping the vma
> -		 * lock.
> -		 */
> -		ret = hmm_vma_fault(addr, end, required_fault, walk);
> -		hugetlb_vma_lock_read(vma);
> -		return ret;
> +		return hmm_record_fault(addr, end, required_fault, walk);

Yes, that looks much better, as discussed. The downside is another vma_lookup()
in hmm_do_fault().

>  	}
>  
>  	pfn = pte_pfn(entry) + ((start & ~hmask) >> PAGE_SHIFT);
> @@ -637,6 +621,44 @@ static const struct mm_walk_ops hmm_walk_ops = {
>  	.walk_lock	= PGWALK_RDLOCK,
>  };
>  
> +/*
> + * hmm_do_fault - fault in a range recorded by a walk callback
> + *
> + * Called from the outer loop in hmm_range_fault() after a callback
> + * returned HMM_FAULT_PENDING.  At this point we hold only mmap_lock;
> + * the page-table spinlock and any hugetlb_vma_lock acquired by the walk
> + * framework have already been released by the unwind.
> + *
> + * Returns -EBUSY on success (all pages faulted, caller should re-walk).
> + * Returns a negative errno on failure.
> + */
> +static int hmm_do_fault(struct mm_struct *mm,
> +			struct hmm_vma_walk *hmm_vma_walk)
> +{
> +	unsigned long addr = hmm_vma_walk->last;
> +	unsigned long end = hmm_vma_walk->end;
> +	unsigned int required_fault = hmm_vma_walk->required_fault;

end and required_fault could be const.


Reviewed-by: David Hildenbrand (Arm) <david@kernel.org>


-- 
Cheers,

David

  reply	other threads:[~2026-07-21 15:19 UTC|newest]

Thread overview: 14+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-15 18:15 [PATCH v9 0/8] mm/hmm: Add mmap lock-drop support for userfaultfd-backed mappings Stanislav Kinsburskii
2026-07-15 18:15 ` [PATCH v9 1/8] mm/hmm: move page fault handling out of walk callbacks Stanislav Kinsburskii
2026-07-21 15:19   ` David Hildenbrand (Arm) [this message]
2026-07-15 18:16 ` [PATCH v9 2/8] mm/hmm: add hmm_range_fault_unlocked_timeout() for mmap lock-drop support Stanislav Kinsburskii
2026-07-21 15:42   ` David Hildenbrand (Arm)
2026-07-21 16:49     ` Stanislav Kinsburskii
2026-07-15 18:16 ` [PATCH v9 3/8] selftests/mm: add HMM test for mmap lock-dropping faults Stanislav Kinsburskii
2026-07-15 18:16 ` [PATCH v9 4/8] mshv: Use hmm_range_fault_unlocked_timeout() for region faults Stanislav Kinsburskii
2026-07-15 18:16 ` [PATCH v9 5/8] drm/nouveau: Use hmm_range_fault_unlocked_timeout() for SVM faults Stanislav Kinsburskii
2026-07-15 18:16 ` [PATCH v9 6/8] RDMA/umem: Use hmm_range_fault_unlocked_timeout() for ODP faults Stanislav Kinsburskii
2026-07-15 18:16 ` [PATCH v9 7/8] accel/amdxdna: Use hmm_range_fault_unlocked_timeout() for range population Stanislav Kinsburskii
2026-07-15 18:16 ` [PATCH v9 8/8] drm/gpusvm: Use hmm_range_fault_unlocked_timeout() for range faults Stanislav Kinsburskii
2026-07-21  0:49   ` Matthew Brost
2026-07-15 21:17 ` [PATCH v9 0/8] mm/hmm: Add mmap lock-drop support for userfaultfd-backed mappings Andrew Morton

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=0b9be5b3-93aa-407f-b83d-409bec4b55e1@kernel.org \
    --to=david@kernel.org \
    --cc=airlied@gmail.com \
    --cc=akhilesh@ee.iitb.ac.in \
    --cc=akpm@linux-foundation.org \
    --cc=corbet@lwn.net \
    --cc=dakr@kernel.org \
    --cc=decui@microsoft.com \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=haiyangz@microsoft.com \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=jgg@ziepe.ca \
    --cc=kees@kernel.org \
    --cc=kys@microsoft.com \
    --cc=leon@kernel.org \
    --cc=liam@infradead.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-hyperv@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=linux-rdma@vger.kernel.org \
    --cc=lizhi.hou@amd.com \
    --cc=ljs@kernel.org \
    --cc=longli@microsoft.com \
    --cc=lyude@redhat.com \
    --cc=maarten.lankhorst@linux.intel.com \
    --cc=mamin506@gmail.com \
    --cc=mhocko@suse.com \
    --cc=mripard@kernel.org \
    --cc=nouveau@lists.freedesktop.org \
    --cc=ogabbay@kernel.org \
    --cc=oleg@redhat.com \
    --cc=rppt@kernel.org \
    --cc=shuah@kernel.org \
    --cc=simona@ffwll.ch \
    --cc=skhan@linuxfoundation.org \
    --cc=skinsburskii@gmail.com \
    --cc=surenb@google.com \
    --cc=tzimmermann@suse.de \
    --cc=vbabka@kernel.org \
    --cc=wei.liu@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox