Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: mpenttil@redhat.com
Cc: dri-devel@lists.freedesktop.org, intel-xe@lists.freedesktop.org
Subject: Re: [PATCH v14 03/12] mm/hmm: preparations for HMM to participate in migration
Date: Tue, 22 Sep 2026 05:49:35 +0000	[thread overview]
Message-ID: <20260922054936.26AB11F000FF@smtp.kernel.org> (raw)
In-Reply-To: <20260922053421.4092027-4-mpenttil@redhat.com>

Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
- [High] mm: hmm: migration range start address truncated on page walk retry
- [High] mm: hmm: unbalanced MMU notifier invalidation when dev_private_owner is NULL
- [High] mm: hmm: NULL pointer dereference in hmm_range_fault_unlocked_timeout
--

commit 7ddb4b8635599a4339659239ade3b1bbc78ceec7
Author: Mika Penttilä <mpenttil@redhat.com>

    mm/hmm: preparations for HMM to participate in migration
    
    For migration to happen after hmm_range_fault(), the vma has to
    retrieved. Luckily, pagewalk already resolves that for us, so
    just have to save it for further use. For migration, mmap_lock
    has to be hold for the whole operation, so vma stays stable.

> diff --git a/mm/hmm.c b/mm/hmm.c
> index 2f1e98c6b6440..be39d2e5403ab 100644
> --- a/mm/hmm.c
> +++ b/mm/hmm.c

[ ... ]

> @@ -400,6 +407,57 @@ static int hmm_vma_handle_absent_pmd(struct mm_walk *walk, unsigned long start,
>  }
>  #endif  /* CONFIG_ARCH_HAS_PMD_SOFTLEAVES */
>  
> +static int hmm_vma_capture_migrate_range(unsigned long start,
> +					 unsigned long end,
> +					 struct mm_walk *walk)
> +{
> +	struct hmm_vma_walk *hmm_vma_walk = walk->private;
> +	struct hmm_range *range = hmm_vma_walk->range;
> +
> +	if (!hmm_select_migrate(range))
> +		return 0;
> +
> +	if (hmm_vma_walk->vma && (hmm_vma_walk->vma != walk->vma))
> +		return -ERANGE;
> +
> +	hmm_vma_walk->vma = walk->vma;
> +	hmm_vma_walk->start = start;

[Severity: High]
Is the tracked start address unconditionally overwritten on every walk restart?

If the walk requires a retry, returning -EBUSY, the outer loop restarts
walk_page_range() from the intermediate faulting address. On retry, could this
clobber the tracking start with the intermediate address, leading to a
truncated start address?

> +	hmm_vma_walk->end = end;
> +
> +	if (end - start > range->end - range->start)
> +		return -ERANGE;
> +
> +	if (!hmm_vma_walk->mmu_range.owner) {

[Severity: High]
Could using the owner field as an initialization flag be problematic here?

If range->dev_private_owner is legitimately NULL, mmu_range.owner becomes
NULL. If a page fault retry occurs, the check passes again, causing a
duplicate mmu_notifier_invalidate_range_start(). Can this cause unbalanced
MMU notifier invalidation?

> +		mmu_notifier_range_init_owner(&hmm_vma_walk->mmu_range, MMU_NOTIFY_MIGRATE, 0,
> +					      walk->vma->vm_mm, start, end,
> +					      range->dev_private_owner);
> +		mmu_notifier_invalidate_range_start(&hmm_vma_walk->mmu_range);
> +	}
> +
> +	return 0;
> +}
> +
> +static void hmm_vma_post_range_fault(struct hmm_vma_walk *hmm_vma_walk)
> +{
> +
> +	struct hmm_range *range = hmm_vma_walk->range;
> +
> +	if (hmm_select_migrate(range) &&
> +	    hmm_vma_walk->mmu_range.owner) {

[Severity: High]
Are we missing a call to mmu_notifier_invalidate_range_end() if the owner is
NULL?

Assuming a NULL owner means the invalidation was never started could
permanently leak the invalidation if dev_private_owner is legitimately NULL.

> +		/*
> +		 *  The migrate_vma path has the following initialized,
> +		 *  so take care of fault path below.
> +		 */
> +		if (range->notifier) {
> +			hmm_fill_migrate_vma(range,
> +					     hmm_vma_walk->vma,
> +					     hmm_vma_walk->start,

[Severity: High]
Will this propagate the truncated start address back to the caller's migration
struct?

If the walk was retried and the start address was clobbered in
hmm_vma_capture_migrate_range(), it could cause the migration setup to
miscalculate the page count and read from the wrong array offset.

> +					     hmm_vma_walk->end);
> +		}
> +		mmu_notifier_invalidate_range_end(&hmm_vma_walk->mmu_range);
> +	}
> +}

[ ... ]

> @@ -694,19 +757,31 @@ static int hmm_range_fault_locked(struct hmm_range *range, bool *locked)
>  {
>  	struct hmm_vma_walk hmm_vma_walk = {
>  		.range = range,
> -		.locked = locked,
> +		.mmlocked = locked,
>  		.last = range->start,
>  	};
> -	struct mm_struct *mm = range->notifier->mm;
> +	/*
> +	 *  Could be serving a device fault or come from migrate
> +	 *  entry point. For the former we have not resolved the vma
> +	 *  yet, and the latter we don't have a notifier (but have a vma).
> +	 *
> +	 */
> +	struct mm_struct *mm = hmm_range_fault_mm(range);

[Severity: High]
How does this API change affect hmm_range_fault_unlocked_timeout()?

The patch modifies the HMM API to allow range->notifier to be NULL for
migration entry points, handling it gracefully here with
hmm_range_fault_mm(). However, looking at hmm_range_fault_unlocked_timeout()
in mm/hmm.c:

int hmm_range_fault_unlocked_timeout(struct hmm_range *range,
                                     unsigned long timeout)
{
        struct mm_struct *mm = range->notifier->mm;

It unconditionally dereferences range->notifier->mm at its very first line. Can
this trigger a NULL pointer dereference if driver code passes a migration range
with a NULL notifier to the unlocked variant?

>  	int ret;
>  
> +	/* Migration is incompatible with mmap lock drop */
> +	if (locked && hmm_select_migrate(range))
> +		return -EINVAL;
> +

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260922053421.4092027-1-mpenttil@redhat.com?part=3

  reply	other threads:[~2026-09-22  5:49 UTC|newest]

Thread overview: 35+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-22  5:34 [PATCH 00/12] [PATCH v14 00/12] migrate on fault for device pages mpenttil
2026-09-22  5:34 ` [PATCH v14 01/12] mm/Kconfig: changes for " mpenttil
2026-09-22  5:44   ` sashiko-bot
2026-09-22 22:27   ` Balbir Singh
2026-09-23  5:42     ` Mika Penttilä
2026-09-22  5:34 ` [PATCH v14 02/12] mm: add helper to convert HMM pfn to migrate pfn mpenttil
2026-09-22  5:34 ` [PATCH v14 03/12] mm/hmm: preparations for HMM to participate in migration mpenttil
2026-09-22  5:49   ` sashiko-bot [this message]
2026-09-22  5:34 ` [PATCH v14 04/12] mm/hmm: do the plumbing " mpenttil
2026-09-22  5:50   ` sashiko-bot
2026-09-22  5:34 ` [PATCH v14 05/12] mm/hmm: implement folio split for migrate needs in HMM pagewalk mpenttil
2026-09-22  5:47   ` sashiko-bot
2026-09-22  5:34 ` [PATCH v14 06/12] mm/hmm: migrate collection in HMM pagewalk - pte level mpenttil
2026-09-22  5:47   ` sashiko-bot
2026-09-22  5:34 ` [PATCH v14 07/12] mm/hmm: migrate collection in HMM pagewalk - pmd level mpenttil
2026-09-22  5:50   ` sashiko-bot
2026-09-22  5:34 ` [PATCH v14 08/12] mm/hmm: add lazy MMU mode support for migration in HMM pagewalk mpenttil
2026-09-22  5:51   ` sashiko-bot
2026-09-22  5:34 ` [PATCH v14 09/12] mm/hmm: implement rollback for device page " mpenttil
2026-09-22  5:34 ` [PATCH v14 10/12] mm: enable device page migration from " mpenttil
2026-09-22  5:58   ` sashiko-bot
2026-09-22  5:34 ` [PATCH v14 11/12] lib/test_hmm: add a new testcase for the migrate on fault mpenttil
2026-09-22  6:09   ` sashiko-bot
2026-09-22  5:34 ` [PATCH v14 12/12] Documentation/mm/hmm: document migration through hmm_range_fault() mpenttil
2026-09-22  5:41 ` ✗ CI.checkpatch: warning for Migrate on fault for device pages (rev6) Patchwork
2026-09-22  5:43 ` ✓ CI.KUnit: success " Patchwork
2026-09-22  6:00 ` ✗ CI.checksparse: warning " Patchwork
2026-09-22  7:08 ` ✓ Xe.CI.BAT: success " Patchwork
2026-09-22 15:03 ` ✗ Xe.CI.FULL: failure " Patchwork
2026-09-23  2:27 ` [PATCH 00/12] [PATCH v14 00/12] migrate on fault for device pages Andrew Morton
2026-09-23  5:29   ` Mika Penttilä
2026-09-23 21:19     ` Andrew Morton
2026-09-23 23:24       ` Jason Gunthorpe
2026-09-24  0:14         ` Mika Penttilä
2026-09-24  0:10       ` Mika Penttilä

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260922054936.26AB11F000FF@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=mpenttil@redhat.com \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox