Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: tilak.tirumalesh.tangudu@intel.com
Cc: intel-xe@lists.freedesktop.org
Subject: Re: [PATCH 1/3] drm/xe/ggtt: stop mapping unmapped GGTT pages to scratch
Date: Mon, 24 Aug 2026 16:46:59 +0000	[thread overview]
Message-ID: <20260824164659.7471D1F000E9@smtp.kernel.org> (raw)
In-Reply-To: <20260824163458.3129935-2-tilak.tirumalesh.tangudu@intel.com>

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] Writing faulting PTEs for recycled holes but scratch PTEs for virgin holes causes non-deterministic spurious hardware faults depending on memory allocation history.
--

--- Patch [1]: [PATCH 1/3] drm/xe/ggtt: stop mapping unmapped GGTT pages to scratch ---
commit 240e3aa087f4e7c9c8232bdeba69ae3fe654d294
Author: Tangudu Tilak Tirumalesh <tilak.tirumalesh.tangudu@intel.com>

drm/xe/ggtt: stop mapping unmapped GGTT pages to scratch

This commit updates xe_ggtt_clear() to take a use_scratch parameter, leaving
the init-time clear to use scratch PTEs while forcing PTE=0 during node 
removal on multi-queue platforms. This aims to expose stale hardware TLB 
entries as observable faults.

> diff --git a/drivers/gpu/drm/xe/xe_ggtt.c b/drivers/gpu/drm/xe/xe_ggtt.c
> index 8ec23862477fc..4cddd69f0e992 100644
> --- a/drivers/gpu/drm/xe/xe_ggtt.c
> +++ b/drivers/gpu/drm/xe/xe_ggtt.c

[ ... ]

> @@ -257,23 +258,36 @@ static u64 xe_ggtt_get_pte(struct xe_ggtt *ggtt, u64 addr)
>  	return readq(&ggtt->gsm[addr >> XE_PTE_SHIFT]);
>  }
>  
> -static void xe_ggtt_clear(struct xe_ggtt *ggtt, u64 start, u64 size)
> +/**
> + * xe_ggtt_clear() - Clear a range of GGTT page-table entries
> + * @ggtt: the &xe_ggtt to clear
> + * @start: start of the range (GGTT address)
> + * @size: size of the range in bytes
> + * @use_scratch: true fills the range with the scratch PTE (init-time clear so
> + *		 engines walking free holes read zeros safely); false writes
> + *		 PTE=0 so a stale HW TLB entry faults instead of silently
> + *		 landing on scratch.
> + */

[Severity: High]
Does this explanation imply that hardware prefetchers speculatively walk
into unmapped ranges?

If engines legitimately walk free holes (as stated here for the init-time
clear), the hardware prefetcher cannot distinguish between virgin holes and
recycled holes.

[ ... ]

> @@ -473,12 +487,19 @@ static void ggtt_node_fini(struct xe_ggtt_node *node)
>  static void ggtt_node_remove(struct xe_ggtt_node *node)
>  {
>  	struct xe_ggtt *ggtt = node->ggtt;
> -	bool bound;
> +	bool use_scratch, bound;
> +
> +	/*
> +	 * Addition of a secondary queue while a multi-queue context is live on
> +	 * HW requires proper GGTT TLB clearing, so use faulting PTEs here.
> +	 */
> +	use_scratch = !xe_gt_has_multi_queue(ggtt->tile->primary_gt);
>  
>  	mutex_lock(&ggtt->lock);
>  	bound = ggtt->flags & XE_GGTT_FLAGS_ONLINE;
>  	if (bound)
> -		xe_ggtt_clear(ggtt, xe_ggtt_node_addr(node), xe_ggtt_node_size(node));
> +		xe_ggtt_clear(ggtt, xe_ggtt_node_addr(node),
> +			      xe_ggtt_node_size(node), use_scratch);

[Severity: High]
Can writing PTE=0 here cause spurious GPU or IOMMU page faults on
multi-queue platforms?

Since a prefetching engine might speculatively read past its buffer boundary
into an adjacent recycled hole, stripping the scratch PTE protection for
recycled memory while keeping it for virgin memory creates a situation
where faults depend entirely on memory allocation history.

Could this lead to engine hangs or system instability when a prefetching
engine hits the PTE=0 in a recently freed node?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260824163458.3129935-1-tilak.tirumalesh.tangudu@intel.com?part=1

  reply	other threads:[~2026-08-24 16:47 UTC|newest]

Thread overview: 13+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-24 16:34 [PATCH 0/3] drm/xe/ggtt: fix stale GGTT mappings on unmap tilak.tirumalesh.tangudu
2026-08-24 16:34 ` [PATCH 1/3] drm/xe/ggtt: stop mapping unmapped GGTT pages to scratch tilak.tirumalesh.tangudu
2026-08-24 16:46   ` sashiko-bot [this message]
2026-08-24 18:17   ` Niranjana Vishwanathapura
2026-08-24 16:34 ` [PATCH 2/3] drm/xe/tlb_inval: add xe_tlb_inval_ggtt_full() GGTT invalidation helper tilak.tirumalesh.tangudu
2026-08-24 18:20   ` Niranjana Vishwanathapura
2026-08-24 16:34 ` [PATCH 3/3] drm/xe/ggtt: invalidate engine GGTT TLBs for multi-queue GTs tilak.tirumalesh.tangudu
2026-08-24 16:48   ` sashiko-bot
2026-08-24 18:25   ` Niranjana Vishwanathapura
2026-08-24 20:44 ` ✓ CI.KUnit: success for drm/xe/ggtt: fix stale GGTT mappings on unmap (rev2) Patchwork
2026-08-24 21:24 ` ✓ Xe.CI.BAT: " Patchwork
  -- strict thread matches above, loose matches on Subject: below --
2026-08-21  7:42 [PATCH 0/3] drm/xe/ggtt: fix stale GGTT mappings on unmap tilak.tirumalesh.tangudu
2026-08-21  7:42 ` [PATCH 1/3] drm/xe/ggtt: stop mapping unmapped GGTT pages to scratch tilak.tirumalesh.tangudu
2026-08-24  0:32   ` Niranjana Vishwanathapura

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260824164659.7471D1F000E9@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=sashiko-reviews@lists.linux.dev \
    --cc=tilak.tirumalesh.tangudu@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox