Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Matthew Brost <matthew.brost@intel.com>
To: Matthew Auld <matthew.auld@intel.com>
Cc: intel-xe@lists.freedesktop.org,
	"Thomas Hellström" <thomas.hellstrom@linux.intel.com>,
	"Linus Torvalds" <torvalds@linux-foundation.org>,
	"Stuart Summers" <stuart.summers@intel.com>,
	"Rodrigo Vivi" <rodrigo.vivi@intel.com>
Subject: Re: [PATCH v3 5/5] drm/xe/vram: add early VRAM health check
Date: Wed, 2 Sep 2026 12:43:16 -0700	[thread overview]
Message-ID: <aph8VNG33AzI/bcV@gsse-cloud1.jf.intel.com> (raw)
In-Reply-To: <20260902124117.918018-12-matthew.auld@intel.com>

On Wed, Sep 02, 2026 at 01:41:23PM +0100, Matthew Auld wrote:
> During early probe, use the last page as a canary for BAR sizing, CCS
> sizing, identity map setup etc. If something is wrong the last page is
> where we will likely find it. Hit it with everything we have. For now
> this is gated behind a debug config option, so shouldn't trigger on
> production.
> 
> Main motivation is around CCS sizing where on some BMG cards the CCS
> offset is programmed misaligned, for whatever reason, and our handling
> of that was busted, as found by Linus, leading to some amount of CCS
> storage getting pulled into the allocator as normal VRAM.
> 
> Nothing in our CI farm has such a misaligned offset it would seem,
> however I did get this to pop on my b570, which does also have the
> misaligned CCS offset:
> 
>   Tile 0: Running VRAM memtest...
>   Tile 0: VRAM bounds overlap CCS region! VRAM sizing is incorrect.
> 
> With the fix from Linus, this goes away:
> 
>   Tile 0: Running VRAM memtest...
>   Tile 0: VRAM memtest completed.
> 
> For the VRAM health check itself, this adds:
> 
>   - CPU access to the last page (BAR).
>   - GPU access to the last page (identity map).
>   - CCS overlap check. This one is more involved, but overall idea is
>     fill the last page with a known pattern, and also save the CCS state
>     for the first 4M of VRAM to some scratch memory. We then zero the CCS
>     storage for that same range, all using the proper CCS copy instruction.
>     At this point we readback the last page, and check if the pattern we
>     wrote changed. Finally we restore the CCS state. This works since CCS
>     1:1 maps with VRAM, so the start of the raw CCS, should map to the
>     start of VRAM. This successfully catches the issue that Linus found
>     and fixed.
> 
> Assisted-by: Gemini:gemini-3.1-pro-preview
> Signed-off-by: Matthew Auld <matthew.auld@intel.com>
> Cc: Thomas Hellström <thomas.hellstrom@linux.intel.com>
> Cc: Linus Torvalds <torvalds@linux-foundation.org>
> Cc: Stuart Summers <stuart.summers@intel.com>
> Cc: Matthew Brost <matthew.brost@intel.com>

Reviewed-by: Matthew Brost <matthew.brost@intel.com>

> Cc: Rodrigo Vivi <rodrigo.vivi@intel.com>
> ---
>  drivers/gpu/drm/xe/xe_device.c     |   8 +
>  drivers/gpu/drm/xe/xe_migrate.c    |  64 ++++++++
>  drivers/gpu/drm/xe/xe_migrate.h    |   6 +
>  drivers/gpu/drm/xe/xe_tile_types.h |   4 +
>  drivers/gpu/drm/xe/xe_vram.c       | 241 +++++++++++++++++++++++++++++
>  drivers/gpu/drm/xe/xe_vram.h       |  10 ++
>  6 files changed, 333 insertions(+)
> 
> diff --git a/drivers/gpu/drm/xe/xe_device.c b/drivers/gpu/drm/xe/xe_device.c
> index fa98b2d12204..8583b2e9ecf4 100644
> --- a/drivers/gpu/drm/xe/xe_device.c
> +++ b/drivers/gpu/drm/xe/xe_device.c
> @@ -1051,6 +1051,10 @@ int xe_device_probe(struct xe_device *xe)
>  	if (err)
>  		return err;
>  
> +	err = xe_vram_reserve_memtest_bo(xe);
> +	if (err)
> +		return err;
> +
>  	for_each_tile(tile, xe, id) {
>  		err = xe_tile_init(tile);
>  		if (err)
> @@ -1067,6 +1071,10 @@ int xe_device_probe(struct xe_device *xe)
>  			return err;
>  	}
>  
> +	err = xe_vram_memtest(xe);
> +	if (err)
> +		return err;
> +
>  	err = xe_pagefault_init(xe);
>  	if (err)
>  		return err;
> diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
> index 0bf000d7c901..ff45c24d8889 100644
> --- a/drivers/gpu/drm/xe/xe_migrate.c
> +++ b/drivers/gpu/drm/xe/xe_migrate.c
> @@ -2650,3 +2650,67 @@ void xe_migrate_job_lock_assert(struct xe_exec_queue *q)
>  #if IS_ENABLED(CONFIG_DRM_XE_KUNIT_TEST)
>  #include "tests/xe_migrate.c"
>  #endif
> +
> +#if IS_ENABLED(CONFIG_DRM_XE_DEBUG_MEM)
> +int xe_migrate_debug_ccs_overlap(struct xe_migrate *m,
> +				 struct xe_bo *scratch_bo,
> +				 bool write_to_ccs)
> +{
> +	struct xe_device *xe = tile_to_xe(m->tile);
> +	struct xe_gt *gt = m->tile->primary_gt;
> +	struct dma_fence *fence;
> +	struct xe_bb *bb;
> +	struct xe_sched_job *job;
> +	u64 first_page_dpa, clear_L0_ofs, scratch_dpa, scratch_L0_ofs;
> +
> +	if (!xe_device_has_flat_ccs(xe))
> +		return -EINVAL;
> +
> +	first_page_dpa = xe_vram_region_dpa_base(m->tile->mem.vram);
> +	clear_L0_ofs = xe_migrate_vram_ofs(xe, first_page_dpa, true);
> +
> +	scratch_dpa = xe_bo_addr(scratch_bo, 0, XE_PAGE_SIZE);
> +	scratch_L0_ofs = xe_migrate_vram_ofs(xe, scratch_dpa, false);
> +
> +	bb = xe_bb_new(gt, EMIT_COPY_CCS_DW + 1, xe->info.has_usm);
> +	if (IS_ERR(bb)) {
> +		drm_warn(&xe->drm, "Failed to create bb for VRAM overlap check\n");
> +		return PTR_ERR(bb);
> +	}
> +
> +	/* 4MB payload = 8KB CCS metadata */
> +	if (write_to_ccs) {
> +		emit_copy_ccs(gt, bb, clear_L0_ofs, true,
> +			      scratch_L0_ofs, false, SZ_4M);
> +	} else {
> +		emit_copy_ccs(gt, bb, scratch_L0_ofs, false,
> +			      clear_L0_ofs, true, SZ_4M);
> +	}
> +
> +	bb->cs[bb->len++] = MI_BATCH_BUFFER_END;
> +
> +	job = xe_bb_create_migration_job(m->q, bb,
> +					 xe_migrate_batch_base(m, xe->info.has_usm),
> +					 0);
> +	if (!IS_ERR(job)) {
> +		xe_sched_job_add_migrate_flush(job, MI_FLUSH_DW_CCS);
> +
> +		mutex_lock(&m->job_mutex);
> +		xe_sched_job_arm(job);
> +
> +		fence = dma_fence_get(&job->drm.s_fence->finished);
> +		xe_sched_job_push(job);
> +		mutex_unlock(&m->job_mutex);
> +
> +		dma_fence_wait(fence, false);
> +		dma_fence_put(fence);
> +	} else {
> +		drm_warn(&xe->drm, "Failed to create job for VRAM overlap check\n");
> +		xe_bb_free(bb, NULL);
> +		return PTR_ERR(job);
> +	}
> +
> +	xe_bb_free(bb, NULL);
> +	return 0;
> +}
> +#endif
> diff --git a/drivers/gpu/drm/xe/xe_migrate.h b/drivers/gpu/drm/xe/xe_migrate.h
> index c3a268b01768..a9acc62f78f0 100644
> --- a/drivers/gpu/drm/xe/xe_migrate.h
> +++ b/drivers/gpu/drm/xe/xe_migrate.h
> @@ -182,4 +182,10 @@ static inline void xe_migrate_job_lock_assert(struct xe_exec_queue *q)
>  void xe_migrate_job_lock(struct xe_migrate *m, struct xe_exec_queue *q);
>  void xe_migrate_job_unlock(struct xe_migrate *m, struct xe_exec_queue *q);
>  
> +#if IS_ENABLED(CONFIG_DRM_XE_DEBUG_MEM)
> +int xe_migrate_debug_ccs_overlap(struct xe_migrate *m,
> +				 struct xe_bo *scratch_bo,
> +				 bool write_to_ccs);
> +#endif
> +
>  #endif
> diff --git a/drivers/gpu/drm/xe/xe_tile_types.h b/drivers/gpu/drm/xe/xe_tile_types.h
> index 0048100ccb72..e1368c04846a 100644
> --- a/drivers/gpu/drm/xe/xe_tile_types.h
> +++ b/drivers/gpu/drm/xe/xe_tile_types.h
> @@ -97,6 +97,10 @@ struct xe_tile {
>  		 * Only main GT has page reclaim list allocations.
>  		 */
>  		struct xe_sa_manager *reclaim_pool;
> +#if IS_ENABLED(CONFIG_DRM_XE_DEBUG_MEM)
> +		/** @mem.memtest_bo: VRAM overlap check BO */
> +		struct xe_bo *memtest_bo;
> +#endif
>  	} mem;
>  
>  	/** @sriov: tile level virtualization data */
> diff --git a/drivers/gpu/drm/xe/xe_vram.c b/drivers/gpu/drm/xe/xe_vram.c
> index 04d831b101bd..dcedd8cfd731 100644
> --- a/drivers/gpu/drm/xe/xe_vram.c
> +++ b/drivers/gpu/drm/xe/xe_vram.c
> @@ -17,8 +17,11 @@
>  #include "xe_device.h"
>  #include "xe_force_wake.h"
>  #include "xe_gt_mcr.h"
> +#include "xe_map.h"
> +#include "xe_migrate.h"
>  #include "xe_mmio.h"
>  #include "xe_sriov.h"
> +#include "xe_tile.h"
>  #include "xe_tile_sriov_vf.h"
>  #include "xe_ttm_vram_mgr.h"
>  #include "xe_vram.h"
> @@ -407,3 +410,241 @@ resource_size_t xe_vram_region_actual_physical_size(const struct xe_vram_region
>  	return vram ? vram->actual_physical_size : 0;
>  }
>  EXPORT_SYMBOL_IF_KUNIT(xe_vram_region_actual_physical_size);
> +
> +#if IS_ENABLED(CONFIG_DRM_XE_DEBUG_MEM)
> +static void memtest_bo_cleanup(void *arg)
> +{
> +	struct xe_device *xe = arg;
> +
> +	xe_vram_free_memtest_bos(xe);
> +}
> +
> +int xe_vram_reserve_memtest_bo(struct xe_device *xe)
> +{
> +	struct xe_tile *tile;
> +	u8 id;
> +
> +	if (IS_SRIOV_VF(xe))
> +		return 0;
> +
> +	for_each_tile(tile, xe, id) {
> +		u64 vram_size;
> +
> +		if (!tile->mem.vram)
> +			continue;
> +
> +		if (tile->mem.vram->io_size < tile->mem.vram->usable_size) {
> +			drm_info(&xe->drm,
> +				 "Tile %d: Small-BAR system detected, skipping VRAM memtest\n",
> +				 id);
> +			continue;
> +		}
> +
> +		vram_size = tile->mem.vram->usable_size;
> +
> +		tile->mem.memtest_bo = xe_bo_create_pin_map_at_novm(xe, tile, SZ_64K,
> +								    vram_size - SZ_64K,
> +								    ttm_bo_type_kernel,
> +								    XE_BO_FLAG_VRAM_IF_DGFX(tile),
> +								    0, false);
> +		if (IS_ERR(tile->mem.memtest_bo)) {
> +			drm_warn(&xe->drm, "Tile %d: Failed to reserve memtest BO\n", id);
> +			tile->mem.memtest_bo = NULL;
> +			continue;
> +		}
> +
> +		drm_info(&xe->drm, "Tile %d: Reserved memtest BO at offset 0x%llx\n",
> +			 id, vram_size - SZ_64K);
> +	}
> +
> +	return devm_add_action_or_reset(xe->drm.dev, memtest_bo_cleanup, xe);
> +}
> +
> +void xe_vram_free_memtest_bos(struct xe_device *xe)
> +{
> +	struct xe_tile *tile;
> +	u8 id;
> +
> +	for_each_tile(tile, xe, id) {
> +		if (tile->mem.memtest_bo) {
> +			xe_bo_unpin_map_no_vm(tile->mem.memtest_bo);
> +			tile->mem.memtest_bo = NULL;
> +		}
> +	}
> +}
> +
> +int xe_vram_memtest(struct xe_device *xe)
> +{
> +	struct xe_tile *tile;
> +	u8 id;
> +	int err = 0;
> +
> +	if (IS_SRIOV_VF(xe))
> +		return 0;
> +
> +	for_each_tile(tile, xe, id) {
> +		struct xe_bo *last_page_bo = tile->mem.memtest_bo;
> +		struct dma_fence *fence;
> +		bool overlap = false;
> +		int i;
> +		u8 val;
> +
> +		if (!last_page_bo || !tile->migrate)
> +			continue;
> +
> +		drm_info(&xe->drm, "Tile %d: Running VRAM memtest...\n", id);
> +
> +		/* CPU write and readback first and last byte of the last page */
> +		xe_map_wr(xe, &last_page_bo->vmap, 0, u8, 0xA5);
> +		xe_map_wr(xe, &last_page_bo->vmap, SZ_64K - 1, u8, 0x5A);
> +
> +		val = xe_map_rd(xe, &last_page_bo->vmap, 0, u8);
> +		if (drm_WARN(&xe->drm, val != 0xA5,
> +			     "Tile %d: CPU memtest failed at offset 0 (expected 0xA5, got 0x%02x)\n",
> +			     id, val)) {
> +			err = -EIO;
> +			goto unpin;
> +		}
> +
> +		val = xe_map_rd(xe, &last_page_bo->vmap, SZ_64K - 1, u8);
> +		if (drm_WARN(&xe->drm, val != 0x5A,
> +			     "Tile %d: CPU memtest failed at offset 65535 (expected 0x5A, got 0x%02x)\n",
> +			     id, val)) {
> +			err = -EIO;
> +			goto unpin;
> +		}
> +
> +		/* Non-CCS access via GPU on the last page */
> +		xe_bo_lock(last_page_bo, false);
> +		fence = xe_migrate_clear(tile->migrate, last_page_bo,
> +					 last_page_bo->ttm.resource,
> +					 XE_MIGRATE_CLEAR_FLAG_BO_DATA);
> +		xe_bo_unlock(last_page_bo);
> +
> +		if (!IS_ERR(fence)) {
> +			dma_fence_wait(fence, false);
> +			dma_fence_put(fence);
> +		} else {
> +			err = PTR_ERR(fence);
> +			goto unpin;
> +		}
> +
> +		val = xe_map_rd(xe, &last_page_bo->vmap, 0, u8);
> +		if (drm_WARN(&xe->drm, val != 0x00,
> +			     "Tile %d: GPU memtest clear failed at offset 0 (expected 0x00, got 0x%02x)\n",
> +			     id, val)) {
> +			err = -EIO;
> +			goto unpin;
> +		}
> +
> +		/*
> +		 * Check for CCS overlap on the root tile.
> +		 *
> +		 * TODO: maybe extend if we ever get multi-tile + CCS. Pay
> +		 * special attention to the l2 flush below. Currently that is
> +		 * hard coded to the root tile.
> +		 */
> +		if (!id && xe_device_has_flat_ccs(xe) &&
> +		    GRAPHICS_VERx100(xe) >= 2000) {
> +			struct xe_bo *scratch_bo_before;
> +			struct xe_bo *scratch_bo_after;
> +
> +			scratch_bo_before = xe_bo_create_pin_map_novm(xe, tile, SZ_64K,
> +								      ttm_bo_type_kernel,
> +								      XE_BO_FLAG_VRAM_IF_DGFX(tile),
> +								      false);
> +			if (IS_ERR(scratch_bo_before)) {
> +				err = PTR_ERR(scratch_bo_before);
> +				goto unpin;
> +			}
> +
> +			scratch_bo_after = xe_bo_create_pin_map_novm(xe, tile, SZ_64K,
> +								     ttm_bo_type_kernel,
> +								     XE_BO_FLAG_VRAM_IF_DGFX(tile),
> +								     false);
> +			if (IS_ERR(scratch_bo_after)) {
> +				xe_bo_unpin_map_no_vm(scratch_bo_before);
> +				err = PTR_ERR(scratch_bo_after);
> +				goto unpin;
> +			}
> +
> +			/* Save original CCS metadata for PA 0 + */
> +			err = xe_migrate_debug_ccs_overlap(tile->migrate, scratch_bo_before, false);
> +			if (err) {
> +				xe_bo_unpin_map_no_vm(scratch_bo_before);
> +				xe_bo_unpin_map_no_vm(scratch_bo_after);
> +				goto unpin;
> +			}
> +
> +			/*
> +			 * Fill last page. If there is CCS overlap in the last
> +			 * page this will snag the raw CCS storage.
> +			 */
> +			xe_map_memset(xe, &last_page_bo->vmap, 0, 0x5A, SZ_64K);
> +			xe_device_wmb(xe);
> +
> +			/*
> +			 * Global invalidation. Some BMG SKUs will cache the BAR
> +			 * writes in the GPU side VRAM cache. Make sure above
> +			 * writes are fully flushed out to VRAM, so this is
> +			 * hopefully more well behaved with the CCS unit, if
> +			 * there is indeed CCS overlap with normal VRAM. Since
> +			 * there is a separate CCS cache, the CCS unit might not
> +			 * respect the GPU VRAM cache for CCS accesses, so opt
> +			 * for being super careful here.
> +			 */
> +			xe_device_l2_flush(xe, true);
> +
> +			/* Use GPU to clear CCS state for PA 0 */
> +			xe_map_memset(xe, &scratch_bo_after->vmap, 0, 0x00, SZ_64K);
> +			err = xe_migrate_debug_ccs_overlap(tile->migrate, scratch_bo_after, true);
> +			if (err) {
> +				xe_bo_unpin_map_no_vm(scratch_bo_before);
> +				xe_bo_unpin_map_no_vm(scratch_bo_after);
> +				goto unpin;
> +			}
> +			/*
> +			 * Global invalidation. Ensure CCS caches really are
> +			 * nuked and the raw CCS data is visible in VRAM, for
> +			 * the below access.
> +			 */
> +			xe_device_l2_flush(xe, true);
> +
> +			/* Check if last_page_bo was corrupted by the GPU CCS clear */
> +			for (i = 0; i < SZ_64K; i += 8) {
> +				u64 payload = xe_map_rd(xe, &last_page_bo->vmap, i, u64);
> +
> +				if (payload != 0x5A5A5A5A5A5A5A5AULL) {
> +					overlap = true;
> +					break;
> +				}
> +			}
> +
> +			/* Restore original CCS metadata for PA 0 + */
> +			err = xe_migrate_debug_ccs_overlap(tile->migrate, scratch_bo_before, true);
> +			if (err)
> +				drm_warn(&xe->drm, "Failed to restore CCS metadata\n");
> +
> +			xe_bo_unpin_map_no_vm(scratch_bo_before);
> +			xe_bo_unpin_map_no_vm(scratch_bo_after);
> +		}
> +
> +		if (drm_WARN(&xe->drm, overlap,
> +			     "Tile %d: VRAM bounds overlap CCS region! VRAM sizing is incorrect.\n",
> +			     id)) {
> +			err = -EINVAL;
> +			goto unpin;
> +		}
> +
> +		drm_info(&xe->drm, "Tile %d: VRAM memtest completed.\n", id);
> +
> +unpin:
> +		if (err)
> +			break;
> +	}
> +
> +	xe_vram_free_memtest_bos(xe);
> +
> +	return err;
> +}
> +#endif
> diff --git a/drivers/gpu/drm/xe/xe_vram.h b/drivers/gpu/drm/xe/xe_vram.h
> index dd1c8bf17922..38425d81f777 100644
> --- a/drivers/gpu/drm/xe/xe_vram.h
> +++ b/drivers/gpu/drm/xe/xe_vram.h
> @@ -23,4 +23,14 @@ resource_size_t xe_vram_region_dpa_base(const struct xe_vram_region *vram);
>  resource_size_t xe_vram_region_usable_size(const struct xe_vram_region *vram);
>  resource_size_t xe_vram_region_actual_physical_size(const struct xe_vram_region *vram);
>  
> +#if IS_ENABLED(CONFIG_DRM_XE_DEBUG_MEM)
> +int xe_vram_reserve_memtest_bo(struct xe_device *xe);
> +void xe_vram_free_memtest_bos(struct xe_device *xe);
> +int xe_vram_memtest(struct xe_device *xe);
> +#else
> +static inline int xe_vram_reserve_memtest_bo(struct xe_device *xe) { return 0; }
> +static inline void xe_vram_free_memtest_bos(struct xe_device *xe) {}
> +static inline int xe_vram_memtest(struct xe_device *xe) { return 0; }
> +#endif
> +
>  #endif
> -- 
> 2.55.0
> 

  reply	other threads:[~2026-09-02 19:43 UTC|newest]

Thread overview: 14+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-02 12:41 [PATCH v3 0/5] VRAM health check + CCS fix Matthew Auld
2026-09-02 12:41 ` [PATCH v3 1/5] drm/xe/migrate: support 4K PTEs for identity map Matthew Auld
2026-09-02 12:41 ` [PATCH v3 2/5] drm/xe/vram: report FLAT_CCS base misalignment Matthew Auld
2026-09-02 19:42   ` Matthew Brost
2026-09-02 12:41 ` [PATCH v3 3/5] drm/xe/vram: revamp CPU VRAM mapping Matthew Auld
2026-09-02 13:12   ` sashiko-bot
2026-09-02 14:36     ` Matthew Auld
2026-09-02 19:51   ` Matthew Brost
2026-09-02 12:41 ` [PATCH v3 4/5] drm/xe: add force option for global invalidation Matthew Auld
2026-09-02 12:41 ` [PATCH v3 5/5] drm/xe/vram: add early VRAM health check Matthew Auld
2026-09-02 19:43   ` Matthew Brost [this message]
2026-09-02 13:28 ` ✓ CI.KUnit: success for VRAM health check + CCS fix (rev3) Patchwork
2026-09-02 14:23 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-03  0:36 ` ✗ Xe.CI.FULL: failure " Patchwork

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=aph8VNG33AzI/bcV@gsse-cloud1.jf.intel.com \
    --to=matthew.brost@intel.com \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=matthew.auld@intel.com \
    --cc=rodrigo.vivi@intel.com \
    --cc=stuart.summers@intel.com \
    --cc=thomas.hellstrom@linux.intel.com \
    --cc=torvalds@linux-foundation.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox