From: Matthew Brost <matthew.brost@intel.com>
To: Matthew Auld <matthew.auld@intel.com>
Cc: intel-xe@lists.freedesktop.org,
"Thomas Hellström" <thomas.hellstrom@linux.intel.com>,
"Linus Torvalds" <torvalds@linux-foundation.org>,
"Stuart Summers" <stuart.summers@intel.com>,
"Rodrigo Vivi" <rodrigo.vivi@intel.com>
Subject: Re: [PATCH v3 5/5] drm/xe/vram: add early VRAM health check
Date: Wed, 2 Sep 2026 12:43:16 -0700 [thread overview]
Message-ID: <aph8VNG33AzI/bcV@gsse-cloud1.jf.intel.com> (raw)
In-Reply-To: <20260902124117.918018-12-matthew.auld@intel.com>
On Wed, Sep 02, 2026 at 01:41:23PM +0100, Matthew Auld wrote:
> During early probe, use the last page as a canary for BAR sizing, CCS
> sizing, identity map setup etc. If something is wrong the last page is
> where we will likely find it. Hit it with everything we have. For now
> this is gated behind a debug config option, so shouldn't trigger on
> production.
>
> Main motivation is around CCS sizing where on some BMG cards the CCS
> offset is programmed misaligned, for whatever reason, and our handling
> of that was busted, as found by Linus, leading to some amount of CCS
> storage getting pulled into the allocator as normal VRAM.
>
> Nothing in our CI farm has such a misaligned offset it would seem,
> however I did get this to pop on my b570, which does also have the
> misaligned CCS offset:
>
> Tile 0: Running VRAM memtest...
> Tile 0: VRAM bounds overlap CCS region! VRAM sizing is incorrect.
>
> With the fix from Linus, this goes away:
>
> Tile 0: Running VRAM memtest...
> Tile 0: VRAM memtest completed.
>
> For the VRAM health check itself, this adds:
>
> - CPU access to the last page (BAR).
> - GPU access to the last page (identity map).
> - CCS overlap check. This one is more involved, but overall idea is
> fill the last page with a known pattern, and also save the CCS state
> for the first 4M of VRAM to some scratch memory. We then zero the CCS
> storage for that same range, all using the proper CCS copy instruction.
> At this point we readback the last page, and check if the pattern we
> wrote changed. Finally we restore the CCS state. This works since CCS
> 1:1 maps with VRAM, so the start of the raw CCS, should map to the
> start of VRAM. This successfully catches the issue that Linus found
> and fixed.
>
> Assisted-by: Gemini:gemini-3.1-pro-preview
> Signed-off-by: Matthew Auld <matthew.auld@intel.com>
> Cc: Thomas Hellström <thomas.hellstrom@linux.intel.com>
> Cc: Linus Torvalds <torvalds@linux-foundation.org>
> Cc: Stuart Summers <stuart.summers@intel.com>
> Cc: Matthew Brost <matthew.brost@intel.com>
Reviewed-by: Matthew Brost <matthew.brost@intel.com>
> Cc: Rodrigo Vivi <rodrigo.vivi@intel.com>
> ---
> drivers/gpu/drm/xe/xe_device.c | 8 +
> drivers/gpu/drm/xe/xe_migrate.c | 64 ++++++++
> drivers/gpu/drm/xe/xe_migrate.h | 6 +
> drivers/gpu/drm/xe/xe_tile_types.h | 4 +
> drivers/gpu/drm/xe/xe_vram.c | 241 +++++++++++++++++++++++++++++
> drivers/gpu/drm/xe/xe_vram.h | 10 ++
> 6 files changed, 333 insertions(+)
>
> diff --git a/drivers/gpu/drm/xe/xe_device.c b/drivers/gpu/drm/xe/xe_device.c
> index fa98b2d12204..8583b2e9ecf4 100644
> --- a/drivers/gpu/drm/xe/xe_device.c
> +++ b/drivers/gpu/drm/xe/xe_device.c
> @@ -1051,6 +1051,10 @@ int xe_device_probe(struct xe_device *xe)
> if (err)
> return err;
>
> + err = xe_vram_reserve_memtest_bo(xe);
> + if (err)
> + return err;
> +
> for_each_tile(tile, xe, id) {
> err = xe_tile_init(tile);
> if (err)
> @@ -1067,6 +1071,10 @@ int xe_device_probe(struct xe_device *xe)
> return err;
> }
>
> + err = xe_vram_memtest(xe);
> + if (err)
> + return err;
> +
> err = xe_pagefault_init(xe);
> if (err)
> return err;
> diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
> index 0bf000d7c901..ff45c24d8889 100644
> --- a/drivers/gpu/drm/xe/xe_migrate.c
> +++ b/drivers/gpu/drm/xe/xe_migrate.c
> @@ -2650,3 +2650,67 @@ void xe_migrate_job_lock_assert(struct xe_exec_queue *q)
> #if IS_ENABLED(CONFIG_DRM_XE_KUNIT_TEST)
> #include "tests/xe_migrate.c"
> #endif
> +
> +#if IS_ENABLED(CONFIG_DRM_XE_DEBUG_MEM)
> +int xe_migrate_debug_ccs_overlap(struct xe_migrate *m,
> + struct xe_bo *scratch_bo,
> + bool write_to_ccs)
> +{
> + struct xe_device *xe = tile_to_xe(m->tile);
> + struct xe_gt *gt = m->tile->primary_gt;
> + struct dma_fence *fence;
> + struct xe_bb *bb;
> + struct xe_sched_job *job;
> + u64 first_page_dpa, clear_L0_ofs, scratch_dpa, scratch_L0_ofs;
> +
> + if (!xe_device_has_flat_ccs(xe))
> + return -EINVAL;
> +
> + first_page_dpa = xe_vram_region_dpa_base(m->tile->mem.vram);
> + clear_L0_ofs = xe_migrate_vram_ofs(xe, first_page_dpa, true);
> +
> + scratch_dpa = xe_bo_addr(scratch_bo, 0, XE_PAGE_SIZE);
> + scratch_L0_ofs = xe_migrate_vram_ofs(xe, scratch_dpa, false);
> +
> + bb = xe_bb_new(gt, EMIT_COPY_CCS_DW + 1, xe->info.has_usm);
> + if (IS_ERR(bb)) {
> + drm_warn(&xe->drm, "Failed to create bb for VRAM overlap check\n");
> + return PTR_ERR(bb);
> + }
> +
> + /* 4MB payload = 8KB CCS metadata */
> + if (write_to_ccs) {
> + emit_copy_ccs(gt, bb, clear_L0_ofs, true,
> + scratch_L0_ofs, false, SZ_4M);
> + } else {
> + emit_copy_ccs(gt, bb, scratch_L0_ofs, false,
> + clear_L0_ofs, true, SZ_4M);
> + }
> +
> + bb->cs[bb->len++] = MI_BATCH_BUFFER_END;
> +
> + job = xe_bb_create_migration_job(m->q, bb,
> + xe_migrate_batch_base(m, xe->info.has_usm),
> + 0);
> + if (!IS_ERR(job)) {
> + xe_sched_job_add_migrate_flush(job, MI_FLUSH_DW_CCS);
> +
> + mutex_lock(&m->job_mutex);
> + xe_sched_job_arm(job);
> +
> + fence = dma_fence_get(&job->drm.s_fence->finished);
> + xe_sched_job_push(job);
> + mutex_unlock(&m->job_mutex);
> +
> + dma_fence_wait(fence, false);
> + dma_fence_put(fence);
> + } else {
> + drm_warn(&xe->drm, "Failed to create job for VRAM overlap check\n");
> + xe_bb_free(bb, NULL);
> + return PTR_ERR(job);
> + }
> +
> + xe_bb_free(bb, NULL);
> + return 0;
> +}
> +#endif
> diff --git a/drivers/gpu/drm/xe/xe_migrate.h b/drivers/gpu/drm/xe/xe_migrate.h
> index c3a268b01768..a9acc62f78f0 100644
> --- a/drivers/gpu/drm/xe/xe_migrate.h
> +++ b/drivers/gpu/drm/xe/xe_migrate.h
> @@ -182,4 +182,10 @@ static inline void xe_migrate_job_lock_assert(struct xe_exec_queue *q)
> void xe_migrate_job_lock(struct xe_migrate *m, struct xe_exec_queue *q);
> void xe_migrate_job_unlock(struct xe_migrate *m, struct xe_exec_queue *q);
>
> +#if IS_ENABLED(CONFIG_DRM_XE_DEBUG_MEM)
> +int xe_migrate_debug_ccs_overlap(struct xe_migrate *m,
> + struct xe_bo *scratch_bo,
> + bool write_to_ccs);
> +#endif
> +
> #endif
> diff --git a/drivers/gpu/drm/xe/xe_tile_types.h b/drivers/gpu/drm/xe/xe_tile_types.h
> index 0048100ccb72..e1368c04846a 100644
> --- a/drivers/gpu/drm/xe/xe_tile_types.h
> +++ b/drivers/gpu/drm/xe/xe_tile_types.h
> @@ -97,6 +97,10 @@ struct xe_tile {
> * Only main GT has page reclaim list allocations.
> */
> struct xe_sa_manager *reclaim_pool;
> +#if IS_ENABLED(CONFIG_DRM_XE_DEBUG_MEM)
> + /** @mem.memtest_bo: VRAM overlap check BO */
> + struct xe_bo *memtest_bo;
> +#endif
> } mem;
>
> /** @sriov: tile level virtualization data */
> diff --git a/drivers/gpu/drm/xe/xe_vram.c b/drivers/gpu/drm/xe/xe_vram.c
> index 04d831b101bd..dcedd8cfd731 100644
> --- a/drivers/gpu/drm/xe/xe_vram.c
> +++ b/drivers/gpu/drm/xe/xe_vram.c
> @@ -17,8 +17,11 @@
> #include "xe_device.h"
> #include "xe_force_wake.h"
> #include "xe_gt_mcr.h"
> +#include "xe_map.h"
> +#include "xe_migrate.h"
> #include "xe_mmio.h"
> #include "xe_sriov.h"
> +#include "xe_tile.h"
> #include "xe_tile_sriov_vf.h"
> #include "xe_ttm_vram_mgr.h"
> #include "xe_vram.h"
> @@ -407,3 +410,241 @@ resource_size_t xe_vram_region_actual_physical_size(const struct xe_vram_region
> return vram ? vram->actual_physical_size : 0;
> }
> EXPORT_SYMBOL_IF_KUNIT(xe_vram_region_actual_physical_size);
> +
> +#if IS_ENABLED(CONFIG_DRM_XE_DEBUG_MEM)
> +static void memtest_bo_cleanup(void *arg)
> +{
> + struct xe_device *xe = arg;
> +
> + xe_vram_free_memtest_bos(xe);
> +}
> +
> +int xe_vram_reserve_memtest_bo(struct xe_device *xe)
> +{
> + struct xe_tile *tile;
> + u8 id;
> +
> + if (IS_SRIOV_VF(xe))
> + return 0;
> +
> + for_each_tile(tile, xe, id) {
> + u64 vram_size;
> +
> + if (!tile->mem.vram)
> + continue;
> +
> + if (tile->mem.vram->io_size < tile->mem.vram->usable_size) {
> + drm_info(&xe->drm,
> + "Tile %d: Small-BAR system detected, skipping VRAM memtest\n",
> + id);
> + continue;
> + }
> +
> + vram_size = tile->mem.vram->usable_size;
> +
> + tile->mem.memtest_bo = xe_bo_create_pin_map_at_novm(xe, tile, SZ_64K,
> + vram_size - SZ_64K,
> + ttm_bo_type_kernel,
> + XE_BO_FLAG_VRAM_IF_DGFX(tile),
> + 0, false);
> + if (IS_ERR(tile->mem.memtest_bo)) {
> + drm_warn(&xe->drm, "Tile %d: Failed to reserve memtest BO\n", id);
> + tile->mem.memtest_bo = NULL;
> + continue;
> + }
> +
> + drm_info(&xe->drm, "Tile %d: Reserved memtest BO at offset 0x%llx\n",
> + id, vram_size - SZ_64K);
> + }
> +
> + return devm_add_action_or_reset(xe->drm.dev, memtest_bo_cleanup, xe);
> +}
> +
> +void xe_vram_free_memtest_bos(struct xe_device *xe)
> +{
> + struct xe_tile *tile;
> + u8 id;
> +
> + for_each_tile(tile, xe, id) {
> + if (tile->mem.memtest_bo) {
> + xe_bo_unpin_map_no_vm(tile->mem.memtest_bo);
> + tile->mem.memtest_bo = NULL;
> + }
> + }
> +}
> +
> +int xe_vram_memtest(struct xe_device *xe)
> +{
> + struct xe_tile *tile;
> + u8 id;
> + int err = 0;
> +
> + if (IS_SRIOV_VF(xe))
> + return 0;
> +
> + for_each_tile(tile, xe, id) {
> + struct xe_bo *last_page_bo = tile->mem.memtest_bo;
> + struct dma_fence *fence;
> + bool overlap = false;
> + int i;
> + u8 val;
> +
> + if (!last_page_bo || !tile->migrate)
> + continue;
> +
> + drm_info(&xe->drm, "Tile %d: Running VRAM memtest...\n", id);
> +
> + /* CPU write and readback first and last byte of the last page */
> + xe_map_wr(xe, &last_page_bo->vmap, 0, u8, 0xA5);
> + xe_map_wr(xe, &last_page_bo->vmap, SZ_64K - 1, u8, 0x5A);
> +
> + val = xe_map_rd(xe, &last_page_bo->vmap, 0, u8);
> + if (drm_WARN(&xe->drm, val != 0xA5,
> + "Tile %d: CPU memtest failed at offset 0 (expected 0xA5, got 0x%02x)\n",
> + id, val)) {
> + err = -EIO;
> + goto unpin;
> + }
> +
> + val = xe_map_rd(xe, &last_page_bo->vmap, SZ_64K - 1, u8);
> + if (drm_WARN(&xe->drm, val != 0x5A,
> + "Tile %d: CPU memtest failed at offset 65535 (expected 0x5A, got 0x%02x)\n",
> + id, val)) {
> + err = -EIO;
> + goto unpin;
> + }
> +
> + /* Non-CCS access via GPU on the last page */
> + xe_bo_lock(last_page_bo, false);
> + fence = xe_migrate_clear(tile->migrate, last_page_bo,
> + last_page_bo->ttm.resource,
> + XE_MIGRATE_CLEAR_FLAG_BO_DATA);
> + xe_bo_unlock(last_page_bo);
> +
> + if (!IS_ERR(fence)) {
> + dma_fence_wait(fence, false);
> + dma_fence_put(fence);
> + } else {
> + err = PTR_ERR(fence);
> + goto unpin;
> + }
> +
> + val = xe_map_rd(xe, &last_page_bo->vmap, 0, u8);
> + if (drm_WARN(&xe->drm, val != 0x00,
> + "Tile %d: GPU memtest clear failed at offset 0 (expected 0x00, got 0x%02x)\n",
> + id, val)) {
> + err = -EIO;
> + goto unpin;
> + }
> +
> + /*
> + * Check for CCS overlap on the root tile.
> + *
> + * TODO: maybe extend if we ever get multi-tile + CCS. Pay
> + * special attention to the l2 flush below. Currently that is
> + * hard coded to the root tile.
> + */
> + if (!id && xe_device_has_flat_ccs(xe) &&
> + GRAPHICS_VERx100(xe) >= 2000) {
> + struct xe_bo *scratch_bo_before;
> + struct xe_bo *scratch_bo_after;
> +
> + scratch_bo_before = xe_bo_create_pin_map_novm(xe, tile, SZ_64K,
> + ttm_bo_type_kernel,
> + XE_BO_FLAG_VRAM_IF_DGFX(tile),
> + false);
> + if (IS_ERR(scratch_bo_before)) {
> + err = PTR_ERR(scratch_bo_before);
> + goto unpin;
> + }
> +
> + scratch_bo_after = xe_bo_create_pin_map_novm(xe, tile, SZ_64K,
> + ttm_bo_type_kernel,
> + XE_BO_FLAG_VRAM_IF_DGFX(tile),
> + false);
> + if (IS_ERR(scratch_bo_after)) {
> + xe_bo_unpin_map_no_vm(scratch_bo_before);
> + err = PTR_ERR(scratch_bo_after);
> + goto unpin;
> + }
> +
> + /* Save original CCS metadata for PA 0 + */
> + err = xe_migrate_debug_ccs_overlap(tile->migrate, scratch_bo_before, false);
> + if (err) {
> + xe_bo_unpin_map_no_vm(scratch_bo_before);
> + xe_bo_unpin_map_no_vm(scratch_bo_after);
> + goto unpin;
> + }
> +
> + /*
> + * Fill last page. If there is CCS overlap in the last
> + * page this will snag the raw CCS storage.
> + */
> + xe_map_memset(xe, &last_page_bo->vmap, 0, 0x5A, SZ_64K);
> + xe_device_wmb(xe);
> +
> + /*
> + * Global invalidation. Some BMG SKUs will cache the BAR
> + * writes in the GPU side VRAM cache. Make sure above
> + * writes are fully flushed out to VRAM, so this is
> + * hopefully more well behaved with the CCS unit, if
> + * there is indeed CCS overlap with normal VRAM. Since
> + * there is a separate CCS cache, the CCS unit might not
> + * respect the GPU VRAM cache for CCS accesses, so opt
> + * for being super careful here.
> + */
> + xe_device_l2_flush(xe, true);
> +
> + /* Use GPU to clear CCS state for PA 0 */
> + xe_map_memset(xe, &scratch_bo_after->vmap, 0, 0x00, SZ_64K);
> + err = xe_migrate_debug_ccs_overlap(tile->migrate, scratch_bo_after, true);
> + if (err) {
> + xe_bo_unpin_map_no_vm(scratch_bo_before);
> + xe_bo_unpin_map_no_vm(scratch_bo_after);
> + goto unpin;
> + }
> + /*
> + * Global invalidation. Ensure CCS caches really are
> + * nuked and the raw CCS data is visible in VRAM, for
> + * the below access.
> + */
> + xe_device_l2_flush(xe, true);
> +
> + /* Check if last_page_bo was corrupted by the GPU CCS clear */
> + for (i = 0; i < SZ_64K; i += 8) {
> + u64 payload = xe_map_rd(xe, &last_page_bo->vmap, i, u64);
> +
> + if (payload != 0x5A5A5A5A5A5A5A5AULL) {
> + overlap = true;
> + break;
> + }
> + }
> +
> + /* Restore original CCS metadata for PA 0 + */
> + err = xe_migrate_debug_ccs_overlap(tile->migrate, scratch_bo_before, true);
> + if (err)
> + drm_warn(&xe->drm, "Failed to restore CCS metadata\n");
> +
> + xe_bo_unpin_map_no_vm(scratch_bo_before);
> + xe_bo_unpin_map_no_vm(scratch_bo_after);
> + }
> +
> + if (drm_WARN(&xe->drm, overlap,
> + "Tile %d: VRAM bounds overlap CCS region! VRAM sizing is incorrect.\n",
> + id)) {
> + err = -EINVAL;
> + goto unpin;
> + }
> +
> + drm_info(&xe->drm, "Tile %d: VRAM memtest completed.\n", id);
> +
> +unpin:
> + if (err)
> + break;
> + }
> +
> + xe_vram_free_memtest_bos(xe);
> +
> + return err;
> +}
> +#endif
> diff --git a/drivers/gpu/drm/xe/xe_vram.h b/drivers/gpu/drm/xe/xe_vram.h
> index dd1c8bf17922..38425d81f777 100644
> --- a/drivers/gpu/drm/xe/xe_vram.h
> +++ b/drivers/gpu/drm/xe/xe_vram.h
> @@ -23,4 +23,14 @@ resource_size_t xe_vram_region_dpa_base(const struct xe_vram_region *vram);
> resource_size_t xe_vram_region_usable_size(const struct xe_vram_region *vram);
> resource_size_t xe_vram_region_actual_physical_size(const struct xe_vram_region *vram);
>
> +#if IS_ENABLED(CONFIG_DRM_XE_DEBUG_MEM)
> +int xe_vram_reserve_memtest_bo(struct xe_device *xe);
> +void xe_vram_free_memtest_bos(struct xe_device *xe);
> +int xe_vram_memtest(struct xe_device *xe);
> +#else
> +static inline int xe_vram_reserve_memtest_bo(struct xe_device *xe) { return 0; }
> +static inline void xe_vram_free_memtest_bos(struct xe_device *xe) {}
> +static inline int xe_vram_memtest(struct xe_device *xe) { return 0; }
> +#endif
> +
> #endif
> --
> 2.55.0
>
next prev parent reply other threads:[~2026-09-02 19:43 UTC|newest]
Thread overview: 14+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-02 12:41 [PATCH v3 0/5] VRAM health check + CCS fix Matthew Auld
2026-09-02 12:41 ` [PATCH v3 1/5] drm/xe/migrate: support 4K PTEs for identity map Matthew Auld
2026-09-02 12:41 ` [PATCH v3 2/5] drm/xe/vram: report FLAT_CCS base misalignment Matthew Auld
2026-09-02 19:42 ` Matthew Brost
2026-09-02 12:41 ` [PATCH v3 3/5] drm/xe/vram: revamp CPU VRAM mapping Matthew Auld
2026-09-02 13:12 ` sashiko-bot
2026-09-02 14:36 ` Matthew Auld
2026-09-02 19:51 ` Matthew Brost
2026-09-02 12:41 ` [PATCH v3 4/5] drm/xe: add force option for global invalidation Matthew Auld
2026-09-02 12:41 ` [PATCH v3 5/5] drm/xe/vram: add early VRAM health check Matthew Auld
2026-09-02 19:43 ` Matthew Brost [this message]
2026-09-02 13:28 ` ✓ CI.KUnit: success for VRAM health check + CCS fix (rev3) Patchwork
2026-09-02 14:23 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-03 0:36 ` ✗ Xe.CI.FULL: failure " Patchwork
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aph8VNG33AzI/bcV@gsse-cloud1.jf.intel.com \
--to=matthew.brost@intel.com \
--cc=intel-xe@lists.freedesktop.org \
--cc=matthew.auld@intel.com \
--cc=rodrigo.vivi@intel.com \
--cc=stuart.summers@intel.com \
--cc=thomas.hellstrom@linux.intel.com \
--cc=torvalds@linux-foundation.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox