From: "Ghimiray, Himal Prasad" <himal.prasad.ghimiray@intel.com>
To: Tejas Upadhyay <tejas.upadhyay@intel.com>,
<intel-xe@lists.freedesktop.org>
Cc: <rodrigo.vivi@intel.com>, <michal.wajdeczko@intel.com>
Subject: Re: [PATCH V18 14/14] drm/xe: Add fault-inject based VRAM page offline injection
Date: Thu, 27 Aug 2026 12:40:00 +0530 [thread overview]
Message-ID: <2c698f6d-0a30-4bc9-9ca9-e12a2e01ff67@intel.com> (raw)
In-Reply-To: <20260826135136.204044-30-tejas.upadhyay@intel.com>
On 26-08-2026 19:21, Tejas Upadhyay wrote:
> Add a fault-inject based debugfs interface for testing VRAM page
> offlining. This replaces the previous standalone debugfs approach
> with the standard kernel fault-inject infrastructure.
>
> Two debugfs entries are created under the xe debugfs root for
> CRI platforms:
> - inject_mempage_offline/: Standard fault-inject knobs (probability,
> times, interval, etc.) created by fault_create_debugfs_attr().
> Without CONFIG_FAULT_INJECTION_DEBUG_FS, the stub returns
> ERR_PTR(-ENODEV) and no knobs are created, making the trigger
> effectively a no-op.
> - inject_mempage_offline_trigger: Write a PFN value to inject a
> specific page, or write "0" to auto-pick the last unallocated
> VRAM page
>
> The trigger accepts:
> - "0" : auto-pick last unallocated page
> - "0xPFN" : inject fault at a specific PFN address
>
> Usage:
> echo 100 > inject_mempage_offline/probability
> echo 1 > inject_mempage_offline/times
> echo 0 > inject_mempage_offline_trigger
>
> probability: likelihood of should_fail() returning true (0-100)
> times: number of times injection is allowed (-1 for unlimited)
>
> v5(Sashiko):
> - exclude SRIOV and remove dpa_base addition, already absolute dpa
> v4(Himal):
> - Use xe_fault_mempage_offline() instead of IS_ENABLED() +
> direct should_fail(). CONFIG_FAULT_INJECTION_DEBUG_FS is now
> an implicit requirement for the trigger to function.
> v3(Himal):
> - Use FAULT_ACTION
> v2(sashiko):
> - use cond_resched()
> - validate input first and fix addr < 0 case
> - validate vr, move block, found var as local to scope_guard
Seems reclaiming the page at addr is left to unbind/bind of driver.
Need doc for that? Multiple fault injection at random offset might make
contigous vram allocation problematic.
>
> Signed-off-by: Tejas Upadhyay <tejas.upadhyay@intel.com>
> ---
> drivers/gpu/drm/xe/xe_debugfs.c | 47 +++++++++++++++++++++++++
> drivers/gpu/drm/xe/xe_debugfs.h | 2 ++
> drivers/gpu/drm/xe/xe_ttm_vram_mgr.c | 52 ++++++++++++++++++++++++++++
> drivers/gpu/drm/xe/xe_ttm_vram_mgr.h | 1 +
> 4 files changed, 102 insertions(+)
>
> diff --git a/drivers/gpu/drm/xe/xe_debugfs.c b/drivers/gpu/drm/xe/xe_debugfs.c
> index e19eafffbb08..003abf4af2e3 100644
> --- a/drivers/gpu/drm/xe/xe_debugfs.c
> +++ b/drivers/gpu/drm/xe/xe_debugfs.c
> @@ -45,12 +45,18 @@
> DECLARE_FAULT_ATTR(gt_reset_failure);
> DECLARE_FAULT_ATTR(inject_csc_hw_error);
> DECLARE_FAULT_ATTR(wedge_cold_reset);
> +DECLARE_FAULT_ATTR(inject_mempage_offline);
>
> static bool csc_hw_error_available(struct xe_device *xe)
> {
> return !IS_SRIOV_VF(xe) && xe->info.platform == XE_BATTLEMAGE;
> }
>
> +static bool is_crescent_island_pf(struct xe_device *xe)
> +{
> + return !IS_SRIOV_VF(xe) && xe->info.platform == XE_CRESCENTISLAND;
> +}
> +
> /*
> * Fault injection table. Each entry registers a debugfs attribute; add a
> * matching FAULT_ACTION() below for every entry added here.
> @@ -67,6 +73,9 @@ static struct {
> .is_visible = csc_hw_error_available },
> { .name = "wedge_cold_reset",
> .attr = &wedge_cold_reset },
> + { .name = "inject_mempage_offline",
> + .attr = &inject_mempage_offline,
> + .is_visible = is_crescent_island_pf },
> };
>
> /*
> @@ -82,6 +91,39 @@ bool xe_fault_##name(void) \
> FAULT_ACTION(gt_reset, gt_reset_failure)
> FAULT_ACTION(csc_hw_error, inject_csc_hw_error)
> FAULT_ACTION(wedge_cold_reset, wedge_cold_reset)
> +FAULT_ACTION(mempage_offline, inject_mempage_offline)
> +
> +static ssize_t inject_mempage_offline_trigger(struct file *f,
> + const char __user *ubuf,
> + size_t size, loff_t *pos)
> +{
> + struct xe_device *xe = file_inode(f)->i_private;
> + struct xe_tile *tile = xe_device_get_root_tile(xe);
> + struct xe_vram_region *vr = tile->mem.vram;
> + u64 pfn;
> + int ret;
> +
> + if (!vr)
> + return -ENODEV;
> +
> + ret = kstrtou64_from_user(ubuf, size, 0, &pfn);
> + if (ret)
> + return ret;
> +
> + if (!xe_fault_mempage_offline())
> + return size;
> +
> + if (pfn == 0)
> + return xe_ttm_vram_inject_fault(xe) ?: size;
> +
> + /* User provided PFN - convert to DPA and inject */
> + return xe_ttm_vram_handle_addr_fault(xe, pfn << PAGE_SHIFT) ?: size;
> +}
> +
> +static const struct file_operations inject_mempage_offline_fops = {
> + .owner = THIS_MODULE,
> + .write = inject_mempage_offline_trigger,
> +};
>
> static void xe_fault_inject_debugfs_register(struct xe_device *xe,
> struct dentry *root)
> @@ -96,6 +138,11 @@ static void xe_fault_inject_debugfs_register(struct xe_device *xe,
> fault_create_debugfs_attr(xe_fault_inject_entry[i].name, root,
> xe_fault_inject_entry[i].attr);
> }
> +
> + if (is_crescent_island_pf(xe)) {
> + debugfs_create_file("inject_mempage_offline_trigger", 0200,
> + root, xe, &inject_mempage_offline_fops);
> + }
> }
>
> static void read_residency_counter(struct xe_device *xe, struct xe_mmio *mmio,
> diff --git a/drivers/gpu/drm/xe/xe_debugfs.h b/drivers/gpu/drm/xe/xe_debugfs.h
> index 0dcd28fd7dc0..88d91c78036b 100644
> --- a/drivers/gpu/drm/xe/xe_debugfs.h
> +++ b/drivers/gpu/drm/xe/xe_debugfs.h
> @@ -14,11 +14,13 @@ struct xe_device;
> bool xe_fault_gt_reset(void);
> bool xe_fault_csc_hw_error(void);
> bool xe_fault_wedge_cold_reset(void);
> +bool xe_fault_mempage_offline(void);
> void xe_debugfs_register(struct xe_device *xe);
> #else
> static inline bool xe_fault_gt_reset(void) { return false; }
> static inline bool xe_fault_csc_hw_error(void) { return false; }
> static inline bool xe_fault_wedge_cold_reset(void) { return false; }
> +static inline bool xe_fault_mempage_offline(void) { return false; }
> static inline void xe_debugfs_register(struct xe_device *xe) { }
> #endif
>
> diff --git a/drivers/gpu/drm/xe/xe_ttm_vram_mgr.c b/drivers/gpu/drm/xe/xe_ttm_vram_mgr.c
> index 8f583f1631bf..c54ad017725f 100644
> --- a/drivers/gpu/drm/xe/xe_ttm_vram_mgr.c
> +++ b/drivers/gpu/drm/xe/xe_ttm_vram_mgr.c
> @@ -895,6 +895,58 @@ int xe_ttm_vram_handle_addr_fault(struct xe_device *xe, u64 addr)
> }
> EXPORT_SYMBOL(xe_ttm_vram_handle_addr_fault);
>
> +/**
> + * xe_ttm_vram_inject_fault - Inject a VRAM page fault for testing
> + * @xe: xe device instance
> + *
> + * Picks the last unallocated VRAM page and reports it as faulted
> + * via xe_ttm_vram_handle_addr_fault(). Used by the fault-inject
> + * debugfs interface for testing page offlining.
> + *
> + * Return: 0 on success, negative error code on failure.
> + */
> +int xe_ttm_vram_inject_fault(struct xe_device *xe)
> +{
> + struct xe_tile *tile = xe_device_get_root_tile(xe);
> + struct xe_vram_region *vr = tile->mem.vram;
> + struct xe_ttm_vram_mgr *vram_mgr = &vr->ttm;
> + struct gpu_buddy *mm = &vram_mgr->mm;
> + u64 addr;
> +
> + if (vr->actual_physical_size < SZ_4K)
> + return -ENOSPC;
> +
> + addr = vr->actual_physical_size - SZ_4K;
> + while (addr < vr->actual_physical_size) {
> + struct gpu_buddy_block *block;
> + bool found = false;
> +
> + scoped_guard(mutex, &vram_mgr->lock) {
> + block = gpu_buddy_allocated_addr_to_block(mm, addr);
> + if (!block)
> + found = true;
> + }
> +
> + /*
> + * Intentional race window: xe_ttm_vram_handle_addr_fault()
> + * re-acquires vram_mgr->lock internally, so we cannot hold
> + * it here. A concurrent allocation claiming this page between
> + * the two calls is an acceptable false negative for this
> + * test-only path.
> + */
> + if (found)
> + return xe_ttm_vram_handle_addr_fault(xe, addr + vr->dpa_base);
> +
> + cond_resched();
> + if (addr == 0)
> + break;
> + addr -= SZ_4K;
> + }
> +
> + return -ENOSPC;
> +}
> +EXPORT_SYMBOL(xe_ttm_vram_inject_fault);
> +
> static int vram_bad_pages_show(struct seq_file *m, void *unused)
> {
> struct xe_device *xe = m->private;
> diff --git a/drivers/gpu/drm/xe/xe_ttm_vram_mgr.h b/drivers/gpu/drm/xe/xe_ttm_vram_mgr.h
> index f354c26c4257..8878e36292b2 100644
> --- a/drivers/gpu/drm/xe/xe_ttm_vram_mgr.h
> +++ b/drivers/gpu/drm/xe/xe_ttm_vram_mgr.h
> @@ -33,6 +33,7 @@ void xe_ttm_vram_get_used(struct ttm_resource_manager *man,
> u64 *used, u64 *used_visible);
>
> int xe_ttm_vram_handle_addr_fault(struct xe_device *xe, u64 addr);
> +int xe_ttm_vram_inject_fault(struct xe_device *xe);
> void xe_ttm_vram_debugfs_init(struct xe_device *xe, struct dentry *root);
> static inline struct xe_ttm_vram_mgr_resource *
> to_xe_ttm_vram_mgr_resource(struct ttm_resource *res)
next prev parent reply other threads:[~2026-08-27 7:10 UTC|newest]
Thread overview: 56+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-26 13:51 [PATCH V18 00/14] Add memory page offlining support Tejas Upadhyay
2026-08-26 13:51 ` [PATCH V18 01/14] drm/xe: Link VRAM object with gpu buddy Tejas Upadhyay
2026-08-26 22:31 ` Andi Shyti
2026-08-26 13:51 ` [PATCH V18 02/14] drm/xe: Link LRC BO and its execution Queue Tejas Upadhyay
2026-08-26 22:34 ` Andi Shyti
2026-08-26 13:51 ` [PATCH V18 03/14] drm/xe: Extend BO purge to handle vram pages as well Tejas Upadhyay
2026-08-26 14:07 ` sashiko-bot
2026-08-26 22:42 ` Andi Shyti
2026-08-27 6:17 ` Upadhyay, Tejas
2026-08-27 14:40 ` Andi Shyti
2026-08-27 14:48 ` Upadhyay, Tejas
2026-08-28 5:25 ` Upadhyay, Tejas
2026-08-28 7:39 ` Andi Shyti
2026-08-28 17:29 ` Upadhyay, Tejas
2026-08-26 13:51 ` [PATCH V18 04/14] drm/xe/bo: Make xe_bo_is_user() public Tejas Upadhyay
2026-08-26 22:44 ` Andi Shyti
2026-08-26 13:51 ` [PATCH V18 05/14] drm/xe: Guard teardown paths against purged BOs Tejas Upadhyay
2026-08-26 14:12 ` sashiko-bot
2026-08-27 6:08 ` Ghimiray, Himal Prasad
2026-08-27 8:27 ` Upadhyay, Tejas
2026-08-26 13:51 ` [PATCH V18 06/14] drm/xe/vram: Extract buddy alloc and free helpers Tejas Upadhyay
2026-08-26 22:50 ` Andi Shyti
2026-08-26 13:51 ` [PATCH V18 07/14] drm/xe/vram: Add page offline data structures and lifecycle Tejas Upadhyay
2026-08-26 23:09 ` Andi Shyti
2026-08-27 6:19 ` Ghimiray, Himal Prasad
2026-08-26 13:51 ` [PATCH V18 08/14] drm/xe/vram: Add VRAM page offline fault handler Tejas Upadhyay
2026-08-26 14:05 ` sashiko-bot
2026-08-26 13:51 ` [PATCH V18 09/14] drm/xe/configfs: Add bad_page_reservation attribute Tejas Upadhyay
2026-08-27 6:42 ` Ghimiray, Himal Prasad
2026-08-27 15:00 ` Michal Wajdeczko
2026-08-28 17:48 ` Upadhyay, Tejas
2026-08-26 13:51 ` [PATCH V18 10/14] drm/xe/ras: Cache bad_page_reservation policy at init Tejas Upadhyay
2026-08-26 14:11 ` sashiko-bot
2026-08-27 6:45 ` Ghimiray, Himal Prasad
2026-08-26 13:51 ` [PATCH V18 11/14] drm/xe/vram: Check bad_page_reservation policy in fault handler Tejas Upadhyay
2026-08-26 14:08 ` sashiko-bot
2026-08-27 6:46 ` Ghimiray, Himal Prasad
2026-08-27 15:04 ` Michal Wajdeczko
2026-09-02 7:14 ` Mallesh, Koujalagi
2026-08-26 13:51 ` [PATCH V18 12/14] drm/xe: Expose bad VRAM pages via debugfs Tejas Upadhyay
2026-08-26 14:13 ` sashiko-bot
2026-08-27 15:16 ` Michal Wajdeczko
2026-08-28 19:06 ` Upadhyay, Tejas
2026-08-31 3:28 ` Iddamsetty, Aravind
2026-08-28 15:04 ` Rodrigo Vivi
2026-08-26 13:51 ` [PATCH V18 13/14] drm/xe/uapi: Expose ban reason in EXEC_QUEUE_GET_PROPERTY_BAN Tejas Upadhyay
2026-08-26 14:20 ` sashiko-bot
2026-08-27 18:26 ` Andi Shyti
2026-08-28 5:31 ` Upadhyay, Tejas
2026-08-26 13:51 ` [PATCH V18 14/14] drm/xe: Add fault-inject based VRAM page offline injection Tejas Upadhyay
2026-08-27 7:10 ` Ghimiray, Himal Prasad [this message]
2026-08-27 8:23 ` Upadhyay, Tejas
2026-08-26 14:37 ` ✗ CI.checkpatch: warning for Add memory page offlining support (rev21) Patchwork
2026-08-26 14:39 ` ✓ CI.KUnit: success " Patchwork
2026-08-26 15:21 ` ✓ Xe.CI.BAT: " Patchwork
2026-08-26 19:01 ` ✓ Xe.CI.FULL: " Patchwork
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=2c698f6d-0a30-4bc9-9ca9-e12a2e01ff67@intel.com \
--to=himal.prasad.ghimiray@intel.com \
--cc=intel-xe@lists.freedesktop.org \
--cc=michal.wajdeczko@intel.com \
--cc=rodrigo.vivi@intel.com \
--cc=tejas.upadhyay@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.