Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Tejas Upadhyay" <tejas.upadhyay@intel.com>
Cc: intel-xe@lists.freedesktop.org
Subject: Re: [PATCH V21 04/15] drm/xe: Handle NULL resource and allow purging of VRAM pages
Date: Thu, 03 Sep 2026 17:08:55 +0000	[thread overview]
Message-ID: <20260903170856.060271F000E9@smtp.kernel.org> (raw)
In-Reply-To: <20260903161553.528932-21-tejas.upadhyay@intel.com>

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:

Pre-existing issues:
- [High] Missing NULL check for bo->ttm.resource in xe_vm_snapshot_capture(), which will cause a kernel panic when devcoredump attempts to snapshot a purged BO.
- [High] xe_bo_migrate() and xe_bo_is_mem_type() unconditionally dereference bo->ttm.resource without checking for NULL, leading to a crash when called on purged BOs.
--

commit 1ae4c5706e8f483f5e1f01922281841438651876
Author: Tejas Upadhyay <tejas.upadhyay@intel.com>

drm/xe: Handle NULL resource and allow purging of VRAM pages

Add NULL checks for the buffer object's resource pointer in
xe_bo_is_vram() and xe_bo_is_stolen() to prevent potential NULL
pointer dereferences when checking memory location during page
offlining.

Additionally, remove the `!ttm_bo->ttm` check from
xe_ttm_bo_purge(). This restriction is stripped away to
explicitly allow the driver to purge VRAM pages during the
memory page offline sequence.

> diff --git a/drivers/gpu/drm/xe/xe_bo.c b/drivers/gpu/drm/xe/xe_bo.c
> index 8f93e09e51bb8..d30a8411d2a33 100644
> --- a/drivers/gpu/drm/xe/xe_bo.c
> +++ b/drivers/gpu/drm/xe/xe_bo.c

[ ... ]

> @@ -929,9 +932,6 @@ int xe_ttm_bo_purge(struct ttm_buffer_object *ttm_bo, struct ttm_operation_ctx *
>  
>  	xe_bo_assert_held(bo);
>  
> -	if (!ttm_bo->ttm)
> -		return 0;
> -

[Severity: High]
This isn't a bug introduced by this patch, but removing this restriction allows
VRAM BOs to be purged, which exposes a missing NULL check in
xe_vm_snapshot_capture().

If a userspace application maps a BO in a VM, marks its VMA as DONTNEED, and
memory pressure purges the BO, bo->ttm.resource becomes NULL.

If a GPU hang then occurs, devcoredump is triggered and xe_vm_snapshot_capture()
iterates over the VM's vma_list, unconditionally accessing
bo->ttm.resource->mem_type:

drivers/gpu/drm/xe/xe_vm.c:xe_vm_snapshot_capture() {
	...
	drm_gpuvm_for_each_va(gpuva, &vm->gpuvm) {
		...
		if (bo) {
			snap->snap[i].cpu_caching = bo->cpu_caching;
			snap->snap[i].bo = xe_bo_get(bo);
			snap->snap[i].bo_ofs = xe_vma_bo_offset(vma);
			switch (bo->ttm.resource->mem_type) {
			case XE_PL_SYSTEM:
			...
}

Will this cause a kernel panic when devcoredump attempts to snapshot a purged
BO?


[Severity: High]
This is another pre-existing issue, but expanding purgeability to VRAM BOs
makes it reachable for them as well.

If a BO is marked DONTNEED and purged via memory pressure
(bo->ttm.resource = NULL), and then utilized in a path that triggers migration
(such as pinning it to a KMS plane or exporting/pinning it via dma-buf), both
xe_bo_migrate() and xe_bo_is_mem_type() unconditionally dereference
bo->ttm.resource.

drivers/gpu/drm/xe/xe_bo.c:xe_bo_migrate() {
	...
	xe_bo_assert_held(bo);
	tctx = tctx ? tctx : &ctx;

	if (bo->ttm.resource->mem_type == mem_type)
		return 0;
	...
}

drivers/gpu/drm/xe/xe_bo.h:xe_bo_is_mem_type() {
	...
	xe_bo_assert_held(bo);
	return bo->ttm.resource->mem_type == mem_type;
}

Can this lead to a kernel panic when these functions are called on purged BOs?

>  	if (!xe_bo_madv_is_dontneed(bo))
>  		return 0;
>

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260903161553.528932-17-tejas.upadhyay@intel.com?part=4

  reply	other threads:[~2026-09-03 17:08 UTC|newest]

Thread overview: 25+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-03 16:15 [PATCH V21 00/15] Add memory page offlining support Tejas Upadhyay
2026-09-03 16:15 ` [PATCH V21 01/15] drm/xe: Link VRAM resource allocation with gpu buddy Tejas Upadhyay
2026-09-03 16:15 ` [PATCH V21 02/15] drm/xe: Link LRC BO and its execution queue with safe lifetime rules Tejas Upadhyay
2026-09-03 16:15 ` [PATCH V21 03/15] drm/xe: Export xe_ttm_bo_purge() Tejas Upadhyay
2026-09-03 16:15 ` [PATCH V21 04/15] drm/xe: Handle NULL resource and allow purging of VRAM pages Tejas Upadhyay
2026-09-03 17:08   ` sashiko-bot [this message]
2026-09-03 16:15 ` [PATCH V21 05/15] drm/xe/bo: Make xe_bo_is_user() public Tejas Upadhyay
2026-09-03 16:15 ` [PATCH V21 06/15] drm/xe: Guard teardown paths against purged BOs Tejas Upadhyay
2026-09-03 16:55   ` sashiko-bot
2026-09-03 16:16 ` [PATCH V21 07/15] drm/xe/vram: Extract buddy allocation and free helpers Tejas Upadhyay
2026-09-03 16:16 ` [PATCH V21 08/15] drm/xe/vram: Add page offline data structures and lifecycle Tejas Upadhyay
2026-09-03 16:16 ` [PATCH V21 09/15] drm/xe/vram: Add VRAM page offline fault handler Tejas Upadhyay
2026-09-03 17:23   ` sashiko-bot
2026-09-03 16:16 ` [PATCH V21 10/15] drm/xe/configfs: Add disable_vram_page_offline attribute Tejas Upadhyay
2026-09-03 16:16 ` [PATCH V21 11/15] drm/xe/ras: Cache disable_vram_page_offline policy at init Tejas Upadhyay
2026-09-03 16:16 ` [PATCH V21 12/15] drm/xe/vram: Check disable_vram_page_offline policy in fault handler Tejas Upadhyay
2026-09-03 16:16 ` [PATCH V21 13/15] drm/xe: Expose bad VRAM pages via debugfs Tejas Upadhyay
2026-09-03 16:16 ` [PATCH V21 14/15] drm/xe/uapi: Expose ban reason in EXEC_QUEUE_GET_PROPERTY_BAN Tejas Upadhyay
2026-09-03 17:46   ` sashiko-bot
2026-09-03 16:16 ` [PATCH V21 15/15] drm/xe: Add fault-inject based VRAM page offline injection Tejas Upadhyay
2026-09-03 16:23 ` ✗ CI.checkpatch: warning for Add memory page offlining support (rev25) Patchwork
2026-09-03 16:25 ` ✓ CI.KUnit: success " Patchwork
2026-09-03 17:46 ` ✗ Xe.CI.BAT: failure " Patchwork
2026-09-04  4:33 ` [PATCH V21 00/15] Add memory page offlining support Matthew Brost
2026-09-04  5:18 ` ✗ Xe.CI.FULL: failure for Add memory page offlining support (rev25) Patchwork

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260903170856.060271F000E9@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=sashiko-reviews@lists.linux.dev \
    --cc=tejas.upadhyay@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox