From: Matthew Brost <matthew.brost@intel.com>
To: Arvind Yadav <arvind.yadav@intel.com>
Cc: <intel-xe@lists.freedesktop.org>,
<himal.prasad.ghimiray@intel.com>,
<thomas.hellstrom@linux.intel.com>,
Sashiko <sashiko-bot@kernel.org>
Subject: Re: [PATCH] drm/xe/vm: Handle purged BOs in VM snapshots
Date: Thu, 3 Sep 2026 12:16:15 -0700 [thread overview]
Message-ID: <apnHf2QWL+FOYiD7@gsse-cloud1.jf.intel.com> (raw)
In-Reply-To: <20260903090630.3857182-1-arvind.yadav@intel.com>
On Thu, Sep 03, 2026 at 02:36:30PM +0530, Arvind Yadav wrote:
> A purged BO can remain in a VM after its TTM resource has been freed.
> xe_vm_snapshot_capture() can then dereference a NULL resource while
> capturing a devcoredump.
>
> Record -ENODATA for the affected VMA instead of dereferencing
> the resource.
>
> Fixes: ad9843aac91a ("drm/xe/madvise: Implement purgeable buffer object support")
> Reported-by: Sashiko <sashiko-bot@kernel.org>
> Closes: https://sashiko.dev/#/patchset/20260831064942.315720-17-tejas.upadhyay%40intel.com?part=4
> Cc: Thomas Hellström <thomas.hellstrom@linux.intel.com>
> Cc: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
> Cc: Matthew Brost <matthew.brost@intel.com>
> Signed-off-by: Arvind Yadav <arvind.yadav@intel.com>
> ---
> drivers/gpu/drm/xe/xe_vm.c | 28 +++++++++++++++++-----------
> 1 file changed, 17 insertions(+), 11 deletions(-)
>
> diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c
> index 19b3d0be7928..77a851d9e373 100644
> --- a/drivers/gpu/drm/xe/xe_vm.c
> +++ b/drivers/gpu/drm/xe/xe_vm.c
> @@ -4666,17 +4666,23 @@ struct xe_vm_snapshot *xe_vm_snapshot_capture(struct xe_vm *vm)
> snap->snap[i].cpu_caching = bo->cpu_caching;
> snap->snap[i].bo = xe_bo_get(bo);
> snap->snap[i].bo_ofs = xe_vma_bo_offset(vma);
> - switch (bo->ttm.resource->mem_type) {
> - case XE_PL_SYSTEM:
> - case XE_PL_TT:
> - snap->snap[i].uapi_mem_region = 0;
> - break;
> - case XE_PL_VRAM0:
> - snap->snap[i].uapi_mem_region = 1;
> - break;
> - case XE_PL_VRAM1:
> - snap->snap[i].uapi_mem_region = 2;
> - break;
> + snap->snap[i].uapi_mem_region = -1;
> +
> + if (!bo->ttm.resource) {
I think technically this can change at any moment, so this is a TOCTOU
issue with the switch below.
The value is only stable while holding the dma-resv lock, which we can't
take here because we're in the signaling path.
To at least avoid a NULL pointer dereference, could we do something
like:
res = READ_ONCE(bo->ttm.resource);
if (!res)
error;
else
switch (res->mem_type)
KASAN could still complain if res is freed after the read, but the
kernel wouldn't explode, which is objectively better than both the
current situation and this patch.
Maybe someone has a better solution that fully closes this race? I'm
actually spotting a few other issues in xe_vm_snapshot_capture_delayed()
where eviction can race as well and those should be cleaned up too.
Perhaps we should open a broader Jira covering all devcoredump paths
that access BOs which may be moving, and clean up all of these issues in
a separate series.
Matt
> + snap->snap[i].data = ERR_PTR(-ENODATA);
> + } else {
> + switch (bo->ttm.resource->mem_type) {
> + case XE_PL_SYSTEM:
> + case XE_PL_TT:
> + snap->snap[i].uapi_mem_region = 0;
> + break;
> + case XE_PL_VRAM0:
> + snap->snap[i].uapi_mem_region = 1;
> + break;
> + case XE_PL_VRAM1:
> + snap->snap[i].uapi_mem_region = 2;
> + break;
> + }
> }
> } else if (xe_vma_is_userptr(vma)) {
> struct mm_struct *mm =
> --
> 2.43.0
>
next prev parent reply other threads:[~2026-09-03 19:16 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-03 9:06 [PATCH] drm/xe/vm: Handle purged BOs in VM snapshots Arvind Yadav
2026-09-03 9:15 ` ✓ CI.KUnit: success for " Patchwork
2026-09-03 10:01 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-03 19:16 ` Matthew Brost [this message]
2026-09-04 4:46 ` [PATCH] " Yadav, Arvind
2026-09-03 20:51 ` ✗ Xe.CI.FULL: failure for " Patchwork
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=apnHf2QWL+FOYiD7@gsse-cloud1.jf.intel.com \
--to=matthew.brost@intel.com \
--cc=arvind.yadav@intel.com \
--cc=himal.prasad.ghimiray@intel.com \
--cc=intel-xe@lists.freedesktop.org \
--cc=sashiko-bot@kernel.org \
--cc=thomas.hellstrom@linux.intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.