dri-devel Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH 1/2] RFC: drm/xe: Fix pinned list UAF panic
@ 2026-10-06 17:57 FNU VISHWANATHA
  2026-10-06 17:57 ` [PATCH 2/2] RFC: drm/xe: Prevent pinned_link double add FNU VISHWANATHA
  2026-10-06 18:15 ` [PATCH 1/2] RFC: drm/xe: Fix pinned list UAF panic sashiko-bot
  0 siblings, 2 replies; 4+ messages in thread
From: FNU VISHWANATHA @ 2026-10-06 17:57 UTC (permalink / raw)
  To: intel-xe; +Cc: dri-devel, Kornel Dulęba

From: Kornel Dulęba <korneld@google.com>

A kernel panic caused by a list corruption was observed when the IPU6
driver is running under heavy load. The panic would reproduce
consistently within a few minutes of the test suite running, though
never during the same test. The corruption would always happen in
list_add_tail called from xe_bo_pin_external:

list_add corruption. prev->next should be next (ffff8bb9862912c8), but was ffff8bb9e811e2a8. (prev=ffff8bb9e811e2a8)

Having prev->next == prev, suggests that node was re-initialized without
being removed from the list first. This indicates that the BO might have
been destroyed while it was still pinned.
Looking into xe_ttm_bo_destroy and its caller ttm_bo_release, reveals
that the latter supports a case where bo->pin_count > 0, albeit with a
WARN_ON_ONCE. Since ttm_bo_release sets bo->pin_count to 0 in that case,
add a complementary change to also remove bo->pinned_link from the list.
With this applied, the test progresses much further, with the newly
added warning seen in dmesg.

Test: atest CtsCameraTestCases

Signed-off-by: Kornel Dulęba <korneld@google.com>
---
 drivers/gpu/drm/xe/xe_bo.c | 5 +++++
 1 file changed, 5 insertions(+)

diff --git a/drivers/gpu/drm/xe/xe_bo.c b/drivers/gpu/drm/xe/xe_bo.c
index f2634d048cd4..c5989db950f7 100644
--- a/drivers/gpu/drm/xe/xe_bo.c
+++ b/drivers/gpu/drm/xe/xe_bo.c
@@ -1688,6 +1688,11 @@ static void xe_ttm_bo_destroy(struct ttm_buffer_object *ttm_bo)
 	if (bo->parent_obj)
 		xe_bo_put(bo->parent_obj);
 
+	spin_lock(&xe->pinned.lock);
+	if (WARN_ON_ONCE(!list_empty(&bo->pinned_link)))
+		list_del_init(&bo->pinned_link);
+	spin_unlock(&xe->pinned.lock);
+
 	mutex_lock(&xe->mem_access.vram_userfault.lock);
 	if (!list_empty(&bo->vram_userfault_link))
 		list_del(&bo->vram_userfault_link);
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 4+ messages in thread

* [PATCH 2/2] RFC: drm/xe: Prevent pinned_link double add
  2026-10-06 17:57 [PATCH 1/2] RFC: drm/xe: Fix pinned list UAF panic FNU VISHWANATHA
@ 2026-10-06 17:57 ` FNU VISHWANATHA
  2026-10-06 18:16   ` sashiko-bot
  2026-10-06 18:15 ` [PATCH 1/2] RFC: drm/xe: Fix pinned list UAF panic sashiko-bot
  1 sibling, 1 reply; 4+ messages in thread
From: FNU VISHWANATHA @ 2026-10-06 17:57 UTC (permalink / raw)
  To: intel-xe; +Cc: dri-devel, Kornel Dulęba

From: Kornel Dulęba <korneld@google.com>

When running CtsCameraTestCases on a memory constrained, 4GB device the
following list corruption oops was consistently observed:

list_del corruption. prev->next should be ffff9ebe1ce42ea8, but was ffff9ebdf741fea8. (prev=ffff9ebe1ce43aa8)

After adding printfs in multiple places it was concluded that
bo->pinned_link was already a part of a list when xe_bo_pin_external
tried adding it to xe->pinned.late.external.
To fix this add a list_empty check and a WARN_ONCE in case the list is
in fact not empty.
With this addition the CtsCameraTestCases can now complete with no
kernel crashes.

Test: atest CtsCameraTestCases
Signed-off-by: Kornel Dulęba <korneld@google.com>
---
 drivers/gpu/drm/xe/xe_bo.c | 9 ++++++++-
 1 file changed, 8 insertions(+), 1 deletion(-)

diff --git a/drivers/gpu/drm/xe/xe_bo.c b/drivers/gpu/drm/xe/xe_bo.c
index c5989db950f7..cc5a4ff64791 100644
--- a/drivers/gpu/drm/xe/xe_bo.c
+++ b/drivers/gpu/drm/xe/xe_bo.c
@@ -2762,7 +2762,14 @@ int xe_bo_pin_external(struct xe_bo *bo, bool in_place, struct drm_exec *exec)
 		}
 
 		spin_lock(&xe->pinned.lock);
-		list_add_tail(&bo->pinned_link, &xe->pinned.late.external);
+		if (list_empty(&bo->pinned_link)) {
+			list_add_tail(&bo->pinned_link,
+				      &xe->pinned.late.external);
+		} else {
+			WARN_ONCE(true,
+				  "pinned_link already added to a list, bo->ttm.pin_count: %d\n",
+				  bo->ttm.pin_count);
+		}
 		spin_unlock(&xe->pinned.lock);
 	}
 
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 4+ messages in thread

* Re: [PATCH 1/2] RFC: drm/xe: Fix pinned list UAF panic
  2026-10-06 17:57 [PATCH 1/2] RFC: drm/xe: Fix pinned list UAF panic FNU VISHWANATHA
  2026-10-06 17:57 ` [PATCH 2/2] RFC: drm/xe: Prevent pinned_link double add FNU VISHWANATHA
@ 2026-10-06 18:15 ` sashiko-bot
  1 sibling, 0 replies; 4+ messages in thread
From: sashiko-bot @ 2026-10-06 18:15 UTC (permalink / raw)
  To: FNU VISHWANATHA; +Cc: dri-devel, intel-xe

Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
- [High] xe: use-after-free race in xe_bo_apply_to_pinned() due to 0-refcount BOs
- [High] xe: shrinker accounting leak when pinned BO is destroyed
- [High] xe: unconditional panic due to NULL bdev in ___xe_bo_create error path
--

commit 2ba63aa4ec480f34e429a19652b5cfeb92d01204
Author: Kornel Dulęba <korneld@google.com>
RFC: drm/xe: Fix pinned list UAF panic

This commit attempts to fix a list corruption panic observed when a
buffer object is destroyed while still pinned. It modifies
xe_ttm_bo_destroy() to check and remove the buffer object from the
pinned list before it is freed.

> diff --git a/drivers/gpu/drm/xe/xe_bo.c b/drivers/gpu/drm/xe/xe_bo.c
> index 6921b6967330d..98c7c17cc46cc 100644
> --- a/drivers/gpu/drm/xe/xe_bo.c
> +++ b/drivers/gpu/drm/xe/xe_bo.c
> @@ -1885,6 +1885,11 @@ static void xe_ttm_bo_destroy(struct ttm_buffer_object *ttm_bo)
>  	if (bo->parent_obj)
>  		xe_bo_put(bo->parent_obj);
>  
> +	spin_lock(&xe->pinned.lock);
> +	if (WARN_ON_ONCE(!list_empty(&bo->pinned_link)))
> +		list_del_init(&bo->pinned_link);
> +	spin_unlock(&xe->pinned.lock);

[Severity: High]
Does this expose a use-after-free race with xe_bo_apply_to_pinned()?

When ttm_bo_release() drops the object's kref to 0, the object remains
on the pinned list until xe_ttm_bo_destroy() acquires the xe->pinned.lock
and removes it here.

Concurrently, xe_bo_apply_to_pinned() can iterate over the pinned list,
find the object, and unconditionally call xe_bo_get(bo). Since the refcount
is already 0, xe_bo_get() will call kref_get() on a 0-refcount object,
which triggers a refcount underflow warning and leads to a use-after-free
panic.

Should xe_bo_apply_to_pinned() use xe_bo_get_unless_zero() to safely
skip objects undergoing destruction?

[Severity: High]
Does removing the buffer object from the pinned list here bypass shrinker
accounting?

When a buffer object is pinned via xe_bo_pin_external(), it calls
xe_ttm_tt_account_subtract() to remove its pages from the shrinker's
shrinkable_pages pool.

By handling the cleanup directly in xe_ttm_bo_destroy() instead of the
normal xe_bo_unpin_external() path, xe_ttm_tt_account_add() is never
called to return the pages to the shrinker.

Furthermore, for native buffer objects, the teardown path in
xe_ttm_tt_unpopulate() will still execute and call
xe_ttm_tt_account_subtract() again. Because the addition was missed and a
double-subtraction occurs, this creates a permanent negative delta for
shrinkable_pages. Due to unsigned integer math, this underflow converts
into a massive positive value, potentially forcing the kernel into an
infinite memory reclaim loop.

[Severity: High]
Can this spin_lock() cause a kernel panic if reached from the
___xe_bo_create() error path?

In ___xe_bo_create(), if __xe_bo_placement_for_flags() fails, the error
path directly calls xe_ttm_bo_destroy(&bo->ttm).

At this point, ttm_bo_init_reserved() has not yet been called, meaning
bo->ttm.bdev is NULL. At the top of xe_ttm_bo_destroy(), it computes:

    struct xe_device *xe = ttm_to_xe_device(ttm_bo->bdev);

This yields an invalid pointer (NULL - offsetof(struct xe_device, ttm)).
When the new code unconditionally attempts to lock xe->pinned.lock, it
will dereference this invalid xe pointer and trigger a kernel panic.

Should the error path in ___xe_bo_create() free the buffer object directly
or pass the xe device context explicitly, rather than relying on the
uninitialized bdev?

> +
>  	mutex_lock(&xe->mem_access.vram_userfault.lock);
>  	if (!list_empty(&bo->vram_userfault_link))
>  		list_del(&bo->vram_userfault_link);

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261006175741.3582911-1-fnu.vishwanatha@intel.com?part=1

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH 2/2] RFC: drm/xe: Prevent pinned_link double add
  2026-10-06 17:57 ` [PATCH 2/2] RFC: drm/xe: Prevent pinned_link double add FNU VISHWANATHA
@ 2026-10-06 18:16   ` sashiko-bot
  0 siblings, 0 replies; 4+ messages in thread
From: sashiko-bot @ 2026-10-06 18:16 UTC (permalink / raw)
  To: FNU VISHWANATHA; +Cc: dri-devel, intel-xe

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] drm/xe: Masked list corruption causes system suspend failure
--

commit 2ec4b5f1ac44438a30593cce5cd58f19ac02f193
Author: Kornel Dulęba <korneld@google.com>

RFC: drm/xe: Prevent pinned_link double add

The commit attempts to fix a list double-add by checking if the pinned link
is empty and adding a warning. It resolves kernel crashes on memory
constrained devices running CtsCameraTestCases.

> diff --git a/drivers/gpu/drm/xe/xe_bo.c b/drivers/gpu/drm/xe/xe_bo.c
> index 98c7c17cc46cc..a5af53b36a5f2 100644
> --- a/drivers/gpu/drm/xe/xe_bo.c
> +++ b/drivers/gpu/drm/xe/xe_bo.c
> @@ -3166,7 +3166,14 @@ int xe_bo_pin_external(struct xe_bo *bo, bool in_place, struct drm_exec *exec)
>  		}
>  
>  		spin_lock(&xe->pinned.lock);
> -		list_add_tail(&bo->pinned_link, &xe->pinned.late.external);
> +		if (list_empty(&bo->pinned_link)) {
> +			list_add_tail(&bo->pinned_link,
> +				      &xe->pinned.late.external);
> +		} else {
> +			WARN_ONCE(true,
> +				  "pinned_link already added to a list, bo->ttm.pin_count: %d\n",
> +				  bo->ttm.pin_count);
> +		}

[Severity: High]
Does this list_empty() check mask the root cause of a desynchronized
lifecycle when a buffer object is pinned by multiple subsystems, such as
external and display?

If xe_bo_unpin_external() is called while the buffer object is still pinned
elsewhere, the pin_count is greater than 1, and it skips removing the
buffer object from xe->pinned.late.external:

drivers/gpu/drm/xe/xe_bo.c:xe_bo_unpin_external() {
    ...
	if (bo->ttm.pin_count == 1 && !list_empty(&bo->pinned_link))
		list_del_init(&bo->pinned_link);
    ...
}

If a subsequent display unpin happens, it calls ttm_bo_unpin() directly and
drops the pin_count to 0, completely bypassing the external pinned list
removal:

drivers/gpu/drm/xe/display/xe_fb_pin.c:__xe_unpin_fb_vma() {
    ...
	ttm_bo_unpin(&vma->bo->ttm);
    ...
}

This leaves the unpinned buffer object stranded on the pinned list. During
system suspend, power management eviction encounters this stranded unpinned
buffer object, triggering a warning and returning -EINVAL:

drivers/gpu/drm/xe/xe_bo.c:xe_bo_evict_pinned() {
    ...
		if (WARN_ON(!xe_bo_is_pinned(bo))) {
			ret = -EINVAL;
			break;
		}
    ...
}

Will this cause power management eviction to abort and break system suspend?

>  		spin_unlock(&xe->pinned.lock);
>  	}
>

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261006175741.3582911-1-fnu.vishwanatha@intel.com?part=2

^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2026-10-06 18:16 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-06 17:57 [PATCH 1/2] RFC: drm/xe: Fix pinned list UAF panic FNU VISHWANATHA
2026-10-06 17:57 ` [PATCH 2/2] RFC: drm/xe: Prevent pinned_link double add FNU VISHWANATHA
2026-10-06 18:16   ` sashiko-bot
2026-10-06 18:15 ` [PATCH 1/2] RFC: drm/xe: Fix pinned list UAF panic sashiko-bot

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox