Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Matthew Auld <matthew.auld@intel.com>
To: intel-xe@lists.freedesktop.org
Cc: stable@vger.kernel.org,
	"Thomas Hellström" <thomas.hellstrom@linux.intel.com>,
	"Matthew Brost" <matthew.brost@intel.com>
Subject: [PATCH 1/2] drm/xe/dma-buf: keep non-p2p imported buffers in system memory
Date: Mon, 28 Sep 2026 17:48:22 +0100	[thread overview]
Message-ID: <20260928164820.1237049-5-matthew.auld@intel.com> (raw)
In-Reply-To: <20260928164820.1237049-4-matthew.auld@intel.com>

When an exported buffer is mapped by a foreign device lacking peer-to-peer
DMA support (such as an integrated GPU for display offload),
xe_dma_buf_map() migrates the buffer to XE_PL_TT so that system memory
pages can be accessed by the importer.

However, since commit 5c87fee3c96c ("drm/xe: Attempt to bring bos back to
VRAM after eviction"), XE_PL_TT is marked with TTM_PL_FLAG_FALLBACK in
the buffer's placement. When DMABUF_MOVE_NOTIFY was enabled by default
(meaning dynamic attachments are no longer pinned on map), the next
xe_bo_validate() during render submission sees TT as a fallback placement
and attempts to migrate the buffer back to VRAM.

On the next frame, the foreign importer accesses the buffer, requiring
another migration to TT, resulting in a continuous ping-pong between
VRAM and system memory every frame and causing severe rendering performance
degradations.

To fix this, introduce xe_bo_migrate_tt_sticky() which records a
temporary sticky placement in XE_PL_TT on successful migration. Subsequent
validations use this sticky placement to keep the buffer in system memory
until the non-p2p attachment is detached or memory pressure forces a
fallback to the default placement.

User is reporting what looks to be exactly this, with horrible
performance when DMABUF_MOVE_NOTIFY was enabled by default.

Assisted-by: LLM
Link: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/9415
Fixes: 5c87fee3c96c ("drm/xe: Attempt to bring bos back to VRAM after eviction")
Signed-off-by: Matthew Auld <matthew.auld@intel.com>
Cc: <stable@vger.kernel.org> # v6.12+
Cc: Thomas Hellström <thomas.hellstrom@linux.intel.com>
Cc: Matthew Brost <matthew.brost@intel.com>
---
 drivers/gpu/drm/xe/xe_bo.c       | 59 ++++++++++++++++++++++++++++++--
 drivers/gpu/drm/xe/xe_bo.h       |  3 ++
 drivers/gpu/drm/xe/xe_bo_types.h |  4 +++
 drivers/gpu/drm/xe/xe_dma_buf.c  | 16 ++++++++-
 4 files changed, 79 insertions(+), 3 deletions(-)

diff --git a/drivers/gpu/drm/xe/xe_bo.c b/drivers/gpu/drm/xe/xe_bo.c
index 6921b6967330..1c1eab0a7434 100644
--- a/drivers/gpu/drm/xe/xe_bo.c
+++ b/drivers/gpu/drm/xe/xe_bo.c
@@ -3324,6 +3324,9 @@ int xe_bo_validate(struct xe_bo *bo, struct xe_vm *vm, bool allow_res_evict,
 		.no_wait_gpu = false,
 		.gfp_retry_mayfail = true,
 	};
+	struct ttm_placement *placement = bo->sticky_placement.num_placement ?
+					  &bo->sticky_placement :
+					  &bo->placement;
 	int ret;
 
 	if (xe_bo_is_pinned(bo))
@@ -3340,7 +3343,12 @@ int xe_bo_validate(struct xe_bo *bo, struct xe_vm *vm, bool allow_res_evict,
 	xe_vm_set_validating(vm, allow_res_evict);
 	trace_xe_bo_validate(bo);
 	xe_validation_assert_exec(xe_bo_device(bo), exec, &bo->ttm.base);
-	ret = ttm_bo_validate(&bo->ttm, &bo->placement, &ctx);
+	ret = ttm_bo_validate(&bo->ttm, placement, &ctx);
+	if (ret && ret != -EINTR && ret != -ERESTARTSYS &&
+	    bo->sticky_placement.num_placement) {
+		xe_bo_reset_sticky_placement(bo);
+		ret = ttm_bo_validate(&bo->ttm, &bo->placement, &ctx);
+	}
 	xe_vm_clear_validating(vm, allow_res_evict);
 
 	return ret;
@@ -3891,6 +3899,7 @@ int xe_bo_migrate(struct xe_bo *bo, u32 mem_type, struct ttm_operation_ctx *tctx
 	};
 	struct ttm_placement placement;
 	struct ttm_place requested;
+	int ret;
 
 	xe_bo_assert_held(bo);
 	tctx = tctx ? tctx : &ctx;
@@ -3922,7 +3931,53 @@ int xe_bo_migrate(struct xe_bo *bo, u32 mem_type, struct ttm_operation_ctx *tctx
 
 	if (!tctx->no_wait_gpu)
 		xe_validation_assert_exec(xe_bo_device(bo), exec, &bo->ttm.base);
-	return ttm_bo_validate(&bo->ttm, &placement, tctx);
+	ret = ttm_bo_validate(&bo->ttm, &placement, tctx);
+	if (!ret)
+		xe_bo_reset_sticky_placement(bo);
+	return ret;
+}
+
+/**
+ * xe_bo_migrate_tt_sticky - Migrate an object to TT and record its placement as sticky
+ * @bo: The buffer object to migrate.
+ * @tctx: The ttm_operation_ctx to use for migration, or NULL for default.
+ * @exec: The drm_exec transaction to use for exhaustive eviction.
+ *
+ * Like xe_bo_migrate() to XE_PL_TT, but on success records the resulting placement
+ * so that subsequent validations try to keep the object in TT instead of falling
+ * back to the default placement. The stickiness is removed if validation falls
+ * back to the default placement (e.g. under memory pressure), on any non-sticky
+ * migration, or via an explicit call to xe_bo_reset_sticky_placement().
+ *
+ * Return: 0 on success. Negative error code on failure.
+ */
+int xe_bo_migrate_tt_sticky(struct xe_bo *bo,
+			    struct ttm_operation_ctx *tctx,
+			    struct drm_exec *exec)
+{
+	int ret;
+
+	ret = xe_bo_migrate(bo, XE_PL_TT, tctx, exec);
+	if (!ret) {
+		xe_place_from_ttm_type(XE_PL_TT, &bo->sticky_place);
+		bo->sticky_placement = (struct ttm_placement){
+			.num_placement = 1,
+			.placement = &bo->sticky_place,
+		};
+	}
+	return ret;
+}
+
+/**
+ * xe_bo_reset_sticky_placement - Reset sticky placement for an object
+ * @bo: The buffer object whose sticky placement should be cleared.
+ *
+ * Clear any sticky placement recorded by xe_bo_migrate_tt_sticky(),
+ * returning subsequent validations to the default placement.
+ */
+void xe_bo_reset_sticky_placement(struct xe_bo *bo)
+{
+	bo->sticky_placement.num_placement = 0;
 }
 
 /**
diff --git a/drivers/gpu/drm/xe/xe_bo.h b/drivers/gpu/drm/xe/xe_bo.h
index 861b1be231de..7327628070f2 100644
--- a/drivers/gpu/drm/xe/xe_bo.h
+++ b/drivers/gpu/drm/xe/xe_bo.h
@@ -434,6 +434,9 @@ bool xe_bo_can_migrate(struct xe_bo *bo, u32 mem_type);
 
 int xe_bo_migrate(struct xe_bo *bo, u32 mem_type, struct ttm_operation_ctx *ctc,
 		  struct drm_exec *exec);
+int xe_bo_migrate_tt_sticky(struct xe_bo *bo, struct ttm_operation_ctx *ctc,
+			    struct drm_exec *exec);
+void xe_bo_reset_sticky_placement(struct xe_bo *bo);
 int xe_bo_evict(struct xe_bo *bo, struct drm_exec *exec);
 
 int xe_bo_evict_pinned(struct xe_bo *bo);
diff --git a/drivers/gpu/drm/xe/xe_bo_types.h b/drivers/gpu/drm/xe/xe_bo_types.h
index 8ec4a01a0092..253a1dba55a7 100644
--- a/drivers/gpu/drm/xe/xe_bo_types.h
+++ b/drivers/gpu/drm/xe/xe_bo_types.h
@@ -56,6 +56,10 @@ struct xe_bo {
 	struct ttm_place placements[XE_BO_MAX_PLACEMENTS];
 	/** @placement: current placement for this BO */
 	struct ttm_placement placement;
+	/** @sticky_placement: target placement from forced migration */
+	struct ttm_placement sticky_placement;
+	/** @sticky_place: place for sticky_placement */
+	struct ttm_place sticky_place;
 	/** @ggtt_node: Array of GGTT nodes if this BO is mapped in the GGTTs */
 	struct xe_ggtt_node *ggtt_node[XE_MAX_TILES_PER_DEVICE];
 	/** @vmap: iosys map of this buffer */
diff --git a/drivers/gpu/drm/xe/xe_dma_buf.c b/drivers/gpu/drm/xe/xe_dma_buf.c
index 6a85b292dee7..b973579a69b8 100644
--- a/drivers/gpu/drm/xe/xe_dma_buf.c
+++ b/drivers/gpu/drm/xe/xe_dma_buf.c
@@ -59,6 +59,20 @@ static void xe_dma_buf_detach(struct dma_buf *dmabuf,
 			      struct dma_buf_attachment *attach)
 {
 	struct drm_gem_object *obj = attach->dmabuf->priv;
+	struct xe_bo *bo = gem_to_xe_bo(obj);
+	bool has_non_p2p = false;
+	struct dma_buf_attachment *a;
+
+	dma_resv_lock(dmabuf->resv, NULL);
+	list_for_each_entry(a, &dmabuf->attachments, node) {
+		if (!a->peer2peer) {
+			has_non_p2p = true;
+			break;
+		}
+	}
+	if (!has_non_p2p)
+		xe_bo_reset_sticky_placement(bo);
+	dma_resv_unlock(dmabuf->resv);
 
 	xe_pm_runtime_put(to_xe_device(obj->dev));
 }
@@ -129,7 +143,7 @@ static struct sg_table *xe_dma_buf_map(struct dma_buf_attachment *attach,
 
 	if (!xe_bo_is_pinned(bo)) {
 		if (!attach->peer2peer)
-			r = xe_bo_migrate(bo, XE_PL_TT, NULL, exec);
+			r = xe_bo_migrate_tt_sticky(bo, NULL, exec);
 		else
 			r = xe_bo_validate(bo, NULL, false, exec);
 		if (r)
-- 
2.55.0


  reply	other threads:[~2026-09-28 16:49 UTC|newest]

Thread overview: 13+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-28 16:48 [PATCH 0/2] sticky TT Matthew Auld
2026-09-28 16:48 ` Matthew Auld [this message]
2026-09-28 22:01   ` [PATCH 1/2] drm/xe/dma-buf: keep non-p2p imported buffers in system memory Matthew Brost
2026-09-29  7:56     ` Thomas Hellström
2026-09-29 16:11       ` Matthew Brost
2026-10-02 12:40   ` Thomas Hellström
2026-09-28 16:48 ` [PATCH 2/2] drm/xe/vm: make PREFETCH to TT sticky Matthew Auld
2026-10-02 12:43   ` Thomas Hellström
2026-10-02 13:32     ` Matthew Auld
2026-09-28 17:30 ` ✗ CI.checkpatch: warning for sticky TT Patchwork
2026-09-28 17:32 ` ✓ CI.KUnit: success " Patchwork
2026-09-28 18:54 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-28 23:28 ` ✗ Xe.CI.FULL: failure " Patchwork

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260928164820.1237049-5-matthew.auld@intel.com \
    --to=matthew.auld@intel.com \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=matthew.brost@intel.com \
    --cc=stable@vger.kernel.org \
    --cc=thomas.hellstrom@linux.intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox