From: Matthew Brost <matthew.brost@intel.com>
To: "Thomas Hellström" <thomas.hellstrom@linux.intel.com>
Cc: Matthew Auld <matthew.auld@intel.com>,
<intel-xe@lists.freedesktop.org>, <stable@vger.kernel.org>
Subject: Re: [PATCH 1/2] drm/xe/dma-buf: keep non-p2p imported buffers in system memory
Date: Tue, 29 Sep 2026 09:11:08 -0700 [thread overview]
Message-ID: <arvjHLdLi3Vk5lDe@gsse-cloud1.jf.intel.com> (raw)
In-Reply-To: <70ae41bf314d518215dda1f062b609f3dfac73e4.camel@linux.intel.com>
On Tue, Sep 29, 2026 at 09:56:10AM +0200, Thomas Hellström wrote:
> On Mon, 2026-09-28 at 15:01 -0700, Matthew Brost wrote:
> > On Mon, Sep 28, 2026 at 05:48:22PM +0100, Matthew Auld wrote:
> > > When an exported buffer is mapped by a foreign device lacking peer-
> > > to-peer
> > > DMA support (such as an integrated GPU for display offload),
> > > xe_dma_buf_map() migrates the buffer to XE_PL_TT so that system
> > > memory
> > > pages can be accessed by the importer.
> > >
> > > However, since commit 5c87fee3c96c ("drm/xe: Attempt to bring bos
> > > back to
> > > VRAM after eviction"), XE_PL_TT is marked with TTM_PL_FLAG_FALLBACK
> > > in
> > > the buffer's placement. When DMABUF_MOVE_NOTIFY was enabled by
> > > default
> > > (meaning dynamic attachments are no longer pinned on map), the next
> > > xe_bo_validate() during render submission sees TT as a fallback
> > > placement
> > > and attempts to migrate the buffer back to VRAM.
> > >
> > > On the next frame, the foreign importer accesses the buffer,
> > > requiring
> > > another migration to TT, resulting in a continuous ping-pong
> > > between
> > > VRAM and system memory every frame and causing severe rendering
> > > performance
> > > degradations.
> > >
> > > To fix this, introduce xe_bo_migrate_tt_sticky() which records a
> > > temporary sticky placement in XE_PL_TT on successful migration.
> > > Subsequent
> > > validations use this sticky placement to keep the buffer in system
> > > memory
> > > until the non-p2p attachment is detached or memory pressure forces
> > > a
> > > fallback to the default placement.
> > >
> > > User is reporting what looks to be exactly this, with horrible
> > > performance when DMABUF_MOVE_NOTIFY was enabled by default.
> > >
> >
> > The patch looks functionally correct but couldn't we just clear
> > TTM_PL_FLAG_FALLBACK in xe_bo_migrate for certain callers? Then
> > restore
> > it in other places?
> >
> > Matt
>
> Hi,
>
> I recommended Matt to use a separate placement because the TTM flags
> are hints anyway, I think, and code becomes easier to understand if we
> make this more explicit. But that was just my opinion.
>
I don't have a strong opinion here, so separate placement works for me.
> But anyway I see there are a couple of places we migrate where we might
> need to adjust the current behaviour accordingly:
>
> * We have migration to TT for dma-buf CPU access.
> * We have migration to TT for atomic access, and then implicit
> assumptions that it will be migrated to VRAM again for GPU access, I
> suppose.
> * We have prefetch back to VRAM.
> * In what situations do we want to make eviction to TT sticky?
> * If we remove stickyness, should we trigger a rebind?
>
> This is some serious technical debt we have. We probably want to put
> together a design document on this, but for the time being we need to
> make sure we fix that imminent dma-buf bouncing problem.
>
+1, as it's good to audit everything, document the design, and clean up
any technical debt.
I'm fine with merging the series as-is. Functionally, I believe
everything is correct and should address the problem at hand. Any
remaining cleanup can be handled in a follow-up series, unless there is
something on your list that is a hard blocker. I haven't spent much time
thinking about the technical debt items yet.
Matt
> /Thomas
>
>
> >
> > > Assisted-by: LLM
> > > Link:
> > > https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/9415
> > > Fixes: 5c87fee3c96c ("drm/xe: Attempt to bring bos back to VRAM
> > > after eviction")
> > > Signed-off-by: Matthew Auld <matthew.auld@intel.com>
> > > Cc: <stable@vger.kernel.org> # v6.12+
> > > Cc: Thomas Hellström <thomas.hellstrom@linux.intel.com>
> > > Cc: Matthew Brost <matthew.brost@intel.com>
> > > ---
> > > drivers/gpu/drm/xe/xe_bo.c | 59
> > > ++++++++++++++++++++++++++++++--
> > > drivers/gpu/drm/xe/xe_bo.h | 3 ++
> > > drivers/gpu/drm/xe/xe_bo_types.h | 4 +++
> > > drivers/gpu/drm/xe/xe_dma_buf.c | 16 ++++++++-
> > > 4 files changed, 79 insertions(+), 3 deletions(-)
> > >
> > > diff --git a/drivers/gpu/drm/xe/xe_bo.c
> > > b/drivers/gpu/drm/xe/xe_bo.c
> > > index 6921b6967330..1c1eab0a7434 100644
> > > --- a/drivers/gpu/drm/xe/xe_bo.c
> > > +++ b/drivers/gpu/drm/xe/xe_bo.c
> > > @@ -3324,6 +3324,9 @@ int xe_bo_validate(struct xe_bo *bo, struct
> > > xe_vm *vm, bool allow_res_evict,
> > > .no_wait_gpu = false,
> > > .gfp_retry_mayfail = true,
> > > };
> > > + struct ttm_placement *placement = bo-
> > > >sticky_placement.num_placement ?
> > > + &bo->sticky_placement :
> > > + &bo->placement;
> > > int ret;
> > >
> > > if (xe_bo_is_pinned(bo))
> > > @@ -3340,7 +3343,12 @@ int xe_bo_validate(struct xe_bo *bo, struct
> > > xe_vm *vm, bool allow_res_evict,
> > > xe_vm_set_validating(vm, allow_res_evict);
> > > trace_xe_bo_validate(bo);
> > > xe_validation_assert_exec(xe_bo_device(bo), exec, &bo-
> > > >ttm.base);
> > > - ret = ttm_bo_validate(&bo->ttm, &bo->placement, &ctx);
> > > + ret = ttm_bo_validate(&bo->ttm, placement, &ctx);
> > > + if (ret && ret != -EINTR && ret != -ERESTARTSYS &&
> > > + bo->sticky_placement.num_placement) {
> > > + xe_bo_reset_sticky_placement(bo);
> > > + ret = ttm_bo_validate(&bo->ttm, &bo->placement,
> > > &ctx);
> > > + }
> > > xe_vm_clear_validating(vm, allow_res_evict);
> > >
> > > return ret;
> > > @@ -3891,6 +3899,7 @@ int xe_bo_migrate(struct xe_bo *bo, u32
> > > mem_type, struct ttm_operation_ctx *tctx
> > > };
> > > struct ttm_placement placement;
> > > struct ttm_place requested;
> > > + int ret;
> > >
> > > xe_bo_assert_held(bo);
> > > tctx = tctx ? tctx : &ctx;
> > > @@ -3922,7 +3931,53 @@ int xe_bo_migrate(struct xe_bo *bo, u32
> > > mem_type, struct ttm_operation_ctx *tctx
> > >
> > > if (!tctx->no_wait_gpu)
> > > xe_validation_assert_exec(xe_bo_device(bo), exec,
> > > &bo->ttm.base);
> > > - return ttm_bo_validate(&bo->ttm, &placement, tctx);
> > > + ret = ttm_bo_validate(&bo->ttm, &placement, tctx);
> > > + if (!ret)
> > > + xe_bo_reset_sticky_placement(bo);
> > > + return ret;
> > > +}
> > > +
> > > +/**
> > > + * xe_bo_migrate_tt_sticky - Migrate an object to TT and record
> > > its placement as sticky
> > > + * @bo: The buffer object to migrate.
> > > + * @tctx: The ttm_operation_ctx to use for migration, or NULL for
> > > default.
> > > + * @exec: The drm_exec transaction to use for exhaustive eviction.
> > > + *
> > > + * Like xe_bo_migrate() to XE_PL_TT, but on success records the
> > > resulting placement
> > > + * so that subsequent validations try to keep the object in TT
> > > instead of falling
> > > + * back to the default placement. The stickiness is removed if
> > > validation falls
> > > + * back to the default placement (e.g. under memory pressure), on
> > > any non-sticky
> > > + * migration, or via an explicit call to
> > > xe_bo_reset_sticky_placement().
> > > + *
> > > + * Return: 0 on success. Negative error code on failure.
> > > + */
> > > +int xe_bo_migrate_tt_sticky(struct xe_bo *bo,
> > > + struct ttm_operation_ctx *tctx,
> > > + struct drm_exec *exec)
> > > +{
> > > + int ret;
> > > +
> > > + ret = xe_bo_migrate(bo, XE_PL_TT, tctx, exec);
> > > + if (!ret) {
> > > + xe_place_from_ttm_type(XE_PL_TT, &bo-
> > > >sticky_place);
> > > + bo->sticky_placement = (struct ttm_placement){
> > > + .num_placement = 1,
> > > + .placement = &bo->sticky_place,
> > > + };
> > > + }
> > > + return ret;
> > > +}
> > > +
> > > +/**
> > > + * xe_bo_reset_sticky_placement - Reset sticky placement for an
> > > object
> > > + * @bo: The buffer object whose sticky placement should be
> > > cleared.
> > > + *
> > > + * Clear any sticky placement recorded by
> > > xe_bo_migrate_tt_sticky(),
> > > + * returning subsequent validations to the default placement.
> > > + */
> > > +void xe_bo_reset_sticky_placement(struct xe_bo *bo)
> > > +{
> > > + bo->sticky_placement.num_placement = 0;
> > > }
> > >
> > > /**
> > > diff --git a/drivers/gpu/drm/xe/xe_bo.h
> > > b/drivers/gpu/drm/xe/xe_bo.h
> > > index 861b1be231de..7327628070f2 100644
> > > --- a/drivers/gpu/drm/xe/xe_bo.h
> > > +++ b/drivers/gpu/drm/xe/xe_bo.h
> > > @@ -434,6 +434,9 @@ bool xe_bo_can_migrate(struct xe_bo *bo, u32
> > > mem_type);
> > >
> > > int xe_bo_migrate(struct xe_bo *bo, u32 mem_type, struct
> > > ttm_operation_ctx *ctc,
> > > struct drm_exec *exec);
> > > +int xe_bo_migrate_tt_sticky(struct xe_bo *bo, struct
> > > ttm_operation_ctx *ctc,
> > > + struct drm_exec *exec);
> > > +void xe_bo_reset_sticky_placement(struct xe_bo *bo);
> > > int xe_bo_evict(struct xe_bo *bo, struct drm_exec *exec);
> > >
> > > int xe_bo_evict_pinned(struct xe_bo *bo);
> > > diff --git a/drivers/gpu/drm/xe/xe_bo_types.h
> > > b/drivers/gpu/drm/xe/xe_bo_types.h
> > > index 8ec4a01a0092..253a1dba55a7 100644
> > > --- a/drivers/gpu/drm/xe/xe_bo_types.h
> > > +++ b/drivers/gpu/drm/xe/xe_bo_types.h
> > > @@ -56,6 +56,10 @@ struct xe_bo {
> > > struct ttm_place placements[XE_BO_MAX_PLACEMENTS];
> > > /** @placement: current placement for this BO */
> > > struct ttm_placement placement;
> > > + /** @sticky_placement: target placement from forced
> > > migration */
> > > + struct ttm_placement sticky_placement;
> > > + /** @sticky_place: place for sticky_placement */
> > > + struct ttm_place sticky_place;
> > > /** @ggtt_node: Array of GGTT nodes if this BO is mapped
> > > in the GGTTs */
> > > struct xe_ggtt_node *ggtt_node[XE_MAX_TILES_PER_DEVICE];
> > > /** @vmap: iosys map of this buffer */
> > > diff --git a/drivers/gpu/drm/xe/xe_dma_buf.c
> > > b/drivers/gpu/drm/xe/xe_dma_buf.c
> > > index 6a85b292dee7..b973579a69b8 100644
> > > --- a/drivers/gpu/drm/xe/xe_dma_buf.c
> > > +++ b/drivers/gpu/drm/xe/xe_dma_buf.c
> > > @@ -59,6 +59,20 @@ static void xe_dma_buf_detach(struct dma_buf
> > > *dmabuf,
> > > struct dma_buf_attachment *attach)
> > > {
> > > struct drm_gem_object *obj = attach->dmabuf->priv;
> > > + struct xe_bo *bo = gem_to_xe_bo(obj);
> > > + bool has_non_p2p = false;
> > > + struct dma_buf_attachment *a;
> > > +
> > > + dma_resv_lock(dmabuf->resv, NULL);
> > > + list_for_each_entry(a, &dmabuf->attachments, node) {
> > > + if (!a->peer2peer) {
> > > + has_non_p2p = true;
> > > + break;
> > > + }
> > > + }
> > > + if (!has_non_p2p)
> > > + xe_bo_reset_sticky_placement(bo);
> > > + dma_resv_unlock(dmabuf->resv);
> > >
> > > xe_pm_runtime_put(to_xe_device(obj->dev));
> > > }
> > > @@ -129,7 +143,7 @@ static struct sg_table *xe_dma_buf_map(struct
> > > dma_buf_attachment *attach,
> > >
> > > if (!xe_bo_is_pinned(bo)) {
> > > if (!attach->peer2peer)
> > > - r = xe_bo_migrate(bo, XE_PL_TT, NULL,
> > > exec);
> > > + r = xe_bo_migrate_tt_sticky(bo, NULL,
> > > exec);
> > > else
> > > r = xe_bo_validate(bo, NULL, false, exec);
> > > if (r)
> > > --
> > > 2.55.0
> > >
next prev parent reply other threads:[~2026-09-29 16:11 UTC|newest]
Thread overview: 13+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-28 16:48 [PATCH 0/2] sticky TT Matthew Auld
2026-09-28 16:48 ` [PATCH 1/2] drm/xe/dma-buf: keep non-p2p imported buffers in system memory Matthew Auld
2026-09-28 22:01 ` Matthew Brost
2026-09-29 7:56 ` Thomas Hellström
2026-09-29 16:11 ` Matthew Brost [this message]
2026-10-02 12:40 ` Thomas Hellström
2026-09-28 16:48 ` [PATCH 2/2] drm/xe/vm: make PREFETCH to TT sticky Matthew Auld
2026-10-02 12:43 ` Thomas Hellström
2026-10-02 13:32 ` Matthew Auld
2026-09-28 17:30 ` ✗ CI.checkpatch: warning for sticky TT Patchwork
2026-09-28 17:32 ` ✓ CI.KUnit: success " Patchwork
2026-09-28 18:54 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-28 23:28 ` ✗ Xe.CI.FULL: failure " Patchwork
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=arvjHLdLi3Vk5lDe@gsse-cloud1.jf.intel.com \
--to=matthew.brost@intel.com \
--cc=intel-xe@lists.freedesktop.org \
--cc=matthew.auld@intel.com \
--cc=stable@vger.kernel.org \
--cc=thomas.hellstrom@linux.intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox