From: Mike Lothian <mike@fireburn.co.uk>
To: amd-gfx@lists.freedesktop.org
Cc: alexander.deucher@amd.com, christian.koenig@amd.com,
kevinyang.wang@amd.com, dri-devel@lists.freedesktop.org,
stable@vger.kernel.org, Mike Lothian <mike@fireburn.co.uk>
Subject: [PATCH v2] drm/amdgpu: don't migrate a dma-buf into VRAM while runtime suspended
Date: Wed, 9 Sep 2026 10:46:09 +0100 [thread overview]
Message-ID: <20260909094609.1541-1-mike@fireburn.co.uk> (raw)
In-Reply-To: <20260909020854.58462-1-mike@fireburn.co.uk>
amdgpu_dma_buf_map() adds VRAM to the allowed domains for a peer2peer
attachment, so ttm_bo_validate() can migrate the buffer from GTT into
VRAM. While the exporting device is runtime suspended its SDMA rings
are down and the move fails:
amdgpu: Move buffer fallback to memcpy unavailable
An importer on a second GPU reaches this holding no runtime PM
reference on the exporter, e.g. a compositor on the APU submitting a
frame that references a buffer exported by an idle dGPU:
amdgpu_cs_ioctl -> amdgpu_cs_parser_bos -> amdgpu_cs_bo_validate
-> ttm_bo_validate -> amdgpu_bo_move -> dma_buf_map_attachment
-> amdgpu_dma_buf_map -> ttm_bo_validate -> amdgpu_bo_move
Only migrate into VRAM while holding the exporter awake.
pm_runtime_get_if_active() takes a reference only when the device is
already active and never resumes it, so it cannot deadlock against the
reservation held across these callbacks. That deadlock is why
commit 030631e97b20 ("drm/amdgpu: revert "take runtime pm reference
when we attach a buffer" v2") removed the pm_runtime_get_sync() from
the attach callback.
If the device is suspended or suspending the buffer stays in GTT, which
remains accessible while the GPU is powered down. A negative return
means runtime PM is disabled, so the device cannot suspend and VRAM
stays usable.
Fixes: 030631e97b20 ("drm/amdgpu: revert "take runtime pm reference when we attach a buffer" v2")
Cc: stable@vger.kernel.org
Signed-off-by: Mike Lothian <mike@fireburn.co.uk>
Assisted-by: Claude:Opus-5 [Claude Code]
---
v2: hold the exporter with pm_runtime_get_if_active() across the
validate instead of testing adev->mman.buffer_funcs_enabled. The
v1 check was racy - the device could suspend between the test and
ttm_bo_validate(), so amdgpu_bo_move() could still see the rings
torn down. Reported by Sashiko AI review.
Is this the failure that commit c52feb436539 ("drm/amdgpu: Disable
runtime PM for externally attached dGPUs") was working around? That
commit explains how to detect external attachment but not what breaks,
so I can't tell which.
If it is the same thing, could the pci_is_thunderbolt_attached() ||
dev_is_removable() check there be narrowed or dropped on top of this,
so eGPU users keep runtime PM?
I can't test that here: this box hits the bug by missing that check.
The dGPU is on an oculink port off a native AMD root port, so
pci_is_thunderbolt_attached() is false, dev_is_removable() is empty,
and runtime PM stays enabled.
Reproduced on a HawkPoint APU [1002:1900] driving the display with a
Navi 48 [Radeon AI PRO R9700] [1002:7551] on oculink for render
offload. Without the patch kwin_wayland hits the call chain above
within a minute of the dGPU autosuspending and the desktop stops
repainting until it resumes.
drivers/gpu/drm/amd/amdgpu/amdgpu_dma_buf.c | 16 ++++++++++++++--
1 file changed, 14 insertions(+), 2 deletions(-)
diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_dma_buf.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_dma_buf.c
index b33c300e26e2..c89846f266d3 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_dma_buf.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_dma_buf.c
@@ -43,6 +43,7 @@
#include <linux/dma-buf.h>
#include <linux/dma-fence-array.h>
#include <linux/pci-p2pdma.h>
+#include <linux/pm_runtime.h>
static const struct dma_buf_attach_ops amdgpu_dma_buf_attach_ops;
@@ -189,14 +190,25 @@ static struct sg_table *amdgpu_dma_buf_map(struct dma_buf_attachment *attach,
/* move buffer into GTT or VRAM */
struct ttm_operation_ctx ctx = { false, false };
unsigned int domains = AMDGPU_GEM_DOMAIN_GTT;
+ int pm_ref = 0;
if (bo->preferred_domains & AMDGPU_GEM_DOMAIN_VRAM &&
attach->peer2peer) {
- bo->flags |= AMDGPU_GEM_CREATE_CPU_ACCESS_REQUIRED;
- domains |= AMDGPU_GEM_DOMAIN_VRAM;
+ /*
+ * Only migrate into VRAM while the exporter is held
+ * awake. A negative return means runtime PM is
+ * disabled, so it cannot suspend either.
+ */
+ pm_ref = pm_runtime_get_if_active(adev_to_drm(adev)->dev);
+ if (pm_ref) {
+ bo->flags |= AMDGPU_GEM_CREATE_CPU_ACCESS_REQUIRED;
+ domains |= AMDGPU_GEM_DOMAIN_VRAM;
+ }
}
amdgpu_bo_placement_from_domain(bo, domains);
r = ttm_bo_validate(&bo->tbo, &bo->placement, &ctx);
+ if (pm_ref > 0)
+ pm_runtime_put_autosuspend(adev_to_drm(adev)->dev);
if (r)
return ERR_PTR(r);
}
--
2.55.0
next prev parent reply other threads:[~2026-09-09 9:46 UTC|newest]
Thread overview: 11+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-09 2:08 [PATCH] drm/amdgpu: don't migrate a dma-buf into VRAM while runtime suspended Mike Lothian
2026-09-09 2:23 ` sashiko-bot
2026-09-09 9:46 ` Mike Lothian [this message]
2026-09-09 12:46 ` [PATCH v2] " Christian König
2026-09-10 0:15 ` Mike Lothian
2026-09-10 9:27 ` Christian König
2026-09-11 18:38 ` [PATCH v3] drm/amdgpu: hold a runtime PM reference for P2P dma-buf attachments Mike Lothian
2026-09-11 18:46 ` sashiko-bot
2026-09-11 23:29 ` [PATCH v4] " Mike Lothian
2026-09-14 10:51 ` [PATCH v3] " Christian König
2026-09-14 14:11 ` Alex Deucher
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260909094609.1541-1-mike@fireburn.co.uk \
--to=mike@fireburn.co.uk \
--cc=alexander.deucher@amd.com \
--cc=amd-gfx@lists.freedesktop.org \
--cc=christian.koenig@amd.com \
--cc=dri-devel@lists.freedesktop.org \
--cc=kevinyang.wang@amd.com \
--cc=stable@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.