* [PATCH V2 00/31] Rework GPU TLB invalidation
@ 2026-09-01 20:10 Alex Deucher
2026-09-01 20:10 ` [PATCH 01/31] drm/amdgpu/gmc9: disallow gfxoff around TLB flushes Alex Deucher
` (31 more replies)
0 siblings, 32 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
GMC 9-12 use KIQ or MES for TLB invalidations to avoid using MMIO which would
require disallowing GFXOFF. KIQ and MES are management queues however and if
they hang, they cannot be recovered by a queue reset since they are the
mechanisms which handle queue resets. Since using KIQ or MES will exit GFXOFF
anyway, explicitly disallow it on the MMIO path and use that. Next, switch to
using SDMA for TLB invalidations. SDMA 4.4.x and newer have special packets
specifically for this purpose. If SDMA hangs while doing the invalidation for
some reason, it's easier to reset the SDMA queue than KIQ or MES. Finally,
most of the TLB invalidation code between GMC 9 through 12 was identical, so
move it to common GMC helpers and remove the IP specific code. If the SMDA
and MMIO pathes prove to be stable, the KIQ pathes can be removed in the future
to further simplify things. SDMA 4.x could also be updated to support PASID
invalidation via SDMA using either the new packet (SDMA 4.4.x) or via
REG_WRITE/REG_WAIT packets (SDMA 4.0.x).
Code is available on this branch as well:
https://gitlab.freedesktop.org/agd5f/linux/-/commits/tlb_inv_rework?ref_type=heads
V2:
- Add missing hub callbacks in gmc9 hubs
Alex Deucher (31):
drm/amdgpu/gmc9: disallow gfxoff around TLB flushes
drm/amdgpu/gmc10: disallow gfxoff around TLB flushes
drm/amdgpu/gmc11: disallow gfxoff around TLB flushes
drm/amdgpu/gmc12: disallow gfxoff around TLB flushes
drm/amdgpu/gmc9: set vmhub funcs for gfxhub and mmhub
drm/amdgpu: add a gmc flag for using MMIO for TLB flush
drm/amdgpu/gmc9: use MMIO for TLB flushes
drm/amdgpu/gmc10: use MMIO for TLB flushes
drm/amdgpu/gmc11: use MMIO for TLB flushes
drm/amdgpu/gmc12: use MMIO for TLB flushes
drm/amdgpu: add a buffer funcs callback for TLB invalidation
drm/amdgpu/sdma5.0: add tlb invalidation buffer func callback
drm/amdgpu/sdma5.2: add tlb invalidation buffer func callback
drm/amdgpu/sdma6: add tlb invalidation buffer func callback
drm/amdgpu/sdma7: add tlb invalidation buffer func callback
drm/amdgpu: simplify amdgpu_gmc_flush_gpu_tlb()
drm/amdgpu: add tlb invalidation method enum
drm/amdgpu: plumb tlb inv method in amdgpu_gmc_flush_gpu_tlb_gart()
drm/amdgpu: uplevel reset check in amdgpu_gmc_flush_gpu_tlb_gart()
drm/amdgpu/gmc: add new callback to lookup vmid to pasid mapping
drm/amdgpu: add a gmc callback for the inv semaphore
drm/amdgpu/gmc: rework pasid flushing
drm/amdgpu/gmc9: use SDMA for gart TLB invalidation
drm/amdgpu/gmc10: use SDMA for gart TLB invalidation
drm/amdgpu/gmc11: use SDMA for gart TLB invalidation
drm/amdgpu/gmc12: use SDMA for gart TLB invalidation
drm/amdgpu/gmc10: use SDMA for pasid TLB invalidation
drm/amdgpu/gmc11: use SDMA for pasid TLB invalidation
drm/amdgpu/gmc12: use MES or SDMA for pasid TLB invalidation
drm/amdgpu/gmc12: drop MES tlb inv in gmc callbacks
drm/amdgpu/gmc: add helpers for various tlb inv functions
drivers/gpu/drm/amd/amdgpu/amdgpu_gart.c | 2 +-
drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c | 486 +++++++++++++++++++----
drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 33 +-
drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h | 18 +
drivers/gpu/drm/amd/amdgpu/gfx_v10_0.c | 2 +
drivers/gpu/drm/amd/amdgpu/gfxhub_v1_0.c | 33 ++
drivers/gpu/drm/amd/amdgpu/gfxhub_v1_2.c | 33 ++
drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c | 202 +---------
drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c | 207 +---------
drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c | 250 ++----------
drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c | 225 +----------
drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c | 135 ++-----
drivers/gpu/drm/amd/amdgpu/mes_v12_0.c | 4 +
drivers/gpu/drm/amd/amdgpu/mes_v12_1.c | 4 +
drivers/gpu/drm/amd/amdgpu/mmhub_v1_0.c | 33 ++
drivers/gpu/drm/amd/amdgpu/mmhub_v1_7.c | 32 ++
drivers/gpu/drm/amd/amdgpu/mmhub_v1_8.c | 32 ++
drivers/gpu/drm/amd/amdgpu/mmhub_v9_4.c | 32 ++
drivers/gpu/drm/amd/amdgpu/sdma_v5_0.c | 49 +++
drivers/gpu/drm/amd/amdgpu/sdma_v5_2.c | 49 +++
drivers/gpu/drm/amd/amdgpu/sdma_v6_0.c | 49 +++
drivers/gpu/drm/amd/amdgpu/sdma_v7_0.c | 48 +++
22 files changed, 943 insertions(+), 1015 deletions(-)
--
2.55.0
^ permalink raw reply [flat|nested] 43+ messages in thread
* [PATCH 01/31] drm/amdgpu/gmc9: disallow gfxoff around TLB flushes
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-02 7:10 ` Christian König
2026-09-01 20:10 ` [PATCH 02/31] drm/amdgpu/gmc10: " Alex Deucher
` (30 subsequent siblings)
31 siblings, 1 reply; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
We need to disallow gfxoff if we touch GC MMIO registers.
At the moment we use KIQ or MES for TLB flushes so
no intended functional change.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c | 7 +++++++
1 file changed, 7 insertions(+)
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
index b46b87291c512..80f1cf1f21736 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
@@ -808,6 +808,10 @@ static void gmc_v9_0_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
return;
}
+ /* disabllow gfxoff when we invalidate */
+ if (vmhub < AMDGPU_MMHUB0(0))
+ amdgpu_gfx_off_ctrl(adev, false);
+
/* This path is needed before KIQ/MES/GFXOFF are set up */
spin_lock(&adev->gmc.invalidate_lock);
@@ -873,6 +877,9 @@ static void gmc_v9_0_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
spin_unlock(&adev->gmc.invalidate_lock);
+ if (vmhub < AMDGPU_MMHUB0(0))
+ amdgpu_gfx_off_ctrl(adev, true);
+
if (j < adev->usec_timeout)
return;
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 02/31] drm/amdgpu/gmc10: disallow gfxoff around TLB flushes
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
2026-09-01 20:10 ` [PATCH 01/31] drm/amdgpu/gmc9: disallow gfxoff around TLB flushes Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 03/31] drm/amdgpu/gmc11: " Alex Deucher
` (29 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
We need to disallow gfxoff if we touch GC MMIO registers.
At the moment we use KIQ or MES for TLB flushes so
no intended functional change.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c | 7 +++++++
1 file changed, 7 insertions(+)
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
index 75a552782c979..8d88651b18ef9 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
@@ -265,6 +265,10 @@ static void gmc_v10_0_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
/* This path is needed before KIQ/MES/GFXOFF are set up */
hub_ip = (vmhub == AMDGPU_GFXHUB(0)) ? GC_HWIP : MMHUB_HWIP;
+ /* disabllow gfxoff when we invalidate */
+ if (hub_ip == GC_HWIP)
+ amdgpu_gfx_off_ctrl(adev, false);
+
spin_lock(&adev->gmc.invalidate_lock);
/*
* It may lose gpuvm invalidate acknowldege state across power-gating
@@ -313,6 +317,9 @@ static void gmc_v10_0_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
spin_unlock(&adev->gmc.invalidate_lock);
+ if (hub_ip == GC_HWIP)
+ amdgpu_gfx_off_ctrl(adev, true);
+
if (i >= adev->usec_timeout)
dev_err(adev->dev, "Timeout waiting for VM flush hub: %d!\n",
vmhub);
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 03/31] drm/amdgpu/gmc11: disallow gfxoff around TLB flushes
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
2026-09-01 20:10 ` [PATCH 01/31] drm/amdgpu/gmc9: disallow gfxoff around TLB flushes Alex Deucher
2026-09-01 20:10 ` [PATCH 02/31] drm/amdgpu/gmc10: " Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 04/31] drm/amdgpu/gmc12: " Alex Deucher
` (28 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
We need to disallow gfxoff if we touch GC MMIO registers.
At the moment we use KIQ or MES for TLB flushes so
no intended functional change.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c | 7 +++++++
1 file changed, 7 insertions(+)
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
index f454aff831b03..41ebc6ae182cc 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
@@ -253,6 +253,10 @@ static void gmc_v11_0_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
/* This path is needed before KIQ/MES/GFXOFF are set up */
hub_ip = (vmhub == AMDGPU_GFXHUB(0)) ? GC_HWIP : MMHUB_HWIP;
+ /* disabllow gfxoff when we invalidate */
+ if (hub_ip == GC_HWIP)
+ amdgpu_gfx_off_ctrl(adev, false);
+
spin_lock(&adev->gmc.invalidate_lock);
/*
* It may lose gpuvm invalidate acknowldege state across power-gating
@@ -306,6 +310,9 @@ static void gmc_v11_0_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
spin_unlock(&adev->gmc.invalidate_lock);
+ if (hub_ip == GC_HWIP)
+ amdgpu_gfx_off_ctrl(adev, true);
+
if (i >= adev->usec_timeout)
dev_err(adev->dev, "Timeout waiting for VM flush ACK!\n");
}
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 04/31] drm/amdgpu/gmc12: disallow gfxoff around TLB flushes
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (2 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 03/31] drm/amdgpu/gmc11: " Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 05/31] drm/amdgpu/gmc9: set vmhub funcs for gfxhub and mmhub Alex Deucher
` (27 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
We need to disallow gfxoff if we touch GC MMIO registers.
At the moment we use KIQ or MES for TLB flushes so
no intended functional change.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c | 8 +++++++-
1 file changed, 7 insertions(+), 1 deletion(-)
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
index 2c7e3eae7b237..3b7764b94de19 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
@@ -327,8 +327,14 @@ static void gmc_v12_0_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
return;
}
+ /* disabllow gfxoff when we invalidate */
+ if (vmhub == AMDGPU_GFXHUB(0))
+ amdgpu_gfx_off_ctrl(adev, false);
+
gmc_v12_0_flush_vm_hub(adev, vmid, vmhub, 0);
- return;
+
+ if (vmhub == AMDGPU_GFXHUB(0))
+ amdgpu_gfx_off_ctrl(adev, true);
}
/**
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 05/31] drm/amdgpu/gmc9: set vmhub funcs for gfxhub and mmhub
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (3 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 04/31] drm/amdgpu/gmc12: " Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 06/31] drm/amdgpu: add a gmc flag for using MMIO for TLB flush Alex Deucher
` (26 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
This was handled in gmc_v9_0 directly. Set the pointers
so that we can use rely on the callbacks in the generic gmc
helpers.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/gfxhub_v1_0.c | 33 ++++++++++++++++++++++++
drivers/gpu/drm/amd/amdgpu/gfxhub_v1_2.c | 33 ++++++++++++++++++++++++
drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c | 21 +--------------
drivers/gpu/drm/amd/amdgpu/mmhub_v1_0.c | 33 ++++++++++++++++++++++++
drivers/gpu/drm/amd/amdgpu/mmhub_v1_7.c | 32 +++++++++++++++++++++++
drivers/gpu/drm/amd/amdgpu/mmhub_v1_8.c | 32 +++++++++++++++++++++++
drivers/gpu/drm/amd/amdgpu/mmhub_v9_4.c | 32 +++++++++++++++++++++++
7 files changed, 196 insertions(+), 20 deletions(-)
diff --git a/drivers/gpu/drm/amd/amdgpu/gfxhub_v1_0.c b/drivers/gpu/drm/amd/amdgpu/gfxhub_v1_0.c
index bfe247b1a333c..f4e6c8b6c2cf9 100644
--- a/drivers/gpu/drm/amd/amdgpu/gfxhub_v1_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gfxhub_v1_0.c
@@ -410,6 +410,37 @@ static void gfxhub_v1_0_set_fault_enable_default(struct amdgpu_device *adev,
WREG32_SOC15(GC, 0, mmVM_L2_PROTECTION_FAULT_CNTL, tmp);
}
+static void
+gfxhub_v1_0_print_l2_protection_fault_status(struct amdgpu_device *adev,
+ uint32_t status)
+{
+ /* this is handled in gmc_v9_0.c already */
+}
+
+static uint32_t gfxhub_v1_0_get_invalidate_req(unsigned int vmid,
+ uint32_t flush_type)
+{
+ u32 req = 0;
+
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ,
+ PER_VMID_INVALIDATE_REQ, 1 << vmid);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, FLUSH_TYPE, flush_type);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PTES, 1);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE0, 1);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE1, 1);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE2, 1);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L1_PTES, 1);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ,
+ CLEAR_PROTECTION_FAULT_STATUS_ADDR, 0);
+
+ return req;
+}
+
+static const struct amdgpu_vmhub_funcs gfxhub_v1_0_vmhub_funcs = {
+ .print_l2_protection_fault_status = gfxhub_v1_0_print_l2_protection_fault_status,
+ .get_invalidate_req = gfxhub_v1_0_get_invalidate_req,
+};
+
static void gfxhub_v1_0_init(struct amdgpu_device *adev)
{
struct amdgpu_vmhub *hub = &adev->vmhub[AMDGPU_GFXHUB(0)];
@@ -439,6 +470,8 @@ static void gfxhub_v1_0_init(struct amdgpu_device *adev)
hub->eng_distance = mmVM_INVALIDATE_ENG1_REQ - mmVM_INVALIDATE_ENG0_REQ;
hub->eng_addr_distance = mmVM_INVALIDATE_ENG1_ADDR_RANGE_LO32 -
mmVM_INVALIDATE_ENG0_ADDR_RANGE_LO32;
+
+ hub->vmhub_funcs = &gfxhub_v1_0_vmhub_funcs;
}
const struct amdgpu_gfxhub_funcs gfxhub_v1_0_funcs = {
diff --git a/drivers/gpu/drm/amd/amdgpu/gfxhub_v1_2.c b/drivers/gpu/drm/amd/amdgpu/gfxhub_v1_2.c
index fbdf46070b38b..fd216e1670abe 100644
--- a/drivers/gpu/drm/amd/amdgpu/gfxhub_v1_2.c
+++ b/drivers/gpu/drm/amd/amdgpu/gfxhub_v1_2.c
@@ -539,6 +539,37 @@ static void gfxhub_v1_2_set_fault_enable_default(struct amdgpu_device *adev,
gfxhub_v1_2_xcc_set_fault_enable_default(adev, value, xcc_mask);
}
+static void
+gfxhub_v1_2_print_l2_protection_fault_status(struct amdgpu_device *adev,
+ uint32_t status)
+{
+ /* this is handled in gmc_v9_0.c already */
+}
+
+static uint32_t gfxhub_v1_2_get_invalidate_req(unsigned int vmid,
+ uint32_t flush_type)
+{
+ u32 req = 0;
+
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ,
+ PER_VMID_INVALIDATE_REQ, 1 << vmid);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, FLUSH_TYPE, flush_type);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PTES, 1);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE0, 1);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE1, 1);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE2, 1);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L1_PTES, 1);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ,
+ CLEAR_PROTECTION_FAULT_STATUS_ADDR, 0);
+
+ return req;
+}
+
+static const struct amdgpu_vmhub_funcs gfxhub_v1_2_vmhub_funcs = {
+ .print_l2_protection_fault_status = gfxhub_v1_2_print_l2_protection_fault_status,
+ .get_invalidate_req = gfxhub_v1_2_get_invalidate_req,
+};
+
static void gfxhub_v1_2_xcc_init(struct amdgpu_device *adev, uint32_t xcc_mask)
{
struct amdgpu_vmhub *hub;
@@ -577,6 +608,8 @@ static void gfxhub_v1_2_xcc_init(struct amdgpu_device *adev, uint32_t xcc_mask)
hub->eng_addr_distance =
regVM_INVALIDATE_ENG1_ADDR_RANGE_LO32 -
regVM_INVALIDATE_ENG0_ADDR_RANGE_LO32;
+
+ hub->vmhub_funcs = &gfxhub_v1_2_vmhub_funcs;
}
}
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
index 80f1cf1f21736..73b96efd45484 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
@@ -705,25 +705,6 @@ static void gmc_v9_0_set_irq_funcs(struct amdgpu_device *adev)
}
}
-static uint32_t gmc_v9_0_get_invalidate_req(unsigned int vmid,
- uint32_t flush_type)
-{
- u32 req = 0;
-
- req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ,
- PER_VMID_INVALIDATE_REQ, 1 << vmid);
- req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, FLUSH_TYPE, flush_type);
- req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PTES, 1);
- req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE0, 1);
- req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE1, 1);
- req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE2, 1);
- req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L1_PTES, 1);
- req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ,
- CLEAR_PROTECTION_FAULT_STATUS_ADDR, 0);
-
- return req;
-}
-
/**
* gmc_v9_0_use_invalidate_semaphore - judge whether to use semaphore
*
@@ -785,7 +766,7 @@ static void gmc_v9_0_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
BUG_ON(vmhub >= AMDGPU_MAX_VMHUBS);
hub = &adev->vmhub[vmhub];
- inv_req = gmc_v9_0_get_invalidate_req(vmid, flush_type);
+ inv_req = hub->vmhub_funcs->get_invalidate_req(vmid, flush_type);
sem = hub->vm_inv_eng0_sem + hub->eng_distance * eng;
req = hub->vm_inv_eng0_req + hub->eng_distance * eng;
ack = hub->vm_inv_eng0_ack + hub->eng_distance * eng;
diff --git a/drivers/gpu/drm/amd/amdgpu/mmhub_v1_0.c b/drivers/gpu/drm/amd/amdgpu/mmhub_v1_0.c
index 243eabda06077..46b0fc94007cc 100644
--- a/drivers/gpu/drm/amd/amdgpu/mmhub_v1_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/mmhub_v1_0.c
@@ -464,6 +464,37 @@ static void mmhub_v1_0_set_fault_enable_default(struct amdgpu_device *adev, bool
WREG32_SOC15(MMHUB, 0, mmVM_L2_PROTECTION_FAULT_CNTL, tmp);
}
+static uint32_t mmhub_v1_0_get_invalidate_req(unsigned int vmid,
+ uint32_t flush_type)
+{
+ u32 req = 0;
+
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ,
+ PER_VMID_INVALIDATE_REQ, 1 << vmid);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, FLUSH_TYPE, flush_type);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PTES, 1);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE0, 1);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE1, 1);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE2, 1);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L1_PTES, 1);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ,
+ CLEAR_PROTECTION_FAULT_STATUS_ADDR, 0);
+
+ return req;
+}
+
+static void
+mmhub_v1_0_print_l2_protection_fault_status(struct amdgpu_device *adev,
+ uint32_t status)
+{
+ /* this is handled in gmc_v9_0.c already */
+}
+
+static const struct amdgpu_vmhub_funcs mmhub_v1_0_vmhub_funcs = {
+ .print_l2_protection_fault_status = mmhub_v1_0_print_l2_protection_fault_status,
+ .get_invalidate_req = mmhub_v1_0_get_invalidate_req,
+};
+
static void mmhub_v1_0_init(struct amdgpu_device *adev)
{
struct amdgpu_vmhub *hub = &adev->vmhub[AMDGPU_MMHUB0(0)];
@@ -493,6 +524,8 @@ static void mmhub_v1_0_init(struct amdgpu_device *adev)
hub->eng_distance = mmVM_INVALIDATE_ENG1_REQ - mmVM_INVALIDATE_ENG0_REQ;
hub->eng_addr_distance = mmVM_INVALIDATE_ENG1_ADDR_RANGE_LO32 -
mmVM_INVALIDATE_ENG0_ADDR_RANGE_LO32;
+
+ hub->vmhub_funcs = &mmhub_v1_0_vmhub_funcs;
}
static void mmhub_v1_0_update_medium_grain_clock_gating(struct amdgpu_device *adev,
diff --git a/drivers/gpu/drm/amd/amdgpu/mmhub_v1_7.c b/drivers/gpu/drm/amd/amdgpu/mmhub_v1_7.c
index 2adee2b94c37d..84d1dc32728bd 100644
--- a/drivers/gpu/drm/amd/amdgpu/mmhub_v1_7.c
+++ b/drivers/gpu/drm/amd/amdgpu/mmhub_v1_7.c
@@ -448,6 +448,37 @@ static void mmhub_v1_7_set_fault_enable_default(struct amdgpu_device *adev, bool
WREG32_SOC15(MMHUB, 0, regVM_L2_PROTECTION_FAULT_CNTL, tmp);
}
+static void
+mmhub_v1_7_print_l2_protection_fault_status(struct amdgpu_device *adev,
+ uint32_t status)
+{
+ /* this is handled in gmc_v9_0.c already */
+}
+
+static uint32_t mmhub_v1_7_get_invalidate_req(unsigned int vmid,
+ uint32_t flush_type)
+{
+ u32 req = 0;
+
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ,
+ PER_VMID_INVALIDATE_REQ, 1 << vmid);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, FLUSH_TYPE, flush_type);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PTES, 1);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE0, 1);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE1, 1);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE2, 1);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L1_PTES, 1);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ,
+ CLEAR_PROTECTION_FAULT_STATUS_ADDR, 0);
+
+ return req;
+}
+
+static const struct amdgpu_vmhub_funcs mmhub_v1_7_vmhub_funcs = {
+ .print_l2_protection_fault_status = mmhub_v1_7_print_l2_protection_fault_status,
+ .get_invalidate_req = mmhub_v1_7_get_invalidate_req,
+};
+
static void mmhub_v1_7_init(struct amdgpu_device *adev)
{
struct amdgpu_vmhub *hub = &adev->vmhub[AMDGPU_MMHUB0(0)];
@@ -476,6 +507,7 @@ static void mmhub_v1_7_init(struct amdgpu_device *adev)
hub->eng_addr_distance = regVM_INVALIDATE_ENG1_ADDR_RANGE_LO32 -
regVM_INVALIDATE_ENG0_ADDR_RANGE_LO32;
+ hub->vmhub_funcs = &mmhub_v1_7_vmhub_funcs;
}
static void mmhub_v1_7_update_medium_grain_clock_gating(struct amdgpu_device *adev,
diff --git a/drivers/gpu/drm/amd/amdgpu/mmhub_v1_8.c b/drivers/gpu/drm/amd/amdgpu/mmhub_v1_8.c
index 47d07cd25fc41..4c9f41905b650 100644
--- a/drivers/gpu/drm/amd/amdgpu/mmhub_v1_8.c
+++ b/drivers/gpu/drm/amd/amdgpu/mmhub_v1_8.c
@@ -577,6 +577,37 @@ static void mmhub_v1_8_set_fault_enable_default(struct amdgpu_device *adev, bool
}
}
+static void
+mmhub_v1_8_print_l2_protection_fault_status(struct amdgpu_device *adev,
+ uint32_t status)
+{
+ /* this is handled in gmc_v9_0.c already */
+}
+
+static uint32_t mmhub_v1_8_get_invalidate_req(unsigned int vmid,
+ uint32_t flush_type)
+{
+ u32 req = 0;
+
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ,
+ PER_VMID_INVALIDATE_REQ, 1 << vmid);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, FLUSH_TYPE, flush_type);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PTES, 1);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE0, 1);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE1, 1);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE2, 1);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L1_PTES, 1);
+ req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ,
+ CLEAR_PROTECTION_FAULT_STATUS_ADDR, 0);
+
+ return req;
+}
+
+static const struct amdgpu_vmhub_funcs mmhub_v1_8_vmhub_funcs = {
+ .print_l2_protection_fault_status = mmhub_v1_8_print_l2_protection_fault_status,
+ .get_invalidate_req = mmhub_v1_8_get_invalidate_req,
+};
+
static void mmhub_v1_8_init(struct amdgpu_device *adev)
{
struct amdgpu_vmhub *hub;
@@ -610,6 +641,7 @@ static void mmhub_v1_8_init(struct amdgpu_device *adev)
regVM_INVALIDATE_ENG0_REQ;
hub->eng_addr_distance = regVM_INVALIDATE_ENG1_ADDR_RANGE_LO32 -
regVM_INVALIDATE_ENG0_ADDR_RANGE_LO32;
+ hub->vmhub_funcs = &mmhub_v1_8_vmhub_funcs;
}
}
diff --git a/drivers/gpu/drm/amd/amdgpu/mmhub_v9_4.c b/drivers/gpu/drm/amd/amdgpu/mmhub_v9_4.c
index fe0710b55c3ac..af7958931b152 100644
--- a/drivers/gpu/drm/amd/amdgpu/mmhub_v9_4.c
+++ b/drivers/gpu/drm/amd/amdgpu/mmhub_v9_4.c
@@ -536,6 +536,37 @@ static void mmhub_v9_4_set_fault_enable_default(struct amdgpu_device *adev, bool
}
}
+static void
+mmhub_v9_4_print_l2_protection_fault_status(struct amdgpu_device *adev,
+ uint32_t status)
+{
+ /* this is handled in gmc_v9_0.c already */
+}
+
+static uint32_t mmhub_v9_4_get_invalidate_req(unsigned int vmid,
+ uint32_t flush_type)
+{
+ u32 req = 0;
+
+ req = REG_SET_FIELD(req, VML2VC0_VM_INVALIDATE_ENG0_REQ,
+ PER_VMID_INVALIDATE_REQ, 1 << vmid);
+ req = REG_SET_FIELD(req, VML2VC0_VM_INVALIDATE_ENG0_REQ, FLUSH_TYPE, flush_type);
+ req = REG_SET_FIELD(req, VML2VC0_VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PTES, 1);
+ req = REG_SET_FIELD(req, VML2VC0_VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE0, 1);
+ req = REG_SET_FIELD(req, VML2VC0_VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE1, 1);
+ req = REG_SET_FIELD(req, VML2VC0_VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE2, 1);
+ req = REG_SET_FIELD(req, VML2VC0_VM_INVALIDATE_ENG0_REQ, INVALIDATE_L1_PTES, 1);
+ req = REG_SET_FIELD(req, VML2VC0_VM_INVALIDATE_ENG0_REQ,
+ CLEAR_PROTECTION_FAULT_STATUS_ADDR, 0);
+
+ return req;
+}
+
+static const struct amdgpu_vmhub_funcs mmhub_v9_4_vmhub_funcs = {
+ .print_l2_protection_fault_status = mmhub_v9_4_print_l2_protection_fault_status,
+ .get_invalidate_req = mmhub_v9_4_get_invalidate_req,
+};
+
static void mmhub_v9_4_init(struct amdgpu_device *adev)
{
struct amdgpu_vmhub *hub[MMHUB_NUM_INSTANCES] = {
@@ -584,6 +615,7 @@ static void mmhub_v9_4_init(struct amdgpu_device *adev)
mmVML2VC0_VM_INVALIDATE_ENG0_REQ;
hub[i]->eng_addr_distance = mmVML2VC0_VM_INVALIDATE_ENG1_ADDR_RANGE_LO32 -
mmVML2VC0_VM_INVALIDATE_ENG0_ADDR_RANGE_LO32;
+ hub[i]->vmhub_funcs = &mmhub_v9_4_vmhub_funcs;
}
}
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 06/31] drm/amdgpu: add a gmc flag for using MMIO for TLB flush
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (4 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 05/31] drm/amdgpu/gmc9: set vmhub funcs for gfxhub and mmhub Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 07/31] drm/amdgpu/gmc9: use MMIO for TLB flushes Alex Deucher
` (25 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
No intended functional change.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 2 ++
drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c | 5 ++++-
drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c | 5 ++++-
drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c | 5 ++++-
drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c | 6 +++---
drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c | 5 ++++-
6 files changed, 21 insertions(+), 7 deletions(-)
diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
index c0884797dd544..f8f4df37d5805 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
@@ -367,6 +367,8 @@ struct amdgpu_gmc {
bool flush_pasid_uses_kiq;
bool override_pte;
+
+ bool use_mmio_for_tlb_flush;
};
#define amdgpu_gmc_emit_flush_gpu_tlb(r, vmid, addr) (r)->adev->gmc.gmc_funcs->emit_flush_gpu_tlb((r), (vmid), (addr))
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
index 8d88651b18ef9..7ea1c4b878e97 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
@@ -256,7 +256,7 @@ static void gmc_v10_0_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
* properly under bare metal
*/
if (adev->gfx.kiq[0].ring.sched.ready && !adev->enable_mes &&
- (amdgpu_sriov_runtime(adev) || !amdgpu_sriov_vf(adev))) {
+ !adev->gmc.use_mmio_for_tlb_flush) {
amdgpu_gmc_fw_reg_write_reg_wait(adev, req, ack, inv_req,
1 << vmid, GET_INST(GC, 0));
return;
@@ -883,6 +883,9 @@ static int gmc_v10_0_sw_init(struct amdgpu_ip_block *ip_block)
if (r)
return r;
+ adev->gmc.use_mmio_for_tlb_flush =
+ !(amdgpu_sriov_runtime(adev) || !amdgpu_sriov_vf(adev));
+
return 0;
}
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
index 41ebc6ae182cc..fc730504be5f0 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
@@ -244,7 +244,7 @@ static void gmc_v11_0_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
* properly under bare metal
*/
if ((adev->gfx.kiq[0].ring.sched.ready || adev->mes.ring[0].sched.ready) &&
- (amdgpu_sriov_runtime(adev) || !amdgpu_sriov_vf(adev))) {
+ !adev->gmc.use_mmio_for_tlb_flush) {
amdgpu_gmc_fw_reg_write_reg_wait(adev, req, ack, inv_req,
1 << vmid, GET_INST(GC, 0));
return;
@@ -865,6 +865,9 @@ static int gmc_v11_0_sw_init(struct amdgpu_ip_block *ip_block)
if (r)
return r;
+ adev->gmc.use_mmio_for_tlb_flush =
+ !(amdgpu_sriov_runtime(adev) || !amdgpu_sriov_vf(adev));
+
return 0;
}
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
index 3b7764b94de19..3b1fe15fa1c01 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
@@ -315,7 +315,7 @@ static void gmc_v12_0_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
* properly under bare metal
*/
if ((adev->gfx.kiq[0].ring.sched.ready || adev->mes.ring[0].sched.ready) &&
- (amdgpu_sriov_runtime(adev) || !amdgpu_sriov_vf(adev))) {
+ !adev->gmc.use_mmio_for_tlb_flush) {
struct amdgpu_vmhub *hub = &adev->vmhub[vmhub];
const unsigned eng = 17;
u32 inv_req = hub->vmhub_funcs->get_invalidate_req(vmid, flush_type);
@@ -968,6 +968,9 @@ static int gmc_v12_0_sw_init(struct amdgpu_ip_block *ip_block)
if (r)
return r;
+ adev->gmc.use_mmio_for_tlb_flush =
+ !(amdgpu_sriov_runtime(adev) || !amdgpu_sriov_vf(adev));
+
return 0;
}
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c b/drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c
index 5fe43f7eab29d..6c0d2689cc05d 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c
@@ -375,9 +375,9 @@ static void gmc_v12_1_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
/* This is necessary for SRIOV as well as for GFXOFF to function
* properly under bare metal
*/
- if (((adev->gfx.kiq[inst].ring.sched.ready ||
- adev->mes.ring[MES_PIPE_INST(inst, 0)].sched.ready) &&
- (amdgpu_sriov_runtime(adev) || !amdgpu_sriov_vf(adev)))) {
+ if ((adev->gfx.kiq[inst].ring.sched.ready ||
+ adev->mes.ring[MES_PIPE_INST(inst, 0)].sched.ready) &&
+ !adev->gmc.use_mmio_for_tlb_flush) {
struct amdgpu_vmhub *hub = &adev->vmhub[vmhub];
const unsigned eng = 17;
u32 inv_req = hub->vmhub_funcs->get_invalidate_req(vmid, flush_type);
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
index 73b96efd45484..23de2caf13646 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
@@ -780,7 +780,7 @@ static void gmc_v9_0_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
* properly under bare metal
*/
if (adev->gfx.kiq[inst].ring.sched.ready &&
- (amdgpu_sriov_runtime(adev) || !amdgpu_sriov_vf(adev))) {
+ !adev->gmc.use_mmio_for_tlb_flush) {
uint32_t req = hub->vm_inv_eng0_req + hub->eng_distance * eng;
uint32_t ack = hub->vm_inv_eng0_ack + hub->eng_distance * eng;
@@ -2037,6 +2037,9 @@ static int gmc_v9_0_sw_init(struct amdgpu_ip_block *ip_block)
if (amdgpu_is_multi_aid(adev))
amdgpu_gmc_sysfs_init(adev);
+ adev->gmc.use_mmio_for_tlb_flush =
+ !(amdgpu_sriov_runtime(adev) || !amdgpu_sriov_vf(adev));
+
return 0;
}
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 07/31] drm/amdgpu/gmc9: use MMIO for TLB flushes
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (5 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 06/31] drm/amdgpu: add a gmc flag for using MMIO for TLB flush Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 08/31] drm/amdgpu/gmc10: " Alex Deucher
` (24 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
KIQ handles scheduling on GC 9 so it can get delayed if
there are a lot of queues active.
KIQ was used to avoid disallowing GFXOFF, but since the
KIQ will wake GFX anyway, just disallow it when we
flush and use MMIO on vega20 and older.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
index 23de2caf13646..a91d0ddb6e729 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
@@ -2038,7 +2038,8 @@ static int gmc_v9_0_sw_init(struct amdgpu_ip_block *ip_block)
amdgpu_gmc_sysfs_init(adev);
adev->gmc.use_mmio_for_tlb_flush =
- !(amdgpu_sriov_runtime(adev) || !amdgpu_sriov_vf(adev));
+ !amdgpu_sriov_runtime(adev) &&
+ (amdgpu_ip_version(adev, GC_HWIP, 0) <= IP_VERSION(9, 4, 0));
return 0;
}
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 08/31] drm/amdgpu/gmc10: use MMIO for TLB flushes
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (6 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 07/31] drm/amdgpu/gmc9: use MMIO for TLB flushes Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 09/31] drm/amdgpu/gmc11: " Alex Deucher
` (23 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
KIQ handles scheduling on GC 10 so it can get delayed if
there are a lot of queues active.
KIQ was used to avoid disallowing GFXOFF, but since the
KIQ will wake GFX anyway, just disallow it when we
flush and use MMIO.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
index 7ea1c4b878e97..23d9fe995a2b5 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
@@ -884,7 +884,7 @@ static int gmc_v10_0_sw_init(struct amdgpu_ip_block *ip_block)
return r;
adev->gmc.use_mmio_for_tlb_flush =
- !(amdgpu_sriov_runtime(adev) || !amdgpu_sriov_vf(adev));
+ !amdgpu_sriov_runtime(adev);
return 0;
}
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 09/31] drm/amdgpu/gmc11: use MMIO for TLB flushes
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (7 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 08/31] drm/amdgpu/gmc10: " Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 10/31] drm/amdgpu/gmc12: " Alex Deucher
` (22 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
MES only has one pipe on GC 11 which handles scheduling
so it can get delayed if there are a lot of queues active.
MES was used to avoid disallowing GFXOFF, but since the
MES will wake GFX anyway, just disallow it when we
flush and use MMIO.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
index fc730504be5f0..098e1340554c5 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
@@ -866,7 +866,7 @@ static int gmc_v11_0_sw_init(struct amdgpu_ip_block *ip_block)
return r;
adev->gmc.use_mmio_for_tlb_flush =
- !(amdgpu_sriov_runtime(adev) || !amdgpu_sriov_vf(adev));
+ !amdgpu_sriov_runtime(adev);
return 0;
}
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 10/31] drm/amdgpu/gmc12: use MMIO for TLB flushes
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (8 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 09/31] drm/amdgpu/gmc11: " Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 11/31] drm/amdgpu: add a buffer funcs callback for TLB invalidation Alex Deucher
` (21 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
MES has a dedicated pipe for these sort of tasks
on GC12 so it shouldn't be an issue.
MES is used to avoid disallowing GFXOFF, but since the
MES will wake GFX anyway to execute the flush, just disallow
it when we flush and use MMIO for consistency with
other chips.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c | 14 ++++++++++++--
1 file changed, 12 insertions(+), 2 deletions(-)
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
index 3b1fe15fa1c01..cffc818880d12 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
@@ -968,8 +968,18 @@ static int gmc_v12_0_sw_init(struct amdgpu_ip_block *ip_block)
if (r)
return r;
- adev->gmc.use_mmio_for_tlb_flush =
- !(amdgpu_sriov_runtime(adev) || !amdgpu_sriov_vf(adev));
+ switch (amdgpu_ip_version(adev, GC_HWIP, 0)) {
+ case IP_VERSION(12, 0, 0):
+ case IP_VERSION(12, 0, 1):
+ default:
+ adev->gmc.use_mmio_for_tlb_flush =
+ !amdgpu_sriov_runtime(adev);
+ break;
+ case IP_VERSION(12, 1, 0):
+ adev->gmc.use_mmio_for_tlb_flush =
+ !(amdgpu_sriov_runtime(adev) || !amdgpu_sriov_vf(adev));
+ break;
+ }
return 0;
}
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 11/31] drm/amdgpu: add a buffer funcs callback for TLB invalidation
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (9 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 10/31] drm/amdgpu/gmc12: " Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 12/31] drm/amdgpu/sdma5.0: add tlb invalidation buffer func callback Alex Deucher
` (20 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
Use this interface to issue TLB invalidations using
SDMA.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h | 18 ++++++++++++++++++
1 file changed, 18 insertions(+)
diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h b/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h
index 4f4e56022c970..4ab92d287675a 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h
@@ -155,6 +155,23 @@ struct amdgpu_buffer_funcs {
uint64_t dst_offset,
/* number of byte to fill */
uint32_t byte_count);
+
+ /* number of dw to reserve per operation */
+ unsigned tlb_inv_num_dw;
+
+ /* used for buffer clearing */
+ void (*emit_tlb_inv)(struct amdgpu_device *adev,
+ struct amdgpu_ib *ib,
+ /* vmid to target */
+ unsigned int vmid,
+ /* vmhub to target */
+ u32 vmhub,
+ /* inv eng to target */
+ u32 eng,
+ /* flush type */
+ u32 flush_type,
+ /* XCC to target */
+ u32 xcc_inst);
};
int amdgpu_sdma_reset_engine(struct amdgpu_device *adev, uint32_t instance_id,
@@ -162,6 +179,7 @@ int amdgpu_sdma_reset_engine(struct amdgpu_device *adev, uint32_t instance_id,
#define amdgpu_emit_copy_buffer(adev, ib, s, d, b, t) (adev)->mman.buffer_funcs->emit_copy_buffer((ib), (s), (d), (b), (t))
#define amdgpu_emit_fill_buffer(adev, ib, s, d, b) (adev)->mman.buffer_funcs->emit_fill_buffer((ib), (s), (d), (b))
+#define amdgpu_emit_tlb_inv(adev, ib, v, h, e, f, x) (adev)->mman.buffer_funcs->emit_tlb_inv((adev), (ib), (v), (h), (e), (f), (x))
struct amdgpu_sdma_instance *
amdgpu_sdma_get_instance_from_ring(struct amdgpu_ring *ring);
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 12/31] drm/amdgpu/sdma5.0: add tlb invalidation buffer func callback
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (10 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 11/31] drm/amdgpu: add a buffer funcs callback for TLB invalidation Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 13/31] drm/amdgpu/sdma5.2: " Alex Deucher
` (19 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
Will be used for TLB invalidation.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/sdma_v5_0.c | 49 ++++++++++++++++++++++++++
1 file changed, 49 insertions(+)
diff --git a/drivers/gpu/drm/amd/amdgpu/sdma_v5_0.c b/drivers/gpu/drm/amd/amdgpu/sdma_v5_0.c
index 97fee70dc2f64..8cb579ecf522e 100644
--- a/drivers/gpu/drm/amd/amdgpu/sdma_v5_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/sdma_v5_0.c
@@ -2030,6 +2030,52 @@ static void sdma_v5_0_emit_fill_buffer(struct amdgpu_ib *ib,
ib->ptr[ib->length_dw++] = byte_count - 1;
}
+/**
+ * sdma_v5_0_emit_tlb_inv - Invalidate TLB using the sDMA engine
+ *
+ * @adev: amdgpu device structure
+ * @ib: indirect buffer to fill
+ * @vmid: vmid to target
+ * @vmhub: vmhub to target
+ * @eng: invalidation engine to use
+ * @flush_type: type of flush (lightweight, heavyweight)
+ * @xcc_inst: XCC to target
+ *
+ * Invalidate TLB using the DMA engine.
+ */
+static void sdma_v5_0_emit_tlb_inv(struct amdgpu_device *adev,
+ struct amdgpu_ib *ib,
+ unsigned int vmid,
+ u32 vmhub,
+ u32 eng,
+ u32 flush_type,
+ u32 xcc_inst)
+{
+ struct amdgpu_vmhub *hub = &adev->vmhub[vmhub];
+ u32 inv_req = hub->vmhub_funcs->get_invalidate_req(vmid, flush_type);
+ u32 mmhub_eng, gfxhub_eng;
+
+ if (AMDGPU_IS_GFXHUB(vmhub)) {
+ mmhub_eng = 0x1f;
+ gfxhub_eng = eng;
+ } else {
+ mmhub_eng = eng;
+ gfxhub_eng = 0x1f;
+ }
+
+ /* Trigger invalidation. */
+ ib->ptr[ib->length_dw++] =
+ (SDMA_PKT_VM_INVALIDATION_HEADER_OP(SDMA_OP_POLL_REGMEM) |
+ SDMA_PKT_VM_INVALIDATION_HEADER_SUB_OP(SDMA_SUBOP_VM_INVALIDATION) |
+ SDMA_PKT_VM_INVALIDATION_HEADER_GFX_ENG_ID(gfxhub_eng) |
+ SDMA_PKT_VM_INVALIDATION_HEADER_MM_ENG_ID(mmhub_eng));
+ ib->ptr[ib->length_dw++] = inv_req;
+ ib->ptr[ib->length_dw++] = 0xFFFFFFFF;
+ ib->ptr[ib->length_dw++] =
+ (SDMA_PKT_VM_INVALIDATION_ADDRESSRANGEHI_INVALIDATEACK(1 << vmid) |
+ SDMA_PKT_VM_INVALIDATION_ADDRESSRANGEHI_ADDRESSRANGEHI(0x1F));
+}
+
static const struct amdgpu_buffer_funcs sdma_v5_0_buffer_funcs = {
.copy_max_bytes = 0x400000,
.copy_num_dw = 7,
@@ -2038,6 +2084,9 @@ static const struct amdgpu_buffer_funcs sdma_v5_0_buffer_funcs = {
.fill_max_bytes = 0x400000,
.fill_num_dw = 5,
.emit_fill_buffer = sdma_v5_0_emit_fill_buffer,
+
+ .tlb_inv_num_dw = 4,
+ .emit_tlb_inv = sdma_v5_0_emit_tlb_inv,
};
static void sdma_v5_0_set_buffer_funcs(struct amdgpu_device *adev)
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 13/31] drm/amdgpu/sdma5.2: add tlb invalidation buffer func callback
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (11 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 12/31] drm/amdgpu/sdma5.0: add tlb invalidation buffer func callback Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 14/31] drm/amdgpu/sdma6: " Alex Deucher
` (18 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
Will be used for TLB invalidation.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/sdma_v5_2.c | 49 ++++++++++++++++++++++++++
1 file changed, 49 insertions(+)
diff --git a/drivers/gpu/drm/amd/amdgpu/sdma_v5_2.c b/drivers/gpu/drm/amd/amdgpu/sdma_v5_2.c
index 35cdf6c149f86..a10086de35957 100644
--- a/drivers/gpu/drm/amd/amdgpu/sdma_v5_2.c
+++ b/drivers/gpu/drm/amd/amdgpu/sdma_v5_2.c
@@ -2038,6 +2038,52 @@ static void sdma_v5_2_emit_fill_buffer(struct amdgpu_ib *ib,
ib->ptr[ib->length_dw++] = byte_count - 1;
}
+/**
+ * sdma_v5_2_emit_tlb_inv - Invalidate TLB using the sDMA engine
+ *
+ * @adev: amdgpu device structure
+ * @ib: indirect buffer to fill
+ * @vmid: vmid to target
+ * @vmhub: vmhub to target
+ * @eng: invalidation engine to use
+ * @flush_type: type of flush (lightweight, heavyweight)
+ * @xcc_inst: XCC to target
+ *
+ * Invalidate TLB using the DMA engine.
+ */
+static void sdma_v5_2_emit_tlb_inv(struct amdgpu_device *adev,
+ struct amdgpu_ib *ib,
+ unsigned int vmid,
+ u32 vmhub,
+ u32 eng,
+ u32 flush_type,
+ u32 xcc_inst)
+{
+ struct amdgpu_vmhub *hub = &adev->vmhub[vmhub];
+ u32 inv_req = hub->vmhub_funcs->get_invalidate_req(vmid, flush_type);
+ u32 mmhub_eng, gfxhub_eng;
+
+ if (AMDGPU_IS_GFXHUB(vmhub)) {
+ mmhub_eng = 0x1f;
+ gfxhub_eng = eng;
+ } else {
+ mmhub_eng = eng;
+ gfxhub_eng = 0x1f;
+ }
+
+ /* Trigger invalidation. */
+ ib->ptr[ib->length_dw++] =
+ (SDMA_PKT_VM_INVALIDATION_HEADER_OP(SDMA_OP_POLL_REGMEM) |
+ SDMA_PKT_VM_INVALIDATION_HEADER_SUB_OP(SDMA_SUBOP_VM_INVALIDATION) |
+ SDMA_PKT_VM_INVALIDATION_HEADER_GFX_ENG_ID(gfxhub_eng) |
+ SDMA_PKT_VM_INVALIDATION_HEADER_MM_ENG_ID(mmhub_eng));
+ ib->ptr[ib->length_dw++] = inv_req;
+ ib->ptr[ib->length_dw++] = 0xFFFFFFFF;
+ ib->ptr[ib->length_dw++] =
+ (SDMA_PKT_VM_INVALIDATION_ADDRESSRANGEHI_INVALIDATEACK(1 << vmid) |
+ SDMA_PKT_VM_INVALIDATION_ADDRESSRANGEHI_ADDRESSRANGEHI(0x1F));
+}
+
static const struct amdgpu_buffer_funcs sdma_v5_2_buffer_funcs = {
.copy_max_bytes = 1 << 30,
.copy_num_dw = 7,
@@ -2046,6 +2092,9 @@ static const struct amdgpu_buffer_funcs sdma_v5_2_buffer_funcs = {
.fill_max_bytes = 1 << 30, /* HW supports 1 << 30, but PAL uses 1 << 22 */
.fill_num_dw = 5,
.emit_fill_buffer = sdma_v5_2_emit_fill_buffer,
+
+ .tlb_inv_num_dw = 4,
+ .emit_tlb_inv = sdma_v5_2_emit_tlb_inv,
};
static void sdma_v5_2_set_buffer_funcs(struct amdgpu_device *adev)
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 14/31] drm/amdgpu/sdma6: add tlb invalidation buffer func callback
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (12 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 13/31] drm/amdgpu/sdma5.2: " Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 15/31] drm/amdgpu/sdma7: " Alex Deucher
` (17 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
Will be used for TLB invalidation.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/sdma_v6_0.c | 49 ++++++++++++++++++++++++++
1 file changed, 49 insertions(+)
diff --git a/drivers/gpu/drm/amd/amdgpu/sdma_v6_0.c b/drivers/gpu/drm/amd/amdgpu/sdma_v6_0.c
index 303fd7d1b7c88..7b84c996e2359 100644
--- a/drivers/gpu/drm/amd/amdgpu/sdma_v6_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/sdma_v6_0.c
@@ -1846,6 +1846,52 @@ static void sdma_v6_0_emit_fill_buffer(struct amdgpu_ib *ib,
ib->ptr[ib->length_dw++] = byte_count - 1;
}
+/**
+ * sdma_v6_0_emit_tlb_inv - Invalidate TLB using the sDMA engine
+ *
+ * @adev: amdgpu device structure
+ * @ib: indirect buffer to fill
+ * @vmid: vmid to target
+ * @vmhub: vmhub to target
+ * @eng: invalidation engine to use
+ * @flush_type: type of flush (lightweight, heavyweight)
+ * @xcc_inst: XCC to target
+ *
+ * Invalidate TLB using the DMA engine.
+ */
+static void sdma_v6_0_emit_tlb_inv(struct amdgpu_device *adev,
+ struct amdgpu_ib *ib,
+ unsigned int vmid,
+ u32 vmhub,
+ u32 eng,
+ u32 flush_type,
+ u32 xcc_inst)
+{
+ struct amdgpu_vmhub *hub = &adev->vmhub[vmhub];
+ u32 inv_req = hub->vmhub_funcs->get_invalidate_req(vmid, flush_type);
+ u32 mmhub_eng, gfxhub_eng;
+
+ if (AMDGPU_IS_GFXHUB(vmhub)) {
+ mmhub_eng = 0x1f;
+ gfxhub_eng = eng;
+ } else {
+ mmhub_eng = eng;
+ gfxhub_eng = 0x1f;
+ }
+
+ /* Trigger invalidation. */
+ ib->ptr[ib->length_dw++] =
+ (SDMA_PKT_VM_INVALIDATION_HEADER_OP(SDMA_OP_POLL_REGMEM) |
+ SDMA_PKT_VM_INVALIDATION_HEADER_SUB_OP(SDMA_SUBOP_VM_INVALIDATION) |
+ SDMA_PKT_VM_INVALIDATION_HEADER_GFX_ENG_ID(gfxhub_eng) |
+ SDMA_PKT_VM_INVALIDATION_HEADER_MM_ENG_ID(mmhub_eng));
+ ib->ptr[ib->length_dw++] = inv_req;
+ ib->ptr[ib->length_dw++] = 0xFFFFFFFF;
+ ib->ptr[ib->length_dw++] =
+ (SDMA_PKT_VM_INVALIDATION_ADDRESSRANGEHI_INVALIDATEACK(1 << vmid) |
+ SDMA_PKT_VM_INVALIDATION_ADDRESSRANGEHI_ADDRESSRANGEHI(0x1F));
+}
+
static const struct amdgpu_buffer_funcs sdma_v6_0_buffer_funcs = {
.copy_max_bytes = 1 << 30,
.copy_num_dw = 7,
@@ -1854,6 +1900,9 @@ static const struct amdgpu_buffer_funcs sdma_v6_0_buffer_funcs = {
.fill_max_bytes = 1 << 30,
.fill_num_dw = 5,
.emit_fill_buffer = sdma_v6_0_emit_fill_buffer,
+
+ .tlb_inv_num_dw = 4,
+ .emit_tlb_inv = sdma_v6_0_emit_tlb_inv,
};
static void sdma_v6_0_set_buffer_funcs(struct amdgpu_device *adev)
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 15/31] drm/amdgpu/sdma7: add tlb invalidation buffer func callback
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (13 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 14/31] drm/amdgpu/sdma6: " Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 16/31] drm/amdgpu: simplify amdgpu_gmc_flush_gpu_tlb() Alex Deucher
` (16 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
Will be used for TLB invalidation.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/sdma_v7_0.c | 48 ++++++++++++++++++++++++++
1 file changed, 48 insertions(+)
diff --git a/drivers/gpu/drm/amd/amdgpu/sdma_v7_0.c b/drivers/gpu/drm/amd/amdgpu/sdma_v7_0.c
index d5552f206e4dc..dab2d08d8766b 100644
--- a/drivers/gpu/drm/amd/amdgpu/sdma_v7_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/sdma_v7_0.c
@@ -1796,6 +1796,52 @@ static void sdma_v7_0_emit_fill_buffer(struct amdgpu_ib *ib,
ib->ptr[ib->length_dw++] = byte_count - 1;
}
+/**
+ * sdma_v7_0_emit_tlb_inv - Invalidate TLB using the sDMA engine
+ *
+ * @adev: amdgpu device structure
+ * @ib: indirect buffer to fill
+ * @vmid: vmid to target
+ * @vmhub: vmhub to target
+ * @eng: invalidation engine to use
+ * @flush_type: type of flush (lightweight, heavyweight)
+ * @xcc_inst: XCC to target
+ *
+ * Invalidate TLB using the DMA engine.
+ */
+static void sdma_v7_0_emit_tlb_inv(struct amdgpu_device *adev,
+ struct amdgpu_ib *ib,
+ unsigned int vmid,
+ u32 vmhub,
+ u32 eng,
+ u32 flush_type,
+ u32 xcc_inst)
+{
+ struct amdgpu_vmhub *hub = &adev->vmhub[vmhub];
+ u32 inv_req = hub->vmhub_funcs->get_invalidate_req(vmid, flush_type);
+ u32 mmhub_eng, gfxhub_eng;
+
+ if (AMDGPU_IS_GFXHUB(vmhub)) {
+ mmhub_eng = 0x1f;
+ gfxhub_eng = eng;
+ } else {
+ mmhub_eng = eng;
+ gfxhub_eng = 0x1f;
+ }
+
+ /* Trigger invalidation. */
+ ib->ptr[ib->length_dw++] =
+ (SDMA_PKT_VM_INVALIDATION_HEADER_OP(SDMA_OP_POLL_REGMEM) |
+ SDMA_PKT_VM_INVALIDATION_HEADER_SUB_OP(SDMA_SUBOP_VM_INVALIDATION) |
+ SDMA_PKT_VM_INVALIDATION_HEADER_GFX_ENG_ID(gfxhub_eng) |
+ SDMA_PKT_VM_INVALIDATION_HEADER_MM_ENG_ID(mmhub_eng));
+ ib->ptr[ib->length_dw++] = inv_req;
+ ib->ptr[ib->length_dw++] = 0xFFFFFFFF;
+ ib->ptr[ib->length_dw++] =
+ (SDMA_PKT_VM_INVALIDATION_ADDRESSRANGEHI_INVALIDATEACK(1 << vmid) |
+ SDMA_PKT_VM_INVALIDATION_ADDRESSRANGEHI_ADDRESSRANGEHI(0x1F));
+}
+
static const struct amdgpu_buffer_funcs sdma_v7_0_buffer_funcs = {
.copy_max_bytes = 1 << 30,
.copy_num_dw = 8,
@@ -1803,6 +1849,8 @@ static const struct amdgpu_buffer_funcs sdma_v7_0_buffer_funcs = {
.fill_max_bytes = 1 << 30,
.fill_num_dw = 5,
.emit_fill_buffer = sdma_v7_0_emit_fill_buffer,
+ .tlb_inv_num_dw = 4,
+ .emit_tlb_inv = sdma_v7_0_emit_tlb_inv,
};
static void sdma_v7_0_set_buffer_funcs(struct amdgpu_device *adev)
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 16/31] drm/amdgpu: simplify amdgpu_gmc_flush_gpu_tlb()
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (14 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 15/31] drm/amdgpu/sdma7: " Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 17/31] drm/amdgpu: add tlb invalidation method enum Alex Deucher
` (15 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
It's only used for GART TLB flushes so rename it and
remove additional parameters that are never used.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/amdgpu_gart.c | 2 +-
drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c | 16 ++++++----------
drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 4 ++--
3 files changed, 9 insertions(+), 13 deletions(-)
diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gart.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_gart.c
index c4c21dbbbdbf8..0db997c8fa03a 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gart.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gart.c
@@ -502,7 +502,7 @@ void amdgpu_gart_invalidate_tlb(struct amdgpu_device *adev)
up_read(&adev->reset_domain->sem);
}
for_each_set_bit(i, adev->vmhubs_mask, AMDGPU_MAX_VMHUBS)
- amdgpu_gmc_flush_gpu_tlb(adev, 0, i, 0);
+ amdgpu_gmc_flush_gpu_tlb_gart(adev, i);
}
/**
diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c
index 2f6d20c00ce29..6af4ed2170b06 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c
@@ -714,8 +714,8 @@ int amdgpu_gmc_allocate_vm_inv_eng(struct amdgpu_device *adev)
return 0;
}
-void amdgpu_gmc_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
- uint32_t vmhub, uint32_t flush_type)
+void amdgpu_gmc_flush_gpu_tlb_gart(struct amdgpu_device *adev,
+ uint32_t vmhub)
{
struct amdgpu_ring *ring;
struct amdgpu_vmhub *hub = &adev->vmhub[vmhub];
@@ -725,7 +725,7 @@ void amdgpu_gmc_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
ring = to_amdgpu_ring(adev->mman.buffer_funcs_scheds[0]);
- if (!hub->sdma_invalidation_workaround || vmid ||
+ if (!hub->sdma_invalidation_workaround ||
!adev->mman.buffer_funcs_enabled || !adev->ib_pool_ready ||
!ring->sched.ready) {
/*
@@ -736,15 +736,11 @@ void amdgpu_gmc_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
return;
if (adev->gmc.flush_tlb_needs_extra_type_2)
- adev->gmc.gmc_funcs->flush_gpu_tlb(adev, vmid,
+ adev->gmc.gmc_funcs->flush_gpu_tlb(adev, 0,
vmhub, 2);
- if (adev->gmc.flush_tlb_needs_extra_type_0 && flush_type == 2)
- adev->gmc.gmc_funcs->flush_gpu_tlb(adev, vmid,
- vmhub, 0);
-
- adev->gmc.gmc_funcs->flush_gpu_tlb(adev, vmid, vmhub,
- flush_type);
+ adev->gmc.gmc_funcs->flush_gpu_tlb(adev, 0, vmhub,
+ 0);
up_read(&adev->reset_domain->sem);
return;
}
diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
index f8f4df37d5805..23954e3180a8c 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
@@ -441,8 +441,8 @@ int amdgpu_gmc_handle_retry_fault(struct amdgpu_device *adev,
bool write_fault);
int amdgpu_gmc_ras_sw_init(struct amdgpu_device *adev);
int amdgpu_gmc_allocate_vm_inv_eng(struct amdgpu_device *adev);
-void amdgpu_gmc_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
- uint32_t vmhub, uint32_t flush_type);
+void amdgpu_gmc_flush_gpu_tlb_gart(struct amdgpu_device *adev,
+ uint32_t vmhub);
int amdgpu_gmc_flush_gpu_tlb_pasid(struct amdgpu_device *adev, uint16_t pasid,
uint32_t flush_type, bool all_hub,
uint32_t inst);
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 17/31] drm/amdgpu: add tlb invalidation method enum
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (15 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 16/31] drm/amdgpu: simplify amdgpu_gmc_flush_gpu_tlb() Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 18/31] drm/amdgpu: plumb tlb inv method in amdgpu_gmc_flush_gpu_tlb_gart() Alex Deucher
` (14 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
Use it to track invalidation method.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 10 ++++++++++
1 file changed, 10 insertions(+)
diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
index 23954e3180a8c..204cf1c360896 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
@@ -98,6 +98,13 @@ enum amdgpu_memory_partition {
#define AMDGPU_GMC121_FAULT_SOURCE_DATA_WRITE 0x200000
#define AMDGPU_GMC121_FAULT_SOURCE_DATA_EXE 0x100000
+enum amdgpu_tlb_inv_method {
+ AMDGPU_TLB_INV_METHOD_MMIO = 0,
+ AMDGPU_TLB_INV_METHOD_SDMA,
+ AMDGPU_TLB_INV_METHOD_KIQ,
+ AMDGPU_TLB_INV_METHOD_MES
+};
+
/*
* GMC page fault information
*/
@@ -369,6 +376,9 @@ struct amdgpu_gmc {
bool override_pte;
bool use_mmio_for_tlb_flush;
+
+ enum amdgpu_tlb_inv_method gart_inv_method;
+ enum amdgpu_tlb_inv_method pasid_inv_method;
};
#define amdgpu_gmc_emit_flush_gpu_tlb(r, vmid, addr) (r)->adev->gmc.gmc_funcs->emit_flush_gpu_tlb((r), (vmid), (addr))
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 18/31] drm/amdgpu: plumb tlb inv method in amdgpu_gmc_flush_gpu_tlb_gart()
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (16 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 17/31] drm/amdgpu: add tlb invalidation method enum Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 19/31] drm/amdgpu: uplevel reset check " Alex Deucher
` (13 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
Plumb the invalidation method into amdgpu_gmc_flush_gpu_tlb_gart().
Keep the current behavior.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c | 16 +++++++++++++---
1 file changed, 13 insertions(+), 3 deletions(-)
diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c
index 6af4ed2170b06..779f9b0974d2e 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c
@@ -719,15 +719,25 @@ void amdgpu_gmc_flush_gpu_tlb_gart(struct amdgpu_device *adev,
{
struct amdgpu_ring *ring;
struct amdgpu_vmhub *hub = &adev->vmhub[vmhub];
+ bool use_mmio;
struct dma_fence *fence;
struct amdgpu_job *job;
int r;
ring = to_amdgpu_ring(adev->mman.buffer_funcs_scheds[0]);
- if (!hub->sdma_invalidation_workaround ||
- !adev->mman.buffer_funcs_enabled || !adev->ib_pool_ready ||
- !ring->sched.ready) {
+ switch (adev->gmc.gart_inv_method) {
+ case AMDGPU_TLB_INV_METHOD_MMIO:
+ default:
+ use_mmio = !hub->sdma_invalidation_workaround;
+ break;
+ case AMDGPU_TLB_INV_METHOD_SDMA:
+ use_mmio = false;
+ break;
+ }
+
+ if (!adev->mman.buffer_funcs_enabled ||
+ !adev->ib_pool_ready || !ring->sched.ready || use_mmio) {
/*
* A GPU reset should flush all TLBs anyway, so no need to do
* this while one is ongoing.
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 19/31] drm/amdgpu: uplevel reset check in amdgpu_gmc_flush_gpu_tlb_gart()
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (17 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 18/31] drm/amdgpu: plumb tlb inv method in amdgpu_gmc_flush_gpu_tlb_gart() Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 20/31] drm/amdgpu/gmc: add new callback to lookup vmid to pasid mapping Alex Deucher
` (12 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
Check for both the MMIO and SDMA pathes.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c | 16 +++++++++-------
1 file changed, 9 insertions(+), 7 deletions(-)
diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c
index 779f9b0974d2e..8a975eddd75c7 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c
@@ -724,6 +724,13 @@ void amdgpu_gmc_flush_gpu_tlb_gart(struct amdgpu_device *adev,
struct amdgpu_job *job;
int r;
+ /*
+ * A GPU reset should flush all TLBs anyway, so no need to do
+ * this while one is ongoing.
+ */
+ if (!down_read_trylock(&adev->reset_domain->sem))
+ return;
+
ring = to_amdgpu_ring(adev->mman.buffer_funcs_scheds[0]);
switch (adev->gmc.gart_inv_method) {
@@ -738,13 +745,6 @@ void amdgpu_gmc_flush_gpu_tlb_gart(struct amdgpu_device *adev,
if (!adev->mman.buffer_funcs_enabled ||
!adev->ib_pool_ready || !ring->sched.ready || use_mmio) {
- /*
- * A GPU reset should flush all TLBs anyway, so no need to do
- * this while one is ongoing.
- */
- if (!down_read_trylock(&adev->reset_domain->sem))
- return;
-
if (adev->gmc.flush_tlb_needs_extra_type_2)
adev->gmc.gmc_funcs->flush_gpu_tlb(adev, 0,
vmhub, 2);
@@ -778,11 +778,13 @@ void amdgpu_gmc_flush_gpu_tlb_gart(struct amdgpu_device *adev,
dma_fence_wait(fence, false);
dma_fence_put(fence);
+ up_read(&adev->reset_domain->sem);
return;
error_alloc:
mutex_unlock(&adev->mman.default_entity.lock);
+ up_read(&adev->reset_domain->sem);
dev_err(adev->dev, "Error flushing GPU TLB using the SDMA (%d)!\n", r);
}
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 20/31] drm/amdgpu/gmc: add new callback to lookup vmid to pasid mapping
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (18 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 19/31] drm/amdgpu: uplevel reset check " Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-02 6:26 ` Zhang, Jesse(Jie)
2026-09-01 20:10 ` [PATCH 21/31] drm/amdgpu: add a gmc callback for the inv semaphore Alex Deucher
` (11 subsequent siblings)
31 siblings, 1 reply; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
Look up the mapping so we know which vmid to flush.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 7 ++++++-
drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c | 6 ++++--
drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c | 6 ++++--
drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c | 6 ++++--
drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c | 1 +
drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c | 6 ++++--
6 files changed, 23 insertions(+), 9 deletions(-)
diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
index 204cf1c360896..27785b1b38ccd 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
@@ -171,6 +171,10 @@ struct amdgpu_gmc_funcs {
/* Change the VMID -> PASID mapping */
void (*emit_pasid_mapping)(struct amdgpu_ring *ring, unsigned vmid,
unsigned pasid);
+ /* look up the vmids for the pasid */
+ bool (*get_vmid_pasid_mapping_info)(struct amdgpu_device *adev,
+ uint8_t vmid, uint8_t inst,
+ uint16_t *p_pasid);
/* enable/disable PRT support */
void (*set_prt)(struct amdgpu_device *adev, bool enable);
/* get the pde for a given mc addr */
@@ -384,7 +388,8 @@ struct amdgpu_gmc {
#define amdgpu_gmc_emit_flush_gpu_tlb(r, vmid, addr) (r)->adev->gmc.gmc_funcs->emit_flush_gpu_tlb((r), (vmid), (addr))
#define amdgpu_gmc_emit_pasid_mapping(r, vmid, pasid) (r)->adev->gmc.gmc_funcs->emit_pasid_mapping((r), (vmid), (pasid))
#define amdgpu_gmc_get_vm_pde(adev, level, dst, flags) (adev)->gmc.gmc_funcs->get_vm_pde((adev), (level), (dst), (flags))
-#define amdgpu_gmc_get_vm_pte(adev, vm, bo, vm_flags, pte_flags) \
+#define amdgpu_gmc_get_vmid_pasid_mapping_info(adev, v, i, p) (adev)->gmc.gmc_funcs->get_vmid_pasid_mapping_info((a), (v), (i), (p))
+#define amdgpu_gmc_get_vm_pte(adev, vm, bo, vm_flags, pte_flags) \
((adev)->gmc.gmc_funcs->get_vm_pte((adev), (vm), (bo), (vm_flags), \
(pte_flags)))
#define amdgpu_gmc_override_vm_pte_flags(adev, vm, addr, pte_flags) \
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
index 23d9fe995a2b5..0028639448956 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
@@ -204,7 +204,8 @@ static bool gmc_v10_0_use_invalidate_semaphore(struct amdgpu_device *adev,
static bool gmc_v10_0_get_atc_vmid_pasid_mapping_info(
struct amdgpu_device *adev,
- uint8_t vmid, uint16_t *p_pasid)
+ uint8_t vmid, uint8_t inst,
+ uint16_t *p_pasid)
{
uint32_t value;
@@ -346,7 +347,7 @@ static void gmc_v10_0_flush_gpu_tlb_pasid(struct amdgpu_device *adev,
for (vmid = 1; vmid < AMDGPU_NUM_VMID; vmid++) {
bool valid;
- valid = gmc_v10_0_get_atc_vmid_pasid_mapping_info(adev, vmid,
+ valid = gmc_v10_0_get_atc_vmid_pasid_mapping_info(adev, vmid, 0,
&queried);
if (!valid || queried != pasid)
continue;
@@ -555,6 +556,7 @@ static const struct amdgpu_gmc_funcs gmc_v10_0_gmc_funcs = {
.flush_gpu_tlb_pasid = gmc_v10_0_flush_gpu_tlb_pasid,
.emit_flush_gpu_tlb = gmc_v10_0_emit_flush_gpu_tlb,
.emit_pasid_mapping = gmc_v10_0_emit_pasid_mapping,
+ .get_vmid_pasid_mapping_info = gmc_v10_0_get_atc_vmid_pasid_mapping_info,
.get_vm_pde = gmc_v10_0_get_vm_pde,
.get_vm_pte = gmc_v10_0_get_vm_pte,
.get_vbios_fb_size = gmc_v10_0_get_vbios_fb_size,
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
index 098e1340554c5..b34bd7881726d 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
@@ -200,7 +200,8 @@ static bool gmc_v11_0_use_invalidate_semaphore(struct amdgpu_device *adev,
static bool gmc_v11_0_get_vmid_pasid_mapping_info(
struct amdgpu_device *adev,
- uint8_t vmid, uint16_t *p_pasid)
+ uint8_t vmid, uint8_t inst,
+ uint16_t *p_pasid)
{
*p_pasid = RREG32(SOC15_REG_OFFSET(OSSSYS, 0, regIH_VMID_0_LUT) + vmid) & 0xffff;
@@ -338,7 +339,7 @@ static void gmc_v11_0_flush_gpu_tlb_pasid(struct amdgpu_device *adev,
for (vmid = 1; vmid < 16; vmid++) {
bool valid;
- valid = gmc_v11_0_get_vmid_pasid_mapping_info(adev, vmid,
+ valid = gmc_v11_0_get_vmid_pasid_mapping_info(adev, vmid, 0,
&queried);
if (!valid || queried != pasid)
continue;
@@ -546,6 +547,7 @@ static const struct amdgpu_gmc_funcs gmc_v11_0_gmc_funcs = {
.flush_gpu_tlb_pasid = gmc_v11_0_flush_gpu_tlb_pasid,
.emit_flush_gpu_tlb = gmc_v11_0_emit_flush_gpu_tlb,
.emit_pasid_mapping = gmc_v11_0_emit_pasid_mapping,
+ .get_vmid_pasid_mapping_info = gmc_v11_0_get_vmid_pasid_mapping_info,
.get_vm_pde = gmc_v11_0_get_vm_pde,
.get_vm_pte = gmc_v11_0_get_vm_pte,
.get_vbios_fb_size = gmc_v11_0_get_vbios_fb_size,
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
index cffc818880d12..9179dc0787f13 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
@@ -196,7 +196,8 @@ static bool gmc_v12_0_use_invalidate_semaphore(struct amdgpu_device *adev,
static bool gmc_v12_0_get_vmid_pasid_mapping_info(
struct amdgpu_device *adev,
- uint8_t vmid, uint16_t *p_pasid)
+ uint8_t vmid, uint8_t inst,
+ uint16_t *p_pasid)
{
*p_pasid = RREG32(SOC15_REG_OFFSET(OSSSYS, 0, regIH_VMID_0_LUT) + vmid) & 0xffff;
@@ -374,7 +375,7 @@ static void gmc_v12_0_flush_gpu_tlb_pasid(struct amdgpu_device *adev,
for (vmid = 1; vmid < 16; vmid++) {
bool valid;
- valid = gmc_v12_0_get_vmid_pasid_mapping_info(adev, vmid,
+ valid = gmc_v12_0_get_vmid_pasid_mapping_info(adev, vmid, 0,
&queried);
if (!valid || queried != pasid)
continue;
@@ -581,6 +582,7 @@ static const struct amdgpu_gmc_funcs gmc_v12_0_gmc_funcs = {
.flush_gpu_tlb_pasid = gmc_v12_0_flush_gpu_tlb_pasid,
.emit_flush_gpu_tlb = gmc_v12_0_emit_flush_gpu_tlb,
.emit_pasid_mapping = gmc_v12_0_emit_pasid_mapping,
+ .get_vmid_pasid_mapping_info = gmc_v12_0_get_vmid_pasid_mapping_info,
.get_vm_pde = gmc_v12_0_get_vm_pde,
.get_vm_pte = gmc_v12_0_get_vm_pte,
.get_vbios_fb_size = gmc_v12_0_get_vbios_fb_size,
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c b/drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c
index 6c0d2689cc05d..3fa1ec3dca273 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c
@@ -668,6 +668,7 @@ static const struct amdgpu_gmc_funcs gmc_v12_1_gmc_funcs = {
.flush_gpu_tlb_pasid = gmc_v12_1_flush_gpu_tlb_pasid,
.emit_flush_gpu_tlb = gmc_v12_1_emit_flush_gpu_tlb,
.emit_pasid_mapping = gmc_v12_1_emit_pasid_mapping,
+ .get_vmid_pasid_mapping_info = gmc_v12_1_get_vmid_pasid_mapping_info,
.get_vm_pde = gmc_v12_1_get_vm_pde,
.get_vm_pte = gmc_v12_1_get_vm_pte,
.query_mem_partition_mode = &amdgpu_gmc_query_memory_partition,
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
index a91d0ddb6e729..0019706b0b2a9 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
@@ -727,7 +727,8 @@ static bool gmc_v9_0_use_invalidate_semaphore(struct amdgpu_device *adev,
}
static bool gmc_v9_0_get_atc_vmid_pasid_mapping_info(struct amdgpu_device *adev,
- uint8_t vmid, uint16_t *p_pasid)
+ uint8_t vmid, uint8_t inst,
+ uint16_t *p_pasid)
{
uint32_t value;
@@ -889,7 +890,7 @@ static void gmc_v9_0_flush_gpu_tlb_pasid(struct amdgpu_device *adev,
bool valid;
valid = gmc_v9_0_get_atc_vmid_pasid_mapping_info(adev, vmid,
- &queried);
+ inst, &queried);
if (!valid || queried != pasid)
continue;
@@ -1302,6 +1303,7 @@ static const struct amdgpu_gmc_funcs gmc_v9_0_gmc_funcs = {
.flush_gpu_tlb_pasid = gmc_v9_0_flush_gpu_tlb_pasid,
.emit_flush_gpu_tlb = gmc_v9_0_emit_flush_gpu_tlb,
.emit_pasid_mapping = gmc_v9_0_emit_pasid_mapping,
+ .get_vmid_pasid_mapping_info = gmc_v9_0_get_atc_vmid_pasid_mapping_info,
.get_vm_pde = gmc_v9_0_get_vm_pde,
.get_vm_pte = gmc_v9_0_get_vm_pte,
.override_vm_pte_flags = gmc_v9_0_override_vm_pte_flags,
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 21/31] drm/amdgpu: add a gmc callback for the inv semaphore
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (19 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 20/31] drm/amdgpu/gmc: add new callback to lookup vmid to pasid mapping Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 22/31] drm/amdgpu/gmc: rework pasid flushing Alex Deucher
` (10 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
Used to determine whether or not to use the invalidation
semaphore for a particular hub.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 3 +++
drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c | 1 +
drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c | 1 +
drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c | 1 +
drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c | 1 +
drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c | 1 +
6 files changed, 8 insertions(+)
diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
index 27785b1b38ccd..f20f08630408e 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
@@ -175,6 +175,9 @@ struct amdgpu_gmc_funcs {
bool (*get_vmid_pasid_mapping_info)(struct amdgpu_device *adev,
uint8_t vmid, uint8_t inst,
uint16_t *p_pasid);
+ /* determine whether or not to use the inv semaphore */
+ bool (*use_invalidate_semaphore)(struct amdgpu_device *adev,
+ uint32_t vmhub);
/* enable/disable PRT support */
void (*set_prt)(struct amdgpu_device *adev, bool enable);
/* get the pde for a given mc addr */
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
index 0028639448956..4bb10a89ce582 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
@@ -557,6 +557,7 @@ static const struct amdgpu_gmc_funcs gmc_v10_0_gmc_funcs = {
.emit_flush_gpu_tlb = gmc_v10_0_emit_flush_gpu_tlb,
.emit_pasid_mapping = gmc_v10_0_emit_pasid_mapping,
.get_vmid_pasid_mapping_info = gmc_v10_0_get_atc_vmid_pasid_mapping_info,
+ .use_invalidate_semaphore = gmc_v10_0_use_invalidate_semaphore,
.get_vm_pde = gmc_v10_0_get_vm_pde,
.get_vm_pte = gmc_v10_0_get_vm_pte,
.get_vbios_fb_size = gmc_v10_0_get_vbios_fb_size,
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
index b34bd7881726d..acab9f4180da6 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
@@ -548,6 +548,7 @@ static const struct amdgpu_gmc_funcs gmc_v11_0_gmc_funcs = {
.emit_flush_gpu_tlb = gmc_v11_0_emit_flush_gpu_tlb,
.emit_pasid_mapping = gmc_v11_0_emit_pasid_mapping,
.get_vmid_pasid_mapping_info = gmc_v11_0_get_vmid_pasid_mapping_info,
+ .use_invalidate_semaphore = gmc_v11_0_use_invalidate_semaphore,
.get_vm_pde = gmc_v11_0_get_vm_pde,
.get_vm_pte = gmc_v11_0_get_vm_pte,
.get_vbios_fb_size = gmc_v11_0_get_vbios_fb_size,
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
index 9179dc0787f13..889c0005fec3d 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
@@ -583,6 +583,7 @@ static const struct amdgpu_gmc_funcs gmc_v12_0_gmc_funcs = {
.emit_flush_gpu_tlb = gmc_v12_0_emit_flush_gpu_tlb,
.emit_pasid_mapping = gmc_v12_0_emit_pasid_mapping,
.get_vmid_pasid_mapping_info = gmc_v12_0_get_vmid_pasid_mapping_info,
+ .use_invalidate_semaphore = gmc_v12_0_use_invalidate_semaphore,
.get_vm_pde = gmc_v12_0_get_vm_pde,
.get_vm_pte = gmc_v12_0_get_vm_pte,
.get_vbios_fb_size = gmc_v12_0_get_vbios_fb_size,
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c b/drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c
index 3fa1ec3dca273..d6be7c092a22d 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c
@@ -669,6 +669,7 @@ static const struct amdgpu_gmc_funcs gmc_v12_1_gmc_funcs = {
.emit_flush_gpu_tlb = gmc_v12_1_emit_flush_gpu_tlb,
.emit_pasid_mapping = gmc_v12_1_emit_pasid_mapping,
.get_vmid_pasid_mapping_info = gmc_v12_1_get_vmid_pasid_mapping_info,
+ .use_invalidate_semaphore = gmc_v12_1_use_invalidate_semaphore,
.get_vm_pde = gmc_v12_1_get_vm_pde,
.get_vm_pte = gmc_v12_1_get_vm_pte,
.query_mem_partition_mode = &amdgpu_gmc_query_memory_partition,
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
index 0019706b0b2a9..3c13a6920171a 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
@@ -1304,6 +1304,7 @@ static const struct amdgpu_gmc_funcs gmc_v9_0_gmc_funcs = {
.emit_flush_gpu_tlb = gmc_v9_0_emit_flush_gpu_tlb,
.emit_pasid_mapping = gmc_v9_0_emit_pasid_mapping,
.get_vmid_pasid_mapping_info = gmc_v9_0_get_atc_vmid_pasid_mapping_info,
+ .use_invalidate_semaphore = gmc_v9_0_use_invalidate_semaphore,
.get_vm_pde = gmc_v9_0_get_vm_pde,
.get_vm_pte = gmc_v9_0_get_vm_pte,
.override_vm_pte_flags = gmc_v9_0_override_vm_pte_flags,
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 22/31] drm/amdgpu/gmc: rework pasid flushing
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (20 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 21/31] drm/amdgpu: add a gmc callback for the inv semaphore Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-02 3:04 ` Zhang, Jesse(Jie)
2026-09-01 20:10 ` [PATCH 23/31] drm/amdgpu/gmc9: use SDMA for gart TLB invalidation Alex Deucher
` (9 subsequent siblings)
31 siblings, 1 reply; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
Split out all of the various flush methods and use
the new pasid flush method enum to determine which
one to use.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c | 247 +++++++++++++++++++-----
drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c | 2 +
drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c | 2 +
drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c | 2 +
4 files changed, 201 insertions(+), 52 deletions(-)
diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c
index 8a975eddd75c7..2800eebe50649 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c
@@ -788,9 +788,66 @@ void amdgpu_gmc_flush_gpu_tlb_gart(struct amdgpu_device *adev,
dev_err(adev->dev, "Error flushing GPU TLB using the SDMA (%d)!\n", r);
}
-int amdgpu_gmc_flush_gpu_tlb_pasid(struct amdgpu_device *adev, uint16_t pasid,
- uint32_t flush_type, bool all_hub,
- uint32_t inst)
+static int amdgpu_gmc_flush_gpu_tlb_sdma_helper(struct amdgpu_device *adev,
+ uint16_t pasid, uint32_t flush_type,
+ bool all_hub, uint32_t inst)
+{
+ struct amdgpu_ring *ring;
+ struct dma_fence *fence;
+ struct amdgpu_job *job;
+ uint16_t queried;
+ /* Use register 17 for GART */
+ u32 eng = 17;
+ int vmid, r, i, ndw;
+
+ /* flush hdp cache */
+ amdgpu_device_flush_hdp(adev, NULL);
+
+ ndw = ALIGN(adev->mman.buffer_funcs->tlb_inv_num_dw * 16 * AMDGPU_MAX_VMHUBS, 8);
+ ring = to_amdgpu_ring(adev->mman.buffer_funcs_scheds[0]);
+
+ mutex_lock(&adev->mman.default_entity.lock);
+ r = amdgpu_job_alloc_with_ib(ring->adev, &adev->mman.default_entity.base,
+ AMDGPU_FENCE_OWNER_UNDEFINED,
+ ndw * 4, AMDGPU_IB_POOL_IMMEDIATE,
+ AMDGPU_KERNEL_JOB_ID_VM_UPDATE,
+ &job);
+ if (r)
+ goto exit;
+
+ for (vmid = 1; vmid < 16; vmid++) {
+ bool valid;
+
+ valid = adev->gmc.gmc_funcs->get_vmid_pasid_mapping_info(adev, vmid, inst,
+ &queried);
+ if (!valid || queried != pasid)
+ continue;
+
+ if (all_hub) {
+ for_each_set_bit(i, adev->vmhubs_mask, AMDGPU_MAX_VMHUBS)
+ amdgpu_emit_tlb_inv(adev, &job->ibs[0], vmid, i, eng,
+ flush_type, inst);
+ } else {
+ amdgpu_emit_tlb_inv(adev, &job->ibs[0], vmid, AMDGPU_GFXHUB(inst), eng,
+ flush_type, inst);
+ }
+ }
+ amdgpu_ring_pad_ib(ring, &job->ibs[0]);
+ fence = amdgpu_job_submit(job);
+ mutex_unlock(&adev->mman.default_entity.lock);
+
+ dma_fence_wait(fence, false);
+ dma_fence_put(fence);
+
+exit:
+ mutex_unlock(&adev->mman.default_entity.lock);
+
+ return r;
+}
+
+static int amdgpu_gmc_flush_gpu_tlb_kiq_helper(struct amdgpu_device *adev,
+ uint16_t pasid, uint32_t flush_type,
+ bool all_hub, uint32_t inst)
{
struct amdgpu_ring *ring = &adev->gfx.kiq[inst].ring;
struct amdgpu_kiq *kiq = &adev->gfx.kiq[inst];
@@ -798,6 +855,110 @@ int amdgpu_gmc_flush_gpu_tlb_pasid(struct amdgpu_device *adev, uint16_t pasid,
int r, cnt = 0;
uint32_t seq;
+ /* 2 dwords flush + 8 dwords fence */
+ ndw = kiq->pmf->invalidate_tlbs_size + 8;
+
+ if (adev->gmc.flush_tlb_needs_extra_type_2)
+ ndw += kiq->pmf->invalidate_tlbs_size;
+
+ if (adev->gmc.flush_tlb_needs_extra_type_0)
+ ndw += kiq->pmf->invalidate_tlbs_size;
+
+ spin_lock(&adev->gfx.kiq[inst].ring_lock);
+ r = amdgpu_ring_alloc(ring, ndw);
+ if (r) {
+ spin_unlock(&adev->gfx.kiq[inst].ring_lock);
+ return r;
+ }
+ if (adev->gmc.flush_tlb_needs_extra_type_2)
+ kiq->pmf->kiq_invalidate_tlbs(ring, pasid, 2, all_hub);
+
+ if (flush_type == 2 && adev->gmc.flush_tlb_needs_extra_type_0)
+ kiq->pmf->kiq_invalidate_tlbs(ring, pasid, 0, all_hub);
+
+ kiq->pmf->kiq_invalidate_tlbs(ring, pasid, flush_type, all_hub);
+ r = amdgpu_fence_emit_polling(ring, &seq, MAX_KIQ_REG_WAIT);
+ if (r) {
+ amdgpu_ring_undo(ring);
+ spin_unlock(&adev->gfx.kiq[inst].ring_lock);
+ return r;
+ }
+
+ amdgpu_ring_commit(ring);
+ spin_unlock(&adev->gfx.kiq[inst].ring_lock);
+
+ r = amdgpu_fence_wait_polling(ring, seq, MAX_KIQ_REG_WAIT);
+
+ might_sleep();
+ while (r < 1 && cnt++ < MAX_KIQ_REG_TRY &&
+ !amdgpu_reset_pending(adev->reset_domain)) {
+ msleep(MAX_KIQ_REG_BAILOUT_INTERVAL);
+ r = amdgpu_fence_wait_polling(ring, seq, MAX_KIQ_REG_WAIT);
+ }
+
+ if (cnt > MAX_KIQ_REG_TRY) {
+ dev_err(adev->dev, "timeout waiting for kiq fence\n");
+ r = -ETIME;
+ } else
+ r = 0;
+
+ return r;
+}
+
+static int amdgpu_gmc_flush_gpu_tlb_mes_helper(struct amdgpu_device *adev,
+ uint16_t pasid, uint32_t flush_type,
+ bool all_hub, uint32_t inst)
+{
+ struct mes_inv_tlbs_pasid_input input = {0};
+ int r;
+
+ input.xcc_id = inst;
+ input.pasid = pasid;
+ input.flush_type = flush_type;
+
+ /* MES will invalidate hubs for the device(including slave xcc)
+ * from master, ignore request from slave
+ */
+ if (!amdgpu_gfx_is_master_xcc(adev, inst))
+ return -EINVAL;
+
+ input.hub_id = AMDGPU_GFXHUB(0);
+ amdgpu_mes_lock(&adev->mes);
+ r = adev->mes.funcs->invalidate_tlbs_pasid(&adev->mes, &input);
+ amdgpu_mes_unlock(&adev->mes);
+ if (r)
+ return r;
+
+ if (all_hub) {
+ /* invalidate mm_hub */
+ if (test_bit(AMDGPU_MMHUB0(0), adev->vmhubs_mask)) {
+ input.hub_id = AMDGPU_MMHUB0(0);
+ amdgpu_mes_lock(&adev->mes);
+ r = adev->mes.funcs->invalidate_tlbs_pasid(&adev->mes, &input);
+ amdgpu_mes_unlock(&adev->mes);
+ if (r)
+ return r;
+ }
+ if (test_bit(AMDGPU_MMHUB1(0), adev->vmhubs_mask)) {
+ input.hub_id = AMDGPU_MMHUB1(0);
+ amdgpu_mes_lock(&adev->mes);
+ r = adev->mes.funcs->invalidate_tlbs_pasid(&adev->mes, &input);
+ amdgpu_mes_unlock(&adev->mes);
+ if (r)
+ return r;
+ }
+ }
+ return 0;
+}
+
+int amdgpu_gmc_flush_gpu_tlb_pasid(struct amdgpu_device *adev, uint16_t pasid,
+ uint32_t flush_type, bool all_hub,
+ uint32_t inst)
+{
+ struct amdgpu_ring *ring;
+ bool use_mmio = false;
+ int r;
+
/*
* A GPU reset should flush all TLBs anyway, so no need to do
* this while one is ongoing.
@@ -805,8 +966,38 @@ int amdgpu_gmc_flush_gpu_tlb_pasid(struct amdgpu_device *adev, uint16_t pasid,
if (!down_read_trylock(&adev->reset_domain->sem))
return 0;
- if (!adev->gmc.flush_pasid_uses_kiq || !ring->sched.ready) {
+ switch (adev->gmc.pasid_inv_method) {
+ case AMDGPU_TLB_INV_METHOD_MMIO:
+ default:
+ use_mmio = true;
+ break;
+ case AMDGPU_TLB_INV_METHOD_SDMA:
+ ring = to_amdgpu_ring(adev->mman.buffer_funcs_scheds[0]);
+ if (!ring->sched.ready)
+ use_mmio = true;
+ else
+ r = amdgpu_gmc_flush_gpu_tlb_sdma_helper(adev, pasid, flush_type,
+ all_hub, inst);
+ break;
+ case AMDGPU_TLB_INV_METHOD_KIQ:
+ ring = &adev->gfx.kiq[inst].ring;
+ if (!adev->gmc.flush_pasid_uses_kiq || !ring->sched.ready)
+ use_mmio = true;
+ else
+ r = amdgpu_gmc_flush_gpu_tlb_kiq_helper(adev, pasid, flush_type,
+ all_hub, inst);
+ break;
+ case AMDGPU_TLB_INV_METHOD_MES:
+ ring = &adev->mes.ring[MES_PIPE_INST(inst, 0)];
+ if (!ring->sched.ready)
+ use_mmio = true;
+ else
+ r = amdgpu_gmc_flush_gpu_tlb_mes_helper(adev, pasid, flush_type,
+ all_hub, inst);
+ break;
+ }
+ if (use_mmio) {
if (!adev->gmc.gmc_funcs->flush_gpu_tlb_pasid) {
r = 0;
goto error_unlock_reset;
@@ -825,54 +1016,6 @@ int amdgpu_gmc_flush_gpu_tlb_pasid(struct amdgpu_device *adev, uint16_t pasid,
adev->gmc.gmc_funcs->flush_gpu_tlb_pasid(adev, pasid,
flush_type, all_hub,
inst);
- r = 0;
- } else {
- /* 2 dwords flush + 8 dwords fence */
- ndw = kiq->pmf->invalidate_tlbs_size + 8;
-
- if (adev->gmc.flush_tlb_needs_extra_type_2)
- ndw += kiq->pmf->invalidate_tlbs_size;
-
- if (adev->gmc.flush_tlb_needs_extra_type_0)
- ndw += kiq->pmf->invalidate_tlbs_size;
-
- spin_lock(&adev->gfx.kiq[inst].ring_lock);
- r = amdgpu_ring_alloc(ring, ndw);
- if (r) {
- spin_unlock(&adev->gfx.kiq[inst].ring_lock);
- goto error_unlock_reset;
- }
- if (adev->gmc.flush_tlb_needs_extra_type_2)
- kiq->pmf->kiq_invalidate_tlbs(ring, pasid, 2, all_hub);
-
- if (flush_type == 2 && adev->gmc.flush_tlb_needs_extra_type_0)
- kiq->pmf->kiq_invalidate_tlbs(ring, pasid, 0, all_hub);
-
- kiq->pmf->kiq_invalidate_tlbs(ring, pasid, flush_type, all_hub);
- r = amdgpu_fence_emit_polling(ring, &seq, MAX_KIQ_REG_WAIT);
- if (r) {
- amdgpu_ring_undo(ring);
- spin_unlock(&adev->gfx.kiq[inst].ring_lock);
- goto error_unlock_reset;
- }
-
- amdgpu_ring_commit(ring);
- spin_unlock(&adev->gfx.kiq[inst].ring_lock);
-
- r = amdgpu_fence_wait_polling(ring, seq, MAX_KIQ_REG_WAIT);
-
- might_sleep();
- while (r < 1 && cnt++ < MAX_KIQ_REG_TRY &&
- !amdgpu_reset_pending(adev->reset_domain)) {
- msleep(MAX_KIQ_REG_BAILOUT_INTERVAL);
- r = amdgpu_fence_wait_polling(ring, seq, MAX_KIQ_REG_WAIT);
- }
-
- if (cnt > MAX_KIQ_REG_TRY) {
- dev_err(adev->dev, "timeout waiting for kiq fence\n");
- r = -ETIME;
- } else
- r = 0;
}
error_unlock_reset:
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
index 4bb10a89ce582..e7c529620199b 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
@@ -974,6 +974,8 @@ static int gmc_v10_0_hw_init(struct amdgpu_ip_block *ip_block)
int r;
adev->gmc.flush_pasid_uses_kiq = !amdgpu_emu_mode;
+ if (adev->gmc.flush_pasid_uses_kiq)
+ adev->gmc.pasid_inv_method = AMDGPU_TLB_INV_METHOD_KIQ;
/* The sequence of these two function calls matters.*/
gmc_v10_0_init_golden_registers(adev);
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
index acab9f4180da6..a35c84cc7385f 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
@@ -950,6 +950,8 @@ static int gmc_v11_0_hw_init(struct amdgpu_ip_block *ip_block)
int r;
adev->gmc.flush_pasid_uses_kiq = !amdgpu_emu_mode;
+ if (adev->gmc.flush_pasid_uses_kiq)
+ adev->gmc.pasid_inv_method = AMDGPU_TLB_INV_METHOD_KIQ;
/* The sequence of these two function calls matters.*/
gmc_v11_0_init_golden_registers(adev);
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
index 3c13a6920171a..317b44412b5a6 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
@@ -2162,6 +2162,8 @@ static int gmc_v9_0_hw_init(struct amdgpu_ip_block *ip_block)
int i, r;
adev->gmc.flush_pasid_uses_kiq = true;
+ if (adev->gmc.flush_pasid_uses_kiq)
+ adev->gmc.pasid_inv_method = AMDGPU_TLB_INV_METHOD_KIQ;
/* Vega20+XGMI caches PTEs in TC and TLB. Add a heavy-weight TLB flush
* (type 2), which flushes both. Due to a race condition with
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 23/31] drm/amdgpu/gmc9: use SDMA for gart TLB invalidation
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (21 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 22/31] drm/amdgpu/gmc: rework pasid flushing Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 24/31] drm/amdgpu/gmc10: " Alex Deucher
` (8 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
Use SDMA rather than MMIO on vega20 and older. The avoids
the need to disallow gfxoff when invalidating.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c | 2 ++
1 file changed, 2 insertions(+)
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
index 317b44412b5a6..ad4bb86b8ffbf 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
@@ -2164,6 +2164,8 @@ static int gmc_v9_0_hw_init(struct amdgpu_ip_block *ip_block)
adev->gmc.flush_pasid_uses_kiq = true;
if (adev->gmc.flush_pasid_uses_kiq)
adev->gmc.pasid_inv_method = AMDGPU_TLB_INV_METHOD_KIQ;
+ if (amdgpu_ip_version(adev, GC_HWIP, 0) <= IP_VERSION(9, 4, 0))
+ adev->gmc.gart_inv_method = AMDGPU_TLB_INV_METHOD_SDMA;
/* Vega20+XGMI caches PTEs in TC and TLB. Add a heavy-weight TLB flush
* (type 2), which flushes both. Due to a race condition with
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 24/31] drm/amdgpu/gmc10: use SDMA for gart TLB invalidation
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (22 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 23/31] drm/amdgpu/gmc9: use SDMA for gart TLB invalidation Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 25/31] drm/amdgpu/gmc11: " Alex Deucher
` (7 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
Use SDMA rather than MMIO. The avoids the need to
disallow gfxoff when invalidating.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c | 1 +
1 file changed, 1 insertion(+)
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
index e7c529620199b..7eb34d55565cd 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
@@ -976,6 +976,7 @@ static int gmc_v10_0_hw_init(struct amdgpu_ip_block *ip_block)
adev->gmc.flush_pasid_uses_kiq = !amdgpu_emu_mode;
if (adev->gmc.flush_pasid_uses_kiq)
adev->gmc.pasid_inv_method = AMDGPU_TLB_INV_METHOD_KIQ;
+ adev->gmc.gart_inv_method = AMDGPU_TLB_INV_METHOD_SDMA;
/* The sequence of these two function calls matters.*/
gmc_v10_0_init_golden_registers(adev);
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 25/31] drm/amdgpu/gmc11: use SDMA for gart TLB invalidation
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (23 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 24/31] drm/amdgpu/gmc10: " Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 26/31] drm/amdgpu/gmc12: " Alex Deucher
` (6 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
Use SDMA rather than MMIO. The avoids the need to
disallow gfxoff when invalidating.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c | 1 +
1 file changed, 1 insertion(+)
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
index a35c84cc7385f..46e98fe66dd3e 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
@@ -952,6 +952,7 @@ static int gmc_v11_0_hw_init(struct amdgpu_ip_block *ip_block)
adev->gmc.flush_pasid_uses_kiq = !amdgpu_emu_mode;
if (adev->gmc.flush_pasid_uses_kiq)
adev->gmc.pasid_inv_method = AMDGPU_TLB_INV_METHOD_KIQ;
+ adev->gmc.gart_inv_method = AMDGPU_TLB_INV_METHOD_SDMA;
/* The sequence of these two function calls matters.*/
gmc_v11_0_init_golden_registers(adev);
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 26/31] drm/amdgpu/gmc12: use SDMA for gart TLB invalidation
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (24 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 25/31] drm/amdgpu/gmc11: " Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 27/31] drm/amdgpu/gmc10: use SDMA for pasid " Alex Deucher
` (5 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
Use SDMA rather than MMIO. The avoids the need to
disallow gfxoff when invalidating.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c | 2 ++
1 file changed, 2 insertions(+)
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
index 889c0005fec3d..6ec5798494be1 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
@@ -1061,6 +1061,8 @@ static int gmc_v12_0_hw_init(struct amdgpu_ip_block *ip_block)
int r;
struct amdgpu_device *adev = ip_block->adev;
+ adev->gmc.gart_inv_method = AMDGPU_TLB_INV_METHOD_SDMA;
+
/* The sequence of these two function calls matters.*/
gmc_v12_0_init_golden_registers(adev);
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 27/31] drm/amdgpu/gmc10: use SDMA for pasid TLB invalidation
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (25 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 26/31] drm/amdgpu/gmc12: " Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 28/31] drm/amdgpu/gmc11: " Alex Deucher
` (4 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
Use SDMA rather than MMIO. The avoids the need to
disallow gfxoff when invalidating.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c | 3 +--
1 file changed, 1 insertion(+), 2 deletions(-)
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
index 7eb34d55565cd..e716a86ba2913 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
@@ -974,8 +974,7 @@ static int gmc_v10_0_hw_init(struct amdgpu_ip_block *ip_block)
int r;
adev->gmc.flush_pasid_uses_kiq = !amdgpu_emu_mode;
- if (adev->gmc.flush_pasid_uses_kiq)
- adev->gmc.pasid_inv_method = AMDGPU_TLB_INV_METHOD_KIQ;
+ adev->gmc.pasid_inv_method = AMDGPU_TLB_INV_METHOD_SDMA;
adev->gmc.gart_inv_method = AMDGPU_TLB_INV_METHOD_SDMA;
/* The sequence of these two function calls matters.*/
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 28/31] drm/amdgpu/gmc11: use SDMA for pasid TLB invalidation
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (26 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 27/31] drm/amdgpu/gmc10: use SDMA for pasid " Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 29/31] drm/amdgpu/gmc12: use MES or " Alex Deucher
` (3 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
Use SDMA rather than MMIO. The avoids the need to
disallow gfxoff when invalidating.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c | 3 +--
1 file changed, 1 insertion(+), 2 deletions(-)
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
index 46e98fe66dd3e..8b95a1281886b 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
@@ -950,8 +950,7 @@ static int gmc_v11_0_hw_init(struct amdgpu_ip_block *ip_block)
int r;
adev->gmc.flush_pasid_uses_kiq = !amdgpu_emu_mode;
- if (adev->gmc.flush_pasid_uses_kiq)
- adev->gmc.pasid_inv_method = AMDGPU_TLB_INV_METHOD_KIQ;
+ adev->gmc.pasid_inv_method = AMDGPU_TLB_INV_METHOD_SDMA;
adev->gmc.gart_inv_method = AMDGPU_TLB_INV_METHOD_SDMA;
/* The sequence of these two function calls matters.*/
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 29/31] drm/amdgpu/gmc12: use MES or SDMA for pasid TLB invalidation
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (27 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 28/31] drm/amdgpu/gmc11: " Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 30/31] drm/amdgpu/gmc12: drop MES tlb inv in gmc callbacks Alex Deucher
` (2 subsequent siblings)
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
Use SDMA or MES rather than MMIO. The avoids the need to
disallow gfxoff when invalidating. Default to SDMA and
then use MES if the MES firmware supports it.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c | 1 +
drivers/gpu/drm/amd/amdgpu/mes_v12_0.c | 4 ++++
drivers/gpu/drm/amd/amdgpu/mes_v12_1.c | 4 ++++
3 files changed, 9 insertions(+)
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
index 6ec5798494be1..67eb1b595abe0 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
@@ -1061,6 +1061,7 @@ static int gmc_v12_0_hw_init(struct amdgpu_ip_block *ip_block)
int r;
struct amdgpu_device *adev = ip_block->adev;
+ adev->gmc.pasid_inv_method = AMDGPU_TLB_INV_METHOD_SDMA;
adev->gmc.gart_inv_method = AMDGPU_TLB_INV_METHOD_SDMA;
/* The sequence of these two function calls matters.*/
diff --git a/drivers/gpu/drm/amd/amdgpu/mes_v12_0.c b/drivers/gpu/drm/amd/amdgpu/mes_v12_0.c
index ee72afdb0a1d3..dd68db14c43f6 100644
--- a/drivers/gpu/drm/amd/amdgpu/mes_v12_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/mes_v12_0.c
@@ -2097,6 +2097,10 @@ static int mes_v12_0_hw_init(struct amdgpu_ip_block *ip_block)
adev->gfx.kiq[0].ring.sched.ready = false;
adev->mes.ring[0].sched.ready = true;
+ if (adev->enable_uni_mes &&
+ (adev->mes.sched_version & AMDGPU_MES_VERSION_MASK) >= 0x84)
+ adev->gmc.pasid_inv_method = AMDGPU_TLB_INV_METHOD_MES;
+
return 0;
failure:
diff --git a/drivers/gpu/drm/amd/amdgpu/mes_v12_1.c b/drivers/gpu/drm/amd/amdgpu/mes_v12_1.c
index f1098118d5c05..5eac7615efff5 100644
--- a/drivers/gpu/drm/amd/amdgpu/mes_v12_1.c
+++ b/drivers/gpu/drm/amd/amdgpu/mes_v12_1.c
@@ -1962,6 +1962,10 @@ static int mes_v12_1_hw_init(struct amdgpu_ip_block *ip_block)
return r;
}
+ if (adev->enable_uni_mes &&
+ (adev->mes.sched_version & AMDGPU_MES_VERSION_MASK) >= 0x6f)
+ adev->gmc.pasid_inv_method = AMDGPU_TLB_INV_METHOD_MES;
+
return 0;
}
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 30/31] drm/amdgpu/gmc12: drop MES tlb inv in gmc callbacks
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (28 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 29/31] drm/amdgpu/gmc12: use MES or " Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 31/31] drm/amdgpu/gmc: add helpers for various tlb inv functions Alex Deucher
2026-09-02 7:07 ` [PATCH V2 00/31] Rework GPU TLB invalidation Christian König
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
It's now handled directly in the front end functions.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c | 16 ---------------
drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c | 28 --------------------------
2 files changed, 44 deletions(-)
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
index 67eb1b595abe0..4bb8c1e4f335b 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
@@ -356,22 +356,6 @@ static void gmc_v12_0_flush_gpu_tlb_pasid(struct amdgpu_device *adev,
uint16_t queried;
int vmid, i;
- if (adev->enable_uni_mes && adev->mes.ring[AMDGPU_MES_SCHED_PIPE].sched.ready &&
- (adev->mes.sched_version & AMDGPU_MES_VERSION_MASK) >= 0x84) {
- struct mes_inv_tlbs_pasid_input input = {0};
- input.pasid = pasid;
- input.flush_type = flush_type;
- input.hub_id = AMDGPU_GFXHUB(0);
- /* MES will invalidate all gc_hub for the device from master */
- adev->mes.funcs->invalidate_tlbs_pasid(&adev->mes, &input);
- if (all_hub) {
- /* Only need to invalidate mm_hub now, gfx12 only support one mmhub */
- input.hub_id = AMDGPU_MMHUB0(0);
- adev->mes.funcs->invalidate_tlbs_pasid(&adev->mes, &input);
- }
- return;
- }
-
for (vmid = 1; vmid < 16; vmid++) {
bool valid;
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c b/drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c
index d6be7c092a22d..2c980dc02c262 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c
@@ -411,34 +411,6 @@ static void gmc_v12_1_flush_gpu_tlb_pasid(struct amdgpu_device *adev,
uint16_t queried;
int vmid, i;
- if (adev->enable_uni_mes && adev->mes.ring[0].sched.ready &&
- (adev->mes.sched_version & AMDGPU_MES_VERSION_MASK) >= 0x6f) {
- struct mes_inv_tlbs_pasid_input input = {0};
- input.xcc_id = inst;
- input.pasid = pasid;
- input.flush_type = flush_type;
-
- /* MES will invalidate hubs for the device(including slave xcc) from master, ignore request from slave */
- if (!amdgpu_gfx_is_master_xcc(adev, inst))
- return;
-
- input.hub_id = AMDGPU_GFXHUB(0);
- adev->mes.funcs->invalidate_tlbs_pasid(&adev->mes, &input);
-
- if (all_hub) {
- /* invalidate mm_hub */
- if (test_bit(AMDGPU_MMHUB0(0), adev->vmhubs_mask)) {
- input.hub_id = AMDGPU_MMHUB0(0);
- adev->mes.funcs->invalidate_tlbs_pasid(&adev->mes, &input);
- }
- if (test_bit(AMDGPU_MMHUB1(0), adev->vmhubs_mask)) {
- input.hub_id = AMDGPU_MMHUB1(0);
- adev->mes.funcs->invalidate_tlbs_pasid(&adev->mes, &input);
- }
- }
- return;
- }
-
for (vmid = 1; vmid < 16; vmid++) {
bool valid;
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* [PATCH 31/31] drm/amdgpu/gmc: add helpers for various tlb inv functions
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (29 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 30/31] drm/amdgpu/gmc12: drop MES tlb inv in gmc callbacks Alex Deucher
@ 2026-09-01 20:10 ` Alex Deucher
2026-09-02 7:07 ` [PATCH V2 00/31] Rework GPU TLB invalidation Christian König
31 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-01 20:10 UTC (permalink / raw)
To: amd-gfx, christian.koenig; +Cc: Alex Deucher
gmc9-12 use the same logic for almost all of these, so move it to
helpers and remove the gmc specific functions.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c | 195 +++++++++++++++++++++
drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 7 +
drivers/gpu/drm/amd/amdgpu/gfx_v10_0.c | 2 +
drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c | 199 +--------------------
drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c | 204 +---------------------
drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c | 219 +-----------------------
drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c | 195 +--------------------
drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c | 92 +---------
8 files changed, 221 insertions(+), 892 deletions(-)
diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c
index 2800eebe50649..185290a392bec 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c
@@ -951,6 +951,201 @@ static int amdgpu_gmc_flush_gpu_tlb_mes_helper(struct amdgpu_device *adev,
return 0;
}
+/**
+ * amdgpu_gmc_flush_gpu_tlb_pasid_helper - tlb flush via pasid
+ *
+ * @adev: amdgpu_device pointer
+ * @pasid: pasid to be flush
+ * @flush_type: the flush type
+ * @all_hub: flush all hubs
+ * @inst: is used to select which instance of KIQ to use for the invalidation
+ *
+ * A helper to flush the TLB for the requested pasid using other callbacks.
+ */
+void amdgpu_gmc_flush_gpu_tlb_pasid_helper(struct amdgpu_device *adev,
+ uint16_t pasid, uint32_t flush_type,
+ bool all_hub, uint32_t inst)
+{
+ uint16_t queried;
+ int vmid, i;
+
+ for (vmid = 1; vmid < 16; vmid++) {
+ bool valid;
+
+ valid = adev->gmc.gmc_funcs->get_vmid_pasid_mapping_info(adev, vmid, inst,
+ &queried);
+ if (!valid || queried != pasid)
+ continue;
+
+ if (all_hub) {
+ for_each_set_bit(i, adev->vmhubs_mask, AMDGPU_MAX_VMHUBS)
+ adev->gmc.gmc_funcs->flush_gpu_tlb(adev, vmid, i,
+ flush_type);
+ } else {
+ adev->gmc.gmc_funcs->flush_gpu_tlb(adev, vmid, AMDGPU_GFXHUB(inst),
+ flush_type);
+ }
+ }
+}
+
+/**
+ * amdgpu_gmc_flush_gpu_tlb_helper - gart tlb flush callback
+ *
+ * @adev: amdgpu_device pointer
+ * @vmid: vm instance to flush
+ * @vmhub: which hub to flush
+ * @flush_type: the flush type
+ *
+ * Flush the TLB for the requested page table.
+ */
+void amdgpu_gmc_flush_gpu_tlb_helper(struct amdgpu_device *adev, uint32_t vmid,
+ uint32_t vmhub, uint32_t flush_type)
+{
+ bool use_semaphore = adev->gmc.gmc_funcs->use_invalidate_semaphore(adev, vmhub);
+ struct amdgpu_vmhub *hub = &adev->vmhub[vmhub];
+ u32 inv_req = hub->vmhub_funcs->get_invalidate_req(vmid, flush_type);
+ /* Use register 17 for GART */
+ const unsigned int eng = 17;
+ unsigned char hub_ip;
+ u32 sem, req, ack;
+ unsigned int i;
+ u32 tmp, inst;
+
+ if (AMDGPU_IS_GFXHUB(vmhub) && !adev->gfx.is_poweron)
+ return;
+
+ sem = hub->vm_inv_eng0_sem + hub->eng_distance * eng;
+ req = hub->vm_inv_eng0_req + hub->eng_distance * eng;
+ ack = hub->vm_inv_eng0_ack + hub->eng_distance * eng;
+
+
+ if (vmhub >= AMDGPU_MMHUB0(0))
+ inst = 0;
+ else
+ inst = vmhub;
+
+ /* flush hdp cache */
+ amdgpu_device_flush_hdp(adev, NULL);
+
+ /* This is necessary for SRIOV as well as for GFXOFF to function
+ * properly under bare metal
+ */
+ if ((adev->gfx.kiq[inst].ring.sched.ready ||
+ adev->mes.ring[MES_PIPE_INST(inst, 0)].sched.ready) &&
+ !adev->gmc.use_mmio_for_tlb_flush) {
+ amdgpu_gmc_fw_reg_write_reg_wait(adev, req, ack, inv_req,
+ 1 << vmid, inst);
+ return;
+ }
+
+ /* This path is needed before KIQ/MES/GFXOFF are set up */
+ hub_ip = AMDGPU_IS_GFXHUB(vmhub) ? GC_HWIP : MMHUB_HWIP;
+
+ /* disabllow gfxoff when we invalidate */
+ if (hub_ip == GC_HWIP)
+ amdgpu_gfx_off_ctrl(adev, false);
+
+ spin_lock(&adev->gmc.invalidate_lock);
+ /*
+ * It may lose gpuvm invalidate acknowldege state across power-gating
+ * off cycle, add semaphore acquire before invalidation and semaphore
+ * release after invalidation to avoid entering power gated state
+ * to WA the Issue
+ */
+
+ /* TODO: It needs to continue working on debugging with semaphore for GFXHUB as well. */
+ if (use_semaphore) {
+ for (i = 0; i < adev->usec_timeout; i++) {
+ /* a read return value of 1 means semaphore acuqire */
+ tmp = RREG32_RLC_NO_KIQ(sem, hub_ip);
+ if (tmp & 0x1)
+ break;
+ udelay(1);
+ }
+
+ if (i >= adev->usec_timeout)
+ DRM_ERROR("Timeout waiting for sem acquire in VM flush!\n");
+ }
+
+ WREG32_RLC_NO_KIQ(req, inv_req, hub_ip);
+
+ /* Wait for ACK with a delay.*/
+ for (i = 0; i < adev->usec_timeout; i++) {
+ tmp = RREG32_RLC_NO_KIQ(ack, hub_ip);
+ tmp &= 1 << vmid;
+ if (tmp)
+ break;
+
+ udelay(1);
+ }
+
+ /* TODO: It needs to continue working on debugging with semaphore for GFXHUB as well. */
+ if (use_semaphore)
+ WREG32_RLC_NO_KIQ(sem, 0, hub_ip);
+
+ /* Issue additional private vm invalidation to MMHUB */
+ if ((vmhub != AMDGPU_GFXHUB(0)) &&
+ (hub->vm_l2_bank_select_reserved_cid2) &&
+ !amdgpu_sriov_vf(adev)) {
+ inv_req = RREG32_NO_KIQ(hub->vm_l2_bank_select_reserved_cid2);
+ /* bit 25: RSERVED_CACHE_PRIVATE_INVALIDATION */
+ inv_req |= (1 << 25);
+ /* Issue private invalidation */
+ WREG32_NO_KIQ(hub->vm_l2_bank_select_reserved_cid2, inv_req);
+ /* Read back to ensure invalidation is done*/
+ RREG32_NO_KIQ(hub->vm_l2_bank_select_reserved_cid2);
+ }
+
+ spin_unlock(&adev->gmc.invalidate_lock);
+
+ if (hub_ip == GC_HWIP)
+ amdgpu_gfx_off_ctrl(adev, true);
+
+ if (i >= adev->usec_timeout)
+ dev_err(adev->dev, "Timeout waiting for VM flush ACK!\n");
+}
+
+uint64_t amdgpu_gmc_emit_flush_gpu_tlb_helper(struct amdgpu_ring *ring,
+ unsigned vmid, uint64_t pd_addr)
+{
+ bool use_semaphore =
+ ring->adev->gmc.gmc_funcs->use_invalidate_semaphore(ring->adev,
+ ring->vm_hub);
+ struct amdgpu_vmhub *hub = &ring->adev->vmhub[ring->vm_hub];
+ uint32_t req = hub->vmhub_funcs->get_invalidate_req(vmid, 0);
+ unsigned eng = ring->vm_inv_eng;
+
+ if (use_semaphore)
+ /* a read return value of 1 means semaphore acuqire */
+ amdgpu_ring_emit_reg_wait(ring,
+ hub->vm_inv_eng0_sem +
+ hub->eng_distance * eng, 0x1, 0x1);
+
+ amdgpu_ring_emit_wreg(ring, hub->ctx0_ptb_addr_lo32 +
+ (hub->ctx_addr_distance * vmid),
+ lower_32_bits(pd_addr));
+
+ amdgpu_ring_emit_wreg(ring, hub->ctx0_ptb_addr_hi32 +
+ (hub->ctx_addr_distance * vmid),
+ upper_32_bits(pd_addr));
+
+ amdgpu_ring_emit_reg_write_reg_wait(ring, hub->vm_inv_eng0_req +
+ hub->eng_distance * eng,
+ hub->vm_inv_eng0_ack +
+ hub->eng_distance * eng,
+ req, 1 << vmid);
+
+ if (use_semaphore)
+ /*
+ * add semaphore release after invalidation,
+ * write with 0 means semaphore release
+ */
+ amdgpu_ring_emit_wreg(ring, hub->vm_inv_eng0_sem +
+ hub->eng_distance * eng, 0);
+
+ return pd_addr;
+}
+
int amdgpu_gmc_flush_gpu_tlb_pasid(struct amdgpu_device *adev, uint16_t pasid,
uint32_t flush_type, bool all_hub,
uint32_t inst)
diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
index f20f08630408e..22814dd451430 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
@@ -459,6 +459,13 @@ int amdgpu_gmc_handle_retry_fault(struct amdgpu_device *adev,
bool write_fault);
int amdgpu_gmc_ras_sw_init(struct amdgpu_device *adev);
int amdgpu_gmc_allocate_vm_inv_eng(struct amdgpu_device *adev);
+void amdgpu_gmc_flush_gpu_tlb_pasid_helper(struct amdgpu_device *adev,
+ uint16_t pasid, uint32_t flush_type,
+ bool all_hub, uint32_t inst);
+void amdgpu_gmc_flush_gpu_tlb_helper(struct amdgpu_device *adev, uint32_t vmid,
+ uint32_t vmhub, uint32_t flush_type);
+uint64_t amdgpu_gmc_emit_flush_gpu_tlb_helper(struct amdgpu_ring *ring,
+ unsigned vmid, uint64_t pd_addr);
void amdgpu_gmc_flush_gpu_tlb_gart(struct amdgpu_device *adev,
uint32_t vmhub);
int amdgpu_gmc_flush_gpu_tlb_pasid(struct amdgpu_device *adev, uint16_t pasid,
diff --git a/drivers/gpu/drm/amd/amdgpu/gfx_v10_0.c b/drivers/gpu/drm/amd/amdgpu/gfx_v10_0.c
index 5033f85d31022..654bf7f780d59 100644
--- a/drivers/gpu/drm/amd/amdgpu/gfx_v10_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gfx_v10_0.c
@@ -7497,6 +7497,8 @@ static int gfx_v10_0_hw_init(struct amdgpu_ip_block *ip_block)
int r;
struct amdgpu_device *adev = ip_block->adev;
+ adev->gfx.is_poweron = true;
+
if (!amdgpu_emu_mode)
gfx_v10_0_init_golden_registers(adev);
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
index e716a86ba2913..35397b44c9dfc 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
@@ -223,195 +223,6 @@ static bool gmc_v10_0_get_atc_vmid_pasid_mapping_info(
* by the amdgpu vm/hsa code.
*/
-/**
- * gmc_v10_0_flush_gpu_tlb - gart tlb flush callback
- *
- * @adev: amdgpu_device pointer
- * @vmid: vm instance to flush
- * @vmhub: vmhub type
- * @flush_type: the flush type
- *
- * Flush the TLB for the requested page table.
- */
-static void gmc_v10_0_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
- uint32_t vmhub, uint32_t flush_type)
-{
- bool use_semaphore = gmc_v10_0_use_invalidate_semaphore(adev, vmhub);
- struct amdgpu_vmhub *hub = &adev->vmhub[vmhub];
- u32 inv_req = hub->vmhub_funcs->get_invalidate_req(vmid, flush_type);
- /* Use register 17 for GART */
- const unsigned int eng = 17;
- unsigned char hub_ip = 0;
- u32 sem, req, ack;
- unsigned int i;
- u32 tmp;
-
- sem = hub->vm_inv_eng0_sem + hub->eng_distance * eng;
- req = hub->vm_inv_eng0_req + hub->eng_distance * eng;
- ack = hub->vm_inv_eng0_ack + hub->eng_distance * eng;
-
- /* flush hdp cache */
- amdgpu_device_flush_hdp(adev, NULL);
-
- /* This is necessary for SRIOV as well as for GFXOFF to function
- * properly under bare metal
- */
- if (adev->gfx.kiq[0].ring.sched.ready && !adev->enable_mes &&
- !adev->gmc.use_mmio_for_tlb_flush) {
- amdgpu_gmc_fw_reg_write_reg_wait(adev, req, ack, inv_req,
- 1 << vmid, GET_INST(GC, 0));
- return;
- }
-
- /* This path is needed before KIQ/MES/GFXOFF are set up */
- hub_ip = (vmhub == AMDGPU_GFXHUB(0)) ? GC_HWIP : MMHUB_HWIP;
-
- /* disabllow gfxoff when we invalidate */
- if (hub_ip == GC_HWIP)
- amdgpu_gfx_off_ctrl(adev, false);
-
- spin_lock(&adev->gmc.invalidate_lock);
- /*
- * It may lose gpuvm invalidate acknowldege state across power-gating
- * off cycle, add semaphore acquire before invalidation and semaphore
- * release after invalidation to avoid entering power gated state
- * to WA the Issue
- */
-
- /* TODO: It needs to continue working on debugging with semaphore for GFXHUB as well. */
- if (use_semaphore) {
- for (i = 0; i < adev->usec_timeout; i++) {
- /* a read return value of 1 means semaphore acuqire */
- tmp = RREG32_RLC_NO_KIQ(sem, hub_ip);
- if (tmp & 0x1)
- break;
- udelay(1);
- }
-
- if (i >= adev->usec_timeout)
- DRM_ERROR("Timeout waiting for sem acquire in VM flush!\n");
- }
-
- WREG32_RLC_NO_KIQ(req, inv_req, hub_ip);
-
- /*
- * Issue a dummy read to wait for the ACK register to be cleared
- * to avoid a false ACK due to the new fast GRBM interface.
- */
- if ((vmhub == AMDGPU_GFXHUB(0)) &&
- (amdgpu_ip_version(adev, GC_HWIP, 0) < IP_VERSION(10, 3, 0)))
- RREG32_RLC_NO_KIQ(req, hub_ip);
-
- /* Wait for ACK with a delay.*/
- for (i = 0; i < adev->usec_timeout; i++) {
- tmp = RREG32_RLC_NO_KIQ(ack, hub_ip);
- tmp &= 1 << vmid;
- if (tmp)
- break;
-
- udelay(1);
- }
-
- /* TODO: It needs to continue working on debugging with semaphore for GFXHUB as well. */
- if (use_semaphore)
- WREG32_RLC_NO_KIQ(sem, 0, hub_ip);
-
- spin_unlock(&adev->gmc.invalidate_lock);
-
- if (hub_ip == GC_HWIP)
- amdgpu_gfx_off_ctrl(adev, true);
-
- if (i >= adev->usec_timeout)
- dev_err(adev->dev, "Timeout waiting for VM flush hub: %d!\n",
- vmhub);
-}
-
-/**
- * gmc_v10_0_flush_gpu_tlb_pasid - tlb flush via pasid
- *
- * @adev: amdgpu_device pointer
- * @pasid: pasid to be flush
- * @flush_type: the flush type
- * @all_hub: Used with PACKET3_INVALIDATE_TLBS_ALL_HUB()
- * @inst: is used to select which instance of KIQ to use for the invalidation
- *
- * Flush the TLB for the requested pasid.
- */
-static void gmc_v10_0_flush_gpu_tlb_pasid(struct amdgpu_device *adev,
- uint16_t pasid, uint32_t flush_type,
- bool all_hub, uint32_t inst)
-{
- uint16_t queried;
- int vmid, i;
-
- for (vmid = 1; vmid < AMDGPU_NUM_VMID; vmid++) {
- bool valid;
-
- valid = gmc_v10_0_get_atc_vmid_pasid_mapping_info(adev, vmid, 0,
- &queried);
- if (!valid || queried != pasid)
- continue;
-
- if (all_hub) {
- for_each_set_bit(i, adev->vmhubs_mask,
- AMDGPU_MAX_VMHUBS)
- gmc_v10_0_flush_gpu_tlb(adev, vmid, i,
- flush_type);
- } else {
- gmc_v10_0_flush_gpu_tlb(adev, vmid, AMDGPU_GFXHUB(0),
- flush_type);
- }
- }
-}
-
-static uint64_t gmc_v10_0_emit_flush_gpu_tlb(struct amdgpu_ring *ring,
- unsigned int vmid, uint64_t pd_addr)
-{
- bool use_semaphore = gmc_v10_0_use_invalidate_semaphore(ring->adev, ring->vm_hub);
- struct amdgpu_vmhub *hub = &ring->adev->vmhub[ring->vm_hub];
- uint32_t req = hub->vmhub_funcs->get_invalidate_req(vmid, 0);
- unsigned int eng = ring->vm_inv_eng;
-
- /*
- * It may lose gpuvm invalidate acknowldege state across power-gating
- * off cycle, add semaphore acquire before invalidation and semaphore
- * release after invalidation to avoid entering power gated state
- * to WA the Issue
- */
-
- /* TODO: It needs to continue working on debugging with semaphore for GFXHUB as well. */
- if (use_semaphore)
- /* a read return value of 1 means semaphore acuqire */
- amdgpu_ring_emit_reg_wait(ring,
- hub->vm_inv_eng0_sem +
- hub->eng_distance * eng, 0x1, 0x1);
-
- amdgpu_ring_emit_wreg(ring, hub->ctx0_ptb_addr_lo32 +
- (hub->ctx_addr_distance * vmid),
- lower_32_bits(pd_addr));
-
- amdgpu_ring_emit_wreg(ring, hub->ctx0_ptb_addr_hi32 +
- (hub->ctx_addr_distance * vmid),
- upper_32_bits(pd_addr));
-
- amdgpu_ring_emit_reg_write_reg_wait(ring, hub->vm_inv_eng0_req +
- hub->eng_distance * eng,
- hub->vm_inv_eng0_ack +
- hub->eng_distance * eng,
- req, 1 << vmid);
-
- /* TODO: It needs to continue working on debugging with semaphore for GFXHUB as well. */
- if (use_semaphore)
- /*
- * add semaphore release after invalidation,
- * write with 0 means semaphore release
- */
- amdgpu_ring_emit_wreg(ring, hub->vm_inv_eng0_sem +
- hub->eng_distance * eng, 0);
-
- return pd_addr;
-}
-
static void gmc_v10_0_emit_pasid_mapping(struct amdgpu_ring *ring, unsigned int vmid,
unsigned int pasid)
{
@@ -552,9 +363,9 @@ static unsigned int gmc_v10_0_get_vbios_fb_size(struct amdgpu_device *adev)
}
static const struct amdgpu_gmc_funcs gmc_v10_0_gmc_funcs = {
- .flush_gpu_tlb = gmc_v10_0_flush_gpu_tlb,
- .flush_gpu_tlb_pasid = gmc_v10_0_flush_gpu_tlb_pasid,
- .emit_flush_gpu_tlb = gmc_v10_0_emit_flush_gpu_tlb,
+ .flush_gpu_tlb = amdgpu_gmc_flush_gpu_tlb_helper,
+ .flush_gpu_tlb_pasid = amdgpu_gmc_flush_gpu_tlb_pasid_helper,
+ .emit_flush_gpu_tlb = amdgpu_gmc_emit_flush_gpu_tlb_helper,
.emit_pasid_mapping = gmc_v10_0_emit_pasid_mapping,
.get_vmid_pasid_mapping_info = gmc_v10_0_get_atc_vmid_pasid_mapping_info,
.use_invalidate_semaphore = gmc_v10_0_use_invalidate_semaphore,
@@ -957,9 +768,9 @@ static int gmc_v10_0_gart_enable(struct amdgpu_device *adev)
if (!adev->in_s0ix)
adev->gfxhub.funcs->set_fault_enable_default(adev, value);
adev->mmhub.funcs->set_fault_enable_default(adev, value);
- gmc_v10_0_flush_gpu_tlb(adev, 0, AMDGPU_MMHUB0(0), 0);
+ adev->gmc.gmc_funcs->flush_gpu_tlb(adev, 0, AMDGPU_MMHUB0(0), 0);
if (!adev->in_s0ix)
- gmc_v10_0_flush_gpu_tlb(adev, 0, AMDGPU_GFXHUB(0), 0);
+ adev->gmc.gmc_funcs->flush_gpu_tlb(adev, 0, AMDGPU_GFXHUB(0), 0);
drm_info(adev_to_drm(adev), "PCIE GART of %uM enabled (table at 0x%016llX).\n",
(unsigned int)(adev->gmc.gart_size >> 20),
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
index 8b95a1281886b..b05dda4b0b690 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
@@ -208,202 +208,6 @@ static bool gmc_v11_0_get_vmid_pasid_mapping_info(
return !!(*p_pasid);
}
-/**
- * gmc_v11_0_flush_gpu_tlb - gart tlb flush callback
- *
- * @adev: amdgpu_device pointer
- * @vmid: vm instance to flush
- * @vmhub: which hub to flush
- * @flush_type: the flush type
- *
- * Flush the TLB for the requested page table.
- */
-static void gmc_v11_0_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
- uint32_t vmhub, uint32_t flush_type)
-{
- bool use_semaphore = gmc_v11_0_use_invalidate_semaphore(adev, vmhub);
- struct amdgpu_vmhub *hub = &adev->vmhub[vmhub];
- u32 inv_req = hub->vmhub_funcs->get_invalidate_req(vmid, flush_type);
- /* Use register 17 for GART */
- const unsigned int eng = 17;
- unsigned char hub_ip;
- u32 sem, req, ack;
- unsigned int i;
- u32 tmp;
-
- if ((vmhub == AMDGPU_GFXHUB(0)) && !adev->gfx.is_poweron)
- return;
-
- sem = hub->vm_inv_eng0_sem + hub->eng_distance * eng;
- req = hub->vm_inv_eng0_req + hub->eng_distance * eng;
- ack = hub->vm_inv_eng0_ack + hub->eng_distance * eng;
-
- /* flush hdp cache */
- amdgpu_device_flush_hdp(adev, NULL);
-
- /* This is necessary for SRIOV as well as for GFXOFF to function
- * properly under bare metal
- */
- if ((adev->gfx.kiq[0].ring.sched.ready || adev->mes.ring[0].sched.ready) &&
- !adev->gmc.use_mmio_for_tlb_flush) {
- amdgpu_gmc_fw_reg_write_reg_wait(adev, req, ack, inv_req,
- 1 << vmid, GET_INST(GC, 0));
- return;
- }
-
- /* This path is needed before KIQ/MES/GFXOFF are set up */
- hub_ip = (vmhub == AMDGPU_GFXHUB(0)) ? GC_HWIP : MMHUB_HWIP;
-
- /* disabllow gfxoff when we invalidate */
- if (hub_ip == GC_HWIP)
- amdgpu_gfx_off_ctrl(adev, false);
-
- spin_lock(&adev->gmc.invalidate_lock);
- /*
- * It may lose gpuvm invalidate acknowldege state across power-gating
- * off cycle, add semaphore acquire before invalidation and semaphore
- * release after invalidation to avoid entering power gated state
- * to WA the Issue
- */
-
- /* TODO: It needs to continue working on debugging with semaphore for GFXHUB as well. */
- if (use_semaphore) {
- for (i = 0; i < adev->usec_timeout; i++) {
- /* a read return value of 1 means semaphore acuqire */
- tmp = RREG32_RLC_NO_KIQ(sem, hub_ip);
- if (tmp & 0x1)
- break;
- udelay(1);
- }
-
- if (i >= adev->usec_timeout)
- DRM_ERROR("Timeout waiting for sem acquire in VM flush!\n");
- }
-
- WREG32_RLC_NO_KIQ(req, inv_req, hub_ip);
-
- /* Wait for ACK with a delay.*/
- for (i = 0; i < adev->usec_timeout; i++) {
- tmp = RREG32_RLC_NO_KIQ(ack, hub_ip);
- tmp &= 1 << vmid;
- if (tmp)
- break;
-
- udelay(1);
- }
-
- /* TODO: It needs to continue working on debugging with semaphore for GFXHUB as well. */
- if (use_semaphore)
- WREG32_RLC_NO_KIQ(sem, 0, hub_ip);
-
- /* Issue additional private vm invalidation to MMHUB */
- if ((vmhub != AMDGPU_GFXHUB(0)) &&
- (hub->vm_l2_bank_select_reserved_cid2) &&
- !amdgpu_sriov_vf(adev)) {
- inv_req = RREG32_NO_KIQ(hub->vm_l2_bank_select_reserved_cid2);
- /* bit 25: RSERVED_CACHE_PRIVATE_INVALIDATION */
- inv_req |= (1 << 25);
- /* Issue private invalidation */
- WREG32_NO_KIQ(hub->vm_l2_bank_select_reserved_cid2, inv_req);
- /* Read back to ensure invalidation is done*/
- RREG32_NO_KIQ(hub->vm_l2_bank_select_reserved_cid2);
- }
-
- spin_unlock(&adev->gmc.invalidate_lock);
-
- if (hub_ip == GC_HWIP)
- amdgpu_gfx_off_ctrl(adev, true);
-
- if (i >= adev->usec_timeout)
- dev_err(adev->dev, "Timeout waiting for VM flush ACK!\n");
-}
-
-/**
- * gmc_v11_0_flush_gpu_tlb_pasid - tlb flush via pasid
- *
- * @adev: amdgpu_device pointer
- * @pasid: pasid to be flush
- * @flush_type: the flush type
- * @all_hub: flush all hubs
- * @inst: is used to select which instance of KIQ to use for the invalidation
- *
- * Flush the TLB for the requested pasid.
- */
-static void gmc_v11_0_flush_gpu_tlb_pasid(struct amdgpu_device *adev,
- uint16_t pasid, uint32_t flush_type,
- bool all_hub, uint32_t inst)
-{
- uint16_t queried;
- int vmid, i;
-
- for (vmid = 1; vmid < 16; vmid++) {
- bool valid;
-
- valid = gmc_v11_0_get_vmid_pasid_mapping_info(adev, vmid, 0,
- &queried);
- if (!valid || queried != pasid)
- continue;
-
- if (all_hub) {
- for_each_set_bit(i, adev->vmhubs_mask,
- AMDGPU_MAX_VMHUBS)
- gmc_v11_0_flush_gpu_tlb(adev, vmid, i,
- flush_type);
- } else {
- gmc_v11_0_flush_gpu_tlb(adev, vmid, AMDGPU_GFXHUB(0),
- flush_type);
- }
- }
-}
-
-static uint64_t gmc_v11_0_emit_flush_gpu_tlb(struct amdgpu_ring *ring,
- unsigned int vmid, uint64_t pd_addr)
-{
- bool use_semaphore = gmc_v11_0_use_invalidate_semaphore(ring->adev, ring->vm_hub);
- struct amdgpu_vmhub *hub = &ring->adev->vmhub[ring->vm_hub];
- uint32_t req = hub->vmhub_funcs->get_invalidate_req(vmid, 0);
- unsigned int eng = ring->vm_inv_eng;
-
- /*
- * It may lose gpuvm invalidate acknowldege state across power-gating
- * off cycle, add semaphore acquire before invalidation and semaphore
- * release after invalidation to avoid entering power gated state
- * to WA the Issue
- */
-
- /* TODO: It needs to continue working on debugging with semaphore for GFXHUB as well. */
- if (use_semaphore)
- /* a read return value of 1 means semaphore acuqire */
- amdgpu_ring_emit_reg_wait(ring,
- hub->vm_inv_eng0_sem +
- hub->eng_distance * eng, 0x1, 0x1);
-
- amdgpu_ring_emit_wreg(ring, hub->ctx0_ptb_addr_lo32 +
- (hub->ctx_addr_distance * vmid),
- lower_32_bits(pd_addr));
-
- amdgpu_ring_emit_wreg(ring, hub->ctx0_ptb_addr_hi32 +
- (hub->ctx_addr_distance * vmid),
- upper_32_bits(pd_addr));
-
- amdgpu_ring_emit_reg_write_reg_wait(ring, hub->vm_inv_eng0_req +
- hub->eng_distance * eng,
- hub->vm_inv_eng0_ack +
- hub->eng_distance * eng,
- req, 1 << vmid);
-
- /* TODO: It needs to continue working on debugging with semaphore for GFXHUB as well. */
- if (use_semaphore)
- /*
- * add semaphore release after invalidation,
- * write with 0 means semaphore release
- */
- amdgpu_ring_emit_wreg(ring, hub->vm_inv_eng0_sem +
- hub->eng_distance * eng, 0);
-
- return pd_addr;
-}
-
static void gmc_v11_0_emit_pasid_mapping(struct amdgpu_ring *ring, unsigned int vmid,
unsigned int pasid)
{
@@ -543,9 +347,9 @@ static unsigned int gmc_v11_0_get_vbios_fb_size(struct amdgpu_device *adev)
}
static const struct amdgpu_gmc_funcs gmc_v11_0_gmc_funcs = {
- .flush_gpu_tlb = gmc_v11_0_flush_gpu_tlb,
- .flush_gpu_tlb_pasid = gmc_v11_0_flush_gpu_tlb_pasid,
- .emit_flush_gpu_tlb = gmc_v11_0_emit_flush_gpu_tlb,
+ .flush_gpu_tlb = amdgpu_gmc_flush_gpu_tlb_helper,
+ .flush_gpu_tlb_pasid = amdgpu_gmc_flush_gpu_tlb_pasid_helper,
+ .emit_flush_gpu_tlb = amdgpu_gmc_emit_flush_gpu_tlb_helper,
.emit_pasid_mapping = gmc_v11_0_emit_pasid_mapping,
.get_vmid_pasid_mapping_info = gmc_v11_0_get_vmid_pasid_mapping_info,
.use_invalidate_semaphore = gmc_v11_0_use_invalidate_semaphore,
@@ -935,7 +739,7 @@ static int gmc_v11_0_gart_enable(struct amdgpu_device *adev)
value = amdgpu_vm_fault_stop != AMDGPU_VM_FAULT_STOP_ALWAYS;
adev->mmhub.funcs->set_fault_enable_default(adev, value);
- gmc_v11_0_flush_gpu_tlb(adev, 0, AMDGPU_MMHUB0(0), 0);
+ adev->gmc.gmc_funcs->flush_gpu_tlb(adev, 0, AMDGPU_MMHUB0(0), 0);
drm_info(adev_to_drm(adev), "PCIE GART of %uM enabled (table at 0x%016llX).\n",
(unsigned int)(adev->gmc.gart_size >> 20),
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
index 4bb8c1e4f335b..842ea38347367 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
@@ -211,219 +211,6 @@ static bool gmc_v12_0_get_vmid_pasid_mapping_info(
* by the amdgpu vm/hsa code.
*/
-static void gmc_v12_0_flush_vm_hub(struct amdgpu_device *adev, uint32_t vmid,
- unsigned int vmhub, uint32_t flush_type)
-{
- bool use_semaphore = gmc_v12_0_use_invalidate_semaphore(adev, vmhub);
- struct amdgpu_vmhub *hub = &adev->vmhub[vmhub];
- u32 inv_req = hub->vmhub_funcs->get_invalidate_req(vmid, flush_type);
- u32 tmp;
- /* Use register 17 for GART */
- const unsigned eng = 17;
- unsigned int i;
- unsigned char hub_ip = 0;
-
- hub_ip = (vmhub == AMDGPU_GFXHUB(0)) ?
- GC_HWIP : MMHUB_HWIP;
-
- spin_lock(&adev->gmc.invalidate_lock);
- /*
- * It may lose gpuvm invalidate acknowldege state across power-gating
- * off cycle, add semaphore acquire before invalidation and semaphore
- * release after invalidation to avoid entering power gated state
- * to WA the Issue
- */
-
- /* TODO: It needs to continue working on debugging with semaphore for GFXHUB as well. */
- if (use_semaphore) {
- for (i = 0; i < adev->usec_timeout; i++) {
- /* a read return value of 1 means semaphore acuqire */
- tmp = RREG32_RLC_NO_KIQ(hub->vm_inv_eng0_sem +
- hub->eng_distance * eng, hub_ip);
- if (tmp & 0x1)
- break;
- udelay(1);
- }
-
- if (i >= adev->usec_timeout)
- dev_err(adev->dev,
- "Timeout waiting for sem acquire in VM flush!\n");
- }
-
- WREG32_RLC_NO_KIQ(hub->vm_inv_eng0_req + hub->eng_distance * eng, inv_req, hub_ip);
-
- /* Wait for ACK with a delay.*/
- for (i = 0; i < adev->usec_timeout; i++) {
- tmp = RREG32_RLC_NO_KIQ(hub->vm_inv_eng0_ack +
- hub->eng_distance * eng, hub_ip);
- tmp &= 1 << vmid;
- if (tmp)
- break;
-
- udelay(1);
- }
-
- /* TODO: It needs to continue working on debugging with semaphore for GFXHUB as well. */
- if (use_semaphore)
- /*
- * add semaphore release after invalidation,
- * write with 0 means semaphore release
- */
- WREG32_RLC_NO_KIQ(hub->vm_inv_eng0_sem +
- hub->eng_distance * eng, 0, hub_ip);
-
- /* Issue additional private vm invalidation to MMHUB */
- if ((vmhub != AMDGPU_GFXHUB(0)) &&
- (hub->vm_l2_bank_select_reserved_cid2) &&
- !amdgpu_sriov_vf(adev)) {
- inv_req = RREG32_NO_KIQ(hub->vm_l2_bank_select_reserved_cid2);
- /* bit 25: RSERVED_CACHE_PRIVATE_INVALIDATION */
- inv_req |= (1 << 25);
- /* Issue private invalidation */
- WREG32_NO_KIQ(hub->vm_l2_bank_select_reserved_cid2, inv_req);
- /* Read back to ensure invalidation is done*/
- RREG32_NO_KIQ(hub->vm_l2_bank_select_reserved_cid2);
- }
-
- spin_unlock(&adev->gmc.invalidate_lock);
-
- if (i < adev->usec_timeout)
- return;
-
- dev_err(adev->dev, "Timeout waiting for VM flush ACK!\n");
-}
-
-/**
- * gmc_v12_0_flush_gpu_tlb - gart tlb flush callback
- *
- * @adev: amdgpu_device pointer
- * @vmid: vm instance to flush
- * @vmhub: which hub to flush
- * @flush_type: the flush type
- *
- * Flush the TLB for the requested page table.
- */
-static void gmc_v12_0_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
- uint32_t vmhub, uint32_t flush_type)
-{
- if ((vmhub == AMDGPU_GFXHUB(0)) && !adev->gfx.is_poweron)
- return;
-
- /* flush hdp cache */
- amdgpu_device_flush_hdp(adev, NULL);
-
- /* This is necessary for SRIOV as well as for GFXOFF to function
- * properly under bare metal
- */
- if ((adev->gfx.kiq[0].ring.sched.ready || adev->mes.ring[0].sched.ready) &&
- !adev->gmc.use_mmio_for_tlb_flush) {
- struct amdgpu_vmhub *hub = &adev->vmhub[vmhub];
- const unsigned eng = 17;
- u32 inv_req = hub->vmhub_funcs->get_invalidate_req(vmid, flush_type);
- u32 req = hub->vm_inv_eng0_req + hub->eng_distance * eng;
- u32 ack = hub->vm_inv_eng0_ack + hub->eng_distance * eng;
-
- amdgpu_gmc_fw_reg_write_reg_wait(adev, req, ack, inv_req,
- 1 << vmid, GET_INST(GC, 0));
- return;
- }
-
- /* disabllow gfxoff when we invalidate */
- if (vmhub == AMDGPU_GFXHUB(0))
- amdgpu_gfx_off_ctrl(adev, false);
-
- gmc_v12_0_flush_vm_hub(adev, vmid, vmhub, 0);
-
- if (vmhub == AMDGPU_GFXHUB(0))
- amdgpu_gfx_off_ctrl(adev, true);
-}
-
-/**
- * gmc_v12_0_flush_gpu_tlb_pasid - tlb flush via pasid
- *
- * @adev: amdgpu_device pointer
- * @pasid: pasid to be flush
- * @flush_type: the flush type
- * @all_hub: flush all hubs
- * @inst: is used to select which instance of KIQ to use for the invalidation
- *
- * Flush the TLB for the requested pasid.
- */
-static void gmc_v12_0_flush_gpu_tlb_pasid(struct amdgpu_device *adev,
- uint16_t pasid, uint32_t flush_type,
- bool all_hub, uint32_t inst)
-{
- uint16_t queried;
- int vmid, i;
-
- for (vmid = 1; vmid < 16; vmid++) {
- bool valid;
-
- valid = gmc_v12_0_get_vmid_pasid_mapping_info(adev, vmid, 0,
- &queried);
- if (!valid || queried != pasid)
- continue;
-
- if (all_hub) {
- for_each_set_bit(i, adev->vmhubs_mask,
- AMDGPU_MAX_VMHUBS)
- gmc_v12_0_flush_gpu_tlb(adev, vmid, i,
- flush_type);
- } else {
- gmc_v12_0_flush_gpu_tlb(adev, vmid, AMDGPU_GFXHUB(0),
- flush_type);
- }
- }
-}
-
-static uint64_t gmc_v12_0_emit_flush_gpu_tlb(struct amdgpu_ring *ring,
- unsigned vmid, uint64_t pd_addr)
-{
- bool use_semaphore = gmc_v12_0_use_invalidate_semaphore(ring->adev, ring->vm_hub);
- struct amdgpu_vmhub *hub = &ring->adev->vmhub[ring->vm_hub];
- uint32_t req = hub->vmhub_funcs->get_invalidate_req(vmid, 0);
- unsigned eng = ring->vm_inv_eng;
-
- /*
- * It may lose gpuvm invalidate acknowldege state across power-gating
- * off cycle, add semaphore acquire before invalidation and semaphore
- * release after invalidation to avoid entering power gated state
- * to WA the Issue
- */
-
- /* TODO: It needs to continue working on debugging with semaphore for GFXHUB as well. */
- if (use_semaphore)
- /* a read return value of 1 means semaphore acuqire */
- amdgpu_ring_emit_reg_wait(ring,
- hub->vm_inv_eng0_sem +
- hub->eng_distance * eng, 0x1, 0x1);
-
- amdgpu_ring_emit_wreg(ring, hub->ctx0_ptb_addr_lo32 +
- (hub->ctx_addr_distance * vmid),
- lower_32_bits(pd_addr));
-
- amdgpu_ring_emit_wreg(ring, hub->ctx0_ptb_addr_hi32 +
- (hub->ctx_addr_distance * vmid),
- upper_32_bits(pd_addr));
-
- amdgpu_ring_emit_reg_write_reg_wait(ring, hub->vm_inv_eng0_req +
- hub->eng_distance * eng,
- hub->vm_inv_eng0_ack +
- hub->eng_distance * eng,
- req, 1 << vmid);
-
- /* TODO: It needs to continue working on debugging with semaphore for GFXHUB as well. */
- if (use_semaphore)
- /*
- * add semaphore release after invalidation,
- * write with 0 means semaphore release
- */
- amdgpu_ring_emit_wreg(ring, hub->vm_inv_eng0_sem +
- hub->eng_distance * eng, 0);
-
- return pd_addr;
-}
-
static void gmc_v12_0_emit_pasid_mapping(struct amdgpu_ring *ring, unsigned vmid,
unsigned pasid)
{
@@ -562,9 +349,9 @@ static unsigned int gmc_v12_0_get_dcc_alignment(struct amdgpu_device *adev)
}
static const struct amdgpu_gmc_funcs gmc_v12_0_gmc_funcs = {
- .flush_gpu_tlb = gmc_v12_0_flush_gpu_tlb,
- .flush_gpu_tlb_pasid = gmc_v12_0_flush_gpu_tlb_pasid,
- .emit_flush_gpu_tlb = gmc_v12_0_emit_flush_gpu_tlb,
+ .flush_gpu_tlb = amdgpu_gmc_flush_gpu_tlb_helper,
+ .flush_gpu_tlb_pasid = amdgpu_gmc_flush_gpu_tlb_pasid_helper,
+ .emit_flush_gpu_tlb = amdgpu_gmc_emit_flush_gpu_tlb_helper,
.emit_pasid_mapping = gmc_v12_0_emit_pasid_mapping,
.get_vmid_pasid_mapping_info = gmc_v12_0_get_vmid_pasid_mapping_info,
.use_invalidate_semaphore = gmc_v12_0_use_invalidate_semaphore,
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c b/drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c
index 2c980dc02c262..b20088e841349 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c
@@ -281,195 +281,6 @@ static bool gmc_v12_1_use_invalidate_semaphore(struct amdgpu_device *adev,
(!amdgpu_sriov_vf(adev)));
}
-static void gmc_v12_1_flush_vm_hub(struct amdgpu_device *adev, uint32_t vmid,
- unsigned int vmhub, uint32_t flush_type)
-{
- bool use_semaphore = gmc_v12_1_use_invalidate_semaphore(adev, vmhub);
- struct amdgpu_vmhub *hub = &adev->vmhub[vmhub];
- u32 inv_req = hub->vmhub_funcs->get_invalidate_req(vmid, flush_type);
- u32 tmp;
- /* Use register 17 for GART */
- const unsigned eng = 17;
- unsigned int i;
- unsigned char hub_ip = 0;
-
- hub_ip = (AMDGPU_IS_GFXHUB(vmhub)) ?
- GC_HWIP : MMHUB_HWIP;
-
- spin_lock(&adev->gmc.invalidate_lock);
-
- if (use_semaphore) {
- for (i = 0; i < adev->usec_timeout; i++) {
- /* a read return value of 1 means semaphore acuqire */
- tmp = RREG32_RLC_NO_KIQ(hub->vm_inv_eng0_sem + hub->eng_distance * eng, hub_ip);
- if (tmp & 0x1)
- break;
- udelay(1);
- }
-
- if (i >= adev->usec_timeout)
- DRM_ERROR("Timeout waiting for sem acquire in VM flush!\n");
- }
-
- WREG32_RLC_NO_KIQ(hub->vm_inv_eng0_req + hub->eng_distance * eng, inv_req, hub_ip);
-
- /* Wait for ACK with a delay.*/
- for (i = 0; i < adev->usec_timeout; i++) {
- tmp = RREG32_RLC_NO_KIQ(hub->vm_inv_eng0_ack +
- hub->eng_distance * eng, hub_ip);
- tmp &= 1 << vmid;
- if (tmp)
- break;
-
- udelay(1);
- }
-
- if (use_semaphore)
- WREG32_RLC_NO_KIQ(hub->vm_inv_eng0_sem + hub->eng_distance * eng, 0, hub_ip);
-
- /* Issue additional private vm invalidation to MMHUB */
- if (!AMDGPU_IS_GFXHUB(vmhub) &&
- (hub->vm_l2_bank_select_reserved_cid2) &&
- !amdgpu_sriov_vf(adev)) {
- inv_req = RREG32_NO_KIQ(hub->vm_l2_bank_select_reserved_cid2);
- /* bit 25: RSERVED_CACHE_PRIVATE_INVALIDATION */
- inv_req |= (1 << 25);
- /* Issue private invalidation */
- WREG32_NO_KIQ(hub->vm_l2_bank_select_reserved_cid2, inv_req);
- /* Read back to ensure invalidation is done*/
- RREG32_NO_KIQ(hub->vm_l2_bank_select_reserved_cid2);
- }
-
- spin_unlock(&adev->gmc.invalidate_lock);
-
- if (i < adev->usec_timeout)
- return;
-
- dev_err(adev->dev, "Timeout waiting for VM flush ACK!\n");
-}
-
-/**
- * gmc_v12_1_flush_gpu_tlb - gart tlb flush callback
- *
- * @adev: amdgpu_device pointer
- * @vmid: vm instance to flush
- * @vmhub: which hub to flush
- * @flush_type: the flush type
- *
- * Flush the TLB for the requested page table.
- */
-static void gmc_v12_1_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
- uint32_t vmhub, uint32_t flush_type)
-{
- u32 inst;
-
- if (AMDGPU_IS_GFXHUB(vmhub) &&
- !adev->gfx.is_poweron)
- return;
-
- if (vmhub >= AMDGPU_MMHUB0(0))
- inst = 0;
- else
- inst = vmhub;
-
- /* This is necessary for SRIOV as well as for GFXOFF to function
- * properly under bare metal
- */
- if ((adev->gfx.kiq[inst].ring.sched.ready ||
- adev->mes.ring[MES_PIPE_INST(inst, 0)].sched.ready) &&
- !adev->gmc.use_mmio_for_tlb_flush) {
- struct amdgpu_vmhub *hub = &adev->vmhub[vmhub];
- const unsigned eng = 17;
- u32 inv_req = hub->vmhub_funcs->get_invalidate_req(vmid, flush_type);
- u32 req = hub->vm_inv_eng0_req + hub->eng_distance * eng;
- u32 ack = hub->vm_inv_eng0_ack + hub->eng_distance * eng;
-
- amdgpu_gmc_fw_reg_write_reg_wait(adev, req, ack, inv_req,
- 1 << vmid, inst);
- return;
- }
-
- gmc_v12_1_flush_vm_hub(adev, vmid, vmhub, 0);
- return;
-}
-
-/**
- * gmc_v12_1_flush_gpu_tlb_pasid - tlb flush via pasid
- *
- * @adev: amdgpu_device pointer
- * @pasid: pasid to be flush
- * @flush_type: the flush type
- * @all_hub: flush all hubs
- * @inst: is used to select which instance of KIQ to use for the invalidation
- *
- * Flush the TLB for the requested pasid.
- */
-static void gmc_v12_1_flush_gpu_tlb_pasid(struct amdgpu_device *adev,
- uint16_t pasid, uint32_t flush_type,
- bool all_hub, uint32_t inst)
-{
- uint16_t queried;
- int vmid, i;
-
- for (vmid = 1; vmid < 16; vmid++) {
- bool valid;
-
- valid = gmc_v12_1_get_vmid_pasid_mapping_info(adev, vmid, inst,
- &queried);
- if (!valid || queried != pasid)
- continue;
-
- if (all_hub) {
- for_each_set_bit(i, adev->vmhubs_mask,
- AMDGPU_MAX_VMHUBS)
- gmc_v12_1_flush_gpu_tlb(adev, vmid, i,
- flush_type);
- } else {
- gmc_v12_1_flush_gpu_tlb(adev, vmid, AMDGPU_GFXHUB(inst),
- flush_type);
- }
- }
-}
-
-static uint64_t gmc_v12_1_emit_flush_gpu_tlb(struct amdgpu_ring *ring,
- unsigned vmid, uint64_t pd_addr)
-{
- bool use_semaphore = gmc_v12_1_use_invalidate_semaphore(ring->adev, ring->vm_hub);
- struct amdgpu_vmhub *hub = &ring->adev->vmhub[ring->vm_hub];
- uint32_t req = hub->vmhub_funcs->get_invalidate_req(vmid, 0);
- unsigned eng = ring->vm_inv_eng;
-
- if (use_semaphore)
- /* a read return value of 1 means semaphore acuqire */
- amdgpu_ring_emit_reg_wait(ring,
- hub->vm_inv_eng0_sem +
- hub->eng_distance * eng, 0x1, 0x1);
-
- amdgpu_ring_emit_wreg(ring, hub->ctx0_ptb_addr_lo32 +
- (hub->ctx_addr_distance * vmid),
- lower_32_bits(pd_addr));
-
- amdgpu_ring_emit_wreg(ring, hub->ctx0_ptb_addr_hi32 +
- (hub->ctx_addr_distance * vmid),
- upper_32_bits(pd_addr));
-
- amdgpu_ring_emit_reg_write_reg_wait(ring, hub->vm_inv_eng0_req +
- hub->eng_distance * eng,
- hub->vm_inv_eng0_ack +
- hub->eng_distance * eng,
- req, 1 << vmid);
-
- if (use_semaphore)
- /*
- * add semaphore release after invalidation,
- * write with 0 means semaphore release
- */
- amdgpu_ring_emit_wreg(ring, hub->vm_inv_eng0_sem +
- hub->eng_distance * eng, 0);
-
- return pd_addr;
-}
-
static void gmc_v12_1_emit_pasid_mapping(struct amdgpu_ring *ring,
unsigned vmid, unsigned pasid)
{
@@ -636,9 +447,9 @@ static void gmc_v12_1_get_vm_pte(struct amdgpu_device *adev,
}
static const struct amdgpu_gmc_funcs gmc_v12_1_gmc_funcs = {
- .flush_gpu_tlb = gmc_v12_1_flush_gpu_tlb,
- .flush_gpu_tlb_pasid = gmc_v12_1_flush_gpu_tlb_pasid,
- .emit_flush_gpu_tlb = gmc_v12_1_emit_flush_gpu_tlb,
+ .flush_gpu_tlb = amdgpu_gmc_flush_gpu_tlb_helper,
+ .flush_gpu_tlb_pasid = amdgpu_gmc_flush_gpu_tlb_pasid_helper,
+ .emit_flush_gpu_tlb = amdgpu_gmc_emit_flush_gpu_tlb_helper,
.emit_pasid_mapping = gmc_v12_1_emit_pasid_mapping,
.get_vmid_pasid_mapping_info = gmc_v12_1_get_vmid_pasid_mapping_info,
.use_invalidate_semaphore = gmc_v12_1_use_invalidate_semaphore,
diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
index ad4bb86b8ffbf..9cccbc36f1705 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
@@ -868,94 +868,6 @@ static void gmc_v9_0_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
DRM_ERROR("Timeout waiting for VM flush ACK!\n");
}
-/**
- * gmc_v9_0_flush_gpu_tlb_pasid - tlb flush via pasid
- *
- * @adev: amdgpu_device pointer
- * @pasid: pasid to be flush
- * @flush_type: the flush type
- * @all_hub: flush all hubs
- * @inst: is used to select which instance of KIQ to use for the invalidation
- *
- * Flush the TLB for the requested pasid.
- */
-static void gmc_v9_0_flush_gpu_tlb_pasid(struct amdgpu_device *adev,
- uint16_t pasid, uint32_t flush_type,
- bool all_hub, uint32_t inst)
-{
- uint16_t queried;
- int i, vmid;
-
- for (vmid = 1; vmid < 16; vmid++) {
- bool valid;
-
- valid = gmc_v9_0_get_atc_vmid_pasid_mapping_info(adev, vmid,
- inst, &queried);
- if (!valid || queried != pasid)
- continue;
-
- if (all_hub) {
- for_each_set_bit(i, adev->vmhubs_mask,
- AMDGPU_MAX_VMHUBS)
- gmc_v9_0_flush_gpu_tlb(adev, vmid, i,
- flush_type);
- } else {
- gmc_v9_0_flush_gpu_tlb(adev, vmid,
- AMDGPU_GFXHUB(0),
- flush_type);
- }
- }
-}
-
-static uint64_t gmc_v9_0_emit_flush_gpu_tlb(struct amdgpu_ring *ring,
- unsigned int vmid, uint64_t pd_addr)
-{
- bool use_semaphore = gmc_v9_0_use_invalidate_semaphore(ring->adev, ring->vm_hub);
- struct amdgpu_device *adev = ring->adev;
- struct amdgpu_vmhub *hub = &adev->vmhub[ring->vm_hub];
- uint32_t req = gmc_v9_0_get_invalidate_req(vmid, 0);
- unsigned int eng = ring->vm_inv_eng;
-
- /*
- * It may lose gpuvm invalidate acknowldege state across power-gating
- * off cycle, add semaphore acquire before invalidation and semaphore
- * release after invalidation to avoid entering power gated state
- * to WA the Issue
- */
-
- /* TODO: It needs to continue working on debugging with semaphore for GFXHUB as well. */
- if (use_semaphore)
- /* a read return value of 1 means semaphore acuqire */
- amdgpu_ring_emit_reg_wait(ring,
- hub->vm_inv_eng0_sem +
- hub->eng_distance * eng, 0x1, 0x1);
-
- amdgpu_ring_emit_wreg(ring, hub->ctx0_ptb_addr_lo32 +
- (hub->ctx_addr_distance * vmid),
- lower_32_bits(pd_addr));
-
- amdgpu_ring_emit_wreg(ring, hub->ctx0_ptb_addr_hi32 +
- (hub->ctx_addr_distance * vmid),
- upper_32_bits(pd_addr));
-
- amdgpu_ring_emit_reg_write_reg_wait(ring, hub->vm_inv_eng0_req +
- hub->eng_distance * eng,
- hub->vm_inv_eng0_ack +
- hub->eng_distance * eng,
- req, 1 << vmid);
-
- /* TODO: It needs to continue working on debugging with semaphore for GFXHUB as well. */
- if (use_semaphore)
- /*
- * add semaphore release after invalidation,
- * write with 0 means semaphore release
- */
- amdgpu_ring_emit_wreg(ring, hub->vm_inv_eng0_sem +
- hub->eng_distance * eng, 0);
-
- return pd_addr;
-}
-
static void gmc_v9_0_emit_pasid_mapping(struct amdgpu_ring *ring, unsigned int vmid,
unsigned int pasid)
{
@@ -1300,8 +1212,8 @@ static bool gmc_v9_0_need_reset_on_init(struct amdgpu_device *adev)
static const struct amdgpu_gmc_funcs gmc_v9_0_gmc_funcs = {
.flush_gpu_tlb = gmc_v9_0_flush_gpu_tlb,
- .flush_gpu_tlb_pasid = gmc_v9_0_flush_gpu_tlb_pasid,
- .emit_flush_gpu_tlb = gmc_v9_0_emit_flush_gpu_tlb,
+ .flush_gpu_tlb_pasid = amdgpu_gmc_flush_gpu_tlb_pasid_helper,
+ .emit_flush_gpu_tlb = amdgpu_gmc_emit_flush_gpu_tlb_helper,
.emit_pasid_mapping = gmc_v9_0_emit_pasid_mapping,
.get_vmid_pasid_mapping_info = gmc_v9_0_get_atc_vmid_pasid_mapping_info,
.use_invalidate_semaphore = gmc_v9_0_use_invalidate_semaphore,
--
2.55.0
^ permalink raw reply related [flat|nested] 43+ messages in thread
* RE: [PATCH 22/31] drm/amdgpu/gmc: rework pasid flushing
2026-09-01 20:10 ` [PATCH 22/31] drm/amdgpu/gmc: rework pasid flushing Alex Deucher
@ 2026-09-02 3:04 ` Zhang, Jesse(Jie)
0 siblings, 0 replies; 43+ messages in thread
From: Zhang, Jesse(Jie) @ 2026-09-02 3:04 UTC (permalink / raw)
To: Deucher, Alexander, amd-gfx@lists.freedesktop.org,
Koenig, Christian
Cc: Deucher, Alexander
AMD General
> -----Original Message-----
> From: amd-gfx <amd-gfx-bounces@lists.freedesktop.org> On Behalf Of Alex
> Deucher
> Sent: Wednesday, September 2, 2026 4:10 AM
> To: amd-gfx@lists.freedesktop.org; Koenig, Christian
> <Christian.Koenig@amd.com>
> Cc: Deucher, Alexander <Alexander.Deucher@amd.com>
> Subject: [PATCH 22/31] drm/amdgpu/gmc: rework pasid flushing
>
> Split out all of the various flush methods and use the new pasid flush method enum
> to determine which one to use.
>
> Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
> ---
> drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c | 247 +++++++++++++++++++-----
> drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c | 2 +
> drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c | 2 +
> drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c | 2 +
> 4 files changed, 201 insertions(+), 52 deletions(-)
>
> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c
> b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c
> index 8a975eddd75c7..2800eebe50649 100644
> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c
> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c
> @@ -788,9 +788,66 @@ void amdgpu_gmc_flush_gpu_tlb_gart(struct
> amdgpu_device *adev,
> dev_err(adev->dev, "Error flushing GPU TLB using the SDMA (%d)!\n", r); }
>
> -int amdgpu_gmc_flush_gpu_tlb_pasid(struct amdgpu_device *adev, uint16_t pasid,
> - uint32_t flush_type, bool all_hub,
> - uint32_t inst)
> +static int amdgpu_gmc_flush_gpu_tlb_sdma_helper(struct amdgpu_device *adev,
> + uint16_t pasid, uint32_t flush_type,
> + bool all_hub, uint32_t inst)
> +{
> + struct amdgpu_ring *ring;
> + struct dma_fence *fence;
> + struct amdgpu_job *job;
> + uint16_t queried;
> + /* Use register 17 for GART */
> + u32 eng = 17;
> + int vmid, r, i, ndw;
> +
> + /* flush hdp cache */
> + amdgpu_device_flush_hdp(adev, NULL);
> +
> + ndw = ALIGN(adev->mman.buffer_funcs->tlb_inv_num_dw * 16 *
> AMDGPU_MAX_VMHUBS, 8);
> + ring = to_amdgpu_ring(adev->mman.buffer_funcs_scheds[0]);
> +
> + mutex_lock(&adev->mman.default_entity.lock);
> + r = amdgpu_job_alloc_with_ib(ring->adev, &adev-
> >mman.default_entity.base,
> + AMDGPU_FENCE_OWNER_UNDEFINED,
> + ndw * 4, AMDGPU_IB_POOL_IMMEDIATE,
> + AMDGPU_KERNEL_JOB_ID_VM_UPDATE,
> + &job);
> + if (r)
> + goto exit;
> +
> + for (vmid = 1; vmid < 16; vmid++) {
> + bool valid;
> +
> + valid = adev->gmc.gmc_funcs->get_vmid_pasid_mapping_info(adev,
> vmid, inst,
> + &queried);
> + if (!valid || queried != pasid)
> + continue;
> +
> + if (all_hub) {
> + for_each_set_bit(i, adev->vmhubs_mask,
> AMDGPU_MAX_VMHUBS)
> + amdgpu_emit_tlb_inv(adev, &job->ibs[0], vmid, i, eng,
> + flush_type, inst);
> + } else {
> + amdgpu_emit_tlb_inv(adev, &job->ibs[0], vmid,
> AMDGPU_GFXHUB(inst), eng,
> + flush_type, inst);
> + }
> + }
> + amdgpu_ring_pad_ib(ring, &job->ibs[0]);
> + fence = amdgpu_job_submit(job);
> + mutex_unlock(&adev->mman.default_entity.lock);
It should be double unlock, and it unlocks again at exit. we can remove it .
Regards
Jesse
> +
> + dma_fence_wait(fence, false);
> + dma_fence_put(fence);
> +
> +exit:
> + mutex_unlock(&adev->mman.default_entity.lock);
> +
> + return r;
> +}
> +
> +static int amdgpu_gmc_flush_gpu_tlb_kiq_helper(struct amdgpu_device *adev,
> + uint16_t pasid, uint32_t flush_type,
> + bool all_hub, uint32_t inst)
> {
> struct amdgpu_ring *ring = &adev->gfx.kiq[inst].ring;
> struct amdgpu_kiq *kiq = &adev->gfx.kiq[inst]; @@ -798,6 +855,110 @@
> int amdgpu_gmc_flush_gpu_tlb_pasid(struct amdgpu_device *adev, uint16_t pasid,
> int r, cnt = 0;
> uint32_t seq;
>
> + /* 2 dwords flush + 8 dwords fence */
> + ndw = kiq->pmf->invalidate_tlbs_size + 8;
> +
> + if (adev->gmc.flush_tlb_needs_extra_type_2)
> + ndw += kiq->pmf->invalidate_tlbs_size;
> +
> + if (adev->gmc.flush_tlb_needs_extra_type_0)
> + ndw += kiq->pmf->invalidate_tlbs_size;
> +
> + spin_lock(&adev->gfx.kiq[inst].ring_lock);
> + r = amdgpu_ring_alloc(ring, ndw);
> + if (r) {
> + spin_unlock(&adev->gfx.kiq[inst].ring_lock);
> + return r;
> + }
> + if (adev->gmc.flush_tlb_needs_extra_type_2)
> + kiq->pmf->kiq_invalidate_tlbs(ring, pasid, 2, all_hub);
> +
> + if (flush_type == 2 && adev->gmc.flush_tlb_needs_extra_type_0)
> + kiq->pmf->kiq_invalidate_tlbs(ring, pasid, 0, all_hub);
> +
> + kiq->pmf->kiq_invalidate_tlbs(ring, pasid, flush_type, all_hub);
> + r = amdgpu_fence_emit_polling(ring, &seq, MAX_KIQ_REG_WAIT);
> + if (r) {
> + amdgpu_ring_undo(ring);
> + spin_unlock(&adev->gfx.kiq[inst].ring_lock);
> + return r;
> + }
> +
> + amdgpu_ring_commit(ring);
> + spin_unlock(&adev->gfx.kiq[inst].ring_lock);
> +
> + r = amdgpu_fence_wait_polling(ring, seq, MAX_KIQ_REG_WAIT);
> +
> + might_sleep();
> + while (r < 1 && cnt++ < MAX_KIQ_REG_TRY &&
> + !amdgpu_reset_pending(adev->reset_domain)) {
> + msleep(MAX_KIQ_REG_BAILOUT_INTERVAL);
> + r = amdgpu_fence_wait_polling(ring, seq, MAX_KIQ_REG_WAIT);
> + }
> +
> + if (cnt > MAX_KIQ_REG_TRY) {
> + dev_err(adev->dev, "timeout waiting for kiq fence\n");
> + r = -ETIME;
> + } else
> + r = 0;
> +
> + return r;
> +}
> +
> +static int amdgpu_gmc_flush_gpu_tlb_mes_helper(struct amdgpu_device *adev,
> + uint16_t pasid, uint32_t flush_type,
> + bool all_hub, uint32_t inst) {
> + struct mes_inv_tlbs_pasid_input input = {0};
> + int r;
> +
> + input.xcc_id = inst;
> + input.pasid = pasid;
> + input.flush_type = flush_type;
> +
> + /* MES will invalidate hubs for the device(including slave xcc)
> + * from master, ignore request from slave
> + */
> + if (!amdgpu_gfx_is_master_xcc(adev, inst))
> + return -EINVAL;
> +
> + input.hub_id = AMDGPU_GFXHUB(0);
> + amdgpu_mes_lock(&adev->mes);
> + r = adev->mes.funcs->invalidate_tlbs_pasid(&adev->mes, &input);
> + amdgpu_mes_unlock(&adev->mes);
> + if (r)
> + return r;
> +
> + if (all_hub) {
> + /* invalidate mm_hub */
> + if (test_bit(AMDGPU_MMHUB0(0), adev->vmhubs_mask)) {
> + input.hub_id = AMDGPU_MMHUB0(0);
> + amdgpu_mes_lock(&adev->mes);
> + r = adev->mes.funcs->invalidate_tlbs_pasid(&adev->mes,
> &input);
> + amdgpu_mes_unlock(&adev->mes);
> + if (r)
> + return r;
> + }
> + if (test_bit(AMDGPU_MMHUB1(0), adev->vmhubs_mask)) {
> + input.hub_id = AMDGPU_MMHUB1(0);
> + amdgpu_mes_lock(&adev->mes);
> + r = adev->mes.funcs->invalidate_tlbs_pasid(&adev->mes,
> &input);
> + amdgpu_mes_unlock(&adev->mes);
> + if (r)
> + return r;
> + }
> + }
> + return 0;
> +}
> +
> +int amdgpu_gmc_flush_gpu_tlb_pasid(struct amdgpu_device *adev, uint16_t
> pasid,
> + uint32_t flush_type, bool all_hub,
> + uint32_t inst)
> +{
> + struct amdgpu_ring *ring;
> + bool use_mmio = false;
> + int r;
> +
> /*
> * A GPU reset should flush all TLBs anyway, so no need to do
> * this while one is ongoing.
> @@ -805,8 +966,38 @@ int amdgpu_gmc_flush_gpu_tlb_pasid(struct
> amdgpu_device *adev, uint16_t pasid,
> if (!down_read_trylock(&adev->reset_domain->sem))
> return 0;
>
> - if (!adev->gmc.flush_pasid_uses_kiq || !ring->sched.ready) {
> + switch (adev->gmc.pasid_inv_method) {
> + case AMDGPU_TLB_INV_METHOD_MMIO:
> + default:
> + use_mmio = true;
> + break;
> + case AMDGPU_TLB_INV_METHOD_SDMA:
> + ring = to_amdgpu_ring(adev->mman.buffer_funcs_scheds[0]);
> + if (!ring->sched.ready)
> + use_mmio = true;
> + else
> + r = amdgpu_gmc_flush_gpu_tlb_sdma_helper(adev, pasid,
> flush_type,
> + all_hub, inst);
> + break;
> + case AMDGPU_TLB_INV_METHOD_KIQ:
> + ring = &adev->gfx.kiq[inst].ring;
> + if (!adev->gmc.flush_pasid_uses_kiq || !ring->sched.ready)
> + use_mmio = true;
> + else
> + r = amdgpu_gmc_flush_gpu_tlb_kiq_helper(adev, pasid,
> flush_type,
> + all_hub, inst);
> + break;
> + case AMDGPU_TLB_INV_METHOD_MES:
> + ring = &adev->mes.ring[MES_PIPE_INST(inst, 0)];
> + if (!ring->sched.ready)
> + use_mmio = true;
> + else
> + r = amdgpu_gmc_flush_gpu_tlb_mes_helper(adev, pasid,
> flush_type,
> + all_hub, inst);
> + break;
> + }
>
> + if (use_mmio) {
> if (!adev->gmc.gmc_funcs->flush_gpu_tlb_pasid) {
> r = 0;
> goto error_unlock_reset;
> @@ -825,54 +1016,6 @@ int amdgpu_gmc_flush_gpu_tlb_pasid(struct
> amdgpu_device *adev, uint16_t pasid,
> adev->gmc.gmc_funcs->flush_gpu_tlb_pasid(adev, pasid,
> flush_type, all_hub,
> inst);
> - r = 0;
> - } else {
> - /* 2 dwords flush + 8 dwords fence */
> - ndw = kiq->pmf->invalidate_tlbs_size + 8;
> -
> - if (adev->gmc.flush_tlb_needs_extra_type_2)
> - ndw += kiq->pmf->invalidate_tlbs_size;
> -
> - if (adev->gmc.flush_tlb_needs_extra_type_0)
> - ndw += kiq->pmf->invalidate_tlbs_size;
> -
> - spin_lock(&adev->gfx.kiq[inst].ring_lock);
> - r = amdgpu_ring_alloc(ring, ndw);
> - if (r) {
> - spin_unlock(&adev->gfx.kiq[inst].ring_lock);
> - goto error_unlock_reset;
> - }
> - if (adev->gmc.flush_tlb_needs_extra_type_2)
> - kiq->pmf->kiq_invalidate_tlbs(ring, pasid, 2, all_hub);
> -
> - if (flush_type == 2 && adev->gmc.flush_tlb_needs_extra_type_0)
> - kiq->pmf->kiq_invalidate_tlbs(ring, pasid, 0, all_hub);
> -
> - kiq->pmf->kiq_invalidate_tlbs(ring, pasid, flush_type, all_hub);
> - r = amdgpu_fence_emit_polling(ring, &seq, MAX_KIQ_REG_WAIT);
> - if (r) {
> - amdgpu_ring_undo(ring);
> - spin_unlock(&adev->gfx.kiq[inst].ring_lock);
> - goto error_unlock_reset;
> - }
> -
> - amdgpu_ring_commit(ring);
> - spin_unlock(&adev->gfx.kiq[inst].ring_lock);
> -
> - r = amdgpu_fence_wait_polling(ring, seq, MAX_KIQ_REG_WAIT);
> -
> - might_sleep();
> - while (r < 1 && cnt++ < MAX_KIQ_REG_TRY &&
> - !amdgpu_reset_pending(adev->reset_domain)) {
> - msleep(MAX_KIQ_REG_BAILOUT_INTERVAL);
> - r = amdgpu_fence_wait_polling(ring, seq,
> MAX_KIQ_REG_WAIT);
> - }
> -
> - if (cnt > MAX_KIQ_REG_TRY) {
> - dev_err(adev->dev, "timeout waiting for kiq fence\n");
> - r = -ETIME;
> - } else
> - r = 0;
> }
>
> error_unlock_reset:
> diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
> b/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
> index 4bb10a89ce582..e7c529620199b 100644
> --- a/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
> +++ b/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
> @@ -974,6 +974,8 @@ static int gmc_v10_0_hw_init(struct amdgpu_ip_block
> *ip_block)
> int r;
>
> adev->gmc.flush_pasid_uses_kiq = !amdgpu_emu_mode;
> + if (adev->gmc.flush_pasid_uses_kiq)
> + adev->gmc.pasid_inv_method =
> AMDGPU_TLB_INV_METHOD_KIQ;
>
> /* The sequence of these two function calls matters.*/
> gmc_v10_0_init_golden_registers(adev);
> diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
> b/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
> index acab9f4180da6..a35c84cc7385f 100644
> --- a/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
> +++ b/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
> @@ -950,6 +950,8 @@ static int gmc_v11_0_hw_init(struct amdgpu_ip_block
> *ip_block)
> int r;
>
> adev->gmc.flush_pasid_uses_kiq = !amdgpu_emu_mode;
> + if (adev->gmc.flush_pasid_uses_kiq)
> + adev->gmc.pasid_inv_method =
> AMDGPU_TLB_INV_METHOD_KIQ;
>
> /* The sequence of these two function calls matters.*/
> gmc_v11_0_init_golden_registers(adev);
> diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
> b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
> index 3c13a6920171a..317b44412b5a6 100644
> --- a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
> +++ b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
> @@ -2162,6 +2162,8 @@ static int gmc_v9_0_hw_init(struct amdgpu_ip_block
> *ip_block)
> int i, r;
>
> adev->gmc.flush_pasid_uses_kiq = true;
> + if (adev->gmc.flush_pasid_uses_kiq)
> + adev->gmc.pasid_inv_method =
> AMDGPU_TLB_INV_METHOD_KIQ;
>
> /* Vega20+XGMI caches PTEs in TC and TLB. Add a heavy-weight TLB
> flush
> * (type 2), which flushes both. Due to a race condition with
> --
> 2.55.0
^ permalink raw reply [flat|nested] 43+ messages in thread
* RE: [PATCH 20/31] drm/amdgpu/gmc: add new callback to lookup vmid to pasid mapping
2026-09-01 20:10 ` [PATCH 20/31] drm/amdgpu/gmc: add new callback to lookup vmid to pasid mapping Alex Deucher
@ 2026-09-02 6:26 ` Zhang, Jesse(Jie)
0 siblings, 0 replies; 43+ messages in thread
From: Zhang, Jesse(Jie) @ 2026-09-02 6:26 UTC (permalink / raw)
To: Deucher, Alexander, amd-gfx@lists.freedesktop.org,
Koenig, Christian
Cc: Deucher, Alexander
AMD General
> -----Original Message-----
> From: amd-gfx <amd-gfx-bounces@lists.freedesktop.org> On Behalf Of Alex
> Deucher
> Sent: Wednesday, September 2, 2026 4:10 AM
> To: amd-gfx@lists.freedesktop.org; Koenig, Christian
> <Christian.Koenig@amd.com>
> Cc: Deucher, Alexander <Alexander.Deucher@amd.com>
> Subject: [PATCH 20/31] drm/amdgpu/gmc: add new callback to lookup vmid to
> pasid mapping
>
> Look up the mapping so we know which vmid to flush.
>
> Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
> ---
> drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 7 ++++++-
> drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c | 6 ++++--
> drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c | 6 ++++--
> drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c | 6 ++++--
> drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c | 1 +
> drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c | 6 ++++--
> 6 files changed, 23 insertions(+), 9 deletions(-)
>
> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
> b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
> index 204cf1c360896..27785b1b38ccd 100644
> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
> @@ -171,6 +171,10 @@ struct amdgpu_gmc_funcs {
> /* Change the VMID -> PASID mapping */
> void (*emit_pasid_mapping)(struct amdgpu_ring *ring, unsigned vmid,
> unsigned pasid);
> + /* look up the vmids for the pasid */
> + bool (*get_vmid_pasid_mapping_info)(struct amdgpu_device *adev,
> + uint8_t vmid, uint8_t inst,
> + uint16_t *p_pasid);
> /* enable/disable PRT support */
> void (*set_prt)(struct amdgpu_device *adev, bool enable);
> /* get the pde for a given mc addr */
> @@ -384,7 +388,8 @@ struct amdgpu_gmc {
> #define amdgpu_gmc_emit_flush_gpu_tlb(r, vmid, addr) (r)->adev-
> >gmc.gmc_funcs->emit_flush_gpu_tlb((r), (vmid), (addr)) #define
> amdgpu_gmc_emit_pasid_mapping(r, vmid, pasid) (r)->adev->gmc.gmc_funcs-
> >emit_pasid_mapping((r), (vmid), (pasid)) #define amdgpu_gmc_get_vm_pde(adev,
> level, dst, flags) (adev)->gmc.gmc_funcs->get_vm_pde((adev), (level), (dst),
> (flags)) -#define amdgpu_gmc_get_vm_pte(adev, vm, bo, vm_flags, pte_flags) \
> +#define amdgpu_gmc_get_vmid_pasid_mapping_info(adev, v, i, p) (adev)-
> >gmc.gmc_funcs->get_vmid_pasid_mapping_info((a), (v), (i), (p))
Spelling error?
It should be get_vmid_pasid_mapping_info((adev), (v), (i), (p)) // a -> adev
Thanks
Jesse
> +#define amdgpu_gmc_get_vm_pte(adev, vm, bo, vm_flags, pte_flags) \
> ((adev)->gmc.gmc_funcs->get_vm_pte((adev), (vm), (bo), (vm_flags), \
> (pte_flags)))
> #define amdgpu_gmc_override_vm_pte_flags(adev, vm, addr, pte_flags) \
> diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
> b/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
> index 23d9fe995a2b5..0028639448956 100644
> --- a/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
> +++ b/drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c
> @@ -204,7 +204,8 @@ static bool gmc_v10_0_use_invalidate_semaphore(struct
> amdgpu_device *adev,
>
> static bool gmc_v10_0_get_atc_vmid_pasid_mapping_info(
> struct amdgpu_device *adev,
> - uint8_t vmid, uint16_t *p_pasid)
> + uint8_t vmid, uint8_t inst,
> + uint16_t *p_pasid)
> {
> uint32_t value;
>
> @@ -346,7 +347,7 @@ static void gmc_v10_0_flush_gpu_tlb_pasid(struct
> amdgpu_device *adev,
> for (vmid = 1; vmid < AMDGPU_NUM_VMID; vmid++) {
> bool valid;
>
> - valid = gmc_v10_0_get_atc_vmid_pasid_mapping_info(adev, vmid,
> + valid = gmc_v10_0_get_atc_vmid_pasid_mapping_info(adev, vmid,
> 0,
> &queried);
> if (!valid || queried != pasid)
> continue;
> @@ -555,6 +556,7 @@ static const struct amdgpu_gmc_funcs
> gmc_v10_0_gmc_funcs = {
> .flush_gpu_tlb_pasid = gmc_v10_0_flush_gpu_tlb_pasid,
> .emit_flush_gpu_tlb = gmc_v10_0_emit_flush_gpu_tlb,
> .emit_pasid_mapping = gmc_v10_0_emit_pasid_mapping,
> + .get_vmid_pasid_mapping_info =
> +gmc_v10_0_get_atc_vmid_pasid_mapping_info,
> .get_vm_pde = gmc_v10_0_get_vm_pde,
> .get_vm_pte = gmc_v10_0_get_vm_pte,
> .get_vbios_fb_size = gmc_v10_0_get_vbios_fb_size, diff --git
> a/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
> b/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
> index 098e1340554c5..b34bd7881726d 100644
> --- a/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
> +++ b/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
> @@ -200,7 +200,8 @@ static bool gmc_v11_0_use_invalidate_semaphore(struct
> amdgpu_device *adev,
>
> static bool gmc_v11_0_get_vmid_pasid_mapping_info(
> struct amdgpu_device *adev,
> - uint8_t vmid, uint16_t *p_pasid)
> + uint8_t vmid, uint8_t inst,
> + uint16_t *p_pasid)
> {
> *p_pasid = RREG32(SOC15_REG_OFFSET(OSSSYS, 0,
> regIH_VMID_0_LUT) + vmid) & 0xffff;
>
> @@ -338,7 +339,7 @@ static void gmc_v11_0_flush_gpu_tlb_pasid(struct
> amdgpu_device *adev,
> for (vmid = 1; vmid < 16; vmid++) {
> bool valid;
>
> - valid = gmc_v11_0_get_vmid_pasid_mapping_info(adev, vmid,
> + valid = gmc_v11_0_get_vmid_pasid_mapping_info(adev, vmid, 0,
> &queried);
> if (!valid || queried != pasid)
> continue;
> @@ -546,6 +547,7 @@ static const struct amdgpu_gmc_funcs
> gmc_v11_0_gmc_funcs = {
> .flush_gpu_tlb_pasid = gmc_v11_0_flush_gpu_tlb_pasid,
> .emit_flush_gpu_tlb = gmc_v11_0_emit_flush_gpu_tlb,
> .emit_pasid_mapping = gmc_v11_0_emit_pasid_mapping,
> + .get_vmid_pasid_mapping_info =
> gmc_v11_0_get_vmid_pasid_mapping_info,
> .get_vm_pde = gmc_v11_0_get_vm_pde,
> .get_vm_pte = gmc_v11_0_get_vm_pte,
> .get_vbios_fb_size = gmc_v11_0_get_vbios_fb_size, diff --git
> a/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
> b/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
> index cffc818880d12..9179dc0787f13 100644
> --- a/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
> +++ b/drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c
> @@ -196,7 +196,8 @@ static bool gmc_v12_0_use_invalidate_semaphore(struct
> amdgpu_device *adev,
>
> static bool gmc_v12_0_get_vmid_pasid_mapping_info(
> struct amdgpu_device *adev,
> - uint8_t vmid, uint16_t *p_pasid)
> + uint8_t vmid, uint8_t inst,
> + uint16_t *p_pasid)
> {
> *p_pasid = RREG32(SOC15_REG_OFFSET(OSSSYS, 0,
> regIH_VMID_0_LUT) + vmid) & 0xffff;
>
> @@ -374,7 +375,7 @@ static void gmc_v12_0_flush_gpu_tlb_pasid(struct
> amdgpu_device *adev,
> for (vmid = 1; vmid < 16; vmid++) {
> bool valid;
>
> - valid = gmc_v12_0_get_vmid_pasid_mapping_info(adev, vmid,
> + valid = gmc_v12_0_get_vmid_pasid_mapping_info(adev, vmid, 0,
> &queried);
> if (!valid || queried != pasid)
> continue;
> @@ -581,6 +582,7 @@ static const struct amdgpu_gmc_funcs
> gmc_v12_0_gmc_funcs = {
> .flush_gpu_tlb_pasid = gmc_v12_0_flush_gpu_tlb_pasid,
> .emit_flush_gpu_tlb = gmc_v12_0_emit_flush_gpu_tlb,
> .emit_pasid_mapping = gmc_v12_0_emit_pasid_mapping,
> + .get_vmid_pasid_mapping_info =
> gmc_v12_0_get_vmid_pasid_mapping_info,
> .get_vm_pde = gmc_v12_0_get_vm_pde,
> .get_vm_pte = gmc_v12_0_get_vm_pte,
> .get_vbios_fb_size = gmc_v12_0_get_vbios_fb_size, diff --git
> a/drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c
> b/drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c
> index 6c0d2689cc05d..3fa1ec3dca273 100644
> --- a/drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c
> +++ b/drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c
> @@ -668,6 +668,7 @@ static const struct amdgpu_gmc_funcs
> gmc_v12_1_gmc_funcs = {
> .flush_gpu_tlb_pasid = gmc_v12_1_flush_gpu_tlb_pasid,
> .emit_flush_gpu_tlb = gmc_v12_1_emit_flush_gpu_tlb,
> .emit_pasid_mapping = gmc_v12_1_emit_pasid_mapping,
> + .get_vmid_pasid_mapping_info =
> gmc_v12_1_get_vmid_pasid_mapping_info,
> .get_vm_pde = gmc_v12_1_get_vm_pde,
> .get_vm_pte = gmc_v12_1_get_vm_pte,
> .query_mem_partition_mode = &amdgpu_gmc_query_memory_partition,
> diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
> b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
> index a91d0ddb6e729..0019706b0b2a9 100644
> --- a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
> +++ b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
> @@ -727,7 +727,8 @@ static bool gmc_v9_0_use_invalidate_semaphore(struct
> amdgpu_device *adev, }
>
> static bool gmc_v9_0_get_atc_vmid_pasid_mapping_info(struct amdgpu_device
> *adev,
> - uint8_t vmid, uint16_t *p_pasid)
> + uint8_t vmid, uint8_t inst,
> + uint16_t *p_pasid)
> {
> uint32_t value;
>
> @@ -889,7 +890,7 @@ static void gmc_v9_0_flush_gpu_tlb_pasid(struct
> amdgpu_device *adev,
> bool valid;
>
> valid = gmc_v9_0_get_atc_vmid_pasid_mapping_info(adev, vmid,
> - &queried);
> + inst, &queried);
> if (!valid || queried != pasid)
> continue;
>
> @@ -1302,6 +1303,7 @@ static const struct amdgpu_gmc_funcs
> gmc_v9_0_gmc_funcs = {
> .flush_gpu_tlb_pasid = gmc_v9_0_flush_gpu_tlb_pasid,
> .emit_flush_gpu_tlb = gmc_v9_0_emit_flush_gpu_tlb,
> .emit_pasid_mapping = gmc_v9_0_emit_pasid_mapping,
> + .get_vmid_pasid_mapping_info =
> +gmc_v9_0_get_atc_vmid_pasid_mapping_info,
> .get_vm_pde = gmc_v9_0_get_vm_pde,
> .get_vm_pte = gmc_v9_0_get_vm_pte,
> .override_vm_pte_flags = gmc_v9_0_override_vm_pte_flags,
> --
> 2.55.0
^ permalink raw reply [flat|nested] 43+ messages in thread
* Re: [PATCH V2 00/31] Rework GPU TLB invalidation
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
` (30 preceding siblings ...)
2026-09-01 20:10 ` [PATCH 31/31] drm/amdgpu/gmc: add helpers for various tlb inv functions Alex Deucher
@ 2026-09-02 7:07 ` Christian König
2026-09-02 13:06 ` Alex Deucher
31 siblings, 1 reply; 43+ messages in thread
From: Christian König @ 2026-09-02 7:07 UTC (permalink / raw)
To: Alex Deucher, amd-gfx
On 9/1/26 22:10, Alex Deucher wrote:
> GMC 9-12 use KIQ or MES for TLB invalidations to avoid using MMIO which would
> require disallowing GFXOFF. KIQ and MES are management queues however and if
> they hang, they cannot be recovered by a queue reset since they are the
> mechanisms which handle queue resets. Since using KIQ or MES will exit GFXOFF
> anyway, explicitly disallow it on the MMIO path and use that. Next, switch to
> using SDMA for TLB invalidations. SDMA 4.4.x and newer have special packets
> specifically for this purpose.
That packet was just introduced to work around SRIOV issues and proved to cause stability issues as well.
Since Arun now found that on basically all Navi generations the SDMA can hang while doing a TLB invalidation I have to clearly NAK this approach.
In the long run we should do the TLB invalidations on newer HW completely with the MES, the SDMA was always just a workaround we should not push forward.
Regards,
Christian.
> If SDMA hangs while doing the invalidation for
> some reason, it's easier to reset the SDMA queue than KIQ or MES. Finally,
> most of the TLB invalidation code between GMC 9 through 12 was identical, so
> move it to common GMC helpers and remove the IP specific code. If the SMDA
> and MMIO pathes prove to be stable, the KIQ pathes can be removed in the future
> to further simplify things. SDMA 4.x could also be updated to support PASID
> invalidation via SDMA using either the new packet (SDMA 4.4.x) or via
> REG_WRITE/REG_WAIT packets (SDMA 4.0.x).
>
> Code is available on this branch as well:
> https://gitlab.freedesktop.org/agd5f/linux/-/commits/tlb_inv_rework?ref_type=heads
>
> V2:
> - Add missing hub callbacks in gmc9 hubs
>
> Alex Deucher (31):
> drm/amdgpu/gmc9: disallow gfxoff around TLB flushes
> drm/amdgpu/gmc10: disallow gfxoff around TLB flushes
> drm/amdgpu/gmc11: disallow gfxoff around TLB flushes
> drm/amdgpu/gmc12: disallow gfxoff around TLB flushes
> drm/amdgpu/gmc9: set vmhub funcs for gfxhub and mmhub
> drm/amdgpu: add a gmc flag for using MMIO for TLB flush
> drm/amdgpu/gmc9: use MMIO for TLB flushes
> drm/amdgpu/gmc10: use MMIO for TLB flushes
> drm/amdgpu/gmc11: use MMIO for TLB flushes
> drm/amdgpu/gmc12: use MMIO for TLB flushes
> drm/amdgpu: add a buffer funcs callback for TLB invalidation
> drm/amdgpu/sdma5.0: add tlb invalidation buffer func callback
> drm/amdgpu/sdma5.2: add tlb invalidation buffer func callback
> drm/amdgpu/sdma6: add tlb invalidation buffer func callback
> drm/amdgpu/sdma7: add tlb invalidation buffer func callback
> drm/amdgpu: simplify amdgpu_gmc_flush_gpu_tlb()
> drm/amdgpu: add tlb invalidation method enum
> drm/amdgpu: plumb tlb inv method in amdgpu_gmc_flush_gpu_tlb_gart()
> drm/amdgpu: uplevel reset check in amdgpu_gmc_flush_gpu_tlb_gart()
> drm/amdgpu/gmc: add new callback to lookup vmid to pasid mapping
> drm/amdgpu: add a gmc callback for the inv semaphore
> drm/amdgpu/gmc: rework pasid flushing
> drm/amdgpu/gmc9: use SDMA for gart TLB invalidation
> drm/amdgpu/gmc10: use SDMA for gart TLB invalidation
> drm/amdgpu/gmc11: use SDMA for gart TLB invalidation
> drm/amdgpu/gmc12: use SDMA for gart TLB invalidation
> drm/amdgpu/gmc10: use SDMA for pasid TLB invalidation
> drm/amdgpu/gmc11: use SDMA for pasid TLB invalidation
> drm/amdgpu/gmc12: use MES or SDMA for pasid TLB invalidation
> drm/amdgpu/gmc12: drop MES tlb inv in gmc callbacks
> drm/amdgpu/gmc: add helpers for various tlb inv functions
>
> drivers/gpu/drm/amd/amdgpu/amdgpu_gart.c | 2 +-
> drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c | 486 +++++++++++++++++++----
> drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 33 +-
> drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h | 18 +
> drivers/gpu/drm/amd/amdgpu/gfx_v10_0.c | 2 +
> drivers/gpu/drm/amd/amdgpu/gfxhub_v1_0.c | 33 ++
> drivers/gpu/drm/amd/amdgpu/gfxhub_v1_2.c | 33 ++
> drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c | 202 +---------
> drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c | 207 +---------
> drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c | 250 ++----------
> drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c | 225 +----------
> drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c | 135 ++-----
> drivers/gpu/drm/amd/amdgpu/mes_v12_0.c | 4 +
> drivers/gpu/drm/amd/amdgpu/mes_v12_1.c | 4 +
> drivers/gpu/drm/amd/amdgpu/mmhub_v1_0.c | 33 ++
> drivers/gpu/drm/amd/amdgpu/mmhub_v1_7.c | 32 ++
> drivers/gpu/drm/amd/amdgpu/mmhub_v1_8.c | 32 ++
> drivers/gpu/drm/amd/amdgpu/mmhub_v9_4.c | 32 ++
> drivers/gpu/drm/amd/amdgpu/sdma_v5_0.c | 49 +++
> drivers/gpu/drm/amd/amdgpu/sdma_v5_2.c | 49 +++
> drivers/gpu/drm/amd/amdgpu/sdma_v6_0.c | 49 +++
> drivers/gpu/drm/amd/amdgpu/sdma_v7_0.c | 48 +++
> 22 files changed, 943 insertions(+), 1015 deletions(-)
>
^ permalink raw reply [flat|nested] 43+ messages in thread
* Re: [PATCH 01/31] drm/amdgpu/gmc9: disallow gfxoff around TLB flushes
2026-09-01 20:10 ` [PATCH 01/31] drm/amdgpu/gmc9: disallow gfxoff around TLB flushes Alex Deucher
@ 2026-09-02 7:10 ` Christian König
2026-09-02 13:08 ` Alex Deucher
0 siblings, 1 reply; 43+ messages in thread
From: Christian König @ 2026-09-02 7:10 UTC (permalink / raw)
To: Alex Deucher, amd-gfx
On 9/1/26 22:10, Alex Deucher wrote:
> We need to disallow gfxoff if we touch GC MMIO registers.
> At the moment we use KIQ or MES for TLB flushes so
> no intended functional change.
One reason to use the KIQ for TLB invalidation was to avoid enabling/disabling GFXOFF all the time.
So this change here might work around GFXOFF issues, but that only hides the problems we have with GFXOFF and doesn't fix them.
The same is true for all other generations.
Regards,
Christian.
>
> Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
> ---
> drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c | 7 +++++++
> 1 file changed, 7 insertions(+)
>
> diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
> index b46b87291c512..80f1cf1f21736 100644
> --- a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
> +++ b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
> @@ -808,6 +808,10 @@ static void gmc_v9_0_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
> return;
> }
>
> + /* disabllow gfxoff when we invalidate */
> + if (vmhub < AMDGPU_MMHUB0(0))
> + amdgpu_gfx_off_ctrl(adev, false);
> +
> /* This path is needed before KIQ/MES/GFXOFF are set up */
> spin_lock(&adev->gmc.invalidate_lock);
>
> @@ -873,6 +877,9 @@ static void gmc_v9_0_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
>
> spin_unlock(&adev->gmc.invalidate_lock);
>
> + if (vmhub < AMDGPU_MMHUB0(0))
> + amdgpu_gfx_off_ctrl(adev, true);
> +
> if (j < adev->usec_timeout)
> return;
>
^ permalink raw reply [flat|nested] 43+ messages in thread
* Re: [PATCH V2 00/31] Rework GPU TLB invalidation
2026-09-02 7:07 ` [PATCH V2 00/31] Rework GPU TLB invalidation Christian König
@ 2026-09-02 13:06 ` Alex Deucher
2026-09-02 13:20 ` Christian König
0 siblings, 1 reply; 43+ messages in thread
From: Alex Deucher @ 2026-09-02 13:06 UTC (permalink / raw)
To: Christian König; +Cc: Alex Deucher, amd-gfx
On Wed, Sep 2, 2026 at 3:25 AM Christian König <christian.koenig@amd.com> wrote:
>
> On 9/1/26 22:10, Alex Deucher wrote:
> > GMC 9-12 use KIQ or MES for TLB invalidations to avoid using MMIO which would
> > require disallowing GFXOFF. KIQ and MES are management queues however and if
> > they hang, they cannot be recovered by a queue reset since they are the
> > mechanisms which handle queue resets. Since using KIQ or MES will exit GFXOFF
> > anyway, explicitly disallow it on the MMIO path and use that. Next, switch to
> > using SDMA for TLB invalidations. SDMA 4.4.x and newer have special packets
> > specifically for this purpose.
>
> That packet was just introduced to work around SRIOV issues and proved to cause stability issues as well.
>
It was introduced because of gfxoff and having to toggle it to use
MMIO. WIndows uses SDMA pretty much exclusively for TLB
invalidations. It's also what the memory hubs team recommends.
> Since Arun now found that on basically all Navi generations the SDMA can hang while doing a TLB invalidation I have to clearly NAK this approach.
>
> In the long run we should do the TLB invalidations on newer HW completely with the MES, the SDMA was always just a workaround we should not push forward.
>
Well KIQ and MES also seem to be problematic as well. See the thread
from Denis about KIQ and his renoir system. There are also various
reports of MES timeouts doing TLB invalidations.
Alex
> Regards,
> Christian.
>
> > If SDMA hangs while doing the invalidation for
> > some reason, it's easier to reset the SDMA queue than KIQ or MES. Finally,
> > most of the TLB invalidation code between GMC 9 through 12 was identical, so
> > move it to common GMC helpers and remove the IP specific code. If the SMDA
> > and MMIO pathes prove to be stable, the KIQ pathes can be removed in the future
> > to further simplify things. SDMA 4.x could also be updated to support PASID
> > invalidation via SDMA using either the new packet (SDMA 4.4.x) or via
> > REG_WRITE/REG_WAIT packets (SDMA 4.0.x).
> >
> > Code is available on this branch as well:
> > https://gitlab.freedesktop.org/agd5f/linux/-/commits/tlb_inv_rework?ref_type=heads
> >
> > V2:
> > - Add missing hub callbacks in gmc9 hubs
> >
> > Alex Deucher (31):
> > drm/amdgpu/gmc9: disallow gfxoff around TLB flushes
> > drm/amdgpu/gmc10: disallow gfxoff around TLB flushes
> > drm/amdgpu/gmc11: disallow gfxoff around TLB flushes
> > drm/amdgpu/gmc12: disallow gfxoff around TLB flushes
> > drm/amdgpu/gmc9: set vmhub funcs for gfxhub and mmhub
> > drm/amdgpu: add a gmc flag for using MMIO for TLB flush
> > drm/amdgpu/gmc9: use MMIO for TLB flushes
> > drm/amdgpu/gmc10: use MMIO for TLB flushes
> > drm/amdgpu/gmc11: use MMIO for TLB flushes
> > drm/amdgpu/gmc12: use MMIO for TLB flushes
> > drm/amdgpu: add a buffer funcs callback for TLB invalidation
> > drm/amdgpu/sdma5.0: add tlb invalidation buffer func callback
> > drm/amdgpu/sdma5.2: add tlb invalidation buffer func callback
> > drm/amdgpu/sdma6: add tlb invalidation buffer func callback
> > drm/amdgpu/sdma7: add tlb invalidation buffer func callback
> > drm/amdgpu: simplify amdgpu_gmc_flush_gpu_tlb()
> > drm/amdgpu: add tlb invalidation method enum
> > drm/amdgpu: plumb tlb inv method in amdgpu_gmc_flush_gpu_tlb_gart()
> > drm/amdgpu: uplevel reset check in amdgpu_gmc_flush_gpu_tlb_gart()
> > drm/amdgpu/gmc: add new callback to lookup vmid to pasid mapping
> > drm/amdgpu: add a gmc callback for the inv semaphore
> > drm/amdgpu/gmc: rework pasid flushing
> > drm/amdgpu/gmc9: use SDMA for gart TLB invalidation
> > drm/amdgpu/gmc10: use SDMA for gart TLB invalidation
> > drm/amdgpu/gmc11: use SDMA for gart TLB invalidation
> > drm/amdgpu/gmc12: use SDMA for gart TLB invalidation
> > drm/amdgpu/gmc10: use SDMA for pasid TLB invalidation
> > drm/amdgpu/gmc11: use SDMA for pasid TLB invalidation
> > drm/amdgpu/gmc12: use MES or SDMA for pasid TLB invalidation
> > drm/amdgpu/gmc12: drop MES tlb inv in gmc callbacks
> > drm/amdgpu/gmc: add helpers for various tlb inv functions
> >
> > drivers/gpu/drm/amd/amdgpu/amdgpu_gart.c | 2 +-
> > drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c | 486 +++++++++++++++++++----
> > drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 33 +-
> > drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h | 18 +
> > drivers/gpu/drm/amd/amdgpu/gfx_v10_0.c | 2 +
> > drivers/gpu/drm/amd/amdgpu/gfxhub_v1_0.c | 33 ++
> > drivers/gpu/drm/amd/amdgpu/gfxhub_v1_2.c | 33 ++
> > drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c | 202 +---------
> > drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c | 207 +---------
> > drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c | 250 ++----------
> > drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c | 225 +----------
> > drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c | 135 ++-----
> > drivers/gpu/drm/amd/amdgpu/mes_v12_0.c | 4 +
> > drivers/gpu/drm/amd/amdgpu/mes_v12_1.c | 4 +
> > drivers/gpu/drm/amd/amdgpu/mmhub_v1_0.c | 33 ++
> > drivers/gpu/drm/amd/amdgpu/mmhub_v1_7.c | 32 ++
> > drivers/gpu/drm/amd/amdgpu/mmhub_v1_8.c | 32 ++
> > drivers/gpu/drm/amd/amdgpu/mmhub_v9_4.c | 32 ++
> > drivers/gpu/drm/amd/amdgpu/sdma_v5_0.c | 49 +++
> > drivers/gpu/drm/amd/amdgpu/sdma_v5_2.c | 49 +++
> > drivers/gpu/drm/amd/amdgpu/sdma_v6_0.c | 49 +++
> > drivers/gpu/drm/amd/amdgpu/sdma_v7_0.c | 48 +++
> > 22 files changed, 943 insertions(+), 1015 deletions(-)
> >
>
^ permalink raw reply [flat|nested] 43+ messages in thread
* Re: [PATCH 01/31] drm/amdgpu/gmc9: disallow gfxoff around TLB flushes
2026-09-02 7:10 ` Christian König
@ 2026-09-02 13:08 ` Alex Deucher
2026-09-02 13:10 ` Alex Deucher
0 siblings, 1 reply; 43+ messages in thread
From: Alex Deucher @ 2026-09-02 13:08 UTC (permalink / raw)
To: Christian König; +Cc: Alex Deucher, amd-gfx
On Wed, Sep 2, 2026 at 3:35 AM Christian König <christian.koenig@amd.com> wrote:
>
> On 9/1/26 22:10, Alex Deucher wrote:
> > We need to disallow gfxoff if we touch GC MMIO registers.
> > At the moment we use KIQ or MES for TLB flushes so
> > no intended functional change.
>
> One reason to use the KIQ for TLB invalidation was to avoid enabling/disabling GFXOFF all the time.
>
> So this change here might work around GFXOFF issues, but that only hides the problems we have with GFXOFF and doesn't fix them.
>
> The same is true for all other generations.
This was meant to be a short term fix which could more easily be
backported on the way to enabling SDMA later in the series to work
around issues like KIQ dying when doing invalidations.
Alex
>
> Regards,
> Christian.
>
> >
> > Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
> > ---
> > drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c | 7 +++++++
> > 1 file changed, 7 insertions(+)
> >
> > diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
> > index b46b87291c512..80f1cf1f21736 100644
> > --- a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
> > +++ b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
> > @@ -808,6 +808,10 @@ static void gmc_v9_0_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
> > return;
> > }
> >
> > + /* disabllow gfxoff when we invalidate */
> > + if (vmhub < AMDGPU_MMHUB0(0))
> > + amdgpu_gfx_off_ctrl(adev, false);
> > +
> > /* This path is needed before KIQ/MES/GFXOFF are set up */
> > spin_lock(&adev->gmc.invalidate_lock);
> >
> > @@ -873,6 +877,9 @@ static void gmc_v9_0_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
> >
> > spin_unlock(&adev->gmc.invalidate_lock);
> >
> > + if (vmhub < AMDGPU_MMHUB0(0))
> > + amdgpu_gfx_off_ctrl(adev, true);
> > +
> > if (j < adev->usec_timeout)
> > return;
> >
>
^ permalink raw reply [flat|nested] 43+ messages in thread
* Re: [PATCH 01/31] drm/amdgpu/gmc9: disallow gfxoff around TLB flushes
2026-09-02 13:08 ` Alex Deucher
@ 2026-09-02 13:10 ` Alex Deucher
0 siblings, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-02 13:10 UTC (permalink / raw)
To: Christian König; +Cc: Alex Deucher, amd-gfx
On Wed, Sep 2, 2026 at 9:08 AM Alex Deucher <alexdeucher@gmail.com> wrote:
>
> On Wed, Sep 2, 2026 at 3:35 AM Christian König <christian.koenig@amd.com> wrote:
> >
> > On 9/1/26 22:10, Alex Deucher wrote:
> > > We need to disallow gfxoff if we touch GC MMIO registers.
> > > At the moment we use KIQ or MES for TLB flushes so
> > > no intended functional change.
> >
> > One reason to use the KIQ for TLB invalidation was to avoid enabling/disabling GFXOFF all the time.
> >
> > So this change here might work around GFXOFF issues, but that only hides the problems we have with GFXOFF and doesn't fix them.
> >
> > The same is true for all other generations.
>
> This was meant to be a short term fix which could more easily be
> backported on the way to enabling SDMA later in the series to work
> around issues like KIQ dying when doing invalidations.
>
Regardless, we shouldn't be touching these registers while gfxoff is
allowed so even with these patches applied, we'd still be using the
KIQ and MES path if those engines are available. This just makes sure
the fallback case is safe.
Alex
> Alex
>
> >
> > Regards,
> > Christian.
> >
> > >
> > > Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
> > > ---
> > > drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c | 7 +++++++
> > > 1 file changed, 7 insertions(+)
> > >
> > > diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
> > > index b46b87291c512..80f1cf1f21736 100644
> > > --- a/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
> > > +++ b/drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c
> > > @@ -808,6 +808,10 @@ static void gmc_v9_0_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
> > > return;
> > > }
> > >
> > > + /* disabllow gfxoff when we invalidate */
> > > + if (vmhub < AMDGPU_MMHUB0(0))
> > > + amdgpu_gfx_off_ctrl(adev, false);
> > > +
> > > /* This path is needed before KIQ/MES/GFXOFF are set up */
> > > spin_lock(&adev->gmc.invalidate_lock);
> > >
> > > @@ -873,6 +877,9 @@ static void gmc_v9_0_flush_gpu_tlb(struct amdgpu_device *adev, uint32_t vmid,
> > >
> > > spin_unlock(&adev->gmc.invalidate_lock);
> > >
> > > + if (vmhub < AMDGPU_MMHUB0(0))
> > > + amdgpu_gfx_off_ctrl(adev, true);
> > > +
> > > if (j < adev->usec_timeout)
> > > return;
> > >
> >
^ permalink raw reply [flat|nested] 43+ messages in thread
* Re: [PATCH V2 00/31] Rework GPU TLB invalidation
2026-09-02 13:06 ` Alex Deucher
@ 2026-09-02 13:20 ` Christian König
2026-09-02 13:25 ` Alex Deucher
0 siblings, 1 reply; 43+ messages in thread
From: Christian König @ 2026-09-02 13:20 UTC (permalink / raw)
To: Alex Deucher; +Cc: Alex Deucher, amd-gfx
On 9/2/26 15:06, Alex Deucher wrote:
> On Wed, Sep 2, 2026 at 3:25 AM Christian König <christian.koenig@amd.com> wrote:
>>
>> On 9/1/26 22:10, Alex Deucher wrote:
>>> GMC 9-12 use KIQ or MES for TLB invalidations to avoid using MMIO which would
>>> require disallowing GFXOFF. KIQ and MES are management queues however and if
>>> they hang, they cannot be recovered by a queue reset since they are the
>>> mechanisms which handle queue resets. Since using KIQ or MES will exit GFXOFF
>>> anyway, explicitly disallow it on the MMIO path and use that. Next, switch to
>>> using SDMA for TLB invalidations. SDMA 4.4.x and newer have special packets
>>> specifically for this purpose.
>>
>> That packet was just introduced to work around SRIOV issues and proved to cause stability issues as well.
>>
>
> It was introduced because of gfxoff and having to toggle it to use
> MMIO.
No, that was completely unrelated to GFXOFF. This also applies to HW where the SDMA is not even in the GFX domain.
The problem was the SRIOV could interrupt the SDMA while it waited for the ACK and when the VF was scheduled in again the ACK bit was resetted.
So a single packet was introduced which couldn't be interrupted by the hypervisor.
> WIndows uses SDMA pretty much exclusively for TLB
> invalidations. It's also what the memory hubs team recommends.
That is rather interesting.
>> Since Arun now found that on basically all Navi generations the SDMA can hang while doing a TLB invalidation I have to clearly NAK this approach.
>>
>> In the long run we should do the TLB invalidations on newer HW completely with the MES, the SDMA was always just a workaround we should not push forward.
>>
>
> Well KIQ and MES also seem to be problematic as well. See the thread
> from Denis about KIQ and his renoir system. There are also various
> reports of MES timeouts doing TLB invalidations.
The MES is the only instance which knows the process to VMID mapping for user queues.
The hack to read out the PASID->VMID mapping register directly from the IH is not really something we can use in the future.
So the plan was to move all of that into MES in the near term. IIRC the MES even added a new packet for that.
Regards,
Christian.
>
> Alex
>
>> Regards,
>> Christian.
>>
>>> If SDMA hangs while doing the invalidation for
>>> some reason, it's easier to reset the SDMA queue than KIQ or MES. Finally,
>>> most of the TLB invalidation code between GMC 9 through 12 was identical, so
>>> move it to common GMC helpers and remove the IP specific code. If the SMDA
>>> and MMIO pathes prove to be stable, the KIQ pathes can be removed in the future
>>> to further simplify things. SDMA 4.x could also be updated to support PASID
>>> invalidation via SDMA using either the new packet (SDMA 4.4.x) or via
>>> REG_WRITE/REG_WAIT packets (SDMA 4.0.x).
>>>
>>> Code is available on this branch as well:
>>> https://gitlab.freedesktop.org/agd5f/linux/-/commits/tlb_inv_rework?ref_type=heads
>>>
>>> V2:
>>> - Add missing hub callbacks in gmc9 hubs
>>>
>>> Alex Deucher (31):
>>> drm/amdgpu/gmc9: disallow gfxoff around TLB flushes
>>> drm/amdgpu/gmc10: disallow gfxoff around TLB flushes
>>> drm/amdgpu/gmc11: disallow gfxoff around TLB flushes
>>> drm/amdgpu/gmc12: disallow gfxoff around TLB flushes
>>> drm/amdgpu/gmc9: set vmhub funcs for gfxhub and mmhub
>>> drm/amdgpu: add a gmc flag for using MMIO for TLB flush
>>> drm/amdgpu/gmc9: use MMIO for TLB flushes
>>> drm/amdgpu/gmc10: use MMIO for TLB flushes
>>> drm/amdgpu/gmc11: use MMIO for TLB flushes
>>> drm/amdgpu/gmc12: use MMIO for TLB flushes
>>> drm/amdgpu: add a buffer funcs callback for TLB invalidation
>>> drm/amdgpu/sdma5.0: add tlb invalidation buffer func callback
>>> drm/amdgpu/sdma5.2: add tlb invalidation buffer func callback
>>> drm/amdgpu/sdma6: add tlb invalidation buffer func callback
>>> drm/amdgpu/sdma7: add tlb invalidation buffer func callback
>>> drm/amdgpu: simplify amdgpu_gmc_flush_gpu_tlb()
>>> drm/amdgpu: add tlb invalidation method enum
>>> drm/amdgpu: plumb tlb inv method in amdgpu_gmc_flush_gpu_tlb_gart()
>>> drm/amdgpu: uplevel reset check in amdgpu_gmc_flush_gpu_tlb_gart()
>>> drm/amdgpu/gmc: add new callback to lookup vmid to pasid mapping
>>> drm/amdgpu: add a gmc callback for the inv semaphore
>>> drm/amdgpu/gmc: rework pasid flushing
>>> drm/amdgpu/gmc9: use SDMA for gart TLB invalidation
>>> drm/amdgpu/gmc10: use SDMA for gart TLB invalidation
>>> drm/amdgpu/gmc11: use SDMA for gart TLB invalidation
>>> drm/amdgpu/gmc12: use SDMA for gart TLB invalidation
>>> drm/amdgpu/gmc10: use SDMA for pasid TLB invalidation
>>> drm/amdgpu/gmc11: use SDMA for pasid TLB invalidation
>>> drm/amdgpu/gmc12: use MES or SDMA for pasid TLB invalidation
>>> drm/amdgpu/gmc12: drop MES tlb inv in gmc callbacks
>>> drm/amdgpu/gmc: add helpers for various tlb inv functions
>>>
>>> drivers/gpu/drm/amd/amdgpu/amdgpu_gart.c | 2 +-
>>> drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c | 486 +++++++++++++++++++----
>>> drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 33 +-
>>> drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h | 18 +
>>> drivers/gpu/drm/amd/amdgpu/gfx_v10_0.c | 2 +
>>> drivers/gpu/drm/amd/amdgpu/gfxhub_v1_0.c | 33 ++
>>> drivers/gpu/drm/amd/amdgpu/gfxhub_v1_2.c | 33 ++
>>> drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c | 202 +---------
>>> drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c | 207 +---------
>>> drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c | 250 ++----------
>>> drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c | 225 +----------
>>> drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c | 135 ++-----
>>> drivers/gpu/drm/amd/amdgpu/mes_v12_0.c | 4 +
>>> drivers/gpu/drm/amd/amdgpu/mes_v12_1.c | 4 +
>>> drivers/gpu/drm/amd/amdgpu/mmhub_v1_0.c | 33 ++
>>> drivers/gpu/drm/amd/amdgpu/mmhub_v1_7.c | 32 ++
>>> drivers/gpu/drm/amd/amdgpu/mmhub_v1_8.c | 32 ++
>>> drivers/gpu/drm/amd/amdgpu/mmhub_v9_4.c | 32 ++
>>> drivers/gpu/drm/amd/amdgpu/sdma_v5_0.c | 49 +++
>>> drivers/gpu/drm/amd/amdgpu/sdma_v5_2.c | 49 +++
>>> drivers/gpu/drm/amd/amdgpu/sdma_v6_0.c | 49 +++
>>> drivers/gpu/drm/amd/amdgpu/sdma_v7_0.c | 48 +++
>>> 22 files changed, 943 insertions(+), 1015 deletions(-)
>>>
>>
^ permalink raw reply [flat|nested] 43+ messages in thread
* Re: [PATCH V2 00/31] Rework GPU TLB invalidation
2026-09-02 13:20 ` Christian König
@ 2026-09-02 13:25 ` Alex Deucher
2026-09-02 13:38 ` Alex Deucher
2026-09-02 13:57 ` Christian König
0 siblings, 2 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-02 13:25 UTC (permalink / raw)
To: Christian König; +Cc: Alex Deucher, amd-gfx
On Wed, Sep 2, 2026 at 9:20 AM Christian König <christian.koenig@amd.com> wrote:
>
> On 9/2/26 15:06, Alex Deucher wrote:
> > On Wed, Sep 2, 2026 at 3:25 AM Christian König <christian.koenig@amd.com> wrote:
> >>
> >> On 9/1/26 22:10, Alex Deucher wrote:
> >>> GMC 9-12 use KIQ or MES for TLB invalidations to avoid using MMIO which would
> >>> require disallowing GFXOFF. KIQ and MES are management queues however and if
> >>> they hang, they cannot be recovered by a queue reset since they are the
> >>> mechanisms which handle queue resets. Since using KIQ or MES will exit GFXOFF
> >>> anyway, explicitly disallow it on the MMIO path and use that. Next, switch to
> >>> using SDMA for TLB invalidations. SDMA 4.4.x and newer have special packets
> >>> specifically for this purpose.
> >>
> >> That packet was just introduced to work around SRIOV issues and proved to cause stability issues as well.
> >>
> >
> > It was introduced because of gfxoff and having to toggle it to use
> > MMIO.
>
> No, that was completely unrelated to GFXOFF. This also applies to HW where the SDMA is not even in the GFX domain.
>
> The problem was the SRIOV could interrupt the SDMA while it waited for the ACK and when the VF was scheduled in again the ACK bit was resetted.
>
> So a single packet was introduced which couldn't be interrupted by the hypervisor.
>
> > WIndows uses SDMA pretty much exclusively for TLB
> > invalidations. It's also what the memory hubs team recommends.
>
> That is rather interesting.
> >> Since Arun now found that on basically all Navi generations the SDMA can hang while doing a TLB invalidation I have to clearly NAK this approach.
> >>
> >> In the long run we should do the TLB invalidations on newer HW completely with the MES, the SDMA was always just a workaround we should not push forward.
> >>
> >
> > Well KIQ and MES also seem to be problematic as well. See the thread
> > from Denis about KIQ and his renoir system. There are also various
> > reports of MES timeouts doing TLB invalidations.
>
> The MES is the only instance which knows the process to VMID mapping for user queues.
>
> The hack to read out the PASID->VMID mapping register directly from the IH is not really something we can use in the future.
>
> So the plan was to move all of that into MES in the near term. IIRC the MES even added a new packet for that.
Sure, this patch set retains the use of MES for pasid invalidation
where it's supported (navi4x and newer with new enough firmware). We
could try and get the new packet added to gfx11 MES as well, but then
there is still gfx9 and 10 which still use KIQ.
Alex
>
> Regards,
> Christian.
>
> >
> > Alex
> >
> >> Regards,
> >> Christian.
> >>
> >>> If SDMA hangs while doing the invalidation for
> >>> some reason, it's easier to reset the SDMA queue than KIQ or MES. Finally,
> >>> most of the TLB invalidation code between GMC 9 through 12 was identical, so
> >>> move it to common GMC helpers and remove the IP specific code. If the SMDA
> >>> and MMIO pathes prove to be stable, the KIQ pathes can be removed in the future
> >>> to further simplify things. SDMA 4.x could also be updated to support PASID
> >>> invalidation via SDMA using either the new packet (SDMA 4.4.x) or via
> >>> REG_WRITE/REG_WAIT packets (SDMA 4.0.x).
> >>>
> >>> Code is available on this branch as well:
> >>> https://gitlab.freedesktop.org/agd5f/linux/-/commits/tlb_inv_rework?ref_type=heads
> >>>
> >>> V2:
> >>> - Add missing hub callbacks in gmc9 hubs
> >>>
> >>> Alex Deucher (31):
> >>> drm/amdgpu/gmc9: disallow gfxoff around TLB flushes
> >>> drm/amdgpu/gmc10: disallow gfxoff around TLB flushes
> >>> drm/amdgpu/gmc11: disallow gfxoff around TLB flushes
> >>> drm/amdgpu/gmc12: disallow gfxoff around TLB flushes
> >>> drm/amdgpu/gmc9: set vmhub funcs for gfxhub and mmhub
> >>> drm/amdgpu: add a gmc flag for using MMIO for TLB flush
> >>> drm/amdgpu/gmc9: use MMIO for TLB flushes
> >>> drm/amdgpu/gmc10: use MMIO for TLB flushes
> >>> drm/amdgpu/gmc11: use MMIO for TLB flushes
> >>> drm/amdgpu/gmc12: use MMIO for TLB flushes
> >>> drm/amdgpu: add a buffer funcs callback for TLB invalidation
> >>> drm/amdgpu/sdma5.0: add tlb invalidation buffer func callback
> >>> drm/amdgpu/sdma5.2: add tlb invalidation buffer func callback
> >>> drm/amdgpu/sdma6: add tlb invalidation buffer func callback
> >>> drm/amdgpu/sdma7: add tlb invalidation buffer func callback
> >>> drm/amdgpu: simplify amdgpu_gmc_flush_gpu_tlb()
> >>> drm/amdgpu: add tlb invalidation method enum
> >>> drm/amdgpu: plumb tlb inv method in amdgpu_gmc_flush_gpu_tlb_gart()
> >>> drm/amdgpu: uplevel reset check in amdgpu_gmc_flush_gpu_tlb_gart()
> >>> drm/amdgpu/gmc: add new callback to lookup vmid to pasid mapping
> >>> drm/amdgpu: add a gmc callback for the inv semaphore
> >>> drm/amdgpu/gmc: rework pasid flushing
> >>> drm/amdgpu/gmc9: use SDMA for gart TLB invalidation
> >>> drm/amdgpu/gmc10: use SDMA for gart TLB invalidation
> >>> drm/amdgpu/gmc11: use SDMA for gart TLB invalidation
> >>> drm/amdgpu/gmc12: use SDMA for gart TLB invalidation
> >>> drm/amdgpu/gmc10: use SDMA for pasid TLB invalidation
> >>> drm/amdgpu/gmc11: use SDMA for pasid TLB invalidation
> >>> drm/amdgpu/gmc12: use MES or SDMA for pasid TLB invalidation
> >>> drm/amdgpu/gmc12: drop MES tlb inv in gmc callbacks
> >>> drm/amdgpu/gmc: add helpers for various tlb inv functions
> >>>
> >>> drivers/gpu/drm/amd/amdgpu/amdgpu_gart.c | 2 +-
> >>> drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c | 486 +++++++++++++++++++----
> >>> drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 33 +-
> >>> drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h | 18 +
> >>> drivers/gpu/drm/amd/amdgpu/gfx_v10_0.c | 2 +
> >>> drivers/gpu/drm/amd/amdgpu/gfxhub_v1_0.c | 33 ++
> >>> drivers/gpu/drm/amd/amdgpu/gfxhub_v1_2.c | 33 ++
> >>> drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c | 202 +---------
> >>> drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c | 207 +---------
> >>> drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c | 250 ++----------
> >>> drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c | 225 +----------
> >>> drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c | 135 ++-----
> >>> drivers/gpu/drm/amd/amdgpu/mes_v12_0.c | 4 +
> >>> drivers/gpu/drm/amd/amdgpu/mes_v12_1.c | 4 +
> >>> drivers/gpu/drm/amd/amdgpu/mmhub_v1_0.c | 33 ++
> >>> drivers/gpu/drm/amd/amdgpu/mmhub_v1_7.c | 32 ++
> >>> drivers/gpu/drm/amd/amdgpu/mmhub_v1_8.c | 32 ++
> >>> drivers/gpu/drm/amd/amdgpu/mmhub_v9_4.c | 32 ++
> >>> drivers/gpu/drm/amd/amdgpu/sdma_v5_0.c | 49 +++
> >>> drivers/gpu/drm/amd/amdgpu/sdma_v5_2.c | 49 +++
> >>> drivers/gpu/drm/amd/amdgpu/sdma_v6_0.c | 49 +++
> >>> drivers/gpu/drm/amd/amdgpu/sdma_v7_0.c | 48 +++
> >>> 22 files changed, 943 insertions(+), 1015 deletions(-)
> >>>
> >>
>
^ permalink raw reply [flat|nested] 43+ messages in thread
* Re: [PATCH V2 00/31] Rework GPU TLB invalidation
2026-09-02 13:25 ` Alex Deucher
@ 2026-09-02 13:38 ` Alex Deucher
2026-09-02 13:57 ` Christian König
1 sibling, 0 replies; 43+ messages in thread
From: Alex Deucher @ 2026-09-02 13:38 UTC (permalink / raw)
To: Christian König; +Cc: Alex Deucher, amd-gfx
On Wed, Sep 2, 2026 at 9:25 AM Alex Deucher <alexdeucher@gmail.com> wrote:
>
> On Wed, Sep 2, 2026 at 9:20 AM Christian König <christian.koenig@amd.com> wrote:
> >
> > On 9/2/26 15:06, Alex Deucher wrote:
> > > On Wed, Sep 2, 2026 at 3:25 AM Christian König <christian.koenig@amd.com> wrote:
> > >>
> > >> On 9/1/26 22:10, Alex Deucher wrote:
> > >>> GMC 9-12 use KIQ or MES for TLB invalidations to avoid using MMIO which would
> > >>> require disallowing GFXOFF. KIQ and MES are management queues however and if
> > >>> they hang, they cannot be recovered by a queue reset since they are the
> > >>> mechanisms which handle queue resets. Since using KIQ or MES will exit GFXOFF
> > >>> anyway, explicitly disallow it on the MMIO path and use that. Next, switch to
> > >>> using SDMA for TLB invalidations. SDMA 4.4.x and newer have special packets
> > >>> specifically for this purpose.
> > >>
> > >> That packet was just introduced to work around SRIOV issues and proved to cause stability issues as well.
> > >>
> > >
> > > It was introduced because of gfxoff and having to toggle it to use
> > > MMIO.
> >
> > No, that was completely unrelated to GFXOFF. This also applies to HW where the SDMA is not even in the GFX domain.
> >
> > The problem was the SRIOV could interrupt the SDMA while it waited for the ACK and when the VF was scheduled in again the ACK bit was resetted.
> >
> > So a single packet was introduced which couldn't be interrupted by the hypervisor.
> >
> > > WIndows uses SDMA pretty much exclusively for TLB
> > > invalidations. It's also what the memory hubs team recommends.
> >
> > That is rather interesting.
> > >> Since Arun now found that on basically all Navi generations the SDMA can hang while doing a TLB invalidation I have to clearly NAK this approach.
> > >>
> > >> In the long run we should do the TLB invalidations on newer HW completely with the MES, the SDMA was always just a workaround we should not push forward.
> > >>
> > >
> > > Well KIQ and MES also seem to be problematic as well. See the thread
> > > from Denis about KIQ and his renoir system. There are also various
> > > reports of MES timeouts doing TLB invalidations.
> >
> > The MES is the only instance which knows the process to VMID mapping for user queues.
> >
> > The hack to read out the PASID->VMID mapping register directly from the IH is not really something we can use in the future.
> >
> > So the plan was to move all of that into MES in the near term. IIRC the MES even added a new packet for that.
>
> Sure, this patch set retains the use of MES for pasid invalidation
> where it's supported (navi4x and newer with new enough firmware). We
> could try and get the new packet added to gfx11 MES as well, but then
> there is still gfx9 and 10 which still use KIQ.
>
Regarding gfx 11 mes, there is only a single scheduler queue, similar
to KIQ on previous generations. That micro controller handles both
driver commands and queue scheduling. On gfx12, there are now two MES
queues, one for scheduling and one for misc commands to avoid
contention while scheduling.
This series doesn't actually remove any functionality, it just
provides the option of using:
- KIQ
- SDMA
- MES
- MMIO
and cleans up a lot of duplicate code in the gmc modules. The default
is changed to SDMA or MES depending on what the device supports. We
can adjust the defaults for each gmc generation as needed.
Alex
> Alex
>
> >
> > Regards,
> > Christian.
> >
> > >
> > > Alex
> > >
> > >> Regards,
> > >> Christian.
> > >>
> > >>> If SDMA hangs while doing the invalidation for
> > >>> some reason, it's easier to reset the SDMA queue than KIQ or MES. Finally,
> > >>> most of the TLB invalidation code between GMC 9 through 12 was identical, so
> > >>> move it to common GMC helpers and remove the IP specific code. If the SMDA
> > >>> and MMIO pathes prove to be stable, the KIQ pathes can be removed in the future
> > >>> to further simplify things. SDMA 4.x could also be updated to support PASID
> > >>> invalidation via SDMA using either the new packet (SDMA 4.4.x) or via
> > >>> REG_WRITE/REG_WAIT packets (SDMA 4.0.x).
> > >>>
> > >>> Code is available on this branch as well:
> > >>> https://gitlab.freedesktop.org/agd5f/linux/-/commits/tlb_inv_rework?ref_type=heads
> > >>>
> > >>> V2:
> > >>> - Add missing hub callbacks in gmc9 hubs
> > >>>
> > >>> Alex Deucher (31):
> > >>> drm/amdgpu/gmc9: disallow gfxoff around TLB flushes
> > >>> drm/amdgpu/gmc10: disallow gfxoff around TLB flushes
> > >>> drm/amdgpu/gmc11: disallow gfxoff around TLB flushes
> > >>> drm/amdgpu/gmc12: disallow gfxoff around TLB flushes
> > >>> drm/amdgpu/gmc9: set vmhub funcs for gfxhub and mmhub
> > >>> drm/amdgpu: add a gmc flag for using MMIO for TLB flush
> > >>> drm/amdgpu/gmc9: use MMIO for TLB flushes
> > >>> drm/amdgpu/gmc10: use MMIO for TLB flushes
> > >>> drm/amdgpu/gmc11: use MMIO for TLB flushes
> > >>> drm/amdgpu/gmc12: use MMIO for TLB flushes
> > >>> drm/amdgpu: add a buffer funcs callback for TLB invalidation
> > >>> drm/amdgpu/sdma5.0: add tlb invalidation buffer func callback
> > >>> drm/amdgpu/sdma5.2: add tlb invalidation buffer func callback
> > >>> drm/amdgpu/sdma6: add tlb invalidation buffer func callback
> > >>> drm/amdgpu/sdma7: add tlb invalidation buffer func callback
> > >>> drm/amdgpu: simplify amdgpu_gmc_flush_gpu_tlb()
> > >>> drm/amdgpu: add tlb invalidation method enum
> > >>> drm/amdgpu: plumb tlb inv method in amdgpu_gmc_flush_gpu_tlb_gart()
> > >>> drm/amdgpu: uplevel reset check in amdgpu_gmc_flush_gpu_tlb_gart()
> > >>> drm/amdgpu/gmc: add new callback to lookup vmid to pasid mapping
> > >>> drm/amdgpu: add a gmc callback for the inv semaphore
> > >>> drm/amdgpu/gmc: rework pasid flushing
> > >>> drm/amdgpu/gmc9: use SDMA for gart TLB invalidation
> > >>> drm/amdgpu/gmc10: use SDMA for gart TLB invalidation
> > >>> drm/amdgpu/gmc11: use SDMA for gart TLB invalidation
> > >>> drm/amdgpu/gmc12: use SDMA for gart TLB invalidation
> > >>> drm/amdgpu/gmc10: use SDMA for pasid TLB invalidation
> > >>> drm/amdgpu/gmc11: use SDMA for pasid TLB invalidation
> > >>> drm/amdgpu/gmc12: use MES or SDMA for pasid TLB invalidation
> > >>> drm/amdgpu/gmc12: drop MES tlb inv in gmc callbacks
> > >>> drm/amdgpu/gmc: add helpers for various tlb inv functions
> > >>>
> > >>> drivers/gpu/drm/amd/amdgpu/amdgpu_gart.c | 2 +-
> > >>> drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c | 486 +++++++++++++++++++----
> > >>> drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 33 +-
> > >>> drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h | 18 +
> > >>> drivers/gpu/drm/amd/amdgpu/gfx_v10_0.c | 2 +
> > >>> drivers/gpu/drm/amd/amdgpu/gfxhub_v1_0.c | 33 ++
> > >>> drivers/gpu/drm/amd/amdgpu/gfxhub_v1_2.c | 33 ++
> > >>> drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c | 202 +---------
> > >>> drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c | 207 +---------
> > >>> drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c | 250 ++----------
> > >>> drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c | 225 +----------
> > >>> drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c | 135 ++-----
> > >>> drivers/gpu/drm/amd/amdgpu/mes_v12_0.c | 4 +
> > >>> drivers/gpu/drm/amd/amdgpu/mes_v12_1.c | 4 +
> > >>> drivers/gpu/drm/amd/amdgpu/mmhub_v1_0.c | 33 ++
> > >>> drivers/gpu/drm/amd/amdgpu/mmhub_v1_7.c | 32 ++
> > >>> drivers/gpu/drm/amd/amdgpu/mmhub_v1_8.c | 32 ++
> > >>> drivers/gpu/drm/amd/amdgpu/mmhub_v9_4.c | 32 ++
> > >>> drivers/gpu/drm/amd/amdgpu/sdma_v5_0.c | 49 +++
> > >>> drivers/gpu/drm/amd/amdgpu/sdma_v5_2.c | 49 +++
> > >>> drivers/gpu/drm/amd/amdgpu/sdma_v6_0.c | 49 +++
> > >>> drivers/gpu/drm/amd/amdgpu/sdma_v7_0.c | 48 +++
> > >>> 22 files changed, 943 insertions(+), 1015 deletions(-)
> > >>>
> > >>
> >
^ permalink raw reply [flat|nested] 43+ messages in thread
* Re: [PATCH V2 00/31] Rework GPU TLB invalidation
2026-09-02 13:25 ` Alex Deucher
2026-09-02 13:38 ` Alex Deucher
@ 2026-09-02 13:57 ` Christian König
1 sibling, 0 replies; 43+ messages in thread
From: Christian König @ 2026-09-02 13:57 UTC (permalink / raw)
To: Alex Deucher; +Cc: Alex Deucher, amd-gfx
On 9/2/26 15:25, Alex Deucher wrote:
> On Wed, Sep 2, 2026 at 9:20 AM Christian König <christian.koenig@amd.com> wrote:
>>
>> On 9/2/26 15:06, Alex Deucher wrote:
>>> On Wed, Sep 2, 2026 at 3:25 AM Christian König <christian.koenig@amd.com> wrote:
>>>>
>>>> On 9/1/26 22:10, Alex Deucher wrote:
>>>>> GMC 9-12 use KIQ or MES for TLB invalidations to avoid using MMIO which would
>>>>> require disallowing GFXOFF. KIQ and MES are management queues however and if
>>>>> they hang, they cannot be recovered by a queue reset since they are the
>>>>> mechanisms which handle queue resets. Since using KIQ or MES will exit GFXOFF
>>>>> anyway, explicitly disallow it on the MMIO path and use that. Next, switch to
>>>>> using SDMA for TLB invalidations. SDMA 4.4.x and newer have special packets
>>>>> specifically for this purpose.
>>>>
>>>> That packet was just introduced to work around SRIOV issues and proved to cause stability issues as well.
>>>>
>>>
>>> It was introduced because of gfxoff and having to toggle it to use
>>> MMIO.
>>
>> No, that was completely unrelated to GFXOFF. This also applies to HW where the SDMA is not even in the GFX domain.
>>
>> The problem was the SRIOV could interrupt the SDMA while it waited for the ACK and when the VF was scheduled in again the ACK bit was resetted.
>>
>> So a single packet was introduced which couldn't be interrupted by the hypervisor.
>>
>>> WIndows uses SDMA pretty much exclusively for TLB
>>> invalidations. It's also what the memory hubs team recommends.
>>
>> That is rather interesting.
>>>> Since Arun now found that on basically all Navi generations the SDMA can hang while doing a TLB invalidation I have to clearly NAK this approach.
>>>>
>>>> In the long run we should do the TLB invalidations on newer HW completely with the MES, the SDMA was always just a workaround we should not push forward.
>>>>
>>>
>>> Well KIQ and MES also seem to be problematic as well. See the thread
>>> from Denis about KIQ and his renoir system. There are also various
>>> reports of MES timeouts doing TLB invalidations.
>>
>> The MES is the only instance which knows the process to VMID mapping for user queues.
>>
>> The hack to read out the PASID->VMID mapping register directly from the IH is not really something we can use in the future.
>>
>> So the plan was to move all of that into MES in the near term. IIRC the MES even added a new packet for that.
>
> Sure, this patch set retains the use of MES for pasid invalidation
> where it's supported (navi4x and newer with new enough firmware). We
> could try and get the new packet added to gfx11 MES as well, but then
> there is still gfx9 and 10 which still use KIQ.
Navi 1x already uses the SDMA for TLB invalidations to work around a HW bug. We could trivially enable that for Navi 2x as well.
My take is that when the MES doesn't work reliable for VMID specific invalidation (or rather VMID 0 TLB invalidations since we don't use that for anything else on Navi with kernel queues) then it won't work reliable for PASID based invalidations either.
For GFX9 or rather Vega and MI* products I never heard that they had a problem with KIQ based invalidation and we have a *lot* of those systems in production.
For MI* products switching to the SDMA pretty much sounds like a no-go to me without very extensive testing, especially the KFD/SVM implementation will most likely go boom with that because some HW generations still use the VMID->PASID mapping read out hack.
Switching to the SDMA sounds to my like hiding problems without actually fixing them.
BTW Denis Pisarev reported this morning that his symptoms are gone after updating FW versions.
Regards,
Christian.
>
> Alex
>
>>
>> Regards,
>> Christian.
>>
>>>
>>> Alex
>>>
>>>> Regards,
>>>> Christian.
>>>>
>>>>> If SDMA hangs while doing the invalidation for
>>>>> some reason, it's easier to reset the SDMA queue than KIQ or MES. Finally,
>>>>> most of the TLB invalidation code between GMC 9 through 12 was identical, so
>>>>> move it to common GMC helpers and remove the IP specific code. If the SMDA
>>>>> and MMIO pathes prove to be stable, the KIQ pathes can be removed in the future
>>>>> to further simplify things. SDMA 4.x could also be updated to support PASID
>>>>> invalidation via SDMA using either the new packet (SDMA 4.4.x) or via
>>>>> REG_WRITE/REG_WAIT packets (SDMA 4.0.x).
>>>>>
>>>>> Code is available on this branch as well:
>>>>> https://gitlab.freedesktop.org/agd5f/linux/-/commits/tlb_inv_rework?ref_type=heads
>>>>>
>>>>> V2:
>>>>> - Add missing hub callbacks in gmc9 hubs
>>>>>
>>>>> Alex Deucher (31):
>>>>> drm/amdgpu/gmc9: disallow gfxoff around TLB flushes
>>>>> drm/amdgpu/gmc10: disallow gfxoff around TLB flushes
>>>>> drm/amdgpu/gmc11: disallow gfxoff around TLB flushes
>>>>> drm/amdgpu/gmc12: disallow gfxoff around TLB flushes
>>>>> drm/amdgpu/gmc9: set vmhub funcs for gfxhub and mmhub
>>>>> drm/amdgpu: add a gmc flag for using MMIO for TLB flush
>>>>> drm/amdgpu/gmc9: use MMIO for TLB flushes
>>>>> drm/amdgpu/gmc10: use MMIO for TLB flushes
>>>>> drm/amdgpu/gmc11: use MMIO for TLB flushes
>>>>> drm/amdgpu/gmc12: use MMIO for TLB flushes
>>>>> drm/amdgpu: add a buffer funcs callback for TLB invalidation
>>>>> drm/amdgpu/sdma5.0: add tlb invalidation buffer func callback
>>>>> drm/amdgpu/sdma5.2: add tlb invalidation buffer func callback
>>>>> drm/amdgpu/sdma6: add tlb invalidation buffer func callback
>>>>> drm/amdgpu/sdma7: add tlb invalidation buffer func callback
>>>>> drm/amdgpu: simplify amdgpu_gmc_flush_gpu_tlb()
>>>>> drm/amdgpu: add tlb invalidation method enum
>>>>> drm/amdgpu: plumb tlb inv method in amdgpu_gmc_flush_gpu_tlb_gart()
>>>>> drm/amdgpu: uplevel reset check in amdgpu_gmc_flush_gpu_tlb_gart()
>>>>> drm/amdgpu/gmc: add new callback to lookup vmid to pasid mapping
>>>>> drm/amdgpu: add a gmc callback for the inv semaphore
>>>>> drm/amdgpu/gmc: rework pasid flushing
>>>>> drm/amdgpu/gmc9: use SDMA for gart TLB invalidation
>>>>> drm/amdgpu/gmc10: use SDMA for gart TLB invalidation
>>>>> drm/amdgpu/gmc11: use SDMA for gart TLB invalidation
>>>>> drm/amdgpu/gmc12: use SDMA for gart TLB invalidation
>>>>> drm/amdgpu/gmc10: use SDMA for pasid TLB invalidation
>>>>> drm/amdgpu/gmc11: use SDMA for pasid TLB invalidation
>>>>> drm/amdgpu/gmc12: use MES or SDMA for pasid TLB invalidation
>>>>> drm/amdgpu/gmc12: drop MES tlb inv in gmc callbacks
>>>>> drm/amdgpu/gmc: add helpers for various tlb inv functions
>>>>>
>>>>> drivers/gpu/drm/amd/amdgpu/amdgpu_gart.c | 2 +-
>>>>> drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c | 486 +++++++++++++++++++----
>>>>> drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 33 +-
>>>>> drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h | 18 +
>>>>> drivers/gpu/drm/amd/amdgpu/gfx_v10_0.c | 2 +
>>>>> drivers/gpu/drm/amd/amdgpu/gfxhub_v1_0.c | 33 ++
>>>>> drivers/gpu/drm/amd/amdgpu/gfxhub_v1_2.c | 33 ++
>>>>> drivers/gpu/drm/amd/amdgpu/gmc_v10_0.c | 202 +---------
>>>>> drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c | 207 +---------
>>>>> drivers/gpu/drm/amd/amdgpu/gmc_v12_0.c | 250 ++----------
>>>>> drivers/gpu/drm/amd/amdgpu/gmc_v12_1.c | 225 +----------
>>>>> drivers/gpu/drm/amd/amdgpu/gmc_v9_0.c | 135 ++-----
>>>>> drivers/gpu/drm/amd/amdgpu/mes_v12_0.c | 4 +
>>>>> drivers/gpu/drm/amd/amdgpu/mes_v12_1.c | 4 +
>>>>> drivers/gpu/drm/amd/amdgpu/mmhub_v1_0.c | 33 ++
>>>>> drivers/gpu/drm/amd/amdgpu/mmhub_v1_7.c | 32 ++
>>>>> drivers/gpu/drm/amd/amdgpu/mmhub_v1_8.c | 32 ++
>>>>> drivers/gpu/drm/amd/amdgpu/mmhub_v9_4.c | 32 ++
>>>>> drivers/gpu/drm/amd/amdgpu/sdma_v5_0.c | 49 +++
>>>>> drivers/gpu/drm/amd/amdgpu/sdma_v5_2.c | 49 +++
>>>>> drivers/gpu/drm/amd/amdgpu/sdma_v6_0.c | 49 +++
>>>>> drivers/gpu/drm/amd/amdgpu/sdma_v7_0.c | 48 +++
>>>>> 22 files changed, 943 insertions(+), 1015 deletions(-)
>>>>>
>>>>
>>
^ permalink raw reply [flat|nested] 43+ messages in thread
end of thread, other threads:[~2026-09-02 13:57 UTC | newest]
Thread overview: 43+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-01 20:10 [PATCH V2 00/31] Rework GPU TLB invalidation Alex Deucher
2026-09-01 20:10 ` [PATCH 01/31] drm/amdgpu/gmc9: disallow gfxoff around TLB flushes Alex Deucher
2026-09-02 7:10 ` Christian König
2026-09-02 13:08 ` Alex Deucher
2026-09-02 13:10 ` Alex Deucher
2026-09-01 20:10 ` [PATCH 02/31] drm/amdgpu/gmc10: " Alex Deucher
2026-09-01 20:10 ` [PATCH 03/31] drm/amdgpu/gmc11: " Alex Deucher
2026-09-01 20:10 ` [PATCH 04/31] drm/amdgpu/gmc12: " Alex Deucher
2026-09-01 20:10 ` [PATCH 05/31] drm/amdgpu/gmc9: set vmhub funcs for gfxhub and mmhub Alex Deucher
2026-09-01 20:10 ` [PATCH 06/31] drm/amdgpu: add a gmc flag for using MMIO for TLB flush Alex Deucher
2026-09-01 20:10 ` [PATCH 07/31] drm/amdgpu/gmc9: use MMIO for TLB flushes Alex Deucher
2026-09-01 20:10 ` [PATCH 08/31] drm/amdgpu/gmc10: " Alex Deucher
2026-09-01 20:10 ` [PATCH 09/31] drm/amdgpu/gmc11: " Alex Deucher
2026-09-01 20:10 ` [PATCH 10/31] drm/amdgpu/gmc12: " Alex Deucher
2026-09-01 20:10 ` [PATCH 11/31] drm/amdgpu: add a buffer funcs callback for TLB invalidation Alex Deucher
2026-09-01 20:10 ` [PATCH 12/31] drm/amdgpu/sdma5.0: add tlb invalidation buffer func callback Alex Deucher
2026-09-01 20:10 ` [PATCH 13/31] drm/amdgpu/sdma5.2: " Alex Deucher
2026-09-01 20:10 ` [PATCH 14/31] drm/amdgpu/sdma6: " Alex Deucher
2026-09-01 20:10 ` [PATCH 15/31] drm/amdgpu/sdma7: " Alex Deucher
2026-09-01 20:10 ` [PATCH 16/31] drm/amdgpu: simplify amdgpu_gmc_flush_gpu_tlb() Alex Deucher
2026-09-01 20:10 ` [PATCH 17/31] drm/amdgpu: add tlb invalidation method enum Alex Deucher
2026-09-01 20:10 ` [PATCH 18/31] drm/amdgpu: plumb tlb inv method in amdgpu_gmc_flush_gpu_tlb_gart() Alex Deucher
2026-09-01 20:10 ` [PATCH 19/31] drm/amdgpu: uplevel reset check " Alex Deucher
2026-09-01 20:10 ` [PATCH 20/31] drm/amdgpu/gmc: add new callback to lookup vmid to pasid mapping Alex Deucher
2026-09-02 6:26 ` Zhang, Jesse(Jie)
2026-09-01 20:10 ` [PATCH 21/31] drm/amdgpu: add a gmc callback for the inv semaphore Alex Deucher
2026-09-01 20:10 ` [PATCH 22/31] drm/amdgpu/gmc: rework pasid flushing Alex Deucher
2026-09-02 3:04 ` Zhang, Jesse(Jie)
2026-09-01 20:10 ` [PATCH 23/31] drm/amdgpu/gmc9: use SDMA for gart TLB invalidation Alex Deucher
2026-09-01 20:10 ` [PATCH 24/31] drm/amdgpu/gmc10: " Alex Deucher
2026-09-01 20:10 ` [PATCH 25/31] drm/amdgpu/gmc11: " Alex Deucher
2026-09-01 20:10 ` [PATCH 26/31] drm/amdgpu/gmc12: " Alex Deucher
2026-09-01 20:10 ` [PATCH 27/31] drm/amdgpu/gmc10: use SDMA for pasid " Alex Deucher
2026-09-01 20:10 ` [PATCH 28/31] drm/amdgpu/gmc11: " Alex Deucher
2026-09-01 20:10 ` [PATCH 29/31] drm/amdgpu/gmc12: use MES or " Alex Deucher
2026-09-01 20:10 ` [PATCH 30/31] drm/amdgpu/gmc12: drop MES tlb inv in gmc callbacks Alex Deucher
2026-09-01 20:10 ` [PATCH 31/31] drm/amdgpu/gmc: add helpers for various tlb inv functions Alex Deucher
2026-09-02 7:07 ` [PATCH V2 00/31] Rework GPU TLB invalidation Christian König
2026-09-02 13:06 ` Alex Deucher
2026-09-02 13:20 ` Christian König
2026-09-02 13:25 ` Alex Deucher
2026-09-02 13:38 ` Alex Deucher
2026-09-02 13:57 ` Christian König
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.