* [RFC PATCH 0/2] drm/amdgpu: First-level trap handler for kernel queue VMIDs
@ 2026-09-02 15:06 Srinivasan Shanmugam
2026-09-02 15:06 ` [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure Srinivasan Shanmugam
2026-09-02 15:06 ` [RFC PATCH 2/2] drm/amdgpu: Implement kernel VMID trap handler for GFX10/11/12 Srinivasan Shanmugam
0 siblings, 2 replies; 12+ messages in thread
From: Srinivasan Shanmugam @ 2026-09-02 15:06 UTC (permalink / raw)
To: Christian König, Alex Deucher; +Cc: amd-gfx, Srinivasan Shanmugam
Background
----------
An earlier RFC [1] introduced a new ioctl (AMDGPU_VM_OP_SET_L2_TRAP)
to allow userspace to register a second-level trap handler for user
queues. During review, it was noted that RADV on Vega/Navi hardware
uses kernel queues (context-based submission), not user queues, and
therefore cannot benefit from the second-level trap handler without
first-level trap handler support for kernel queue VMIDs.
Problem
-------
On GFX11+ (MES-based hardware), MES programs SQ_SHADER_TBA/TMA for
user queue VMIDs via the ADD_QUEUE packet's trap_handler_addr field.
However, MES maps kernel queues via ADD_QUEUE with map_legacy_kq=1 but
does NOT program the trap handler for those VMIDs.
On GFX10 and earlier (HWS-based), the driver programs trap handler
registers via SRBM select when KFD queues are set up, but no equivalent
programming exists for the driver-managed kernel queue VMIDs.
Key design decisions:
- MES owns kernel VMIDs but does not program trap handler state
- The driver must program SQ_SHADER_TBA/TMA directly via SRBM select
for kernel queue VMIDs
- Trap handler is a VMID property, not a queue property
- Only kernel queue VMIDs (1..first_kfd_vmid-1) should be programmed
here; user queue VMIDs are handled by MES
Use Cases
---------
- RADV graphics debugging on Vega/Navi/Steam Deck (Valve)
- Future Navi ray tracing features requiring trap handlers on kernel
queues
- Consistent first-level trap handler behavior when userspace switches
between kernel queues and user queues
Design
------
A new vmhub callback (program_kernel_trap_vmids) is added to
amdgpu_vmhub_funcs. Each gfxhub version implements this callback to
write SQ_SHADER_TBA/TMA registers for kernel queue VMIDs. A device-level
TMA BO (kq_tma_bo) is created at trap_init time as the scratch buffer
for kernel queue trap context.
The registers are programmed at two points:
1. amdgpu_trap_init() — on first boot, after ISA and TMA BOs are ready
2. setup_vmid_config() — on GPU resume, after GART registers are restored
GFX versions covered:
- GFX10 (gfxhub_v2_0): mm-prefixed registers, single XCC
- GFX11 (gfxhub_v3_0): reg-prefixed registers, single XCC
- GFX11.5 (gfxhub_v11_5_0): reg-prefixed registers, single XCC
- GFX12 (gfxhub_v12_0): reg-prefixed registers, single XCC
- GFX12.1 (gfxhub_v12_1): reg-prefixed registers, multi-XCC
Open Questions
--------------
1. XCP partitioning: gfxhub_v12_1 programs all XCC instances on each
call. For XCP partition suspend/resume, only a subset of XCCs may
need reprogramming. This will be addressed in a follow-up.
2. MES firmware: ideally MES should honor trap_en regardless of queue
type, but MES APIs are frozen. The driver-side SRBM approach in
this series is the agreed workaround.
This series depends on [1] (second-level trap handler RFC) for the
amdgpu_trap infrastructure (isa_bo, amdgpu_trap_init, amdgpu_trap_alloc).
[1] https://patchwork.freedesktop.org/series/172489/
Srinivasan Shanmugam (2):
drm/amdgpu: Add kernel VMID trap handler infrastructure
drm/amdgpu: Implement kernel VMID trap handler for GFX10/11/12
drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 1 +
drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c | 37 ++++++++++++++++++
drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h | 2 +
drivers/gpu/drm/amd/amdgpu/gfxhub_v11_5_0.c | 43 +++++++++++++++++++++
drivers/gpu/drm/amd/amdgpu/gfxhub_v12_0.c | 34 ++++++++++++++++
drivers/gpu/drm/amd/amdgpu/gfxhub_v12_1.c | 37 ++++++++++++++++++
drivers/gpu/drm/amd/amdgpu/gfxhub_v2_0.c | 34 ++++++++++++++++
drivers/gpu/drm/amd/amdgpu/gfxhub_v3_0.c | 34 ++++++++++++++++
8 files changed, 222 insertions(+)
--
2.34.1
^ permalink raw reply [flat|nested] 12+ messages in thread* [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure 2026-09-02 15:06 [RFC PATCH 0/2] drm/amdgpu: First-level trap handler for kernel queue VMIDs Srinivasan Shanmugam @ 2026-09-02 15:06 ` Srinivasan Shanmugam 2026-09-02 15:28 ` Lazar, Lijo 2026-09-02 15:45 ` Deucher, Alexander 2026-09-02 15:06 ` [RFC PATCH 2/2] drm/amdgpu: Implement kernel VMID trap handler for GFX10/11/12 Srinivasan Shanmugam 1 sibling, 2 replies; 12+ messages in thread From: Srinivasan Shanmugam @ 2026-09-02 15:06 UTC (permalink / raw) To: Christian König, Alex Deucher; +Cc: amd-gfx, Srinivasan Shanmugam MES owns kernel queue VMIDs (1..first_kfd_vmid-1) but does not program SQ_SHADER_TBA/TMA registers for them. Add infrastructure to let the driver program the first-level CWSR trap handler for these VMIDs directly via SRBM select. Add kq_tma_bo — a pinned GTT BO used as device-level TMA scratch for kernel queue VMIDs. Unlike per-process TMA (created in amdgpu_trap_alloc), this is device-level and lives for the lifetime of the device. It is zero-initialized by the OS: the second-level handler address is 0 until userspace calls SET_L2_TRAP. Add amdgpu_trap_program_kernel_vmids() which dispatches to a per-HW vmhub callback, and a new program_kernel_trap_vmids hook in amdgpu_vmhub_funcs for per-GFX-generation register writes. Required for: - RADV graphics debugging on Vega/Navi/Steam Deck (Valve request) - Consistent trap handler behavior when switching between kernel queues and user queues Suggested-by: Christian König <christian.koenig@amd.com> Cc: Alexander Deucher <alexander.deucher@amd.com> Signed-off-by: Srinivasan Shanmugam <srinivasan.shanmugam@amd.com> Change-Id: I0709e788835b69d3d492864de8d5c2d36d0f08c8 --- drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 1 + drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c | 37 ++++++++++++++++++++++++ drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h | 2 ++ 3 files changed, 40 insertions(+) diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h index 3ca187f5ade8..5624a5ab5c62 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h @@ -115,6 +115,7 @@ struct amdgpu_vmhub_funcs { void (*print_l2_protection_fault_status)(struct amdgpu_device *adev, uint32_t status); uint32_t (*get_invalidate_req)(unsigned int vmid, uint32_t flush_type); + void (*program_kernel_trap_vmids)(struct amdgpu_device *adev); }; struct amdgpu_vmhub { diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c index 31f653ec3fb1..0e0aeea0aa2d 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c @@ -256,8 +256,23 @@ int amdgpu_trap_init(struct amdgpu_device *adev) memcpy(ptr, trap_info->isa_buf, trap_info->isa_sz); + /* + * Device-level TMA for kernel queue VMIDs. Pinned GTT — not subject + * to eviction. Zero-initialized by OS: second-level handler address + * is 0 until userspace calls SET_L2_TRAP. + */ + r = amdgpu_bo_create_kernel(adev, AMDGPU_TRAP_TMA_MAX_SIZE, PAGE_SIZE, + AMDGPU_GEM_DOMAIN_GTT, + &trap_info->kq_tma_bo, NULL, NULL); + if (r) { + /* isa_bo freed explicitly; trap_info struct freed by __free */ + amdgpu_bo_free_kernel(&trap_info->isa_bo, NULL, NULL); + return r; + } + amdgpu_trap_cwsr_init_save_area_info(adev, trap_info); adev->trap_info = no_free_ptr(trap_info); + amdgpu_trap_program_kernel_vmids(adev); return 0; } @@ -267,11 +282,33 @@ void amdgpu_trap_fini(struct amdgpu_device *adev) if (!amdgpu_trap_is_enabled(adev)) return; + amdgpu_bo_free_kernel(&adev->trap_info->kq_tma_bo, NULL, NULL); amdgpu_bo_free_kernel(&adev->trap_info->isa_bo, NULL, NULL); kfree(adev->trap_info); adev->trap_info = NULL; } +/** + * amdgpu_trap_program_kernel_vmids - program first-level trap handler for + * kernel queue VMIDs + * @adev: amdgpu device pointer + * + * Programs SQ_SHADER_TBA/TMA for kernel queue VMIDs (1..first_kfd_vmid-1) + * via SRBM select. MES owns these VMIDs but does not program trap handler + * state. Called after trap init and on GPU resume via setup_vmid_config. + */ +void amdgpu_trap_program_kernel_vmids(struct amdgpu_device *adev) +{ + struct amdgpu_vmhub *hub = &adev->vmhub[AMDGPU_GFXHUB(0)]; + + if (!amdgpu_trap_is_enabled(adev)) + return; + if (!hub->vmhub_funcs || !hub->vmhub_funcs->program_kernel_trap_vmids) + return; + + hub->vmhub_funcs->program_kernel_trap_vmids(adev); +} + /* * amdgpu_map_cwsr_trap_handler should be called during amdgpu_vm_init * it maps virtual address amdgpu_trap_tba_vaddr() to this VM, and each diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h index 6d4664469bad..326910f1d94d 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h @@ -48,6 +48,7 @@ struct amdgpu_trap_obj { struct amdgpu_trap_info { /* cwsr isa */ struct amdgpu_bo *isa_bo; + struct amdgpu_bo *kq_tma_bo; /* pinned GTT, device-level TMA for kernel queue VMIDs */ const void *isa_buf; uint32_t isa_sz; /* cwsr size info per XCC*/ @@ -70,6 +71,7 @@ struct amdgpu_trap_usr_addr { }; int amdgpu_trap_init(struct amdgpu_device *adev); +void amdgpu_trap_program_kernel_vmids(struct amdgpu_device *adev); void amdgpu_trap_fini(struct amdgpu_device *adev); int amdgpu_trap_alloc(struct amdgpu_device *adev, struct amdgpu_vm *vm, -- 2.34.1 ^ permalink raw reply related [flat|nested] 12+ messages in thread
* Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure 2026-09-02 15:06 ` [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure Srinivasan Shanmugam @ 2026-09-02 15:28 ` Lazar, Lijo 2026-09-02 15:59 ` SHANMUGAM, SRINIVASAN 2026-09-02 15:45 ` Deucher, Alexander 1 sibling, 1 reply; 12+ messages in thread From: Lazar, Lijo @ 2026-09-02 15:28 UTC (permalink / raw) To: Srinivasan Shanmugam, Christian König, Alex Deucher; +Cc: amd-gfx On 02-Sep-26 8:36 PM, Srinivasan Shanmugam wrote: > MES owns kernel queue VMIDs (1..first_kfd_vmid-1) but does not program > SQ_SHADER_TBA/TMA registers for them. Add infrastructure to let the > driver program the first-level CWSR trap handler for these VMIDs directly > via SRBM select. > > Add kq_tma_bo — a pinned GTT BO used as device-level TMA scratch for > kernel queue VMIDs. Unlike per-process TMA (created in amdgpu_trap_alloc), > this is device-level and lives for the lifetime of the device. It is > zero-initialized by the OS: the second-level handler address is 0 until > userspace calls SET_L2_TRAP. > > Add amdgpu_trap_program_kernel_vmids() which dispatches to a per-HW > vmhub callback, and a new program_kernel_trap_vmids hook in > amdgpu_vmhub_funcs for per-GFX-generation register writes. > > Required for: > - RADV graphics debugging on Vega/Navi/Steam Deck (Valve request) > - Consistent trap handler behavior when switching between kernel > queues and user queues > > Suggested-by: Christian König <christian.koenig@amd.com> > Cc: Alexander Deucher <alexander.deucher@amd.com> > Signed-off-by: Srinivasan Shanmugam <srinivasan.shanmugam@amd.com> > Change-Id: I0709e788835b69d3d492864de8d5c2d36d0f08c8 > --- > drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 1 + > drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c | 37 ++++++++++++++++++++++++ > drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h | 2 ++ > 3 files changed, 40 insertions(+) > > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > index 3ca187f5ade8..5624a5ab5c62 100644 > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > @@ -115,6 +115,7 @@ struct amdgpu_vmhub_funcs { > void (*print_l2_protection_fault_status)(struct amdgpu_device *adev, > uint32_t status); > uint32_t (*get_invalidate_req)(unsigned int vmid, uint32_t flush_type); > + void (*program_kernel_trap_vmids)(struct amdgpu_device *adev); > }; > > struct amdgpu_vmhub { > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > index 31f653ec3fb1..0e0aeea0aa2d 100644 > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > @@ -256,8 +256,23 @@ int amdgpu_trap_init(struct amdgpu_device *adev) > > memcpy(ptr, trap_info->isa_buf, trap_info->isa_sz); > > + /* > + * Device-level TMA for kernel queue VMIDs. Pinned GTT — not subject > + * to eviction. Zero-initialized by OS: second-level handler address > + * is 0 until userspace calls SET_L2_TRAP. How is the exclusivity maintained as a user app doesn't 'own' kernel queue? How is the conflict of different user apps trying to install their own second level handler on a kernel queue handled? Thanks, Lijo > + */ > + r = amdgpu_bo_create_kernel(adev, AMDGPU_TRAP_TMA_MAX_SIZE, PAGE_SIZE, > + AMDGPU_GEM_DOMAIN_GTT, > + &trap_info->kq_tma_bo, NULL, NULL); > + if (r) { > + /* isa_bo freed explicitly; trap_info struct freed by __free */ > + amdgpu_bo_free_kernel(&trap_info->isa_bo, NULL, NULL); > + return r; > + } > + > amdgpu_trap_cwsr_init_save_area_info(adev, trap_info); > adev->trap_info = no_free_ptr(trap_info); > + amdgpu_trap_program_kernel_vmids(adev); > > return 0; > } > @@ -267,11 +282,33 @@ void amdgpu_trap_fini(struct amdgpu_device *adev) > if (!amdgpu_trap_is_enabled(adev)) > return; > > + amdgpu_bo_free_kernel(&adev->trap_info->kq_tma_bo, NULL, NULL); > amdgpu_bo_free_kernel(&adev->trap_info->isa_bo, NULL, NULL); > kfree(adev->trap_info); > adev->trap_info = NULL; > } > > +/** > + * amdgpu_trap_program_kernel_vmids - program first-level trap handler for > + * kernel queue VMIDs > + * @adev: amdgpu device pointer > + * > + * Programs SQ_SHADER_TBA/TMA for kernel queue VMIDs (1..first_kfd_vmid-1) > + * via SRBM select. MES owns these VMIDs but does not program trap handler > + * state. Called after trap init and on GPU resume via setup_vmid_config. > + */ > +void amdgpu_trap_program_kernel_vmids(struct amdgpu_device *adev) > +{ > + struct amdgpu_vmhub *hub = &adev->vmhub[AMDGPU_GFXHUB(0)]; > + > + if (!amdgpu_trap_is_enabled(adev)) > + return; > + if (!hub->vmhub_funcs || !hub->vmhub_funcs->program_kernel_trap_vmids) > + return; > + > + hub->vmhub_funcs->program_kernel_trap_vmids(adev); > +} > + > /* > * amdgpu_map_cwsr_trap_handler should be called during amdgpu_vm_init > * it maps virtual address amdgpu_trap_tba_vaddr() to this VM, and each > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h > index 6d4664469bad..326910f1d94d 100644 > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h > @@ -48,6 +48,7 @@ struct amdgpu_trap_obj { > struct amdgpu_trap_info { > /* cwsr isa */ > struct amdgpu_bo *isa_bo; > + struct amdgpu_bo *kq_tma_bo; /* pinned GTT, device-level TMA for kernel queue VMIDs */ > const void *isa_buf; > uint32_t isa_sz; > /* cwsr size info per XCC*/ > @@ -70,6 +71,7 @@ struct amdgpu_trap_usr_addr { > }; > > int amdgpu_trap_init(struct amdgpu_device *adev); > +void amdgpu_trap_program_kernel_vmids(struct amdgpu_device *adev); > void amdgpu_trap_fini(struct amdgpu_device *adev); > > int amdgpu_trap_alloc(struct amdgpu_device *adev, struct amdgpu_vm *vm, ^ permalink raw reply [flat|nested] 12+ messages in thread
* RE: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure 2026-09-02 15:28 ` Lazar, Lijo @ 2026-09-02 15:59 ` SHANMUGAM, SRINIVASAN 2026-09-02 16:11 ` Lazar, Lijo 0 siblings, 1 reply; 12+ messages in thread From: SHANMUGAM, SRINIVASAN @ 2026-09-02 15:59 UTC (permalink / raw) To: Lazar, Lijo, Koenig, Christian, Deucher, Alexander Cc: amd-gfx@lists.freedesktop.org Public > -----Original Message----- > From: Lazar, Lijo <Lijo.Lazar@amd.com> > Sent: Wednesday, September 2, 2026 8:59 PM > To: SHANMUGAM, SRINIVASAN <SRINIVASAN.SHANMUGAM@amd.com>; > Koenig, Christian <Christian.Koenig@amd.com>; Deucher, Alexander > <Alexander.Deucher@amd.com> > Cc: amd-gfx@lists.freedesktop.org > Subject: Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler > infrastructure > > > > On 02-Sep-26 8:36 PM, Srinivasan Shanmugam wrote: > > MES owns kernel queue VMIDs (1..first_kfd_vmid-1) but does not program > > SQ_SHADER_TBA/TMA registers for them. Add infrastructure to let the > > driver program the first-level CWSR trap handler for these VMIDs > > directly via SRBM select. > > > > Add kq_tma_bo — a pinned GTT BO used as device-level TMA scratch for > > kernel queue VMIDs. Unlike per-process TMA (created in > > amdgpu_trap_alloc), this is device-level and lives for the lifetime of > > the device. It is zero-initialized by the OS: the second-level handler > > address is 0 until userspace calls SET_L2_TRAP. > > > > Add amdgpu_trap_program_kernel_vmids() which dispatches to a per-HW > > vmhub callback, and a new program_kernel_trap_vmids hook in > > amdgpu_vmhub_funcs for per-GFX-generation register writes. > > > > Required for: > > - RADV graphics debugging on Vega/Navi/Steam Deck (Valve request) > > - Consistent trap handler behavior when switching between kernel > > queues and user queues > > > > Suggested-by: Christian König <christian.koenig@amd.com> > > Cc: Alexander Deucher <alexander.deucher@amd.com> > > Signed-off-by: Srinivasan Shanmugam <srinivasan.shanmugam@amd.com> > > Change-Id: I0709e788835b69d3d492864de8d5c2d36d0f08c8 > > --- > > drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 1 + > > drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c | 37 > ++++++++++++++++++++++++ > > drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h | 2 ++ > > 3 files changed, 40 insertions(+) > > > > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > > b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > > index 3ca187f5ade8..5624a5ab5c62 100644 > > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > > @@ -115,6 +115,7 @@ struct amdgpu_vmhub_funcs { > > void (*print_l2_protection_fault_status)(struct amdgpu_device *adev, > > uint32_t status); > > uint32_t (*get_invalidate_req)(unsigned int vmid, uint32_t > > flush_type); > > + void (*program_kernel_trap_vmids)(struct amdgpu_device *adev); > > }; > > > > struct amdgpu_vmhub { > > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > > b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > > index 31f653ec3fb1..0e0aeea0aa2d 100644 > > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > > @@ -256,8 +256,23 @@ int amdgpu_trap_init(struct amdgpu_device *adev) > > > > memcpy(ptr, trap_info->isa_buf, trap_info->isa_sz); > > > > + /* > > + * Device-level TMA for kernel queue VMIDs. Pinned GTT — not subject > > + * to eviction. Zero-initialized by OS: second-level handler address > > + * is 0 until userspace calls SET_L2_TRAP. > > How is the exclusivity maintained as a user app doesn't 'own' kernel queue? How is > the conflict of different user apps trying to install their own second level handler on a > kernel queue handled? As pointed out by Alex: - Each process gets its own per-VM TMA buffer for kernel queues - It is allocated and mapped at a fixed VA in the process's GPUVM when the device is opened, similar to amdgpu_map_static_csa() - Before each job is dispatched to a kernel queue VMID, the driver programs SQ_SHADER_TMA to that process's own TMA VA This way: - App A submits job → TMA = App A's TMA → job runs - App B submits job → TMA = App B's TMA → job runs - No conflict — each process has its own TMA buffer Note: SET_L2_TRAP for kernel queue VMIDs is not part of this series. This series only installs the first-level trap handler. For second-level handler support on kernel queues, I think since: - Each process has its own per-VM TMA buffer (allocated at device open, mapped at a fixed VA in the process's GPUVM) - User calls SET_L2_TRAP → writes second-level handler address into that process's own TMA buffer - When shader crashes → first-level handler reads from that process's TMA → jumps to that process's second-level handler - No conflict — each process has its own TMA with its own second-level handler address The per-VM TMA design is the foundation for this future series. The device-level kq_tma_bo will be removed in v2. For long term — kernel queues are replaced by user queues entirely , where MES already handles this correctly via ADD_QUEUE. Thanks, Srini ^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure 2026-09-02 15:59 ` SHANMUGAM, SRINIVASAN @ 2026-09-02 16:11 ` Lazar, Lijo 2026-09-02 16:55 ` Deucher, Alexander 0 siblings, 1 reply; 12+ messages in thread From: Lazar, Lijo @ 2026-09-02 16:11 UTC (permalink / raw) To: SHANMUGAM, SRINIVASAN, Koenig, Christian, Deucher, Alexander Cc: amd-gfx@lists.freedesktop.org [-- Attachment #1: Type: text/plain, Size: 5502 bytes --] Public I'm not sure how this works, I thought the TMA mapping is per VMID and kernel queues have static VMIDs. Thanks, Lijo ________________________________ From: SHANMUGAM, SRINIVASAN <SRINIVASAN.SHANMUGAM@amd.com> Sent: Wednesday, 02 September 2026 21:29:38 To: Lazar, Lijo <Lijo.Lazar@amd.com>; Koenig, Christian <Christian.Koenig@amd.com>; Deucher, Alexander <Alexander.Deucher@amd.com> Cc: amd-gfx@lists.freedesktop.org <amd-gfx@lists.freedesktop.org> Subject: RE: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure Public > -----Original Message----- > From: Lazar, Lijo <Lijo.Lazar@amd.com> > Sent: Wednesday, September 2, 2026 8:59 PM > To: SHANMUGAM, SRINIVASAN <SRINIVASAN.SHANMUGAM@amd.com>; > Koenig, Christian <Christian.Koenig@amd.com>; Deucher, Alexander > <Alexander.Deucher@amd.com> > Cc: amd-gfx@lists.freedesktop.org > Subject: Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler > infrastructure > > > > On 02-Sep-26 8:36 PM, Srinivasan Shanmugam wrote: > > MES owns kernel queue VMIDs (1..first_kfd_vmid-1) but does not program > > SQ_SHADER_TBA/TMA registers for them. Add infrastructure to let the > > driver program the first-level CWSR trap handler for these VMIDs > > directly via SRBM select. > > > > Add kq_tma_bo — a pinned GTT BO used as device-level TMA scratch for > > kernel queue VMIDs. Unlike per-process TMA (created in > > amdgpu_trap_alloc), this is device-level and lives for the lifetime of > > the device. It is zero-initialized by the OS: the second-level handler > > address is 0 until userspace calls SET_L2_TRAP. > > > > Add amdgpu_trap_program_kernel_vmids() which dispatches to a per-HW > > vmhub callback, and a new program_kernel_trap_vmids hook in > > amdgpu_vmhub_funcs for per-GFX-generation register writes. > > > > Required for: > > - RADV graphics debugging on Vega/Navi/Steam Deck (Valve request) > > - Consistent trap handler behavior when switching between kernel > > queues and user queues > > > > Suggested-by: Christian König <christian.koenig@amd.com> > > Cc: Alexander Deucher <alexander.deucher@amd.com> > > Signed-off-by: Srinivasan Shanmugam <srinivasan.shanmugam@amd.com> > > Change-Id: I0709e788835b69d3d492864de8d5c2d36d0f08c8 > > --- > > drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 1 + > > drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c | 37 > ++++++++++++++++++++++++ > > drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h | 2 ++ > > 3 files changed, 40 insertions(+) > > > > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > > b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > > index 3ca187f5ade8..5624a5ab5c62 100644 > > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > > @@ -115,6 +115,7 @@ struct amdgpu_vmhub_funcs { > > void (*print_l2_protection_fault_status)(struct amdgpu_device *adev, > > uint32_t status); > > uint32_t (*get_invalidate_req)(unsigned int vmid, uint32_t > > flush_type); > > + void (*program_kernel_trap_vmids)(struct amdgpu_device *adev); > > }; > > > > struct amdgpu_vmhub { > > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > > b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > > index 31f653ec3fb1..0e0aeea0aa2d 100644 > > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > > @@ -256,8 +256,23 @@ int amdgpu_trap_init(struct amdgpu_device *adev) > > > > memcpy(ptr, trap_info->isa_buf, trap_info->isa_sz); > > > > + /* > > + * Device-level TMA for kernel queue VMIDs. Pinned GTT — not subject > > + * to eviction. Zero-initialized by OS: second-level handler address > > + * is 0 until userspace calls SET_L2_TRAP. > > How is the exclusivity maintained as a user app doesn't 'own' kernel queue? How is > the conflict of different user apps trying to install their own second level handler on a > kernel queue handled? As pointed out by Alex: - Each process gets its own per-VM TMA buffer for kernel queues - It is allocated and mapped at a fixed VA in the process's GPUVM when the device is opened, similar to amdgpu_map_static_csa() - Before each job is dispatched to a kernel queue VMID, the driver programs SQ_SHADER_TMA to that process's own TMA VA This way: - App A submits job → TMA = App A's TMA → job runs - App B submits job → TMA = App B's TMA → job runs - No conflict — each process has its own TMA buffer Note: SET_L2_TRAP for kernel queue VMIDs is not part of this series. This series only installs the first-level trap handler. For second-level handler support on kernel queues, I think since: - Each process has its own per-VM TMA buffer (allocated at device open, mapped at a fixed VA in the process's GPUVM) - User calls SET_L2_TRAP → writes second-level handler address into that process's own TMA buffer - When shader crashes → first-level handler reads from that process's TMA → jumps to that process's second-level handler - No conflict — each process has its own TMA with its own second-level handler address The per-VM TMA design is the foundation for this future series. The device-level kq_tma_bo will be removed in v2. For long term — kernel queues are replaced by user queues entirely , where MES already handles this correctly via ADD_QUEUE. Thanks, Srini [-- Attachment #2: Type: text/html, Size: 8844 bytes --] ^ permalink raw reply [flat|nested] 12+ messages in thread
* RE: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure 2026-09-02 16:11 ` Lazar, Lijo @ 2026-09-02 16:55 ` Deucher, Alexander 2026-09-03 4:02 ` Lazar, Lijo 0 siblings, 1 reply; 12+ messages in thread From: Deucher, Alexander @ 2026-09-02 16:55 UTC (permalink / raw) To: Lazar, Lijo, SHANMUGAM, SRINIVASAN, Koenig, Christian Cc: amd-gfx@lists.freedesktop.org [-- Attachment #1: Type: text/plain, Size: 6471 bytes --] Public For kernel queues each IB executes with a kernel provided vmid assigned dynamically by the kernel driver. Alex From: Lazar, Lijo <Lijo.Lazar@amd.com> Sent: Wednesday, September 2, 2026 12:12 PM To: SHANMUGAM, SRINIVASAN <SRINIVASAN.SHANMUGAM@amd.com>; Koenig, Christian <Christian.Koenig@amd.com>; Deucher, Alexander <Alexander.Deucher@amd.com> Cc: amd-gfx@lists.freedesktop.org Subject: Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure Public I'm not sure how this works, I thought the TMA mapping is per VMID and kernel queues have static VMIDs. Thanks, Lijo ________________________________ From: SHANMUGAM, SRINIVASAN <SRINIVASAN.SHANMUGAM@amd.com<mailto:SRINIVASAN.SHANMUGAM@amd.com>> Sent: Wednesday, 02 September 2026 21:29:38 To: Lazar, Lijo <Lijo.Lazar@amd.com<mailto:Lijo.Lazar@amd.com>>; Koenig, Christian <Christian.Koenig@amd.com<mailto:Christian.Koenig@amd.com>>; Deucher, Alexander <Alexander.Deucher@amd.com<mailto:Alexander.Deucher@amd.com>> Cc: amd-gfx@lists.freedesktop.org<mailto:amd-gfx@lists.freedesktop.org> <amd-gfx@lists.freedesktop.org<mailto:amd-gfx@lists.freedesktop.org>> Subject: RE: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure Public > -----Original Message----- > From: Lazar, Lijo <Lijo.Lazar@amd.com<mailto:Lijo.Lazar@amd.com>> > Sent: Wednesday, September 2, 2026 8:59 PM > To: SHANMUGAM, SRINIVASAN <SRINIVASAN.SHANMUGAM@amd.com<mailto:SRINIVASAN.SHANMUGAM@amd.com>>; > Koenig, Christian <Christian.Koenig@amd.com<mailto:Christian.Koenig@amd.com>>; Deucher, Alexander > <Alexander.Deucher@amd.com<mailto:Alexander.Deucher@amd.com>> > Cc: amd-gfx@lists.freedesktop.org<mailto:amd-gfx@lists.freedesktop.org> > Subject: Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler > infrastructure > > > > On 02-Sep-26 8:36 PM, Srinivasan Shanmugam wrote: > > MES owns kernel queue VMIDs (1..first_kfd_vmid-1) but does not program > > SQ_SHADER_TBA/TMA registers for them. Add infrastructure to let the > > driver program the first-level CWSR trap handler for these VMIDs > > directly via SRBM select. > > > > Add kq_tma_bo — a pinned GTT BO used as device-level TMA scratch for > > kernel queue VMIDs. Unlike per-process TMA (created in > > amdgpu_trap_alloc), this is device-level and lives for the lifetime of > > the device. It is zero-initialized by the OS: the second-level handler > > address is 0 until userspace calls SET_L2_TRAP. > > > > Add amdgpu_trap_program_kernel_vmids() which dispatches to a per-HW > > vmhub callback, and a new program_kernel_trap_vmids hook in > > amdgpu_vmhub_funcs for per-GFX-generation register writes. > > > > Required for: > > - RADV graphics debugging on Vega/Navi/Steam Deck (Valve request) > > - Consistent trap handler behavior when switching between kernel > > queues and user queues > > > > Suggested-by: Christian König <christian.koenig@amd.com<mailto:christian.koenig@amd.com>> > > Cc: Alexander Deucher <alexander.deucher@amd.com<mailto:alexander.deucher@amd.com>> > > Signed-off-by: Srinivasan Shanmugam <srinivasan.shanmugam@amd.com<mailto:srinivasan.shanmugam@amd.com>> > > Change-Id: I0709e788835b69d3d492864de8d5c2d36d0f08c8 > > --- > > drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 1 + > > drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c | 37 > ++++++++++++++++++++++++ > > drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h | 2 ++ > > 3 files changed, 40 insertions(+) > > > > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > > b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > > index 3ca187f5ade8..5624a5ab5c62 100644 > > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > > @@ -115,6 +115,7 @@ struct amdgpu_vmhub_funcs { > > void (*print_l2_protection_fault_status)(struct amdgpu_device *adev, > > uint32_t status); > > uint32_t (*get_invalidate_req)(unsigned int vmid, uint32_t > > flush_type); > > + void (*program_kernel_trap_vmids)(struct amdgpu_device *adev); > > }; > > > > struct amdgpu_vmhub { > > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > > b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > > index 31f653ec3fb1..0e0aeea0aa2d 100644 > > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > > @@ -256,8 +256,23 @@ int amdgpu_trap_init(struct amdgpu_device *adev) > > > > memcpy(ptr, trap_info->isa_buf, trap_info->isa_sz); > > > > + /* > > + * Device-level TMA for kernel queue VMIDs. Pinned GTT — not subject > > + * to eviction. Zero-initialized by OS: second-level handler address > > + * is 0 until userspace calls SET_L2_TRAP. > > How is the exclusivity maintained as a user app doesn't 'own' kernel queue? How is > the conflict of different user apps trying to install their own second level handler on a > kernel queue handled? As pointed out by Alex: - Each process gets its own per-VM TMA buffer for kernel queues - It is allocated and mapped at a fixed VA in the process's GPUVM when the device is opened, similar to amdgpu_map_static_csa() - Before each job is dispatched to a kernel queue VMID, the driver programs SQ_SHADER_TMA to that process's own TMA VA This way: - App A submits job → TMA = App A's TMA → job runs - App B submits job → TMA = App B's TMA → job runs - No conflict — each process has its own TMA buffer Note: SET_L2_TRAP for kernel queue VMIDs is not part of this series. This series only installs the first-level trap handler. For second-level handler support on kernel queues, I think since: - Each process has its own per-VM TMA buffer (allocated at device open, mapped at a fixed VA in the process's GPUVM) - User calls SET_L2_TRAP → writes second-level handler address into that process's own TMA buffer - When shader crashes → first-level handler reads from that process's TMA → jumps to that process's second-level handler - No conflict — each process has its own TMA with its own second-level handler address The per-VM TMA design is the foundation for this future series. The device-level kq_tma_bo will be removed in v2. For long term — kernel queues are replaced by user queues entirely , where MES already handles this correctly via ADD_QUEUE. Thanks, Srini [-- Attachment #2: Type: text/html, Size: 11940 bytes --] ^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure 2026-09-02 16:55 ` Deucher, Alexander @ 2026-09-03 4:02 ` Lazar, Lijo 2026-09-03 4:03 ` Lazar, Lijo 0 siblings, 1 reply; 12+ messages in thread From: Lazar, Lijo @ 2026-09-03 4:02 UTC (permalink / raw) To: Deucher, Alexander, SHANMUGAM, SRINIVASAN, Koenig, Christian Cc: amd-gfx@lists.freedesktop.org On 02-Sep-26 10:25 PM, Deucher, Alexander wrote: > Public > > > For kernel queues each IB executes with a kernel provided vmid assigned > dynamically by the kernel driver. > In this case, it's a device level TMA 'kq_tma_bo' for first level. That address is programmed in SQ registers. When a job is submitted, the second level handler is picked from what is programmed in kq_tma_bo. Do you mean to say that driver will change that value dynamically based on what is provided by user? Thanks, Lijo > Alex > > *From:*Lazar, Lijo <Lijo.Lazar@amd.com> > *Sent:* Wednesday, September 2, 2026 12:12 PM > *To:* SHANMUGAM, SRINIVASAN <SRINIVASAN.SHANMUGAM@amd.com>; Koenig, > Christian <Christian.Koenig@amd.com>; Deucher, Alexander > <Alexander.Deucher@amd.com> > *Cc:* amd-gfx@lists.freedesktop.org > *Subject:* Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler > infrastructure > > Public > > I'm not sure how this works, I thought the TMA mapping is per VMID and > kernel queues have static VMIDs. > > Thanks, > > Lijo > > ------------------------------------------------------------------------ > > *From:*SHANMUGAM, SRINIVASAN <SRINIVASAN.SHANMUGAM@amd.com > <mailto:SRINIVASAN.SHANMUGAM@amd.com>> > *Sent:* Wednesday, 02 September 2026 21:29:38 > *To:* Lazar, Lijo <Lijo.Lazar@amd.com <mailto:Lijo.Lazar@amd.com>>; > Koenig, Christian <Christian.Koenig@amd.com > <mailto:Christian.Koenig@amd.com>>; Deucher, Alexander > <Alexander.Deucher@amd.com <mailto:Alexander.Deucher@amd.com>> > *Cc:* amd-gfx@lists.freedesktop.org <mailto:amd- > gfx@lists.freedesktop.org> <amd-gfx@lists.freedesktop.org <mailto:amd- > gfx@lists.freedesktop.org>> > *Subject:* RE: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler > infrastructure > > Public > >> -----Original Message----- >> From: Lazar, Lijo <Lijo.Lazar@amd.com <mailto:Lijo.Lazar@amd.com>> >> Sent: Wednesday, September 2, 2026 8:59 PM >> To: SHANMUGAM, SRINIVASAN <SRINIVASAN.SHANMUGAM@amd.com <mailto:SRINIVASAN.SHANMUGAM@amd.com>>; >> Koenig, Christian <Christian.Koenig@amd.com <mailto:Christian.Koenig@amd.com>>; Deucher, > Alexander >> <Alexander.Deucher@amd.com <mailto:Alexander.Deucher@amd.com>> >> Cc: amd-gfx@lists.freedesktop.org <mailto:amd-gfx@lists.freedesktop.org> >> Subject: Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler >> infrastructure >> >> >> >> On 02-Sep-26 8:36 PM, Srinivasan Shanmugam wrote: >> > MES owns kernel queue VMIDs (1..first_kfd_vmid-1) but does not program >> > SQ_SHADER_TBA/TMA registers for them. Add infrastructure to let the >> > driver program the first-level CWSR trap handler for these VMIDs >> > directly via SRBM select. >> > >> > Add kq_tma_bo — a pinned GTT BO used as device-level TMA scratch for >> > kernel queue VMIDs. Unlike per-process TMA (created in >> > amdgpu_trap_alloc), this is device-level and lives for the lifetime of >> > the device. It is zero-initialized by the OS: the second-level handler >> > address is 0 until userspace calls SET_L2_TRAP. >> > >> > Add amdgpu_trap_program_kernel_vmids() which dispatches to a per-HW >> > vmhub callback, and a new program_kernel_trap_vmids hook in >> > amdgpu_vmhub_funcs for per-GFX-generation register writes. >> > >> > Required for: >> > - RADV graphics debugging on Vega/Navi/Steam Deck (Valve request) >> > - Consistent trap handler behavior when switching between kernel >> > queues and user queues >> > >> > Suggested-by: Christian König <christian.koenig@amd.com <mailto:christian.koenig@amd.com>> >> > Cc: Alexander Deucher <alexander.deucher@amd.com <mailto:alexander.deucher@amd.com>> >> > Signed-off-by: Srinivasan Shanmugam <srinivasan.shanmugam@amd.com <mailto:srinivasan.shanmugam@amd.com>> >> > Change-Id: I0709e788835b69d3d492864de8d5c2d36d0f08c8 >> > --- >> > drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 1 + >> > drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c | 37 >> ++++++++++++++++++++++++ >> > drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h | 2 ++ >> > 3 files changed, 40 insertions(+) >> > >> > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h >> > b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h >> > index 3ca187f5ade8..5624a5ab5c62 100644 >> > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h >> > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h >> > @@ -115,6 +115,7 @@ struct amdgpu_vmhub_funcs { >> > void (*print_l2_protection_fault_status)(struct amdgpu_device *adev, >> > uint32_t status); >> > uint32_t (*get_invalidate_req)(unsigned int vmid, uint32_t >> > flush_type); >> > + void (*program_kernel_trap_vmids)(struct amdgpu_device *adev); >> > }; >> > >> > struct amdgpu_vmhub { >> > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c >> > b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c >> > index 31f653ec3fb1..0e0aeea0aa2d 100644 >> > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c >> > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c >> > @@ -256,8 +256,23 @@ int amdgpu_trap_init(struct amdgpu_device *adev) >> > >> > memcpy(ptr, trap_info->isa_buf, trap_info->isa_sz); >> > >> > + /* >> > + * Device-level TMA for kernel queue VMIDs. Pinned GTT — not subject >> > + * to eviction. Zero-initialized by OS: second-level handler address >> > + * is 0 until userspace calls SET_L2_TRAP. >> >> How is the exclusivity maintained as a user app doesn't 'own' kernel queue? How is >> the conflict of different user apps trying to install their own second level handler on a >> kernel queue handled? > > As pointed out by Alex: > > - Each process gets its own per-VM TMA buffer for kernel queues > - It is allocated and mapped at a fixed VA in the process's GPUVM > when the device is opened, similar to amdgpu_map_static_csa() > - Before each job is dispatched to a kernel queue VMID, the driver > programs SQ_SHADER_TMA to that process's own TMA VA > > This way: > - App A submits job → TMA = App A's TMA → job runs > - App B submits job → TMA = App B's TMA → job runs > - No conflict — each process has its own TMA buffer > > Note: SET_L2_TRAP for kernel queue VMIDs is not part of this series. > This series only installs the first-level trap handler. > > For second-level handler support on kernel queues, I think since: > > - Each process has its own per-VM TMA buffer (allocated at device > open, mapped at a fixed VA in the process's GPUVM) > - User calls SET_L2_TRAP → writes second-level handler address > into that process's own TMA buffer > - When shader crashes → first-level handler reads from that > process's TMA → jumps to that process's second-level handler > - No conflict — each process has its own TMA with its own > second-level handler address > > The per-VM TMA design is the foundation for this future series. > The device-level kq_tma_bo will be removed in v2. > > For long term — kernel queues are replaced by user queues entirely > , where MES already handles this correctly via ADD_QUEUE. > > Thanks, > Srini > ^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure 2026-09-03 4:02 ` Lazar, Lijo @ 2026-09-03 4:03 ` Lazar, Lijo 2026-09-03 7:09 ` Christian König 0 siblings, 1 reply; 12+ messages in thread From: Lazar, Lijo @ 2026-09-03 4:03 UTC (permalink / raw) To: Deucher, Alexander, SHANMUGAM, SRINIVASAN, Koenig, Christian Cc: amd-gfx@lists.freedesktop.org On 03-Sep-26 9:32 AM, Lazar, Lijo wrote: > > > On 02-Sep-26 10:25 PM, Deucher, Alexander wrote: >> Public >> >> >> For kernel queues each IB executes with a kernel provided vmid >> assigned dynamically by the kernel driver. >> > > In this case, it's a device level TMA 'kq_tma_bo' for first level. That > address is programmed in SQ registers. When a job is submitted, the > second level handler is picked from what is programmed in kq_tma_bo. Do > you mean to say that driver will change that value dynamically based on > what is provided by user? Do you mean to say that driver will change that value dynamically based on what is provided by user for each job submission? Thanks, Lijo > > Thanks, > Lijo > > > Alex >> >> *From:*Lazar, Lijo <Lijo.Lazar@amd.com> >> *Sent:* Wednesday, September 2, 2026 12:12 PM >> *To:* SHANMUGAM, SRINIVASAN <SRINIVASAN.SHANMUGAM@amd.com>; Koenig, >> Christian <Christian.Koenig@amd.com>; Deucher, Alexander >> <Alexander.Deucher@amd.com> >> *Cc:* amd-gfx@lists.freedesktop.org >> *Subject:* Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap >> handler infrastructure >> >> Public >> >> I'm not sure how this works, I thought the TMA mapping is per VMID and >> kernel queues have static VMIDs. >> >> Thanks, >> >> Lijo >> >> ------------------------------------------------------------------------ >> >> *From:*SHANMUGAM, SRINIVASAN <SRINIVASAN.SHANMUGAM@amd.com >> <mailto:SRINIVASAN.SHANMUGAM@amd.com>> >> *Sent:* Wednesday, 02 September 2026 21:29:38 >> *To:* Lazar, Lijo <Lijo.Lazar@amd.com <mailto:Lijo.Lazar@amd.com>>; >> Koenig, Christian <Christian.Koenig@amd.com >> <mailto:Christian.Koenig@amd.com>>; Deucher, Alexander >> <Alexander.Deucher@amd.com <mailto:Alexander.Deucher@amd.com>> >> *Cc:* amd-gfx@lists.freedesktop.org <mailto:amd- >> gfx@lists.freedesktop.org> <amd-gfx@lists.freedesktop.org <mailto:amd- >> gfx@lists.freedesktop.org>> >> *Subject:* RE: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap >> handler infrastructure >> >> Public >> >>> -----Original Message----- >>> From: Lazar, Lijo <Lijo.Lazar@amd.com <mailto:Lijo.Lazar@amd.com>> >>> Sent: Wednesday, September 2, 2026 8:59 PM >>> To: SHANMUGAM, SRINIVASAN <SRINIVASAN.SHANMUGAM@amd.com >>> <mailto:SRINIVASAN.SHANMUGAM@amd.com>>; >>> Koenig, Christian <Christian.Koenig@amd.com >>> <mailto:Christian.Koenig@amd.com>>; Deucher, >> Alexander >>> <Alexander.Deucher@amd.com <mailto:Alexander.Deucher@amd.com>> >>> Cc: amd-gfx@lists.freedesktop.org <mailto:amd-gfx@lists.freedesktop.org> >>> Subject: Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler >>> infrastructure >>> >>> >>> >>> On 02-Sep-26 8:36 PM, Srinivasan Shanmugam wrote: >>> > MES owns kernel queue VMIDs (1..first_kfd_vmid-1) but does not program >>> > SQ_SHADER_TBA/TMA registers for them. Add infrastructure to let the >>> > driver program the first-level CWSR trap handler for these VMIDs >>> > directly via SRBM select. >>> > >>> > Add kq_tma_bo — a pinned GTT BO used as device-level TMA scratch for >>> > kernel queue VMIDs. Unlike per-process TMA (created in >>> > amdgpu_trap_alloc), this is device-level and lives for the lifetime of >>> > the device. It is zero-initialized by the OS: the second-level handler >>> > address is 0 until userspace calls SET_L2_TRAP. >>> > >>> > Add amdgpu_trap_program_kernel_vmids() which dispatches to a per-HW >>> > vmhub callback, and a new program_kernel_trap_vmids hook in >>> > amdgpu_vmhub_funcs for per-GFX-generation register writes. >>> > >>> > Required for: >>> > - RADV graphics debugging on Vega/Navi/Steam Deck (Valve request) >>> > - Consistent trap handler behavior when switching between kernel >>> > queues and user queues >>> > >>> > Suggested-by: Christian König <christian.koenig@amd.com >>> <mailto:christian.koenig@amd.com>> >>> > Cc: Alexander Deucher <alexander.deucher@amd.com >>> <mailto:alexander.deucher@amd.com>> >>> > Signed-off-by: Srinivasan Shanmugam <srinivasan.shanmugam@amd.com >>> <mailto:srinivasan.shanmugam@amd.com>> >>> > Change-Id: I0709e788835b69d3d492864de8d5c2d36d0f08c8 >>> > --- >>> > drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 1 + >>> > drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c | 37 >>> ++++++++++++++++++++++++ >>> > drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h | 2 ++ >>> > 3 files changed, 40 insertions(+) >>> > >>> > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h >>> > b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h >>> > index 3ca187f5ade8..5624a5ab5c62 100644 >>> > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h >>> > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h >>> > @@ -115,6 +115,7 @@ struct amdgpu_vmhub_funcs { >>> > void (*print_l2_protection_fault_status)(struct amdgpu_device >>> *adev, >>> > uint32_t status); >>> > uint32_t (*get_invalidate_req)(unsigned int vmid, uint32_t >>> > flush_type); >>> > + void (*program_kernel_trap_vmids)(struct amdgpu_device *adev); >>> > }; >>> > >>> > struct amdgpu_vmhub { >>> > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c >>> > b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c >>> > index 31f653ec3fb1..0e0aeea0aa2d 100644 >>> > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c >>> > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c >>> > @@ -256,8 +256,23 @@ int amdgpu_trap_init(struct amdgpu_device *adev) >>> > >>> > memcpy(ptr, trap_info->isa_buf, trap_info->isa_sz); >>> > >>> > + /* >>> > + * Device-level TMA for kernel queue VMIDs. Pinned GTT — not >>> subject >>> > + * to eviction. Zero-initialized by OS: second-level handler >>> address >>> > + * is 0 until userspace calls SET_L2_TRAP. >>> >>> How is the exclusivity maintained as a user app doesn't 'own' kernel >>> queue? How is >>> the conflict of different user apps trying to install their own >>> second level handler on a >>> kernel queue handled? >> >> As pointed out by Alex: >> >> - Each process gets its own per-VM TMA buffer for kernel queues >> - It is allocated and mapped at a fixed VA in the process's GPUVM >> when the device is opened, similar to amdgpu_map_static_csa() >> - Before each job is dispatched to a kernel queue VMID, the driver >> programs SQ_SHADER_TMA to that process's own TMA VA >> >> This way: >> - App A submits job → TMA = App A's TMA → job runs >> - App B submits job → TMA = App B's TMA → job runs >> - No conflict — each process has its own TMA buffer >> >> Note: SET_L2_TRAP for kernel queue VMIDs is not part of this series. >> This series only installs the first-level trap handler. >> >> For second-level handler support on kernel queues, I think since: >> >> - Each process has its own per-VM TMA buffer (allocated at device >> open, mapped at a fixed VA in the process's GPUVM) >> - User calls SET_L2_TRAP → writes second-level handler address >> into that process's own TMA buffer >> - When shader crashes → first-level handler reads from that >> process's TMA → jumps to that process's second-level handler >> - No conflict — each process has its own TMA with its own >> second-level handler address >> >> The per-VM TMA design is the foundation for this future series. >> The device-level kq_tma_bo will be removed in v2. >> >> For long term — kernel queues are replaced by user queues entirely >> , where MES already handles this correctly via ADD_QUEUE. >> >> Thanks, >> Srini >> > ^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure 2026-09-03 4:03 ` Lazar, Lijo @ 2026-09-03 7:09 ` Christian König 2026-09-03 9:34 ` SHANMUGAM, SRINIVASAN 0 siblings, 1 reply; 12+ messages in thread From: Christian König @ 2026-09-03 7:09 UTC (permalink / raw) To: Lazar, Lijo, Deucher, Alexander, SHANMUGAM, SRINIVASAN Cc: amd-gfx@lists.freedesktop.org On 9/3/26 06:03, Lazar, Lijo wrote: > > > On 03-Sep-26 9:32 AM, Lazar, Lijo wrote: >> >> >> On 02-Sep-26 10:25 PM, Deucher, Alexander wrote: >>> Public >>> >>> >>> For kernel queues each IB executes with a kernel provided vmid assigned dynamically by the kernel driver. >>> >> >> In this case, it's a device level TMA 'kq_tma_bo' for first level. That address is programmed in SQ registers. When a job is submitted, the second level handler is picked from what is programmed in kq_tma_bo. Do you mean to say that driver will change that value dynamically based on what is provided by user? > > Do you mean to say that driver will change that value dynamically based on what is provided by user for each job submission? Yeah I agree with Lijo, something doesn't adds up here. As far as I know the same register value is used for both graphics and all compute queues at the same time, so changing this dynamically on each submission won't work (at unless we complete isolate the applications). If I'm not completely mistaken we either need allocate a BO per VM and always map it at the same location or give the location to userspace so that userspace so that the UMD can map it. I don't think we have discussed what the actual plan for that would be. Regards, Christian. > > Thanks, > Lijo > >> >> Thanks, >> Lijo >> >> > Alex >>> >>> *From:*Lazar, Lijo <Lijo.Lazar@amd.com> >>> *Sent:* Wednesday, September 2, 2026 12:12 PM >>> *To:* SHANMUGAM, SRINIVASAN <SRINIVASAN.SHANMUGAM@amd.com>; Koenig, Christian <Christian.Koenig@amd.com>; Deucher, Alexander <Alexander.Deucher@amd.com> >>> *Cc:* amd-gfx@lists.freedesktop.org >>> *Subject:* Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure >>> >>> Public >>> >>> I'm not sure how this works, I thought the TMA mapping is per VMID and kernel queues have static VMIDs. >>> >>> Thanks, >>> >>> Lijo >>> >>> ------------------------------------------------------------------------ >>> >>> *From:*SHANMUGAM, SRINIVASAN <SRINIVASAN.SHANMUGAM@amd.com <mailto:SRINIVASAN.SHANMUGAM@amd.com>> >>> *Sent:* Wednesday, 02 September 2026 21:29:38 >>> *To:* Lazar, Lijo <Lijo.Lazar@amd.com <mailto:Lijo.Lazar@amd.com>>; Koenig, Christian <Christian.Koenig@amd.com <mailto:Christian.Koenig@amd.com>>; Deucher, Alexander <Alexander.Deucher@amd.com <mailto:Alexander.Deucher@amd.com>> >>> *Cc:* amd-gfx@lists.freedesktop.org <mailto:amd- gfx@lists.freedesktop.org> <amd-gfx@lists.freedesktop.org <mailto:amd- gfx@lists.freedesktop.org>> >>> *Subject:* RE: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure >>> >>> Public >>> >>>> -----Original Message----- >>>> From: Lazar, Lijo <Lijo.Lazar@amd.com <mailto:Lijo.Lazar@amd.com>> >>>> Sent: Wednesday, September 2, 2026 8:59 PM >>>> To: SHANMUGAM, SRINIVASAN <SRINIVASAN.SHANMUGAM@amd.com <mailto:SRINIVASAN.SHANMUGAM@amd.com>>; >>>> Koenig, Christian <Christian.Koenig@amd.com <mailto:Christian.Koenig@amd.com>>; Deucher, >>> Alexander >>>> <Alexander.Deucher@amd.com <mailto:Alexander.Deucher@amd.com>> >>>> Cc: amd-gfx@lists.freedesktop.org <mailto:amd-gfx@lists.freedesktop.org> >>>> Subject: Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler >>>> infrastructure >>>> >>>> >>>> >>>> On 02-Sep-26 8:36 PM, Srinivasan Shanmugam wrote: >>>> > MES owns kernel queue VMIDs (1..first_kfd_vmid-1) but does not program >>>> > SQ_SHADER_TBA/TMA registers for them. Add infrastructure to let the >>>> > driver program the first-level CWSR trap handler for these VMIDs >>>> > directly via SRBM select. >>>> > >>>> > Add kq_tma_bo — a pinned GTT BO used as device-level TMA scratch for >>>> > kernel queue VMIDs. Unlike per-process TMA (created in >>>> > amdgpu_trap_alloc), this is device-level and lives for the lifetime of >>>> > the device. It is zero-initialized by the OS: the second-level handler >>>> > address is 0 until userspace calls SET_L2_TRAP. >>>> > >>>> > Add amdgpu_trap_program_kernel_vmids() which dispatches to a per-HW >>>> > vmhub callback, and a new program_kernel_trap_vmids hook in >>>> > amdgpu_vmhub_funcs for per-GFX-generation register writes. >>>> > >>>> > Required for: >>>> > - RADV graphics debugging on Vega/Navi/Steam Deck (Valve request) >>>> > - Consistent trap handler behavior when switching between kernel >>>> > queues and user queues >>>> > >>>> > Suggested-by: Christian König <christian.koenig@amd.com <mailto:christian.koenig@amd.com>> >>>> > Cc: Alexander Deucher <alexander.deucher@amd.com <mailto:alexander.deucher@amd.com>> >>>> > Signed-off-by: Srinivasan Shanmugam <srinivasan.shanmugam@amd.com <mailto:srinivasan.shanmugam@amd.com>> >>>> > Change-Id: I0709e788835b69d3d492864de8d5c2d36d0f08c8 >>>> > --- >>>> > drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 1 + >>>> > drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c | 37 >>>> ++++++++++++++++++++++++ >>>> > drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h | 2 ++ >>>> > 3 files changed, 40 insertions(+) >>>> > >>>> > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h >>>> > b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h >>>> > index 3ca187f5ade8..5624a5ab5c62 100644 >>>> > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h >>>> > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h >>>> > @@ -115,6 +115,7 @@ struct amdgpu_vmhub_funcs { >>>> > void (*print_l2_protection_fault_status)(struct amdgpu_device *adev, >>>> > uint32_t status); >>>> > uint32_t (*get_invalidate_req)(unsigned int vmid, uint32_t >>>> > flush_type); >>>> > + void (*program_kernel_trap_vmids)(struct amdgpu_device *adev); >>>> > }; >>>> > >>>> > struct amdgpu_vmhub { >>>> > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c >>>> > b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c >>>> > index 31f653ec3fb1..0e0aeea0aa2d 100644 >>>> > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c >>>> > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c >>>> > @@ -256,8 +256,23 @@ int amdgpu_trap_init(struct amdgpu_device *adev) >>>> > >>>> > memcpy(ptr, trap_info->isa_buf, trap_info->isa_sz); >>>> > >>>> > + /* >>>> > + * Device-level TMA for kernel queue VMIDs. Pinned GTT — not subject >>>> > + * to eviction. Zero-initialized by OS: second-level handler address >>>> > + * is 0 until userspace calls SET_L2_TRAP. >>>> >>>> How is the exclusivity maintained as a user app doesn't 'own' kernel queue? How is >>>> the conflict of different user apps trying to install their own second level handler on a >>>> kernel queue handled? >>> >>> As pointed out by Alex: >>> >>> - Each process gets its own per-VM TMA buffer for kernel queues >>> - It is allocated and mapped at a fixed VA in the process's GPUVM >>> when the device is opened, similar to amdgpu_map_static_csa() >>> - Before each job is dispatched to a kernel queue VMID, the driver >>> programs SQ_SHADER_TMA to that process's own TMA VA >>> >>> This way: >>> - App A submits job → TMA = App A's TMA → job runs >>> - App B submits job → TMA = App B's TMA → job runs >>> - No conflict — each process has its own TMA buffer >>> >>> Note: SET_L2_TRAP for kernel queue VMIDs is not part of this series. >>> This series only installs the first-level trap handler. >>> >>> For second-level handler support on kernel queues, I think since: >>> >>> - Each process has its own per-VM TMA buffer (allocated at device >>> open, mapped at a fixed VA in the process's GPUVM) >>> - User calls SET_L2_TRAP → writes second-level handler address >>> into that process's own TMA buffer >>> - When shader crashes → first-level handler reads from that >>> process's TMA → jumps to that process's second-level handler >>> - No conflict — each process has its own TMA with its own >>> second-level handler address >>> >>> The per-VM TMA design is the foundation for this future series. >>> The device-level kq_tma_bo will be removed in v2. >>> >>> For long term — kernel queues are replaced by user queues entirely >>> , where MES already handles this correctly via ADD_QUEUE. >>> >>> Thanks, >>> Srini >>> >> > ^ permalink raw reply [flat|nested] 12+ messages in thread
* RE: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure 2026-09-03 7:09 ` Christian König @ 2026-09-03 9:34 ` SHANMUGAM, SRINIVASAN 0 siblings, 0 replies; 12+ messages in thread From: SHANMUGAM, SRINIVASAN @ 2026-09-03 9:34 UTC (permalink / raw) To: Koenig, Christian, Lazar, Lijo, Deucher, Alexander Cc: amd-gfx@lists.freedesktop.org AMD General > -----Original Message----- > From: Koenig, Christian <Christian.Koenig@amd.com> > Sent: Thursday, September 3, 2026 12:39 PM > To: Lazar, Lijo <Lijo.Lazar@amd.com>; Deucher, Alexander > <Alexander.Deucher@amd.com>; SHANMUGAM, SRINIVASAN > <SRINIVASAN.SHANMUGAM@amd.com> > Cc: amd-gfx@lists.freedesktop.org > Subject: Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler > infrastructure > > On 9/3/26 06:03, Lazar, Lijo wrote: > > > > > > On 03-Sep-26 9:32 AM, Lazar, Lijo wrote: > >> > >> > >> On 02-Sep-26 10:25 PM, Deucher, Alexander wrote: > >>> Public > >>> > >>> > >>> For kernel queues each IB executes with a kernel provided vmid assigned > dynamically by the kernel driver. > >>> > >> > >> In this case, it's a device level TMA 'kq_tma_bo' for first level. That address is > programmed in SQ registers. When a job is submitted, the second level handler is > picked from what is programmed in kq_tma_bo. Do you mean to say that driver will > change that value dynamically based on what is provided by user? > > > > Do you mean to say that driver will change that value dynamically based on what > is provided by user for each job submission? > > Yeah I agree with Lijo, something doesn't adds up here. > > As far as I know the same register value is used for both graphics and all compute > queues at the same time, so changing this dynamically on each submission won't > work (at unless we complete isolate the applications). > > If I'm not completely mistaken we either need allocate a BO per VM and always > map it at the same location or give the location to userspace so that userspace so > that the UMD can map it. > > I don't think we have discussed what the actual plan for that would be. Agreed — we cannot reprogram SQ_SHADER_TMA per job submission. Since it is a per-VMID register, it covers all queues in that process at the same time. Swapping it per job would corrupt other running apps. For v2, can we do something like this?: 1. Userspace allocates a TMA BO with VM_ALWAYS_VALID so it is never evicted or moved 2. Userspace passes the TMA GPU VA to the driver via the VM ioctl 3. Driver validates that the BO has VM_ALWAYS_VALID, then programs SQ_SHADER_TMA once 4. Before programming, driver evicts all queues, flushes TLB, writes the register, then restores queues 5. TBA and TMA GPU VA are stored in amdgpu_vm and cleared when VM is destroyed or CLEAR is called For kernel queues: Can we reuse the same userspace-registered TMA BO for kernel queues as well? Since there is only one SQ_SHADER_TMA register per VMID, a separate kernel-internal BO would conflict with what userspace registered. The userspace trap handler shader also needs to find its own TMA to work correctly. If a kernel queue is created before userspace registers the TMA, the driver simply waits — second-level handler activates only after userspace calls the VM ioctl. First-level handler continues to work in the meantime. Should the TMA BO live in VRAM or GTT? VRAM gives faster GPU access during a trap. GTT is easier to manage (a trap handler runs infrequently (only on shader exception) and a trap can fire anytime under memory pressure, and a evicted TMA is worse than a slow TMA.) Best regards, Srini . ^ permalink raw reply [flat|nested] 12+ messages in thread
* RE: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure 2026-09-02 15:06 ` [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure Srinivasan Shanmugam 2026-09-02 15:28 ` Lazar, Lijo @ 2026-09-02 15:45 ` Deucher, Alexander 1 sibling, 0 replies; 12+ messages in thread From: Deucher, Alexander @ 2026-09-02 15:45 UTC (permalink / raw) To: SHANMUGAM, SRINIVASAN, Koenig, Christian Cc: amd-gfx@lists.freedesktop.org, SHANMUGAM, SRINIVASAN Public > -----Original Message----- > From: amd-gfx <amd-gfx-bounces@lists.freedesktop.org> On Behalf Of > Srinivasan Shanmugam > Sent: Wednesday, September 2, 2026 11:07 AM > To: Koenig, Christian <Christian.Koenig@amd.com>; Deucher, Alexander > <Alexander.Deucher@amd.com> > Cc: amd-gfx@lists.freedesktop.org; SHANMUGAM, SRINIVASAN > <SRINIVASAN.SHANMUGAM@amd.com> > Subject: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler > infrastructure > > MES owns kernel queue VMIDs (1..first_kfd_vmid-1) but does not program > SQ_SHADER_TBA/TMA registers for them. Add infrastructure to let the driver > program the first-level CWSR trap handler for these VMIDs directly via SRBM > select. > > Add kq_tma_bo — a pinned GTT BO used as device-level TMA scratch for > kernel queue VMIDs. Unlike per-process TMA (created in amdgpu_trap_alloc), > this is device-level and lives for the lifetime of the device. It is zero-initialized by > the OS: the second-level handler address is 0 until userspace calls > SET_L2_TRAP. There should be one BO per user GPUVM instance, and it needs to be mapped at the same address in every user's GPUVM address space. E.g., this should be allocated and mapped when the user's GPUVM object is created. Alex > > Add amdgpu_trap_program_kernel_vmids() which dispatches to a per-HW > vmhub callback, and a new program_kernel_trap_vmids hook in > amdgpu_vmhub_funcs for per-GFX-generation register writes. > > Required for: > - RADV graphics debugging on Vega/Navi/Steam Deck (Valve request) > - Consistent trap handler behavior when switching between kernel > queues and user queues > > Suggested-by: Christian König <christian.koenig@amd.com> > Cc: Alexander Deucher <alexander.deucher@amd.com> > Signed-off-by: Srinivasan Shanmugam <srinivasan.shanmugam@amd.com> > Change-Id: I0709e788835b69d3d492864de8d5c2d36d0f08c8 > --- > drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 1 + > drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c | 37 > ++++++++++++++++++++++++ > drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h | 2 ++ > 3 files changed, 40 insertions(+) > > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > index 3ca187f5ade8..5624a5ab5c62 100644 > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > @@ -115,6 +115,7 @@ struct amdgpu_vmhub_funcs { > void (*print_l2_protection_fault_status)(struct amdgpu_device *adev, > uint32_t status); > uint32_t (*get_invalidate_req)(unsigned int vmid, uint32_t > flush_type); > + void (*program_kernel_trap_vmids)(struct amdgpu_device *adev); > }; > > struct amdgpu_vmhub { > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > index 31f653ec3fb1..0e0aeea0aa2d 100644 > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > @@ -256,8 +256,23 @@ int amdgpu_trap_init(struct amdgpu_device *adev) > > memcpy(ptr, trap_info->isa_buf, trap_info->isa_sz); > > + /* > + * Device-level TMA for kernel queue VMIDs. Pinned GTT — not > subject > + * to eviction. Zero-initialized by OS: second-level handler address > + * is 0 until userspace calls SET_L2_TRAP. > + */ > + r = amdgpu_bo_create_kernel(adev, > AMDGPU_TRAP_TMA_MAX_SIZE, PAGE_SIZE, > + AMDGPU_GEM_DOMAIN_GTT, > + &trap_info->kq_tma_bo, NULL, NULL); > + if (r) { > + /* isa_bo freed explicitly; trap_info struct freed by __free */ > + amdgpu_bo_free_kernel(&trap_info->isa_bo, NULL, NULL); > + return r; > + } > + > amdgpu_trap_cwsr_init_save_area_info(adev, trap_info); > adev->trap_info = no_free_ptr(trap_info); > + amdgpu_trap_program_kernel_vmids(adev); > > return 0; > } > @@ -267,11 +282,33 @@ void amdgpu_trap_fini(struct amdgpu_device > *adev) > if (!amdgpu_trap_is_enabled(adev)) > return; > > + amdgpu_bo_free_kernel(&adev->trap_info->kq_tma_bo, NULL, > NULL); > amdgpu_bo_free_kernel(&adev->trap_info->isa_bo, NULL, NULL); > kfree(adev->trap_info); > adev->trap_info = NULL; > } > > +/** > + * amdgpu_trap_program_kernel_vmids - program first-level trap handler for > + * kernel queue VMIDs > + * @adev: amdgpu device pointer > + * > + * Programs SQ_SHADER_TBA/TMA for kernel queue VMIDs > +(1..first_kfd_vmid-1) > + * via SRBM select. MES owns these VMIDs but does not program trap > +handler > + * state. Called after trap init and on GPU resume via setup_vmid_config. > + */ > +void amdgpu_trap_program_kernel_vmids(struct amdgpu_device *adev) { > + struct amdgpu_vmhub *hub = &adev- > >vmhub[AMDGPU_GFXHUB(0)]; > + > + if (!amdgpu_trap_is_enabled(adev)) > + return; > + if (!hub->vmhub_funcs || !hub->vmhub_funcs- > >program_kernel_trap_vmids) > + return; > + > + hub->vmhub_funcs->program_kernel_trap_vmids(adev); > +} > + > /* > * amdgpu_map_cwsr_trap_handler should be called during amdgpu_vm_init > * it maps virtual address amdgpu_trap_tba_vaddr() to this VM, and each diff > --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h > b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h > index 6d4664469bad..326910f1d94d 100644 > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h > @@ -48,6 +48,7 @@ struct amdgpu_trap_obj { struct amdgpu_trap_info { > /* cwsr isa */ > struct amdgpu_bo *isa_bo; > + struct amdgpu_bo *kq_tma_bo; /* pinned GTT, device-level > TMA for kernel queue VMIDs */ > const void *isa_buf; > uint32_t isa_sz; > /* cwsr size info per XCC*/ > @@ -70,6 +71,7 @@ struct amdgpu_trap_usr_addr { }; > > int amdgpu_trap_init(struct amdgpu_device *adev); > +void amdgpu_trap_program_kernel_vmids(struct amdgpu_device *adev); > void amdgpu_trap_fini(struct amdgpu_device *adev); > > int amdgpu_trap_alloc(struct amdgpu_device *adev, struct amdgpu_vm *vm, > -- > 2.34.1 ^ permalink raw reply [flat|nested] 12+ messages in thread
* [RFC PATCH 2/2] drm/amdgpu: Implement kernel VMID trap handler for GFX10/11/12 2026-09-02 15:06 [RFC PATCH 0/2] drm/amdgpu: First-level trap handler for kernel queue VMIDs Srinivasan Shanmugam 2026-09-02 15:06 ` [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure Srinivasan Shanmugam @ 2026-09-02 15:06 ` Srinivasan Shanmugam 1 sibling, 0 replies; 12+ messages in thread From: Srinivasan Shanmugam @ 2026-09-02 15:06 UTC (permalink / raw) To: Christian König, Alex Deucher Cc: amd-gfx, Srinivasan Shanmugam, Timur Kristof Implement program_kernel_trap_vmids() in each gfxhub version to write SQ_SHADER_TBA_LO/HI and SQ_SHADER_TMA_LO/HI registers for kernel queue VMIDs (1..first_kfd_vmid-1) using SRBM select. Only kernel queue VMIDs are programmed here. User queue VMIDs (first_kfd_vmid..15) are handled by MES via the ADD_QUEUE packet's trap_handler_addr field and must not be touched by the driver. The TBA address points to the device-level CWSR ISA BO (isa_bo). The TMA address points to the device-level scratch BO (kq_tma_bo). Both are pinned GTT BOs and cannot be evicted. Addresses are stored as addr >> 8 to match the hardware register format (256-byte aligned). TRAP_EN is set in TBA_HI to enable trap handling for each VMID. WARN_ON is used to catch alignment regressions at development time. The null check on gfx.funcs->select_me_pipe_q guards against calls before GFX IP is fully initialized. GFX10 (gfxhub_v2_0): uses mm-prefixed registers. GFX11 (gfxhub_v3_0, gfxhub_v11_5_0): uses reg-prefixed registers. GFX12 (gfxhub_v12_0): uses reg-prefixed registers. GFX12.1 (gfxhub_v12_1): multi-XCC, iterates over all XCC instances. The function is called at two points: 1. amdgpu_trap_init() — first boot, after ISA and TMA BOs are ready 2. setup_vmid_config() — GPU resume, after GART registers are restored Suggested-by: Christian König <christian.koenig@amd.com> Cc: Alexander Deucher <alexander.deucher@amd.com> Cc: Timur Kristof <timur.kristof@valve.com> Signed-off-by: Srinivasan Shanmugam <srinivasan.shanmugam@amd.com> --- drivers/gpu/drm/amd/amdgpu/gfxhub_v11_5_0.c | 43 +++++++++++++++++++++ drivers/gpu/drm/amd/amdgpu/gfxhub_v12_0.c | 34 ++++++++++++++++ drivers/gpu/drm/amd/amdgpu/gfxhub_v12_1.c | 37 ++++++++++++++++++ drivers/gpu/drm/amd/amdgpu/gfxhub_v2_0.c | 34 ++++++++++++++++ drivers/gpu/drm/amd/amdgpu/gfxhub_v3_0.c | 34 ++++++++++++++++ 5 files changed, 182 insertions(+) diff --git a/drivers/gpu/drm/amd/amdgpu/gfxhub_v11_5_0.c b/drivers/gpu/drm/amd/amdgpu/gfxhub_v11_5_0.c index 652eea6eae4a..32d651be9ad3 100644 --- a/drivers/gpu/drm/amd/amdgpu/gfxhub_v11_5_0.c +++ b/drivers/gpu/drm/amd/amdgpu/gfxhub_v11_5_0.c @@ -23,6 +23,7 @@ #include "amdgpu.h" #include "gfxhub_v11_5_0.h" +#include "amdgpu_trap.h" #include "gc/gc_11_5_0_offset.h" #include "gc/gc_11_5_0_sh_mask.h" @@ -290,6 +291,44 @@ static void gfxhub_v11_5_0_disable_identity_aperture(struct amdgpu_device *adev) } +/* + * MES owns kernel VMIDs but does not program trap handler registers. + * Program SQ_SHADER_TBA/TMA directly via SRBM select so the first-level + * CWSR handler is active for kernel queue VMIDs. Required for RADV + * debugging (Valve/Steam Deck) and future Navi ray tracing on kernel queues. + */ +static void gfxhub_v11_5_0_program_kernel_trap_vmids(struct amdgpu_device *adev) +{ + u64 tba_addr = amdgpu_bo_gpu_offset(adev->trap_info->isa_bo); + u64 tma_addr = amdgpu_bo_gpu_offset(adev->trap_info->kq_tma_bo); + int i; + + if (!adev->gfx.funcs || !adev->gfx.funcs->select_me_pipe_q) + return; + + WARN_ON(!IS_ALIGNED(tba_addr, 256)); + WARN_ON(!IS_ALIGNED(tma_addr, 256)); + + mutex_lock(&adev->srbm_mutex); + /* Program VMIDs 1..first_kfd_vmid-1 (kernel queue range only). + * User queue VMIDs (first_kfd_vmid..15) are programmed by MES. + */ + for (i = 1; i < adev->vm_manager.first_kfd_vmid; i++) { + amdgpu_gfx_select_me_pipe_q(adev, 0, 0, 0, i, 0); + WREG32_SOC15(GC, 0, regSQ_SHADER_TBA_LO, + lower_32_bits(tba_addr >> 8)); + WREG32_SOC15(GC, 0, regSQ_SHADER_TBA_HI, + upper_32_bits(tba_addr >> 8) | + (1 << SQ_SHADER_TBA_HI__TRAP_EN__SHIFT)); + WREG32_SOC15(GC, 0, regSQ_SHADER_TMA_LO, + lower_32_bits(tma_addr >> 8)); + WREG32_SOC15(GC, 0, regSQ_SHADER_TMA_HI, + upper_32_bits(tma_addr >> 8)); + } + amdgpu_gfx_select_me_pipe_q(adev, 0, 0, 0, 0, 0); + mutex_unlock(&adev->srbm_mutex); +} + static void gfxhub_v11_5_0_setup_vmid_config(struct amdgpu_device *adev) { struct amdgpu_vmhub *hub = &adev->vmhub[AMDGPU_GFXHUB(0)]; @@ -337,6 +376,9 @@ static void gfxhub_v11_5_0_setup_vmid_config(struct amdgpu_device *adev) } hub->vm_cntx_cntl = tmp; + + if (amdgpu_trap_is_enabled(adev)) + gfxhub_v11_5_0_program_kernel_trap_vmids(adev); } static void gfxhub_v11_5_0_program_invalidation(struct amdgpu_device *adev) @@ -459,6 +501,7 @@ static void gfxhub_v11_5_0_set_fault_enable_default(struct amdgpu_device *adev, static const struct amdgpu_vmhub_funcs gfxhub_v11_5_0_vmhub_funcs = { .print_l2_protection_fault_status = gfxhub_v11_5_0_print_l2_protection_fault_status, .get_invalidate_req = gfxhub_v11_5_0_get_invalidate_req, + .program_kernel_trap_vmids = gfxhub_v11_5_0_program_kernel_trap_vmids, }; static void gfxhub_v11_5_0_init(struct amdgpu_device *adev) diff --git a/drivers/gpu/drm/amd/amdgpu/gfxhub_v12_0.c b/drivers/gpu/drm/amd/amdgpu/gfxhub_v12_0.c index 6cbf837d50dd..ffab4a25ec6c 100644 --- a/drivers/gpu/drm/amd/amdgpu/gfxhub_v12_0.c +++ b/drivers/gpu/drm/amd/amdgpu/gfxhub_v12_0.c @@ -28,6 +28,7 @@ #include "gc/gc_12_0_0_sh_mask.h" #include "soc24_enum.h" #include "soc15_common.h" +#include "amdgpu_trap.h" #define regGCVM_L2_CNTL3_DEFAULT 0x80120007 #define regGCVM_L2_CNTL4_DEFAULT 0x000000c1 @@ -295,6 +296,35 @@ static void gfxhub_v12_0_disable_identity_aperture(struct amdgpu_device *adev) } +static void gfxhub_v12_0_program_kernel_trap_vmids(struct amdgpu_device *adev) +{ + u64 tba_addr = amdgpu_bo_gpu_offset(adev->trap_info->isa_bo); + u64 tma_addr = amdgpu_bo_gpu_offset(adev->trap_info->kq_tma_bo); + int i; + + if (!adev->gfx.funcs || !adev->gfx.funcs->select_me_pipe_q) + return; + + WARN_ON(!IS_ALIGNED(tba_addr, 256)); + WARN_ON(!IS_ALIGNED(tma_addr, 256)); + + mutex_lock(&adev->srbm_mutex); + for (i = 1; i < adev->vm_manager.first_kfd_vmid; i++) { + amdgpu_gfx_select_me_pipe_q(adev, 0, 0, 0, i, 0); + WREG32_SOC15(GC, 0, regSQ_SHADER_TBA_LO, + lower_32_bits(tba_addr >> 8)); + WREG32_SOC15(GC, 0, regSQ_SHADER_TBA_HI, + upper_32_bits(tba_addr >> 8) | + (1 << SQ_SHADER_TBA_HI__TRAP_EN__SHIFT)); + WREG32_SOC15(GC, 0, regSQ_SHADER_TMA_LO, + lower_32_bits(tma_addr >> 8)); + WREG32_SOC15(GC, 0, regSQ_SHADER_TMA_HI, + upper_32_bits(tma_addr >> 8)); + } + amdgpu_gfx_select_me_pipe_q(adev, 0, 0, 0, 0, 0); + mutex_unlock(&adev->srbm_mutex); +} + static void gfxhub_v12_0_setup_vmid_config(struct amdgpu_device *adev) { struct amdgpu_vmhub *hub = &adev->vmhub[AMDGPU_GFXHUB(0)]; @@ -342,6 +372,9 @@ static void gfxhub_v12_0_setup_vmid_config(struct amdgpu_device *adev) } hub->vm_cntx_cntl = tmp; + + if (amdgpu_trap_is_enabled(adev)) + gfxhub_v12_0_program_kernel_trap_vmids(adev); } static void gfxhub_v12_0_program_invalidation(struct amdgpu_device *adev) @@ -464,6 +497,7 @@ static void gfxhub_v12_0_set_fault_enable_default(struct amdgpu_device *adev, static const struct amdgpu_vmhub_funcs gfxhub_v12_0_vmhub_funcs = { .print_l2_protection_fault_status = gfxhub_v12_0_print_l2_protection_fault_status, .get_invalidate_req = gfxhub_v12_0_get_invalidate_req, + .program_kernel_trap_vmids = gfxhub_v12_0_program_kernel_trap_vmids, }; static void gfxhub_v12_0_init(struct amdgpu_device *adev) diff --git a/drivers/gpu/drm/amd/amdgpu/gfxhub_v12_1.c b/drivers/gpu/drm/amd/amdgpu/gfxhub_v12_1.c index 4c2fd1e6616e..de162c5066ea 100644 --- a/drivers/gpu/drm/amd/amdgpu/gfxhub_v12_1.c +++ b/drivers/gpu/drm/amd/amdgpu/gfxhub_v12_1.c @@ -21,6 +21,7 @@ * */ #include "amdgpu.h" +#include "amdgpu_trap.h" #include "amdgpu_xcp.h" #include "gfxhub_v12_1.h" @@ -406,6 +407,38 @@ static void gfxhub_v12_1_xcc_disable_identity_aperture(struct amdgpu_device *ade } } +static void gfxhub_v12_1_program_kernel_trap_vmids(struct amdgpu_device *adev) +{ + u64 tba_addr = amdgpu_bo_gpu_offset(adev->trap_info->isa_bo); + u64 tma_addr = amdgpu_bo_gpu_offset(adev->trap_info->kq_tma_bo); + u32 xcc_mask = GENMASK(NUM_XCC(adev->gfx.xcc_mask) - 1, 0); + int i, j; + + if (!adev->gfx.funcs || !adev->gfx.funcs->select_me_pipe_q) + return; + + WARN_ON(!IS_ALIGNED(tba_addr, 256)); + WARN_ON(!IS_ALIGNED(tma_addr, 256)); + + for_each_inst(j, xcc_mask) { + mutex_lock(&adev->srbm_mutex); + for (i = 1; i < adev->vm_manager.first_kfd_vmid; i++) { + amdgpu_gfx_select_me_pipe_q(adev, 0, 0, 0, i, j); + WREG32_SOC15(GC, GET_INST(GC, j), regSQ_SHADER_TBA_LO, + lower_32_bits(tba_addr >> 8)); + WREG32_SOC15(GC, GET_INST(GC, j), regSQ_SHADER_TBA_HI, + upper_32_bits(tba_addr >> 8) | + (1 << SQ_SHADER_TBA_HI__TRAP_EN__SHIFT)); + WREG32_SOC15(GC, GET_INST(GC, j), regSQ_SHADER_TMA_LO, + lower_32_bits(tma_addr >> 8)); + WREG32_SOC15(GC, GET_INST(GC, j), regSQ_SHADER_TMA_HI, + upper_32_bits(tma_addr >> 8)); + } + amdgpu_gfx_select_me_pipe_q(adev, 0, 0, 0, 0, j); + mutex_unlock(&adev->srbm_mutex); + } +} + static void gfxhub_v12_1_xcc_setup_vmid_config(struct amdgpu_device *adev, uint32_t xcc_mask) { @@ -468,6 +501,9 @@ static void gfxhub_v12_1_xcc_setup_vmid_config(struct amdgpu_device *adev, hub->vm_cntx_cntl = tmp; } + + if (amdgpu_trap_is_enabled(adev)) + gfxhub_v12_1_program_kernel_trap_vmids(adev); } static void gfxhub_v12_1_xcc_program_invalidation(struct amdgpu_device *adev, @@ -751,6 +787,7 @@ static void gfxhub_v12_1_print_l2_protection_fault_status(struct amdgpu_device * static const struct amdgpu_vmhub_funcs gfxhub_v12_1_vmhub_funcs = { .print_l2_protection_fault_status = gfxhub_v12_1_print_l2_protection_fault_status, .get_invalidate_req = gfxhub_v12_1_get_invalidate_req, + .program_kernel_trap_vmids = gfxhub_v12_1_program_kernel_trap_vmids, }; static void gfxhub_v12_1_xcc_init(struct amdgpu_device *adev, uint32_t xcc_mask) diff --git a/drivers/gpu/drm/amd/amdgpu/gfxhub_v2_0.c b/drivers/gpu/drm/amd/amdgpu/gfxhub_v2_0.c index 9ea593e2c719..52a6ee09b036 100644 --- a/drivers/gpu/drm/amd/amdgpu/gfxhub_v2_0.c +++ b/drivers/gpu/drm/amd/amdgpu/gfxhub_v2_0.c @@ -22,6 +22,7 @@ */ #include "amdgpu.h" +#include "amdgpu_trap.h" #include "gfxhub_v2_0.h" #include "gc/gc_10_1_0_offset.h" @@ -280,6 +281,35 @@ static void gfxhub_v2_0_disable_identity_aperture(struct amdgpu_device *adev) } +static void gfxhub_v2_0_program_kernel_trap_vmids(struct amdgpu_device *adev) +{ + u64 tba_addr = amdgpu_bo_gpu_offset(adev->trap_info->isa_bo); + u64 tma_addr = amdgpu_bo_gpu_offset(adev->trap_info->kq_tma_bo); + int i; + + if (!adev->gfx.funcs || !adev->gfx.funcs->select_me_pipe_q) + return; + + WARN_ON(!IS_ALIGNED(tba_addr, 256)); + WARN_ON(!IS_ALIGNED(tma_addr, 256)); + + mutex_lock(&adev->srbm_mutex); + for (i = 1; i < adev->vm_manager.first_kfd_vmid; i++) { + amdgpu_gfx_select_me_pipe_q(adev, 0, 0, 0, i, 0); + WREG32_SOC15(GC, 0, mmSQ_SHADER_TBA_LO, + lower_32_bits(tba_addr >> 8)); + WREG32_SOC15(GC, 0, mmSQ_SHADER_TBA_HI, + upper_32_bits(tba_addr >> 8) | + (1 << SQ_SHADER_TBA_HI__TRAP_EN__SHIFT)); + WREG32_SOC15(GC, 0, mmSQ_SHADER_TMA_LO, + lower_32_bits(tma_addr >> 8)); + WREG32_SOC15(GC, 0, mmSQ_SHADER_TMA_HI, + upper_32_bits(tma_addr >> 8)); + } + amdgpu_gfx_select_me_pipe_q(adev, 0, 0, 0, 0, 0); + mutex_unlock(&adev->srbm_mutex); +} + static void gfxhub_v2_0_setup_vmid_config(struct amdgpu_device *adev) { struct amdgpu_vmhub *hub = &adev->vmhub[AMDGPU_GFXHUB(0)]; @@ -327,6 +357,9 @@ static void gfxhub_v2_0_setup_vmid_config(struct amdgpu_device *adev) } hub->vm_cntx_cntl = tmp; + + if (amdgpu_trap_is_enabled(adev)) + gfxhub_v2_0_program_kernel_trap_vmids(adev); } static void gfxhub_v2_0_program_invalidation(struct amdgpu_device *adev) @@ -428,6 +461,7 @@ static void gfxhub_v2_0_set_fault_enable_default(struct amdgpu_device *adev, static const struct amdgpu_vmhub_funcs gfxhub_v2_0_vmhub_funcs = { .print_l2_protection_fault_status = gfxhub_v2_0_print_l2_protection_fault_status, .get_invalidate_req = gfxhub_v2_0_get_invalidate_req, + .program_kernel_trap_vmids = gfxhub_v2_0_program_kernel_trap_vmids, }; static void gfxhub_v2_0_init(struct amdgpu_device *adev) diff --git a/drivers/gpu/drm/amd/amdgpu/gfxhub_v3_0.c b/drivers/gpu/drm/amd/amdgpu/gfxhub_v3_0.c index 9e6a6e13dec0..287e7ea43287 100644 --- a/drivers/gpu/drm/amd/amdgpu/gfxhub_v3_0.c +++ b/drivers/gpu/drm/amd/amdgpu/gfxhub_v3_0.c @@ -22,6 +22,7 @@ */ #include "amdgpu.h" +#include "amdgpu_trap.h" #include "gfxhub_v3_0.h" #include "gc/gc_11_0_0_offset.h" @@ -287,6 +288,35 @@ static void gfxhub_v3_0_disable_identity_aperture(struct amdgpu_device *adev) } +static void gfxhub_v3_0_program_kernel_trap_vmids(struct amdgpu_device *adev) +{ + u64 tba_addr = amdgpu_bo_gpu_offset(adev->trap_info->isa_bo); + u64 tma_addr = amdgpu_bo_gpu_offset(adev->trap_info->kq_tma_bo); + int i; + + if (!adev->gfx.funcs || !adev->gfx.funcs->select_me_pipe_q) + return; + + WARN_ON(!IS_ALIGNED(tba_addr, 256)); + WARN_ON(!IS_ALIGNED(tma_addr, 256)); + + mutex_lock(&adev->srbm_mutex); + for (i = 1; i < adev->vm_manager.first_kfd_vmid; i++) { + amdgpu_gfx_select_me_pipe_q(adev, 0, 0, 0, i, 0); + WREG32_SOC15(GC, 0, regSQ_SHADER_TBA_LO, + lower_32_bits(tba_addr >> 8)); + WREG32_SOC15(GC, 0, regSQ_SHADER_TBA_HI, + upper_32_bits(tba_addr >> 8) | + (1 << SQ_SHADER_TBA_HI__TRAP_EN__SHIFT)); + WREG32_SOC15(GC, 0, regSQ_SHADER_TMA_LO, + lower_32_bits(tma_addr >> 8)); + WREG32_SOC15(GC, 0, regSQ_SHADER_TMA_HI, + upper_32_bits(tma_addr >> 8)); + } + amdgpu_gfx_select_me_pipe_q(adev, 0, 0, 0, 0, 0); + mutex_unlock(&adev->srbm_mutex); +} + static void gfxhub_v3_0_setup_vmid_config(struct amdgpu_device *adev) { struct amdgpu_vmhub *hub = &adev->vmhub[AMDGPU_GFXHUB(0)]; @@ -334,6 +364,9 @@ static void gfxhub_v3_0_setup_vmid_config(struct amdgpu_device *adev) } hub->vm_cntx_cntl = tmp; + + if (amdgpu_trap_is_enabled(adev)) + gfxhub_v3_0_program_kernel_trap_vmids(adev); } static void gfxhub_v3_0_program_invalidation(struct amdgpu_device *adev) @@ -456,6 +489,7 @@ static void gfxhub_v3_0_set_fault_enable_default(struct amdgpu_device *adev, static const struct amdgpu_vmhub_funcs gfxhub_v3_0_vmhub_funcs = { .print_l2_protection_fault_status = gfxhub_v3_0_print_l2_protection_fault_status, .get_invalidate_req = gfxhub_v3_0_get_invalidate_req, + .program_kernel_trap_vmids = gfxhub_v3_0_program_kernel_trap_vmids, }; static void gfxhub_v3_0_init(struct amdgpu_device *adev) -- 2.34.1 ^ permalink raw reply related [flat|nested] 12+ messages in thread
end of thread, other threads:[~2026-09-03 9:34 UTC | newest] Thread overview: 12+ messages (download: mbox.gz follow: Atom feed -- links below jump to the message on this page -- 2026-09-02 15:06 [RFC PATCH 0/2] drm/amdgpu: First-level trap handler for kernel queue VMIDs Srinivasan Shanmugam 2026-09-02 15:06 ` [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure Srinivasan Shanmugam 2026-09-02 15:28 ` Lazar, Lijo 2026-09-02 15:59 ` SHANMUGAM, SRINIVASAN 2026-09-02 16:11 ` Lazar, Lijo 2026-09-02 16:55 ` Deucher, Alexander 2026-09-03 4:02 ` Lazar, Lijo 2026-09-03 4:03 ` Lazar, Lijo 2026-09-03 7:09 ` Christian König 2026-09-03 9:34 ` SHANMUGAM, SRINIVASAN 2026-09-02 15:45 ` Deucher, Alexander 2026-09-02 15:06 ` [RFC PATCH 2/2] drm/amdgpu: Implement kernel VMID trap handler for GFX10/11/12 Srinivasan Shanmugam
This is an external index of several public inboxes, see mirroring instructions on how to clone and mirror all data and code used by this external index.