From: "Lazar, Lijo" <lijo.lazar@amd.com>
To: "Deucher, Alexander" <Alexander.Deucher@amd.com>,
"SHANMUGAM, SRINIVASAN" <SRINIVASAN.SHANMUGAM@amd.com>,
"Koenig, Christian" <Christian.Koenig@amd.com>
Cc: "amd-gfx@lists.freedesktop.org" <amd-gfx@lists.freedesktop.org>
Subject: Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure
Date: Thu, 3 Sep 2026 09:33:35 +0530 [thread overview]
Message-ID: <de1313dc-62bb-42b6-a22e-4775260fd680@amd.com> (raw)
In-Reply-To: <f2755b35-ef95-4067-b5bd-0a37ee1df457@amd.com>
On 03-Sep-26 9:32 AM, Lazar, Lijo wrote:
>
>
> On 02-Sep-26 10:25 PM, Deucher, Alexander wrote:
>> Public
>>
>>
>> For kernel queues each IB executes with a kernel provided vmid
>> assigned dynamically by the kernel driver.
>>
>
> In this case, it's a device level TMA 'kq_tma_bo' for first level. That
> address is programmed in SQ registers. When a job is submitted, the
> second level handler is picked from what is programmed in kq_tma_bo. Do
> you mean to say that driver will change that value dynamically based on
> what is provided by user?
Do you mean to say that driver will change that value dynamically based
on what is provided by user for each job submission?
Thanks,
Lijo
>
> Thanks,
> Lijo
>
> > Alex
>>
>> *From:*Lazar, Lijo <Lijo.Lazar@amd.com>
>> *Sent:* Wednesday, September 2, 2026 12:12 PM
>> *To:* SHANMUGAM, SRINIVASAN <SRINIVASAN.SHANMUGAM@amd.com>; Koenig,
>> Christian <Christian.Koenig@amd.com>; Deucher, Alexander
>> <Alexander.Deucher@amd.com>
>> *Cc:* amd-gfx@lists.freedesktop.org
>> *Subject:* Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap
>> handler infrastructure
>>
>> Public
>>
>> I'm not sure how this works, I thought the TMA mapping is per VMID and
>> kernel queues have static VMIDs.
>>
>> Thanks,
>>
>> Lijo
>>
>> ------------------------------------------------------------------------
>>
>> *From:*SHANMUGAM, SRINIVASAN <SRINIVASAN.SHANMUGAM@amd.com
>> <mailto:SRINIVASAN.SHANMUGAM@amd.com>>
>> *Sent:* Wednesday, 02 September 2026 21:29:38
>> *To:* Lazar, Lijo <Lijo.Lazar@amd.com <mailto:Lijo.Lazar@amd.com>>;
>> Koenig, Christian <Christian.Koenig@amd.com
>> <mailto:Christian.Koenig@amd.com>>; Deucher, Alexander
>> <Alexander.Deucher@amd.com <mailto:Alexander.Deucher@amd.com>>
>> *Cc:* amd-gfx@lists.freedesktop.org <mailto:amd-
>> gfx@lists.freedesktop.org> <amd-gfx@lists.freedesktop.org <mailto:amd-
>> gfx@lists.freedesktop.org>>
>> *Subject:* RE: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap
>> handler infrastructure
>>
>> Public
>>
>>> -----Original Message-----
>>> From: Lazar, Lijo <Lijo.Lazar@amd.com <mailto:Lijo.Lazar@amd.com>>
>>> Sent: Wednesday, September 2, 2026 8:59 PM
>>> To: SHANMUGAM, SRINIVASAN <SRINIVASAN.SHANMUGAM@amd.com
>>> <mailto:SRINIVASAN.SHANMUGAM@amd.com>>;
>>> Koenig, Christian <Christian.Koenig@amd.com
>>> <mailto:Christian.Koenig@amd.com>>; Deucher,
>> Alexander
>>> <Alexander.Deucher@amd.com <mailto:Alexander.Deucher@amd.com>>
>>> Cc: amd-gfx@lists.freedesktop.org <mailto:amd-gfx@lists.freedesktop.org>
>>> Subject: Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler
>>> infrastructure
>>>
>>>
>>>
>>> On 02-Sep-26 8:36 PM, Srinivasan Shanmugam wrote:
>>> > MES owns kernel queue VMIDs (1..first_kfd_vmid-1) but does not program
>>> > SQ_SHADER_TBA/TMA registers for them. Add infrastructure to let the
>>> > driver program the first-level CWSR trap handler for these VMIDs
>>> > directly via SRBM select.
>>> >
>>> > Add kq_tma_bo — a pinned GTT BO used as device-level TMA scratch for
>>> > kernel queue VMIDs. Unlike per-process TMA (created in
>>> > amdgpu_trap_alloc), this is device-level and lives for the lifetime of
>>> > the device. It is zero-initialized by the OS: the second-level handler
>>> > address is 0 until userspace calls SET_L2_TRAP.
>>> >
>>> > Add amdgpu_trap_program_kernel_vmids() which dispatches to a per-HW
>>> > vmhub callback, and a new program_kernel_trap_vmids hook in
>>> > amdgpu_vmhub_funcs for per-GFX-generation register writes.
>>> >
>>> > Required for:
>>> > - RADV graphics debugging on Vega/Navi/Steam Deck (Valve request)
>>> > - Consistent trap handler behavior when switching between kernel
>>> > queues and user queues
>>> >
>>> > Suggested-by: Christian König <christian.koenig@amd.com
>>> <mailto:christian.koenig@amd.com>>
>>> > Cc: Alexander Deucher <alexander.deucher@amd.com
>>> <mailto:alexander.deucher@amd.com>>
>>> > Signed-off-by: Srinivasan Shanmugam <srinivasan.shanmugam@amd.com
>>> <mailto:srinivasan.shanmugam@amd.com>>
>>> > Change-Id: I0709e788835b69d3d492864de8d5c2d36d0f08c8
>>> > ---
>>> > drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 1 +
>>> > drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c | 37
>>> ++++++++++++++++++++++++
>>> > drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h | 2 ++
>>> > 3 files changed, 40 insertions(+)
>>> >
>>> > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
>>> > b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
>>> > index 3ca187f5ade8..5624a5ab5c62 100644
>>> > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
>>> > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
>>> > @@ -115,6 +115,7 @@ struct amdgpu_vmhub_funcs {
>>> > void (*print_l2_protection_fault_status)(struct amdgpu_device
>>> *adev,
>>> > uint32_t status);
>>> > uint32_t (*get_invalidate_req)(unsigned int vmid, uint32_t
>>> > flush_type);
>>> > + void (*program_kernel_trap_vmids)(struct amdgpu_device *adev);
>>> > };
>>> >
>>> > struct amdgpu_vmhub {
>>> > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
>>> > b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
>>> > index 31f653ec3fb1..0e0aeea0aa2d 100644
>>> > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
>>> > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
>>> > @@ -256,8 +256,23 @@ int amdgpu_trap_init(struct amdgpu_device *adev)
>>> >
>>> > memcpy(ptr, trap_info->isa_buf, trap_info->isa_sz);
>>> >
>>> > + /*
>>> > + * Device-level TMA for kernel queue VMIDs. Pinned GTT — not
>>> subject
>>> > + * to eviction. Zero-initialized by OS: second-level handler
>>> address
>>> > + * is 0 until userspace calls SET_L2_TRAP.
>>>
>>> How is the exclusivity maintained as a user app doesn't 'own' kernel
>>> queue? How is
>>> the conflict of different user apps trying to install their own
>>> second level handler on a
>>> kernel queue handled?
>>
>> As pointed out by Alex:
>>
>> - Each process gets its own per-VM TMA buffer for kernel queues
>> - It is allocated and mapped at a fixed VA in the process's GPUVM
>> when the device is opened, similar to amdgpu_map_static_csa()
>> - Before each job is dispatched to a kernel queue VMID, the driver
>> programs SQ_SHADER_TMA to that process's own TMA VA
>>
>> This way:
>> - App A submits job → TMA = App A's TMA → job runs
>> - App B submits job → TMA = App B's TMA → job runs
>> - No conflict — each process has its own TMA buffer
>>
>> Note: SET_L2_TRAP for kernel queue VMIDs is not part of this series.
>> This series only installs the first-level trap handler.
>>
>> For second-level handler support on kernel queues, I think since:
>>
>> - Each process has its own per-VM TMA buffer (allocated at device
>> open, mapped at a fixed VA in the process's GPUVM)
>> - User calls SET_L2_TRAP → writes second-level handler address
>> into that process's own TMA buffer
>> - When shader crashes → first-level handler reads from that
>> process's TMA → jumps to that process's second-level handler
>> - No conflict — each process has its own TMA with its own
>> second-level handler address
>>
>> The per-VM TMA design is the foundation for this future series.
>> The device-level kq_tma_bo will be removed in v2.
>>
>> For long term — kernel queues are replaced by user queues entirely
>> , where MES already handles this correctly via ADD_QUEUE.
>>
>> Thanks,
>> Srini
>>
>
next prev parent reply other threads:[~2026-09-03 4:03 UTC|newest]
Thread overview: 12+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-02 15:06 [RFC PATCH 0/2] drm/amdgpu: First-level trap handler for kernel queue VMIDs Srinivasan Shanmugam
2026-09-02 15:06 ` [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure Srinivasan Shanmugam
2026-09-02 15:28 ` Lazar, Lijo
2026-09-02 15:59 ` SHANMUGAM, SRINIVASAN
2026-09-02 16:11 ` Lazar, Lijo
2026-09-02 16:55 ` Deucher, Alexander
2026-09-03 4:02 ` Lazar, Lijo
2026-09-03 4:03 ` Lazar, Lijo [this message]
2026-09-03 7:09 ` Christian König
2026-09-03 9:34 ` SHANMUGAM, SRINIVASAN
2026-09-02 15:45 ` Deucher, Alexander
2026-09-02 15:06 ` [RFC PATCH 2/2] drm/amdgpu: Implement kernel VMID trap handler for GFX10/11/12 Srinivasan Shanmugam
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=de1313dc-62bb-42b6-a22e-4775260fd680@amd.com \
--to=lijo.lazar@amd.com \
--cc=Alexander.Deucher@amd.com \
--cc=Christian.Koenig@amd.com \
--cc=SRINIVASAN.SHANMUGAM@amd.com \
--cc=amd-gfx@lists.freedesktop.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.