On 02-Sep-26 10:25 PM, Deucher, Alexander wrote:
Public
For kernel queues each IB executes with a kernel provided vmid assigned
dynamically by the kernel driver.
In this case, it's a device level TMA 'kq_tma_bo' for first level. That
address is programmed in SQ registers. When a job is submitted, the
second level handler is picked from what is programmed in kq_tma_bo. Do
you mean to say that driver will change that value dynamically based on
what is provided by user?
Thanks,
Lijo
> Alex
*From:*Lazar, Lijo <[email protected]>
*Sent:* Wednesday, September 2, 2026 12:12 PM
*To:* SHANMUGAM, SRINIVASAN <[email protected]>; Koenig,
Christian <[email protected]>; Deucher, Alexander
<[email protected]>
*Cc:* [email protected]
*Subject:* Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler
infrastructure
Public
I'm not sure how this works, I thought the TMA mapping is per VMID and
kernel queues have static VMIDs.
Thanks,
Lijo
------------------------------------------------------------------------
*From:*SHANMUGAM, SRINIVASAN <[email protected]
<mailto:[email protected]>>
*Sent:* Wednesday, 02 September 2026 21:29:38
*To:* Lazar, Lijo <[email protected] <mailto:[email protected]>>;
Koenig, Christian <[email protected]
<mailto:[email protected]>>; Deucher, Alexander
<[email protected] <mailto:[email protected]>>
*Cc:* [email protected] <mailto:amd-
[email protected]> <[email protected] <mailto:amd-
[email protected]>>
*Subject:* RE: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler
infrastructure
Public
-----Original Message-----
From: Lazar, Lijo <[email protected] <mailto:[email protected]>>
Sent: Wednesday, September 2, 2026 8:59 PM
To: SHANMUGAM, SRINIVASAN <[email protected]
<mailto:[email protected]>>;
Koenig, Christian <[email protected] <mailto:[email protected]>>; Deucher,
Alexander
<[email protected] <mailto:[email protected]>>
Cc: [email protected] <mailto:[email protected]>
Subject: Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler
infrastructure
On 02-Sep-26 8:36 PM, Srinivasan Shanmugam wrote:
> MES owns kernel queue VMIDs (1..first_kfd_vmid-1) but does not program
> SQ_SHADER_TBA/TMA registers for them. Add infrastructure to let the
> driver program the first-level CWSR trap handler for these VMIDs
> directly via SRBM select.
>
> Add kq_tma_bo — a pinned GTT BO used as device-level TMA scratch for
> kernel queue VMIDs. Unlike per-process TMA (created in
> amdgpu_trap_alloc), this is device-level and lives for the lifetime of
> the device. It is zero-initialized by the OS: the second-level handler
> address is 0 until userspace calls SET_L2_TRAP.
>
> Add amdgpu_trap_program_kernel_vmids() which dispatches to a per-HW
> vmhub callback, and a new program_kernel_trap_vmids hook in
> amdgpu_vmhub_funcs for per-GFX-generation register writes.
>
> Required for:
> - RADV graphics debugging on Vega/Navi/Steam Deck (Valve request)
> - Consistent trap handler behavior when switching between kernel
> queues and user queues
>
> Suggested-by: Christian König <[email protected]
<mailto:[email protected]>>
> Cc: Alexander Deucher <[email protected]
<mailto:[email protected]>>
> Signed-off-by: Srinivasan Shanmugam <[email protected]
<mailto:[email protected]>>
> Change-Id: I0709e788835b69d3d492864de8d5c2d36d0f08c8
> ---
> drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 1 +
> drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c | 37
++++++++++++++++++++++++
> drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h | 2 ++
> 3 files changed, 40 insertions(+)
>
> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
> b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
> index 3ca187f5ade8..5624a5ab5c62 100644
> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
> @@ -115,6 +115,7 @@ struct amdgpu_vmhub_funcs {
> void (*print_l2_protection_fault_status)(struct amdgpu_device *adev,
> uint32_t status);
> uint32_t (*get_invalidate_req)(unsigned int vmid, uint32_t
> flush_type);
> + void (*program_kernel_trap_vmids)(struct amdgpu_device *adev);
> };
>
> struct amdgpu_vmhub {
> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
> b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
> index 31f653ec3fb1..0e0aeea0aa2d 100644
> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
> @@ -256,8 +256,23 @@ int amdgpu_trap_init(struct amdgpu_device *adev)
>
> memcpy(ptr, trap_info->isa_buf, trap_info->isa_sz);
>
> + /*
> + * Device-level TMA for kernel queue VMIDs. Pinned GTT — not subject
> + * to eviction. Zero-initialized by OS: second-level handler address
> + * is 0 until userspace calls SET_L2_TRAP.
How is the exclusivity maintained as a user app doesn't 'own' kernel queue? How
is
the conflict of different user apps trying to install their own second level
handler on a
kernel queue handled?
As pointed out by Alex:
- Each process gets its own per-VM TMA buffer for kernel queues
- It is allocated and mapped at a fixed VA in the process's GPUVM
when the device is opened, similar to amdgpu_map_static_csa()
- Before each job is dispatched to a kernel queue VMID, the driver
programs SQ_SHADER_TMA to that process's own TMA VA
This way:
- App A submits job → TMA = App A's TMA → job runs
- App B submits job → TMA = App B's TMA → job runs
- No conflict — each process has its own TMA buffer
Note: SET_L2_TRAP for kernel queue VMIDs is not part of this series.
This series only installs the first-level trap handler.
For second-level handler support on kernel queues, I think since:
- Each process has its own per-VM TMA buffer (allocated at device
open, mapped at a fixed VA in the process's GPUVM)
- User calls SET_L2_TRAP → writes second-level handler address
into that process's own TMA buffer
- When shader crashes → first-level handler reads from that
process's TMA → jumps to that process's second-level handler
- No conflict — each process has its own TMA with its own
second-level handler address
The per-VM TMA design is the foundation for this future series.
The device-level kq_tma_bo will be removed in v2.
For long term — kernel queues are replaced by user queues entirely
, where MES already handles this correctly via ADD_QUEUE.
Thanks,
Srini