On 02-Sep-26 10:25 PM, Deucher, Alexander wrote:
Public


For kernel queues each IB executes with a kernel provided vmid assigned dynamically by the kernel driver.


In this case, it's a device level TMA 'kq_tma_bo' for first level. That address is programmed in SQ registers. When a job is submitted, the second level handler is picked from what is programmed in kq_tma_bo. Do you mean to say that driver will change that value dynamically based on what is provided by user?

Thanks,
Lijo

 > Alex

*From:*Lazar, Lijo <[email protected]>
*Sent:* Wednesday, September 2, 2026 12:12 PM
*To:* SHANMUGAM, SRINIVASAN <[email protected]>; Koenig, Christian <[email protected]>; Deucher, Alexander <[email protected]>
*Cc:* [email protected]
*Subject:* Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure

Public

I'm not sure how this works, I thought the TMA mapping is per VMID and kernel queues have static VMIDs.

Thanks,

Lijo

------------------------------------------------------------------------

*From:*SHANMUGAM, SRINIVASAN <[email protected] <mailto:[email protected]>>
*Sent:* Wednesday, 02 September 2026 21:29:38
*To:* Lazar, Lijo <[email protected] <mailto:[email protected]>>; Koenig, Christian <[email protected] <mailto:[email protected]>>; Deucher, Alexander <[email protected] <mailto:[email protected]>> *Cc:* [email protected] <mailto:amd- [email protected]> <[email protected] <mailto:amd- [email protected]>> *Subject:* RE: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure

Public

-----Original Message-----
From: Lazar, Lijo <[email protected] <mailto:[email protected]>>
Sent: Wednesday, September 2, 2026 8:59 PM
To: SHANMUGAM, SRINIVASAN <[email protected] 
<mailto:[email protected]>>;
Koenig, Christian <[email protected] <mailto:[email protected]>>; Deucher,
Alexander
<[email protected] <mailto:[email protected]>>
Cc: [email protected] <mailto:[email protected]>
Subject: Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler
infrastructure



On 02-Sep-26 8:36 PM, Srinivasan Shanmugam wrote:
> MES owns kernel queue VMIDs (1..first_kfd_vmid-1) but does not program
> SQ_SHADER_TBA/TMA registers for them. Add infrastructure to let the
> driver program the first-level CWSR trap handler for these VMIDs
> directly via SRBM select.
>
> Add kq_tma_bo — a pinned GTT BO used as device-level TMA scratch for
> kernel queue VMIDs. Unlike per-process TMA (created in
> amdgpu_trap_alloc), this is device-level and lives for the lifetime of
> the device. It is zero-initialized by the OS: the second-level handler
> address is 0 until userspace calls SET_L2_TRAP.
>
> Add amdgpu_trap_program_kernel_vmids() which dispatches to a per-HW
> vmhub callback, and a new program_kernel_trap_vmids hook in
> amdgpu_vmhub_funcs for per-GFX-generation register writes.
>
> Required for:
>    - RADV graphics debugging on Vega/Navi/Steam Deck (Valve request)
>    - Consistent trap handler behavior when switching between kernel
>      queues and user queues
>
> Suggested-by: Christian König <[email protected] 
<mailto:[email protected]>>
> Cc: Alexander Deucher <[email protected] 
<mailto:[email protected]>>
> Signed-off-by: Srinivasan Shanmugam <[email protected] 
<mailto:[email protected]>>
> Change-Id: I0709e788835b69d3d492864de8d5c2d36d0f08c8
> ---
>   drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h  |  1 +
>   drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c | 37
++++++++++++++++++++++++
>   drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h |  2 ++
>   3 files changed, 40 insertions(+)
>
> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
> b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
> index 3ca187f5ade8..5624a5ab5c62 100644
> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
> @@ -115,6 +115,7 @@ struct amdgpu_vmhub_funcs {
>     void (*print_l2_protection_fault_status)(struct amdgpu_device *adev,
>                                              uint32_t status);
>     uint32_t (*get_invalidate_req)(unsigned int vmid, uint32_t
> flush_type);
> +   void (*program_kernel_trap_vmids)(struct amdgpu_device *adev);
>   };
>
>   struct amdgpu_vmhub {
> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
> b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
> index 31f653ec3fb1..0e0aeea0aa2d 100644
> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
> @@ -256,8 +256,23 @@ int amdgpu_trap_init(struct amdgpu_device *adev)
>
>     memcpy(ptr, trap_info->isa_buf, trap_info->isa_sz);
>
> +   /*
> +    * Device-level TMA for kernel queue VMIDs. Pinned GTT — not subject
> +    * to eviction. Zero-initialized by OS: second-level handler address
> +    * is 0 until userspace calls SET_L2_TRAP.

How is the exclusivity maintained as a user app doesn't 'own' kernel queue? How 
is
the conflict of different user apps trying to install their own second level 
handler on a
kernel queue handled?

As pointed out by Alex:

   - Each process gets its own per-VM TMA buffer for kernel queues
   - It is allocated and mapped at a fixed VA in the process's GPUVM
     when the device is opened, similar to amdgpu_map_static_csa()
   - Before each job is dispatched to a kernel queue VMID, the driver
     programs SQ_SHADER_TMA to that process's own TMA VA

This way:
   - App A submits job → TMA = App A's TMA → job runs
   - App B submits job → TMA = App B's TMA → job runs
   - No conflict — each process has its own TMA buffer

Note: SET_L2_TRAP for kernel queue VMIDs is not part of this series.
This series only installs the first-level trap handler.

For second-level handler support on kernel queues, I think since:

   - Each process has its own per-VM TMA buffer (allocated at device
     open, mapped at a fixed VA in the process's GPUVM)
   - User calls SET_L2_TRAP → writes second-level handler address
     into that process's own TMA buffer
   - When shader crashes → first-level handler reads from that
     process's TMA → jumps to that process's second-level handler
   - No conflict — each process has its own TMA with its own
     second-level handler address

The per-VM TMA design is the foundation for this future series.
The device-level kq_tma_bo will be removed in v2.

For long term — kernel queues are replaced by user queues entirely
, where MES already handles this correctly via ADD_QUEUE.

Thanks,
Srini


Reply via email to