Public For kernel queues each IB executes with a kernel provided vmid assigned dynamically by the kernel driver.
Alex From: Lazar, Lijo <[email protected]> Sent: Wednesday, September 2, 2026 12:12 PM To: SHANMUGAM, SRINIVASAN <[email protected]>; Koenig, Christian <[email protected]>; Deucher, Alexander <[email protected]> Cc: [email protected] Subject: Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure Public I'm not sure how this works, I thought the TMA mapping is per VMID and kernel queues have static VMIDs. Thanks, Lijo ________________________________ From: SHANMUGAM, SRINIVASAN <[email protected]<mailto:[email protected]>> Sent: Wednesday, 02 September 2026 21:29:38 To: Lazar, Lijo <[email protected]<mailto:[email protected]>>; Koenig, Christian <[email protected]<mailto:[email protected]>>; Deucher, Alexander <[email protected]<mailto:[email protected]>> Cc: [email protected]<mailto:[email protected]> <[email protected]<mailto:[email protected]>> Subject: RE: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler infrastructure Public > -----Original Message----- > From: Lazar, Lijo <[email protected]<mailto:[email protected]>> > Sent: Wednesday, September 2, 2026 8:59 PM > To: SHANMUGAM, SRINIVASAN > <[email protected]<mailto:[email protected]>>; > Koenig, Christian > <[email protected]<mailto:[email protected]>>; Deucher, > Alexander > <[email protected]<mailto:[email protected]>> > Cc: [email protected]<mailto:[email protected]> > Subject: Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler > infrastructure > > > > On 02-Sep-26 8:36 PM, Srinivasan Shanmugam wrote: > > MES owns kernel queue VMIDs (1..first_kfd_vmid-1) but does not program > > SQ_SHADER_TBA/TMA registers for them. Add infrastructure to let the > > driver program the first-level CWSR trap handler for these VMIDs > > directly via SRBM select. > > > > Add kq_tma_bo — a pinned GTT BO used as device-level TMA scratch for > > kernel queue VMIDs. Unlike per-process TMA (created in > > amdgpu_trap_alloc), this is device-level and lives for the lifetime of > > the device. It is zero-initialized by the OS: the second-level handler > > address is 0 until userspace calls SET_L2_TRAP. > > > > Add amdgpu_trap_program_kernel_vmids() which dispatches to a per-HW > > vmhub callback, and a new program_kernel_trap_vmids hook in > > amdgpu_vmhub_funcs for per-GFX-generation register writes. > > > > Required for: > > - RADV graphics debugging on Vega/Navi/Steam Deck (Valve request) > > - Consistent trap handler behavior when switching between kernel > > queues and user queues > > > > Suggested-by: Christian König > > <[email protected]<mailto:[email protected]>> > > Cc: Alexander Deucher > > <[email protected]<mailto:[email protected]>> > > Signed-off-by: Srinivasan Shanmugam > > <[email protected]<mailto:[email protected]>> > > Change-Id: I0709e788835b69d3d492864de8d5c2d36d0f08c8 > > --- > > drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 1 + > > drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c | 37 > ++++++++++++++++++++++++ > > drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h | 2 ++ > > 3 files changed, 40 insertions(+) > > > > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > > b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > > index 3ca187f5ade8..5624a5ab5c62 100644 > > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h > > @@ -115,6 +115,7 @@ struct amdgpu_vmhub_funcs { > > void (*print_l2_protection_fault_status)(struct amdgpu_device *adev, > > uint32_t status); > > uint32_t (*get_invalidate_req)(unsigned int vmid, uint32_t > > flush_type); > > + void (*program_kernel_trap_vmids)(struct amdgpu_device *adev); > > }; > > > > struct amdgpu_vmhub { > > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > > b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > > index 31f653ec3fb1..0e0aeea0aa2d 100644 > > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c > > @@ -256,8 +256,23 @@ int amdgpu_trap_init(struct amdgpu_device *adev) > > > > memcpy(ptr, trap_info->isa_buf, trap_info->isa_sz); > > > > + /* > > + * Device-level TMA for kernel queue VMIDs. Pinned GTT — not subject > > + * to eviction. Zero-initialized by OS: second-level handler address > > + * is 0 until userspace calls SET_L2_TRAP. > > How is the exclusivity maintained as a user app doesn't 'own' kernel queue? > How is > the conflict of different user apps trying to install their own second level > handler on a > kernel queue handled? As pointed out by Alex: - Each process gets its own per-VM TMA buffer for kernel queues - It is allocated and mapped at a fixed VA in the process's GPUVM when the device is opened, similar to amdgpu_map_static_csa() - Before each job is dispatched to a kernel queue VMID, the driver programs SQ_SHADER_TMA to that process's own TMA VA This way: - App A submits job → TMA = App A's TMA → job runs - App B submits job → TMA = App B's TMA → job runs - No conflict — each process has its own TMA buffer Note: SET_L2_TRAP for kernel queue VMIDs is not part of this series. This series only installs the first-level trap handler. For second-level handler support on kernel queues, I think since: - Each process has its own per-VM TMA buffer (allocated at device open, mapped at a fixed VA in the process's GPUVM) - User calls SET_L2_TRAP → writes second-level handler address into that process's own TMA buffer - When shader crashes → first-level handler reads from that process's TMA → jumps to that process's second-level handler - No conflict — each process has its own TMA with its own second-level handler address The per-VM TMA design is the foundation for this future series. The device-level kq_tma_bo will be removed in v2. For long term — kernel queues are replaced by user queues entirely , where MES already handles this correctly via ADD_QUEUE. Thanks, Srini
