Public
> -----Original Message-----
> From: Lazar, Lijo <[email protected]>
> Sent: Wednesday, September 2, 2026 8:59 PM
> To: SHANMUGAM, SRINIVASAN <[email protected]>;
> Koenig, Christian <[email protected]>; Deucher, Alexander
> <[email protected]>
> Cc: [email protected]
> Subject: Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler
> infrastructure
>
>
>
> On 02-Sep-26 8:36 PM, Srinivasan Shanmugam wrote:
> > MES owns kernel queue VMIDs (1..first_kfd_vmid-1) but does not program
> > SQ_SHADER_TBA/TMA registers for them. Add infrastructure to let the
> > driver program the first-level CWSR trap handler for these VMIDs
> > directly via SRBM select.
> >
> > Add kq_tma_bo — a pinned GTT BO used as device-level TMA scratch for
> > kernel queue VMIDs. Unlike per-process TMA (created in
> > amdgpu_trap_alloc), this is device-level and lives for the lifetime of
> > the device. It is zero-initialized by the OS: the second-level handler
> > address is 0 until userspace calls SET_L2_TRAP.
> >
> > Add amdgpu_trap_program_kernel_vmids() which dispatches to a per-HW
> > vmhub callback, and a new program_kernel_trap_vmids hook in
> > amdgpu_vmhub_funcs for per-GFX-generation register writes.
> >
> > Required for:
> > - RADV graphics debugging on Vega/Navi/Steam Deck (Valve request)
> > - Consistent trap handler behavior when switching between kernel
> > queues and user queues
> >
> > Suggested-by: Christian König <[email protected]>
> > Cc: Alexander Deucher <[email protected]>
> > Signed-off-by: Srinivasan Shanmugam <[email protected]>
> > Change-Id: I0709e788835b69d3d492864de8d5c2d36d0f08c8
> > ---
> > drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h | 1 +
> > drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c | 37
> ++++++++++++++++++++++++
> > drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h | 2 ++
> > 3 files changed, 40 insertions(+)
> >
> > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
> > b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
> > index 3ca187f5ade8..5624a5ab5c62 100644
> > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
> > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
> > @@ -115,6 +115,7 @@ struct amdgpu_vmhub_funcs {
> > void (*print_l2_protection_fault_status)(struct amdgpu_device *adev,
> > uint32_t status);
> > uint32_t (*get_invalidate_req)(unsigned int vmid, uint32_t
> > flush_type);
> > + void (*program_kernel_trap_vmids)(struct amdgpu_device *adev);
> > };
> >
> > struct amdgpu_vmhub {
> > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
> > b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
> > index 31f653ec3fb1..0e0aeea0aa2d 100644
> > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
> > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
> > @@ -256,8 +256,23 @@ int amdgpu_trap_init(struct amdgpu_device *adev)
> >
> > memcpy(ptr, trap_info->isa_buf, trap_info->isa_sz);
> >
> > + /*
> > + * Device-level TMA for kernel queue VMIDs. Pinned GTT — not subject
> > + * to eviction. Zero-initialized by OS: second-level handler address
> > + * is 0 until userspace calls SET_L2_TRAP.
>
> How is the exclusivity maintained as a user app doesn't 'own' kernel queue?
> How is
> the conflict of different user apps trying to install their own second level
> handler on a
> kernel queue handled?
As pointed out by Alex:
- Each process gets its own per-VM TMA buffer for kernel queues
- It is allocated and mapped at a fixed VA in the process's GPUVM
when the device is opened, similar to amdgpu_map_static_csa()
- Before each job is dispatched to a kernel queue VMID, the driver
programs SQ_SHADER_TMA to that process's own TMA VA
This way:
- App A submits job → TMA = App A's TMA → job runs
- App B submits job → TMA = App B's TMA → job runs
- No conflict — each process has its own TMA buffer
Note: SET_L2_TRAP for kernel queue VMIDs is not part of this series.
This series only installs the first-level trap handler.
For second-level handler support on kernel queues, I think since:
- Each process has its own per-VM TMA buffer (allocated at device
open, mapped at a fixed VA in the process's GPUVM)
- User calls SET_L2_TRAP → writes second-level handler address
into that process's own TMA buffer
- When shader crashes → first-level handler reads from that
process's TMA → jumps to that process's second-level handler
- No conflict — each process has its own TMA with its own
second-level handler address
The per-VM TMA design is the foundation for this future series.
The device-level kq_tma_bo will be removed in v2.
For long term — kernel queues are replaced by user queues entirely
, where MES already handles this correctly via ADD_QUEUE.
Thanks,
Srini