Public

> -----Original Message-----
> From: Lazar, Lijo <[email protected]>
> Sent: Wednesday, September 2, 2026 8:59 PM
> To: SHANMUGAM, SRINIVASAN <[email protected]>;
> Koenig, Christian <[email protected]>; Deucher, Alexander
> <[email protected]>
> Cc: [email protected]
> Subject: Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler
> infrastructure
>
>
>
> On 02-Sep-26 8:36 PM, Srinivasan Shanmugam wrote:
> > MES owns kernel queue VMIDs (1..first_kfd_vmid-1) but does not program
> > SQ_SHADER_TBA/TMA registers for them. Add infrastructure to let the
> > driver program the first-level CWSR trap handler for these VMIDs
> > directly via SRBM select.
> >
> > Add kq_tma_bo — a pinned GTT BO used as device-level TMA scratch for
> > kernel queue VMIDs. Unlike per-process TMA (created in
> > amdgpu_trap_alloc), this is device-level and lives for the lifetime of
> > the device. It is zero-initialized by the OS: the second-level handler
> > address is 0 until userspace calls SET_L2_TRAP.
> >
> > Add amdgpu_trap_program_kernel_vmids() which dispatches to a per-HW
> > vmhub callback, and a new program_kernel_trap_vmids hook in
> > amdgpu_vmhub_funcs for per-GFX-generation register writes.
> >
> > Required for:
> >    - RADV graphics debugging on Vega/Navi/Steam Deck (Valve request)
> >    - Consistent trap handler behavior when switching between kernel
> >      queues and user queues
> >
> > Suggested-by: Christian König <[email protected]>
> > Cc: Alexander Deucher <[email protected]>
> > Signed-off-by: Srinivasan Shanmugam <[email protected]>
> > Change-Id: I0709e788835b69d3d492864de8d5c2d36d0f08c8
> > ---
> >   drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h  |  1 +
> >   drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c | 37
> ++++++++++++++++++++++++
> >   drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h |  2 ++
> >   3 files changed, 40 insertions(+)
> >
> > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
> > b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
> > index 3ca187f5ade8..5624a5ab5c62 100644
> > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
> > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
> > @@ -115,6 +115,7 @@ struct amdgpu_vmhub_funcs {
> >     void (*print_l2_protection_fault_status)(struct amdgpu_device *adev,
> >                                              uint32_t status);
> >     uint32_t (*get_invalidate_req)(unsigned int vmid, uint32_t
> > flush_type);
> > +   void (*program_kernel_trap_vmids)(struct amdgpu_device *adev);
> >   };
> >
> >   struct amdgpu_vmhub {
> > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
> > b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
> > index 31f653ec3fb1..0e0aeea0aa2d 100644
> > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
> > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
> > @@ -256,8 +256,23 @@ int amdgpu_trap_init(struct amdgpu_device *adev)
> >
> >     memcpy(ptr, trap_info->isa_buf, trap_info->isa_sz);
> >
> > +   /*
> > +    * Device-level TMA for kernel queue VMIDs. Pinned GTT — not subject
> > +    * to eviction. Zero-initialized by OS: second-level handler address
> > +    * is 0 until userspace calls SET_L2_TRAP.
>
> How is the exclusivity maintained as a user app doesn't 'own' kernel queue? 
> How is
> the conflict of different user apps trying to install their own second level 
> handler on a
> kernel queue handled?

As pointed out by Alex:

  - Each process gets its own per-VM TMA buffer for kernel queues
  - It is allocated and mapped at a fixed VA in the process's GPUVM
    when the device is opened, similar to amdgpu_map_static_csa()
  - Before each job is dispatched to a kernel queue VMID, the driver
    programs SQ_SHADER_TMA to that process's own TMA VA

This way:
  - App A submits job → TMA = App A's TMA → job runs
  - App B submits job → TMA = App B's TMA → job runs
  - No conflict — each process has its own TMA buffer

Note: SET_L2_TRAP for kernel queue VMIDs is not part of this series.
This series only installs the first-level trap handler.

For second-level handler support on kernel queues, I think since:

  - Each process has its own per-VM TMA buffer (allocated at device
    open, mapped at a fixed VA in the process's GPUVM)
  - User calls SET_L2_TRAP → writes second-level handler address
    into that process's own TMA buffer
  - When shader crashes → first-level handler reads from that
    process's TMA → jumps to that process's second-level handler
  - No conflict — each process has its own TMA with its own
    second-level handler address

The per-VM TMA design is the foundation for this future series.
The device-level kq_tma_bo will be removed in v2.

For long term — kernel queues are replaced by user queues entirely
, where MES already handles this correctly via ADD_QUEUE.

Thanks,
Srini

Reply via email to