On 9/3/26 06:03, Lazar, Lijo wrote:
> 
> 
> On 03-Sep-26 9:32 AM, Lazar, Lijo wrote:
>>
>>
>> On 02-Sep-26 10:25 PM, Deucher, Alexander wrote:
>>> Public
>>>
>>>
>>> For kernel queues each IB executes with a kernel provided vmid assigned 
>>> dynamically by the kernel driver.
>>>
>>
>> In this case, it's a device level TMA 'kq_tma_bo' for first level. That 
>> address is programmed in SQ registers. When a job is submitted, the second 
>> level handler is picked from what is programmed in kq_tma_bo. Do you mean to 
>> say that driver will change that value dynamically based on what is provided 
>> by user?
> 
> Do you mean to say that driver will change that value dynamically based on 
> what is provided by user for each job submission?

Yeah I agree with Lijo, something doesn't adds up here.

As far as I know the same register value is used for both graphics and all 
compute queues at the same time, so changing this dynamically on each 
submission won't work (at unless we complete isolate the applications).

If I'm not completely mistaken we either need allocate a BO per VM and always 
map it at the same location or give the location to userspace so that userspace 
so that the UMD can map it.

I don't think we have discussed what the actual plan for that would be.

Regards,
Christian.

> 
> Thanks,
> Lijo
> 
>>
>> Thanks,
>> Lijo
>>
>>   > Alex
>>>
>>> *From:*Lazar, Lijo <[email protected]>
>>> *Sent:* Wednesday, September 2, 2026 12:12 PM
>>> *To:* SHANMUGAM, SRINIVASAN <[email protected]>; Koenig, 
>>> Christian <[email protected]>; Deucher, Alexander 
>>> <[email protected]>
>>> *Cc:* [email protected]
>>> *Subject:* Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler 
>>> infrastructure
>>>
>>> Public
>>>
>>> I'm not sure how this works, I thought the TMA mapping is per VMID and 
>>> kernel queues have static VMIDs.
>>>
>>> Thanks,
>>>
>>> Lijo
>>>
>>> ------------------------------------------------------------------------
>>>
>>> *From:*SHANMUGAM, SRINIVASAN <[email protected] 
>>> <mailto:[email protected]>>
>>> *Sent:* Wednesday, 02 September 2026 21:29:38
>>> *To:* Lazar, Lijo <[email protected] <mailto:[email protected]>>; Koenig, 
>>> Christian <[email protected] <mailto:[email protected]>>; 
>>> Deucher, Alexander <[email protected] 
>>> <mailto:[email protected]>>
>>> *Cc:* [email protected] <mailto:amd- 
>>> [email protected]> <[email protected] <mailto:amd- 
>>> [email protected]>>
>>> *Subject:* RE: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler 
>>> infrastructure
>>>
>>> Public
>>>
>>>> -----Original Message-----
>>>> From: Lazar, Lijo <[email protected] <mailto:[email protected]>>
>>>> Sent: Wednesday, September 2, 2026 8:59 PM
>>>> To: SHANMUGAM, SRINIVASAN <[email protected] 
>>>> <mailto:[email protected]>>;
>>>> Koenig, Christian <[email protected] 
>>>> <mailto:[email protected]>>; Deucher, 
>>> Alexander
>>>> <[email protected] <mailto:[email protected]>>
>>>> Cc: [email protected] <mailto:[email protected]>
>>>> Subject: Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler
>>>> infrastructure
>>>>
>>>>
>>>>
>>>> On 02-Sep-26 8:36 PM, Srinivasan Shanmugam wrote:
>>>> > MES owns kernel queue VMIDs (1..first_kfd_vmid-1) but does not program
>>>> > SQ_SHADER_TBA/TMA registers for them. Add infrastructure to let the
>>>> > driver program the first-level CWSR trap handler for these VMIDs
>>>> > directly via SRBM select.
>>>> >
>>>> > Add kq_tma_bo — a pinned GTT BO used as device-level TMA scratch for
>>>> > kernel queue VMIDs. Unlike per-process TMA (created in
>>>> > amdgpu_trap_alloc), this is device-level and lives for the lifetime of
>>>> > the device. It is zero-initialized by the OS: the second-level handler
>>>> > address is 0 until userspace calls SET_L2_TRAP.
>>>> >
>>>> > Add amdgpu_trap_program_kernel_vmids() which dispatches to a per-HW
>>>> > vmhub callback, and a new program_kernel_trap_vmids hook in
>>>> > amdgpu_vmhub_funcs for per-GFX-generation register writes.
>>>> >
>>>> > Required for:
>>>> >    - RADV graphics debugging on Vega/Navi/Steam Deck (Valve request)
>>>> >    - Consistent trap handler behavior when switching between kernel
>>>> >      queues and user queues
>>>> >
>>>> > Suggested-by: Christian König <[email protected] 
>>>> > <mailto:[email protected]>>
>>>> > Cc: Alexander Deucher <[email protected] 
>>>> > <mailto:[email protected]>>
>>>> > Signed-off-by: Srinivasan Shanmugam <[email protected] 
>>>> > <mailto:[email protected]>>
>>>> > Change-Id: I0709e788835b69d3d492864de8d5c2d36d0f08c8
>>>> > ---
>>>> >   drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h  |  1 +
>>>> >   drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c | 37
>>>> ++++++++++++++++++++++++
>>>> >   drivers/gpu/drm/amd/amdgpu/amdgpu_trap.h |  2 ++
>>>> >   3 files changed, 40 insertions(+)
>>>> >
>>>> > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
>>>> > b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
>>>> > index 3ca187f5ade8..5624a5ab5c62 100644
>>>> > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
>>>> > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.h
>>>> > @@ -115,6 +115,7 @@ struct amdgpu_vmhub_funcs {
>>>> >     void (*print_l2_protection_fault_status)(struct amdgpu_device *adev,
>>>> >                                              uint32_t status);
>>>> >     uint32_t (*get_invalidate_req)(unsigned int vmid, uint32_t
>>>> > flush_type);
>>>> > +   void (*program_kernel_trap_vmids)(struct amdgpu_device *adev);
>>>> >   };
>>>> >
>>>> >   struct amdgpu_vmhub {
>>>> > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
>>>> > b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
>>>> > index 31f653ec3fb1..0e0aeea0aa2d 100644
>>>> > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
>>>> > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_trap.c
>>>> > @@ -256,8 +256,23 @@ int amdgpu_trap_init(struct amdgpu_device *adev)
>>>> >
>>>> >     memcpy(ptr, trap_info->isa_buf, trap_info->isa_sz);
>>>> >
>>>> > +   /*
>>>> > +    * Device-level TMA for kernel queue VMIDs. Pinned GTT — not subject
>>>> > +    * to eviction. Zero-initialized by OS: second-level handler address
>>>> > +    * is 0 until userspace calls SET_L2_TRAP.
>>>>
>>>> How is the exclusivity maintained as a user app doesn't 'own' kernel 
>>>> queue? How is
>>>> the conflict of different user apps trying to install their own second 
>>>> level handler on a
>>>> kernel queue handled?
>>>
>>> As pointed out by Alex:
>>>
>>>    - Each process gets its own per-VM TMA buffer for kernel queues
>>>    - It is allocated and mapped at a fixed VA in the process's GPUVM
>>>      when the device is opened, similar to amdgpu_map_static_csa()
>>>    - Before each job is dispatched to a kernel queue VMID, the driver
>>>      programs SQ_SHADER_TMA to that process's own TMA VA
>>>
>>> This way:
>>>    - App A submits job → TMA = App A's TMA → job runs
>>>    - App B submits job → TMA = App B's TMA → job runs
>>>    - No conflict — each process has its own TMA buffer
>>>
>>> Note: SET_L2_TRAP for kernel queue VMIDs is not part of this series.
>>> This series only installs the first-level trap handler.
>>>
>>> For second-level handler support on kernel queues, I think since:
>>>
>>>    - Each process has its own per-VM TMA buffer (allocated at device
>>>      open, mapped at a fixed VA in the process's GPUVM)
>>>    - User calls SET_L2_TRAP → writes second-level handler address
>>>      into that process's own TMA buffer
>>>    - When shader crashes → first-level handler reads from that
>>>      process's TMA → jumps to that process's second-level handler
>>>    - No conflict — each process has its own TMA with its own
>>>      second-level handler address
>>>
>>> The per-VM TMA design is the foundation for this future series.
>>> The device-level kq_tma_bo will be removed in v2.
>>>
>>> For long term — kernel queues are replaced by user queues entirely
>>> , where MES already handles this correctly via ADD_QUEUE.
>>>
>>> Thanks,
>>> Srini
>>>
>>
> 

Reply via email to