AMD General

> -----Original Message-----
> From: Koenig, Christian <[email protected]>
> Sent: Thursday, September 3, 2026 12:39 PM
> To: Lazar, Lijo <[email protected]>; Deucher, Alexander
> <[email protected]>; SHANMUGAM, SRINIVASAN
> <[email protected]>
> Cc: [email protected]
> Subject: Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler
> infrastructure
>
> On 9/3/26 06:03, Lazar, Lijo wrote:
> >
> >
> > On 03-Sep-26 9:32 AM, Lazar, Lijo wrote:
> >>
> >>
> >> On 02-Sep-26 10:25 PM, Deucher, Alexander wrote:
> >>> Public
> >>>
> >>>
> >>> For kernel queues each IB executes with a kernel provided vmid assigned
> dynamically by the kernel driver.
> >>>
> >>
> >> In this case, it's a device level TMA 'kq_tma_bo' for first level. That 
> >> address is
> programmed in SQ registers. When a job is submitted, the second level handler 
> is
> picked from what is programmed in kq_tma_bo. Do you mean to say that driver 
> will
> change that value dynamically based on what is provided by user?
> >
> > Do you mean to say that driver will change that value dynamically based on 
> > what
> is provided by user for each job submission?
>
> Yeah I agree with Lijo, something doesn't adds up here.
>
> As far as I know the same register value is used for both graphics and all 
> compute
> queues at the same time, so changing this dynamically on each submission won't
> work (at unless we complete isolate the applications).
>
> If I'm not completely mistaken we either need allocate a BO per VM and always
> map it at the same location or give the location to userspace so that 
> userspace so
> that the UMD can map it.
>
> I don't think we have discussed what the actual plan for that would be.


Agreed — we cannot reprogram SQ_SHADER_TMA per job submission. Since it is a 
per-VMID register, it covers all queues in that process at the same time. 
Swapping it per job would corrupt other running apps.

For v2, can we do something like this?:

1. Userspace allocates a TMA BO with VM_ALWAYS_VALID so it is never evicted or 
moved
2. Userspace passes the TMA GPU VA to the driver via the VM ioctl
3. Driver validates that the BO has VM_ALWAYS_VALID, then programs 
SQ_SHADER_TMA once
4. Before programming, driver evicts all queues, flushes TLB, writes the 
register, then restores queues
5. TBA and TMA GPU VA are stored in amdgpu_vm and cleared when VM is destroyed 
or CLEAR is called

For kernel queues:

Can we reuse the same userspace-registered TMA BO for kernel queues as well? 
Since there is only one SQ_SHADER_TMA register per VMID, a separate 
kernel-internal BO would conflict with what userspace registered. The userspace 
trap handler shader also needs to find its own TMA to work correctly.

If a kernel queue is created before userspace registers the TMA, the driver 
simply waits — second-level handler activates only after userspace calls the VM 
ioctl. First-level handler continues to work in the meantime.

Should the TMA BO live in VRAM or GTT? VRAM gives faster GPU access during a 
trap. GTT is easier to manage (a trap handler runs infrequently (only on shader 
exception) and a trap can fire anytime under memory pressure, and a evicted TMA 
is worse than a slow TMA.)

Best regards,
Srini
.

Reply via email to