AMD General > -----Original Message----- > From: Koenig, Christian <[email protected]> > Sent: Thursday, September 3, 2026 12:39 PM > To: Lazar, Lijo <[email protected]>; Deucher, Alexander > <[email protected]>; SHANMUGAM, SRINIVASAN > <[email protected]> > Cc: [email protected] > Subject: Re: [RFC PATCH 1/2] drm/amdgpu: Add kernel VMID trap handler > infrastructure > > On 9/3/26 06:03, Lazar, Lijo wrote: > > > > > > On 03-Sep-26 9:32 AM, Lazar, Lijo wrote: > >> > >> > >> On 02-Sep-26 10:25 PM, Deucher, Alexander wrote: > >>> Public > >>> > >>> > >>> For kernel queues each IB executes with a kernel provided vmid assigned > dynamically by the kernel driver. > >>> > >> > >> In this case, it's a device level TMA 'kq_tma_bo' for first level. That > >> address is > programmed in SQ registers. When a job is submitted, the second level handler > is > picked from what is programmed in kq_tma_bo. Do you mean to say that driver > will > change that value dynamically based on what is provided by user? > > > > Do you mean to say that driver will change that value dynamically based on > > what > is provided by user for each job submission? > > Yeah I agree with Lijo, something doesn't adds up here. > > As far as I know the same register value is used for both graphics and all > compute > queues at the same time, so changing this dynamically on each submission won't > work (at unless we complete isolate the applications). > > If I'm not completely mistaken we either need allocate a BO per VM and always > map it at the same location or give the location to userspace so that > userspace so > that the UMD can map it. > > I don't think we have discussed what the actual plan for that would be.
Agreed — we cannot reprogram SQ_SHADER_TMA per job submission. Since it is a per-VMID register, it covers all queues in that process at the same time. Swapping it per job would corrupt other running apps. For v2, can we do something like this?: 1. Userspace allocates a TMA BO with VM_ALWAYS_VALID so it is never evicted or moved 2. Userspace passes the TMA GPU VA to the driver via the VM ioctl 3. Driver validates that the BO has VM_ALWAYS_VALID, then programs SQ_SHADER_TMA once 4. Before programming, driver evicts all queues, flushes TLB, writes the register, then restores queues 5. TBA and TMA GPU VA are stored in amdgpu_vm and cleared when VM is destroyed or CLEAR is called For kernel queues: Can we reuse the same userspace-registered TMA BO for kernel queues as well? Since there is only one SQ_SHADER_TMA register per VMID, a separate kernel-internal BO would conflict with what userspace registered. The userspace trap handler shader also needs to find its own TMA to work correctly. If a kernel queue is created before userspace registers the TMA, the driver simply waits — second-level handler activates only after userspace calls the VM ioctl. First-level handler continues to work in the meantime. Should the TMA BO live in VRAM or GTT? VRAM gives faster GPU access during a trap. GTT is easier to manage (a trap handler runs infrequently (only on shader exception) and a trap can fire anytime under memory pressure, and a evicted TMA is worse than a slow TMA.) Best regards, Srini .
