On 8/31/26 13:39, Timur Kristóf wrote:
> On 2026. augusztus 28., péntek 6:47:34 közép-európai nyári idő Arunpravin
> Paneer Selvam wrote:
>> Clear-on-release only runs on VRAM, which amdgpu_ttm_map_buffer() reaches
>> via its direct MC address without programming a GART window, yet the wipe
>> still forces a VMID 0 flush.
>
> Makes sense.
> We don't need the VM flush when we are not changing the page tables.
>
> I agree with the patch, just would like to ask a few questions to better
> understand the underlying problem:
>
>> On GFX11 (e.g. Navi33) that spurious SDMA
>> flush can wedge the engine
>
> What is happening when the SDMA engine is wedged?
As far as Arun has investigate the UTCL1 request queue (which is part of the
memory interface of the SDMA) is in a deadlock, but we haven't quite figured
out why yet.
> Can it be recovered by an
> SDMA queue reset?
Most likely no. The UTCL1 is the translation and memory request queue between
SDMA and the core memory hub. To reset that one you need to reset both ends and
the core memory hub usually needs a full ASIC reset for that.
> Is it just a hang, or can it cause other issues such as page faults?
Good question we honestly don't know at this point. The HW guys need to find
the root cause first.
>> only flush when a GART window is actually used.
>
> Does that mean that there is still a risk of the wedge when the GART windows
> are used?
Yes, and that is actually not limited to the GART windows. It looks like every
time we map something into any VM it can happen that the SDMA crashes when
there are concurrent operations ongoing.
It's just that the GART flushes triggered by the SDMA made that scenario much
more likely than anything else.
> Can you remind me when/why we need the GART windows exactly?
Basically every time we want to copy something from system memory to VRAM with
the kernel.
Regards,
Christian.
>
>> Fixes: a68c7eaa7a8f ("drm/amdgpu: Enable clear page functionality")
>> Cc: [email protected]
>> Cc: Christian König <[email protected]>
>> Signed-off-by: Arunpravin Paneer Selvam <[email protected]>
>
> Reviewed-by: Timur Kristóf <[email protected]>
>
> Thanks and best regards,
> Timur
>
>> ---
>> drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.c | 5 ++++-
>> 1 file changed, 4 insertions(+), 1 deletion(-)
>>
>> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.c
>> b/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.c index
>> 6c07cee8e8777..2e6c98c2f1efa 100644
>> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.c
>> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.c
>> @@ -2578,6 +2578,7 @@ int amdgpu_ttm_clear_buffer(struct
>> amdgpu_ttm_buffer_entity *entity, struct amdgpu_device *adev =
>> amdgpu_ttm_adev(bo->tbo.bdev);
>> struct dma_fence *fence = NULL;
>> struct amdgpu_res_cursor dst;
>> + bool vm_needs_flush;
>> int r;
>>
>> if (!entity)
>> @@ -2585,6 +2586,8 @@ int amdgpu_ttm_clear_buffer(struct
>> amdgpu_ttm_buffer_entity *entity,
>>
>> amdgpu_res_first(bo->tbo.resource, 0, amdgpu_bo_size(bo), &dst);
>>
>> + vm_needs_flush = bo->tbo.resource->start ==
> AMDGPU_BO_INVALID_OFFSET;
>> +
>> mutex_lock(&entity->lock);
>> while (dst.remaining) {
>> struct dma_fence *next;
>> @@ -2605,7 +2608,7 @@ int amdgpu_ttm_clear_buffer(struct
>> amdgpu_ttm_buffer_entity *entity,
>>
>> r = amdgpu_ttm_fill_mem(adev, entity,
>> 0, to, cur_size, resv,
>> - &next, true,
> k_job_id);
>> + &next, vm_needs_flush,
> k_job_id);
>> if (r)
>> goto error;
>
>
>
>