On 03-Aug-26 3:05 PM, Jesse Zhang wrote:
Creating a user queue can race with a GPU reset. While recovery holds
reset_domain->sem for write, MES is unresponsive, so the ADD_QUEUE from
amdgpu_userq_map_helper() times out (-110) and an otherwise valid queue
create fails:

   amdgpu: MES(0) failed to respond to msg=ADD_QUEUE
   [drm:mes_userq_map [amdgpu]] *ERROR* Failed to map queue in HW, err (-110)
   amdgpu: [drm] *ERROR* ... Failed to map Queue
   amdgpu: [drm] *ERROR* ... Failed to create usermode queue

Take reset_domain->sem for read around the map so it runs only once MES
is back up. This mirrors amdgpu_userq_cleanup() and honors the
userq_mutex -> reset_domain->sem order; the reset path never takes
userq_mutex, so there is no deadlock.

Signed-off-by: Jesse Zhang <[email protected]>
---
  drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c | 5 +++++
  1 file changed, 5 insertions(+)

diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c 
b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
index 652599f08990..770635ab5298 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
@@ -745,7 +745,12 @@ amdgpu_userq_create(struct drm_file *filp, union 
drm_amdgpu_userq *args)
        if (!adev->userq_halt_for_enforce_isolation ||
            ((queue->queue_type != AMDGPU_HW_IP_GFX) &&
             (queue->queue_type != AMDGPU_HW_IP_COMPUTE))) {
+               /* Serialize the map against an in-progress GPU reset (MES is
+                * unresponsive during recovery), matching 
amdgpu_userq_cleanup().
+                */
+               down_read(&adev->reset_domain->sem);

If a reset is ongoing and a userq operation comes in the middle, there is no need to wait till reset semaphore is available. The ongoing process i going to terminate sooner or later after the reset.

This should only try to acquire the lock and return an error if not able to.

Thanks,
Lijo

                r = amdgpu_userq_map_helper(queue);
+               up_read(&adev->reset_domain->sem);
                if (r) {
                        drm_file_err(uq_mgr->file, "Failed to map Queue\n");
                        trace_amdgpu_userq_create_end(queue, r);

Reply via email to