On Tue, Sep 22, 2026 at 09:35:29AM +0200, Christian KKKnig wrote:
> On 9/22/26 08:08, Zhu, Lingshan wrote:
> > On 9/1/2026 4:50 PM, Zhu, Lingshan wrote:
> > 
> >> On 8/31/2026 4:21 PM, Christian König wrote:
> >>
> >>> On 8/28/26 17:59, Zhu, Lingshan wrote:
> >>>> On 8/28/2026 9:08 PM, Christian König wrote:
> >>>>
> >>>>> On 8/28/26 11:53, Zhu Lingshan wrote:
> >>>>>> This commit introduces a new helper
> >>>>>> amdgpu_lookup_queue_by_doorbell which helps
> >>>>>> look up a user queue with the given doorbell id
> >>>>>> in a xarray.
> >>>>>>
> >>>>>> This function takes a kref of the user space queue.
> >>>> Hello Christian
> >>>>
> >>>> Thanks for your comments.
> >>>>
> >>>>> Well absolutely clear NAK to the whole approach.
> >>>>>
> >>>>> This is the nonsense Sunil and I have worked quite hard to remove and 
> >>>>> we certainly shouldn't repeat such mistakes.
> >>>>>
> >>>>> When the userq needs to be used from interrupt context we need to hold 
> >>>>> the xa_lock_irqsave() or otherwise we don't have any guarantee that the 
> >>>>> userq, userq_mgr or associated fpriv went out of scope.
> >>>> Holding the spin lock by xa_lock_irqsave() can surely avoid racing with 
> >>>> the destruction process, however, it does not apply to all scenarios, 
> >>>> for example, you can not hold spin lock in mes_userq_reset_queue(),
> >>>> because it calls either amdgpu_mes_reset_queue_mmio or 
> >>>> amdgpu_mes_reset_queue_mmio, both of them acquire the MES mutex through 
> >>>> amdgpu_mes_lock.
> >>> Yeah which is exactly the reason why mes_userq_reset_queue() should *NOT* 
> >>> be called from non IOCTL context.
> >> I think it is not about whether called from IOCTL, it is a common racing 
> >> we should fix, and holding a kref is a low haning fruit.
> >>
> >>>> Another thing, out of the topic is, holding xa_lock does not guarantee 
> >>>> fpriv/userq_mgr alive, for example, when drm_device->unplugged is true, 
> >>>> all amdgpu teardown paths in amdgpu_drm_release are skipped,
> >>>> and the fpriv/userq_mgr is freed, no matter whether holding the xa spin 
> >>>> lock.
> >>> That would clearly be a massive bug. Those objects still need to be 
> >>> cleaned up independent of device hot plug.
> >> I agree, when unplugged == true, means can not access any HW registers, so 
> >> this bug deserve another series to fix.
> >>
> >>>> So IMHO since we have userq->kref, lets use it to maintain the lifecycle 
> >>>> of the queues. 
> >>>>
> >>>>> Grabbing references from this side would obviously result in circle 
> >>>>> dependencies.
> >>>> I am not sure, we should use the lock/unlock and kref_put/get in pairs 
> >>>> in sequence, can you name some circle dependencies or AB-BA lockings as 
> >>>> examples?
> >>> That is not AB-BA locking, but circle dependencies. E.g. A reference B, B 
> >>> referencing C, C referencing A again.
> >> A kref is an atomic counter, not a lock, that means we can try hold more 
> >> than one kref in a thread, and a kref not
> >> depend on another. As far as I can see, current amdgpu driver does not 
> >> have such problems, do you see any occurrences?
> > 
> > Hello Christian,
> > 
> > I have not heard back from you for three weeks, I wonder whether you have 
> > any further comments for this series,
> > or shall I make any improvements?
> 
> Well just completely drop that series. As I said this approach is a 
> fundamental no-go from my side.
> 
> As far as I can see we have solved the problems at hand and the rules how to 
> handle the user queues should be pretty clear by now.
> 

Assuming every kref_get() has a matching kref_put(), and the queue object
itself does not participate in a circular ownership chain, I don't
immediately see how taking a temporary queue reference would cause
correctness issues.

Could you please point us to the historical issue or the corresponding
commits that led to removing this approach previously?

Thanks,
Ray

> Regards,
> Christian.
> 
> > 
> > Thanks
> > Lingshan 
> > 
> >> Thanks
> >> Lingshan
> >>
> >>> Regards,
> >>> Christian.
> >>>
> >>>> Thanks
> >>>> Lingshan
> >>>>
> >>>>> Regards,
> >>>>> Christian.
> >>>>>
> >>>>>> Signed-off-by: Zhu Lingshan <[email protected]>
> >>>>>> ---
> >>>>>>  drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c | 30 +++++++++++++++++++++++
> >>>>>>  drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h |  2 ++
> >>>>>>  2 files changed, 32 insertions(+)
> >>>>>>
> >>>>>> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c 
> >>>>>> b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
> >>>>>> index 0a816b3c5ff9..e0639f844a8e 100644
> >>>>>> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
> >>>>>> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
> >>>>>> @@ -609,6 +609,36 @@ struct amdgpu_usermode_queue 
> >>>>>> *amdgpu_userq_get(struct amdgpu_userq_mgr *uq_mgr,
> >>>>>>        return queue;
> >>>>>>  }
> >>>>>>  
> >>>>>> +/**
> >>>>>> + * amdgpu_lookup_queue_by_doorbell - look up a user queue by doorbell
> >>>>>> + * @xa: user queue XArray indexed by doorbell
> >>>>>> + * @doorbell: doorbell index
> >>>>>> + *
> >>>>>> + * Return: A queue with the doorbell indexed, or NULL if no such a 
> >>>>>> queue found.
> >>>>>> + *
> >>>>>> + * This function increases kref of the queue, the caller
> >>>>>> + * must release the reference with amdgpu_userq_put().
> >>>>>> + */
> >>>>>> +struct amdgpu_usermode_queue *
> >>>>>> +amdgpu_lookup_queue_by_doorbell(struct xarray *xa, u32 doorbell)
> >>>>>> +{
> >>>>>> +      struct amdgpu_usermode_queue *queue;
> >>>>>> +      unsigned long flags;
> >>>>>> +
> >>>>>> +      xa_lock_irqsave(xa, flags);
> >>>>>> +      queue = xa_load(xa, doorbell);
> >>>>>> +      if (!queue)
> >>>>>> +              goto out_unlock;
> >>>>>> +
> >>>>>> +      if (!kref_get_unless_zero(&queue->refcount))
> >>>>>> +              queue = NULL;
> >>>>>> +
> >>>>>> +out_unlock:
> >>>>>> +      xa_unlock_irqrestore(xa, flags);
> >>>>>> +
> >>>>>> +      return queue;
> >>>>>> +}
> >>>>>> +
> >>>>>>  void amdgpu_userq_put(struct amdgpu_usermode_queue *queue)
> >>>>>>  {
> >>>>>>        if (queue)
> >>>>>> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h 
> >>>>>> b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h
> >>>>>> index 6412a7f7b6ef..8fc73862f64e 100644
> >>>>>> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h
> >>>>>> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h
> >>>>>> @@ -151,6 +151,8 @@ struct amdgpu_db_info {
> >>>>>>  };
> >>>>>>  
> >>>>>>  struct amdgpu_usermode_queue *amdgpu_userq_get(struct 
> >>>>>> amdgpu_userq_mgr *uq_mgr, u32 qid);
> >>>>>> +struct amdgpu_usermode_queue *
> >>>>>> +amdgpu_lookup_queue_by_doorbell(struct xarray *xa, u32 doorbell);
> >>>>>>  void amdgpu_userq_put(struct amdgpu_usermode_queue *queue);
> >>>>>>  
> >>>>>>  int amdgpu_userq_ioctl(struct drm_device *dev, void *data, struct 
> >>>>>> drm_file *filp);
> 

Reply via email to