p-szucs opened a new pull request, #8747:
URL: https://github.com/apache/hadoop/pull/8747

   <!--
     Thanks for sending a pull request!
       1. If this is your first time, please read our contributor guidelines: 
https://cwiki.apache.org/confluence/display/HADOOP/How+To+Contribute
       2. Make sure your PR title starts with JIRA issue id, e.g., 
'HADOOP-17799. Your PR title ...'.
   -->
   
   ### Description of PR
   **Fix possible deadlock when using legacy dynamic queues**
   
   The deadlock happens when we are using ManagedParentQueues with periodically 
calling the scheduling monitor, that uses 
GuaranteedOrZeroCapacityOverTimePolicy. In this case a refresQueues operation 
can collide with this thread if the timing is unfortunate, because both 
operations lock the policy and the ManagedParentQueue as well, but in a 
different order.
   
   **SchedulingMonitor:**
   
   - locks the policy in GuaranteedOrZeroCapacityOverTimePolicy's 
updateLeafQueueState()
   - then iterates through the child queues, locking the ManagedParentQueue
   
   **RefreshQueues operation:**
   
   - locks the queue by calling ManagedParentQueue.reinitialize
   - locks the policy by calling 
GuaranteedOrZeroCapacityOverTimePolicy.computeQueueManagementChanges - 
updateTemplateAbsoluteCapacities
   
   Because of the reversed order of the locks, with an unfortunate timing it 
can happen that both threads can acquire their first lock at the same time, and 
after that they are waiting on each other to release those.
   
   The fix changes the order of the locking, so both threads locks the queue 
first, then the monitoring policy, so the deadlock can not happen.
   
   Contains content generated by Claude
   
   ### How was this patch tested?
   Tested in a cluster with a ManagedParentQueue in the queue structure, and 
frequent scheduling monitoring, also requesting queue configuration changes 
frequently. This way I could reproduce the deadlock consistently.
   
   During my testing I could not reproduce the deadlock with the fix.
   
   ### For code changes:
   
   - [x] Does the title of this PR start with the corresponding JIRA issue id 
(e.g. 'HADOOP-17799. Your PR title ...')?
   - [ ] Object storage: Have the integration tests been executed and the 
endpoint
         declared according to the connector-specific documentation? *Note: 
Automated CI
         testing doesn't cover all cases so manual testing with cloud storage 
is still
         required.*
   - [ ] If adding new dependencies to the code, are these dependencies 
licensed in a way that is compatible for inclusion under [ASF 
2.0](http://www.apache.org/legal/resolved.html#category-a)?
   - [ ] If applicable, have you updated the `LICENSE`, `LICENSE-binary`, 
`NOTICE-binary` files?
   
   ### AI Tooling
   
   If an AI tool was used:
   
   - [x] The PR includes the phrase "Contains content generated by <tool>"
         where <tool> is the name of the AI tool used.
   - [x] My use of AI contributions follows the ASF legal policy
         https://www.apache.org/legal/generative-tooling.html
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to