Metastarx opened a new issue, #11305:
URL: https://github.com/apache/rocketmq/issues/11305

   ### Background
   
   `LmqPrefixIndex.forEachLmqByPrefix()` 
(broker/src/main/java/org/apache/rocketmq/broker/lite/LmqPrefixIndex.java:78-96)
 takes `rwLock.readLock()` at line 85 and then runs the caller's visitor at 
line 90 while the lock is still held; the lock is only released in the 
`finally` block after the whole traversal finishes.
   
   The visitor is caller code that can do unbounded work. 
`AbstractLiteLifecycleManager.forEachLiteTopicByPrefix` (154-163) passes a 
lambda that calls `getMaxOffsetInQueue(lmqName)` for every entry, and 
`LiteEventDispatcher.doFullDispatchForWildcardGroup` (272-284) runs 
`ConsumerOffsetManager.queryOffset()`, subscription lookups, event-queue offers 
and long-poll wakeups for every matched lmq. A wildcard group matches every lmq 
of a parent topic, so the read lock is held for the entire scan instead of a 
short copy of a few keys.
   
   Two consequences:
   
   1. `onLmqCreate` (94-95) and `onLmqDelete` (101-102) need the write lock, 
and `onLmqCreate` is reached from `NotifyMessageArrivingListener.arriving` on 
the store's `ReputMessageService` thread. While a wildcard dispatch is 
iterating, the reput thread blocks on the write lock, so consume-queue 
dispatch, long-poll notification and pop triggers stop advancing for every 
topic on the broker. `ReentrantReadWriteLock` is non-fair by default, so a 
stream of prefix scans can additionally starve the writer.
   2. The lock cannot be upgraded, so any callback that transitively calls 
`add()`/`remove()` on the same thread requests the write lock while holding the 
read lock and hangs forever. `cleanByParentTopic` (224-230) already carries a 
collect-then-delete workaround with a "nesting causes deadlock" comment, and 
the javadoc of `AbstractLiteLifecycleManager` (139, 152) documents the same 
restriction, so the invariant is enforced only by convention.
   
   ### How to reproduce
   
   Unit level: build an `LmqPrefixIndex`, add two lmq names under one parent 
topic, and call `forEachLmqByPrefix(prefix, name -> { index.remove(name); 
return true; })` from a single thread. The call never returns.
   
   Load level: start a broker with a lite parent topic that holds many lmqs, 
register a wildcard lite group, and trigger a full dispatch. `jstack` on the 
reput thread shows it parked in `LmqPrefixIndex.add` waiting for the write 
lock, and `dispatchBehindBytes` keeps growing.
   
   ### Expected
   
   A traversal should work on a snapshot so that caller code never runs while 
the index lock is held.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to