zmuxuny opened a new issue, #6179: URL: https://github.com/apache/rocketmq-dashboard/issues/6179
### Before Creating the Bug Report - [x] I have searched the [open issues](https://github.com/apache/rocketmq-dashboard/issues) of this repository and believe that this is not a duplicate. - [x] This is a defect in RocketMQ Studio, not a usage question and not a defect in another Apache RocketMQ repository. - [x] I can reproduce this on the current `master` branch, or I have stated the exact version I am running below. ### Studio Version rocketmq-studio@5e4c39b0, verified 2026-10-10. ### Runtime Environment Focused JUnit tests use cloud provider aggregate row shapes and pass collected metrics through AlertRuleEvaluator and AlertStateMachine. Synthetic fixtures only; no real production incident or live cloud operation is claimed. ### Connected RocketMQ Cluster ALIYUN and TENCENT provider shapes, represented by mocks. No live cluster was changed. ### Build Toolchain _No response_ ### Describe the Bug CloudRocketMqBusinessMetricsCollector publishes consumer.lag.max_queue as AVAILABLE by taking the maximum of provider progress rows. Those rows are not queues: - AliyunConverters.toQueueProgressRows returns one row per topic, or a broker="total" fallback when only group totals exist. - TencentInstanceProvider.getGroupProgress returns one row per topic subscription. - Both mark broker/consumer offsets as QueueProgressVO.UNKNOWN_OFFSET; topic rows have synthetic broker="topic:<name>" and queueId=0. The native metric catalog labels the value "Consumer lag max queue". For example, a topic with two queues each at lag 40 has aggregate lag 80, so an alert for queue maximum >50 fires despite neither queue exceeding 50. This is an illustrative data-shape scenario, not a measured production incident. ### Steps to Reproduce Focused JUnit regressions use the provider's actual aggregate row shape and run the metric through AlertRuleEvaluator and AlertStateMachine: 1. ALIYUN topic aggregate 80: max-queue >50 evaluates true. 2. TENCENT topic aggregate 80: same false-positive condition. 3. Aliyun group-total fallback: queue maximum is AVAILABLE rather than UNSUPPORTED. 4. A later aggregate zero causes an existing FIRING queue-maximum alert to transition to RESOLVED without a queue-level measurement. Command: cd server && mvn -B -ntp -Dtest=CloudRocketMqBusinessMetricsCollectorTest test ### What Did You Expect to See? Emit existing MetricAvailability.UNSUPPORTED with null value and a diagnostic reason when the cloud API only supplies aggregate data. Retain this sample and its consumerGroup labels, metricKeys and catalog entries so existing alert state is preserved. Keep real consumer total and topic backlog unchanged; genuine group RPC failures continue to report UNAVAILABLE. ### What Did You See Instead? Against unchanged production code, 6 tests ran: 4 new regression instances failed, 2 existing tests passed, no errors. The failures demonstrate false-positive queue-threshold evaluation, incorrect AVAILABLE classification, and false resolution of an existing queue-maximum alert. ### Additional Context The evidence and proposed scope received independent code/test review. This requires no new API or metric. Duplicate check: all current open PRs and historical searches for CloudRocketMqBusinessMetricsCollector / max_queue cloud found no matching fix. Older collector coverage PRs #3325/#3444 do not cover aggregate granularity; #4674 addresses scope; #5068 concerns missing Tencent lag, a separate issue. AI assistance was used for source inspection and deterministic regression authoring. ### Are You Willing to Submit a Pull Request? - [x] Yes, I am willing to submit a pull request. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
