zmuxuny opened a new issue, #6194:
URL: https://github.com/apache/rocketmq-dashboard/issues/6194

   ### Before Creating the Bug Report
   
   - [x] I have searched the [open 
issues](https://github.com/apache/rocketmq-dashboard/issues) of this repository 
and believe that this is not a duplicate.
   
   - [x] This is a defect in RocketMQ Studio, not a usage question and not a 
defect in another Apache RocketMQ repository.
   
   - [x] I can reproduce this on the current `master` branch, or I have stated 
the exact version I am running below.
   
   
   ### Studio Version
   
   rocketmq-studio@5e4c39b053c35a5598cc79ff331b8a54e649d7b1
   
   ### Runtime Environment
   
   Java/Maven regression with actual Tencent provider, collector and native 
rule-test logic; mocked SDK responses, repositories, registry and client 
factory.
   
   ### Connected RocketMQ Cluster
   
   Synthetic Tencent group with retained ClusterIdV4; no cloud account 
contacted.
   
   ### Build Toolchain
   
   _No response_
   
   ### Describe the Bug
   
   ## Problem
   
   A cloud consumer group's known cluster scope is lost when its progress 
lookup fails. As a result, testing a native lag rule with that cluster selected 
produces an empty sample list instead of the group's UNAVAILABLE diagnostic.
   
   On `rocketmq-studio` at `5e4c39b0`, successful samples in 
`CloudRocketMqBusinessMetricsCollector` carry `group.getClusterId()`. The three 
samples in the per-group exception path are instead created with null 
clusterId. `NativeAlertRuleTestService` filters those samples using 
`NativeAlertRuleScopeMatcher`, so a rule selecting the group's actual cluster 
excludes the failure samples.
   
   This is observable with actual Tencent provider mapping: 
`ConsumeGroupItem.clusterIdV4` becomes `ConsumerGroupVO.clusterId`, including 
migrated groups that retain their original cluster identity. The group is known 
even when its subsequent `DescribeTopicListByGroup` request fails.
   
   ## Reproduction and regression evidence
   
   A synthetic test exercises the real `TencentInstanceProvider` -> 
`CloudRocketMqBusinessMetricsCollector` -> `NativeAlertRuleTestService` logic 
with mocked SDK responses, repositories, registry and client factory:
   
   1. DescribeConsumerGroupList returns group `orders` with ClusterIdV4 
`legacy-cluster`.
   2. DescribeTopicListByGroup throws a progress-read failure.
   3. Test a numeric native rule scoped to `legacy-cluster` and `orders`.
   
   Actual: `samples=[]`. Expected: one UNAVAILABLE sample, null numeric value, 
and the known consumer-group identity. The same failure remains visible when 
the cluster selector is omitted.
   
   Against unchanged production code, `TencentCloudLagFailureScopeTest` ran 5 
tests: the 3 scoped cases (consumer total lag, queue maximum lag, topic 
backlog) failed, and 2 compatibility cases passed. No errors. The tests 
explicitly verify the real provider retains ClusterIdV4 before testing the 
collector.
   
   Equivalent command: `cd server && mvn -B -ntp 
-Dtest=TencentCloudLagFailureScopeTest test`. For the reproduction, previously 
built trunk dependencies were reused, the unchanged collector and new test were 
compiled with javac, then Maven `surefire:test` was run. No cloud account was 
contacted.
   
   ## Proposed correction and limits
   
   Preserve the known group clusterId in the three per-group failure samples. 
Keep whole-instance failure markers unscoped with empty labels; preserve null 
provider clusterIds and successful measurements.
   
   Credit: yyqdbngt's unmerged, stale-closed PR #4674 identified the missing 
scope. Its original alert-resolution claim does not apply to current trunk: 
current reconciliation already preserves numeric alert state using unavailable 
group labels, and business lag rules do not accept the UNAVAILABLE operator. 
This follow-up is deliberately limited to accurate diagnostic visibility and 
sample identity, not a claim that it fixes false RESOLVED events or missing 
unavailable-condition alerts.
   
   This is separate from #6184, which corrects successful cloud aggregate 
measurements being misreported as per-queue maxima. This issue concerns 
exception-path cluster identity.
   
   
   ### Steps to Reproduce
   
   DescribeConsumerGroupList returns orders in legacy-cluster; 
DescribeTopicListByGroup throws; test a numeric native lag rule selecting 
legacy-cluster and orders. Detailed fixture and test command above.
   
   ### What Did You Expect to See?
   
   One UNAVAILABLE diagnostic with null numeric value and the known group 
cluster identity.
   
   ### What Did You See Instead?
   
   samples=[] for cluster-scoped tests; the same failure is visible without a 
cluster selector. Three scoped regressions fail; two compatibility cases pass.
   
   ### Additional Context
   
   AI assistance was used. Credit yyqdbngt / #4674 for the original 
missing-scope fix; current-trunk impact is limited as explained above.
   
   ### Are You Willing to Submit a Pull Request?
   
   - [x] Yes, I am willing to submit a pull request.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to