Tsukikage7 opened a new issue, #3562:
URL: https://github.com/apache/dubbo-go/issues/3562

   **What would you like to be added**:
   
   ## Summary / 概要
   
   I would like to volunteer to coordinate and implement most of the remaining 
work needed to complete dubbo-go observability across Metrics, OpenTelemetry 
Tracing, structured logging, metadata diagnostics, end-to-end 
samples/documentation, and integration with dubbo-admin.
   
   我希望主推进 dubbo-go 剩余的大部分可观测性工作,包括统一语义、Metrics、OpenTelemetry 
Trace、结构化日志、Metadata 诊断、端到端 Samples/Docs,以及后续与 dubbo-admin 的联动。
   
   This is an implementation and delivery plan under the existing observability 
umbrella issue #3292. It does not replace #3292, #3337, #3338, or #3356, and it 
does not take ownership away from contributors already working on open PRs. The 
purpose is to make the remaining scope explicit, coordinate parallel work, 
avoid duplicate code, and deliver an end-to-end operational experience.
   
   ## Relationship to existing work
   
   | Area | Existing issue/PR | Current boundary |
   | --- | --- | --- |
   | Overall observability | #3292 | Existing umbrella and capability 
inventory; this issue proposes the complete execution plan. |
   | Metrics | #3337 | Coordinate before implementing overlapping metric names 
or labels. |
   | OpenTelemetry diagnostics | #3338 | Preserve assigned work and implement 
only agreed remaining parts. |
   | OTel span names and base attributes | #3551 | Open PR by @jiaming2li; 
review and build on it instead of duplicating it. |
   | Metadata observability | #3356 | Existing assigned issue; coordinate with 
its contributors. |
   | Metadata mapping metrics/errors | #3463 | Open PR by @Rachit-Gandhi; 
review and build on it instead of duplicating it. |
   | Community ownership context | Discussion #3534 and #3374 | Weekly notes 
list observability ownership and the plan to connect with Admin. |
   | AI-ready observability | Discussion #3454 | Future extension; should not 
block the base observability closure. |
   
   ## Desired outcome
   
   A user should be able to move through one coherent troubleshooting path:
   
   ```text
   Prometheus/Grafana signal
     -> affected service/method/error category
     -> OpenTelemetry trace
     -> correlated structured logs
     -> RPC, registry, metadata, routing, retry, timeout, rejection, or codec 
root cause
   ```
   
   The same operation or failure should use compatible semantics across 
Metrics, Trace, and Log. High-cardinality diagnostic context belongs in 
traces/logs, while Prometheus labels remain bounded and stable.
   
   ## Workstream 1: shared observability semantic contract
   
   I would like to own the cross-signal contract and compatibility rules.
   
   - [ ] Inventory existing metric names, labels, units, and cardinality.
   - [ ] Inventory current span names, semantic attributes, events, error 
recording, and propagation.
   - [ ] Inventory CtxLogger, zap, logrus, access-log, and trace-correlation 
behavior.
   - [ ] Inventory metadata mapping/report/RPC/revision/cache/fallback 
diagnostic points.
   - [ ] Publish a field mapping table for Metrics, Trace, Log, and Metadata.
   - [ ] Define a shared error taxonomy: `error_type`, `error_code`, and 
`error_source`.
   - [ ] Define low-cardinality and high-cardinality boundaries.
   - [ ] Map standard fields to OpenTelemetry semantic conventions where 
available.
   - [ ] Use a stable `dubbo.*` namespace only for Dubbo-specific trace 
attributes.
   - [ ] Document backward-compatibility rules for existing metric names, 
configuration, and logger behavior.
   
   Candidate shared dimensions:
   
   - `application_name`, `application_version`
   - `side`, `protocol`, `interface`, `method`, `group`, `version`
   - `error_type`, `error_code`, `error_source`
   - `metadata_operation`, `metadata_source`, `metadata_result`
   - `metadata_storage_type`, `metadata_revision`, `service_key`, 
`registry_id`, `cache_hit`
   
   `trace_id`, `span_id`, revision values, request IDs, full URLs, exception 
messages, payloads, and attachment values must not become default Prometheus 
labels.
   
   ## Workstream 2: shared RPC error taxonomy and structured logging
   
   This is a candidate first implementation area because it can connect the 
existing signals without duplicating #3551 or #3463.
   
   - [ ] Extract a stable error classification API from current RPC metrics 
logic.
   - [ ] Cover timeout, cancellation, limit/rejection, service unavailable, 
business failure, network failure, codec/serialization failure, and unknown 
failure.
   - [ ] Define structured mappings for Triple/gRPC and Dubbo protocol errors.
   - [ ] Avoid independent error-string parsing in Metrics, Trace, and Log.
   - [ ] Align zap and logrus structured fields.
   - [ ] Verify `trace_id`, `span_id`, and `trace_flags` behavior in CtxLogger.
   - [ ] Verify that trace integration does not break existing logger controls 
such as dynamic log-level updates.
   - [ ] Align `record-error-to-span` with the shared error taxonomy.
   
   ## Workstream 3: Metrics completion and standardization
   
   I would like to implement most of the remaining Metrics work after 
coordinating with the current contributors on #3337.
   
   - [ ] Document built-in metric names, labels, units, and cardinality.
   - [ ] Keep the existing metrics event bus and Prometheus adapter.
   - [ ] Define a stable provider/consumer RED model.
   - [ ] Extend RPC error classification to Dubbo protocol, network, and codec 
failures where possible.
   - [ ] Classify registry, metadata, and config-center metrics into basic and 
detailed layers.
   - [ ] Keep detailed/high-cardinality metrics opt-in.
   - [ ] Verify Prometheus pull, Pushgateway push, graceful shutdown, and 
zombie-metric cleanup.
   - [ ] Update PromQL examples and Grafana dashboards.
   - [ ] Document compatibility impact for renamed or deprecated metrics.
   
   Important open decision: the `side` label may duplicate information already 
encoded in metric names and can approximately double RPC time-series count. The 
contract should measure this cost before adding it by default, as already 
raised in #3337.
   
   ## Workstream 4: OpenTelemetry Trace completion
   
   PR #3551 should remain the source of truth for its current span-name and 
base-attribute changes. I would like to implement or coordinate the remaining 
tracing work after it is reviewed/merged.
   
   - [ ] Review #3551 against the shared semantic contract.
   - [ ] Add stable `error_type`, `error_code`, and `error_source` attributes.
   - [ ] Use `RecordError` and span status consistently.
   - [ ] Add low-noise diagnostics for retries, timeouts, rejection, codec 
failures, registry lookups, and metadata failures.
   - [ ] Add propagation tests for W3C Trace Context, B3, baggage, and Dubbo 
attachment carriers.
   - [ ] Cover consumer-to-provider propagation for Triple and Dubbo protocol 
paths.
   - [ ] Verify sampling and stdout/Jaeger/Zipkin/OTLP HTTP/OTLP gRPC exporters.
   - [ ] Document the OpenTracing-to-Otel migration boundary.
   
   Important open architecture decision from #3338: retry/load-balance/router 
selection currently occurs above the OTel filter span. We should agree whether 
to keep a single upper logical consumer span before implementing those 
diagnostic events, so we do not create incompatible span topology.
   
   ## Workstream 5: Metadata observability completion
   
   PR #3463 should remain the source of truth for the metadata mapping metrics 
and MetadataError code it already implements. I would like to complete the 
remaining diagnostic path after coordinating with its contributors.
   
   - [ ] Review #3463 against the shared operation/source/result/error contract.
   - [ ] Cover mapping listen, revision calculation, cache hit/miss, and 
fallback.
   - [ ] Cover metadata-report loading, RPC metadata loading, URL construction, 
and revision mismatch.
   - [ ] Keep revision, service key, registry ID, and similar diagnostic values 
in Trace/Log by default.
   - [ ] Keep Metrics limited to bounded operation/source/result/error 
dimensions.
   - [ ] Add failure-path, fallback, cache, and concurrency tests.
   
   This workstream does not include redesigning metadata lifecycle, revision 
algorithms, or mapping consistency.
   
   ## Workstream 6: end-to-end samples, documentation, and Admin integration
   
   - [ ] Add a runnable end-to-end observability example in dubbo-go-samples.
   - [ ] Demonstrate a Metrics -> Trace -> Log -> root-cause workflow.
   - [ ] Provide Prometheus/Grafana and Jaeger/Tempo examples.
   - [ ] Document Loki/Alloy-style log collection without coupling those 
backends to core.
   - [ ] Define which stable metrics and diagnostic links dubbo-admin can 
consume.
   - [ ] Coordinate trace deep links, log correlation, and metadata diagnostics 
with dubbo-admin.
   - [ ] Keep AI-specific telemetry such as TTFT, stream chunk interval, token 
counts, fallback frequency, and cost attribution as a follow-up extension.
   
   ## Non-goals
   
   - Replacing the metrics event bus with OpenTelemetry Metrics.
   - Introducing OpenTelemetry Logs SDK into core.
   - Hard-coding Grafana, Loki, Tempo, Jaeger, Alloy, or another backend into 
core.
   - Redesigning metadata lifecycle or registry architecture.
   - Adding unbounded values to default Prometheus labels.
   - Reimplementing code already present in #3551 or #3463.
   - Delivering all changes in one large PR.
   
   ## Collaboration and conflict-avoidance rules
   
   - Before each implementation PR, check Apache main and all related open PRs 
again.
   - List overlapping issues, PRs, files, and contributors in every 
implementation PR.
   - Prefer reviewing or extending an existing PR when the same files and 
behavior are already being changed.
   - After an upstream PR merges, synchronize main before continuing dependent 
work.
   - Split work by independently reviewable concern.
   - Keep code, tests, samples/docs, Admin integration, and release status 
separately verifiable.
   - Do not claim existing assigned issues without confirmation from their 
current contributors and maintainers.
   
   ## Proposed PR sequence
   
   1. Shared semantic contract and current-state audit.
   2. Shared RPC error taxonomy.
   3. Logger structured-field and correlation consistency.
   4. Metrics names/labels/error-classification completion.
   5. Trace error/events/propagation completion.
   6. Remaining metadata observability path.
   7. End-to-end Prometheus/Grafana -> Trace -> Log sample.
   8. Admin integration contract and documentation.
   
   The actual split should be adjusted after maintainers and the contributors 
on #3337, #3338/#3551, and #3356/#3463 confirm the boundaries.
   
   ## Acceptance criteria
   
   - [ ] Maintainers and current contributors have confirmed the major 
ownership boundaries.
   - [ ] A shared semantic contract is documented and reused by implementation 
code.
   - [ ] Metrics, Trace, and Log do not maintain conflicting error taxonomies.
   - [ ] Metrics names/labels/units/cardinality are documented and bounded.
   - [ ] Trace behavior covers stable span semantics, error recording, 
diagnostic events, and propagation tests.
   - [ ] Logger correlation is consistent across zap/logrus without breaking 
existing controls.
   - [ ] Metadata mapping/report/RPC/revision/cache/fallback has a diagnosable 
path.
   - [ ] At least one end-to-end Metrics -> Trace -> Log/Metadata 
troubleshooting scenario is verified.
   - [ ] Compatibility and migration impact is documented.
   - [ ] Changes are delivered through small, independently reviewable PRs with 
explicit tests.
   
   ## Request to maintainers and current contributors
   
   I would like to take primary responsibility for coordinating and 
implementing most of the remaining work described above. Before coding 
overlapping areas, please help confirm:
   
   1. Whether this execution plan is suitable as a meta issue under #3292.
   2. Which remaining workstreams I can formally own.
   3. Which parts should stay with the current assignees/contributors.
   4. Whether the proposed Admin integration boundary belongs in this effort.
   5. Whether the proposed PR sequence is appropriate.
   
   Once the boundaries are confirmed, I will update this checklist with the 
agreed owners and start from the shared semantic contract and the first 
non-overlapping implementation slice.
   
   **Why is this needed**:
   
   dubbo-go already has substantial observability capabilities, but they are 
not yet expressed as one stable operational contract. Metrics, Trace, Log, and 
Metadata diagnostics can expose different dimensions and error semantics, 
making production troubleshooting harder and increasing the risk of 
high-cardinality metrics or incompatible follow-up implementations.
   
   A coordinated implementation plan is needed to reuse the work already in 
progress, prevent duplicate PRs, keep compatibility visible, and deliver a 
complete troubleshooting workflow rather than isolated observability features.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to