nzw921rx opened a new issue, #12086: URL: https://github.com/apache/seatunnel/issues/12086
### Code of Conduct - [x] I agree to follow this project's [Code of Conduct](https://www.apache.org/foundation/policies/conduct) ### Search before asking - [x] I had searched in the [issues](https://github.com/apache/seatunnel/issues?q=is%3Aissue+label%3A%22bug%22) and found no similar issues. ### Describe the proposal [STIP-32](https://github.com/apache/seatunnel/issues/11598) established a reproducible benchmark and diagnostics framework for SeaTunnel. The next stage is to use that framework to continuously discover, investigate, and optimize real production bottlenecks. This umbrella issue tracks benchmark-driven performance investigations and defines the expected evidence standard for performance optimization work, especially for Zeta engine hot paths, checkpoint/state storage, WAL, queue/concurrency paths, serialization, recovery, and other performance-sensitive runtime code. The purpose is to make performance optimization **evidence-driven rather than intuition-driven**. Benchmark execution and diagnostics are documented in the [Zeta benchmark guide](https://seatunnel.apache.org/docs/engines/zeta/benchmark). > **No evidence, no performance claim.** > > A performance optimization should prove both the bottleneck and the improvement. --- ## Why this standard is needed A benchmark score changing after a code modification is not sufficient by itself to establish a valid optimization. For example: ```text Benchmark anomaly ↓ Inspect code ↓ Find something that looks expensive ↓ Change several paths ↓ Benchmark becomes faster ``` This only establishes correlation. It does not necessarily prove: * which production path caused the original problem; * whether the problem came from CPU, blocking/IO, lock contention, allocation/GC, benchmark fixture, or execution environment; * which individual change produced the improvement; * whether average latency improved while variance/CV remained unchanged; * whether throughput improved by trading away durability, correctness, memory, tail latency, or failure semantics. For engine/runtime optimization work, we should establish a reproducible causal chain instead. --- ## Expected investigation flow ```text Benchmark symptom ↓ Reproduce baseline ↓ Select appropriate profiler(s) ↓ Identify production call chain / hotspot ↓ Form root-cause hypothesis ↓ Make the smallest targeted change ↓ Equivalent before/after benchmark ↓ Before/after profiling ↓ Correctness / durability / failure verification ↓ Conclusion ``` --- ## Performance optimization evidence standard ### 1. Reproduce the original problem Before changing production code, reproduce the reported behavior under controlled conditions. Record the relevant baseline metrics, for example: * Score / mean latency / throughput * Error / confidence interval * Coefficient of variation (CV) * Allocation (`B/op`, allocation rate) * GC count/time * Latency distribution or other workload-specific metrics when relevant Baseline and candidate results should use the same: * machine / runner * JDK * JVM options * JMH parameters * benchmark parameters * runtime/storage configuration whenever possible. Absolute results collected from different GitHub-hosted runner CPU models should not be treated as an equivalent before/after comparison. --- ### 2. Profile the original implementation Choose diagnostics based on the symptom instead of mechanically running every profiler. Typical examples: * CPU hotspot → CPU profile * Blocking / IO / sleep / park → Wall profile * Lock contention → Lock profile plus CPU/Wall evidence * Allocation / GC pressure → GC/allocation profile * JVM/runtime behavior requiring deeper analysis → JFR * Storage latency → Wall/storage call-chain evidence plus latency/variance measurement * Atomic/CAS hot path → CPU profile plus focused benchmark where appropriate The objective is to identify the actual production call chain responsible for the measured cost or variance. The goal is not: ```text Run every profiler ``` The goal is: ```text Observed symptom ↓ Form hypothesis ↓ Choose the appropriate diagnostic tool ↓ Collect sufficient evidence ``` --- ### 3. Establish the root-cause hypothesis The investigation should clearly connect the profiling evidence to the benchmark symptom. Finding an expensive-looking method in the code is not enough. For example, if a change claims to reduce CV, the evidence should explain why that path causes **variance**, rather than only average latency. Important distinctions include: ```text high average latency != high latency variance high CPU != lock contention high Wall time != high CPU usage allocation != retained-memory growth ``` A suspected root cause should remain a hypothesis until profiling and controlled measurement support it. --- ### 4. Prefer the smallest targeted production change A performance experiment should change the confirmed bottleneck first. Avoid mixing several unrelated optimizations into the same experiment unless each one has independent evidence and measurable attribution. Otherwise, even if the final benchmark improves, reviewers cannot determine which change produced the improvement. For example: ```text CV: 20% → 8% ``` is not sufficient if the same PR simultaneously changes: ```text WAL sync RequestFuture semantics exception handling allocation behavior serialization timeout behavior ``` because the improvement cannot be attributed to one confirmed cause. If multiple independent problems are discovered, prefer separate follow-up issues/PRs where practical. --- ### 5. Provide equivalent before/after measurements Every performance optimization PR should provide an equivalent before/after comparison under the same environment. Depending on the issue, report at least the relevant subset of: | Metric | Before | After | Change | | --------------------------------- | -----: | ----: | -----: | | Score / throughput / mean latency | | | | | Error / confidence interval | | | | | CV | | | | | Allocation (`B/op`) | | | | | Allocation rate | | | | | GC count/time | | | | The comparison should answer both: 1. Did performance improve? 2. Did the original reported symptom improve? For example: ```text Mean latency: 300 us/op → 180 us/op ``` does not automatically prove that: ```text CV: 20% → ? ``` has improved. Lower average latency alone does not prove that a high-variance problem has been resolved. --- ### 6. Profile again after the change When profiling was used to establish the bottleneck, run the same profiler again after the optimization. The previously identified hotspot should: * disappear; * shrink materially; * move to another expected path; * or otherwise show evidence consistent with the claimed optimization. Benchmark improvement alone demonstrates correlation. Before/after profiling helps establish causality. A complete performance investigation should ideally demonstrate: ```text Before: business path ↓ confirmed hotspot ↓ large CPU / Wall / allocation contribution After: same business path ↓ hotspot significantly reduced ↓ benchmark symptom also improves ``` --- ### 7. Verify correctness and runtime semantics independently Performance must not be improved by weakening behavior. For engine/runtime paths, verify relevant semantics such as: * checkpoint correctness * WAL/state durability * recovery behavior * concurrency/thread-safety guarantees * timeout/failure propagation * queue lifecycle/backpressure behavior * serialization compatibility * state consistency * public/runtime semantics affected by the change Performance and correctness should be treated as separate acceptance requirements. For semantics-sensitive changes such as: ```text hsync hflush WAL completion checkpoint persistence recovery concurrency failure handling ``` tests should cover the actual execution path whose semantics are being changed. For example, local-file visibility alone should not be treated as proof of HDFS durability semantics. --- ### 8. Keep investigation conclusions in the PR until they are established facts One investigation hypothesis should not immediately become permanent benchmark documentation or JavaDoc. Prefer documenting the initial evidence and conclusions in the issue/PR first. Long-term benchmark documentation should be updated when the behavior has been sufficiently established and future contributors need to understand it as a stable characteristic of the workload or runtime. The preferred progression is: ```text Investigation hypothesis ↓ Profiling evidence ↓ Controlled before/after verification ↓ Confirmed conclusion ↓ Long-term documentation if necessary ``` rather than: ```text Initial hypothesis ↓ Immediately document as permanent fact ``` --- ## Expected performance PR structure A performance optimization PR should make the following causal chain easy for reviewers to verify: ```text 1. Observed benchmark symptom 2. Baseline results 3. Profiling evidence 4. Confirmed production hotspot 5. Root-cause explanation 6. Targeted code change 7. Equivalent before/after results 8. Before/after profiling evidence 9. Correctness / durability / failure verification 10. Final conclusion ``` The PR should contain the evidence needed for review and reproduction. Reviewers should not have to run the missing benchmark or profiler work in order to determine whether the author's stated root cause is valid. --- ## Review standard by optimization type The required evidence should be proportional to the risk of the change. ### Focused low-risk micro optimization Usually sufficient: * focused benchmark * equivalent before/after results * correctness coverage Example: ```text small isolated utility single implementation change no runtime semantic change ``` --- ### Runtime hot-path optimization Normally expected: * benchmark baseline * relevant profiling evidence * root-cause explanation * targeted production change * equivalent before/after benchmark * correctness coverage * before/after profiling when needed to establish causality Examples: ```text queue serialization metrics scheduler hot collection operations atomic/CAS paths ``` --- ### Semantics-sensitive engine optimization For areas such as: ```text checkpoint WAL state storage recovery concurrency failure handling timeout semantics ``` normally expected: * complete performance evidence chain * before/after profiling where relevant * equivalent benchmark comparison * explicit correctness verification * durability/recovery verification where applicable * failure-semantics verification where applicable These areas should have a higher acceptance standard because a performance improvement may otherwise silently change runtime guarantees. --- ## What is not required This standard does **not** mean every performance PR must run: ```text CPU + Wall + Lock + GC + JFR ``` The objective is not to maximize the amount of profiling output. The objective is to collect: > **the minimum sufficient evidence that proves the claimed causal relationship.** For example: ```text CPU-bound problem → CPU profile may be sufficient IO/blocking problem → Wall profile may be more useful than CPU Allocation problem → GC/allocation profile may be the primary evidence Concurrency problem → CPU + Lock/Wall may be required ``` The profiler should be selected according to the problem. --- ## Relationship with STIP-32 [STIP-32](https://github.com/apache/seatunnel/issues/11598) established the measurement and diagnostic infrastructure. This umbrella tracks the next stage: ```text STIP-32 ↓ Reliable benchmark ↓ Detect abnormal behavior ↓ Performance investigation ↓ Profiling ↓ Root cause ↓ Production optimization ↓ Equivalent before/after verification ↓ New baseline ``` ### Task list - [ ] https://github.com/apache/seatunnel/issues/12058 — Investigate and optimize checkpoint state-store latency variance - [ ] https://github.com/apache/seatunnel/issues/12059 — Investigate and improve intermediate queue benchmark stability - [ ] https://github.com/apache/seatunnel/issues/12060 — Investigate and improve finished JobDAG store benchmark stability - [ ] https://github.com/apache/seatunnel/issues/12061 — Investigate and improve IMap job-storage benchmark stability - [ ] https://github.com/apache/seatunnel/issues/12063 — Investigate and improve job lifecycle growth benchmark stability - [ ] https://github.com/apache/seatunnel/issues/12074 — Challenge: Optimize Debezium JSON serialization and deserialization performance ### Are you willing to submit PR? - [ ] Yes I am willing to submit a PR! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
