nzw921rx opened a new issue, #12086:
URL: https://github.com/apache/seatunnel/issues/12086

   ### Code of Conduct
   
   - [x] I agree to follow this project's [Code of 
Conduct](https://www.apache.org/foundation/policies/conduct)
   
   
   ### Search before asking
   
   - [x] I had searched in the 
[issues](https://github.com/apache/seatunnel/issues?q=is%3Aissue+label%3A%22bug%22)
 and found no similar issues.
   
   
   ### Describe the proposal
   
   [STIP-32](https://github.com/apache/seatunnel/issues/11598) established a 
reproducible benchmark and diagnostics framework for SeaTunnel.
   
   The next stage is to use that framework to continuously discover, 
investigate, and optimize real production bottlenecks.
   
   This umbrella issue tracks benchmark-driven performance investigations and 
defines the expected evidence standard for performance optimization work, 
especially for Zeta engine hot paths, checkpoint/state storage, WAL, 
queue/concurrency paths, serialization, recovery, and other 
performance-sensitive runtime code.
   
   The purpose is to make performance optimization **evidence-driven rather 
than intuition-driven**.
   
   Benchmark execution and diagnostics are documented in the [Zeta benchmark 
guide](https://seatunnel.apache.org/docs/engines/zeta/benchmark).
   
   > **No evidence, no performance claim.**
   >
   > A performance optimization should prove both the bottleneck and the 
improvement.
   
   ---
   
   ## Why this standard is needed
   
   A benchmark score changing after a code modification is not sufficient by 
itself to establish a valid optimization.
   
   For example:
   
   ```text
   Benchmark anomaly
       ↓
   Inspect code
       ↓
   Find something that looks expensive
       ↓
   Change several paths
       ↓
   Benchmark becomes faster
   ```
   
   This only establishes correlation.
   
   It does not necessarily prove:
   
   * which production path caused the original problem;
   * whether the problem came from CPU, blocking/IO, lock contention, 
allocation/GC, benchmark fixture, or execution environment;
   * which individual change produced the improvement;
   * whether average latency improved while variance/CV remained unchanged;
   * whether throughput improved by trading away durability, correctness, 
memory, tail latency, or failure semantics.
   
   For engine/runtime optimization work, we should establish a reproducible 
causal chain instead.
   
   ---
   
   ## Expected investigation flow
   
   ```text
   Benchmark symptom
       ↓
   Reproduce baseline
       ↓
   Select appropriate profiler(s)
       ↓
   Identify production call chain / hotspot
       ↓
   Form root-cause hypothesis
       ↓
   Make the smallest targeted change
       ↓
   Equivalent before/after benchmark
       ↓
   Before/after profiling
       ↓
   Correctness / durability / failure verification
       ↓
   Conclusion
   ```
   
   ---
   
   ## Performance optimization evidence standard
   
   ### 1. Reproduce the original problem
   
   Before changing production code, reproduce the reported behavior under 
controlled conditions.
   
   Record the relevant baseline metrics, for example:
   
   * Score / mean latency / throughput
   * Error / confidence interval
   * Coefficient of variation (CV)
   * Allocation (`B/op`, allocation rate)
   * GC count/time
   * Latency distribution or other workload-specific metrics when relevant
   
   Baseline and candidate results should use the same:
   
   * machine / runner
   * JDK
   * JVM options
   * JMH parameters
   * benchmark parameters
   * runtime/storage configuration
   
   whenever possible.
   
   Absolute results collected from different GitHub-hosted runner CPU models 
should not be treated as an equivalent before/after comparison.
   
   ---
   
   ### 2. Profile the original implementation
   
   Choose diagnostics based on the symptom instead of mechanically running 
every profiler.
   
   Typical examples:
   
   * CPU hotspot → CPU profile
   * Blocking / IO / sleep / park → Wall profile
   * Lock contention → Lock profile plus CPU/Wall evidence
   * Allocation / GC pressure → GC/allocation profile
   * JVM/runtime behavior requiring deeper analysis → JFR
   * Storage latency → Wall/storage call-chain evidence plus latency/variance 
measurement
   * Atomic/CAS hot path → CPU profile plus focused benchmark where appropriate
   
   The objective is to identify the actual production call chain responsible 
for the measured cost or variance.
   
   The goal is not:
   
   ```text
   Run every profiler
   ```
   
   The goal is:
   
   ```text
   Observed symptom
       ↓
   Form hypothesis
       ↓
   Choose the appropriate diagnostic tool
       ↓
   Collect sufficient evidence
   ```
   
   ---
   
   ### 3. Establish the root-cause hypothesis
   
   The investigation should clearly connect the profiling evidence to the 
benchmark symptom.
   
   Finding an expensive-looking method in the code is not enough.
   
   For example, if a change claims to reduce CV, the evidence should explain 
why that path causes **variance**, rather than only average latency.
   
   Important distinctions include:
   
   ```text
   high average latency != high latency variance
   
   high CPU != lock contention
   
   high Wall time != high CPU usage
   
   allocation != retained-memory growth
   ```
   
   A suspected root cause should remain a hypothesis until profiling and 
controlled measurement support it.
   
   ---
   
   ### 4. Prefer the smallest targeted production change
   
   A performance experiment should change the confirmed bottleneck first.
   
   Avoid mixing several unrelated optimizations into the same experiment unless 
each one has independent evidence and measurable attribution.
   
   Otherwise, even if the final benchmark improves, reviewers cannot determine 
which change produced the improvement.
   
   For example:
   
   ```text
   CV: 20% → 8%
   ```
   
   is not sufficient if the same PR simultaneously changes:
   
   ```text
   WAL sync
   RequestFuture semantics
   exception handling
   allocation behavior
   serialization
   timeout behavior
   ```
   
   because the improvement cannot be attributed to one confirmed cause.
   
   If multiple independent problems are discovered, prefer separate follow-up 
issues/PRs where practical.
   
   ---
   
   ### 5. Provide equivalent before/after measurements
   
   Every performance optimization PR should provide an equivalent before/after 
comparison under the same environment.
   
   Depending on the issue, report at least the relevant subset of:
   
   | Metric                            | Before | After | Change |
   | --------------------------------- | -----: | ----: | -----: |
   | Score / throughput / mean latency |        |       |        |
   | Error / confidence interval       |        |       |        |
   | CV                                |        |       |        |
   | Allocation (`B/op`)               |        |       |        |
   | Allocation rate                   |        |       |        |
   | GC count/time                     |        |       |        |
   
   The comparison should answer both:
   
   1. Did performance improve?
   2. Did the original reported symptom improve?
   
   For example:
   
   ```text
   Mean latency:
   300 us/op → 180 us/op
   ```
   
   does not automatically prove that:
   
   ```text
   CV:
   20% → ?
   ```
   
   has improved.
   
   Lower average latency alone does not prove that a high-variance problem has 
been resolved.
   
   ---
   
   ### 6. Profile again after the change
   
   When profiling was used to establish the bottleneck, run the same profiler 
again after the optimization.
   
   The previously identified hotspot should:
   
   * disappear;
   * shrink materially;
   * move to another expected path;
   * or otherwise show evidence consistent with the claimed optimization.
   
   Benchmark improvement alone demonstrates correlation.
   
   Before/after profiling helps establish causality.
   
   A complete performance investigation should ideally demonstrate:
   
   ```text
   Before:
   business path
       ↓
   confirmed hotspot
       ↓
   large CPU / Wall / allocation contribution
   
   After:
   same business path
       ↓
   hotspot significantly reduced
       ↓
   benchmark symptom also improves
   ```
   
   ---
   
   ### 7. Verify correctness and runtime semantics independently
   
   Performance must not be improved by weakening behavior.
   
   For engine/runtime paths, verify relevant semantics such as:
   
   * checkpoint correctness
   * WAL/state durability
   * recovery behavior
   * concurrency/thread-safety guarantees
   * timeout/failure propagation
   * queue lifecycle/backpressure behavior
   * serialization compatibility
   * state consistency
   * public/runtime semantics affected by the change
   
   Performance and correctness should be treated as separate acceptance 
requirements.
   
   For semantics-sensitive changes such as:
   
   ```text
   hsync
   hflush
   WAL completion
   checkpoint persistence
   recovery
   concurrency
   failure handling
   ```
   
   tests should cover the actual execution path whose semantics are being 
changed.
   
   For example, local-file visibility alone should not be treated as proof of 
HDFS durability semantics.
   
   ---
   
   ### 8. Keep investigation conclusions in the PR until they are established 
facts
   
   One investigation hypothesis should not immediately become permanent 
benchmark documentation or JavaDoc.
   
   Prefer documenting the initial evidence and conclusions in the issue/PR 
first.
   
   Long-term benchmark documentation should be updated when the behavior has 
been sufficiently established and future contributors need to understand it as 
a stable characteristic of the workload or runtime.
   
   The preferred progression is:
   
   ```text
   Investigation hypothesis
       ↓
   Profiling evidence
       ↓
   Controlled before/after verification
       ↓
   Confirmed conclusion
       ↓
   Long-term documentation if necessary
   ```
   
   rather than:
   
   ```text
   Initial hypothesis
       ↓
   Immediately document as permanent fact
   ```
   
   ---
   
   ## Expected performance PR structure
   
   A performance optimization PR should make the following causal chain easy 
for reviewers to verify:
   
   ```text
   1. Observed benchmark symptom
   
   2. Baseline results
   
   3. Profiling evidence
   
   4. Confirmed production hotspot
   
   5. Root-cause explanation
   
   6. Targeted code change
   
   7. Equivalent before/after results
   
   8. Before/after profiling evidence
   
   9. Correctness / durability / failure verification
   
   10. Final conclusion
   ```
   
   The PR should contain the evidence needed for review and reproduction.
   
   Reviewers should not have to run the missing benchmark or profiler work in 
order to determine whether the author's stated root cause is valid.
   
   ---
   
   ## Review standard by optimization type
   
   The required evidence should be proportional to the risk of the change.
   
   ### Focused low-risk micro optimization
   
   Usually sufficient:
   
   * focused benchmark
   * equivalent before/after results
   * correctness coverage
   
   Example:
   
   ```text
   small isolated utility
   single implementation change
   no runtime semantic change
   ```
   
   ---
   
   ### Runtime hot-path optimization
   
   Normally expected:
   
   * benchmark baseline
   * relevant profiling evidence
   * root-cause explanation
   * targeted production change
   * equivalent before/after benchmark
   * correctness coverage
   * before/after profiling when needed to establish causality
   
   Examples:
   
   ```text
   queue
   serialization
   metrics
   scheduler
   hot collection operations
   atomic/CAS paths
   ```
   
   ---
   
   ### Semantics-sensitive engine optimization
   
   For areas such as:
   
   ```text
   checkpoint
   WAL
   state storage
   recovery
   concurrency
   failure handling
   timeout semantics
   ```
   
   normally expected:
   
   * complete performance evidence chain
   * before/after profiling where relevant
   * equivalent benchmark comparison
   * explicit correctness verification
   * durability/recovery verification where applicable
   * failure-semantics verification where applicable
   
   These areas should have a higher acceptance standard because a performance 
improvement may otherwise silently change runtime guarantees.
   
   ---
   
   ## What is not required
   
   This standard does **not** mean every performance PR must run:
   
   ```text
   CPU + Wall + Lock + GC + JFR
   ```
   
   The objective is not to maximize the amount of profiling output.
   
   The objective is to collect:
   
   > **the minimum sufficient evidence that proves the claimed causal 
relationship.**
   
   For example:
   
   ```text
   CPU-bound problem
   → CPU profile may be sufficient
   
   IO/blocking problem
   → Wall profile may be more useful than CPU
   
   Allocation problem
   → GC/allocation profile may be the primary evidence
   
   Concurrency problem
   → CPU + Lock/Wall may be required
   ```
   
   The profiler should be selected according to the problem.
   
   ---
   
   ## Relationship with STIP-32
   
   [STIP-32](https://github.com/apache/seatunnel/issues/11598) established the 
measurement and diagnostic infrastructure.
   
   This umbrella tracks the next stage:
   
   ```text
   STIP-32
       ↓
   Reliable benchmark
       ↓
   Detect abnormal behavior
       ↓
   Performance investigation
       ↓
   Profiling
       ↓
   Root cause
       ↓
   Production optimization
       ↓
   Equivalent before/after verification
       ↓
   New baseline
   ```
   
   
   
   ### Task list
   
   - [ ] https://github.com/apache/seatunnel/issues/12058 — Investigate and 
optimize checkpoint state-store latency variance
   - [ ] https://github.com/apache/seatunnel/issues/12059 — Investigate and 
improve intermediate queue benchmark stability
   - [ ] https://github.com/apache/seatunnel/issues/12060 — Investigate and 
improve finished JobDAG store benchmark stability
   - [ ] https://github.com/apache/seatunnel/issues/12061 — Investigate and 
improve IMap job-storage benchmark stability
   - [ ] https://github.com/apache/seatunnel/issues/12063 — Investigate and 
improve job lifecycle growth benchmark stability
   - [ ] https://github.com/apache/seatunnel/issues/12074 — Challenge: Optimize 
Debezium JSON serialization and deserialization performance
   
   ### Are you willing to submit PR?
   
   - [ ] Yes I am willing to submit a PR!


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to