nzw921rx opened a new issue, #12058:
URL: https://github.com/apache/seatunnel/issues/12058

   ## Description
   
   This is a focused investigation and optimization task for the latency 
variance observed in checkpoint state-store benchmarks.
   
   It is not a performance regression report and does not attempt to compare 
the performance of Java 8 with Java 11.
   
   The benchmark results show significant sample-to-sample variance within 
individual runs:
   
   | Benchmark | Java 8 CV | Java 11 CV |
   | --- | ---: | ---: |
   | `checkpointIdAtomicIncrement` | 14.40% | 12.71% |
   | `checkpointOverviewIncrementalUpdate` | 20.88% | 19.72% |
   
   For `checkpointOverviewIncrementalUpdate`, the observed results cover a wide 
range:
   
   - Java 8: approximately `198–440 us/op`
   - Java 11: approximately `247–454 us/op`
   
   The upper end is approximately 1.8–2.2 times the lower end.
   
   Benchmark run:
   
   https://github.com/apache/seatunnel/actions/runs/33722870811
   
   The Java 8 and Java 11 benchmarks ran on different GitHub Actions runner CPU 
models, so their absolute scores should not be compared directly. However, 
similar within-run variance appears on both runners and is worth investigating.
   
   The goals of this issue are to:
   
   1. Identify which part of the measured path causes the variance.
   2. Determine whether it comes from the production implementation, benchmark 
fixture, or execution environment.
   3. Implement a focused optimization if profiling confirms an optimizable 
production bottleneck.
   4. Preserve checkpoint correctness and state-store durability.
   
   ## Relevant benchmarks
   
   The investigation should focus on:
   
   ```text
   CheckpointStorageBenchmark.checkpointIdAtomicIncrement
   CheckpointStorageBenchmark.checkpointOverviewIncrementalUpdate
   ```
   
   The following benchmark may be used as an end-to-end reference:
   
   ```text
   CheckpointStorageBenchmark.checkpointPersistenceTransaction
   ```
   
   ## Profiling tools
   
   SeaTunnel provides the `Benchmarks Diagnostics` workflow for running one 
exact benchmark method with CPU, wall-clock, lock, GC, or JFR profiling:
   
   
https://github.com/apache/seatunnel/actions/workflows/benchmarks_diagnostics.yml
   
   The benchmark can also be run directly from the GitHub Actions page by 
selecting **Benchmarks Diagnostics**, clicking **Run workflow**, and providing 
the target branch, tag, commit SHA, or trusted PR number together with the 
exact benchmark method.
   
   The generated artifacts include JMH results, profiling summaries, flame 
graphs, and optional JFR recordings.
   
   ## Local profiling
   
   Build the benchmark module:
   
   ```bash
   ./mvnw -Pbenchmark -pl seatunnel-benchmarks -am -DskipTests package
   ```
   
   Run CPU profiling:
   
   ```bash
   bash tools/benchmarks/profile_benchmarks.sh profile cpu \
     --repository . \
     --benchmark 
'CheckpointStorageBenchmark.checkpointOverviewIncrementalUpdate$'
   ```
   
   Replace `cpu` with `wall`, `lock`, or `gc` to run another profiler:
   
   ```bash
   bash tools/benchmarks/profile_benchmarks.sh profile wall \
     --repository . \
     --benchmark 
'CheckpointStorageBenchmark.checkpointOverviewIncrementalUpdate$'
   ```
   
   Capture a JFR recording:
   
   ```bash
   bash tools/benchmarks/profile_benchmarks.sh capture jfr \
     --repository . \
     --benchmark 
'CheckpointStorageBenchmark.checkpointOverviewIncrementalUpdate$'
   ```
   
   The same commands can be used for:
   
   ```text
   CheckpointStorageBenchmark.checkpointIdAtomicIncrement$
   ```
   
   CPU, wall-clock, and lock profiling require `ASYNC_PROFILER_HOME`. The 
GitHub Actions workflow installs the required profiler automatically.
   
   ## Expected outcome
   
   The investigation should provide:
   
   - Profiling evidence identifying the source of the variance.
   - A conclusion on whether the source is the production implementation, 
benchmark fixture, or execution environment.
   - Before-and-after benchmark results if an optimization is implemented.
   - Results collected with the same JDK, runner, benchmark arguments, and 
state-store configuration.
   - Confirmation that checkpoint correctness and durability remain unchanged.
   
   The implementation should remain focused on the confirmed source of the 
variance. Unrelated state-store refactoring is outside the scope of this issue.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to