nzw921rx commented on PR #12081: URL: https://github.com/apache/seatunnel/pull/12081#issuecomment-5535716662
I think we should complete the performance evidence chain before concluding that this issue has been resolved. For this kind of performance optimization, I suggest following this process: 1. Reproduce the original problem locally using the same JDK, JVM options, JMH parameters, state-store configuration, and machine environment. Record the baseline latency, Error, CV, and GC/allocation metrics if relevant. 2. Profile the original implementation using CPU, Wall, Lock, GC, or JFR as appropriate. We should identify the exact production call chain responsible for the observed latency or variance. 3. Establish the root-cause hypothesis based on profiling evidence. For example, if redundant WAL sync is considered the root cause, we should first show that the measured path actually spends significant time there and explain why it causes variance. It is also important to distinguish between reducing average latency and reducing CV. Lower average latency does not automatically prove that the variance problem has been solved. 4. Make the smallest targeted production change based on the confirmed bottleneck. Avoid mixing several unrelated optimizations into one experiment, otherwise we cannot determine which change actually produced the improvement. 5. Run an equivalent before/after benchmark on the same machine and environment, and provide at least: * Mean latency * Error / confidence interval * CV * Allocation rate / B/op if relevant * GC time/count if relevant 6. Run the same profiler again after the change and compare it with the baseline. The previously identified hotspot should disappear or become significantly smaller. 7. Verify correctness and durability independently. If we change `hsync` / `hflush` behavior, the test should cover the actual HDFS/filesystem execution path whose durability semantics are being changed. The complete evidence chain should ideally be: Benchmark symptom → Baseline measurement → CPU / Wall / Lock / GC profiling → Confirmed hotspot → Root-cause hypothesis → Targeted optimization → Equivalent before/after benchmark → Before/after profiling → Correctness / durability verification → Conclusion With this evidence in the PR description, reviewers can understand the causal relationship, reproduce the experiment locally, and verify that the optimization actually addresses the original latency-variance/CV problem rather than only improving the average score. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
