nzw921rx commented on PR #12248:
URL: https://github.com/apache/seatunnel/pull/12248#issuecomment-5632743268

   Hello, I have some thoughts that I would like to discuss with you regarding 
the original design idea of the report:
   
   A. I first asked myself several questions:
   
   1. How can we use the most concise daily report to intuitively identify 
problems at a glance?
      My answer to myself is that the fewer metrics we focus on, and the more 
those metrics can truly express meaningful information and provide long-term 
value, the more appropriate they are to be reflected in the report. The 
detailed process should be handled by downloading the detailed artifacts for 
investigation once a problem has been identified.
   
   2. How much runner fluctuation do we consider acceptable?
      Based on my own long-term observation, fluctuations within 5% are very 
common.
   
   3. Are different JDKs comparable?
      They are not. The servers they run on almost always have different 
CPUs/models/core counts/memory for each run, so the scores between different 
JDKs are naturally not comparable, and Error/CV fluctuations are also related 
to machine load.
   
   B. Have we currently found that there are somewhat too many scenarios in the 
long-term `benchmarks_core` report?
   
   Yes, they can be simplified. I think the long-term value lies in stable 
benchmarks, for example:
   
   CheckpointingTime.checkpointSingleInput
   CheckpointingTime.checkpointSingleInput
   Pipeline.sourceSink
   Pipeline.sourceTransformSink
   
   They have been very stable recently, with both Error/CV around 1%. I think 
these are suitable to remain in the long-term report, and they help us identify 
regressions.
   
   Summary:
   
   I think the intention of this PR is very clear. It wants people who see CV 
fluctuations to further diagnose the fluctuation based on the details, but 
there is a maintainability issue here, as well as the question of what kind of 
problem it ultimately provides value in identifying. In fact, it is trying to 
identify the case where only one fork is high while the others are very stable, 
but here I want to give an example: in my observation, the case where one fork 
is high has almost never occurred.
   
   1.Pipeline.sourceTransformSink
   This JMH benchmark has been very stable during my observations over the past 
few weeks and has not shown this kind of problem, so this further reduces the 
intention of the new report to reduce investigation costs.
   
   2.IntermediateQueue.disruptorRecordHandoff
   This JMH benchmark has been consistently unstable during my observations 
over the past few weeks, which is intended to demonstrate that the current CV 
metric already fully expresses this situation.
   
   I look forward to your reply. I would like to seriously discuss this with 
you, and I also look forward to your suggestions.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to