SEZ9 commented on issue #12063:
URL: https://github.com/apache/seatunnel/issues/12063#issuecomment-5578286139

   Thanks @Rangsh for following the investigation-first order and for the clear 
write-up. Attributing the variance mainly to the shared growth fixture is a 
plausible finding, and the three points you listed (WAL replay through 
`FileMapStore.loadAll` in iteration tear-down, fixed keys reused across growth 
iterations, and the double write-through in 
`JobHistoryService.storeFinishedPipelineMetrics`) give us something concrete to 
verify. I have not reviewed the PR you linked yet; I will look at it, but a few 
things would help me confirm the analysis and keep the scope tight:
   
   1. Evidence for the fixture attribution: please attach (here or in the PR) 
the JMH JSON plus the GC / JFR or wall-clock data that shows the per-iteration 
cost growing across SingleShot samples with the old tear-down, and the 
corresponding data after the fixture change. That is what lets us separate 
fixture noise from the storage path itself.
   2. Same-runner before/after: as you noted, local scores are not comparable 
to the GitHub Actions baseline. Before we settle the production part, I would 
like to see before/after for both `runningJobGrowth` and 
`completedJobHistoryGrowth` with the same runner class, JDK, arguments and 
storage configuration as the original report, including 
`initialStoredJobCount=1000`, not only `initialStoredJobCount=0`.
   3. Moving the durable MapStore reload sample from iteration tear-down to 
trial tear-down: please explain in the PR how the benchmark still detects a 
failure to persist or reload job state between iterations, so we are not 
trading variance for lost coverage.
   4. For the single TTL `put` change: please spell out in the PR description 
why the `computeIfAbsent` + `put` collapse cannot change the observed state for 
a concurrently finishing job or for an existing entry, and point to the unit 
test that covers both the new-job and existing-job cases. Keeping running and 
completed job-state durability unchanged is the constraint I care most about 
here.
   
   If the same-runner data does not show a measurable reduction in CV from the 
production change on its own, I would prefer to land the fixture stabilization 
first and discuss the production change separately. Happy to iterate on this 
once the evidence is up.
   
   <!-- streview-comment:890 -->


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to