nzw921rx commented on PR #12130: URL: https://github.com/apache/seatunnel/pull/12130#issuecomment-5565251399
> @nzw921rx happy to split them out. But need to make a decision on something as it changes how the comparison works. > > The two benchmarks are `finishedJobsFirstPageOld`, which calls the existing `getJobsByStateJson(state)`, and `finishedJobsFirstPageNew`, which calls the `getJobsByStateJson(state, start, rows)` overload this PR adds. The second will not compile against dev. The first does compile, but this PR only refactors that method so it shares its filter and sort with the new overload, it does not make it faster. So a benchmark PR run against dev alone would show no improvement, which I do not think is what you are after. > > Two ways to get a real baseline comparison: > > 1. Benchmark the servlet rather than the service, driving `FinishedJobsServlet.doGet` with `page` and `rows` set. On dev that path builds the whole listing and slices afterwards, so the same benchmark gets faster once this PR merges and the before and after is genuine. It also measures what callers actually hit. Needs a small fake request and response in the harness. > 2. Land the benchmark PR after this one, where both methods exist and a single run compares them directly. > > I lean towards 1, since it stays useful as a regression check on the endpoint afterwards. Happy to do either. Which would you prefer? I understand that if it is a newly added method, it may not be possible to compare through a stable dev baseline, and there may not be a need to submit a benchmark PR. So, I have a question to ask you? The specific steps of your benchmark testing, I understand that the bottlenecks we have tested so far are all in IMAP. It seems that pagination cannot speed up IMAP operations, resulting in fluctuations and large confidence intervals for CV errors. Do you have identified any issues with flame graphs, JFR graphs, or optimized GC? I am very puzzled about the optimization effect described in the PR, as it seems to have improved many times faster. However, I am more confused about which aspect of improvement has been sacrificed. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
