andygrove opened a new pull request, #6308:
URL: https://github.com/apache/datafusion-comet/pull/6308

   ## Which issue does this PR close?
   
   N/A
   
   ## Rationale for this change
   
   The README and the TPC-DS benchmark page show 1.0.0 results measured on 
Spark 3.5.8. This updates them for Comet 1.1.0, which is about to be released, 
measured on Spark 4.2.0.
   
   ## What changes are included in this PR?
   
   - New result files in `benchmarks/results/1.1.0/` (`spark-tpcds.json`, 
`comet-tpcds.json`), in the same format as the 1.0.0 ones.
   - New charts in `docs/source/_static/images/benchmark-results/1.1.0/`, 
generated with `benchmarks/tpc/generate-comparison.py`.
   - The README and `benchmark-results/tpc-ds.md` now use the 1.1.0 charts. The 
TPC-DS page now states the Spark and Comet versions and that each query was run 
twice (mean reported).
   - New README headline: **~1.8x speedup over Spark 4.2, ~45% cost savings** 
(was ~2x and ~50%). The speedup ratio drops because Spark itself got much 
faster between 3.5.8 and 4.2.0; Comet's own time also improved.
   
   | | Spark | Comet | Speedup |
   |---|---:|---:|---:|
   | 1.0.0 (Spark 3.5.8) | 1410.4s | 637.5s | 2.21x |
   | 1.1.0 (Spark 4.2.0) | 966.4s | 524.7s | 1.84x |
   
   Across the 103 queries, the geometric mean speedup is 1.66x. Comet is slower 
than Spark on 4 queries, most notably q54 (also slower in 1.0.0) and q68 (1.4s 
in 1.0.0, 3.9s in 1.1.0).
   
   Spark 4.2 support is still experimental. I used it as the headline because 
it's the newest Spark version, but I'm happy to switch the headline comparison 
to 4.1 if people prefer.
   
   ## How are these changes tested?
   
   The benchmarks ran TPC-DS SF1000 (Parquet) on the same EKS setup as the 
1.0.0 results (`r6i.24xlarge`, data in S3), with the configuration documented 
on the page:
   - Spark 4.2.0 (Scala 2.13), Comet built from `branch-1.1` (`ee3f239`).
   - 32 executors × 16 cores.
   - 2 iterations per query, no errors.
   
   I checked from each run's Spark environment that both used the documented 
settings, and that Comet was fully disabled in the Spark run.
   
   `prettier --check` passes on the changed Markdown.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to