Neuw84 commented on PR #6270: URL: https://github.com/apache/datafusion-comet/pull/6270#issuecomment-5913466670
We tested this PR together with #6268 on TPC-DS at SF1000: Parquet on S3, a Spark 4.1 cluster on EKS, 8 executors x 13 cores on m5.4xlarge, one availability zone, AQE on with 300 shuffle partitions. The build was main `b58b2f3a` with #6268 (`588c029f`) and #6270 (`b3f4f058`) merged on top (HEAD `0f22d064`), with the native library built for x86-64-v3. Comet ran with native scan, exec and shuffle enabled. **q64 now returns every row.** It gives 12,185 rows, the same as vanilla Spark, with the same checksum (`bce69949`). It ran in 63.8 s with 249/249 operators on Comet. Before this fix, Comet lost rows on q64 at this scale (#6264). All 103 queries complete, and row counts match vanilla Spark on every query. Checksums match on all but q65, which has ties in its ORDER BY and differs the same way for every engine we compare. Thanks for the fix. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
