Neuw84 commented on PR #6270:
URL: 
https://github.com/apache/datafusion-comet/pull/6270#issuecomment-5913466670

   We tested this PR together with #6268 on TPC-DS at SF1000: Parquet on S3, a 
Spark 4.1 cluster on EKS, 8 executors x 13 cores on m5.4xlarge, one 
availability zone, AQE on with 300 shuffle partitions.
   
   The build was main `b58b2f3a` with #6268 (`588c029f`) and #6270 (`b3f4f058`) 
merged on top (HEAD `0f22d064`), with the native library built for x86-64-v3. 
Comet ran with native scan, exec and shuffle enabled.
   
   **q64 now returns every row.** It gives 12,185 rows, the same as vanilla 
Spark, with the same checksum (`bce69949`). It ran in 63.8 s with 249/249 
operators on Comet. Before this fix, Comet lost rows on q64 at this scale 
(#6264).
   
   All 103 queries complete, and row counts match vanilla Spark on every query. 
Checksums match on all but q65, which has ties in its ORDER BY and differs the 
same way for every engine we compare.
   
   Thanks for the fix.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to